Syntax for Subpicture Signaling in Video Bitstreams
Sub-picture based coding techniques and encoder-side preprocessing address the challenges of managing sub-pictures and slices in video coding standards, enhancing decoding efficiency and quality.
Patent Information
- Application Number
- JP2024191324
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-10-02
- Filing Date
- 2024-10-31
- Publication Date
- 2025-12-09
- Estimated Expiration
- 2040-10-09
AI Technical Summary
Existing video coding standards face challenges in efficiently handling the increasing bandwidth demands for digital video, particularly in managing sub-pictures and slices within video bitstreams, which affect decoding efficiency and quality.
Implementing sub-picture based coding techniques that allow for flexible tiling, controlling in-loop filtering across sub-picture boundaries, and using encoder-side preprocessing for temporal filtering to enhance decoding efficiency and quality.
Improves decoding efficiency and quality by optimizing the handling of sub-pictures and slices within video bitstreams, reducing bandwidth requirements and enhancing video processing capabilities.
Smart Images

Figure 0007783383000080 
Figure 0007783383000081 
Figure 0007783383000082
Abstract
Description
[Technical Field]
[0001] This application is a division of Patent Application No. 2023-121558, which is a division of Patent Application No. 2022-520358 based on International Patent Application No. PCT / CN2020 / 119932 filed on October 9, 2020, which claims priority to and the benefit of International Patent Application No. PCT / CN2019 / 109809 filed on October 2, 2019. The entire disclosure of the aforementioned application is incorporated by reference as part of the disclosure of this application.
[0002] This document relates to video and image coding and decoding techniques. [Background technology]
[0003] Digital video accounts for the largest bandwidth usage on the Internet and other digital communication networks, and the bandwidth demands for digital video usage are expected to continue to increase as the number of connected user devices capable of receiving and displaying video increases. Summary of the Invention
[0004] The disclosed techniques may be used by embodiments of video or image decoders or encoders in which sub-picture based coding or decoding is performed.
[0005] In one exemplary aspect, a method of video processing is disclosed. The method includes performing a conversion between a video including one or more pictures and a bitstream representation of the video. The bitstream representation is required to conform to format rules that specify that each picture is coded as one or more slices, and the format rules prohibit a sample in a picture from being uncovered by any of the one or more slices.
[0006] In yet another aspect, a method of video processing is disclosed. The method includes determining a manner of signaling information for one or more slices in a picture according to rules related to a number of tiles or bricks in the picture for conversion between a picture of the video and a bitstream representation of the video. The method also includes performing the conversion based on the determination.
[0007] In yet another aspect, a method of video processing is disclosed. The method includes performing a conversion between a picture of a video and a bitstream representation of the video according to rules. A picture is coded in the bitstream representation as one or more slices, and the rules specify whether or how addresses of slices of the picture are included in the bitstream representation.
[0008] In yet another aspect, a method of video processing is disclosed. The method includes, for a conversion between a picture of the video and a bitstream representation of the video, determining whether a syntax element indicating a filter operation that accesses samples across multiple bricks in the picture is enabled to be included in the bitstream representation based on a number of tiles or a number of bricks in the picture. The method also includes performing the conversion based on the determination.
[0009] In yet another aspect, a method of video processing is disclosed. The method includes performing a conversion between a picture of a video and a bitstream representation of the video, the picture including one or more sub-pictures, the number of the one or more sub-pictures being indicated by a syntax element in the bitstream representation.
[0010] In yet another aspect, a method of video processing is disclosed. The method includes performing a conversion between a picture of a video that includes one or more sub-pictures and a bitstream representation of the video. The bitstream representation conforms to format rules that specify that information about the sub-pictures is included in the bitstream representation based on at least one of: (1) one or more corner locations of the sub-pictures; or (2) dimensions of the sub-pictures.
[0011] In yet another aspect, a method of video processing is disclosed. The method includes determining that a reference picture resampling tool is enabled for converting between a picture of a video and a bitstream representation of the video by dividing the picture into one or more sub-pictures. The method also includes performing the conversion based on the determination.
[0012] In yet another aspect, a method of video processing is disclosed. The method includes performing a conversion between a video picture including one or more sub-pictures that contain one or more slices and a bitstream representation of the video. The bitstream representation conforms to format rules that specify that for the sub-pictures and slices, if an index identifying the sub-picture is included in the slice's header, then an address field for the slice indicates an address of the slice within the sub-picture.
[0013] In yet another aspect, a method of video processing is disclosed that includes determining, for a video block in a first video region of a video, whether a location at which a temporal motion vector predictor is determined for a transformation between the video block and a bitstream representation of the current video block using an affine mode is within a second video region, and performing the transformation based on the determination.
[0014] In yet another example aspect, a method of video processing is disclosed that includes determining, for a video block in a first video region of a video, whether a location from which integer samples in a reference picture are fetched for a conversion between the video block and a bitstream representation of a current video block is within a second video region, where the reference picture is not used in an interpolation process during the conversion, and performing the conversion based on the determination.
[0015] In yet another example aspect, a method of video processing is disclosed that includes, for a video block in a first video domain of a video, determining whether a location from which reconstructed luma sample values are fetched for a conversion between the video block and a bitstream representation of a current video block is within a second video domain, and performing the conversion based on the determination.
[0016] In yet another exemplary aspect, a method of video processing is disclosed that includes determining, for a video block in a first video region of a video, whether a location within a second video region where a check for splitting, depth derivation, or split flag signaling for the video block is performed during a conversion between the video block and a bitstream representation of the current video block is within the second video region, and performing the conversion based on the determination.
[0017] In yet another exemplary aspect, a method of video processing is disclosed. The method includes performing a conversion between a video including one or more video pictures including one or more video blocks and a coded representation of the video, wherein the coded representation complies with coding syntax requirements that the conversion does not use sub-picture coding / decoding and dynamic resolution coding / decoding tools, or reference picture resampling tools within a video unit.
[0018] In yet another example aspect, a method of video processing is disclosed. The method includes performing a conversion between a video including one or more video pictures including one or more video blocks and a coded representation of the video, wherein the coded representation complies with a coding syntax requirement that a first syntax element, subpic_grid_idx[i][j], is not greater than a second syntax element, max_subpics_minus1.
[0019] In yet another exemplary aspect, the above-described method may be implemented by a video encoder device including a processor.
[0020] In yet another exemplary aspect, the above-described methods may be implemented by a video decoder device including a processor.
[0021] In yet another exemplary aspect, the above-described methods may be embodied in the form of processor-executable instructions and stored on a computer-readable program medium.
[0022] These and other aspects are further described in this document. [Brief explanation of the drawings]
[0023] [Figure 1] 1 shows an example of region constraints in temporal motion vector prediction (TMVP) and sub-block TMVP.
[0024] [Figure 2] 1 illustrates an example of a hierarchical motion prediction scheme.
[0025] [Figure 3] FIG. 1 is a block diagram of an example hardware platform for implementing the techniques described herein.
[0026] [Figure 4] 1 shows a flowchart of an exemplary method for video processing.
[0027] [Figure 5] An example of a picture with an 18x12 luma CTU partitioned into 12 tiles and 3 raster scan slices (reference) is shown.
[0028] [Figure 6] An example of a picture with an 18x12 luma CTU partitioned into 24 tiles and 9 rectangular slices (reference) is shown.
[0029] [Figure 7] An example of a picture partitioned into 4 tiles, 11 bricks, and 4 rectangular slices (for reference) is shown.
[0030] [Figure 8] FIG. 1 is a block diagram illustrating an example video processing system capable of implementing various techniques disclosed herein.
[0031] [Figure 9] 1 is a block diagram illustrating an example video coding system.
[0032] [Figure 10] FIG. 2 is a block diagram illustrating an encoder according to some embodiments of the present disclosure.
[0033] [Figure 11] FIG. 2 is a block diagram illustrating a decoder according to some embodiments of the present disclosure.
[0034] [Figure 12] 1 is a flowchart representation of a method for video processing in accordance with the present technology.
[0035] [Figure 13] 10 is a flowchart representation of another method for video processing in accordance with the present technology.
[0036] [Figure 14] 10 is a flowchart representation of another method for video processing in accordance with the present technology.
[0037] [Figure 15] 10 is a flowchart representation of another method for video processing in accordance with the present technology.
[0038] [Figure 16] 10 is a flowchart representation of another method for video processing in accordance with the present technology.
[0039] [Figure 17] 10 is a flowchart representation of another method for video processing in accordance with the present technology.
[0040] [Figure 18] 10 is a flowchart representation of another method for video processing in accordance with the present technology.
[0041] [Figure 19] 10 is a flowchart representation of yet another method for video processing in accordance with the present technology. DETAILED DESCRIPTION OF THE INVENTION
[0042] This document provides various techniques that may be used by decoders of images or video bitstreams to improve the quality of decompressed or decoded digital video or images. For simplicity, the term "video" is used herein to include both a series of pictures (traditionally called a video) and individual images. Furthermore, video encoders may also implement these techniques during the encoding process to reconstruct decoded frames to be used for further encoding.
[0043] Section headings are used in this document for ease of understanding, but do not limit the embodiments and techniques to the corresponding section, and therefore, embodiments from one section can be combined with embodiments from other sections.
[0044] 1. Overview
[0045] This document relates to video coding technology. Specifically, it relates to palette coding using a basic color-based representation in video coding. This may be applied to existing video coding standards such as HEVC or to a finalized standard (Versatile Video Coding). This may also be applicable to future video coding standards or video codecs.
[0046] 2. Initial discussion
[0047] Video coding standards have evolved primarily through the development of well-known ITU-T and ISO / IEC standards. ITU-T developed H.261 and H.263, while ISO / IEC developed MPEG-1 and MPEG-4 Visual. The two organizations jointly developed H.262 / MPEG-2 Video, H.264 / MPEG-4 Advanced Video Coding (AVC), and H.265 / HEVC standards [1, 2]. Starting with H.262, video coding standards have been based on hybrid video coding architectures that utilize temporal prediction plus transform coding. To explore future video coding technologies beyond HEVC, the Joint Video Exploration Team (JVET) was jointly established by VCEG and MPEG in 2015. Since then, many new methods have been adopted by JVET and incorporated into reference software named the Joint Exploration Model (JEM). In April 2018, the Joint Video Expert Team (JVET) was established by VCEG (Q6 / 16) and ISO / IEC JTC1 SC29 / WG11 (MPEG) to work on the VVC standard, which aims to reduce the bitrate by 50% compared to HEVC.
[0048] 2.1 Region Constraints in TMVP and Sub-Block TMVP in VVC
[0049] Figure 1 shows exemplary region constraints in TMVP and sub-block TMVP. In TMVP and sub-block TMVP, we are constrained to be able to fetch temporal MVs only from a collocated column of CTU+4x4 blocks, as shown in Figure 1.
[0050] 2.2 Exemplary Subpictures
[0051] In some embodiments, a sub-picture based coding technique based on a flexible tiling approach can be implemented. An overview of the sub-picture based coding technique includes the following.
[0052] (1) A picture can be divided into sub-pictures.
[0053] (2) An indication of the presence of a sub-picture is shown in the SPS along with other sequence-level information about the sub-picture.
[0054] (3) Whether a subpicture is treated as a picture in the decoding process (except for in-loop filtering operations) can be controlled by the bitstream.
[0055] (4) Whether in-loop filtering across subpicture boundaries is disabled can be controlled by the bitstream for each subpicture. The DBF, SAO, and ALF processes are updated to control in-loop filtering operations across subpicture boundaries.
[0056] (5) For simplicity, as a starting point, the subpicture width, height, horizontal offset, and vertical offset are signaled in units of luma samples in SPS. The subpicture boundaries are constrained to be slice boundaries.
[0057] (6) Treating subpictures as pictures in the decoding process (except for in-loop filtering operations) is specified by a slight update to the coding_tree_unit() syntax, and is updated to the following decoding process:
[0058] - Derivation process for (advanced) temporal luma motion vector prediction
[0059] - Luma sample bilinear interpolation process
[0060] - Luma sample 8-tap interpolation filtering process
[0061] - Chroma sample interpolation process
[0062] (7) Subpicture IDs are explicitly specified in the SPS and included in the tile group header to enable extraction of the subpicture sequence without the need to modify the VCL NAL units.
[0063] (8) Output Subpicture Set (OSPS) is proposed to specify the norm extraction and matching points for a subpicture and its set.
[0064] 2.3 Exemplary Subpictures in Versatile Video Coding Sequence Parameter Set RBSP Syntax [Table 1] subpics_present_flag equal to 1 indicates that subpicture parameters are currently present in the SPS RBSP syntax. subpics_present_flag equal to 0 indicates that subpicture parameters are not currently present in the SPS RBSP syntax. NOTE 2 – When a bitstream is the result of a sub-bitstream extraction process and contains only a subset of the subpictures of the input bitstream to the sub-bitstream extraction process, it may be necessary to set the value of subpics_present_flag equal to 1 in the RBSP of the SPS. max_subpics_minus1+1 specifies the maximum number of subpictures that can exist in the CVS. max_subpics_minus1 is in the range 0 to 254. The value 255 is reserved for future use by ITU-T|ISO / IEC. subpic_grid_col_width_minus1+1 specifies the width of each element of the subpicture identifier grid in units of 4 samples. The syntax element is Ceil(Log2(pic_width_max_in_luma_samples / 4)) bits in length. The variable NumSubPicGridCols is derived as follows:
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
[0065] 2.4 Exemplary Encoder-Only GOP-Based Temporal Filter
[0066] In some embodiments, an encoder-only temporal filter may be implemented. Filtering is performed on the encoder side as a preprocessing step. Source pictures before and after the selected picture to be coded are read, and a block-based motion compensation method for the selected picture is applied to those source pictures. Samples in the selected picture are temporally filtered using the motion-compensated sample values.
[0067] The overall filter strength is set according to the temporal sublayer and QP of the selected picture: only pictures in temporal sublayers 0 and 1 are filtered, with pictures in layer 0 being filtered more strongly than pictures in layer 1. The per-sample filter strength is adjusted according to the difference between the sample values of the selected picture and the juxtaposed samples of the motion compensated picture, so that small differences between the motion compensated picture and the selected picture are filtered more strongly than large differences.
[0068] GOP-based temporal filtering
[0069] The temporal filter is applied immediately after reading the picture and before encoding. The steps are described in more detail below.
[0070] Action 1: Read the picture by the encoder
[0071] Action 2: If the picture is low enough in the coding hierarchy, it is filtered before encoding. Otherwise, the picture is coded without filtering. RA pictures with POC%8==0 are filtered in the same way as LD pictures with POC%4==0. AI pictures are not filtered.
[0072] The overall filter strength So is set relative to RA according to the following formula:
number
[0073] where n is the number of images to be loaded.
[0074] For the LD case, so(n)=0.95 is used.
[0075] Action 3: The two pictures before and / or after the selected picture (hereafter called the original picture) are read. In edge cases, for example, if it is the first picture or close to the last picture, only the available picture is read.
[0076] Operation 4: The motion of the pictures read before and after the original picture is estimated for each 8x8 picture block.
[0077] A hierarchical motion estimation scheme is used, with layers L0, L1, and L2 shown in Figure 2. Subsampled pictures. For example, L1 in Figure 1 is generated by averaging each 2x2 block for all read pictures and the original picture. L2 is derived from L1 using the same subsampling method.
[0078] Figure 2 shows an example of different layers of hierarchical motion estimation: L0 is the original resolution; L1 is a subsampled version of L0; and L2 is a subsampled version of L1.
[0079] First, motion estimation is performed for each 16x16 block in L2. The squared difference is calculated for each selected motion vector, and the motion vector corresponding to the smallest difference is selected. The selected motion vector is then used as the initial value when estimating motion in L1. The same is then done to estimate motion in L0. As a final step, sub-pixel motion is estimated for each 8x8 block by using an interpolation filter on L0.
[0080] A VTM 6-tap interpolation filter can be used. [Table 2]
[0081] Step 5: Motion compensation is applied to pictures before and after the original picture according to the best matching motion for each block, e.g., so that the sample coordinates of the original picture in each block have the best matching coordinates in the referenced image.
[0082] Action 6: Process the samples one by one for the luma and chroma channels as described in the following steps.
[0083] Action 7: The new sample value In is calculated using the following formula:
number
[0084] where Io is the sample value of the original sample, Ir(i) is the intensity of the corresponding sample in motion-compensated picture i, and wr(i,a) is the weight of motion-compensated picture i when the number of available motion-compensated pictures is a.
[0085] For the luma channel, the weight wr(i,a) is defined as follows:
number
[0086] where:
number
[0087] For all other cases of i and a,
number
[0088] For the luma channel, the weight wr(i,a) is defined as follows:
number
[0089] where sc=0.55 and σc=30.
[0090] Action 8: The filter is applied to the current sample, and the resulting sample values are stored separately.
[0091] Act 9: The filtered picture is encoded.
[0092] 2.5 Example Picture Partitions (Tiles, Bricks, Slices)
[0093] In some implementations, a picture is divided into one or more tile rows and one or more tile columns. A tile is a sequence of CTUs that covers a rectangular area of the picture.
[0094] A tile is divided into one or more bricks, each consisting of multiple rows of CTUs within the tile.
[0095] A tile that is not partitioned into multiple bricks is also called a brick, but a brick that is a true subset of a tile is not called a tile.
[0096] A slice contains multiple tiles of a picture, or multiple bricks of a tile.
[0097] A subpicture contains one or more slices that together cover a rectangular area of the picture.
[0098] Two modes of slicing are supported: raster scan slice mode and rectangular slice mode. In raster scan slice mode, a slice contains a sequence of tiles in the tile raster scan of a picture. In rectangular slice mode, a slice contains multiple bricks of a picture that collectively form a rectangular region of the picture. The bricks within a rectangular slice are in the order of the brick raster scan of the slice.
[0099] FIG. 5 shows an example of raster scan slice partitioning of a picture, where the picture is divided into 12 tiles and 3 raster scan slices.
[0100] FIG. 6 shows an example of rectangular slice partitioning of a picture, where the picture is divided into 24 tiles (6 tile rows and 4 tile columns) and 9 rectangular slices.
[0101] Figure 7 shows an example of a picture partitioned into tiles, bricks, and rectangular slices, where the picture is divided into four tiles (two tile rows and two tile columns), 11 bricks (the top left tile contains one brick, the top right tile contains five bricks, the bottom left tile contains two bricks, and the bottom right tile contains three bricks), and four rectangular slices. Picture Parameter Set RBSP Syntax [Table 3] single_tile_in_pic_flag equal to TIFF0007783383000034.tif2111681 specifies that there is only one tile in each picture that references the PPS. single_tile_in_pic_flag equal to 0 specifies that there are multiple tiles in each picture that reference the PPS. NOTE - If there are no further brick splits within the tile, the whole tile is called a brick. If a picture contains only a single tile with no further brick splits, it is called a single brick. A requirement for bitstream compatibility is that the value of single_tile_in_pic_flag shall be the same for all PPSs referenced by coded pictures in the CVS. uniform_tile_spacing_flag equal to 1 specifies that tile column boundaries, as well as tile row boundaries, are uniformly distributed across the picture and are signaled using the syntax elements tile_cols_width_minus1 and tile_rows_height_minus1. uniform_tile_spacing_flag equal to 0 specifies that tile column boundaries, as well as tile row boundaries, may or may not be uniformly distributed across the picture and are signaled using the syntax elements num_tile_columns_minus1 and num_tile_rows_minus1 and the list of syntax element pairs tile_column_width_minus1[i] and tile_row_height_minus1[i]. When not present, the value of uniform_tile_spacing_flag is inferred to be equal to 1. tile_cols_width_minus1+1 specifies the width in CTB units of tile columns excluding the rightmost tile column of the picture when uniform_tile_spacing_flag is equal to 1. The value of tile_cols_width_minus1 shall be in the range 0 to PicWidthInCtbsY-1 inclusive. If not present, the value of tile_cols_width_minus1 is inferred to be equal to PicWidthInCtbsY-1. tile_rows_height_minus1+1 specifies the height in CTB units of the tile rows of the picture, excluding the bottom tile row, when uniform_tile_spacing_flag is equal to 1. The value of tile_rows_height_minus1 shall be in the range 0 to PicHeightInCtbsY-1 inclusive. If not present, the value of tile_rows_height_minus1 is inferred to be equal to PicHeightInCtbsY-1. num_tile_columns_minus1+1 specifies the number of tile columns into which the picture is partitioned when uniform_tile_spacing_flag is equal to 0. The value of num_tile_columns_minus1 shall be in the range 0 to PicWidthInCtbsY-1, inclusive. If single_tile_in_pic_flag is equal to 1, the value of num_tile_columns_minus1 is inferred to be equal to 0. Otherwise, when uniform_tile_spacing_flag is equal to 1, the value of num_tile_columns_minus1 is inferred as specified in Section 6.5.1. num_tile_rows_minus1+1 specifies the number of tile rows into which the picture is partitioned when uniform_tile_spacing_flag is equal to 0. The value of num_tile_rows_minus1 shall be in the range 0 to PicHeightInCtbsY-1, inclusive. If single_tile_in_pic_flag is equal to 1, the value of num_tile_rows_minus1 is inferred to be equal to 0. Otherwise, when uniform_tile_spacing_flag is equal to 1, the value of num_tile_rows_minus1 is inferred as specified in Section 6.5.1. The variable NumTilesInPic is set equal to (num_tile_columns_minus1+1)*(num_tile_rows_minus1+1). When single_tile_in_pic_flag is equal to 0, NumTilesInPic shall be greater than 1. tile_column_width_minus1[i]+1 specifies the width of the i-th tile column in CTB units. tile_row_height_minus1[i]+1 specifies the height of the i-th tile row in CTB units. brick_splitting_present_flag equal to 1 specifies that one or more tiles of a picture referencing a PPS may be split into two or more bricks. brick_splitting_present_flag equal to 0 specifies that tiles of a picture referencing a PPS shall not be split into two or more bricks. num_tiles_in_pic_minus1+1 specifies the number of tiles in each picture that refer to the PPS. The value of num_tiles_in_pic_minus1 shall be equal to NumTilesInPic-1. When not present, the value of num_tiles_in_pic_minus1 is inferred to be equal to NumTilesInPic-1. brick_split_flag[i] equal to 1 specifies that the i-th tile is split into two or more bricks. brick_split_flag[i] equal to 0 specifies that the i-th tile is not split into two or more bricks. When not present, the value of brick_split_flag[i] is inferred to be equal to 0. In some embodiments, a PPS parsing dependency on SPS is introduced by adding the syntactic condition "if(RowHeight[i]>1)" (e.g., similarly for uniform_brick_spacing_flag[i]). uniform_brick_spacing_flag[i] equal to 1 specifies that horizontal brick borders are uniformly distributed across the i-th tile and are signaled using the syntax element brick_height_minus1[i]. uniform_brick_spacing_flag[i] equal to 0 specifies that horizontal brick borders may or may not be uniformly distributed across the i-th tile and are signaled using the syntax element num_brick_rows_minus2[i] and the list of syntax elements brick_row_height_minus1[i][j]. When not present, the value of uniform_brick_spacing_flag[i] is inferred to be equal to 1. brick_height_minus1[i]+1 specifies the height in CTB units of the brick row excluding the bottom brick of the ith tile when uniform_brick_spacing_flag[i] is equal to 1. If present, the value of brick_height_minus1 shall be in the range 0 to RowHeight[i]-2 inclusive. If absent, the value of brick_height_minus1[i] is inferred to be equal to RowHeight[i]-1. num_brick_rows_minus2[i] specifies the number of bricks into which the ith tile is partitioned when uniform_brick_spacing_flag[i] is equal to 0. When present, the value of num_brick_rows_minus2[i] shall be in the range 0 to RowHeight[i]-2, inclusive. When brick_split_flag[i] is equal to 0, the value of num_brick_rows_minus2[i] is inferred to be equal to -1. Otherwise, when uniform_brick_spacing_flag[i] is equal to 1, the value of num_brick_rows_minus2[i] is inferred as specified in Section 6.5.1. brick_row_height_minus1[i][j]+1 specifies the height of the jth brick in the ith tile in CTB units when uniform_tile_spacing_flag is equal to 0. The following variables are derived: the values of num_tile_columns_minus1 and num_tile_rows_minus1 are inferred when uniform_tile_spacing_flag is equal to 1, and for each i in the range 0 to NumTilesInPic-1 (inclusive), the value of num_brick_rows_minus2[i] is inferred when uniform_brick_spacing_flag[i] is equal to 1 by invoking the CTB raster and brick scan conversion process specified in Section 6.5.1. - A list RowHeight[j] in the range 0 to num_tile_rows_minus1, specifying the height of the jth tile row in CTB units - CtbAddrRsToBs[ctbAddrRs], a list of ctbAddrRs in the range 0 to PicSizeInCtbsY-1 (inclusive) that specifies the conversion of CTB addresses in the CTB raster scan of the picture to CTB addresses in the brick scan. - A list CtbAddrBsToRs[ctbAddrBs] of ctbAddrBs in the range 0 to PicSizeInCtbsY-1 (inclusive) that specifies the conversion from CTB addresses in brick scan to CTB addresses in the CTB raster scan of the picture. - BrickId[ctbAddrBs] is a list of ctbAddrBs in the range 0 to PicSizeInCtbsY-1 (inclusive) that specifies the conversion to CTB brick IDs in brick scanning. - A list NumCtusInBrick[brickIdx] for brickIdx in the range 0 to NumBricksInPic-1 (inclusive) that specifies the conversion from brick index to the number of CTUs in the brick - A list FirstCtbAddrBs[brickIdx] for brickIdx in the range 0 to NumBricksInPic-1 (inclusive) that specifies the conversion from brick ID to CTB address in brick traversal of the first CTB in the brick single_brick_per_slice_flag equal to 1 specifies that each slice referencing this PPS contains one brick. single_brick_per_slice_flag equal to 0 specifies that each slice referencing this PPS contains multiple bricks. When not present, the value of single_brick_per_slice_flag is inferred to be equal to 1. rect_slice_flag equal to 0 specifies that the bricks in each slice are in raster scan order and slice information is not signaled in the PPS. rect_slice_flag equal to 1 specifies that the bricks in each slice cover a rectangular area of the picture and slice information is signaled in the PPS. When brick_splitting_present_flag is equal to 1, the value of rect_slice_flag shall be equal to 1. When not present, rect_slice_flag is inferred to be equal to 1. num_slices_in_pic_minus1+1 specifies the number of slices in each picture that refer to the PPS. The value of num_slices_in_pic_minus1 shall be in the range 0 to NumBricksInPic-1, inclusive. When not present and single_brick_per_slice_flag is equal to 1, the value of num_slices_in_pic_minus1 is inferred to be equal to NumBricksInPic-1. bottom_right_brick_idx_length_minus1+1 specifies the number of bits used to represent the information element bottom_right_brick_idx_delta[i]. The value of bottom_right_brick_idx_length_minus1 shall be in the range 0 to Ceil(Log2(NumBricksInPic))-1, inclusive. When i is greater than 0, bottom_right_brick_idx_delta[i] specifies the difference between the brick index of the brick located in the bottom right corner of the i-th slice and the brick index of the bottom right corner of the (i-1)-th slice. bottom_right_brick_idx_delta[0] specifies the brick index of the bottom right corner of the 0th slice. When single_brick_per_slice_flag is equal to 1, the value of bottom_right_brick_idx_delta[i] is inferred to be equal to 1. The value of BottomRightBrickIdx[num_slices_in_pic_minus1] is inferred to be equal to NumBricksInPic-1. The length of the bottom_right_brick_idx_delta[i] syntax element is bottom_right_brick_idx_length_minus1+1 bits. brick_idx_delta_sign_flag[i] equal to 1 indicates the positive sign of bottom_right_brick_idx_delta[i]. sign_bottom_right_brick_idx_delta[i] equal to 0 indicates the negative sign of bottom_right_brick_idx_delta[i]. A requirement for bitstream conformance is that a slice shall only contain a contiguous sequence of multiple full tiles, or full bricks of one tile. The variables TopLeftBrickIdx[i], BottomRightBrickIdx[i], NumBricksInSlice[i], and BricksToSliceMap[j], which specify the brick index of the brick located in the top-left corner of the i-th slice, the brick index of the brick located in the bottom-right corner of the i-th slice, the number of bricks in the i-th slice, and the mapping of bricks to slices, are derived as follows:
number
number
number
[0102] 3. Examples of Technical Problems Solved by the Disclosed Embodiments
[0103] (1) There are some designs that may violate the subpicture constraint.
[0104] A. TMVP in affine constructed candidates may fetch MVs in collocated pictures from the extent of the current subpicture.
[0105] B. When deriving gradients in BDOF (Bi-Directional Optical Flow) and PROF (Prediction Refinement Optical Flow), two extended rows and two extended columns of integer reference samples need to be fetched. These reference samples may be outside the scope of the current subpicture.
[0106] C. When deriving the chroma residual scaling factor in LMCS (luma mapping chroma scaling), the accessed reconstructed luma samples may be outside the range of the current sub-picture.
[0107] D. Neighboring blocks may be outside the range of the current subpicture when deriving the luma intra prediction mode, reference samples for intra prediction, reference samples for CCLM, availability of neighboring blocks for spatial neighbor candidates for merge / AMVP / CIIP / IBC / LMCS, quantization parameters, the CABAC initialization process, ctxInc derivation using the top-left syntax element, and ctxIncfor for the syntax element mtt_split_cu_vertical_flag. Display of the subpicture may lead to a subpicture with an incomplete CTU. The CTU partition and CU split process may need to take the incomplete CTU into account.
[0108] (2) The signaled syntax elements for sub-pictures can be arbitrarily large, which can cause overflow problems.
[0109] (3) The representation of a subpicture may lead to a non-rectangular subpicture.
[0110] (4) Currently, subpictures and subpicture grids are defined in units of 4 samples. The length of the syntax element depends on the picture height divided by 4. However, because pic_width_in_luma_samples and pic_height_in_luma_samples are currently integer multiples of Max(8,MinCbSizeY), the subpicture grid may need to be defined in units of 8 samples.
[0111] (5) The SPS constructs pic_width_max_in_luma_samples and pic_height_max_in_luma_samples may need to be limited to 8 or more.
[0112] (6) The interaction between reference picture resampling / scalability and subpictures is not considered in the current design.
[0113] (7) In temporal filtering, samples across different sub-pictures may be required.
[0114] (8) When signaling a slice, information may sometimes be inferred without signaling.
[0115] (9) It is possible that not all of the defined slices can cover the entire picture or sub-picture.
[0116] 4. Exemplary Techniques and Embodiments
[0117] The following detailed list should be considered as examples for explaining general concepts. These items should not be construed narrowly. Further, these items can be combined in any manner. Hereinafter, a temporal filter is used to represent a filter that requires samples in other pictures. Max(x, y) returns the larger of x and y. Min(x, y) returns the smaller of x and y. 1. The position (named position RB) where a temporal MV predictor is fetched in a picture to generate an affine motion candidate must be within the required sub-picture. Assuming that the upper left corner coordinates of the required sub-picture are (xTL, yTL) and the lower right coordinates of the required sub-picture are (xBR, yBR). a. In one example, the required sub-picture is the sub-picture that covers the current block. b. In one example, if the position RB with coordinates (x, y) is outside the required sub-picture, the temporal MV predictor is treated as unavailable. i. In one example, the position RB is outside the required sub-picture if x > xBR. ii. In one example, the position RB is outside the required sub-picture if y > yBR. iii. In one example, the position RB is outside the required sub-picture if x < xTL. iv. In one example, when y < yTL, the position RB is outside the required subpicture. c. In one example, when the position RB is outside the required subpicture, replacement of RB is utilized. i. Alternatively, further, the replacement position shall be within the required subpicture. d. In one example, the position RB is clipped within the required subpicture. i. In one example, x is clipped as x = Min(x, xBR). ii. In one example, x is clipped as y = Min(y, yBR). iii. In one example, x is clipped as x = Max(x, xTL). iv. In one example, x is clipped as y = Max(y, yTL). e. In one example, the position RB may be the lower right position within the corresponding block of the current block in the juxtaposed picture. f. The proposed method may be utilized in other coding tools that require accessing motion information from a picture different from the current picture. g. In one example, whether the above method is applied (e.g., whether the position RB must be within the required subpicture (e.g., to be as claimed in 1.a and / or 1.b)) may depend on one or more syntax elements signaled in the VPS / DPS / SPS / PPS / APS / slice header / tile group header. For example, the syntax element may be subpic_treated_as_pic_flag[SubPicIdx], where SubPicIdx is the subpicture index of the subpicture covering the current block. 2. The position (named position S) where integer samples are fetched by a reference not used in the interpolation process must be within the required subpicture. It is assumed that the upper left corner coordinates of the required subpicture are (xTL, yTL) and the lower right coordinates of the required subpicture are (xBR, yBR). a. In one example, the necessary subpicture is the subpicture covering the current block. b. In one example, if the position S with coordinates (x, y) is outside the necessary subpicture, the reference sample is treated as unavailable. i. In one example, if x > xBR, the position S is outside the necessary subpicture. ii. In one example, if y > yBR, the position S is outside the necessary subpicture. iii. In one example, if x < xTL, the position S is outside the necessary subpicture. iv. In one example, if y < yTL, the position S is outside the necessary subpicture. c. In one example, the position S is clipped within the necessary subpicture. i. In one example, x is clipped as x = Min(x, xBR). ii. In one example, x is clipped as y = Min(y, yBR). iii. In one example, x is clipped as x = Max(x, xTL). iv. In one example, x is clipped as y = Max(y, yTL). d. In one example, whether the position S must be within the required subpicture (e.g., to be as claimed in 2.a and / or 2.b) may depend on one or more syntax elements signaled in the VPS / DPS / SPS / PPS / APS / slice header / tile group header. For example, the syntax element may be subpic_treated_as_pic_flag[SubPicIdx], where SubPicIdx is the subpicture index of the subpicture covering the current block. e. In one example, the fetched integer sample is used to generate gradients in BDOF and / or PORF. 3. The position where the reconstituted luma sample value is fetched (named position R) may be within the required subpicture. Assuming that the upper left corner coordinates of the required subpicture are (xTL, yTL) and the lower right coordinates of the required subpicture are (xBR, yBR). a. In one example, the required subpicture is the subpicture covering the current block. b. In one example, if the position R with coordinates (x, y) is outside the required subpicture, the reference sample is treated as unavailable. i. In one example, if x > xBR, the position R is outside the required subpicture. ii. In one example, if y > yBR, the position R is outside the required subpicture. iii. In one example, if x < xTL, the position R is outside the required subpicture. iv. In one example, if y < yTL, the position R is outside the required subpicture. c. In one example, the position R is clipped within the required subpicture. i. In one example, x is clipped as x = Min(x, xBR). ii. In one example, y is clipped as y = Min(y, yBR). iii. In one example, x is clipped as x = Max(x, xTL). iv. In one example, y is clipped as y = Max(y, yTL). d. In one example, whether the position R must be within the required subpicture (e.g., to be as claimed in 4.a and / or 4.b) may depend on one or more syntax elements signaled in the VPS / DPS / SPS / PPS / APS / slice header / tile group header. For example, the syntax element may be subpic_treated_as_pic_flag[SubPicIdx], where SubPicIdx is the subpicture index of the subpicture covering the current block. e. In one example, the fetched luma sample is used to derive a scaling factor for the chroma component in the LMCS. 4. The position of the picture boundary check (named position N) for BT / TT / QT split, BT / TT / QT depth derivation, and / or signaling of the CU split flag must be within the required subpicture. Assuming that the upper left corner coordinates of the required subpicture are (xTL, yTL) and the lower right coordinates of the required subpicture are (xBR, yBR). a. In one example, the required subpicture is the subpicture that covers the current block. b. In one example, if the position N with coordinates (x, y) is outside the required subpicture, the reference sample is treated as unavailable. i. In one example, if x > xBR, the position N is outside the required subpicture. ii. In one example, if y > yBR, the position N is outside the required subpicture. iii. In one example, if x < xTL, the position N is outside the required subpicture. iv. In one example, if y < yTL, the position N is outside the required subpicture. c. In one example, the position N is clipped within the required subpicture. i. In one example, x is clipped as x = Min(x, xBR). ii. In one example, x is clipped as y = Min(y, yBR). iii. In one example, x is clipped as x = Max(x, xTL). d. In one example, x is clipped as y = Max(y, yTL). In one example, whether position N must be within the requested subpicture (e.g., to be as claimed in 5.a and / or 5.b) may depend on one or more syntax elements signaled in the VPS / DPS / SPS / PPS / APS / slice header / tile group header. For example, the syntax element may be subpic_treated_as_pic_flag[SubPicIdx], where SubPicIdx is the subpicture index of the subpicture covering the current block. 5. The History-based Motion Vector Prediction (HMVP) table may be reset before decoding a new sub-picture within a picture. In one example, the HMVP table used for IBC coding may be reset. b. In one example, the HMVP table used for inter-coding may be reset. c. In one example, the HMVP table used for intra-coding may be reset. 6. The subpicture syntax element may be defined in units of N (N=8, 32, etc.) samples. a. In one example, the width of each element of the subpicture identifier grid in N samples b. In one example, the height of each element of the subpicture identifier grid in N samples c. In one example, N is set to the width and / or height of the CTU. 7. The picture width and picture height syntax elements may be limited to K (K>=8). In one example, picture width may need to be limited to 8 or more. b. In one example, picture height may need to be limited to 8 or more. 8. A conforming bitstream is one in which ARC (Adaptive resolution conversion) / DRC (Dynamic resolution conversion) / RPR (Reference picture resampling) are prohibited from being enabled for one video unit (e.g., a sequence). a. In one example, signaling to enable sub-picture coding may be under conditions that prohibit ARC / DRC / RPR. i. In one example, when subpictures are enabled, such that subpics_present_flag is equal to 1, pic_width_in_luma_samples for all pictures for which this SPS is active is equal to max_width_in_luma_samples. b. Alternatively, sub-picture coding and ARC / DRC / RPR may both be enabled for one video unit (eg, a sequence). i. In one example, a conforming bitstream satisfies that a sub-picture downsampled by ARC / DRC / RPR is still in the form of K CTUs in width and M CTUs in height (where K and M are both integers). ii. In one example, a conforming bitstream shall satisfy that for subpictures that are not located on a picture boundary (e.g., the right and / or bottom boundary), the subpictures downsampled by ARC / DRC / RPR are still in the form of K CTUs in width and M CTUs in height (where K and M are both integers). iii. In one example, the CTU size may be adaptively changed based on the picture resolution. 1) In one example, the maximum CTU size may be signaled in the SPS. For each lower-resolution picture, the CTU size may be changed accordingly based on the reduced resolution. 2) In one example, the CTU size may be signaled at the SPS and PPS, and / or sub-picture level. 9. The syntax elements subpic_grid_col_width_minus1 and subpic_grid_row_height_minus1 may be constrained. a. In one example, subpic_grid_col_width_minus1 must be less than or equal to T1 (or less than T1). b. In one example, subpic_grid_row_height_minus1 must be less than or equal to T2 (or less than T2). c. In one example, in a conforming bitstream, subpic_grid_col_width_minus1 and / or subpic_grid_row_height_minus1 must obey constraints such as item 3.a or 3.b. d. In one example, T1 in 3.a and / or T2 in 3.b may depend on the profile / level / tier of the video coding standard. e. In one example, T1 in 3.a may depend on the picture width. For example, T1 is equal to pic_width_max_in_luma_samples / 4 or pic_width_max_in_luma_samples / 4+Off, where Off can be 1, 2, -1, -2, etc. f. In one example, T2 in 3.b may depend on the picture width. For example, T1 is equal to pic_height_max_in_luma_samples / 4 or pic_height_max_in_luma_samples / 4-1+Off, where Off can be 1, 2, -1, -2, etc. 10. The boundary between two subpictures is constrained to be the boundary between two CTUs. In other words, a CTU cannot be covered by multiple sub-pictures. b. In one example, the unit of subpic_grid_col_width_minus1 may be the CTU width (32, 64, 128, etc.) instead of 4 as in VVC. The subpicture grid width should be (subpic_grid_col_width_minus1+1)*CTU width. c. In one example, the units of subpic_grid_col_height_minus1 may be CTU heights (32, 64, 128, etc.) instead of 4 as in VVC. The subpicture grid height should be (subpic_grid_col_height_minus1+1)*CTU height. d. In one example, in a conforming bitstream, if the sub-picture approach is applied, constraints must be met. 11. The shape of the subpicture is constrained to be rectangular. In one example, in a conforming bitstream, if the sub-picture approach is applied, constraints must be met. b. Sub-pictures may contain only rectangular slices. For example, in conforming bitstreams, when the sub-picture approach is applied, constraints must be met. 12. Two sub-pictures are constrained to not be overlapping. In one example, in a conforming bitstream, if the sub-picture approach is applied, constraints must be met. b. Alternatively, the two sub-pictures may overlap each other. 13. It is constrained that any position in a picture must be covered by only one sub-picture. In one example, in a conforming bitstream, if the sub-picture approach is applied, constraints must be met. b. Alternatively, one sample may not belong to any sub-picture. c. Alternatively, one ample may belong to multiple sub-pictures. 14. Sub-pictures defined in an SPS that are mapped to all resolutions presented in the same sequence may be constrained to follow the constrained positions and / or sizes described above. a. In one example, the width and height of a sub-picture defined in an SPS mapped to a resolution presented in the same sequence should be an integer multiple of N luma samples (such as 8, 16, 32). b. In one example, sub-pictures may be defined for a particular layer and may be mapped to other layers. For example, a sub-picture may be defined for the layer with the highest resolution in the sequence. ii. For example, a sub-picture may be defined for the layer with the lowest resolution in the sequence. iii. The layer for which the sub-picture is defined may be signaled in the SPS / VPS / PPS / PPS / slice header. c. In one example, when both sub-pictures and different resolutions are applied, all resolutions (eg, width or height) may be integer multiples of a given resolution. d. In one example, the width and / or height of a sub-picture defined in an SPS may be an integer multiple (eg, M) of the CTU size. e. Alternatively, sub-pictures and different resolutions within a sequence may not be allowed at the same time. 15. Subpictures may only be applied to specific layers. In one example, sub-pictures defined in an SPS may only be applied to the highest resolution layer in a sequence. b. In one example, sub-pictures defined in an SPS may only apply to the layer with the lowest temporal ID in the sequence. c. To which layer a sub-picture may be applied may be indicated by one or more syntax elements in the SPS / VPS / PPS. d. The layers to which sub-pictures cannot be applied may be indicated by one or more syntax elements in the SPS / VPS / PPS. 16. In one example, the position and / or dimensions of a subpicture may be signaled without using subpic_grid_idx. In one example, the top left position of the sub-picture may be signaled. b. In one example, the bottom right position of the sub-picture may be signaled. c. In one example, the width of the sub-picture may be signaled. d. In one example, the height of the sub-picture may be signaled. 17. For a temporal filter, when performing temporal filtering of samples, only samples in the same subpicture to which the current sample belongs may be used. The required samples may be in the same picture to which the current sample belongs, or in other pictures. 18. In one example, whether and / or how to apply a partitioning method (QT, horizontal BT, vertical BT, horizontal TT, vertical TT, no split, etc.) may depend on whether the current block (or partition) crosses one or more boundaries of a sub-picture. a. In one example, the picture boundary handling method for partitioning in VVC can also be applied when picture boundaries are replaced with sub-picture boundaries. b. In one example, whether to parse a syntax element (e.g., a flag) representing a partitioning method (QT, horizontal BT, vertical BT, horizontal TT, vertical TT, no split, etc.) may depend on whether the current block (or partition) crosses one or more boundaries of a subpicture. 19. Instead of dividing a picture into multiple sub-pictures with independent coding of each sub-picture, it is proposed to divide the picture into at least two sets of sub-regions, the first set containing multiple sub-pictures and the second set containing all the remaining samples. In one example, the samples in the second set are not in any subpicture. b. Alternatively, the second set may also be encoded / decoded based on the first set of information. c. In one example, a default value may be utilized to mark whether a sample / MxK sub-region belongs to the second set. i. In one example, the default value may be set equal to (max_subpics_minus1+K), where K is an integer greater than 1. ii. A default value is assigned to subpic_grid_idx[i][j] to indicate that the grid belongs to the second set. 20. It is proposed that the syntax element subpic_grid_idx[i][j] cannot be larger than max_subpics_minus1. a. For example, in a conforming bitstream, the constraint is that subpic_grid_idx[i][j] cannot be greater than max_subpics_minus1. b. For example, the codeword that codes subpic_grid_idx[i][j] cannot be larger than max_subpics_minus1. 21. It is proposed that any integer between 0 and max_subpics_minus1 must equal at least one subpic_grid_idx[i][j]. 22. The IBC virtual buffer may be reset before decoding a new sub-picture within a picture. In one example, all samples in the IBC virtual buffer may be reset to -1. 23. The palette entry list may be reset before decoding a new sub-picture within a picture. In one example, PredictorPaletteSize may be set equal to 0 before decoding a new sub-picture within a picture. 24. Whether to signal slice information (eg, the number of slices and / or the range of slices) may depend on the number of tiles and / or the number of bricks. a. In one example, if the number of bricks in a picture is 1, then num_slices_in_pic_minus1 is not signaled and is inferred to be 0. b. In one example, if the number of bricks in a picture is 1, slice information (e.g., the number of slices and / or the range of the slices) may not need to be signaled. c. In one example, if the number of bricks in a picture is 1, the number of slices may be inferred to be 1. A slice covers the entire picture. In one example, if the number of bricks in a picture is 1, single_brick_per_slice_flag is not signaled and is inferred to be 0. i. Alternatively, if the number of bricks in a picture is 1, then single_brick_per_slice_flag is not signaled and is inferred to be 0. d. An exemplary syntax design is as follows: [Table 4] 25. Whether slice_address is signaled may be decoupled from whether the slice is signaled as rectangular (e.g., whether rect_slice_flag is 0 or 1). a. An exemplary syntax design is as follows: [Table 5] 26.Whether to signal slice_address may depend on the number of slices when the slice is signaled as a rectangle. [Table 6] 27.Whether to signal num_bricks_in_slice_minus1 may depend on slice_address and / or the number of bricks in the picture. a. An exemplary syntax design is as follows: [Table 7] 28.Whether to signal loop_filter_across_bricks_enabled_flag may depend on the number of tiles and / or the number of bricks in the picture. In one example, if the number of bricks is less than 2, the loop_filter_across_bricks_enabled_flag is not signaled. b. An exemplary syntax design is as follows: [Table 8] 29. A requirement for bitstream conformance is that all slices of a picture must cover the entire picture. a. This requirement must be met when the slice is signaled to be rectangular (e.g., when rect_slice_flag is equal to 1). 30. A requirement for bitstream conformance is that all slices of a subpicture must cover the entire subpicture. a. This requirement must be met when the slice is signaled to be rectangular (e.g., when rect_slice_flag is equal to 1). 31. A requirement for bitstream compatibility is that a slice cannot overlap multiple sub-pictures. 32. A requirement for bitstream compatibility is that tiles cannot overlap with multiple sub-pictures. 33. A requirement for bitstream compatibility is that a brick cannot overlap multiple subpictures. In the following discussion, a basic unit block (BUB) having dimensions CW×CH is a rectangular region. For example, a BUB may be a coding tree block (CTB). 34. In one example, the number of sub-pictures (denoted as N) may be signaled. a. For conformance bitstreams, if subpictures are used (eg, if subpics_present_flag is equal to 1), there may be required to be at least two subpictures in a picture. b. Alternatively, N minus d (ie, Nd) may be signaled, where d is an integer such as 0, 1, or 2. c. For example, Nd may be coded with a fixed length coding, eg, u(x). In one example, x may be a fixed number, such as 8. ii. In one example, x or x-dx may be signaled before Nd is signaled, where dx is an integer such as 0, 1, or 2. The signaled x may not be greater than the maximum value of a conforming bitstream. iii. In one example, x may be derived on the fly. 1) For example, x may be derived as a function of the total number of BUBs in the picture (denoted as M), e.g., x = Ceil(log2(M + d0)) + d1, where d0 and d1 are two integers, such as -2, -1, 0, 1, 2, etc. 2) M may be derived as M=Ceiling(W / CW)×Ceiling(H / CH), where W and H represent the width and height of the picture, and CW and CH represent the width and height of the BUB. d. For example, Nd may be coded with a unary code or a truncated unary code. e. In one example, the maximum allowed value for Nd may be a fixed number. i. Alternatively, the maximum allowed value of Nd may be derived as a function of the total number of BUBs in the picture (denoted as M), for example, x = Ceil(log2(M + d0)) + d1, where d0 and d1 are two integers, such as -2, -1, 0, 1, 2, etc. 35. In one example, a subpicture may be signaled by an indication of its selected position (e.g., top left / top right / bottom left / bottom right) and / or one or more of its width and / or height. a. In one example, the location of the top left corner of a sub-picture may be signaled at the granularity of a basic unit block having dimensions CW×CH. i. For example, the column index (denoted as Col) for the BUB of the top left BUB of the subpicture may be signaled. 1) For example, Col-d may be signaled, where d is an integer such as 0, 1, or 2. a) Alternatively, d may be equal to Col of the previously coded sub-picture added by d1, where d1 is an integer such as -1, 0 or 1. b) The code of Col-d may be signaled. ii. For example, the row index (denoted as Row) for the BUB of the top left BUB of the sub-picture may be signaled. 1) For example, Row-d may be signaled, where d is an integer such as 0, 1, or 2. a) Alternatively, d may be equal to the Row of the previously coded sub-picture added by d1, where d1 is an integer such as -1, 0 or 1. b) The code of Row-d may be signaled. iii. The above-mentioned row / column index (denoted as Row) may be represented in a CTB (Codling Tree Block), e.g., the x or y coordinate for the top left position of the picture may be divided by the CTB size and signaled. iv. In one example, whether to signal the location of a sub-picture may depend on the sub-picture index. 1) In one example, the top left position may not be signaled for the first sub-picture in a picture. a) Alternatively, the top left position may also be inferred to be, for example, (0,0). 2) In one example, the top left position may not be signaled for the last sub-picture in a picture. a) The top left position may be inferred depending on previously signaled sub-picture information. b. In one example, the indication of the width / height / selected position of the subpicture may be signaled in truncated unary / truncated binary / unary / fixed length / K-th EG coding (e.g., K=0,1,2,3). c. In one example, the width of a sub-picture may be signaled at the granularity of a BUB with dimensions CW×CH. i. For example, the number of columns of BUBs in a sub-picture (denoted as W) may be signaled. ii. For example, Wd may be signaled, where d is an integer such as 0, 1, or 2. 1) Alternatively, d may be equal to W of the previously coded sub-picture added by d1, where d1 is an integer such as -1, 0 or 1. 2) The sign of Wd may be signaled. d. In one example, the height of a sub-picture may be signaled at the granularity of a BUB with dimensions CW x CH. i. For example, the number of rows of BUBs in a sub-picture (denoted as H) may be signaled. ii. For example, Hd may be signaled, where d is an integer such as 0, 1, or 2. 1) Alternatively, d may be equal to H of the previously coded sub-picture added by d1, where d1 is an integer such as -1, 0 or 1. 2) The sign of Hd may be signaled. e. In one example, Col-d may be coded with a fixed length coding, eg, u(x). In one example, x may be a fixed number, such as 8. ii. In one example, x or x-dx may be signaled before Col-d is signaled, where dx is an integer such as 0, 1, or 2. The signaled x may not be greater than the maximum value of a conforming bitstream. iii. In one example, x may be derived on the fly. 1) For example, x may be derived as a function of the total number of BUB columns in the picture (denoted as M), e.g., x = Ceil(log2(M + d0)) + d1, where d0 and d1 are two integers, such as -2, -1, 0, 1, 2, etc. 2) M may be derived as M=Ceiling(W / CW), where W represents the width of the picture and CW represents the width of the BUB. f. In one example, Row-d may be coded with a fixed length coding, eg, u(x). In one example, x may be a fixed number, such as 8. ii. In one example, x or x-dx may be signaled before Row-d is signaled, where dx is an integer such as 0, 1, or 2. The signaled x may not be greater than the maximum value of a conforming bitstream. iii. In one example, x may be derived on the fly. 1) For example, x may be derived as a function of the total number of BUB rows in the picture (denoted as M), e.g., x = Ceil(log2(M + d0)) + d1, where d0 and d1 are two integers, such as -2, -1, 0, 1, 2, etc. 2) M may be derived as M=Ceiling(H / CH), where H represents the height of the picture and CH represents the height of the BUB. g. In one example, Wd may be coded with a fixed length coding, eg, u(x). In one example, x may be a fixed number, such as 8. ii. In one example, x or x-dx may be signaled before Wd is signaled, where dx is an integer such as 0, 1, or 2. The signaled x may not be greater than the maximum value of a conforming bitstream. iii. In one example, x may be derived on the fly. 1) For example, x may be derived as a function of the total number of BUB columns in the picture (denoted as M), e.g., x = Ceil(log2(M + d0)) + d1, where d0 and d1 are two integers, such as -2, -1, 0, 1, 2, etc. 2) M may be derived as M=Ceiling(W / CW), where W represents the width of the picture and CW represents the width of the BUB. h. In one example, Hd may be coded with a fixed length coding, eg, u(x). In one example, x may be a fixed number, such as 8. ii. In one example, x or x-dx may be signaled before Hd is signaled, where dx is an integer such as 0, 1, or 2. The signaled x does not have to be greater than the maximum value for a conforming bitstream. iii. In one example, x may be derived on the fly. 1) For example, x may be derived as a function of the total number of BUB rows in the picture (denoted as M), e.g., x = Ceil(log2(M + d0)) + d1, where d0 and d1 are two integers, such as -2, -1, 0, 1, 2, etc. 2) M may be derived as M=Ceiling(H / CH), where H represents the height of the picture and CH represents the height of the BUB. i. Col-d and / or Row-d may be signaled for all sub-pictures. i. Alternatively, Col-d and / or Row-d may not be signaled for all sub-pictures. 1) Col-d and / or Row-d may not be signaled if the number of sub-pictures is less than 2 (equal to 1). 2) For example, Col-d and / or Row-d may not be signaled for the first sub-picture (eg, sub-picture index (or sub-picture ID) equals 0). a) When not signaled, they may be inferred to be 0. 3) For example, Col-d and / or Row-d may not be signaled for the last sub-picture (eg, sub-picture index (or sub-picture ID) equals NumSubPics-1). a) When not signaled, they may be inferred depending on the position and dimensions of sub-pictures that have already been signaled. jW-d and / or Hd may be signaled for all sub-pictures. i. Alternatively, Wd and / or Hd may not be signaled for all sub-pictures. 1) Wd and / or Hd may not be signaled if the number of sub-pictures is less than 2 (equal to 1). 2) For example, Wl-d and / or Hd may not be signaled for the last sub-picture (eg, sub-picture index (or sub-picture ID) equals NumSubPics-1). a) When not signaled, they may be inferred depending on the position and dimensions of sub-pictures that have already been signaled. k. In the above items, the BUB may be a CTB (Coding Tree Block). 36. In one example, sub-picture information should be signaled after CTB size information (e.g., log2_ctu_size_minus5) has already been signaled. 37. subpic_treated_as_pic_flag[i] does not have to be signaled for each subpicture. Instead, one subpic_treated_as_pic_flag is signaled to control whether a subpicture is treated as a picture for all subpictures. 38. loop_filter_across_subpic_enabled_flag[i] does not have to be signaled for each subpicture, instead one loop_filter_across_subpic_enabled_flag is signaled for all subpictures to control whether the loop filter can be applied across subpictures. 39. subpic_treated_as_pic_flag[i] (subpic_treated_as_pic_flag) and / or loop_filter_across_subpic_enabled_flag[i] (loop_filter_across_subpic_enabled_flag) may be conditionally signaled. a. In one example, subpic_treated_as_pic_flag[i] and / or loop_filter_across_subpic_enabled_flag[i] may not be signaled if the number of subpictures is less than 2 (equal to 1). 40. RPR may be applied when sub-pictures are used. a. In one example, the scaling ratio in the RPR may be constrained to be set exclusively when sub-pictures are used, such as {1:1, 1:2 and / or 2:1} or {1:1, 1:2 and / or 2:1, 1:4 and / or 4:1}, {1:1, 1:2 and / or 2:1, 1:4 and / or 4:1, 1:8 and / or 8:1}. b. In one example, if picture A and picture B have different resolutions, the CTB size of picture A and the CTB size of picture B may be different. c. In one example, suppose that picture A has sub-picture SA of dimensions SAW x SAH and picture B has sub-picture SB of dimensions SBW x SBH, where SA corresponds to SB, and the scaling ratios between picture A and picture B are Rw and Rh along the horizontal and vertical directions: i.SAW / SBW or SBW / SAW must be equal to Rw. ii. SAH / SBH or SBH / SAH must be equal to Rh. 41. When subpictures are used (e.g., when sub_picts_present_flag is true), the subpicture index (or subpicture ID) may be signaled in the slice header, and the slice address is interrupted as an address in the subpicture rather than the entire picture.
[0118] 5. Implementation
[0119] In the following embodiment, newly added text is in bold italics, and deleted text is marked with [[]].
[0120] 5.1 Embodiment 1: Sub-picture Constraints for Affinely Structured Merge Candidates 8.5.5.6 Derivation Process for Constructed Affine Control Point Motion Vector Merge Candidates The inputs to this process are: - Luma position (xCb, yCb) that specifies the top-left sample of the current luma coding block relative to the top-left luma sample of the current picture - Two variables cbWidth and cbHeight that specify the width and height of the current luma coding block - Availability flags availableA0, availableA1, availableA2, availableB0, availableB1, availableB2, availableB3 - Sample positions (xNbA0,yNbA0), (xNbA1,yNbA1), (xNbA2,yNbA2), (xNbB0,yNbB0), (xNbB1,yNbB1), (xNbB2,yNbB2), and (xNbB3,yNbB3) The output of this process is as follows: - Availability flag of configured affine control point motion vector merge candidates availableFlagConstK(K=1..6) - Reference index refIdxLXConstK (K=1..6, X is 0 or 1) - Prediction list usage flag predFlagLXConstK (K=1.6, X is 0 or 1), - Affine motion model index ModelIdcConstK (K=1..6) - Biprediction weight index bcwIdxConstK (K=1..6) - The constructed affine control point motion vector cpMvLXConstK[cpIdx] (cpIdx=0..2, K=1..6, X is 0 or 1) ... The fourth (juxtaposed bottom-right) control point motion vector cpMvLXCorner[3], reference index refIdxLXCorner[3], prediction list usage flag predFlagLXCorner[3], and availability flag avalableFlagCorner[3] (X is 0 and 1) are derived as follows: - Reference index for temporal merge candidate refIdxLXCorner[3] (X is 0 or 1), set equal to 0. The variables mvLXCol and availableFlagLXCol (where X is 0 or 1) are derived as follows: If slice_temporal_mvp_enabled_flag is equal to 0, then both components of mvLXCol are set equal to 0 and availableFlagLXCol is set equal to 0. - Otherwise (rect_slice_flag is equal to 1), the following applies:
number
[0121] 5.2 Embodiment 2: Sub-picture Constraints for Affinely Structured Merging Candidates 8.5.5.6 Derivation Process for Constructed Affine Control Point Motion Vector Merge Candidates The inputs to this process are: - Luma position (xCb, yCb) that specifies the top-left sample of the current luma coding block relative to the top-left luma sample of the current picture - Two variables cbWidth and cbHeight that specify the width and height of the current luma coding block - Availability flags availableA0, availableA1, availableA2, availableB0, availableB1, availableB2, availableB3 - Sample positions (xNbA0,yNbA0), (xNbA1,yNbA1), (xNbA2,yNbA2), (xNbB0,yNbB0), (xNbB1,yNbB1), (xNbB2,yNbB2), and (xNbB3,yNbB3) The output of this process is as follows: - Availability flag of configured affine control point motion vector merge candidates availableFlagConstK(K=1..6) - Reference index refIdxLXConstK (K=1..6, X is 0 or 1) - Prediction list usage flag predFlagLXConstK (K=1.6, X is 0 or 1), - Affine motion model index ModelIdcConstK (K=1..6) - Biprediction weight index bcwIdxConstK (K=1..6) - The constructed affine control point motion vector cpMvLXConstK[cpIdx] (cpIdx=0..2, K=1..6, X is 0 or 1) ... The fourth (juxtaposed bottom-right) control point motion vector cpMvLXCorner[3], reference index refIdxLXCorner[3], prediction list usage flag predFlagLXCorner[3], and availability flag avalableFlagCorner[3] (X is 0 and 1) are derived as follows: - Reference index for temporal merge candidate refIdxLXCorner[3] (X is 0 or 1), set equal to 0. The variables mvLXCol and availableFlagLXCol (where X is 0 or 1) are derived as follows: If slice_temporal_mvp_enabled_flag is equal to 0, then both components of mvLXCol are set equal to 0 and availableFlagLXCol is set equal to 0. - Otherwise (rect_slice_flag is equal to 1), the following applies:
number
[0122] 5.3 Embodiment 3: Fetching Integer Samples Under Subpicture Constraints 8.5.6.3.3 Luma Integer Sample Fetch Process The inputs to this process are: - Luma position in full sample units (xInt L ,yInt L ) - Luma reference sample array refPicLXL The output of this process is the predicted luma sample value predSampleLXL. The variable shift is set equal to Max(2,14-BitDepthY). The variable picW is set equal to pic_width_in_luma_samples, and the variable picH is set equal to pic_height_in_luma_samples. The luma position (xInt, yInt) in full sample units is derived as follows: [Outside 2] TIFF0007783383000046.tif14168
number
number
number
[0123] 5.4 Embodiment 4: Derivation of the variable invAvgLuma in chroma residual scaling for LMCS 8.7.5.3 Image reconstruction by luma-dependent chroma residual scaling process for chroma samples The inputs to this process are: - Chroma position (xCurr, yCurr) of the top-left sample of the current chroma transformation block relative to the top-left chroma sample of the current picture - nCurrSw variable that specifies the chroma conversion block width - nCurrSh variable that specifies the chroma conversion block height - The variable tuCbfChroma that specifies the coded block flag of the current chroma transformation block - (nCurrSw) x (nCurrSh) array predSamples specifying the chroma prediction samples for the current block - resamples, a (nCurrSw) x (nCurrSh) array specifying the chroma residual samples for the current block The output of this process is the reconstructed chroma picture sample array recSamples The variable sizeY is set equal to Min(CtbSizeY,64). The reconstructed chroma picture samples recSamples are derived as follows for i=0..nCurrSw-1, j=0..nCurrSh-1: - ... - Otherwise, the following applies: - ... The variable cursPic specifies the array of reconstructed luma samples in the current picture. For the derivation of the variable varScale, the following ordered steps are applied: 1. The variable invAvgLuma is derived as follows: The array recLuma[i] (i=0..(2*sizeY-1)) and the variable cnt are derived as follows: The variable cnt is set equal to 0. [Outside 4] TIFF0007783383000051.tif15168
number
number
number
[0124] 5.5 Embodiment 5: Example of defining subpicture elements in units of N samples (N=8, 32, etc.) other than four samples 7.4.3.3 Sequence Parameter Set RBSP Semantics subpic_grid_col_width_minus1+1 specifies the width of each element in the subpicture identifier grid. [Outside 7] TIFF0007783383000057.tif8130Specified in sample units. The length of the syntax element is Ceil(Log2(pic_width_max_in_luma_samples / [Outside 8] TIFF0007783383000058.tif8130)) bits. The variable NumSubPicGridCols is derived as follows:
number
number
number
[0125] 5.6 Embodiment 6: Picture width and picture height are limited to 8 or more. 7.4.3.3 Sequence Parameter Set RBSP Semantics pic_width_max_in_luma_samples specifies the maximum width, in units of luma samples, of each decoded picture that references an SPS. pic_width_max_in_luma_samples must not be equal to 0 and must be greater than or equal to [[MinCbSizeY]]. [Outside 10] It must be an integer multiple of TIFF0007783383000063.tif11130. pic_height_max_in_luma_samples specifies the maximum height, in units of luma samples, of each decoded picture that references the SPS. pic_height_max_in_luma_samples must not be equal to 0 and must be greater than or equal to [[MinCbSizeY]] [Outside 11] It must be an integer multiple of TIFF0007783383000064.tif10130.
[0126] 5.7 Implementation 7: Subpicture Boundary Check for BT / TT / QT Split, BT / TT / QT Depth Derivation, and / or CU Split Flag Signaling 6.4.2 Allowed Binary Splitting Processes The variable allowBtSplit is derived as follows: - ... Otherwise, if all of the following conditions are true, allowBtSplit is set equal to FALSE: - btSplit equals SPLIT_BT_VER - y0+cbHeight is [[pic_height_in_luma_samples]] [Outside 12] Larger than TIFF0007783383000065.tif16168. Otherwise, if all of the following conditions are true, allowBtSplit is set equal to FALSE: - btSplit is equal to SPLIT_BT_VER. - cbHeight is greater than MaxTbSizeY. - x0+cbWidth is [[pic_width_in_luma_samples]] [Outside 13] Larger than TIFF0007783383000066.tif14167. Otherwise, if all of the following conditions are true, allowBtSplit is set equal to FALSE: - btSplit equals SPLIT_BT_HOR. - cbWidth is greater than MaxTbSizeY. - y0+cbHeight is [[pic_height_in_luma_samples]] [Outside 14] Larger than TIFF0007783383000067.tif14167. Otherwise, if all of the following conditions are true, allowBtSplit is set equal to FALSE: - x0+cbWidth is [[pic_width_in_luma_samples]] [Outside 15] Larger than TIFF0007783383000068.tif14167. - y0+cbHeight is [[pic_height_in_luma_samples]] [Outside 16] Larger than TIFF0007783383000069.tif15167. - cbWidth is larger than minQtSize. Otherwise, if all of the following conditions are true, allowBtSplit is set equal to FALSE: - btSplit equals SPLIT_BT_HOR. - x0+cbWidth is [[pic_width_in_luma_samples]] [Outside 17] Larger than TIFF0007783383000070.tif15167. -y0+cbHeight is [[pic_height_in_luma_samples]] [Outside 18] Smaller than TIFF0007783383000071.tif15167. 6.4.2 Allowed Ternary Split Processes The variable allowTtSplit is derived as follows: allowTtSplit is set equal to FALSE if one or more of the following conditions are true: - cbSize is smaller than 2*MinTtSizeY. - cbWidth is greater than Min(MaxTbSizeY,maxTtSize). - cbHeight is greater than Min(MaxTbSize,maxTtSize). - mttDepth is greater than or equal to maxMttDepth. - x0+cbWidth is [[pic_width_in_luma_samples]] [Outside 19] Larger than TIFF0007783383000072.tif15167. - y0+cbHeight is [[pic_height_in_luma_samples]] [Outside 20] Larger than TIFF0007783383000073.tif14167. - treeType is equal to DUAL_TREE_CHROMA and (cbWidth / SubWidthC)*(cbHeight / SubHeightC) is less than or equal to 32. - treeType is equal to DUAL_TREE_CHROMA and modeType is equal to MODE_TYPE_INTRA. Otherwise, allowTtSplit is set equal to TRUE. 7.3.8.2 Coding Tree Unit Syntax [Table 9] 7.3.8.4 Coding Tree Syntax [Table 10]
[0127] 5.8 Embodiment 8: Example of subpicture definition [Table 11]
[0128] 5.9 Embodiment 9: Example of Subpicture Definition [Table 12]
[0129] 5.10 Embodiment 10: Example of subpicture definition [Table 13]
[0130] 5.11 Embodiment 11: Example of subpicture definition [Table 14]
[0131] 3 is a block diagram of a video processing device 300. The device 300 may be used to implement one or more of the methods described herein. The device 300 may be embodied in a smartphone, a tablet, a computer, an Internet of Things (IoT) receiver, etc. The device 300 may include one or more processors 302, one or more memories 304, and video processing hardware 306. The processor 302 may be configured to implement one or more methods described herein. The memory(s) 304 may be used to store data and code used to implement the methods and techniques described herein. The video processing hardware 306 may be used to implement some of the techniques described herein in a hardware circuit.
[0132] FIG. 4 is a flow chart of a method 400 for processing video. The method 400 includes determining (402) for a video block in a first video region of a video whether a location at which a temporal motion vector predictor is determined for a transformation between the video block and a bitstream representation of the current video block using an affine mode is within a second video region, and performing the transformation based on the determination (404).
[0133] The following solution may be implemented as a preferred solution in some embodiments.
[0134] The following solutions may be implemented in conjunction with additional techniques described in the items listed in the previous section (e.g., item 1).
[0135] 1. A method of video processing, comprising: determining, for a video block in a first video region of a video, whether a location at which a temporal motion vector predictor is determined for a transformation between the video block and a bitstream representation of the current video block using an affine mode is within a second video region; and performing the transformation based on the determination.
[0136] 2. The method according to solution 1, wherein the video block is covered by a first region and a second region.
[0137] 3. A method according to any one of solutions 1 to 2, wherein if the position of the temporal motion vector predictor is outside the second video region, the temporal motion vector predictor is marked as unavailable and is not used in the transformation.
[0138] The following solutions may be implemented in conjunction with additional techniques described in the items listed in the previous section (e.g., item 2).
[0139] 4. A method of video processing, comprising: determining, for a video block in a first video region of a video, whether a location from which an integer sample in a reference picture is fetched for a conversion between the video block and a bitstream representation of the current video block is within a second video region, wherein the reference picture is not used in an interpolation process during the conversion; and performing the conversion based on the determination.
[0140] 5. The method according to solution 4, wherein the video block is covered by a first region and a second region.
[0141] 6. The method according to any one of Solutions 4 to 5, wherein if the position of the sample is outside the second video region, the sample is marked as unavailable and is not used in the transformation.
[0142] The following solutions may be implemented in conjunction with additional techniques described in the items listed in the previous section (e.g., item 3).
[0143] 7. A method comprising: for a video block in a first video region of a video, determining whether a location from which reconstructed luma sample values are fetched for a conversion between the video block and a bitstream representation of the current video block is within a second video region; and performing the conversion based on the determination.
[0144] 8. The method of solution 7, where the luma sample is covered by the first region and the second region.
[0145] 9. The method according to any one of Solutions 7 to 8, wherein if the position of the luma sample is outside the second video area, the luma sample is marked as unavailable and is not used in the transformation.
[0146] The following solutions may be implemented in conjunction with additional techniques described in the items listed in the previous section (e.g., item 4).
[0147] 10. A method comprising: for a video block in a first video region of a video, determining whether a location within a second video region where a check regarding splitting, depth derivation, or split flag signaling for the video block is performed during conversion between the video block and a bitstream representation of the current video block is located; and performing the conversion based on the determination.
[0148] 11. The method of solution 10, wherein the location is covered by a first region and a second region.
[0149] 12. The method according to any one of Solutions 10 to 11, wherein if the position is outside the second video region, the luma sample is marked as unavailable and is not used in the transformation.
[0150] The following solutions may be implemented in conjunction with additional techniques described in the items listed in the previous section (e.g., item 8).
[0151] 13. A method of video processing, comprising performing a conversion between a video including one or more video pictures including one or more video blocks and a coded representation of the video, wherein the coded representation complies with coding syntax requirements that the conversion does not use sub-picture coding / decoding and dynamic resolution coding / decoding tools, or intra-video unit reference picture resampling tools.
[0152] 14. The method according to solution 13, wherein a video unit corresponds to a sequence of one or more video pictures.
[0153] 15. The method according to any one of solutions 13-14, wherein the dynamic resolution conversion coding / decoding tool comprises an adaptive resolution conversion coding / decoding tool.
[0154] 16. The method according to any one of solutions 13-14, wherein the dynamic resolution conversion coding / decoding tool comprises a dynamic resolution conversion coding / decoding tool.
[0155] 17. The method according to any one of solutions 13 to 16, wherein the coded representation indicates that the video unit complies with coding syntax requirements.
[0156] 18. The method of solution 17, wherein the coded representation indicates that the video unit uses sub-picture coding.
[0157] 19. The method according to solution 17, wherein the coded representation indicates that the video unit uses dynamic resolution conversion coding / decoding tools or reference picture resampling tools.
[0158] The following solutions may be implemented in conjunction with additional techniques described in the items listed in the previous section (e.g., item 10).
[0159] 20. The method according to any one of solutions 1 to 19, wherein the second video area includes a video sub-picture and the boundary between the second video area and another video area is also the boundary between two coding tree units.
[0160] 21. The method according to any one of solutions 1 to 19, wherein the second video region includes a video sub-picture and the boundary between the second video region and another video region is also the boundary between two coding tree units.
[0161] The following solutions may be implemented in conjunction with additional techniques described in the items listed in the previous section (e.g., item 11).
[0162] 22. The method according to any one of solutions 1 to 21, wherein the first video area and the second video area comprise rectangular shapes.
[0163] The following solutions may be implemented in conjunction with additional techniques described in the items listed in the previous section (e.g., item 12).
[0164] 23. The method according to any one of solutions 1 to 22, wherein the first video area and the second video area do not overlap.
[0165] The following solutions may be implemented in conjunction with additional techniques described in the items listed in the previous section (e.g., item 13).
[0166] 24. The method according to any one of solutions 1 to 23, wherein the video picture is divided into video regions such that a pixel in the video picture is covered by only one video region.
[0167] The following solutions may be implemented in conjunction with additional techniques described in the items listed in the previous section (e.g., item 15).
[0168] 25. The method according to any one of solutions 1 to 24, wherein the video picture is split into a first video area and a second video area due to the video picture being present in a particular layer of the video sequence.
[0169] The following solutions may be implemented in conjunction with additional techniques described in the items listed in the previous section (e.g., item 10).
[0170] 26. A method of video processing, comprising: performing a conversion between a video including one or more video pictures including one or more video blocks and a coded representation of the video, wherein the coded representation complies with a coding syntax requirement that a first syntax element, subpic_grid_idx[i][j], is not greater than a second syntax element, max_subpics_minus1.
[0171] 27. The method of solution 26, wherein the codeword representing the first syntax element is not larger than the codeword representing the second syntax element.
[0172] 28. The method according to any one of solutions 1 to 27, wherein the first video region comprises a video subpicture.
[0173] 29. The method according to any one of solutions 1 to 28, wherein the second video area comprises a video sub-picture.
[0174] 30. The method according to any one of solutions 1 to 29, wherein the conversion comprises encoding the video into a coded representation.
[0175] 31. The method according to any one of solutions 1 to 29, wherein the conversion includes decoding the coded representation to generate pixel values of the video.
[0176] 32. A video decoding device including a processor configured to implement the methods defined in one or more of solutions 1 to 31.
[0177] 33. A video encoding device including a processor configured to implement the methods defined in one or more of solutions 1 to 31.
[0178] A computer program product having stored thereon computer code which, when executed by a processor, causes the processor to implement a method according to any one of solutions 1 to 31.
[0179] A method, apparatus or system described herein.
[0180] 8 is a block diagram illustrating an example video processing system 800 capable of implementing various techniques disclosed herein. Various implementations may include some or all of the components of system 800. System 800 may include an input 802 for receiving video content. The video content may be received in raw or uncompressed format, e.g., 8- or 10-bit multi-component pixel values, or may be in a compressed or encoded format. Input 802 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, passive optical network (PON), and wireless interfaces such as Wi-Fi or cellular interfaces.
[0181] System 800 may include a coding component 804 that may implement various coding or encoding methods described herein. The coding component 804 may reduce the average bitrate of the video from the input 802 to the output of the coding component 804 to generate a coded representation of the video. Thus, the coding techniques are sometimes referred to as video compression or video transcoding techniques. The output of the coding component 804 may be stored or transmitted via a connected communication, as represented by component 806. The stored or communicated bitstream (or coded) representation of the video received at the input 802 may be used by component 808 to generate pixel values or displayable video that are sent to a display interface 810. The process of generating user-viewable video from the bitstream representation is sometimes referred to as video decompression. Furthermore, while certain video processing operations are referred to as “coding” operations or tools, it will be understood that the coding tools or operations are used in an encoder, or that corresponding decoding tools or operations that reverse the results of the coding are performed by a decoder.
[0182] Examples of peripheral bus interfaces or display interfaces include universal serial bus (USB), high definition multimedia interface (HDMI), or display port, etc. Examples of storage interfaces include serial advanced technology attachment (SATA), PCI, IDE interfaces, etc. The techniques described herein may be embodied in various electronic devices such as mobile phones, laptops, smartphones, or other devices capable of performing digital data processing and / or video display.
[0183] FIG. 9 is a block diagram illustrating an example video coding system 100 that can utilize the techniques of this disclosure.
[0184] 9, video coding system 100 may include source device 110 and destination device 120. Source device 110 generates encoded video data, which may be referred to as a video encoding device. Destination device 120 may decode the encoded video data generated by source device 110, which may be referred to as a video decoding device.
[0185] The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.
[0186] The video source 112 may include a source such as a video capture device, an interface for receiving video data from a video content provider, and / or a computer graphics system for generating video data, or a combination of such sources. The video data may include one or more pictures. The video encoder 114 encodes the video data from the video source 112 to generate a bitstream. The bitstream may include a sequence of bits forming a coded representation of the video data. The bitstream may include coded pictures and associated data. A coded picture is a coded representation of a picture. The associated data may include sequence parameter sets, picture parameter sets, and other syntax structures. The input / output interface 116 may include a modulator / demodulator (modem) and / or a transmitter. The coded video data may be transmitted directly to the destination device 120 via the I / O interface 116 over the network 130a. The coded video data may also be stored on a storage medium / server 130b for access by the destination device 120.
[0187] The destination device 120 may include an I / O interface 126 , a video decoder 124 , and a display device 122 .
[0188] I / O interface 126 may include a receiver and / or a modem. I / O interface 126 may obtain encoded video data from source device 110 or storage medium / server 130b. Video decoder 124 may decode the encoded video data. Display device 122 may display the decoded video data to a user. Display device 122 may be integrated with destination device 120 or may be external to destination device 120 configured to interface with an external display device.
[0189] Video encoder 114 and video decoder 124 may operate according to a video compression standard, such as the High Efficiency Video Coding (HEVC) standard, the Versatile Video Coding (VVC) standard, and other current and / or future standards.
[0190] FIG. 10 is a block diagram illustrating an example of a video encoder 200, which may be video encoder 114 in system 100 shown in FIG.
[0191] Video encoder 200 may be configured to perform any or all of the techniques of this disclosure. In the example of FIG. 10, video encoder 200 includes multiple functional components. The techniques described in this disclosure may be shared among various components of video encoder 200. In some examples, a processor may be configured to perform any or all of the techniques described in this disclosure.
[0192] The functional components of the video encoder 200 may include a partition unit 201, a prediction unit 202, which may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra prediction unit 206, a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213, and an entropy coding unit 214.
[0193] In other examples, video encoder 200 may include more, fewer, or different functional components. In one example, prediction unit 202 may include an intra block copy (IBC) unit. The IBC unit may perform prediction in an IBC mode, where at least one reference picture is the picture in which the current video block is located.
[0194] Furthermore, some components, such as the motion estimation unit 204 and the motion compensation unit 205, may be highly integrated, but are depicted separately in the example of FIG. 5 for illustrative purposes.
[0195] Partition unit 201 may partition a picture into one or more video blocks. Video encoder 200 and video decoder 300 may support a variety of video block sizes.
[0196] The mode select unit 203 selects one of the coding modes, intra or inter, based on, for example, an error result, and provides the resulting intra- or inter-coded block to the residual generation unit 207 to generate residual block data and to the reconstruction unit 212 to reconstruct a coded block for use as a reference picture. In some examples, the mode select unit 203 may select a combined intra- and inter-prediction (CIIP) mode, in which prediction is based on an inter-prediction signal and an intra-prediction signal. The mode select unit 203 may also select a resolution (e.g., sub-pixel or integer-pixel precision) for the motion vector for the block in the case of inter-prediction.
[0197] To perform inter prediction on the current video block, motion estimation unit 204 may generate motion information for the current video block by comparing one or more reference frames from buffer 213 with the current video frame. Motion compensation unit 205 may determine a predicted video block for the current video block based on the motion information and decoded samples of pictures from buffer 213 other than the picture associated with the current video block.
[0198] Motion estimation unit 204 and motion compensation unit 205 may perform different operations on the current video block depending on, for example, whether the current video block is in an I slice, a P slice, or a B slice.
[0199] In some examples, motion estimation unit 204 may perform unidirectional prediction on the current video block, and motion estimation unit 204 may search reference pictures in list 0 or list 1 for a reference video block for the current video block. Motion estimation unit 204 may then generate a reference index indicating a reference picture in list 0 or list 1 that includes the reference video block and a motion vector that indicates a spatial displacement between the current video block and the reference video block. Motion estimation unit 204 may output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 205 may generate a predictive video block for the current block based on the reference video block indicated by the motion information of the current video block.
[0200] In another example, motion estimation unit 204 may perform bidirectional prediction on the current video block, and motion estimation unit 204 may search reference pictures in list 0 for a reference video block for the current video block and may also search reference pictures in list 1 for another reference video block for the current video block. Motion estimation unit 204 may then generate reference indexes indicating the reference pictures in lists 0 and 1 that include the reference video blocks and motion vectors that indicate spatial displacements between the reference video blocks and the current video block. Motion estimation unit 204 may output the reference index and motion vector of the current video block as motion information for the current video block. Motion compensation unit 205 may generate a predictive video block for the current block based on the reference video block indicated by the motion information of the current video block.
[0201] In some examples, the motion estimation unit 204 may output a full set of motion information for the decoding process of the decoder.
[0202] In some examples, motion estimation unit 204 may not output a complete set of motion information for the current video. Rather, motion estimation unit 204 may signal the motion information of the current video block by reference to motion information of other video blocks. For example, motion estimation unit 204 may determine that the motion information of the current video block is sufficiently similar to the motion information of neighboring video blocks.
[0203] In one example, motion estimation unit 204 may indicate, in a syntax structure associated with the current video block, a value that indicates to video decoder 300 that the current video block has the same motion information as another video block.
[0204] In another example, motion estimation unit 204 may identify another video block and a motion vector difference (NVD) in a syntax structure associated with the current video block. The motion vector difference indicates a difference between the motion vector of the current video block and the motion vector of the indicated video block. Video decoder 300 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[0205] As mentioned above, video encoder 200 may predictively signal motion vectors. Two examples of predictive signaling techniques that may be implemented by video encoder 200 include advanced motion vector prediction (AMVP) and merge mode signaling.
[0206] Intra prediction unit 206 may perform intra prediction on the current video block. When intra prediction unit 206 makes an intra prediction on the current video block, intra prediction unit 206 may generate predictive data for the current video block based on decoded samples of other video blocks in the same picture. The predictive data for the current video block may include a predictive video block and various syntax elements.
[0207] Residual generation unit 207 may generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) the prediction video block of the current video block from the current video block. The residual data for the current video block may include residual video blocks that correspond to different sample components of the samples in the current video block.
[0208] In other examples, for example in skip mode, residual data for the current video block may not exist and residual generation unit 207 may not perform a subtraction operation.
[0209] Transform processing unit 208 may generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video block associated with the current video block.
[0210] After transform processing unit 208 generates the transform coefficient video block associated with the current video block, quantization unit 209 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0211] Inverse quantization unit 210 and inverse transform unit 211 may apply inverse quantization and inverse transform, respectively, to the transform coefficient video block to reconstruct a residual video block from the transform coefficient video block. Reconstruction unit 212 may add the reconstructed residual video block to corresponding samples from one or more prediction video blocks generated by prediction unit 202 to generate a reconstructed video block related to the current block for storage in buffer 213.
[0212] After reconstruction unit 212 reconstructs the video blocks, a loop filtering operation may be performed to reduce video blocking artifacts within the video blocks.
[0213] The entropy encoding unit 214 may receive data from other functional components of the video encoder 200. When the entropy encoding unit 214 receives the data, the entropy encoding unit 214 performs one or more entropy encoding operations to generate entropy-coded data and outputs a bitstream that includes the entropy-coded data.
[0214] FIG. 11 is a block diagram illustrating an example of a video decoder 300, which may be video decoder 124 in system 100 shown in FIG.
[0215] Video decoder 300 may be configured to perform any or all of the techniques of this disclosure. In the example of FIG. 11, video decoder 300 includes multiple functional components. The techniques described in this disclosure may be shared among various components of video decoder 300. In some examples, a processor may be configured to perform any or all of the techniques described in this disclosure.
[0216] 11, video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. Video decoder 300 may, in some examples, perform a decoding path that is generally the reverse of the encoding path described with respect to video encoder 200 (e.g., FIG. 10).
[0217] The entropy decoding unit 301 may retrieve an encoded bitstream. The encoded bitstream may include entropy-coded video data (e.g., encoded blocks of video data). The entropy decoding unit 301 may decode the entropy-coded video data, and from the entropy-decoded video data, the motion compensation unit 302 may determine motion information including motion vectors, motion vector precision, reference picture list indexes, and other motion information. The motion compensation unit 302 may determine such information by, for example, implementing AMVP and merge mode.
[0218] The motion compensation unit 302 may generate motion-compensated blocks and may optionally perform interpolation based on an interpolation filter, and an identifier for the interpolation filter used with sub-pixel precision may be included in the syntax element.
[0219] Motion compensation unit 302 may calculate interpolated values for sub-integer pixels of the reference block using interpolation filters used by video encoder 200 during encoding of the video block. Motion compensation unit 302 determines the interpolation filters used by video encoder 200 according to received syntax information and uses the interpolation filters to generate the prediction block.
[0220] The motion compensation unit 302 may use some of the syntax information to determine the size of the blocks used to encode the frames and / or slices of the encoded video sequence, partitioning information describing how each macroblock of a picture of the encoded video sequence is partitioned, a mode indicating how each partition is coded, one or more reference frames (and reference frame lists) for each inter-coded block, and other information for decoding the encoded video sequence.
[0221] The intra prediction unit 303 may form a prediction block from spatially adjacent blocks, for example, using an intra prediction mode received in the bitstream. The inverse quantization unit 303 inverse quantizes, or dequantizes, the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 301. The inverse transform unit 303 applies an inverse transform.
[0222] Reconstruction unit 306 may sum the residual blocks with corresponding prediction blocks generated by motion compensation unit 202 or intra prediction unit 303 to form decoded blocks. If desired, a deblocking filter may be applied to filter the decoded blocks to remove blockiness artifacts. The decoded video blocks are then stored in buffer 307, which also provides reference blocks for subsequent motion compensation / intra prediction to generate decoded video for presentation on a display device.
[0223] 12 is a flowchart representation of a method for video processing in accordance with the present technology. The method 1200 includes, at operation 1210, performing a conversion between a video including one or more pictures and a bitstream representation of the video. The bitstream representation is required to conform to format rules that specify that each picture be coded as one or more slices. The format rules prohibit a sample in a picture from being uncovered by any of the one or more slices.
[0224] In some embodiments, the picture includes multiple subpictures, and the formatting rules further specify that multiple slices of a subpicture of the picture cover the entire subpicture. In some embodiments, the bitstream representation includes a syntax element indicating that the multiple slices have a rectangular shape. In some embodiments, a slice overlaps at most one subpicture of the picture. In some embodiments, a tile of video overlaps at most one subpicture of the picture. In some embodiments, a brick of video overlaps at most one subpicture of the picture.
[0225] 13 is a flowchart representation of a method for video processing in accordance with the present technology. The method 1300 includes, at operation 1310, determining a manner of signaling information for one or more slices in a picture according to rules related to the number of tiles or bricks in the picture for conversion between a picture of the video and a bitstream representation of the video. The method 1300 includes, at operation 1320, performing the conversion based on the determination.
[0226] In some embodiments, the information of one or more slices of a picture includes at least the number of one or more slices or the range of one or more slices. In some embodiments, the rule specifies that if the number of bricks in a picture is 1, the information of one or more slices in a picture is excluded in the bitstream representation. In some embodiments, the number of one or more slices in a picture is considered to be 1. In some embodiments, a slice covers the entire picture. In some embodiments, the number of bricks per slice is considered to be 1.
[0227] 14 is a flowchart representation of a method for video processing in accordance with the present technology. The method 1400 includes, at operation 1410, performing a conversion between a picture of a video and a bitstream representation of the video according to a rule. A picture is coded in the bitstream representation as one or more slices. The rule specifies whether or how addresses of slices of a picture are included in the bitstream representation.
[0228] In some embodiments, the rules specify that signaling the address of a slice is independent from signaling whether the slice has a rectangular shape. In some embodiments, the rules specify that signaling the address of a slice depends on the number of slices in a picture in the case of a slice as a rectangular shape. In some embodiments, whether to signal the number of bricks in a slice is based at least in part on the address of the slice. In some embodiments, whether to signal the number of bricks in a slice is further based on the number of bricks in a picture.
[0229] 15 is a flowchart representation of a method for video processing in accordance with the present technology. At operation 1510, method 1500 includes, for conversion between a picture of video and a bitstream representation of the video, determining whether a syntax element indicating a filter operation that accesses samples across multiple bricks in the picture is enabled to be included in the bitstream representation based on the number of tiles or bricks in the picture. At operation 1520, method 1500 includes performing the conversion based on the determination. In some embodiments, a single element is omitted in the bitstream representation if the number of bricks in the picture is less than two.
[0230] 16 is a flowchart representation of a method for video processing in accordance with the present technology. The method 1600 includes, at operation 1610, performing a conversion between a picture of the video and a bitstream representation of the video. The picture includes one or more subpictures, the number of which is indicated by a syntax element in the bitstream representation.
[0231] In some embodiments, the number of sub-pictures is represented as N, and the syntax element includes a value of (Nd), where N and d are integers. In some embodiments, d is 0, 1, or 2. In some embodiments, the value of (Nd) is coded using a unary coding scheme or a truncated unary coding scheme. In some embodiments, the value of (Nd) is coded to have a fixed length using a fixed-length coding scheme. In some embodiments, the fixed length is 8 bits. In some embodiments, the fixed length is represented as x-dx, where x and dx are positive integers, where x is less than or equal to a maximum value determined based on a matching rule, and dx is 0, 1, or 2. In some embodiments, x-dx is signaled. In some embodiments, the fixed length is determined based on the number of basic unit blocks in the picture. In some embodiments, the fixed length is represented as x, and the number of basic unit blocks is represented as M, where x = Ceil(log2(M+d0)) + d1, and d0 and d1 are integers.
[0232] In some embodiments, the maximum value of (Nd) is predetermined. In some embodiments, the number of basic unit blocks in a picture is represented as M, where M is an integer, and the maximum value of Nd is Md. In some embodiments, the number of basic unit blocks in a picture is represented as M, where M is an integer, and the maximum value of Nd is Md. In some embodiments, the basic unit block comprises a slice. In some embodiments, the basic unit block comprises a coding tree unit. At least one of d0 or d1 is −2, −1, 0, 1, or 2.
[0233] 17 is a flowchart representation of a method for video processing in accordance with the present technology. The method 1700 includes, at operation 1710, performing a conversion between a picture of a video including one or more subpictures and a bitstream representation of the video. The bitstream representation conforms to formatting rules that specify that information about the subpictures is included in the bitstream representation based on at least one of: (1) one or more corner locations of the subpictures; or (2) dimensions of the subpictures.
[0234] In some embodiments, the format rules further specify that information about the subpicture is placed after information about the size of the coding tree block in the bitstream representation. In some embodiments, the location of one or more corners of the subpicture is signaled in the bitstream representation using the granularity of the blocks in the subpicture. In some implementations, the block comprises an elementary unit block or a coding tree block.
[0235] In some embodiments, a block index is used to indicate a corner location of a subpicture. In some embodiments, the index includes a column index or a row index. In some embodiments, a block index is represented as I, and a value of (Id) is signaled in the bitstream representation, where d is 0, 1, or 2. In some embodiments, d is determined based on the index of a previously coded subpicture and another integer d1, where d1 is equal to −1, 0, or 1. In some embodiments, the sign of (Id) is signaled in the bitstream representation.
[0236] In some embodiments, whether the position of a subpicture is signaled in the bitstream representation is based on the subpicture's index. In some embodiments, if the subpicture's position is the first subpicture in a picture, the subpicture's position is omitted in the bitstream representation. In some embodiments, the subpicture's top-left position is determined to be (0,0). In some embodiments, if the subpicture's position is the last subpicture in a picture, the subpicture's position is omitted in the bitstream representation. In some embodiments, the subpicture's top-left position is determined based on information about previously transformed subpictures. In some implementations, the information about the subpicture includes the subpicture's top-left position and dimensions. The information is coded using a truncated unary coding scheme, a truncated binary coding scheme, a unary coding scheme, a fixed-length coding scheme, or a K-th EG coding scheme.
[0237] In some embodiments, the location of one or more corners of a subpicture is signaled in the bitstream representation using the granularity of blocks in the subpicture. In some embodiments, the dimension includes the width or height of the subpicture. In some embodiments, the dimension is expressed as the number of columns or the number of rows of blocks in the subpicture. In some embodiments, the index of a block is represented as I, and a value of (Id) is signaled in the bitstream representation, where d is 0, 1, or 2. In some embodiments, d is determined based on the index of a previously coded subpicture and another integer d1, where d1 is equal to −1, 0, or 1. In some embodiments, the sign of (Id) is signaled in the bitstream representation.
[0238] In some embodiments, the value of (Nd) is coded to have a fixed length using a fixed-length coding scheme. In some embodiments, the fixed length is 8 bits. In some embodiments, the fixed length is expressed as x-dx, where x and dx are positive integers. x is less than or equal to a maximum value determined based on a matching rule, and dx is 0, 1, or 2. In some embodiments, the fixed length is determined based on the number of basic unit blocks in the picture. In some implementations, the fixed length is expressed as x, the number of basic unit blocks is expressed as M, where x = Ceil(log2(M+d0)) + d1, and d0 and d1 are integers. At least one of d0 or d1 is -2, -1, 0, 1, or 2. In some embodiments, (Id) is signaled for all sub-pictures of a picture. In some embodiments, (Id) is signaled for a subset of sub-pictures of a picture. The value of (Id) is omitted in the bitstream representation if the number of subpictures is equal to 1. The value of (Id) is omitted in the bitstream representation if the subpicture is the first subpicture of the picture. The value of (Id) is determined to be 0. The value of (Id) is omitted in the bitstream representation if the subpicture is the last subpicture of the picture. The value of (Id) is determined based on information about previously transformed subpictures.
[0239] In some embodiments, a single syntax element is conditionally included in the bitstream representation to indicate whether all sub-pictures of a picture are treated as a picture. In some embodiments, a single syntax element is conditionally included in the bitstream representation to indicate whether a loop filter is applicable to all sub-pictures across the sub-pictures of a picture. In some embodiments, the single element is omitted in the bitstream representation if the number of sub-pictures of a picture is equal to 1.
[0240] 18 is a flowchart representation of a method for video processing in accordance with the present technology. The method 1800 includes, at operation 1810, determining that a reference picture resampling tool is enabled for converting between a picture of a video and a bitstream representation of the video by dividing the picture into one or more sub-pictures. The method 1800 also includes, at operation 1820, performing the conversion based on the determination.
[0241] In some embodiments, the scaling ratios used in the reference picture resampling coding tool are determined based on a set of ratios. In some embodiments, the set of ratios includes at least one of: {1:1, 1:2}, {1:1, 2:1}, {1:1, 1:2, 1:4}, {1:1, 1:2, 4:1}, {1:1, 2:1, 1:4}, {1:1, 2:1, 4:1}, {1:1, 1:2, 1:4, 1:8}, {1:1, 1:2, 1:4, 8:1}, {1:1, 1:2, 4:1, 1:8}, {1:1, 1:2, 4:1, 8:1}, {1:1, 2:1, 1:4, 1:8}, {1:1, 2:1, 1:4, 8:1}, {1:1, 2:1, 4:1, 1:8}, or {1:1, 2:1, 4:1, 8:1}. In some embodiments, when the resolution of a picture is different from the resolution of a second picture, the size of the coding tree block of the picture is different from the size of the coding tree block of the second picture. In some embodiments, a sub-picture of a picture has dimensions of SAW x SAH, and a sub-picture of a second picture of a video has dimensions of SBW x SBH. The scaling ratio between the picture and the second picture is Rw and Rh along the horizontal and vertical directions. SAW / SBW or SBW / SAW is Rw, and SAH / SBH or SBH / SAH is Rh.
[0242] 19 is a flowchart representation of a method for video processing in accordance with the present technology. The method 1900 includes, at act 1910, performing a conversion between a video picture including one or more sub-pictures that include one or more slices and a bitstream representation of the video. The bitstream representation conforms to format rules that specify that for sub-pictures and slices, if an index identifying the sub-picture is included in the slice's header, then the address field for the slice indicates the address of the slice within the sub-picture.
[0243] In some embodiments, the conversion generates a video from the bitstream representation.In some embodiments, the conversion generates a bitstream representation from the video.
[0244] Some embodiments of the disclosed techniques include making a determination or decision to enable a video processing tool or mode. In one example, when a video processing tool or mode is enabled, an encoder uses or implements the tool or mode in processing blocks of video, but may not necessarily modify the resulting bitstream based on the use of the tool or mode. That is, conversion from blocks of video to a bitstream representation of video uses the video processing tool or mode when enabled based on the determination or decision. In another example, when a video processing tool or mode is enabled, a decoder processes the bitstream knowing that the bitstream has been modified based on the video processing tool or mode. That is, conversion from a bitstream representation of video to blocks of video is performed using the video processing tool or mode enabled based on the determination or decision.
[0245] Some embodiments of the disclosed techniques include making a determination or decision to disable a video processing tool or mode. In one example, when a video processing tool or mode is disabled, an encoder does not use the tool or mode when converting blocks of video to a bitstream representation of video. In another example, when a video processing tool or mode is disabled, a decoder processes the bitstream knowing that the bitstream has not been modified using the video processing tool or mode that was enabled based on the determination or decision.
[0246] The disclosed and other solutions, examples, embodiments, modules, and functional operations described in this document can be implemented in digital electronic circuitry, or computer software, firmware, or hardware, including the structures disclosed in this document and their structural equivalents, or in one or more combinations of these. The disclosed and other embodiments can be implemented as one or more computer program products, e.g., one or more modules of computer program instructions encoded on a computer-readable medium for executing or controlling the operation of a data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition of matter affecting a machine-readable propagated signal, or a combination of one or more of these. The term "data processing apparatus" encompasses all apparatuses, devices, and machines for processing data, including, by way of example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, an apparatus can include code that creates an execution environment for the computer program in question, e.g., code constituting processor firmware, a protocol stack, a database management system, an operating system, or one or more combinations of these. A propagated signal is an artificially generated signal, for example, a machine-generated electrical, optical, or electromagnetic signal generated to encode information for transmission to a suitable receiver device.
[0247] A computer program (also known as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, and can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored as part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files storing one or more modules, subprograms, or portions of code). A computer program can be deployed to run on one computer or one site, or distributed across multiple sites and interconnected by a communications network.
[0248] The processes and logic flows described herein may be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows may also be performed by, and an apparatus may be implemented as, special purpose logic circuitry, such as an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit).
[0249] Processors suitable for executing a computer program include, by way of example, both general-purpose and special-purpose microprocessors, and any one or more processors of any kind of digital computer. Typically, a processor will receive instructions and data from a read-only memory or a random-access memory, or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include, or be operatively coupled to receive and / or transfer data from, one or more mass storage devices for storing data, e.g., magnetic, magneto-optical, or optical disks. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, by way of example, semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices; magnetic, magneto-optical, or optical disks, e.g., internal or removable hard disks; and CD-ROM and DVD-ROM disks. The processor and memory may be supplemented by, or incorporated in, special-purpose logic circuitry.
[0250] While this patent document contains many details, these should not be construed as limitations on the scope of any subject matter or what may be claimed, but rather as descriptions of features that may be specific to particular embodiments of a particular technique. Certain features that are described in this patent document in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable subcombination. Furthermore, while features may be described above as acting in a particular combination and initially claimed by themselves, one or more features from a claimed combination can, in some cases, be deleted from the combination, and the claimed combination may be directed to subcombinations or variations of the subcombination.
[0251] Similarly, although the figures depict acts in a particular order, this should not be understood as requiring such acts to be performed in a particular order or sequentially, or that all of the depicted acts be performed, to achieve desired results. Furthermore, the separation of various system components in the embodiments described in this patent document should not be understood as requiring such separation in all embodiments.
[0252] Only a few implementations and examples have been described, and other implementations, extensions and variations can be made based on what is described and shown in this patent document.
Claims
1. 1. A method of video processing, comprising: partitioning the picture into one or more slices and one or more sub-pictures according to bitstream compatibility requirements for conversion between a current video block of a picture of a video and a bitstream of the video; performing the transformation based at least on the partitioning; the bitstream conformance requirement specifies that the union of the one or more slices covers an entire picture; a sub-picture of the picture overlaps with at least one tile of the picture; A method, wherein a first syntax element is included in the bitstream that indicates whether the one or more slices have a rectangular shape.
2. The method of claim 1 , wherein the bitstream conformance requirement further specifies that the union of the one or more sub-pictures resulting from the partitioning of the picture covers the entire picture.
3. a sub-picture of the one or more sub-pictures is partitioned into one or more slices; The method of claim 2 , wherein the bitstream compatibility requirement further specifies that the union of the one or more sub-pictures resulting from the partitioning of the sub-picture covers the entire sub-picture.
4. The method of claim 2 , wherein one slice of the picture overlaps with at most one sub-picture of the picture.
5. The method of any one of claims 1 to 4, wherein the transforming comprises encoding the current video block into the bitstream.
6. The method of any one of claims 1 to 4, wherein the conversion comprises decoding from the bitstream to the current video block.
7. 1. An apparatus for processing video data, comprising: a processor; and a non-transitory memory having instructions, the instructions, when executed by the processor, causing the processor to: partitioning the picture into one or more slices and one or more sub-pictures according to bitstream compatibility requirements for conversion between a current video block of a picture of a video and a bitstream of the video; performing the transformation based at least on the partitioning; the bitstream conformance requirement specifies that the union of the one or more slices covers an entire picture; a sub-picture of the picture overlaps with at least one tile of the picture; 11. An apparatus, wherein a first syntax element is included in the bitstream that indicates whether the one or more slices have a rectangular shape.
8. A non-transitory computer-readable storage medium storing instructions, the instructions causing a processor to: partitioning the picture into one or more slices and one or more sub-pictures according to bitstream compatibility requirements for conversion between a current video block of a picture of a video and a bitstream of the video; performing the transformation based at least on the partitioning; the bitstream conformance requirement specifies that the union of the one or more slices covers an entire picture; a sub-picture of the picture overlaps with at least one tile of the picture; 10. A non-transitory computer-readable storage medium, wherein a first syntax element is included in the bitstream that indicates whether the one or more slices have a rectangular shape.
9. 1. A method for storing a video bitstream produced by a method performed by a video processing device, comprising: For a current video block of a picture of the video, partitioning the picture into one or more slices and one or more sub-pictures according to bitstream compatibility requirements; generating the bitstream based at least on the partitioning; storing the bitstream on a non-transitory computer-readable recording medium; the bitstream conformance requirement specifies that the union of the one or more slices covers an entire picture; a sub-picture of the picture overlaps with at least one tile of the picture; A method, wherein a first syntax element is included in the bitstream that indicates whether the one or more slices have a rectangular shape.