Coding of Pictures Containing Slices and Tiles
By deriving and signaling slice indices and partitioning information in higher-level video units, the method addresses decoder crashes and incomplete partitioning in video coding standards, ensuring accurate and efficient video decoding.
Patent Information
- Application Number
- JP2023204674
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-03-31
- Filing Date
- 2023-12-04
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2041-02-22
AI Technical Summary
Existing video coding standards face issues in signaling of sub-pictures, tiles, and slices, leading to incorrect derivation of slice data, potential decoder crashes, and incomplete partitioning information, especially in rectangular slice modes and non-rectangular modes.
The proposed methods include deriving picture-level and sub-picture-level slice indices, signaling tile and slice partitioning information in higher-level video units, and ensuring consistent decoding order to prevent decoder crashes and ensure complete partitioning information.
This approach ensures accurate and complete signaling of video partitioning, preventing decoder crashes and enabling efficient video decoding by addressing the issues in existing video coding standards.
Smart Images

Figure 0007701429000082 
Figure 0007701429000083 
Figure 0007701429000084
Abstract
Description
Technical Field
[0001] This application The present invention This is a divisional application of Japanese Patent Application No. 2022-549961, based on International Patent Application No. PCT / CN2021 / 077220 filed on February 22, 2021, which timely claims the priority and benefits of International Patent Application No. PCT / CN2020 / 076158 filed on February 21, 2020 and International Patent Application No. PCT / CN2020 / 082283 filed on March 31, 2020. The entire disclosure of the above applications is incorporated herein by reference as part of the disclosure of this specification.
[0002] relates to the encoding and decoding of images and videos.
Background Art
[0003] Digital videos occupy the largest bandwidth usage in the Internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying videos increases, the bandwidth demand for the use of digital videos is expected to continue to increase.
Summary of the Invention
[0004] The present application discloses a technique that can be used by a video encoder and a decoder to process a coded representation of a video using control information useful for decoding the coded representation.
[0005] In one exemplary aspect, a video decoding method is disclosed. The method includes performing a conversion between a video picture comprising one or more tiles and one or more rectangular slices and a bitstream of the video according to a rule. The rule defines updating a variable indicating a tile index only for slices having an index smaller than a value obtained by subtracting 1 from the number of slices in the video picture for repeatedly determining information regarding the one or more rectangular slices.
[0006] In another exemplary aspect, a video decoding method is disclosed. The method includes performing a conversion between a video picture including one or more sub-pictures and a bitstream of the video. The conversion follows a rule that defines omitting syntax elements in a sequence parameter set indicating a plurality of sub-pictures in the video picture when a maximum picture width and a maximum picture height are less than or equal to dimensions of a coding tree block.
[0007] In another exemplary aspect, a video decoding method is disclosed. The method includes performing a conversion between a video picture including one or more tiles and a bitstream of the video. The conversion complies with a rule that, when the width of the video picture is less than or equal to the size of a coding tree block, a syntax element indicating the column width of a plurality of explicitly provided tiles is omitted in the bitstream.
[0008] In another exemplary aspect, a video decoding method is disclosed. The method includes performing a conversion between a video picture including one or more tiles and a bitstream of the video. The conversion complies with a rule that, when the height of the video picture is less than or equal to the size of a coding tree block, a syntax element indicating the number of row heights of explicitly provided tiles is omitted in the bitstream.
[0009] In another exemplary aspect, a video decoding method is disclosed. The method includes performing a conversion between a video picture including one or more tiles and a bitstream of the video. The conversion complies with a rule that, when the column widths of a plurality of explicitly provided tiles are equal to the picture width in coding tree block units, one or more syntax elements indicating the column widths of one or more tiles are omitted in the bitstream.
[0010] In another exemplary aspect, a video decoding method is disclosed. The method includes performing a conversion between a video picture including one or more tiles and a bitstream of the video. The conversion complies with a rule that, when the row heights of a plurality of explicitly provided tiles are equal to the picture height in coding tree block units, one or more syntax elements indicating the row heights of one or more tiles are omitted in the bitstream.
[0011] In another exemplary aspect, a video decoding method is disclosed. The method includes performing a conversion between a video picture including one or more slices and a bitstream of the video. The conversion complies with a rule that specifies that slice partition information is included in the bitstream.
[0012] In another exemplary aspect, a video decoding method is disclosed. The method includes performing a conversion between a video tile including one or more rectangular slices and a bitstream of the video based on a rule. The rule specifies determining a uniform slice height based on a first syntax element that specifies the height of a rectangular slice in a video tile in units of rows of coding tree units, and a second syntax element that specifies the number of heights of slices explicitly provided in the video tile.
[0013] In another exemplary aspect, a video decoding method is disclosed. The method includes performing a conversion between a video picture including one or more tiles and a bitstream of the video according to a rule. The rule specifies that a syntax element is equal to or greater than the dimensions of a column or row of tiles with uniform dimensions. The syntax element indicates the dimensions in coding tree block units excluding the total dimensions of the widths of the columns or the heights of the rows of tiles explicitly provided.
[0014] In another exemplary aspect, a video decoding method is disclosed. The method includes performing a conversion between a video tile including one or more slices and a bitstream of the video according to a rule. The rule specifies that a syntax element is greater than or equal to the height of a uniform slice. The syntax element indicates the height in units of coding tree blocks excluding the total height of the heights of a plurality of slices explicitly provided.
[0015] In another exemplary aspect, a video decoding method is disclosed. The method includes performing a conversion between a video and a bitstream of the video according to a rule. The rule specifies using a syntax element for the conversion to indicate the maximum number of affine merge candidates permitted in a sub-block based merge candidate list.
[0016] In another exemplary aspect, a video decoding method is disclosed. The method includes performing a conversion between a video including one or more video pictures and a coded representation of the video, each video picture including one or more sub-pictures each including one or more slices, the coded representation conforming to format rules, the format rules specifying that when a rectangular slice mode is enabled for a video picture, a picture level slice index for each slice in each sub-picture of the video picture is derived without explicit signaling in the coded representation, and the format rules specifying that a plurality of coding tree units in each slice are derivable from the picture level slice index.
[0017] In another exemplary aspect, another video processing method is disclosed. The method includes performing a conversion between a video including one or more video pictures and a coded representation of the video, each video picture including one or more sub-pictures each including one or more slices, the coded representation conforming to format rules, the format rules specifying that a sub-picture level slice index can be derived based on information in the coded representation without signaling the sub-picture level slice index in the coded representation.
[0018] In another exemplary aspect, another video processing method is disclosed. The method includes performing a conversion between a video including one or more video pictures and a coded representation of the video, each video picture including one or more sub-pictures each including one or more slices and / or one or more tiles, the coded representation conforming to format rules, and the conversion conforming to constraint rules.
[0019] In another exemplary aspect, another video processing method is disclosed. The method includes performing a conversion between a video including one or more video pictures, each video picture including one or more tiles and / or one or more slices, and its coded representation, the coded representation conforming to format rules, the format rules defining that a field at the video picture level conveys information regarding the division of slices and / or tiles in the video picture.
[0020] In another exemplary aspect, another video processing method is disclosed. The method includes performing a conversion between a video including one or more pictures and a coded representation of the video, the conversion conforming to a division rule that makes the minimum number of slices for dividing the video picture a function of whether rectangular division is used for dividing the video picture.
[0021] In another exemplary aspect, another video processing method is disclosed. The method includes performing a conversion between a video slice of a video region of a video and a coded representation of the video, the coded representation conforming to format rules, the format rules defining that the coded representation signals the video slice based on the top left position of the video slice, and the format rules defining that the coded representation signals the height and / or width of the video slice in the division information signaled at the video unit level.
[0022] In another exemplary aspect, another video processing method is disclosed. The method includes performing a conversion between a video including a video picture and a coded representation of the video, the coded representation conforming to format rules, the format rules defining omitting signaling of the difference between the tile index of the first tile in a rectangular slice and the tile index of the first tile in the next rectangular slice.
[0023] In another exemplary aspect, another video processing method is disclosed. This method includes performing a conversion between a video and a coded representation of the video, the coded representation conforming to format rules that define signaling of information used to derive the number of columns or rows of tiles in the video picture based on the relationship between the width of the video picture and the size of coding tree units.
[0024] In another exemplary aspect, another video processing method is disclosed. This method includes performing a conversion between a video including one or more video pictures and a coded representation of the video, the coded representation conforming to format rules that define that tile layout information is included in the coded representation for video pictures including tiles with uniform spacing and tiles with non-uniform spacing.
[0025] In yet another exemplary aspect, a video encoder device is disclosed. This video encoder includes a processor configured to implement the method described above.
[0026] In yet another exemplary aspect, a video decoder device is disclosed. This video decoder includes a processor configured to implement the method described above.
[0027] In yet another exemplary aspect, a computer-readable medium storing code is disclosed. This code is implemented in a form of code executable by a processor to perform one of the methods described herein.
[0028] These and other features are described throughout this document.
Brief Description of the Drawings
[0029]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Figure 22
Figure 23
Mode for Carrying Out the Invention
[0030] In this specification, chapter headings are used for ease of understanding, and the technology and the applicability of the embodiments described in each chapter are not limited to that chapter only. Further, the term H.266 is used in a certain description only for ease of understanding and is not used to limit the scope of the disclosed technology. Thus, the technology described in this specification is applicable to other video codec protocols and designs as well.
[0031] 1. Overview This specification relates to video coding techniques. Specifically, it relates to signaling of sub-pictures, tiles, and slices. This idea may be applied, individually or in various combinations, to multi-layer video coding, for example, any video coding standard or non-standard video codec that supports Versatile Video Coding (VVC) currently under development. 2. Abbreviations APS <Adaptation Parameter Set> Adaptation Parameter Set AU <Access Unit> Access Unit AUD <Access Unit Delimiter> Access Unit Delimiter AVC <Advanced Video Coding> Advanced Video Coding CLVS <Coded Layer Video Sequence> Coded Layer Video Sequence CPB <Coded Picture Buffer> Coded Picture Buffer CRA <Clean Random Access> Clean Random Access CTU <Coding Tree Unit> Coding Tree Unit CVS <Coded Video Sequence> Coded Video Sequence DPB <Decoded Picture Buffer> Decoded Picture Buffer DPS <Decoding Parameter Set> Decoding Parameter Set EOB <End Of Bitstream> End Of Bitstream EOS <End Of Sequence> End of sequence GDR <Gradual Decoding Refresh> Gradual decoding refresh HEVC <High Efficiency Video Coding> High efficiency video coding HRD <Hypothetical Reference Decoder> Hypothetical reference decoder IDR <Instantaneous Decoding Refresh> Instantaneous decoding refresh JEM <Joint Exploration Model> Joint exploration model MCTS <Motion-Constrained Tile Sets> Motion-constrained tile sets NAL <Network Abstraction Layer> Network abstraction layer OLS <output layer set> Output layer set PH<Picture Header> Picture header PPS<Picture Parameter Set> Picture parameter set PTL<Profile,Tier and Level> Profile, tier and level PU<Picture Unit> Picture unit RBSP<Raw Byte Sequence Payload> Raw byte sequence payload SEI<Supplemental Enhancement Information> Supplemental enhancement information SPS<Sequence Parameter Set> Sequence parameter set SVC<Scalable Video Coding> Scalable video coding VCL <video coding layer> Video Coding Layer VPS <video parameter set> Image parameter set VTM <VVC Test Model> VVC test model VUI <video usability information> Video availability information VVC <Versatile Video Coding> Versatile video coding
[0032] 3. Introduction of Video Coding Video coding standards have mainly evolved through the development of well-known ITU-T and ISO / IEC standards. ITU-T developed H.261 and H.263, ISO / IEC developed MPEG-1 and MPEG-4 Visual, and both organizations jointly developed the H.262 / MPEG-2 Video, H.264 / MPEG-4 AVC (Advanced Video Coding), and H.265 / HEVC standards. Since H.262, video coding standards have been based on a hybrid video coding structure that utilizes temporal prediction and transform coding. To explore future video coding technologies beyond HEVC, in 2015, VCEG and MPEG jointly established the JVET (Joint Video Exploration Team). Since then, many new methods have been adopted by the JVET and incorporated into the reference software called JEM (Joint Exploration Model). The JVET meets once a quarter, and the new coding standard aims to reduce the bitrate by 50% compared to HEVC. At the JVET meeting in April 2018, the new video coding standard was officially named "Versatile Video Coding (VVC)", and at that time, the first version of the VVC Test Model (VTM) was released. Since efforts to contribute to the standardization of VVC continue, at every JVET meeting, new coding technologies are adopted into the VVC standard. After each meeting, the VVC working draft and the test model VTM are updated. The VVC project is currently aiming for technical completion (FDIS) at the meeting in July 2020.
[0033] 3.1. Picture Partitioning Scheme in HEVC HEVC has four different picture partitioning schemes: regular slices, dependent slices, tiles, and WPP (Wavefront Parallel Processing). By applying these, it becomes possible to match the maximum transfer unit (MTU) size, perform parallel processing, and reduce end-to-end latency.
[0034] Regular slices are similar to those in H.264 / AVC. Each regular slice is encapsulated in its own NAL unit, and picture prediction (intra-sample prediction, motion information prediction, coding mode prediction) and entropy coding dependencies across slice boundaries are disabled. Thus, one regular slice can be reconstructed independently of other regular slices within the same picture (however, there may still be interdependencies due to the loop filtering operation).
[0035] Regular slices are the only tool that can be used for parallelization and can be used in a nearly identical form in H.264 / AVC as well. Parallelization based on regular slices does not require much inter-processor or inter-core communication (when decoding a predicted picture, except for data sharing between processors or cores for motion compensation, it is usually much heavier than data sharing between processors or cores for intra-picture prediction). However, for the same reason, using regular slices can result in a significant coding overhead due to the bit cost of the slice header and the lack of prediction across slice boundaries. Furthermore, regular slices (in contrast to other tools described later) also function as a key mechanism for splitting the bitstream to adapt to the MTU size requirements due to the intra-picture independence of regular slices and each regular slice being encapsulated in its own NAL unit. In many cases, the goals of parallelization and MTU size matching impose conflicting requirements on the slice layout in the picture. By realizing such a situation, the following parallelization tools have been developed.
[0036] A dependent slice has a short slice header and enables the bitstream to be partitioned at tree block boundaries without interrupting any intra-picture prediction. Basically, a dependent slice reduces end-to-end delay by fragmenting a regular slice into multiple NAL units and enabling a part of the regular slice to be sent before the encoding of the entire regular slice is completed.
[0037] In WPP, a picture is divided into a single row of coding tree blocks (CTBs). Entropy decoding and prediction are permitted to use data from CTBs in other partitions. Parallel processing is possible by parallel decoding of CTB rows, with the start of decoding of one CTB row being delayed by only two CTBs, so that data regarding the CTB upper-right of the target CTB is surely available before the target CTB is decoded. By using this staggered start (which looks like a wavefront when represented in a graph), it is possible to parallelize a picture using as many processing devices / cores as the number of CTB rows it contains. Since intra-picture prediction between neighboring tree block rows within one picture is permitted, the inter-processor / inter-core communication necessary to enable intra-picture prediction can be sufficient. The WPP partition does not result in the generation of additional NAL units as compared to the case where it is not applied. Thus, WPP is not a tool for MTU size matching. However, when MTU size matching is required, a regular slice can be used with WPP with a certain coding overhead.
[0038] A tile defines horizontal and vertical boundaries that divide a picture into tile columns and rows. The tile columns extend from the top to the bottom of the picture. Similarly, the tile rows extend from the left to the right of the picture. The number of tiles in a picture can be obtained simply by multiplying the number of tile columns by the number of tile rows.
[0039] The scan order of CTBs is changed to be local within one tile (in the order of raster scan of CTBs in one tile), and then, following the order of raster scan of tiles in one picture, the top-left CTB of the next tile is decoded. Similar to normal slices, tiles break picture-internal prediction dependencies and entropy decoding dependencies. However, they do not need to be included in individual NAL units (the same as WPP in this regard), and thus, tiles cannot be used for MTU size matching. Each tile may be processed by one processor / core, and in the communication between processors / cores required for picture-internal prediction between processing units, the decoding of neighboring tiles is limited to the transmission of a shared slice header and the sharing related to loop filtering of reconstructed samples and metadata when one slice spans two or more tiles. When one slice contains two or more tiles or WPP segments, the entry point byte offset of each tile or WPP segment other than the first one in the slice is signaled in the slice header.
[0040] For simplicity of explanation, in HEVC, restrictions regarding the application of four different picture partitioning methods are defined. For most of the profiles specified in HEVC, a given coded video sequence cannot contain both tiles and waves. For each slice and tile, either or both of the following conditions must be met. 1) All coded tree blocks in one slice belong to the same tile. 2) All coded tree blocks in one tile belong to the same slice. Finally, one wave segment contains exactly one CTB row, and when WPP is used, if one slice starts within one CTB row, it must end in the same CTB row.
[0041] The recent HEVC amendment is the JCT-VC output document JCTVC-AC1005, J. Voice, A. Ramasubramonian, R. Skupin, G.J. Sri, A. Thulaphi, Y.-K. Wang (editors), "Additional Capture Enhancement Information for HEVC (Draft4)", Oct. 24, 2017, available at: http: / / phenix.int-evry.fr / jct / doc_end_user / documents / 29_Macau / wg11 / JCTVC-AC1005-v2.zip. Including this correction, HEVC specifies three MCTS-related SEI (Supplemental Enhancement Information) messages, namely, the Temporal MCTS SEI message, the MCTS Extraction Information Set SEI message, and the MCTS Extraction Information Nest SEI message.
[0042] The Temporal MCTS SEI message indicates the presence of MCTS in the bitstream and signals the MCTS. In each MCTS, the motion vectors are restricted to point to the full-sample positions within the MCTS and the fractional-sample positions that require only the full-sample positions within the MCTS for interpolation, and the use of motion vector candidates for temporal motion vector prediction derived from blocks outside the MCTS is not permitted. Thus, each MCTS has no tiles not included in the MCTS and may be decoded independently.
[0043] The MCTS extraction information set SEI message provides supplementary information that can be used in MCTS sub-bitstream extraction (specified as part of the meaning of the SEI message) and generates a compliant bitstream for the MCTS set. This information is composed of the number of extraction information sets, and each extraction information set defines the number of MCTS sets and contains the RBSP bytes of alternative VPS, SPS, and PPS used in the MCTS sub-bitstream extraction process. When extracting the sub-bitstream by the MCTS sub-bitstream extraction process, it is necessary to rewrite or replace the parameter sets (VPS, SPS, PPS) because one or all of the syntax elements related to the slice address (including first_slice_segment_in_pic_flag and slice_segment_address) need to have different values.
[0044] 3.2. Picture Partitioning in VVC In VVC, one picture is divided into one or more tile rows and one or more tile columns. One tile is a sequence of CTUs that covers a rectangular region of one image. The CTUs in one tile are scanned in raster scan order within that tile.
[0045] One slice contains an integer number of complete tiles or an integer number of consecutive complete CTU rows within the tile of one picture.
[0046] It corresponds to two modes of slices, namely the raster scan slice mode and the rectangular slice mode. In the raster scan slice mode, one slice contains a sequence of one complete tile in the tile raster scan of one picture. In the rectangular slice mode, one slice contains either the number of complete tiles that collectively form a rectangular region of the picture or the number of consecutive complete CTU rows of one tile that collectively form a rectangular region of the picture. The tiles within the rectangular slice are scanned in the order of the tile raster scan within the rectangular region corresponding to that slice.
[0047] One sub-picture includes one or more slices that collectively cover the rectangular area of one picture.
[0048] Figure 1 shows an example of the raster scan slice division of a picture, and the picture is divided into 12 tiles and 3 raster scan slices.
[0049] Figure 2 shows an example of the rectangular slice division of a picture, and the picture is divided into 24 tiles (6 tile columns and 4 tile rows) and 9 rectangular slices.
[0050] Figure 3 shows an example of a picture divided into tiles and rectangular slices, and this picture is divided into 4 tiles (2 tile columns and 2 tile rows) and 4 rectangular slices.
[0051] Figure 4 shows an example of dividing one picture into sub-pictures. One picture is divided into 18 tiles. The 12 tiles on the left each contain one slice of a 4×4 CTU, and the 6 tiles on the right each contain slices stacked vertically of a 2×2 CTU, resulting in a total of 24 slices and 24 sub-pictures of different dimensions (each slice is one sub-picture). 3.3 Signal Notification of SPS / PPS / Picture header / Slice header in VVC 7.3.2.3 Sequence Parameter Set RBSP Syntax [Table 1] [Table 2] [Table 3]
Table 4
Table 5
Table 6
Table 7
Table 8
Table 9
Table 10
Table 11
Table 12
Table 13
Table 14
Table 15
Table 16
Table 17
Table 18
Table 19
Table 20
Table 21
Table 22
Table 23
[0052] 3.5 Color Space and Chroma Subsampling A color space, also known as a color model (or color system), is an abstract mathematical model that simply describes a range of colors as a numerical tuple, typically three or four values or color components (e.g., RGB). Basically, a color space is a refined coordinate system and subspace.
[0053] In the case of video compression, the most frequently used color spaces are YCbCr and RGB.
[0054] YCbCr, Y’CbCr, or YPb / Cb Pr / Cr, also described as YCBCR or Y’CBCR, is a family of color spaces used as part of the pipeline video of color images and digital photo systems. Y’ is the luminance component, and CB and CR are the chroma components of the blue difference and red difference. Y’ (with a prime) is distinguished from Y which is luminance, meaning that the light intensity is non-linearly encoded based on gamma-corrected RGB primaries.
[0055] Chroma subsampling is a method of encoding pictures by taking advantage of the fact that the human visual system is less perceptive to color differences than to luminance, and implementing it so that the chroma information has a lower resolution than the luminance information.
[0056] 3.5.1. 4:4:4 Each of the three Y’CbCr components has the same sample rate, and thus there is no chroma subsampling. This scheme may be used in high-end film scanners and cinematic post-production.
[0057] 3.5.2. 4:2:2 Two chroma components are sampled at half the sample rate of luminance, the horizontal chroma resolution is halved, and the vertical chroma resolution does not change. This results in little or no visual difference and can reduce the bandwidth of an uncompressed video signal by a factor of 1 / 3. Examples of the nominal vertical and horizontal positions of the 4:2:2 color format are shown, for example, in Figure 5 of the VVC working draft.
[0058] 3.5.3. 4:2:0 In 4:2:0, the horizontal sampling is twice that of 4:1:1, but in this scheme, the Cb and Cr channels are sampled only on every other line, so the vertical resolution is halved. Thus, the data rate is the same. Cb and Cr are each subsampled by a factor of 2 in both the horizontal and vertical directions. There are three variants of the 4:2:0 scheme with different horizontal and vertical positions. ● In MPEG-2, Cb and Cr are co-located horizontally. Cb and Cr are located between pixels in the vertical direction (located between grids). ● In JPEG / JFIF, H.261, and MPEG-1, Cb and Cr are located between grids in the middle of alternating luminance samples. ● In 4:2:0 DV, Cb and Cr are co-located horizontally. Vertically, they are co-located alternately. Values of SubWidthC and SubHeightC derived from chroma_format_idc and separate_colour_plane_flag in Table 3-1 [Table 24]
[0059] 4. Examples of technical problems to be solved by the disclosed embodiments Existing designs related to the signaling of SPS / PPS / picture header / slice header in VVC have the following problems. 1) When rect_slice_flag is 1, according to the current VVC text, it is as follows. a. slice_address represents the slice index at the sub-picture level of the slice. b. The slice index at the sub-picture level is defined as the index of the slice with respect to the list arranged in the order in which the slices of the sub-picture are signaled in the PPS. c. The slice index at the picture level is defined as the index of the slice with respect to the list arranged in the order in which the slices within the picture are signaled in the PPS. d. For two slices belonging to two different sub-pictures, the one with the smaller sub-picture index has an earlier decoding order. For two slices belonging to the same sub-picture, the one with the smaller slice index at the sub-picture level has an earlier decoding order. e. And the derivation of the variable NumCtusInCurrSlice that defines the number of CTUs in the current slice according to Equation 117 of the current VVC text assumes that the increasing order of the picture-level slice index values is the same as the decoding order of the slices. However, when a single tile is divided and several slices are generated as a result, it may conflict with the above points. In the example shown in FIG. 6, when a picture is divided into two tiles by a vertical tile boundary and each of the two tiles is divided into two slices by the same horizontal boundary across the entire picture, the upper two slices are included in the first sub-picture and the lower two slices are included in the second sub-picture. In this case, according to the current VVC text, the picture-level slice index values of the four slices in the slice raster scan order are 0, 2, 1, 3, and the decoding order index values of the four slices in the slice raster scan order are 0, 1, 2, 3. As a result, the derivation of NumCtusInCurrSlice becomes incorrect, problems occur in the analysis of slice data, the decoded sample values become incorrect, and there is a high possibility that the decoder will crash. 2) There are two types of slice signaling methods. In the rectangular mode, all slice partitioning information is signaled in the PPS. In the non-rectangular mode, since part of the slice partitioning information is notified in the slice header, it is not possible to know the complete slice partitioning of the picture before analyzing all the slices of the picture in this mode. 3) In the rectangular mode, slices can be signaled arbitrarily by setting tile_idx_delta. A malicious bitstream may crash the decoder with this mechanism. 4) In an embodiment, when i is equal to num_slices_in_pic_minus1, tile_idx_delta[i] that is not initialized 5) The size of the Merge Estimation Region (MER) can be up to a minimum of 4×4. However, if the signaled MER size is smaller than the minimum CU size, it is meaningless. 6) From the syntax table, it can be seen that the width of the "num_exp_tile_columns_minus1"-th tile column is presented as tile_column_width_minus1[num_exp_tile_columns_minus1]. However, in terms of semantics, using tile_column_width_minus1[num_exp_tile_columns_minus1], the width of the tile columns with indexes greater than or equal to num_exp_tile_columns_minus1 defined in Section 6.5.1 is derived. That is, the width of the 'num_exp_tile_columns_minus1'-th tile column can be reset. The same applies to the height of the 'num_exp_tile_columns_minus1'-th tile row.
[0060] Figure 6 shows an example of picture segmentation. The solid line 602 represents the boundary of the tile, the dashed line 604 represents the boundary of the slice, and the dashed line 606 represents the boundary of the sub-picture. The figure shows the picture level index, decoding order index, sub-picture level index, and the indexes of the sub-picture and tile for four slices.
[0061] 5. Exemplary Embodiments and Techniques To solve the above-mentioned problems and the like, the following methods are disclosed. The present invention should be regarded as an example for explaining general concepts and should not be construed in a narrow sense. Furthermore, the present invention may be applied individually or arbitrarily combined. 1. For slices in the rectangular slice mode (i.e., when rect_slice_flag is equal to 1), the picture level slice index for each slice of each sub-picture is derived, and the number of CTUs in each slice is derived using the derived value. 2. The sub-picture level slice index can be defined / derived by the following method. a. In one example, the slice index at the sub-picture level is defined as "the index of the slice in the list of slices within the sub-picture in the decoding order when rect_slice_flag is equal to 1". b. Alternatively, the slice index at the sub-picture level is defined as "the index of the slice with respect to the slice list of the sub-picture when rect_slice_flag is equal to 1, defined by the variable SubpicLevelSliceIdx[i] derived by Equation 32 (similar to Embodiment 1), where i is the picture-level slice index of that slice". c. As an example, the sub-picture index of each slice where the picture-level slice index is a specific value is derived. d. As an example, the slice index at the sub-picture level of each slice where the picture-level slice index is a specific value is derived. e. In one example, when rect_slice_flag is equal to 1, the semantics of the slice address are defined as "the slice address is the slice index at the sub-picture level of the slice defined by the variable SubpicLevelSliceIdx[i] derived by Equation 32 (for example, as in Embodiment 1), where i is the picture-level slice index of that slice". 3. The slice index at the sub-picture level of a slice is assigned to the slice within the first sub-picture that contains that slice. The slice index at the sub-picture level of each slice may be stored in an array (for example, SubpicLevelSliceIdx[i] in Embodiment 1) with the picture-level slice index as the index. a. As an example, the slice index at the sub-picture level is a non-negative integer. b. As an example, the value of the slice index at the sub-picture level of a slice is 0 or greater. c. In one example, the value of the slice index at the sub-picture level of a slice is less than N, where N is the number of slices of the sub-picture. d. As an example, when the slice index (denoted as subIdxA) of the first sub-picture level of the first slice (Slice A) is different while the first slice and the second slice (Slice B) are within the same sub-picture, it must be different from the slice index (denoted as subIdxB) of the second sub-picture level of the second slice (Slice B). e. In one example, when the slice index (denoted as subIdxA) of the first sub-picture level of the first slice (Slice A) in the first sub-picture is smaller than the slice index (denoted as subIdxB) of the second sub-picture level of the second slice (Slice B) in the same first sub-picture, IdxA is smaller than IdxB, and IdxA and IdxB respectively represent the slice index of the entire picture (also known as the slice index at the picture level, e.g., sliceIdx) of Slice A and Slice B. f. As an example, the slice index (denoted as subIdxA) of the first sub-picture level of the first slice (Slice A) within the first sub-picture is smaller than the slice index (denoted as subIdxB) of the second sub-picture level of the second slice (Slice B) within the same first sub-picture In the decoding order, Slice A precedes Slice B. . g. As an example, the slice index at the sub-picture level in a sub-picture is derived based on the slice index at the picture level (e.g., sliceIdx). 4. It is proposed to derive a mapping function / table between the slice index at the sub-picture level and the slice index at the picture level within a sub-picture. a. As an example, the two-dimensional array PicLevelSliceIdx[subPicIdx][SubPicLevelSliceIdx] is derived to map the sub-picture level slice index within the sub-picture to the picture level slice index. Here, PicLevelSliceIdx is the picture level slice index of the slice, subPicIdx is the index of the sub-picture, and SubPicLevelSliceIdx indicates the sub-picture level slice index of the slice of the sub-picture. i. In one example, the array NumSlicesInSubpic[subPicIdx] is used for the derivation of PicLevelSliceIdx, where NumSlicesInSubpic[subPicIdx] indicates the number of slices in the sub-picture with an index equal to subPicIdx. 1) As an example, NumSlicesInSubpic[subPicIdx] and PicLevelSliceIdx[subPicIdx][SubPicLevelSliceIdx] are derived in one pass by scanning all slices in the order of the picture level slice index. a. NumSlicesInSubpic[subPicIdx] is set equal to 0 for all valid subPicIdx before processing. b. When checking a slice with a picture level index of S, if it is within a sub-picture with a sub-picture index of P, set PicLevelSliceIdx[P][NumSlicesInSubpic[P]] to S, and then set NumSlicesInSubpic[P] to NumSlicesInSubpic[P]+1. ii. In one example, SliceIdxInPic[subPicIdx][SubPicLevelSliceIdx] is used to derive the picture level slice index (e.g., picLevelSliceIdx), and this is used to derive the number and / or address of CTBs within the slice during the parsing of the slice header. 5. In the compliance bitstream, it is required that one tile is not included in more than two sub-pictures. 6. In the compliance bitstream, it is required that if slice A is within tile A but smaller than tile A, slice B is within tile B but smaller than tile B, and tile A and tile B are different, one sub-picture cannot include the two slices represented by slice A and slice B. 7. It is proposed to signal the tile and / or slice partitioning information of a picture in the associated picture header. a. In one example, whether the tile and / or slice partitioning information of a picture is signaled in the PPS or in the associated picture header is signaled in the associated PPS. b. In one example, it is signaled whether the tile and / or slice partitioning information of a picture is in the associated picture header. i. As an example, if the tile and / or slice partitioning information of a picture is signaled in both the associated PPS and the associated picture header, the tile and / or slice partitioning information of the picture signaled in the picture header will be used. ii. As an example, if the tile and / or slice partitioning information of a picture is signaled in both the associated PPS and the associated picture header, the tile and / or slice partitioning information of the picture signaled in the PPS will be used. c. As an example, in a video unit at a level higher than the picture (such as in the SPS, etc.), it is signaled to indicate whether the tile and / or slice partitioning information of the picture is signaled in the associated PPS or in the associated picture header. 8. When slicing an associated picture in non-rectangular mode, it is proposed to signal the slice partitioning information in a video unit higher than the slice level (e.g., within the PPS and / or picture header). a. As an example, when the related picture is slice - divided in non - rectangular mode, information indicating the number of slices (e.g., num_slices_in_pic_minus1) may be signaled in the upper - level video unit. b. As an example, when the related picture is divided into slices having non - rectangular mode, it is information for indicating the index (or address, or position, or coordinates) of the first block - unit of the slice in the upper - level video unit. For example, the block unit may be a CTU or a tile. c. As an example, when the related picture is slice - divided in non - rectangular mode, it provides information for indicating the number of block units of the slice in the upper - level video unit. For example, the block unit may be a CTU or a tile. d. In one example, when the related picture is divided into slices having non - rectangular mode, the slice - division information (e.g., num_tiles_in_slice_minus1) is not signaled in the slice header. e. As an example, when the related picture is sliced in non - rectangular mode, the slice index is signaled in the slice header. i. As an example, slice_address is interpreted as the picture - level slice index when the related picture is sliced in non - rectangular mode. f. In one example, when the related picture is divided into slices having non - rectangular mode, the division information of each slice within the picture (such as the index of the first block - unit and / or the number of block - units) may be signaled in order in the upper - level video unit. i. As an example, when the related picture is sliced in non - rectangular mode, the slice index may be signaled for each slice of the upper - level video unit. ii. As an example, the division information of each slice is signaled in ascending order of the slice index. 1) As an example, the segmentation information of each slice is signaled in the order of slice 0, slice 1, ..., slice K-1, slice K, slice K+1, ... slice S-2, slice S-1, where K represents the slice index and S represents the number of slices of a picture. iii. As an example, the segmentation information of each slice is signaled in descending order of the slice index. 1) As an example, the segmentation information of each slice is signaled in the order of slice S-2, slice S-1, ..., slice K+1, slice K, slice K-1, ..., slice 1, slice 0, where K represents the slice index and S represents the number of slices of a picture. iv. As an example, when the related picture is sliced in a non-rectangular mode, the index of the first block-unit for a slice may not be signaled in the upper video unit. 1) For example, it is inferred that the index of the first block unit of slice 0 (the slice with a slice index of 0) is 0. 2) For example, the index of the first block unit of slice K (the slice with a slice index equal to K, K>0) is
Number
Number
Number
Number
Table 25
Table 26
[0062] 6. Embodiments In the following embodiments, the added parts are marked with bold, underlined, and italicized characters. The deleted parts are marked within []. 6.1. Embodiment 1: Example of changing the sub-picture level slice index 3 Definitions Picture level slice index: When rect_slice_flag is equal to 1, it indicates the index of the slice with respect to the list of slices in the picture, in the order signaled in the PPS. [[Subpicture level slice index: When rect_slice_flag is equal to 1, it indicates the index of the slice with respect to the list of slices in the subpicture, in the order signaled in the PPS.]]
Chem.
Chem.
Translated
Table 27
Table 28
Table 29
Table 30
Fig.
Fig.
Chemical
Chem.
Table 31
Table 32
Chem.
Chem.
Chem.
Chemical formula
Chem.
Table 33
Chemical formula
Chem.
Chem.
Chem.
Table 34
化
Chemical
Chem.
Chem.
Table 35
Table 36
Table 37
Fig.
Chem.
Chem.
Chem.
Chem.
Chem.
Chem.
Chem.
Chem.
Chem.
Chem.
Table 38
Table 39
Chem.
Table 40
Chem.
Chemical formula
[0063] FIG. 7 is a block diagram showing an exemplary video processing system 1900 in which various techniques disclosed herein may be implemented. Various implementations may include some or all of the components of system 1900. System 1900 may include an input unit 1902 for receiving video content. The video content may be received in an unprocessed or uncompressed format, such as 8- or 10-bit multi-component pixel values, or in a compressed or encoded format. Input unit 1902 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet®, PON® (Passive Optical Network), etc., and wireless interfaces such as Wi-Fi® or a cellular interface.
[0064] System 1900 may include a coding component 1904 that can implement various coding or encoding methods described herein. The coding component 1904 may reduce the average bitrate of the video from the input unit 1902 to the output of the coding component 1904 and generate a coded representation of the video. Thus, this coding technique may be referred to as a video compression or video transcoding technique. The output of the coding component 1904 may be stored or transmitted via a connected communication as represented by component 1906. The bitstream (or coded) representation of the video received, stored, or communicated at the input unit 1902 may be used by component 1908 to generate pixel values or a displayable video that is transmitted to the display interface 1910. The process of generating a video that a user can view from the bitstream representation may be referred to as video decompression (video expansion). Further, although specific video processing operations are referred to as "coding" operations or tools, it will be understood that the coding tools or operations are performed by an encoder and the corresponding decoding tools or operations that reverse the result of the coding are performed by a decoder.
[0065] Examples of a peripheral bus interface or a display interface may include USB (trademark; Universal Serial Bus) or HDMI (trademark; High Definition Multimedia Interface) or DisplayPort, etc. Examples of a storage interface include SATA (Serial Advanced Technology Attachment), PCI, IDE interface, etc. The techniques described herein may be implemented in various electronic devices such as mobile phones, laptop computers, smartphones, or other devices capable of performing digital data processing and / or video display.
[0066] FIG. 8 is a block diagram of a video processing apparatus 3600. The apparatus 3600 may be used to implement one or more of the methods described herein. The apparatus 3600 may be implemented in a smartphone, a tablet, a computer, an Internet of Things (IoT) receiver for a single device, or the like. The apparatus 3600 may include one or more processing devices 3602, one or more memories 3604, and video processing hardware 3606. One or more of the processing devices 3602 may be configured to implement one or more of the methods described herein. The memory or memories 3604 may be used to store data and code used to implement the methods and techniques described herein. The video processing hardware 3606 may be used to implement the techniques described herein in a hardware circuit.
[0067] FIG. 10 is a block diagram showing an exemplary video coding system 100 that may utilize the techniques of the present disclosure.
[0068] As shown in FIG. 10, the video coding system 100 may include a source device 110 and a destination device 120. The source device 110 generates encoded video data and may also be referred to as a video encoding device. The destination device 120 may decode the encoded video data generated by the source device 110 and may be referred to as a video decoding device.
[0069] The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.
[0070] The video source 112 may include a source such as a video capture device, an interface for receiving video data from a video content provider, and / or a computer graphics system for generating video data, or a combination of these sources. The video data may include one or more pictures. The video encoder 114 encodes the video data from the video source 112 and generates a bitstream. The bitstream may include a sequence of bits forming a coded representation of the video data. The bitstream may include coded pictures and associated data. The coded picture is a coded representation of the picture. The associated data may include a sequence parameter set, a picture parameter set, and other syntax structures. The I / O interface 116 may include a modem and / or a transmitter. The encoded video data may be transmitted directly to the destination device 120 via the I / O interface 116 via the network 130a. The encoded video data may be stored in the storage medium / server 130b for access by the destination device 120.
[0071] The destination device 120 may include an I / O interface 126, a video decoder 124, and a display device 122.
[0072] The I / O interface 126 may include a receiver and / or a modem. The I / O interface 126 may obtain the encoded video data from the source device 110 or the storage medium / server 130b. The video decoder 124 may decode the encoded video data. The display device 122 may display the decoded video data to the user. The display device 122 may be integrated with the destination device 120 or may be external to the destination device 120 configured to interface with an external display device.
[0073] The video encoder 114 and the video decoder 124 may operate in accordance with video compression standards such as the HEVC (High Efficiency Video Coding) standard, the VVM (Versatile Video Coding) standard, and other current and / or future standards.
[0074] FIG. 11 is a block diagram showing an example of a video encoder 200, which may be the video encoder 114 in the system 100 shown in FIG. 10.
[0075] The video encoder 200 may be configured to execute any or all of the techniques of the present disclosure. In the example of FIG. 11, the video encoder 200 includes a plurality of functional components. The techniques described in the present disclosure may be shared among various components of the video encoder 200. In some examples, the processor may be configured to perform any or all of the techniques described in the present disclosure.
[0076] The functional components of the video encoder 200 may include a splitting unit 201, a prediction unit 202 that may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra prediction unit 206, a residual generation unit 207, a conversion unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse conversion unit 211, a reconstruction unit 212, a buffer 213, and an entropy encoding unit 214.
[0077] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an intra block copy (IBC) unit. The IBC unit can perform prediction in the IBC mode where at least one reference picture is the picture in which the current video block is located.
[0078] Furthermore, some components such as the motion estimation unit 204 and the motion compensation unit 205 may be highly integrated, but for the sake of explanation, they are shown separately in the example of FIG. 11.
[0079] The splitting unit 201 may split one picture into one or more video blocks. The video encoder 200 and the video decoder 300 may support various video block sizes.
[0080] The mode selection unit 203 may select, for example, one of the intra or inter coding modes based on the error result, supply the obtained intra or inter coding block to the residual generation unit 207 to generate residual block data, and also supply it to the reconstruction unit 212 to reconstruct the coded block for use as a reference picture. In some examples, the mode selection unit 203 may select a CIIP (Combination of Intra and Inter Prediction) mode that performs prediction based on the inter prediction signal and the intra prediction signal. The mode selection unit 203 may select the resolution of the motion vector (e.g., sub-pixel or integer-pixel accuracy) for the block in the case of inter prediction.
[0081] To perform inter prediction on the current video block, the motion estimation unit 204 may generate motion information for the current video block by comparing one or more reference frames from the buffer 213 with the current video block. The motion compensation unit 205 may determine a predicted video block for the current video block based on the motion information of the picture from the buffer 213 other than the picture associated with the current video block and the decoded samples.
[0082] The motion estimation unit 204 and the motion compensation unit 205 may perform different operations on the current video block, for example, based on whether the current video block is an I slice, a P slice, or a B slice.
[0083] In some examples, the motion estimation unit 204 may perform a uni-directional prediction on the current video block, and the motion estimation unit 204 may search for a reference picture in list 0 or list 1 for the reference video block for the current video block. Then, the motion estimation unit 204 may generate a reference index indicating the reference picture in list 0 or list 1, including the reference video block and a motion vector indicating the spatial displacement between the current video block and the reference video block. The motion estimation unit 204 may output the reference index, the prediction direction indicator, and the motion vector as the motion information of the current video block. The motion compensation unit 205 may generate a predicted video block of the current block based on the reference video block indicated by the motion information of the current video block.
[0084] In other examples, the motion estimation unit 204 may perform a bi-directional prediction on the current video block. The motion estimation unit 204 may search for a reference video block for the current video block from among the reference pictures in list 0, and may also search for another reference video block for the current video block from among the reference pictures in list 1. Then, the motion estimation unit 204 may generate a reference index indicating the reference pictures in list 0 and list 1 including the reference video block, and a motion vector indicating the spatial displacement between the reference video block and the current video block. The motion estimation unit 204 may output the reference index and the motion vector of the current video block as the motion information of the current video block. The motion compensation unit 205 may generate a predicted video block of the current video block based on the reference video block indicated by the motion information of the current video block.
[0085] In some examples, the motion estimation unit 204 may output a full set of motion information for decoder decoding processing.
[0086] In some examples, motion estimation unit 204 may not output a full set of motion information for the current video. Rather, motion estimation unit 204 may signal the motion information of the current video block by referring to the motion information of another video block. For example, motion estimation unit 204 may determine that the motion information of the current video block is sufficiently similar to the motion information of neighboring video blocks.
[0087] In one example, motion estimation unit 204 may indicate a value to video decoder 300 indicating that in the syntax structure associated with the current video block, the current video block has the same motion information as another video block.
[0088] In another example, motion estimation unit 204 may identify another video block and an MVD (Motion Vector Difference) in the syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. Video decoder 300 may determine the motion vector of the current video block using the motion vector of the indicated video block and the motion vector difference.
[0089] As described above, video encoder 200 may signal motion vectors predictively. Two examples of predictive signaling techniques that may be implemented by video encoder 200 include AMVP (Advanced Motion Vector Prediction) and merge mode signaling.
[0090] Intra prediction unit 206 may perform intra prediction on the current video block. When intra prediction unit 206 performs intra prediction on the current video block, intra prediction unit 206 may generate prediction data for the current video block based on the decoded samples of other video blocks in the same picture. The prediction data for the current video block may include the predicted video block and various syntax elements.
[0091] The residual generation unit 207 may generate residual data for the current video block by subtracting the predicted video block of the current video block from the current video block (e.g., indicated by a negative sign). The residual data of the current video block may include a residual video block corresponding to different sample components of the samples in the current video block.
[0092] In other examples, for example, in skip mode, there may be no residual data for the current video block, and the residual generation unit 207 may not perform a subtraction operation.
[0093] The transform processing unit 208 may generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video block associated with the current video block.
[0094] After the transform processing unit 208 generates the transform coefficient video block associated with the current video block, the quantization unit 209 may quantize the transform coefficient video block associated with the current video block based on the values of one or more quantization parameters (QPs) associated with the current video block.
[0095] The inverse quantization unit 210 and the inverse transform unit 211 may respectively apply inverse quantization and inverse transform to the transform coefficient video block and reconstruct the residual video block from the transform coefficient video block. The reconstruction unit 212 may add the reconstructed residual video block to the corresponding samples from the one or more predicted video blocks generated by the prediction unit 202 to generate the reconstructed video block associated with the current block for storage in the buffer 213.
[0096] After the reconstruction unit 212 reconstructs the video block, a loop filtering operation may be performed to reduce video blocking artifacts in the video block.
[0097] The entropy encoding unit 214 may receive data from other functional components of the video encoder 200. When receiving data, the entropy encoding unit 214 may perform one or more entropy encoding operations, generate entropy-encoded data, and output a bitstream including the entropy-encoded data.
[0098] FIG. 12 is a block diagram showing an example of a video decoder 300, and this video decoder 300 may be the video decoder 114 in the system 100 shown in FIG. 10.
[0099] The video decoder 300 may be configured to execute any or all of the techniques of the present disclosure. In the embodiment of FIG. 12, the video decoder 300 includes a plurality of functional components. The techniques described in the present disclosure may be shared among various components of the video decoder 300. In some examples, the processor may be configured to perform any or all of the techniques described in the present disclosure.
[0100] In the example of FIG. 12, the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, the video decoder 300 may perform a decoding path that is substantially inverse to the encoding path described with respect to the video encoder 200 (FIG. 11).
[0101] The entropy decoding unit 301 extracts the encoded bitstream. The encoded bitstream may include entropy-coded video data (e.g., an encoded block of video data). The entropy decoding unit 301 decodes the entropy-coded video data, and from the entropy-decoded video data, the motion compensation unit 302 may determine motion information including a motion vector, motion vector precision, reference picture list index, and other motion information. The motion compensation unit 302 may determine such information, for example, by executing AMVP and merge mode.
[0102] The motion compensation unit 302 may generate a motion-compensated block and, in some cases, perform interpolation based on an interpolation filter. An identifier for the interpolation filter used at sub-pixel precision may be included in the syntax element.
[0103] The motion compensation unit 302 may calculate an interpolation value for sub-integer pixels of a reference block using an interpolation filter such as that used by the video encoder 200 during encoding of a video block. The motion compensation unit 302 may determine the interpolation filter used by the video encoder 200 based on the received syntax information and generate a prediction block using this interpolation filter.
[0104] The motion compensation unit 302 may use some of the syntax information for determining the size of the block used for encoding a frame and / or slice of the encoded video sequence, the partitioning information describing how each macroblock of the picture of the encoded video sequence is partitioned, the mode indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-encoded block, and other information for decoding the encoded video sequence.
[0105] The intra prediction unit 303 may form a prediction block from spatially adjacent blocks, for example, using the intra prediction mode received in the bitstream. The inverse quantization unit 303 inverse quantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 301. The inverse transform unit 303 applies an inverse transform.
[0106] The reconstruction unit 306 may sum the residual block and the corresponding prediction block generated by the motion compensation unit 202 or the intra prediction unit 303 to form a decoded block. If desired, a deblocking filter may be applied to filter the decoded block to remove block artifacts. The decoded video block is stored in the buffer 307, and the buffer 307 provides a reference block for subsequent motion compensation / intra prediction and generates the decoded video for display on a display device.
[0107] Next, solutions suitable in some embodiments are enumerated.
[0108] The following solutions show exemplary embodiments of the techniques discussed in the previous chapter (e.g., item 1).
[0109] 1. The video processing method of the present invention (e.g., method 900 described in FIG. 9) includes a step (902) of performing a conversion between a video including one or more video pictures, each video picture including a sub-picture including one or more slices, and a coded representation of the video whose coded representation conforms to a format rule, the format rule stipulating that when the rectangular slice mode is enabled for a video picture, the picture-level slice index for each slice of each sub-picture in the video picture is derived without explicit signaling in the coded representation, and stipulating that the number of coding tree units in each slice is derivable from the picture-level slice index.
[0110] The following solutions show exemplary embodiments of the techniques discussed in the previous chapter (e.g., item 2).
[0111] A video processing method including performing a conversion between a video including one or more pictures, each video picture including one or more sub-pictures having one or more slices, and a coded representation of the video, the coded representation conforming to format rules, the format rules being such that a sub-picture level slice index is derived based on information in the coded representation without signaling the sub-picture level slice index in the coded representation.
[0112] The method according to solution 2, wherein the format rules define that by using a rectangular slice structure, a sub-picture level slice index corresponds to an index to the slice in a list of slices in the sub-picture.
[0113] The method according to solution 2, wherein the format rules define that the sub-picture level slice index is derived from a specific value of a picture level slice index.
[0114] The following solutions show exemplary embodiments of the techniques discussed in the previous chapter (e.g., items 5, 6).
[0115] A video processing method including performing a conversion between a video including one or more video pictures, each video picture including one or more sub-pictures and / or one or more tiles, and a coded representation of the video, the coded representation conforming to format rules, and the conversion conforming to constraint rules.
[0116] The method according to solution 5, wherein the constraint rules define that a tile shall not be included in two or more sub-pictures.
[0117] 7. The method according to solution 5, wherein the constraint rule stipulates that a sub-picture cannot include two slices smaller than the corresponding tile to which the two slices belong.
[0118] The following solutions show exemplary embodiments of the techniques discussed in the previous chapter (e.g., items 7, 8).
[0119] 8. A video processing method including performing a conversion between a video including one or more video pictures, each video picture including one or more tiles and / or one or more slices, and a coded representation of the video, the coded representation conforming to format rules, the format rules stipulating that a field at the video picture level conveys information regarding the division of slices and / or tiles in the video picture.
[0120] 9. The method according to solution 8, wherein the field includes a video picture header.
[0121] 10. The method according to solution 8, wherein the field includes a set of picture parameters.
[0122] 11. The method according to any one of solutions 8 to 10, wherein the format rules stipulate that slice-level slice division information is omitted by including the slice division information of the field at the video picture level.
[0123] The following solutions show exemplary embodiments of the techniques discussed in the previous chapter (e.g., item 9).
[0124] 12. A video processing method including performing a conversion between a video including one or more pictures and a coded representation of the video, the conversion conforming to a division rule that sets the minimum number of slices for dividing a video picture as a function of whether rectangular division is used for dividing the video picture.
[0125] 13. The method according to solution 12, wherein the segmentation rule stipulates to use at least two slices for non-rectangular segmentation and at least one slice for rectangular segmentation.
[0126] 14. The method according to solution 12, wherein the segmentation rule is also a function of the number and / or the number of sub-pictures used to segment the video picture.
[0127] The following solutions show exemplary embodiments of the technology discussed in the previous chapter (e.g., items 10, 11).
[0128] 15. A video processing method including performing a conversion between a video slice of a video area of a video and a coded representation of the video, the coded representation conforming to a format rule, the format rule stipulating that the coded representation signals the video slice based on the upper left position of the video slice, and the format rule stipulating that the coded representation signals the height and / or width of the video slice in the segmentation information signaled at the video unit level.
[0129] 16. The method according to solution 15, wherein the format rule stipulates that the video slices are signaled in the order of the slices defined by the format rule.
[0130] 17. The method according to solution 15, wherein the video area corresponds to a sub-picture and the video unit level corresponds to a video picture.
[0131] The following solutions show exemplary embodiments of the technology discussed in the previous chapter (e.g., item 12).
[0132] 18. A video processing method including performing conversion between a video including video pictures and a coded representation of the video, the coded representation conforming to format rules, the format rules defining omission of signaling of a difference between a tile index of a first tile in a rectangular slice and a tile index of the first tile in a next rectangular slice.
[0133] 19. The method according to solution 18, wherein the difference is derivable from a 0th slice and the rectangular slice in the video picture.
[0134] The following solutions show exemplary embodiments of the technology discussed in the previous chapter (e.g., item 13).
[0135] 20. A video processing method including performing conversion between a video and a coded representation of the video, the coded representation conforming to format rules, the format rules defining control of signaling of information used to derive a number of tile columns or rows in a video picture based on a relationship between a width of the video picture and a size of a coding tree unit.
[0136] 21. The method according to solution 20, wherein the format rules define exclusion of signaling of a plurality of tile rows or a plurality of tile columns when the width of the video picture is less than or equal to the width of the coding tree unit.
[0137] The following solutions show exemplary embodiments of the technology discussed in the previous chapter (e.g., item 16).
[0138] 22. A video processing method including performing conversion between a video including one or more video pictures and a coded representation of the video, the coded representation conforming to format rules, the format rules defining inclusion of tile layout information in the coded representation for a video picture including tiles with uniform intervals and tiles with non-uniform intervals.
[0139] 23. The method according to solution 22, wherein the tile layout information is included in a syntax flag included in a set of picture parameters.
[0140] 24. The method according to any one of solutions 22 to 23, wherein the number of rows or columns of tiles signaled explicitly is of the same order as the number of tiles with non-uniform spacing.
[0141] 25. The method according to any one of solutions 22 to 23, wherein the number of rows or columns of tiles signaled explicitly is of the same order as the number of tiles with uniform spacing.
[0142] 26. The method according to any one of the above solutions, wherein the video region includes a video coding unit.
[0143] 27. The method according to any one of the above solutions, wherein the video region includes a video picture.
[0144] 28. The method according to any one of solutions 1 to 27, wherein the transformation includes encoding the video into the coded representation.
[0145] 29. The method according to any one of solutions 1 to 27, wherein the transformation includes decoding the coded representation to generate pixel values of the video.
[0146] 30. A video decoding apparatus comprising a processing device configured to implement the method according to one or more of solutions 1 to 29.
[0147] 31. A video encoding apparatus comprising a processing device configured to implement the method according to one or more of solutions 1 to 29.
[0148] 32. A computer program product having computer code stored thereon, which when executed by a processing device causes the processing device to implement the method according to any one of solutions 1 to 29.
[0149] 33. The method, apparatus, or system described in this specification.
[0150] FIG. 13 is a flowchart showing a video processing method 1300 according to the present technology. The method 1300 includes, in step 1310, converting a video including one or more tiles and a video including one or more rectangular slices into a bitstream of the video according to a rule. This rule stipulates updating a variable indicating a tile index only for slices having an index smaller than a value obtained by subtracting 1 from the number of rectangular slices in a video picture in order to repeatedly determine information regarding one or more rectangular slices.
[0151] In an embodiment of the present invention, the variable is not updated to determine information on the last rectangular slice in the video picture. In an embodiment of the present invention, the step of determining the information is performed using a set of picture parameters referred to by the video picture. In an embodiment of the present invention, one slice of one or more rectangular slices has a slice index of i, and the information includes at least one of a tile index of a tile including a first coding tree unit in the slice, a width of the slice, or a height of the slice. In an embodiment of the present invention, the information includes at least one of a list of the number of coding tree units in one or more rectangular slices, a top-left tile index of one or more rectangular slices, or a picture raster scan address of a coding tree block within the slice. In an embodiment of the present invention, updating the variable includes adding the difference between (1) a first tile index of a first tile including a first coding tree unit in a first slice with a slice index of (i + 1) and (2) a second tile index of a second tile including a first coding tree unit to be coded in a second slice with a slice index of i. In an embodiment of the present invention, updating the variable can include increasing the width of the slice in units of tile columns. In an embodiment of the present invention, the variable is updated by adding (A - 1)*B, where A represents the height of the slice in units of tile rows and B represents the number of columns of video tiles. In an embodiment of the present invention, the first tile of one or more tiles includes at least one rectangular slice, and the information on the at least one rectangular slice includes the height of each of the at least one rectangular slices.
[0152] FIG. 14 is a flowchart showing a video processing method 1400 according to the present technology. The method 1400 includes, in step 1410, performing a conversion between a video picture including one or more sub-pictures and a bitstream of the video. This conversion follows a rule that specifies omitting syntax elements in a sequence parameter set indicating a plurality of sub-pictures in the video picture when the maximum picture width and the maximum picture height are less than or equal to the dimensions of a coding tree block.
[0153] FIG. 15 is a flowchart showing a video processing method 1500 according to the present technology. The method 1500 includes, in step 1510, performing a conversion between a video picture including one or more tiles and a bitstream of the video. This conversion complies with a rule that specifies omitting, in the bitstream, syntax elements indicating the column widths of a plurality of explicitly provided tiles when the width of the video picture is less than or equal to the dimensions of a coding tree block.
[0154] FIG. 16 is a flowchart showing a video processing method 1600 according to the present technology. The method 1600 includes, in step 1610, performing a conversion between a video picture including one or more tiles and a bitstream of the video. This conversion complies with a rule that specifies omitting, in the bitstream, syntax elements indicating the number of heights of rows of explicitly provided tiles when the height of the video picture is less than or equal to the dimensions of a coding tree block.
[0155] In an embodiment of the present invention, the dimensions of a coding tree block are represented by a parameter CtbSizeY. In an embodiment of the present invention, it is inferred that the number of sub-pictures in a video picture is 0.
[0156] FIG. 17 is a flowchart showing a video processing method 1700 according to the present technology. The method 1700 includes, in step 1710, performing a conversion between a video picture including one or more tiles and a bitstream of the video. This conversion complies with a rule that, when the column width of a plurality of explicitly provided tiles is equal to the picture width in coding tree block units, one or more syntax elements indicating the column width of one or more tiles are omitted in the bitstream. In an embodiment of the present invention, it is inferred that one or more column widths are 0.
[0157] FIG. 18 is a flowchart showing a video processing method 1800 according to the present technology. The method 1800 includes, in step 1810, performing a conversion between a video picture including one or more tiles and a bitstream of the video. This conversion complies with a rule that, when the height of a row of a plurality of explicitly provided tiles is equal to the height of the picture in coding tree block units, one or more syntax elements indicating the height of a row of one or more tiles are omitted in the bitstream. In an embodiment of the present invention, it is inferred that the height of a row of one or more tiles is 0.
[0158] FIG. 19 is a flowchart showing a video processing method 1900 according to the present technology. The method 1900 includes, in step 1910, performing a conversion between a video picture including one or more slices and a bitstream of the video. This conversion complies with a rule that slice partition information is included in the bitstream.
[0159] In an embodiment of the present invention, the slice division information includes the width of the slice. In an embodiment of the present invention, the width of the slice is represented by subtracting X from the width of the slice, where X is a non-negative value. In an embodiment of the present invention, X is equal to 1, and the width of the slice is represented by variable slice_width_minus1[i]. Here, i represents the slice index of the slice. In an embodiment of the present invention, the slice division information includes the height of the slice. In an embodiment of the present invention, the height of the slice is represented by subtracting X from the height of the slice, where X is a non-negative value. In an embodiment of the present invention, X is equal to 1, and the height of the slice is represented by the variable slice_height_minus1[i], where i represents the slice index of the slice.
[0160] In an embodiment of the present invention, whether the slice division information includes the upper left position or the dimensions of the slice is conditional based on the characteristics of the slice. In an embodiment of the present invention, the slice division information includes the upper left position or the dimensions of the slice when the video picture includes one or more rectangular slices. In an embodiment of the present invention, the characteristics of the slice include the index of the slice, the dimensions of the coding tree block related to the slice, the dimensions of the video picture including the slice, or the number of slices in the video picture.
[0161] In an embodiment of the present invention, when the video picture includes one or more sub-pictures, whether the slice division information includes the upper left position or the dimensions of the slice is conditional based on the relationship between the division of the one or more slices and the division of the one or more sub-pictures. In an embodiment of the present invention, the slice division information includes the upper left position or the dimensions of the slice when at least one of the one or more sub-pictures includes two or more slices.
[0162] In an embodiment of the present invention, whether the slice division information includes the position or dimension of the upper left corner of the slice is conditional based on the number of slices in the video picture. In an embodiment of the present invention, the slice division information includes the position or dimension of the upper left corner of the slice to represent the slice when the number of slices of the video picture is greater than 1. In an embodiment of the present invention, the position or dimension of the upper left corner of at least one slice is represented by the unit dimension of the coding tree unit or the unit dimension of the tile. In an embodiment of the present invention, a syntax element is used for conversion to indicate whether the position or dimension of the upper left corner of at least one slice is represented by the unit dimension of the coding tree or the unit dimension of the tile. In an embodiment of the present invention, the syntax element indicates whether the position or dimension of the upper left corner of each slice is represented by the unit dimension of the coding tree unit or the unit dimension of the tile. In an embodiment of the present invention, the syntax element is slice_representated_in_ctb_flag[i], where i represents the index of the slice.
[0163] In an embodiment of the present invention, in the slice section information, at least one of the position or dimension of the upper left corner of the slice can be omitted, and the position or dimension of the upper left corner of the slice is inferred to a default value. In an embodiment of the present invention, the default value of the position of the upper left corner of the slice may include (0, 0). In an embodiment of the present invention, the default value of the dimension of the slice is based on the syntax element slice_representated_in_ctb_flag[i], where i represents the index of the slice. In an embodiment of the present invention, the default value of the width of the slice is equal to (slice_represented_in_ctb_flag[i]? ((pic_width_in_luma_samples + CtbSizeY - 1) >> CtbLog2SizeY) : NumTileColumns) - slice_top_left_x[i] - 1. In an embodiment of the present invention, the default value of the height of the slice is equal to (slice_represented_in_ctb_flag[i]? ((pic_height_in_luma_samples + CtbSizeY - 1) >> CtbLog2SizeY) : NumTileRows) - slice_top_left_y[i] - 1.
[0164] FIG. 20 is a flowchart showing a video processing method 2000 according to the present technique. The method 2000 includes, in step 2010, performing a conversion between a video tile including one or more rectangular slices according to a rule and a video bitstream. This rule defines determining a uniform slice height based on a first syntax element that defines the height of a rectangular slice in a video tile in units of rows of coding tree units and a second syntax element that defines the number of slice heights explicitly provided in the video tile.
[0165] In an embodiment of the present invention, the first syntax element is exp_slice_height_in_ctus_minus1[i], the second syntax element is num_exp_slices_in_tile[i], i is the index of a rectangular slice, and the height of a uniform slice is based on exp_slice_height_in_ctus_minus1[i][num_exp_slices_in_tile[i]-1]. In an embodiment of the present invention, the height of the (num_exp_slices_in_tile[i]-1)-th non-uniform slice is based on exp_slice_height_in_ctus_minus1[i][num_exp_slices_in_tile[i]-1].
[0166] FIG. 21 is a flowchart showing a video processing method 2100 according to the present technology. Method 2100 includes, in step 2110, performing conversion between a video picture having one or more tiles and a video bitstream according to a rule. This rule stipulates that a syntax element is equal to or greater than the dimensions of a column or row of uniform tiles, and this syntax element indicates the coding tree block unit dimension excluding the total dimension of the widths of columns or heights of rows of a plurality of clearly provided tiles.
[0167] In an embodiment of the present invention, the syntax element is firstRemainingWidthInCtbsY. In an embodiment of the present invention, the syntax element is firstRemainingHeightInCtbsY.
[0168] FIG. 22 is a flowchart showing a video processing method 2200 according to the present technology. Method 2200 includes, in step 2210, performing conversion between a video picture having one or more tiles and a video bitstream according to a rule. This rule stipulates that a syntax element is greater than or equal to the height of a uniform slice, and this syntax element indicates the height of a coding tree unit block excluding the total height of a plurality of clearly provided slice heights.
[0169] FIG. 23 is a flowchart showing a video processing method 2300 according to the present technology. Method 2300 includes, in step 2310, performing conversion between a video and a bitstream of the video according to a rule. This rule stipulates using a syntax element for conversion so as to indicate the maximum number of affine merge candidates permitted in a sub-block-based merge candidate list.
[0170] In an embodiment of the present invention, whether a syntax element is signaled for conversion is based on whether an affine prediction tool is enabled. In an embodiment of the present invention, when a syntax element is omitted in a bitstream, it is assumed that the maximum number of affine combination candidates permitted in a sub-block-based merge candidate list is 0. In an embodiment of the present invention, the maximum number of affine merge candidates permitted in a sub-block-based merge candidate list is equal to the difference between 5 and the syntax element. In an embodiment of the present invention, the syntax element is in the range including [0, X], where X is an integer. In an embodiment of the present invention, X is 5.
[0171] In an embodiment of the present invention, when a syntax element is omitted in a bitstream, the syntax element is inferred to be 5. In an embodiment of the present invention, the maximum number of affine merge candidates permitted in a sub-block-based merge candidate list is determined by the syntax element and the maximum number of sub-block-based temporal motion vector prediction (TMVP) merge candidates. In an embodiment of the present invention, the maximum number of affine merge candidates permitted in a sub-block-based merge candidate list is equal to Min(5, (sps_sbtmvp_enabled_flag && ph_temporal_mvp_enable_flag) + 5 - five_minus_max_num_affine_merge_cand), where five_minus_max_num_affine_merge_cand is a syntax element.
[0172] In some embodiments, the conversion includes encoding the video into the bitstream. In some embodiments, the conversion includes decoding the video from the bitstream.
[0173] In the solutions described herein, an encoder can comply with the format rules by generating a coded representation according to the format rules. In the solutions described herein, a decoder may use this format rule to generate a decoded video by parsing the syntax elements in the coded representation while knowing the presence or absence of the syntax elements according to the format rules.
[0174] As used herein, the term "video processing" can refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm may be applied during the conversion from the pixel representation of a video to the corresponding bitstream representation or vice versa. The bitstream representation of the current video block may correspond to bits that are spread to the same or different locations within the bitstream, as defined by the syntax, for example. For example, a macroblock may be encoded in terms of the transformed and coded residual error values and using bits in the header and other fields in the bitstream. Further, during the conversion, the decoder may parse the bitstream based on a determination and with the knowledge that some fields may or may not be present, as explained in the above solutions. Similarly, the encoder may determine whether a particular syntax field should or should not be included and generate a coded representation accordingly by including or excluding the syntax field from the coded representation.
[0175] The disclosed and other solutions, examples, embodiments, modules, and implementations of functional operations described in this specification may be implemented in digital electronic circuits, or in computer software, firmware, or hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them. The disclosed and other embodiments can be implemented as one or more computer program products, i.e., as one or more modules of computer program instructions encoded on a computer-readable medium for being implemented by, or for controlling the operation of, a data processing apparatus. This computer-readable medium may be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition of matter that provides a machine-readable propagated signal, or a combination of one or more of them. The term "data processing apparatus" includes, for example, a programmable processor, a computer, or multiple processors or computers, and all apparatus devices and machines for processing data. This apparatus can include, in addition to hardware, code that creates an execution environment for the computer program, e.g., processor firmware, protocol stack, database management system, operating system, or code that constitutes a combination of one or more of them. A propagated signal is an artificially generated signal, e.g., an electrical, optical, or electromagnetic signal generated by a machine, and is generated for encoding information for transmission to a suitable receiving device.
[0176] A computer program (also referred to as a program, software, software application, script, or code) can be described in any form of programming language, including a compiled language or an interpreted language, and it can be deployed in any form, either as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. The program may be recorded as part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), stored in a single file dedicated to the program, or stored in a plurality of coordinating files (e.g., files that hold one or more modules, subprograms, or portions of code). It is also possible to deploy the computer program to be executed on one computer located at one site, or on multiple computers distributed across multiple sites and interconnected by a communication network.
[0177] The processes and logic flows described herein can be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by special-purpose logic circuits, such as an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit), and the apparatus can also be implemented as special-purpose logic circuits.
[0178] Processors suitable for the execution of a computer program include, for example, both general and special purpose microprocessors, as well as any one or more processors of any kind of digital computer. Generally, a processor receives instructions and data from a read-only memory or a random access memory or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing the instructions and data. Generally, a computer may include one or more mass storage devices for storing data, such as magnetic, magneto-optical disks, or optical disks, or may be operatively coupled to receive data from or transfer data to such mass storage devices. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, EPROM, EEPROM, flash memory devices, magnetic disks, such as internal hard disks or removable disks, magneto-optical disks, and semiconductor memory devices such as CD-ROM and DVD-ROM disks. The processor and the memory may be supplemented by, or incorporated in, application-specific logic circuitry.
[0179] This patent specification includes many details, but these should not be construed as limiting the scope of any subject matter or the scope of the claims. Rather, they should be construed as descriptions of features that may be specific to particular embodiments of a particular technology. Specific features described in the context of separate embodiments in this patent document may be implemented in combination in one example. Conversely, various features described in the context of one example may be implemented separately or in any suitable sub-combination in multiple embodiments. Further, features may be described and initially claimed above as acting in a particular combination, but one or more features from the claimed combination may, in some cases, be excised from the combination, and the claimed combination may be directed to a sub-combination or variation of a sub-combination.
[0180] Similarly, operations are shown in a particular order in the drawings, but this should not be understood as requiring that such operations be performed in the particular order shown or in a sequential order to achieve the desired result, or that all of the operations shown be performed. Also, the separation of the various system components in the examples described in this patent specification should not be understood as requiring such separation in all embodiments.
[0181] Only some implementations and examples are described, and based on the content described and illustrated in this patent document, other embodiments, extensions, and variations are possible.< / video> < / video> < / video> < / output>
Claims
Claim 1 Determining to apply a scanning process to the video picture for conversion between a video having video pictures and a bitstream of the video, wherein the video picture is divided into one or more tiles, one or more slices, and a plurality of coding tree units, In a first scanning process, determining that a first slice height in a video tile including a first slice in a unit of a coding tree unit (CTU) is derived based on a value of a first syntax element corresponding to the first slice or is included in the bitstream, Performing the conversion based on the determination, The first syntax element is included in a picture parameter set in the bitstream in response to a set of conditions being satisfied, The set of conditions is that the first slice is in a rectangular mode, a value of a syntax element defining a difference between a width of the first slice and 1 in a unit of a tile column is equal to 0, and a value of a syntax element defining a difference between a height of the first slice and 1 in a unit of a tile row is equal to 0, The height of the first slice of the video tile including the first slice in a unit of a coding tree unit is derived for the video in response to the value of the first syntax element being a first value, In a second scanning process, when a rectangular slice mode is used for the video picture, determining that a variable indicating a tile index of a tile including a first coding tree unit in a slice having the picture level slice index is updated when the picture level slice index is less than a value of a third syntax element and is not updated when the picture level slice index is greater than or equal to the value of the third syntax element, The third syntax element is included in a picture parameter set referred to by the video picture in the bitstream for deriving information about one or more slices in the video picture, and the number of slices in the video picture is greater than the value of the third syntax element, A method. Claim 2 When the first syntax element does not exist in the bitstream, the first syntax element is presumed to have the first value, The method according to claim 1. Claim 3 When the value of the first syntax element is the first value, the height of the first slice in the video tile including the first slice in the unit of the coating tree unit is RowHeight[TileIdx / NumTileColumns], RowHeight[j] defines the height of the j-th tile row in the unit of the coding tree block, TileIdx defines the tile index of the tile including the first CTU in the first slice, NumTileColumns defines the number of tile columns, The method according to claim 1 or 2.
4. When the first syntax element has a value of N, N is different from the first value, and N second syntax elements that define the value of the slice height in the video tile including the first slice in the unit of the CTU row are each included in the picture parameter set in the bitstream, The method according to any one of claims 1 to 3.
5. When the tile including the second slice is not divided into a plurality of slices, for the second slice having a picture level slice index less than the value of the third syntax element, the width of the second slice in the unit of the tile column is derived from a fourth syntax element corresponding to the second slice in the picture parameter set, and the height of the second slice in the unit of the tile row is derived from a fifth syntax element corresponding to the second slice in the picture parameter set, For the third slice having a picture level slice index greater than or equal to the value of the third syntax element, the width of the third slice in the unit of the tile column is derived from the difference between the number of tile columns and tileX, and the height of the third slice in the unit of the tile row is derived from the difference between the number of tile rows and tileY, tileX and tileY are determined based on the number of tile columns, The method according to claim 1.
6. The slice has a picture level slice index of i (i: integer), Updating the variable is to add the difference between the first tile index of the first tile including the first coding tree unit in the slice having an index of (i + 1) and the tile index of the tile including the first coding tree unit in the slice having an index of i, The method according to claim 5.
7. The difference is defined by a fourth syntax element included in the picture parameter set for the slice, when the index of i is greater than or equal to the value of the third syntax element, the fourth syntax element does not exist in the picture parameter set for the slice having the index of i, updating the variable includes adding the difference in response to the fourth syntax element being included in the picture parameter set, The method according to claim 6.
8. when the fourth syntax element does not exist in the picture parameter set, updating the variable includes adding the width of the slice having the index of i, and the width is expressed in units of tile columns, The method according to claim 7.
9. when tileIdx%B is equal to 0, updating the variable includes adding (A - 1)*B, tileIdx indicates the variable, A is the height of the slice having the index of i in units of tile rows, B is the number of tile columns in the video picture, % is a modulo operation, The method according to claim 8.
10. when the value of the third syntax element is greater than 0, the width and height of the slice are defined by one or more syntax elements included in the bitstream, The method according to any one of claims 4 to 9.
11. the width of the slice is expressed in units of tile columns, and the height of the third slice is expressed in units of tile rows, The method according to claim 10.
12. one plus the value of the third syntax element defines the number of rectangular slices in the video picture, The method according to any one of claims 4 to 11.
13. when the height of the video tile is greater than the sum of the heights of the N slices, the height of the remaining slices in the video tile other than the N slices is determined based on the height of the last slice of the N slices, The method according to any one of claims 4 to 12.
14. the height of one of the N slices in units of CTU rows is equal to one plus the value of one of the N second syntax elements corresponding to the slice, When the difference between the height of the video tile and the sum of the heights of the N slices is about the same as the height of the last slice of the N slices in a coding tree block unit, the height of the remaining one slice of the video tile other than the N slices is set equal to the height of the last slice of the N slices, or, When the difference between the height of the video tile and the sum of the heights of the N slices is about the same as the height of a uniform slice in a coding tree block unit, the height of the uniform slice is used to recursively divide the remaining portion of the video tile other than the N slices to form the one or more uniform slices until the remaining portion of the video tile other than the N slices and the one or more uniform slices have a height less than the height of the uniform slice, when the height of the uniform slice is equal to the height of the last slice of the N slices and the remaining portion of the video tile other than the N slices and the one or more uniform slices have a height less than the height of the uniform slice, the height of the last slice of the video tile is set to the difference between the height of the video tile and the sum of the heights of the N slices and the one or more uniform slices, or, when the difference between the height of the video tile and the sum of the slice heights of the N slices is less than the height of a uniform slice in a coding tree unit, the height of the (N + 1)-th slice of the video tile is set to the difference, The method according to claim 13.
15. the height of the last slice of the N slices is not permitted to be reset, and the value of the last second syntax element of the N second syntax elements is directly used to derive the height of a uniform slice without referring to other information, when the value of the fourth syntax element is 0, the second syntax element does not exist in the bitstream, and the height of the slice in the video tile is derived to be equal to the height of the video tile in a coding tree block unit, The method according to claim 13, which cites claim 5 or claim 7. **Claim 16**: In the third scanning process, further comprising determining that the picture parameter set includes a list of syntax elements indicating the width of the tile columns for P tile columns each having P (P: integer) indexes. The list of syntax elements includes one added to a sixth syntax element that directly defines the width of the P-th tile column of the P tile columns in a coding tree block unit without referring to other information, and the value of the sixth syntax element is used to derive the width of a tile column having an index larger than the P indexes. The method according to any one of claims 1 to 15. **Claim 17** The value of P is indicated by a seventh syntax element included in the picture parameter set. The width for the P-th tile column of the P tile columns in a coding tree block unit is not permitted to be reset. The seventh syntax element is the P-th entry in the list of syntax elements, and the uniform tile column width is set to the width for the P-th tile column of the P tile columns. When the difference between the width of the picture of the luminance component and the total width of the tile columns of the P tile columns in a coding tree block unit is greater than or equal to the width of the P-th tile column of the P tile columns, the width of the (P + 1)-th tile column is set equal to the width for the P-th tile column of the P tile columns. When the difference between the width of the picture of the luminance component and the total width of the tile columns of the P tile columns in a coding tree block unit is less than the width of the P-th tile column of the P tile columns, the width of the (P + 1)-th tile column is set equal to the difference. The picture parameter set further includes a list of syntax elements indicating the height of the tile rows for M tile rows each having M (M: integer) indexes. The list of syntax elements includes one added to a ninth syntax element that directly defines the height of the M-th tile row of the M tile rows in a coding tree block unit without referring to other information, and the value of the ninth syntax element is used to derive the height of a tile row having an index larger than the M indexes. The value of M is indicated by a tenth syntax element included in the picture parameter set, and the height for the M-th tile row of the M tile rows in a coding tree block unit is not allowed to be reset. The tenth syntax element is the M-th entry in the list of the syntax elements, and the height of a uniform tile row is set to the height for the M-th tile row of the M tile rows. When the difference between the height of the picture of the luminance component and the total height of the tile rows of the M tile rows in a coding tree block unit is greater than or equal to the height for the M-th tile row of the M tile rows, the height of the (M + 1)-th tile row is set to be equal to the height for the M-th tile row of the M tile rows. When the difference between the height of the picture of the luminance component and the total height of the tile rows of the M tile rows in a coding tree block unit is less than the height for the M-th tile row of the P tile rows, the width of the height of the (M + 1)-th tile row is set to be equal to the difference. The method according to claim 16.
18. The conversion includes encoding the video into the bitstream. The method according to any one of claims 1 to 17.
19. The conversion includes decoding the video from the bitstream. The method according to any one of claims 1 to 17.
20. An apparatus for processing video data comprising a processor and a non-transitory memory having instructions, wherein when executed by the processor, the instructions cause the processor to determine to apply a scanning process to the video picture for conversion between the video having video pictures and the bitstream of the video, and the video picture is divided into one or more tiles, one or more slices, and a plurality of coding tree units. In a first scanning process, determine that the first slice height in a video tile including a first slice in a unit of a coding tree unit (CTU) is derived based on the value of a first syntax element corresponding to the first slice or included in the bitstream. perform the conversion based on the determination. The first syntax element is included in a picture parameter set in the bitstream in response to a set of conditions being satisfied. The set of conditions is that the first slice is in rectangular mode, the value of the syntax element that defines the difference between the width of the first slice and 1 in units of a tile column is equal to 0, and the value of the syntax element that defines the difference between the height of the first slice and 1 in units of a tile row is equal to 0. The height of the first slice of the video tile including the first slice in units of a coding tree unit of the coding tree unit is derived for the video according to the value of the first syntax element being a first value. In a second scanning process, when a rectangular slice mode is used for the video picture, when a picture level slice index is less than a value of a third syntax element, a variable indicating a tile index of a tile including a first coding tree unit in a slice having the picture level slice index is updated, and when the picture level slice index is greater than or equal to the value of the third syntax element, it is determined that it is not updated. The third syntax element is included in a picture parameter set referred to by the video picture in the bitstream for deriving information on one or more slices in the video picture, and the number of slices in the video picture is greater than the value of the third syntax element. Apparatus. [
21. ] A non-transitory computer-readable storage medium storing instructions, the instructions causing a processor to determine to apply a scanning process to the video picture for conversion of a video having the video picture and a bitstream of the video, the video picture being divided into one or more tiles, one or more slices, and a plurality of coding tree units; in a first scanning process, determine that a first slice height in a video tile including a first slice in units of a coding tree unit (CTU) is derived based on a value of a first syntax element corresponding to the first slice or is included in the bitstream; perform the conversion based on the determination; the first syntax element is included in a picture parameter set in the bitstream according to a set of conditions being satisfied. The set of conditions is that the first slice is in rectangular mode, the value of the syntax element that defines the difference between the width of the first slice and 1 in units of a tile column is equal to 0, and the value of the syntax element that defines the difference between the height of the first slice and 1 in units of a tile row is equal to 0. The height of the first slice of the video tile including the first slice in units of a coding tree unit of the coding tree unit is derived for the video according to the value of the first syntax element being a first value. In a second scanning process, when a rectangular slice mode is used for the video picture, when the picture level slice index is less than the value of a third syntax element, a variable indicating the tile index of a tile including a first coding tree unit in a slice having the picture level slice index is updated, and it is determined that it is not updated when the picture level slice index is greater than or equal to the value of the third syntax element. The third syntax element is included in a picture parameter set referred to by the video picture in the bitstream to derive information regarding one or more slices in the video picture, and the number of slices in the video picture is greater than the value of the third syntax element. Non-transitory computer-readable storage medium. [
22. ] A method for storing a bitstream of a video, the method comprising: determining to apply a scanning process to the video picture for conversion between the video having the video picture and the bitstream of the video, the video picture being divided into one or more tiles, one or more slices, and a plurality of coding tree units; in a first scanning process, determining that the height of a first slice in a video tile including the first slice in units of a coding tree unit (CTU) is derived based on the value of a first syntax element corresponding to the first slice or is included in the bitstream; performing the conversion based on the determination; storing the bitstream in a non-transitory computer-readable storage medium. The first syntax element is included in a picture parameter set in the bitstream according to a set of conditions being satisfied. The one set of conditions is that the first slice is in the rectangular mode, the value of the syntax element that defines the difference between the width of the first slice and 1 in units of a tile column is equal to 0, and the value of the syntax element that defines the difference between the height of the first slice and 1 in units of a tile row is equal to 0. The height of the first slice of the video tile including the first slice in units of a coding tree unit is derived for the video in response to the value of the first syntax element being a first value. In a second scanning process, when the rectangular slice mode is used for the video picture, when the picture level slice index is less than the value of a third syntax element, a variable indicating the tile index of a tile including a first coding tree unit in the slice having the picture level slice index is updated, and when the picture level slice index is greater than or equal to the value of the third syntax element, it is determined that it is not updated. The third syntax element is included in a picture parameter set referred to by the video picture in the bitstream to derive information about one or more slices in the video picture, and the number of slices in the video picture is greater than the value of the third syntax element. Method.