Image segmentation in video encoding and decoding

By dividing video pictures into sub-pictures, strips and slices, and using rules and format regulations to control the conversion between video pictures and bitstreams, the problems of high encoding and decoding overhead and difficulty in parallel processing in the existing technology are solved, and more efficient video encoding and decoding is achieved.

CN115211110BActive Publication Date: 2025-09-16DOUYIN VISION CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202180016200.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-02-21
Filing Date
2021-02-22
Publication Date
2025-09-16
Estimated Expiration
2041-02-22

AI Technical Summary

Technical Problem

Existing video codec standards have problems with high coding and decoding overhead, difficulty in parallel processing and MTU size matching when processing the conversion between video pictures and bitstreams, especially in terms of video picture segmentation and prediction dependency.

Method used

A new video processing method is adopted to divide the video pictures into sub-pictures, strips and slices, and use rules and format regulations to control the conversion between video pictures and bit streams, including rectangular and non-rectangular segmentation, mapping of strip indexes, and omission and signaling notification of slices to achieve efficient codec conversion.

Benefits of technology

It improves encoding and decoding efficiency, reduces encoding and decoding overhead, supports more efficient parallel processing and MTU size matching, and improves video processing performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115211110B_ABST
    Figure CN115211110B_ABST
Patent Text Reader

Abstract

Techniques for video processing, including video encoding, decoding, and transcoding, are described. An example method includes performing conversion between a video picture of a video and a bitstream of the video according to a rule. A video picture includes one or more slices. The rule specifies that, in response to at least one condition being satisfied, a syntax element indicating a difference between slice indices of two rectangular slices is present in the bitstream. A first of the two rectangular slices is denoted as the i-th rectangular slice, where i is an integer.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to and the benefit of International Patent Application No. PCT / CN2020 / 076158, filed on February 21, 2020, under applicable patent law and / or under the Paris Convention. The entire disclosure of the aforementioned application is incorporated by reference into this disclosure for all legal purposes. Technical Field

[0003] This patent document relates to image encoding and decoding and video encoding and decoding. Background Art

[0004] Digital video accounts for the largest use of bandwidth on the Internet and other digital communications networks. As the number of connected user devices capable of receiving and displaying video increases, bandwidth demand for digital video usage is expected to continue to grow. Summary of the Invention

[0005] This document discloses techniques that can be used by video encoders and decoders to process codec representations of video using control information useful for decoding the codec representations.

[0006] In one example aspect, a video processing method is disclosed. The method includes performing conversion between a video picture of a video and a bitstream of the video according to a rule. The video picture includes one or more slices, and the rule specifies that a syntax element indicating a difference between slice indices of two rectangular slices is present in the bitstream in response to at least one condition being satisfied. A first rectangular slice of the two rectangular slices is denoted as an i-th rectangular slice, where i is an integer.

[0007] In another example aspect, a video processing method is disclosed. The method includes performing conversion between a video picture of a video and a bitstream of the video according to a rule. The video picture includes one or more sub-pictures, and each sub-picture includes one or more rectangular slices. The rule provides for deriving a slice index at the sub-picture level for each rectangular slice in each sub-picture to determine the number of codec tree units in each slice.

[0008] In another example aspect, a video processing method is disclosed. The method includes determining a mapping between a sub-picture-level slice index of a slice in a sub-picture and a picture-level slice index of the slice for conversion between a video picture including one or more sub-pictures of the video and a bitstream of the video. The method also includes performing the conversion based on the determination.

[0009] In another example aspect, a video processing method is disclosed. The method includes performing conversion between a video picture of a video and a bitstream of the video according to a rule. The video picture includes one or more sub-pictures, and the rule specifies that a slice of the video is entirely within a single sub-picture of the video picture.

[0010] In another example aspect, a video processing method is disclosed. The method includes performing conversion between a video picture of a video and a bitstream of the video, wherein the video picture includes one or more sub-pictures. The bitstream conforms to a format rule that specifies that information about sub-pictures is included in a syntax structure associated with the picture.

[0011] In another example aspect, a video processing method is disclosed. The method includes converting between a video picture of a video including one or more slices having a non-rectangular shape and a bitstream of the video, determining slice segmentation information for the video picture, and performing the conversion based on the determination.

[0012] In another example aspect, a video processing method is disclosed. The method includes performing conversion between a video picture comprising one or more slices of a video and a bitstream of the video according to a rule. The rule specifies that the number of slices in the video picture is equal to or greater than a minimum number of slices determined based on whether rectangular or non-rectangular segmentation is applied to the video picture.

[0013] In another example aspect, a video processing method is disclosed. The method includes performing conversion between a video picture of a video and a bitstream of the video according to a rule. The video picture includes one or more slices. When slice segmentation information of the video picture is included in a syntax structure of a video unit, the slice is represented by the upper left position and dimensions of the slice.

[0014] In another exemplary aspect, a video processing method is disclosed. The method includes performing conversion between a video picture of a video and a bitstream of the video according to a rule. The video picture includes one or more sub-pictures, and each sub-picture includes one or more slices. The rule specifies how segmentation information for the one or more slices in each sub-picture is presented in the bitstream.

[0015] In another example aspect, a video processing method is disclosed. The method includes performing conversion between a video picture of a video and a bitstream of the video according to a rule. The video picture includes one or more rectangular slices, and each slice includes one or more slices. The rule omits signaling of a difference between a first slice index of a first slice in an i-th rectangular slice and a second slice index of a first slice in an (i+1)th rectangular slice in the bitstream.

[0016] In another example aspect, a video processing method is disclosed. The method includes: for converting between a video picture of a video and a bitstream of the video, determining, in response to a relationship between dimensions of the video picture and dimensions of a codec treeblock, that information regarding the number of columns and rows of slices in a derived video picture is conditionally included in the bitstream. The method also includes performing the conversion based on the determination.

[0017] In another example aspect, a video processing method is disclosed. The method includes performing conversion between a video picture of a video and a bitstream of the video. The video picture includes one or more sub-pictures. The bitstream complies with a format rule that specifies that, if a variable specifying a sub-picture identifier of a sub-picture of a slice is present in the bitstream, there is one and only one syntax element that satisfies a condition that a second variable corresponding to the syntax element is equal to the variable.

[0018] In another example aspect, a video processing method is disclosed. The method includes performing conversion between a video picture of a video and a bitstream of the video. The video picture includes one or more sub-pictures. When non-rectangular partitioning is applied or sub-picture information is omitted in the bitstream, two slices in a slice have different addresses.

[0019] In another exemplary aspect, a video processing method is disclosed. The method includes performing conversion between a video picture of a video and a bitstream of the video according to a rule. The video picture includes one or more slices. The rule specifies that when one or more slices are organized in both uniform and non-uniform spacing, a syntax element is used to indicate the type of slice layout.

[0020] In another example aspect, a video processing method is disclosed. The method includes performing conversion between a video picture of a video and a bitstream of the video according to a rule. The rule specifies whether or how to handle a Merge Estimation Region (MER) size in the conversion depending on a minimum allowed codec block size.

[0021] In another example aspect, a video processing method is disclosed. The method includes performing conversion between a video including at least one video slice and a bitstream of the video according to a rule. The rule specifies deriving heights of slices in codec tree units for the video slice based on a value of a first syntax element in the bitstream, the value of the first syntax element indicating a number of explicitly provided slice heights for slices in the video slice including the slice.

[0022] In another example aspect, a video processing method is disclosed. The method includes performing conversion between a video comprising a video picture and a bitstream of the video according to a rule, the video picture comprising a video slice comprising one or more slices. The rule specifies that a second slice in a slice comprising a first slice in the picture has a height expressed in units of codec tree units. The first slice has a first slice index, and the second slice has a second slice index, the second slice index being determined based on the first slice index and a number of explicitly provided slice heights in the video slice. The height of the second slice is determined based on the first slice index and the second slice index.

[0023] In another example aspect, a video processing method is disclosed. The method includes performing conversion between a video and a bitstream of the video, the video including a video picture including one or more slices. The video picture references a picture parameter set, and the picture parameter set conforms to a format rule that specifies that the picture parameter set includes a list of column widths for N slice columns, where N is an integer. There is an N-1th slice column in the video picture, and the width of the N-1th slice column is equal to the N-1th entry in the list of explicitly included slice column widths plus one codec tree block.

[0024] In another example aspect, a video processing method is disclosed. The method includes performing conversion between a video and a bitstream of the video, the video including a video picture including one or more slices. The video picture references a picture parameter set, and the picture parameter set conforms to a format rule that specifies that the picture parameter set includes a list of row heights for N slice rows, where N is an integer. There is an N-1th slice row in the video picture, and the height of the N-1th slice row is equal to the N-1th entry in the list of explicitly included slice row heights plus one codec tree block.

[0025] In another example aspect, a video processing method is disclosed. The method includes performing a conversion between a video comprising one or more video pictures and a codec representation of the video, wherein each video picture comprises one or more sub-pictures, each sub-picture comprising one or more slices, wherein the codec representation conforms to a format rule; wherein the format rule provides for deriving a picture-level slice index for each slice in each sub-picture in the video picture when rectangular slice mode is enabled, without explicit signaling in the codec representation; wherein the format rule provides for a number of codec tree units in each slice to be derived from the picture-level slice index.

[0026] In another example aspect, another video processing method is disclosed. The method includes performing a conversion between a video comprising one or more video pictures and a codec representation of the video, wherein each video picture comprises one or more sub-pictures, the sub-pictures comprising one or more slices, wherein the codec representation conforms to a format rule, wherein the format rule provides that a sub-picture-level slice index can be derived based on information in the codec representation without signaling the sub-picture-level slice index in the codec representation.

[0027] In another example aspect, another video processing method is disclosed. The method includes performing a conversion between a video comprising one or more video pictures and a codec representation of the video, wherein each video picture comprises one or more sub-pictures and / or one or more slices, wherein the codec representation complies with a format rule, and wherein the conversion complies with a constraint rule.

[0028] In another example aspect, another video processing method is disclosed. The method includes performing conversion between a video comprising one or more video pictures and a codec representation of the video, wherein each video picture comprises one or more slices and / or one or more slices; wherein the codec representation conforms to a format rule; wherein the format rule specifies that a field at the video picture level carries information about segmentation of the slices and / or slices in the video picture.

[0029] In another example aspect, another video processing method is disclosed that includes performing a conversion between a video comprising one or more pictures and a codec representation of the video, wherein the conversion complies with a segmentation rule that determines whether rectangular segmentation is used to segment the video pictures and a minimum number of slices into which the video pictures are segmented.

[0030] In another example aspect, another video processing method is disclosed. The method includes performing conversion between a video slice of a video region of a video and a codec representation of the video, wherein the codec representation conforms to a format rule, wherein the format rule provides that the codec representation signals the video slice based on a top left position of the video slice, wherein the format rule provides that the codec representation signals a height and / or a width of the video slice in segmentation information, wherein the segmentation information is signaled at a video unit level.

[0031] In another example aspect, another video processing method is disclosed. The method includes performing a conversion between a video including video pictures and a codec representation of the video, wherein the codec representation conforms to a format rule, wherein the format rule provides for omitting signaling a difference between a slice index of a first slice in a rectangular slice and a slice index of a first slice in a next rectangular slice.

[0032] In another example aspect, another video processing method is disclosed. The method includes performing conversion between a video and a codec representation of the video, wherein the codec representation conforms to a format rule, wherein the format rule specifies a relationship between a width of a video picture and a size of a codec tree unit to control signaling of information used to derive a number of slice columns or a number of slice rows in the video picture.

[0033] In another example aspect, another video processing method is disclosed. The method includes performing conversion between a video comprising one or more video pictures and a codec representation of the video, wherein the codec representation conforms to a format rule, wherein the format rule specifies including slice layout information in the codec representation of the video pictures comprising evenly spaced slices and non-evenly spaced slices.

[0034] In yet another exemplary aspect, a video encoder apparatus is disclosed. The video encoder includes a processor configured to implement the above method.

[0035] In yet another exemplary aspect, a video decoder apparatus is disclosed. The video decoder includes a processor configured to implement the above method.

[0036] In yet another exemplary aspect, a computer-readable medium having stored thereon code is disclosed. The code is in the form of processor-executable code embodying one of the methods described herein.

[0037] These features and others are described throughout this document. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 An example of raster scan stripe partitioning of a picture is shown, where the picture is divided into 12 slices and 3 raster scan stripes.

[0039] Figure 2 An example of rectangular strip partitioning of a picture is shown, where the picture is divided into 24 slices (6 slice columns and 4 slice rows) and 9 rectangular strips.

[0040] Figure 3 An example of partitioning a picture into slices and rectangular strips is shown, where the picture is divided into 4 slices (2 slice columns and 2 slice rows) and 4 rectangular strips.

[0041] Figure 4 It shows that the picture is partitioned into 18 slices, 24 slices and 24 sub-pictures.

[0042] Figure 5 The nominal vertical and horizontal positions of the 4:2:2 luma and chroma samples in the picture are shown.

[0043] Figure 6An example of picture partitioning is shown. Solid lines 602 represent slice boundaries; dashed lines 604 represent stripe boundaries, and dashed lines 606 represent sub-picture boundaries. The figure indicates the picture-level index, decoding order index, sub-picture-level index, sub-picture index, and slice index of the four slices.

[0044] Figure 7 is a block diagram of an example video processing system.

[0045] Figure 8 It is a block diagram of a video processing device.

[0046] Figure 9 is a flow chart of an example method of video processing.

[0047] Figure 10 is a block diagram illustrating a video encoding and decoding system according to some embodiments of the present disclosure.

[0048] Figure 11 is a block diagram illustrating an encoder according to some embodiments of the present disclosure.

[0049] Figure 12 is a block diagram illustrating a decoder according to some embodiments of the present disclosure.

[0050] Figure 13 is a flowchart representation of a method for video processing according to the present technology.

[0051] Figure 14 is a flowchart representation of another method for video processing according to the present technology.

[0052] Figure 15 is a flowchart representation of another method for video processing according to the present technology.

[0053] Figure 16 is a flowchart representation of another method for video processing according to the present technology.

[0054] Figure 17 is a flowchart representation of another method for video processing according to the present technology.

[0055] Figure 18 is a flowchart representation of another method for video processing according to the present technology.

[0056] Figure 19 is a flowchart representation of another method for video processing according to the present technology.

[0057] Figure 20 is a flowchart representation of another method for video processing according to the present technology.

[0058] Figure 21is a flowchart representation of another method for video processing according to the present technology.

[0059] Figure 22 is a flowchart representation of another method for video processing according to the present technology.

[0060] Figure 23 is a flowchart representation of another method for video processing according to the present technology.

[0061] Figure 24 is a flowchart representation of another method for video processing according to the present technology.

[0062] Figure 25 is a flowchart representation of another method for video processing according to the present technology.

[0063] Figure 26 is a flowchart representation of another method for video processing according to the present technology.

[0064] Figure 27 is a flowchart representation of another method for video processing according to the present technology.

[0065] Figure 28 is a flowchart representation of another method for video processing according to the present technology.

[0066] Figure 29 is a flowchart representation of another method for video processing according to the present technology.

[0067] Figure 30 is a flowchart representation of another method for video processing according to the present technology.

[0068] Figure 31 is a flowchart representation of another method for video processing according to the present technology. DETAILED DESCRIPTION

[0069] The section headings used in this document are intended to facilitate understanding and do not limit the application of the techniques and embodiments disclosed in each section to that section. Furthermore, the use of H.266 terminology in some descriptions is intended to facilitate understanding and is not intended to limit the scope of the disclosed techniques. Thus, the techniques described herein are also applicable to other video codec protocols and designs.

[0070] 1. Overview

[0071] This document relates to video codec technology. Specifically, it is about the signaling of sub-pictures, slices, and slices. These ideas can be applied alone or in various combinations to any video codec standard or non-standard video codec that supports multi-layer video coding, such as the Versatile Video Codec (VVC) under development.

[0072] 2. Abbreviation

[0073] APS Adaptive Parameter Set

[0074] AU Access Unit

[0075] AUD Access Unit Delimiter

[0076] AVC Advanced Video Codec

[0077] CLVS codec layer video sequence

[0078] CPB Codec Picture Buffer

[0079] CRA Clean Random Access

[0080] CTU Codec Tree Unit

[0081] CVS codec video sequence

[0082] DPB decoded picture buffer

[0083] DPS decoding parameter set

[0084] EOB End of bitstream

[0085] EOS sequence end

[0086] GDR Gradual Decode Refresh

[0087] HEVC High-Efficiency Video Codec

[0088] HRD Hypothesized Reference Decoder

[0089] IDR Instant Decode Refresh

[0090] JEM Joint Exploration Model

[0091] MCTS motion-constrained patches

[0092] NAL Network Abstraction Layer

[0093] OLS output layer set

[0094] PH Image Header

[0095] PPS Picture Parameter Set

[0096] PTL grades, tiers, and levels

[0097] PU picture unit

[0098] RBSP Raw Byte Sequence Payload

[0099] SEI Supplemental Enhancement Information

[0100] SPS sequence parameter set

[0101] SVC Scalable Video Codec

[0102] VCL video codec layer

[0103] VPS Video Parameter Set

[0104] VTM VVC Test Model

[0105] VUI Video Availability Information

[0106] VVC multifunctional video codec

[0107] 3. Preliminary Discussion

[0108] Video codec standards have primarily evolved through the development of the renowned ITU-T and ISO / IEC standards. ITU-T produced H.261 and H.263, while ISO / IEC produced MPEG-1 and MPEG-4 Visual. The two organizations jointly produced the H.262 / MPEG-2 Video standard, the H.264 / MPEG-4 Advanced Video Coding (AVC) standard, and the H.265 / HEVC standard. Since H.262, video codec standards have been based on a hybrid video codec architecture that utilizes temporal prediction plus transform coding. To explore future video codec technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Team (JVET) in 2015. Since then, many new approaches have been adopted by JVET and incorporated into reference software called the Joint Exploration Model (JEM). JVET meetings are held quarterly, and the goal for new codec standards is to reduce bitrates by 50% compared to HEVC. At the JVET meeting in April 2018, the new video codec standard was officially named Versatile Video Coding (VVC), and the first version of the VVC Test Model (VTM) was released. As efforts continue to advance VVC standardization, new codec technologies are adopted for the VCC standard at each JVET meeting. The VVC working draft and test model (VTM) are then updated after each meeting. The VVC project's current goal is to achieve technical completion (FDIS) at the July 2020 meeting.

[0109] 3.1. Image Segmentation Scheme in HEVC

[0110] HEVC includes four different picture partitioning schemes, namely normal slice, dependent slice, slice and Wavefront Parallel Processing (WPP), which can be used for Maximum Transfer Unit (MTU) size matching, parallel processing and reduced end-to-end delay.

[0111] Regular slices are similar to those in H.264 / AVC. Each regular slice is encapsulated in its own NAL unit, and intra-picture prediction (intra sample prediction, motion information prediction, codec mode prediction) and entropy codec dependencies are disabled across slice boundaries. Therefore, regular slices can be reconstructed independently of other regular slices in the same picture (although there may still be interdependencies due to loop filtering operations).

[0112] Regular slices are the only tool available for parallelization, and are also available in H.264 / AVC in a nearly identical form. Regular slice-based parallelization does not require much inter-processor or inter-core communication (except for inter-processor or inter-core data sharing for motion compensation when decoding predictive codec pictures, which is generally much more significant due to intra-picture prediction). However, for the same reasons, the use of regular slices incurs significant codec overhead due to the bit cost of the slice header and the lack of prediction across slice boundaries. Furthermore, due to the intra-picture independence of regular slices and the fact that each regular slice is encapsulated in its own NAL unit, regular slices (compared to the other tools mentioned below) also serve as a key mechanism for bitstream segmentation to match MTU size requirements. In many cases, the goals of parallelization and MTU size matching place conflicting demands on the layout of slices within a picture. Recognition of this situation led to the development of the parallelization tools mentioned below.

[0113] Dependent slices have short headers and allow the bitstream to be split at treeblock boundaries without breaking any intra-picture prediction. Basically, dependent slices provide for partitioning a regular slice into multiple NAL units to provide reduced end-to-end latency by allowing part of a regular slice to be sent before coding of the entire regular slice is complete.

[0114] In WPP, a picture is partitioned into individual codec tree block (CTB) rows. Entropy decoding and prediction are allowed to use data from CTBs in other partitions. Parallel processing is possible through parallel decoding of CTB rows, where the start of decoding is delayed by two CTBs, ensuring that data related to the CTBs above and to the right of the target CTB is available before the target CTB is decoded. This staggered start (which, when represented graphically, looks like a wavefront) allows parallelization to utilize as many processors / cores as the picture contains CTB rows. Because intra-picture prediction is allowed between adjacent tree block rows within a picture, the inter-processor / inter-core communication required to implement intra-picture prediction can be substantial. WPP partitioning does not result in the generation of additional NAL units compared to when it is not used, so WPP is not a tool for MTU size matching. However, if MTU size matching is required, regular slices can be used with WPP, but with some codec overhead.

[0115] Slices define the horizontal and vertical boundaries that divide an image into slice columns and slice rows. Slice columns extend from the top to the bottom of the image. Similarly, slice rows extend from the left to the right of the image. The number of slices in an image can be simply derived by multiplying the number of slice columns by the number of slice rows.

[0116] Before decoding the top left CTB of the next slice in the order of the slice raster scan of the picture, the scan order of the CTBs is changed to be local within the slice (in the order of the slice's CTB raster scan). Similar to regular slices, slices break intra-picture prediction dependencies and entropy decoding dependencies. However, they do not need to be included in a single NAL unit (the same as WPP in this respect); therefore, slices cannot be used for MTU size matching. Each slice can be processed by one processor / core, and the inter-processor / inter-core communication required for intra-picture prediction between processing units decoding adjacent slices is limited to transmitting a shared slice header when a slice spans more than one slice, and sharing of reconstruction samples and metadata related to loop filtering. When more than one slice or WPP segment is included in a slice, the entry point byte offset of each slice or WPP segment in the slice except the first slice is signaled in the slice header.

[0117] For simplicity, HEVC has specified restrictions on the application of four different picture partitioning schemes. A given codec video sequence cannot include both slices and wavefronts from most profiles specified in the HEVC standard. For each slice and slice, one or both of the following conditions must be met: 1) all codec treeblocks in a slice belong to the same slice; 2) all codec treeblocks in a slice belong to the same slice. Finally, a wavefront segment contains exactly one CTB row. When using WPP, if a slice starts on a CTB row, it must end on the same CTB row.

[0118] The latest revision to HEVC is specified in the JCT-VC output document JCTVC-AC1005, J. Boyce, A. Ramasubramonian, R. Skupin, G.J. Sullivan, A. Tourapis, Y.-K. Wang (eds.), "HEVC Additional Supplemental Enhancement Information (Draft 4)", publicly released on October 24, 2017: http: / / phenix.int-evry.fr / jct / doc_end_user / documents / 29_Macau / wg11 / JCTVC-AC1005-v2.zip. With the inclusion of this revision, HEVC specifies three types of MCTS-related SEI (Supplemental Enhancement Information) messages: the time-domain MCTS SEI message, the MCTS Extraction Information Set SEI message, and the MCTS Extraction Information Nesting SEI message.

[0119] The temporal MCTS SEI message indicates the presence of MCTS in the bitstream and signals the MCTS. For each MCTS, motion vectors are restricted to pointing to full sample positions within the MCTS and fractional sample positions that only require full sample positions within the MCTS for interpolation, and motion vector candidates derived from blocks outside the MCTS for temporal motion vector prediction are not allowed. In this way, each MCTS can be decoded independently without the presence of slices not included in the MCTS.

[0120] The MCTS extraction information set SEI message provides supplementary information (defined as part of the semantics of the SEI message) that can be used in MCTS sub-bitstream extraction to generate a bitstream that conforms to the MCTS set. The information consists of multiple extraction information sets, each of which defines multiple MCTS sets and contains RBSP bytes that replace the VPS, SPS, and PPS to be used in the MCTS sub-bitstream extraction process. When extracting a sub-bitstream according to the MCTS sub-bitstream extraction process, the parameter sets (VPS, SPS, and PPS) need to be rewritten or replaced because one or all slice address-related syntax elements (including first_slice_segment_in_pic_flag and slice_segment_address) usually need to have different values, so the slice header needs to be slightly updated.

[0121] 3.2. Image Segmentation in VVC

[0122] In VVC, a picture is divided into one or more slice rows and one or more slice columns. A slice is a sequence of CTUs covering a rectangular area of ​​a picture. The CTUs in a slice are scanned in raster scan order within the slice.

[0123] A slice consists of an integer number of complete slices or an integer number of consecutive complete CTU rows within a slice of a picture.

[0124] Two striping modes are supported: raster scan striping mode and rectangular striping mode. In raster scan striping mode, a strip contains a sequence of complete slices in a slice raster scan of a picture. In rectangular striping mode, a strip contains multiple complete slices that together form a rectangular area of ​​the picture, or multiple consecutive complete CTU rows that together form a slice of a rectangular area of ​​the picture. The slices within a rectangular stripe are scanned in slice raster scan order within the rectangular area corresponding to the stripe.

[0125] A sub-picture consists of one or more strips that together cover a rectangular area of ​​the picture.

[0126] Figure 1 An example of raster scan striping partitioning of a picture is shown, where the picture is divided into 12 slices and 3 raster scan strips.

[0127] Figure 2 An example of rectangular strip partitioning of a picture is shown, where the picture is divided into 24 slices (6 slice columns and 4 slice rows) and 9 rectangular strips.

[0128] Figure 3 An example of a picture being divided into slices and rectangular strips is shown, where the picture is divided into 4 slices (2 slice columns and 2 slice rows) and 4 rectangular strips.

[0129] Figure 4 An example of sub-picture partitioning of a picture is shown, where the picture is divided into 18 slices, with the 12 slices on the left hand side each covering a 4×4 CTU strip, and the 6 slices on the right hand side each covering two vertically stacked strips of 2×2 CTUs, resulting in a total of 24 slices and 24 sub-pictures of different dimensions (each slice is a sub-picture).

[0130] 3.3. Signaling of SPS / PPS / Picture Header / Slice Header in VVC

[0131] 7.3.2.3 Sequence Parameter Set RBSP Syntax

[0132]

[0133]

[0134]

[0135]

[0136]

[0137]

[0138]

[0139] 7.3.2.4 Picture Parameter Set RBSP Syntax

[0140]

[0141]

[0142]

[0143]

[0144] 7.3.2.7 Picture header structure syntax

[0145]

[0146]

[0147]

[0148]

[0149]

[0150] 7.3.7.1 General Strip Header Syntax

[0151]

[0152]

[0153]

[0154]

[0155] 3.4. Example Specifications for Slices, Strips, and Sub-Pictures

[0156] 3 Definition

[0157] Picture level slice index: When rect_slice_flag is equal to 1, the slice index of the slice list in the picture (in the order they are signaled in the PPS).

[0158] Sub-picture level slice index: When rect_slice_flag is equal to 1, the slice index of the slice list in the sub-picture (in the order they are signaled in the PPS).

[0159] 6.5.1 CTB Raster Scanning, Slice Scanning, and Sub-Picture Scanning Process

[0160] The variable NumTileColumns that specifies the number of tile columns and the list colWidth[i] that specifies the width of the i-th tile column in CTB units (where i ranges from 0 to NumTileColumn-1, inclusive) are derived as follows:

[0161]

[0162]

[0163] The variable numtierrows that specifies the number of tile rows and the list RowHeight[j] that specifies the height of the j-th tile row in CTB units (where j ranges from 0 to NumTileRows-1, inclusive) are derived as follows:

[0164]

[0165] The variable NumTilesInPic is set equal to NumTileColumns*NumTileRows.

[0166] The list tileColBd[i] that specifies the position of the i-th tile column boundary in CTB units (where i ranges from 0 to NumTileColumns, inclusive) is derived as follows:

[0167] for(tileColBd[0]=0,i=0;i <NumTileColumns;i++)

[0168] tileColBd[i+1]=tileColBd[i]+colWidth[i] (25)

[0169] NOTE 1 – The size of the array tileColBd[] is one greater than the actual number of tile columns in the derivation of CtbToTileColBd[].

[0170] The list tileRowBd[j] that specifies the position of the jth tile row boundary in CTB units (where j ranges from 0 to NumTileRows, inclusive) is derived as follows:

[0171] for(tileRowBd[0]=0,j=0;j <NumTileRows;j++)

[0172] tileRowBd[j+1]=tileRowBd[j]+RowHeight[j] (26)

[0173] NOTE 2 – The size of the array tileRowBd[] in the above derivation is one larger than the actual number of tile rows in the derivation of CtbToTileRowBd[].

[0174] The list CtbToTileColBd[ctbAddrX] that specifies the conversion from horizontal CTB addresses to left tile column boundaries in units of CTBs (where ctbAddrX ranges from 0 to PicWidthInCtbsY, inclusive) is derived as follows:

[0175]

[0176] NOTE 3 – The size of the array CtbToTileColBd[] in the above derivation is one greater than the actual number of picture widths in the CTB signaled in the derived slice_data().

[0177] The list CtbToTileRowBd[ctbAddrY] that specifies the conversion from the vertical CTB address to the top tile column boundary in units of CTB (where ctbAddrY ranges from 0 to PicHeightInCtbsY, inclusive) is derived as follows:

[0178]

[0179] NOTE 4 – The size of the array CtbToTileRowBd[] in the above derivation is one greater than the actual number of picture heights in the CTB signaled in slice_data().

[0180] For rectangular slices, the list NumCtusInSlice[i] specifying the number of CTUs in the i-th slice (where i ranges from 0 to num_slices_in_pic_minus1, inclusive), the list SliceTopLeftTileIdx[i] specifying the index of the top left slice of the slice (where i ranges from 0 to num_slices_in_pic_minus1, inclusive), and the matrix CtbAddrInSlice[i][j] specifying the picture raster scan address of the j-th CTB within the i-th slice (where i ranges from 0 to num_slices_in_pic_minus1, inclusive) are derived as follows:

[0181]

[0182]

[0183]

[0184] Among them, the function AddCtbsToSlice(sliceIdx, startX, stopX, startY, stopY) is defined as follows:

[0185]

[0186] It is a requirement of bitstream conformance that the value of NumCtusInSlice[i], where i ranges from 0 to num_slices_in_pic_minus1, inclusive, shall be greater than 0. Furthermore, it is a requirement of bitstream conformance that the matrix CtbAddrInSlice[i][j], where i ranges from 0 to num_slices_in_pic_minus1, inclusive, and j ranges from 0 to NumCtusInSlice[i]-1, inclusive, shall include all CTB addresses in the range from 0 to PicSizeInCtbsY-1 once and only once.

[0187] The list CtbToSubpicIdx[ctbAddrRs] that specifies the conversion from CTB addresses in the picture raster scan to sub-picture indices (where tbAddrRs ranges from 0 to PicSizeInCtbsY-1, inclusive) is derived as follows:

[0188]

[0189]

[0190] The list NumSlicesInSubpic[i] that specifies the number of rectangular slices in the i-th sub-picture is derived as follows:

[0191]

[0192] 7.3.4.3 Picture Parameter Set RBSP Semantics

[0193] subpic_id_mapping_in_pps_flag equal to 1 specifies that sub-picture ID mapping is signaled in the PPS. subpic_id_mapping_in_pps_flag equal to 0 specifies that sub-picture ID mapping is not signaled in the PPS. If subpic_id_mapping_explicitly_signaled_flag is 0 or subpic_id_mapping_in_sps_flag is 1, then the value of subpic_id_mapping_in_pps_flag shall be 0. Otherwise (subpic_id_mapping_explicitly_signaled_flag is 1 and subpic_id_mapping_in_sps_flag is 0), the value of subpic_id_mapping_in_pps_flag shall be 1.

[0194] pps_num_subpics_minus1 shall be equal to sps_num_subpics_minus1.

[0195] pps_subpic_id_len_minus1 shall be equal to sps_subpic_id_len_minus1.

[0196] pps_subpic_id[i] specifies the sub-picture ID of the i-th sub-picture. The length of the pps_subpic_id[i] syntax element is pps_subpic_id_len_minus1+1 bits.

[0197] For each value of i in the range 0 to sps_num_subpics_minus1 (inclusive), the variable SubpicIdVal[i] is derived as follows:

[0198]

[0199] A requirement for bitstream conformance is that the following two constraints apply:

[0200] – For any two different values ​​of i and j in the range 0 to sps_num_subpics_minus1 (inclusive), SubpicIdVal[i] shall not be equal to SubpicIdVal[j].

[0201] – When the current picture is not the first picture of the CLVS, for each value of i in the range 0 to sps_num_subpics_minus1 (inclusive), if the value of SubpicIdVal[i] is not equal to the value of SubpicIdVal[i] of the previous picture in decoding order in the same layer, the nal_unit_type of all codec slice NAL units of the sub-picture in the current picture with sub-picture index i shall be equal to the specified value in the range IDR_W_RADL to CRA_NUT (inclusive).

[0202] no_pic_partition_flag equal to 1 specifies that no picture partitioning is applied to each picture of the referenced PPS. no_pic_partition_flag equal to 0 specifies that each picture of the referenced PPS may be partitioned into more than one slice or slice.

[0203] One requirement for bitstream conformance is that the value of no_pic_partition_flag shall be the same for all PPSs referenced by a codec picture within a CLVS.

[0204] It is a bitstream conformance requirement that when the value of sps_num_subpics_minus1+1 is greater than 1, the value of no_pic_partition_flag shall not be equal to 1.

[0205] pps_log2_ctu_size_minus5 plus 5 specifies the luma codec treeblock size for each CTU. pps_log2_ctu_size_minus5 shall be equal to sps_log2_ctu_size_minus5.

[0206] num_exp_tile_columns_minus1 plus 1 specifies the number of explicitly provided tile column widths. The value of num_exp_tile_columns_minus1 shall be in the range of 0 to PicWidthInCtbsY-1, inclusive. When no_pic_partition_flag is equal to 1, the value of num_exp_tile_columns_minus1 is inferred to be equal to 0.

[0207] num_exp_tile_rows_minus1 plus 1 specifies the number of explicitly provided tile row heights. The value of num_exp_tile_rows_minus1 shall be in the range of 0 to PicHeightInCtbsY-1, inclusive. When no_pic_partition_flag is equal to 1, the value of num_tile_rows_minus1 is inferred to be equal to 0.

[0208] tile_column_width_minus1[i] plus 1 specifies the width of the i-th tile column in CTBs, where i ranges from 0 to num_exp_tile_columns_minus1-1 (inclusive). tile_column_width_minus1[num_exp_tile_columns_minus1] is used to derive the width of tile columns with indices greater than or equal to num_exp_tile_columns_minus1 as specified in clause 6.5.1. The value of tile_column_width_minus1[i] shall be in the range of 0 to PicWidthInCtbsY-1 (inclusive). When not present, the value of tile_column_width_minus1[0] is inferred to be equal to PicWidthInCtbsY-1.

[0209] tile_row_height_minus1[i] plus 1 specifies the height of the i-th slice row in CTBs, where i ranges from 0 to num_exp_tile_rows_minus1-1 (inclusive). tile_row_height_minus1[num_exp_tile_rows_minus1] is used to derive the height of slice rows with indices greater than or equal to num_exp_tile_rows_minus1 as specified in clause 6.5.1. The value of tile_row_height_minus1[i] shall be in the range of 0 to PicHeightInCtbsY-1 (inclusive). When not present, the value of tile_row_height_minus1[0] is inferred to be equal to PicHeightInCtbsY-1.

[0210] rect_slice_flag equal to 0 specifies that the slices within each slice are in raster scan order and that slice information is not signaled in the PPS. rect_slice_flag equal to 1 specifies that the slices within each slice cover a rectangular area of ​​the picture and that slice information is signaled in the PPS. When not present, rect_slice_flag is inferred to be equal to 1. When subpic_info_present_flag is equal to 1, the value of rect_slice_flag shall be equal to 1.

[0211] single_slice_per_subpic_flag equal to 1 specifies that each sub-picture consists of one and only one rectangular slice. single_slice_per_subpic_flag equal to 0 specifies that each sub-picture may consist of one or more rectangular slices. When single_slice_per_subpic_flag is equal to 1, num_slices_in_pic_minus1 is inferred to be equal to sps_num_subpics_minus1. When not present, the value of single_slice_per_subpic_flag is inferred to be equal to 0.

[0212] num_slices_in_pic_minus1 plus 1 specifies the number of rectangular slices in each picture referencing the PPS. The value of num_slices_in_pic_minus1 shall be in the range of 0 to MaxSlicesPerPicture1, inclusive, where MaxSlicesPerPicture is specified in Annex A. When no_pic_partition_flag is equal to 1, the value of num_slices_in_pic_minus1 is inferred to be equal to 0.

[0213] tile_idx_delta_present_flag equal to 0 specifies that tile_idx_delta values ​​are not present in the PPS, and all rectangular slices in pictures referencing the PPS are specified in raster order according to the process defined in clause 6.5.1. tile_idx_delta_present_flag equal to 1 specifies that tile_idx_delta values ​​may be present in the PPS, and all rectangular slices in pictures referencing the PPS are specified in the order indicated by the tile_idx_delta values. When not present, the value of tile_idx_delta_present_flag is inferred to be equal to 0.

[0214] slice_width_in_tiles_minus1[i] plus 1 specifies the width of the i-th rectangular strip in tile columns. The value of slice_width_in_tiles_minus1[i] should be in the range of 0 to NumTileColumns-1 (inclusive).

[0215] When slice_width_in_tiles_minus1[i] is not present, the following applies:

[0216] – If NumTileColumns is equal to 1, the value of slice_width_in_tiles_minus1[i] is inferred to be equal to 0.

[0217] – Otherwise, the value of slice_width_in_tiles_minus1[i] is inferred as specified in clause 6.5.1.

[0218] slice_height_in_tiles_minus1[i] plus 1 specifies the height of the i-th rectangular slice in tile rows. The value of slice_height_in_tiles_minus1[i] should be in the range of 0 to NumTileRows-1 (inclusive).

[0219] When slice_height_in_tiles_minus1[i] is not present, the following applies:

[0220] – If NumTileRows is equal to 1, or tile_idx_delta_present_flag is equal to 0, and tileIdx % NumTileColumns is greater than 0, then the value of slice_height_in_tiles_minus1[i] is inferred to be equal to 0.

[0221] – Otherwise (NumTileRows is not equal to 1, and tile_idx_delta_present_flag is equal to 1 or tileIdx % NumTileColumns is equal to 0), when tile_idx_delta_present_flag is equal to 1 or tileIdx % NumTileColumns is equal to 0, the value of slice_height_in_tiles_minus1[i] is inferred to be equal to slice_height_in_tiles_minus1[i-1].

[0222] num_exp_slices_in_tile[i] specifies the number of slice heights explicitly provided in the current slice that contains more than one rectangular strip. The value of num_exp_slices_in_tile[i] shall be in the range of 0 to RowHeight[tileY]-1, inclusive, where tileY is the index of the slice row containing the i-th strip. When not present, the value of num_exp_slices_in_tile[i] is inferred to be equal to 0. When num_exp_slices_in_tile[i] is equal to 0, the value of the variable NumSlicesInTile[i] is inferred to be equal to 1.

[0223] exp_slice_height_in_ctus_minus1[j] plus 1 specifies the height of the j-th rectangular slice in the current slice in CTU rows. The value of exp_slice_height_in_ctus_minus1[j] shall be in the range of 0 to RowHeight[tileY]-1 (inclusive), where tileY is the slice row index of the current slice.

[0224] When num_exp_slices_in_tile[i] is greater than 0, the variables NumSlicesInTile[i] and SliceHeightInCtusMinus1[i+k] with k in the range of 0 to NumSlicesInTile[i]-1 are derived as follows:

[0225]

[0226] tile_idx_delta[i] specifies the difference between the tile index of the first tile in the i-th rectangular stripe and the tile index of the first tile in the i+1-th rectangular stripe. The value of tile_idx_delta[i] shall be in the range -NumTilesInPic+1 to NumTilesInPic-1, inclusive. When not present, the value of tile_idx_delta[i] is inferred to be equal to 0. When present, the value of tile_idx_delta[i] shall not be equal to 0.

[0227]

[0228] 7.4.2.4.5 Order of VCL NAL units and their association with coded and decoded pictures

[0229] The order of VCL NAL units within a codec picture is constrained as follows:

[0230] – For any two codec slice NAL units A and B of a codec picture, let subpicIdxA and subpicIdxB be their sub-picture level index values, and sliceAddrA and sliceddrB be their slice_address values.

[0231] – Codec slice NAL unit A shall precede codec slice NAL unit B when any of the following conditions is true:

[0232] –subpicIdxA is less than subpicIdxB.

[0233] –subpicIdxA is equal to subpicIdxB, and sliceAddrA is less than sliceAddrB.

[0234] 7.4.8.1 Common Strip Header Semantics

[0235] The variable CuQpDeltaVal, which specifies the difference between the luma quantization parameter and its prediction for a codec containing cu_qp_delta_abs, is set equal to 0. The variable CuQpDeltaVal specifies the difference between the luma quantization parameter and its prediction for a codec containing cu_chroma_qp_offset_flag. Cr and Qp′ CbCr The value of the variable CuQpOffset to use when the corresponding value of the quantization parameter Cb 、CuQpOffset Cr and CuQpOffset CbCr are all set equal to 0.

[0236] picture_header_in_slice_header_flag equal to 1 specifies that the PH syntax structure is present in the slice header.

[0237] picture_header_in_slice_header_flag equal to 0 specifies that the PH syntax structure is not present in the slice header.

[0238] One requirement for bitstream conformance is that the value of picture_header_in_slice_header_flag shall be the same in all codec slices of CLVS.

[0239] When picture_header_in_slice_header_flag is equal to 1 for a coded slice, one requirement for bitstream conformance is that VCL NAL units with nal_unit_type equal to PH_NUT shall not appear in the CLVS.

[0240] When picture_header_in_slice_header_flag is equal to 0, all coded slices in the current picture shall have picture_header_in_slice_header_flag equal to 0, and the current PU shall have a PH NAL unit.

[0241] slice_subpic_id specifies the sub-picture ID of the sub-picture containing the slice. If slice_subpic_id exists, the value of the variable CurrSubpicIdx is derived such that SubpicIdVal[CurrSubpicIdx] is equal to slice_subpic_id. Otherwise (slice_subpic_id does not exist), CurrSubpicIdx is derived to be 0. The length of slice_subpic_id is sps_subpic_id_len_minus1+1 bits.

[0242] slice_address specifies the slice address of the slice. When not present, the value of slice_address is inferred to be equal to 0. When rect_slice_flag is equal to 1 and NumSlicesInSubpic[CurrSubpicIdx] is equal to 1, the value of slice_address is inferred to be equal to 0.

[0243] If rect_slice_flag is equal to 0, the following applies:

[0244] - The strip address is the raster scan slice index.

[0245] - The length of slice_address is Ceil(Log2(NumTilesInPic)) bits.

[0246] -slice_address values ​​should be in the range of 0 to NumTilesInPic1-1 (inclusive).

[0247] Otherwise (rect_slice_flag is equal to 1), the following applies:

[0248] - The slice address is the sub-picture level slice index of the slice.

[0249] - The length of slice_address is Ceil(Log2(NumSlicesInSubpic[CurrSubpicIdx])) bits.

[0250] -slice_address values ​​should be in the range of 0 to NumSlicesInSubpic[CurrSubpicIdx]-1 (inclusive).

[0251] A requirement for bitstream conformance is that the following constraints apply:

[0252] - If rect_slice_flag is equal to 0 or subpic_info_present_flag is equal to 0, the value of slice_address shall not be equal to the value of slice_address of any other codec slice NAL unit of the same codec picture.

[0253] Otherwise, the pair of slice_subpic_id value and slice_address value shall not be equal to the pair of slice_subpic_id value and slice_address value of any other codec slice NAL unit of the same codec picture.

[0254] - The shape of the slices of a picture shall be such that, when decoded, the entire left boundary and the entire top boundary of each CTU shall consist of the picture boundary or of the boundary of the previously decoded CTU(s).

[0255] sh_extra_bit[i] may be equal to 1 or 0. Decoders conforming to this version of this specification shall ignore the value of sh_extra_bit[i]. Its value shall not affect the conformance of the decoder to the profile specified in this version of the specification.

[0256] num_tiles_in_slice_minus1 plus 1 (when present) specifies the number of tiles in a slice. The value of num_tiles_in_slice_minus1 should be in the range of 0 to NumTilesInPic-1 (inclusive).

[0257] The variable NumCtusInCurrSlice specifies the number of CTUs in the current slice, and the list of picture raster scan addresses CtbAddrInCurrSlice[i] (where i ranges from 0 to NumCtusInCurrSlice-1 (inclusive)) for the i-th CTB in the slice is derived as follows:

[0258]

[0259]

[0260] 3.5. Color Space and Chroma Subsampling

[0261] A color space, also called a color model (or color system), is an abstract mathematical model that simply describes the range of colors as a tuple of numbers, usually 3 or 4 values ​​or color components (such as RGB). Fundamentally, a color space is a elaboration of coordinate systems and subspaces.

[0262] For video compression, the most commonly used color spaces are YCbCr and RGB.

[0263] YCbCr, Y'CbCr, or Y Pb / Cb Pr / Cr (also written as YCBCR or Y'CBCR) is a family of colorspaces used as part of the color image pipeline in video and digital photography systems. Y' is the luma component, and CB and CR are the blue-difference and red-difference chroma components. Y' (with a prime) is distinguished from Y (Y is luma), which means that the light intensity is nonlinearly encoded based on the gamma-corrected RGB primaries.

[0264] Chroma subsampling is the practice of encoding an image at a lower resolution for chroma information than for luminance information by exploiting the fact that the human visual system is less sensitive to color differences than to luminance. 3.5.1. 4:4:4

[0266] Each of the three Y'CbCr components has the same sampling rate, so there is no chroma subsampling. This scheme is sometimes used in high-end film scanners and film post-production. 3.5.2. 4:2:2

[0268] The two chroma components are sampled at half the luma sampling rate: the horizontal chroma resolution is halved, while the vertical chroma resolution remains unchanged. This reduces the bandwidth of the uncompressed video signal by one third, but there is almost no visual difference. Figure 5 Examples of nominal vertical position and nominal horizontal position for a 4:2:2 color format are depicted in . 3.5.3. 4:2:0

[0270] Compared to 4:1:1, horizontal sampling in 4:2:0 is doubled, but vertical resolution is halved because the Cb and Cr channels are sampled only on alternate lines in this scheme. Therefore, the data rate remains the same. Both Cb and Cr are subsampled by a factor of 2 both horizontally and vertically. There are three variants of the 4:2:0 scheme, with different horizontal and vertical positions.

[0271] In MPEG-2, Cb and Cr coexist horizontally. Cb and Cr are located between pixels in the vertical direction (in the gap).

[0272] In JPEG / JFIF, H.261, and MPEG-1, Cb and Cr are located midway between alternate luma samples.

[0273] In 4:2:0 DV, Cb and Cr are co-sited horizontally and vertically on alternate lines.

[0274] Table 3-1. SubWidthC and SubHeightC values ​​derived from chroma_format_idc and separate_colour_plane_flag

[0275]

[0276] 4. Examples of technical problems solved by the disclosed embodiments

[0277] The existing design of SPS / PPS / picture header / slice header signaling in VVC has the following problems:

[0278] 1) According to the current VVC text, when rect_slice_flag is equal to 1, the following applies:

[0279] a.slice_address represents the sub-picture level slice index of the slice.

[0280] b. The sub-picture level slice index is defined as the slice index of the slice list in the sub-picture (in the order they are signaled in the PPS).

[0281] c. Picture level slice index is defined as the slice index of the slice list in the picture (in the order they are signaled in the PPS).

[0282] d. For any two slices belonging to two different sub-pictures, the slice associated with the smaller sub-picture index comes earlier in the decoding order, while for any two slices belonging to the same sub-picture, the slice with the smaller sub-picture level slice index comes earlier in the decoding order.

[0283] e. Assuming that the increasing order of the picture-level slice index values ​​is the same as the decoding order of the slices, the variable NumCtusInCurrSlice that specifies the number of CTUs in the current slice is derived by equation 117 in the current VVC text.

[0284] However, when some stripes are generated by dividing a slice, some of the above aspects may be violated. Figure 6 In the example shown in , when a picture is divided into two slices by a vertical slice boundary, and each of the two slices is divided into two strips by the same horizontal boundary across the entire picture, the upper two strips are included in the first sub-picture, and the lower two strips are included in the second sub-picture. In this case, according to the current VVC text, the picture-level slice index values ​​of the four slices in slice raster scan order will be 0, 2, 1, 3, while the decoding order index values ​​of the four slices in slice raster scan order will be 0, 1, 2, 3. Therefore, the derivation of NumCtusInCurrSlice will be incorrect, and in turn the parsing of the slice data will be problematic, the decoded sample values ​​will be incorrect, and the decoder may crash.

[0285] 2) There are two types of slice signaling methods. In rectangular mode, all slice partitioning information is signaled in the PPS. In non-rectangular mode, partial slice partitioning information is signaled in the slice header. Therefore, in this mode, the complete slice division of the picture cannot be known until all slices of the picture are parsed.

[0286] 3) In rectangular mode, it is possible to arbitrarily signal the slices by setting tile_idx_delta. A bad bitstream may crash the decoder due to this mechanism.

[0287] 4) In some embodiments, when i is equal to num_slices_in_pic_minus1, tile_idx_delta[i] is not initialized.

[0288] Figure 6 An example of picture partitioning is shown. Solid lines 602 represent slice boundaries; dashed lines 604 represent stripe boundaries, and dashed lines 606 represent sub-picture boundaries. The figure shows the picture-level index, decoding order index, sub-picture-level index, and sub-picture and slice indexes for four slices.

[0289] 5. Example Embodiments and Techniques

[0290] To address the above and other issues, the following methods are disclosed. The present invention should be viewed as an example of a general concept and should not be interpreted in a narrow sense. Furthermore, these inventions may be applied individually or in any combination.

[0291] 1. For slices in rectangular slice mode (ie, when rect_slice_flag is equal to 1), the picture-level slice index for each slice in each sub-picture is derived, and the derived value is used to derive the number of CTUs in each slice.

[0292] 2. Sub-picture level strip index can be defined / exported in the following ways:

[0293] a. In one example, the sub-picture level slice index is defined as "the slice index of the list of slices in the sub-picture (in their decoding order) when rect_slice_flag is equal to 1".

[0294] b. Optionally, the sub-picture level slice index is defined as "the slice index of the slice list in the sub-picture when rect_slice_flag is equal to 1, as specified by the variable SubpicLevelSliceIdx[i] derived in Equation 32 (as in Embodiment 1), where i is the picture level slice index of the slice."

[0295] c. In one example, a sub-picture index is derived for each slice with a specific value of the picture-level slice index.

[0296] d. In one example, a sub-picture level slice index is derived for each slice having a specific value of the picture level slice index.

[0297] e. In one example, when rect_slice_flag is equal to 1, the semantics of the slice address is specified as "the slice address is the sub-picture level slice index of the slice specified by the variable SubpicLevelSliceIdx[i] derived in Equation 32 (e.g., as in Embodiment 1), where i is the picture level slice index of the slice."

[0298] 3. The sub-picture level slice index of the slice is assigned to the slice in the first sub-picture containing the slice. The sub-picture level slice index of each slice can be stored in an array indexed by the picture level slice index (e.g., SubpicLevelSliceIdx[i] in embodiment 1).

[0299] a. In one example, the sub-picture level slice index is a non-negative integer.

[0300] b. In one example, the value of the sub-picture level slice index of the slice is greater than or equal to 0.

[0301] c. In one example, the value of the sub-picture level slice index of a slice is less than N, where N is the number of slices in the sub-picture.

[0302] d. In one example, if a first slice (slice A) and a second slice (slice B) are in the same sub-picture but they are different, the first sub-picture level slice index of the first slice (denoted as subIdxA) must be different from the second sub-picture level slice index of the second slice (denoted as subIdxB).

[0303] e. In one example, if a first sub-picture-level slice index (denoted as subIdxA) of a first slice (slice A) in a first sub-picture is less than a second sub-picture-level slice index (denoted as subIdxB) of a second slice (slice B) in the same first sub-picture, then IdxA is less than IdxB, where idxA and idxB represent slice indexes (also referred to as picture-level slice indices, e.g., sliceIdx) of slice A and slice B, respectively, in the entire picture.

[0304] f. In one example, if the first sub-picture level slice index (denoted as subIdxA) of a first slice (slice A) in a first sub-picture is less than the second sub-picture level slice index (denoted as subIdxB) of a second slice (slice B) in the same first sub-picture, then slice A precedes slice B in decoding order.

[0305] g. In one example, the sub-picture level slice index in a sub-picture is derived based on the picture level slice index (eg, sliceIdx).

[0306] 4. It is proposed to derive a mapping function / mapping table between the sub-picture level strip index and the picture level strip index in the sub-picture.

[0307] a. In one example, a two-dimensional array PicLevelSliceIdx[subPicIdx][SubPicLevelSliceIdx] is derived to map sub-picture level slice indices in a sub-picture to picture level slice indices, where PicLevelSliceIdx represents the picture level slice index of the slice, subPicIdx represents the index of the sub-picture, and SubPicLevelSliceIdx represents the sub-picture level slice index of the slice in the sub-picture.

[0308] i. In one example, the array NumSlicesInSubpic[subPicIdx] is used to derive PicLevelSliceIdx, where NumSlicesInSubpic[subPicIdx] represents the number of slices in the subpicture with an index equal to subPicIdx.

[0309] 1) In one example, NumSlicesInSubpic[subPicIdx] and PicLevelSliceIdx[subPicIdx][SubPicLevelSliceIdx] are derived in a single process by scanning all slices in the order of picture-level slice indices.

[0310] a. Before processing, NumSlicesInSubpic[subPicIdx] is set equal to 0 for all valid subPicIdx.

[0311] b. When checking a slice with picture level index equal to S, if it is in a subpicture with subpicture index equal to P, set PicLevelSliceIdx[P][NumSlicesInSubpic[P]] equal to S, then set NumSlicesInSubpic[P] equal to NumSlicesInSubpic[P]+1.

[0312] ii. In one example, SliceIdxInPic[subPicIdx][SubPicLevelSliceIdx] is used to derive the picture level slice index (eg, picLevelSliceIdx), which is then used to derive the number and / or addresses of CTBs in the slice when parsing the slice header.

[0313] 5. The conforming bitstream requires that a slice cannot be located in multiple sub-pictures.

[0314] 6. The conformance bitstream requires that a sub-picture cannot include two slices (denoted as slice A and slice B), where slice A is in slice A but smaller than slice A, and slice B is in slice B but smaller than slice B, and slice A and slice B are different.

[0315] 7. It is proposed that the slice and / or slice partitioning information of a picture may be signaled in the associated picture header.

[0316] a. In one example, the slice and / or slice partitioning information of a picture is signaled in the PPS or in the associated picture header.

[0317] b. In one example, the slice and / or slice partitioning information of a picture is signaled in the picture header whether it is in the associated picture header.

[0318] i. In one example, if the slice and / or slice partitioning information of a picture is signaled in both the associated PPS and the associated picture header, the slice and / or slice partitioning information of the picture signaled in the picture header will be used.

[0319] ii. In one example, if the slice and / or slice partitioning information of a picture is signaled in both the associated PPS and the associated picture header, the slice and / or slice partitioning information of the picture signaled in the PPS will be used.

[0320] c. In one example, signaling is performed in a video unit at a level higher than a picture, such as in an SPS, to indicate whether the slice and / or slice partitioning information of a picture is signaled in an associated PPS or in an associated picture header.

[0321] 8. It is proposed that when an associated picture is divided into slices in a non-rectangular pattern, the slice partitioning information is signaled in a higher-level video unit (such as PPS and / or picture header) than the slice level.

[0322] a. In one example, when the associated picture is divided into slices in a non-rectangular mode, information indicating the number of slices (eg, num_slices_in_pic_minus1) may be signaled in a higher-level video unit.

[0323] b. In one example, when the associated picture is divided into slices in a non-rectangular pattern, information indicating the index (or address, or position, or coordinates) of the first block unit of the slice is signaled in a higher-level video unit. For example, the block unit may be a CTU or a slice.

[0324] c. In one example, when the associated picture is divided into slices in a non-rectangular pattern, information indicating the number of block units of the slice is signaled in a higher-level video unit. For example, the block unit may be a CTU or a slice.

[0325] d. In one example, when the associated picture is divided in slices in a non-rectangular pattern, the slice partitioning information (eg, num_tiles_in_slice_minus1) is not signaled in the slice header.

[0326] e. In one example, when the associated picture is partitioned into slices in a non-rectangular pattern, the slice index is signaled in the slice header.

[0327] i. In one example, when the associated picture is partitioned into slices in a non-rectangular pattern, slice_address is interpreted as a picture-level slice index.

[0328] f. In one example, when an associated picture is divided into slices in a non-rectangular pattern, partitioning information of each slice in the picture (such as the index of the first block unit and / or the number of block units) may be signaled sequentially in a higher-level video unit.

[0329] i. In one example, when the associated picture is divided into slices in a non-rectangular pattern, the index of the slice may be signaled for each slice in the higher-level video unit.

[0330] ii. In one example, the partitioning information of each stripe is signaled in ascending order of the stripe index.

[0331] 1) In one example, the partitioning information of each slice is signaled in the order of slice 0, slice 1, ..., slice K-1, slice K, slice K+1, ..., slice S-2, slice S-1, where K represents the slice index and S represents the number of slices in the picture.

[0332] iii. In one example, the split information of each stripe is signaled in descending order of stripe index.

[0333] 1) In one example, the partitioning information of each slice is signaled in the order of slice S-2, slice S-1, ..., slice K+1, slice K, slice K-1, ..., slice 1, and slice 0, where K represents the slice index and S represents the number of slices in the picture.

[0334] iv. In one example, when the associated picture is partitioned into slices in a non-rectangular pattern, the index of the first block unit of the slice may not be signaled in the higher-level video unit.

[0335] 1) For example, the index of the first block unit of stripe 0 (the stripe with stripe index equal to 0) is inferred to be 0.

[0336] 2) For example, the index of the first block unit of stripe K (the stripe with stripe index equal to K, K>0) is inferred to be Among them, N i Indicates the number of block units in stripe i.

[0337] v. In one example, when the associated picture is divided into slices in a non-rectangular pattern, the index of the first block unit of the slice may not be signaled in the higher-level video unit.

[0338] 1) For example, the index of the first block unit of stripe 0 (the stripe with stripe index equal to 0) is inferred to be 0.

[0339] 2) For example, the index of the first block unit of stripe K (the stripe with stripe index equal to K, K>0) is inferred to be Among them, N i Indicates the number of block units in stripe i.

[0340] vi. In one example, when the associated picture is divided into slices in a non-rectangular pattern, the number of block units of the slice may not be signaled in the higher-level video unit.

[0341] 1) When there is only one slice in a picture and there are M block units in the picture, the number of block units of slice 0 is M.

[0342] 2) For example, the number of block units of stripe K (the stripe with stripe index equal to 0) is inferred to be T K+1 -T K , where T K represents the index of the first block unit of stripe K when K < S - 1, where S is the number of stripes in the picture and S > 1.

[0343] 3) For example, the number of block units of stripe S - 1 is inferred to be M - where S is the number of stripes in the picture, S > 1, and M is the number of block units in the picture,

[0344] vii. In one example, when partitioning an associated picture into stripes in a non - rectangular mode, the segmentation information of one or more stripes may not be signaled in a higher - level video unit.

[0345] 1) In one example, the segmentation information of one or more stripes not signaled in a higher - level video unit can be inferred from the segmentation information of other stripes to be signaled.

[0346] 2) In one example, the segmentation information of the last C stripes may not be signaled. For example, C equals 1.

[0347] 3) For example, the number of block units of stripe S - 1 is not signaled, where S is the number of stripes in the picture and S > 1.

[0348] a. For example, the number of block units of stripe S - 1 is inferred to be where there are M block units in the picture.

[0349] 9. It is proposed that the minimum number of stripes in a picture can be different depending on whether rectangular segmentation or non - rectangular segmentation is applied.

[0350] a. In one example, if the non - rectangular segmentation mode is applied, the picture is partitioned into at least two stripes, while if the rectangular segmentation mode is applied, the picture is partitioned into at least one stripe.

[0351] i. For example, if the non - rectangular segmentation mode is applied, num_slices_in_pic_minus2 plus 2 that specifies the number of stripes in the picture can be signaled.

[0352] b. In one example, if the non - rectangular segmentation mode is applied, the picture is partitioned into at least one stripe, while if the rectangular segmentation mode is applied, the picture is partitioned into at least one stripe.

[0353] i. For example, if rectangular partitioning mode is applied, num_slices_in_pic_minus2 plus 2, which specifies the number of slices in a picture, may be signaled.

[0354] c. In one example, when a picture is not divided into sub-pictures or is divided into only one sub-picture, the minimum number of slices in the picture may be different depending on whether rectangular partitioning or non-rectangular partitioning is applied.

[0355] 10. It is proposed that when partitioning information is signaled in a video unit such as a PPS or a picture header, a slice is represented by the top left position and the width / height of the slice.

[0356] a. In one example, the index / position / coordinates of the top left block unit (such as a CTU or slice) of the slice are signaled, and / or the width measured in video units (such as a CTU or slice), and / or the height measured in video units (such as a CTU or slice).

[0357] b. In one example, the top left position and width / height information of each stripe is signaled sequentially.

[0358] i. For example, the information of the top left position and width / height of each slice is signaled in ascending order of slice index, such as 0, 1, 2, ..., S-1, where S is the number of slices in the picture.

[0359] 11. It is proposed to signal the partitioning information (such as position / width / height) of slices in sub-pictures in video units such as SPS / PPS / picture header.

[0360] a. In one example, the slice partitioning information of each sub-image is signaled sequentially.

[0361] i. For example, the slice partitioning information of each sub-image is signaled in ascending order of the sub-image index.

[0362] b. In one example, the partitioning information (such as position / width / height) of each slice in the sub-image is signaled sequentially.

[0363] i. In one example, the partitioning information (eg, position / width / height) of each slice in a sub-picture is signaled in ascending order of the sub-picture level slice index.

[0364] 12. It is proposed that the difference between the tile index of the first tile in the i-th rectangular slice and the tile index of the first tile in the i+1-th rectangular slice (denoted as tile_idx_delta[i]) is derived instead of signaling.

[0365] a. In one example, based on the rectangular strips from the 0th rectangular strip to the ith rectangular strip, a slice index of the first slice in the (i+1)th rectangular strip is derived.

[0366] b. In one example, the tile index of the first tile in the (i+1)th rectangular strip is derived as the minimum index of tiles that are not within the rectangular strips from the 0th rectangular strip to the i-th rectangular strip.

[0367] 13. It is proposed to signal information used to derive the number of slice columns / slice rows (eg, NumTileColumns or NumTileRows) based on the relationship between the width of the picture and the size of the CTU.

[0368] a. For example, if the width of the picture is less than or equal to the size or width of the CTU, num_exp_tile_columns_minus1 and / or tile_column_width_minus1 may not be signaled.

[0369] b. For example, if the height of the picture is less than or equal to the size or height of the CTU, num_exp_tile_rows_minus1 and / or tile_row_height_minus1 may not be signaled.

[0370] 14. When slice_subpic_id exists, there must be only one CurrSubpicIdx that satisfies SubpicIdVal[CurrSubpicIdx] equal to slice_subpic_id.

[0371] 15. If rect_slice_flag is equal to 0 or subpic_info_present_flag is equal to 0, the value of slice_address+i (where i is in the range of 0 to num_tiles_in_slice_minus1, inclusive) shall not be equal to the value of slice_address+j (where j is in the range of 0 to num_tiles_in_slice_minus1, inclusive) of any other coded slice NAL unit of the same coded picture, where i is in that range.

[0372] 16. In the case where there are both uniformly spaced and non-uniformly spaced slices in a picture, a syntax element may be signaled in the PPS (or SPS) to specify the type of slice layout.

[0373] a. In one example, a syntax flag may be signaled in the PPS to specify whether the slice layout is non-uniform spacing followed by uniform spacing, or uniform spacing followed by non-uniform spacing.

[0374] b. For example, whenever there are non-uniformly spaced tiles, the number of explicitly provided tile columns / rows (e.g., num_exp_tile_columns_minus1, num_exp_tile_rows_minus1) may be no less than the total number of non-uniform tiles.

[0375] c. For example, whenever there are uniformly spaced tiles, the number of explicitly provided tile columns / rows (e.g., num_exp_tile_columns_minus1, num_exp_tile_rows_minus1) may be less than or equal to the total number of uniform tiles.

[0376] d. If the tile layout is similar to uniformly spaced followed by non-uniformly spaced (i.e., the picture starts with uniformly spaced tiles and ends with multiple non-uniformly spaced tiles),

[0377] i. In one example, the widths of the tile columns of the non-uniformly spaced tiles located behind the picture can be first assigned in reverse order (i.e., the order of tile indices is equal to NumTileColumns, NumTileColumns - 1, NumTileColumns - 2,...), and then the widths of the tile columns of the uniformly spaced tiles located in front of the picture can be implicitly derived in reverse order (i.e., the order of tile indices is equal to NumTileColumns - T, NumTileColumns - T - 1,..., 2, 1, 0, where T represents the number of non-uniform tile columns).

[0378] 17. The syntax element specifying the difference between the representative tile indices of two rectangular stripes can be used only when the condition is true, where one of the two rectangular stripes is the i-th stripe (e.g., tile_idx_delta[i]).

[0379] a. In one example, the condition is (i < num_slices_in_pic_minus1), where num_slices_in_pic_minus1 plus 1 represents the number of stripes in the picture.

[0380] b. In one example, the condition is (i!= num_slices_in_pic_minus1), where num_slices_in_pic_minus1 plus 1 represents the number of stripes in the picture.

[0381] 18. Whether and / or how the Merge Estimation Region (MER) size is signaled or interpreted or limited (e.g., signaled by log2_parallel_merge_level_minus2) may depend on the minimum allowed coding block size (e.g., signaled / expressed as log2_min_luma_coding_block_size_minus2 and / or MinCbSizeY).

[0382] a. In one example, the size of the MER is required to be no smaller than the minimum allowed codec block size.

[0383] i. For example, log2_parallel_merge_level_minus2 is required to be equal to or greater than log2_min_luma_coding_block_size_minus2.

[0384] ii. For example, log2_parallel_merge_level_minus2 is required to be in the range of log2_min_luma_coding_block_size_minus2 to CtbLog2SizeY–2.

[0385] b. In one example, the difference between Log2(MER size) and Log2(MinCbSizeY) is signaled, which is denoted as log2_parallel_merge_level_minus_log2_mincb.

[0386] i. For example, log2_parallel_merge_level_minus_log2_mincb is coded by unary code (ue).

[0387] ii. For example, log2_parallel_merge_level_minus_log2_mincb is required to be in the range of 0 to CtbLog2SizeY-log2_min_luma_coding_block_size_minus2–2.

[0388] iii. For example, Log2ParMrgLevel=log2_parallel_merge_level_minus_log2_mincb+log2_min_luma_coding_block_size_minus2+2, where Log2ParMrgLevel is used to control the MER size.

[0389] 19. It is proposed that when num_exp_slices_in_tile[i] is equal to 0, the slice height of the i-th slice in units of CTU rows is derived, for example, represented as sliceHeightInCtus[i].

[0390] a. In one example, when num_exp_slices_in_tile[i] is equal to 0, sliceHeightInCtus[i] is derived to be equal to RowHeight[SliceTopLeftTileIdx[i] / NumTileColumns].

[0391] 20. It is proposed that the num_exp_slices_in_tile[i]-1th slice in the slice containing the i-th slice in the picture is always present, and the height is always exp_slice_height_in_ctus_minus1[i][num_exp_slices_in_tile[i]-1]+1 CTU rows.

[0392] a. Optionally, the num_exp_slices_in_tile[i]-1th slice in the slice containing the i-th slice in the picture may or may not exist, and the height is less than or equal to exp_slice_height_in_ctus_minus1[i][num_exp_slices_in_tile[i]-1]+1 CTU rows.

[0393] 21. It is proposed that during the derivation of information of rectangular slices, the variable tileIdx is only updated for slices with picture-level slice indices less than num_slices_in_pic_minus1, ie, the variable tileIdx is not updated for the last slice in each picture of the reference PPS.

[0394] 22. It is proposed that the num_exp_tile_columns_minus1-th tile column always exists in the picture of the reference PPS, and the width is always tile_column_width_minus1[num_exp_tile_columns_minus1]+1 CTB.

[0395] 23. It is proposed that the num_exp_tile_rows_minus1th tile row always exists in the picture of the reference PPS, and the height is always tile_column_height_minus1[num_exp_tile_rows_minus1]+1 CTBs.

[0396] 24. It is proposed that when the maximum picture width and the maximum picture height are both less than CtbSizeY, the signaling notification of the syntax element sps_num_subpics_minus1 can be skipped.

[0397] a. Optionally, in addition, when the above condition is true, the value of sps_num_subpics_minus1 is inferred to be equal to 0.

[0398] 25. It is proposed that when the picture width is not greater than CtbSizeY, the signaling of the syntax element num_exp_tile_columns_minus1 can be skipped.

[0399] b. Optionally, in addition, when the above condition is true, the value of num_exp_tile_columns_minus1 is inferred to be equal to 0.

[0400] 26. It is proposed that when the picture height is not greater than CtbSizeY, the signaling notification of the syntax element num_exp_tile_rows_minus1 can be skipped.

[0401] c. Optionally, in addition, when the above condition is true, the value of num_exp_tile_row_minus1 is inferred to be equal to 0.

[0402] 27. It is proposed that when num_exp_tile_columns_minus1 is equal to PicWidthInCtbsY-1, the signaling of the syntax element tile_column_width_minus1[i] can be skipped, and i ranges from 0 to num_exp_tile_columns_minus1 (inclusive).

[0403] d. Optionally, in addition, the value of tile_column_width_minus1[i] is inferred to be equal to 0.

[0404] 28. It is proposed that when num_exp_tile_rows_minus1 is equal to PicHeightInCtbsY-1, the signaling of the syntax element tile_row_height_minus1[i] can be skipped, and i ranges from 0 to num_exp_tile_rows_minus1 (inclusive).

[0405] e. Optionally, in addition, the value of tile_row_height_minus1[i] is inferred to be equal to 0.

[0406] 29. The height of uniform slices that partition a tile is indicated by the last entry of exp_slice_height_in_ctus_minus1[] which indicates the height of a slice in the tile. Non-uniform slices are slices below the explicitly signaled slices. For example: uniformSliceHeight = exp_slice_height_in_ctus_minus1[i][num_exp_slices_in_tile[i]-1]+1

[0407] 30. It is proposed that the width of the 'num_exp_tile_columns_minus1'th tile column is not allowed to be reset, ie the width can be derived directly using the parsed value from the bitstream (eg, represented by tile_column_width_minus1[num_exp_tile_columns_minus1]) without referring to other information.

[0408] a. In one example, the width of the 'num_exp_tile_columns_minus1'th tile column is directly set to tile_column_width_minus1[num_exp_tile_columns_minus1] plus 1. Optionally, in addition, tile_column_width_minus1[num_exp_tile_columns_minus1] is used to derive the width of tile columns with indices greater than num_exp_tile_columns_minus1, for example as specified in clause 6.5.1.

[0409] b. Similarly, no reset is allowed for the height of the num_exp_tile_columns_minus1th tile row, i.e., the height can be derived directly using the parsed value from the bitstream (e.g., represented by tile_row_height_minus1[num_exp_tile_columns_minus1]) without referring to other information.

[0410] i. In one example, the height of the 'num_exp_tile_columns_minus1'th tile row is directly set to tile_row_height_minus1[num_exp_tile_columns_minus1] plus 1. Optionally, in addition, tile_row_height_minus1[num_exp_tile_columns_minus1] is used to derive the height of tile rows with indices greater than num_exp_tile_columns_minus1, for example as specified in clause 6.5.1.

[0411] 31. It is proposed that the height of the num_exp_slices_in_tile[i]–1th slice in a slice is not allowed to be reset, that is, the height can be derived directly using the parsed value from the bitstream (e.g., represented by exp_slice_height_in_ctus_minus1[i][num_exp_slices_in_tile[i]-1]) without referring to other information.

[0412] a. In one example, the height of the num_exp_slices_in_tile[i]–1th slice in a slice is directly set to exp_slice_height_in_ctus_minus1[i][num_exp_slices_in_tile[i]–1] plus 1. Optionally, in addition, exp_slice_height_in_ctus_minus1[i][num_exp_slices_in_tile[i]–1] is used to derive the height of slices with indices greater than num_exp_slices_in_tile[i]-1.

[0413] 6. Examples

[0414] In the following examples, added parts are marked as bold text, underlined text and italic text. Deleted parts are marked in [[ ]].

[0415] 6.1. Example 1: Example sub-picture level strip index change

[0416] 3 Definition

[0417] : When rect_slice_flag is equal to 1, [[one]] slice index in the slice list of the picture (in the order signaled in the PPS).

[0418] : When rect_slice_flag is equal to 1, the slice index of the slice list in the sub-picture (in the order signaled in the PPS). ]]

[0419]

[0420] 6.5.1 CTB Raster Scanning, Slice Scanning, and Sub-Image Scanning Process

[0421]

[0422] [[List NumSlicesInSubpic[i], specifies the number of rectangular slices in the i-th subpic,]] The deduction is as follows:

[0423]

[0424]

[0425] 7.4.8.1 Common Strip Header Semantics

[0426]

[0427] Specifies the slice address of the slice. When not present, the value of slice_address is inferred to be equal to 0. When rect_slice_flag is equal to 1 and NumSlicesInSubpic[CurrSubpicIdx] is equal to 1, the value of slice_address is inferred to be equal to 0.

[0428] If rect_slice_flag is equal to 0, the following applies:

[0429] - The strip address is the raster scan slice index.

[0430] - The length of slice_address is Ceil(Log2(NumTilesInPic)) bits.

[0431] -slice_address values ​​should be in the range of 0 to NumTilesInPic-1 (inclusive).

[0432] Otherwise (rect_slice_flag is equal to 1), the following applies:

[0433] - Stripe address is The sub-picture level slice index of the slice of , where i is the picture level slice index of the slice.

[0434] - The length of slice_address is Ceil(Log2(NumSlicesInSubpic[CurrSubpicIdx])) bits.

[0435] -slice_address values ​​should be in the range of 0 to NumSlicesInSubpic[CurrSubpicIdx]-1 (inclusive).

[0436] A requirement for bitstream conformance is that the following constraints apply:

[0437] - If rect_slice_flag is equal to 0 or subpic_info_present_flag is equal to 0, the value of slice_address shall not be equal to the value of slice_address of any other codec slice NAL unit of the same codec picture.

[0438] Otherwise, the pair of slice_subpic_id value and slice_address value shall not be equal to the pair of slice_subpic_id value and slice_address value of any other codec slice NAL unit of the same codec picture.

[0439] - The shape of the slices of a picture shall be such that, when decoded, the entire left boundary and the entire top boundary of each CTU shall consist of the picture boundary or of the boundary of the previously decoded CTU(s).

[0440]

[0441] Add 1 (when present) to specify the number of tiles in the slice. The value of num_tiles_in_slice_minus1 should be in the range of 0 to NumTilesInPic-1 (inclusive).

[0442] The variable NumCtusInCurrSlice that specifies the number of CTUs in the current slice and the list CtbAddrInCurrSlice[i] that specifies the picture raster scan address of the i-th CTB in the slice (where i ranges from 0 to NumCtusInCurrSlice-1 (inclusive)) are derived as follows:

[0443]

[0444] 6.2. Example 2: Signaling strips in PPS for non-rectangular patterns

[0445] 7.3.2.4 Picture Parameter Set RBSP Syntax

[0446]

[0447]

[0448] 7.3.7.1 General Strip Header Syntax

[0449]

[0450] 7.4.3.4 Picture Parameter Set RBSP Semantics

[0451] Specifies the number of [[rectangular]] slices in each picture of the referenced PPS, plus 1. The value of num_slices_in_pic_minus1 shall be in the range of 0 to MaxSlicesPerPicture-1, inclusive, where MaxSlicesPerPicture is specified in Annex A. When no_pic_partition_flag is equal to 1, the value of num_slices_in_pic_minus1 is inferred to be equal to 0.

[0452]

[0453]

[0454] 7.4.8.1 Common Strip Header Semantics

[0455]

[0456] Slice address of the slice. When not present, the value of slice_address is inferred to be equal to 0. When rect_slice_flag is equal to 1 and NumSlicesInSubpic[CurrSubpicIdx] is equal to 1, the value of slice_address is inferred to be equal to 0. When rect_slice_flag is equal to 0 and NumSlicesInPic is equal to 1, the value of slice_address is inferred to be equal to 0.

[0457] If rect_slice_flag is equal to 0, the following applies:

[0458] - The strip address is the picture-level strip index of the [[raster scan slice index]] strip.

[0459] -The length of slice_address is Ceil(Log2(NumSlicesInPic [[NumTilesInPic]])) positions.

[0460] -slice_address value should be between 0 and The value is in the range of [[NumTilesInPic]] - 1 (inclusive).

[0461] Otherwise (rect_slice_flag is equal to 1), the following applies:

[0462] - The slice address is the sub-picture level slice index of the slice.

[0463] - The length of slice_address is Ceil(Log2(NumSlicesInSubpic[CurrSubpicIdx])) bits.

[0464] - The value of slice_address should be in the range of 0 to NumSlicesInSubpic[CurrSubpicIdx]-1 (inclusive).

[0465] A requirement for bitstream conformance is that the following constraints apply:

[0466] -[[If rect_slice_flag is equal to 0 or subpic_info_present_flag is equal to 0, the value of slice_address shall not be equal to the value of slice_address of any other codec slice NAL unit of the same codec picture.

[0467] -otherwise,]] Then the pair of slice_subpic_id value and slice_address value shall not be equal to the pair of slice_subpic_id value and slice_address value of any other codec slice NAL unit of the same codec picture.

[0468] - The shape of the slices of a picture shall be such that, when decoded, the entire left boundary and the entire top boundary of each CTU shall consist of the picture boundary or of the boundary of the previously decoded CTU(s).

[0469]

[0470] The variable NumCtusInCurrSlice that specifies the number of CTUs in the current slice and the list CtbAddrInCurrSlice[i] that specifies the picture raster scan address of the i-th CTB in the slice (where i ranges from 0 to NumCtusInCurrSlice-1 (inclusive)) are derived as follows:

[0471]

[0472]

[0473] 6.3. Example 3: Signaling Slices Conditioned by Image Dimension

[0474] 7.3.2.4 Picture Parameter Set RBSP Syntax

[0475]

[0476]

[0477] 6.4. Example 4: Example 1 on the semantics of tile_column_width_minus1 and tile_row_height_minus1

[0478] 7.4.3.4 Picture Parameter Set RBSP Semantics

[0479]

[0480] [i] Plus 1 provision The width of the i-th tile column in CTBs, where i ranges from 0 to num_exp_tile_columns_minus1-1, inclusive. tile_column_width_minus1[num_exp_tile_columns_minus1] is used to derive the width of tile columns with indices greater than or equal to num_exp_tile_columns_minus1 as specified in clause 6.5.1. The value of tile_column_width_minus1[i] shall be in the range of 0 to PicWidthInCtbsY-1, inclusive. When not present, the value of tile_column_width_minus1[0] is inferred to be equal to PicWidthInCtbsY1-1.

[0481] [i] Plus 1 provision The height of the i-th tile row in CTBs, where i ranges from 0 to num_exp_tile_rows_minus1-1, inclusive. tile_row_height_minus1[num_exp_tile_rows_minus1] is used to derive the height of tile rows whose indices are greater than or equal to num_exp_tile_rows_minus1 as specified in clause 6.5.1. The value of tile_row_height_minus1[i] shall be in the range of 0 to PicHeightInCtbsY-1, inclusive. When not present, the value of tile_row_height_minus1[0] is inferred to be equal to PicHeightInCtbsY1-1.

[0482]

[0483] 6.5. Example 5: Example 2 on the semantics of tile_column_width_minus1 and tile_row_height_minus1

[0484] 7.4.3.4 Picture Parameter Set RBSP Semantics

[0485]

[0486] [i] plus 1 specifies the width of the i-th slice column in CTB units, and the range of i is 0 to num_exp_tile_columns_minus1–1 (inclusive). tile_column_width_minus1[num_exp_tile_columns_minus1] is used to derive the width of tile columns with indices greater than or equal to num_exp_tile_columns_minus1 as specified in clause 6.5.1. The value of tile_column_width_minus1[i] shall be in the range 0 to PicWidthInCtbsY-1, inclusive. When not present, the value of tile_column_width_minus1[0] is inferred to be equal to PicWidthInCtbsY-1.

[0487] [i] plus 1 specifies the height of the i-th slice row in CTB units, i ranges from 0 to num_exp_tile_rows_minus1–1 (inclusive). tile_row_height_minus1[num_exp_tile_rows_minus1] is used to derive the height of tile rows with indices greater than or equal to num_exp_tile_rows_minus1 as specified in clause 6.5.1. The value of tile_row_height_minus1[i] shall be in the range 0 to PicHeightInCtbsY-1, inclusive. When not present, the value of tile_row_height_minus1[0] is inferred to be equal to PicHeightInCtbsY-1.

[0488]

[0489] 6.6. Example 6: Example Derivation of CTUs in a Strip

[0490] 6.5 Scanning Process

[0491] 6.5.1 CTB Raster Scanning, Slice Scanning, and Sub-Image Scanning Process

[0492]

[0493] For rectangular slices, the list NumCtusInSlice[i] specifying the number of CTUs in the i-th slice, where i ranges from 0 to num_slices_in_pic_minus1 (inclusive), the list SliceTopLeftTileIdx[i] specifying the index of the top left slice of the slice, where i ranges from 0 to num_slices_in_pic_minus1 (inclusive), and the matrix CtbAddrInSlice[i][j] specifying the picture raster scan address of the j-th CTB in the i-th slice, where i ranges from 0 to num_slices_in_pic_minus1 (inclusive) and j ranges from 0 to NumCtusInSlice[i]-1 (inclusive), are derived as follows:

[0494]

[0495]

[0496] 6.7. Example 7: Signaling of MER Size

[0497] 7.3.2.3 Sequence Parameter Set RBSP Syntax

[0498]

[0499] 7.4.3.3 Sequence Parameter Set RBSP Semantics

[0500] Adding log2_min_luma_coding_block_size_minus2+2 specifies the value of the variable Log2ParMrgLevel, which is used in the derivation process of spatial merge candidates as specified in clause 8.5.2.3, the derivation process of motion vectors and reference indices in sub-block merge mode as specified in clause 8.5.5.2, and is used to control the call of the update process of the history-based motion vector prediction value list in clause 8.5.2.1. The value of should be in the range of 0 to CtbLog2SizeY-log2_min_luma_coding_block_size_minus2-2 (inclusive). The derivation of the variable Log2ParMrgLevel is as follows:

[0501] Log2ParMrgLevel=log2_parallel_merge_level_minus2+log2_min_luma_coding_block_size_minus2+2 (68)

[0502] 6.8. Example 8: Signaling for rectangular strips

[0503] 6.5.1 CTB Raster Scanning, Slice Scanning, and Sub-Image Scanning Process

[0504]

[0505] The list ctbToSubpicIdx[ctbAddrRs] that specifies the conversion from CTB addresses in the picture raster scan to sub-picture indices (where ctbAddrRs ranges from 0 to PicSizeInCtbsY-1 (inclusive)) is derived as follows:

[0506]

[0507] When rect_slice_flag is equal to 1, a list NumCtusInSlice[i] specifying the number of CTUs in the i-th slice, where i ranges from 0 to num_slices_in_pic_minus1, inclusive; a list SliceTopLeftTileIdx[i] specifying the slice index of the slice containing the first CTU in the slice, where i ranges from 0 to num_slices_in_pic_minus1, inclusive; a matrix CtbAddrInSlice[i][j] specifying the picture raster scan address of the j-th CTB in the i-th slice, where i ranges from 0 to num_slices_in_pic_minus1, inclusive and j ranges from 0 to NumCtusInSlice[i]-1, inclusive; and The derivation is as follows:

[0508]

[0509]

[0510]

[0511] It is a requirement of bitstream conformance that the value of NumCtusInSlice[i], where i ranges from 0 to num_slices_in_pic_minus1, inclusive, shall be greater than 0. Furthermore, it is a requirement of bitstream conformance that the matrix CtbAddrInSlice[i][j], where i ranges from 0 to num_slices_in_pic_minus1, inclusive, and j ranges from 0 to NumCtusInSlice[i]-1, inclusive, shall include each of all CTB addresses in the range 0 to PicSizeInCtbsY-1, inclusive, once and only once.

[0512]

[0513] 7.3.2.4 Picture Parameter Set RBSP Syntax

[0514]

[0515]

[0516] 7.4.3.4 Picture parameter set semantics

[0517]

[0518] When tile_idx_delta_present_flag is equal to 1, it specifies that the tile_idx_delta[i] syntax element may be present in the PPS, and all rectangular slices in the pictures referencing the PPS are specified with increasing values ​​of i in the order indicated by the value of tile_idx_delta[i]. When not present, the value of tile_idx_delta_present_flag is inferred to be equal to 0.

[0519] [i] plus 1 specifies the width of the i-th rectangular slice in units of tile columns. The value of slice_width_in_tiles_minus1[i] shall be in the range of 0 to NumTileColumns-1 (inclusive).

[0520] When i is less than num_slices_in_pic_minus1 and NumTileColumns is equal to 1, the value of slice_width_in_tiles_minus1[i] is inferred to be equal to 0.

[0521] When num_exp_slices_in_tiles_MINUS1[i] is equal to 0, [i] plus 1 specifies the height of the i-th rectangular slice in tile rows. The value of slice_height_in_tiles_minus1[i] shall be in the range of 0 to NumTileRows-1 (inclusive).

[0522] When i is less than num_slices_in_pic_minus1 and slice_height_in_tiles_minus1[i] does not exist, it is inferred to be equal to NumTileRows==1?0:slice_height_in_tiles_MINUS1[i-1].

[0523] [i] specifies the number of explicitly provided slice heights for the slice in the slice containing the i-th slice (i.e., the slice with slice index equal to SliceTopLeftTileIdx[i]). The value of num_exp_slices_in_tile[i] should be in the range of 0 to RowHeight[SliceTopLeftTileIdx[i] / NumTileColumns] - 1, inclusive. When not present, the value of num_exp_slices_in_tile[i] is inferred to be equal to 0.

[0524]

[0525] Regulation The value of tile_idx_delta[i] shall be in the range -NumTilesInPic+1 to NumTilesInPic-1 (inclusive). When not present, the value of tile_idx_delta[i] is inferred to be equal to 0. When present, the value of tile_idx_delta[i] shall not be equal to 0.

[0526]

[0527] 6.9. Example 9: Signaling for rectangular strips

[0528] 6.5.1 CTB Raster Scanning, Slice Scanning, and Sub-Image Scanning Process

[0529]

[0530] When rect_slice_flag is equal to 1, a list NumCtusInSlice[i] that specifies the number of CTUs in the i-th slice (where i ranges from 0 to num_slices_in_pic_minus1 (inclusive)), a list SliceTopLeftTileIdx[i] that specifies the slice index of the slice containing the first CTU in the slice (where i ranges from 0 to num_slices_in_pic_minus1 (inclusive)), a list The matrix CtbAddrInSlice[i][j] of the picture raster scan addresses of the j-th CTB, where i ranges from 0 to num_slices_in_pic_minus1 (inclusive) and j ranges from 0 to NumCtusInSlice[i]-1 (inclusive), and the variable NumSlicesInTile[i] that specifies the number of slices in the slice containing the i-th slice (i.e., the slice with slice index equal to SliceTopLeftTileIdx[i]), are derived as follows:

[0531]

[0532]

[0533]

[0534] Optionally, above, the following line:

[0535]

[0536] Modify to the following;

[0537]

[0538] 6.10. Example 10: Signaling of Sub-Pictures and Slices

[0539] 6.5.1 CTB Raster Scanning, Slice Scanning, and Sub-Image Scanning Process

[0540] The variable NumTileColumns that specifies the number of tile columns and the list colWidth[i] that specifies the width of the i-th tile column in CTB units (where i ranges from 0 to NumTileColumns-1 (inclusive)) are derived as follows:

[0541]

[0542] The variable NumTileRows that specifies the number of tile rows and the list RowHeight[j] that specifies the height of the j-th tile row in CTB units (where j ranges from 0 to NumTileRows-1 (inclusive)) are derived as follows:

[0543]

[0544]

[0545] 7.3.2.3 Sequence Parameter Set RBSP Syntax

[0546]

[0547] 7.3.2.4 Picture Parameter Set RBSP Syntax

[0548]

[0549] 7.4.3.4 Picture parameter set semantics

[0550]

[0551] Add 1 to specify the number of explicitly provided tile column widths. The value of num_exp_tile_columns_minus1 should be in the range of 0 to PicWidthInCtbsY-1 (inclusive).

[0552] Add 1 to specify the number of explicitly provided tile row heights. The value of num_exp_tile_rows_minus1 should be in the range of 0 to PicHeightInCtbsY-1 (inclusive).

[0553] [i] plus 1 specifies the width of the i-th slice column in CTB units, and the range of i is 0 to (Inclusive). tile_column_width_minus1[num_exp_tile_columns_minus1] is used to derive the index The width of the tile column of num_exp_tile_columns_minus1 as specified in clause 6.5.1. The value of tile_column_width_minus1[i] shall be in the range 0 to PicWidthInCtbsY-1 (inclusive). When not present, the value of tile_column_width_minus1[i] is inferred to be equal to

[0554] tile_row_height_minus1[i] plus 1 specifies the height of the i-th tile row in CTB units, i ranges from 0 to (Inclusive). tile_row_height_minus1[num_exp_tile_rows_minus1] is used to derive the index The height of a tile row of num_exp_tile_rows_minus1 as specified in clause 6.5.1. The value of tile_row_height_minus1[i] shall be in the range 0 to PicHeightInCtbsY-1 (inclusive). When not present, the value of tile_row_height_minus1[i] is inferred to be equal to

[0555]

[0556] Figure 7 is a block diagram illustrating an example video processing system 1900 in which the various techniques disclosed herein may be implemented. Various implementations may include some or all of the components of system 1900. System 1900 may include an input 1902 for receiving video content. The video content may be received in a raw or uncompressed format (e.g., 8-bit or 10-bit multi-component pixel values), or may be received in a compressed or encoded format. Input 1902 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces (e.g., Ethernet, passive optical network (PON), etc.), and wireless interfaces (e.g., Wi-Fi or cellular interfaces).

[0557] System 1900 may include a codec component 1904 that implements the various codecs or encoding methods described in this document. The codec component 1904 can reduce the average bit rate of the video from the input 1902 to the output of the codec component 1904 to generate a codec representation of the video. Therefore, codec technology is sometimes referred to as video compression or video transcoding technology. The output of the codec component 1904 can be stored or transmitted via a communication connected to the component 1906. The stored or transmitted bitstream (or codec) representation of the video received at the input 1902 can be used by component 1908 to generate pixel values ​​or displayable video sent to the display interface 1910. The process of generating a user-viewable video based on the bitstream is sometimes referred to as video decompression. In addition, although some video processing operations are referred to as "codec" operations or tools, it will be understood that the codec tools or operations are used at the encoder, and the corresponding decoding tools or operations that reverse the codec results will be performed by the decoder.

[0558] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB), High-Definition Multimedia Interface (HDMI), or DisplayPort, etc. Examples of storage interfaces include SATA (Serial Advanced Technology Attachment), PCI, IDE interfaces, etc. The technology described in this document can be embodied in various electronic devices, such as mobile phones, laptop computers, smartphones, or other devices capable of performing digital data processing and / or video display.

[0559] Figure 8 36 is a block diagram of a video processing device 3600. Device 3600 can be used to implement one or more methods described herein. Device 3600 can be embodied in a smartphone, tablet computer, computer, Internet of Things (IoT) receiver, etc. Device 3600 may include one or more processors 3602, one or more memories 3604, and video processing hardware 3606. Processor(s) 3602 can be configured to implement one or more methods described herein. One or more memories 3604 can be used to store data and code used to implement the methods and techniques described herein. Video processing hardware 3606 can be used to implement some of the techniques described in this document in hardware circuitry.

[0560] Figure 10 is a block diagram illustrating an example video coding system 100 that may utilize the techniques of this disclosure.

[0561] like Figure 10 As shown, the video codec system 100 may include a source device 110 and a destination device 120. The source device 110 generates encoded video data, which may be referred to as a video encoding device. The destination device 120 may decode the encoded video data generated by the source device 110, which may be referred to as a video decoding device.

[0562] Source device 110 may include a video source 112 , a video encoder 114 , and an input / output (I / O) interface 116 .

[0563] The video source 112 may include a source such as a video capture device, an interface for receiving video data from a video content provider, and / or a computer graphics system for generating video data, or a combination of these sources. The video data may include one or more pictures. The video encoder 114 encodes the video data from the video source 112 to generate a bitstream. The bitstream may include a sequence of bits that form a codec representation of the video data. The bitstream may include a codec picture and associated data. The codec picture is a codec representation of the picture. The associated data may include a sequence parameter set, a picture parameter set, and other syntax structures. The I / O interface 116 may include a modulator / demodulator (modem) and / or a transmitter. The encoded video data may be sent directly to the destination device 120 via the network 130a via the I / O interface 116. The encoded video data may also be stored on a storage medium / server 130b for access by the destination device 120.

[0564] Destination device 120 may include an I / O interface 126 , a video decoder 124 , and a display device 122 .

[0565] I / O interface 126 may include a receiver and / or a modem. I / O interface 126 may obtain encoded video data from source device 110 or storage medium / server 130b. Video decoder 124 may decode the encoded video data. Display device 122 may display the decoded video data to a user. Display device 122 may be integrated with destination device 120 or may be located external to destination device 120, with destination device 120 configured to interface with an external display device.

[0566] The video encoder 114 and the video decoder 124 may operate according to a video compression standard, such as the High Efficiency Video Codec (HEVC) standard, the Versatile Video Codec (VVM) standard, and other current and / or future standards.

[0567] Figure 11 is a block diagram illustrating an example of a video encoder 200, which may be Figure 10 The video encoder 114 in the system 100 is shown.

[0568] Video encoder 200 may be configured to perform any or all of the techniques of this disclosure. Figure 11 In the example of , video encoder 200 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of video encoder 200. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.

[0569] The functional components of the video encoder 200 may include a segmentation unit 201, a prediction unit 202 which may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205 and an intra-frame prediction unit 206, a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213 and an entropy coding unit 214.

[0570] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an intra block copy (IBC) unit. The IBC unit may perform prediction in IBC mode, where at least one reference picture is a picture in which the current video block is located.

[0571] Furthermore, some components (e.g., motion estimation unit 204 and motion compensation unit 205) may be highly integrated, but for purposes of explanation, are not described in detail in the preceding text. Figure 11 In the examples, they are respectively represented.

[0572] The partitioning unit 201 may partition a picture into one or more video blocks. The video encoder 200 and the video decoder 300 may support various video block sizes.

[0573] The mode selection unit 203 can, for example, select one of the coding modes (intra or inter) based on the error result, and provide the resulting intra or inter coded block to the residual generation unit 207 to generate residual block data, and to the reconstruction unit 212 to reconstruct the coded block for use as a reference picture. In some examples, the mode selection unit 203 can select a combination of intra prediction and inter prediction (CIIP) modes, where the prediction is based on an inter prediction signal and an intra prediction signal. The mode selection unit 203 can also select the resolution of the motion vector for the block in the case of inter prediction (e.g., sub-pixel precision or integer pixel precision).

[0574] To perform inter-frame prediction on the current video block, the motion estimation unit 204 may generate motion information for the current video block by comparing the current video block with one or more reference frames from the buffer 213. The motion compensation unit 205 may determine a predicted video block for the current video block based on the motion information and decoded samples of pictures other than the picture associated with the current video block from the buffer 213.

[0575] For example, motion estimation unit 204 and motion compensation unit 205 may perform different operations on the current video block depending on whether the current video block is in an I slice, a P slice, or a B slice.

[0576] In some examples, motion estimation unit 204 may perform unidirectional prediction on the current video block, and motion estimation unit 204 may search for a reference video block for the current video block in the reference pictures in list 0 or list 1. Motion estimation unit 204 may then generate a reference index indicating the reference picture in list 0 or list 1 containing the reference video block, and a motion vector indicating the spatial displacement between the current video block and the reference video block. Motion estimation unit 204 may output the reference index, the prediction direction indicator, and the motion vector as motion information for the current video block. Motion compensation unit 205 may generate a predicted video block for the current block based on the reference video block indicated by the motion information for the current video block.

[0577] In other examples, the motion estimation unit 204 may perform bidirectional prediction on the current video block. The motion estimation unit 204 may search for a reference video block for the current video block in the reference pictures in list 0 and may also search for another reference video block for the current video block in the reference pictures in list 1. The motion estimation unit 204 may then generate reference indexes indicating the reference pictures in list 0 and list 1 containing the reference video block, and a motion vector indicating the spatial displacement between the reference video block and the current video block. The motion estimation unit 204 may output the reference index and motion vector for the current video block as motion information for the current video block. The motion compensation unit 205 may generate a predicted video block for the current video block based on the reference video block indicated by the motion information of the current video block.

[0578] In some examples, motion estimation unit 204 may output a complete set of motion information for use in the decoding process of a decoder.

[0579] In some examples, motion estimation unit 204 may not output a complete set of motion information for the current video. Instead, motion estimation unit 204 may reference motion information of another video block to signal the motion information of the current video block. For example, motion estimation unit 204 may determine that the motion information of the current video block is sufficiently similar to the motion information of a neighboring video block.

[0580] In one example, motion estimation unit 204 may indicate, in a syntax structure associated with the current video block, a value that indicates to video decoder 300 that the current video block has the same motion information as another video block.

[0581] In another example, the motion estimation unit 204 can identify another video block and a motion vector difference (MVD) in a syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 300 can use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.

[0582] As discussed above, the video encoder 200 may predictively signal motion vectors.Two examples of predictive signaling techniques that may be implemented by the video encoder 200 include advanced motion vector prediction (AMVP) and merge mode signaling.

[0583] The intra-frame prediction unit 206 can perform intra-frame prediction on the current video block. When the intra-frame prediction unit 206 performs intra-frame prediction on the current video block, the intra-frame prediction unit 206 can generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block may include a predicted video block and various syntax elements.

[0584] The residual generation unit 207 may generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) the predicted video block of the current video block from the current video block. The residual data for the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.

[0585] In other examples, for the current video block, there may be no residual data of the current video block, for example, in skip mode, the residual generation unit 207 may not perform the subtraction operation.

[0586] Transform processing unit 208 may generate one or more transform coefficient video blocks for a current video block by applying one or more transforms to the residual video block associated with the current video block.

[0587] After transform processing unit 208 generates a transform coefficient video block associated with the current video block, quantization unit 209 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values ​​associated with the current video block.

[0588] The inverse quantization unit 210 and the inverse transform unit 211 may apply inverse quantization and inverse transform, respectively, to the transform coefficient video block to reconstruct a residual video block from the transform coefficient video block. The reconstruction unit 212 may add the reconstructed residual video block to corresponding samples of one or more predicted video blocks generated by the prediction unit 202 to generate a reconstructed video block associated with the current block for storage in the buffer 213.

[0589] After the reconstruction unit 212 reconstructs the video block, a loop filtering operation may be performed to reduce video block artifacts in the video block.

[0590] The entropy coding unit 214 may receive data from other functional components of the video encoder 200. When the entropy coding unit 214 receives the data, the entropy coding unit 214 may perform one or more entropy coding operations to generate entropy-coded data and output a bitstream including the entropy-coded data.

[0591] Figure 12 is a block diagram illustrating an example of a video decoder 300, which may be Figure 10 The video decoder 114 in the system 100 is shown.

[0592] Video decoder 300 may be configured to perform any or all of the techniques of this disclosure. Figure 12 In the example of FIG, video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of video decoder 300. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.

[0593] exist Figure 12 In the example of FIG. 3 , the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, the video decoder 300 can perform encoding passes generally similar to those described with respect to the video encoder 200 ( Figure 11 )The opposite decoding pass.

[0594] The entropy decoding unit 301 can retrieve a coded bitstream. The coded bitstream can include entropy-coded video data (e.g., coded blocks of video data). The entropy decoding unit 301 can decode the entropy-coded video data, and the motion compensation unit 302 can determine motion information based on the entropy-decoded video data, including motion vectors, motion vector precision, reference picture list index, and other motion information. For example, the motion compensation unit 302 can determine this information by performing AMVP and merge modes.

[0595] The motion compensation unit 302 may generate a motion compensated block, possibly performing interpolation based on an interpolation filter. An identifier of an interpolation filter to be used with sub-pixel precision may be included in the syntax element.

[0596] The motion compensation unit 302 may calculate interpolated values ​​of sub-integer pixels of the reference block using the interpolation filter used by the video encoder 200 during encoding of the video block. The motion compensation unit 302 may determine the interpolation filter used by the video encoder 200 based on received syntax information and use the interpolation filter to generate a prediction block.

[0597] The motion compensation unit 302 may use some syntax information to determine the size of blocks used to encode frames and / or slices of the coded video sequence, partitioning information describing how each macroblock of a picture of the coded video sequence is partitioned, a mode indicating how each partition is to be encoded, one or more reference frames (and reference frame lists) to use for each inter-frame coded block, and other information used to decode the coded video sequence.

[0598] The intra prediction unit 303 can form a prediction block based on spatially adjacent blocks using, for example, an intra prediction mode received in the bitstream. The inverse quantization unit 303 inverse quantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 301. The inverse transform unit 303 applies an inverse transform.

[0599] The reconstruction unit 306 can add the residual block to the corresponding prediction block generated by the motion compensation unit 202 or the intra prediction unit 303 to form a decoded block. If necessary, a deblocking filter can also be applied to the decoded block to remove blocking artifacts. The decoded video block is then stored in the buffer 307, which provides reference blocks for subsequent motion compensation / intra prediction and also produces decoded video for presentation on a display device.

[0600] A list of preferred solutions for some embodiments is provided below.

[0601] The following solution illustrates an example embodiment of the technique discussed in the previous section (eg, item 1).

[0602] 1. A video processing method (e.g., Figure 9 ), comprising: performing a conversion (902) between a video comprising one or more video pictures and a codec representation of the video, wherein each video picture comprises one or more sub-pictures, the sub-picture comprising one or more slices, wherein the codec representation conforms to a format rule; wherein the format rule provides that, when rectangular slice mode is enabled for the video picture, a picture-level slice index for each slice in each sub-picture in the video picture is derived without explicit signaling in the codec representation; wherein the format rule provides that the number of codec tree units in each slice is derivable from the picture-level slice index.

[0603] The following solution illustrates an example embodiment of the technique discussed in the previous section (eg, item 2).

[0604] 2. A video processing method, comprising: performing conversion between a video comprising one or more video pictures and a codec representation of the video, wherein each video picture comprises one or more sub-pictures, each sub-picture comprising one or more strips, wherein the codec representation complies with format rules; wherein the format rules specify that a sub-picture level strip index can be derived based on information in the codec representation without the need for signaling the sub-picture level strip index in the codec representation.

[0605] 3. The method of solution 2, wherein the format rule stipulates that due to the use of a rectangular slice structure, the sub-picture level slice index corresponds to the index of the slice in the list of slices in the sub-picture.

[0606] 4. The method of solution 2, wherein the format rules specify that the sub-picture level slice index is derived from a specific value of the picture level slice index.

[0607] The following solutions illustrate example embodiments of the techniques discussed in the previous section (eg, items 5, 6).

[0608] 5. A video processing method, comprising: performing conversion between a video comprising one or more video pictures and a codec representation of the video, wherein each video picture comprises one or more sub-pictures and / or one or more slices, wherein the codec representation complies with format rules; wherein the conversion complies with constraint rules.

[0609] 6. The method of solution 5, wherein the constraint rule stipulates that a slice cannot be in more than one sub-picture.

[0610] 7. The method of solution 5, wherein the constraint rule stipulates that a sub-picture cannot include two slices that are smaller than the corresponding slice to which the two slices belong.

[0611] The following solutions illustrate example embodiments of the techniques discussed in the previous section (eg, items 7, 8).

[0612] 8. A video processing method, comprising: performing conversion between a video comprising one or more video pictures and a codec representation of the video, wherein each video picture comprises one or more slices and / or one or more strips; wherein the codec representation complies with format rules; wherein the format rules specify that fields at the video picture level carry information about the segmentation of strips and / or slices in the video picture.

[0613] 9. The method of solution 8, wherein the field includes a video picture header.

[0614] 10. The method of solution 8, wherein the field includes a picture parameter set.

[0615] 11. The method of any of solutions 8-10, wherein the format rule specifies omitting slice segmentation information at the slice level by including the slice segmentation information in a field at the video picture level.

[0616] The following solution illustrates an example embodiment of the technique discussed in the previous section (eg, item 9).

[0617] 12. A video processing method comprising: performing a conversion between a video comprising one or more pictures and a codec representation of the video, wherein the conversion complies with a segmentation rule that determines whether a minimum number of strips into which the video picture is segmented is determined based on whether rectangular segmentation is used to segment the video picture.

[0618] 13. The method of solution 12, wherein the segmentation rule specifies using at least two stripes for non-rectangular segmentation and at least one stripe for rectangular segmentation.

[0619] 14. The method of solution 12, wherein the segmentation rule is also a function of whether and / or how many sub-pictures are used to segment the video picture.

[0620] The following solutions illustrate example embodiments of the techniques discussed in the previous section (eg, items 10, 11).

[0621] 15. A video processing method, comprising: performing conversion between a video strip of a video area of ​​a video and a codec representation of the video; wherein the codec representation complies with a format rule; wherein the format rule stipulates that the codec representation signals the video strip based on the upper left position of the video strip, wherein the format rule stipulates that the codec representation signals the height and / or width of the video strip in segmentation information signaled at the video unit level.

[0622] 16. The method of solution 15, wherein the format rule dictates that the video slices are signaled in the order of the slices defined by the format rule.

[0623] 17. The method of solution 15, wherein the video region corresponds to a sub-picture, and wherein the video unit level corresponds to a video picture.

[0624] The following solution illustrates an example embodiment of the technique discussed in the previous section (eg, item 12).

[0625] 18. A video processing method comprising: performing conversion between a video including video pictures and a codec representation of the video; wherein the codec representation complies with a format rule; wherein the format rule dictates omitting signaling of a difference between a slice index of a first slice in a rectangular strip and a slice index of a first slice in a next rectangular strip.

[0626] 19. The method of solution 18, wherein the difference is derived from the zeroth strip and the rectangular strip in the video picture.

[0627] The following solution illustrates an example embodiment of the technique discussed in the previous section (eg, item 13).

[0628] 20. A video processing method comprising: performing conversion between a video and a codec representation of the video, wherein the codec representation conforms to a format rule, wherein the format rule specifies a relationship between a width of a video picture and a size of a codec tree unit to control signaling notification of information used to derive a number of slice columns or a number of slice rows in a video picture.

[0629] 21. The method of solution 20, wherein the format rule specifies excluding signaling the number of slice rows or the number of slice columns when the width of the video picture is less than or equal to the width of the codec tree unit.

[0630] The following solution shows an example embodiment of the technique discussed in the previous section (e.g., item 16).

[0631] 22. A video processing method comprising: performing conversion between a video comprising one or more video pictures and an encoded representation of the video, wherein the encoded representation complies with format rules, wherein the format rules specify including slice layout information in the encoded representation of the video pictures comprising evenly spaced slices and non-evenly spaced slices.

[0632] 23. The method of solution 22, wherein the slice layout information is included in a syntax flag contained in a picture parameter set.

[0633] 24. The method of any of solutions 22-23, wherein the number of tile rows or the number of tile columns explicitly signaled is not less than the number of non-uniformly spaced tiles.

[0634] 25. The method of any of solutions 22-23, wherein the number of tile rows or the number of tile columns explicitly signaled is not less than the number of evenly spaced tiles.

[0635] 26. The method of any of the above solutions, wherein the video area includes a video codec unit.

[0636] 27. The method of any of the above solutions, wherein the video area comprises a video picture.

[0637] 28. The method of any of solutions 1 to 27, wherein converting comprises encoding the video into a codec representation.

[0638] 29. The method of any of solutions 1 to 27, wherein converting comprises decoding the codec representation to generate pixel values ​​of the video.

[0639] 30. A video decoding device comprising a processor configured to implement the method described in one or more of solutions 1 to 29.

[0640] 31. A video encoding apparatus comprising a processor configured to implement the method described in one or more of solutions 1 to 29.

[0641] 32. A computer program product having computer code stored thereon, which, when executed by a processor, causes the processor to implement the method described in any one of solutions 1 to 29.

[0642] 33. A method, apparatus or system as described in this document.

[0643] Figure 13 13 is a flowchart representation of a video processing method according to the present technology. The method 1300 includes, at operation 1310, performing conversion between a video picture of a video and a bitstream of the video according to a rule. The video picture includes one or more slices. The rule provides that, in response to a condition being met, a syntax element indicating a difference between slice indices of two rectangular slices is signaled, wherein one of the slice indices is denoted as i, where i is an integer.

[0644] In some embodiments, a second rectangular slice of the two rectangular slices is denoted as an (i+1)th rectangular slice, and the syntax element indicates a difference between a first slice index of a first slice containing a first codec tree unit in the (i+1)th rectangular slice and a second slice index of a second slice containing the first codec tree unit in the (i)th rectangular slice. In some embodiments, at least one condition satisfied includes i being less than (the number of rectangular slices in the video picture - 1). In some embodiments, at least one condition satisfied includes i not being equal to (the number of rectangular slices in the video picture - 1).

[0645] Figure 14 14 is a flow chart representation of a video processing method according to the present technology. The method 1400 includes, at operation 1410, performing conversion between a video picture of a video and a bitstream of the video according to a rule. The video picture includes one or more sub-pictures, and each sub-picture includes one or more rectangular slices. The rule provides for deriving a slice index at the sub-picture level for each rectangular slice in each sub-picture to determine the number of codec tree units in each slice.

[0646] In some embodiments, a sub-picture-level slice index is determined based on the decoding order of corresponding slices in the list of slices in the sub-picture. In some embodiments, if a first slice index at the sub-picture level of a first slice is less than a second slice index at the sub-picture level of a second slice, the first slice is processed before the second slice according to the decoding order. In some embodiments, a variable SubpicLevelSliceIdx is used to represent the sub-picture-level slice index. In some embodiments, a slice address is determined based on the sub-picture-level slice index. In some embodiments, a sub-picture-level slice index of a slice is determined based on the first sub-picture that includes the slice. In some embodiments, the sub-picture-level slice index is a non-negative integer. In some embodiments, the sub-picture-level slice index is greater than or equal to 0 and less than N, where N is the number of slices in the sub-picture.

[0647] In some embodiments, when a first slice and a second slice are different, a first slice index of the first slice at a sub-picture level is different from a second slice index of the second slice at a sub-picture level, wherein the first slice and the second slice are in the same sub-picture. In some embodiments, when a first slice index at a sub-picture level of the first slice is less than a second slice index at a sub-picture level of the second slice, a first slice index at a picture level of the first slice is less than a second slice index at a picture level of the second slice. In some embodiments, a slice index at a sub-picture level of a slice is determined based on a slice index at a picture level of the slice.

[0648] Figure 15 15 is a flowchart representation of a video processing method according to the present technology. The method 1500 includes, at operation 1510, determining a mapping relationship between a sub-picture-level slice index of a slice in a sub-picture and a picture-level slice index of the slice for conversion between a video picture of a video including one or more sub-pictures and a bitstream of the video. The method 1500 also includes, at operation 1520, performing the conversion based on the determination.

[0649] In some embodiments, the mapping relationship is represented as a two-dimensional array indexed using a picture-level slice index and a sub-picture index of a sub-picture. In some embodiments, the picture-level slice index is determined based on an array indicating the number of slices in each of one or more sub-pictures. In some embodiments, the two-dimensional array and the array indicating the number of slices in each of the one or more sub-pictures are determined based on a process of scanning all slices in the order of the picture-level slice index. In some embodiments, the mapping relationship is used to determine the picture-level slice index of a slice, and the number of codec treeblocks and / or the addresses of the codec treeblocks in the slice are determined based on the picture-level slice index of the slice.

[0650] Figure 1616 is a flow chart representation of a video processing method according to the present technology. Method 1600 includes, at operation 1610, performing conversion between a video picture of a video and a bitstream of the video according to a rule. A video picture includes one or more sub-pictures. The rule stipulates that a slice of the video is entirely within a single sub-picture of the video picture.

[0651] In some embodiments, the first stripe is located in a first slice, the first stripe is smaller than the first slice, the second stripe is located in a second slice, the second stripe is smaller than the second slice, the first slice is different from the second slice, and the first stripe and the second stripe are located in different stripes of the sub-picture.

[0652] Figure 17 17 is a flowchart representation of a video processing method according to the present technology. Method 1700 includes, at operation 1710, performing conversion between a video picture of a video and a bitstream of the video. A video picture includes one or more sub-pictures. The bitstream conforms to a format rule that specifies that information about the sub-pictures is included in a syntax structure associated with the picture.

[0653] In some embodiments, the syntax structure includes a picture header or a picture parameter set. In some embodiments, the video unit includes a flag indicating whether information about the partitioned picture is included in the syntax structure. In some embodiments, the video unit includes a picture header, a picture parameter set, or a sequence parameter set. In some embodiments, where information about the partitioned picture is included in the picture header and the picture parameter set, the information included in the picture header is used for conversion.

[0654] Figure 18 18 is a flowchart representation of a video processing method according to the present technology. The method 1800 includes, at operation 1810, determining slice segmentation information for a video picture of a video including one or more slices having a non-rectangular shape and a bitstream of the video. The method 1800 also includes, at operation 1820, performing the conversion based on the determination.

[0655] In some embodiments, slice segmentation information is stored in a video unit of a video, wherein the video unit includes a picture parameter set or a picture header. In some embodiments, the bitstream conforms to a format rule that specifies that slice segmentation information for a video picture be included in a syntax structure associated with a video unit including one or more slices. In some embodiments, the slice segmentation information for a video picture includes a value indicating the number of slices in the picture. In some embodiments, the slice segmentation information for a video picture includes a value indicating an index of a block unit within a slice. In some embodiments, the value indicating the index of the block unit within a slice is omitted from the slice segmentation information for the video picture. In some embodiments, the slice segmentation information for a video picture includes a value indicating the number of block units within a slice. In some embodiments, the value indicating the number of block units within a slice is omitted from the slice segmentation information for the video picture. In some embodiments, the block unit is a codec tree unit or a slice. In some embodiments, the slice segmentation information for a video picture is omitted from the slice header. In some embodiments, a slice index for a slice is included in the slice header, and a picture-level slice index for the slice is determined based on an address of the slice.

[0656] In some embodiments, the slice segmentation information for each of the one or more slices is organized in an implementation in a syntax structure associated with the video unit. In some embodiments, the slice segmentation information for each of the one or more slices is included in the syntax structure in ascending order. In some embodiments, the slice segmentation information for each of the one or more slices is included in the syntax structure in descending order.

[0657] In some embodiments, slice splitting information for at least one of the one or more slices is omitted from the bitstream. In some embodiments, slice splitting information for at least one of the one or more slices is inferred from slice splitting information for other slices included in the bitstream. In some embodiments, the number of block units of slice S-1 is not included in the bitstream, where S represents the number of slices in the video picture and S is greater than 1.

[0658] Figure 19 1 is a flowchart representation of a video processing method according to the present technology. The method 1900 includes, at operation 1910, performing conversion between a video picture of a video including one or more slices and a bitstream of the video according to a rule. The rule stipulates that the number of slices in the video picture is equal to or greater than a minimum number of slices determined based on whether rectangular or non-rectangular partitioning is applied to the video picture.

[0659] In some embodiments, when non-rectangular partitioning is applied, the minimum number of slices is two, while when rectangular partitioning is applied, the minimum number of slices is one. In some embodiments, when non-rectangular partitioning is applied, the minimum number of slices is one, while when rectangular partitioning is applied, the minimum number of slices is one. In some embodiments, the minimum number of slices is further determined based on the number of sub-pictures in the video picture.

[0660] Figure 20 2 is a flow chart representation of a video processing method according to the present technology. The method 2000 includes, at operation 2010, performing conversion between a video picture of a video and a bitstream of the video according to a rule. The video picture includes one or more slices. When slice segmentation information of the video picture is included in the syntax structure of the video unit, the slice is represented by the upper left position and dimensions of the slice.

[0661] In some embodiments, the upper left position and dimensions of a slice are indicated using the upper left position of a block unit within the slice, the dimensions of the block unit, and the dimensions of the slice measured using the dimensions of the block unit. In some embodiments, the upper left position of each slice and the dimensions of each slice are included in the syntax structure in sequence.

[0662] Figure 21 2 is a flow chart representation of a video processing method according to the present technology. Method 2100 includes, at operation 2110, performing conversion between a video picture of a video and a bitstream of the video according to a rule. A video picture includes one or more sub-pictures, and each sub-picture includes one or more slices. The rule specifies how segmentation information for the one or more slices in each sub-picture is present in the bitstream.

[0663] In some embodiments, the video unit includes a sequence parameter set, a picture parameter set, or a picture header. In some embodiments, the partitioning information of the one or more sub-pictures is arranged in ascending order based on a sub-picture index. In some embodiments, the partitioning information of the one or more slices in each sub-picture is arranged in ascending order based on a sub-picture-level slice index.

[0664] Figure 22 22 is a flowchart representation of a video processing method according to the present technology. The method 2200 includes, at operation 2210, performing conversion between a video picture of a video and a bitstream of the video according to a rule. The video picture includes one or more rectangular slices, and each slice includes one or more slices. The rule specifies omitting signaling of a difference between a first slice index of a first slice in the i-th rectangular slice and a second slice index of a first slice in the (i+1)th rectangular slice in the bitstream.

[0665] In some embodiments, the second tile index of the first tile in the (i+1)th rectangular strip is derived based on the rectangular strips from 0th to the (i)th index. In some embodiments, the second tile index of the first tile in the (i+1)th rectangular strip is derived as a minimum tile index that is outside the range defined by the rectangular strips from 0th to the (i)th index.

[0666] Figure 23 23 is a flowchart representing a video processing method according to the present technology. The method 2300 includes, at operation 2310, determining, for conversion between a video picture of a video and a bitstream of the video, conditionally including in the bitstream information for deriving the number of columns and the number of rows of slices in the video picture in response to a relationship between the dimensions of the video picture and the dimensions of a codec treeblock. The method 2300 also includes, at operation 2320, performing the conversion based on the determination.

[0667] In some embodiments, if the width of the video picture is less than or equal to the width of the codec treeblock, this information is omitted. In some embodiments, if the height of the video picture is less than or equal to the height of the codec treeblock, this information is omitted.

[0668] Figure 24 24 is a flowchart representation of a video processing method according to the present technology. Method 2400 includes, at operation 2410, performing conversion between a video picture of a video and a bitstream of the video. The video picture includes one or more sub-pictures. The bitstream conforms to a format rule that specifies that, if a variable specifying a sub-picture identifier for a sub-picture of a slice is present in the bitstream, there is one and only one syntax element that satisfies a condition that a second variable corresponding to the syntax element is equal to the variable.

[0669] In some embodiments, a variable is denoted as slice_subpic_id and a second variable is denoted as SubpicIdVal.The syntax element corresponding to this syntax element is denoted as SubpicIdVal[CurrSubpicIdx].

[0670] Figure 25 25 is a flowchart representation of a video processing method according to the present technology. The method 2500 includes, at operation 2510, performing conversion between a video picture of a video and a bitstream of the video. A video picture includes one or more sub-pictures. When non-rectangular partitioning is applied or sub-picture information is omitted in the bitstream, two slices in a slice have different addresses.

[0671] In some embodiments, a first syntax flag rect_slice_flag indicates whether non-rectangular partitioning is applied, and a second syntax flag subpic_info_present_info indicates whether sub-picture information is present in the bitstream.

[0672] Figure 26 26 is a flowchart representation of a video processing method according to the present technology. Method 2600 includes, at operation 2610, performing conversion between a video picture of a video and a bitstream of the video according to a rule. A video picture includes one or more slices. The rule specifies that when one or more slices are organized in both uniform and non-uniform spacing, a syntax element is used to indicate the type of slice layout.

[0673] In some embodiments, a syntax element included in a picture parameter set indicates whether uniform spacing is followed by non-uniform spacing or non-uniform spacing is followed by uniform spacing in a slice layout. In some embodiments, the slice layout includes non-uniformly spaced slices, wherein the number of explicitly indicated slice columns or slice rows is equal to or greater than the total number of non-uniformly spaced slices. In some embodiments, the slice layout includes uniformly spaced slices, and the number of explicitly indicated slice columns or slice rows is equal to or greater than the total number of uniformly spaced slices. In some embodiments, in the slice layout, uniform spacing is followed by non-uniform spacing, and the dimensions of the non-uniformly spaced slices are first determined according to a first reverse order, and then the dimensions of the uniformly spaced slices are subsequently determined according to a second reverse order. In some embodiments, the dimensions include the width of the slice columns or the height of the slice rows.

[0674] Figure 27 27 is a flowchart representation of a video processing method according to the present technology. The method 2700 includes, at operation 2710, performing conversion between a video picture of a video and a bitstream of the video according to a rule. The rule specifies whether or how to handle a Merge Estimation Region (MER) size in the conversion depending on a minimum allowed codec block size.

[0675] In some embodiments, the MER size is equal to or greater than the minimum allowed codec block size. In some embodiments, the MER size is indicated by MER_size, the minimum allowed codec block size is indicated by MinCbSizeY, and the difference between Log2(MER_size) and Log2(MinCbSizeY) is included in the bitstream.

[0676] Figure 2828 is a flowchart representation of a video processing method according to the present technology. The method 2800 includes, at operation 2810, performing conversion between a video including at least one video slice and a bitstream of the video according to a rule. The rule specifies deriving heights of slices in codec tree units for slices in the video slice based on a value of a first syntax element in the bitstream, the value of the first syntax element indicating a number of explicitly provided slice heights for slices in the video slice including the slice.

[0677] In some embodiments, in response to the value of the first syntax element being equal to 0, the height of a slice in a video slice in units of codec tree units is derived from information of the video. In some embodiments, the height of a slice in a video slice in units of codec tree units is represented as sliceHeightInCtus[i]. sliceHeightInCtus[i] is equal to the value of row heightRowHeight[SliceTopLeftTileIdx[i] / NumTileColumns], where SliceTopLeftTileIdx specifies the slice index of the slice that includes the first codec tree unit in the slice, NumTileColumns specifies the number of slice columns, and the value of row heightRowHeight[j] specifies the height of the j-th slice row in units of codec tree blocks.

[0678] In some embodiments, a height of a slice in codec tree units in a video slice is derived from information about the video in response to the height of the slice in the video slice not being present in the bitstream. In some embodiments, the height of the slice in codec tree units in the video slice is derived based on a second syntax element present in the bitstream indicating the height of the slice in the video slice. In some embodiments, in response to the first syntax element being equal to 0, the video slice including the slice is not divided into a plurality of slices.

[0679] Figure 29 29 is a flowchart representation of a video processing method according to the present technology. Method 2900 includes, at operation 2910, performing conversion between a video comprising a video picture and a bitstream of the video according to a rule, the video picture comprising a video slice comprising one or more slices. The rule provides that a second slice in a slice comprising a first slice in the picture has a height expressed in units of codec tree units. The first slice has a first slice index, and the second slice has a second slice index, the second slice index being determined based on the first slice index and a number of slice heights explicitly provided in the video slice. The height of the second slice is determined based on the first slice index and the second slice index.

[0680] In some embodiments, the first slice index is denoted as i and the second slice index is denoted as num_exp_slices_in_tile[i]-1, where num_exp_slices_in_tile specifies the number of slice heights explicitly provided in the video slice. The height of the second slice is determined based on exp_slice_height_in_ctus_minus1[i][num_exp_slices_in_tile[i]-1]+1, where exp_slice_height_in_ctus_minus1 specifies the slice height in units of codec tree units in the video slice. In some embodiments, the height of the uniform slices that divide the video slice is indicated by the last entry of exp_slice_height_in_ctus_minus1. In some embodiments, the second slice is always present in the picture. In some embodiments, resetting the height of the second slice for a conversion is not allowed. In some embodiments, the height of the third slice with an index greater than num_exp_slices_in_tile[i]-1 is determined based on exp_slice_height_in_ctus_minus1[i][num_exp_slices_in_tile[i]-1]. In some embodiments, the second slice is not present in the video slice. In some embodiments, the height of the second slice is less than or equal to exp_slice_height_in_ctus_minus1[i][num_exp_slices_in_tile[i]-1]+1.

[0681] Figure 30 3000 is a flowchart representation of a video processing method according to the present technology. The method 3000 includes, at operation 3010, performing conversion between a video comprising a video picture and a bitstream of the video, the video picture comprising one or more slices. The video picture refers to a picture parameter set. The picture parameter set conforms to a format rule that specifies that the picture parameter set comprises a list of column widths for N slice columns, where N is an integer. There is an N-1th slice column in the video picture, and the width of the N-1th slice column is equal to the N-1th entry in the list of explicitly included slice column widths plus one codec tree block.

[0682] In some embodiments, N-1 is represented as num_exp_tile_columns_minus1. In some embodiments, the width of a slice column in units of codec treeblocks is determined based on tile_column_width_minus1[num_exp_tile_columns_minus1]+1, where tile_column_width_minus1 specifies the width of a slice column in units of codec treeblocks. In some embodiments, the width of a slice column in units of codec treeblocks is not allowed to be reset and is determined only based on tile_column_width_minus1[num_exp_tile_columns_minus1]. In some embodiments, the width of a slice column in units of codec treeblocks is equal to tile_column_width_minus1[num_exp_tile_columns_minus1]+1. In some embodiments, the width of a second slice column having an index greater than num_exp_tile_columns_minus1 is determined based on tile_column_width_minus1[num_exp_tile_columns_minus1].

[0683] Figure 31 is a flowchart representation of a video processing method according to the present technology. The method 3100 includes, at operation 3110, performing conversion between a video including a video picture and a bitstream of the video, the video picture including one or more slices. The video picture refers to a picture parameter set. The picture parameter set conforms to a format rule that specifies that the picture parameter set includes a list of row heights of N slice rows, where N is an integer. There is an N-1th slice row in the video picture, and the height of the N-1th slice row is equal to the N-1th entry in the list of explicitly included slice row heights plus the number of codec tree blocks.

[0684] In some embodiments, N-1 is represented as num_exp_tile_rows_minus1. In some embodiments, the height of a slice row in units of codec tree blocks is determined based on tile_row_height_minus1[num_exp_tile_row_minus1]+1, where tile_row_height_minus1 specifies the height of a slice row in units of codec tree blocks. In some embodiments, the height of a slice row in units of codec tree blocks is not allowed to be reset and is determined only based on tile_row_height_minus1[num_exp_tile_row_minus1]. In some embodiments, the height of a slice row in units of codec tree blocks is equal to tile_row_height_minus1[num_exp_tile_row_minus1]+1. In some embodiments, the height of a second slice row having an index greater than num_exp_tile_rows_minus1 is determined based on tile_row_height_minus1[num_exp_tile_row_minus1].

[0685] In some embodiments, converting includes encoding the video into a bitstream. In some embodiments, converting includes decoding the video from the bitstream.

[0686] In the solution described herein, an encoder can conform to the format rules by generating a codec representation according to the format rules. In the solution described herein, a decoder can use the format rules to parse syntax elements in the codec representation and understand the presence and absence of syntax elements according to the format rules to produce decoded video.

[0687] In this document, the term "video processing" may refer to video encoding, video decoding, video compression, or video decompression. For example, during conversion from a pixel representation of a video to a corresponding bitstream, a video compression algorithm may be applied, or vice versa. The bitstream for a current video block may, for example, correspond to bits that are juxtaposed or interspersed at different locations within the bitstream, as defined by the syntax. For example, a macroblock may be encoded based on a transformed and encoded error residual, and also using bits from a header and other fields in the bitstream. Furthermore, during conversion, a decoder may parse the bitstream knowing that some fields may or may not be present, based on determinations as described in the above solution. Similarly, an encoder may determine whether to include or not include certain syntax fields, and generate the codec representation accordingly by including or excluding the syntax fields from the codec representation.

[0688] The disclosed solutions and other solutions, examples, embodiments, modules, and functional operations described in this document may be implemented in digital electronic circuitry, or in computer software, firmware, or hardware, including the structures disclosed in this document and their structural equivalents, or any combination thereof. The disclosed embodiments and other embodiments may be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer-readable medium for execution by or to control the operation of a data processing apparatus. The computer-readable medium may be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition of matter that effects a machine-readable propagated signal, or one or more combinations thereof. The term "data processing apparatus" encompasses all devices, equipment, and machines for processing data, including, for example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, the apparatus may also include code that creates an execution environment for the computer program in question, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, or any combination thereof. A propagated signal is an artificially generated signal, such as a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to a suitable receiver device.

[0689] A computer program (also referred to as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, and can also be deployed in any form, including as a standalone program or a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored in a file portion that holds other programs or data (e.g., one or more scripts stored in a markup language document), a single file dedicated to the relevant program, or multiple coordinated files (e.g., files that store one or more modules, subroutines, or code portions). A computer program can be deployed to execute on a single computer, or on multiple computers located at one site or distributed across multiple sites and interconnected by a communications network.

[0690] The processes and logic flows described in this document can be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by, and apparatus can also be implemented as, special purpose logic circuitry, such as an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit).

[0691] For example, processors suitable for executing computer programs include general-purpose and special-purpose microprocessors, as well as any one or more processors of any type of digital computer. Typically, a processor will receive instructions and data from read-only memory or random access memory, or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include or be operatively coupled to receive data from or transfer data to one or more mass storage devices for storing data (e.g., magnetic disks, magneto-optical disks, or optical disks). However, a computer does not require such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and storage devices, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and memory may be supplemented by or incorporated into dedicated logic circuitry.

[0692] Although this patent document contains many details, these details should not be interpreted as limitations on the scope of any subject matter or claimed content, but rather as descriptions of features unique to specific embodiments of specific technologies. Certain features described in this patent document in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, the various features described in the context of a single embodiment may also be implemented in multiple embodiments individually or in any suitable subcombination. In addition, although the above-mentioned features may be described as working in a particular combination, or even initially claimed to be so, in some cases, one or more features may be deleted from the claimed combination, and the claimed combination may refer to a subcombination or a variant of a subcombination.

[0693] Similarly, while operations are depicted in a particular order in the accompanying drawings, this should not be understood as requiring that such operations be performed in the particular order or sequential order shown, or that all illustrated operations be performed, in order to achieve desired results. Furthermore, the separation of various system components in the embodiments described in this patent document should not be understood as requiring such separation in all embodiments.

[0694] Only a few implementations and examples are described, and other implementations, enhancements, and variations can be made based on what is described and illustrated in this patent document.

Claims

1. A video processing method, comprising: performing conversion between a first video picture of a video and a bitstream of the video according to a first rule, wherein the first video picture includes one or more slices, wherein the first rule provides that, in response to at least one condition being satisfied, a syntax element indicating a difference between slice indices of two rectangular slices is present in the bitstream, wherein a first rectangular slice of the two rectangular slices is denoted as an i-th rectangular slice, where i is an integer; The method further comprises: performing conversion between a second video picture of the video and a bitstream of the video according to a second rule, wherein the second video picture includes one or more sub-pictures, and each sub-picture includes one or more rectangular slices, The second rule provides for deriving a slice index at the sub-picture level for each rectangular slice in each sub-picture to determine the number of codec tree units in each slice; wherein the slice index at the sub-picture level is determined based on a decoding order of corresponding slices in a slice list in the sub-picture; The slice index at the sub-picture level of the slice is determined based on the slice index at the picture level of the slice.

2. The method according to claim 1, wherein A second rectangular slice of the two rectangular slices is denoted as an (i+1)th rectangular slice, and the syntax element indicates a difference between a first slice index of a first slice containing a first codec tree unit in the (i+1)th rectangular slice and a second slice index of a second slice containing the first codec tree unit in the (i)th rectangular slice.

3. The method according to claim 1 or 2, wherein: The at least one condition that is satisfied includes i being less than (the number of rectangular strips in the first video picture - 1).

4. The method according to claim 1 or 2, wherein The at least one condition that is satisfied includes i not being equal to (the number of rectangular strips in the first video picture - 1).

5. A video processing method, comprising: performing conversion between a second video picture of the video and a bitstream of the video according to a second rule, wherein the second video picture includes one or more sub-pictures, and each sub-picture includes one or more rectangular slices, The second rule provides for deriving a slice index at the sub-picture level for each rectangular slice in each sub-picture to determine the number of codec tree units in each slice; wherein the slice index at the sub-picture level is determined based on a decoding order of corresponding slices in a slice list in the sub-picture; The slice index at the sub-picture level of the slice is determined based on the slice index at the picture level of the slice.

6. The method according to claim 5, wherein: In a case where a first slice index of a first slice at the sub-picture level is less than a second slice index of a second slice at the sub-picture level, the first slice is processed before the second slice according to the decoding order.

7. The method according to claim 5 or 6, wherein: The variable SubpicLevelSliceIdx is used to represent the slice index of the sub-picture level.

8. The method according to claim 5 or 6, wherein: A slice address is determined based on the slice index at the sub-picture level.

9. The method according to claim 5 or 6, wherein: The slice index of the sub-picture level of a slice is determined based on a first sub-picture including the slice.

10. The method according to claim 5 or 6, wherein: The slice index at the sub-picture level is a non-negative integer.

11. The method according to claim 10, wherein: The slice index at the sub-picture level is greater than or equal to 0 and less than N, where N is the number of slices in the sub-picture.

12. The method according to claim 5 or 6, wherein: In a case where a first slice and a second slice are different, a first slice index of the first slice at the sub-picture level is different from a second slice index of the second slice at the sub-picture level, wherein the first slice and the second slice are in the same sub-picture.

13. The method according to claim 5 or 6, wherein: When the first slice index of the sub-picture level of the first slice is smaller than the second slice index of the sub-picture level of the second slice, the first slice index of the picture level of the first slice is smaller than the second slice index of the picture level of the second slice.

14. A video processing method, comprising: performing conversion between a third video picture of the video and a bitstream of the video according to a third rule, The third video picture includes one or more rectangular strips, and each strip includes one or more slices, and The third rule specifies omitting signaling of a difference between a first slice index of a first slice in an i-th rectangular slice and a second slice index of a first slice in an (i+1)-th rectangular slice in the bitstream; The method further comprises: performing conversion between a second video picture of the video and a bitstream of the video according to a second rule, wherein the second video picture includes one or more sub-pictures, and each sub-picture includes one or more rectangular slices, The second rule provides for deriving a slice index at the sub-picture level for each rectangular slice in each sub-picture to determine the number of codec tree units in each slice; wherein the slice index at the sub-picture level is determined based on a decoding order of corresponding slices in a slice list in the sub-picture; The slice index at the sub-picture level of the slice is determined based on the slice index at the picture level of the slice.

15. The method according to claim 14, wherein The second slice index of the first slice in the (i+1)th rectangular slice is derived based on rectangular slices from the 0th to the i-th index.

16. The method according to claim 14, wherein The second tile index of the first tile in the (i+1)th rectangular strip is derived as a minimum tile index that is outside the range defined by rectangular strips from the 0th to the i-th index.

17. The method according to any one of claims 1-2, 5-6, 14-16, wherein The converting includes encoding the video into the bitstream.

18. The method according to any one of claims 1-2, 5-6, 14-16, wherein The converting includes decoding the video from the bitstream.

19. A method for storing a video bitstream, comprising: generating a bitstream of the video from a first video picture comprising one or more slices; as well as storing the bitstream in a non-transitory computer-readable recording medium, wherein the generating complies with a first rule, the first rule providing that a syntax element indicating a difference between slice indices of two rectangular slices is present in the bitstream responsive to at least one condition being satisfied, wherein a first rectangular slice of the two rectangular slices is denoted as an i-th rectangular slice, where i is an integer; The method further comprises: generating a bitstream for the video from a second video picture comprising one or more sub-pictures, wherein each sub-picture comprises one or more rectangular slices; and wherein the generating complies with a second rule, wherein the second rule provides for deriving a slice index at a sub-picture level for each rectangular slice in each sub-picture to determine the number of codec tree units in each slice; wherein the slice index at the sub-picture level is determined based on a decoding order of corresponding slices in a slice list in the sub-picture; The slice index at the sub-picture level of the slice is determined based on the slice index at the picture level of the slice.

20. A method for storing a video bitstream, comprising: generating a bitstream for the video from a third video picture comprising one or more rectangular slices, wherein each slice comprises one or more slices; and storing the bitstream in a non-transitory computer-readable recording medium, wherein the generating complies with a third rule, the third rule omitting signaling of a difference between a first slice index of a first slice in an i-th rectangular strip and a second slice index of a first slice in an (i+1)-th rectangular strip in the bitstream; The method further comprises: generating a bitstream for the video from a second video picture comprising one or more sub-pictures, wherein each sub-picture comprises one or more rectangular slices; and wherein the generating complies with a second rule, wherein the second rule provides for deriving a slice index at a sub-picture level for each rectangular slice in each sub-picture to determine the number of codec tree units in each slice; wherein the slice index at the sub-picture level is determined based on a decoding order of corresponding slices in a slice list in the sub-picture; The slice index at the sub-picture level of the slice is determined based on the slice index at the picture level of the slice.

21. A video coding and decoding apparatus comprising a processor configured to implement the method of one of claims 1 to 20.

22. A computer program product having computer code stored thereon, which, when executed by a processor, causes the processor to implement the method of any one of claims 1 to 20.

23. A non-transitory computer-readable storage medium storing instructions, the instructions causing a processor to implement the method of one of claims 1 to 20.

24. A non-transitory computer-readable recording medium storing a bit stream of a video generated by a method performed by a video processing apparatus, wherein: The method comprises: generating a bitstream of the video from a first video picture according to a first rule, wherein the first video picture includes one or more slices, wherein the first rule provides that, in response to at least one condition being satisfied, a syntax element indicating a difference between slice indices of two rectangular slices is present in the bitstream, wherein a first rectangular slice of the two rectangular slices is denoted as an i-th rectangular slice, where i is an integer; The method further comprises: generating a bitstream of the video from a second video picture according to a second rule, wherein the second video picture includes one or more sub-pictures, and each sub-picture includes one or more rectangular slices, The second rule provides for deriving a slice index at the sub-picture level for each rectangular slice in each sub-picture to determine the number of codec tree units in each slice; wherein the slice index at the sub-picture level is determined based on a decoding order of corresponding slices in a slice list in the sub-picture; The slice index at the sub-picture level of the slice is determined based on the slice index at the picture level of the slice.

25. A non-transitory computer-readable recording medium storing a bit stream of a video generated by a method performed by a video processing apparatus, wherein: The method comprises: generating a bitstream of the video from a third video picture according to a third rule, The third video picture includes one or more rectangular strips, and each strip includes one or more slices, and The third rule specifies omitting signaling of a difference between a first slice index of a first slice in an i-th rectangular slice and a second slice index of a first slice in an (i+1)-th rectangular slice in the bitstream; The method further comprises: generating a bitstream of the video from a second video picture according to a second rule, wherein the second video picture includes one or more sub-pictures, and each sub-picture includes one or more rectangular slices, The second rule provides for deriving a slice index at the sub-picture level for each rectangular slice in each sub-picture to determine the number of codec tree units in each slice; wherein the slice index at the sub-picture level is determined based on a decoding order of corresponding slices in a slice list in the sub-picture; The slice index at the sub-picture level of the slice is determined based on the slice index at the picture level of the slice.