Indication of slices in a video picture

By segmenting video images into sub-images, strips, and slices, and optimizing the video encoding and decoding process using codec tree units and format rules, the inefficiency problem in existing technologies is solved, achieving more efficient video transmission and parallel processing.

CN115552899BActive Publication Date: 2026-01-06DOUYIN VISION CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202180016194.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-02-21
Filing Date
2021-02-22
Publication Date
2026-01-06
Estimated Expiration
2041-02-22

AI Technical Summary

Technical Problem

Existing video encoding and decoding technologies suffer from inefficiency and inability to effectively match the maximum transmission unit size when processing the conversion between video images and bitstreams, resulting in excessive encoding and decoding overhead and difficulty in achieving parallel processing.

Method used

A novel video processing method is adopted, which segments video images into sub-images, strips, and slices, and then uses codec tree units and format rules for conversion, including the mapping of strip indexes and the determination of slice segmentation information. The codec process is optimized to adapt to different segmentation methods and codec tree block size requirements.

Benefits of technology

It improves the efficiency and flexibility of the video encoding and decoding process, supports better parallel processing and MTU size matching, reduces encoding and decoding overhead, and enhances video transmission performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115552899B_ABST
    Figure CN115552899B_ABST
Patent Text Reader

Abstract

Techniques for video processing including video coding, video decoding, and video transcoding are described. One example method includes performing a conversion between a video and a bitstream of the video, the video comprising a video picture comprising one or more slices. The video picture references a picture parameter set, and the picture parameter set conforms to a format rule, the format rule specifying that the picture parameter set comprises a list of column widths for N slice columns, where N is an integer. An (N-1)th slice column is present in the video picture, and a width of the (N-1)th slice column is equal to an (N-1)th entry in the explicitly included list of slice column widths plus a number of coding tree blocks.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference of related applications

[0002] Pursuant to applicable patent law and / or the rules of the Paris Convention, this application claims priority and interest in International Patent Application No. PCT / CN2020 / 076158, filed on February 21, 2020. For all legal purposes, the entire disclosure of the aforementioned application is incorporated herein by reference as a part of the disclosure. Technical Field

[0003] This patent document relates to image encoding and decoding as well as video encoding and decoding. Background Technology

[0004] Digital video accounts for the largest share of bandwidth usage on the Internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth demand for digital video is expected to continue to grow. Summary of the Invention

[0005] This document discloses techniques that can be used by video encoders and decoders to process the encoded and decoded representation of video using control information useful for decoding the encoded and decoded representation.

[0006] In one example aspect, a video processing method is disclosed. The method includes performing a conversion between a video image and a bitstream of the video according to rules. The video image comprises one or more stripes, and in response to satisfying at least one condition, the rule specifies that a syntax element indicating the difference between slice indices of two rectangular stripes exists in the bitstream. The first rectangular strip of the two rectangular stripes is represented as the i-th rectangular stripe, where i is an integer.

[0007] In another example, a video processing method is disclosed. This method includes performing a conversion between video frames and a bitstream of a video according to rules. A video frame comprises one or more sub-frames, and each sub-frame comprises one or more rectangular stripes. The rules specify deriving the stripe index at the sub-frame level for each rectangular stripe in each sub-frame to determine the number of codec tree units in each stripe.

[0008] In another example, a video processing method is disclosed. The method includes a conversion between a video image comprising one or more sub-images and a bitstream of the video, determining a mapping between a sub-image-level stripe index of a stripe in a sub-image and a picture-level stripe index of that stripe. The method also includes performing the conversion based on this determination.

[0009] In another example, a video processing method is disclosed. This method includes performing a conversion between video frames and a bitstream of the video according to a rule. The video frame comprises one or more sub-frames, and the rule stipulates that a slice of the video is entirely contained within a single sub-frame of the video frame.

[0010] In another example, a video processing method is disclosed. This method includes performing a conversion between video images and a bitstream of the video, wherein the video images comprise one or more sub-images. The bitstream conforms to a format rule that specifies that information about the segmented images is included in a syntactic structure associated with the images.

[0011] In another example, a video processing method is disclosed. This method includes converting a video image comprising one or more stripes having a non-rectangular shape to a bitstream of the video, determining stripe segmentation information of the video image. The method also includes performing the conversion based on this determination.

[0012] In another example, a video processing method is disclosed. This method includes performing a conversion between a video image comprising one or more stripes and a bitstream of the video, according to a rule. The rule specifies that the number of stripes in the video image is equal to or greater than a minimum number of stripes determined based on whether rectangular or non-rectangular segmentation is applied to the video image.

[0013] In another example, a video processing method is disclosed. This method includes performing a conversion between video images and a bitstream of the video according to rules. The video images comprise one or more stripes. Where the stripe segmentation information of the video images is included in the grammatical structure of the video unit, the stripes are represented by their top-left position and dimension.

[0014] In another example, a video processing method is disclosed. This method includes performing a conversion between video frames and a video bitstream according to rules. A video frame comprises one or more sub-frames, and each sub-frame comprises one or more stripes. The rules specify how the segmentation information of the one or more stripes in each sub-frame is present in the bitstream.

[0015] In another example, a video processing method is disclosed. This method includes performing a conversion between a video image and a video bitstream according to a rule. The video image comprises one or more rectangular strips, and each strip comprises one or more slices. The rule specifies that signaling notifications regarding the difference between the first slice index of the first slice in the i-th rectangular strip and the second slice index of the first slice in the (i+1)-th rectangular strip are omitted from the bitstream.

[0016] In another example, a video processing method is disclosed. The method includes: converting between video frames and a bitstream of the video; and, in response to a relationship between the dimensions of the video frames and the dimensions of codec tree blocks, determining that information regarding the number of columns and rows of slices in the derived video frames is conditionally included in the bitstream. The method also includes performing the conversion based on this determination.

[0017] In another example, a video processing method is disclosed. The method includes performing a conversion between video images and a bitstream of the video. The video images include one or more sub-images. The bitstream conforms to a format rule that states that, where a variable specifying a sub-image identifier including striped sub-images exists in the bitstream, there exists one and only one syntax element that satisfies the condition that a second variable corresponding to that syntax element is equal to that variable.

[0018] In another example, a video processing method is disclosed. This method includes performing a conversion between video frames and a video bitstream. The video frames comprise one or more sub-frames. When non-rectangular segmentation is applied or sub-frame information is omitted from the bitstream, two slices within a stripe have different addresses.

[0019] In another example, a video processing method is disclosed. This method includes performing a conversion between video frames and a video bitstream according to rules. The video frames comprise one or more slices. The rules specify that, in cases where one or more slices are organized with both uniform and non-uniform intervals, syntax elements are used to indicate the type of slice layout.

[0020] In another example, a video processing method is disclosed. This method includes performing a conversion between video frames and a video bitstream according to rules. These rules specify whether and how the Merge Estimation Region (MER) size is handled during the conversion depends on the minimum allowed codec block size.

[0021] In another example, a video processing method is disclosed. The method includes performing a conversion between a video and a video bitstream comprising at least one video slice, according to a rule. The rule specifies the height of a stripe in the video slice, in units of a codec tree unit, based on the value of a first syntax element in the bitstream, the value of which indicates the number of explicitly provided stripe heights in the video slice including that stripe.

[0022] In another example, a video processing method is disclosed. This method includes performing a conversion between a video comprising a video image and a video bitstream, according to a rule. The video image comprises a video slice containing one or more stripes. The rule specifies that a second stripe in a slice comprising a first stripe from the image has a height represented in units of a codec tree unit. The first stripe has a first stripe index, and the second stripe has a second stripe index, which is determined based on the first stripe index and the number of stripe heights explicitly provided in the video slice. The height of the second stripe is determined based on the first and second stripe indices.

[0023] In another example, a video processing method is disclosed. This method includes performing a conversion between a video and a video bitstream, the video comprising a video picture containing one or more slices. The video picture references a picture parameter set, and the picture parameter set conforms to a format rule specifying that the picture parameter set includes a list of column widths of N slice columns, where N is an integer. The video picture contains an (N-1)th slice column, and the width of the (N-1)th slice column is equal to the (N-1)th entry in the explicitly included list of slice column widths plus a codec tree block.

[0024] In another example, a video processing method is disclosed. This method includes performing a conversion between a video and a video bitstream, the video comprising a video picture containing one or more slices. The video picture references a picture parameter set, and the picture parameter set conforms to a format rule specifying that the picture parameter set includes a list of row heights of N slice rows, where N is an integer. There exists an (N-1)th slice row in the video picture, and the height of the (N-1)th slice row is equal to the (N-1)th entry in the explicitly included list of slice row heights plus a codec tree block.

[0025] In another example, a video processing method is disclosed. The method includes performing a conversion between a video comprising one or more video images and a codec representation of the video, wherein each video image comprises one or more sub-images, each sub-image comprising one or more stripes, wherein the codec representation conforms to a format rule; wherein the format rule specifies that, when a rectangular stripe mode is enabled for the video image, a picture-level stripe index for each stripe in each sub-image of the video image is derived without explicit signaling notification in the codec representation; wherein the format rule specifies that the number of codec tree units in each stripe can be derived from the picture-level stripe index.

[0026] In another example, a different video processing method is disclosed. This method includes performing a conversion between a video comprising one or more video images and a codec representation of the video, wherein each video image comprises one or more sub-images, each sub-image comprising one or more stripes, wherein the codec representation conforms to a format rule; wherein the format rule specifies that sub-image level stripe indices can be derived based on information in the codec representation without signaling notification of the sub-image level stripe indices in the codec representation.

[0027] In another example, a different video processing method is disclosed. This method includes performing a conversion between a video comprising one or more video images and a codec representation of the video, wherein each video image comprises one or more sub-images and / or one or more slices, wherein the codec representation conforms to format rules, and wherein the conversion conforms to constraint rules.

[0028] In another example, a different video processing method is disclosed. This method includes performing a conversion between a video comprising one or more video images and a codec representation of the video, wherein each video image comprises one or more slices and / or one or more stripes; wherein the codec representation conforms to a format rule; and wherein the format rule specifies that video image-level fields carry information about the segmentation of stripes and / or slices in the video image.

[0029] In another example, a different video processing method is disclosed. This method includes performing a conversion between a video comprising one or more images and a codec representation of the video, wherein the conversion conforms to a segmentation rule that the minimum number of strips into which the video images are segmented depends on whether rectangular segmentation is used to segment the video images.

[0030] In another example, a different video processing method is disclosed. This method includes performing a conversion between video stripes of a video region and a codec representation of the video; wherein the codec representation conforms to a format rule; wherein the format rule specifies that the codec representation signals the video stripe based on the upper-left position of the video stripe, and wherein the format rule specifies that the codec representation signals the height and / or width of the video stripe in segmentation information, which is signaled at the video unit level.

[0031] In another example, a different video processing method is disclosed. This method includes performing a conversion between a video comprising video images and a codec representation of the video; wherein the codec representation conforms to a format rule; wherein the format rule specifies omitting the difference between the slice index of the first slice in a signaling notification rectangular strip and the slice index of the first slice in a next rectangular strip.

[0032] In another example, a different video processing method is disclosed. This method includes performing a conversion between a video and its codec representation, wherein the codec representation conforms to a format rule specifying the relationship between the width of a video frame and the size of a codec tree unit, and controlling signaling notifications for deriving information about the number of slice columns or slice rows in the video frame.

[0033] In another example, a different video processing method is disclosed. This method includes performing a conversion between a video comprising one or more video images and a video codec representation, wherein the codec representation conforms to a format rule specifying that slice layout information be included in the codec representation of the video images comprising uniformly spaced slices and non-uniformly spaced slices.

[0034] In yet another example, a video encoder apparatus is disclosed. The video encoder includes a processor configured to implement the methods described above.

[0035] In yet another example, a video decoder apparatus is disclosed. The video decoder includes a processor configured to implement the methods described above.

[0036] In yet another example, a computer-readable medium on which code is stored is disclosed. This code embodies one of the methods described herein in the form of processor-executable code.

[0037] These and other features are described in this document. Attached Figure Description

[0038] Figure 1 An example of raster scan strip segmentation of an image is shown, where the image is divided into 12 slices and 3 raster scan strips.

[0039] Figure 2 An example of rectangular strip segmentation of an image is shown, where the image is divided into 24 slices (6 slice columns and 4 slice rows) and 9 rectangular strips.

[0040] Figure 3 An example is shown where an image is divided into slices and rectangular strips, where the image is divided into 4 slices (2 slice columns and 2 slice rows) and 4 rectangular strips.

[0041] Figure 4 The image is shown to be divided into 18 slices, 24 strips, and 24 sub-images.

[0042] Figure 5 The nominal vertical and nominal horizontal positions of the 4:2:2 luminance and chrominance samples in the image are shown.

[0043] Figure 6An example of image segmentation is shown. Solid line 602 represents the boundary of a slice; dashed line 604 represents the boundary of a strip; and dashed line 606 represents the boundary of a sub-image. The figure indicates the image-level index, decoding order index, sub-image-level index, and the indices of the sub-images and slices for the four strips.

[0044] Figure 7 This is a block diagram of an example video processing system.

[0045] Figure 8 This is a block diagram of a video processing device.

[0046] Figure 9 This is a flowchart of an example method for video processing.

[0047] Figure 10 This is a block diagram illustrating a video encoding / decoding system according to some embodiments of the present disclosure.

[0048] Figure 11 This is a block diagram illustrating an encoder according to some embodiments of the present disclosure.

[0049] Figure 12 This is a block diagram illustrating a decoder according to some embodiments of the present disclosure.

[0050] Figure 13 This is a flowchart representation of a method for video processing according to the present technology.

[0051] Figure 14 This is a flowchart representation of another method for video processing according to the present technology.

[0052] Figure 15 This is a flowchart representation of another method for video processing according to the present technology.

[0053] Figure 16 This is a flowchart representation of another method for video processing according to the present technology.

[0054] Figure 17 This is a flowchart representation of another method for video processing according to the present technology.

[0055] Figure 18 This is a flowchart representation of another method for video processing according to the present technology.

[0056] Figure 19 This is a flowchart representation of another method for video processing according to the present technology.

[0057] Figure 20 This is a flowchart representation of another method for video processing according to the present technology.

[0058] Figure 21This is a flowchart representation of another method for video processing according to the present technology.

[0059] Figure 22 This is a flowchart representation of another method for video processing according to the present technology.

[0060] Figure 23 This is a flowchart representation of another method for video processing according to the present technology.

[0061] Figure 24 This is a flowchart representation of another method for video processing according to the present technology.

[0062] Figure 25 This is a flowchart representation of another method for video processing according to the present technology.

[0063] Figure 26 This is a flowchart representation of another method for video processing according to the present technology.

[0064] Figure 27 This is a flowchart representation of another method for video processing according to the present technology.

[0065] Figure 28 This is a flowchart representation of another method for video processing according to the present technology.

[0066] Figure 29 This is a flowchart representation of another method for video processing according to the present technology.

[0067] Figure 30 This is a flowchart representation of another method for video processing according to the present technology.

[0068] Figure 31 This is a flowchart representation of another method for video processing according to the present technology. Detailed Implementation

[0069] The use of section headings in this document is for ease of understanding and does not limit the application of the technologies and embodiments disclosed in each section to that section only. Furthermore, the use of H.266 terminology in some specifications is merely for ease of understanding and not to limit the scope of the disclosed technologies. Thus, the technologies described herein are also applicable to other video codec protocols and designs.

[0070] 1. Overview

[0071] This document relates to video codec technology. Specifically, it concerns signaling for subpictures, slices, and stripes. These ideas can be applied individually or in various combinations to any video codec standard or non-standard video codec that supports multi-layer video codecs, such as the Universal Video Codec (VVC) currently under development.

[0072] 2. Abbreviation

[0073] APS Adaptive Parameter Set

[0074] AU Access Unit

[0075] AUD Access Unit Separator

[0076] AVC Advanced Video Codec

[0077] CLVS codec layer video sequence

[0078] CPB image buffer

[0079] CRA Clean Random Access

[0080] CTU (Codec Tree Unit)

[0081] CVS codec video sequence

[0082] DPB Decoding Image Buffer

[0083] DPS Decoding Parameter Set

[0084] EOB End of Bitstream

[0085] End of EOS sequence

[0086] GDR Gradual Decoding and Refresh

[0087] HEVC High-Efficiency Video Encoding and Decoding

[0088] HRD Assumption Reference Decoder

[0089] IDR Instant Decoding and Refresh

[0090] JEM Joint Exploration Model

[0091] MCTS motion-constrained plate group

[0092] NAL Network Abstraction Layer

[0093] OLS Output Layer Set

[0094] PH image header

[0095] PPS Image Parameter Set

[0096] PTL levels, tiers, and grades

[0097] PU Image Unit

[0098] RBSP raw byte sequence payload

[0099] SEI Supplemental Enhancement Information

[0100] SPS Sequence Parameter Set

[0101] SVC Scalable Video Codec

[0102] VCL (Video Codec Layer)

[0103] VPS Video Parameter Set

[0104] VTM VVC Test Model

[0105] VUI Video Availability Information

[0106] VVC Multi-Functional Video Encoding and Decoding

[0107] 3. Preliminary Discussion

[0108] Video codec standards have primarily evolved through the development of well-known ITU-T and ISO / IEC standards. ITU-T developed H.261 and H.263, while ISO / IEC developed MPEG-1 and MPEG-4 Visual. These two organizations jointly developed the H.262 / MPEG-2 video standard, the H.264 / MPEG-4 Advanced Video Coding (AVC) standard, and the H.265 / HEVC standard. Since H.262, video codec standards have been based on a hybrid video codec architecture, employing temporal prediction plus transform coding. To explore future video codec technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Team (JVET) in 2015. Since then, many new methods have been adopted by JVET and applied to reference software called the Joint Exploration Model (JEM). JVET meetings are held quarterly, and the goal of new codec standards is to reduce the bitrate by 50% compared to HEVC. At the JVET meeting in April 2018, the new video codec standard was officially named Versatile Video Coding (VVC), and the first version of the VVC Test Model (VTM) was released at that time. With ongoing efforts to standardize VVC, each JVET meeting adopts a new codec technology for the VVC standard. The VVC working draft and test model VTM are then updated after each meeting. The current goal of the VVC project is to achieve Technical Completion (FDIS) at the July 2020 meeting.

[0109] 3.1. Image Segmentation Schemes in HEVC

[0110] HEVC includes four different image segmentation schemes: regular striping, dependent striping, slice, and wavefront parallel processing (WPP), which can be used for maximum transfer unit (MTU) size matching, parallel processing, and reduced end-to-end latency.

[0111] Regular slices are similar to those in H.264 / AVC. Each regular slice is encapsulated in its own NAL unit, and intra-picture prediction (intra-sample prediction, motion information prediction, and encoding / decoding mode prediction) and entropy encoding / decoding dependencies across slice boundaries are disabled. Therefore, regular slices can be reconstructed independently of other regular slices within the same picture (although interdependencies may still exist due to loop filtering operations).

[0112] Regular striping is the only tool available for parallelization, and it is available in almost the same form in H.264 / AVC. Parallelization based on regular striping requires minimal inter-processor or inter-core communication (except for inter-processor or inter-core data sharing for motion compensation when decoding predictive encoded / decoded images, which is generally far more important than inter-processor or inter-core data sharing due to intra-image prediction). However, for the same reason, using regular striping results in significant encoding / decoding overhead due to the bit cost of the stripe header and the lack of prediction across stripe boundaries. Furthermore, due to the intra-image independence of regular striping and the fact that each regular stripe is encapsulated in its own NAL unit, regular striping (compared to other tools mentioned below) also serves as a key mechanism for bitstream segmentation to match MTU size requirements. In many cases, the goals of parallelization and MTU size matching present conflicting requirements for stripe layout in the image. This recognition led to the development of the parallelization tools mentioned below.

[0113] Dependency striping has a short header and allows the bitstream to be split at tree block boundaries without disrupting any in-picture predictions. Essentially, dependency striping provides the option to divide a regular stripe into multiple NAL units to reduce end-to-end latency by allowing a portion of the regular stripe to be sent before the entire regular stripe's encoding is complete.

[0114] In WPP, an image is segmented into individual codec tree block (CTB) rows. Entropy decoding and prediction are allowed to use data from CTBs in other segments. Parallel processing is possible through parallel decoding of CTB rows, where the start of decoding a CTB row is delayed by two CTBs, ensuring that data associated with the CTBs above and to the right of the object CTB is available before the object CTB is decoded. Using this staggered start (which looks like a wavefront when graphically represented), parallelization can use as many processors / cores as the number of CTB rows contained in the image. Because intra-image prediction between adjacent tree block rows within an image is allowed, the inter-processor / inter-core communication required to implement intra-image prediction can be substantial. WPP segmentation does not result in the generation of additional NAL units compared to when it is not applied, therefore WPP is not a tool for MTU size matching. However, if MTU size matching is required, regular striping can be used with WPP, but with some encoding / decoding overhead.

[0115] A slice defines the horizontal and vertical boundaries that divide an image into slice columns and slice rows. Slice columns extend from the top to the bottom of the image. Similarly, slice rows extend from the left to the right of the image. The number of slices in an image can be simply derived by multiplying the number of slice columns by the number of slice rows.

[0116] Before decoding the top-left CTB of the next slice in the order of slice raster scans of the image, the scan order of the CTBs is changed to be local within the slice (in the order of slice CTB raster scans). Similar to regular stripes, slices break the intra-image prediction dependency and the entropy decoding dependency. However, they do not need to be included in a single NAL unit (the same as WPP in this respect); therefore, slices cannot be used for MTU size matching. Each slice can be processed by one processor / core, and the inter-processor / inter-core communication required for intra-image prediction between processing units decoding adjacent slices is limited to transmitting a shared stripe header when a stripe spans more than one slice, and sharing related to cyclic filtering for reconstructing samples and metadata. When a stripe includes more than one slice or WPP segment, the entry point byte offset of each slice or WPP segment in the stripe, except for the first slice, is signaled in the stripe header.

[0117] For simplicity, HEVC specifies limitations on the application of four different image segmentation schemes. A given codec video sequence cannot include most slices and wavefronts specified in the HEVC standard. For each strip and slice, one or both of the following conditions must be met: 1) All codec tree blocks in a strip belong to the same slice; 2) All codec tree blocks in a slice belong to the same strip. Finally, a wavefront segment contains exactly one CTB line, and when using WPP, if a strip starts at a CTB line, it must end at the same CTB line.

[0118] The latest revisions to HEVC are specified in the JCT-VC output document JCTVC-AC1005, published on October 24, 2017, by J. Boyce, A. Ramasubramonian, R. Skupin, G. J. Sullivan, A. Tourapis, and Y.-K. Wang (editors): "HEVC Additional Supplemental Enhancement Information (Draft 4)": http: / / phenix.int-evry.fr / jct / doc_end_user / documents / 29_Macau / wg11 / JCTVC-AC1005-v2.zip. With this revision, HEVC defines three types of MCTS-related SEI (Supplemental Enhancement Information) messages: i.e., domain MCTS SEI messages, MCTS extracted information set SEI messages, and MCTS extracted information nested SEI messages.

[0119] The temporal MCTS SEI message indicates the presence of an MCTS in the bitstream, and signaling notifies the MCTS. For each MCTS, motion vectors are restricted to pointing to full-sample locations within the MCTS and fractional-sample locations that require interpolation only from full-sample locations within the MCTS, and motion vector candidates derived from blocks outside the MCTS for temporal motion vector prediction are not allowed. In this way, each MCTS can be decoded independently, and there are no slices not included in the MCTS.

[0120] The MCTS Extraction Information Set (SEI) message provides supplementary information (defined as part of the semantics of the SEI message) that can be used in MCTS sub-bitstream extraction to generate a bitstream conforming to the MCTS set. This information consists of multiple extraction information sets, each defining multiple MCTS sets and containing RBSP bytes that will replace the VPS, SPS, and PPS during the MCTS sub-bitstream extraction process. When extracting a sub-bitstream according to the MCTS sub-bitstream extraction process, the parameter sets (VPS, SPS, and PPS) need to be rewritten or replaced because one or all slice address-related syntax elements (including first_slice_segment_in_pic_flag and slice_segment_address) typically need to have different values, so the slice header needs to be slightly updated.

[0121] 3.2. Image Segmentation in VVC

[0122] In VVC, an image is divided into one or more slice rows and one or more slice columns. A slice is a sequence of CTUs covering a rectangular area of ​​the image. The CTUs in a slice are scanned in raster scan order within that slice.

[0123] A strip consists of an integer number of complete slices or images, and an integer number of consecutive complete CTU lines within a slice.

[0124] Two stripe modes are supported: raster scan stripe mode and rectangular stripe mode. In raster scan stripe mode, a stripe contains a sequence of complete slices in a sheet raster scan of an image. In rectangular stripe mode, a stripe contains multiple complete slices that together form a rectangular area of ​​the image, or multiple consecutive complete CTU rows of a single slice that together form a rectangular area of ​​the image. Slices within a rectangular stripe are scanned in sheet raster scan order within the rectangular area corresponding to that stripe.

[0125] A sub-image contains one or more stripes that collectively cover a rectangular area of ​​the image.

[0126] Figure 1 An example of raster scan strip segmentation of an image is shown, where the image is divided into 12 slices and 3 raster scan strips.

[0127] Figure 2 An example of rectangular strip segmentation of an image is shown, where the image is divided into 24 slices (6 slice columns and 4 slice rows) and 9 rectangular strips.

[0128] Figure 3 An example is shown where an image is divided into slices and rectangular strips, where the image is divided into 4 slices (2 slice columns and 2 slice rows) and 4 rectangular strips.

[0129] Figure 4 An example of sub-image partitioning of an image is shown, where the image is divided into 18 slices. The 12 slices on the left each cover a 4×4 CTU strip, and the 6 slices on the right each cover two vertically stacked 2×2 CTU strips, resulting in a total of 24 strips and 24 sub-images of different dimensions (each strip is a sub-image).

[0130] 3.3. Signaling of SPS / PPS / Image Header / Strip Header in VVC

[0131] 7.3.2.3 Sequence Parameter Set (RBSP) Syntax

[0132]

[0133]

[0134]

[0135]

[0136]

[0137]

[0138]

[0139] 7.3.2.4 Image Parameter Set RBSP Syntax

[0140]

[0141]

[0142]

[0143]

[0144] 7.3.2.7 Image Header Structure Syntax

[0145]

[0146]

[0147]

[0148]

[0149]

[0150] 7.3.7.1 General Strip Header Syntax

[0151]

[0152]

[0153]

[0154]

[0155] 3.4. Example Specifications for Pieces, Strips, and Sub-images

[0156] 3 Definitions

[0157] Image-level slice index: When rect_slice_flag equals 1, the slice index of the list of slices in the image (in the order they are signaled in PPS).

[0158] Sub-image level slice index: When rect_slice_flag equals 1, the slice index of the list of slices in the sub-image (in the order in which they are signaled in PPS).

[0159] 6.5.1 CTB raster scanning, slice scanning, and sub-image scanning process

[0160] The derivation of the variable NumTileColumns, which specifies the number of slice columns, and the list colWidth[i], which specifies the width of the i-th slice column in CTB units (where i ranges from 0 to NumTileColumn-1, inclusive), is as follows:

[0161]

[0162]

[0163] The derivation of the variable numtierrows, which specifies the number of slice rows, and the list RowHeight[j], which specifies the height of the j-th slice row in CTB units (where j ranges from 0 to NumTileRows-1, inclusive), is as follows:

[0164]

[0165] The variable NumTilesInPic is set to be equal to NumTileColumns * NumTileRows.

[0166] The derivation of the list tileColBd[i] (where i ranges from 0 to NumTileColumns, including end values) specifying the positions of the i-th tile column boundary in CTB units is as follows:

[0167] for(tileColBd[0]=0,i=0;i <NumTileColumns;i++)

[0168] tileColBd[i+1]=tileColBd[i]+colWidth[i] (25)

[0169] Note 1 – The size of the array tileColBd[] is 1 greater than the actual number of slices in the derivation of CtbToTileColBd[].

[0170] The derivation of the list of tileRowBd[j] (where j ranges from 0 to NumTileRows, including end values) specifying the position of the j-th tile row boundary in CTB units is as follows:

[0171] for(tileRowBd[0]=0,j=0;j <NumTileRows;j++)

[0172] tileRowBd[j+1]=tileRowBd[j]+RowHeight[j] (26)

[0173] Note 2 – The size of the array tileRowBd[] in the above derivation is 1 greater than the actual number of slice rows in the derivation of CtbToTileRowBd[].

[0174] The derivation of the list CtbToTileColBd[ctbAddrX] (where ctbAddrX ranges from 0 to PicWidthInCtbsY, inclusive) specifying the conversion from a horizontal CTB address to the left slice column boundary in CTB units is as follows:

[0175]

[0176] Note 3 – The size of the array CtbToTileColBd[] in the above derivation is 1 greater than the actual image width in the CTB of the slice_data() signaling in the derivation.

[0177] The derivation of the list CtbToTileRowBd[ctbAddrY] (where ctbAddrY ranges from 0 to PicHeightInCtbsY, including end values) specifying the conversion from the vertical CTB address to the top slice row boundary in CTB units is as follows:

[0178]

[0179] Note 4 – The size of the array CtbToTileRowBd[] in the above derivation is 1 greater than the actual number of image heights in the CTB of the slice_data() signaling.

[0180] For rectangular stripes, the derivation of the list NumCtusInSlice[i] (where i ranges from 0 to num_slices_in_pic_minus1, inclusive), the list SliceTopLeftTileIdx[i] (where i ranges from 0 to num_slices_in_pic_minus1, inclusive), the matrix CtbAddrInSlice[i][j] (where i ranges from 0 to num_slices_in_pic_minus1, inclusive), and the matrix CtbAddrInSlice[i][j] (where i ranges from 0 to num_slices_in_pic_minus1, inclusive) and j ranges from 0 to NumCtusInSlice[i]-1, inclusive) for the number of CTUs in the i-th stripe is as follows:

[0181]

[0182]

[0183]

[0184] The function AddCtbsToSlice(sliceIdx,startX,stopX,startY,stopY) is defined as follows:

[0185]

[0186] One requirement for bitstream consistency is that the value of NumCtusInSlice[i] (where i ranges from 0 to num_slices_in_pic_minus1 (inclusive)) should be greater than 0. Furthermore, another requirement for bitstream consistency is that the matrix CtbAddrInSlice[i][j] (where i ranges from 0 to num_slices_in_pic_minus1 (inclusive) and j ranges from 0 to NumCtusInSlice[i]-1 (inclusive)) should include all CTB addresses ranging from 0 to PicSizeInCtbsY-1 exactly once.

[0187] The derivation of the list CtbToSubpicIdx[ctbAddrRs] (where tbAddrRs ranges from 0 to PicSizeInCtbsY-1, inclusive) specifying the conversion from CTB addresses in image raster scans to subpic indices is as follows:

[0188]

[0189]

[0190] The derivation of the list NumSlicesInSubpic[i], which specifies the number of rectangular stripes in the i-th subpic, is as follows:

[0191]

[0192] 7.3.4.3 Image Parameter Set RBSP Semantics

[0193] `subpic_id_mapping_in_pps_flag` equal to 1 specifies that signaling notification of the subpicture ID mapping is performed in the PPS. `subpic_id_mapping_in_pps_flag` equal to 0 specifies that signaling notification of the subpicture ID mapping is not performed in the PPS. If `subpic_id_mapping_explicitly_signaled_flag` is 0 or `subpic_id_mapping_in_sps_flag` is 1, then the value of `subpic_id_mapping_in_pps_flag` should be 0. Otherwise (if `subpic_id_mapping_explicitly_signaled_flag` is 1 and `subpic_id_mapping_in_sps_flag` is 0), the value of `subpic_id_mapping_in_pps_flag` should be 1.

[0194] pps_num_subpics_minus1 should be equal to sps_num_subpics_minus1.

[0195] pps_subpic_id_len_minus1 should be equal to sps_subpic_id_len_minus1.

[0196] pps_subpic_id[i] specifies the subpick ID of the i-th subpick. The length of the pps_subpic_id[i] syntax element is pps_subpic_id_len_minus1+1 bits.

[0197] For each value of i in the range from 0 to sps_num_subpics_minus1 (inclusive), the derivation of the variable SubpicIdVal[i] is as follows:

[0198] for(i=0;i<=sps_num_subpics_minus1;i++)

[0199] if(subpic_id_mapping_explicitly_signalled_flag)

[0200] SubpicIdVal[i]=subpic_id_mapping_in_pps_flag? pps_subpic_id[i]:sps_subpic_id[i](80)

[0201] else

[0202] SubpicIdVal[i]=i

[0203] One requirement for bitstream consistency is the application of the following two constraints:

[0204] – For any two distinct values ​​of i and j in the range from 0 to sps_num_subpics_minus1 (inclusive), SubpicIdVal[i] should not be equal to SubpicIdVal[j].

[0205] – When the current image is not the first image of CLVS, for each i value in the range of 0 to sps_num_subpics_minus1 (inclusive of end values), if the value of SubpicIdVal[i] is not equal to the value of SubpicIdVal[i] of the previous image in the same layer in the decoding order, then the nal_unit_type of all codec strip NAL units of the subpics in the current image with subpic index i should be equal to a specific value in the range of IDR_W_RADL to CRA_NUT (inclusive of end values).

[0206] `no_pic_partition_flag` equal to 1 indicates that no image partitioning is applied to each image in the reference PPS. `no_pic_partition_flag` equal to 0 indicates that each image in the reference PPS can be divided into more than one slice or strip.

[0207] One requirement for bitstream consistency is that the value of no_pic_partition_flag should be the same for all PPS referenced by the CLVS encoding / decoding images.

[0208] One requirement for bitstream consistency is that when the value of sps_num_subpics_minus1+1 is greater than 1, the value of no_pic_partition_flag should not be equal to 1.

[0209] The value pps_log2_ctu_size_minus5 plus 5 specifies the luma codec tree block size for each CTU. pps_log2_ctu_size_minus5 should be equal to sps_log2_ctu_size_minus5.

[0210] The increment of 1 in `num_exp_tile_columns_minus1` specifies the number of tile column widths explicitly provided. The value of `num_exp_tile_columns_minus1` should be in the range of 0 to `PicWidthInCtbsY-1` (inclusive). When `no_pic_partition_flag` equals 1, the value of `num_exp_tile_columns_minus1` is inferred to be 0.

[0211] The increment of 1 in `num_exp_tile_rows_minus1` specifies the number of tile row heights explicitly provided. The value of `num_exp_tile_rows_minus1` should be in the range of 0 to `PicHeightInCtbsY-1` (inclusive). When `no_pic_partition_flag` is equal to 1, the value of `num_tile_rows_minus1` is inferred to be equal to 0.

[0212] The increment of `tile_column_width_minus1[i]` by 1 specifies the width of the i-th tile column in CTB units, where i ranges from 0 to `num_exp_tile_columns_minus1-1` (inclusive). `tile_column_width_minus1[num_exp_tile_columns_minus1]` is used to derive the width of tile columns whose index is greater than or equal to `num_exp_tile_columns_minus1` as specified in Clause 6.5.1. The value of `tile_column_width_minus1[i]` should be in the range of 0 to `PicWidthInCtbsY-1` (inclusive). When it does not exist, the value of `tile_column_width_minus1[0]` is inferred to be equal to `PicWidthInCtbsY-1`.

[0213] `tile_row_height_minus1[i]` incremented by 1 specifies the height of the i-th slice row in CTB units, where i ranges from 0 to `num_exp_tile_rows_minus1-1` (inclusive). `tile_row_height_minus1[num_exp_tile_rows_minus1]` is used to deduce the height of slice rows whose index is greater than or equal to `num_exp_tile_rows_minus1` as specified in Clause 6.5.1. The value of `tile_row_height_minus1[i]` should be in the range of 0 to `PicHeightInCtbsY-1` (inclusive). When it does not exist, the value of `tile_row_height_minus1[0]` is deduced to be equal to `PicHeightInCtbsY-1`.

[0214] A `rect_slice_flag` value of 0 specifies that slices within each slice are in the raster scan order, and slice information is not signaled in the PPS. A `rect_slice_flag` value of 1 specifies that slices within each slice cover a rectangular area of ​​the image, and slice information is signaled in the PPS. When it does not exist, `rect_slice_flag` is inferred to be equal to 1. When `subpic_info_present_flag` is equal to 1, the value of `rect_slice_flag` should be equal to 1.

[0215] A single_slice_per_subpic_flag value of 1 indicates that each subpicture consists of one and only one rectangular stripe. A single_slice_per_subpic_flag value of 0 indicates that each subpicture can consist of one or more rectangular stripes. When single_slice_per_subpic_flag is 1, num_slices_in_pic_minus1 is inferred to be equal to sps_num_subpics_minus1. When it does not exist, the value of single_slice_per_subpic_flag is inferred to be 0.

[0216] The increment of 1 in `num_slices_in_pic_minus1` specifies the number of rectangular stripes in each image of the reference PPS. The value of `num_slices_in_pic_minus1` should be in the range of 0 to `MaxSlicesPerPicture1` (inclusive), where `MaxSlicesPerPicture` is specified in Appendix A. When `no_pic_partition_flag` equals 1, the value of `num_slices_in_pic_minus1` is inferred to be 0.

[0217] A tile_idx_delta_present_flag value of 0 indicates that the tile_idx_delta value does not exist in the PPS, and all rectangular stripes in the image referencing the PPS are specified in raster order according to the procedure defined in Clause 6.5.1. A tile_idx_delta_present_flag value of 1 indicates that the tile_idx_delta value may exist in the PPS, and all rectangular stripes in the image referencing the PPS are specified in the order indicated by the tile_idx_delta value. When it does not exist, the value of tile_idx_delta_present_flag is inferred to be equal to 0.

[0218] The value of slice_width_in_tiles_minus1[i] plus 1 specifies the width of the i-th rectangular strip in units of slice columns. The value of slice_width_in_tiles_minus1[i] should be in the range of 0 to NumTileColumns-1 (inclusive of the end value).

[0219] The following applies when slice_width_in_tiles_minus1[i] does not exist:

[0220] – If NumTileColumns equals 1, then the value of slice_width_in_tiles_minus1[i] is inferred to be equal to 0.

[0221] Otherwise, the value of slice_width_in_tiles_minus1[i] is inferred as specified in Clause 6.5.1.

[0222] The value of slice_height_in_tiles_minus1[i] plus 1 specifies the height of the i-th rectangular strip in units of slice rows. The value of slice_height_in_tiles_minus1[i] should be in the range of 0 to NumTileRows-1 (inclusive).

[0223] The following applies when slice_height_in_tiles_minus1[i] does not exist:

[0224] – If NumTileRows equals 1, or tile_idx_delta_present_flag equals 0, and tileIdx%NumTileColumns is greater than 0, then the value of slice_height_in_tiles_minus1[i] is inferred to be equal to 0.

[0225] - Otherwise (NumTileRows is not equal to 1, and tile_idx_delta_present_flag is equal to 1 or tileIdx%NumTileColumns is equal to 0), when tile_idx_delta_present_flag is equal to 1 or tileIdx%NumTileColumns is equal to 0, the value of slice_height_in_tiles_minus1[i] is inferred to be equal to slice_height_in_tiles_minus1[i-1].

[0226] `num_exp_slices_in_tile[i]` specifies the number of slice heights explicitly provided in the current slice containing more than one rectangular slice. The value of `num_exp_slices_in_tile[i]` should be in the range of 0 to `RowHeight[tileY]-1` (inclusive), where `tileY` is the slice row index containing the `i`-th slice. When it does not exist, the value of `num_exp_slices_in_tile[i]` is inferred to be equal to 0. When `num_exp_slices_in_tile[i]` is equal to 0, the value of the variable `NumSlicesInTile[i]` is inferred to be equal to 1.

[0227] The increment of 1 in `exp_slice_height_in_ctus_minus1[j]` specifies the height of the j-th rectangular stripe in the current slice, in CTU rows. The value of `exp_slice_height_in_ctus_minus1[j]` should be in the range of 0 to `RowHeight[tileY]-1` (inclusive), where `tileY` is the slice row index of the current slice.

[0228] When num_exp_slices_in_tile[i] is greater than 0, the derivation of SliceHeightInCtusMinus1[i+k], where the range of variables NumSlicesInTile[i] and k is from 0 to NumSlicesInTile[i]-1, is as follows:

[0229]

[0230] `tile_idx_delta[i]` specifies the difference between the tile index of the first tile in the i-th rectangular strip and the tile index of the first tile in the (i+1)-th rectangular strip. The value of `tile_idx_delta[i]` should be in the range of -NumTilesInPic+1 to NumTilesInPic-1 (inclusive). When it does not exist, the value of `tile_idx_delta[i]` is inferred to be equal to 0. When it exists, the value of `tile_idx_delta[i]` should not be equal to 0.

[0231] ...

[0232] 7.4.2.4.5 The order of VCL NAL units and their relationship with the encoded / decoded image.

[0233] The order of VCL NAL units within the encoded / decoded image is constrained as follows:

[0234] – For any two NAL units A and B of a codec image, let subpicIdxA and subpicIdxB be their subpic-level index values, and sliceAddrA and sliceddrB be their slice_address values.

[0235] – Codec stripe NAL unit A should precede codec stripe NAL unit B if any of the following conditions are true:

[0236] –subpicIdxA is less than subpicIdxB.

[0237] –subpicIdxA equals subpicIdxB, sliceAddrA is less than sliceAddrB.

[0238] 7.4.8.1 General Strip Header Semantics

[0239] The variable CuQpDeltaVal, which specifies the difference between the luminance quantization parameter and its prediction for the codec unit containing cu_qp_delta_abs, is set to 0. It is also specified that when determining the Qp′ and Qp′ of the codec unit containing cu_chroma_qp_offset_flag...Cr and Qp′ CbCr The variable CuQpOffset is used to determine the value of the quantized parameter. Cb CuQpOffset Cr and CuQpOffset CbCr All of them were set to 0.

[0240] A value of 1 for picture_header_in_slice_header_flag indicates that the PH syntax structure exists in the slice header.

[0241] A value of 0 for picture_header_in_slice_header_flag indicates that the PH syntax structure does not exist in the slice header.

[0242] One requirement for bitstream consistency is that the value of picture_header_in_slice_header_flag should be the same across all CLVS codec slices.

[0243] When the picture_header_in_slice_header_flag of the encoded / decoded slice is equal to 1, one requirement for bitstream consistency is that there should be no VCL NAL units in CLVS with nal_unit_type equal to PH_NUT.

[0244] When picture_header_in_slice_header_flag equals 0, all encoded and decoded stripes in the current picture should have picture_header_in_slice_header_flag equal to 0, and the current PU should have PHNAL units.

[0245] `slice_subpic_id` specifies the subpick ID of the subpick containing the slice. If `slice_subpic_id` exists, the value of the variable `CurrSubpicIdx` is deduced such that `SubpicIdVal[CurrSubpicIdx]` equals `slice_subpic_id`. Otherwise (if `slice_subpic_id` does not exist), `CurrSubpicIdx` is deduced to be 0. The length of `slice_subpic_id` is `sps_subpic_id_len_minus1+1` bits.

[0246] `slice_address` specifies the slice address. When it does not exist, the value of `slice_address` is inferred to be 0. When `rect_slice_flag` is equal to 1 and `NumSlicesInSubpic[CurrSubpicIdx]` is equal to 1, the value of `slice_address` is inferred to be 0.

[0247] If rect_slice_flag equals 0, the following applies:

[0248] - The stripe address is the raster scan chip index.

[0249] The length of -slice_address is Ceil(Log2(NumTilesInPic)) bits.

[0250] The value of -slice_address should be in the range of 0 to NumTilesInPic1-1 (inclusive).

[0251] Otherwise (rect_slice_flag equals 1), the following applies:

[0252] - The stripe address is the sub-image level stripe index of the stripe.

[0253] The length of -slice_address is Ceil(Log2(NumSlicesInSubpic[CurrSubpicIdx])) bits.

[0254] The value of -slice_address should be in the range of 0 to NumSlicesInSubpic[CurrSubpicIdx]-1 (inclusive).

[0255] One requirement for bitstream consistency is the application of the following constraints:

[0256] - If rect_slice_flag is equal to 0 or subpic_info_present_flag is equal to 0, then the value of slice_address should not be equal to the value of slice_address of any other codec strip NAL unit of the same codec image.

[0257] Otherwise, the slice_subpic_id and slice_address values ​​should not be equal to the slice_subpic_id and slice_address values ​​of any other codec strip NAL unit of the same codec image.

[0258] - The shape of the image stripes should be such that, when decoded, the entire left and top boundaries of each CTU should be formed by the image boundaries or by the boundaries of one or more previously decoded CTUs.

[0259] sh_extra_bit[i] can be equal to 1 or 0. Decoders of this version conforming to this specification should ignore the value of sh_extra_bit[i]. Its value does not affect the level of conformity of the decoder as specified in this version of the specification.

[0260] The increment of 1 in num_tiles_in_slice_minus1 (if present) specifies the number of slices in the strip. The value of num_tiles_in_slice_minus1 should be in the range of 0 to NumTilesInPic-1 (inclusive).

[0261] The variable NumCtusInCurrSlice specifies the number of CTUs in the current slice, and the derivation of the list CtbAddrInCurrSlice[i] (where i ranges from 0 to NumCtusInCurrSlice-1 (inclusive)) of the image raster scan addresses of the i-th CTU in the slice is as follows:

[0262]

[0263]

[0264] The derivation of variables SubpicLeftBoundaryPos, SubpicTopBoundaryPos, SubpicRightBoundaryPos, and SubpicBotBoundaryPos is as follows:

[0265]

[0266] 3.5. Color Spaces and Chromaticity Subsampling

[0267] A color space, also known as a color model (or color system), is an abstract mathematical model that simply describes a range of colors as tuples of numbers, typically consisting of 3 or 4 values ​​or color components (e.g., RGB). Essentially, a color space is an explanation of a coordinate system and its subspaces.

[0268] For video compression, the most commonly used color spaces are YCbCr and RGB.

[0269] YCbCr, Y′CbCr, or Y Pb / Cb Pr / Cr (also written as YCBCR or Y'CBCR) are a family of color spaces used as part of the color image pipeline in video and digital photography systems. Y′ is the luminance component, and CB and CR are the blue-difference and red-difference chromaticity components, respectively. Y′ (with an apostrophe) is distinguished from Y (where Y is luminance), meaning that the light intensity is non-linearly encoded based on gamma-corrected RGB primary colors.

[0270] Chromatic subsampling is a practice of encoding images by taking advantage of the fact that the human visual system is less sensitive to color differences than to brightness, thereby achieving a lower resolution for chromatic information than for luminance information. 3.5.1. 4:4:4

[0272] Each of the three Y'CbCr components has the same sampling rate, therefore there is no chromaticity subsampling. This scheme is sometimes used in high-end film scanners and film post-production. 3.5.2. 4:2:2

[0274] The two chroma components are sampled at half the luminance sampling rate: the horizontal chroma resolution is halved, while the vertical chroma resolution remains unchanged. This reduces the bandwidth of the uncompressed video signal by one-third, but with almost no visual difference. (In the VVC working draft...) Figure 5 The text depicts examples of the nominal vertical and nominal horizontal positions for a 4:2:2 color format. 3.5.3. 4:2:0

[0276] Compared to 4:1:1, the horizontal sampling in 4:2:0 is doubled, but the vertical resolution is halved because the Cb and Cr channels are sampled only on each alternating line in this scheme. Therefore, the data rate is the same. Both Cb and Cr are subsampled by a factor of 2 in both the horizontal and vertical directions. There are three variations of the 4:2:0 scheme with different horizontal and vertical positions.

[0277] In MPEG-2, Cb and Cr coexist horizontally. Cb and Cr are located between pixels in the vertical direction (in the gaps).

[0278] • In JPEG / JFIF, H.261, and MPEG-1, Cb and Cr are located in the middle between alternating luminance samples.

[0279] In a 4:2:0DV configuration, Cb and Cr are co-sited in the horizontal direction. In the vertical direction, they are co-sited on alternating lines.

[0280] Table 3-1. SubWidthC and SubHeightC values ​​derived from chroma_format_idc and separate_colour_plane_flag.

[0281]

[0282] 4. Examples of technical problems solved by the disclosed embodiments

[0283] The existing design of signaling for SPS / PPS / image headers / strip headers in VVC has the following problems:

[0284] 1) Based on the current VVC text, when rect_slice_flag equals 1, the following applies:

[0285] a.slice_address represents the sub-picture level slice index of the slice.

[0286] b. The sub-image level stripe index is defined as the stripe index of the list of stripes in the sub-image (in the order in which they are signaled in the PPS).

[0287] c. The image-level stripe index is defined as the stripe index of the list of stripes in the image (in the order in which they are signaled in the PPS).

[0288] d. For any two stripes belonging to two different sub-images, the stripe associated with the smaller sub-image index is decoded earlier, while for any two stripes belonging to the same sub-image, the stripe with the smaller sub-image level stripe index is decoded earlier.

[0289] e. Furthermore, assuming that the order in which the image-level stripe index values ​​are increased is the same as the order in which the stripes are decoded, the variable NumCtusInCurrSlice, which specifies the number of CTUs in the current stripe, is derived using Equation 117 in the current VVC text.

[0290] However, when some stripes are generated by dividing a slice, some of the above aspects may be violated. As Figure 6In the example shown, when an image is divided into two slices by a vertical slice boundary, and each of these two slices is further divided into two stripes by the same horizontal boundary spanning the entire image, the upper two stripes are included in the first sub-image, and the lower two stripes are included in the second sub-image. In this case, according to the current VVC text, the image-level stripe index values ​​for the four stripes in stripe raster scan order would be 0, 2, 1, 3, while the decoding order index values ​​for the four stripes in stripe raster scan order would be 0, 1, 2, 3. Therefore, the derivation of NumCtusInCurrSlice would be incorrect, and conversely, the parsing of the stripe data would be problematic, the decoded sample values ​​would be incorrect, and the decoder might crash.

[0291] 2) There are two types of stripe signaling notification methods. In rectangular mode, all stripe segmentation information is signaled in the PPS. In non-rectangular mode, partial stripe segmentation information is signaled in the stripe header. Therefore, in this mode, the complete stripe division of the image cannot be known until all stripes of the image are parsed.

[0292] 3) In rectangular mode, arbitrary signaling notification stripes can be configured by setting `tile_idx_delta`. A poor bitstream could cause the decoder to crash due to this mechanism.

[0293] 4) In some embodiments, tile_idx_delta[i] is not initialized when i equals num_slices_in_pic_minus1.

[0294] Figure 6 An example of image partitioning is shown. Solid line 602 represents the boundary of a patch; dashed line 604 represents the boundary of a strip; and dashed line 606 represents the boundary of a sub-image. The figure shows the image-level index, decoding order index, sub-image-level index, and indices of sub-images and patches for the four strips.

[0295] 5. Example Implementations and Technologies

[0296] To address the aforementioned and other problems, methods summarized below are disclosed. This invention should be considered as an example of interpreting general concepts and not interpreted in a narrow sense. Furthermore, these inventions can be applied individually or in any combination.

[0297] 1. For stripes in rectangular stripe pattern (i.e., when rect_slice_flag equals 1), export the image-level stripe index for each stripe in each sub-image, and the exported value is used to export the number of CTUs in each stripe.

[0298] 2. Sub-image level stripe indexes can be defined / exported in the following ways:

[0299] a. In one example, the sub-image level slice index is defined as "the slice index of the list of slices in the sub-image (in their decoding order) when rect_slice_flag equals 1".

[0300] b. Alternatively, the subpic-level slice index is defined as "the slice index of the list of slices in the subpic when rect_slice_flag equals 1, as specified by the variable SubpicLevelSliceIdx[i] derived in Equation 32 (as in Example 1), where i is the picture-level slice index of the slice."

[0301] c. In one example, export the sub-image index for each stripe with a specific value of the image-level stripe index.

[0302] d. In one example, derive the sub-picture-level stripe index for each stripe with a specific value of the picture-level stripe index.

[0303] e. In one example, when rect_slice_flag equals 1, the semantics of the slice address are defined as “the slice address is the subpic-level slice index of the slice as defined by the variable SubpicLevelSliceIdx[i] derived in Equation 32 (e.g., as in Example 1), where i is the picture-level slice index of the slice.”

[0304] 3. The subpico-level slice index of a slice is assigned to the slice in the first subpico containing that slice. The subpico-level slice index of each slice can be stored in an array indexed by the picture-level slice index (e.g., SubpicLevelSliceIdx[i] in Example 1).

[0305] a. In one example, the sub-image level strip index is a non-negative integer.

[0306] b. In one example, the value of the sub-picture level stripe index of the stripe is greater than or equal to 0.

[0307] c. In one example, the value of the sub-image level stripe index of a stripe is less than N, where N is the number of stripes in the sub-image.

[0308] d. In one example, if the first stripe (strip A) and the second stripe (strip B) are in the same sub-image but they are different, then the first sub-image level stripe index of the first stripe (denoted as subIdxA) must be different from the second sub-image level stripe index of the second stripe (denoted as subIdxB).

[0309] e. In one example, if the first sub-image level slice index (denoted as subIdxA) of the first stripe (strip A) in the first sub-image is less than the second sub-image level slice index (denoted as subIdxB) of the second stripe (strip B) in the same first sub-image, then IdxA is less than IdxB, where idxA and idxB represent the slice indices of stripe A and stripe B in the entire image (also known as image-level slice indices, such as sliceIdx), respectively.

[0310] f. In one example, if the first sub-image level slice index (denoted as subIdxA) of the first stripe (strip A) in the first sub-image is less than the second sub-image level slice index (denoted as subIdxB) of the second stripe (strip B) in the same first sub-image, then stripe A is preceding stripe B in the decoding order.

[0311] g. In one example, the sub-image level stripe index in a sub-image is derived based on the image level stripe index (e.g., sliceIdx).

[0312] 4. Propose a function / mapping table for deriving a mapping between sub-image level stripe indices and image level stripe indices in sub-images.

[0313] a. In one example, a two-dimensional array PicLevelSliceIdx[subPicIdx][SubPicLevelSliceIdx] is exported to map subpicture-level slice indices in subpictures to picture-level slice indices, where PicLevelSliceIdx represents the picture-level slice index of a slice, subPicIdx represents the index of a subpicture, and SubPicLevelSliceIdx represents the subpicture-level slice index of a slice in a subpicture.

[0314] i. In one example, the array NumSlicesInSubpic[subPicIdx] is used to derive PicLevelSliceIdx, where NumSlicesInSubpic[subPicIdx] represents the number of slices in the subpicture with an index equal to subPicIdx.

[0315] 1) In one example, NumSlicesInSubpic[subPicIdx] and PicLevelSliceIdx[subPicIdx][SubPicLevelSliceIdx] are exported in a single process by scanning all stripes in order of image-level stripe index.

[0316] a. Before processing, NumSlicesInSubpic[subPicIdx] is set to 0 for all valid subPicIdx.

[0317] b. When checking a stripe with a picture level index equal to S, if it is in a subpicture with a subpicture index equal to P, then set PicLevelSliceIdx[P][NumSlicesInSubpic[P]] to equal S, and then set NumSlicesInSubpic[P] to equal NumSlicesInSubpic[P]+1.

[0318] ii. In one example, SliceIdxInPic[subPicIdx][SubPicLevelSliceIdx] is used to derive the image-level slice index (e.g., picLevelSliceIdx), which is then used to derive the number and / or address of CTBs in the slice when parsing the slice header.

[0319] 5. The consistency bitstream requirement stipulates that a slice cannot be located in multiple sub-pictures.

[0320] 6. The consistency bitstream requirement is that a sub-picture cannot contain two stripes (denoted as stripe A and stripe B), where stripe A is in slice A but smaller than slice A, and stripe B is in slice B but smaller than slice B, and slice A and slice B are different.

[0321] 7. Propose that the image slice and / or strip segmentation information can be signaled in the associated image header.

[0322] a. In one example, is the slice and / or strip segmentation information of the image signaled in the associated PPS or in the associated image header?

[0323] b. In one example, the signaling in the image header indicates whether the slice and / or strip segmentation information of the image is in the associated image header.

[0324] i. In one example, if the slice and / or stripe segmentation information of the image is signaled in both the associated PPS and the associated image header, then the slice and / or stripe segmentation information of the image signaled in the image header will be used.

[0325] ii. In one example, if the slice and / or stripe segmentation information of the image is signaled in both the associated PPS and the associated image header, the slice and / or stripe segmentation information of the image signaled in the PPS will be used.

[0326] c. In one example, signaling notification is provided in a video unit (such as in an SPS) at a level higher than the picture to indicate whether the slice and / or strip segmentation information of the picture is signaled in the associated PPS or in the associated picture header.

[0327] 8. A method is proposed that when associated images are divided into stripes in a non-rectangular pattern, the strip segmentation information is signaled in a higher-level video unit (such as PPS and / or image header) above the strip level.

[0328] a. In one example, when an associated image is divided into stripes in a non-rectangular pattern, signaling information indicating the number of stripes can be provided in a higher-level video unit (e.g., num_slices_in_pic_minus1).

[0329] b. In one example, when associated images are divided into stripes in a non-rectangular pattern, signaling in a higher-level video unit indicates information about the index (or address, or location, or coordinates) of the first block unit in the strip. For example, a block unit could be a CTU or a slice.

[0330] c. In one example, when associated images are divided into stripes in a non-rectangular pattern, signaling at a higher-level video unit indicates the number of block units in the stripe. For example, a block unit could be a CTU or a slice.

[0331] d. In one example, when an associated image is divided into stripes in a non-rectangular pattern, strip segmentation information (e.g., num_tiles_in_slice_minus1) is not signaled in the stripe header.

[0332] e. In one example, when an associated image is divided into stripes in a non-rectangular pattern, the stripe index is signaled in the stripe header.

[0333] i. In one example, when the associated image is divided into stripes in a non-rectangular pattern, slice_address is interpreted as the image-level stripe index.

[0334] f. In one example, when an associated image is divided into strips in a non-rectangular pattern, segmentation information for each strip in the image (such as the index of the first block and / or the number of block units) can be sequentially signaled in higher-level video units.

[0335] i. In one example, when associated images are divided into stripes in a non-rectangular pattern, the index of the strip can be signaled for each strip in a higher-level video unit.

[0336] ii. In one example, the segmentation information for each stripe is notified by ascending signaling of the stripe index.

[0337] 1) In one example, the segmentation information of each strip is notified by signaling in the order of strip 0, strip 1, ..., strip K-1, strip K, strip K+1, ..., strip S-2, strip S-1, where K represents the strip index and S represents the number of strips in the image.

[0338] iii. In one example, the segmentation information for each stripe is notified by descending signaling of the stripe index.

[0339] 1) In one example, the segmentation information of each strip is signaled in the order of strip S-2, strip S-1, ..., strip K+1, strip K, strip K-1, ..., strip 1, strip 0, where K represents the strip index and S represents the number of strips in the image.

[0340] iv. In one example, when associated images are divided into strips in a non-rectangular pattern, the index of the first block of the strip may not be signaled in higher-level video units.

[0341] 1) For example, the index of the first block cell of stripe 0 (the stripe with a stripe index of 0) is inferred to be 0.

[0342] 2) For example, the index of the first block cell of stripe K (the stripe whose index is equal to K, K>0) is inferred as Where, N i This indicates the number of block units in stripe i.

[0343] v. In one example, when associated images are divided into strips in a non-rectangular pattern, the index of the first block of the strip may not be signaled in higher-level video units.

[0344] 1) For example, the index of the first block cell of stripe 0 (the stripe with a stripe index of 0) is inferred to be 0.

[0345] 2) For example, the index of the first block cell of stripe K (the stripe whose index is equal to K, K>0) is inferred as Where, N i This indicates the number of block units in stripe i.

[0346] vi. In one example, when an associated image is divided into stripes in a non-rectangular pattern, the number of block units of the strip may not be signaled in the higher-level video unit.

[0347] 1) When there is only one stripe in the image and there are M block units in the image, the number of block units in stripe 0 is M.

[0348] 2) For example, the number of block units of stripe K (the stripe with stripe index equal to 0) is inferred to be T K+1 -T K , where T K represents the index of the first block unit of stripe K when K < S–1, where S is the number of stripes in the picture and S > 1.

[0349] 3) For example, the number of block units of stripe S-1 is inferred to be , where S is the number of stripes in the picture, S > 1, and M is the number of block units in the picture,

[0350] vii. In one example, when the associated picture is partitioned into stripes in a non-rectangular mode, the segmentation information of one or more stripes may not be signaled in a higher-level video unit.

[0351] 1) In one example, the segmentation information of one or more stripes not signaled in a higher-level video unit can be inferred from the segmentation information of other stripes to be signaled.

[0352] 2) In one example, the segmentation information of the last C stripes may not be signaled. For example, C equals 1.

[0353] 3) For example, the number of block units of stripe S-1 is not signaled, where S is the number of stripes in the picture and S > 1.

[0354] a. For example, the number of block units of stripe S-1 is inferred to be , where there are M block units in the picture.

[0355] 9. It is proposed that the minimum number of stripes in a picture can be different depending on whether rectangular segmentation or non-rectangular segmentation is applied.

[0356] a. In one example, if the non-rectangular segmentation mode is applied, the picture is partitioned into at least two stripes, while if the rectangular segmentation mode is applied, the picture is partitioned into at least one stripe.

[0357] i. For example, if the non-rectangular segmentation mode is applied, num_slices_in_pic_minus2 plus 2 that specifies the number of stripes in the picture can be signaled.

[0358] b. In one example, if the non-rectangular segmentation mode is applied, the picture is partitioned into at least one stripe, while if the rectangular segmentation mode is applied, the picture is partitioned into at least one stripe.

[0359] i. For example, if a rectangular segmentation mode is applied, the signaling can specify the number of slices in the image as num_slices_in_pic_minus2 plus 2.

[0360] c. In one example, when an image is not divided into sub-images or is divided into only one sub-image, the minimum number of stripes in the image can vary depending on whether rectangular segmentation or non-rectangular segmentation is applied.

[0361] 10. It is proposed that when signaling notification of segmentation information in video units such as PPS or image headers, the stripe is represented by the upper left position and the width / height of the stripe.

[0362] a. In one example, the index / position / coordinate of the top left block unit (such as a CTU or slice) of the signaling notification strip, and / or the width measured in video units (such as CTUs or slices), and / or the height measured in video units (such as CTUs or slices).

[0363] b. In one example, the top-left position and width / height information of each strip are signaled sequentially.

[0364] i. For example, the top-left position and width / height information of each stripe are signaled in ascending order of the stripe index, such as 0, 1, 2, ..., S-1, where S is the number of stripes in the image.

[0365] 11. Propose the segmentation information (such as position / width / height) of stripes in sub-pictures in video units such as SPS / PPS / picture headers, using signaling notification.

[0366] a. In one example, the strip segmentation information for each subgraph is signaled sequentially.

[0367] i. For example, the stripe segmentation information for each subgraph is signaled in ascending order of the subgraph index.

[0368] b. In one example, the segmentation information of each strip in the subgraph (such as position /

[0369] Width / height) sequential signaling notification.

[0370] i. In one example, the segmentation information (e.g., position / width / height) of each stripe in the subimage is signaled in ascending order of the subimage-level stripe index.

[0371] 12. Instead of signaling notification, derive the difference (denoted as tile_idx_delta[i]) between the tile index of the first piece in the i-th rectangular strip and the tile index of the first piece in the (i+1)-th rectangular strip.

[0372] a. In one example, based on the rectangular stripes from the 0th rectangular stripe to the ith rectangular stripe, derive the slice index of the first slice in the (i+1)th rectangular stripe.

[0373] b. In one example, the slice index of the first slice in the (i+1)th rectangular strip is derived as the minimum index of the slice that is not in the rectangular strip from the 0th rectangular strip to the 1st rectangular strip.

[0374] 13. A method was proposed to signal the number of tile columns / rows (e.g., NumTileColumns or NumTileRows) to be exported based on the relationship between the width of the image and the size of the CTU.

[0375] a. For example, if the width of the image is less than or equal to the size or width of the CTU, then signaling to num_exp_tile_columns_minus1 and / or tile_column_width_minus1 is not required.

[0376] b. For example, if the height of the image is less than or equal to the size or height of the CTU, then signaling to num_exp_tile_rows_minus1 and / or tile_row_height_minus1 is not required.

[0377] 14. When slice_subpic_id exists, there must be one and only one CurrSubpicIdx that satisfies SubpicIdVal[CurrSubpicIdx] equal to slice_subpic_id.

[0378] 15. If rect_slice_flag equals 0 or subpic_info_present_flag equals 0, then the value of slice_address+i (where i is in the range from 0 to num_tiles_in_slice_minus1 (inclusive)) should not be equal to the value of slice_address+j of any other codec's NAL unit of the same codec picture (where j is in the range from 0 to num_tiles_in_slice_minus1 (inclusive)), where i is in that range.

[0379] 16. When there are both uniform and non-uniform spaced pieces in an image, the type of piece layout can be specified in the signaling notification syntax element in PPS (or SPS).

[0380] a. In one example, a signaling notification syntax flag can be used in the PPS to specify whether the slice layout is a non-uniform interval followed by a uniform interval, or a uniform interval followed by a non-uniform interval.

[0381] b. For example, whenever there are non-uniformly spaced tiles, the number of explicitly provided tile columns / rows (e.g., num_exp_tile_columns_minus1, num_exp_tile_rows_minus1) can be no less than the total number of non-uniform tiles.

[0382] c. For example, whenever there are uniformly spaced tiles, the number of explicitly provided tile columns / rows (e.g., num_exp_tile_columns_minus1, num_exp_tile_rows_minus1) can be less than or equal to the total number of uniform tiles.

[0383] d. If the tile layout is similar to uniformly spaced followed by non-uniformly spaced (i.e., the picture starts with uniformly spaced tiles and ends with multiple non-uniformly spaced tiles),

[0384] i. In one example, the widths of the tile columns of the non-uniformly spaced tiles located behind the picture can be first assigned in reverse order (i.e., the order of tile indices is equal to NumTileColumns, NumTileColumns - 1, NumTileColumns - 2,...), and then the widths of the tile columns of the uniformly spaced tiles located in front of the picture can be implicitly derived in reverse order (i.e., the order of tile indices is equal to NumTileColumns - T, NumTileColumns - T - 1,...,

[0385] 2, 1, 0, where T represents the number of non-uniform tile columns).

[0386] 17. The syntax element that specifies the difference between the representative tile indices of two rectangular stripes can be used only when the condition is true, where one of the two rectangular stripes is the i-th stripe (e.g., tile_idx_delta[i]).

[0387] a. In one example, the condition is (i < num_slices_in_pic_minus1), where num_slices_in_pic_minus1 plus 1 represents the number of stripes in the picture.

[0388] b. In one example, the condition is (i!= num_slices_in_pic_minus1), where num_slices_in_pic_minus1 plus 1 represents the number of stripes in the picture.

[0389] 18. Whether and / or how signaling notifies or interprets or restricts the Merge Estimation Region (MER) size (e.g., notified by log2_parallel_merge_level_minus2 signaling) may depend on the minimum allowed codec block size (e.g., signaled as log2_min_luma_coding_block_size_minus2 and / or MinCbSizeY).

[0390] a. In one example, the size of the MER is required to be no smaller than the minimum allowed codec block size.

[0391] i. For example, it is required that log2_parallel_merge_level_minus2 should be equal to or greater than log2_min_luma_coding_block_size_minus2.

[0392] ii. For example, require log2_parallel_merge_level_minus2 to be in the range of log2_min_luma_coding_block_size_minus2 to CtbLog2SizeY–2.

[0393] b. In one example, the signaling informs the difference between Log2(MER size) and Log2(MinCbSizeY), which is denoted as log2_parallel_merge_level_minus_log2_mincb.

[0394] i. For example, log2_parallel_merge_level_minus_log2_mincb is encoded and decoded by unary codes (ue).

[0395] ii. For example, it is required that log2_parallel_merge_level_minus_log2_mincb should be in the range of 0 to CtbLog2SizeY-log2_min_luma_coding_block_size_minus2–2.

[0396] iii. For example, Log2ParMrgLevel = log2_parallel_merge_level_minus_log2_mincb + log2_min_luma_coding_block_size_minus2 + 2, where Log2ParMrgLevel is used to control the MER size.

[0397] 19. When num_exp_slices_in_tile[i] equals 0, derive the slice height of the i-th slice in units of CTU rows, for example, denoted as sliceHeightInCtus[i].

[0398] a. In one example, when num_exp_slices_in_tile[i] equals 0, sliceHeightInCtus[i] is derived as equal to RowHeight[SliceTopLeftTileIdx[i] / NumTileColumns].

[0399] 20. Propose that the num_exp_slices_in_tile[i]-1th slice in the slice containing the i-th slice in the image always exists, and the height is always exp_slice_height_in_ctus_minus1[i][num_exp_slices_in_tile[i]-1]+1 CTU rows.

[0400] a. Optionally, the num_exp_slices_in_tile[i]-1th slice in the slice containing the i-th slice in the image may or may not exist, and its height is less than or equal to exp_slice_height_in_ctus_minus1[i][num_exp_slices_in_tile[i]-1]+1 CTU rows.

[0401] 21. It is proposed that during the export of rectangular strip information, the variable tileIdx is only updated for stripes with image-level stripe indices less than num_slices_in_pic_minus1, that is, the variable tileIdx is not updated for the last stripe in each image of the reference PPS.

[0402] 22. It is proposed that the num_exp_tile_columns_minus1th tile column always exists in the reference PPS image, and the width is always tile_column_width_minus1[num_exp_tile_columns_minus1]+1 CTB.

[0403] 23. It is proposed that the num_exp_tile_rows_minus1th tile row always exists in the reference PPS image, and its height is always tile_column_height_minus1[num_exp_tile_rows_minus1]+1 CTBs.

[0404] 24. It is proposed that when the maximum image width and maximum image height are both less than CtbSizeY, the signaling notification of the syntax element sps_num_subpics_minus1 can be skipped.

[0405] a. Alternatively, when the above condition is true, the value of sps_num_subpics_minus1 is inferred to be equal to 0.

[0406] 25. Propose that when the image width is not greater than CtbSizeY, the signaling notification of the syntax element num_exp_tile_columns_minus1 can be skipped.

[0407] b. Alternatively, when the above condition is true, the value of num_exp_tile_columns_minus1 is inferred to be equal to 0.

[0408] 26. Propose that when the image height is not greater than CtbSizeY, the signaling notification of the syntax element num_exp_tile_rows_minus1 can be skipped.

[0409] c. Alternatively, when the above condition is true, the value of num_exp_tile_row_minus1 is inferred to be equal to 0.

[0410] 27. It is proposed that when num_exp_tile_columns_minus1 equals PicWidthInCtbsY-1, the signaling notification of the syntax element tile_column_width_minus1[i] can be skipped, where i ranges from 0 to num_exp_tile_columns_minus1 (inclusive).

[0411] d. Alternatively, the value of tile_column_width_minus1[i] is inferred to be equal to 0.

[0412] 28. It is proposed that when num_exp_tile_rows_minus1 equals PicHeightInCtbsY-1, the signaling notification of the syntax element tile_row_height_minus1[i] can be skipped, where i ranges from 0 to num_exp_tile_rows_minus1 (inclusive).

[0413] e. Alternatively, the value of tile_row_height_minus1[i] is inferred to be equal to 0.

[0414] 29. The height of a uniformly sliced ​​tile is indicated by the last entry of exp_slice_height_in_ctus_minus1[], which indicates the height of the tiles in that tile. A non-uniformly sliced ​​tile is the tile below the tile explicitly signaled. For example: uniformSliceHeight = exp_slice_height_in_ctus_minus1[i][num_exp_slices_in_tile[i]-1]+1

[0415] 30. Proposes that the width of the 'num_exp_tile_columns_minus1'th tile column should not be reset, that is, the width can be derived directly from the parsed value from the bitstream (e.g., represented by tile_column_width_minus1[num_exp_tile_columns_minus1]) without referring to other information.

[0416] a. In one example, the width of the 'num_exp_tile_columns_minus1'th tile column is directly set to tile_column_width_minus1[num_exp_tile_columns_minus1] plus 1. Alternatively, tile_column_width_minus1[num_exp_tile_columns_minus1] is used to derive the width of tile columns whose index is greater than, for example, num_exp_tile_columns_minus1 as specified in Clause 6.5.1.

[0417] b. Similarly, the height of the “num_exp_tile_columns_minus1”th tile row is not allowed to be reset; that is, the height can be derived directly from the parsed value from the bitstream (e.g., represented by tile_row_height_minus1[num_exp_tile_columns_minus1]) without referring to other information.

[0418] i. In one example, the height of the 'num_exp_tile_columns_minus1'th tile row is directly set to tile_row_height_minus1[num_exp_tile_columns_minus1] plus 1.

[0419] Alternatively, in addition,

[0420] `tile_row_height_minus1[num_exp_tile_columns_minus1]` is used to export the height of tile rows whose index is greater than, for example, `num_exp_tile_columns_minus1` as specified in Clause 6.5.1.

[0421] 31. It is proposed that the height of the num_exp_slices_in_tile[i]–1th slice in the slice should not be reset, that is, the height can be derived directly from the parsed value from the bitstream (e.g., represented by exp_slice_height_in_ctus_minus1[i][num_exp_slices_in_tile[i]-1]) without referring to other information.

[0422] a. In one example, the height of the num_exp_slices_in_tile[i]–1th slice in the slice is directly set to exp_slice_height_in_ctus_minus1[i][num_exp_slices_in_tile[i]–1] plus 1. Alternatively, in addition,

[0423] exp_slice_height_in_ctus_minus1[i][num_exp_slices_in_tile[i]–1] is used to derive the height of stripes with indices greater than num_exp_slices_in_tile[i]-1.

[0424] 6. Example

[0425] In the following embodiments, the added portions are marked as bold text, underlined text, and italic text. The deleted portions are marked within [[]].

[0426] 6.1. Example 1: Example of sub-image level strip index variation

[0427] 3 Definitions

[0428] Image-level slice index: When rect_slice_flag equals 1, the [[one]] slice index in the list of slices in the image (in the order of signaling notifications in PPS).

[0429] [[Sub-image level slice index: When rect_slice_flag equals 1, the slice index in the list of slices in the sub-image (in the order of signaling notifications in PPS).]]

[0430]

[0431] 6.5.1 CTB raster scanning, slice scanning, and sub-image scanning process

[0432] ...

[0433] [[List NumSlicesInSubpic[i], specifying the number of rectangular stripes in the i-th subpicture]] The following conclusion is drawn:

[0434]

[0435]

[0436] 7.4.8.1 General Strip Header Semantics

[0437] ...

[0438] `slice_address` specifies the slice address. When it does not exist, the value of `slice_address` is inferred to be 0. When `rect_slice_flag` is equal to 1 and `NumSlicesInSubpic[CurrSubpicIdx]` is equal to 1, the value of `slice_address` is inferred to be 0.

[0439] If rect_slice_flag equals 0, the following applies:

[0440] - The stripe address is the raster scan chip index.

[0441] The length of -slice_address is Ceil(Log2(NumTilesInPic)) bits.

[0442] The value of -slice_address should be in the range of 0 to NumTilesInPic-1 (inclusive).

[0443] Otherwise (rect_slice_flag equals 1), the following applies:

[0444] -Strip address is Sub-image level stripe index of the stripe

[0445] The length of -slice_address is Ceil(Log2(NumSlicesInSubpic[CurrSubpicIdx])) bits.

[0446] The value of -slice_address should be in the range of 0 to NumSlicesInSubpic[CurrSubpicIdx]-1 (inclusive).

[0447] One requirement for bitstream consistency is the application of the following constraints:

[0448] - If rect_slice_flag is equal to 0 or subpic_info_present_flag is equal to 0, then the value of slice_address should not be equal to the value of slice_address of any other codec strip NAL unit of the same codec image.

[0449] Otherwise, the slice_subpic_id and slice_address values ​​should not be equal to the slice_subpic_id and slice_address values ​​of any other codec strip NAL unit of the same codec image.

[0450] - The shape of the image stripes should be such that, when decoded, the entire left and top boundaries of each CTU should be formed by the image boundaries or by the boundaries of one or more previously decoded CTUs.

[0451] ...

[0452] The increment of 1 in num_tiles_in_slice_minus1 (if present) specifies the number of slices in the strip. The value of num_tiles_in_slice_minus1 should be in the range of 0 to NumTilesInPic-1 (inclusive).

[0453] The derivation of the variable NumCtusInCurrSlice, which specifies the number of CTUs in the current slice, and the list CtbAddrInCurrSlice[i], which specifies the raster scan address of the i-th CTU in the slice (where i ranges from 0 to NumCtusInCurrSlice-1 (inclusive)), is as follows:

[0454]

[0455] 6.2. Example 2: Signaling stripes in non-rectangular PPS

[0456] 7.3.2.4 Image Parameter Set RBSP Syntax

[0457]

[0458]

[0459]

[0460] 7.3.7.1 General Strip Header Syntax

[0461]

[0462] 7.4.3.4 Image Parameter Set RBSP Semantics

[0463] `num_slices_in_pic_minus1` plus 1 specifies the number of rectangular stripes in each picture of the reference PPS. The value of `num_slices_in_pic_minus1` should be in the range of 0 to `MaxSlicesPerPicture-1` (inclusive), where `MaxSlicesPerPicture` is specified in Appendix A. When `no_pic_partition_flag` equals 1, the value of `num_slices_in_pic_minus1` is inferred to be equal to...

[0464]

[0465] 7.4.8.1 General Strip Header Semantics

[0466] ...

[0467] `slice_address` specifies the slice address. If it doesn't exist, the value of `slice_address` is inferred to be 0. The value of `slice_address` is inferred to be 0 when `rect_slice_flag` is 1 and `NumSlicesInSubpic[CurrSubpicIdx]` is 1. The value of `slice_address` is inferred to be 0 when `rect_slice_flag` is 0 and `NumSlicesInPic` is 1.

[0468] If rect_slice_flag equals 0, the following applies:

[0469] - The stripe address is [[raster scan chip index]]

[0470] The length of -slice_address is [[NumTilesInPic]])) position.

[0471] The value of -slice_address should be between 0 and 1. Within the range of (including endpoints).

[0472] Otherwise (rect_slice_flag equals 1), the following applies:

[0473] - The stripe address is the sub-image level stripe index of the stripe.

[0474] The length of -slice_address is Ceil(Log2(NumSlicesInSubpic[CurrSubpicIdx])) bits.

[0475] The value of -slice_address should be in the range of 0 to NumSlicesInSubpic[CurrSubpicIdx]-1 (inclusive).

[0476] One requirement for bitstream consistency is the application of the following constraints:

[0477] If rect_slice_flag is equal to 0 or subpic_info_present_flag is equal to 0, then the value of slice_address should not be equal to the value of slice_address of any other codec strip NAL unit of the same codec image.

[0478] -otherwise, Therefore, the slice_subpic_id and slice_address values ​​should not be equal to the slice_subpic_id and slice_address values ​​of any other NAL unit of the same codec image.

[0479] - The shape of the image stripes should be such that, when decoded, the entire left and top boundaries of each CTU should be formed by the image boundaries or by the boundaries of one or more previously decoded CTUs.

[0480] ...

[0481] The derivation of the variable NumCtusInCurrSlice, which specifies the number of CTUs in the current slice, and the list CtbAddrInCurrSlice[i], which specifies the raster scan address of the i-th CTU in the slice (where i ranges from 0 to NumCtusInCurrSlice-1 (inclusive)), is as follows:

[0482]

[0483] 6.3. Example 3: Signaling slice based on image dimensions

[0484] 7.3.2.4 Image Parameter Set RBSP Syntax

[0485]

[0486]

[0487] 6.4. Example 4: Example 1 regarding the semantics of tile_column_width_minus1 and tile_row_height_minus1

[0488] 7.4.3.4 Image Parameter Set RBSP Semantics

[0489] ...

[0490] tile_columns_width_minus1[i] plus 1 specifies the value. The width of the i-th tile column in CTB units, where i ranges from 0 to num_exp_tile_columns_minus1-1 (inclusive). `tile_column_width_minus1[num_exp_tile_columns_minus1]` is used to derive the width of tile columns whose index is greater than or equal to num_exp_tile_columns_minus1 as specified in Clause 6.5.1. The value of `tile_column_width_minus1[i]` should be in the range of 0 to PicWidthInCtbsY-1 (inclusive). When it does not exist, the value of `tile_column_width_minus1[0]` is inferred to be equal to PicWidthInCtbsY1-1.

[0491] tile_row_height_minus1[i] +1 specifies the value. The height of the i-th tile row in CTB units, where i ranges from 0 to num_exp_tile_rows_minus1-1 (inclusive). tile_row_height_minus1[num_exp_tile_rows_minus1] is used to derive the height of tile rows whose index is greater than or equal to num_exp_tile_rows_minus1 as specified in Clause 6.5.1. The value of tile_row_height_minus1[i] should be in the range of 0 to PicHeightInCtbsY-1 (inclusive). When it does not exist, the value of tile_row_height_minus1[0] is inferred to be equal to PicHeightInCtbsY1-1.

[0492] ...

[0493] 6.5. Example 5: Example 2 regarding the semantics of tile_column_width_minus1 and tile_row_height_minus1

[0494] 7.4.3.4 Image Parameter Set RBSP Semantics

[0495] ...

[0496] The increment of `tile_column_width_minus1[i]` specifies the width of the i-th tile column in CTB units, where i ranges from 0 to 1. `num_exp_tile_columns_minus1–1` (inclusive of end values). `tile_column_width_minus1[num_exp_tile_columns_minus1]` is used to derive the width of tile columns whose index is greater than or equal to `num_exp_tile_columns_minus1` as specified in Clause 6.5.1. The value of `tile_column_width_minus1[i]` should be in the range of 0 to `PicWidthInCtbsY-1` (inclusive of end values). When it does not exist, the value of `tile_column_width_minus1[0]` is inferred to be equal to `PicWidthInCtbsY-1`.

[0497] `tile_row_height_minus1[i]` incremented by 1 specifies the height of the i-th tile row in CTB units, where i ranges from 0 to 1. num_exp_tile_rows_minus1–1 ) (Including end values). `tile_row_height_minus1[num_exp_tile_rows_minus1]` is used to derive the height of slice rows whose index is greater than or equal to `num_exp_tile_rows_minus1` as specified in Clause 6.5.1. The value of `tile_row_height_minus1[i]` should be in the range of 0 to `PicHeightInCtbsY-1` (inclusive of end values). When it does not exist, the value of `tile_row_height_minus1[0]` is inferred to be equal to `PicHeightInCtbsY-1`.

[0498] ...

[0499] 6.6. Example 6: Example Derivation of CTU in a Strip

[0500] 6.5 Scanning Process

[0501] 6.5.1 CTB raster scanning, slice scanning, and sub-image scanning process

[0502] ...

[0503] For rectangular stripes, the derivation of the list NumCtusInSlice[i] (where i ranges from 0 to num_slices_in_pic_minus1 (inclusive)) specifying the number of CTUs in the i-th strip, the list SliceTopLeftTileIdx[i] (where i ranges from 0 to num_slices_in_pic_minus1 (inclusive)) specifying the index of the top-left corner slice of the strip, and the matrix CtbAddrInSlice[i][j] (where i ranges from 0 to num_slices_in_pic_minus1 (inclusive)) specifying the image raster scan address of the j-th CTB in the i-th strip (where i ranges from 0 to num_slices_in_pic_minus1 (inclusive)) specifying the image raster scan address of the j-th CTB in the i-th strip is as follows:

[0504]

[0505]

[0506] 6.7. Example 7: Signaling regarding MER dimensions

[0507] 7.3.2.3 Sequence Parameter Set (RBSP) Syntax

[0508] seq_parameter_set_rbsp(){ descriptor sps_seq_parameter_set_id u(4) … log2_parallel_merge_level_minus_log2_mincb ue(v) … }

[0509] 7.4.3.3 Sequence Parameter Set (RBSP) Semantics

[0510] The value of the variable `Log2ParMrgLevel` is specified by `log2_parallel_merge_level_minus_log2_mincb` plus `log2_min_luma_coding_block_size_minus2+2`. This variable is used in the derivation process of spatial merge candidates as specified in Clause 8.5.2.3, the derivation process of motion vectors and reference indices in the sub-block merge mode as specified in Clause 8.5.5.2, and controls the invocation of the update process for the historical motion vector prediction value list in Clause 8.5.2.1. The value of `log2_parallel_merge_level_minus_log2_mincb` should be in the range of 0 to `CtbLog2SizeY - log2_min_luma_coding_block_size_minus2-2` (inclusive). The derivation of the variable `Log2ParMrgLevel` is as follows:

[0511] Log2ParMrgLevel=log2_parallel_merge_level_minus2+log2_min_luma_coding_block_size_minus2+2 (68)

[0512] 6.8. Example 8: Signaling regarding rectangular stripes

[0513] 6.5.1 CTB raster scanning, slice scanning, and sub-image scanning process

[0514] ...

[0515] The derivation of the list ctbToSubpicIdx[ctbAddrRs] (where ctbAddrRs ranges from 0 to PicSizeInCtbsY-1 (inclusive)) specifying the conversion from CTB addresses in image raster scans to subpic indices is as follows:

[0516]

[0517] When `rect_slice_flag` equals 1, it specifies the list `NumCtusInSlice[i]` containing the number of CTUs in the i-th slice (where i ranges from 0 to `num_slices_in_pic_minus1` (inclusive)), the list `SliceTopLeftTileIdx[i]` containing the slice index of the first CTU in the slice (where i ranges from 0 to `num_slices_in_pic_minus1` (inclusive)), the matrix `CtbAddrInSlice[i][j]` containing the image raster scan address of the j-th CTB in the i-th slice (where i ranges from 0 to `num_slices_in_pic_minus1` (inclusive) and j ranges from 0 to `NumCtusInSlice[i]-1` (inclusive)), and... The derivation is as follows:

[0518]

[0519]

[0520]

[0521] One requirement for bitstream consistency is that the value of NumCtusInSlice[i] (where i ranges from 0 to num_slices_in_pic_minus1 (inclusive)) should be greater than 0. Furthermore, another requirement for bitstream consistency is that the matrix CtbAddrInSlice[i][j] (where i ranges from 0 to num_slices_in_pic_minus1 (inclusive) and j ranges from 0 to NumCtusInSlice[i]-1 (inclusive)) should include each of all CTB addresses ranging from 0 to PicSizeInCtbsY-1 (inclusive) once and only once.

[0522] ...

[0523] 7.3.2.4 Image Parameter Set RBSP Syntax

[0524]

[0525] 7.4.3.4 Image Parameter Set Semantics

[0526] ...

[0527] A `tile_idx_delta_present_flag` value of 0 indicates that the `tile_idx_delta[i]` syntax element does not exist in PPS, and all images in PPS are divided into rectangular strip rows and columns in strip raster order. A `tile_idx_delta_present_flag` value of 1 indicates that the `tile_idx_delta[i]` syntax element can exist in PPS, and all rectangular stripes in the images in PPS are specified in increments of `i` according to the value of `tile_idx_delta[i]`. When it does not exist, the value of `tile_idx_delta_present_flag` is inferred to be 0.

[0528] The value of slice_width_in_tiles_minus1[i] plus 1 specifies the width of the i-th rectangular strip in units of slice columns. The value of slice_width_in_tiles_minus1[i] should be in the range of 0 to NumTileColumns-1 (inclusive of the end value).

[0529] When i is less than num_slices_in_pic_minus1 and NumTileColumns equals 1, the value of slice_width_in_tiles_minus1[i] is inferred to be equal to 0.

[0530] When num_exp_slices_in_tiles_MINUS1[i] equals 0, incrementing slice_height_in_tiles_minus1[i] by 1 specifies the height of the i-th rectangular strip in units of slice rows. The value of slice_height_in_tiles_minus1[i] should be in the range of 0 to NumTileRows-1 (inclusive).

[0531] When i is less than num_slices_in_pic_minus1 and slice_height_in_tiles_minus1[i] does not exist, it is inferred to be equal to NumTileRows==1?0:slice_height_in_tiles_MINUS1[i-1].

[0532] `num_exp_slices_in_tile[i]` specifies the number of explicitly provided stripe heights in the slice containing the i-th stripe (i.e., the slice whose slice index is equal to `SliceTopLeftTileIdx[i]`). The value of `num_exp_slices_in_tile[i]` should be in the range of 0 to `RowHeight[SliceTopLeftTileIdx[i] / NumTileColumns]-1` (inclusive of end values). When it does not exist, the value of `num_exp_slices_in_tile[i]` is inferred to be equal to 0.

[0533]

[0534] tile_idx_delta[i] specifies The value of tile_idx_delta[i] should be in the range of -NumTilesInPic+1 to NumTilesInPic-1 (inclusive). When it does not exist, the value of tile_idx_delta[i] is inferred to be equal to 0. When it exists, the value of tile_idx_delta[i] should not be equal to 0.

[0535] ...

[0536] 6.9. Example 9: Signaling regarding rectangular stripes

[0537] 6.5.1 CTB raster scanning, slice scanning, and sub-image scanning process

[0538] ...

[0539] When rect_slice_flag equals 1, it specifies the list NumCtusInSlice[i] (where i ranges from 0 to num_slices_in_pic_minus1 (inclusive)) of the number of CTUs in the i-th slice, the list SliceTopLeftTileIdx[i] (where i ranges from 0 to num_slices_in_pic_minus1 (inclusive)) of the slice index containing the first CTU in the slice, and the list SliceTopLeftTileIdx[i] of the slice index containing the first CTU in the slice. The derivation of the matrix CtbAddrInSlice[i][j] of the raster scan addresses of the j-th CTB (where i ranges from 0 to num_slices_in_pic_minus1 (inclusive) and j ranges from 0 to NumCtusInSlice[i]-1 (inclusive)) and the variable NumSlicesInTile[i] which specifies the number of slices in the slice containing the i-th slice (i.e., the slice whose slice index is equal to SliceTopLeftTileIdx[i]) is as follows:

[0540]

[0541]

[0542]

[0543] Alternatively, above, the following lines:

[0544]

[0545] The changes are as follows;

[0546]

[0547] 6.10. Example 10: Signaling regarding sub-pictures and slices

[0548] 6.5.1 CTB raster scanning, slice scanning, and sub-image scanning process

[0549] The derivation of the variable NumTileColumns, which specifies the number of slice columns, and the list colWidth[i], which specifies the width of the i-th slice column in CTB units (where i ranges from 0 to NumTileColumns-1 (inclusive)), is as follows:

[0550]

[0551] The derivation of the variable NumTileRows, which specifies the number of tile rows, and the list RowHeight[j], which specifies the height of the j-th tile row in CTB units (where j ranges from 0 to NumTileRows-1 (inclusive)), is as follows:

[0552]

[0553] 7.3.2.3 Sequence Parameter Set (RBSP) Syntax

[0554]

[0555] 7.3.2.4 Image Parameter Set RBSP Syntax

[0556]

[0557] 7.4.3.4 Image Parameter Set Semantics

[0558] ...

[0559] The value of num_exp_tile_columns_minus1 plus 1 specifies the number of tile column widths explicitly provided. The value of num_exp_tile_columns_minus1 should be in the range of 0 to PicWidthInCtbsY-1 (inclusive).

[0560] The increment of 1 in num_exp_tile_rows_minus1 specifies the number of tile row heights explicitly provided. The value of num_exp_tile_rows_minus1 should be in the range of 0 to PicHeightInCtbsY-1 (inclusive).

[0561] The increment of `tile_column_width_minus1[i]` specifies the width of the i-th tile column in CTB units, where i ranges from 0 to 1. (Including end values). `tile_column_width_minus1[num_exp_tile_columns_minus1]` is used to export the index. The width of the tile column as specified in Clause 6.5.1, num_exp_tile_columns_minus1. The value of tile_column_width_minus1[i] should be in the range of 0 to PicWidthInCtbsY-1 (inclusive). When it does not exist, the value of tile_column_width_minus1[i] is inferred to be equal to

[0562] `tile_row_height_minus1[i]` incremented by 1 specifies the height of the i-th tile row in CTB units, where i ranges from 0 to 1. (Including end values). `tile_row_height_minus1[num_exp_tile_rows_minus1]` is used to export the index. The height of the tile row is num_exp_tile_rows_minus1 as specified in Clause 6.5.1. The value of tile_row_height_minus1[i] should be in the range of 0 to PicHeightInCtbsY-1 (inclusive). When it does not exist, the value of tile_row_height_minus1[i] is inferred to be equal to

[0563] ...

[0564] Figure 7 This is a block diagram illustrating an example video processing system 1900, in which various techniques disclosed herein can be implemented. Various implementations may include some or all of the components of system 1900. System 1900 may include an input 1902 for receiving video content. The video content may be received in a raw or uncompressed format (e.g., 8-bit or 10-bit multi-component pixel values), or in a compressed or encoded format. Input 1902 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces (e.g., Ethernet, Passive Optical Networking (PON), etc.) and wireless interfaces (e.g., Wi-Fi or cellular interfaces).

[0565] System 1900 may include codec component 1904, which can implement the various codec or encoding methods described in this document. Codec component 1904 can reduce the average bit rate of the video from input 1902 to the output of codec component 1904 to produce a codec representation of the video. Therefore, codec techniques are sometimes referred to as video compression or video transcoding techniques. The output of codec component 1904 can be stored or transmitted via communication through the connection represented by component 1906. The stored or transmitted bitstream (or codec) representation of the video received at input 1902 can be used by component 1908 to generate pixel values ​​or displayable video to be sent to display interface 1910. The process of generating a user-viewable video from the bitstream is sometimes referred to as video decompression. Furthermore, although some video processing operations are referred to as “codec” operations or tools, it will be understood that codec tools or operations are used at the encoder, and the corresponding decoding tools or operations that inversely represent the codec results will be performed by the decoder.

[0566] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB), High Definition Multimedia Interface (HDMI), or DisplayPort. Examples of storage interfaces include SATA (Serial Advanced Technology Accessory), PCI, IDE, etc. The technologies described in this document can be found in a variety of electronic devices, such as mobile phones, laptops, smartphones, or other devices capable of performing digital data processing and / or video display.

[0567] Figure 8 This is a block diagram of a video processing apparatus 3600. Apparatus 3600 can be used to implement one or more methods described herein. Apparatus 3600 can be embodied in smartphones, tablets, computers, Internet of Things (IoT) receivers, etc. Apparatus 3600 may include one or more processors 3602, one or more memories 3604, and video processing hardware 3606. The processors(multiple) 3602 can be configured to implement one or more methods described in this document. The one or more memories 3604 can be used to store data and code used to implement the methods and techniques described herein. The video processing hardware 3606 can be used to implement some of the techniques described in this document in hardware circuitry.

[0568] Figure 10 This is a block diagram illustrating an example video codec system 100 that can utilize the techniques disclosed herein.

[0569] like Figure 10 As shown, the video encoding / decoding system 100 may include a source device 110 and a destination device 120. The source device 110 generates encoded video data, which may be referred to as a video encoding device. The destination device 120 can decode the encoded video data generated by the source device 110, which may be referred to as a video decoding device.

[0570] The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.

[0571] Video source 112 may include sources such as video capture devices, interfaces for receiving video data from video content providers, and / or computer graphics systems for generating video data, or combinations of these sources. Video data may include one or more images. Video encoder 114 encodes the video data from video source 112 to generate a bitstream. The bitstream may include a sequence of bits forming a codec representation of the video data. The bitstream may include codec images and associated data. A codec image is a codec representation of an image. Associated data may include sequence parameter sets, image parameter sets, and other syntax structures. I / O interface 116 may include a modulator / demodulator (modem) and / or a transmitter. Encoded video data can be transmitted directly to destination device 120 via network 130a through I / O interface 116. Encoded video data may also be stored on storage medium / server 130b for access by destination device 120.

[0572] Destination device 120 may include I / O interface 126, video decoder 124 and display device 122.

[0573] I / O interface 126 may include a receiver and / or a modem. I / O interface 126 may acquire encoded video data from source device 110 or storage medium / server 130b. Video decoder 124 may decode the encoded video data. Display device 122 may display the decoded video data to a user. Display device 122 may be integrated with destination device 120, or may be located external to destination device 120, which is configured to interact with an external display device.

[0574] The video encoder 114 and the video decoder 124 can operate according to video compression standards such as the High Efficiency Video Codec (HEVC) standard, the Universal Video Codec (VVM) standard, and other current and / or further standards.

[0575] Figure 11 This is a block diagram illustrating an example of a video encoder 200, which can be... Figure 10 The video encoder 114 in the system 100 shown.

[0576] The video encoder 200 can be configured to perform any or all of the technologies disclosed herein. Figure 11 In the example, the video encoder 200 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video encoder 200. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.

[0577] The functional components of the video encoder 200 may include a segmentation unit 201, a prediction unit 202 (which may include a mode selection unit 203), a motion estimation unit 204, a motion compensation unit 205, an intra-frame prediction unit 206, a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213, and an entropy coding unit 214.

[0578] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an intra-block copy (IBC) unit. The IBC unit can perform prediction in IBC mode, where at least one reference picture is the picture containing the current video block.

[0579] Furthermore, some components (such as motion estimation unit 204 and motion compensation unit 205) may be highly aggregated, but for illustrative purposes, in Figure 11 The examples represent the examples respectively.

[0580] The segmentation unit 201 can segment an image into one or more video blocks. The video encoder 200 and the video decoder 300 can support various video block sizes.

[0581] The mode selection unit 203 can, for example, select one of the encoding / decoding modes (intra-frame or inter-frame) based on the error result, and provide the resulting intra-frame or inter-frame encoded / decoded block to the residual generation unit 207 to generate residual block data, and to the reconstruction unit 212 to reconstruct the coded block for use as a reference picture. In some examples, the mode selection unit 203 can select a combination of intra-frame prediction and inter-frame prediction (CIIP) modes, where the prediction is based on the inter-frame prediction signal and the intra-frame prediction signal. The mode selection unit 203 can also select the resolution of the motion vector for the block (e.g., sub-pixel precision or integer pixel precision) in the case of inter-frame prediction.

[0582] To perform inter-frame prediction on the current video block, motion estimation unit 204 can generate motion information for the current video block by comparing one or more reference frames from buffer 213 with the current video block. Motion compensation unit 205 can determine the predicted video block for the current video block based on motion information and decoded samples from images other than those associated with the current video block from buffer 213.

[0583] For example, motion estimation unit 204 and motion compensation unit 205 can perform different operations on the current video block, depending on whether the current video block is in an I-band, P-band, or B-band.

[0584] In some examples, motion estimation unit 204 can perform unidirectional prediction on the current video block, and can search for a reference video block for the current video block in the reference images of list 0 or list 1. Then, motion estimation unit 204 can generate a reference index indicating the reference image in list 0 or list 1 containing the reference video block, and a motion vector indicating the spatial displacement between the current video block and the reference video block. Motion estimation unit 204 can output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current block based on the reference video block indicated by the motion information of the current video block.

[0585] In other examples, motion estimation unit 204 can perform bidirectional prediction on the current video block. Motion estimation unit 204 can search for a reference video block for the current video block in the reference images in list 0, and also search for another reference video block for the current video block in the reference images in list 1. Then, motion estimation unit 204 can generate reference indices indicating the reference images in lists 0 and 1 containing the reference video blocks, and motion vectors indicating the spatial displacement between the reference video blocks and the current video block. Motion estimation unit 204 can output the reference index and motion vector of the current video block as the motion information of the current video block. Motion compensation unit 205 can generate a predicted video block for the current video block based on the reference video blocks indicated by the motion information of the current video block.

[0586] In some examples, the motion estimation unit 204 can output a complete set of motion information for the decoder's decoding processing.

[0587] In some examples, motion estimation unit 204 may not output the complete set of motion information for the current video. Instead, motion estimation unit 204 can signal the motion information of the current video block by referencing the motion information of another video block. For example, motion estimation unit 204 may determine that the motion information of the current video block is sufficiently similar to the motion information of neighboring video blocks.

[0588] In one example, the motion estimation unit 204 may instruct the video decoder 300, within the syntax structure associated with the current video block, to indicate that the current video block has the same motion information as another video block.

[0589] In another example, motion estimation unit 204 can identify another video block and motion vector difference (MVD) within the syntactic structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. Video decoder 300 can use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.

[0590] As discussed above, the video encoder 200 can predictively signal motion vectors. Two examples of predictive signaling techniques that can be implemented by the video encoder 200 include Advanced Motion Vector Prediction (AMVP) and merge pattern signaling.

[0591] Intra-prediction unit 206 can perform intra-prediction on the current video block. When intra-prediction unit 206 performs intra-prediction on the current video block, it can generate prediction data for the current video block based on decoded samples from other video blocks in the same frame. The prediction data for the current video block may include the predicted video block and various syntax elements.

[0592] The residual generation unit 207 can generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) the predicted video block from the current video block. The residual data for the current video block can include residual video blocks corresponding to different sample components in the current video block.

[0593] In other examples, there may be no residual data for the current video block. For example, in skip mode, the residual generation unit 207 may not perform the subtraction operation.

[0594] The transform processing unit 208 can generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video blocks associated with the current video block.

[0595] After the transform processing unit 208 generates a transform coefficient video block associated with the current video block, the quantization unit 209 can quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values ​​associated with the current video block.

[0596] The inverse quantization unit 210 and the inverse transform unit 211 can apply inverse quantization and inverse transform to the transform coefficient video block, respectively, to reconstruct the residual video block based on the transform coefficient video block. The reconstruction unit 212 can add the reconstructed residual video block to the corresponding samples of one or more predicted video blocks generated by the prediction unit 202 to produce a reconstructed video block associated with the current block, which is then stored in the buffer 213.

[0597] After the video block is reconstructed by reconstruction unit 212, a loop filtering operation can be performed to reduce video block artifacts in the video block.

[0598] Entropy encoding unit 214 can receive data from other functional components of video encoder 200. When entropy encoding unit 214 receives data, it can perform one or more entropy encoding operations to generate entropy encoded data and output a bit stream including the entropy encoded data.

[0599] Figure 12 This is a block diagram illustrating an example of a video decoder 300, which can be... Figure 10 The video decoder 114 in the system 100 shown.

[0600] The video decoder 300 can be configured to perform any or all of the technologies disclosed herein. Figure 12 In the example, the video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video decoder 300. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.

[0601] exist Figure 12 In the example, the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra-frame prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, the video decoder 300 can perform encoding passes typically described with respect to the video encoder 200. Figure 11 The opposite decoding iteration.

[0602] The entropy decoding unit 301 can retrieve the encoded bitstream. The encoded bitstream may include entropy-coded video data (e.g., encoded blocks of video data). The entropy decoding unit 301 can decode the entropy-coded video data, and the motion compensation unit 302 can determine motion information based on the entropy-decoded video data. This motion information includes motion vectors, motion vector precision, reference image list index, and other motion information. For example, the motion compensation unit 302 can determine this information by executing AMVP and merge modes.

[0603] The motion compensation unit 302 can generate motion compensation blocks, possibly performing interpolation based on an interpolation filter. The syntax elements may include identifiers for the interpolation filter used at sub-pixel precision.

[0604] The motion compensation unit 302 can use the interpolation filter used by the video encoder 200 during video block encoding to calculate the interpolation of sub-integer pixels of the reference block. The motion compensation unit 302 can determine the interpolation filter used by the video encoder 200 based on the received syntax information, and use the interpolation filter to generate the prediction block.

[0605] The motion compensation unit 302 may use some syntax information to determine the size of the blocks used to encode the frames and / or stripes of the encoded video sequence, segmentation information describing how each macroblock of the image of the encoded video sequence is segmented, a mode indicating how each partition is encoded, one or more reference frames (and a list of reference frames) for each inter-frame coded block, and other information for decoding the encoded video sequence.

[0606] Intra-prediction unit 303 can use, for example, an intra-prediction mode received in the bitstream to form prediction blocks based on spatially adjacent blocks. Dequantization unit 303 dequantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by entropy decoding unit 301. Inverse transform unit 303 applies an inverse transform.

[0607] The reconstruction unit 306 can add the residual block to the corresponding prediction block generated by the motion compensation unit 202 or the intra-frame prediction unit 303 to form a decoded block. If necessary, a deblocking filter can also be applied to filter the decoded block to remove block artifacts. The decoded video block is then stored in a buffer 307, which provides a reference block for subsequent motion compensation / intra-frame prediction and also generates decoded video for presentation on a display device.

[0608] Below is a list of preferred solutions for some embodiments.

[0609] The following solutions illustrate example embodiments of the techniques discussed in the previous section (e.g., item 1).

[0610] 1. A video processing method (e.g., Figure 9 The method 900 described herein includes: performing a conversion between a video and a video codec representation comprising one or more video images (902), wherein each video image comprises one or more sub-images, each sub-image comprising one or more stripes, wherein the codec representation conforms to a format rule; wherein the format rule specifies that, when rectangular stripe mode is enabled for the video image, the image-level stripe index of each stripe in each sub-image of the video image is derived without explicit signaling notification in the codec representation; wherein the format rule specifies that the number of codec tree units in each stripe can be derived from the image-level stripe index.

[0611] The following solutions illustrate example embodiments of the techniques discussed in the previous section (e.g., item 2).

[0612] 2. A video processing method, comprising: performing a conversion between a video comprising one or more video images and a video codec representation, wherein each video image comprises one or more sub-images, each sub-image comprising one or more stripes, wherein the codec representation conforms to a format rule; wherein the format rule specifies that a sub-image level stripe index can be derived based on information in the codec representation without signaling notification of the sub-image level stripe index in the codec representation.

[0613] 3. Solution 2's method, wherein the formatting rules stipulate that, due to the use of rectangular stripe structures, the sub-image level stripe index corresponds to the stripe index in the list of stripes in the sub-image.

[0614] 4. Solution 2, in which the formatting rules specify that the sub-picture level stripe index is derived from a specific value of the picture level stripe index.

[0615] The following solutions illustrate example embodiments of the techniques discussed in the previous section (e.g., items 5 and 6).

[0616] 5. A video processing method, comprising: performing a conversion between a video comprising one or more video images and a video codec representation, wherein each video image comprises one or more sub-images and / or one or more segments, wherein the codec representation conforms to format rules; wherein the conversion conforms to constraint rules.

[0617] 6. Solution 5, where the constraint rule stipulates that a piece cannot be in more than one sub-picture.

[0618] 7. The method of Solution 5, wherein the constraint rule stipulates that a sub-image cannot include two stripes smaller than the corresponding patches to which the two stripes belong.

[0619] The following solutions illustrate example embodiments of the techniques discussed in the previous section (e.g., items 7 and 8).

[0620] 8. A video processing method, comprising: performing a conversion between a video comprising one or more video images and a video codec representation, wherein each video image comprises one or more slices and / or one or more stripes; wherein the codec representation conforms to a format rule; wherein the format rule specifies that video image-level fields carry information about the segmentation of stripes and / or slices in the video image.

[0621] 9. The method of Solution 8, wherein the fields include video image headers.

[0622] 10. The method of Solution 8, wherein the fields include a set of image parameters.

[0623] 11. The method of any one of solutions 8-10, wherein the formatting rule specifies the omission of stripe-level stripe segmentation information by including stripe segmentation information in the video image-level field.

[0624] The following solutions illustrate example embodiments of the techniques discussed in the previous section (e.g., item 9).

[0625] 12. A video processing method, comprising: performing a conversion between a video comprising one or more images and a video codec representation, wherein the conversion conforms to a segmentation rule based on whether rectangular segmentation is used to segment the video images, wherein the minimum number of strips into which the video images are segmented is determined.

[0626] 13. The method of Solution 12, wherein the segmentation rule specifies the use of at least two stripes for non-rectangular segments and at least one stripe for rectangular segments.

[0627] 14. The method of Solution 12, wherein the segmentation rule is also a function of whether and / or how many sub-images are used to segment video images.

[0628] The following solutions illustrate example embodiments of the techniques discussed in the previous section (e.g., items 10 and 11).

[0629] 15. A video processing method, comprising: performing a conversion between a video stripe of a video region and a codec representation of the video; wherein the codec representation conforms to a format rule; wherein the format rule specifies that the codec representation signals the video stripe based on the upper left position of the video stripe, wherein the format rule specifies that the codec representation signals the height and / or width of the video stripe in segmentation information signaled at the video unit level.

[0630] 16. The method of Solution 15, wherein the format rules specify that video stripes are signaled in the order of stripes defined by the format rules.

[0631] 17. The method of Solution 15, wherein a video region corresponds to a sub-picture, and a video unit level corresponds to a video picture.

[0632] The following solutions illustrate example embodiments of the techniques discussed in the previous section (e.g., item 12).

[0633] 18. A video processing method, comprising: performing a conversion between a video including video images and a video codec representation; wherein the codec representation conforms to a format rule; wherein the format rule specifies omitting the difference between the slice index of a first slice in a signaling notification rectangular strip and the slice index of a first slice in a next rectangular strip.

[0634] 19. The method of Solution 18, wherein the difference can be derived from the zeroth strip and the rectangular strip in the video image.

[0635] The following solutions illustrate example embodiments of the techniques discussed in the previous section (e.g., item 13).

[0636] 20. A video processing method, comprising: performing a conversion between a video and a video codec representation, wherein the codec representation conforms to a format rule, wherein the format rule specifies a relationship between the width of a video frame and the size of a codec tree unit, and controlling signaling notification for deriving information about the number of slice columns or slice rows in the video frame.

[0637] 21. The method of Solution 20, wherein the format rules specify that the number of signaling notification slice lines or slice columns shall be excluded when the width of the video image is less than or equal to the width of the codec tree unit.

[0638] The following solutions illustrate example embodiments of the techniques discussed in the previous section (e.g., item 16).

[0639] 22. A video processing method, comprising: performing a conversion between a video and a encoded representation of a video comprising one or more video images, wherein the encoded representation conforms to a format rule, wherein the format rule specifies that the encoded representation of the video images comprising uniformly spaced segments and non-uniformly spaced segments includes segment layout information.

[0640] 23. The method of Solution 22, wherein the piece layout information includes syntax tags contained in the picture parameter set.

[0641] 24. The method of any one of solutions 22-23, wherein the number of slice rows or slice columns explicitly signaled is not less than the number of non-uniformly spaced slices.

[0642] 25. The method of any one of solutions 22-23, wherein the number of slice rows or slice columns explicitly signaled is not less than the number of uniformly spaced slices.

[0643] 26. The method of any of the above solutions, wherein the video region includes a video encoding / decoding unit.

[0644] 27. The method of any of the above solutions, wherein the video region includes video images.

[0645] 28. The method of any one of solutions 1 to 27, wherein the conversion includes encoding the video into a codec representation.

[0646] 29. The method of any one of solutions 1 to 27, wherein the transformation includes decoding the codec representation to generate pixel values ​​of the video.

[0647] 30. A video decoding apparatus, comprising a processor configured to implement the method described in one or more of solutions 1 to 29.

[0648] 31. A video encoding apparatus, comprising a processor configured to implement the method described in one or more of solutions 1 to 29.

[0649] 32. A computer program product having computer code stored thereon, which, when executed by a processor, causes the processor to implement the method described in any one of solutions 1 to 29.

[0650] 33. A method, apparatus, or system described in this document.

[0651] Figure 13 This is a flowchart representation of a video processing method according to the present technology. Method 1300 includes, in operation 1310, performing a conversion between a video picture and a video bitstream according to a rule. The video picture includes one or more stripes. The rule stipulates that, in response to a condition being met, a syntax element indicating the difference between the slice indices of two rectangular stripes is signaled, wherein one of the strip indices is represented as i, where i is an integer.

[0652] In some embodiments, the second rectangular stripe of two rectangular stripes is represented as the (i+1)th rectangular stripe, and the syntax element indicates the difference between the first slice index of the first slice containing the first codec tree unit in the (i+1)th rectangular stripe and the second slice index of the second slice containing the first codec tree unit in the ith rectangular stripe. In some embodiments, at least one condition is satisfied, including i being less than (the number of rectangular stripes in the video frame - 1). In some embodiments, at least one condition is satisfied, including i not being equal to (the number of rectangular stripes in the video frame - 1).

[0653] Figure 14 This is a flowchart representation of a video processing method according to the present technology. Method 1400 includes, in operation 1410, performing a conversion between video frames and a video bitstream according to a rule. A video frame includes one or more sub-frames, and each sub-frame includes one or more rectangular stripes. The rule specifies deriving the stripe index at the sub-frame level for each rectangular stripe in each sub-frame to determine the number of codec tree units in each stripe.

[0654] In some embodiments, the sub-image-level slice index is determined based on the decoding order of the corresponding slices in the list of slices in the sub-image. In some embodiments, if the first slice index at the sub-image level of the first slice is less than the second slice index at the sub-image level of the second slice, the first slice is processed before the second slice according to the decoding order. In some embodiments, the variable SubpicLevelSliceIdx is used to represent the sub-image-level slice index. In some embodiments, the slice address is determined based on the sub-image-level slice index. In some embodiments, the sub-image-level slice index of the slice is determined based on the first sub-image including the slice. In some embodiments, the sub-image-level slice index is a non-negative integer. In some embodiments, the sub-image-level slice index is greater than or equal to 0 and less than N, where N is the number of slices in the sub-image.

[0655] In some embodiments, when the first band and the second band are different, the first band index at the sub-image level is different from the second band index at the sub-image level, wherein the first band and the second band are in the same sub-image. In some embodiments, when the first band index at the sub-image level of the first band is less than the second band index at the sub-image level of the second band, the first band index at the image level of the first band is less than the second band index at the image level of the second band. In some embodiments, the sub-image level band index of a band is determined based on the image level band index of that band.

[0656] Figure 15 This is a flowchart representation of a video processing method according to the present technology. Method 1500 includes, in operation 1510, determining a mapping relationship between a sub-picture-level stripe index of a stripe in a sub-picture and a picture-level stripe index of that stripe for a conversion between a video picture comprising one or more sub-pictures and a video bitstream. Method 1500 further includes, in operation 1520, performing a conversion based on the determination.

[0657] In some embodiments, the mapping is represented as a two-dimensional array indexed using the image-level stripe index of the sub-image and the sub-image index. In some embodiments, the image-level stripe index is determined based on an array indicating the number of stripes in each of one or more sub-images. In some embodiments, the two-dimensional array and the array indicating the number of stripes in each of one or more sub-images are determined based on a process of scanning all stripes in order of image-level stripe index. In some embodiments, the image-level stripe index of a strip is determined using the mapping, and the number of codec tree blocks and / or the addresses of codec tree blocks in the strip are determined based on the image-level stripe index of the strip.

[0658] Figure 16This is a flowchart representation of a video processing method according to the present technology. Method 1600 includes, in operation 1610, performing a conversion between a video picture and a video bitstream according to a rule. The video picture includes one or more sub-pictures. The rule stipulates that a slice of video is entirely contained within a single sub-picture of the video picture.

[0659] In some embodiments, the first stripe is located in the first piece and is smaller than the first piece. The second stripe is located in the second piece and is smaller than the second piece. The first piece is different from the second piece, and the first stripe and the second stripe are located in different stripes of the sub-image.

[0660] Figure 17 This is a flowchart representation of a video processing method according to the present technology. Method 1700 includes, in operation 1710, performing a conversion between video frames and a video bitstream. The video frames include one or more sub-frames. The bitstream conforms to a format rule that specifies that information about the segmented frames is included in a syntax structure associated with the frames.

[0661] In some embodiments, the syntax structure includes an image header or an image parameter set. In some embodiments, the video unit includes a flag indicating whether information about the segmented image is included in the syntax structure. In some embodiments, the video unit includes an image header, an image parameter set, or a sequence parameter set. In some embodiments, when information about the segmented image is included in both the image header and the image parameter set, the information included in the image header is used for transformation.

[0662] Figure 18 This is a flowchart representation of a video processing method according to the present technology. Method 1800 includes, in operation 1810, determining striping information of the video picture for a conversion between a video picture comprising one or more stripes having a non-rectangular shape and a bitstream of the video. Method 1800 further includes, in operation 1820, performing a conversion based on the determination.

[0663] In some embodiments, strip segmentation information is stored in video units of a video, wherein a video unit includes a set of picture parameters or a picture header. In some embodiments, the bitstream conforms to a format rule that specifies that strip segmentation information of video pictures is included in a syntax structure associated with a video unit comprising one or more stripes. In some embodiments, the strip segmentation information of a video picture includes a value indicating the number of stripes in the picture. In some embodiments, the strip segmentation information of a video picture includes a value indicating the index of a block unit within a stripe. In some embodiments, the value indicating the index of a block unit within a stripe is omitted from the strip segmentation information of the video picture. In some embodiments, the strip segmentation information of a video picture includes a value indicating the number of block units within a stripe. In some embodiments, the value indicating the number of block units within a stripe is omitted from the strip segmentation information of the video picture. In some embodiments, a block unit is a codec tree unit or a slice. In some embodiments, the strip segmentation information of the video picture is omitted in the stripe header. In some embodiments, the stripe index of a stripe is included in the stripe header, and the picture-level stripe index of the stripe is determined based on the address of the stripe.

[0664] In some embodiments, the strip segmentation information of each of one or more strips is organized in an implementation of a syntax structure associated with a video unit. In some embodiments, the strip segmentation information of each of one or more strips is included in the syntax structure in ascending order. In some embodiments, the strip segmentation information of each of one or more strips is included in the syntax structure in descending order.

[0665] In some embodiments, the slice segmentation information of at least one slice in one or more slices is omitted from the bitstream. In some embodiments, the slice segmentation information of at least one slice in one or more slices is inferred from the slice segmentation information of other slices included in the bitstream. In some embodiments, the number of block units of slice S-1 is not included in the bitstream, where S represents the number of slices in the video frame, and S is greater than 1.

[0666] Figure 19 This is a flowchart representation of a video processing method according to the present technology. Method 1900 includes, in operation 1910, performing a conversion between a video picture comprising one or more stripes and a video bitstream according to a rule. The rule specifies that the number of stripes in the video picture is equal to or greater than a minimum number of stripes determined based on whether rectangular segmentation or non-rectangular segmentation is applied to the video picture.

[0667] In some embodiments, the minimum number of stripes is two when non-rectangular segmentation is applied, and one when rectangular segmentation is applied. In some embodiments, the minimum number of stripes is one when non-rectangular segmentation is applied, and one when rectangular segmentation is applied. In some embodiments, the minimum number of stripes is also determined based on the number of sub-images in the video image.

[0668] Figure 20 This is a flowchart representation of a video processing method according to the present technology. Method 2000 includes, in operation 2010, performing a conversion between a video image and a video bitstream according to rules. The video image comprises one or more stripes. Where the strip segmentation information of the video image is included in the syntax structure of the video unit, the stripes are represented by the top-left position and dimension of the stripes.

[0669] In some embodiments, the top-left position and dimension of a strip are indicated by the top-left position of a block cell within the strip, the dimension of the block cell, and the dimension of the strip as measured using the dimension of the block cell. In some embodiments, the top-left position and dimension of each strip are included sequentially in the syntax structure.

[0670] Figure 21 This is a flowchart representation of a video processing method according to the present technology. Method 2100 includes, in operation 2110, performing a conversion between video frames and a video bitstream according to rules. A video frame includes one or more sub-frames, and each sub-frame includes one or more stripes. The rules specify how the segmentation information of the one or more stripes in each sub-frame exists in the bitstream.

[0671] In some embodiments, a video unit includes a sequence parameter set, a picture parameter set, or a picture header. In some embodiments, the segmentation information of one or more sub-pictures is arranged in ascending order based on the sub-picture index. In some embodiments, the segmentation information of one or more stripes in each sub-picture is arranged in ascending order based on the sub-picture level stripe index.

[0672] Figure 22 This is a flowchart representation of a video processing method according to the present technology. Method 2200 includes, in operation 2210, performing a conversion between video frames and a video bitstream according to a rule. The video frames comprise one or more rectangular strips, and each strip comprises one or more slices. The rule specifies omitting signaling in the bitstream the difference between the first slice index of the first slice in the i-th rectangular strip and the second slice index of the first slice in the (i+1)-th rectangular strip.

[0673] In some embodiments, the second slice index of the first piece in the (i+1)th rectangular strip is derived based on the rectangular strip from index 0 to index 1. In some embodiments, the second slice index of the first piece in the (i+1)th rectangular strip is derived as a minimum slice index that is outside the range defined by the rectangular strip from index 0 to index 1.

[0674] Figure 23 This is a flowchart representation of a video processing method according to the present technology. Method 2300 includes, in operation 2310, a conversion between a video picture and a bitstream of the video, conditionally including information for determining the number of columns and rows of slices in the video picture, in response to the relationship between the dimensions of the video picture and the dimensions of the codec tree block, in the bitstream. Method 2300 also includes, in operation 2320, performing the conversion based on the determination.

[0675] In some embodiments, this information is omitted when the width of the video image is less than or equal to the width of the codec tree block. In some embodiments, this information is omitted when the height of the video image is less than or equal to the height of the codec tree block.

[0676] Figure 24 This is a flowchart representation of a video processing method according to the present technology. Method 2400 includes, in operation 2410, performing a conversion between a video picture and a video bitstream. The video picture includes one or more subpictures. The bitstream conforms to a format rule that specifies that, where a variable specifying a subpicture identifier including a striped subpicture exists in the bitstream, there exists one and only one syntax element that satisfies the condition that a second variable corresponding to that syntax element is equal to that variable.

[0677] In some embodiments, the variable is represented as slice_subpic_id, and the second variable is represented as SubpicIdVal. The syntax element corresponding to this syntax element is represented as SubpicIdVal[CurrSubpicIdx].

[0678] Figure 25 This is a flowchart representation of a video processing method according to the present technology. Method 2500 includes, in operation 2510, performing a conversion between video frames and a video bitstream. A video frame includes one or more sub-frames. When non-rectangular segmentation is applied or sub-frame information is omitted from the bitstream, two slices in a strip have different addresses.

[0679] In some embodiments, the first syntax flag rect_slice_flag indicates whether non-rectangular segmentation is applied, and the second syntax flag subpic_info_present_info indicates whether subpic information exists in the bitstream.

[0680] Figure 26 This is a flowchart representation of a video processing method according to the present technology. Method 2600 includes, in operation 2610, performing a conversion between video frames and a video bitstream according to rules. The video frames include one or more slices. The rules specify that, in the case where one or more slices are organized with both uniform and non-uniform intervals, syntax elements are used to indicate the type of slice layout.

[0681] In some embodiments, a syntax element in the image parameter set indicates whether a uniform interval is followed by a non-uniform interval or a non-uniform interval is followed by a uniform interval in a piece layout. In some embodiments, the piece layout includes pieces with non-uniform intervals, wherein the number of explicitly indicated piece columns or piece rows is equal to or greater than the total number of pieces with non-uniform intervals. In some embodiments, the piece layout includes pieces with uniform intervals, and the number of explicitly indicated piece columns or piece rows is equal to or greater than the total number of pieces with uniform intervals. In some embodiments, in a piece layout where a uniform interval is followed by a non-uniform interval, the dimensions of the non-uniformly interval pieces are first determined according to a first inversion order, and then the dimensions of the uniformly interval pieces are subsequently determined according to a second inversion order. In some embodiments, the dimensions include the width of a piece column or the height of a piece row.

[0682] Figure 27 This is a flowchart representation of a video processing method according to the present technology. Method 2700 includes, in operation 2710, performing a conversion between a video frame and a video bitstream according to a rule. The rule specifies whether or how the Merge Estimated Region (MER) size is processed during the conversion depends on the minimum allowed codec block size.

[0683] In some embodiments, the MER size is equal to or greater than the minimum allowed codec block size. In some embodiments, the MER size is indicated by MER_size, the minimum allowed codec block size is indicated by MinCbSizeY, and the difference between Log2(MER_size) and Log2(MinCbSizeY) is included in the bitstream.

[0684] Figure 28This is a flowchart representation of a video processing method according to the present technology. Method 2800 includes, in operation 2810, performing a conversion between a video and a video bitstream comprising at least one video slice, according to a rule. The rule specifies the height of a stripe in the video slice in units of a codec tree unit based on the value of a first syntax element in the bitstream, the value of the first syntax element indicating the number of explicitly provided stripe heights for the stripes in the video slice including that stripe.

[0685] In some embodiments, in response to the value of the first syntax element being equal to 0, the height of the slice in the video slice, in units of codec tree units, is derived from the video information. In some embodiments, the height of the slice in the video slice, in units of codec tree units, is represented as sliceHeightInCtus[i]. sliceHeightInCtus[i] is equal to the value of row height RowHeight[SliceTopLeftTileIdx[i] / NumTileColumns], where SliceTopLeftTileIdx specifies the slice index of the slice including the first codec tree unit in the slice, NumTileColumns specifies the number of slice columns, and the value of row height RowHeight[j] specifies the height of the j-th slice row in units of codec tree blocks.

[0686] In some embodiments, the height of a stripe in a video slice, in units of codec tree units, is derived from the video information in response to the absence of a stripe height in the bitstream. In some embodiments, the height of a stripe in a video slice, in units of codec tree units, is derived from a second syntax element present in the bitstream that indicates the height of a stripe in the video slice. In some embodiments, in response to a first syntax element being equal to 0, a video slice including a stripe is not divided into multiple stripes.

[0687] Figure 29 This is a flowchart representation of a video processing method according to the present technology. Method 2900 includes, in operation 2910, performing a conversion between a video and a video bitstream comprising a video picture, the video picture comprising a video slice containing one or more stripes, according to a rule. The rule specifies that a second stripe in a slice comprising a first stripe of the picture has a height represented in units of a codec tree unit. The first stripe has a first stripe index, and the second stripe has a second stripe index, the second stripe index being determined based on the first stripe index and the number of stripe heights explicitly provided in the video slice. The height of the second stripe is determined based on the first stripe index and the second stripe index.

[0688] In some embodiments, the first stripe index is represented as i, and the second stripe index is represented as num_exp_slices_in_tile[i]-1, where num_exp_slices_in_tile specifies the number of stripe heights explicitly provided in the video slice. The height of the second stripe is determined based on exp_slice_height_in_ctus_minus1[i][num_exp_slices_in_tile[i]-1]+1, where exp_slice_height_in_ctus_minus1 specifies the stripe height in units of codec tree units in the video slice. In some embodiments, the height of the uniform stripes dividing the video slice is indicated by the last entry of exp_slice_height_in_ctus_minus1. In some embodiments, the second stripe is always present in the picture. In some embodiments, resetting the height of the second stripe for a transition is not allowed. In some embodiments, the height of the third strip with an index greater than num_exp_slices_in_tile[i]-1 is determined based on exp_slice_height_in_ctus_minus1[i][num_exp_slices_in_tile[i]-1]. In some embodiments, the second strip does not exist in the video slice. In some embodiments, the height of the second strip is less than or equal to exp_slice_height_in_ctus_minus1[i][num_exp_slices_in_tile[i]-1]+1.

[0689] Figure 30 This is a flowchart representation of a video processing method according to the present technology. Method 3000 includes, in operation 3010, performing a conversion between a video and a video bitstream comprising a video picture, the video picture comprising one or more slices. The video picture refers to a picture parameter set. The picture parameter set conforms to a format rule specifying that the picture parameter set includes a column width list of N slice columns, where N is an integer. The video picture contains an (N-1)th slice column, and the width of the (N-1)th slice column is equal to the (N-1)th entry in the explicitly included slice column width list plus a codec tree block.

[0690] In some embodiments, N-1 is represented as num_exp_tile_columns_minus1. In some embodiments, the width of a slice column in units of codec tree blocks is determined based on tile_column_width_minus1[num_exp_tile_columns_minus1]+1, where tile_column_width_minus1 specifies the width of the slice column in units of codec tree blocks. In some embodiments, the width of a slice column in units of codec tree blocks cannot be reset and is determined solely based on tile_column_width_minus1[num_exp_tile_columns_minus1]. In some embodiments, the width of a slice column in units of codec tree blocks is equal to tile_column_width_minus1[num_exp_tile_columns_minus1]+1. In some embodiments, the width of a second slice column with an index greater than num_exp_tile_columns_minus1 is determined based on tile_column_width_minus1[num_exp_tile_columns_minus1].

[0691] Figure 31 This is a flowchart representation of a video processing method according to the present technology. Method 3100 includes, in operation 3110, performing a conversion between a video and a video bitstream, including video pictures, wherein the video pictures include one or more slices. A video picture refers to a picture parameter set. The picture parameter set conforms to a format rule specifying that the picture parameter set includes a list of row heights of N slice rows, where N is an integer. There exists an (N-1)th slice row in the video picture, and the height of the (N-1)th slice row is equal to the (N-1)th entry in the explicitly included list of slice row heights plus the number of codec tree blocks.

[0692] In some embodiments, N-1 is represented as num_exp_tile_rows_minus1. In some embodiments, the height of a slice row in units of codec tree blocks is determined based on tile_row_height_minus1[num_exp_tile_row_minus1]+1, where tile_row_height_minus1 specifies the height of a slice row in units of codec tree blocks. In some embodiments, the height of a slice row in units of codec tree blocks cannot be reset and is determined solely based on tile_row_height_minus1[num_exp_tile_row_minus1]. In some embodiments, the height of a slice row in units of codec tree blocks is equal to tile_row_height_minus1[num_exp_tile_row_minus1]+1. In some embodiments, the height of a second slice row with an index greater than num_exp_tile_rows_minus1 is determined based on tile_row_height_minus1[num_exp_tile_row_minus1].

[0693] In some embodiments, the conversion includes encoding the video into a bitstream. In some embodiments, the conversion includes decoding the video from the bitstream.

[0694] In the solution described herein, the encoder conforms to the format rules by generating a codec representation based on those rules. In the solution described herein, the decoder uses the format rules to parse the syntax elements in the codec representation and determines the presence or absence of these elements to generate the decoded video.

[0695] In this document, the term "video processing" can refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm can be applied during the conversion from the pixel representation of a video to the corresponding bitstream, and vice versa. As defined in the syntax, the bitstream of the current video block can, for example, correspond to bits juxtaposed or scattered at different positions within the bitstream. For example, a macroblock can be encoded based on the residual error values ​​after transformation and encoding / decoding, and also using bits from the header and other fields in the bitstream. Furthermore, during the conversion, the decoder can parse the bitstream based on determinations as described in the solutions above, knowing that some fields may or may not be present. Similarly, the encoder can determine whether certain syntax fields are included or excluded, and generate the codec representation accordingly by including or excluding syntax fields from the codec representation.

[0696] The solutions and other solutions, examples, embodiments, modules, and functional operations disclosed in this document can be implemented in digital electronic circuits, or computer software, firmware, or hardware, including the structures disclosed in this document and their structural equivalents, or combinations of one or more of the foregoing. The disclosed embodiments and other embodiments can be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer-readable medium for execution by or control of the operation of a data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition of substances affecting machine-readable propagation signals, or one or more combinations thereof. The term "data processing apparatus" includes all means, devices, and machines for processing data, such as programmable processors, computers, or multiple processors or computers. In addition to hardware, the apparatus may also include code that creates an execution environment for the computer program in question, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination thereof. Propagation signals are artificially generated signals, such as machine-generated electrical, optical, or electromagnetic signals, which are generated to encode information for transmission to a suitable receiver device.

[0697] Computer programs (also known as programs, software, software applications, scripts, or code) can be written in any programming language, including compiled or interpreted languages, and can be deployed in any form, including as standalone programs or modules, components, subroutines, or other units suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored in a file portion that holds other programs or data (e.g., one or more scripts stored in a markup language document), a single file dedicated to a related program, or multiple coordination files (e.g., a file storing one or more modules, subroutines, or code portions). A computer program can be deployed to execute on a single computer, or on multiple computers located at one site or distributed across multiple sites and interconnected through a communication network.

[0698] The processes and logic flows described in this document can be executed by one or more programmable processors, which execute one or more computer programs to perform functions by manipulating input data and generating output. The processes and logic flows can also be executed by dedicated logic circuitry, and the devices can be implemented as dedicated logic circuitry, such as FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits).

[0699] For example, processors suitable for executing computer programs include general-purpose and special-purpose microprocessors, as well as any one or more processors in any type of digital computer. Typically, a processor receives instructions and data from read-only memory or random access memory, or both. The basic components of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include, or be operatively coupled to, receiving or transferring data to one or more mass storage devices (e.g., magnetic disks, magneto-optical disks, or optical disks) for storing data. However, a computer does not require such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM optical disks. The processor and memory may be supplemented by or incorporated into special-purpose logic circuitry.

[0700] Although this patent document contains numerous details, these details should not be construed as limiting the scope of any subject matter or claimed content, but rather as descriptions of features characteristic of specific embodiments of a particular technology. Certain features described in the context of individual embodiments in this patent document may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments. Furthermore, although the foregoing features may be described as functioning in a particular combination, or even initially claimed to be so, in some cases one or more features may be removed from the claimed combination, and the claimed combination may refer to a sub-combination or a variation of a sub-combination.

[0701] Similarly, although these operations are described in a specific order in the accompanying drawings, this should not be construed as requiring that such operations be performed in the specific order or sequence shown, or requiring that all shown operations be performed to obtain the desired result. Furthermore, the separation of various system components in the embodiments described in this patent document should not be construed as requiring such separation in all embodiments.

[0702] Only some implementations and examples are described, and other implementations, enhancements and variations may be made based on the content described and illustrated in this patent document.

Claims

1. A method of processing video data, comprising: a first conversion between a video comprising a video picture of a reference picture parameter set and a bitstream of the video, determining that a scan process is applied to the video picture, wherein the video picture is partitioned into one or more slices, one or more tiles and a plurality of coding tree units, in the first scan process, determining that the picture parameter set comprises a list of syntax elements respectively indicating slice column widths of N slice columns with N indices, wherein N is an integer, performing the first conversion based on the determining; wherein the list of syntax elements comprises a first syntax element, a value of the first syntax element directly specifies, without reference to other information, a width of an Nth slice column of the N slice columns in units of coding tree blocks, and wherein the value of the first syntax element is used to derive widths of slice columns with indices greater than the N indices; wherein a number of slice heights explicitly provided for slices in a video slice comprising rectangular tiles in the video picture is equal to M, and M is an integer not less than 0; wherein when a difference between a height of the video slice in units of coding tree blocks and a sum of slice heights of the M tiles is less than a uniform tile height, a slice height of an (M+1)th tile of the video slice is set to the difference; wherein in a second scan process, when a rectangular tile mode is used for the video picture, determining that a variable indicating a slice index of a slice containing a first coding tree unit in a tile with a picture level tile index is updated when the picture level tile index is less than a value of a fifth syntax element, and the variable is not updated when the picture level tile index is not less than the value of the fifth syntax element, wherein the fifth syntax element is included in a picture parameter set referred by the video picture in the bitstream to derive information about one or more tiles in the video picture, a value of the fifth syntax element plus one specifies a number of rectangular tiles in each video picture referring to the picture parameter set, and a number of tiles in the video picture is greater than the value of the fifth syntax element.

2. The method of claim 1, wherein, a value of N is indicated by a second syntax element included in the picture parameter set.

3. The method of claim 1, wherein, the first syntax element is an Nth entry in the list of syntax elements, and a uniform slice column width is set to a width of the Nth slice column of the N slice columns, and wherein when a difference between a picture width in units of coding tree blocks of a luma component and a sum of slice column widths of the N slice columns is not less than the width of the Nth slice column of the N slice columns, a width of an (N+1)th slice column is set to be equal to the width of the Nth slice column of the N slice columns.

4. The method of claim 1, wherein, when a difference between a picture width in units of coding tree blocks of a luma component and a sum of slice column widths of the N slice columns is less than the width of the Nth slice column of the N slice columns, a width of an (N+1)th slice column is set to be equal to the difference.

5. The method of claim 1, wherein, the picture parameter set further comprises a list of syntax elements respectively indicating slice row heights of Q slice rows with Q indices, wherein Q is an integer, wherein the list of syntax elements includes a third syntax element, a value of the third syntax element directly specifies, without reference to other information, a height of a Q-th slice row of the Q slice rows in units of coding tree blocks, and wherein the value of the third syntax element is used to derive heights of slice rows with indices greater than the Q indices.

6. The method of claim 5, wherein, a value of Q is indicated by a fourth syntax element included in the picture parameter set.

7. The method of claim 5, wherein, the third syntax element is a Q-th entry in the list of syntax elements, and a uniform slice row height is set to the height of the Q-th slice row of the Q slice rows, and wherein when a difference between a picture height of the luma component in units of coding tree blocks and a sum of slice row heights of the Q slice rows is not less than the height of the Q-th slice row of the Q slice rows, a height of a Q+1-th slice row is set to be equal to the height of the Q-th slice row of the Q slice rows.

8. The method of claim 5, wherein, when the difference between the picture height of the luma component in units of coding tree blocks and the sum of the slice row heights of the Q slice rows is less than the height of the Q-th slice row of the Q slice rows, the height of the Q+1-th slice row is set to be equal to the difference.

9. The method of claim 1, wherein, the first conversion includes encoding the video into the bitstream.

10. The method of claim 1, wherein, the first conversion includes decoding the video from the bitstream.

11. An apparatus for processing video data, comprising a processor and a non-transitory memory having instructions thereon, wherein, the instructions, when executed by the processor, cause the processor to: a first conversion between a video and a bitstream of the video, the video comprising a video picture that references a picture parameter set, determine that a scan process is applied to the video picture, wherein the video picture is partitioned into one or more slices, one or more tiles, and a plurality of coding tree units, in the first scan process, determine that the picture parameter set includes a list of syntax elements that respectively indicate slice column widths of N slice columns with N indices, where N is an integer, perform the first conversion based on the determination; wherein the list of syntax elements includes a first syntax element, a value of the first syntax element directly specifies, without reference to other information, a width of an N-th slice column of the N slice columns in units of coding tree blocks, and wherein the value of the first syntax element is used to derive widths of slice columns with indices greater than the N indices. wherein a number of slice heights explicitly provided for slices in a video slice of the video picture that includes rectangular slices is equal to M, and M is an integer that is not less than 0; wherein when a difference between a height of the video slice in units of coding tree blocks and a sum of slice heights of M slices is less than a uniform slice height, a slice height of a (M+1)-th slice of the video slice is set to be the difference; wherein in a second scan process, when a rectangular slice mode is used for the video picture, it is determined that a variable indicating a slice index of a slice containing a first coding tree unit in a slice with a picture level slice index is updated when the picture level slice index is less than a value of a fifth syntax element, and the variable is not updated when the picture level slice index is not less than the value of the fifth syntax element, wherein the fifth syntax element is included in a picture parameter set referred to by the video picture in the bitstream to derive information about one or more slices in the video picture, a value of the fifth syntax element plus one specifies a number of rectangular slices in each video picture referring to the picture parameter set, and a number of slices in the video picture is greater than the value of the fifth syntax element.

12. The apparatus of claim 11, wherein, A value of N is indicated by a second syntax element included in the picture parameter set.

13. The apparatus of claim 11, the first syntax element is an Nth entry in a list of the syntax elements, and uniform slice column width is set to a width of an Nth slice column of the N slice columns, and wherein when a difference between a picture width in units of coding tree blocks of a luma component and a sum of slice column widths of the N slice columns is not less than a width of the Nth slice column of the N slice columns, a width of an N+1th slice column is set to be equal to the width of the Nth slice column of the N slice columns.

14. A non-transitory computer-readable storage medium storing instructions to cause a processor to: A first conversion between a video of video pictures including a reference picture parameter set and a bitstream of the video, determines that a scan process is applied to the video picture, wherein, the video picture is partitioned into one or more slices, one or more slices, and a plurality of coding tree units, in a first scanning process, it is determined that the picture parameter set includes a list of syntax elements that respectively indicate slice column widths of N slice columns having N indices, where N is an integer, perform the first conversion based on the determination; wherein the list of syntax elements includes a first syntax element, a value of the first syntax element plus one directly specifies a width in units of coding tree blocks of an Nth slice column of the N slice columns without referring to other information, and wherein the value of the first syntax element is used to derive widths of slice columns having indices greater than the N indices; wherein a number of slice heights explicitly provided for slices in a video slice including rectangular slices in the video picture is equal to M, and M is an integer not less than 0; wherein when a difference between a height of a video slice in units of coding tree blocks and a sum of slice heights of M slices is less than a uniform slice height, a slice height of an (M+1)th slice of the video slice is set to be equal to the difference; wherein in a second scanning process, when a rectangular slice mode is used for the video picture, it is determined that a variable indicating a slice index of a slice containing a first coding tree unit in a slice having a picture level slice index is updated when the picture level slice index is less than a value of a fifth syntax element, and the variable is not updated when the picture level slice index is not less than the value of the fifth syntax element, wherein the fifth syntax element is included in a picture parameter set referred to by the video picture in the bitstream to derive information about one or more slices in the video picture, a value of the fifth syntax element plus one specifies a number of rectangular slices in each video picture referring to the picture parameter set, and a number of slices in the video picture is greater than the value of the fifth syntax element.

15. The non-transitory computer-readable storage medium of claim 14, wherein, A value of N is indicated by a second syntax element included in the picture parameter set.

16. A method of storing a bitstream of a video, comprising: for a video of a video picture comprising a picture parameter set, determining that a scan process is applied to the video picture, wherein the video picture is partitioned into one or more slices, one or more tiles, and a plurality of coding tree units, in a first scan process, determining that the picture parameter set comprises a list of syntax elements respectively indicating slice column widths of N slice columns with N indices, where N is an integer, generating the bitstream based on the determining; and storing the bitstream into a non-transitory computer-readable storage medium; wherein the list of syntax elements comprises a first syntax element, a value of the first syntax element directly specifies, without reference to other information, a width in coding tree blocks of an Nth slice column of the N slice columns with the value of the first syntax element plus 1, and wherein the value of the first syntax element is used to derive widths of slice columns with indices greater than the N indices; wherein a number of slice heights explicitly provided for slices in a video slice comprising rectangular tiles in the video picture is equal to M, and M is an integer not less than 0; wherein when a difference between a height in coding tree blocks of the video slice and a sum of slice heights of M tiles is less than a uniform tile height, a slice height of an (M+1)th tile of the video slice is set to the difference; wherein in a second scan process, when a rectangular tile mode is used for the video picture, determining that a variable indicating a slice index of a slice containing a first coding tree unit in a tile with a picture level tile index is updated when the picture level tile index is less than a value of a fifth syntax element, and the variable is not updated when the picture level tile index is not less than the value of the fifth syntax element, wherein the fifth syntax element is included in a picture parameter set referenced by the video picture in the bitstream to derive information about one or more tiles in the video picture, a value of the fifth syntax element plus 1 specifies a number of rectangular tiles in each video picture referencing the picture parameter set, and a number of tiles in the video picture is greater than the value of the fifth syntax element.