Signaling for Wavefront Parallel Processing

By adopting wavefront parallel processing technology in video encoding and decoding, the picture is divided into CTB rows and parallel decoding, the contradiction between parallel processing and MTU size matching in the prior art is solved, and the encoding and decoding efficiency and resource utilization of multi-core processors are improved.

CN114930830BActive Publication Date: 2025-08-29DOUYIN CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202180008704.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-01-09
Filing Date
2021-01-08
Publication Date
2025-08-29
Estimated Expiration
2041-01-08

AI Technical Summary

Technical Problem

The existing video codec standards have contradictions in parallel processing and MTU size matching, resulting in large overhead for encoding and codec and difficulty in effectively utilizing multi-core processor resources.

Method used

Wavefront Parallel Processing (WPP) technology is used to segment the pictures into CTB rows in parallel decoding, allowing intra-picture prediction and entropy decoding between CTB rows, reducing inter-processor communication, and signaling the layout information of the stripes and slices to optimize processor utilization.

Benefits of technology

It realizes more efficient video encoding and decoding, reduces codec overhead, improves resource utilization of multi-core processors, and reduces end-to-end latency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114930830B_ABST
    Figure CN114930830B_ABST
Patent Text Reader

Abstract

Methods, apparatus, and systems for video coding and encoding are described, including constraints, restrictions, and signaling for sub-pictures, slices, and slices. An example video processing method includes performing conversion between a picture of a video and a bitstream of the video, wherein the bitstream conforms to a format rule, wherein the format rule specifies a first syntax element that enables (a) synchronization of context variables before decoding a codec tree unit (CTU) in the picture and (b) storage of the context variables after decoding the CTU, wherein the first syntax element is signaled in a sequence parameter set (SPS) associated with the picture.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to and the benefit of U.S. Provisional Patent Application No. 62 / 959,108, filed on January 9, 2020, in a timely manner, under applicable patent laws and / or the Paris Convention. The entire disclosure of the foregoing application is incorporated herein by reference and made a part of the disclosure of this application for all legal purposes. Technical Field

[0003] This application document relates to image and video encoding and decoding. Background Art

[0004] Digital video accounts for the largest use of bandwidth on the Internet and other digital communications networks. As the number of connected user devices capable of receiving and displaying video increases, bandwidth demand for digital video usage is expected to continue to grow. Summary of the Invention

[0005] This document discloses a video encoder and decoder for video encoding and decoding, respectively, and includes constraints, limitations and signaling for sub-pictures, slices and slices.

[0006] In one example aspect, a video processing method is disclosed. The method includes performing conversion between a video and a bitstream of the video, wherein the bitstream includes one or more access units according to a format rule, and wherein the format rule specifies an order in which a first message and a second message for an operation point (OP) appear in the access unit (AU) such that the first message precedes the second message in decoding order.

[0007] In another example aspect, a video processing method is disclosed. The method includes performing conversion between a video and a bitstream of the video, wherein the bitstream includes one or more access units according to a format rule, and wherein the format rule specifies an order in which a plurality of messages for an operation point (OP) appear in the access units such that a first message of the plurality of messages precedes a second message of the plurality of messages in decoding order.

[0008] In another example aspect, a video processing method is disclosed. The method includes performing conversion between a video including a picture and a bitstream of the video, wherein the bitstream conforms to a format rule, wherein the format rule specifies whether to signal an indication of a first flag at the beginning of a picture header associated with the picture, wherein the first flag indicates whether the picture is an intra random access point (IRAP) picture or a gradual decoding refresh (GDR) picture.

[0009] In yet another example aspect, a video processing method is disclosed. The method includes performing conversion between a video comprising one or more pictures and a bitstream of the video, wherein the bitstream conforms to a format rule, wherein the format rule does not permit a picture of the one or more pictures to be coded to include a codec slice network abstraction layer (NAL) unit having a progressive decoding refresh type and associated with a flag indicating that the picture includes the NAL unit of a hybrid type.

[0010] In yet another example aspect, a video processing method is disclosed. The method includes performing conversion between a video comprising one or more pictures and a bitstream of the video, wherein the bitstream conforms to a format rule, wherein the format rule allows a picture of the one or more pictures to be encoded to include a codec slice network abstraction layer (NAL) unit having a progressive decoding refresh type and associated with a flag indicating that the picture does not include the NAL unit of a hybrid type.

[0011] In yet another example aspect, a video processing method is disclosed. The method includes performing conversion between a picture of a video and a bitstream of the video, wherein the bitstream conforms to a format rule, the format rule specifying whether to signal a first syntax element in a picture parameter set (PPS) associated with the picture, wherein the picture includes one or more slices having a slice type, wherein the first syntax element indicates that the slice type is signaled in a picture header due to the first syntax element being equal to 0, otherwise the first syntax element indicates that the slice type is signaled in a slice header.

[0012] In yet another example aspect, a video processing method is disclosed, the method comprising: performing conversion between a picture of a video and a bitstream of the video according to a rule, wherein the conversion includes a loop filtering process, and wherein the rule specifies that a total number of vertical virtual boundaries and a total number of horizontal virtual boundaries associated with the loop filtering operation be signaled at a picture level or a sequence level.

[0013] In yet another example aspect, a video processing method is disclosed. The method includes performing conversion between a video comprising one or more pictures and a bitstream of the video, wherein the bitstream conforms to a format rule, wherein the format rule conditionally allows encoding and decoding of a picture in a layer by using a reference picture from another layer based on a first syntax element, the first syntax element indicating whether the reference picture from the other layer is present in the bitstream, and wherein the first syntax element is conditionally signaled in the bitstream based on a second syntax element indicating whether an identifier of a parameter set associated with the picture is not equal to 0.

[0014] In another example aspect, a video processing method is disclosed. The method includes performing conversion between a picture of a video and a bitstream of the video, wherein the bitstream conforms to a format rule, wherein the format rule specifies a first syntax element that causes (a) synchronization of context variables before decoding a codec tree unit (CTU) in the picture and (b) storage of the context variables after decoding the CTU, wherein the first syntax element is signaled in a sequence parameter set (SPS) associated with the picture.

[0015] In yet another example aspect, a video processing method is disclosed. The method includes performing conversion between a picture of a video and a bitstream of the video, wherein the bitstream conforms to a format rule, wherein the format rule specifies a syntax element indicating whether an entry point offset for a slice or a slice-specific codec tree unit (CTU) row is signaled in a slice header of the picture, and wherein the syntax element is signaled in a sequence parameter set (SPS) associated with the picture.

[0016] In yet another example aspect, a video processing method is disclosed, the method comprising: performing conversion between a video and a bitstream of the video according to a rule, wherein the rule stipulates that a first syntax element indicating a number of parameters for an output layer set (OLS) hypothesized reference decoder (HRD) in a video parameter set (VPS) associated with the video is less than a first preset threshold.

[0017] In yet another example aspect, a video processing method is disclosed. The method includes performing conversion between a video and a bitstream of the video according to a rule, wherein the rule stipulates that a syntax element indicating a number of profile / tier / level (PTL) syntax structures in a video parameter set (VPS) associated with the video is less than a preset threshold.

[0018] In yet another example aspect, a video processing method is disclosed, the method comprising: performing conversion between a video and a bitstream of the video according to a rule, wherein the rule specifies a first syntax element indicating that a number of decoded picture buffer parameter syntax structures in a video parameter set (VPS) must be less than or equal to a second syntax element indicating a number of layers specified by the VPS.

[0019] In yet another example aspect, a video processing method is disclosed, comprising: performing conversion between a video and a bitstream of the video according to a rule, wherein the rule allows a decoder to obtain a terminal Network Abstraction Layer (NAL) unit by signaling in the bitstream or providing it by external means.

[0020] In yet another example aspect, a video processing method is disclosed. The method includes performing conversion between a video and a bitstream of the video, wherein the bitstream conforms to a format rule, and wherein the format rule restricts the bitstream to include only one sub-picture per layer due to a syntax element being equal to 0, indicating that each layer is configured to use inter-layer prediction.

[0021] In yet another example aspect, a video processing method is disclosed. The method includes performing conversion between a video and a bitstream of the video according to a rule, wherein the rule provides for performing a sub-bitstream extraction process to generate a sub-bitstream for decoding, wherein the sub-bitstream extraction process is configured to extract a sub-bitstream having a target highest temporal identifier from the bitstream, and wherein the rule provides for removing, during the extraction process, a padding data unit and a padding supplemental enhancement information (SEI) message in a video codec layer (VCL) network abstraction layer (NAL) unit associated with the VCL NAL unit when the VCL NAL unit is removed.

[0022] In yet another example aspect, a video processing method is disclosed, the method comprising: performing conversion between a video unit of a video and a bitstream of the video, wherein the bitstream conforms to a format rule, wherein the format rule specifies that the bitstream includes a first syntax element, the first syntax element indicating whether the video unit is encoded or decoded in a lossy mode or a lossless mode, and wherein signaling a second syntax element indicates selective inclusion of escaped samples in a palette mode applied to the video unit based on a value of the first syntax element.

[0023] In another exemplary aspect, a video encoding apparatus is disclosed, wherein the video encoding apparatus includes a processor, wherein the processor is configured to execute the above method.

[0024] In another exemplary aspect, a video decoding apparatus is disclosed, wherein the video decoding apparatus includes a processor, wherein the processor is configured to execute the above method.

[0025] In yet another exemplary aspect, a computer-readable medium having code stored thereon is disclosed. The codec implements the above method in the form of processor-executable code.

[0026] These and other features are described in this document. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 An example of partitioning a picture using luma codec tree units (CTUs) is shown.

[0028] Figure 2Another example of partitioning a picture using luma CTUs is shown.

[0029] Figure 3 An example of picture segmentation is shown.

[0030] Figure 4 Another example of picture segmentation is shown.

[0031] Figure 5 is a block diagram of an example video processing system in which the disclosed technology may be implemented.

[0032] Figure 6 is a block diagram of an example hardware platform for video processing.

[0033] Figure 7 is a block diagram illustrating a video encoding and decoding system according to some embodiments of the present disclosure.

[0034] Figure 8 is a block diagram illustrating an encoder according to some embodiments of the present disclosure.

[0035] Figure 9 is a block diagram illustrating a decoder according to some embodiments of the present disclosure.

[0036] Figure 10-26 A flow chart illustrating an example method of video processing is shown. DETAILED DESCRIPTION

[0037] Throughout this document, section headings are used for ease of understanding and do not limit the applicability of the techniques and embodiments disclosed in each section to that section. Furthermore, the use of H.266 terminology in some descriptions is for ease of understanding and is not intended to limit the scope of the disclosed techniques. Therefore, the techniques described herein are also applicable to other video codec protocols and designs.

[0038] 1. Introduction

[0039] This article relates to video codec technology. Specifically, it discusses the signaling of sub-pictures, slices, and stripes. These concepts can be applied, alone or in various combinations, to any video codec standard or non-standard video codec that supports multi-layer video codecs, such as the Versatile Video Codec (VVC) under development.

[0040] 2. Abbreviation

[0041] APS adaptive parameter set

[0042] AU access unit

[0043] AUD Access Unit Delimiter

[0044] AVC Advanced Video Codec

[0045] CLVS codec layer video sequence

[0046] CPB codec picture buffer

[0047] CRA Clear Random Access

[0048] CTU Codec Tree Unit

[0049] CVS codec video sequence

[0050] DPB decoded picture buffer

[0051] DPS decoding parameter set

[0052] EOB End of bitstream

[0053] EOS sequence ends

[0054] GDR Progressive Decode Refresh

[0055] HEVC High-Efficiency Video Codec

[0056] HRD Virtual Reference Decoder

[0057] IDR instant decoding refresh

[0058] JEM Joint Exploration Model

[0059] MCTS motion constraint set

[0060] NAL Network Abstraction Layer

[0061] OLS output layer set

[0062] PH picture header

[0063] PPS picture parameter set

[0064] PTL Profiles, Tiers, and Levels

[0065] PU picture unit

[0066] RBSP Raw Byte Sequence Payload

[0067] SEI Supplemental Enhancement Information

[0068] SPS sequence parameter set

[0069] SVC Scalable Video Codec

[0070] VCL video codec layer

[0071] VPS Video Parameter Set

[0072] VTM VVC test model

[0073] VUI Video Availability Information

[0074] VVC multifunctional video codec

[0075] 3. Preliminary Discussion

[0076] Video codec standards have primarily evolved through the development of the well-known ITU-T and ISO / IEC standards. ITU-T developed H.261 and H.263, ISO / IEC developed MPEG-1 and MPEG-4 Video, and the two organizations jointly developed H.262 / MPEG-2 Video, H.264 / MPEG-4 Advanced Video Codec (AVC), and H.265 / HEVC. Since H.262, video codec standards have been based on a hybrid video codec architecture that utilizes temporal prediction plus transform coding. To explore future video codec technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Team (JVET) in 2015. Since then, JVET has adopted many new methods and implemented them in reference software called the Joint Exploration Model (JEM). JVET meetings are held simultaneously every quarter, and the goal of the new codec standard is to reduce the bit rate by 50% compared to HEVC. The new video codec standard was officially named Versatile Video Coding (VVC) at the JVET meeting in April 2018, with the release of the first version of the VVC Test Model (VTM). With the ongoing progress of VVC standardization, new codec technologies are adopted into the VVC standard at each JVET meeting. The VVC working draft and test model (VTM) are updated after each meeting. The VVC project is currently aiming for technical completion (FDIS) at the July 2020 meeting.

[0077] 3.1. Image Segmentation Scheme in HEVC

[0078] HEVC includes four different picture partitioning schemes, namely regular slices, dependent slices, tiles, and wavefront parallel processing (WPP), which can be applied to maximum transmission unit (MTU) size matching, parallel processing, and reduced end-to-end delay.

[0079] Regular slices are similar to those in H.264 / AVC. Each regular slice is encapsulated in its own NAL unit, and intra-picture prediction (intra sample prediction, motion information prediction, codec mode prediction) and entropy codec dependencies are disabled across slice boundaries. Therefore, regular slices can be reconstructed independently of other regular slices in the same picture (although interdependencies may still exist due to loop filtering operations).

[0080] Regular slices are the only tool available for parallelization, and are also available in H.264 / AVC in a nearly identical form. Regular slice-based parallelization does not require much inter-processor or inter-core communication (except for inter-processor or inter-core data sharing for motion compensation when decoding predictive codec pictures, which is typically much more expensive than inter-processor or inter-core data sharing due to intra-picture prediction). However, for the same reasons, using regular slices can incur significant codec overhead due to the bit cost of the slice header and the loss of prediction across slice boundaries. Furthermore, due to their intra-picture independence and the fact that each regular slice is encapsulated in its own NAL, regular slices (compared to the other tools mentioned below) can also serve as a key mechanism for bitstream segmentation to match MTU size requirements. In many cases, the goals of parallelization and MTU size matching conflict with respect to slice layout within a picture. The realization of such situations led to the development of the parallelization tools mentioned below.

[0081] Dependent slices have short slice headers and allow the bitstream to be partitioned at treeblock boundaries without breaking any intra-picture prediction. Essentially, dependent slices split a regular slice into multiple NAL units, reducing end-to-end latency by allowing part of a regular slice to be sent before coding of the entire regular slice is complete.

[0082] In WPP, a picture is partitioned into a single row of codec treeblocks (CTBs). This allows entropy decoding and prediction to use data from CTBs in other partitions. Parallel processing is enabled by parallel decoding of CTB rows, where the start of decoding a CTB row is delayed by two CTBs to ensure that data associated with CTBs above and to the right of the subject CTB is available before the subject CTB being decoded. This staggered start (graphically represented as a wavefront) allows parallelization across as many processors / cores as there are CTB rows in a picture. Because intra-picture prediction is allowed between adjacent treeblock rows within a picture, the inter-processor / inter-core communication required to implement intra-picture prediction can be substantial. WPP partitioning does not result in the generation of additional NAL units compared to when WPP partitioning was not applied, so WPP is not a tool for MTU size matching. However, if MTU size matching is required, regular slices can be used with WPP, albeit with some codec overhead.

[0083] Slices define the horizontal and vertical boundaries that divide an image into slice columns and slice rows. Slice columns extend from the top to the bottom of the image. Similarly, slice rows extend from the left to the right of the image. The number of slices in an image can be simply calculated by multiplying the number of slice columns by the number of slice rows.

[0084] Before decoding the top left CTB of the next slice in the order of the slice raster scan of one picture, the CTB scan order is changed to the local scan order within the slice (in the order of the slice's CTB raster scan). Similar to regular slices, slices break intra-picture prediction dependencies and entropy decoding dependencies. However, they do not need to be contained in separate NAL units (same as WPP in this respect); therefore, slices cannot be used for MTU size matching. Each slice can be processed by one processor / core, and when a slice spans multiple slices, the inter-processor / inter-core communication required to decode intra-picture prediction between processing units of adjacent slices is limited to the transmission of a shared slice header and loop filtering related to the sharing of reconstruction samples and metadata. When a slice contains more than one slice or WPP segment, the entry point byte offset of each slice or WPP segment in the slice except the first is signaled in the slice header.

[0085] For simplicity, HEVC specifies restrictions on the application of four different picture partitioning schemes. For most profiles specified in HEVC, a given codec video sequence cannot contain both slices and wavefronts. For each slice and slice, one or both of the following conditions must be met: 1) all codec tree blocks in a slice belong to the same slice; 2) all codec tree blocks in a slice belong to the same slice. Finally, a wavefront segment contains exactly one CTB row, and when using WPP, if a slice starts within a CTB row, it must end in the same CTB row.

[0086] The latest revision of HEVC is specified in the JCTVC output document JCTVC-AC1005, "HEVC Additional Supplemental Enhancement Information (Draft 4)", published on October 24, 2017 at http: / / phenix.intevry.fr / jct / doc_end_user / documents- / 29_Macau / wg11 / JCTVC-AC1005-v2.zip. Within this revision, HEVC specifies three SEI messages related to MCT: the Time-Domain MCTS SEI message, the MCTS Extraction Information Set SEI message, and the MCTS Extraction Information Nesting SEI message.

[0087] The temporal MCTSs SEI message indicates the presence of MCTSs in the bitstream and signals the MCTSs. For each MCTS, motion vectors are restricted to pointing to full sample positions within the MCTS and fractional sample positions that only require full sample positions within the MCTS for interpolation, and motion vector candidates predicted by temporal motion vectors from blocks outside the MCTS are not allowed. In this way, each MCTS can be independently decoded in the absence of slices not included in the MCTS.

[0088] The MCTSs extraction information set SEI message provides supplementary information that can be used for MCTS sub-bitstream extraction (specified as part of the semantics of the SEI message) to generate a bitstream that conforms to the MCTS set. The information consists of multiple extraction information sets, each of which defines multiple MCTS sets and contains RBSP bytes for replacement VPSs, SPSs, and PPSs to be used in the MCTS sub-bitstream extraction process. When extracting a sub-bitstream according to the MCTS sub-bitstream extraction process, the parameter sets (VPSs, SPSs, and PPSs) need to be rewritten or replaced, and the slice header needs to be slightly updated because one or all of the slice address-related syntax elements (including first_slice_segment_in_pic_flag and slice_segment_address) usually need to have different values.

[0089] 3.2. Image Segmentation in VVC

[0090] In VVC, a picture is divided into one or more slice rows and one or more slice columns. A slice is a sequence of CTUs that covers a rectangular area of ​​the picture. The CTUs in a slice are scanned in raster scan order within the slice.

[0091] A slice consists of an integer number of complete slices or an integer number of consecutive complete CTU rows within a picture.

[0092] Two striping modes are supported: raster scan striping mode and rectangular striping mode. In raster scan striping mode, a stripe contains a complete sequence of stripes in a stripe raster scan of a picture. In rectangular striping mode, a stripe contains multiple complete slices that together form a rectangular area of ​​the picture, or multiple consecutive complete CTU rows that together form a slice of a rectangular area of ​​the picture. Stripes within a rectangular stripe are scanned in stripe raster scan order within the rectangular area corresponding to the stripe.

[0093] A sub-picture consists of one or more strips that together cover a rectangular area of ​​the picture.

[0094] Figure 1 An example of raster scan striping partitioning of a picture is shown, where the picture is partitioned into 12 slices and 3 raster scan strips.

[0095] Figure 2 An example of rectangular strip partitioning of a picture is shown, where the picture is partitioned into 24 slices (6 slice columns and 4 slice rows) and 9 rectangular slices.

[0096] Figure 3 An example of a picture partitioned into slices and rectangular strips is shown, where the picture is partitioned into 4 slices (2 slice columns and 2 slice rows) and 4 rectangular strips.

[0097] Figure 4 An example of sub-picture partitioning of a picture is shown, where the picture is partitioned into 18 slices, 12 slices on the left, each covering a strip of 4x4 CTUs, and 6 slices on the right, each covering 2 vertically stacked strips of 2x2 CTUs, resulting in a total of 24 strips and 24 sub-pictures of different dimensions (each strip is a sub-picture).

[0098] 3.3. Signaling Notification of VVC Sub-Pictures, Slices, and Stripes

[0099] In the latest VVC draft text, sub-picture information is signaled in the SPS. The sub-picture information includes the sub-picture layout (i.e., the number of sub-pictures per picture and the position and size of each picture) and other sequence-level sub-picture information. The order of sub-pictures signaled in the SPS defines the sub-picture index. The sub-picture ID list of each sub-picture can be explicitly signaled in the SPS or PPS, for example.

[0100] Slices in VVC are conceptually the same as in HEVC, ie, each picture is partitioned into slice columns and slice rows, but have a different syntax for signaling slices in the PPS.

[0101] In VVC, the slice mode is also signaled in the PPS. When the slice mode is rectangular slice mode, the slice layout of each picture (i.e., the number of slices per picture and the position and size of each slice) is signaled in the PPS. The order of the rectangular slices within a picture signaled in the PPS defines the picture level slice index. The sub-picture level slice index is defined as the order of the slices within a sub-picture in ascending order of their picture level slice indices. The position and size of the rectangular slices are sent / derived based on the sub-picture position and size signaled in the SPS (when each sub-picture contains only one slice), or based on the slice position and size signaled in the PPS (when the sub-picture may contain multiple slices). When the slice mode is raster scan strip mode, similar to HEVC, the slice layout within the picture is signaled in different details in the slices themselves.

[0102] The SPS, PPS and slice header and semantics in the latest VVC draft text that is most relevant to the present invention are as follows.

[0103] 7.3.2.3 Sequence Parameter Set RBSP Syntax

[0104]

[0105]

[0106] 7.4.3.3 Sequence Parameter Set RBSP Semantics ...

[0108] subpics_present_flag equal to 1 indicates that the sub-picture parameters are present in the SPS RBSP syntax. subpics_present_flag equal to 0 indicates that the sub-picture parameters are not present in the SPS RBSP syntax.

[0109] NOTE 2 – When the bitstream is the result of a sub-bitstream extraction process and contains only a subset of the sub-pictures of the input bitstream of the sub-bitstream extraction process, it may be necessary to set the value of subpics_present_flag in the RBSP of the SPS equal to 1.

[0110] sps_num_subpics_minus1 plus 1 refers to the number of sub-pictures. sps_num_subpics_minus1 should be in the range of 0 to 254. When not present, the value of sps_num_subpics_minus1 is inferred to be equal to 0.

[0111] subpic_ctu_top_left_x[i] refers to the horizontal position of the top left corner CTU of the i-th sub-picture in units of CtbSizeY. The length of the syntax element is Ceil(Log2(pic_width_max_in_luma_samples / CtbSizeY)) bits. When it is not present, the value of subpic_ctu_top_left_x[i] is inferred to be equal to 0.

[0112] subpic_ctu_top_left_y[i] specifies the vertical position of the top-left corner of the CTU of the i-th sub-picture in units of CtbSizeY. The length of this syntax element is Ceil(Log2(pic_height_max_in_luma_samples / CtbSizeY)) bits. When it is not present, the value of subpic_ctu_top_left_y[i] is inferred to be zero.

[0113] subpic_width_minus1[i] plus 1 specifies the width of the i-th sub-picture in units of CtbSizeY. The length of this syntax element is Ceil(Log2(pic_width_max_in_luma_samples / CtbSizeY)) bits. When not present, the value of subpic_width_minus1[i] is inferred to be equal to Ceil(pic_width_max_in_luma_samples / CtbSizeY)-1.

[0114] subpic_height_minus1[i] plus 1 specifies the height of the i-th sub-picture in units of CtbSizeY. The length of this syntax element is Ceil(Log2(pic_height_max_in_luma_samples / CtbSizeY)) bits. When not present, the value of subpic_height_minus1[i] is inferred to be equal to Ceil(pic_height_max_in_luma_samples / CtbSizeY)-1.

[0115] subpic_treated_as_pic_flag[i] equal to 1 indicates that the i-th subpicture of each codec picture in the CLVS is treated as a picture in the decoding process except for loop filtering operations. subpic_treatment_as_pic_flag[i] equal to 0 indicates that the i-th subpicture of each codec picture in the CLVS is not treated as a picture in the decoding process except for loop filtering operations. When it is not present, the value of subpic_treatment_as_pic_flag[i] is inferred to be 0.

[0116] loop_filter_across_subpic_enabled_flag[i] is equal to 1, which means that loop filtering operations can be performed across the boundary of the i-th sub-picture in each codec picture in the CLVS. loop_filter_across_subpic_enabled_flag[i] is equal to 0, which means that loop filtering operations are not performed across the boundary of the i-th sub-picture in each codec picture in the CLVS. When it is not present, the value of loop_filter_across_subpic_enabled_pic_flag[i] is inferred to be equal to 1.

[0117] The following constraints apply to bitstream conformance requirements:

[0118] For any two sub-pictures subpicA and subpicB, when the sub-picture index of subpicA is less than the sub-slice index of subpicB, any codec slice NAL unit of subPicA shall precede any codec slice NAL unit of subPicB in decoding order.

[0119] The shape of the sub-pictures should be such that, when each sub-picture is decoded, its entire left and upper borders consist of either picture boundaries or the boundaries of previously decoded sub-pictures.

[0120] sps_subpic_id_present_flag equal to 1 indicates that the sub-picture ID map is present in the SPS. sps_subpic_id_present_flag equal to 0 indicates that the sub-picture ID map is not present in the SPS.

[0121] sps_subpic_id_signalling_present_flag equal to 1 indicates that the sub-picture ID mapping is signaled in the SPS. sps_subpic_id_signalling_present_flag equal to 0 indicates that the sub-picture ID mapping is not signaled in the SPS. When not present, the value of sps_subpic_id_signalling_present_flag is inferred to be equal to 0.

[0122] sps_subpic_id_len_minus1 plus 1 refers to the number of bits used to represent the syntax element sps_subpic_id[i]. The value of sps_subpic_id_len_minus1 shall be in the range of 0 to 15, inclusive.

[0123] sps_subpic_id[i] refers to the sub-picture ID of the i-th sub-picture. The length of the sps_subpic_id[i] syntax element is sps_subpic_id_len_minus1+1 bits. When not present, and when sps_subpic_id_present_flag is equal to 0, the value of sps_subpic_id[i] is inferred to be equal to i, for each i in the range from 0 to sps_num_subpics_minus1, inclusive. ...

[0125] 7.3.2.4 Picture Parameter Set RBSP Syntax

[0126]

[0127]

[0128]

[0129] 7.4.3.4 Picture Parameter Set RBSP Semantics ...

[0131] pps_subpic_id_signalling_present_flag equal to 1 indicates that the sub-picture ID mapping is signaled in the PPS. pps_subpic_id_signalling_present_flag equal to 0 indicates that the sub-picture ID mapping is not signaled in the PPS. When sps_subpic_id_present_flag is 0 or sps_subpic_id_signalling_present_flag is equal to 1, pps_subpic_id_signalling_present_flag shall be equal to 0.

[0132] pps_num_subpics_minus1 plus 1 refers to the number of sub-pictures in the coded picture that references the PPS.

[0133] It is a bitstream conformance requirement that the value of pps_num_subpic_minus1 shall be equal to sps_num_subpics_minus1.

[0134] pps_subpic_id_len_minus1 plus 1 refers to the number of bits used to represent the syntax element pps_subpic_id[i]. The value of pps_subpic_id_len_minus1 shall be in the range of 0 to 15, inclusive.

[0135] The value of pps_subpic_id_len_minus1 for codec pictures referenced in CLVS shall be the same for all PPSs, which is a requirement for bitstream conformance.

[0136] pps_subpic_id[i] refers to the sub-picture ID of the i-th sub-picture. The length of the pps_subpic_id[i] syntax element is pps_subpic_id_len_minus1+1 bits.

[0137] no_pic_partition_flag equal to 1 indicates that picture partitioning is not applied to each picture referencing a PPS. no_pic_partition_flag equal to 0 indicates that each picture referencing a PPS may be partitioned into multiple slices or slices.

[0138] The value of no_pic_partition_flag shall be the same for all PPSs referenced by a codec picture within a CLVS, which is a requirement for bitstream conformance.

[0139] When the value of sps_num_subpics_minus1+1 is greater than 1, the value of no_pic_partition_flag cannot be equal to 1, which is a requirement for bitstream consistency.

[0140] pps_log2_ctu_size_minus5 plus 5 refers to the luma codec treeblock size of each CTU. pps_log2_ctu_size_minus5 should be equal to sps_log2_ctu_size_minus5.

[0141] num_exp_tile_columns_minus1 plus 1 refers to the number of explicitly provided tile column widths. The value of num_exp_tile_columns_minus1 shall be in the range of 0 to PicWidthInCtbsY-1, inclusive. When no_pic_partition_flag is equal to 1, the value of num_exp_tile_columns_minus1 is inferred to be equal to 0.

[0142] num_exp_tile_rows_minus1 plus 1 refers to the number of explicitly provided tile row heights. The value of num_exp_tile_rows_minus1 shall be in the range of 0 to PicHeightInCtbsY-1, inclusive. When no_pic_partition_flag is equal to 1, the value of num_exp_tile_columns_minus1 is inferred to be equal to 0.

[0143] tile_column_width_minus1[i] plus 1 refers to the width of the i-th tile column in the unit of CTBs, i in the range of 0 to num_exp_tile_columns_minus1-1, inclusive. tile_column_width_minus1[num_exp_tile_columns_minus1] is used to derive the width of tile columns with indices greater than or equal to num_exp_tile_columns_minus1, as described in clause 6.5.1. When it is not present, tile_column_width_minus1[0] is inferred to be equal to PicWidthInCtbsY-1.

[0144] tile_row_height_minus1[i] plus 1 refers to the height of the i-th tile row in units of CTBs, where i is in the range of 0 to num_exp_tile_rows_minus1-1, inclusive. tile_row_height_minus1[num_exp_tile_rows_minus1] is used to derive the height of tile rows with indices greater than or equal to num_exp_tile_rows_minus1, as described in clause 6.5.1. When it is not present, the value of tile_row_height_minus1[0] is inferred to be equal to PicHeightInCtbsY-1.

[0145] rect_slice_flag equal to 0 means that the slices within each slice are arranged in raster scan order and the slice information is not signaled in the PPS. rect_slice_flag equal to 1 means that the slices within each slice cover a rectangular area of ​​the picture and the slice information is signaled in the PPS. When it is not present, rect_slice_flag is inferred to be equal to 1. When subpics_present_flag is equal to 1, the value of rect_slice_flag shall be equal to 1.

[0146] single_slice_per_subpic_flag equal to 1 indicates that each sub-picture consists of one and only one rectangular slice. single_slice_per_subpic_flag equal to 0 indicates that each sub-picture may contain one or more rectangular slices. When subpics_present_flag is equal to 0, single_slice_per_subpic_flag shall be equal to 0. When single_slice_per_subpic_flag is equal to 1, num_slices_in_pic_minus1 is inferred to be equal to sps_num_subpics_minus1.

[0147] num_slices_in_pic_minus1 plus 1 refers to the number of rectangular slices in each picture that reference the PPS. The value of num_slices_in_pic_minus1 shall be in the range of 0 to MaxSlicesPerPicture-1, inclusive, where MaxSlicesPerPicture is specified in Appendix A. When no_pic_partition_flag is equal to 1, the value of num_slices_in_pic_minus1 is inferred to be equal to 0.

[0148] tile_idx_delta_present_flag equal to 0 means that tile_idx_delta values ​​are not present in the PPS, and all rectangular slices in pictures referencing the PPS are specified in raster order according to the procedure defined in clause 6.5.1. tile_idx_delta_present_flag equal to 1 means that tile_idx_delta values ​​may be present in the PPS, and all rectangular slices in pictures referencing the PPS are specified in the order indicated by the tile_idx_delta values.

[0149] slice_width_in_tiles_minus1[i] plus 1 specifies the width of the i-th rectangular strip in cells of tile columns. The value of slice_width_in_tiles_minus1[i] shall be in the range of 0 to NumTileColumns-1, inclusive. When it is not present, the value of slice_width_in_tiles_minus1[i] shall be inferred as specified in Section 6.5.1.

[0150] slice_height_in_tiles_minus1[i] plus 1 refers to the height of the i-th rectangular slice in units of tile rows. The value of slice_height_in_tiles_minus1[i] shall be in the range of 0 to NumTileRows-1, inclusive. When it is not present, the value of slice_height_in_tiles_minus1[i] shall be inferred as specified in Section 6.5.1.

[0151] num_slices_in_tile_minus1[i] plus 1 refers to the number of slices in the current slice, applicable when the i-th slice contains a subset of CTU rows from a single slice. The value of num_slices_in_tile_minus1[i] should be in the range of 0 to RowHeight[tileY]-1, inclusive, where tileY is the index of the slice row containing the i-th slice. When it is not present, the value of num_slices_in_tile_minus1[i] is inferred to be equal to 0.

[0152] slice_height_in_ctu_minus1[i] plus 1 refers to the height of the i-th rectangular slice in units of CTU rows, applicable to the case where the i-th slice contains a subset of CTU rows from a single slice. The value of slice_height_in_ctu_minus1[i] should be in the range of 0 to RowHeight[tileY]-1, inclusive, where tileY is the index of the slice row containing the i-th slice.

[0153] tile_idx_delta[i] refers to the tile index difference between the i-th rectangular strip and the (i+1)-th rectangular strip. The value of tile_idx_delta[i] should be in the range of –NumTilesInPic+1 to NumTilesInPic-1, inclusive. When it is not present, the value of tile_idx_delta[i] is inferred to be equal to 0. In all other cases, the value of tile_idx_delta[i] should not be equal to 0.

[0154] When loop_filter_across_tiles_enabled_flag is set to 1, loop filtering operations can be performed across slice boundaries in pictures that reference the PPS. When loop_filter_across_tiles_enabled_flag is set to 0, loop filtering operations are not performed across slice boundaries in pictures that reference the PPS. Loop filtering operations include deblocking filter, sample adaptive offset filter, and adaptive loop filter operations. When loop_filter_across_tiles_enabled_flag is not present, the value of loop_filter_across_tiles_enabled_flag is inferred to be 1.

[0155] When loop_filter_across_slices_enabled_flag is set to 1, loop filtering operations can be performed across slice boundaries in pictures that reference the PPS. When loop_filter_across_slice_enabled_flag is set to 0, loop filtering operations are not performed across slice boundaries in pictures that reference the PPS. Loop filtering operations include deblocking filter, sample adaptive offset filter, and adaptive loop filter operations. When loop_filter_across_slices_enabled_flag is not present, the value of loop_filter_across_slices_enabled_flag is inferred to be 0.

[0156] 7.3.7.1 General Strip Header Syntax

[0157]

[0158] 7.4.8.1 General Strip Header Semantics ...

[0160] slice_subpic_id refers to the sub-picture identifier of the sub-picture containing the slice. If slice_subpic_id exists, the value of the variable SubPicIdx is derived so that SubpicIdList[SubPicIdx] is equal to slice_subpic_id. Otherwise (slice_subpic_id does not exist), the variable SubPicIdx is derived to be equal to 0. The length of slice_subpic_id, in bits, is derived as follows:

[0161] - If sps_subpic_id_signalling_present_flag is equal to 1, the length of slice_subpic_id is equal to sps_subpic_id_len_minus1+1.

[0162] Otherwise, if ph_subpic_id_signalling_present_flag is equal to 1, the length of slice_subpic_id is equal to ph_subpic_id_len_minus1+1.

[0163] Otherwise, if pps_subpic_id_signalling_present_flag is equal to 1, the length of slice_subpic_id is equal to pps_subpic_id_len_minus1+1.

[0164] Otherwise, the length of slice_subpic_id is equal to Ceil(Log2(sps_num_subpics_minus1+1)).

[0165] slice_address refers to the slice address of the slice. When not present, the value of slice_address is inferred to be equal to 0.

[0166] If rect_slice_flag is equal to 0, the following applies:

[0167] —The strip address is the raster scan slice index.

[0168] - The length of slice_address is Ceil(Log2(NumTilesInPic)) bits.

[0169] — The value of slice_address should be in the range of 0 to NumTilesInPic-1, inclusive.

[0170] Otherwise (rect_slice_flag is equal to 1), the following applies:

[0171] —The slice address is the slice index of the slice in the SubPicIdx-th sub-picture.

[0172] —The length of slice_address is Ceil(Log2(NumSlicesInSubpic[SubPicIdx])) bits.

[0173] —The value of slice_address shall be in the range of 0 to NumSlicesInSubpic[SubPicIdx]-1, inclusive.

[0174] The following constraints apply to bitstream conformance requirements:

[0175] - If rect_slice_flag is equal to 0 or subpics_present_flag is equal to 0, the value of slice_address shall not be equal to the value of slice_address of any other codec slice NAL unit of the same codec picture.

[0176] — Otherwise, the pair of slice_subpic_id and slice_address values ​​shall not be equal to the pair of slice_subpic_id and slice_address values ​​of any other codec slice NAL unit of the same codec picture.

[0177] — When rect_slice_flag is equal to 0, the slices of the picture shall be arranged in ascending order of their slice_address values.

[0178] - The shape of a picture slice shall be such that the entire left and upper boundaries of each CTU, when decoded, shall consist of either a picture boundary or a boundary of a previously decoded CTU.

[0179] num_tiles_in_slice_minus1 plus 1 (if present) refers to the number of tiles in the slice. The value of num_tiles_in_slice_minus1 should be in the range of 0 to NumTilesInPic-1, inclusive.

[0180] The variable NumCtuInCurrSlice refers to the number of CTUs in the current slice, and the list CtbAddrInCurrSlice[i] refers to the picture raster scan address of the i-th CTB in the slice, where i ranges from 0 to NumCtuInCurrSlice-1, inclusive, as derived as follows:

[0181]

[0182]

[0183] 3.4 Example of JVET-Q0075

[0184] Escape samples are used to handle unusual situations in palette mode.

[0185] The binarization of escaped samples is EG3 in the current VTM. However, for uniformly distributed signaling notifications, fixed-length binarization may outperform EG3 in both distortion and bit rate measurements.

[0186] JVET-Q0075 proposes to use fixed-length binarization for escaped samples and also modify the quantization and dequantization processes accordingly.

[0187] In the proposed method, the maximum bit depth of escaped samples depends on the quantization parameter, which is derived as follows.

[0188] max(1,bitDepth–(max(QpPrimeTsMin,Qp)–4) / 6)

[0189] Here, bitDepth is the internal bit depth, Qp is the current quantization parameter (QP), QpPrimeTsMin is the minimum QP for transform skip blocks, and max is the operation that obtains the larger value between the two inputs.

[0190] Furthermore, only shift operations are required during the dequantization of the escaped samples. Let escapeVal be the decoded escape value and recon be the reconstructed value of the escaped sample. The derivation of these values ​​is as follows.

[0191] shift=min(bitDepth–1,(max(QpPrimeTsMin,Qp)–4) / 6)

[0192] recon=(escapeVal< <shift)

[0193] The distortion of the reconstructed value is guaranteed to be always less than or equal to the current design distortion, for example, ((escapeVal*levelScale[qP%6])<<(qP / 6)+32)>>6.

[0194] At the encoder, quantization is implemented as follows:

[0195] escapeVal=(p+(1<<(shift–1)))>>shift

[0196] escapeVal=clip3(0,(1< <bd)–1,escapeVal)

[0197] Compared with the current design that uses EG3 and an inverse quantization table, one addition, one multiplication, and two shift operations for quantization, the method proposed in this application is much simpler and only requires one shift operation.

[0198] 3.5 Example of JVET-Q0294

[0199] In order to achieve efficient compression in mixed lossy and lossless codecs, JVET-Q0294 recommends signaling a flag at each codec tree unit (CTU) to indicate whether the CTU is coded in lossless or lossy mode. If the CTU is lossless, an additional CTU-level flag is signaled to specify the residual codec method used for the CTU, either normal residual codec or transform skip residual codec.

[0200] 4. Examples of technical problems solved by this solution

[0201] The existing design of signaling for sub-pictures, slices, and stripes in VVC has the following problems:

[0202] 1) The codec for sps_num_subpics_minus1 is u(8), which means that each picture cannot have more than 256 subpictures. However, in some applications, the maximum number of subpictures per picture may need to be greater than 256.

[0203] 2) It is allowed that subpics_present_flag is equal to 0 and sps_subpic_id_present_flag is equal to 1. However, this does not make sense because subpics_present_flag equal to 0 means that CLVS has no information about sub-pictures at all.

[0204] 3) For each sub-picture, a list of sub-picture IDs can be signaled in the picture header (PH). However, when the list of sub-picture IDs is signaled in the PH, all PHs will need to be changed when a subset of sub-pictures is extracted from the bitstream. This is undesirable.

[0205] 4) Currently, when the sub-picture ID is indicated as explicitly signaled, the sub-picture ID may not be signaled anywhere via sps_subpic_id_present_flag (or subpic_ids_explicitly_signalled_flag if the name of the syntax element is changed to 1). This is problematic because when the sub-picture ID is indicated as explicitly signaled, the sub-picture ID needs to be explicitly signaled in the SPS or PPS.

[0206] 5) When the sub-picture ID is not explicitly signaled, the slice header syntax element slice_subpic_id still needs to be signaled whenever subpics_present_flag is equal to 1, including when sps_num_subpics_minus1 is equal to 0. However, the length of slice_subpic_id is currently specified as Ceil(Log2(sps_num_subpics_minus1+1)) bits, which is 0 bits when sps_num_subpics_minus1 is equal to 0. This is problematic because any existing syntax element cannot be 0 bits long.

[0207] 6) The sub-picture layout, including the number, size, and position of sub-pictures, remains unchanged throughout the CLVS. Even if the sub-picture ID is not explicitly signaled in the SPS or PPS, the sub-picture ID length still needs to be signaled for the sub-picture ID syntax element in the slice header.

[0208] 7) Whenever rect_slice_flag is equal to 1, the syntax element slice_address is signaled in the slice header and specifies the slice index within the sub-picture that contains the slice, including when the number of slices within the sub-picture (i.e., NumSlicesInSubpic[SubPicIdx]) is equal to 1. However, currently, when rect_slice_flag is equal to 1, the length of slice_address is specified to be Ceil(Log2(NumSlicesInSubpic[SubPicIdx])) bits, and the length of slice_address is 0 bits when NumSlicesInSubpic[SubPicIdx] is equal to 1. This is problematic because any existing syntax element cannot be 0 bits in length.

[0209] 8) There is redundancy between the syntax elements no_pic_partition_flag and pps_num_subpics_minus1, although the latest VVC text has the following constraint: when sps_num_subpics_minus1 is greater than 0, the value of no_pic_partition_flag shall be equal to 1.

[0210] 9) Within CLVS, the sub-picture ID value for a particular sub-picture position or index may vary from picture to picture. When this happens, in principle, the sub-picture cannot use inter-layer prediction by referencing reference pictures in the same layer. However, there is currently a lack of constraints in the current VVC specification that prohibit this practice.

[0211] 10) In the current VVC design, reference pictures can be from different layers to support various applications, such as scalable video codec and multi-view video codec. If sub-pictures exist in different layers, it is necessary to study whether to allow or not inter-layer prediction.

[0212] 5. Example Techniques and Embodiments

[0213] To address the above and other issues, the following methods are disclosed. The present invention should be considered as an example to explain the general concept and should not be interpreted narrowly. In addition, these inventions can be applied alone or in combination in any way.

[0214] The following abbreviations have the same meanings as in JVET-P1001-vE.

[0215] BP (buffer period), BP SEI (supplemental enhancement information),

[0216] PT (Picture Timing), PT SEI,

[0217] AU (Access Unit),

[0218] OP (operating point),

[0219] DUI (Decoding Unit Information), DUI SEI,

[0220] NAL (Network Layer),

[0221] NUT (NAL unit type),

[0222] GDR (Gradual Decoder Refresh),

[0223] SLI (Sub-Picture Level Information), SLI SEI.

[0224] 1) To solve the first problem, change the codec of sps_num_subpics_minus1 from u(8) to ue(v) so that each picture can have more than 256 subpictures.

[0225] a. In addition, the value of sps_num_subpics_minus1 is limited to the range of 0 to Ceil(pic_width_max_in_luma_samples÷CtbSizeY)*Ceil(pic_height_max_in_-luma_samples÷CtbSizeY)-1.

[0226] b. In addition, the number of sub-images per image is further restricted in the definition of the level.

[0227] 2) To solve the second problem, the condition for signaling the syntax element sps_subpic_id_present_flag is set to "if(subpics_present_flag)", that is, when subpics_present_flag is equal to 0, the sps_subpic_id_present_flag syntax element is not signaled, and when it does not exist, it is inferred that the value of sps_subpic_id_present_flag is equal to 0.

[0228] a. Alternatively, when subpics_present_flag is equal to 0, the syntax element sps_subpic_id_present_flag is still signaled, but its value needs to be equal to 0 when subpics_present_flag is equal to 0.

[0229] b. In addition, the names of the syntax elements subpics_present_flag and sps_subpic_id_present_flag are changed to subpic_info_present_flag and subpic_ids_explicitly_signalled_flag, respectively.

[0230] 3) To solve the third problem, the signaling of the sub-picture ID in the PH syntax is removed. Therefore, for i in the range of 0 to sps_num_subpics_minus1 (inclusive), the list SubpicIdList[i] is derived as follows:

[0231]

[0232] 4) To solve the fourth problem, when a sub-picture is indicated for explicit signaling, the sub-picture ID is signaled in the SPS or PPS.

[0233] a. Implemented by adding the following constraint: If subpic_ids_explicitly_signalled_flag is 0 or subpic_ids_in_sps_flag is equal to 1, then subpic_ids_in_pps_flag shall be equal to 0. Otherwise (subpic_ids_explicitly_signalled_flag is 1 or subpic_ids_in_sps_flag is equal to 0), subpic_ids_in_pps_flag shall be equal to 1.

[0234] 5) To address the fifth and sixth issues, the length of the sub-picture ID is signaled in the SPS regardless of the value of the SPS flag sps_subpic_id_present_flag (or renamed subpic_ids_explicitly_signalled_flag). Although when the sub-picture ID is also explicitly signaled in the PPS, the length can also be signaled in the PPS to avoid parsing the PPS's dependency on the SPS. In this case, the length also specifies the length of the sub-picture ID in the slice header, even if the sub-picture ID is not explicitly signaled in the SPS or PPS. Therefore, when it is present, the length of slice_subpic_id is also specified by the sub-picture ID length signaled in the SPS.

[0235] 6) Alternatively, to address the fifth and sixth issues, a flag is added to the SPS syntax with a value of 1 to specify the presence of the sub-picture ID length in the SPS syntax. The presence of this flag is independent of the value of the flag indicating whether the sub-picture ID is explicitly signaled in the SPS or PPS. When subpic_ids_explicitly_signalled_flag is equal to 0, the value of this flag can be equal to 1 or 0, but when subpic_ids_explicitly_signalled_flag is equal to 1, the value of this flag must be equal to 1. When this flag is equal to 0, that is, the sub-picture length does not exist, the length of slice_subpic_id is specified as Max(Ceil(Log2(sps_num_subpics_minus1+1)), 1) bits (instead of Ceil(Log2(sps_num_subpics_minus1+1)) bits in the latest VVC draft text).

[0236] a. Alternatively, this flag is present only when subpic_ids_explicitly_signalled_flag is equal to 0, and is inferred to be equal to 1 when subpic_ids_explicitly_signalled_flag is equal to 1.

[0237] 7) To solve the seventh problem, when rect_slice_flag is equal to 1, the length of slice_address is Max(Ceil(Log2(NumSlicesInSubpic[SubPicIdx])), 1) bits.

[0238] a. Alternatively, further, when rect_slice_flag is equal to 0, the length of slice_address is specified as Max(Ceil(Log2(NumTilesInPic)), 1) bits instead of Ceil(Log2(NumTilesInPic)) bits.

[0239] 8) To address the eighth issue, the condition for signaling no_pic_partition_flag is set to “if (subpic_ids_in_pps_flag&&pps_num_subpics-_minus1>0)”, and the following inference is added: when it does not exist, the value of no_pic_partition_flag is inferred to be equal to 1.

[0240] a. Alternatively, move the sub-picture ID syntax (all four syntax elements) after the slice and slice syntax in the PPS, for example, immediately before the syntax element entropy_coding_sync_enabled_flag, and then set the condition for signaling pps_num_subpics_minus1 to "if (no_pic_partition_flag)".

[0241] 9) To address the ninth issue, the following constraint is specified: for each specific sub-picture index (or equivalently, sub-picture position), when the sub-picture ID value changes between a picture picA and the previous picture at the same layer in decoding order, unless picA is the first picture of a CLVS, the sub-picture at picA shall only contain codec slice NAL units with nal_unit_type equal to IDR_W_RADL, IDR_N_LP or CRA_NUT.

[0242] a. Alternatively, the above constraint applies only to sub-picture indices where the value of subpic_treatment_as_pic_flag[i] is equal to 1.

[0243] b. Alternatively, for 9 and 9a above, change "IDR_W_RADL, IDR_N_LP, or CRA_NUT" to "IDR_W_RADL, IDR_N_LP, CRA_NUT, RSV_IRAP_11, or RSV_IRAP_12".

[0244] c. Alternatively, the sub-picture at picA may contain other types of codec slice NAL units, however, these codec slice NAL units only use one or more of inter-layer prediction, intra-layer block copy (IBC) prediction, and palette mode prediction.

[0245] Alternatively, a first video unit (e.g., slice, tile, etc.) in a sub-picture of picA can refer to a second video unit in the previous picture. The second video unit and the first video unit can be in a sub-picture with the same sub-picture index, although their sub-picture IDs may be different. The sub-picture index is a unique number assigned to a sub-picture and cannot be changed in CLVS.

[0246] 10) For a specific sub-picture index (or equivalently, sub-picture position), it is possible to signal in the bitstream which sub-pictures (identified by a layer ID value together with the sub-picture index or sub-picture ID value) are used as reference pictures.

[0247] 11) For the multi-layer case, inter-layer prediction (ILR) of sub-pictures from different layers is allowed when certain conditions are met (e.g., may depend on the number of sub-pictures, the location of sub-pictures), and is disabled when certain conditions are not met.

[0248] a. In one example, even when two sub-pictures in two layers have the same sub-picture index value but different sub-picture ID values, inter-layer prediction may still be allowed when certain conditions are met.

[0249] i. In one example, some of the conditions are "if two layers are associated with different view order index / view order ID values".

[0250] b. If two sub-pictures have the same sub-picture index, the first sub-picture in the first layer and the second sub-picture in the second layer may be constrained to be at collocated locations and / or have reasonable widths / heights.

[0251] c. If the first sub-picture can refer to the second reference sub-picture, the first sub-picture in the first layer and the second sub-picture in the second layer may be restricted to be in a collocated position and / or a reasonable width / height.

[0252] 12) An indication of whether the current sub-picture can use inter-layer prediction (ILP) from sample values ​​and / or other values ​​(e.g., motion information and / or coding mode information) associated with a region or sub-picture of a reference layer is signaled in the bitstream, for example, in the VPS / DPS / SPS / PPS / APS / sequence header / picture header.

[0253] a. In one example, the reference region or sub-picture of the reference layer is at least one collocated sample containing samples within the current sub-picture.

[0254] b. In one example, the reference area or sub-picture of the reference layer is outside the collocated area of ​​the current sub-picture.

[0255] c. In one example, the indication is signaled in one or more SEI messages.

[0256] d. In one example, regardless of whether the reference layer has multiple sub-pictures, and when multiple sub-pictures exist in one or more reference layers, regardless of whether the picture is partitioned into sub-pictures such that each sub-picture in the current picture is a corresponding sub-picture in the reference picture covering the collocated region, and regardless of whether the corresponding / collocated sub-pictures have the same sub-picture ID value as the current sub-picture, the indication is signaled.

[0257] 13) When a BP SEI message and a PT SEI message applicable to a specific OP appear in an AU, the BP SEI message shall precede the PT SEI message in decoding order.

[0258] 14) When a BP SEI message and a DUI SEI message applicable to a specific OP appear in an AU, the BP SEI message shall precede the DUI SEI message in decoding order.

[0259] 15) When the PT SEI message and the DUI SEI message applicable to a specific OP appear in an AU, the PT SEI message should precede the DUI SEI message in the decoding order.

[0260] 16) An indication of whether a picture is an IRAP / GDR picture flag is signaled at the beginning of the picture header, and no_output_of_prior_pics_flag may be signaled based on the indication.

[0261] An example grammar design is as follows:

[0262]

[0263] irap_or_gdr_pic_flag equal to 1 specifies that the picture associated with the PH is an IRAP or GDR picture. irap_or_gdr_pic_flag equal to 0 specifies that the picture associated with the PH is neither an IRAP picture nor a GDR picture.

[0264] 17) A picture with the value of mixed_nalu_types_in_pic_flag equal to 0 is not allowed to not contain a codec slice NAL unit with nal_unit_type equal to GDR_NUT.

[0265] 18) Signaling syntax element (i.e., mixed_slice_types_in_pic_flag) in PPS. If mixed_slice_types_in_pic_flag is equal to 0, the slice type (B, P, or I) is coded in the PH. Otherwise, the slice type is coded in the SHs. Syntax values ​​associated with unused slice types are further skipped in the picture header. The syntax element mixed_slice_types_in_pic_flag is conditionally signaled as follows:

[0266]

[0267] 19) In SPS or PPS, a maximum of N1 (e.g., 3) vertical virtual borders and a maximum of N2 (e.g., 3) horizontal virtual borders are signaled. In PH, a maximum of N3 (e.g., 1) additional vertical borders and a maximum of N4 (e.g., 1) additional horizontal virtual borders are signaled, with the restriction that the total number of vertical virtual borders should be less than or equal to N1, and the total number of horizontal virtual borders should be less than or equal to N2.

[0268] 20) The syntax element inter_layer_ref_pics_present_flag in the SPS may be conditionally signaled as follows:

[0269]

[0270] 21) The syntax elements entropy_coding_sync_enabled_flag and entry_point_offsets_present_flag are signaled in the SPS instead of the PPS.

[0271] 22) The values ​​of vps_num_ptls_minus1 and num_ols_hrd_params_minus1 are required to be less than the value T. For example, T can be equal to TotalNumOlss as specified in JVET-2001-vE.

[0272] a. The difference between T and vps_num_ptls_minus1 or (vps_num_ptls_minus1+1) may be signaled.

[0273] b. The difference between T and hrd_params_minus1 or (hrd_params_minus1+1) may be signaled.

[0274] c. The difference may be signaled via unary codec or Exponential-Golomb codec.

[0275] 23) The value of vps_num_dpb_params_minus1 must be less than or equal to vps_max_layers_minus1.

[0276] a. The difference between T and vps_num_ptls_minus1 or (vps_num_ptls_minus1+1) may be signaled.

[0277] b. The difference can be signaled via unary codec or Exponential-Golomb codec.

[0278] 24) Allows the EOB NAL unit to be provided to the decoder either by inclusion in the bitstream or by external means.

[0279] 25) Allows EOS NAL units to be provided to the decoder either by inclusion in the bitstream or by external means.

[0280] 26) For each layer with layer index i, when vps_independent_layer_flag[i] is equal to 0, each picture in this layer shall contain only one sub-picture.

[0281] 27) During the sub-bitstream extraction process, specify that whenever a VCL NAL unit is removed, the padding data units associated with the VCL NAL unit are also removed, and all padding SEI messages in the SEI NAL units associated with the VCL NAL unit are removed.

[0282] 28) Restriction When an SEI NAL unit contains a padding SEI message, the SEI NAL unit shall not contain any other SEI messages that are not padding SEI messages.

[0283] a. Alternatively, restrict that when an SEI NAL unit contains a padding SEI message, the SEI NAL unit shall contain any other SEI message.

[0284] 29) Constraints A VCL NAL unit shall have at most one associated padding NAL unit.

[0285] 30) Constraint: A VCL NAL unit shall have at most one associated padding NAL unit.

[0286] 31) It is recommended to signal a flag to indicate whether a video unit is coded in lossless or lossy mode, where the video unit can be a CU or CTU, and the signaling of escape samples in palette mode can depend on the flag.

[0287] a. In one example, the binarization method of escaped samples in palette mode can depend on a flag.

[0288] b. In one example, the determination of the codec context in arithmetic coding of escape samples in palette mode may depend on a flag.

[0289] 6. Examples

[0290]

[0291] 6.1. First embodiment

[0292] 7.3.2.3 Sequence Parameter Set RBSP Syntax

[0293]

[0294] 7.4.3.3 Sequence Parameter Set RBSP Semantics ...

[0296]

[0297]

[0298] NOTE 2 – When the bitstream is the result of a sub-bitstream extraction process and contains only a subset of the sub-pictures of the input bitstream of the sub-bitstream extraction process, it may be necessary to include The value of is set to 1.

[0299] sps_num_subpics_minus1 plus 1 refers to the number of sub-pictures.

[0300]

[0301] When it is not present, the value of sps_num_subpics_minus1 is inferred to be equal to 0.

[0302] subpic_ctu_top_left_x[i] refers to the horizontal position of the top left CTU of the i-th sub-picture in units of CtbSizeY. The length of the syntax element is Ceil(Log2(pic_width_max_in_luma_samples / CtbSizeY)) bits. When it is not present, the value of subpic_ctu_top_left_x[i] is inferred to be equal to 0.

[0303] subpic_ctu_top_left_y[i] refers to the vertical position of the top left CTU of the i-th subpicture in units of CtbSizeY. The length of this syntax element is Ceil(Log2(pic_height_max_in_luma_samples / CtbSizeY)) bits. When it is not present, the value of subpic_ctu_top_left_y[i] is inferred to be equal to 0.

[0304] subpic_width_minus1[i] plus 1 specifies the width of the i-th sub-picture in units of CtbSizeY. The length of this syntax element is Ceil(Log2(pic_width_max_in_luma_samples / CtbSizeY)) bits. When not present, the value of subpic_width_minus1[i] is inferred to be equal to Ceil(pic_width_max_in_luma_samples / CtbSizeY)-1.

[0305] subpic_height_minus1[i] plus 1 refers to the height of the i-th sub-picture in units of CtbSizeY. The length of this syntax element is Ceil(Log2(pic_height_max_in_luma_samples / CtbSizeY)) bits. When it is not present, the value of subpic_height_minus1[i] is inferred to be equal to Ceil(pic_height_max_in_luma_samples / CtbSizeY)-1.

[0306] subpic_treated_as_pic_flag[i] equal to 1 indicates that the i-th subpicture of each codec picture in the CLVS is treated as a picture during decoding without loop filtering. subpic_treatment_as_pic_flag[i] equal to 0 indicates that the i-th subpicture of each codec picture in the CLVS is not treated as a picture during decoding without loop filtering. When it is not present, the value of subpic_treatment_as_pic_flag[i] is inferred to be 0.

[0307] loop_filter_across_subpic_enabled_flag[i] is equal to 1, which means that loop filtering operations can be performed across the boundary of the i-th sub-picture in each codec picture in the CLVS. loop_filter_across_subpic_enabled_flag[i] is equal to 0, which means that loop filtering operations are not performed across the boundary of the i-th sub-picture in each codec picture in the CLVS. When it is not present, the value of loop_filter_across_subpic_enabled_pic_flag[i] is inferred to be equal to 1.

[0308] The following constraints apply to bitstream conformance requirements:

[0309] —For any two sub-pictures subpicA and subpicB, when the sub-picture index of subpicA is less than the sub-picture index of subpicB, any codec slice NAL unit of subPicA shall take precedence over any codec slice NAL unit of subPicB in decoding order.

[0310] - The shape of the sub-pictures shall be such that, when each sub-picture is decoded, its entire left and upper boundaries consist of either picture boundaries or boundaries of previously decoded sub-pictures.

[0311]

[0312] sps_subpic_id[i] refers to the sub-picture ID of the i-th sub-picture. The length of the sps_subpic_id[i] syntax element is sps_subpic_id_len_minus1+1 bits. ...

[0314] 7.3.2.4 Picture Parameter Set RBSP Syntax

[0315]

[0316]

[0317]

[0318] 7.4.3.4 Picture Parameter Set RBSP Semantics ...

[0320]

[0321] pps_num_subpics_minus1 shall be equal to sps_num_subpics_minus1.

[0322] pps_subpic_id_len_minus1 shall be equal to sps_subpic_id_len_minus1.

[0323] pps_subpic_id[i] refers to the sub-picture ID of the i-th sub-picture. The length of the pps_subpic_id[i] syntax element is pps_subpic_id_len_minus1+1 bits.

[0324]

[0325]

[0326] A bitstream conformance requirement is that for any i and j in the range 0 to sps_num_subpics_minus1 (inclusive), when i is less than j, SubpicIdList[i] shall be less than SubpicIdList[j]. ...

[0328] rect_slice_flag equal to 0 means that the slices within each slice are arranged in raster scan order and the slice information is not signaled in the PPS. rect_slice_flag equal to 1 means that the slices within each slice cover a rectangular area of ​​the picture and the slice information is signaled in the PPS. When it is not present, rect_slice_flag is inferred to be 1. When equal to 1, the value of rect_slice_flag shall be equal to 1.

[0329] single_slice_per_subpic_flag is equal to 1, which means that each sub-picture consists of one and only one rectangular slice. single_slice_per_subpic_flag is equal to 0, which means that each sub-picture can contain one or more rectangular slices. When equal to 0, single_slice_per_subpic_flag shall be equal to 0. When single_slice_per_subpic_flag is equal to 1, num_slices_in_pic_minus1 is inferred to be equal to sps_num_subpics_minus1. ...

[0331] 7.3.7.1 General Strip Header Syntax

[0332]

[0333]

[0334] 7.4.8.1 General Strip Header Semantics ...

[0336] slice_subpic_id refers to the sub-picture ID of the sub-picture containing the slice.

[0337] When it is not present, the value of slice_subpic_id is inferred to be equal to 0.

[0338] The variable SubPicIdx is derived such that SubpicIdList[SubPicIdx] is equal to the value of slice_subpic_id.

[0339] slice_address refers to the slice address of the slice. If it does not exist, the value of slice_address is inferred to be equal to 0.

[0340] If rect_slice_flag is equal to 0, the following applies:

[0341] —The strip address is the raster scan slice index.

[0342] - The length of slice_address is Ceil(Log2(NumTilesInPic)) bits.

[0343] — The value of slice_address should be in the range of 0 to NumTilesInPic-1, inclusive.

[0344] Otherwise (rect_slice_flag is equal to 1), the following applies:

[0345] — The slice address is the sub-picture level slice index of the slice.

[0346] —The length of slice_address is bit.

[0347] —The value of slice_address shall be in the range of 0 to NumSlicesInSubpic[SubPicIdx]-1, inclusive.

[0348] The following constraints apply to bitstream conformance requirements:

[0349] — If rect_slice_flag is equal to 0 or If equal to 0, the value of slice_address shall not be equal to the value of slice_address of any other codec slice NAL unit of the same codec picture.

[0350] — Otherwise, the pair of slice_subpic_id and slice_address values ​​shall not be equal to the pair of slice_subpic_id and slice_address values ​​of any other codec slice NAL unit of the same codec picture.

[0351] — When rect_slice_flag is equal to 0, the slices of the picture shall be arranged in ascending order of their slice_address values.

[0352] - The shape of the picture slice shall be such that the entire left and upper boundaries of each CTU, when decoded, shall consist of a picture boundary or of the boundaries of the previously decoded CTU(s). ...

[0354] Figure 5 A block diagram of an example video processing system 500 is shown that can implement various techniques of the present disclosure. Various implementations can include some or all of the components of system 500. System 500 can include an input 502 for receiving video content. The video content can be received in a raw or uncompressed format (e.g., 8 or 10 bit multi-component pixel values), or can be received in a compressed or encoded format. Input 502 can represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, a passive optical network (PON), and wireless interfaces such as Wi-Fi or a cellular interface.

[0355] System 500 may include a codec component 504 that can implement the various codecs or encoding methods described in this disclosure. Codec component 504 can reduce the average bit rate of the video from input 502 to the output of codec component 504 to produce a codec representation of the video. Therefore, codec technology is sometimes referred to as video compression or video transcoding technology. The output of codec component 504 can be stored or transmitted via a communication connection such as represented by component 506. The stored or communicated bitstream (or codec) representation of the video received at input 502 can be used by component 508, which is used to generate pixel values ​​or displayable video sent to display interface 510. The process of generating a user-viewable video from the bitstream representation is sometimes referred to as video decompression. In addition, although some video processing operations are referred to as "codec" operations or tools, it should be understood that the encoding tools or operations are used at the encoder, and the corresponding decoding tools or operations that reverse the encoding results will be performed by the decoder.

[0356] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB), High-Definition Multimedia Interface (HDMI), or DisplayPort, etc. Examples of storage interfaces include SATA (Serial Advanced Technology Attachment), PCI, IDE interfaces, etc. The technology described in this disclosure may be implemented in various electronic devices, such as mobile phones, laptop computers, smartphones, or other devices capable of performing digital data processing and / or video display.

[0357] Figure 6 6 is a block diagram of a video processing device 600. Device 600 can be used to implement one or more methods described in this disclosure. Device 600 can be located in a smartphone, tablet computer, computer, Internet of Things (IoT) receiver, etc. Device 600 may include one or more processors 602, one or more memories 604, and video processing hardware 606. Processor 602 can be configured to implement one or more methods described in this disclosure. Memory 604 can be used to store data and code for implementing the methods and techniques described in this disclosure. Video processing hardware 606 can be used in hardware circuits to implement some of the techniques described in this disclosure. In some embodiments, hardware 606 can be partially or entirely in processor 602, such as a graphics processor.

[0358] Figure 7 is a block diagram illustrating an example video coding system 100 that may utilize the techniques of this disclosure.

[0359] like Figure 7 As shown, the video encoding and decoding system 100 may include a source device 110 and a target device 120. The source device 110 may generate encoded video data and may be referred to as a video encoding device. The target device 120 may decode the encoded video data generated by the source device 110 and may be referred to as a video decoding device.

[0360] Source device 110 may include a video source 112 , a video encoder 114 , and an input / output (I / O) interface 116 .

[0361] The video source 112 may include, for example, a source of a video capture device, an interface for receiving video data from a video content provider, and / or a computer graphics system for generating video data, or a combination of these sources. The video data may include one or more pictures. The video encoder 114 encodes the video data from the video source 112 to generate a bitstream. The bitstream may include a sequence of bits that form a codec representation of the video data. The bitstream may include a codec picture and associated data. The codec picture is a codec representation of the picture. Associated data may include a sequence parameter set, a picture parameter set, and other syntax structures. The I / O interface 116 may include a modulator / demodulator (modem) and / or a transmitter. The encoded video data may be sent directly to the target device 120 via the network 130a via the I / O interface 116. The encoded video data may also be stored on a storage medium / server 130b for access by the target device 120.

[0362] Target device 120 may include an I / O interface 126 , a video decoder 124 , and a display device 122 .

[0363] I / O interface 126 may include a receiver and / or a modem. I / O interface 126 may obtain encoded video data from source device 110 or storage medium / server 130b. Video decoder 124 may decode the encoded video data. Display device 122 may display the decoded video data to a user. Display device 122 may be integrated with target device 120 or may be external to target device 120, with target device 120 configured to interface with an external display device.

[0364] The video encoder 114 and the video decoder 124 may operate according to a video compression standard, such as the High Efficiency Video Coding (HEVC) standard, the Versatile Video Coding (VVM) standard, and other current and / or other standards.

[0365] Figure 8 is a block diagram illustrating an example of a video encoder 200, which may be Figure 7 The video encoder 114 in the system 100 is illustrated in FIG.

[0366] Video encoder 200 can be configured to perform any or all of the techniques of this disclosure. Figure 8 In the example shown, video encoder 200 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of video encoder 200. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.

[0367] The functional components of the video encoder 200 may include a segmentation unit 201, a prediction unit 202 which may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205 and an intra-frame prediction unit 206, a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213 and an entropy coding unit 214.

[0368] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an intra block copy (IBC) unit. The IBC unit may perform prediction in an IBC mode where at least one reference picture is a picture in which the current video block is located.

[0369] Furthermore, some components such as the motion estimation unit 204 and the motion compensation unit 205 may be highly integrated but are not shown for the purpose of description. Figure 8 are represented separately in the example.

[0370] The partitioning unit 201 may partition a picture into one or more video blocks. The video encoder 200 and the video decoder 300 may support various video block sizes.

[0371] The mode selection unit 203 may select one of the codec modes (intra or inter) based on the error result, for example, and provide the resulting intra or inter codec block to the residual generation unit 207 to generate residual block data and the reconstruction unit 212 to reconstruct the codec block for use as a reference picture. In some examples, the mode selection unit 203 may select a combination of intra and inter prediction (CIIP) modes, where the prediction is based on an inter prediction signal and an intra prediction signal. In the case of inter prediction, the mode selection unit 203 may also select a resolution for the motion vector for the block (e.g., sub-pixel or integer pixel precision).

[0372] To perform inter-frame prediction on the current video block, the motion estimation unit 204 may generate motion information of the current video block by comparing the current video block with one or more reference frames from the buffer 213. The motion compensation unit 205 may determine a predicted video block for the current video block based on the motion information and decoded samples of pictures other than the picture associated with the current video block from the buffer 213.

[0373] Motion estimation unit 204 and motion compensation unit 205 may perform different operations on the current video block, eg, depending on whether the current video block is in an I slice, a P slice, or a B slice.

[0374] In some examples, motion estimation unit 204 may perform unidirectional prediction on the current video block, and motion estimation unit 204 may search for a reference picture in list 0 or 1 for the reference video block of the current video block. Motion estimation unit 204 may then generate a reference index indicating the reference picture in list 0 or list 1 that contains the reference video block and a motion vector indicating the spatial displacement between the current video block and the reference video block. Motion estimation unit 204 may output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 205 may generate a predicted video block for the current block based on the reference video block indicated by the motion information of the current video block.

[0375] In other examples, the motion estimation unit 204 may perform bidirectional prediction on the current video block. The motion estimation unit 204 may search for reference pictures in List 0 for the reference video block of the current video block, and may also search for reference pictures in List 1 for the reference video block of the current video block. The motion estimation unit 204 may then generate reference indexes indicating reference pictures in List 0 and List 1, which include reference video blocks and motion vectors indicating spatial displacements between the reference video blocks and the current video block. The motion estimation unit 204 may output the reference index and motion vector of the current video block as motion information for the current video block. The motion compensation unit 205 may generate a predicted video block for the current video block based on the reference video block indicated by the motion information of the current video block.

[0376] In some examples, motion estimation unit 204 may output a complete motion information set for use in a decoding process by a decoder.

[0377] In some examples, motion estimation unit 204 may not output a complete set of motion information for the current video. Instead, motion estimation unit 204 may reference motion information of another video block to signal the motion information for the current video block. For example, motion estimation unit 204 may determine that the motion information for the current video block is sufficiently similar to the motion information for a neighboring video block.

[0378] In one example, motion estimation unit 204 may indicate, in a syntax structure associated with the current video block, a value that indicates to video decoder 300 that the current video block has the same motion information as another video block.

[0379] In another example, motion estimation unit 204 may identify another video block and a motion vector difference (MVD) in a syntax structure associated with the current video block. The motion vector difference represents the difference between the motion vector of the current video block and the motion vector of the indicated video block. Video decoder 300 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.

[0380] As described above, the video encoder 200 may predictively signal motion vectors. Two examples of predictive signaling techniques that may be implemented by the video encoder 200 include advanced motion vector prediction (AMVP) and merge mode signaling.

[0381] The intra-frame prediction unit 206 can perform intra-frame prediction on the current video block. When the intra-frame prediction unit 206 performs intra-frame prediction on the current video block, the intra-frame prediction unit 206 can generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block may include a predicted video block and various syntax elements.

[0382] The residual generation unit 207 may generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) the predicted video block of the current video block from the current video block. The residual data for the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.

[0383] In other examples, for the current video block, such as in skip mode, there may be no residual data for the current video block, and the residual generation unit 207 may not perform the subtraction operation.

[0384] Transform processing unit 208 may generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video block associated with the current video block.

[0385] After transform processing unit 208 generates a transform coefficient video block associated with the current video block, quantization unit 209 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values ​​associated with the current video block.

[0386] The inverse quantization unit 210 and the inverse transform unit 211 may apply inverse quantization and inverse transform, respectively, to the transform coefficient video block to reconstruct a residual video block from the transform coefficient video block. The reconstruction unit 212 may add the reconstructed residual video block to corresponding samples of one or more predicted video blocks generated by the prediction unit 202 to generate a reconstructed video block associated with the current block stored in the buffer 213.

[0387] After the reconstruction unit 212 reconstructs the video block, a loop filtering operation may be performed to reduce video block artifacts in the video block.

[0388] The entropy coding unit 214 may receive data from other functional components of the video encoder 200. When the entropy coding unit 214 receives the data, the entropy coding unit 214 may perform one or more entropy coding operations to generate entropy-coded data and output a bitstream including the entropy-coded data.

[0389] Figure 9 is a block diagram illustrating an example of a video decoder 300, which may be Figure 7 The video decoder 114 in the system 100 is shown.

[0390] Video decoder 300 may be configured to perform any or all of the techniques of this disclosure. Figure 8 In the example shown, video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of video decoder 300. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.

[0391] In such Figure 9 In the example shown, the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra-frame prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, the video decoder 300 can perform operations generally related to the video encoder 200 ( Figure 8 ) describes the decoding channel opposite to the encoding channel.

[0392] The entropy decoding unit 301 can obtain an encoded bitstream. The encoded bitstream can include entropy-encoded video data (e.g., encoded video data blocks). The entropy decoding unit 301 can decode the entropy-encoded video data, and based on the entropy-decoded video data, the motion compensation unit 302 can determine motion information including motion vectors, motion vector precision, reference picture list index, and other motion information. For example, the motion compensation unit 302 can determine such information by performing AMVP and merge modes.

[0393] The motion compensation unit 302 may generate a motion compensated block, possibly performing interpolation based on an interpolation filter. An identifier for the interpolation filter at sub-pixel precision may be included in a syntax element.

[0394] Motion compensation unit 302 may calculate interpolated values ​​for sub-integer pixels of a reference block using interpolation filters used by video encoder 20 during encoding of the video block. Motion compensation unit 302 may determine the interpolation filters used by video encoder 200 from received syntax information and use the interpolation filters to produce a prediction block.

[0395] The motion compensation unit 302 may use some syntax information to determine the size of blocks used to encode frames and / or slices of the coded video sequence, partitioning information describing how each macroblock of a picture of the coded video sequence is partitioned, a mode indicating how each partition is encoded, one or more reference frames (and a list of reference frames) for each inter-frame coded block, and other information for decoding the coded video sequence.

[0396] The intra prediction unit 303 can form a prediction block from spatially neighboring blocks using, for example, an intra prediction mode received in the bitstream. The inverse quantization unit 303 inversely quantizes, i.e., dequantizes, the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 301. The inverse transform unit 303 applies an inverse transform.

[0397] The reconstruction unit 306 may add the residual block to the corresponding prediction block generated by the motion compensation unit 202 or the intra prediction unit 303 to form a decoded block. If necessary, a deblocking filter may be applied to the decoded block to remove blocking artifacts. The decoded video block is then stored in a buffer 307, which provides reference blocks for subsequent motion compensation / intra prediction and also produces decoded video for presentation on a display device.

[0398] Figure 10-26 An example method for implementing the above technical solution is shown, for example, Figure 5-9 The embodiment shown.

[0399] Figure 10 A flow chart of an example method 1000 for video processing is shown. The method 1000 includes, at operation 1010, performing conversion between a video and a bitstream of the video, the bitstream including one or more access units according to a format rule, and the format rule dictating an order in which a first message and a second message for an operation point OP appear in the access unit AU such that the first message precedes the second message in decoding order.

[0400] Figure 11 A flow chart of an example method 1100 for video processing is shown. The method 1000 includes, at operation 1110, performing conversion between a video and a bitstream of the video, the bitstream including one or more access units according to a format rule, and the format rule dictating an order in which a plurality of messages for an operation point OP appear in the access units such that a first message of the plurality of messages precedes a second message of the plurality of messages in a decoding order.

[0401] Figure 12 A flow chart of an example method 1200 for video processing is shown. The method 1200 includes, at operation 1210, performing conversion between a video including a picture and a bitstream of the video, the bitstream conforming to a format rule, the format rule specifying whether an indication of a first flag is signaled at the beginning of a picture header associated with the picture, the first flag indicating whether the picture is an intra random access point (IRAP) picture or a gradual decoding refresh (GDR) picture.

[0402] Figure 13A flow chart of an example method 1300 for video processing is shown. The method 1300 includes, at operation 1310, performing conversion between a video including one or more pictures and a bitstream of the video, the bitstream conforming to a format rule that does not allow a picture of the one or more pictures to be coded to include a codec slice network abstraction layer NAL unit having a progressive decoding refresh type and associated with a flag indicating that the picture includes a NAL unit of a mixed type.

[0403] Figure 14 A flow chart of an example method 1400 for video processing is shown. The method 1400 includes, at operation 1410, performing conversion between a video including one or more pictures and a bitstream of the video, the bitstream conforming to a format rule, the format rule allowing a picture of the one or more pictures to be encoded to include a codec slice network abstraction layer NAL unit having a progressive decoding refresh type and associated with a flag indicating that the picture does not include a NAL unit of a hybrid type.

[0404] Figure 15 A flow chart of an example method 1500 for video processing is shown. The method 1500 includes, at operation 1510, performing conversion between a picture of a video and a bitstream of the video, the bitstream conforming to a format rule, the format rule specifying whether a first syntax element is signaled in a picture parameter set (PPS) associated with the picture, the picture including one or more slices having a slice type, the first syntax element indicating that the slice type is signaled in a picture header due to the first syntax element being equal to 0, otherwise the first syntax element indicating that the slice type is signaled in a slice header.

[0405] Figure 16 A flow chart of an example method 1600 for video processing is shown. The method 1600 includes, at operation 1610, performing conversion between a picture of a video and a bitstream of the video according to a rule, the conversion including a loop filtering process, and the rule dictating that a total number of vertical virtual boundaries and a total number of horizontal virtual boundaries associated with the loop filtering operation be signaled at a picture level or a sequence level.

[0406] Figure 17 A flowchart of an example method 1700 for video processing is shown. The method 1700 includes, at operation 1710, performing conversion between a video including one or more pictures and a bitstream of the video, the bitstream conforming to a format rule, the format rule conditionally allowing encoding and decoding of a picture in a layer by using a reference picture from another layer based on a first syntax element, the first syntax element indicating whether the reference picture from the other layer is present in the bitstream, and conditionally signaling the first syntax element in the bitstream based on a second syntax element indicating whether an identifier of a parameter set associated with the picture is not equal to 0.

[0407] Figure 18 A flowchart of an example method 1800 for video processing is shown. The method 1800 includes, at operation 1810, performing conversion between a picture of a video and a bitstream of the video, the bitstream conforming to a format rule, the format rule specifying a first syntax element, the first syntax element enabling (a) synchronization of context variables before decoding a codec tree unit (CTU) in a picture and (b) storage of the context variables after decoding the CTU, the first syntax element being signaled in a sequence parameter set (SPS) associated with the picture.

[0408] Figure 19 A flow chart of an example method 1900 for video processing is shown. The method 1900 includes, at operation 1910, performing conversion between a picture of a video and a bitstream of the video, the bitstream conforming to a format rule, the format rule specifying a syntax element, the syntax element indicating whether an entry point offset for a slice or a slice-specific codec tree unit (CTU) row is signaled in a slice header of the picture, and the syntax element is signaled in a sequence parameter set (SPS) associated with the picture.

[0409] Figure 20 A flowchart of an example method 2000 for video processing is shown. The method 2000 includes, at operation 2010, performing conversion between a video and a bitstream of the video according to a rule, the rule stipulating that a first syntax element is less than a first preset threshold, the first syntax element indicating the number of parameters for an output layer set OLS hypothetical reference decoder HRD in a video parameter set VPS associated with the video.

[0410] Figure 21 A flow chart of an example method 2100 for video processing is shown. The method 2100 includes, at operation 2110, performing conversion between a video and a bitstream of the video according to a rule, wherein the rule stipulates that a syntax element is less than a preset threshold, the syntax element indicating the number of profile / tier / level PTL syntax structures in a video parameter set (VPS) associated with the video.

[0411] Figure 22 A flow chart of an example method 2200 for video processing is shown. The method 2200 includes, at operation 2210, performing conversion between a video and a bitstream of the video according to a rule, the rule specifying a first syntax element, the first syntax element indicating that the number of decoded picture buffer parameter syntax structures in a video parameter set (VPS) must be less than or equal to a second syntax element, the second syntax element indicating the number of layers specified by the VPS.

[0412] Figure 23A flow chart of an example method 2300 for video processing is shown. The method 2300 includes, at operation 2310, performing conversion between a video and a bitstream of the video according to rules that allow a decoder to obtain terminal network abstraction layer NAL units by signaling in the bitstream or providing them by external means.

[0413] Figure 24 A flow chart of an example method 2400 for video processing is shown. The method 2400 includes, at operation 2410, performing conversion between a video and a bitstream of the video, the bitstream conforming to a format rule, and the format rule restricting the bitstream to include only one sub-picture per layer due to a syntax element being equal to 0, indicating that each layer is configured to use inter-layer prediction.

[0414] Figure 25 A flowchart of an example method 2500 for video processing is shown. The method 2500 includes, at operation 2510, performing conversion between a video and a bitstream of the video according to a rule, the rule specifying that a sub-bitstream extraction process be performed to generate a sub-bitstream for decoding, wherein the sub-bitstream extraction process is configured to extract a sub-bitstream having a target highest temporal identifier from the bitstream, and the rule specifying that during the extraction process, when removing a video codec layer (VCL) network abstraction layer (NAL) unit, padding data units and padding supplemental enhancement information (SEI) messages in a supplemental enhancement information (SEI) unit associated with the VCL NAL unit are also deleted.

[0415] Figure 26 A flowchart of an example method 2600 for video processing is shown. The method 2600 includes, at operation 2610, performing conversion between a video unit of a video and a bitstream of the video, the bitstream conforming to a format rule, the format rule specifying that the bitstream includes a first syntax element, the first syntax element indicating whether the video unit is encoded or decoded in a lossy mode or a lossless mode, and signaling a second syntax element indicating selective inclusion of escaped samples in a palette mode applied to the video unit based on a value of the first syntax element.

[0416] The following provides a list of preferred solutions for some embodiments.

[0417] A1. A video processing method comprising: performing conversion between a video and a bitstream of the video, wherein the bitstream comprises one or more access units according to a format rule, and wherein the format rule specifies the order in which a first message and a second message for an operation point OP appear in the access unit AU so that the first message precedes the second message in decoding order.

[0418] A2. The method according to solution A1, wherein the first message comprises a Buffer Period (BP) Supplemental Enhancement Information (SEI) message, and the second message comprises a Picture Timing (PI) SEI message.

[0419] A3. The method of solution A1, wherein the first message comprises a Buffering Period (BP) Supplemental Enhancement Information (SEI) message, and the second message comprises a Decoding Unit (DUI) SEI message.

[0420] A4. The method of solution A1, wherein the first message comprises a Picture Timing (PT) Supplemental Enhancement Information (SEI) message, and the second message comprises a Decoding Unit (DUI) SEI message.

[0421] A5. A video processing method, comprising: performing conversion between a video and a bit stream of the video, wherein the bit stream includes one or more access units according to a format rule, and wherein the format rule specifies the order in which multiple messages for an operation point OP appear in the access unit so that a first message of the multiple messages precedes a second message of the multiple messages in a decoding order.

[0422] A6. The method of solution A5, wherein the plurality of messages includes a buffer period (BP) SEI message, a decoding unit information (DUI) SEI message, a picture timing (PT) SEI message, and a sub-picture level information (SLI) SEI message.

[0423] A7. The method according to solution A6, wherein the decoding order is SLI SEI message, BP SEI message, PT SEI message and DUI message.

[0424] Another list of preferred solutions for some embodiments is provided next.

[0425] B1. A video processing method, comprising: performing conversion between a video including a picture and a bitstream of the video, wherein the bitstream complies with a format rule, wherein the format rule specifies whether an indication of a first flag is signaled at the beginning of a picture header associated with the picture, wherein the first flag indicates whether the picture is an intra-frame random access point (IRAP) picture or a progressive decoding refresh (GDR) picture.

[0426] B2. The method according to solution B1, wherein an IRAP picture is a picture such that when decoding of the bitstream starts from this picture, this picture and all subsequent pictures in output order can be correctly decoded.

[0427] B3. The method according to solution B1 or B2, wherein the IRAP picture comprises only I slices.

[0428] B4. The method according to solution B1, wherein a GDR picture is a picture for which the associated recovery point picture and all subsequent pictures in decoding order and output order can be correctly decoded when decoding of the bitstream starts from this picture.

[0429] B5. The method according to solution B1, wherein the format rule further specifies whether a second flag is signaled in the bitstream based on the indication.

[0430] B6. The method according to solution B5, wherein the first flag is irap_or_gdr_pic_flag.

[0431] B7. The method according to solution B6, wherein the second flag is no_output_of_prior_pics_flag.

[0432] B8. The method according to any of the solutions B1-7, wherein the first flag equal to 1 specifies whether the picture is an IRAP picture or a GDR picture.

[0433] B9. The method according to any of the solutions B1-7, wherein the first flag equal to 0 specifies that the picture is neither an IRAP picture nor a GDR picture.

[0434] B10. The method according to any of the solutions B1-9, wherein each Network Access Layer NAL of a GDR picture has a nal_unit_type syntax element equal to GDR_NUT.

[0435] Another list of preferred solutions for some embodiments is provided next.

[0436] C1. A video processing method, comprising: performing conversion between a video comprising one or more pictures and a bitstream of the video, wherein the bitstream conforms to a format rule, wherein the format rule does not allow encoding or decoding of a picture in the one or more pictures to include a codec slice network abstraction layer (NAL) unit with a progressive decoding refresh type and to be associated with a flag indicating that the picture includes a NAL unit of a mixed type.

[0437] C2. A video processing method, comprising: performing conversion between a video comprising one or more pictures and a bitstream of the video, wherein the bitstream conforms to a format rule, wherein the format rule allows encoding and decoding of a picture in the one or more pictures to include a codec slice network abstraction layer (NAL) unit with a progressive decoding refresh type and to be associated with a flag indicating that the picture does not include a NAL unit of a mixed type.

[0438] C3. Method according to solution C1 or C2, wherein the flag is mixed_nalu_types_in_pic_flag.

[0439] C4. Method according to solution C3, wherein the flag is in the picture parameter set PPS.

[0440] C5. The method according to any of the solutions C1-C4, wherein the picture is a Progressive Decoding Refresh (GDR) picture, and wherein each slice or NAL unit in the picture has nal_unit_type equal to GDR_NUT.

[0441] C6. A video processing method, comprising: performing conversion between a picture of a video and a bitstream of the video, wherein the bitstream conforms to a format rule, the format rule specifying whether a first syntax element is signaled in a picture parameter set (PPS) associated with the picture, wherein the picture includes one or more slices having a slice type, wherein since the first syntax element is equal to 0, the first syntax element indicates that the slice type is signaled in a picture header, otherwise the first syntax element indicates that the slice type is signaled in a slice header.

[0442] C7. Method according to solution C6, wherein the first syntax element is mixed_slice_types_in_pic_flag.

[0443] C8. Method according to solution C6 or C7, wherein the first syntax element is signaled due to the second syntax element being equal to 0 and / or the third syntax element being greater than 0.

[0444] C9. The method according to solution C8, wherein the second syntax element specifies at least one characteristic of one or more slices in one or more slices.

[0445] C10. The method according to solution C9, wherein the second syntax element equal to 0 specifies that the one or more slices are in raster scan order, and slice information is not included in the PPS.

[0446] C11. Method according to solution C9, wherein the second syntax element equal to 0 specifies that one or more slices cover a rectangular area of ​​the picture, and slice information is signaled in the PPS.

[0447] C12. Method according to any of solutions C9-C11, wherein the second syntax element is rect_slice_flag.

[0448] C13. Method according to solution C8, wherein the third syntax element specifies the number of rectangular slices in the picture referring to the PPS.

[0449] C14. Method according to solution C13, wherein the third syntax element is num_slices_in_pic_minus1.

[0450] C15. A video processing method, comprising: performing conversion between a picture of a video and a bitstream of the video according to a rule, wherein the conversion includes a loop filtering process, and wherein the rule stipulates that the total number of vertical virtual boundaries and the total number of horizontal virtual boundaries associated with the loop filtering operation are signaled at the picture level or the sequence level.

[0451] C16. A method according to solution C15, wherein the total number of vertical virtual boundaries includes a first number N1 of vertical virtual boundaries signaled in a picture parameter set (PPS) or in a sequence parameter set (SPS) and a second number N3 of additional vertical virtual boundaries signaled in a picture header (PH), and wherein the total number of horizontal virtual boundaries includes a third number N2 of horizontal virtual boundaries signaled in the PPS or SPS and a fourth number N4 of additional horizontal virtual boundaries signaled in the PH.

[0452] C17. The method according to solution C16, wherein N1+N3≤N1 and N2+N4≤N2.

[0453] C18. The method according to solution C16 or C17, wherein N1=3, N2=3, N3=1, and N4=1.

[0454] Another list of preferred solutions for some embodiments is provided next.

[0455] D1. A video processing method, comprising: performing conversion between a video comprising one or more pictures and a bitstream of the video, wherein the bitstream conforms to a format rule, wherein the format rule conditionally allows encoding and decoding of a picture within a layer by using a reference picture from another layer based on a first syntax element, the first syntax element indicating whether the reference picture from the other layer is present in the bitstream, and wherein the first syntax element is conditionally signaled in the bitstream based on a second syntax element indicating whether an identifier of a parameter set associated with the picture is not equal to 0.

[0456] D2. The method according to solution D1, wherein the first syntax element is inter_layer_ref_pics_present_flag.

[0457] D3. The method according to solution D1 or D2, wherein the first syntax element specifies whether inter-layer prediction is enabled and whether inter-layer reference pictures can be used.

[0458] D4. The method according to solution D1, wherein the first syntax element is sps_inter_layer_prediction_enabled_flag.

[0459] D5. The method according to any of the solutions D1-D4, wherein the first syntax element is in a sequence parameter set (SPS).

[0460] D6. The method according to any of the solutions D1-D4, wherein the parameter set is a video parameter set VPS.

[0461] D7. The method according to solution D6, wherein the second syntax element is sps_video_parameter_set_id.

[0462] D8. The method according to solution D6 or D7, wherein the parameter set provides an identifier for the VPS as a reference for other syntax elements.

[0463] D9. The method according to solution D7, wherein the second syntax element is in a sequence parameter set SPS.

[0464] Another list of preferred solutions for some embodiments is provided next.

[0465] E1. A video processing method, comprising: performing conversion between a picture of a video and a bitstream of the video, wherein the bitstream conforms to a format rule, wherein the format rule specifies a first syntax element, the first syntax element enabling (a) synchronization processing of context variables before decoding a codec tree unit (CTU) in the picture and (b) storage processing of the context variables after decoding the CTU, wherein the first syntax element is signaled in a sequence parameter set (SPS) associated with the picture.

[0466] E2. The method according to solution E1, wherein the picture parameter set PPS associated with the picture does not include the first syntax element.

[0467] E3. Method according to solution E1 or E2, wherein the CTU comprises a first codec tree block (CTB) in a row of codec tree blocks in each slice of the picture.

[0468] E4. The method according to any of the solutions E1-E3, wherein the first syntax element is entropy_coding_sync_enabled_flag.

[0469] E5. The method according to any of solutions E1-E3, wherein the first syntax element is sps_entropy_coding_sync_enabled_flag.

[0470] E6. A method according to any of solutions E1-E5, wherein the format rule specifies a second syntax element, the second syntax element indicating whether there is signaling of an entry point offset for a slice or a slice-specified CTU row in a slice header of the picture, and wherein the second syntax element is signaled in an SPS associated with the picture.

[0471] E7. A video processing method comprising: performing conversion between a picture of a video and a bitstream of the video, wherein the bitstream conforms to a format rule, wherein the format rule specifies a syntax element, the syntax element indicating whether there is signaling of an entry point offset for a slice or a slice-specified codec tree unit (CTU) row in a slice header of the picture, and wherein the syntax element is signaled in a sequence parameter set (SPS) associated with the picture.

[0472] E8. The method according to solution E7, wherein the format rule allows signaling of the entry point offset based on a syntax element with value equal to 1.

[0473] E9. The method according to solution E7 or E8, wherein the picture parameter set PPS associated with the picture does not include syntax elements.

[0474] E10. The method according to any of the solutions E7-E9, wherein the syntax element is entry_point_offsets_present_flag.

[0475] E11. The method according to any of the solutions E7-E9, wherein the syntax element is sps_entry_point_offsets_present_flag.

[0476] Another list of preferred solutions for some embodiments is provided next.

[0477] F1. A video processing method, comprising: performing conversion between a video and a bitstream of the video according to a rule, wherein the rule stipulates that a first syntax element is less than a first preset threshold, and the first syntax element indicates the number of parameters for an output layer set OLS hypothetical reference decoder HRD in a video parameter set VPS associated with the video.

[0478] F2. The method according to solution 1, wherein the first syntax element is vps_num_ols_timing_hrd_params_minus1.

[0479] F3. A method according to solution F1 or F2, wherein the first preset threshold is the number of multi-layer output layer sets (denoted as NumMultiLayerOlss) minus 1.

[0480] F4. The method according to any of the solutions F1-F3, wherein the first syntax element is inferred to be 0 when not signaled in the bitstream.

[0481] F5. A method according to any of the solutions F1-F4, wherein the rule stipulates that a second syntax element is less than a second preset threshold, the second syntax element indicating the number of profile / tier / level PTL syntax structures in the video parameter set VPS associated with the video.

[0482] F6. A video processing method, comprising: performing conversion between a video and a bitstream of the video according to a rule, wherein the rule specifies a syntax element, and the syntax element indicates that the number of profile / tier / level PTL syntax structures in a video parameter set VPS associated with the video is less than a preset threshold.

[0483] F7. Method according to solution F6, wherein the syntax element is vps_num_ptls_minus1.

[0484] F8. A method according to solution F6 or F7, wherein the preset threshold is the total number of output layer sets (denoted as TotalNumOlss).

[0485] F9. The method according to any of the solutions F1-F8, wherein the difference between the preset threshold and the syntax element is signaled in the bitstream.

[0486] F10. The method of solution F9, wherein the difference is signaled using a unary codec.

[0487] F11. The method according to solution F9, wherein the difference is signaled using Exponential-Golomb EG codec.

[0488] F12. A video processing method, comprising: performing conversion between a video and a bitstream of the video according to a rule, wherein the rule specifies a first syntax element, the first syntax element indicating that the number of decoded picture buffer parameter syntax structures in a video parameter set VPS must be less than or equal to a second syntax element, the second syntax element indicating the number of layers specified by the VPS.

[0489] F13. Method according to solution F12, wherein the difference between the first syntax element and a preset threshold is signaled in the bitstream.

[0490] F14. Method according to solution F12, wherein the difference between the second syntax element and a preset threshold is signaled in the bitstream.

[0491] F15. Method according to solution F13 or F14, wherein the difference is signaled using a unary codec.

[0492] F16. Method according to solution F13 or F14, wherein the difference is signaled using Exponential-Golomb EG codec.

[0493] F17. A video processing method, comprising: performing conversion between a video and a bitstream of the video according to a rule, wherein the rule allows a decoder to obtain a terminal network abstraction layer NAL unit by signaling in the bitstream or providing it by external means.

[0494] F18. The method of solution F17, wherein the terminating NAL unit is a terminating EOB of a bitstream NAL unit.

[0495] F19. The method of solution F9, wherein the terminal NAL unit is a terminal EOS of a sequence NAL unit.

[0496] F20. The method according to any of the solutions F17-F19, wherein the external means comprises a parameter set.

[0497] F21. A video processing method, comprising: performing conversion between a video and a bitstream of the video, wherein the bitstream conforms to a format rule, and wherein, due to a syntax element being equal to 0, the format rule restricts each layer in the bitstream to include only one sub-picture, indicating that each layer is configured to use inter-layer prediction.

[0498] F22. Method according to solution F21, wherein the syntax element is vps_independent_layer_flag.

[0499] Another list of preferred solutions for some embodiments is provided next.

[0500] G1. A video processing method, comprising: performing conversion between a video and a bitstream of the video according to rules, wherein the rules stipulate that a sub-bitstream extraction process is implemented to generate a sub-bitstream for decoding, wherein the sub-bitstream extraction process is configured to extract a sub-bitstream with a target highest temporal identifier from the bitstream, and wherein the rules stipulate that during the extraction process, when removing a video codec layer (VCL) network abstraction layer (NAL) unit, a padding data unit and a padding supplemental enhancement information (SEI) message in a supplemental enhancement information (VCL) unit associated with the VCL NAL unit are also deleted.

[0501] G2. Method according to solution G1, wherein the VCL NAL units are removed based on an identifier of the layer to which the VCL NAL belongs.

[0502] G3. The method according to solution G2, wherein the identifier is nuh_layer_id.

[0503] G4. The method of solution G1, wherein the bitstream conforms to a format rule that specifies that a SEI NAL unit including a SEI message with a padding payload does not include a SEI message with a payload different from the padding payload.

[0504] G5. The method according to solution G1, wherein the bitstream conforms to a format rule that stipulates that a SEI NAL unit including a SEI message with a padding payload is configured not to include another SEI message with any other payload.

[0505] G6. The method according to solution G1, wherein a VCL NAL unit has at most one associated padding NAL unit.

[0506] G7. A video processing method, comprising: performing conversion between a video unit of a video and a bitstream of the video, wherein the bitstream conforms to a format rule, wherein the format rule stipulates that the bitstream includes a first syntax element, the first syntax element indicates whether the video unit is encoded or decoded in a lossy mode or a lossless mode, and wherein signaling notifies a second syntax element indicating selective inclusion of escaped samples in a palette mode applied to the video unit based on the value of the first syntax element.

[0507] G8. The method according to solution G7, wherein the binarization method of the escaped samples is based on the first syntax element.

[0508] G9. The method according to solution G7, wherein the context coding in the arithmetic coding of the escape sample is determined based on a flag.

[0509] G10. The method according to any of solutions G7-G9, wherein the video unit is a codec unit CU or a codec tree unit CTU.

[0510] Another list of preferred solutions for some embodiments is provided next.

[0511] P1. A video processing method, comprising performing a conversion between a picture of a video and a codec representation of the video, wherein the number of sub-pictures in the picture is included as a field in the codec representation, and the bit width of the field depends on the value of the number of sub-pictures.

[0512] P2. The method according to solution P1, wherein the field indicates the number of sub-pictures using the codeword.

[0513] P3. The method according to solution P2, wherein the codewords comprise Golomb codewords.

[0514] P4. The method according to any of solutions P1 to P3, wherein the value of the number of sub-pictures is restricted to be less than or equal to the integer number of codec treeblocks that fit within the picture.

[0515] P5. Method according to any of the solutions P1 to P4, wherein the field depends on a codec level associated with the codec representation.

[0516] P6. A video processing method, comprising performing a conversion between a video region of a video and a codec representation of the video, wherein the codec representation conforms to a format rule, wherein the format rule specifies that a syntax element indicating a sub-picture identifier is omitted because the video region does not contain any sub-pictures.

[0517] P7. The method according to solution P6, wherein the codec representation comprises a field with a value of 0, the field with a value of 0 indicating that the video region does not include any sub-picture.

[0518] P8. A video processing method comprising performing a conversion between a video region of a video and a codec representation of the video, wherein the codec representation conforms to a format rule, wherein the format rule specifies omitting identifiers of sub-pictures in the video region at a video region header level of the codec representation.

[0519] P9. A method according to solution P8, wherein the codec indicates that sub-pictures are identified numerically according to the order in which they are listed in the video region header.

[0520] P10. A video processing method, comprising performing conversion between a video region of a video and a codec representation of the video, wherein the codec representation conforms to a format rule, wherein the format rule specifies an identifier of a sub-picture in the video region and / or the length of the identifier of the sub-picture at a sequence parameter set level or a picture parameter set level.

[0521] P11. The method according to solution P10, wherein the length is included at the picture parameter set level.

[0522] P12. A video processing method, comprising performing conversion between a video region of a video and a codec representation of the video, wherein the codec representation conforms to a format rule, wherein the format rule specifies including a field in the codec representation at a video sequence level to indicate whether a sub-picture identifier length field is included in the codec representation at the video sequence level.

[0523] P13. The method of solution P12, wherein the format rule specifies setting the field to 1 in case another field in the codec representation indicates that a length identifier of the video region is included in the codec representation.

[0524] P14. A video processing method, comprising performing a conversion between a video region of a video and a codec representation of the video, wherein the codec representation complies with a format rule, and wherein the format rule provides for including in the codec representation an indication of whether the video region can be used as a reference picture.

[0525] P15. Method according to solution P14, wherein the indication comprises a layer ID and an index or ID value associated with the video region.

[0526] P16. A video processing method comprising performing a conversion between a video region of a video and a codec representation of the video, wherein the codec representation conforms to a format rule, and wherein the format rule provides for including an indication in the codec representation to indicate whether the video region can use inter-layer prediction (ILP) from multiple sample values ​​associated with the video region of a reference layer.

[0527] P17. The method according to solution P16, wherein the indication is included at sequence level, picture level or video level.

[0528] P18. Method according to solution P16, wherein the video region of the reference layer comprises at least one collocated sample of the samples within the video region.

[0529] P19. The method according to solution P16, wherein the indication is included in one or more Supplemental Enhancement Information (SEI) messages.

[0530] P20. A video processing method, comprising, when a first message and a second message of an operation point in an access unit for determining an application video exist in a codec representation of the video, configuring a codec representation in which the first message is superior to the second message in a decoding order, for example; and based on the configuration, performing conversion between a video area and a codec representation of the video.

[0531] P21. The method according to solution P20, wherein the first message comprises a Buffer Period (BP) Enhancement Supplemental Information (SEI) message, and the second message comprises a Picture Timing (PT) SEI message.

[0532] P22. The method according to solution P20, wherein the first message comprises a Buffer Period (BP) Supplemental Enhancement Information (SEI) message, and the second message comprises a Decoding Unit (DUI) SEI message.

[0533] P23. The method according to solution P20, wherein the first message comprises a Picture Timing (PT) Supplemental Enhancement Information (SEI) message, and the second message comprises a Decoding Unit (DUI) SEI message.

[0534] P24. The method of any of the above schemes, wherein the video area includes a sub-picture of the video.

[0535] P25. A method according to any of the above schemes, wherein the conversion comprises parsing and decoding the codec representation to generate the video.

[0536] P26. A method according to any of the above schemes, wherein the conversion comprises encoding the video to generate a codec representation.

[0537] P27. A video decoding device comprising a processor configured to execute any one or more methods of schemes P1 to P26.

[0538] P28. A video encoding device comprising a processor configured to execute any one or more methods of schemes P1 to P26.

[0539] P29. A computer program product having computer code stored thereon, which, when executed, causes a processor to perform the method of any one of schemes P1 to P26.

[0540] Another list of preferred solutions for some embodiments is provided next.

[0541] O1. A method according to any of the above schemes, wherein the conversion comprises decoding the video from the bitstream.

[0542] O2. A method according to any of the above schemes, wherein the conversion comprises encoding the video into a bitstream.

[0543] O3. The method of any of the above schemes, wherein the conversion includes generating a bitstream from the video, and wherein the method further includes: storing the bitstream in a non-transitory computer-readable recording medium.

[0544] O4. A method for storing a bit stream representing a video in a computer-readable recording medium, comprising: generating a bit stream from a video according to the method of any of the above schemes; and writing the bit stream into the computer-readable recording medium.

[0545] O5. A video processing device, comprising a processor, wherein the processor is configured to execute the method of any of the above schemes.

[0546] O6. A computer-readable medium having instructions stored thereon, wherein the instructions, when executed, cause a processor to perform the method of any of the above schemes.

[0547] O7. A computer-readable medium storing a bit stream generated by the method of any of the above schemes.

[0548] O8. A video processing device for storing a bit stream, wherein the video processing device is configured to execute the method of any of the above schemes.

[0549] O9. A bit stream generated using the method in any of the above schemes, wherein the bit stream is stored on a computer-readable medium.

[0550] The disclosures and other solutions, examples, embodiments, modules, and functional operations described in this document can be implemented in digital electronic circuitry, or in computer software, firmware, or hardware, including the structures disclosed in this specification and their structural equivalents, or any combination thereof. The disclosures and other embodiments disclosed in this specification can be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a tangible, non-volatile computer-readable medium, for execution by a data processing apparatus or to control the operation of the data processing apparatus. A computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a storage device, a composition of matter that effects a machine-readable propagated signal, or any combination thereof. The term "data processing unit" or "data processing apparatus" includes all devices, equipment, and machines for processing data, including, for example, a programmable processor, a computer, or a plurality of processors or computers. In addition to hardware, the apparatus may also include code that creates an execution environment for a computer program, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, or any combination thereof. A propagated signal is an artificially generated signal, such as a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to a suitable receiver device.

[0551] A computer program (also referred to as a program, software, software application, script, or code) may be written in any form of programming language (including compiled or interpreted languages) and may be deployed in any form, including as a standalone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program may be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program, or in multiple coordinated files (e.g., files storing one or more modules, subroutines, or portions of code). A computer program may be deployed for execution on one or more computers, located at one site or distributed across multiple sites and interconnected by a communications network.

[0552] The processes and logic flows described in this specification can be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by, and apparatus can be implemented as, special-purpose logic circuitry, such as an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit).

[0553] For example, processors suitable for executing computer programs include general-purpose and special-purpose microprocessors, as well as any one or more of any type of digital computer. Typically, a processor will receive instructions and data from a read-only memory or a random-access memory, or both. The essential components of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include one or more mass storage devices for storing data, such as magnetic, magneto-optical, or optical disks, or be operatively coupled to one or more mass storage devices to receive data from them or transfer data to one or more mass storage devices, or both. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of nonvolatile memory, media, and storage devices, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal or removable hard disks; magneto-optical disks; and CDROM and DVDROM disks. The processor and memory may be supplemented by, or incorporated into, special-purpose logic circuitry.

[0554] While this patent document contains many specifics, they should not be construed as limitations on the scope of any invention or the claims, but rather as descriptions of features for particular embodiments of particular inventions. Certain features described in this patent document in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various functions described in the context of a single embodiment can also be implemented separately in multiple embodiments, or in any suitable subcombination. Furthermore, while the features described above may be described as functioning in certain combinations, or even initially claimed to be so, in some cases one or more features in a claim combination may be removed from the combination, and a claim combination may be directed to a subcombination or variations of a subcombination.

[0555] Likewise, while operations are depicted in a particular order in the drawings, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, in order to achieve desired results. Furthermore, the separation of various system components in the embodiments described in this patent document should not be understood as requiring such separation in all embodiments.

[0556] Only a few implementations and examples are described, and other implementations, enhancements, and variations can be made based on what is described and illustrated in this patent document.

Claims

1. A video processing method, comprising: Perform conversion between pictures of a video and a bitstream of said video, wherein the bitstream complies with the format rules, The format rule specifies a first syntax element that enables (a) synchronization of context variables before decoding a codec tree unit (CTU) in the picture and (b) storage of the context variables after decoding the CTU. The first syntax element is transmitted via a signal in a sequence parameter set SPS associated with the picture, and the first syntax element is not included in a picture parameter set PPS associated with the picture.

2. The method according to claim 1, wherein The CTU includes a first codec tree block (CTB) in a row of codec tree blocks in each slice of the picture.

3. The method according to claim 1, wherein The first syntax element is entropy_coding_sync_enabled_flag.

4. The method according to claim 1, wherein The first syntax element is sps_entropy_coding_sync_enabled_flag.

5. The method according to claim 1, wherein The format rule specifies a second syntax element that indicates whether an entry point offset for a slice or a slice-specified CTU row is signaled in a slice header of the picture, and wherein the second syntax element is signaled in the SPS associated with the picture.

6. The method according to any one of claims 1 to 5, wherein The converting includes decoding the video from the bitstream.

7. The method according to any one of claims 1 to 5, wherein The converting includes encoding the video into the bitstream.

8. A video processing method, comprising: Perform conversion between pictures of a video and a bitstream of said video, wherein the bitstream complies with the format rules, wherein the format rule specifies a syntax element indicating whether there is signaling of an entry point offset for a slice or a slice-specified codec tree unit (CTU) row in a slice header of the picture, and The syntax element is transmitted via a signal in a sequence parameter set SPS associated with the picture, and a picture parameter set PPS associated with the picture does not include the syntax element.

9. The method according to claim 8, wherein The format rule allows the entry point offset to be signaled based on the syntax element having a value equal to 1.

10. The method according to claim 8, wherein The syntax element is entry_point_offsets_present_flag.

11. The method according to claim 8, wherein The syntax element is sps_entry_point_offsets_present_flag.

12. The method according to any one of claims 8 to 11, wherein: The converting includes decoding the video from the bitstream.

13. The method according to any one of claims 8 to 11, wherein: The converting includes encoding the video into the bitstream.

14. The method according to any one of claims 8 to 11, wherein: The converting comprises generating the bitstream from the video, and wherein the method further comprises: The bit stream is stored in a non-transitory computer-readable recording medium.

15. A method of storing a bitstream representing a video in a computer-readable recording medium, comprising: Generating the bitstream from the video according to the method of any one or more of claims 1-11; as well as The bit stream is written to the computer-readable recording medium.

16. A video processing device, comprising a processor, wherein: The processor is configured to execute the method according to any one of claims 1 to 15.

17. A computer-readable medium having instructions stored thereon, wherein: When executed, the instructions cause a processor to perform the method according to any one of claims 1-15.

18. A computer-readable medium storing a bit stream, wherein the bit stream is generated by the method according to any one of claims 1 to 15 executed by a video processing device.

Citation Information

Patent Citations

  • Wavefront parallel processing for video coding

    CN104221381A