Methods, apparatuses, media, and methods of storing a bitstream of a video processing

By modifying the processing rules of the EOS NAL unit, the problem of improper processing of the EOS NAL unit in VVC text was solved, and the correct decoding of multi-layer bitstreams and temporal scalability were achieved, improving the flexibility and efficiency of video encoding and decoding.

CN115699724BActive Publication Date: 2025-12-05DOUYIN CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202180036440.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-05-22
Filing Date
2021-05-21
Publication Date
2025-12-05
Estimated Expiration
2041-05-21

AI Technical Summary

Technical Problem

The existing VVC text processing design for EOS NAL units is confusing and has interoperability issues, leading to improper processing of layer-specific EOS NAL units and affecting temporal scalability and correct decoding of bitstreams.

Method used

By modifying the processing rules of EOS NAL units, it is allowed that the nuh_layer_id of the EOS NAL unit is different from the nuh_layer_id of the associated VCL NAL unit, and multiple EOS NAL units can be contained in one PU. The association order of PU and AU is adjusted to ensure the correct position of the EOS NAL unit in the PU, and the decoding of multi-layer bitstreams is supported.

Benefits of technology

It solves the mess and interoperability problems in EOS NAL unit processing, realizes correct decoding of multi-layer bitstreams and temporal scalability, and improves the flexibility and efficiency of video encoding and decoding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115699724B_ABST
    Figure CN115699724B_ABST
Patent Text Reader

Abstract

Examples of video encoding methods and apparatuses and video decoding methods and apparatuses are described. An example method of video processing includes performing a conversion between a video and a bitstream of the video. The bitstream conforms to a format rule. The bitstream includes one or more layers including one or more picture units (PUs). The format rule specifies, in response to a first PU in a layer of the bitstream following, in a decoding order, a sequence end network abstraction layer (EOS NAL) unit in the layer, setting a variable of the first PU to a particular value, where the variable indicates whether the first PU is a coded layer video sequence start (CLVSS) PU.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross Reference to Related Applications

[0002] This application is based on International Patent Application No. PCT / US2021 / 033724 filed May 21, 2021, which claims priority to and the benefit of U.S. Provisional Patent Application No. 63 / 029,334 filed May 22, 2020. All of the above-identified applications are hereby incorporated by reference in their entirety into this document. TECHNICAL FIELD

[0003] This patent document relates to image and video coding and decoding. BACKGROUND

[0004] Digital video plays an important role in a wide range of applications, such as cultural, educational, and social activities. These applications include advertising, video conferencing, entertainment, and security. The demand for high-quality video is increasing, which can be attributed to the increasing popularity of video streaming, video conferencing, and social media. As the demand for high-quality video increases, the amount of data used to stream video also increases. SUMMARY

[0005] This document discloses techniques that can be used by video encoders and decoders to perform video encoding or decoding.

[0006] In one example aspect, a method of video processing is disclosed. The method includes performing a conversion between a video and a bitstream of the video, wherein the bitstream conforms to a format rule, wherein the bitstream comprises one or more layers comprising one or more picture units (PUs), wherein the format rule specifies, in response to a first PU in a layer of the bitstream following a sequence end network abstraction layer (EOS NAL) unit in the layer in a decoding order, setting a variable of the first PU to a particular value, wherein the variable indicates whether the first PU is a coded layer video sequence start (CLVSS) PU.

[0007] In another example aspect, another method of video processing is disclosed. The method includes performing a conversion between a video and a bitstream of the video, wherein the bitstream comprises one or more layers comprising one or more picture units (PUs) according to a format rule; wherein the format rule specifies that a PU of a particular layer following a sequence end network abstraction layer (EOS NAL) unit of the particular layer is a particular type of PU. In some embodiments, the particular type of PU is one of an intra random access point (IRAP) type or a gradual decoding refresh (GDR) type. In some embodiments, wherein the particular type of PU is a coded layer video sequence start (CLVSS) PU.

[0008] In another example aspect, another video processing method is disclosed. The method includes performing a conversion between a video and a bitstream of the video, wherein the bitstream includes one or more layers comprising one or more picture units (PUs) according to a format rule, wherein the format rule specifies that, when present, an end of sequence (EOS) raw byte sequence payload (RBSP) syntax structure specifies that a next subsequent PU in the bitstream that belongs to a same layer as an EOS network abstraction layer (NAL) unit in a decoding order is of a particular PU type from an intra random access point (IRAP) PU type or a gradual decoding refresh (GDR) PU type. In some embodiments, the particular PU type is the IRAP PU type. In some embodiments, the particular PU type is the GDR PU type.

[0009] In another example aspect, another video processing method is disclosed. The method includes performing a conversion between a video and a bitstream of the video, wherein the bitstream includes one or more layers comprising one or more network abstraction layer (NAL) units according to a format rule, wherein the format rule specifies that a first layer identifier in a header of an end of sequence network abstraction layer (EOS NAL) unit needs to be equal to a second layer identifier of one of the one or more layers in the bitstream. In some embodiments, the format rule further allows more than one EOS NAL unit to be included in a picture unit (PU). In some embodiments, wherein the format rule specifies that the first layer identifier of the EOS NAL unit needs to be less than or equal to a third layer identifier of a video coding layer (VCL) NAL unit associated with the EOS NAL unit.

[0010] In another example aspect, another video processing method is disclosed. The method includes performing a conversion between a video and a bitstream of the video, wherein the bitstream includes one or more layers comprising one or more picture units (PUs) comprising one or more network abstraction layer (NAL) units according to a format rule, wherein the format rule specifies that, in response to a first NAL unit in a PU indicating an end of sequence, the first NAL unit is a last NAL unit among all NAL units in the PU other than another NAL unit indicating another end of sequence or another NAL unit indicating an end of bitstream if present. In some embodiments, the another NAL unit is an end of sequence (EOS) NAL unit. In some embodiments, the another NAL unit is an end of bitstream (EOB) NAL unit.

[0011] In yet another example aspect, a video encoder apparatus is disclosed. The video encoder includes a processor configured to implement a method recited above.

[0012] In yet another example aspect, a video decoder apparatus is disclosed. The video decoder includes a processor configured to implement a method described above.

[0013] In yet another example aspect, a computer readable medium having code stored thereon is disclosed. The code embodies one of the methods described herein in the form of processor-executable code.

[0014] In yet another example aspect, a method of storing a bitstream to a computer readable medium is disclosed. The bitstream is generated using a method described above.

[0015] These and other features are described throughout this document. BRIEF DESCRIPTION OF DRAWINGS

[0016] Figure 1 is a block diagram illustrating a video coding system according to some embodiments of the disclosure.

[0017] Figure 2 is a block diagram of an example hardware platform for video processing.

[0018] Figure 3 is a flowchart of an example method of video processing.

[0019] Figure 4 is a block diagram illustrating an example video coding system.

[0020] Figure 5 is a block diagram illustrating an encoder according to some embodiments of the disclosure.

[0021] Figure 6 is a block diagram illustrating a decoder according to some embodiments of the disclosure.

[0022] Figures 7 to 11 is a flowchart illustrating a method of various video processing. DETAILED DESCRIPTION

[0023] The use of section headings in this document is for ease of understanding and does not limit the applicability of techniques and embodiments disclosed in each section only to that section. Also, the use of H.266 terminology in some descriptions is merely for ease of understanding and is not intended to limit the scope of the disclosed techniques. Thus, the techniques described herein are applicable to other video codec protocols and designs as well.

[0024] 1. INTRODUCTION

[0025] This document relates to video coding technology. In particular, it is about the handling of EOS NAL units in video coding, especially in the context of multi-layer and multi-sublayer. These ideas can be applied individually or in various combinations to any video coding standard or non-standard video codec that supports multi-layer video coding (e.g., the Versatile Video Coding (VVC) that is under development).

[0026] 2Abbreviations

[0027] APS adaptation parameter set

[0028] AU access unit

[0029] AUD access unit delimiter

[0030] AVC advanced video coding

[0031] CLVS coded layer video sequence

[0032] CPB coded picture buffer

[0033] CRA clean random access

[0034] CTU coding tree unit

[0035] CVS coded video sequence

[0036] DCI decoding capability information

[0037] DPB decoded picture buffer

[0038] EOB end of bitstream

[0039] EOS end of sequence

[0040] GDR gradual decoding refresh

[0041] HEVC high efficiency video coding

[0042] HRD hypothetical reference decoder

[0043] IDR instantaneous decoding refresh

[0044] ILP inter-layer prediction

[0045] ILRP inter-layer reference picture

[0046] JEM joint exploration model

[0047] LTRP long-term reference picture

[0048] MCTS motion-constrained tile set

[0049] NAL network abstraction layer

[0050] OLS output layer set

[0051] PH picture header

[0052] PPS picture parameter set

[0053] PTL profile, tier, level

[0054] PU picture unit

[0055] RAP random access point

[0056] RBSP raw byte sequence payload

[0057] SEI supplemental enhancement information

[0058] SPS sequence parameter set

[0059] STRP short-term reference picture

[0060] SVC scalable video coding

[0061] VCL video coding layer

[0062] VPS video parameter set

[0063] VTM VVC test model

[0064] VUI video usability information

[0065] VVC versatile video coding

[0066] 3 preliminary discussion

[0067] Video coding standards have evolved mainly through the development of the well-known ITU-T and ISO / IEC standards. The ITU-T produced H.261 and H.263 standards, ISO / IEC produced the MPEG-1 and MPEG-4 Visual standards, and the two organizations jointly produced the H.262 / MPEG-2 Video and H.264 / MPEG-4 Advanced Video Coding (AVC) and H.265 / HEVC [1] standards. Starting with H.262, the video coding standards are based on the hybrid video coding structure wherein temporal prediction plus transform coding are utilized. To explore future video coding technologies beyond HEVC, the Joint Video Exploration Team (JVET) was founded by VCEG and MPEG jointly in 2015. Since then, the JVET has been incorporating many new methods into the reference software named Joint Exploration Model (JEM) [2]. Currently, the JVET meeting is held every quarter, and the new coding standard targets at 50% bitrate reduction compared to HEVC. The new video coding standard was officially named as Versatile Video Coding (VVC) in the April 2018 JVET meeting, and the first version of VVC test model (VTM) was also released at that time. With continuous effort contributing to VVC standardization, new coding techniques are being adopted into the VVC standard with every JVET meeting. The working draft of VVC and the test model VTM are updated after every meeting. The current goal of VVC project is to achieve the functional completion (FDIS) in the July 2020 meeting.

[0068] 3.1. Reference picture management and reference picture lists (RPLs)

[0069] Reference picture management is a core functionality necessary for any video coding scheme that uses inter prediction. It manages the storage and removal of reference pictures in the decoded picture buffer (DPB) and places the reference pictures in their correct order in the RPLs.

[0070] Reference picture management, including reference picture marking and removal from the decoded picture buffer (DPB) and reference picture list construction (RPLC) in HEVC is different from AVC. Instead of the reference picture marking mechanism based on a sliding window plus adaptive memory management control operation (MMCO) in AVC, HEVC specifies a reference picture management and marking mechanism based on a so-called reference picture set (RPS), thus RPLC is based on the RPS mechanism. An RPS consists of a set of reference pictures associated with a picture, which includes all reference pictures that precede the associated picture in decoding order and that can be used for inter prediction of the associated picture or any picture following the associated picture in decoding order. The reference picture set consists of five reference picture lists. The first three lists contain all reference pictures that can be used for inter prediction of the current picture and for inter prediction of one or more pictures following the current picture in decoding order. The other two lists consist of all reference pictures that are not used for inter prediction of the current picture but can be used for inter prediction of one or more pictures following the current picture in decoding order. The RPS provides an "intra-coded" signaling of the DPB state instead of an "inter-coded" signaling as in AVC, mainly to improve error resilience. The RPLC process in HEVC is based on the RPS by signaling an index to a subset of the RPS for each reference index; this process is simpler than the RPLC process in AVC.

[0071] Reference picture management in VVC is more similar to HEVC than to AVC, but is simpler and more robust. As in those standards, two RPLs, List 0 and List 1, are derived, but they are not based on the reference picture set concept used in HEVC or the automatic sliding window process used in AVC; instead, they are signaled more directly. For an RPL, reference pictures are listed as active and inactive entries, and only active entries can be used as reference indices in inter prediction of CTUs of the current picture. Inactive entries indicate other pictures to be kept in the DPB for reference by other pictures arriving later in the bitstream.

[0072] 3.2. Random Access and Its Support in HEVC and VVC

[0073] Random access refers to accessing and decoding a bitstream starting from a picture that is not the first picture of the bitstream in decoding order. To support tuning and channel switching in broadcast / multicast and multi-party video conferencing, search in local playback and streaming, and stream adaptation in streaming, a bitstream needs to include frequent random access points, which are usually intra-coded pictures, but can also be inter-coded pictures (e.g., in the case of gradual decoding refresh).

[0074] HEVC includes signaling of intra random access points (IRAP) pictures in the NAL unit header by means of NAL unit types. Three types of IRAP pictures are supported, namely instantaneous decoding refresh (IDR), clean random access (CRA), and broken link access (BLA) pictures. IDR pictures constrain the inter-picture prediction structure to not reference any picture before the current group of pictures (GOP), commonly referred to as a closed-GOP random access point. CRA pictures are less constrained by allowing certain pictures to reference pictures before the current GOP, all of which will be discarded in case of a random access. CRA pictures are commonly referred to as open-GOP random access points. BLA pictures typically originate from splicing of two bitstreams or parts thereof at a CRA picture, e.g., during stream switching. To better enable systems to use IRAP pictures, a total of six different NAL units are defined to signal properties of IRAP pictures, which can be used to better match the stream access point types as defined in the ISO base media file format (ISOBMFF) [6], which is used for random access support in dynamic adaptive streaming over HTTP (DASH) [7].

[0075] VVC supports three types of IRAP pictures, two types of IDR pictures (one type with or another type without associated RADL pictures), and one type of CRA picture. These are essentially the same as in HEVC. The BLA picture type in HEVC is not included in VVC, mainly for two reasons: i) the basic functionality of BLA pictures can be achieved by a CRA picture plus an end of sequence NAL unit, whose presence indicates that the following pictures start a new CVS in a single-layer bitstream. ii) during the development of VVC, it was desired to specify fewer NAL unit types than HEVC, as indicated by the use of 5 bits instead of 6 bits for the NAL unit type field in the NAL unit header.

[0076] Another key difference in random access support between VVC and HEVC is that GDR is supported in a more normative way in VVC. In GDR, the decoding of a bitstream can start from an inter-coded picture, although at the beginning not the entire picture region can be correctly decoded, but after multiple pictures, the entire picture region will be correct. AVC and HEVC also support GDR, using the recovery point SEI message to signal the GDR random access point and the recovery point. In VVC, a new NAL unit type is specified to indicate a GDR picture, and the recovery point is signaled in the picture header syntax structure. It is allowed for a CVS and a bitstream to start with a GDR picture. This means that it is allowed for an entire bitstream to contain only inter-coded pictures, without a single intra-coded picture. The main benefit of specifying GDR support in this way is that it provides a consistent behavior for GDR. GDR enables an encoder to smooth the bit rate of a bitstream by distributing intra-coded slices or blocks in multiple pictures instead of intra-coding an entire picture, allowing significant reduction of end-to-end delay, which is considered more important today than before as ultra-low delay applications such as wireless display, online gaming, drone-based applications, etc. become more popular.

[0077] Another GDR-related feature in VVC is virtual boundary signaling. The boundary between the refreshed region (i.e., correctly decoded region) and the un-refreshed region at a picture between a GDR picture and its recovery point can be signaled as a virtual boundary, and when signaled, no in-loop filtering across the boundary will be applied, so there will be no decoding mismatch for some samples at or near the boundary. This can be useful when the application determines to display the correctly decoded region during the GDR process.

[0078] IRAP pictures and GDR pictures can be collectively referred to as random access point (RAP) pictures.

[0079] 3.3. Picture resolution change within a sequence

[0080] In AVC and HEVC, the spatial resolution of a picture cannot change unless a new sequence with a new SPS starts with an IRAP picture. VVC allows changing the picture resolution within a sequence at a location without encoding an IRAP picture, which is always intra-coded. This feature is sometimes referred to as reference picture resampling (RPR) because when the reference picture has a different resolution from the current picture being decoded, the feature requires resampling of the reference picture used for inter prediction.

[0081] The scaling ratio is limited to be greater than or equal to 1 / 2 (2x downsampling from the reference picture to the current picture) and less than or equal to 8 (8x upsampling). Three sets of resampling filters with different frequency cutoffs are specified to handle various scaling ratios between the reference picture and the current picture. The three sets of resampling filters are applied to scaling ratios ranging from 1 / 2 to 1 / 1.75, from 1 / 1.75 to 1 / 1.25, and from 1 / 1.25 to 8, respectively. Each set of resampling filters has 16 taps for luma and 32 taps for chroma, which is the same as the case of the motion compensation interpolation filter. The regular MC interpolation process is actually a special case of the resampling process, where the scaling ratio ranges from 1 / 1.25 to 8. The horizontal and vertical scaling ratios are derived based on the picture width and height and the left, right, top, and bottom scaling offsets specified for the reference picture and the current picture.

[0082] Other aspects of the VVC design that support this feature differ from HEVC include: i) picture resolution and corresponding conformance window are signaled in PPS instead of SPS, while the maximum picture resolution is signaled in SPS. ii) For single-layer bitstreams, each picture store (a slot in the DPB for storing one decoded picture) occupies the buffer size required to store a decoded picture with the maximum picture resolution.

[0083] 3.4. Scalable Video Coding (SVC) in General and in VVC

[0084] Scalable Video Coding (SVC, sometimes also referred to as scalability in video coding) refers to video coding using a base layer (BL) (sometimes referred to as a reference layer (RL)) and one or more scalable enhancement layers (ELs). In SVC, the base layer can carry video data with a base level of quality. The one or more enhancement layers can carry additional video data to support, for example, higher spatial, temporal, and / or signal-to-noise ratio (SNR) levels. An enhancement layer can be defined relative to a previously coded layer. For example, a bottom layer can serve as a BL, while a top layer can serve as an EL. An intermediate layer can serve as an EL or an RL, or both. For example, an intermediate layer (e.g., a layer that is neither the lowest layer nor the highest layer) can be an EL of the layers below the intermediate layer (e.g., the base layer or any intervening enhancement layers) and, at the same time, serve as an RL of the one or more enhancement layers above the intermediate layer. Similarly, in the Multiview or 3D extension of the HEVC standard, there can be multiple views, and information of one view can be used to code (e.g., encode or decode) information of another view (e.g., motion estimation, motion vector prediction, and / or other redundancies).

[0085] In SVC, the parameters used by the encoder or decoder are grouped into parameter sets based on the level of coding (e.g., video level, sequence level, picture level, slice level, etc.) at which they can be utilized. For example, parameters that can be used by one or more coded video sequences of different layers in a bitstream can be included in a video parameter set (VPS), and parameters that can be used by one or more pictures in a coded video sequence can be included in a sequence parameter set (SPS). Similarly, parameters that are used by one or more slices in a picture can be included in a picture parameter set (PPS), and other parameters that are specific to a single slice can be included in a slice header. Similarly, an indication of which parameter set(s) a particular layer uses at a given time can be provided at various levels of coding.

[0086] Thanks to the support of reference picture resampling (RPR) in VVC, the support of bitstreams containing multiple layers (e.g., two layers with SD and HD resolutions in VVC) can be designed without any additional signaling of the level of coding tools, as the upsampling needed for spatial scalability support can just use the RPR upsampling filter. However, for scalability support, high level syntax changes (compared to not supporting scalability) are needed. Scalability support is specified in VVC version 1. The design of VVC scalability has been made as friendly as possible to the single-layer decoder design. The decoding capability for multi-layer bitstreams is specified in a way as if there is only a single layer in the bitstream. For example, the decoding capability such as DPB size is specified in a way independent of the number of layers in the bitstream to be decoded. Basically, a decoder designed for single-layer bitstreams can decode multi-layer bitstreams with not much change. Compared to the design of multi-layer extensions of AVC and HEVC, the HLS aspect is significantly simplified at the expense of some flexibility. For example, an IRAP AU needs to contain pictures of every layer present in the CVS.

[0087] 3.5. Parameter sets

[0088] AVC, HEVC, and VVC specify parameter sets. The types of parameter sets include SPS, PPS, APS, and VPS. SPS and PPS are supported by all of AVC, HEVC, and VVC. VPS is introduced from HEVC and is included in both HEVC and VVC. APS is not included in AVC or HEVC, but is included in the latest VVC draft text.

[0089] SPS is designed to carry sequence-level header information, and PPS is designed to carry picture-level header information that does not change frequently. With SPS and PPS, the infrequently changing information does not need to be repeated for each sequence or picture, thus avoiding the redundant signaling of this information. Furthermore, the use of SPS and PPS enables the out-of-band transmission of important header information, thus not only avoiding the need for redundant transmission, but also improving error resilience.

[0090] The VPS is introduced to carry sequence-level header information that is common to all layers in a multi-layer bitstream.

[0091] The APS is introduced to carry picture-level or slice-level information that requires a considerable number of bits to code, can be shared by multiple pictures, and can have considerable variation across a sequence.

[0092] 4 Technical problems solved by the technical solutions disclosed

[0093] The existing design for handling EOS NAL units in the latest VVC text (in JVET-R2001-vA / vl0) has the following problems:

[0094] 1) In clause 3 (definitions), as part of the definition of a CLVS picture, the phrase “the first PU in the layer of the bitstream following the EOS NAL unit in decoding order” is problematic, because the EOS NAL unit is layer-specific and the EOS NAL unit only applies to the layer with nuh layer id equal to the nuh layer id of the EOS NAL unit. Therefore, this will cause confusion and interoperability issues.

[0095] 2) In clause 7.4.2.2 (NAL unit header semantics), it is specified that when nal_unit_type is equal to PH NUT, EOS NUT or FD NUT, nuh layer id shall be equal to the nuh layer id of the associated VCL NAL unit. However, this does not fully enable temporal scalability, for example the operation of extracting a temporal subset of a temporally scalable bitstream while preserving the EOS NAL units of each layer in the extracted output. For example, assume there are two layers with nuh layer id equal to 0 and 1, and each layer has two sub-layers with Temporalld equal to 0 and 1. In an AU n with n greater than 0 and Temporalld equal to 1, there is one EOS NAL unit in each PU, and the two EOS NAL units have nuh layer id equal to 0 and 1. Note that any EOS NAL unit is required to have Temporalld equal to 0. By an extraction process that preserves only the lowest sub-layer in each layer, the NAL units with Temporalld equal to 1 would be removed, and thus, the two EOS NAL units in AU n

[0096] would become part of the PUs with nuh layer id equal to 1 in AU n-1. This would violate the rule that the nuh layer id of an EOS NAL unit shall be equal to the nuh layer id of the associated VCL NAL unit. Therefore, it is needed to allow the nuh layer id of an EOS NAL unit to be different from the nuh layer id of the associated VCL NAL unit, and also to allow one PU to contain multiple EOS NAL units.

[0097] 3) In clause 7.4.2.4.3 (Order of PUs and their association with AUs), it is specified that when present, the next PU of a particular layer following a PU that belongs to the same layer and contains an EOS NAL unit shall be a CLVSS PU. However, as mentioned above, it is needed to allow the nuh layer id of an EOS NAL unit

[0098] to be different from the nuh layer id of the associated VCL NAL unit. Therefore, the constraint here needs to be changed accordingly.

[0099] 4) In clause 7.4.2.4.4 (Order of NAL units and coded pictures and their association with PUs), it is specified that when an EOS NAL unit is present in a PU, it shall be the last NAL unit in the PU other than EOB NAL

[0100] the last NAL unit among all NAL units other than the units, if any, of the current PU. However, as mentioned above, it is needed to allow one PU to contain more than one EOS NAL unit. Therefore, this constraint needs to be changed accordingly here.

[0101] 5) In clause 7.4.3.10 (End of sequence RBSP semantics), it is specified that when present, the EOS RBSP

[0102] specifies that the current PU is the last PU in the CLVS in decoding order and the next subsequent PU in decoding order in the bitstream, if any, is an IRAP or GDR PU. However, as mentioned above, a PU can contain EOS NAL units of different layers, therefore, this constraint needs to be changed accordingly.

[0103] 5 List of technical solutions and embodiments

[0104] To address the above and other issues, methods summarized as follows are disclosed. These items should be considered as examples to explain the general concepts and should not be interpreted in a narrow way. Moreover, these items can be used individually or in any combination.

[0105] 1) To address issue 1, in clause 3 (Definitions), in the note that is part of the definition of a CLVS picture, change the phrase “the first PU in the layer of the bitstream in decoding order following the EOS NAL unit” to “the first PU in the layer of the bitstream in decoding order following the EOS NAL unit In this layer in the CLVS”.

[0106] 2) To address issue 2, instead of requiring that the nuh layer id of the EOS NAL unit is equal to the nuh layer id of the associated VCL NAL unit, it is specified that the nuh layer id of the EOS NAL unit

[0107] shall be equal to one of the nuh layer id values of the layers present in the CVS.

[0108] a. Furthermore, in one example, it is allowed that one PU contains more than one EOS NAL unit.

[0109] b. Furthermore, in one example, the value of the nuh layer id of the EOS NAL unit needs to be less than or equal to the nuh layer id of the associated VCL NAL unit.

[0110] 3) To address issue 3, it is specified that when present, the next PU of a particular layer following the EOS NAL units belonging to the same layer shall be an IRAP or GDR PU.

[0111] a. Alternatively, it is specified that when present, the next PU of a particular layer following a CLVSS PU belonging to the same layer shall be a CLVSS PU.

[0112] 4) To address issue 4, it is specified that when an EOS NAL unit is present in a PU, it shall be the last NAL unit among all NAL units in the PU except other EOS NAL units (when present) or EOB NAL units (when present).

[0113] 5) To address issue 4, it is specified that when present, the EOS RBSP specifies that the next subsequent PU (if any) in the bitstream that belongs to the same layer as the EOS NAL unit in decoding order is an IRAP or GDR PU.

[0114] 6 Embodiments

[0115] Below are some example embodiments of some aspects of the invention summarized in Section 5 of the previous section that can be applied to the VVC specification. The changed text is based on the latest VVC text in JVET-R2001-vA / vlO. Most of the added or modified relevant parts are highlighted in Bold underlined boldface italics, and parts of the deleted parts are highlighted in [[double square brackets in boldface italics]]. There can be other editorial nature changes that are not highlighted.

[0116] 6.1. First Embodiment

[0117] This embodiment addresses issues 1, 2, 2a, 2b, 3, 4, and 5.

[0118] 3 Definition ...

[0120] coded layer video sequence (CLVS): a sequence of PUs with the same nuh layer id value, consisting of CLVSS PUs in decoding order, followed by zero or more PUs that are not CLVSS PUs, including all subsequent PUs up to, but not including, any subsequent PU that is a CLVSS PU.

[0121] NOTE - A CLVSS PU can be an IDR PU, a CRA PU, or a GDR PU. For each IDR PU, the value of NoOutputBeforeRecoveryFlag is equal to 1, and each CRA PU has HandleCraAsClvsStartFlag equal to 1, and each CRA or GDR PU is the first PU in the layer of the bitstream in decoding order or the first PU in the layer of the bitstream in decoding order In this layer following an EOS NAL unit. ...

[0123] 7.4.2.2 NAL unit header semantics ...

[0125] nuh layer id specifies the identifier of the layer to which a VCL NAL unit belongs or the identifier of the layer to which a non-VCL NAL unit applies. The value of nuh layer id shall be in the range of 0 to 55, inclusive. Other values of nuh layer id are reserved for future use by ITU-T | ISO / IEC.

[0126] The value of nuh layer id shall be the same for all VCL NAL units of a coded picture. The value of nuh layer id for a coded picture or a PU is the value of nuh layer id for the VCL NAL units of the coded picture or the PU.

[0127] When nal unit type is equal to PH NUT, [[EOS NUT,]] or FD NUT, nuh layer id shall be equal to the nuh layer id of the associated VCL NAL unit.

[0128] When nal_unit_type is equal to EOS NUT, nuh layer id shall be equal to one of the nuh layer id values of the layers present in the CVS. EOS NAL unit

[0129] NOTE 1 - The value of nuh layer id for DCI, VPS, AUD, and EOB NAL units is not constrained. ...

[0131] 7.4.2.4.3 Order of PUs and their association with AUs ...

[0133] The requirement of bitstream conformance is that, when present, the [[PUs belonging to the same layer [[and containing EOS NAL units]]]] IRAP or GDR the next PU of the particular layer shall be PU other EOS NAL units (when present) or [[a CLVSS PU that is either an IRAP PU with NoOutputBeforeRecoveryFlag equal to 1 or a GDR PU with NoOutputBeforeRecoveryFlag equal to 1.]] ...

[0135] 7.4.2.4.4 Order of NAL units and coded pictures and their association with PUs

[0136] A PU consists of zero or one PH NAL unit, one coded picture (including one or more VCL NAL units), and zero or more other non-VCL NAL units. The association of VCL NAL units to a coded picture is described in clause 7.4.2.4.5.

[0137] When a picture consists of multiple VCL NAL units, there shall be a PH NAL unit in the PU.

[0138] A VCL NAL unit is the first VCL NAL unit of a picture when sh_picture_header_in_slice_header_flag of the VCL NAL unit is equal to 1 or the VCL NAL unit is the first VCL NAL unit after a PH NAL unit.

[0139] The order of non-VCL NAL units (except AUD and EOB NAL units) within a PU shall follow the below constraints:

[0140] - When a PH NAL unit is present in the PU, it shall be located before the first VCL NAL unit of the PU.

[0141] - When any DCI NAL unit, VPS NAL unit, SPS NAL unit, PPS NAL unit, prefix SEI NAL unit, NAL unit with nal_unit_type equal to RSV_NVCL_26, or NAL unit with nal_unit_type in the range of UNSPEC_28..UNSPEC_29 is present in the PU, it shall not be located after the last VCL NAL unit of the PU.

[0142] - When any DCI NAL unit, VPS NAL unit, SPS NAL unit, or PPS NAL unit is present in the PU, it shall be located before the PH NAL unit (if present) of the PU and shall be located before the first VCL NAL unit of the PU.

[0143] - NAL units with nal_unit_type equal to SUFFIX_SEI_NUT, FD_NUT, or RSV_NVCL_27, or in the range of UNSPEC_30..UNSPEC_31 in the PU shall not be located before the first VCL NAL unit of the PU.

[0144] - When any prefix APS NAL unit is present in the PU, it shall be located before the first VCL unit of the PU.

[0145] - When any suffix APS NAL units are present in a PU, they shall follow the last VCL unit of the PU.

[0146] - When an EOS NAL unit is present in a PU, it shall be the last NAL unit among all NAL units within the PU except When present in the bitstream, an EOS NAL unit is considered to belong or be in the layer with nuh layer id equal to the nuh layer id of the EOS NAL unit. the EOB NAL unit (when present).

[0147] 7.4.3.10 End of sequence RBSP semantics

[0148] the same layer as the EOS NAL unit Figure 1

[0149] When present, the EOS RBSP specifies [[the current PU is the last PU in the CLVS in decoding order and]] that the next subsequent PU (if any) is an IRAP or GDR PU. Figure 2

[0150] The SODB and the RBSP syntax content of the EOS RBSP are empty.

[0151] Figure 4 is a block diagram of an example video processing system 1900 that can implement the various techniques disclosed herein. Various implementations can include some or all of the components of system 1900. System 1900 can include an input 1902 for receiving video content. The video content can be received in a raw or uncompressed format (e.g., 8 or 10 bit multi-component pixel values) or can be received in a compressed or encoded format. Input 1902 can represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, passive optical network (PON), etc., and wireless interfaces such as Wi-Fi or cellular interfaces.

[0152] ​The system 1900 can include a coding component 1904 that can implement various coding or encoding methods described in this document. The coding component 1904 can reduce the average bitrate of a video from an input 1902 to an output of the coding component 1904 to produce a coded representation of the video. The coding techniques are thus sometimes called video compression or video transcoding techniques. The output of the coding component 1904 can be stored or transmitted via a communication connection, as represented by the component 1906. The stored or communicated bitstream (or coded) representation of the video received at the input 1902 can be used by a component 1908 to generate pixel values or displayable video that is sent to a display interface 1910. The process of generating user- visible video from a bitstream representation is sometimes called video decompression. Also, although certain video processing operations are referred to as “coding” operations or tools, it should be understood that a coding tool or operation is used at an encoder, and a corresponding decoding tool or operation that reverses the result of the coding will be performed by a decoder.

[0153] Examples of peripheral bus interfaces or display interfaces can include Universal Serial Bus (USB) or High Definition Multimedia Interface (HDMI) or DisplayPort, etc. Examples of storage interfaces include SATA (Serial Advanced Technology Attachment), PCI, IDE interfaces, etc. The techniques described in this document can be implemented in various electronic devices such as mobile phones, laptops, smart phones, or other devices capable of processing digital data and / or displaying video.

[0154] Figure 4 is a block diagram of a video processing device 3600. The device 3600 can be used to implement one or more methods described herein. The device 3600 can be implemented in a smart phone, a tablet, a computer, an Internet of Things (IoT) receiver, etc. The device 3600 can include one or more processors 3602, one or more memories 3604, and video processing circuitry 3606. The processor(s) 3602 can be configured to implement one or more methods described in this document. The memory(ies) 3604 can be used for storing data and code used for implementing the methods and techniques described herein. The video processing circuitry 3606 can be used to implement, in hardware circuitry, some of the techniques described in this document.

[0155] Figure 5 is a block diagram illustrating an example video coding system 100 that can utilize the techniques of this disclosure.

[0156] As Figure 4As shown, video coding system 100 can include a source device 110 and a destination device 120. Source device 110 generates encoded video data, and can be referred to as a video encoding device. Destination device 120 can decode the encoded video data generated by source device 110, and can be referred to as a video decoding device.

[0157] Source device 110 can include a video source 112, a video encoder 114, and an input / output (EO) interface 116.

[0158] Video source 112 can include a source such as a video capture device, an interface to receive video data from a video content provider, and / or a computer graphics system for generating video data, or a combination of such sources. Video data can comprise one or more pictures. Video encoder 114 encodes video data from video source 112 to generate a bitstream. The bitstream can include a sequence of bits that form a coded representation of the video data. The bitstream can include coded pictures and associated data. A coded picture is a coded representation of a picture. Associated data can include sequence parameter sets, picture parameter sets, and other syntax structures. EO interface 116 can include a modulator / demodulator (modem) and / or a transmitter. Encoded video data can be transmitted directly to destination device 120 via EO interface 116 through network 130a. The encoded video data can also be stored onto storage medium / server 130b for access by destination device 120.

[0159] Destination device 120 can include an EO interface 126, a video decoder 124, and a display device 122.

[0160] EO interface 126 can include a receiver and / or a modem. EO interface 126 can acquire encoded video data from source device 110 or storage medium / server 130b. Video decoder 124 can decode the encoded video data. Display device 122 can display the decoded video data to a user. Display device 122 can be integrated with destination device 120, or can be external to destination device 120, configured to interface with destination device 120.

[0161] Video encoder 114 and video decoder 124 can operate according to a video compression standard, such as the High Efficiency Video Coding (HEVC) standard, Versatile Video Coding (VVC) standard, and other current and / or further standards.

[0162] Figure 5 is a block diagram illustrating an example of a video encoder 200, which can be Figure 5 video encoder 114 in system 100 shown in FIG. 1.

[0163] Video encoder 200 can be configured to perform any or all of the techniques of this disclosure. In Figure 6 In an example, video encoder 200 includes a plurality of functional components. The techniques described in this disclosure can be shared among the various components of video encoder 200. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.

[0164] The functional components of video encoder 200 can include partition unit 201, prediction unit 202, which can include mode select unit 203, motion estimation unit 204, motion compensation unit 205, and intra-prediction unit 206, residual generation unit 207, transform unit 208, quantization unit 209, inverse quantization unit 210, inverse transform unit 211, reconstruction unit 212, buffer 213, and entropy encoding unit 214.

[0165] In other examples, video encoder 200 can include more, less, or different functional components. In one example, prediction unit 202 can include an intra-block copy (IBC) unit. The IBC unit can make predictions in IBC mode, in which at least one reference picture is the picture in which the current video block resides.

[0166] Furthermore, some components, such as motion estimation unit 204 and motion compensation unit 205, can be highly integrated, but are represented separately in Figure 4 examples for purposes of explanation.

[0167] Partition unit 201 can partition a picture into one or more video blocks. Video encoder 200 and video decoder 300 can support various video block sizes.

[0168] Mode select unit 203 can select one of the coding modes (e.g., intra or inter), for example, based on the residual results, and provide the resulting intra or inter coded block to residual generation unit 207 to generate residual block data and to reconstruction unit 212 to reconstruct the coded block for use as a reference picture. In some examples, mode select unit 203 can select a combined intra and inter prediction (CIIP) mode, in which the prediction is based on both an inter prediction signal and an intra prediction signal. Mode select unit 203 can also select a resolution for a motion vector for a block in the case of inter prediction (e.g., sub-pixel or integer-pixel precision).

[0169] To perform inter prediction for a current video block, motion estimation unit 204 can generate motion information for the current video block by comparing one or more reference frames from buffer 213 to the current video block. Motion compensation unit 205 can determine a predicted video block for the current video block based on the motion information and decoded samples from pictures in buffer 213 (rather than the picture associated with the current video block).

[0170] Motion estimation unit 204 and motion compensation unit 205 can perform different operations for a current video block, e.g., depending on whether the current video block is in an I slice, a P slice, or a B slice.

[0171] In some examples, motion estimation unit 204 can perform uni-prediction for a current video block, and motion estimation unit 204 can search for a reference video block for the current video block in reference pictures in List 0 or List 1. Motion estimation unit 204 can then generate a reference index indicating a reference picture in List 0 or List 1 that contains the reference video block and a motion vector indicating a spatial displacement between the current video block and the reference video block. Motion estimation unit 204 can output the reference index, a prediction direction indicator, and the motion vector as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current block based on the reference video block indicated by the motion information for the current video block.

[0172] In other examples, motion estimation unit 204 can perform bi-prediction for a current video block, and motion estimation unit 204 can search for a reference video block for the current video block in reference pictures in List 0 and also search for another reference video block for the current video block in reference pictures in List 1. Motion estimation unit 204 can then generate a reference index indicating a reference picture in List 0 or List 1 that contains the reference video block and a motion vector indicating a spatial displacement between the reference video block and the current video block. Motion estimation unit 204 can output the reference index and the motion vector for the current video block as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current video block based on the reference video block indicated by the motion information for the current video block.

[0173] In some examples, motion estimation unit 204 can output a full set of motion information for a current video block for decoding processing at a decoder.

[0174] In some examples, motion estimation unit 204 can not output a full set of motion information for a current video block. Instead, motion estimation unit 204 can signal motion information for the current video block with reference to motion information for another video block. For example, motion estimation unit 204 can determine that the motion information for the current video block is sufficiently similar to the motion information for a neighboring video block.

[0175] In one example, the motion estimation unit 204 can indicate, in a syntax structure associated with the current video block, a value that indicates to the video decoder 300 that the current video block has the same motion information as another video block.

[0176] In another example, the motion estimation unit 204 can identify, in a syntax structure associated with the current video block, another video block and a motion vector difference (MVD). The motion vector difference indicates a difference between a motion vector of the current video block and a motion vector that indicates the video block. The video decoder 300 can use the motion vector that indicates the video block and the motion vector difference to determine the motion vector of the current video block.

[0177] As discussed above, the video encoder 200 can predictively signal motion vectors. Two examples of predictively signaling techniques that can be implemented by the video encoder 200 include advanced motion vector prediction (AMVP) and Merge mode signaling.

[0178] The intra prediction unit 206 can perform intra prediction on the current video block. When the intra prediction unit 206 performs intra prediction on the current video block, the intra prediction unit 206 can generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block can include a predicted video block and various syntax elements.

[0179] The residual generation unit 207 can generate residual data for the current video block by subtracting (e.g., represented by a minus sign) the prediction video block(s) for the current video block from the current video block. The residual data for the current video block can include residual video blocks that correspond to different sample components of samples in the current video block.

[0180] In other examples, such as in skip mode, there can not be residual data for the current video block for the current video block, and the residual generation unit 207 can not perform the subtraction operation.

[0181] The transform processing unit 208 can generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video blocks associated with the current video block.

[0182] After the transform processing unit 208 generates the transform coefficient video blocks associated with the current video block, the quantization unit 209 can quantize the transform coefficient video blocks associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.

[0183] The inverse quantization unit 210 and the inverse transform unit 211 can apply inverse quantization and inverse transform, respectively, to the transform coefficient video block to reconstruct a residual video block from the transform coefficient video block. The reconstruction unit 212 can add the reconstructed residual video block to corresponding samples from one or more prediction video blocks generated by the prediction unit 202 to produce a reconstructed video block associated with the current block for storage in the buffer 213.

[0184] After the reconstruction unit 212 reconstructs the video block, an in-loop filtering operation can be performed to reduce video blockiness artifacts in the video block.

[0185] The entropy encoding unit 214 can receive data from other functional components of the video encoder 200. When the entropy encoding unit 214 receives data, the entropy encoding unit 214 can perform one or more entropy encoding operations to generate entropy encoded data and output a bitstream that includes the entropy encoded data.

[0186] Figure 6 is a block diagram illustrating an example of a video decoder 300 that can be Figure 6 the video decoder 124 in the system 100 illustrated in FIG.

[0187] The video decoder 300 can be configured to perform any or all of the techniques of the present disclosure. In Figure 5 examples, the video decoder 300 includes a plurality of functional components. The techniques described in the present disclosure can be shared among the various components of the video decoder 300. In some examples, a processor can be configured to perform any or all of the techniques described in the present disclosure.

[0188] In Figure 3 examples, the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, the video decoder 300 can perform a decoding process generally reciprocal to the encoding process described with respect to the video encoder 200 Figure 7 ) above.

[0189] The entropy decoding unit 301 can retrieve an encoded bitstream. The encoded bitstream can include entropy encoded video data (e.g., encoded blocks of video data). The entropy decoding unit 301 can decode the entropy encoded video and, from the entropy decoded video data, the motion compensation unit 302 can determine motion information including motion vectors, motion vector precisions, reference picture list indices, and other motion information. The motion compensation unit 302 may, for example, determine such information by performing AMVP and Merge modes.

[0190] Motion compensation unit 302 can generate the motion compensated block, possibly interpolating based on an interpolation filter. An identifier of the interpolation filter to be used with sub-pixel precision can be included in the syntax elements.

[0191] Motion compensation unit 302 can use the interpolation filter used by video encoder 200 during encoding of the video block to calculate interpolated values for sub-integer pixels of the reference block. Motion compensation unit 302 can determine the interpolation filter used by video encoder 200 from the received syntax information and use the interpolation filter to generate the prediction block.

[0192] Motion compensation unit 302 can use some of the syntax information to determine the size of the blocks used to encode the frame(s) and / or slice(s) of the coded video sequence, partitioning information describing how each macroblock of a picture of the coded video sequence is partitioned, a mode indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-coded block, and other information used in decoding the coded video sequence.

[0193] Intra prediction unit 303 can use, for example, intra prediction modes received in the bitstream to form the prediction block from spatial neighboring blocks. Inverse quantization unit 303 inverse quantizes, i.e., de-quantizes, quantized video block coefficients provided in the bitstream and decoded by entropy decoding unit 301. Inverse transform unit 303 applies an inverse transform.

[0194] Reconstruction unit 306 can sum the residual block with the corresponding prediction block, generated by motion compensation unit 202 or intra prediction unit 303, to form a decoded block. If desired, a deblocking filter can also be applied to the filtered decoded block in order to remove blockiness artifacts. The decoded video block is then stored in buffer 307, which provides reference blocks for subsequent motion compensation / intra prediction, and also produces decoded video for presentation on a display device.

[0195] Next, a list of preferred solutions of some embodiments is provided.

[0196] The following solutions show example embodiments of the techniques discussed in the previous section (e.g., item 1).

[0197] 1. A method of video processing (e.g., method 600 of Figure 8 includes performing (602) a conversion between a video comprising one or more video pictures and a bitstream representation of the video, wherein the coded representation conforms to a format rule, wherein the format rule specifies that a first picture unit (PU) in a layer of the bitstream follows, in decoding order, a sequence end network abstraction layer (EOS NAL) unit in the layer.

[0198] The following solutions show example embodiments of the techniques discussed in the previous section (e.g., item 2).

[0199] 2. A method of video processing, comprising performing a conversion between a video comprising one or more video pictures and a bitstream representation of the video, wherein the bitstream representation comprises a coded video sequence with one or more video layers, wherein the bitstream representation conforms to a format rule that specifies a layer id of a sequence end network abstraction layer (EOS NAL) unit is equal to another layer id of one of the video layers in the coded video sequence.

[0200] 3. The method of solution 1, wherein the format rule further allows for more than one EOS NAL unit to be included in a picture unit (PU).

[0201] The following solutions show example embodiments of the techniques discussed in the previous section (e.g., item 1).

[0202] 4. A method of video processing, comprising performing a conversion between a video comprising one or more video pictures and a bitstream representation of the video, wherein the bitstream representation comprises a coded video sequence with one or more video layers, wherein the bitstream representation conforms to a format rule that specifies a next picture unit of a particular layer after a sequence end network abstraction layer (EOS NAL) unit belonging to the same layer is an intra random access point or a gradual decoding refresh picture unit.

[0203] The following solutions show example embodiments of the techniques discussed in the previous section (e.g., item 4).

[0204] 5. A method of video processing, comprising performing a conversion between a video comprising one or more video pictures and a bitstream representation of the video, wherein the bitstream representation comprises a coded video sequence with one or more video layers, wherein the bitstream representation conforms to a format rule that specifies an EOS NAL unit in a picture unit is the last NAL unit among all NAL units other than other EOS NAL units or EOB NAL units within the PU.

[0205] The following solutions show example embodiments of the techniques discussed in the previous section (e.g., item 5).

[0206] 6. A method of video processing, comprising performing a conversion between a video comprising one or more video pictures and a bitstream representation of the video, wherein the bitstream representation comprises a coded video sequence having one or more video layers, wherein the bitstream representation conforms to a format rule that specifies that an EOS RBSP specifies that a next subsequent PU in the bitstream that belongs to the same layer as the EOS NAL unit must be an IRAP or GDR PU in decoding order.

[0207] 7. The method of any of solutions 1-6, wherein performing the conversion comprises encoding the video into the coded representation.

[0208] 8. The method of any of solutions 1-6, wherein performing the conversion comprises parsing and decoding the coded representation to generate the video.

[0209] 9. A video decoding apparatus comprising a processor configured to implement a method recited in one or more of solutions 1-8.

[0210] 10. A video encoding apparatus comprising a processor configured to implement a method recited in one or more of solutions 1-8.

[0211] 11. A computer program product having computer code stored thereon, the code, when executed by a processor, causing the processor to implement a method recited in any of solutions 1-8.

[0212] 12. A method, apparatus or system described in this document.

[0213] Some preferred embodiments are described below.

[0214] In some embodiments (e.g., see item 1 in section 5), a method of video processing (e.g., Figure 9The method 700 depicted in includes performing (702) a conversion between a video and a bitstream of the video, wherein the bitstream conforms to a format rule, wherein the bitstream includes one or more layers that include one or more picture units (PUs), wherein the format rule specifies, in response to a first PU in a layer of the bitstream following, in a decoding order, a sequence end network abstraction layer (EOS NAL) unit in the layer, a variable of the first PU is set to a particular value, wherein the variable indicates whether the first PU is a coded layer video sequence start (CLVSS) PU. In some embodiments, the first PU is an instantaneous decoding refresh PU. In some embodiments, the first PU is a clean random access PU and another variable of the clean random access PU is set to indicate that the clean random access PU is treated as a CLVSS PU. In some embodiments, the first PU is a clean random access PU. In some embodiments, the first PU is a gradual decoding refresh PU. In some embodiments, the first PU is a first PU in the layer in the decoding order. In some embodiments, the variable corresponds to NoOutputBeforeRecoveryFlag.

[0215] In some embodiments (e.g., see item 3 in section 5), a method of video processing (e.g., VPS Figure 10 The method 800 depicted in includes performing (802) a conversion between a video and a bitstream of the video, wherein the bitstream includes one or more layers that include one or more picture units (PUs) according to a format rule; wherein the format rule specifies that a PU of a particular layer following, in a decoding order, a sequence end network abstraction layer (EOS NAL) unit of the particular layer is a particular type of PU. In some embodiments, the particular type of PU is one of an intra random access point (IRAP) type or a gradual decoding refresh (GDR) type. In some embodiments, wherein the particular type of PU is a coded layer video sequence start (CLVSS) PU.

[0216] In some embodiments (e.g., see item 5 in section 5), a method of video processing (e.g., VPS Figure 11The method 900 depicted in the figure includes performing 902 a conversion between a video and a bitstream of the video, wherein the bitstream includes one or more layers comprising one or more picture units (PUs) according to a format rule, wherein the format rule specifies that, when present, an end of sequence (EOS) raw byte sequence payload (RBSP) syntax structure specifies that a next subsequent PU in the bitstream that belongs to a same layer as an EOS network abstraction layer (NAL) unit in decoding order is of a particular PU type from an intra random access point (IRAP) PU type or a gradual decoding refresh (GDR) PU type. In some embodiments, the particular PU type is the IRAP PU type. In some embodiments, the particular PU type is the GDR PU type.

[0217] In some embodiments (e.g., see item 2 in section 5), a method of video processing (e.g., ​ The method 1000 depicted in the figure includes performing 1002 a conversion between a video and a bitstream of the video, wherein the bitstream includes one or more layers comprising one or more network abstraction layer (NAL) units according to a format rule, wherein the format rule specifies that a first layer identifier in a header of an end of sequence network abstraction layer (EOS NAL) unit needs to be equal to a second layer identifier of one of the one or more layers in the bitstream. In some embodiments, the format rule also allows more than one EOS NAL unit to be included in a picture unit (PU). In some embodiments, wherein the format rule specifies that the first layer identifier of the EOS NAL unit needs to be less than or equal to a third layer identifier of a video coding layer (VCL) NAL unit associated with the EOS NAL unit.

[0218] In some embodiments (e.g., see item 4 in section 5), a method of video processing (e.g., ​ The method 1100 depicted in the figure includes performing 1102 a conversion between a video and a bitstream of the video, wherein the bitstream includes one or more layers comprising one or more picture units (PUs) comprising one or more network abstraction layer (NAL) units according to a format rule, wherein the format rule specifies that, in response to a first NAL unit in a PU indicating an end of sequence, the first NAL unit is a last NAL unit among all NAL units within the PU other than another NAL unit indicating another end of sequence or another NAL unit indicating an end of bitstream if present. In some embodiments, the another NAL unit is an end of sequence (EOS) NAL unit. In some embodiments, the another NAL unit is an end of bitstream (EOB) NAL unit.

[0219] In the above disclosed embodiments, the PU can have a format comprising a picture header NAL unit, a coded picture comprising one or more video coding layer NAL units, and zero or more non-video coding layer NAL units.

[0220] In the above disclosed embodiments, performing the conversion comprises encoding the video into a bitstream.

[0221] In the above disclosed embodiments, performing the conversion comprises decoding the video from the bitstream.

[0222] In some embodiments, a video decoding apparatus comprising a processor can be configured to implement the methods described in any of the above disclosed embodiments.

[0223] In some embodiments, a video encoding apparatus comprising a processor can be configured to implement the above described methods.

[0224] In some embodiments, a computer program product can have computer code stored thereon, which, when executed by a processor, causes the processor to implement the above disclosed methods.

[0225] In some embodiments, a method of bitstream generation, comprising generating a bitstream according to the method of any one or more of the above claims; and storing the bitstream on a computer readable program medium.

[0226] In some embodiments, a non-transitory computer readable recording medium can store a bitstream of a video generated by a method performed by a video processing apparatus, wherein the method comprises generating a bitstream according to the methods disclosed herein.

[0227] The disclosed and other solutions, examples, embodiments, modules and functional operations described herein can be realized in digital electronic circuitry, or in a computer software, firmware, or hardware, including the structural equivalents of such software, firmware, or hardware, or in combinations of one or more of them. The disclosed and other embodiments can be realized as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer readable medium for execution by, or to control the operation of, data processing apparatus. The computer readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition of matter effecting a machine-readable propagated signal, or a combination of one or more of them. The term "data processing apparatus" encompasses all apparatus, devices, and machines for processing data, including by way of example a programmable processor, a computer, or multiple processors or computers. The apparatus can include, in addition to hardware, code that creates an execution environment for the computer program in question, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them. A propagated signal is an artificially generated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to suitable receiver apparatus.

[0228] A computer program (also known as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files that store one or more modules, sub programs, or portions of code). A computer program can be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and are interconnected by a communication network.

[0229] The processes and logic flows described in this document can be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by special purpose logic circuitry, e.g., an FPGA or an ASIC, and apparatuses can be implemented as special purpose logic circuitry (e.g., an FPGA or an ASIC).

[0230] Processors suitable for the execution of a computer program include, by way of example, both general and special purpose microprocessors, and any one or more processors of any kind of digital computer. Generally, a processor will receive instructions and data from a read-only memory or a random access memory or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g., magnetic, magneto-optical disks, or optical disks. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media and memory devices, including by way of example semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.

[0231] Although the patent document contains many details, these should not be construed as limiting the subject matter or the scope of any patent covering embodiments, but as a description of features specific to certain embodiments of the technology. Certain features described in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment can also be implemented separately or in any suitable subcombination. Moreover, although features can be described above as acting in certain combinations and even initially claimed as such, one or more features from a claimed combination can in some cases be excised from the combination and the claimed combination can be directed to a subcombination or variation of a subcombination.

[0232] Similarly, while operations are depicted in the drawings in a particular order, this should not be understood as requiring such an order, or the order in which such operations are illustrated, and the order of the operations can be changed, including the order in which the operations are performed.

[0233] Only a few implementations and examples are described and other implementations, enhancements and variations can be made based on what is described and illustrated in this patent document.

Claims

1. A method of video processing, comprising: performing a conversion between a video and a bitstream of the video according to a first format rule, wherein the bitstream includes one or more layers that include one or more picture units (PUs), wherein the first format rule specifies that, responsive to a first PU in a layer of the bitstream following a sequence end network abstraction layer (EOS NAL) unit in the layer in a decoding order, a variable of the first PU is set to a particular value, wherein the variable indicates whether the first PU is a coded layer video sequence start (CLVSS) PU, wherein the first PU is a clean random access PU and another variable of the clean random access PU is set to indicate that the clean random access PU is treated as the CLVSS PU.

2. The method of claim 1, wherein, the first PU is replaced by the clean random access PU as an instantaneous decoding refresh PU.

3. The method of claim 1, wherein, the first PU is replaced by the clean random access PU as a gradual decoding refresh PU.

4. The method of claim 1, wherein, the first PU is replaced by the clean random access PU as a first PU in the layer in the decoding order.

5. The method of claim 1, wherein, the variable corresponds to NoOutputBeforeRecoveryFlag.

6. The method of claim 1, wherein, performing the conversion according to a second format rule, wherein the second format rule specifies that a PU of a particular layer is of a particular type, the PU of the particular layer is following a sequence end network abstraction layer (EOS NAL) unit in the particular layer.

7. The method of claim 6, wherein, the particular type of PU is one of an intra random access point (IRAP) type or a gradual decoding refresh (GDR) type.

8. The method of claim 6, wherein, the particular type of PU is a coded layer video sequence start (CLVSS) PU.

9. The method of claim 1, wherein, performing the conversion according to a third format rule, wherein the third format rule specifies that, when present, a sequence end (EOS) raw byte sequence payload (RBSP) syntax structure specifies that a next subsequent PU in the bitstream following the EOS network abstraction layer (NAL) unit in a same layer as the EOS NAL unit in a decoding order is of a particular type of PU from an intra random access point (IRAP) PU type or a gradual decoding refresh (GDR) PU type.

10. The method of claim 9, wherein, the particular type of PU is the IRAP PU type.

11. The method of claim 9, wherein, the particular type of PU is the GDR PU type.

12. The method of claim 1, wherein, the PU includes a picture header NAL unit, a coded picture including one or more video coding layer NAL units, and zero or more non-video coding layer NAL units.

13. The method of claim 1, wherein, performing the conversion includes encoding the video into the bitstream.

14. The method of claim 1, wherein, performing the conversion includes decoding the video from the bitstream.

15. An apparatus for processing video data, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to: perform a conversion between a video and a bitstream of the video according to a first format rule, wherein the bitstream includes one or more layers that include one or more picture units (PUs), wherein the first format rule specifies that, in response to a first PU in a layer of the bitstream that follows a sequence end network abstraction layer (EOS NAL) unit in the layer in decoding order, a variable of the first PU is set to a particular value, wherein the variable indicates whether the first PU is a coded layer video sequence start (CLVSS) PU, wherein the first PU is a clean random access PU and another variable of the clean random access PU is set to indicate that the clean random access PU is treated as the CLVSS PU.

16. The apparatus of claim 15, wherein, performing the conversion according to a second format rule or a third format rule, wherein the second format rule specifies that a PU of a particular layer is of a particular type of PU, the PU of the particular layer being located after a sequence end network abstraction layer (EOS NAL) unit of the particular layer; wherein the third format rule specifies that, when present, a sequence end (EOS) raw byte sequence payload (RBSP) syntax structure specifies that a next subsequent PU in the bitstream in decoding order that belongs to a same layer as the EOS network abstraction layer (NAL) unit is of a particular type of PU from an intra random access point (IRAP) PU type or a gradual decoding refresh (GDR) PU type.

17. A non-transitory computer-readable storage medium storing instructions that cause a processor to: perform a conversion between a video and a bitstream of the video according to a first format rule, wherein, the bitstream comprising one or more layers, the one or more layers comprising one or more picture units (PUs), wherein the first format rule specifies that, in response to a first PU in a layer of the bitstream that follows a sequence end network abstraction layer (EOS NAL) unit in the layer in decoding order, a variable of the first PU is set to a particular value, wherein the variable indicates whether the first PU is a coded layer video sequence start (CLVSS) PU, wherein the first PU is a clean random access PU and another variable of the clean random access PU is set to indicate that the clean random access PU is treated as the CLVSS PU. 18.A non-transitory computer-readable recording medium storing a bitstream of a video generated by a method performed by a video processing apparatus, wherein, The method comprises: generating a bitstream of the video according to a first format rule, the bitstream comprising one or more layers, the one or more layers comprising one or more picture units (PUs), wherein the first format rule specifies that, in response to a first PU in a layer of the bitstream that follows a sequence end network abstraction layer (EOS NAL) unit in the layer in decoding order, a variable of the first PU is set to a particular value, wherein the variable indicates whether the first PU is a coded layer video sequence start (CLVSS) PU, wherein the first PU is a clean random access PU and another variable of the clean random access PU is set to indicate that the clean random access PU is treated as the CLVSS PU.

19. A method for storing a bitstream of a video, comprising: generating a bitstream of the video according to a first format rule; and storing the bitstream in a non-transitory computer-readable recording medium, wherein the bitstream includes one or more layers including one or more picture units (PUs), wherein the first format rule specifies that, in response to a first PU following a sequence end network abstraction layer (EOS NAL) unit in a layer of the bitstream in decoding order, a variable of the first PU is set to a particular value, wherein the variable indicates whether the first PU is a coded layer video sequence start (CLVSS) PU, wherein the first PU is a clean random access PU and another variable of the clean random access PU is set to indicate that the clean random access PU is treated as the CLVSS PU.