Reference Picture Order Constraints

By enforcing specific constraints on video layer organization and reference picture management, the video processing methods address bandwidth challenges and enhance decoding efficiency in video coding standards like VVC, supporting scalable video coding and seamless random access.

JP7723134B2Active Publication Date: 2025-08-13BYTEDANCE INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024045473
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-03-19
Filing Date
2024-03-21
Publication Date
2025-08-13
Estimated Expiration
2041-03-16

AI Technical Summary

Technical Problem

Existing video encoding and decoding technologies face challenges in managing bandwidth demands and supporting efficient random access, sub-layer switching, and scalability, particularly in the context of evolving video coding standards like VVC, which require improved bitstream syntax to enhance performance.

Method used

Implementing video processing methods that enforce specific constraints and rules on the organization of video layers, pictures, and reference picture lists to improve decoding efficiency and support random access points, including intra-random access points, gradual decoder refresh, and reference picture management, adhering to format rules that ensure proper output ordering and layer-specific constraints.

Benefits of technology

Enhances decoding efficiency and supports scalable video coding by optimizing bitstream syntax for single-layer decoders, reducing bandwidth requirements and enabling seamless random access and sub-layer switching, thereby improving video processing performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007723134000012
    Figure 0007723134000012
  • Figure 0007723134000013
    Figure 0007723134000013
  • Figure 0007723134000014
    Figure 0007723134000014
Patent Text Reader

Abstract

To provide methods and devices for video processing such as video encoding or decoding.SOLUTION: An example video processing method 910 includes a step 912 of performing a conversion between a video having one or more video layers comprising a current picture comprising a current slice and a bitstream of the video according to a rule. The rule specifies a condition under which a reference picture for the current slice is disallowed to have an active entry that refers to a picture that precedes, in decoding order or output order, an intra random access point picture associated with the current picture.SELECTED DRAWING: Figure 9A
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS Book This application claims priority to and benefit of U.S. Provisional Patent Application No. 62 / 992,046, filed March 19, 2020. main stretch This is a divisional application of Japanese Patent Application No. 2022-556157 based on International Patent Application No. PCT / US2021 / 022584 filed on March 16, 2021. The entire disclosure of the above application is three By lighting Here It is used as a reference.

[0002] This patent specification relates to image and video encoding and decoding. [Background technology]

[0003] Digital video accounts for the largest bandwidth usage on the Internet and other digital communications networks, and the bandwidth demands for digital video use are expected to continue to grow as the number of connected user devices capable of receiving and displaying video increases. Summary of the Invention

[0004] This specification discloses techniques that can be used by video encoders and decoders to process coded representations of video using bitstream syntax that provides improved performance. The disclosed methods may be used by devices that perform video processing, such as video encoding, video decoding, or video transcoding.

[0005] In one exemplary aspect, a video processing method is disclosed that includes performing a conversion between video having one or more video layers including one or more video pictures and a coded representation of the video, the coded representation being organized according to a rule that specifies that a first video picture and a second picture that is an intra-random access point picture of the second picture are constrained to belong to the same video layer.

[0006] In another exemplary aspect, another video processing method is disclosed. The method includes performing a conversion between video having one or more video layers including one or more video pictures and a coded representation of the video, where the coded representation complies with format rules that specify that a trailing picture in the coded representation following a first type of picture that is an intra random access point is also allowed to be associated with a second type of picture that includes a progressive decoding refresh picture.

[0007] In another exemplary aspect, another video processing method is disclosed that includes performing a conversion between video having one or more video layers that include one or more video pictures and a coded representation of the video, where the coded representation complies with format rules that specify constraints on the output order of pictures that precede an intra random access point in decoding order such that the output order is only applicable to pictures in the same video layer.

[0008] In another exemplary aspect, another video processing method is disclosed. The method includes performing a conversion between video having one or more video layers including one or more video pictures and a coded representation of the video, where the coded representation complies with format rules that specify the constraints that (1) a trailing picture must follow an associated Intra Random Access Point (IRAP) picture or Gradual Decoder Refresh (GDR) picture in output order, or (2) a picture having the same layer ID as that of a GDR picture must precede the GDR picture and all associated pictures of the GDR picture in output order.

[0009] In another exemplary aspect, another video processing method is disclosed. The method includes performing a conversion between video having one or more video layers including one or more video pictures and a coded representation of the video, wherein the conversion complies with a rule that an order constraint is applicable to a picture, an Intra Random Access Point (IRAP) picture, and a non-leading picture only if the picture, the IRAP picture, and the non-leading picture are in the same layer, the rule being one of (a) a first rule specifying field sequence values and decoding order, and (b) an order of leading and / or non-leading pictures of a layer.

[0010] In another example aspect, another video processing method is disclosed. The method includes performing a conversion between video having one or more video layers including one or more video pictures and a coded representation of the video, the conversion conforming to rules that specify an order of leading pictures, Random Access Decodable Leading (RADL) pictures, and Random Access Skipped Leading (RASL) pictures associated with Gradual Decoding Refresh (GDR) pictures.

[0011] In another exemplary aspect, another video processing method is disclosed. The method includes performing a conversion between video having one or more video layers including one or more video pictures and a coded representation of the video, wherein the conversion complies with rules specifying that a reference picture list constraint of a clean random access picture is limited to one layer.

[0012] In another exemplary aspect, another video processing method is disclosed that includes performing a conversion between video having one or more video layers including one or more video pictures and a coded representation of the video, the conversion conforming to rules that specify conditions under which a current picture is allowed to reference entries in a reference picture list generated by a decoding process to generate unavailable reference pictures.

[0013] In another exemplary aspect, another video processing method is disclosed. The method includes performing a conversion between video having one or more video layers including one or more video pictures and a coded representation of the video, wherein the conversion conforms to an ordering rule between a current picture and a reference picture list corresponding to the current picture.

[0014] In another exemplary aspect, another video processing method is disclosed. The method includes performing a conversion between a video having one or more video layers including one or more video pictures and a video bitstream in accordance with a format rule, wherein the format rule specifies that a first video picture and a second picture, the first video picture being an associated intra-random access point picture of the second picture, are constrained to belong to the same video layer.

[0015] In another exemplary aspect, another video processing method is disclosed. The method includes performing a conversion between video having one or more video layers including one or more video pictures and a video bitstream in accordance with format rules, the format rules specifying that a trailing picture in the bitstream is allowed to be associated with a progressive decoding refresh picture.

[0016] In another exemplary aspect, another video processing method is disclosed. The method includes performing a conversion between a video having one or more video layers including one or more video pictures and a video bitstream in accordance with format rules, where the format rules specify that constraints on an output order of pictures preceding an intra random access point in decoding order are applicable to pictures within the same video layer.

[0017] In another exemplary aspect, another video processing method is disclosed. The method includes performing a conversion between a video having one or more video layers including one or more video pictures and a video bitstream in accordance with a format rule, where the format rule specifies a constraint that (1) a trailing picture follows an associated intra random access point picture or gradual decoder refresh picture in output order, or (2) a picture having the same Network Abstraction Layer (NAL) unit header layer identifier as a gradual decoder refresh picture precedes the gradual decoder refresh picture and all associated pictures of the gradual decoder refresh picture in output order.

[0018] In another exemplary aspect, another video processing method is disclosed. The method includes performing a conversion between a video having one or more video layers including one or more video pictures and a video bitstream according to rules, the rules specifying that a constraint applies to a decoding order of pictures associated with an intra random access point and a non-leading picture only if the picture, the intra random access point picture, and the non-leading picture are in the same layer.

[0019] In another exemplary aspect, another video processing method is disclosed. The method includes performing a conversion between video having one or more video layers including one or more video pictures and a video bitstream according to rules, the rules specifying an order of leading pictures, random access decodable leading pictures, and random access-skipped leading pictures associated with progressive decoding refresh pictures.

[0020] In another exemplary aspect, another video processing method is disclosed. The method includes performing a conversion between a video having one or more video layers including one or more video pictures and a video bitstream in accordance with rules, the rules specifying layer-specific constraints on reference picture lists for slices of clean random access pictures.

[0021] In another exemplary aspect, another video processing method is disclosed. The method includes performing a conversion between video having one or more video layers including one or more video pictures and a video bitstream according to rules, where the rules specify a condition that no pictures generated by the decoding process for generating an unavailable reference picture referenced by an active entry in a reference picture list of a current slice of a current picture exist.

[0022] In another exemplary aspect, another video processing method is disclosed. The method includes performing a conversion between video having one or more video layers including one or more video pictures and a video bitstream according to rules, where the rules specify a condition that no pictures generated by the decoding process to generate an unavailable reference picture referenced by an entry in a reference picture list of a current slice of the current picture exist.

[0023] In another exemplary aspect, another video processing method is disclosed. The method includes performing a conversion between a video having one or more video layers including a current picture that includes a current slice and a video bitstream according to rules, where the rules specify a condition that a reference picture list of the current slice is not allowed to have an active entry that points to a picture that precedes, in decoding order or output order, an intra random access point picture associated with the current picture.

[0024] In another exemplary aspect, another video processing method is disclosed. The method includes performing a conversion between a video having one or more video layers including a current picture that includes a current slice and a video bitstream according to rules, where the rules specify a condition that a reference picture list of the current slice is not allowed to have an entry that points to a picture that precedes, in decoding order or output order, an intra random access point picture associated with the current picture.

[0025] In yet another exemplary aspect, a video encoder apparatus is disclosed, the video encoder comprising a processor configured to implement the above-described method.

[0026] In yet another exemplary aspect, a video decoder apparatus is disclosed, the video decoder comprising a processor configured to implement the above-described method.

[0027] In yet another exemplary aspect, a computer readable medium having code stored thereon is disclosed, the code performing one of the methods described herein in the form of processor executable code.

[0028] These and other features are described throughout this document. [Brief explanation of the drawings]

[0029] [Figure 1] FIG. 1 is a block diagram illustrating an example of a video processing system. [Figure 2] FIG. 2 is a block diagram of the video processing device. [Figure 3] FIG. 3 is a flow chart of an exemplary method of video processing. [Figure 4] FIG. 4 is a block diagram illustrating a video coding system according to some embodiments of this disclosure. [Figure 5] FIG. 5 is a block diagram illustrating an encoder according to some embodiments of the present invention. [Figure 6] FIG. 6 is a block diagram illustrating a decoder according to some embodiments of the present invention. [Figure 7A] FIG. 7A shows a flowchart of an exemplary method of video processing according to some implementations of the disclosed technology. [Figure 7B] FIG. 7B shows a flowchart of an exemplary method of video processing according to some implementations of the disclosed technology. [Figure 7C] FIG. 7C shows a flowchart of an exemplary method of video processing according to some implementations of the disclosed technology. [Figure 7D] FIG. 7D shows a flowchart of an exemplary method of video processing according to some implementations of the disclosed technology. [Figure 7E] FIG. 7E shows a flowchart of an exemplary method of video processing according to some implementations of the disclosed technology. [Figure 7F] FIG. 7F shows a flowchart of an exemplary method of video processing according to some implementations of the disclosed technology. [Figure 7G] FIG. 7G shows a flowchart of an exemplary method of video processing according to some implementations of the disclosed technology. [Figure 8A] FIG. 8A shows a flowchart of an exemplary method of video processing according to some implementations of the disclosed technology. [Figure 8B] FIG. 8B shows a flowchart of an exemplary method of video processing according to some implementations of the disclosed technology. [Figure 9A]FIG. 9A shows a flowchart of an exemplary method of video processing according to some implementations of the disclosed technology. [Figure 9B] FIG. 9B shows a flowchart of an exemplary method of video processing according to some implementations of the disclosed technology. DETAILED DESCRIPTION OF THE INVENTION

[0030] Section headings are used herein for ease of understanding and are not intended to limit the applicability of the techniques and embodiments described in each section to that section alone. Furthermore, the term H.266 is used in certain descriptions for ease of understanding only and is not intended to limit the scope of the disclosed techniques. Thus, the techniques described herein are applicable to other video codec protocols and designs.

[0031] 1. Summary of the invention This specification relates to video coding technologies. In particular, it relates to various aspects for supporting random access, sub-layer switching, and scalability, including the definition of different types of pictures, their relationships in decoding order, output order, and prediction relationships. The concepts may be applied, individually or in various combinations, to any video coding standard or non-standard video codec that supports multi-layer video coding, such as the currently developed Versatile Video Coding (VVC).

[0032] 2. Abbreviation APS Adaptation Parameter Set AU Access Unit AUD Access Unit Delimiter AVC Advanced Video Coding CLVS Coded Layer Video Sequence CPB Coded Picture Buffer CRA Clean Random Access CTU Coding Tree Unit CVS Coded Video Sequence DCI Decoding Capability Information DPB Decoded Picture Buffer EOB End Of Bitstream EOS End Of Sequence GDR Gradual Decoding Refresh HEVC High Efficiency Video Coding HRD Hypothetical Reference Decoder IDR Instantaneous Decoding Refresh JEM Joint Exploration Model MCTS Motion-Constrained Tile Sets NAL Network Abstraction Layer OLS Output Layer Set PH Picture Header PPS Picture Parameter Set PTL Profile,Tier and Level PU Picture Unit RADL Random Access Decodable Leading (Picture) RAP Random Access Point RASL Random Access Skipped Leading (Picture) RBSP Raw Byte Sequence Payload RPL Reference Picture List SEI Supplemental Enhancement Information SPS Sequence Parameter Set STSA Step-wise Temporal Sublayer Access SVC Scalable Video Coding VCL Video Coding Layer VPS Video Parameter Set VTM VVC Test Model VUI Video Usability Information VVC Versatile Video Coding

[0033] 3. Initial Negotiation Video coding standards have evolved primarily through the development of well-known ITU-T and ISO / IEC standards. ITU-T developed H.261 and H.263, while ISO / IEC developed MPEG-1 and MPEG-4 Visual. The two organizations jointly developed H.262 / MPEG-2 Video, H.264 / MPEG-4 AVC (Advanced Video Coding), and H.265 / HEVC. Since H.262, video coding standards have been based on hybrid video coding architectures that utilize temporal prediction and transform coding. To explore future video coding technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Team (JVET) in 2015. Since then, many new methods have been adopted by the JVET and incorporated into reference software called the Joint Exploration Model (JEM). JVET meets quarterly, and the new coding standard aims to achieve a 50% bitrate reduction compared to HEVC. At the April 2018 JVET meeting, the new video coding standard was officially named "VVC (Versatile Video Coding)" and the first version of the VVC Test Model (VTM) was released at that time. As efforts to contribute to the standardization of VVC continue, new coding techniques are adopted for the VVC standard at every JVET meeting. After each meeting, the VVC Working Draft and Test Model VTM are updated. The VVC project is currently aiming for Technical Finalization (FDIS) at the July 2020 meeting.

[0034] 3.1 SVC (Scalable Video Coding) in general and VVC Scalable Video Coding (SVC), sometimes referred to as scalability in video coding, refers to video coding in which a Base Layer (BL) (sometimes called a Reference Layer (RL)) and one or more scalable Enhancement Layers (EL) are used. In SVC, the Base Layer can carry video data for a basic level of quality. One or more Enhancement Layers can carry additional video data, for example, to support higher spatial, temporal, and / or Signal-to-Noise (SNR) levels. Enhancement layers may be defined relative to a previous, coded layer. For example, a bottom layer can function as a BL and a top layer as an EL. Intermediate layers can function as either an EL or an RL, or both. For example, an intermediate layer (e.g., a layer that is neither the lowest nor the top layer) may be an EL for a layer below it, e.g., a base layer or any intervening enhancement layer, and simultaneously serve as a RL for one or more enhancement layers above it. Similarly, in multiview or 3D extensions of the HEVC standard, there may be multiple views, and information from one view may be utilized to code (e.g., encode or decode) information from another view (e.g., motion estimation, motion vector prediction, and / or other redundancy).

[0035] In SVC, parameters used by an encoder or decoder are grouped into parameter sets based on the coding level at which they can be utilized (e.g., video level, sequence level, picture level, slice level, etc.). For example, parameters that can be utilized by a coded video sequence at different layers in a bitstream may be included in a Video Parameter Set (VPS), and parameters utilized by one or more pictures in a coded video sequence may be included in a Sequence Parameter Set (SPS). Similarly, parameters utilized by one or more slices in a picture may be included in a Picture Parameter Set (PPS), and other parameters specific to one slice may be included in a slice header. Similarly, an indication of which parameter set a particular layer is using at a given time may be provided at various coding levels.

[0036] Thanks to VVC's support for Reference Picture Resampling (RPR), the upsampling required for spatial scalability support can be performed simply using an RPR upsampling filter. This allows a bitstream containing multiple layers, such as two layers for SD and HD resolutions in VVC, to be designed without requiring additional signal processing-level coding tools. Nevertheless, scalability support requires high-level syntax changes (compared to a case without scalability support). Scalability support is specified in VVC Version 1. Unlike scalability support in any previous video coding standard, including extensions to AVC and HEVC, the scalability design of VVC has been optimized for single-layer decoders. The decoding capabilities of multi-layer bitstreams are specified as if the bitstream contained only one layer. For example, decoding capabilities such as DPB size are specified in a manner independent of the number of layers in the bitstream being decoded. Essentially, a decoder designed for a single-layer bitstream does not require significant modifications to be able to decode a multi-layer bitstream. Compared to the multi-layer extension designs of AVC and HEVC, aspects of HLS have been significantly simplified at the expense of some flexibility: for example, an IRAP AU is required to contain a picture in each of the layers present in the CVS.

[0037] 3.2. Random Access and its Support in HEVC and VVC Random access refers to accessing and decoding a bitstream starting from a picture that is not the first picture of the bitstream in decoding order. To support tuning and channel switching in broadcast / multicast and multi-party video conferencing, local playback and seeking in streaming, and stream adaptation in streaming, the bitstream needs to contain frequent random access points, which are typically intra-coded pictures but may also be inter-coded pictures (e.g., in the case of gradual decoding refresh).

[0038] HEVC includes signaling of Intra Random Access Point (IRAP) pictures in the header of an NAL unit via the NAL unit type. Three types of IRAP pictures are supported: Instantaneous Decoder Refresh (IDR), Clean Random Access (CRA), and Broken Link Access (BLA). IDR pictures constrain the inter-picture prediction structure to not reference any pictures before the current Group-Of-Picture (GOP) and are traditionally called closed GOP random access points. CRA pictures relax the restrictions by allowing a picture to reference pictures before the current GOP, all of which are discarded in the case of random access. CRA pictures are traditionally called open GOP random access points. BLA pictures are typically generated by splicing two bitstreams or parts of them at a CRA picture, for example, during a stream switch. To enable better system usage of IRAP pictures, a total of six different NAL units are defined to signal properties of IRAP pictures that can be used to better suit the types of stream access points, such as those defined in ISOBMFF (ISO Base Media File Format), used for random access support in dynamic adaptive streaming over HTTP (DASH).

[0039] VVC supports three types of IRAP pictures, two types of IDR pictures (one type with an associated RADL picture and the other type without an associated RADL picture), and one type of CRA picture, which are basically the same as HEVC. The BLA picture type in HEVC is not included in VVC mainly for two reasons. i) The basic functionality of a BLA picture can be achieved by adding an end-of-sequence NAL unit to a CRA picture, the presence of which indicates that the following picture starts a new CVS in the single-layer bitstream. ii) In the development of VVC, it was desirable to specify fewer NAL unit types than HEVC, as indicated by using 5 bits instead of 6 bits for the NAL unit type field in the NAL unit header.

[0040] Another important difference in random access support between VVC and HEVC is that VVC supports GDR in a more prescriptive way. In GDR, bitstream decoding can start with an inter-coded picture, and initially, the entire picture region cannot be correctly decoded, but after several pictures, the entire picture region can be correctly decoded. AVC and HEVC also support GDR using a recovery point SEI message for signaling GDR random access points and recovery points. In VVC, a new NAL unit type is specified to indicate GDR pictures, and recovery points are signaled in the picture header syntax structure. CVS and bitstreams can start with a GDR picture. This means that an entire bitstream can contain only inter-coded pictures, without a single intra-coded picture. The main advantage of specifying GDR support in this way is that it provides GDR-compatible operation. GDR allows the encoder to smooth the bitstream bitrate by distributing intra-coded slices or blocks across multiple pictures rather than intra-coding the entire picture, thereby enabling a significant reduction in end-to-end delay, which is more important today than ever before as ultra-low latency applications such as wireless display, online gaming, and drone-based applications become more common.

[0041] Another GDR-related feature in VVC is virtual boundary signaling. The boundary between a refreshed region (i.e., a correctly decoded region) and an unrefreshed region in a picture between a GDR picture and its recovery point may be signaled as a virtual boundary, and when signaled, in-loop filtering across the boundary is not applied, thus preventing decoding inconsistencies for some samples near the boundary. This can be useful if an application decides to display correctly decoded regions during GDR processing.

[0042] IRAP pictures and GDR pictures can be collectively called RAP (Random Access Point) pictures.

[0043] 3.3 Reference Picture Management and RPL (Reference Picture List) Reference picture management is a core functionality required for any video coding scheme that uses inter prediction: it manages the storage and removal of reference pictures into and from the Decoded Picture Buffer (DPB), and places reference pictures in the proper order within the RPL.

[0044] HEVC's reference picture management, which includes reference picture marking and removal from the Decoded Picture Buffer (DPB) and Reference Picture List Construction (RPLC), differs from that of AVC. Instead of AVC's reference picture marking mechanism based on a sliding window plus adaptive Memory Management Control Operation (MMCO), HEVC specifies a reference picture management and marking mechanism based on the so-called Reference Picture Set (RPS). As a result, RPLC is based on the RPS mechanism. An RPS consists of a set of reference pictures associated with a picture, consisting of all reference pictures preceding the associated picture in decoding order, and may be used for inter-prediction of the associated picture or any picture following the associated picture in decoding order. A reference picture set consists of five lists of reference pictures. The first three lists contain all reference pictures that may be used in inter-prediction of the current picture and of one or more pictures following the current picture in decoding order. The other two lists consist of all reference pictures that are not used in the inter prediction of the current picture but may be used in the inter prediction of one or more pictures that follow the current picture in decoding order. RPS provides "intra-coded" signaling of the DPB status instead of "inter-coded" signaling as in AVC, mainly to improve error resilience. RPLC processing in HEVC is based on RPS by signaling an index to an RPS subset of each reference index, which is simpler than the RPLC processing in AVC.

[0045] Reference picture management in VVC is more similar to HEVC than AVC, but somewhat simpler and more robust. As in these standards, two RPLs, list0 and list1, are derived, but these are signaled more directly rather than based on the reference picture set concept used in HEVC or the automatic sliding window processing used in AVC. Reference pictures are listed as either active or inactive entries for the RPL, and only active entries may be used as reference indexes in inter-prediction of CTUs of the current picture. Inactive entries indicate other pictures that should be kept in the DPB to reference other pictures arriving later in the bitstream.

[0046] 3.4 Parameter Set AVC, HEVC, and VVC specify parameter sets. Parameter set types include SPS, PPS, APS, and VPS. SPS and PPS are supported by all of AVC, HEVC, and VVC. VPS was introduced in HEVC and is included in both HEVC and VVC. APS was not included in AVC or HEVC, but is included in the text of recent VVC drafts.

[0047] The SPS is designed to carry sequence-level header information, while the PPS is designed to carry picture-level header information that does not change frequently. Using the SPS and PPS avoids redundant signaling of frequently changing information because it does not need to be repeated for each sequence or picture. Furthermore, using the SPS and PPS enables out-of-band transmission of important header information, thereby not only avoiding the need for redundant transmission but also improving error resilience.

[0048] VPS was introduced to carry sequence-level header information that is common to all layers of a multi-layer bitstream.

[0049] APS was introduced to carry such picture-level or slice-level information, which requires a significant number of bits to code, is shared by multiple pictures, and can exist in a large number of different variations in the sequence.

[0050] 3.5 Related Definitions in VVC The relevant definition in the recent VVC text JVET-Q2001-vE / v15) is as follows: Associated IRAP picture (of a particular picture): The previous IRAP picture in decoding order (if any) has the same value nuh_layer_id as the particular picture. CRA (Clean Random Access) PU: A PU in which the coded picture is a CRA picture. CRA (Clean Random Access) picture: An IRAP picture in which the nal_unit_type of each VCL NAL unit is CRA_NUT. CVS (Coded Video Sequence): A sequence of AUs that consists of a CVSS AU followed by zero or more AUs that are not CVSS AUs, including all subsequent AUs (but not including any subsequent AUs that are CVSS AUs), in decoding order. CVSS (Coded Video Sequence Start) AU: An AU in which there is a PU in each layer of the CVS, and the coded picture of each PU is a CLVSS picture. GDR (Gradual Decoding Refresh) AU: An AU in which each coded picture in this PU is a GDR picture. GDR (Gradual Decoding Refresh) PU: A PU in which the coded picture is a GDR picture. GDR (Gradual Decoding Refresh) picture: A picture whose NAL unit nal_unit_type is GDR_NUT. IDR (Instantaneous Decoding Refresh) PU: A PU in which the coded picture is an IDR picture. IDR (Instantaneous Decoding Refresh) picture: An IRAP picture where the nal_unit_type of each VCL NAL unit is IDR_W_RADL or IDR_N_LP. IRAP (Intra Random Access Point) AU: An AU in which a PU exists in each layer of a CVS, and the coded picture of each PU is an IRAP picture. IRAP (Intra Random Access Point) PU: A PU in which the coded picture is an IRAP picture. IRAP (Intra Random Access Picture): A coded picture in the range from IDR_W_RADL to CRA_NUT, where all VCL NAL units have the same value of nal_unit_type. Leading picture: A picture that is in the same layer as the associated IRAP picture and precedes the associated IRAP picture in output order. RADL (Random Access Decodable Leading) PU: A PU in which the coded picture is a RADL picture. RADL (Random Access Decodable Leading) picture: A picture in which the nal_unit_type of each VCL NAL unit is RADL_NUT. RASL (Random Access Skipped Leading) PU: A PU in which the coded picture is a RASL picture. RASL (Random Access Skipped Leading) picture: A picture in which the nal_unit_type of each VCL NAL unit is RASL_NUT. STSA (Step-wise Temporal Sublayer Access) PU: A PU in which the coded picture is an STSA picture. STSA (Step-wise Temporal Sublayer Access) picture: A picture in which the nal_unit_type of each VCL NAL unit is STSA_NUT. NOTE - An STSA picture shall not use a picture with the same TemporalId as the STSA picture for inter-prediction reference. A picture following the STSA picture in decoding order with the same TemporalId as the STSA picture shall not use a picture preceding the STSA picture in decoding order with the same TemporalId as the STSA picture for inter-prediction reference. An STSA picture enables up-switching from the sub-layer immediately below it to the sub-layer containing the STSA picture. The TemporalId of an STSA picture shall be greater than 0. Trailing picture: A non-IRAP picture that follows the associated IRAP picture in output order and is not an STSA picture. NOTE - A trailing picture associated with an IRAP picture also follows the IRAP picture in decoding order. Pictures that follow the associated IRAP picture in output order and precede the associated IRAP picture in decoding order are not allowed.

[0051] 3.6. NAL Unit Header Syntax and Semantics in VVC In the recent VVC text (JVET-Q2001-vE / v15), the NAL unit header syntax and semantics are as follows:

[0052] 7.3.1.2 NAL unit header syntax

[0053] [Table 1]

[0054] 7.4.2.2. NAL Unit Header Semantics forbidden_zero_bit shall be equal to 0. nuh_reserved_zero_bit shall be equal to 0. A value of 1 for nuh_reserved_zero_bit may be specified by ITU-T|ISO / IEC in the future. Decoders shall ignore (i.e., remove from the bitstream and discard) NAL units with nuh_reserved_zero_bit equal to 1. nuh_layer_id specifies the identifier of the layer to which a VCL NAL unit belongs or to which a non-VCL NAL unit applies. The value of nuh_layer_id shall be in the range from 0 to 55. Other values of nuh_layer_id are reserved for future use by ITU-T|ISO / IEC. The value of nuh_layer_id shall be the same for all VCL NAL units of one coded picture. The value of nuh_layer_id of a coded picture or PU is the value of nuh_layer_id of the VCL NAL units of the coded picture or PU. The values of nuh_layer_id for AUD, PH, EOS, and FD NAL units are constrained as follows: - If nal_unit_type is equal to AUD_NUT, nuh_layer_id shall be equal to vps_layer_id[0]. Alternatively, if nal_unit_type is equal to PH_NUT, EOS_NUT, or FD_NUT, nuh_layer_id shall be equal to the nuh_layer_id of the associated VCL NAL unit. NOTE 1 - The values of nuh_layer_id in DCI, VPS, and EOB NAL units are not constrained.

[0055] The value of nal_unit_type shall be the same for all pictures in a CVSS AU. nal_unit_type specifies the NAL unit type, i.e., the type of RBSP data structure contained in the NAL unit, as specified in Table 5. NAL units that are in the range UNSPEC_28..UNSPEC_31 and have a nal_unit_type with unspecified semantics shall have no effect on the decoding process specified in this specification. NOTE 2 - NAL unit types within the range UNSPEC_28...UNSPEC_31 may be used as determined by the application. This specification does not specify the decoding process for these values of nal_unit_type. Because different applications may use these NAL unit types for different purposes, special care must be taken in the design of encoders that generate NAL units with these nal_unit_type values, and in the design of decoders that interpret the content of NAL units with these nal_unit_type values. This specification does not define the management of these values. These nal_unit_type values may only be suitable for use in contexts where "collisions" of usage (i.e., different definitions of the meaning of the NAL unit content for the same nal_unit_type value) are not important or possible, or are defined or managed by a controlling application or transport specification, or by controlling the environment in which the bitstream is distributed.

[0056] For purposes other than determining the number of data in a DU of the bitstream (as specified in Annex C), a decoder SHALL ignore (remove from the bitstream and discard) the content of all NAL units that use reserved values of nal_unit_type. NOTE 3 - This requirement allows for future definition of extensions that conform to this specification.

[0057] Table 5 - NAL unit type codes and NAL unit type classes

[0058] [Table 2]

[0059] [Table 3]

[0060] NOTE 4 – A CRA (Clean Random Access) picture may have an associated RASL or RADL picture present in the bitstream. NOTE 5 - An IDR (Instantaneous Decoding Refresh) picture with nal_unit_type equal to IDR_N_LP does not have an associated leading picture present in the bitstream. An IDR picture with nal_unit_type equal to IDR_W_RADL does not have an associated RADL picture present in the bitstream, but may have an RADL picture associated with it.

[0061] The value of nal_unit_type shall be the same for all VCL NAL units of a subpicture: the subpicture is considered to have the same NAL unit type as the VCL NAL units of the subpicture. For any particular picture's VCL NAL units, the following applies: - If mixed_nalu_types_in_pic_flag is equal to 0, the value of nal_unit_type should be the same for all VCL NAL units of a picture, and the picture or PU is considered to have the same NAL unit type as the coded slice NAL unit of the picture or PU. Otherwise (mixed_nalu_types_in_pic_flag is equal to 1), the picture has at least two subpictures and the VCL NAL units of the picture SHOULD have exactly two different nal_unit_type values as follows: The VCL NAL units of at least one subpicture of the picture SHOULD all have a particular value of nal_unit_type equal to STSA_NUT, RADL_NUT, RASL_NUT, IDR_W_RADL, IDR_N_LP, or CRA_NUT, while the VCL NAL units of other subpictures in the picture SHOULD all have a different value of nal_unit_type equal to TRAIL_NUT, RADL_NUT, or RASL_NUT.

[0062] For single-layer bitstreams, the following constraints apply: Each picture is considered to be related to the previous IRAP picture in decoding order, except for the first picture in the bitstream. If the picture is the leading picture of an IRAP picture, it is a RADL or RASL picture. If a picture is the trailing picture of an IRAP picture, it shall not be a RADL or RASL picture. - RASL pictures associated with IDR pictures shall not be included in the bitstream. - RADL pictures associated with IDR pictures with nal_unit_type equal to IDR_N_LP shall not be included in the bitstream. NOTE 6 - As long as each parameter set is available (in the bitstream or by external means not specified in this specification) at the time it is referenced, random access can be performed at the position of the IRAP PU by discarding all PUs preceding the IRAP PU (and by correctly decoding the IRAP picture and all subsequent non-RASL pictures in decoding order). A picture that precedes an IRAP picture in decoding order shall precede the IRAP picture in output order and shall precede the RADL picture associated with the IRAP picture in output order. - The RASL picture associated with a CRA picture shall precede the RADL picture associated with the CRA picture in output order. The RASL picture associated with a CRA picture shall follow in output order the IRAP picture that precedes the CRA picture in decoding order. If field_seq_flag is equal to 0 and the current picture is equal to a leading picture associated with an IRAP picture, it shall precede all non-leading pictures associated with the same IRAP picture in decoding order. Otherwise, if picA and picB are the first and last leading pictures, respectively, associated with an IRAP picture in decoding order, there shall be at most one non-leading picture preceding picA in decoding order, and there shall be no non-leading pictures between picA and picB in decoding order.

[0063] nuh_temporal_id_plus1-1 specifies the temporal identifier of the NAL unit. The value of nuh_temporal_id_plus1 shall not be equal to 0. The variable TemporalId is derived as follows: TemporalId=nuh_temporal_id_plus1-1 (36) If nal_unit_type is in the range IDR_W_RADL to RSV_IRAP_12, TemporalId shall be equal to 0. If nal_unit_type is equal to STSA_NUT and vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] is equal to 1, then TemporalId shall not be equal to 0. The value of TemporalId shall be the same for all VCL NAL units of an AU. The value of TemporalId of a coded picture, PU, or AU is the value of TemporalId of the VCL NAL units of the coded picture, PU, or AU. The value of TemporalId of a sub-layer representation is the maximum value of TemporalId of all VCL NAL units in the sub-layer representation.

[0064] The values of TemporalId for non-VCL NAL units are constrained as follows: - If nal_unit_type is equal to DCI_NUT, VPS_NUT, VPS_NUT, or SPS_NUT, TemporalId shall be equal to 0 and the TemporalId of the AU containing the NAL unit shall be equal to 0. Otherwise, if nal_unit_type is equal to PH_NUT, TemporalId shall be the TemporalId of the PU containing the NAL unit. - Otherwise, if nal_unit_type is equal to EOS_NUT or EOB_NUT, TemporalId shall be equal to 0. Otherwise, if nal_unit_type is equal to AUD_NUT, FD_NUT, PREFIX_SEI_NUT, or SUFFIX_SEI_NUT, TemporalId shall be the TemporalId of the AU containing the NAL unit. Otherwise, if nal_unit_type is equal to PPS_NUT, PREFIX_APS_NUT, or SUFFIX_APS_NUT, TemporalId shall be greater than or equal to the TemporalId of the PU containing the NAL unit. NOTE 7 - If the NAL unit is a non-VCL NAL unit, the value of TemporalId is equal to the minimum of the TemporalId values of all AUs to which the non-VCL NAL unit applies. If nal_unit_type is equal to PPS_NUT, PREFIX_APS_NUT, or SUFFIX_APS_NUT, TemporalId may be greater than or equal to the TemporalId of the containing AU, all PPSs and APSs may be included at the start of the bitstream (e.g., if they are transported out-of-band, the receiver places them at the beginning of the bitstream), and the first coded picture has TemporalId equal to 0.

[0065] 3.7. Syntax and Semantics of Picture Header Structure in VVC In the latest VVC text (JVET-Q2001-vE / v15), the syntax and semantics of the picture header structure most relevant to the present invention are as follows:

[0066] 7.3.2.7 Picture Header Structure Syntax

[0067] [Table 4]

[0068] 7.4.3.7 Picture Header Structure Semantics A PH syntax structure contains information that is common to all slices of the coded picture associated with the PH syntax structure. gdr_or_irap_pic_flag equal to 1 specifies that the current picture is a GDR or IRAP picture. gdr_or_irap_pic_flag equal to 0 specifies that the current picture may or may not be a GDR or IRAP picture. gdr_pic_flag equal to 1 specifies that the picture associated with the PH is a GDR picture. gdr_pic_flag equal to 0 specifies that the picture associated with the PH is not a GDR picture. If not present, the value of gdr_pic_flag is inferred to be equal to 0. If gdr_enabled_flag is equal to 0, the value of gdr_pic_flag shall be equal to 0. NOTE 1 - If gdr_or_irap_pic_flag is equal to 1 and gdr_pic_flag is equal to 0, the picture associated with the PH is an IRAP picture. ...

[0069] ph_pic_order_cnt_lsb specifies the picture order count modulo MaxPicOrderCntLsb for the current picture. The length of the ph_pic_order_cnt_lsb syntax element is log2_max_pic_order_cnt_lsb_minus4+4 bits. The value of ph_pic_order_cnt_lsb shall be in the range of 0 to MaxPicOrderCntLsb-1. The no_output_of_prior_pics_flag affects the output of the previously decoded picture in the DPB after decoding of a CLVSS picture that is not the first picture in the bitstream, as specified in Annex C. recovery_poc_cnt specifies the recovery point of a decoded picture in output order. If the current picture is a GDR picture associated with a PH, and there is a picture picA following the current GDR picture in decoding order in a CLVS with a PicOrderCntVal equal to the PicOrderCntVal of the current GDR picture plus the value of recovery_poc_cnt, then picture picA is called the recovery point picture. Otherwise, the first picture in output order with a PicOrderCntVal greater than the PicOrderCntVal of the current picture plus the value of recovery_poc_cnt is called the recovery point picture. The recovery point picture shall not precede the current GDR picture in decoding order. The value of recovery_poc_cnt shall be in the range from 0 to MaxPicOrderCntLsb-1.

[0070] If the current picture is a GDR picture, the variable RpPicOrderCntVal is derived as follows: RpPicOrderCntVal=PicOrderCntVal+recovery_poc_cnt (81) NOTE 2 - If gdr_enabled_flag is equal to 1 and the PicOrderCntVal of the current picture is greater than or equal to the RpPicOrderCntVal of the associated GDR picture, then the current and subsequent decoded pictures in output order exactly match the corresponding pictures produced by starting the decoding process from the previous IRAP picture (if any) that precedes the associated GDR picture in decoding order. ...

[0071] 3.8. RPL Constraints in VVC In a recent VVC text (JVET-Q2001-vE / v15), the constraints on RPL in VVC are as follows (as part of the decoding process in VVC clause 8.3.2 Reference Picture List Construction):

[0072] 8.3.2 Decoding process for reference picture list construction ... For each i equal to 0 or 1, the first NumRefIdxActive[i] entry in RefPicList[i] is referred to as the active entry in RefPicList[i], and the other entries in RefPicList[i] are referred to as inactive entries in RefPicList[i]. NOTE 2 - A particular picture may be referenced by both an entry in RefPicList[0] and an entry in RefPicList[1], and a particular picture may be referenced by multiple entries in RefPicList[0] or by multiple entries in RefPicList[1]. NOTE 3 - The active entries of RefPicList[0] and the active entries of RefPicList[1] collectively refer to all reference pictures that may be used for inter-prediction of the current picture and one or more pictures that follow the current picture in decoding order. The inactive entries of RefPicList[0] and the inactive entries of RefPicList[1] collectively refer to all reference pictures that are not used for inter-prediction of the current picture but that may be used in inter-prediction for one or more pictures that follow the current picture in decoding order. NOTE 4 - RefPicList[0] or RefPicList[1] may have one or more entries equal to "No Reference Picture" because the corresponding picture is not present in the DPB. Each inactive entry in RefPicList[0] or RefPicList[0] equal to "No Reference Picture" should be ignored. For each active entry in RefPicList[0] or RefPicList[1] equal to "No Reference Picture", an unintended picture loss should be inferred.

[0073] The requirements for bitstream conformance are that the following constraints apply: - For each i equal to 0 or 1, num_ref_entries[i][RplsIdx[i]] must not be less than NumRefIdxActive[i]. The picture referenced by each active entry in RefPicList[0] or RefPicList[1] shall be included in the DPB and shall be less than or equal to the TemporalId of the current picture. The picture referenced by each entry in RefPicList[0] or RefPicList[1] shall not be the current picture and shall have non_reference_picture_flag equal to 0. An STRP entry in RefPicList[0] or RefPicList[1] of a slice of a picture and an LTRP entry in RefPicList[0] or RefPicList[1] of the same slice or a different slice of the same picture shall not refer to the same picture. - The difference between the PicOrderCntVal of the current picture and the PicOrderCntVal of the picture referenced by the entry is 2 24 Assume that there are no LTRP entries in RefPicList[0] or RefPicList[1]. setOfRefPics is the unique set of pictures referenced by all entries in RefPicList[0] with the same nuh_layer_id as the current picture and all entries in RefPicList[1] with the same nuh_layer_id as the current picture. The number of pictures in setOfRefPics must be less than or equal to MaxDpbSize-1, where MaxDpbSize is as specified in section A.4.2, and setOfRefPics is the same for all slices of a picture.

[0074] - If the current slice has nal_unit_type equal to STSA_NUT, there shall be no active entries in RefPicList[0] or RefPicList[1] with TemporalId equal to that of the current picture and nuh_layer_id equal to that of the current picture. If the current picture is a picture that follows, in decoding order, an STSA picture with TemporalId equal to that of the current picture and nuh_layer_id equal to that of the current picture, then no picture that precedes the STSA picture, in decoding order, with TemporalId equal to that of the current picture and nuh_layer_id equal to that of the current picture shall be included as an active entry in RefPicList[0] or RefPicList[1]. If the current picture is a CRA picture, the IRAP picture (if any) referenced by an entry in RefPicList[0] or RefPicList[1] that precedes it in decoding order shall not have a preceding picture in output order or decoding order.

[0075] - If the current picture is a trailing picture, then no pictures referenced by active entries in RefPicList[0] or RefPicList[1] that were generated by the decoding process to generate reference pictures that are unavailable for the IRAP picture associated with the current picture shall exist. - If the current picture is a trailing picture that follows, in both decoding order and output order, one or more leading pictures associated with the same IRAP picture, then no pictures referenced by entries in RefPicList[0] or RefPicList[1] that were generated by the decoding process to generate reference pictures shall be unavailable for the IRAP picture associated with the current picture. If the current picture is a recovery point picture or a picture that follows the recovery point picture in output order, there shall be no entries in RefPicList[0] or RefPicList[1] that contain pictures generated by the decoding process to generate unavailable reference pictures for the GDR picture of the recovery point picture.

[0076] If the current picture is a trailing picture, then no pictures referenced by active entries in RefPicList[0] or RefPicList[1] shall precede the associated IRAP picture in output order or decoding order. - If the current picture is a trailing picture that follows, in both decoding order and output order, one or more leading pictures associated with the same IRAP picture, then no pictures referenced by entries in RefPicList[0] or RefPicList[1] shall precede the associated IRAP picture in output order or decoding order. - If the current picture is a RADL picture, there shall be no active entries in RefPicList[0] or RefPicList[1] such that: ○RASL Picture ○ Pictures generated by the decoding process to generate unavailable reference pictures ○ A picture that precedes the associated IRAP picture in decoding order

[0077] The picture referenced by each ILRP entry in RefPicList[0] or RefPicList[1] of a slice of the current picture shall be in the same AU as the current picture. The picture referenced by each ILRP entry in RefPicList[0] or RefPicList[1] of a slice of the current picture shall be present in the DPB and have a nuh_layer_id smaller than that of the current picture. Each ILRP entry in RefPicList[0] or RefPicList[1] of a slice shall be an active entry. ...

[0078] 4. Technical Problem Addressed by the Disclosed Technical Solution The existing design in the recent VVC text (JVET-Q2001-vE / v15) has the following problems: 1) The definition of the associated IRAP picture should be updated so that the associated IRAP picture of a particular picture belongs to the same layer as the particular picture. 2) The current definition of a trailing picture is as follows: Trailing picture: A non-IRAP picture that follows the associated IRAP picture in output order and is not an STSA picture. Therefore, a bitstream must have an IRAP picture for trailing pictures to exist, and if a bitstream does not have an IRAP picture, the NAL unit type value TRAIL_NUT cannot be used. However, non-STSA pictures associated with GDR pictures must use the NAL unit type value TRAIL_NUT. 3) The existing constraint on the output order of pictures before an IRAP picture in decoding order needs to be specified to apply only to pictures within a layer. 4) There are no restrictions on the relative decoding order and output order between a GDR picture and the pictures before and after it in decoding order. 5) The existing constraints on the decoding order of pictures associated with IRAP pictures and some non-leading pictures need to be specified to apply only to pictures within a layer.

[0079] 6) Currently, leading pictures, RADL pictures, and RASL pictures associated with GDR pictures are not supported. 7) The existing constraints on RPL for CRA pictures need to be specified to apply only to pictures within a layer. 8) For STSA pictures, trailing pictures associated with GDR pictures, and GDR pictures with NoOutputBeforeRecoveryFlag equal to 0, there is a missing constraint on the active entries in the RPL that are not generated by the decoding process to generate unavailable reference pictures. 9) In the case of STSA pictures, IDR pictures, CRA pictures with NoOutputBeforeRecoveryFlag equal to 0, etc., there is a lack of constraints on entries in the RPL that are not generated by the decoding process to generate unavailable reference pictures. 10) For an STSA picture, there is a lack of a constraint on the active entry in the RPL not to precede the associated IRAP picture in output order or decoding order. 11) For an STSA picture, there is a lack of a constraint on its entry in the RPL not to precede the associated IRAP picture in output order or decoding order.

[0080] 5. Examples of implementations and technical solutions In order to solve the above-mentioned problems, the following method is disclosed. The present invention should be regarded as an example for explaining a general concept, and should not be interpreted in a narrow sense. Furthermore, the present invention may be applied individually or in any combination. 1) To solve problem 1, the definition of the associated IRAP picture is updated so that the associated IRAP picture of a particular picture belongs to the same layer as the particular picture. 2) To solve problem 2, the definition of a trailing picture may be updated so that the trailing picture is associated with a GDR picture. a. Further, adding a definition of an associated GDR picture and updating the definition of an associated IRAP picture, specifying that each picture of a layer other than the first picture in the bitstream is associated with the previous IRAP or GDR picture of the same layer, whichever is closer in decoding order. b. Additionally, a constraint is added requiring trailing pictures to follow the output order of the associated IRAP or GDR pictures.

[0081] 3) To solve issue 3, the existing constraint on the output order of pictures preceding an IRAP picture in decoding order is updated to only impose restrictions on pictures within a layer. a. In one example, this constraint is specified as follows: Any picture that precedes in decoding order an IRAP picture with nuh_layer_id equal to nuh_layerId and has nuh_layer_Id equal to a particular value layerId precedes the IRAP picture and its associated RADL picture in output order. 4) To solve problem 4, add one or more of the following constraints: a. The trailing picture follows the associated IRAP or GDR picture in output order. b.layerId precedes in decoding order a GDR picture with nuh_layer_id equal to the specified value layerId, and any picture with nuh_layer_id equal to the specified value layerId precedes the GDR picture and its associated picture in output order.

[0082] 5) To solve issue 5, the existing constraints on the decoding order of pictures associated with IRAP pictures and some non-leading pictures are updated to only impose restrictions on pictures within one layer. a. In one example, this constraint is specified as follows: If the current picture with field_seq_flag equal to 0 and nuh_layer_id equal to a particular value layerId is a leading picture associated with an IRAP picture, then it shall precede, in decoding order, all non-leading pictures associated with the same IRAP picture. Otherwise, if picA and picB are the first and last leading pictures associated with an IRAP picture, respectively, in decoding order, there shall be at most one non-leading picture with nuh_layer_id equal to layerId that precedes picA in decoding order, and there shall be no non-leading pictures with nuh_layer_id equal to layerId between picA and picB in decoding order. b. In another example, this constraint is specified as follows: If field_seq_flag is equal to 0 and the current picture is equal to a leading picture associated with an IRAP picture, it shall precede all non-leading pictures associated with the same IRAP picture in decoding order. Otherwise, if picA and picB are the first and last leading pictures associated with an IRAP picture, respectively, in decoding order, there shall be at most one non-leading picture associated with the IRAP picture that precedes picA in decoding order, and there shall be no non-leading pictures associated with the IRAP picture between picA and picB in decoding order.

[0083] 6) To solve problem 6, a leading picture, a RADL picture, and a RASL picture associated with a GDR picture are defined and specified. a. The leading picture associated with a GDR picture is the picture that follows the GDR picture in decoding order and precedes it in output order. b. The RADL picture associated with the GDR picture is the leading picture associated with the GDR picture and has a nal_unit_type equal to RADL_NUT. c. The RASL picture associated with the GDR picture is the leading picture associated with the GDR picture and has nal_unit_type equal to RASL_NUT.

[0084] 7) To solve issue 7, the existing constraint on the RPL of CRA pictures is updated to impose the restriction only on pictures within a layer. In one example, this constraint is defined as follows: If the current picture has nuh_layer_id equal to a particular value layerId and is a CRA picture, then there shall be no pictures referenced by entries in RefPicList[0] or RefPicList[1] that precede, in output order or decoding order, the preceding IRAP picture (if any) that has nuh_layer_id equal to layerId in decoding order.

[0085] 8) To solve problem 8, the following constraints are specified. If the current picture with nuh_layer_id equal to a specific value layerId is not a RASL picture associated with a CRA picture with NoOutputBeforeRecoveryFlag equal to 1, a GDR picture with NoOutputBeforeRecoveryFlag equal to 1, or a reconstructed picture of a GDR picture with NoOutputBeforeRecoveryFlag equal to 1 and num_layer_id equal to layerId, it is specified that there are no pictures referenced by active entries in RefPicList[0] or RefPicList[1] generated in the decoding process to generate unavailable reference pictures.

[0086] 9) To solve problem 9, the following constraints are specified. It is specified that if the current picture with nuh_layer_id equal to a specific value layerId is not a CRA picture with NoOutputBeforeRecoveryFlag equal to 1, a picture preceding in decoding order a leading picture associated with the same CRA picture with NoOutputBeforeRecoveryFlag equal to 1, a leading picture associated with a CRA picture with NoOutputBeforeRecoveryFlag equal to 1, a GDR picture with NoOutputBeforeRecoveryFlag equal to 1, or a reconstructed picture of a GDR picture with NoOutputBeforeRecoveryFlag equal to 1 and nuh_layer_id equal to layerId, there are no pictures referenced by entries in RefPicList[0] or RefPicList[1] generated in the decoding process to generate unavailable reference pictures.

[0087] 10) To solve problem 10, the following constraints are specified. If the current picture is associated with an IRAP picture and follows the IRAP picture in output order, there shall be no pictures referenced by active entries in RefPicList[0] or RefPicList[1] that precede the associated IRAP picture in output or decoding order. 11) To solve problem 11, the following constraints are specified. If the current picture is associated with an IRAP picture, follows the IRAP picture in output order, and follows a leading picture associated with the same IRAP picture in both decoding order and output order, then no pictures referenced by entries in RefPicList[0] or RefPicList[1] shall precede the associated IRAP picture in output order or decoding order.

[0088] 6. Implementation form Below are some example embodiments for some of the inventive aspects summarized in Chapter 5 above, applicable to the VVC specification. The modified text is based on the latest VVC text in JVET-Q2001-vE / v15. The most relevant parts that have already been added or modified are highlighted in bold italics, and some of the parts that have been deleted are marked with double brackets (e.g., [[a]] indicates the deletion of the letter "a"). There are some other changes that are not highlighted, as they are editable in nature.

[0089] 6.1. First embodiment This embodiment is for items 1, 2, 3, 4, 5, and 5a.

[0090] 3 Definition ...

[0091]

number

[0092] 7.4.2.2. NAL Unit Header Semantics ...

[0093]

number

[0094]

number

[0095] 7.4.3.7 Picture Header Structure Semantics ...

[0096]

number

[0097] [[If the current picture is a GDR picture, the variable RpPicOrderCntVal is derived as follows: RpPicOrderCntVal=PicOrderCntVal+recovery_poc_cnt (81)]] NOTE 2 - If gdr_enabled_flag is equal to 1 and the PicOrderCntVal of the current picture is greater than or equal to the recoveryPointPocVal[[RpPicOrderCntVal]] of the associated GDR picture, then the current and subsequent decoded pictures in output order exactly match the corresponding pictures produced by starting the decoding process from the previous IRAP picture (if any) that precedes the associated GDR picture in decoding order. ...

[0098] 8.3.2 Decoding process for reference picture list construction ...

[0099]

number

[0100]

number

[0101]

number

[0102] - If the current picture is a RADL picture, there shall be no active entries in RefPicList[0] or RefPicList[1] such that: ○RASL Picture [[Pictures generated by the decoding process to generate unavailable reference pictures]] ○ The picture that precedes the associated IRAP picture in decoding order -...

[0103] 1 is a block diagram illustrating an example video processing system 1900 in which various techniques disclosed herein may be implemented. Various implementations may include some or all of the components of system 1900. System 1900 may include an input 1902 for receiving video content. The video content may be received in a raw or uncompressed format, e.g., 8- or 10-bit multi-component pixel values, or in a compressed or encoded format. Input 1902 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, PON (Passive Optical Network), and wireless interfaces such as Wi-Fi or cellular interfaces.

[0104] System 1900 may include a coding component 1904 capable of implementing various coding or encoding methods described herein. Coding component 1904 may reduce the average bit rate of video from input 1902 to the output of coding component 1904, generating a coded representation of the video. Thus, coding techniques are sometimes referred to as video compression or video transcoding techniques. The output of coding component 1904 may be stored or transmitted via a connected communication, as represented by component 1906. The bitstream (or coded) representation of the video received at input 1902, stored, or communicated, may be used by component 1908 to generate pixel values or displayable video that are transmitted to display interface 1910. The process of generating user-viewable video from the bitstream representation is sometimes referred to as video unpacking. Furthermore, while certain video processing operations are referred to as “coding” operations or tools, it will be understood that the coding tools or operations are performed by an encoder and corresponding decoding tools or operations that reverse the results of the coding are performed by a decoder.

[0105] Examples of peripheral bus interfaces or display interfaces may include USB (Universal Serial Bus), HDMI (High Definition Multimedia Interface), DisplayPort, etc. Examples of storage interfaces include SATA (Serial Advanced Technology Attachment), PCI, IDE interfaces, etc. The techniques described herein may be implemented in various electronic devices such as mobile phones, laptops, smartphones, or other devices capable of digital data processing and / or video display.

[0106] 2 is a block diagram of a video processing device 3600. The device 3600 may be used to implement one or more of the methods described herein. The device 3600 may be implemented in a smartphone, tablet, computer, Internet of Things (IoT) receiver, etc. The device 3600 may include one or more processors 3602, one or more memories 3604, and video processing hardware 3606. The one or more processors 3602 may be configured to implement one or more methods described herein. The one or more memories 3604 may be used to store data and code used to implement the methods and techniques described herein. The video processing hardware 3606 may be used to implement the techniques described herein in a hardware circuit.

[0107] FIG. 4 is a block diagram illustrating an example video coding system 100 that may utilize techniques of this disclosure.

[0108] 4, video coding system 100 may include source device 110 and destination device 120. Source device 110 generates encoded video data and may also be referred to as a video encoding device. Destination device 120 decodes the encoded video data generated by source device 110 and may also be referred to as a video decoding device.

[0109] The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.

[0110] The video source 112 may include a source such as a video capture device, an interface for receiving video data from a video content provider, and / or a computer graphics system for generating video data, or a combination of these sources. The video data may include one or more pictures. The video encoder 114 encodes the video data from the video source 112 and generates a bitstream. The bitstream may include a sequence of bits forming a coded representation of the video data. The bitstream may include coded pictures and associated data. A coded picture is a coded representation of a picture. The associated data may include sequence parameter sets, picture parameter sets, and other syntax structures. The I / O interface 116 may include a modulator-demodulator (modem) and / or transmitter. The coded video data may be transmitted directly to the destination device 120 via the I / O interface 116 over the network 130a. The coded video data may be stored on a recording medium / server 130b for access by the destination device 120.

[0111] Destination device 120 may include an I / O interface 126 , a video decoder 124 , and a display device 122 .

[0112] I / O interface 126 may include a receiver and / or a modem. I / O interface 126 may obtain encoded video data from source device 110 or storage medium / server 130b. Video decoder 124 may decode the encoded video data. Display device 122 may display the decoded video data to a user. Display device 122 may be integrated with destination device 120 or may be external to destination device 120 configured to interface with an external display device.

[0113] Video encoder 114 and video decoder 124 may operate according to video compression standards such as the High Efficiency Video Coding (HEVC) standard, the Versatile Video Coding (VVC) standard, and other current and / or future standards.

[0114] FIG. 5 is a block diagram illustrating an example of a video encoder 200, which may be video encoder 114 in system 100 shown in FIG.

[0115] Video encoder 200 may be configured to perform any or all of the techniques of this disclosure. In the example of FIG. 5, video encoder 200 includes multiple functional components. Techniques described in this disclosure may be shared among various components of video encoder 200. In some examples, a processor may be configured to perform any or all of the techniques described in this disclosure.

[0116] Functional components of the video encoder 200 may include a partitioning unit 201, a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205, and a prediction unit 202 which may include an intra prediction unit 206, a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213, and an entropy coding unit 214.

[0117] In other examples, video encoder 200 may include more, fewer, or different functional components. In one example, prediction unit 202 may include an Intra Block Copy (IBC) unit. The IBC unit can perform prediction in an IBC mode in which at least one reference picture is the picture in which the current video block is located.

[0118] Furthermore, some components, such as the motion estimator 204 and the motion compensator 205, may be highly integrated, but are represented separately in the example of FIG. 5 for illustrative purposes.

[0119] Divider 201 may divide a picture into one or more video blocks. Video encoder 200 and video decoder 300 may support various video block sizes.

[0120] The mode selection unit 203 may, for example, select one of intra or inter coding modes based on the error result, provide the resulting intra- or inter-coded block to the residual generation unit 207, generate residual block data, and provide it to the reconstruction unit 212 to reconstruct the coded block for use as a reference picture. In some examples, the mode selection unit 203 may select a Combination of Intra and Inter Prediction (CIIP) mode in which prediction is performed based on an inter prediction signal and an intra prediction signal. In addition, in the case of inter prediction, the mode selection unit 203 may select a resolution (e.g., sub-pixel or integer pixel accuracy) of the motion vector of the block.

[0121] To perform inter prediction on the current video block, motion estimation unit 204 may generate motion information for the current video block by comparing the current video block to one or more reference frames from buffer 213. Motion compensation unit 205 may determine a prediction video block for the current video block based on the motion information and decoded samples of pictures from buffer 213 other than the picture associated with the current video block.

[0122] Motion estimation unit 204 and motion compensation unit 205 may, for example, perform different operations on the current video block depending on whether the current video block is an I-slice, a P-slice, or a B-slice.

[0123] In some examples, motion estimator 204 may perform unidirectional prediction on the current video block, and motion estimator 204 may search reference pictures in list 0 or list 1 for a reference video block for the current video block. Motion estimator 204 may then generate a reference index indicating the reference picture in list 0 or list 1, including the reference video block and a motion vector indicating a spatial displacement between the current video block and the reference video block. Motion estimator 204 may output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 205 may generate a prediction video block for the current block based on the reference video block indicated by the motion information of the current video block.

[0124] In another example, motion estimator 204 may bidirectionally predict the current video block, and motion estimator 204 may search for a reference video block from among the reference pictures in list 0 to obtain the current video block, and may also search for another reference video block from among the reference pictures in list 1 to obtain the current video block. Motion estimator 204 may then generate reference indices indicating the reference pictures in lists 0 and 1 that contain the reference video blocks, and motion vectors indicating spatial displacements between the reference video blocks and the current video block. Motion estimator 204 may output the reference index and motion vector for the current video block as motion information for the current video block. Motion compensation unit 205 may generate a prediction video block for the current video block based on the reference video block indicated by the motion information of the current video block.

[0125] In some examples, the motion estimator 204 may output a full set of motion information for decoding processing in a decoder.

[0126] In some examples, the motion estimator 204 may not output a full set of motion information for the current picture. Rather, the motion estimator 204 may signal the motion information of the current video block by reference to the motion information of another video block. For example, the motion estimator 204 may determine that the motion information of the current video block is sufficiently similar to the motion information of a neighboring video block.

[0127] In one example, motion estimator 204 may indicate, in a syntax structure associated with the current video block, a value that indicates to video decoder 300 that the current video block has the same motion information as another video block.

[0128] In another example, motion estimation unit 204 may identify another video block and a motion vector difference (MVD) in a syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. Video decoder 300 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.

[0129] As mentioned above, video encoder 200 may predictively signal motion vectors. Two examples of predictive signaling techniques that may be implemented by video encoder 200 include Advanced Motion Vector Prediction (AMVP) and merge mode signaling.

[0130] Intra predictor 206 may perform intra prediction on the current video block. When intra predictor 206 intra predicts the current video block, intra predictor 206 may generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block may include a predicted video block and various syntax elements.

[0131] Residual generator 207 may generate residual data for the current video block by subtracting (e.g., as indicated by a minus sign) the predicted video block of the current video block from the current video block. The residual data for the current video block may include residual video blocks that correspond to different sample components of the samples in the current video block.

[0132] In other examples, for example in skip mode, there may be no residual data for the current video block, and residual generator 207 may not perform the subtraction operation.

[0133] Transform processor 208 may generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video block associated with the current video block.

[0134] After the transform processor 208 generates the transform coefficient video block associated with the current video block, the quantizer 209 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.

[0135] The inverse quantization unit 210 and the inverse transform unit 211 may apply inverse quantization and inverse transform, respectively, to the transform coefficient video block to reconstruct a residual video block from the transform coefficient video block. The reconstruction unit 212 may add the reconstructed residual video block to corresponding samples from one or more prediction video blocks generated by the prediction unit 202 to generate a reconstructed video block associated with the current block, which is stored in the buffer 213.

[0136] After reconstructor 212 reconstructs the video block, a loop filtering operation may be performed to reduce video blocking artifacts in the video block.

[0137] The entropy encoder 214 may receive data from other functional components of the video encoder 200. Once the entropy encoder 214 receives the data, the entropy encoder 214 may perform one or more entropy encoding operations to generate entropy-coded data and output a bitstream that includes the entropy-coded data.

[0138] FIG. 6 is a block diagram illustrating an example of a video decoder 300, which may be the video decoder 114 in the system 100 shown in FIG.

[0139] Video decoder 300 may be configured to perform any or all of the techniques of this disclosure. In the example of Figure 6, video decoder 300 includes multiple functional components. Techniques described in this disclosure may be shared among various components of video decoder 300. In some examples, a processor may be configured to perform any or all of the techniques described in this disclosure.

[0140] 6, video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, and a reconstruction unit 306, as well as a buffer 307. Video decoder 300 may, in some examples, perform a decoding path that is approximately the reverse of the encoding path described with respect to video encoder 200 (FIG. 5).

[0141] The entropy decoding unit 301 retrieves an encoded bitstream. The encoded bitstream may include entropy-coded video data (e.g., encoded blocks of video data). The entropy decoding unit 301 decodes the entropy-coded video data, and from the entropy-decoded video data, the motion compensation unit 302 may determine motion information including motion vectors, motion vector precision, reference picture list indexes, and other motion information. The motion compensation unit 302 may determine such information by, for example, implementing AMVP and merge mode.

[0142] The motion compensation unit 302 may generate motion-compensated blocks, possibly performing interpolation based on an interpolation filter, and an identifier for the interpolation filter to be used with sub-pixel accuracy may be included in the syntax element.

[0143] Motion compensation unit 302 may calculate interpolated values for sub-integer pixels of the reference block using an interpolation filter such as that used by video encoder 200 during encoding of the video block. Motion compensation unit 302 may determine the interpolation filter used by video encoder 200 based on received syntax information and use the interpolation filter to generate the prediction block.

[0144] The motion compensation unit 302 may use syntax information to determine the size of the blocks used to code the frames and / or slices of the coded video sequence, partitioning information that describes how each macroblock of a picture of the coded video sequence is divided, a mode that indicates how each division is coded, one or more reference frames (and reference frame lists) for each inter-coded block, and some other information to decode the coded video sequence.

[0145] The intra predictor 303 may form a prediction block from spatially adjacent blocks, for example, using an intra prediction mode received in the bitstream. The inverse quantizer 303 inverse quantizes the quantized video block coefficients provided in the bitstream and decoded by the entropy decoder 301. The inverse transformer 303 applies an inverse transform.

[0146] The reconstruction unit 306 may sum the residual block with the corresponding prediction block generated by the motion compensation unit 202 or the intra prediction unit 303 to form a decoded block. If desired, a deblocking filter may be applied to filter the decoded block to remove block artifacts. The decoded video block is stored in a buffer 307, which provides reference blocks for subsequent motion compensation / intra prediction and also generates decoded video for display on a display device.

[0147] Next, preferred examples in some embodiments will be listed.

[0148] The following items illustrate exemplary embodiments of the techniques described in the previous section: The following items illustrate exemplary embodiments of the techniques described in the previous section (e.g., item 1).

[0149] 1. A video processing method (e.g., method 3000 shown in FIG. 3 ) that includes performing (3002) a conversion between video having one or more video layers including one or more video pictures and a coded representation of the video, wherein the coded representation is configured according to a rule that specifies that a first video picture and a second picture, the first video picture being an intra random access point picture of the second picture, are constrained to belong to the same video layer.

[0150] The following items present exemplary embodiments of the techniques described in the previous section (eg, item 2).

[0151] 2. A video processing method comprising performing a conversion between video having one or more video layers including one or more video pictures and a coded representation of the video, wherein the coded representation complies with a format rule specifying that a trailing picture in the coded representation following a picture of a first type that is an intra random access point is also allowed to be associated with a picture of a second type that includes a progressive decoding refresh picture.

[0152] 3. The method of item 2, further specifying that the format rules are specified such that, for each layer, each picture of a layer other than the first picture of the layer in the bitstream is associated in decoding order with the previous intra random access point or gradual decoder refresh picture of the same layer, whichever is closer.

[0153] The following items present exemplary embodiments of the techniques described in the previous section (eg, item 3).

[0154] 4. A video processing method comprising performing a conversion between video having one or more video layers containing one or more video pictures and a coded representation of the video, wherein the coded representation complies with format rules specifying constraints on the output order of pictures preceding an intra random access point in decoding order such that the output order is only applicable to pictures in the same video layer.

[0155] 5. The method described in item 1, wherein the constraint specifies that any picture with nuh_layer_id equal to a particular value layerId that precedes in decoding order an intra random access point picture with nuh_layer_id equal to layerId must precede in output order the intra random access point picture and all associated random access decodable leading pictures.

[0156] The following items present exemplary embodiments of the techniques described in the previous section (eg, item 4).

[0157] 6. A video processing method comprising: performing a conversion between video having one or more video layers containing one or more video pictures and a coded representation of the video, wherein the coded representation complies with a format rule specifying the constraint that (1) a trailing picture must follow an associated Intra Random Access Point (IRAP) picture or Gradual Decoder Refresh (GDR) picture in output order, or (2) a picture having the same layer ID as a GDR picture must precede the GDR picture and all of the GDR picture's associated pictures in output order.

[0158] The following items present exemplary embodiments of the techniques described in the previous section (eg, item 5).

[0159] 7. Performing a transformation between a video having one or more video layers containing one or more video pictures and a coded representation of the video, wherein the transformation complies with a rule that an order constraint is applicable to a picture, an Intra Random Access Point (IRAP) picture, and a non-leading picture only if the picture, the IRAP picture, and the non-leading picture are in the same layer, the rule comprising: (a) a first rule specifying the values and decoding order of the field sequence, or (b) Layer leading and / or non-leading picture ordering.

[0160] The following items present exemplary embodiments of the techniques described in the previous section (eg, item 6).

[0161] 8. A video processing method comprising performing a conversion between video having one or more video layers containing one or more video pictures and a coded representation of the video, wherein the conversion complies with rules governing the ordering of leading pictures, Random Access Decodable Leading (RADL) pictures, and Random Access Skipped Leading (RASL) pictures associated with Gradual Decoding Refresh (GDR) pictures.

[0162] The following items present exemplary embodiments of the techniques described in the previous section (eg, item 7).

[0163] 9. A video processing method comprising performing a conversion between video having one or more video layers containing one or more video pictures and a coded representation of the video, wherein the conversion complies with rules specifying that reference picture list constraints for clean random access pictures are layer-limited.

[0164] 10. The method of item 9, wherein the constraint specifies that for a layer having a clean random access picture, a preceding intra random access point picture in decoding or output order is not referenced by an entry in the reference picture list.

[0165] The following items present exemplary embodiments of the techniques described in the previous section (eg, item 8).

[0166] 11. A video processing method comprising performing a transformation between video having one or more video layers containing one or more video pictures and a coded representation of the video, wherein the transformation complies with rules specifying the conditions under which a current picture is allowed to reference entries in a reference picture list generated by a decoding process to generate unavailable reference pictures.

[0167] 12. The method described in item 11, wherein the condition is that the current picture is a Random Access Skipped Leading (RASL) picture associated with a Clean Random Access (CRA) picture with NoOutputBeforeRecoveryFlag equal to 1, a Gradual Decoder Refresh (GDR) picture with NoOutputBeforeRecoveryFlag equal to 1, or a restored picture of a GDR picture with NoOutputBeforeRecoveryFlag equal to 1.

[0168] The following items present exemplary embodiments of the techniques described in the previous sections (eg, items 9, 10, and 11).

[0169] 13. A video processing method comprising performing a conversion between a video having one or more video layers containing one or more video pictures and a coded representation of the video, wherein the conversion complies with an ordering rule between a current picture and a reference picture list corresponding to the current picture.

[0170] 14. The method according to item 13, wherein the rule specifies that if the current picture having nuh_layer_id equal to a particular value layerId is not a CRA (Clean Random Access) picture having NoOutputBeforeRecoveryFlag equal to 1, a picture preceding in coding order a leading picture associated with the same CRA picture having NoOutputBeforeRecoveryFlag equal to 1, a leading picture associated with a CRA picture having NoOutputBeforeRecoveryFlag equal to 1, a GDR picture having NoOutputBeforeRecoveryFlag equal to 1, or a reconstructed picture of a GDR picture with NoOutputBeforeRecoveryFlag equal to 1 and nuh_layer_id equal to layerId, then no picture shall be referenced by an entry in RefPicList[0] or RefPicList[1] generated by the decoding process to generate an unavailable reference picture.

[0171] 15. The method of item 13, wherein the rule specifies that if the current picture is associated with an IRAP (Intra Random Access Point) picture and follows the IRAP picture in output order, there shall be no pictures referenced by active entries in RefPicList[0] or RefPicList[1] that precede the associated IRAP picture in output order or decoding order.

[0172] 16. The method of item 13, wherein the rule specifies that if the current picture is associated with an IRAP (Intra Random Access Point) picture, follows the IRAP picture in output order, and follows a leading picture associated with the same IRAP picture in both decoding order and output order, then no picture referenced by an entry in RefPicList[0] or RefPicList[1] precedes the associated IRAP picture in output order or decoding order.

[0173] 17. The method of any of items 1 to 16, wherein the conversion includes encoding the video into a coded representation.

[0174] 18. The method of any of items 1 to 16, wherein the transforming includes decoding the coded representation to generate pixel values for the image.

[0175] 19. A video decoding device comprising a processor configured to implement the methods described in one or more of items 1 to 18.

[0176] 20. A video encoding device comprising a processor configured to implement the methods described in one or more of items 1 to 18.

[0177] 21. A computer program product having computer code stored therein, the code being executed by a processor, causing the processor to implement the method according to any one of items 1 to 18.

[0178] 22. A method, apparatus or system described herein.

[0179] The second section presents illustrative examples of the techniques discussed in the previous section (eg, items 1-7).

[0180] 1. A method of video processing (e.g., method 710 shown in FIG. 7A), comprising performing 712 a conversion between a video having one or more video layers including one or more video pictures and a video bitstream in accordance with a format rule, wherein the format rule specifies that a first video picture and a second picture, which is an associated intra random access point picture of the second picture, are constrained to belong to the same video layer.

[0181] 2. The method of item 1, wherein there are no gradual decoding refresh pictures in the same video layer between the first video picture and the second picture in decoding order.

[0182] 3. The method described in item 1, wherein the first video picture and the second picture have the same identifier of the layer to which the video coding layer network abstraction layer unit belongs and the same identifier of the layer to which the non-video coding layer network abstraction layer unit is applied.

[0183] 4. A method according to any one of items 1 to 3, in which there are no gradual decoding refresh pictures with the same identifier between the first video picture and the second video picture in decoding order.

[0184] 5. The method of item 1, wherein the format rules further specify that the trailing picture follows the associated intra random access point picture or gradual decoding refresh picture in output order.

[0185] 6. A video processing method (e.g., method 720 shown in FIG. 7B) that includes performing 722 a conversion between video having one or more video layers including one or more video pictures and a video bitstream in accordance with format rules, the format rules specifying that trailing pictures in the bitstream are allowed to be associated with progressively decoded refresh pictures.

[0186] 7. The method of item 6, wherein the trailing picture is a picture in which each video coding layer network abstraction layer unit has a trail network abstraction layer unit type.

[0187] 8. The method according to item 6 or 7, wherein the trailing picture is allowed to be associated with an intra-random access point picture.

[0188] 9. A method according to any of items 6 to 8, wherein a trailing picture associated with an intra random access point picture or a progressive decoding refresh picture follows the intra random access point picture or the progressive decoding refresh picture in decoding order.

[0189] 10. A method according to any of items 6 to 8, in which a picture that follows the associated intra random access point picture in output order and precedes the associated intra random access point picture in decoding order is not permitted.

[0190] 11. The method of item 6, further specifying that the format rules are specified such that, for each layer, each picture of a layer other than the first picture of the layer in the bitstream is associated in decoding order with the nearest previous intra random access point or gradual decoder refresh picture of the same layer.

[0191] 12. The method of item 6 or 11, wherein the trailing picture is required to follow the associated intra random access point or gradual decoder refresh picture in output order.

[0192] 13. A video processing method (e.g., method 730 shown in FIG. 7C ), comprising performing 732 a conversion between video having one or more video layers including one or more video pictures and a video bitstream in accordance with format rules, wherein the format rules specify that constraints on the output order of pictures preceding an intra random access point in decoding order are applicable to pictures within the same video layer.

[0193] 14. The method described in item 13, wherein the constraint specifies that any picture having a NAL (Network Abstraction Layer) unit header layer identifier equal to a particular value and preceding an intra random access point picture having a NAL unit header layer identifier equal to the particular value in decoding order is required to precede the intra random access point picture and all associated random access decodable leading pictures in output order.

[0194] 15. The method according to item 14, wherein the NAL (Network Abstraction Layer) unit header layer identifier is nuh_layer_id.

[0195] 16. The method of item 14, wherein the NAL (Network Abstraction Layer) unit header layer identifier specifies an identifier of the layer to which the video coding layer Network Abstraction Layer unit belongs, or an identifier of the layer to which the non-video coding layer Network Abstraction Layer unit applies.

[0196] 17. A video processing method (e.g., method 740 shown in FIG. 7D) that includes performing 742 a conversion between video having one or more video layers including one or more video pictures and a video bitstream in accordance with a format rule, the format rule specifying the constraint that (1) a trailing picture follows an associated intra random access point picture or gradual decoder refresh picture in output order, or (2) a picture having the same NAL (Network Abstraction Layer) unit header layer identifier as a gradual decoder refresh picture precedes the gradual decoder refresh picture and all associated pictures of the gradual decoder refresh picture in output order.

[0197] 18. A video processing method (e.g., method 750 shown in FIG. 7E) that includes performing 752 a conversion between video having one or more video layers including one or more video pictures and a video bitstream in accordance with rules, the rules specifying that constraints on the decoding order of pictures associated with intra random access points and non-leading pictures apply only if the pictures, intra random access point pictures, and non-leading pictures are in the same layer.

[0198] 19. The method described in item 18, wherein the constraint specifies that if a picture having a field sequence flag value equal to 0 and a NAL (Network Abstraction Layer) unit header layer identifier picture equal to a specific value is a leading picture associated with an intra random access point picture, the picture precedes, in decoding order, all non-leading pictures associated with the intra random access point picture.

[0199] Item 20. The method of item 19, wherein a value of the field sequence flag equal to 20.0 indicates that the coded layer video sequence conveys pictures representing frames.

[0200] 21. The method of item 18, wherein the constraint specifies that if the value of the field sequence flag is equal to 0 and the picture is a leading picture associated with an intra random access point picture, the picture precedes, in decoding order, all non-leading pictures associated with the intra random access point picture.

[0201] 22. A video processing method (e.g., method 760 shown in FIG. 7F) that includes performing 762 a conversion between video having one or more video layers including one or more video pictures and a video bitstream according to rules, the rules specifying an order of leading pictures, random access decodable leading pictures, and random access skip leading pictures associated with progressive decoding refresh pictures.

[0202] 23. The method of item 22, wherein a leading picture associated with a gradual decoding refresh picture follows the gradual decoding refresh picture in decoding order and precedes the gradual decoding refresh picture in output order.

[0203] 24. The method described in item 22, wherein a random access decodable leading picture associated with a gradual decoding refresh picture is a leading picture associated with the gradual decoding refresh picture and has a NAL (Network Abstraction Layer) unit type corresponding to a coded slice of the random access decodable leading picture.

[0204] 25. The method described in item 22, wherein a random access decodable leading picture associated with a gradual decoding refresh picture is a leading picture associated with the gradual decoding refresh picture and has a NAL (Network Abstraction Layer) unit type corresponding to a coded slice of the random access decodable leading picture.

[0205] 26. A video processing method (e.g., method 770 shown in FIG. 7G) that includes performing 772 a conversion between video having one or more video layers including one or more video pictures and a video bitstream in accordance with rules, the rules specifying that constraints on reference picture lists for slices of clean random access pictures are limited to layers.

[0206] 27. The method of item 26, wherein the constraint specifies that for a layer having a clean random access picture, the preceding intra random access point picture in decoding or output order is not referenced by an entry in the reference picture list.

[0207] 28. The method of any one of items 1 to 27, wherein the conversion includes encoding the video into a bitstream.

[0208] 29. The method of any one of items 1 to 27, wherein the conversion includes decoding the video from a bitstream.

[0209] 30. A method according to any one of items 1 to 27, wherein the converting includes generating a bitstream from the video, and the method further includes storing the bitstream in a non-transitory computer-readable recording medium.

[0210] 31. A video processing device comprising a processor configured to implement the methods described in one or more of items 1 to 30.

[0211] 32. A method for storing a video bitstream, comprising the method of any one of items 1 to 30 and further comprising storing the bitstream on a non-transitory computer-readable recording medium.

[0212] 33. A computer-readable medium storing program code that, when executed, causes a processor to implement the method described in one or more of items 1 to 30.

[0213] 34. A computer-readable medium storing a bitstream generated according to any of the methods described above.

[0214] 35. A video processing device for storing a bitstream representation, configured to implement the methods described in one or more of items 1 to 30.

[0215] The third section presents exemplary implementations of the techniques described in the previous chapter (eg, items 8 and 9).

[0216] 1. A video processing method (e.g., method 810 shown in FIG. 8A) that includes performing 810 a conversion between a video having one or more video layers including one or more video pictures and a video bitstream in accordance with rules, the rules specifying a condition that a picture generated by a decoding process for generating an unavailable reference picture is not referenced by an active entry in a reference picture list of a current slice of a current picture.

[0217] 2. The method according to item 1, wherein the active entries correspond to entries available for use as reference indexes in inter prediction of the current picture.

[0218] 3. The method described in item 1, wherein the condition is that the current picture having a NAL (Network Abstraction Layer) unit header layer identifier equal to a specific value is not a random access skip leading picture associated with a clean random access picture having a variable indicating no pre-restored output equal to 1, a gradual decoder refresh picture having a variable equal to 1, or a restored picture of a gradual decoder refresh picture having a variable equal to 1 and a NAL unit header layer identifier equal to the specific value.

[0219] 4. A video processing method (e.g., method 820 shown in FIG. 8B) that includes performing 822 a conversion between a video having one or more video layers including one or more video pictures and a video bitstream in accordance with rules, the rules specifying a condition that a picture generated by a decoding process for generating an unavailable reference picture is not referenced by an entry in the reference picture list of a current slice of a current picture.

[0220] 5. The method described in item 4, wherein the condition is that the current picture having a NAL (Network Abstraction Layer) unit header layer identifier equal to a specific value is not a clean random access picture having a variable indicating no pre-reconstruction output equal to 1, a picture preceding a leading picture associated with a clean random access picture having a variable equal to 1 in decoding order, or a reconstructed picture of a gradual decoder refresh picture having a variable equal to 1 and a NAL unit header layer identifier equal to the specific value.

[0221] 6. The method of any one of items 1 to 5, wherein the conversion includes encoding the video into a bitstream.

[0222] 7. The method of any one of items 1 to 5, wherein the conversion includes decoding the video from a bitstream.

[0223] 8. The method of any of items 1 to 5, wherein the converting includes generating a bitstream from the video, and the method further includes storing the bitstream in a non-transitory computer-readable recording medium.

[0224] 9. A video processing device comprising a processor configured to implement the methods described in one or more of items 1 to 8.

[0225] 10. A method for storing a video bitstream, comprising the method of any one of items 1 to 8, and further comprising storing the bitstream in a non-transitory computer-readable recording medium.

[0226] 11. A computer-readable medium storing program code that, when executed, causes a processor to implement the methods described in one or more of items 1 to 8.

[0227] 12. A computer-readable medium storing a bitstream generated according to any of the methods described above.

[0228] 13. A video processing device for storing a bitstream representation, configured to implement the methods described in one or more of items 1 to 12.

[0229] The fourth set of sections presents exemplary embodiments of the techniques discussed in the previous section (eg, items 10 and 11).

[0230] 1. A video processing method (e.g., method 910 shown in FIG. 9A ), comprising: performing 912 a conversion between a video having one or more video layers including a current picture including a current slice and a video bitstream in accordance with rules, the rules specifying a condition that the reference picture list of the current slice is not permitted to have an active entry that references a picture that precedes, in decoding order or output order, an intra random access point picture associated with the current picture.

[0231] 2. The method according to item 1, wherein the active entries correspond to entries available for use as reference indexes in inter prediction of the current picture.

[0232] 3. The method according to item 1 or 2, wherein the rule specifies the conditions under which the reference picture list of the current picture is not allowed to have an active entry that references a picture that precedes, in decoding order, the intra random access point picture associated with the current picture.

[0233] 4. A method according to any of items 1 to 3, wherein the rule specifies conditions under which the reference picture list of the current picture is not allowed to have an active entry that references a picture that precedes, in output order, the intra random access point picture associated with the current picture.

[0234] 5. The method according to any one of items 1 to 4, wherein the condition is that the current picture is associated with an intra random access point picture and follows the intra random access point picture in decoding order and / or output order.

[0235] 6. A method according to any one of items 1 to 4, wherein the current picture follows in decoding order and / or output order an intra random access point picture having the same value as the identifier of the layer to which the video coding layer network abstraction layer unit belongs or the identifier of the layer to which the non-video coding layer network abstraction layer unit applies.

[0236] 7. A video processing method (method 920 shown in FIG. 9B) that includes performing 922 a conversion between a video having one or more video layers including a current picture including a current slice and a video bitstream in accordance with rules, the rules specifying a condition that the reference picture list of the current slice is not allowed to have an active entry pointing to a picture that precedes, in decoding order or output order, an intra random access point picture associated with the current picture.

[0237] 8. The method of item 7, wherein the rules specify conditions under which the reference picture list of the current picture is not allowed to have an entry that references a picture that precedes, in decoding order, the intra random access point picture associated with the current picture.

[0238] 9. The method of item 7 or 8, wherein the rule specifies conditions under which the reference picture list of the current picture is not allowed to have an entry that references a picture that precedes, in output order, the intra random access point picture associated with the current picture.

[0239] 10. A method according to any of items 7 to 9, wherein the intra random access point picture is associated with zero or more leading pictures, and the condition is that the current picture is associated with the intra random access point picture, follows the intra random access point picture in decoding order and / or output order, and follows zero or more leading pictures associated with the intra random access point picture in both decoding order and output order.

[0240] 11. A method according to any of items 7 to 9, wherein the intra random access point picture is associated with zero or more leading pictures, and the condition is that the current picture follows, in decoding order and / or output order, an intra random access point picture having the same value as the identifier of the layer to which the video coding layer network abstraction layer unit belongs or the identifier of the layer to which the non-video coding layer network abstraction layer unit is applied, and zero or more leading pictures.

[0241] 12. The method of any of items 1 to 11, wherein the conversion includes encoding the video into a bitstream.

[0242] 13. The method according to any one of items 1 to 11, wherein the conversion includes decoding the video from a bitstream.

[0243] 14. A method according to any one of items 1 to 11, wherein the converting includes generating a bitstream from the video, and the method further includes storing the bitstream in a non-transitory computer-readable recording medium.

[0244] 15. A video processing device comprising a processor configured to implement the methods described in one or more of items 1 to 14.

[0245] 16. A method for storing a video bitstream, comprising the method of any one of items 1 to 14 and further comprising storing the bitstream in a non-transitory computer-readable recording medium.

[0246] 17. A computer-readable medium storing program code that, when executed, causes a processor to implement the method described in one or more of items 1 to 14.

[0247] 18. A computer-readable medium storing a bitstream generated according to any of the methods described above.

[0248] 19. A video processing device for storing a bitstream representation, configured to implement the methods described in one or more of items 1 to 14.

[0249] As used herein, the term "video processing" may refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm may be applied during the conversion of a pixel representation of video to a corresponding bitstream representation, or vice versa. The bitstream representation of a current video block may correspond to bits spread to the same or different locations in the bitstream, e.g., as specified by the syntax. For example, a macroblock may be encoded in terms of transformed and coded error residual values and using bits in a header and other fields in the bitstream. Furthermore, during the conversion, a decoder may parse the bitstream with the knowledge that some fields may or may not be present based on a decision, as described in the solution above. Similarly, an encoder may determine whether a particular syntax field should or should not be included and generate a coded representation accordingly by including or excluding the syntax field in the coded representation.

[0250] Implementations of the disclosed and other solutions, examples, embodiments, modules, and functional operations described herein, including the structures disclosed herein and their structural equivalents, may be implemented in digital electronic circuitry, or in computer software, firmware, or hardware, or in one or more combinations thereof. The disclosed and other embodiments may be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer-readable medium for implementation by or controlling the operation of a data processing apparatus. The computer-readable medium may be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition of matter providing a machine-readable propagated signal, or one or more combinations thereof. The term "data processing apparatus" includes all apparatuses, devices, and machines for processing data, including, for example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, the apparatus may include code that creates an execution environment for the computer program, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, or one or more combinations thereof. A propagated signal is an artificially generated signal, for example, a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to an appropriate receiving device.

[0251] A computer program (also referred to as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, and can be deployed in any form, including as a stand-alone program or as modules, components, subroutines, or other units suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program may be recorded as part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), may be stored in a single file dedicated to the program, or may be stored in multiple coordinating files (e.g., files containing one or more modules, subprograms, or portions of code). A computer program can be deployed to run on one computer located at one site or on multiple computers distributed across multiple sites and interconnected by a communications network.

[0252] The processing and logic flows described herein may be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processing and logic flows may also be performed by, and devices may be implemented as, special purpose logic circuitry, such as a Field Programmable Gate Array (FPGA) or an Application Specific Integrated Circuit (ASIC).

[0253] Processors suitable for executing a computer program include, for example, both general-purpose and special-purpose microprocessors, and any one or more processors of any kind of digital computer. Typically, a processor receives instructions and data from a read-only memory or a random-access memory, or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will include one or more mass storage devices, e.g., magnetic, magneto-optical, or optical disks, for storing data, or be operatively coupled to receive data from or transfer data to these mass storage devices. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, EPROMs, EEPROMs, flash storage devices, magnetic disks, e.g., internal or removable disks, magneto-optical disks, and semiconductor storage devices such as CD-ROM and DVD-ROM disks. The processor and memory may be supplemented by, or incorporated in, special-purpose logic circuitry.

[0254] While this patent specification contains many details, these should not be construed as limiting the scope of any subject matter or the scope of the claims, but rather as descriptions of features that may be specific to particular embodiments of a particular technology. Certain features described in this patent document in the context of separate embodiments may also be implemented in combination in a single example. Conversely, various features described in the context of a single example may also be implemented in multiple embodiments separately or in any suitable subcombination. Furthermore, while features may be described above as acting in a particular combination and initially claimed as such, one or more features from a claimed combination may, in some cases, be extracted from the combination, and the claimed combination may be directed to subcombinations or variations of the subcombination.

[0255] Similarly, although operations are shown in a particular order in the figures, this should not be understood as requiring such operations to be performed in the particular order or sequential order shown, or that all of the operations shown be performed, to achieve desired results. Also, the separation of various system components in the examples described in this patent specification should not be understood as requiring such separation in all embodiments.

[0256] Only a few implementations and examples have been described; other embodiments, extensions and variations are possible based on the content described and illustrated in this patent document.

Claims

1. 1. A method of video processing, comprising: performing a conversion between a video having one or more video layers including a current picture including a current slice of the current picture according to a first rule and a bitstream of the video; The first rule is that the current picture having a Network Abstraction Layer (NAL) unit header layer identifier equal to a first particular value is associated with a first clean random access picture having a first variable equal to 1 indicating no pre-recovery output, a second clean random access picture having a second variable equal to 1 indicating no pre-recovery output, and a picture preceding in decoding order all leading pictures associated with the second clean random access picture having the second variable equal to 1 indicating no pre-recovery output, and a third clean random access picture having a third variable equal to 1 indicating no pre-recovery output. a first condition including that a picture generated by the decoding process for generating an unavailable reference picture is not referenced by an entry in a reference picture list of the current slice of the current picture if the picture is not a recovering picture of a leading picture attached to the current slice, a first gradual decoding refresh picture having a fourth variable equal to 1 indicating that there is no output before recovery, or a second gradual decoding refresh picture having a fifth variable equal to 1 indicating that there is no output before recovery and the NAL unit header layer identifier equal to the first particular value.

2. The conversion is performed according to a second rule, 2. The method of claim 1 , wherein the second rule specifies that if a picture having a NAL unit header layer identifier equal to a particular value layerId is a particular picture, then in output order or in the decoding order, no picture referenced by an entry in RefPicList[0] or RefPicList[1] occurs before a preceding Intra Random Access Point (IRAP) picture in the decoding order that has a NAL unit header layer identifier equal to layerId in the decoding order.

3. The method of claim 2 , wherein the particular picture is a clean random access picture.

4. The transformation is performed according to a third rule:

4. The method of claim 1, wherein the third rule specifies that a first video picture and a second picture that is an associated intra random access point of a second picture are constrained to belong to the same video layer.

5. The method of claim 4 , wherein there are no gradual decoding refresh pictures in the same video layer between the first video picture and the second video picture in the decoding order.

6. 5. The method of claim 4, wherein an identifier of a layer to which a video coding layer network abstraction layer unit belongs or an identifier of a layer to which a non-video coding layer network abstraction layer unit is applied, which is included in the first video picture, is the same as an identifier of a layer to which a video coding layer network abstraction layer unit belongs or an identifier of a layer to which a non-video coding layer network abstraction layer unit is applied, which is included in the second picture.

7. The transformation is performed according to the fourth rule, 7. The method of claim 1, wherein the fourth rule specifies that a trailing picture in the bitstream is allowed to be associated with a progressively decoded refresh picture or an intra-random access point picture.

8. The method of claim 7 , wherein the trailing picture is a picture in which each video coding layer network abstraction layer unit has a trail network abstraction layer unit type.

9. 8. The method of claim 7, wherein the trailing picture associated with an intra random access point picture or a gradual decoding refresh picture follows the intra random access point picture or the gradual decoding refresh picture in the decoding order or output order.

10. The method of claim 7 , wherein a picture that follows the associated intra random access point picture in output order and precedes the associated intra random access point picture in decoding order is not allowed.

11. 8. The method of claim 7, wherein the fourth rule further specifies that, for each layer, each picture of the layer, except for a first picture of the layer in the bitstream, is specified to be associated with the previous intra-layer access point picture of the same layer in the decoding order or a closer progressive decoding refresh picture.

12. The method of claim 1 , wherein the conversion comprises encoding the video into the bitstream.

13. The method of claim 1 , wherein the converting comprises decoding the video from the bitstream.

14. 1. An apparatus for processing video data, comprising a processor and a non-transitory memory having instructions, The instructions, when executed by the processor, cause the processor to: performing a conversion between a video having one or more video layers including a current picture including a current slice of the current picture and a bitstream of the video according to a first rule; The first rule is that the current picture having a Network Abstraction Layer (NAL) unit header layer identifier equal to a first particular value is associated with a first clean random access picture having a first variable equal to 1 indicating no pre-recovery output, a second clean random access picture having a second variable equal to 1 indicating no pre-recovery output, and a picture preceding in decoding order all leading pictures associated with the second clean random access picture having the second variable equal to 1 indicating no pre-recovery output, and a third clean random access picture having a third variable equal to 1 indicating no pre-recovery output. the picture generated by the decoding process for generating an unavailable reference picture is not referenced by an entry in a reference picture list of the current slice of the current picture if the picture is not a recovering picture of a leading picture attached to the current slice, a first gradual decoding refresh picture having a fourth variable equal to 1 indicating no pre-recovery output, or a second gradual decoding refresh picture having a fifth variable equal to 1 indicating no pre-recovery output and the NAL unit header layer identifier equal to the first particular value.

15. The processor performing a conversion between a video having one or more video layers including a current picture including a current slice of the current picture and a bitstream of the video according to a first rule; Execute The first rule specifies that the current picture having a Network Abstraction Layer (NAL) unit header layer identifier equal to a first particular value is associated with a first clean random access picture having a first variable equal to 1 indicating no pre-recovery output, a second clean random access picture having a second variable equal to 1 indicating no pre-recovery output, and a picture preceding, in decoding order, all leading pictures associated with the second clean random access picture having the second variable equal to 1 indicating no pre-recovery output, and a leading picture associated with a third clean random access picture having a third variable equal to 1 indicating no pre-recovery output. a fourth variable equal to 1 indicating that there is no output before recovery, or a recovering picture of a second gradual decoding refresh picture having a fifth variable equal to 1 indicating that there is no output before recovery, and the NAL unit header layer identifier equal to the first particular value; and a picture generated by the decoding process for generating an unavailable reference picture is not referenced by an entry in a reference picture list of the current slice of the current picture.

16. 1. A method for storing a video bitstream, comprising: generating the bitstream for the video having one or more video layers including a current picture that includes a current slice of the current picture according to a first rule; storing the bitstream on a non-transitory computer-readable recording medium; and The first rule is that a current picture having a Network Abstraction Layer (NAL) unit header layer identifier equal to a first particular value is associated with a first clean random access picture having a first variable equal to 1 indicating no pre-recovery output, a second clean random access picture having a second variable equal to 1 indicating no pre-recovery output, and a picture preceding in decoding order all leading pictures associated with the second clean random access picture having the second variable equal to 1 indicating no pre-recovery output, and a third clean random access picture having a third variable equal to 1 indicating no pre-recovery output. a first condition including that a picture generated by the decoding process for generating an unavailable reference picture is not referenced by an entry in a reference picture list of the current slice of the current picture if the picture generated by the decoding process for generating an unavailable reference picture is not a leading picture with a fourth variable equal to 1 indicating that there is no output before recovery, a first gradual decoding refresh picture with a fourth variable equal to 1 indicating that there is no output before recovery, or a recovering picture of a second gradual decoding refresh picture with a fifth variable equal to 1 indicating that there is no output before recovery, and the NAL unit header layer identifier equal to the first particular value.