Signaling of picture information in access units

By changing the semantics related to RASL images and HRDs in the VVC specification from image-specific to AU-specific, the interoperability problem in the VVC specification was solved, and the stability and decoding accuracy of the video codec were improved.

CN115699765BActive Publication Date: 2026-08-04DOUYIN CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
DOUYIN CO LTD
Filing Date
2021-05-21
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

The existing VVC specification has interoperability issues with the design of discardable images and access units, including unclear definitions of RASL images, incorrect semantics of HRD operations, and inaccurate derivation of image-specific variables, which may cause the decoder to crash or decode incorrectly.

Method used

The definition of RASL images, HRD-related semantics, and variable derivation were changed from image-specific to access unit (AU)-specific to ensure that flags and constraints conform to AU characteristics, including the accuracy of RASL image referencing relationships and HRD operations.

Benefits of technology

It solves interoperability issues, improves the stability and decoding accuracy of video codecs, and avoids decoder crashes and erroneous decoding behavior.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115699765B_ABST
    Figure CN115699765B_ABST
Patent Text Reader

Abstract

Methods, systems, and devices are disclosed for signaling picture information in access units in video bitstream processing. An example method of video processing includes performing a conversion between a video comprising one or more pictures and a bitstream of the video, wherein the bitstream conforms to a format rule, wherein the format rule specifies that a first syntax element and a second syntax element, when included in the bitstream, are access unit (AU) specific, wherein, responsive to a current AU not being a first AU in the bitstream in decoding order, the first syntax element indicates that a nominal coded picture buffer (CPB) removal time of the current AU is determined relative to (a) a nominal CPB removal time of a previous AU associated with a buffering period (BP) supplemental enhancement information (SEI) message or (b) a nominal CPB removal time of the current AU, and wherein, responsive to the current AU not being the first AU in the bitstream in decoding order, the second syntax element specifies a CPB removal delay delta value relative to the nominal CPB removal time of the previous AU.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications

[0002] Pursuant to the applicable patent law and / or rules of the Paris Convention, this application aims to promptly claim priority and benefit to U.S. Provisional Patent Application No. 63 / 029,321, filed May 22, 2020. For all purposes required by law, the entire disclosure of the aforementioned application is incorporated herein by reference as part of the disclosure of this application. Technical Field

[0003] This patent document relates to image and video encoding and decoding. Background Technology

[0004] Digital video consumes the largest share of bandwidth in the internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth demand for digital video is expected to continue to grow. Summary of the Invention

[0005] This article discloses techniques that can be used by video encoders and decoders to perform video encoding or decoding.

[0006] In an example, a video processing method is disclosed. The method includes performing a conversion between a video comprising one or more images and a video bitstream, wherein the bitstream conforms to a format rule specifying that a Picture Timing (PT) Supplemental Enhancement Information (SEI) message, when included in the bitstream, is Access Unit (AU) specific, and wherein each of the one or more images, which is a Random Access Skip Precedence (RASL) image, includes only an RASL Network Abstraction Layer Unit (NUT) type.

[0007] In another example, a different video processing method is disclosed. This method includes performing a conversion between a video and a video bitstream comprising one or more images, wherein the bitstream conforms to a format rule that allows the use of a Random Access Skip-Before (RASL) image as a reference sub-image to predict a juxtaposed RADL image in a RADL image associated with the same Completely Random Access (CRA) image as the RASL image.

[0008] In another example, another video processing method is disclosed. This method includes performing a conversion between a video and a video bitstream comprising one or more images, wherein the bitstream conforms to a format rule specifying the images associated with a first flag and the derivation in the decoding process of image sequence counting based on a second flag, wherein the image associated with the first flag is a previous image in decoding order, the previous image having (i) a first identifier identical to a stripe or image header referencing a reference image list syntax structure, (ii) a second flag and a second identifier equal to zero, and (iii) an image type different from Random Access Skip Prep (RASL) images and Random Access Decodeable Prep (RADL) images, wherein the first flag indicates the presence of a third flag in the bitstream, wherein the second flag indicates whether the current image is used as a reference image, and wherein the third flag is used to determine the value of one or more most significant bits of the image sequence count value of a long-term reference image.

[0009] In another example, a different video processing method is disclosed. This method includes performing a conversion between a video and a video bitstream comprising one or more images, wherein the bitstream conforms to a format rule specifying that variables used to determine the timing of removing a decoding unit (DU) or decoding the DU are access unit (AU) specific and derived based on a flag indicating whether the current image is permitted to be used as a reference image.

[0010] In another example, another video processing method is disclosed. The method includes performing a conversion between a video and a video bitstream comprising one or more images, wherein the bitstream conforms to a format rule specifying that a Buffer Period Supplement Enhancement Information (SEI) message and a Picture Timing SEI message are access unit (AU) specific when included in the bitstream. A first variable associated with the Buffer Period SEI message and a second variable associated with the Buffer Period SEI message and the Picture Timing SEI message are derived based on a flag indicating whether the current image is permitted to be used as a reference image. The first variable indicates that the access unit includes (i) an identifier equal to zero, and (ii) an image that is not a Random Access Skip Preceding (RASL) image or a Random Access Decodable Preceding (RADL) image and the flag is equal to zero. The second variable indicates that the current AU is not the first AU in decoding order, and that the previous AU in decoding order includes (i) an identifier equal to zero, and (ii) an image that is not a Random Access Skip Preceding (RASL) image or a Random Access Decodable Preceding (RADL) image and the flag is equal to zero.

[0011] In another example, another video processing method is disclosed. This method includes performing a conversion between a video and a video bitstream comprising one or more images, wherein the bitstream conforms to a format rule specifying that the derivation of a first variable and a second variable associated with a first image and a second image is based on a flag, wherein the first image is the current image and the second image is a previous image in decoding order, the previous image (i) including a first identifier equal to zero, (ii) including a flag equal to zero, and (iii) not a Random Access Skip Precedence (RASL) image or a Random Access Decodeable Precedence (RADL) image, and wherein the first variable and the second variable are respectively the maximum and minimum values ​​of the image order counts for each of the following images for which the second identifier is equal to the second identifier of the first image: (i) the first image, (ii) the second image, (iii) one or more short-term reference images referenced by all entries in the reference image list of the first image, and (iv) each image that has been output, the codec image buffer (CPB) removal time of which is less than the CPB removal time of the first image and the decode image buffer (DPB) output time is greater than or equal to the CP removal time of the first image.

[0012] In another example, another video processing method is disclosed. This method includes performing a conversion between a video and a video bitstream comprising one or more images, the bitstream conforming to a format rule specifying that flags and syntax elements included in the bitstream are access unit (AU) specific, wherein, in response to the current AU not being the first AU in the bitstream in decoding order, the flag indicates that the nominal codec picture buffer (CPB) removal time of the current AU is determined relative to (a) the nominal CPB removal time of a previous AU associated with a Buffer Period Supplemental Enhancement Information (SEI) message or (b) the nominal CPB removal time of the current AU, and wherein, in response to the current AU not being the first AU in the bitstream in decoding order, the syntax element specifies a CPB removal delay increment value relative to the nominal CPB removal time of the current AU.

[0013] In another example, another video processing method is disclosed. This method includes performing a conversion between a video and a video bitstream comprising one or more images, the bitstream conforming to a format rule specifying multiple variables and a Picture Timing Supplement Enhancement Information (SEI) message that is Access Unit (AU) specific when included in the bitstream. The Picture Timing SEI message includes multiple syntax elements, a first variable among multiple variables indicating whether the current AU is associated with a buffered time-period SEI message, and a second and third variable among multiple variables associated with an indication of whether the current AU is an AU for initializing a hypothetical reference decoder (HRD). The first syntax element among multiple syntax elements specifies the number of clock ticks to wait before outputting one or more decoded images of the AU from the decoded image buffer (DPB) after the AU is removed from the codec picture buffer (CPB). The second syntax element among multiple syntax elements specifies the number of sub-clock ticks to wait before outputting one or more decoded images of the AU from the DPB after the last decoded unit (DU) in the AU is removed from the CPB. The third syntax element among multiple syntax elements specifies the number of element image time-period intervals occupied by the one or more decoded images of the current AU for use in displaying the model.

[0014] In another example, a different video processing method is disclosed. This method includes performing a conversion between a video and a video bitstream comprising one or more images, wherein the bitstream conforms to a format rule specifying that syntax elements associated with a decoded picture buffer (DPB) are access unit (AU) specific when included in the bitstream, and wherein the syntax element specifies the number of sub-clock ticks to wait before outputting one or more decoded pictures of the AU from the DPB after the last decoded unit (DU) in the AU is removed from the codec-decoder picture buffer (CPB).

[0015] In another example, another video processing method is disclosed. This method includes performing a conversion between a video comprising one or more images and a video bitstream, wherein the bitstream conforms to a format rule specifying that a flag is Access Unit (AU) specific when included in the bitstream, wherein the value of the flag is based on whether the associated AU is an Intra-Frame Random Access Point (IRAP) AU or a Progressive Decoding Refresh (GDR) AU, and the value of the flag specifies (i) the presence of a syntax element in a Buffered Period Supplemental Enhancement Information (SEI) message and (ii) the presence of alternative timing information in the current buffered period's picture timing SEI message.

[0016] In another example, a different video processing method is disclosed. This method includes performing a conversion between a video and a video bitstream comprising one or more images, wherein the bitstream conforms to a format rule specifying that the value of a first syntax element is based on a flag indicating whether the temporal distance between the output times of consecutive images of a hypothetical reference decoder (HRD) is constrained, the variable identifying the highest temporal sublayer to be decoded, and the first syntax element specifying the number of element image time intervals occupied by one or more decoded images of the current AU for use in displaying the model.

[0017] In another example aspect, a video encoder apparatus is disclosed. This video encoder includes a processor configured as described above.

[0018] In another example, a video decoder apparatus is disclosed. This video decoder includes a processor configured as described above.

[0019] In another example, a computer-readable medium on which code is stored is disclosed. This code embodies one of the methods described herein in the form of processor-executable code.

[0020] These and other features will be described in this document. Attached Figure Description

[0021] Figure 1 This is a block diagram illustrating an example video processing system in which the various techniques disclosed herein can be implemented.

[0022] Figure 2 This is a block diagram of an example hardware platform used for video processing.

[0023] Figure 3 This is a block diagram illustrating an example video codec system that can implement some embodiments of the present disclosure.

[0024] Figure 4 This is a block diagram illustrating an example of an encoder that can implement some embodiments of the present disclosure.

[0025] Figure 5 This is a block diagram illustrating examples of decoders that can implement some embodiments of the present disclosure.

[0026] Figures 6 to 16 A flowchart of an example method for video processing is shown. Detailed Implementation

[0027] Chapter headings are used in this document for ease of understanding and not to limit the applicability of the techniques and embodiments disclosed in each chapter to that chapter. Furthermore, the use of H.266 terminology in some descriptions is merely for ease of understanding and not to limit the scope of the disclosed techniques. Therefore, the techniques described herein are also applicable to other video codec protocols and designs.

[0028] 1. Introduction

[0029] This article relates to video codec techniques. Specifically, it concerns the semantics of handling discardable images / AUs and HRD-related SEI messages in video codecs. Examples of discardable images that can be discarded in certain situations include RASL images, RADL images, and images where ph_non_ref_pic_flag equals 1. HRD-related SEI messages include BP, PT, and DUISEI messages. These concepts can be applied individually or in various combinations to any video codec standard or non-standard video codec that supports multi-layer video codecs, such as the Multi-Functional Video Codec (VVC) under development.

[0030] 2. Abbreviations

[0031] APS: Adaptive Parameter Set

[0032] AU: Access Unit

[0033] AUD: Access Unit Delimiter

[0034] AVC: Advanced Video Coding

[0035] CLVS: Coded Layer Video Sequence

[0036] CPB: Coded Picture Buffer

[0037] CRA: Clean Random Access

[0038] CTU: Coding Tree Unit

[0039] CVS: Coded Video Sequence

[0040] DCI: Decoding Capability Information

[0041] DPB: Decoded Picture Buffer

[0042] EOB: End of Bitstream

[0043] EOS: End of Sequence

[0044] GDR: Gradual Decoding Refresh

[0045] HEVC: High Efficiency Video Coding

[0046] HRD: Hypothetical Reference Decoder

[0047] IDR: Instantaneous Decoding Refresh

[0048] ILP: Inter-Layer Prediction

[0049] ILRP: Inter-Layer Reference Picture

[0050] JEM: Joint Exploration Model

[0051] LTRP: Long-Term Reference Picture

[0052] MCTS: Motion-Constrained Tile Sets

[0053] NAL: Network Abstraction Layer

[0054] OLS: Output Layer Set

[0055] PH: Picture Header

[0056] PPS: Picture Parameter Set

[0057] PTL: Profile, Tier, and Level

[0058] PU: Picture Unit

[0059] RAP: Random Access Point

[0060] RADL: Random Access Decodable Leading Picture

[0061] RASL: Random Access Skipped Leading Picture

[0062] RBSP: Raw Byte Sequence Payload

[0063] SEI: Supplemental Enhancement Information

[0064] SPS: Sequence Parameter Set

[0065] STRP: Short-Term Reference Picture

[0066] SVC: Scalable Video Coding

[0067] VCL: Video Coding Layer

[0068] VPS: Video Parameter Set

[0069] VTM: VVC Test Model

[0070] VUI: Video Usability Information

[0071] VVC: Versatile Video Coding

[0072] 3. Preliminary Exploration

[0073] Video codec standards have primarily evolved through the development of well-known ITU-T and ISO / IEC standards. ITU-T developed the H.261 and H.263 video standards, while ISO / IEC developed the MPEG-1 and MPEG-4 video standards. These two organizations jointly developed the H.262 / MPEG-2 video standard and the H.264 / MPEG-4 Advanced Video Codec (AVC) and H.265 / HEVC standards. Since H.262, video codec standards have been based on a hybrid video codec architecture, utilizing time prediction plus transform coding. To explore future video codec technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Group (JVET) in 2015. Since then, JVET has adopted many new methods and incorporated them into reference software called the Joint Exploration Model (JEM). JVET meetings are held concurrently quarterly, and the goal of the new codec standard is to reduce the bitrate by 50% compared to HEVC. The new video codec standard was officially named Multifunctional Video Codec (VVC) at the JVET meeting in April 2018, and the first version of the VVC Test Model (VTM) was released at that time. Due to ongoing efforts to standardize VVC, new codec technologies have been adopted into the VVC standard at every JVET meeting. The VVC working draft and test model VTM are updated after each meeting. The current goal of the VVC project is to achieve Technical Finality (FDIS) at the meeting in July 2020. 3.1. Reference Picture Management and Reference Picture List (RPL)

[0074] Reference picture management is a core function required by any video codec scheme that uses inter-frame prediction. Reference picture management manages the storage and removal of reference pictures in the decoded picture buffer (DPB) and puts reference pictures into the RPL in their correct order.

[0075] HEVC's reference picture management differs from AVC's, including reference picture flagging and removal from the decoded picture buffer (DPB) and reference picture list construction (RPLC). Instead of AVC's sliding window-based reference picture flagging mechanism with Adaptive Memory Management Control Operations (MMCO), HEVC specifies a reference picture management and flagging mechanism based on a so-called Reference Picture Set (RPS), and therefore RPLC is based on the RPS mechanism. An RPS consists of a reference picture set associated with a picture (consisting of all reference pictures preceding the associated picture in decoding order), which can be used for inter-frame prediction of the associated picture or any picture following the associated picture in decoding order. The reference picture set consists of five reference picture lists. The first three lists contain all reference pictures that can be used for inter-frame prediction of the current picture and for inter-frame prediction of one or more pictures following the current picture in decoding order. The other two lists consist of all reference pictures that are not used for inter-frame prediction of the current picture but can be used for inter-frame prediction of one or more pictures following the current picture in decoding order. RPS provides "intra-frame encoding / decoding" signaling in the DPB state, instead of "inter-frame encoding / decoding" signaling as in AVC, primarily to improve error resilience. HEVC's RPLC procedure is based on RPS, notifying the index by signaling a subset of RPS for each reference index; this process is simpler than the RPLC procedure in AVC.

[0076] VVC's reference picture management is more similar to HEVC than AVC's, but simpler and more robust. As in those standards, two Reference Picture Sets (RPLs), List 0 and List 1, are derived, but they are not based on the reference picture set concept used in HEVC or the automatic sliding window process used in AVC; instead, they are communicated more directly by signaling. Reference pictures used for RPLs are listed as active and inactive entries, and only active entries can be used as reference indices for inter-frame prediction of the current picture's CTU. Invalid entries indicate other pictures to be saved in the DPB for reference by other pictures arriving later in the bitstream.

[0077] 3.2. Random Access and its Support in HEVC and VVC

[0078] Random access refers to accessing and decoding the bitstream starting with the image that is not the first image in the decoding order. To support tuning and channel switching in broadcast / multicast and multi-party video conferencing, searching in local playback and streaming, and stream adaptation in streaming, the bitstream needs to include frequent random access points. These access points are typically intra-frame codec images, but can also be inter-frame codec images (e.g., in the case of progressive decoding refresh).

[0079] HEVC includes signaling for Intra-Frame Random Access Point (IRAP) pictures in the NAL unit header via NAL unit type. Three types of IRAP pictures are supported: Instantaneous Decoder Refresh (IDR), Full Random Access (CRA), and Broken Link Access (BLA) pictures. IDR pictures constrain the inter-frame picture prediction structure to not reference any pictures preceding the current Group of Pictures (GOP), and are often referred to as Closed GOP IRAP pictures. CRA pictures are less restrictive by allowing some pictures to reference pictures preceding the current GOP; in the case of random access, all pictures are discarded. CRA pictures are often referred to as Open GOP IRAP pictures. BLA pictures are typically derived from the concatenation of two bitstreams or a portion thereof in a CRA picture, for example, during stream switching. To enable better system utilization of IRAP pictures, a total of six different NAL units are defined to signal the properties of IRAP pictures. This can be used to better match the stream access point types defined in the ISO Basic Media File Format (ISOBMFF), which are used for random access support in Dynamic Adaptive Streaming (DASH) over HTTP.

[0080] VVC supports three types of IRAP pictures, two types of IDR pictures (one type has an associated RADL picture, and the other does not), and one type of CRA picture. These are essentially the same as HEVC. The BLA picture type in HEVC is not included in VVC, primarily for two reasons: i) The basic functionality of a BLA picture can be achieved by adding a sequence NAL unit end to a CRA picture, the presence of which indicates that a new CVS begins in a single-layer bitstream. ii) During the development of VVC, it was desirable to specify fewer NAL unit types than in HEVC, as shown by using 5 bits instead of 6 bits for the NAL unit type field in the NAL unit header.

[0081] Another key difference in random access support between VVC and HEVC is the more canonical support for GDR in VVC. In GDR, bitstream decoding can begin with an inter-frame encoded picture, although not the entire picture region may be correctly decoded initially, but after several pictures, the entire picture region will be correctly decoded. AVC and HEVC also support GDR, using Recovery Point (SEI) messages to signal GDR random access points and recovery points. In VVC, a new NAL unit type is specified to indicate GDR pictures, and recovery points are signaled in the picture header syntax structure. This allows CVS and bitstreams to begin with GDR pictures. This means that an entire bitstream can contain only inter-frame encoded pictures, without any single intra-frame encoded pictures. The main benefit of specifying GDR support in this way is providing consistent behavior for GDR. GDR enables encoders to smooth the bitrate of a bitstream by distributing intra-frame encoded stripes or blocks across multiple images, rather than intra-frame encoding and decoding the entire image, thus significantly reducing end-to-end latency. This is considered more important than ever today as ultra-low latency applications such as wireless displays, online gaming, and drone-based applications become increasingly popular.

[0082] Another GDR-related feature in VVC is virtual boundary signaling. The boundary between the refreshed area (i.e., the correctly decoded area) and the unrefreshed area at the GDR image and its recovery point can be signaled as a virtual boundary. When signaled, loop filtering across the boundary will not be applied, thus preventing decoding mismatches in samples at or near the boundary. This is useful when the application decides to display the correctly decoded area during the GDR process.

[0083] IRAP images and GDR images can be collectively referred to as Random Access Point (RAP) images.

[0084] 3.3. Image resolution variations within a sequence

[0085] In AVC and HEVC, the spatial resolution of a picture cannot be changed unless a new sequence with a new SPS begins with an IRAP picture. VVC allows changing the picture resolution within a sequence at locations where IRAP pictures are not encoded; IRAP pictures are always intra-frame encoded and decoded. This feature is sometimes called Reference Picture Resampling (RPR) because it requires resampling the reference picture used for inter-frame prediction when the reference picture has a different resolution than the current picture being decoded.

[0086] The scaling ratio is limited to greater than or equal to 1 / 2 (2x downsampling from the reference image to the current image) and less than or equal to 8 (8x upsampling). Three sets of resampling filters with different frequency cutoffs are specified to handle various scaling ratios between the reference and current images. The three sets of resampling filters are applied to scaling ratios ranging from 1 / 2 to 1 / 1.75, from 1 / 1.75 to 1 / 1.25, and from 1 / 1.25 to 8, respectively. Each set of resampling filters has 16 phases for luma and 32 phases for chroma, the same as in motion-compensated interpolation filters. In fact, the normal MC interpolation process is a special case of the resampling process where the scaling ratio ranges from 1 / 1.25 to 8. The horizontal and vertical scaling ratios are derived based on the image width and height, as well as the left, right, top, and bottom scaling offsets specified for the reference and current images.

[0087] Other aspects of the VVC design that support this feature differ from HEVC include: i) Picture resolution and the corresponding consistency window are signaled in the PPS instead of the SPS, where the maximum picture resolution is signaled in the SPS. ii) For a single-layer bitstream, each picture storage (the slot in the DPB used to store a decoded picture) occupies the buffer size required to store the decoded picture with the maximum picture resolution.

[0088] 3.4. Scalable Video Codec (SVC) in General and VVC

[0089] Scalable Video Coding (SVC, sometimes also called Scalability in Video Coding) refers to video coding and decoding using a base layer (BL) (sometimes called a reference layer (RL)) and one or more scalable enhancement layers (EL). In SVC, the base layer can carry video data with a basic quality level. One or more enhancement layers can carry additional video data to support, for example, higher spatial, temporal, and / or signal-to-noise ratio (SNR) levels. Enhancement layers can be defined relative to previously encoded layers. For example, the bottom layer can be used as a BL, while the top layer can be used as an EL. Intermediate layers can be used as ELs or RLs, or both. For example, an intermediate layer (e.g., a layer that is neither the lowest nor the highest layer) can be an EL of the layer below the intermediate layer (such as a base layer or any intermediate enhancement layer) and simultaneously used as an RL of one or more enhancement layers above the intermediate layer. Similarly, in the multi-view or 3D extension of the HEVC standard, there can be multiple views, and information from one view can be used to codec (e.g., encode or decode) information from another view (e.g., motion estimation, motion vector prediction, and / or other redundancy).

[0090] In SVC, parameters used by the encoder or decoder are grouped into parameter sets based on the codec level (e.g., video level, sequence level, picture level, stripe level, etc.), and these sets may be utilized at that codec level. For example, parameters that can be utilized by one or more codec video sequences at different layers in a bitstream can be included in the Video Parameter Set (VPS), and parameters that can be utilized by one or more pictures in a codec video sequence can be included in the Sequence Parameter Set (SPS). Similarly, parameters used by one or more stripes in a picture can be included in the Picture Parameter Set (PPS), and other parameters specific to a single strip can be included in the stripe header. Likewise, indications of which parameter set(s) a particular layer uses at a given time can be provided at various codec levels.

[0091] Because of VVC's support for Reference Picture Resampling (RPR), support for bitstreams containing multiple layers (e.g., two layers with SD and HD resolutions in VVC) can be designed without any additional signal processing-level codec tools, as the upsampling required for spatial scalability support can be achieved using only RPR upsampling filters. However, scalability support requires a higher level of syntax changes (compared to no scalability support). Scalability support is specified in VVC version 1. Unlike scalability support in any earlier video codec standards (including extensions to AVC and HEVC), VVC's scalability is designed to be as friendly as possible to single-layer decoder designs. The decoding capability of a multi-layer bitstream is specified as if there were only one layer in the bitstream. For example, decoding capabilities such as the DPB size are specified in a way that is independent of the number of layers in the bitstream to be decoded. Essentially, a decoder designed for a single-layer bitstream does not require many changes to be able to decode multi-layer bitstreams. Compared to the multi-layer extension designs of AVC and HEVC, HLS is significantly simplified at the expense of some flexibility. For example, IRAPU requires a picture of every layer present in CVS.

[0092] 3.5. Parameter Set

[0093] AVC, HEVC, and VVC specify parameter sets. Parameter set types include SPS, PPS, APS, and VPS. All versions of AVC, HEVC, and VVC support SPS and PPS. VPS was introduced with HEVC and is included in both HEVC and VVC. APS is not included in AVC or HEVC, but is included in the latest VVC draft text.

[0094] The Sequence-Level Prefix (SPS) is designed to carry sequence-level header information, while the Picture-Level Prefix (PPS) is designed to carry infrequently changing picture-level header information. Using SPS and PPS, infrequently changing information does not need to be repeated for each sequence or picture, thus avoiding redundant signaling. Furthermore, the use of SPS and PPS enables out-of-band transmission of important header information, thereby not only avoiding the need for redundant transmission but also improving error resilience.

[0095] The VPS was introduced to carry sequence-level header information shared by all layers in a multi-layer bitstream.

[0096] The purpose of APS is to carry such image-level or stripe-level information, which requires a considerable number of bits to encode and decode, can be shared by multiple images, and can have a considerable number of different variations in the sequence.

[0097] 4. The technical problem solved by the disclosed technical solution

[0098] The existing design for handling discardable images and AUs in the latest VVC text (JVET-R2001-vA / v10) has the following problems:

[0099] 1) In Clause 3 (Definitions), the sentence “RASL images shall not be used as reference images for the decoding process of non-RASL images” as part of the definition of RASL images is problematic because it no longer applies in some cases, and therefore it causes confusion and interoperability issues.

[0100] 2) In Clause D.4.2 (Picture Timing SEI Message Semantics), the constraint that pt_cpb_alt_timing_info_present_flag equals 0 is specified in a picture-specific manner. However, since HRD operations are based on OLS, this semantics is technically incorrect and will cause interoperability issues.

[0101] 3) The derivation of prevTid0Pic in the semantics of delta_poc_msb_cycle_present_flag[i][j] and the decoding process of picture order count (POC) does not take into account the value of ph_non_ref_pic_flag. This will cause problems when pictures with ph_non_ref_pic_flag equal to 1 are discarded, because the subsequently derived POC value for the reference picture for the picture and / or signaling notification may be incorrect, and unexpected erroneous decoding behavior may occur, including decoder crashes.

[0102] 4) In Clause C.2.3 (Timing of DU Removal and DU Decoding), the variable prevNonDiscardablePic is specified in a picture-specific manner, and the value of ph_non_ref_pic_flag is not considered. Therefore, similar issues to those described above may occur.

[0103] 5) Similar issues to question 4 apply to notDiscardablePic and prevNonDiscardablePic in clause D.3.2 (buffered time SEI message semantics) and prevNonDiscardablePic in clause D.4.2 (picture time-series SEI message syntax).

[0104] 6) The semantics of bp_concatenation_flag and bp_cpb_removal_delay_delta_minus1 are specified in a picture-specific manner. However, since HRD operations are based on OLS, the semantics are technically incorrect and will cause interoperability issues.

[0105] 7) In the constraints on the value of bp_alt_cpb_params_present_flag, there are two issues with the semantics of bp_alt_cpb_params_present_flag: First, it is specified in an image-specific way, while it should be AU-specific; second, it only considers IRAP images, while GDR images also need to be considered.

[0106] 8) In Clause D.4.2 (Picture Sequence SEI Message Semantics), the semantics of variables BpResetFlag, CpbRemovalDelayMsb[i], and CpbRemovalDelayVal[i], as well as the syntax elements pt_dpb_output_delay, pt_dpb_output_du_delay, and pt_display_elemental_periods_minus1, are specified in a picture-specific manner. However, since HRD operations are based on OLS, the semantics are technically incorrect and could cause interoperability issues.

[0107] 9) In Clause C.4 (Bitstream Consistency), the derivation of maxPicOrderCnt and minPicOrderCnt does not take into account the value of ph_non_ref_pic_flag. This will cause problems when pictures with ph_non_ref_pic_flag equal to 1 are discarded, because the subsequently derived POC values ​​for reference pictures for pictures and / or signaling notifications may be incorrect, and unexpected decoding errors may occur, including decoder crashes.

[0108] 10) In the semantics of pt_display_elemental_periods_minus1, the syntax element fixed_pic_rate_within_cvs_flag[TemporalId] is used. However, since the semantics should be described in the context of the highest TemporalId value of the target, as in other semantics of BP, PT, and DUI SEI messages, fixed_pic_rate_within_cvs_flag[Htid] should be used instead.

[0109] 11) The semantics of dui_dpb_output_du_delay are specified in a picture-specific manner. However, since HRD operations are based on OLS, the semantics are technically incorrect and will cause interoperability issues.

[0110] 5. List of technical solutions and implementation examples

[0111] To address the aforementioned and other issues, methods outlined below are disclosed. These items should be considered as examples for interpreting general concepts, and not interpreted narrowly. Furthermore, these items can be applied individually or in combination in any way.

[0112] 1) To address issue 1, in Clause 3 (Definitions), the sentence in the comment that is part of the definition of RASL images, “RASL images are not used as reference images for the decoding process of non-RASL images”, is changed to “RASL images are not used as reference images for the decoding process of non-RASL images, unless RADL sub-images in the RASL image can be used for inter-frame prediction of juxtaposed RADL sub-images in RADL images associated with the same CRA image as the RASL image, when they exist.”

[0113] 2) To address issue 2, in clause D.4.2 (Picture Timing SEI Message Semantics), the description of the constraint that pt_cpb_alt_timing_info_present_flag equals 0 is changed from picture-specific to AU-specific, and the following is added: Here, RASL pictures only contain RASL NUTs.

[0114] a. In one example, the constraint is specified as follows: when all images in the associated AU are RASL images with pps_mixed_nalu_types_in_pic_flag equal to 0, the value of pt_cpb_alt_timing_info_present_flag should be equal to 0.

[0115] b. In another example, the constraint is specified as follows: when all pictures in the associated AU are RASL pictures, where for RASL pictures, pps_mixed_nalu_types_in_pic_flag is equal to 0 and pt_cpb_alt_timing_info_present_flag should be equal to 0.

[0116] c. In another example, the constraint is specified as follows: when all pictures in the associated AU are RASL pictures containing VCL NAL units with nal_unit_type equal to RASL_NUT, the value of pt_cpb_alt_timing_info_present_flag should be equal to 0. 3) To address issue 3, ph_non_ref_pic_flag is added to the derivation of the semantics of delta_poc_msb_cycle_present_flag[i][j] and the picture order count in the decoding process of prevTid0Pic. 4) To address issue 4, in clause C.2.3 (Timing of DU Removal and Decoding of DU), the specification of prevNonDiscardablePic is changed from picture-specific to AU-specific, including renaming it to prevNonDiscardableAu, and adding ph_non_ref_pic_flag to the derivation of the same variable.

[0117] 5) To address issue 5, for notDiscardablePic and prevNonDiscardablePic in Clause D.3.2 (buffered time SEI message semantics) and prevNonDiscardablePic in Clause D.4.2 (picture time SEI message semantics), the specifications of these variables are changed from picture-specific to AU-specific. This includes renaming them to notDiscardableAu and prevNonDiscardableAu respectively, and adding ph_non_ref_pic_flag to the derivation of these two variables.

[0118] 6) To resolve issue 6, the semantic descriptions of bp_concatenation_flag and bp_cpb_removal_delay_delta_minus1 were changed from image-specific to AU-specific.

[0119] 7) To address issue 7, the constraints on the value of bp_alt_cpb_params_present_flag are specified in an AU-specific manner, and it is stipulated that the value of bp_alt_cpb_params_present_flag also depends on whether the associated AU is a GDR AU.

[0120] 8) To address problem 8, the semantics of the variables BpResetFlag, CpbRemovalDelayMsb[i], and CpbRemovalDelayVal[i], as well as the semantics of the syntax elements pt_dpb_output_delay, pt_dpb_output_du_delay, and pt_display_elemental_periods_minus1 in clause D.4.2 (Semantics of Picture Sequence SEI Messages), were changed from picture-specific to AU-specific.

[0121] 9) To address issue 9, in Clause C.4 (Bitstream Consistency), in the derivation of maxPicOrderCnt and minPicOrderCnt, add ph_non_ref_pic_flag to "the previous picture in decoding order whose TemporalId is equal to 0 and is not a RASL or RADL picture".

[0122] 10) To address issue 10, use fixed_pic_rate_within_cvs_flag[Htid] instead of fixed_pic_rate_within_cvs_flag[TemporalId] to specify the semantics of pt_display_elemental_periods_minus1.

[0123] 11) To address problem 11, specify the semantics of dui_dpb_output_du_delay in an AU-specific manner.

[0124] 6. Example

[0125] The following are some example embodiments of some aspects of the invention outlined in Section 5 of the previous article, which can be applied to the VVC specification. The modified text is based on the latest VVC text in JVET-R2001-vA / v10. Most relevant additions or modifications are indicated in bold, underlined, and italic, for example, "Using A..." Some deleted parts are marked with italics and strikethrough, for example, "based on..." "B". There may be other editable changes, which are not highlighted.

[0126] 6.1. First Embodiment

[0127] This embodiment applies to projects 1 through 9.

[0128] 3 Definitions ...

[0130] Random access skip-before (RASL) images: encoded / decoded images in which at least one VCL NAL unit has a nal_unit_type equal to RASL_NUT and the nal_unit_type of all other VCL NAL units is equal to RASL_NUT or RADL_NUT.

[0131] Note – All RASL images are preceding images of their associated CRA images. When the NoOutputBeforeRecoveryFlag of the associated CRA image is equal to 1, the RASL image is not output and may not be correctly decoded because it may contain references to images that do not exist in the bitstream. RASL images are not used as reference images for decoding non-RASL images. When sps_field_seq_flag equals 0, all RASL images, when present, are decoded before all non-preceding images of the same associated CRA image in the decoding order. ...

[0133] 7.4.9 Semantics of the Reference Image List ...

[0135] A value of 1 for delta_poc_msb_cycle_present_flag[i][j] indicates the existence of delta_poc_msb_cycle_lt[i][j]. A value of 0 for delta_poc_msb_cycle_present_flag[i][j] indicates the non-existence of delta_poc_msb_cycle_lt[i][j].

[0136] Set prevTid0Pic to nuh_layer_id (which is the same as the current image) and TemporalId. Previous images in decoding order that are equal to 0 and are not RASL or RADL images. Let setOfPrevPocVals be a set including the following:

[0137] –prevTid0Pic's PicOrderCntVal,

[0138] – PicOrderCntVal for each image referenced by an entry in RefPicList[0] or RefPicList[1] of prevTid0Pic and whose nuh_layer_id is the same as the current image.

[0139] -PicOrderCntVal for each image that follows prevTid0Pic in decoding order, has the same nuh_layer_id as the current image, and precedes the current image in decoding order.

[0140] When the modulus MaxPicOrderCntL equals more than one value in sbsetOfPrevPocVals of PocLsbLt[i][j], the value of delta_poc_msb_cycle_present_flag[i][j] should be equal to 1. ...

[0142] 8.3.1 Decoding process of image sequential counting ...

[0144] When ph_poc_msb_cycle_present_flag equals 0 and the current image is not a CLVSS image, the variables prevPicOrderCntLsb and prevPicOrderCntMsb are derived as follows:

[0145] – Set prevTid0Pic to nuh_layer_id, which is the same as the current image, and TemporalId. Previous images in decoding order that are equal to 0 and are not RASL or RADL images.

[0146] The variable prevPicOrderCntLsb is set to equal ph_pic_order_cnt_lsb of prevTid0Pic.

[0147] The variable prevPicOrderCntMsb is set to be equal to the PicOrderCntMsb of prevTid0Pic. ...

[0149] C.2.3 Timing of DU Removal and DU Decoding ...

[0151] The nominal removal time of AU n from CPB is specified as follows:

[0152] – If AU n is an AU with n equal to 0 (the AU that initializes the HRD), then the nominal removal time of the AU from the CPB is determined by the following equation:

[0153] AuNominalRemovalTime[0]=InitCpbRemovalDelay[Htid][ScIdx]÷90000(C.9)

[0154] –Otherwise, the following applies:

[0155] – When AU n is the first AU of the BP of the uninitialized HRD, the following applies:

[0156] The nominal removal time of AU n from CPB is specified as follows:

[0157]

[0158]

[0159] AuNominalRemovalTime[first] [InPrevBuffPeriod] is the nominal removal time of the first AU of the previous BP, AuNominalRemovalTime[prevNonDiscardable] [This is a TemporalId value equal to 0] Not a RASL or RADL image In decoding order The nominal removal time, AuCpbRemovalDelayVal, is the value of CpbRemovalDelayVal[Htid] derived from pt_cpb_removal_delay_minus1[Htid] and pt_cpb_removal_delay_delta_idx[Htid] in the PT SEI message, and bp_cpb_removal_delay_delta_val[pt_cpb_removal_delay_delta_idx[Htid]] in the BP SEI message associated with AU n and concatenationFlag and CpbRemovalDelayDeltaMinus1, as specified in Clause C.1, are the values ​​of the syntax elements bp_concatenation_flag and bp_cpb_removal_delay_delta_minus1 in the BP SEI message associated with AU n, as specified in Clause C.1.

[0160] After the derivation of the nominal CPB removal time and before the derivation of the DPB output time of access unit n, the variables DpbDelayOffset and CpbDelayOffset are derived as follows:

[0161] – If one or more of the following conditions are true, DpbDelayOffset is set to the value of the PT SEI message syntax element dpb_delay_offset[Htid] equal to AU n+1, and CpbDelayOffset is set to the value of the PT SEI message syntax element cpb_delay_offset[Htid] equal to AUn+1, wherein the PT SEI message containing the syntax element is selected as specified in Clause C.1:

[0162] –AU n’s UseAltCpbParamsFlag is equal to 1.

[0163] –DefaultInitCpbParamsFlag equals 0.

[0164] Otherwise, both DpbDelayOffset and CpbDelayOffset are set to 0.

[0165] – When AU n is not the first AU of BP, the nominal removal time of AU n from CPB is determined by the following equation:

[0166] AuNominalRemovalTime[n]=AuNominalRemovalTime[first InCurrBuffPeriod]+ClockTick*(AuCpbRemovalDelayVal-CpbDelayOffset)(C.11)

[0167] AuNominalRemovalTime[first] [InCurrBuffPeriod] is the nominal removal time of the first AU of the current BP, and AuCpbRemovalDelayVal is the value of CpbRemovalDelayVal[OpTid] derived from pt_cpb_removal_delay_minus1[OpTid] and pt_cpb_removal_delay_delta_idx[OpTid] in the PT SEI message and bp_cpb_removal_delay_delta_val[pt_cpb_removal_delay_delta_idx[OpTid]] in the BP SEI message associated with AU n as specified in Clause C.1. ...

[0169] C.4 Bitstream Consistency ...

[0171] Set currPicLayerId to equal the nuh_layer_id of the current image.

[0172] For each current image, let the variables maxPicOrderCnt and minPicOrderCnt be set to the maximum and minimum values ​​of the PicOrderCntVal values ​​for the following images whose nuh_layer_id is equal to currPicLayerId:

[0173] – Current image.

[0174] –TemporalId Previous images in decoding order that are equal to 0 and are not RASL or RADL images.

[0175] – The STRP referenced by all entries in RefPicList[0] and all entries in RefPicList[1] of the current image.

[0176] –PictureOutputFlag equals 1, AuCpbRemovalTime[n] is less than AuCpbRemovalTime[currPic], and DpbOutputTime[n] is greater than or equal to AuCpbRemovalTime[currPic], where currPic is the current image. ...

[0178] D.3.2 Buffer Period SEI Message Semantics ...

[0180] When a BP SEI message exists TemporalId equals 0 and Not a RASL or RADL image

[0181] when Bitstream in decoding order season TemporalId equals 0 and Not a RASL or RADL image In decoding order

[0182] The existence of BP SEI messages is defined as follows:

[0183] – If NalHrdBpPresentFlag equals 1 or VclHrdBpPresentFlag equals 1, then the following applies to each AU in CVS:

[0184] – If the AU is an IRAP or GDR AU, the BP SEI message applicable to the operation point should be associated with the AU.

[0185] Otherwise, if AU The BP SEI message applicable to the operation point can be associated with an AU or not.

[0186] Otherwise, the AU should not be associated with the BP SEI message applicable to the operation point.

[0187] Otherwise (NalHrdBpPresentFlag and VclHrdBpPresentFlag are both equal to 0), there is no AU associated with the BP SEI message in CVS.

[0188] Note 1 – For some applications, frequent BP SEI messages may be required (e.g., for...).

[0189] Random access or bitstream concatenation). ...

[0191] The value of `bp_alt_cpb_params_present_flag` equal to 1 specifies the presence of the syntax element `bp_use_alt_cpb_params_flag` in the BP SEI message and the presence of alternative timing information in the PT SEI message of the current BP. When it does not exist, the value of `bp_alt_cpb_params_present_flag` is inferred to be 0. At that time, the value of bp_alt_cpb_params_present_flag should be equal to 0. ...

[0193] when Bitstream in decoding order At that time, bp_concatenation_flag indicates The nominal CPB removal time is relative to the BP SEI message. Determined by the nominal CPB removal time or relative to The nominal CP removal time is used to determine this. ...

[0195] when Bitstream in decoding order bp_cpb_removal_delay_delta_minus1 plus 1 specifies the distance relative to the specified distance. The nominal CPB removal time is the CPB removal delay increment value. The length of this syntax element is bp_cpb_removal_delay_length_minus1+1 bits.

[0196] when BP SEI message And bp_concatenation_flag equals 0 and Bitstream in decoding order The requirement for bitstream consistency applies to the following constraints:

[0197] -if If not associated with BP SEI messages, then pt_cpb_removal_delay_minus1 should be equal to The pt_cpb_removal_delay_minus1 plus bp_cpb_removal_delay_delta_minus1+1.

[0198] -otherwise, pt_cpb_removal_delay_minus1 should be equal to bp_cpb_removal_delay_delta_minus1.

[0199] Note 2 – When BP SEI message Furthermore, when bp_concatenation_flag equals 1, it is not used. pt_cpb_removal_delay_minus1. In some cases, the constraints specified above can be easily resolved by simply adding them at the splicing point. The value of bp_concatenation_flag in the BP SEI message is changed from 0 to 1 to concatenate the bitstream (using a properly designed reference structure). When bp_concatenation_flag equals 0, the constraints specified above allow the decoder to check whether the constraints are satisfied as a detection. The method of loss. ...

[0201] D.4.2 Image Sequential SEI Message Semantics ...

[0203] `pt_cpb_alt_timing_info_present_flag` equal to 1 indicates that the syntax elements `pt_nal_cpb_alt_initial_removal_delay_delta[i][j]`, `pt_nal_cpb_alt_initial_removal_offset_delta[i][j]`, `pt_nal_cpb_delay_offset[i]`, `pt_nal_dpb_delay_offset[i]`, `pt_vcl_cpb_alt_initial_removal_delay_delta[i][j]`, `pt_vcl_cpb_alt_initial_removal_offset_delta[i][j]`, `pt_vcl_cpb_delay_offset[i]`, and `pt_vcl_dpb_delay_offset[i]` may exist in the PT SEI message. `pt_cpb_alt_timing_info_present_flag` equal to 0 indicates that these syntax elements do not exist in the PT SEI message. At that time, the value of pt_cpb_alt_timing_info_present_flag should be equal to 0.

[0204] Note 1 – For following the decoding order in For more than one subsequent AU, the value of pt_cpb_alt_timing_info_present_flag may be equal to 1. However, the alternative timing only applies if pt_cpb_alt_timing_info_present_flag is equal to 1 and follows the decoding order. The first AU that followed. ...

[0206] `pt_vcl_dpb_delay_offset[i]` specifies the offset to be used in deriving the DPB output time of the IRAP AU associated with the BP SEI message, when the AU associated with the PT SEI message directly follows the IRAP AU associated with the BP SEI message in decoding order for the i-th sublayer of the VCL HRD. The length of `pt_vcl_dpb_delay_offset[i]` is `bp_dpb_output_delay_length_minus1+1` bits. When it does not exist, the value of `pt_vcl_dpb_delay_offset[i]` is inferred to be equal to 0.

[0207] The variable BpResetFlag is derived as follows:

[0208] -if When associated with a BP SEI message, BpResetFlag is set to 1.

[0209] Otherwise, BpResetFlag is set to 0. ...

[0211] `pt_cpb_removal_delay_delta_idx[i]` specifies the index of the CPB removal increment, which applies to the Htid of `i` in the list `bp_cpb_rmoval_dellta_val[j]` (inclusive) of `bp_num_cpb_emoval_deley_deltas_minus1`, ranging from 0 to the endpoint. The length of `pt_cpb_removal_delay_delta_idx[i]` is Ceil(Log2(bp_num_cpb_emoval_deley_deltas_minus1+1)) bits. The value of `pt_cpb_removal_delay_delta_idx[i]` is inferred to be 0 when `pt_cpb_removal_delay_delta_idx[i]` does not exist and `pt_cpb_removal_delay_delta_enabled_flag[i]` is equal to 1.

[0212] The variables CpbRemovalDelayMsb[i] and CpbRemovalDelayVal[i] are derived as follows:

[0213] – If the current AU is the AU that initializes the HRD, then CpbRemovalDelayMsb[i] and CpbRemovalDelayVal[i] are both set to 0, and the value of cpbRemovalDelayValTmp[i] is set to pt_cpb_removal_delay_minus1[i]+1.

[0214] Otherwise, let TemporalId equals 0 Not a RASL or RADL image In decoding order for Let prevCpbRemovalDelayMinus1[i], prevCpbRemovalDelayMsb[i], and prevBpResetFlag be set to the values ​​of cpbRemovalDeliveValTmp[i]-1, CpbRemovalDelayMsb[i], and BpResetFlag, respectively, and the following applies: ...

[0216] Pt_dpb_output_delay is used for calculation The DPB output time. It specifies the time after the AU is removed from the CPB before outputting from the DPB. How many clock ticks were waited for?

[0217] Note 2 – When the decoded image is still marked as “for short-term reference” or “for long-term reference”, the decoded image will not be removed from the DPB at output.

[0218] The length of pt_dpb_output_delay is bp_dpb_output_delay_length_minus1+1 bits. When max_dec_pic_buffering_minus1[Htid] equals 0, the value of pt_dpb_output_delay should be equal to 0.

[0219] The output time derived from the pt_dpb_output_delay of any image output from a decoder that conforms to the output timing should precede the output time derived from the pt_dpb_output_delay of all images in any subsequent CVS in the decoding order.

[0220] The order of image output based on the value of this syntax element should be the same as the order based on the value of PicOrderCntVal.

[0221] Images that did not pass the "collision" process are because they were decoded in the order that ph_no_output_of_prior_pics_flag is equal to 1 or inferred to be equal to 1. Previously, the output time derived from pt_dpb_output_delay should increase as the PicOrderCntVal value increases relative to all images within the same CVS.

[0222] When DecodingUnitHrdFlag equals 1, pt_dpb_output_du_delay is used for calculation. The DPB output time. It specifies the time after the last AU is removed from the CPB before outputting from the DPB. How many clock beats did we have to wait beforehand? ...

[0224] When sps_field_seq_flag equals 0 and fixed_pic_rate_within_cvs_flag When pt_display_elemental_periods_minus1 equals 1, increment 1 to indicate this. This refers to the number of time intervals between the element images displayed by the model.

[0225] When fixed_pic_rate_within_cvs_flag When pt_display_elemental_periods_minus1 is equal to 0 or sps_field_seq_flag is equal to 1, the value of pt_display_elemental_periods_minus1 should be equal to 0.

[0226] When sps_field_seq_flag equals 0 and fixed_pic_rate_within_cvs_flag When equal to 1, a value greater than 0 for pt_display_elemental_periods_minus1 can be used to indicate the frame repetition period of a display using a fixed frame refresh interval equal to DpbOutputElementalInterval[n], as given in Equation 112. ...

[0228] D.5.2DU Information SEI Message Semantics ...

[0230] When DecodingUnitHrdFlag equals 1 and bp_du_dpb_params_in_pic_timing_sei_flag equals 0, dui_dpb_output_du_delay is used for calculation. The DPB output time. It specifies the time after the last AU is removed from the CPB before outputting from the DPB. The number of sub-clock ticks to wait beforehand. When not present, the value of `dui_dpb_output_du_delay` is inferred to be equal to `pt_dpb_output_du_delay`. The length of the syntax element `dui_dpb_output_du_delay` is given in bits by `bp_dpb_output_delay_du_length_minus1+1`.

[0231] The requirement for bitstream consistency is that all DU information SEI messages associated with the same AU, applicable to the same point of operation, and with bp_du_dpb_params_in_pic_timing_sei_flag equal to 0 should have the same dui_dpb_output_du_delay value.

[0232] The output time derived from the dui_dpb_output_du_delay of any image output from a decoder that conforms to the output timing should precede the output time derived from the dui_dpb_output_du_delay of all images in any subsequent CVS in the decoding order.

[0233] The order of image output based on the value of this syntax element should be the same as the order based on the value of PicOrderCntVal.

[0234] For images that were not output through the "collision" process, this is because they were decoded before ph_no_output_of_prior_pics_flag was equal to 1 or inferred to be equal to 1. The output time derived from dui_dpb_output_du_delay should increase as the PicOrderCntVal value increases relative to all images within the same CVS.

[0235] For any two images in CVS, the difference in output time when DecodingUnitHrdFlag equals 1 should be the same as the difference in output time when Decoding UnitHrd Flag is 0. ...

[0237] Figure 1 This is a block diagram illustrating an example video processing system 1000 that can implement the various techniques disclosed herein. Various implementations may include some or all of the components of system 1000. System 1000 may include an input 1002 for receiving video content. The video content may be received in a raw or uncompressed format, such as 8 or 10-bit multi-component pixel values, or it may be in a compressed or encoded format. Input 1002 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet and Passive Optical Network (PON) and wireless interfaces such as Wi-Fi or cellular interfaces.

[0238] System 1000 may include a codec component 1004 capable of implementing the various codec or encoding methods described herein. Codec component 1004 can reduce the average bit rate of the video from input 1002 to the output of codec component 1004 to produce a codec representation of the video. Therefore, codec techniques are sometimes referred to as video compression or video transcoding techniques. As indicated by component 1006, the output of codec component 1004 can be stored or transmitted via a connected communication. The stored or communicatively transmitted bitstream (or codec) representation of the video received at input 1002 can be used by component 1008 to generate pixel values ​​or displayable video, which is then sent to display interface 1010. The process of generating a user-viewable video from the bitstream representation is sometimes referred to as video decompression. Furthermore, although some video processing operations are referred to as "codec" operations or tools, it should be understood that codec tools or operations are used at the encoder, and the corresponding decoding tools or operations that reverse the codec results are performed by the decoder.

[0239] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB), High Definition Multimedia Interface (HDMI), or DisplayPort. Examples of storage interfaces include SATA (Serial Advanced Technology Accessory), PCI, IDE, and so on. The technologies described herein can be found in a variety of electronic devices, such as mobile phones, laptops, smartphones, or other devices capable of performing digital data processing and / or video display.

[0240] Figure 2 This is a block diagram of a video processing apparatus 2000. The apparatus 2000 can be used to implement one or more methods described herein. The apparatus 2000 can be embodied in a smartphone, tablet computer, computer, Internet of Things (IoT) receiver, etc. The apparatus 2000 may include one or more processors 2002, one or more memories 2004, and video processing hardware 2006. The processor 2002 can be configured to implement one or more methods described herein (e.g., ...). Figures 6 to 9 Multiple memories 2004 may be used to store data and code for implementing the methods and techniques described herein. Video processing hardware 2006 may be used to implement some of the techniques described herein in hardware circuitry. In some embodiments, hardware 2006 may be partially or wholly located within one or more processors 2002, such as a graphics processor.

[0241] Figure 3 This is a block diagram illustrating an example video codec system 100 that can utilize the techniques disclosed herein. Figure 3As shown, the video encoding / decoding system 100 may include a source device 110 and a destination device 120. The source device 110 generates encoded video data and may be referred to as a video encoding device. The destination device 120 can decode the encoded video data generated by the source device 110, and the source device may be referred to as a video decoding device. The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.

[0242] Video source 112 may include sources such as video capture devices, interfaces for receiving video data from video content providers, and / or computer graphics systems for generating video data, or combinations thereof. Video data may include one or more images. Video encoder 114 encodes the video data from video source 112 to generate a bitstream. The bitstream may include a sequence of bits forming a codec representation of the video data. The bitstream may include codec images and associated data. A codec image is a codec representation of an image. Associated data may include sequence parameter sets, image parameter sets, and other syntax structures. I / O interface 116 may include a modulator / demodulator (modem) and / or a transmitter. Encoded video data may be transmitted directly to destination device 120 via network 130a through I / O interface 116. Encoded video data may also be stored on storage media / server 130b for access by destination device 120.

[0243] The target device 120 may include an I / O interface 126, a video decoder 124, and a display device 122.

[0244] I / O interface 126 may include a receiver and / or a modem. I / O interface 126 may acquire encoded video data from source device 110 or storage medium / server 130b. Video decoder 124 may decode the encoded video data. Display device 122 may display the decoded video data to a user. Display device 122 may be integrated with destination device 120, or it may be external to destination device 120, which is configured to interface with an external display device.

[0245] The video encoder 114 and the video decoder 124 can operate according to video compression standards, such as the High Efficiency Video Codec (HEVC) standard, the Multi-Functional Video Codec (VVM) standard, and other current and / or further standards.

[0246] Figure 4 This is a block diagram illustrating an example of a video encoder 200. The video encoder can be... Figure 3 The video encoder 114 in the system 100 shown.

[0247] The video encoder 200 can be configured to perform any or all of the techniques disclosed herein. Figure 4 In the example, the video encoder 200 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video encoder 200. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.

[0248] The functional components of the video encoder 200 may include a segmentation unit 201, a prediction unit 202 (which may include a mode selection unit 203), a motion estimation unit 204, a motion compensation unit 205, an intra-frame prediction unit 206, a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213, and an entropy coding unit 214.

[0249] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an intra-block copy (IBC) unit. The IBC unit can perform prediction in IBC mode, where at least one reference picture is the picture containing the current video block.

[0250] Furthermore, some components, such as the motion estimation unit 204 and the motion compensation unit 205, can be highly integrated, but for interpretive purposes, in Figure 4 The examples are shown separately.

[0251] The segmentation unit 201 can segment an image into one or more video blocks. The video encoder 200 and the video decoder 300 can support various video block sizes.

[0252] The mode selection unit 203 may, for example, select one of a plurality of encoding / decoding modes (intra-frame or inter-frame) based on the error result, and provide the resulting intra-frame or inter-frame encoded / decoded block to the residual generation unit 207 to generate residual block data, and to the reconstruction unit 212 to reconstruct the coded block for use as a reference picture. In some examples, the mode selection unit 203 may select a combination of intra-frame and inter-frame prediction (CIIP) modes, wherein the prediction is based on the inter-frame prediction signal and the intra-frame prediction signal. In the case of inter-frame prediction, the mode selection unit 203 may also select the resolution of the motion vector for the block (e.g., sub-pixel or integer pixel precision).

[0253] To perform inter-frame prediction on the current video block, motion estimation unit 204 can generate motion information for the current video block by comparing one or more reference frames from buffer 213 with the current video block. Motion compensation unit 205 can determine the predicted video block for the current video block based on motion information and decoded samples from images other than those associated with the current video block from buffer 213.

[0254] The motion estimation unit 204 and the motion compensation unit 205 can perform different operations on the current video block, for example, depending on whether the current video block is in an I-band, P-band, or B-band.

[0255] In some examples, motion estimation unit 204 can perform unidirectional prediction on the current video block, and can search for a reference video block for the current video block in the reference images of list 0 or list 1. Motion estimation unit 204 can then generate a reference index indicating a reference image in list 0 or list 1 containing the reference video block, and a motion vector indicating the spatial displacement between the current video block and the reference video block. Motion estimation unit 204 can output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current block based on the reference video block indicated by the motion information of the current video block.

[0256] In other examples, motion estimation unit 204 can perform bidirectional prediction on the current video block. Motion estimation unit 204 can search for a reference video block for the current video block in the reference images in list 0, and can also search for another reference video block for the current video block in the reference images in list 1. Motion estimation unit 204 can then generate a reference index and a motion vector, where the reference index indicates the reference images in lists 0 and 1 containing the reference video block, and the motion vector indicates the spatial displacement between the reference video block and the current video block. Motion estimation unit 204 can output the reference index and motion vector of the current video block as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current block based on the reference video blocks indicated by the motion information of the current video block.

[0257] In some examples, the motion estimation unit 204 can output a complete set of motion information for the decoder's decoding processing.

[0258] In some examples, motion estimation unit 204 may not output the complete set of motion information for the current video. Instead, motion estimation unit 204 may refer to the motion information of another video block to signal the motion information of the current video block. For example, motion estimation unit 204 may determine that the motion information of the current video block is sufficiently similar to the motion information of neighboring video blocks.

[0259] In one example, the motion estimation unit 204 may indicate a value in the syntax structure associated with the current video block that indicates to the video decoder 300 that the current video block has the same motion information as another video block.

[0260] In another example, motion estimation unit 204 can identify another video block and motion vector difference (MVD) within the syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. Video decoder 300 can use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.

[0261] As discussed above, the video encoder 200 can predictively signal motion vectors. Two examples of predictive signaling techniques that can be implemented by the video encoder 200 include Advanced Motion Vector Prediction (AMVP) and merge pattern signaling.

[0262] Intra-prediction unit 206 can perform intra-prediction on the current video block. When intra-prediction unit 206 performs intra-prediction on the current video block, it can generate prediction data for the current video block based on decoded samples from other video blocks in the same frame. The prediction data for the current video block can include the predicted video block and various syntax elements.

[0263] The residual generation unit 207 can generate residual data for the current video block by subtracting (e.g., indicated by a negative sign) the predicted video block of the current video block from the current video block. The residual data for the current video block can include residual video blocks corresponding to different sample components of the samples in the current video block.

[0264] In other examples, the current video block may not have residual data for the current video block, for example, in skip mode, and the residual generation unit 207 may not perform the subtraction operation.

[0265] The transform processing unit 208 can generate one or more transform coefficient video blocks of the current video block by applying one or more transforms to the residual video block associated with the current video block.

[0266] After the transform processing unit 208 generates a transform coefficient video block associated with the current video block, the quantization unit 209 can quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values ​​associated with the current video block.

[0267] Inverse quantization unit 210 and inverse transform unit 211 can apply inverse quantization and inverse transform to the transform coefficient video block, respectively, to reconstruct the residual video block from the transform coefficient video block. Reconstruction unit 212 can add the reconstructed residual video block to the corresponding samples of one or more predicted video blocks generated by prediction unit 202 to generate a reconstructed video block associated with the current block, which is stored in buffer 213.

[0268] After the video block is reconstructed by the reconstruction unit 212, a loop filtering operation can be performed to reduce the video block effect in the video block.

[0269] Entropy encoding unit 214 can receive data from other functional components of video encoder 200. When entropy encoding unit 214 receives data, it can perform one or more entropy encoding operations to generate entropy encoded data and output a bit stream including the entropy encoded data.

[0270] Figure 5 This is a block diagram illustrating an example of a video decoder 300, which can be... Figure 3 The video decoder 114 in the system 100 shown.

[0271] The video encoder 300 can be configured to perform any or all of the technologies disclosed herein. Figure 5 In the example, the video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video decoder 300. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.

[0272] exist Figure 5 In the example, video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra-frame prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, video decoder 300 can perform functions typically associated with video encoder 200. Figure 4 The decoding process is the inverse of the encoding process described.

[0273] Entropy decoding unit 301 can retrieve encoded bitstreams. The encoded bitstreams may include entropy-encoded video data (e.g., encoded blocks of video data). Entropy decoding unit 301 can decode the entropy-encoded video data, and motion compensation unit 302 can determine motion information from the entropy-encoded video data, including motion vectors, motion vector precision, reference image list index, and other motion information. Motion compensation unit 302 may determine this information, for example, by executing AMVP and merge modes.

[0274] The motion compensation unit 302 can generate motion compensation blocks and can perform interpolation based on an interpolation filter. The identifier of the interpolation filter to be used at sub-pixel precision can be included in the syntax element.

[0275] The motion compensation unit 302 can use interpolation filters, such as those used by the video encoder 200 during the encoding of video blocks, to calculate interpolated values ​​for sub-integer pixels of the reference block. The motion compensation unit 302 can determine the interpolation filter used by the video encoder 200 based on the received syntax information and use the interpolation filter to generate a prediction block.

[0276] The motion compensation unit 302 may use some syntax information to determine the size of the blocks of frames and / or stripes used to encode the encoded video sequence, segmentation information describing how each macroblock of the picture of the encoded video sequence is segmented, a mode indicating how each segment is encoded, one or more reference frames (and a list of reference frames) for each inter-frame coded block, and other information for decoding the encoded video sequence.

[0277] Intra-prediction unit 303 can use, for example, an intra-prediction mode received in the bitstream to form prediction blocks from spatially adjacent blocks. Inverse quantization unit 303 inverse quantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by entropy decoding unit 301. Inverse transform unit 303 applies an inverse transform.

[0278] The reconstruction unit 306 can add the residual block to the corresponding prediction block generated by the motion compensation unit 202 or the intra-frame prediction unit 303 to form a decoded block. If necessary, a deblocking filter can also be applied to filter the decoded block to remove block artifacts. The decoded video block is then stored in a buffer 307, which provides a reference block for subsequent motion compensation / intra-frame prediction and also generates decoded video for presentation on a display device.

[0279] Figures 6 to 10 It shows that it can be implemented, for example Figures 1 to 5 The above-described technical solutions are illustrated in the embodiments shown.

[0280] Figure 6 A flowchart of an example method 600 for video processing is shown. Method 600 includes, at operation 610, performing a conversion between a video and a video bitstream comprising one or more images, the bitstream conforming to a format rule specifying that Picture Timing (PT) Supplemental Enhancement Information (SEI) messages included in the bitstream are Access Unit (AU) specific, and that each of the one or more images is a Random Access Skip Precedence (RASL) image, including only RASL Network Abstraction Layer Unit Type (NUT).

[0281] Figure 7A flowchart of an example method 700 for video processing is shown. Method 700 includes, at operation 710, performing a conversion between a video and a video bitstream comprising one or more images, the bitstream conforming to a format rule that allows the use of a random access skip-before (RASL) image as a reference sub-image for predicting a juxtaposed RADL image in a RADL image associated with the same fully random access (CRA) image as the RASL image.

[0282] Figure 8 A flowchart of an example method 800 for video processing is shown. Method 800 includes, at operation 810, performing a conversion between a video and a video bitstream comprising one or more images, the bitstream conforming to a format rule specifying that the derivation of an image associated with a first flag during the decoding of image sequence counting is based on a second flag, the image associated with the first flag being a previous image in decoding order, the previous image having (i) a first identifier identical to a stripe or image header of a reference image list syntax structure, (ii) a second identifier and a second flag equal to zero, and (iii) an image type different from Random Access Skip Precursor (RASL) images and Random Access Decodeable Precursor (RADL) images, the first flag indicating the presence of a third flag in the bitstream, the second flag indicating whether the current image is used as a reference image, and the third flag used to determine the value of one or more most significant bits of the image sequence count value of a long-term reference image.

[0283] Figure 9 A flowchart of an example method 900 for video processing is shown. Method 900 includes, at operation 910, performing a conversion between a video and a video bitstream comprising one or more images, the bitstream conforming to a format rule specifying that variables used to determine the timing of a decoding unit (DU) removing or decoding a DU are access unit (AU) specific and derived based on a flag indicating whether the current image is allowed to be used as a reference image.

[0284] Figure 10A flowchart of an example method 1000 for video processing is shown. Method 1000 includes, at operation 1010, performing a conversion between a video and a video bitstream comprising one or more images, the bitstream conforming to a format rule specifying that Buffered Period Supplemental Enhancement Information (SEI) messages and Picture Timing SEI messages, when included in the bitstream, are access unit (AU) specific. A first variable associated with the Buffered Period SEI message and a second variable associated with the Buffered Period SEI message and Picture Timing SEI message are derived based on flags indicating whether the current image is allowed to be used as a reference image. The first variable indicates that the access unit includes (i) an identifier equal to zero and (ii) an image that is not a Random Access Skip Preceding (RASL) image or a Random Access Decodeable Preceding (RADL) image and whose flag is equal to zero. The second variable indicates that the current AU is not the first AU in decoding order, and the previous AU in decoding order includes (i) an identifier equal to zero and (ii) an image that is not a Random Access Skip Preceding (RASL) image or a Random Access Decodeable Preceding (RADL) image and whose flag is equal to zero.

[0285] Figure 11 A flowchart of an example method 1100 for video processing is shown. Method 1100 includes, at operation 1110, performing a conversion between a video and a video bitstream comprising one or more images, the bitstream conforming to a format rule specifying that the derivation of a first variable and a second variable associated with a first image and a second image is based on a flag, the first image being the current image and the second image being a previous image in decoding order, the previous image (i) including a first identifier equal to zero, (ii) including a flag equal to zero, and (iii) not a Random Access Skip Precedence (RASL) image or a Random Access Decodeable Precedence (RADL) image, and the first variable and the second variable being the maximum and minimum values, respectively, of the image sequence counts of each of the following images for which the second identifier is equal to the second identifier of the first image: (i) the first image, (ii) the second image, (iii) one or more short-term reference images referenced by all entries in the reference image list of the first image, and (iv) each image that has been output, the code-decode image buffer (CPB) removal time of which is less than the CPB removal time of the first image and the decode image buffer (DPB) output time is greater than or equal to the CP removal time of the first image.

[0286] Figure 12A flowchart of an example method 1200 for video processing is shown. Method 1200 includes, at operation 1210, performing a conversion between a video and a video bitstream comprising one or more pictures, the bitstream conforming to a format rule specifying that flags and syntax elements are Access Unit (AU) specific when included in the bitstream, in response to the current AU not being the first AU in the bitstream in decoding order, the flag indicating that the nominal codec picture buffer (CPB) removal time of the current AU is determined relative to (a) the nominal CPB removal time of a previous AU associated with a Buffer Period Supplemental Enhancement Information (SEI) message or (b) the nominal CPB removal time of the current AU, and in response to the current AU not being the first AU in the bitstream in decoding order, the syntax element specifying a CPB removal delay increment value relative to the nominal CPB removal time of the current AU.

[0287] Figure 13 A flowchart of an example method 1300 for video processing is shown. Method 1300 includes, at operation 1310, performing a conversion between a video and a video bitstream comprising one or more images, the bitstream conforming to a format rule specifying that multiple variables and a Picture Timing Supplement Enhancement Information (SEI) message, when included in the bitstream, are Access Unit (AU) specific. The Picture Timing SEI message includes multiple syntax elements, a first variable among multiple variables indicating whether the current AU is associated with a buffered time-period SEI message, a second and a third variable among multiple variables associated with an indication of whether the current AU is an AU for initializing a hypothetical reference decoder (HRD), a first syntax element among multiple syntax elements specifying the number of clock ticks to wait before outputting one or more decoded images of the AU from the decoded image buffer (DPB) after removing the AU from the codec image buffer (CPB), a second syntax element among multiple syntax elements specifying the number of sub-clock ticks to wait before outputting one or more decoded images of the AU from the DPB after removing the last decoded unit (DU) from the AU from the CPB, and a third syntax element among multiple syntax elements specifying the number of element image time-period intervals occupied by the one or more decoded images of the current AU for use in displaying the model.

[0288] Figure 14 A flowchart of an example method 1400 for video processing is shown. Method 1400 includes, at operation 1410, performing a conversion between a video and a video bitstream comprising one or more images, the bitstream conforming to a format rule specifying that syntax elements associated with a decoded picture buffer (DPB) are access unit (AU) specific when included in the bitstream, and the syntax elements specifying the number of sub-clock ticks to wait before outputting one or more decoded pictures of the AU from the DPB after the last decoded unit (DU) in the AU is removed from the codec picture buffer (CPB).

[0289] Figure 15 A flowchart of an example method 1500 for video processing is shown. Method 1500 includes, at operation 1510, performing a conversion between a video and a video bitstream comprising one or more pictures, the bitstream conforming to a format rule specifying that a flag is Access Unit (AU) specific when included in the bitstream, the value of the flag being based on whether the associated AU is an Intra-Frame Random Access Point (IRAP) AU or a Progressive Decoding Refresh (GDR) AU, and the value of the flag specifying (i) the presence of a syntax element in a buffer-time supplemental enhancement information (SEI) message and (ii) the presence of alternative timing information in a picture timing SEI message of the current buffer time.

[0290] Figure 16 A flowchart of an example method 1600 for video processing is shown. Method 1600 includes, at operation 1610, performing a conversion between a video and a video bitstream comprising one or more images, the bitstream conforming to a format rule specifying that the value of a first syntax element is based on a flag and a variable, the flag indicating whether the time distance between the output times of consecutive images of a hypothetical reference decoder (HRD) is constrained, the variable identifying the highest time sublayer to be decoded, and the first syntax element specifying the number of element image time intervals occupied by one or more decoded images of the current AU for use in displaying the model.

[0291] The following is a list of preferred solutions for some embodiments.

[0292] A1. A method for video processing, comprising: performing a conversion between a video comprising one or more images and a video bitstream, wherein the bitstream conforms to a format rule, wherein the format rule specifies that a Picture Timing Supplemental Enhancement Information (SEI) message is Access Unit (AU) specific when included in the bitstream, and wherein each of the one or more images is a Random Access Skip Precedence (RASL) image, comprising only an RASL Network Abstraction Layer Unit Type (NUT).

[0293] A2. According to the method of solution A1, wherein each picture in the associated AU is a RASL picture with a first flag equal to zero and a second flag equal to zero, wherein the first flag indicates whether each picture referencing the Picture Parameter Set (PPS) has more than one Video Codec Layer (VCL) Network Abstraction Layer (NAL) unit, and at least two of the more than one VCL NAL unit are of different types, and wherein the second flag indicates whether one or more syntax elements related to timing information are allowed to exist in the picture timing SEI message.

[0294] A3. According to the method of solution A2, the first flag is pps_mixed_nalu_types_in_pic_flag, and the second flag is pt_cpb_alt_timing_info_present_flag.

[0295] A4. According to the method of solution A2, wherein one or more syntax elements include at least one of the following: a first syntax element indicating the alternative initial CPB removal delay increment for the i-th sublayer of the j-th codec picture buffer (CPB) of the NAL hypothetical reference decoder (HRD) in 90kHz clock units; a second syntax element indicating the alternative initial CPB removal offset increment for the i-th sublayer of the j-th CPB of the NAL HRD in 90kHz clock units; a third syntax element indicating, for the i-th sublayer of the NAL HRD, when the AU associated with the PT SEI message directly follows the AU associated with the Buffer Period (BP) SEI message in decoding order, the offset to be used in the derivation of the nominal CPB removal time of the AU associated with the PT SEI message and one or more subsequent AUs in decoding order; a fourth syntax element indicating, for the i-th sublayer of the NAL HRD, when the AU associated with the PT SEI message directly follows the Intra-Frame Random Access Point (IRAP) AU associated with the BP SEI message in decoding order, the offset to be used in the derivation of the nominal CPB removal time of the AU associated with the PT SEI message and one or more subsequent AUs in decoding order; and a fourth syntax element indicating, for the i-th sublayer of the NAL HRD, when the AU associated with the PT SEI message directly follows the Intra-Frame Random Access Point (IRAP) AU associated with the BP SEI message in decoding order, the offset to be used in the derivation of the nominal CPB removal time of the AU associated with the PT SEI message and one or more subsequent AUs in decoding order. The fifth syntax element indicates the alternative initial CPB removal delay increment for the i-th sublayer of the j-th CPB of the VCL HRD, in 90kHz clock units; the sixth syntax element indicates the alternative initial CPB removal offset increment for the i-th sublayer of the j-th CPB of the VCL HRD, in 90kHz clock units; the seventh syntax element indicates the offset to be used in the derivation of the nominal CPB removal time of the AU associated with the PT SEI message and one or more subsequent AUs in decoding order for the i-th sublayer of the VCL HRD; the eighth syntax element indicates the offset to be used in the derivation of the DPB output time of the IRAP AU associated with the BP SEI message for the i-th sublayer of the VCL HRD, when the AU associated with the PT SEI message directly follows the AU associated with the BP SEI message in decoding order; and the eighth syntax element indicates the offset to be used in the derivation of the DPB output time of the IRAP AU associated with the BP SEI message for the i-th sublayer of the VCL HRD, when the AU associated with the PT SEI message directly follows the IRAP AU associated with the BP SEI message in decoding order.

[0296] A5. According to the method of solution A1, in response to each picture in the associated AU being a RASL picture comprising a Video Codec Layer (VCL) Network Abstraction Layer (NAL) unit, each of which is a RASL NUT, a flag equal to zero is provided, wherein the flag indicates the presence of one or more syntax elements associated with timing information in the picture timing SEI message.

[0297] A6. According to the method in solution A5, where the flag is

[0298] pt_cpb_alt_timing_info_present_flag.

[0299] A7. A video processing method comprising performing a conversion between a video and a video bitstream comprising one or more images, wherein the bitstream conforms to a format rule, wherein the format rule permits the use of a random access skip-before (RASL) image as a reference subimage for predicting a juxtaposed RADL image in a RADL image associated with the same fully random access (CRA) image as the RASL image.

[0300] The following is another list of preferred solutions for some of the embodiments.

[0301] B1. A video processing method comprising: performing a conversion between a video and a video bitstream comprising one or more pictures, wherein the bitstream conforms to a format rule, wherein the format rule specifies that the pictures associated with a first flag and the derivation in the decoding process of picture sequence counting are based on a second flag, wherein the picture associated with the first flag is a previous picture in decoding order, the previous picture having (i) a first identifier that is the same as a stripe or picture header referencing a reference picture list syntax structure, (ii) a second flag and a second identifier equal to zero, and (iii) a picture type different from a Random Access Skip Prep (RASL) picture and a Random Access Decodeable Prep (RADL) picture, wherein the first flag indicates the presence of a third flag in the bitstream, wherein the second flag indicates whether the current picture is used as a reference picture, and wherein the third flag is used to determine the value of one or more most significant bits of a picture sequence count value of a long-term reference picture.

[0302] B2. According to the method of solution B1, where the first identifier is the identifier of the layer and the second identifier is the time identifier.

[0303] B3. According to the method of solution B1, where the first identifier is a syntax element and the second identifier is a variable.

[0304] B4. According to the method of any one of solutions B1 to B3, wherein the first flag is delta_poc_msb_cycle_present_flag, the second flag is ph_non_ref_pic_flag, and the third flag is delta_poc_msb_cycle_present_flag, and wherein the first identifier is nuh_layer_id, and the second identifier is TemporalId.

[0305] B5. A video processing method comprising: performing a conversion between a video and a video bitstream comprising one or more images, wherein the bitstream conforms to a format rule, wherein the format rule specifies that variables used to determine the timing of removing a decoding unit (DU) or decoding the DU are access unit (AU) specific and derived based on a flag indicating whether the current image is permitted to be used as a reference image.

[0306] B6. According to the method in solution B5, where the variable is prevNonDiscardableAu and the flag is ph_non_ref_pic_flag.

[0307] B7. Following the approach in solution B6, where ph_non_ref_pic_flag equals 1, it specifies that the current image is never used as a reference image.

[0308] B8. Following the approach in solution B6, where ph_non_ref_pic_flag equals zero, it specifies whether the current image can or cannot be used as a reference image.

[0309] B9. A method of video processing, comprising: performing a conversion between a video and a video bitstream comprising one or more pictures, wherein the bitstream conforms to a format rule, wherein the format rule specifies that a Buffer Period Supplement Enhancement Information (SEI) message and a Picture Timing SEI message are access unit (AU) specific when included in the bitstream, wherein a first variable associated with the Buffer Period SEI message and a second variable associated with the Buffer Period SEI message and the Picture Timing SEI message are derived based on a flag indicating whether the current picture is allowed to be used as a reference picture, wherein the first variable indicates that the access unit includes (i) an identifier equal to zero, and (ii) a picture that is not a Random Access Skip Preceding (RASL) picture or a Random Access Decodable Preceding (RADL) picture and the flag is equal to zero, and wherein the second variable indicates that the current AU is not the first AU in decoding order, and the previous AU in decoding order includes (i) an identifier equal to zero, and (ii) a picture that is not a Random Access Skip Preceding (RASL) picture or a Random Access Decodable Preceding (RADL) picture and the flag is equal to zero.

[0310] B10. According to the method of solution B9, where the identifier is a time identifier.

[0311] B11. According to the method of solution B9, the first variable is notDiscardableAu, the second variable is prevNonDiscardableAu, the flag is ph_non_ref_pic_flag, and the identifier is TemporalId.

[0312] B12. A method of video processing, comprising: performing a conversion between a video and a video bitstream comprising one or more images, wherein the bitstream conforms to a format rule, wherein the format rule specifies that the derivation of a first variable and a second variable associated with a first image and a second image is based on a flag, wherein the first image is a current image and the second image is a previous image in decoding order, the previous image (i) including a first identifier equal to zero, (ii) including a flag equal to zero, and (iii) not a Random Access Skip Precedence (RASL) image or a Random Access Decodeable Precedence (RADL) image, and wherein the first variable and the second variable are respectively the maximum and minimum values ​​of the image sequence counts of each of the following images for which the second identifier is equal to the second identifier of the first image: (i) the first image, (ii) the second image, (iii) one or more short-term reference images referenced by all entries in a reference image list of the first image, and (iv) each image that has been output, wherein the codec image buffer (CPB) removal time of the image is less than the CPB removal time of the first image and the decode image buffer (DPB) output time is greater than or equal to the CP removal time of the first image.

[0313] B13. According to the method of solution B12, where the first variable indicates the maximum value of the image sequence count and the second variable indicates the minimum value of the image sequence count.

[0314] B14. According to the method of solution B12, the flag indicates whether the current image is allowed to be used as a reference image.

[0315] B15. According to the method of solution B12, where the first identifier is a time identifier and the second identifier is a layer identifier.

[0316] B16. The method of any one of solutions B12 to B15, wherein the first variable is maxPicOrderCnt, the second variable is minPicOrderCnt, the first identifier is TemporalId, the second identifier is nuh_layer_id, and the flag is ph_non_ref_pic_flag.

[0317] The following is another list of preferred solutions for some of the embodiments.

[0318] C1. A method of video processing, comprising: performing a conversion between a video and a video bitstream comprising one or more images, wherein the bitstream conforms to a format rule, wherein the format rule specifies that flags and syntax elements are access unit (AU) specific when included in the bitstream, wherein in response to the current AU not being the first AU in the bitstream in decoding order, the flag indicates that the nominal codec picture buffer (CPB) removal time of the current AU is determined relative to (a) the nominal CPB removal time of a previous AU associated with a buffer duration supplemental enhancement information (SEI) message or (b) the nominal CPB removal time of the current AU, and wherein in response to the current AU not being the first AU in the bitstream in decoding order, the syntax element specifies a CPB removal delay increment value relative to the nominal CPB removal time of the current AU.

[0319] C2. According to the method of solution C1, the length of the syntax element is indicated in the syntax structure of the buffered time SEI message.

[0320] C3. According to the method of solution C1, the length of the syntax element is (bp_cpb_removal_delay_length_minus1+1) bits.

[0321] C4. According to the method of any one of solutions C1 to C3, where the flag is bp_concatenation_flag and the syntax element is bp_cpb_removal_delay_delta_minus1.

[0322] C5. A video processing method comprising: performing a conversion between a video and a video bitstream comprising one or more images, the bitstream conforming to a format rule specifying that a plurality of variables and a Picture Timing Supplement Enhancement Information (SEI) message are Access Unit (AU) specific when included in the bitstream, wherein the Picture Timing SEI message comprises a plurality of syntax elements, wherein a first variable of the plurality of variables indicates whether the current AU is associated with a buffered time interval SEI message, wherein a second and a third variable of the plurality of variables are associated with an indication of whether the current AU is an AU for initializing a hypothetical reference decoder (HRD), wherein the first syntax element of the plurality of syntax elements specifies the number of clock ticks to wait before outputting one or more decoded images of the AU from the decoded image buffer (DPB) after the AU is removed from the code-decode image buffer (CPB), wherein the second syntax element of the plurality of syntax elements specifies the number of sub-clock ticks to wait before outputting one or more decoded images of the AU from the DPB after the last decoded unit (DU) in the AU is removed from the CPB, and wherein the third syntax element of the plurality of syntax elements specifies the number of element image time intervals occupied by the one or more decoded images of the current AU for use in displaying a model.

[0323] C6. According to the method in solution C5, the first variable is BpResetFlag, the second variable is CpbRemovalDelayMsb, and the third variable is CpbRemovalDelayVal.

[0324] C7. According to the method of solution C5, the first syntax element is pt_dpb_output_delay, the second syntax element is pt_dpb_output_du_delay, and the third syntax element is pt_display_elemental_periods_minus1.

[0325] C8. A method of video processing, comprising: performing a conversion between a video comprising one or more pictures and a video bitstream, wherein the bitstream conforms to a format rule, wherein the format rule specifies that a syntax element associated with a decoded picture buffer (DPB) is access unit (AU) specific when included in the bitstream, and wherein the syntax element specifies a number of sub-clock beats to wait before outputting one or more decoded pictures of the AU from the DPB after the last decoded unit (DU) in the AU is removed from the codec-decoder-picture buffer (CPB).

[0326] C9. According to the method in solution C8, this syntax element is used to calculate the DPB output time.

[0327] C10. According to the method of solution C8, where the syntax element is dui_dpb_output_du_delay.

[0328] The following is another list of preferred solutions for some of the embodiments.

[0329] D1. A video processing method comprising: performing a conversion between a video comprising one or more pictures and a video bitstream, wherein the bitstream conforms to a format rule, wherein the format rule specifies that a flag is Access Unit (AU) specific when included in the bitstream, wherein the value of the flag is based on whether the associated AU is an Intra-Frame Random Access Point (IRAP) AU or a Progressive Decoding Refresh (GDR) AU, and wherein the value of the flag specifies (i) the presence of a syntax element in a buffer-time supplemental enhancement information (SEI) message and (ii) the presence of alternative timing information in a picture timing SEI message of the current buffer time.

[0330] D2. According to the method of solution D1, the value of the flag is equal to zero in response to the associated AU not being an IRAP AU or a GDR AU.

[0331] D3. According to the method of solution D1, the value of the flag is inferred to be zero in response to the flag not being included in the bitstream.

[0332] D4. According to the method of solution D1, where the value of the flag is 1, it indicates that the syntax element exists in the buffered SEI message.

[0333] D5. According to any one of the solutions D through D4, where the flag is bp_alt_cpb_params_present_flag and the syntax element is bp_use_alt_cpb_params_flag.

[0334] The following is another list of preferred solutions for some of the embodiments.

[0335] E1. A method for video processing, comprising: performing a conversion between a video and a video bitstream comprising one or more images, wherein the bitstream conforms to a format rule, wherein the format rule specifies that the value of a first syntax element is based on a flag and a variable, the flag indicating whether the time distance between the output times of consecutive images of a hypothetical reference decoder (HRD) is constrained, the variable identifying the highest time sublayer to be decoded, and wherein the first syntax element specifies the number of element image time intervals occupied by one or more decoded images of the current AU for use in displaying a model.

[0336] E2. According to the approach of solution E1, the variable identifies the highest time sublayer to be decoded.

[0337] E3. Based on the method of solution E1 or E2, where the variable is Htid.

[0338] E4. According to the approach of solution E1, the flag is a second syntax element included in the output layer set (OLS) timing and HRD parameter syntax structure.

[0339] E5. According to the method of solution E1, the first syntax element is included in the Picture Temporal Supplement Enhancement Information (SEI) message.

[0340] E6. According to the method of any one of solutions E1 to E5, where the flag is fixed_pic_rate_within_cvs_flag, the first syntax element is pt_display_elemental_periods_minus1, and the variable is Htid.

[0341] The following applies to one or more of the aforementioned solutions.

[0342] O1. The method according to any of the aforementioned solutions, wherein the conversion includes decoding video from a bitstream.

[0343] O2. The method according to any of the foregoing solutions, wherein the conversion includes encoding the video into a bitstream.

[0344] O3. A method for storing a bitstream representing a video into a computer-readable recording medium, comprising generating a bitstream from the video according to the method described in any one or more of the foregoing solutions; and storing the bitstream in a computer-readable recording medium.

[0345] O4. A video processing apparatus, comprising a processor configured to implement any one or more of the aforementioned solutions.

[0346] O5. A computer-readable medium having instructions stored thereon that, when executed, cause a processor to perform one or more of the methods described in the foregoing solutions.

[0347] O6. A computer-readable medium storing a bit stream generated according to any one or more of the foregoing solutions.

[0348] O7. A video processing apparatus for storing bitstreams, wherein the video processing apparatus is configured to implement the method described in any one or more of the foregoing solutions.

[0349] The following is another list of preferred solutions for some embodiments.

[0350] P1. A video processing method, comprising: performing a conversion between a video and a video codec representation comprising one or more video images, wherein the codec representation conforms to a format rule, wherein the format rule permits the use of a random access skip-before (RASL) image as a reference sub-image for predicting a juxtaposed RADL image in a RADL image associated with the same fully random access image as the RASL image.

[0351] P2. A video processing method, comprising: performing a conversion between a video and a video codec representation comprising one or more video images, wherein the codec representation conforms to a format rule, wherein the format rule specifies that a picture timing supplementation enhancement information message is access unit specific when included in the codec representation and wherein the corresponding random access skip before (RASL) picture must include a RASL network abstraction layer unit type (NUT).

[0352] P3. The method described in solution P1 or P2, wherein performing the conversion includes parsing and decoding the codec representation to generate video.

[0353] P4. The method described in solution P1 or P2, wherein performing the conversion includes encoding the video into a codec representation.

[0354] P5. A video decoding apparatus, including a processor configured to implement one or more of the methods described in solutions P1 to P4.

[0355] P6. A video encoding apparatus, including a processor configured to implement one or more of the methods described in solutions P1 to P4.

[0356] P7. A computer program product having computer code stored thereon, which, when executed by a processor, causes the processor to implement the method of any one of solutions P1 to P4.

[0357] In this document, the term "video processing" can refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm may be applied during the conversion from a pixel representation of a video to its corresponding bitstream representation, and vice versa. The bitstream representation (or simply, the bitstream) of the current video block may, for example, correspond to bits juxtaposed or scattered at different locations within the bitstream, as defined in the syntax. For example, macroblocks may be encoded based on transform and encoding / decoding error residuals, and may also utilize bits in the header and other fields in the bitstream.

[0358] The disclosed and other solutions, examples, embodiments, modules, and functional operations described in this document can be implemented in digital electronic circuits, or in computer software, firmware, or hardware, including the structures disclosed herein and their structural equivalents, or in combinations thereof. The disclosed and other embodiments can be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer-readable medium for execution by or control of the operation of a data processing apparatus. The computer-readable medium can be a combination of a machine-readable storage device, a machine-readable storage substrate, a memory device, a substance that implements a machine-readable propagating signal, or a combination thereof. The term "data processing apparatus" includes all means, devices, and machines for processing data, including, for example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, the apparatus may include code that creates an execution environment for the computer program in question, for example, code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination thereof. The propagating signal is an artificially generated signal, such as a machine-generated electrical, optical, or electromagnetic signal, which is generated to encode information for transmission to a suitable receiver device.

[0359] Computer programs (also known as programs, software, software applications, scripts, or code) can be written in any programming language, including compiled or interpreted languages, and can be deployed in any form, including as standalone programs or as modules, components, subroutines, or other units suited to a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored as a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), as a single file dedicated to the program in question, or as portions of multiple collaborative files (e.g., a file storing portions of one or more modules, subroutines, or code). Computer programs can be deployed to execute on a single computer or on multiple computers located in one place or distributed across multiple locations and interconnected via a communication network.

[0360] The processes and logic flows described in this document can be executed by one or more programmable processors that execute one or more computer programs to perform functions by manipulating input data and generating output. The processes and logic flows can also be executed by dedicated logic circuitry, and the devices can be implemented as dedicated logic circuitry, such as FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits).

[0361] For example, processors suitable for executing computer programs include both general-purpose and special-purpose microprocessors, as well as any one or more processors in any kind of digital computer. Typically, the processor receives instructions and data from read-only memory or random access memory, or both. The basic components of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include, or be operatively coupled to, one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, to receive data from or transfer data to, or both. However, a computer does not need to have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD-ROMs and DVD-ROMs. The processor and memory may be supplemented or incorporated therein by dedicated logic circuitry.

[0362] While this patent document contains numerous details, these details should not be construed as limiting the scope of any subject matter or claimed content, but rather as descriptions of features characteristic of specific embodiments of a particular art. Certain features described in this patent document within the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments. Furthermore, although features may be described above as functioning in certain combinations, and even initially claimed in this way, one or more features from a claimed combination may be removed from that combination in some cases, and the claimed combination may be for sub-combinations or variations thereof.

[0363] Similarly, although operations are described in a specific order in the accompanying drawings, this should not be construed as requiring these operations to be performed in the specific order or sequence shown, or requiring all illustrated operations to be performed to obtain the desired result. Furthermore, the separation of various system components in the embodiments described in this patent document should not be construed as requiring such separation in all embodiments.

[0364] Only some implementation methods and examples are described, and other implementation methods, enhancements and variations can be made based on the content described and shown in this patent document.

Claims

1. A video processing method, comprising: Perform a conversion between a video containing one or more images and the bitstream of that video. The bitstream conforms to the format rules. The format rules specify that the first and second syntax elements, when included in the bitstream, are access unit (AU) specific. Wherein, when the current AU is not the first AU in the bitstream in decoding order, the first syntax element indicates whether the nominal codec picture buffer (CPB) removal time of the current AU is determined relative to (a) the nominal CPB removal time of the previous AU associated with the Buffer Period (BP) Supplemental Enhancement Information (SEI) message, or relative to (b) the nominal CPB removal time of AUprevNonDiscardableAu, and Wherein, when the current AU is not the first AU in the bitstream in decoding order, the second syntax element specifies a CPB removal delay increment value relative to the nominal CPB removal time of the AU prevNonDiscardableAu.

2. The method of claim 1, wherein, The length of the second syntax element is indicated by the third syntax element in the syntax structure of the BP SEI message.

3. The method of claim 2, wherein, The third syntax element is bp_cpb_removal_delay_length_minus1, and the length of the second syntax element is (bp_cpb_removal_delay_length_minus1 + 1) bits.

4. The method of claim 1, wherein, The first syntax element is bp_concatenation_flag, and the second syntax element is bp_cpb_removal_delay_delta_minus1.

5. The method of claim 1, wherein, The format rules specify multiple variables used in the semantics of the Picture Timing Supplementation and Enhancement Information (SEI) message and multiple syntax elements present in the SEI message. When included in the bitstream, these are specific to the Access Unit (AU). Among these variables, the first variable is the Buffer Period (BP) reset flag, and the value of the first variable is based on whether the current AU is associated with a BP SEI message. Among these variables, the second and third variables are associated with the CPB removal delay, and the values ​​of the second and third variables are based on whether the current AU is the AU that initializes the hypothetical reference decoder (HRD). The fourth syntax element among the plurality of syntax elements specifies the number of clock ticks to wait between the removal of the AU from the codec image buffer (CPB) and the output of one or more decoded images of the AU from the decoded image buffer (DPB). Wherein, when DecodingUnitHrdFlag equals 1, the fifth syntax element among the plurality of syntax elements specifies the number of sub-clock beats to wait between the removal of the last decoding unit (DU) from the CPB and the output of one or more decoded images of the AU from the DPB, and The sixth syntax element among the plurality of syntax elements specifies the number of element image period intervals occupied by one or more decoded images of the current AU for displaying the model.

6. The method of claim 5, wherein, The first variable is BpResetFlag, the second variable is CpbRemovalDelayMsb, the third variable is CpbRemovalDelayVal, the fourth syntax element is pt_dpb_output_delay, the fifth syntax element is pt_dpb_output_du_delay, and the sixth syntax element is pt_display_elemental_periods_minus1.

7. The method of claim 1, wherein, The format rules specify that the seventh syntax element used to calculate the output time of the decoded picture buffer (DPB) included in the decoding unit (DU) information SEI message is access unit (AU) specific when included in the bitstream, and The seventh syntax element specifies the number of sub-clock beats to wait between the removal of the last DU from the CPB and the output of one or more decoded pictures of the AU from the DPB.

8. The method of claim 7, wherein, The seventh syntax element is dui_dpb_output_du_delay.

9. The method of claim 1, wherein, The conversion includes decoding video from a bitstream.

10. The method of claim 1, wherein, The conversion includes encoding the video into a bitstream.

11. An apparatus for processing video data, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to: Perform a conversion between a video containing one or more images and the bitstream of that video. wherein The bitstream conforms to the format rules. The format rules specify that the first and second syntax elements, when included in the bitstream, are access unit (AU) specific. Wherein, when the current AU is not the first AU in the bitstream in decoding order, the first syntax element indicates whether the nominal codec picture buffer (CPB) removal time of the current AU is determined relative to (a) the nominal CPB removal time of the previous AU associated with the Buffer Period (BP) Supplemental Enhancement Information (SEI) message, or relative to (b) the nominal CPB removal time of AUprevNonDiscardableAu, and Wherein, when the current AU is not the first AU in the bitstream in decoding order, the second syntax element specifies a CPB removal delay increment value relative to the nominal CPB removal time of the AU prevNonDiscardableAu.

12. The apparatus of claim 11, wherein, The length of the second syntax element is indicated by the third syntax element in the syntax structure of the BP SEI message. Wherein, the third syntax element is bp_cpb_removal_delay_length_minus1, and the length of the second syntax element is (bp_cpb_removal_delay_length_minus1 + 1) bits, and The first syntax element is bp_concatenation_flag, and the second syntax element is bp_cpb_removal_delay_delta_minus1.

13. The apparatus of claim 11, wherein, The format rules specify multiple variables used in the semantics of the Picture Timing Supplemental Enhancement (SEI) message and multiple syntax elements present in the picture timing SEI message. When included in the bitstream, these are specific to the Access Unit (AU). Among these variables, the first variable is the Buffer Period (BP) reset flag, and the value of the first variable is based on whether the current AU is associated with a BP SEI message. Among these variables, the second and third variables are associated with the CPB removal delay, and the values ​​of the second and third variables are based on whether the current AU is the AU that initializes the hypothetical reference decoder (HRD). The fourth syntax element among the plurality of syntax elements specifies the number of clock ticks to wait between the removal of the AU from the codec image buffer (CPB) and the output of one or more decoded images of the AU from the decoded image buffer (DPB). Wherein, when DecodingUnitHrdFlag equals 1, the fifth syntax element among the plurality of syntax elements specifies the number of sub-clock beats to wait between the removal of the last decoding unit (DU) from the CPB and the output of one or more decoded images of the AU from the DPB, and Wherein, the sixth syntax element among the plurality of syntax elements specifies the number of element image period intervals occupied by one or more decoded images of the current AU for displaying the model, and Wherein, the first variable is BpResetFlag, the second variable is CpbRemovalDelayMsb, the third variable is CpbRemovalDelayVal, the fourth syntax element is pt_dpb_output_delay, the fifth syntax element is pt_dpb_output_du_delay, and the sixth syntax element is pt_display_elemental_periods_minus1.

14. The apparatus of claim 11, wherein, The format rules specify that the seventh syntax element used to calculate the output time of the decoded picture buffer (DPB) included in the decoding unit (DU) information SEI message is access unit (AU) specific when included in the bitstream, and The seventh syntax element specifies the number of sub-clock beats to wait between the removal of the last DU from the CPB and the output of one or more decoded pictures of the AU from the DPB. The seventh syntax element is dui_dpb_output_du_delay.

15. A non-transitory computer-readable storage medium for storing instructions, said instructions causing a processor to: Perform a conversion between a video containing one or more images and the bitstream of that video. wherein The bitstream conforms to the format rules. The format rules specify that the first and second syntax elements, when included in the bitstream, are access unit (AU) specific. Wherein, when the current AU is not the first AU in the bitstream in decoding order, the first syntax element indicates whether the nominal codec picture buffer (CPB) removal time of the current AU is determined relative to (a) the nominal CPB removal time of the previous AU associated with the Buffer Period (BP) Supplemental Enhancement Information (SEI) message, or relative to (b) the nominal CPB removal time of AUprevNonDiscardableAu, and Wherein, when the current AU is not the first AU in the bitstream in decoding order, the second syntax element specifies a CPB removal delay increment value relative to the nominal CPB removal time of the AU prevNonDiscardableAu.

16. The non-transitory computer-readable storage medium of claim 15, wherein, The length of the second syntax element is indicated by the third syntax element in the syntax structure of the BP SEI message. Wherein, the third syntax element is bp_cpb_removal_delay_length_minus1, and the length of the second syntax element is (bp_cpb_removal_delay_length_minus1 + 1) bits, and Wherein, the first syntax element is bp_concatenation_flag, and the second syntax element is bp_cpb_removal_delay_delta_minus1. The format rules specify multiple variables used in the semantics of the Picture Timing Supplemental Enhancement (SEI) message and multiple syntax elements present in the picture timing SEI message. When included in the bitstream, these are specific to the Access Unit (AU). Among these variables, the first variable is the Buffer Period (BP) reset flag, and the value of the first variable is based on whether the current AU is associated with a BP SEI message. Among these variables, the second and third variables are associated with the CPB removal delay, and the values ​​of the second and third variables are based on whether the current AU is the AU that initializes the hypothetical reference decoder (HRD). The fourth syntax element among the plurality of syntax elements specifies the number of clock ticks to wait between the removal of the AU from the codec image buffer (CPB) and the output of one or more decoded images of the AU from the decoded image buffer (DPB). Wherein, when DecodingUnitHrdFlag equals 1, the fifth syntax element among the plurality of syntax elements specifies the number of sub-clock beats to wait between the removal of the last decoding unit (DU) from the CPB and the output of one or more decoded images of the AU from the DPB. Wherein, the sixth syntax element among the plurality of syntax elements specifies the number of element image period intervals occupied by one or more decoded images of the current AU for displaying the model, and Wherein, the first variable is BpResetFlag, the second variable is CpbRemovalDelayMsb, the third variable is CpbRemovalDelayVal, the fourth syntax element is pt_dpb_output_delay, the fifth syntax element is pt_dpb_output_du_delay, and the sixth syntax element is pt_display_elemental_periods_minus1.

17. The non-transitory computer-readable storage medium of claim 15, wherein, The format rules specify that the seventh syntax element used to calculate the output time of the decoded picture buffer (DPB) included in the decoding unit (DU) information SEI message is access unit (AU) specific when included in the bitstream. The seventh syntax element specifies the number of sub-clock beats to wait between the removal of the last DU from the CPB and the output of one or more decoded pictures of the AU from the DPB. The seventh syntax element is dui_dpb_output_du_delay.

18. A non-transitory computer-readable recording medium storing instructions and a bitstream, wherein the instructions, when executed by a processor, implement a method to generate the bitstream, wherein, The method includes: Generate a bitstream of video containing one or more images. The bitstream conforms to the format rules. The format rules specify that the first and second syntax elements, when included in the bitstream, are access unit (AU) specific. Wherein, when the current AU is not the first AU in the bitstream in decoding order, the first syntax element indicates whether the nominal codec picture buffer (CPB) removal time of the current AU is determined relative to (a) the nominal CPB removal time of the previous AU associated with the Buffer Period (BP) Supplemental Enhancement Information (SEI) message, or relative to (b) the nominal CPB removal time of AUprevNonDiscardableAu, and Wherein, when the current AU is not the first AU in the bitstream in decoding order, the second syntax element specifies a CPB removal delay increment value relative to the nominal CPB removal time of the AU prevNonDiscardableAu. 19.The non-transitory computer-readable recording medium of claim 18, wherein, The length of the second syntax element is indicated by the third syntax element in the syntax structure of the BP SEI message. The third syntax element is bp_cpb_removal_delay_length_minus1, and the length of the second syntax element is (bp_cpb_removal_delay_length_minus1 + 1) bits. Wherein, the first syntax element is bp_concatenation_flag, and the second syntax element is bp_cpb_removal_delay_delta_minus1. The format rules specify multiple variables used in the semantics of the Picture Timing Supplemental Enhancement (SEI) message and multiple syntax elements present in the picture timing SEI message. When included in the bitstream, these are specific to the Access Unit (AU). Among these variables, the first variable is the Buffer Period (BP) reset flag, and the value of the first variable is based on whether the current AU is associated with a BP SEI message. Among these variables, the second and third variables are associated with the CPB removal delay, and the values ​​of the second and third variables are based on whether the current AU is the AU that initializes the hypothetical reference decoder (HRD). The fourth syntax element among the plurality of syntax elements specifies the number of clock ticks to wait between the removal of the AU from the codec image buffer (CPB) and the output of one or more decoded images of the AU from the decoded image buffer (DPB). Wherein, when DecodingUnitHrdFlag equals 1, the fifth syntax element among the plurality of syntax elements specifies the number of sub-clock beats to wait between the removal of the last decoding unit (DU) from the CPB and the output of one or more decoded images of the AU from the DPB. Wherein, the sixth syntax element among the plurality of syntax elements specifies the number of element image period intervals occupied by one or more decoded images of the current AU for displaying the model, and Wherein, the first variable is BpResetFlag, the second variable is CpbRemovalDelayMsb, the third variable is CpbRemovalDelayVal, the fourth syntax element is pt_dpb_output_delay, the fifth syntax element is pt_dpb_output_du_delay, and the sixth syntax element is pt_display_elemental_periods_minus1. 20.The non-transitory computer-readable recording medium of claim 18, wherein, The format rules specify that the seventh syntax element used to calculate the output time of the decoded picture buffer (DPB) included in the decoding unit (DU) information SEI message is access unit (AU) specific when included in the bitstream, and The seventh syntax element specifies the number of sub-clock beats to wait between the removal of the last DU from the CPB and the output of one or more decoded pictures of the AU from the DPB. The seventh syntax element is dui_dpb_output_du_delay.

21. A method for storing a video bitstream, comprising: Generate a bitstream of video containing one or more images; as well as The bitstream is stored in a computer-readable recording medium. The bitstream conforms to the format rules. The format rules specify that the first and second syntax elements, when included in the bitstream, are access unit (AU) specific. Wherein, when the current AU is not the first AU in the bitstream in decoding order, the first syntax element indicates whether the nominal codec picture buffer (CPB) removal time of the current AU is determined relative to (a) the nominal CPB removal time of the previous AU associated with the Buffer Period (BP) Supplemental Enhancement Information (SEI) message, or relative to (b) the nominal CPB removal time of AUprevNonDiscardableAu, and Wherein, when the current AU is not the first AU in the bitstream in decoding order, the second syntax element specifies a CPB removal delay increment value relative to the nominal CPB removal time of the AU prevNonDiscardableAu.

22. A video processing method, comprising: Perform a conversion between a video containing one or more images and the bitstream of that video. The bitstream conforms to the format rules. The format rules specify that multiple variables and image timing supplementation enhancement information (SEI) messages, when included in the bitstream, are access unit (AU) specific. The image time-series SEI message includes multiple syntax elements. Among these variables, the first variable indicates whether the current AU is associated with a buffered periodic SEI message. Among these variables, the second and third variables are associated with an indication of whether the current AU is the AU that initializes the hypothetical reference decoder (HRD). The first syntax element among the plurality of syntax elements specifies the number of clock ticks to wait between the removal of the AU from the codec image buffer (CPB) and the output of one or more decoded images of the AU from the decoded image buffer (DPB). Wherein, the second syntax element among the plurality of syntax elements specifies the number of sub-clock beats to wait between the removal of the last decoding unit (DU) from the CPB and the output of one or more decoded pictures of the AU from the DPB, and The third syntax element among the plurality of syntax elements specifies the number of element image period intervals occupied by one or more decoded images of the current AU for displaying the model.

23. The method of claim 22, wherein, The first variable is BpResetFlag, the second variable is CpbRemovalDelayMsb, and the third variable is CpbRemovalDelayVal.

24. The method of claim 22, wherein, The first syntax element is pt_dpb_output_delay, the second syntax element is pt_dpb_output_du_delay, and the third syntax element is pt_display_elemental_periods_minus1.

25. A video processing method, comprising: Perform a conversion between a video containing one or more images and the bitstream of that video. The bitstream conforms to the format rules. The format rules specify that the syntax elements associated with the Decoded Picture Buffer (DPB) are access unit (AU) specific when included in the bitstream, and The syntax element specifies the number of sub-clock beats to wait between the removal of the last decoded unit (DU) from the codec picture buffer (CPB) in the AU and the output of one or more decoded pictures of the AU from the DPB.

26. The method of claim 25, wherein, The syntax elements are used to calculate the DPB output time.

27. The method of claim 25, wherein, The syntax element is dui_dpb_output_du_delay.

28. The method of any one of claims 22-27, wherein, The conversion includes decoding the video from the bitstream.

29. The method of any one of claims 22 to 27, wherein, The conversion includes encoding the video into the bitstream.

30. A method for storing a bitstream representing video to a computer-readable recording medium, comprising: The method according to any one of claims 22 to 27 generates the bitstream from the video; as well as The bitstream is stored in the computer-readable recording medium.

31. A video processing apparatus comprising a processor configured to perform the method of any one of claims 22 to 30.

32. A computer-readable medium having instructions stored thereon, which, when executed, cause a processor to perform the method of any one of claims 22 to 30.