Picture Header Constraints for Multi-Layer Video Coding and Decoding
By enforcing specific constraints on syntax elements in the bitstream format, the issues of incorrect POC derivation, inconsistent end-to-end delay, and discontinuous content layers in VVC are resolved, enhancing the reliability and efficiency of multi-layer video coding.
Patent Information
- Application Number
- CN202180042176.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-06-12
- Filing Date
- 2021-06-11
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2041-06-11
AI Technical Summary
In multi-layer video encoding and decoding, the POC derivation in the prior art is incorrect, the GDR feature is not suitable for low-end to end delay applications, the EOS NAL unit causes inter-layer content discontinuity, and the possible output image in the bitstream may not be available.
By adjusting the constraints of syntax elements, ensuring the correctness of POC derivation, limiting GDR images to be used only for low-end delay applications, specifying that the subsequent pictures of the EOS NAL unit are CLVSS pictures, and requiring at least one output picture in the bitstream to ensure the consistency of the bitstream.
It solves the problems of POC derivation errors, GDR features inapplicable and inter-layer content discontinuity, ensures the correct decoding and output of the bitstream, and is suitable for low-latency applications of multi-layer video encoding and decoding.
Smart Images

Figure CN115918067B_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This application is filed to timely claim the priority and benefit of U.S. Provisional Patent Application No. US 63 / 038,601, filed on June 12, 2020. The entire disclosure of the above application is incorporated by reference as part of the disclosure of this application. Technical Field
[0003] This patent document relates to image and video encoding and decoding. Background Art
[0004] Digital video accounts for the largest bandwidth usage on the Internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth demand for digital video use is expected to continue to grow. Summary of the Invention
[0005] This document discloses techniques for picture header constraints for multi - layer coding and decoding, which video encoders and decoders can use to perform video encoding, decoding, or processing.
[0006] In one example aspect, a video processing method is disclosed. The method includes performing a conversion between a video including one or more pictures and a bitstream of the video, where the bitstream conforms to format rules, where the format rules specify a constraint on a value of a first syntax element, the first syntax element specifying whether a second syntax element exists in a picture header syntax structure of a current picture, and where the second syntax element specifies a value of a most significant bit (MSB) period of a picture order count (POC) of the current picture.
[0007] In another example aspect, a video processing method is disclosed. The method includes performing a conversion between a video including one or more pictures and a bitstream of the video, where the bitstream conforms to format rules, where the format rules specify a derivation of a picture order count (POC) in the absence of a syntax element, and where the syntax element specifies a value of a most significant bit (MSB) period of a POC of the current picture.
[0008] In yet another example aspect, a video processing method is disclosed. The method includes performing a conversion between a video and a bitstream of the video, where the bitstream includes an access unit AU, the access unit AU includes a picture according to a rule, and the rule specifies that a gradual decoding refresh (GDR) picture is not allowed in the bitstream in response to an output order of the AU being different from a decoding order of the AU.
[0009] In yet another example aspect, a video processing method is disclosed. The method includes performing a conversion between a video and a bitstream of the video, wherein according to format rules, the bitstream includes multiple layers in multiple access units (AUs), the multiple access units (AUs) include one or more pictures, and wherein the format rules specify that in response to the presence of a sequence end (EOS) network abstraction layer (NAL) unit of a first layer in a first access unit (AU) in the bitstream, subsequent pictures of each of one or more higher layers of the first layer in AUs after the first AU in the bitstream are codec layer video sequence start (CLVSS) pictures.
[0010] In yet another example aspect, a video processing method is disclosed. The method includes performing a conversion between a video and a bitstream of the video, wherein according to format rules, the bitstream includes multiple layers in multiple access units (AUs), the multiple access units (AUs) include one or more pictures, and wherein the format rules specify that in response to a first picture in a first access unit being a codec layer video sequence start (CLVSS) picture and a second picture being a CLVSS picture, the codec layer video sequence start (CLVSS) picture is a clean random access (CRA) picture or a gradual decoding refresh (GDR) picture.
[0011] In yet another example aspect, a video processing method is disclosed. The method includes performing a conversion between a video including one or more pictures and a bitstream of the video according to rules, wherein the rules specify that the bitstream includes at least a first picture to be output, wherein the first picture is in an output layer, wherein the first picture includes a syntax element equal to one, and wherein the syntax element affects the decoded picture output and removal process associated with a hypothetical reference decoder (HRD).
[0012] In yet another example aspect, a video encoder device is disclosed. The video encoder includes a processor configured to implement the above method.
[0013] In yet another example aspect, a video decoder device is disclosed. The video decoder includes a processor configured to implement the above method.
[0014] In yet another example aspect, a computer-readable medium storing code is disclosed. The code embodies one of the methods described herein in the form of processor-executable code.
[0015] These and other features are described throughout this document. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 is a block diagram showing an example video processing system implementing various techniques disclosed herein.
[0017] Figure 2It is a block diagram of an example hardware platform for video processing.
[0018] Figure 3 It is a block diagram showing an example video codec system that can implement some embodiments of the present disclosure.
[0019] Figure 4 It is a block diagram showing an example encoder that can implement some embodiments of the present disclosure.
[0020] Figure 5 It is a block diagram showing an example decoder that can implement some embodiments of the present disclosure.
[0021] Figures 6 to 11 It shows a flowchart of an example method for video processing. Detailed Description
[0022] The chapter titles are used in this document for ease of understanding and do not limit the applicability of the technologies and embodiments disclosed in each chapter to that chapter only. In addition, the use of H.266 terms in some descriptions is for ease of understanding only and does not limit the scope of the disclosed technologies. Therefore, the technologies described herein are also applicable to other video codec protocols and designs.
[0023] 1. Introduction
[0024] This document relates to video codec technology. Specifically, it is about defining levels and bitstream consistency for video codecs that support both single-layer video coding and decoding and multi-layer video coding and decoding. It can be applied to any video codec standard or non-standard video codec that supports single-layer video coding and decoding and multi-layer video coding and decoding, such as the Versatile Video Coding (VVC) being developed.
[0025] 2. Abbreviations
[0026] APS Adaptive Parameter Set
[0027] AU Access Unit
[0028] AUD Access Unit Delimiter
[0029] AVC Advanced Video Coding
[0030] CLVS Coding and Decoding Layer Video Sequence
[0031] CLVSS Coding and Decoding Layer Video Sequence Start
[0032] CPB Coding and Decoding Picture Buffer
[0033] CRA Clean Random Access
[0034] CTU Coding and Decoding Tree Unit
[0035] CVS Encoded / Decoded Video Sequence
[0036] DCI Decoding Capability Information
[0037] DPB Decoded Picture Buffer
[0038] EOB End of Bitstream
[0039] EOS End of Sequence
[0040] GDR Gradual Decoding Refresh
[0041] HEVC High Efficiency Video Coding
[0042] HRD Hypothetical Reference Decoder
[0043] IDR Instantaneous Decoding Refresh
[0044] ILP Inter-Layer Prediction
[0045] ILRP Inter-Layer Reference Picture
[0046] JEM Joint Exploration Model
[0047] LTRP Long-Term Reference Picture
[0048] MCTS Motion-Constrained Tile Set
[0049] NAL Network Abstraction Layer
[0050] OLS Output Layer Set
[0051] PH Picture Header
[0052] POC Picture Order Count
[0053] PPS Picture Parameter Set
[0054] PTL Profile, Tier, Level
[0055] PU Picture Unit
[0056] RAP Random Access Point
[0057] RBSP Raw Byte Sequence Payload
[0058] SEI Supplementary Enhancement Information
[0059] SLI Sub-Picture Level Information
[0060] SPS Sequence Parameter Set
[0061] STRP Short-Term Reference Picture
[0062] SVC Scalable Video Coding
[0063] VCL Video Coding Layer
[0064] VPS Video Parameter Set
[0065] VTM VVC Test Model
[0066] VUI Video Usability Information
[0067] VVC Versatile Video Coding
[0068] 3. Preliminary Discussion
[0069] Video coding standards have evolved mainly through the development of the well-known ITU-T and ISO / IEC standards. ITU-T produced the H.261 and H.263 standards, ISO / IEC produced the MPEG-1 and MPEG-4 Visual standards, and the two organizations jointly produced the H.262 / MPEG-2 video standard, the H.264 / MPEG-4 Advanced Video Coding (AVC) standard, and the H.265 / HEVC standard. Since H.262, video coding standards have been based on a hybrid video coding structure, which utilizes temporal prediction plus transform coding. To explore future video coding technologies beyond HEVC, the Joint Video Exploration Team (JVET) was jointly established by VCEG and MPEG in 2015. Since then, JVET has adopted many new methods and incorporated them into a reference software called the Joint Exploration Model (JEM). JVET meetings are held quarterly simultaneously. The goal of the new coding standard is to reduce the bitrate by 50% compared to HEVC. The new video coding standard was officially named Versatile Video Coding (VVC) at the JVET meeting in April 2018, and the first version of the VVC Test Model (VTM) was also released at that time. With continuous efforts in VVC standardization, new coding technologies have been adopted into the VVC standard at each JVET meeting. The working draft of VVC and the test model VTM are updated after each meeting. The VVC project now aims to be technically completed (FDIS) at the meeting in July 2020.
[0070] 3.1. Random Access in HEVC and VVC and Its Support
[0071] Random access refers to accessing and decoding a bitstream starting from a picture that is not the first picture of the bitstream in decoding order. To support tuning and channel switching in broadcast / multicast and multiparty video conferencing, searching in local playback and streaming, and stream adaptation in streaming, the bitstream needs to include randomly accessible points that are close together, which are usually intra-coded pictures, but can also be inter-coded pictures (e.g., in the case of gradual decoding refresh).
[0072] HEVC includes signaling of Intra Random Access Point (IRAP) pictures in the NAL unit header via the NAL unit type. Three types of IRAP pictures are supported, namely Instantaneous Decoder Refresh (IDR), Clean Random Access (CRA), and Broken Link Access (BLA) pictures. IDR pictures constrain the inter-picture prediction structure to not reference any pictures prior to the current Group of Pictures (GOP), which is traditionally referred to as a closed GOP random access point. By allowing a particular picture to reference pictures prior to the current GOP, CRA pictures are less restrictive, where in the case of random access, all pictures are discarded. CRA pictures are traditionally referred to as open GOP random access points. BLA pictures typically result from the concatenation of two bitstreams or parts thereof at a CRA picture, e.g., during stream switching. To better enable the system to use IRAP pictures, a total of six different NAL units are defined to signal the attributes of IRAP pictures, which can be used to better match the stream access point types defined in the ISO Base Media File Format (ISOBMFF) [7], which is used for random access support in HTTP-based Dynamic Adaptive Streaming over HTTP (DASH) [8].
[0073] VVC supports three types of IRAP pictures, two types of IDR pictures (one type with or the other type without associated RADL pictures), and one type of CRA picture. These are essentially the same as in HEVC. The BLA picture type in HEVC is not included in VVC, mainly for two reasons: i) The basic functionality of BLA pictures can be achieved by a CRA picture plus a Sequence End NAL unit, the presence of which indicates that subsequent pictures start a new CVS in a single-layer bitstream. ii) During the development of VVC, it was desired to specify fewer NAL unit types than in HEVC, as indicated by using 5 bits instead of 6 bits for the NAL unit type field in the NAL unit header.
[0074] Another key difference between VVC and HEVC in terms of random access support is that GDR is supported in a more standardized way in VVC. In GDR, the decoding of the bitstream can start from an inter-coded picture, and although not the entire picture area can be correctly decoded at the beginning, after multiple pictures, the entire picture area will be correct. AVC and HEVC also support GDR by using recovery point SEI messages to signal GDR random access points and recovery points. In VVC, a new NAL unit type is defined for the indication of GDR pictures, and the recovery point is signaled in the picture header syntax structure. It is allowed for the CVS and the bitstream to start with GDR pictures. This means that it is allowed for the entire bitstream to contain only inter-coded pictures without a single intra-coded picture. The main benefit of specifying GDR support in this way is to provide consistent behavior for GDR. GDR enables the encoder to smooth the bitrate of the bitstream by distributing intra-coded stripes or blocks over multiple pictures, as opposed to intra-coding the entire picture, thus allowing a significant reduction in end-to-end latency, which is considered more important today than before as ultra-low latency applications such as wireless display, online gaming, and drone-based applications become more popular.
[0075] Another GDR-related feature in VVC is virtual boundary signaling. The boundary between the refreshed area (i.e., the correctly decoded area) and the non-refreshed area at the picture between a GDR picture and its recovery point can be signaled as a virtual boundary, and when signaled, loop filtering across the boundary will not be applied, so there will be no decoding mismatch for some samples at or near the boundary. This can be useful when the application determines to display the correctly decoded area during the GDR process.
[0076] IRAP pictures and GDR pictures can be collectively referred to as random access point (RAP) pictures.
[0077] 3.2. Picture Resolution Change within a Sequence
[0078] In AVC and HEVC, the spatial resolution of a picture cannot be changed unless a new sequence with a new SPS starts with an IRAP picture. VVC allows changing the picture resolution within a sequence at positions where IRAP pictures are not encoded, and IRAP pictures are always intra-coded. This feature is sometimes referred to as reference picture resampling (RPR) because it requires resampling of the reference pictures used for inter prediction when the reference pictures have a different resolution from the current picture being decoded.
[0079] The scaling ratio is limited to be greater than or equal to 1 / 2 (2x downsampling from the reference picture to the current picture), and less than or equal to 8 (8x upsampling). Three sets of resampling filters with different frequency cut-offs are defined to handle various scaling ratios between the reference picture and the current picture. The three sets of resampling filters are applied to scaling ratios in the ranges from 1 / 2 to 1 / 1.75, from 1 / 1.75 to 1 / 1.25, and from 1 / 1.25 to 8, respectively. Each set of resampling filters has 16 phases for luminance and 32 phases for chrominance, which is the same as the case of the motion compensation interpolation filter. In fact, the normal MC interpolation process is a special case of the resampling process, where the scaling ratio ranges from 1 / 1.25 to 8. The horizontal and vertical scaling ratios are derived based on the picture width and height and the left, right, top, and bottom scaling offsets defined for the reference picture and the current picture.
[0080] Other aspects of the VVC design that support this feature and are different from HEVC include: i) the picture resolution and the corresponding consistency window signaled in the PPS instead of in the SPS, while the maximum picture resolution is signaled in the SPS. ii) For a single-layer bitstream, each picture storage (the slot in the DPB for storing a decoded picture) occupies the buffer size required to store a decoded image with the maximum picture resolution.
[0081] 3.3. General and Scalable Video Coding (SVC) in VVC
[0082] Scalable Video Coding (SVC, sometimes also referred to as scalability in video coding) refers to video coding that uses a base layer (BL) (sometimes called a reference layer (RL)) and one or more scalable enhancement layers (ELs). In SVC, the base layer can carry video data with a base quality level. One or more enhancement layers can carry additional video data to support, for example, higher spatial, temporal, and / or signal-to-noise ratio (SNR) levels. The enhancement layers can be defined relative to previously encoded layers. For example, the bottom layer can be used as the BL, and the top layer can be used as the EL. An intermediate layer can be used as an EL or an RL, or both. For example, an intermediate layer (e.g., a layer that is neither the lowest nor the highest layer) can be an EL of the layer below it (e.g., the base layer or any intervening enhancement layer) and at the same time be an RL for one or more enhancement layers above it. Similarly, in the Multiview or 3D extension of the HEVC standard, there may be multiple views, and the information of one view can be used to code (e.g., encode or decode) the information of another view (e.g., motion estimation, motion vector prediction, and / or other redundancies).
[0083] In SVC, the parameters used by the encoder or decoder are grouped into parameter sets based on the coding levels where they can be used (e.g., video level, sequence level, picture level, slice level, etc.). For example, the parameters that can be used by one or more coded video sequences in different layers of the bitstream can be included in the Video Parameter Set (VPS), and the parameters that can be used by one or more pictures in the coded video sequence are included in the Sequence Parameter Set (SPS). Similarly, the parameters used by one or more slices in a picture can be included in the Picture Parameter Set (PPS), and other parameters specific to a single slice can be included in the slice header. Similarly, an indication of which (which) parameter set a particular layer uses at a given time can be provided at various coding levels.
[0084] Due to the support for Reference Picture Resampling (RPR) in VVC, it is possible to design the support for bitstreams containing multiple layers without the need for any additional signaling of processing-level coding tools. For example, in VVC, two layers with SD and HD resolutions, since the upsampling required for spatial scalability support can use only the RPR upsampling filter. However, to support scalability, higher-level syntax changes are required (compared to non-scalability support). Scalability support is specified in VVC version 1. Different from the scalability support in any earlier video coding standard, including the extensions of AVC and HEVC, the design of VVC scalability has been made as friendly as possible to the single-layer decoder design. The decoding capabilities of the multi-layer bitstream are specified in a way as if there were only a single layer in the bitstream. For example, the decoding capabilities such as the DPB size are specified in a way independent of the number of layers in the bitstream to be decoded. Basically, a decoder designed for a single-layer bitstream does not need much modification to be able to decode a multi-layer bitstream. Compared with the design of the multi-layer extensions of AVC and HEVC, the HLS aspect has been significantly simplified at the expense of some flexibility. For example, the IRAP AU needs to contain pictures of each layer present in the CVS.
[0085] 3.4. Parameter Sets
[0086] AVC, HEVC, and VVC specify parameter sets. The types of parameter sets include SPS, PPS, APS, and VPS. All of AVC, HEVC, and VVC support SPS and PPS. VPS was introduced starting from HEVC and is included in both HEVC and VVC. APS is not included in AVC or HEVC but is included in the latest VVC draft text.
[0087] The SPS is designed to carry sequence-level header information, and the PPS is designed to carry picture-level header information that does not change frequently. Using the SPS and PPS, there is no need to repeat the information that does not change frequently for each sequence or picture, so redundant signaling of this information can be avoided. In addition, using the SPS and PPS enables out-of-band transmission of important header information, thus not only avoiding the need for redundant transmission but also improving the error recovery ability.
[0088] The VPS is introduced to carry sequence-level header information common to all layers in a multi-layer bitstream.
[0089] The APS is introduced to carry such picture-level or slice-level information that requires a considerable number of bits to be encoded and decoded, can be shared by multiple pictures, and can have many different variations in the sequence.
[0090] 4. Technical Problems Solved by the Disclosed Technical Solutions
[0091] There are the following problems with the latest designs of POC, GDR, EOS, and still picture profiles in VVC:
[0092] 1) When vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] is equal to 0 and there is a picture in the current AU in the reference layer of the current layer, it is required that ph_poc_msb_cycle_present_flag should be equal to 0. However, such a picture in the reference layer can be removed through the general sub-bitstream extraction process specified in Clause C.6. Therefore, the POC derivation will be incorrect.
[0093] 2) The value of ph_poc_msb_cycle_present_flag is used in the POC derivation process, and the flag may not exist and there is no inferred value in this case.
[0094] 3) The GDR feature is mainly used for low end-to-end delay applications. Therefore, when the bitstream is encoded in a way that is not suitable for low end-to-end delay applications, it makes sense not to allow its use.
[0095] 4) When the EOS NAL unit of a layer exists in the AU of a multi-layer bitstream, this will mean that there is a search operation to jump to that AU, or that AU is a bitstream splicing point. For either of these two cases, for the same content, this layer is discontinuous while in another layer of the same bitstream, the content is continuous, which is meaningless regardless of whether there is an inter-layer dependency between the layers.
[0096] 5) The bitstream may not have a picture to output. This should not be allowed, typically for all profiles, or only for the still picture profile.
[0097] 5. List of Examples and Solutions
[0098] To solve the above problems and other problems, methods summarized as follows are disclosed. These items should be regarded as examples for explaining general concepts and should not be interpreted in a narrow sense. In addition, these items can be applied individually or in any combination.
[0099] 1) To solve Problem 1, when vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] is equal to 0 and there is a picture in the current AU in the reference layer of the current layer, instead of requiring ph_poc_msb_cycle_present_flag to be equal to 0, under more strict conditions, the value of ph_poc_msb_cycle_present_flag can be required to be equal to 0.
[0100] a. In one example, when vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] is equal to 0 and there is an ILRP entry in RefPicList[0] or RefPicList[1] of the strip of the current picture, the value of ph_poc_msb_cycle_present_flag is required to be equal to 0.
[0101] b. In one example, when vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] is equal to 0 and there is a picture with nuh_layer_id equal to refpicLayerId in the current AU in the reference layer of the current layer and has a TemporalId less than or equal to Max(0, vps_max_tid_il_ref_pics_plus1[currLayerIdx][refLayerIdx] - 1), the value of ph_poc_msb_cycle_present_flag is required to be equal to 0, where currLayerIdx and refLayerIdx are equal to GeneralLayerIdx[nuh_layer_id] and GeneralLayerIdx[refpicLayerId], respectively.
[0102] c. In one example, the value of ph_poc_msb_cycle_present_flag never needs to be equal to 0.
[0103] 2) To solve Problem 2, instead of using "ph_poc_msb_cycle_present_flag equals 1 (0)" in the POC derivation process, "ph_poc_msb_cycle_val exists (does not exist)" is used.
[0104] 3) To solve Problem 3, assume that GDR pictures are only used for low-end-to-end latency applications, and when the output order of the AU is different from the decoding order, GDR pictures may not be allowed.
[0105] a. In one example, it is required that when sps_gdr_enabled_flag equals 1, the decoding order and the output order of all pictures in the CLVS must be the same. Note that this constraint will also require the decoding order and the output order of the AU to be the same in the multi-layer bitstream, because all pictures within the AU need to be consecutive in the decoding order, and all pictures within the AU have the same output order.
[0106] b. In one example, it is required that when sps_gdr_enabled_flag equals 1 for the SPS referred to by the pictures in the CVS, the decoding order and the output order of all AUs in the CVS should be the same.
[0107] c. In one example, it is required that when sps_gdr_enabled_flag equals 1 for the SPS referred to by the pictures, the decoding order and the output order of all AUs in the bitstream should be the same.
[0108] d. In one example, it is required that when sps_gdr_enabled_flag equals 1 for the SPS existing in the bitstream, the decoding order and the output order of all AUs in the bitstream should be the same.
[0109] e. In one example, it is required that when sps_gdr_enabled_flag equals 1 for the SPS of the bitstream (provided in the bitstream or by external means), the decoding order and the output order of all AUs in the bitstream should be the same.
[0110] 4) To solve Problem 4, when the EOS NAL unit of a layer exists in the AU of a multi-layer bitstream, it is required that the next picture in each layer of all or specific higher layers is a CLVSS picture.
[0111] a. In one example, it is stipulated that when AU auA contains an EOS NAL unit in layer layerA, for each layer layerB that exists in the CVS and has layerA as the reference layer, the first picture in layerB in the AU following auA in decoding order should be a CLVSS picture.
[0112] b. Alternatively, in one example, it is stipulated that when AU auA contains an EOS NAL unit in layer layerA, for each layer layerB that exists in the CVS and is higher than layerA, the first picture in layerB in the AU following auA in decoding order should be a CLVSS picture.
[0113] c. Alternatively, in one example, it is stipulated that when a picture in AU auA is a CLVSS picture and this CLVSS picture is a CRA or GDR picture, for each layer layerA that exists in the CVS, if there is a picture picA of layerA in auA, then picA should be a CLVSS picture; otherwise (if there is no picture of layerA in auA), the first picture of layerA in the AU following auA in decoding order should be a CLVSS picture.
[0114] d. Alternatively, in one example, it is stipulated that when a picture in layer layerB in AU auA is a CLVSS picture and this CLVSS picture is a CRA or GDR picture, for each layer layerA that exists in the CVS and is higher than layerB, if there is a picture picA of layerA in auA, then picA should be a CLVSS picture; otherwise (if there is no picture of layerA in auA), the first picture of layerA in the AU following auA in decoding order should be a CLVSS picture.
[0115] e. Alternatively, in one example, it is stipulated that when a picture in layer layerB in AU auA is a CLVSS picture and this CLVSS picture is a CRA or GDR picture, for each layer layerA that exists in the CVS and has layerB as the reference layer, if there is a picture picA of layerA in auA, then picA should be a CLVSS picture; otherwise (if there is no picture of layerA in auA), the first picture of layerA in the AU following auA in decoding order should be a CLVSS picture.
[0116] f. Alternatively, in one example, it is stipulated that when an EOS NAL unit exists in the AU, for each layer that exists in the CVS, an EOS NAL unit should exist in the AU.
[0117] g. Alternatively, in one example, it is stipulated that when there is an EOS NAL unit in layer layerB in the AU, for each layer in the CVS that is higher than layerB, there should be an EOS NAL unit in the AU.
[0118] h. Alternatively, in one example, it is stipulated that when there is an EOS NAL unit in layer layerB in the AU, for each layer in the CVS that has layerB as the reference layer, there should be an EOS NAL unit in the AU.
[0119] i. Alternatively, in one example, it is stipulated that when the picture in the AU is a CLVSS picture and the CLVSS picture is a CRA or GDR picture, all pictures in the AU should be CLVSS pictures.
[0120] j. Alternatively, in one example, it is stipulated that when the picture in layer layerB in the AU is a CLVSS picture and the CLVSS picture is a CRA or GDR picture, the pictures in the AU in all layers higher than layerB should be CLVSS pictures.
[0121] k. Alternatively, in one example, it is stipulated that when the picture in layer layerB of the AU is a CLVSS picture and the CLVSS picture is a CRA or GDR picture, the pictures in the AU in all layers that have layerB as the reference layer should be CLVSS pictures.
[0122] l. Alternatively, in one example, it is stipulated that when the picture in the AU is a CLVSS picture and the CLVSS picture is a CRA or GDR picture, the AU should have pictures of each layer existing in the CVS, and all pictures in the AU should be CLVSS pictures.
[0123] m. Alternatively, in one example, it is stipulated that when the picture in layer layerB in the AU is a CLVSS picture and the CLVSS picture is a CRA or GDR picture, the AU should have pictures of each layer in the CVS that has layerB as the reference layer, and all pictures in the AU should be CLVSS pictures.
[0124] n. Alternatively, in one example, it is stipulated that when the picture in layer layerB in the AU is a CLVSS picture and the CLVSS picture is a CRA or GDR picture, the AU should have pictures of each layer in the CVS that has layerB as the reference layer, and all pictures in the AU should be CLVSS pictures.
[0125] 5) To solve Problem 5, it is stipulated that the bitstream should have at least one output picture.
[0126] a. In one example, it is stipulated that when the bitstream contains only one picture, the picture shall have a ph_pic_output_flag equal to 1.
[0127] b. In one example, it is stipulated that the bitstream shall have at least one picture in the output layer with a ph_pic_output_flag equal to 1.
[0128] c. In the example, any one of the above constraints is stipulated as part of the definition of one or more still picture profiles, for example, the main 10 still picture profile and the main 4:4:4 10 still picture profile.
[0129] d. In the example, any one of the above constraints is stipulated not to be part of the profile definition, so it applies to any profile.
[0130] 6. Embodiments
[0131] The following are some example embodiments of some inventive aspects summarized in Section 5 above, which can be applied to the VVC specification. The changed text is based on the latest VVC text in JVET-S0152-v5. Most of the relevant parts added or modified are shown in bold, underlined and italic, such as "using A and B", while some deleted parts are shown in italic and enclosed in bold double brackets, such as "based on [[A and]] B".
[0132] 6.1. First Embodiment
[0133] This embodiment is directed to Items 1 and 5 and some of their sub-items.
[0134] 7.4.3.7 Picture Header Structure Semantics ...
[0136] A ph_poc_msb_cycle_present_flag equal to 1 stipulates that the syntax element ph_poc_msb_cycle_val exists in the PH. A ph_poc_msb_cycle_present_flag equal to 0 stipulates that the syntax element ph_poc_msb_cycle_val does not exist in the PH. When vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] is equal to 0 and [[in the current AU in the reference layer of the current layer]] in RefPicList[0] or RefPicList[1] of the stripes of the current picture there is [[a picture]] ILRP entry , the value of ph_poc_msb_cycle_present_flag shall be equal to 0. ...
[0138] The ph_pic_output_flag affects the decoded picture output and removal processes specified in Annex C. When the ph_pic_output_flag is not present, it is inferred to be equal to 1.
[0139] The requirement for bitstream conformance is that there should be at least one picture in the bitstream at the output layer and ph_pic_ output_flag is equal to 1.
[0140] Note 5 – There is no picture in the bitstream with ph_non_ref_pic_flag equal to 1 and ph_pic_output_flag equal to 0. ...
[0142] 8.3.1 Decoding process of picture order count ...
[0144] When when ph_poc_msb_cycle_val does not exist [[ph_poc_msb_cycle_present_flag is equal to 0]] and the current picture is not a CLVSS picture, the variables prevPicOrderCntLsb and prevPicOrderCntMsb are derived as follows:
[0145] – Let prevTid0Pic be the previous picture in decoding order with nuh_layer_id equal to that of the current picture, TemporalId and ph_non_ref_pic_flag both equal to 0, and the previous picture not being a RASL or RADL picture.
[0146] – The variable prevPicOrderCntLsb is set to be equal to the ph_pic_order_cnt_lsb of prevTid0Pic.
[0147] – The variable prevPicOrderCntMsb is set to be equal to the PicOrderCntMsb of prevTid0Pic.
[0148] The variable PicOrderCntMsb of the current picture is derived as follows:
[0149] – If ph_poc_msb_cycle_val exists [[_ph_poc_msb_cycle_present_flag is equal to 1]], PicOrderCntMsb is set to be equal to ph_poc_msb_cycle_val * MaxPicOrderCntLsb.
[0150] – Otherwise ( ph_poc_msb_cycle_val does not existIf [[ph_poc_msb_cycle_present_flag]] is equal to 0 and the current picture is a CLVSS picture, then PicOrderCntMsb is set to be equal to 0. ...
[0152] 7.4.3.3 Sequence parameter set RBSP semantics ...
[0154] When sps_gdr_enabled_flag is equal to 1, it specifies that GDR pictures are enabled and may exist in CLVS. When sps_gdr_enabled_flag is equal to 0, it specifies that GDR pictures are disabled and do not exist in CLVS.
[0155] When sps_gdr_enabled_flag is equal to 1, the decoding order and output order of all pictures in the CLVS should be the same.
[0156] Note – Please note that the above constraint also requires that the decoding order and output order of the AUs containing pictures in the CLVS be the same in the multi-layer bitstream, because all pictures within an AU need to be consecutive in decoding order, and all pictures within an AU have the same output order. ...
[0158] 7.4.3.10 Sequence end RBSP semantics ...
[0160] When AU auA contains an EOS NAL unit in layer layerA, for each layer layerB that exists in the CVS and uses layerA as a reference layer, the first picture in layerB in the AU that is in decoding order after auA should be a CLVSS picture. Figure 1 ...
[0162] Figure 2 FIG. 1000 is a block diagram of an example video processing system 1000 that can implement the various techniques disclosed herein. Various implementations may include some or all of the components in system 1000. System 1000 may include an input 1002 for receiving video content. The video content may be received in a raw or uncompressed format (e.g., 8- or 10-bit multi-component pixel values), or it may be received in a compressed or encoded format. Input 1002 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces (such as Ethernet, passive optical network (PON), etc.) and wireless interfaces (such as Wi-Fi or cellular interfaces).
[0163] System 1000 may include a codec component 1004 that can implement various encoding / decoding or coding methods described in this document. The codec component 1004 may reduce the average bitrate of a video from the input 1002 to the output of the codec component 1004 to produce a coded representation of the video. Thus, encoding / decoding techniques are sometimes referred to as video compression or video transcoding techniques. The output of the codec component 1004 may be stored or transmitted via a connected communication, as represented by component 1006. The stored or communicated bitstream (or coded) representation of the video received at the input 1002 may be used by component 1008 to generate pixel values or a displayable video that is sent to the display interface 1010. The process of generating a user-visible video from the bitstream representation is sometimes referred to as video decompression. Additionally, although some video processing operations are referred to as "encoding / decoding" operations or tools, it should be understood that encoding / decoding tools or operations are used at the encoder, and corresponding decoding tools or operations that invert the results of the encoding / decoding will be performed by the decoder.
[0164] Examples of a peripheral bus interface or a display interface may include Universal Serial Bus (USB), High-Definition Multimedia Interface (HDMI), Displayport, etc. Examples of a storage interface include Serial Advanced Technology Attachment (SATA), PCI, IDE interface, etc. The techniques described in this document may be implemented in various electronic devices, such as mobile phones, laptop computers, smartphones, or other devices capable of digital data processing and / or video display.
[0165] Figures 6 - 9 is a block diagram of a video processing apparatus 2000. The apparatus 2000 may be used to implement one or more of the methods described herein. The apparatus 2000 may be implemented in a smartphone, a tablet computer, a computer, an Internet of Things (IoT) receiver, etc. The apparatus 2000 may include one or more processors 2002, one or more memories 2004, and video processing hardware 2006. The (one or more) processors 2002 may be configured to implement one or more of the methods described in this document (e.g., in Figure 3 ). The (one or more) memories 2004 may be used to store data and code for implementing the methods and techniques described herein. The video processing hardware 2006 may be used to implement some of the techniques described in this document in hardware circuitry. In some embodiments, the hardware 2006 may be partially or fully within one or more of the processors 2002, such as a graphics processor.
[0166] Figure 3 is a block diagram showing an example video codec system 100 that can utilize the techniques of the present disclosure. As Figure 3As shown, the video encoding and decoding system 100 may include a source device 110 and a destination device 120. The source device 110 generates encoded video data, which may be referred to as a video encoding device. The destination device 120 may decode the encoded video data generated by the source device 110, and the destination device 120 may be referred to as a video decoding device. The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.
[0167] The video source 112 may include sources such as a video capture device, an interface for receiving video data from a video content provider, and / or a computer graphics system that generates video data, or a combination of these sources. The video data may include one or more pictures. The video encoder 114 encodes the video data from the video source 112 to generate a bitstream. The bitstream may include a sequence of bits that form a coded representation of the video data. The bitstream may include coded pictures and associated data. A coded picture is a coded representation of a picture. The associated data may include sequence parameter sets, picture parameter sets, and other syntax elements. The I / O interface 116 includes a modulator / demodulator (modem) and / or a transmitter. The encoded video data may be directly sent to the destination device 120 via the I / O interface 116 over the network 130a. The encoded video data may also be stored on a storage medium / server 130b for access by the destination device 120.
[0168] The destination device 120 may include an I / O interface 126, a video decoder 124, and a display device 122.
[0169] The I / O interface 126 may include a receiver and / or a modem. The I / O interface 126 may obtain the encoded video data from the source device 110 or the storage medium / server 130b. The video decoder 124 may decode the encoded video data. The display device 122 may display the decoded video data to a user. The display device 122 may be integrated with the destination device 120 or may be external to the destination device 120 configured to interface with an external display device.
[0170] The video encoder 114 and the video decoder 124 may operate according to video compression standards such as the High Efficiency Video Coding (HEVC) standard, the Versatile Video Coding (VVC) standard, and other current and / or other standards.
[0171] Figure 4 is a block diagram showing an example of a video encoder 200, and the video encoder 200 may be Figure 3 the video encoder 114 in the system 100 shown in
[0172] Video encoder 200 may be configured to perform any or all of the techniques of the present disclosure. In Figure 4 the example, video encoder 200 includes a plurality of functional components. The techniques described in the present disclosure may be shared among the various components of video encoder 200. In some examples, a processor may be configured to perform any or all of the techniques described in the present disclosure.
[0173] The functional components of video encoder 200 may include a segmentation unit 201, a prediction unit 202 (which may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra prediction unit 206), a residual generation unit 207, a transformation unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transformation unit 211, a reconstruction unit 212, a buffer 213, and an entropy coding unit 214.
[0174] In other examples, video encoder 200 may include more, fewer, or different functional components. In one example, prediction unit 202 may include an intra block copy (IBC) unit. The IBC unit may perform prediction in an IBC mode, where at least one reference picture is the picture in which the current video block is located.
[0175] In addition, some components such as motion estimation unit 204 and motion compensation unit 205 may be highly integrated, but are shown separately in the Figure 4 example for purposes of explanation.
[0176] Segmentation unit 201 may segment a picture into one or more video blocks. Video encoder 200 and video decoder 300 may support various video block sizes.
[0177] Mode selection unit 203 may select, for example, one of an intra or inter coding mode based on error results, and provide the resulting intra or inter coded block to residual generation unit 207 to generate residual block data and to reconstruction unit 212 to reconstruct the coded block for use as a reference picture. In some examples, mode selection unit 203 may select a combination of intra and inter prediction (CIIP) mode, where the prediction is based on an inter prediction signal and an intra prediction signal. Mode selection unit 203 may also select the resolution of the motion vector (e.g., sub-pixel or integer-pixel accuracy) for a block in the case of inter prediction.
[0178] To perform inter prediction for a current video block, motion estimation unit 204 may generate motion information for the current video block by comparing one or more reference frames from buffer 213 with the current video block. Motion compensation unit 205 may determine a predicted video block for the current video block based on motion information from a picture in buffer 213 (rather than the picture associated with the current video block) and decoded samples.
[0179] The motion estimation unit 204 and the motion compensation unit 205 can perform different operations for the current video block. For example, the operations performed depend on whether the current video block is in an I-slice, a P-slice, or a B-slice.
[0180] In some examples, the motion estimation unit 204 can perform unidirectional prediction of the current video block, and the motion estimation unit 204 can search for a reference video block of the current video block in the reference pictures of list 0 or list 1. The motion estimation unit 204 can then generate a reference index indicating that the reference picture of list 0 or list 1 contains the reference video block and a motion vector indicating the spatial displacement between the current video block and the reference video block. The motion estimation unit 204 can output the reference index, the prediction direction indicator, and the motion vector as the motion information of the current video block. The motion compensation unit 205 can generate a predicted video block of the current block based on the reference video block indicated by the motion information of the current video block.
[0181] In other examples, the motion estimation unit 204 can perform bidirectional prediction of the current video block. The motion estimation unit 204 can search for a reference video block of the current video block in the reference pictures of list 0 and can also search for another reference video block of the current video block in the reference pictures of list 1. The motion estimation unit 204 can then generate a reference index indicating that the reference picture of list 0 or list 1 contains the reference video block and a motion vector indicating the spatial displacement between the reference video block and the current video block. The motion estimation unit 204 can output the reference index and the motion vector of the current video block as the motion information of the current video block. The motion compensation unit 205 can generate a predicted video block of the current video block based on the reference video block indicated by the motion information of the current video block.
[0182] In some examples, the motion estimation unit 204 can output the entire set of motion information for the decoding process of the decoder.
[0183] In some examples, the motion estimation unit 204 may not output the entire set of motion information of the current video. Instead, the motion estimation unit 204 can signal the motion information of the current video block by referring to the motion information of another video block. For example, the motion estimation unit 204 can determine that the motion information of the current video block is similar enough to the motion information of an adjacent video block.
[0184] In one example, the motion estimation unit 204 can indicate in the syntax structure associated with the current video block: indicating to the video decoder 300 that the current video block has a value of the same motion information as another video block.
[0185] In another example, the motion estimation unit 204 may identify another video block and a motion vector difference (MVD) in the syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector indicating the video block. The video decoder 300 may use the motion vector indicating the video block and the motion vector difference to determine the motion vector of the current video block.
[0186] As discussed above, the video encoder 200 may predictively signal motion vectors. Two examples of predictive signaling techniques that may be implemented by the video encoder 200 include advanced motion vector prediction (AMVP) and merge mode signaling.
[0187] The intra prediction unit 206 may perform intra prediction on the current video block. When the intra prediction unit 206 performs intra prediction on the current video block, the intra prediction unit 206 may generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block may include a predicted video block and various syntax elements.
[0188] The residual generation unit 207 may generate residual data for the current video block by subtracting (e.g., denoted by a minus sign) the (multiple) predicted video blocks of the current video block from the current video block. The residual data for the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.
[0189] In other examples, such as in the skip mode, there may be no residual data for the current video block, and the residual generation unit 207 may not perform the subtraction operation.
[0190] The transform processing unit 208 may generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video block associated with the current video block.
[0191] After the transform processing unit 208 generates the transform coefficient video block associated with the current video block, the quantization unit 209 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0192] The inverse quantization unit 210 and the inverse transform unit 211 may apply inverse quantization and inverse transform to the transform coefficient video block, respectively, to reconstruct the residual video block from the transform coefficient video block. The reconstruction unit 212 may add the reconstructed residual video block to the corresponding samples of one or more predicted video blocks generated by the prediction unit 202 to produce a reconstructed video block associated with the current block for storage in the buffer 213.
[0193] After reconstructing a video block in the reconstruction unit 212, a loop filtering operation may be performed to reduce blockiness artifacts in the video block.
[0194] The entropy coding unit 214 may receive data from other functional components of the video encoder 200. When the entropy coding unit 214 receives data, the entropy coding unit 214 may perform one or more entropy coding operations to generate entropy-coded data and output a bitstream including the entropy-coded data.
[0195] Figure 5 is a block diagram showing an example of a video decoder 300, which may be Figure 3 the video decoder 114 in the system 100 shown in
[0196] The video decoder 300 may be configured to perform any or all of the techniques of the present disclosure. In Figure 5 the example of, the video decoder 300 includes a plurality of functional components. The techniques described in the present disclosure may be shared among various components of the video decoder 300. In some examples, a processor may be configured to perform any or all of the techniques described in the present disclosure.
[0197] In Figure 5 the example of, the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, and a reconstruction unit 306 and a buffer 307. In some examples, the video decoder 300 may perform a decoding process that is generally inverse to the encoding process described with respect to the video encoder 200 ( Figure 4 ).
[0198] The entropy decoding unit 301 may retrieve the encoded bitstream. The encoded bitstream may include entropy-coded video data (e.g., encoded blocks of video data). The entropy decoding unit 301 may decode the entropy-coded video, and based on the entropy-decoded video data, the motion compensation unit 302 may determine motion information including a motion vector, motion vector precision, reference picture list index, and other motion information. The motion compensation unit 302 may determine such information, for example, by performing AMVP and merge mode.
[0199] The motion compensation unit 302 may generate a motion-compensated block, possibly performing interpolation based on an interpolation filter. The identifier of the interpolation filter to be used with sub-pixel precision may be included in a syntax element.
[0200] The motion compensation unit 302 may use the interpolation filter used by the video encoder 200 during the encoding of a video block to calculate interpolated values for sub-integer pixels of a reference block. The motion compensation unit 302 may determine the interpolation filter used by the video encoder 200 based on received syntax information and use the interpolation filter to generate a prediction block.
[0201] The motion compensation unit 302 may use some syntax information to determine: the size of the blocks used to encode the (multiple) frames and / or (multiple) slices of the encoded video sequence, the partitioning information that describes how each macroblock of a picture of the encoded video sequence is partitioned, the mode indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-coded block, and other information for decoding the encoded video sequence.
[0202] The intra prediction unit 303 may form a prediction block from spatially adjacent blocks using, for example, the intra prediction mode received in the bitstream. The inverse quantization unit 303 inverse quantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 301. The inverse transform unit 303 applies an inverse transform.
[0203] The reconstruction unit 306 may sum the corresponding prediction block generated by the motion compensation unit 202 or the intra prediction unit 303 and the residual block to form a decoded block. As desired, the deblocking filter may also be applied to filter the decoded block to remove block effect artifacts. The decoded video block is then stored in the buffer 307, which provides reference blocks for subsequent motion compensation / intra prediction and also generates the decoded video for presentation on a display device.
[0204] Figures 6 - 11 Shows an example method for implementing the above technical solution in an embodiment such as Figures 1 - 5 shown.
[0205] Figure 6 A flowchart showing an example method 600 for video processing is shown. Method 600 includes, at operation 610, performing a conversion between a video including one or more pictures and a bitstream of the video, the bitstream conforming to format rules that specify constraints on the values of a first syntax element that specifies whether a second syntax element exists in the picture header syntax structure of the current picture, and the second syntax element specifies the value of the most significant bit (MSB) period of the picture order count (POC) of the current picture.
[0206] Figure 7A flowchart showing an example method 700 for video processing. Method 700 includes, at operation 710, performing a conversion between a video including one or more pictures and a bitstream of the video, the bitstream conforming to format rules that specify the derivation of a picture order count (POC) in the absence of a syntax element, where the syntax element specifies the value of a most significant bit (MSB) period of the POC of the current picture.
[0207] Figure 8 A flowchart showing an example method 800 for video processing. Method 800 includes, at operation 810, performing a conversion between a video and a bitstream of the video according to rules, the bitstream including access units (AUs), each access unit including a picture, the rules specifying that in response to the output order of an AU being different from its decoding order, a gracefully degraded refresh (GDR) picture is not allowed in the bitstream.
[0208] Figure 9 A flowchart showing an example method 900 for video processing. Method 900 includes, at operation 910, performing a conversion between a video and a bitstream of the video according to format rules, the bitstream including multiple layers in multiple access units (AUs), the multiple access units including one or more pictures, the format rules specifying that in response to the presence of an end-of-sequence (EOS) network abstraction layer (NAL) unit of a first layer in a first access unit (AU) in the bitstream, each subsequent picture in one or more higher layers of the first layer in AUs after the first AU in the bitstream is a coded layer video sequence start (CLVSS) picture.
[0209] Figure 10 A flowchart showing an example method 1000 for video processing. Method 1000 includes, at operation 1010, performing a conversion between a video and a bitstream of the video according to format rules, the bitstream including multiple layers in multiple access units (AUs), the multiple access units including one or more pictures, the format rules specifying that in response to a first picture in a first access unit being a coded layer video sequence start (CLVSS) picture and a second picture being a CLVSS picture, the coded layer video sequence start (CLVSS) picture is a clean random access (CRA) picture or a gracefully degraded refresh (GDR) picture.
[0210] Figure 11 A flowchart showing an example method 1100 for video processing. Method 1100 includes, at operation 1110, performing a conversion between a video including one or more pictures and a bitstream of the video according to rules, the rules specifying that the bitstream includes at least a first picture output, the first picture being in an output layer, the first picture including a syntax element equal to one, and the syntax element affecting the decoded picture output and removal process associated with a hypothetical reference decoder (HRD).
[0211] The following solutions show example embodiments of the techniques discussed in the previous section (e.g., Items 1-5).
[0212] Next, a list of preferred solutions for some embodiments is provided.
[0213] A1. A method for video processing, comprising: performing a conversion between a video including one or more pictures and a bitstream of the video, wherein the bitstream conforms to format rules, wherein the format rules specify a constraint on a value of a first syntax element, the first syntax element specifying whether a second syntax element exists in a picture header syntax structure of a current picture, and wherein the second syntax element specifies a value of a most significant bit (MSB) period of a picture order count (POC) of the current picture.
[0214] A2. The method according to Solution A1, wherein in response to a value of a flag being equal to zero and an inter-layer reference picture (ILRP) entry being in a reference picture list of a strip of the current picture, the value of the first syntax element is equal to zero, and wherein the flag specifies whether an index layer uses inter-layer prediction.
[0215] A3. The method according to Solution A2, wherein the reference picture list includes a first reference picture list (RefPicList[0]) or a second reference picture list (RefPicList[1]).
[0216] A4. The method according to Solution A2, wherein a value of the first syntax element being equal to zero specifies that the second syntax element does not exist in the picture header syntax structure.
[0217] A5. The method according to Solution A2, wherein a value of the flag being equal to zero specifies that the index layer is allowed to use inter-layer prediction.
[0218] A6. The method according to Solution A1, wherein in response to a value of the flag being equal to zero and the picture having (i) a first identifier equal to a second identifier in a current access unit (AU) in a reference layer of the current layer, and (ii) a third identifier less than or equal to a threshold, wherein the flag specifies whether an index layer uses inter-layer prediction, wherein the first identifier specifies a layer to which a video coding layer (VCL) network abstraction layer (NAL) unit belongs, wherein the second identifier specifies a layer to which a reference picture belongs, wherein the third identifier is a temporal identifier, and wherein the threshold is based on the second syntax element, the second syntax element specifying whether a picture in the index layer that is neither an intra random access picture (IRAP) picture nor a gradual decoding refresh (GDR) picture is used as an inter-layer reference picture (IRLP) for decoding a picture in the index layer.
[0219] A7. The method according to solution A6, wherein the first identifier is nuh_layer_id, the second identifier is refpicLayerId, and the third identifier is TemporalId, and wherein the second syntax element is vps_max_tid_il_ref_pics_plus1.
[0220] A8. The method according to solution A1, wherein the first syntax element never needs to be zero.
[0221] A9. The method according to any one of solutions A2 to A8, wherein the first syntax element is ph_poc_msb_cycle_present_flag, the flag is vps_independent_layer_flag, and wherein the second syntax element is ph_poc_msb_cycle_val.
[0222] A10. A method for video processing, comprising: performing a conversion between a video including one or more pictures and a bitstream of the video, wherein the bitstream conforms to format rules, wherein the format rules specify the derivation of the picture order count (POC) in the absence of a syntax element, and wherein the syntax element specifies the value of the most significant bit (MSB) cycle of the POC of the current picture.
[0223] A11. The method according to solution A10, wherein the syntax element is ph_poc_msb_cycle_val.
[0224] A12. A method for video processing, comprising: performing a conversion between a video and a bitstream of the video according to rules, wherein the bitstream includes access units (AUs), and each access unit (AU) includes a picture, and wherein the rules specify that in response to the output order of an AU being different from its decoding order, a gradual decoding refresh (GDR) picture is not allowed in the bitstream.
[0225] A13. The method according to solution A12, wherein in response to the flag being equal to one, the output order and the decoding order of all pictures in the codec layer video sequence (CLVS) are the same, and wherein the flag specifies whether a GDR picture is enabled.
[0226] A14. The method according to solution A12, wherein in response to the flag in the sequence parameter set (SPS) referred to by a picture in the codec video sequence (CVS) being equal to one, the output order and the decoding order of the AU are the same, and wherein the flag specifies whether a GDR picture is enabled.
[0227] A15. The method according to solution A12, wherein in response to the flag of the sequence parameter set (SPS) referred to by the picture being equal to one, the output order and the decoding order of the AU are the same, and wherein the flag specifies whether the GDR picture is enabled.
[0228] A16. The method according to solution A12, wherein in response to the flag of the sequence parameter set (SPS) in the bitstream being equal to one, the output order and the decoding order of the AU are the same, and wherein the flag specifies whether the GDR picture is enabled.
[0229] A17. The method according to any one of solutions A13 to A16, wherein the flag is sps_gdr_enabled_flag.
[0230] Next, another list of preferred solutions of some embodiments is provided.
[0231] B1. A method for video processing, comprising: performing conversion between a video and a bitstream of the video according to format rules, wherein the bitstream includes multiple layers in multiple access units (AUs), the multiple access units (AUs) include one or more pictures, and wherein the format rules specify that in response to the presence of an end-of-sequence (EOS) network abstraction layer (NAL) unit of the first layer in the first access unit (AU) in the bitstream, each subsequent picture in one or more higher layers of the first layer in the AUs after the first AU in the bitstream is a codec layer video sequence start (CLVSS) picture.
[0232] B2. The method according to solution B1, wherein the format rules further specify that for the second layer using the first layer as a reference layer, the first picture in decoding order is a CLVSS picture, and the second layer is present in a codec video sequence (CVS) including the first layer.
[0233] B3. The method according to solution B1, wherein one or more higher layers include all or specific higher layers.
[0234] B4. The method according to solution B1, wherein the format rules further specify that for the second layer being a layer higher than the first layer, the first picture in decoding order is a CLVSS picture, and the second layer is present in a codec video sequence (CVS) including the first layer.
[0235] B5. The method according to solution B1, wherein the format rules further specify that the EOS NAL unit is present in each layer of the codec video sequence (CVS) in the bitstream.
[0236] B6. The method according to solution B1, wherein the formatting rule further specifies that a second layer, which is higher than the first layer, includes an EOS NAL unit, and the second layer is present in a coded video sequence (CVS) including the first layer.
[0237] B7. The method according to solution B1, wherein the formatting rule further specifies that a second layer using the first layer as a reference layer includes an EOS NAL unit, and the second layer is present in a coded video sequence (CVS) including the first layer.
[0238] B8. A method for video processing, comprising: performing conversion between a video and a bitstream of the video according to a formatting rule, wherein the bitstream includes multiple layers in multiple access units (AUs), the multiple access units (AUs) include one or more pictures, and the formatting rule specifies that in response to a first picture in a first access unit being a coded layer video sequence start (CLVSS) picture, a second picture being a CLVSS picture, and the coded layer video sequence start (CLVSS) picture being a clean random access (CRA) picture or a gradual decoding refresh (GDR) picture.
[0239] B9. The method according to solution B8, wherein the second picture is a picture for the layer in the first access unit.
[0240] B10. The method according to solution B8, wherein the first layer includes a first picture, and the second picture is a picture in a second layer that is higher than the first layer.
[0241] B11. The method according to solution B8, wherein the first layer includes a first picture, and the second picture is a picture in a second layer that uses the first layer as a reference layer.
[0242] B12. The method according to solution B8, wherein the second picture is the first picture in decoding order in a second access unit after the first access unit.
[0243] B13. The method according to solution B8, wherein the second picture is any picture in the first access unit.
[0244] B14. The method according to any one of solutions B1 to B13, wherein the CLVSS picture is a coded picture having a flag equal to one, the coded picture being an instantaneous random access point (IRAP) picture or a GDR picture, and wherein the flag equal to one indicates that when determining that an associated picture includes a reference to a picture that does not exist in the bitstream, the decoder does not output the associated picture.
[0245] Next, another list of preferred solutions for some embodiments is provided.
[0246] C1. A method for video processing, comprising: performing a conversion between a video including one or more pictures and a bitstream of the video according to a rule, wherein the rule stipulates that the bitstream includes at least a first picture to be output, wherein the first picture is in an output layer, wherein the first picture includes a syntax element equal to one, and wherein the syntax element affects a decoded picture output and removal process associated with a Hypothetical Reference Decoder (HRD).
[0247] C2. The method according to solution C1, wherein the rule applies to all profiles and the bitstream is allowed to conform to any profile.
[0248] C3. The method according to solution C2, wherein the syntax element is ph_pic_output_flag.
[0249] C4. The method according to solution C2, wherein the profile is the main 10 still picture profile or the main 4:4:4 10 still picture profile.
[0250] The following list of solutions applies to each of the solutions listed above.
[0251] O1. The method according to any of the foregoing solutions, wherein the conversion includes decoding the video from the bitstream.
[0252] O2. The method according to any of the foregoing solutions, wherein the conversion includes encoding the video into a bitstream.
[0253] O3. A method for storing a bitstream representing a video in a computer-readable recording medium, comprising generating a bitstream from the video according to the method described in any one or more of the foregoing solutions, and storing the bitstream in the computer-readable recording medium.
[0254] O4. A video processing apparatus, comprising a processor configured to implement the method described in any one or more of the foregoing solutions.
[0255] O5. A computer-readable medium having instructions stored thereon, which when executed cause the processor to implement the method described in one or more of the foregoing solutions.
[0256] O6. A computer-readable medium storing a bitstream generated according to any one or more of the foregoing solutions.
[0257] O7. A video processing apparatus for storing a bitstream, wherein the video processing apparatus is configured to implement the method described in any one or more of the foregoing solutions.
[0258] Next, another list of preferred solutions for some embodiments is provided.
[0259] P1. A video processing method includes performing a conversion between a video including one or more pictures and an encoded / decoded representation of the video, wherein the encoded / decoded representation complies with format rules, and the format rules specify constraints on values of syntax elements that indicate the presence of the most significant bit period of the picture order count in pictures of the video.
[0260] P2. The method according to solution P1, wherein the format rules specify that when an independent value flag is set to a zero value and at least one strip of a picture uses an inter-layer reference picture in its reference list, the value of the syntax element is 0.
[0261] P3. The method according to any one of solutions P1 to P2, wherein the format rules specify that a zero value of a syntax element is indicated by not including the syntax element in the encoded / decoded representation.
[0262] P4. A video processing method includes performing a conversion between a video including one or more pictures and an encoded / decoded representation of the video, wherein the conversion complies with specified rules that do not allow progressive decoding refresh pictures in cases where the output order of access units is different from the decoding order of access units.
[0263] P5. A video processing method includes performing a conversion between a video including a video layer containing one or more video pictures and an encoded / decoded representation of the video, wherein the encoded / decoded representation complies with format rules, and the format rules specify that if a first Network Abstraction Layer unit (NAL) indicating the end of a video sequence exists in the access units of the layer, each next picture of each higher layer in the encoded / decoded representation must have an encoded / decoded layer video sequence start type.
[0264] P6. The method according to solution P5, wherein the format rules further specify that the first picture in decoding order of a second layer using the layer as a reference layer should have an encoded / decoded layer video sequence start type.
[0265] P7. The method according to any one of solutions P1 to P5, wherein performing the conversion includes encoding the video to generate the encoded / decoded representation.
[0266] P8. The method according to any one of solutions P1 to P5, wherein performing the conversion includes parsing and decoding the encoded / decoded representation to generate the video.
[0267] P9. A video decoding device includes a processor configured to implement the method described in one or more of solutions P1 to P8.
[0268] P10. A video encoding device includes a processor configured to implement the method described in one or more of solutions P1 to P8.
[0269] P11. A computer program product having computer code stored thereon which, when executed by a processor, causes the processor to implement the method described in any one of Solutions P1 to P8.
[0270] In this document, the term "video processing" may refer to video encoding, video decoding, video compression or video decompression. For example, during the conversion from the pixel representation of a video to the corresponding bitstream representation, a video compression algorithm may be applied, and vice versa. As defined by the syntax, the bitstream representation (or simply bitstream) of the current video block may (for example) correspond to bits co-located or scattered at different positions within the bitstream. For example, a macroblock may be encoded based on the transform and codec error residual values and also using bits in the header and other fields in the bitstream.
[0271] The disclosures and other solutions, examples, embodiments, modules and functional operations described in this document may be implemented in digital electronic circuits or in computer software, firmware or hardware, including the structures disclosed herein and their structural equivalents, or a combination of one or more of them. The disclosed and other embodiments may be implemented as one or more computer program products encoded on a computer-readable medium, i.e., one or more computer program instruction modules, for execution by, or to control the operation of, a data processing apparatus. The computer-readable medium may be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition affecting a machine-readable propagated signal, or a combination of one or more of them. The term "data processing apparatus" encompasses all apparatus, devices and machines for processing data, including, for example, programmable processors, computers or multiple processors or computers. In addition to the hardware, the apparatus may also include code that creates an execution environment for the computer program being discussed, for example, code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them. A propagated signal is an artificially generated signal, for example, a machine-generated electrical, optical or electromagnetic signal which is generated to encode information for transmission to a suitable receiver device.
[0272] A computer program (also known as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, and can be deployed in any form, including as a stand-alone program or as modules, components, subroutines, or other units suitable for use in a computing environment. A computer program does not have to correspond to a file in a file system. The program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple co-operating files (e.g., files that store one or more modules, subroutines, or portions of code). A computer program can be deployed to execute on one computer or on multiple computers, which may be located at one site or distributed across multiple sites and interconnected by a communication network.
[0273] The processes and logical flows described in this document can be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logical flows can also be performed by, and apparatus can be implemented as, special purpose logic circuitry, e.g., a field programmable gate array (FPGA) or an application specific integrated circuit (ASIC).
[0274] Processors suitable for the execution of a computer program include, by way of example, both general and special purpose microprocessors, and any one or more processors of any type of digital computer. Generally, a processor will receive instructions and data from a read only memory or a random access memory or both. The essential elements of a computer are a processor for performing the instructions and one or more memory devices for storing the instructions and data. Generally, a computer will also include one or more mass storage devices for storing data (e.g., magnetic, magneto-optical disks, or optical disks), or operatively coupled to receive data from or transfer data to a mass storage device (e.g., magnetic, magneto-optical disks, or optical disks), or both. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including by way of example semiconductor memory devices (e.g., EPROM, EEPROM, and flash memory devices); magnetic disks (e.g., internal hard disks or removable disks); magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.
[0275] Although this patent document contains many details, these details should not be construed as limitations on any subject matter or the scope of what can be claimed, but rather as descriptions of features of particular embodiments of a particular technology. In this patent document, certain features that are described in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, the various features described in the context of a single embodiment may also be implemented separately in multiple embodiments or in various suitable sub-combinations. In addition, although features may be described as acting in certain combinations and even initially claimed as such, in some cases one or more features from a claimed combination may be excluded from the combination, and the claimed combination may be directed to a sub-combination or variation of a sub-combination.
[0276] Similarly, although operations are depicted in the drawings in a particular order, this should not be construed as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, to achieve a desired result. In addition, the separation of various system components in the embodiments described in this patent document should not be construed as requiring such separation in all embodiments.
[0277] Only a few implementations and examples have been described, and other implementations, enhancements, and variations can be made based on what is described and illustrated in this patent document.
Claims
1. A method for video processing, comprising: Performing a conversion between a video including one or more pictures and a bitstream of the video, wherein the bitstream conforms to format rules, wherein the format rules specify constraints on the value of a first syntax element that specifies whether a second syntax element exists in a picture header syntax structure of a current picture, and wherein the second syntax element specifies the value of the most significant bit (MSB) period of the picture order count (POC) of the current picture, wherein in response to a flag value being equal to zero and an inter-layer reference picture (ILRP) entry being in a reference picture list of a slice of the current picture, the value of the first syntax element is equal to zero, and wherein the flag specifies whether an index layer uses inter-layer prediction.
2. The method according to claim 1, wherein The reference picture list includes a first reference picture list RefPicList[0] or a second reference picture list RefPicList[1].
3. The method according to claim 1, wherein, The value of the first syntax element being equal to zero specifies that the second syntax element does not exist in the picture header syntax structure.
4. The method according to claim 1, wherein The flag value being equal to zero specifies that the index layer is allowed to use the inter-layer prediction.
5. The method according to claim 1, wherein, The first syntax element is ph_poc_msb_cycle_present_flag, and wherein the second syntax element is ph_poc_msb_cycle_val.
6. The method according to claim 1, wherein The flag is vps_independent_layer_flag.
7. The method according to claim 1, wherein, The conversion is performed according to rules, wherein the bitstream includes access units (AUs), and each AU includes a picture, wherein the rules specify that in response to the output order of an AU being different from its decoding order, a gradual decoding refresh (GDR) picture is not allowed in the bitstream.
8. The method according to claim 7, wherein, In response to a first flag being equal to one, the output order and the decoding order of all pictures in a codec layer video sequence (CLVS) are the same, and wherein the first flag specifies whether a GDR picture is enabled.
9. The method according to claim 7, wherein In response to a second flag of a sequence parameter set (SPS) referred to by a picture in a coded video sequence (CVS) being equal to one, the output order and the decoding order of the AU are the same, and wherein the second flag specifies whether a GDR picture is enabled.
10. The method according to claim 7, wherein, In response to a third flag of a sequence parameter set (SPS) referred to by a picture being equal to one, the output order and the decoding order of the AU are the same, and wherein the third flag specifies whether a GDR picture is enabled.
11. The method according to claim 7, wherein, In response to a fourth flag of a sequence parameter set (SPS) in the bitstream being equal to one, the output order and the decoding order of the AU are the same, and wherein the fourth flag specifies whether a GDR picture is enabled.
12. The method according to claim 8, wherein, The first flag is sps_gdr_enabled_flag.
13. The method according to any one of claims 1-12, wherein, The conversion includes decoding the video from the bitstream.
14. The method according to any one of claims 1-12, wherein, The conversion includes encoding the video into the bitstream.
15. An apparatus for processing video data, comprising a processor and a non-transitory memory having instructions thereon, wherein, The instructions, when executed by the processor, cause the processor to: Perform a conversion between a video including one or more pictures and a bitstream of the video, wherein the bitstream conforms to format rules, Wherein, the format rule stipulates a constraint on the value of a first syntax element, the first syntax element stipulating whether a second syntax element exists in the picture header syntax structure of a current picture, and wherein, the second syntax element stipulates the value of the most significant bit (MSB) cycle of the picture order count (POC) of the current picture, wherein, in response to the value of a flag being equal to zero and an inter-layer reference picture (ILRP) entry being in the reference picture list of the strip of the current picture, the value of the first syntax element is equal to zero, and wherein the flag stipulates whether an index layer uses inter-layer prediction.
16. The device according to claim 15, wherein, The reference picture list includes a first reference picture list RefPicList[0] or a second reference picture list RefPicList[1].
17. The apparatus according to claim 15, wherein, The value of the first syntax element being equal to zero stipulates that the second syntax element does not exist in the picture header syntax structure.
18. The device according to claim 15, wherein The value of the flag being equal to zero stipulates that the index layer is allowed to use the inter-layer prediction.
19. The device according to claim 15, wherein The first syntax element is ph_poc_msb_cycle_present_flag, and wherein, the second syntax element is ph_poc_msb_cycle_val.
20. The device according to claim 15, wherein, The flag is vps_independent_layer_flag.
21. The device according to any one of claims 15 - 20, wherein, The conversion includes decoding the video from the bitstream.
22. The device according to any one of claims 15-20, wherein, The conversion includes encoding the video into the bitstream.
23. A non-transitory computer-readable storage medium storing instructions that cause a processor to: perform a conversion between a video including one or more pictures and a bitstream of the video, Among them, the bitstream conforming to format rules, wherein, the format rule stipulates a constraint on the value of a first syntax element, the first syntax element stipulating whether a second syntax element exists in the picture header syntax structure of a current picture, and wherein, the second syntax element stipulates the value of the most significant bit (MSB) cycle of the picture order count (POC) of the current picture, wherein, in response to the value of a flag being equal to zero and an inter-layer reference picture (ILRP) entry being in the reference picture list of the strip of the current picture, the value of the first syntax element is equal to zero, and wherein the flag stipulates whether an index layer uses inter-layer prediction.
24. A non-transitory computer-readable recording medium storing a bitstream of a video generated by a method executed by a video processing apparatus, wherein, The method includes: generating the bitstream of the video, wherein, the bitstream conforms to format rules, wherein, the format rule stipulates a constraint on the value of a first syntax element, the first syntax element stipulating whether a second syntax element exists in the picture header syntax structure of a current picture, and wherein, the second syntax element stipulates the value of the most significant bit (MSB) cycle of the picture order count (POC) of the current picture, wherein, in response to the value of a flag being equal to zero and an inter-layer reference picture (ILRP) entry being in the reference picture list of the strip of the current picture, the value of the first syntax element is equal to zero, and wherein the flag stipulates whether an index layer uses inter-layer prediction.
25. A method of storing a bitstream of a video, including: generating the bitstream of the video according to format rules, and storing the bitstream into a non-transitory computer-readable recording medium, wherein the format rule specifies a constraint on a value of a first syntax element that specifies whether a second syntax element exists in a picture header syntax structure of a current picture, and wherein the second syntax element specifies a value of a most significant bit (MSB) period of a picture order count (POC) of the current picture, wherein in response to a value of a flag being equal to zero and an inter-layer reference picture (ILRP) entry being in a reference picture list of a slice of the current picture, the value of the first syntax element is equal to zero, and wherein the flag specifies whether an index layer uses inter-layer prediction.
26. A method of storing a bitstream representing a video in a computer-readable recording medium, comprising: generating the bitstream from the video according to the method of any one of claims 1-14; and storing the bitstream in the computer-readable recording medium.
27. A video processing apparatus, comprising a processor configured to implement the method of any one of claims 1-14.
28. A computer-readable medium having instructions stored thereon, which when executed cause a processor to implement the method of any one of claims 1-14.
Citation Information
Patent Citations
Coding method, coding device, decoding method and decoding device for picture order count, and electronic equipment
CN104754347A
Method for encoding / decoding image and device using same
CN105765978A
Cited By
Constraints on picture output ordering in a video bitstream
US12574555B2