Constraints on picture output order in video bitstreams
By modifying the VVC specification, issues such as POC derivation errors, inapplicability of GDR features, and discontinuity of content between layers were resolved, ensuring that output images are present in the bitstream. This enabled low-end latency applications and inter-layer content continuity, thereby improving video encoding and decoding efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- DOUYIN CO LTD
- Filing Date
- 2021-06-11
- Publication Date
- 2026-08-04
AI Technical Summary
In VVC, the design of POC, GDR, EOS and still image levels has problems such as incorrect POC derivation, GDR features not being suitable for low-latency applications, discontinuity of content between layers, and the possibility of no image output in the bitstream.
By modifying the VVC specification, specifying the POC derivation conditions, restricting the use of GDR images, ensuring the continuity of images within the AU and in the bitstream, requiring the existence of EOS NAL units, and ensuring that there is at least one output image in the bitstream.
The POC derivation error was resolved, ensuring that GDR is suitable for low-latency applications, achieving continuity of content between layers and output of images in the bitstream, and improving the efficiency and consistency of video encoding and decoding.
Smart Images

Figure CN115885512B_ABST
Abstract
Description
[0001] Cross-reference of related applications
[0002] This application is filed to promptly claim priority and benefit to U.S. Provisional Patent Application No. 63 / 038,601, filed June 12, 2020. The entire disclosure of the aforementioned application is incorporated herein by reference as a part of the disclosure of this application. Technical Field
[0003] The patent document relates to image and video encoding and decoding. Background Technology
[0004] Digital video accounts for the largest share of bandwidth usage on the Internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth demand for digital video is expected to continue to grow. Summary of the Invention
[0005] This document discloses techniques for constraining the output order of images in a video bitstream, which video codecs and decoders can use to perform video encoding, decoding, or processing.
[0006] In one example aspect, a video processing method is disclosed. The method includes performing a conversion between a video comprising one or more images and a bitstream of the video, wherein the bitstream conforms to a format rule, wherein the format rule specifies constraints on the value of a first syntax element, the first syntax element specifying the presence of a second syntax element in the image header syntax structure of the current image, and wherein the second syntax element specifying the value of the most significant bit (MSB) period of the image sequence count (POC) of the current image.
[0007] In another example, a video processing method is disclosed. The method includes performing a conversion between a video comprising one or more images and a video bitstream, wherein the bitstream conforms to a format rule specifying the derivation of the Picture Order Count (POC) in the absence of syntax elements, and wherein the syntax elements specify the value of the most significant bit (MSB) period of the POC for the current image.
[0008] In another example, a video processing method is disclosed. The method includes performing a conversion between a video and a video bitstream, wherein the bitstream includes access units (AUs), and the access units (AUs) include pictures according to rules, wherein the rules specify that the output order of the AUs is different from the decoding order of the AUs, and step-by-step decoding refresh (GDR) pictures are not allowed in the bitstream.
[0009] In another example, a video processing method is disclosed. The method includes performing a conversion between a video and a video bitstream, wherein, according to a format rule, the bitstream comprises multiple layers in a plurality of access units (AUs), each AU comprising one or more pictures, wherein the format rule specifies that, in response to the presence of a first-layer End-of-Sequence (EOS) Network Abstraction Layer (NAL) unit in a first access unit (AU) of the bitstream, a subsequent picture of each of one or more higher layers in the first layer of AUs following the first AU in the bitstream is a codec layer Video Sequence Start (CLVSS) picture.
[0010] In another example, a video processing method is disclosed. The method includes performing a conversion between a video and a bitstream of the video, wherein, according to a format rule, the bitstream comprises multiple layers in a plurality of access units (AUs), each AU comprising one or more pictures, wherein the format rule specifies that, in response to a first picture in a first access unit, a codec layer video sequence start (CLVSS) picture, a second picture is a CLVSS picture, and the codec layer video sequence start (CLVSS) picture is either a clean random access (CRA) picture or a progressive decode refresh (GDR) picture.
[0011] In another example, a video processing method is disclosed. This method includes performing a conversion between a video and a video bitstream comprising one or more images according to rules, wherein the rules specify that the bitstream includes at least a first output image, wherein the first image is in the output layer, wherein the first image includes a syntax element equal to one, and wherein the syntax element affects the decoded image output and removal process associated with a hypothetical reference decoder (HRD).
[0012] In yet another example, a video encoder apparatus is disclosed. The video encoder includes a processor configured to implement the methods described above.
[0013] In yet another example, a video decoder apparatus is disclosed. The video decoder includes a processor configured to implement the methods described above.
[0014] In yet another example, a computer-readable medium on which code is stored is disclosed. This code embodies one of the methods described herein in the form of processor-executable code.
[0015] These and other features are described throughout this document. Attached Figure Description
[0016] Figure 1 This is a block diagram illustrating an example video processing system that implements the various techniques disclosed herein.
[0017] Figure 2This is a block diagram of an example hardware platform used for video processing.
[0018] Figure 3 This is a block diagram illustrating an example video codec system that can implement some embodiments of the present disclosure.
[0019] Figure 4 This is a block diagram illustrating an example encoder that can implement some embodiments of the present disclosure.
[0020] Figure 5 This is a block diagram illustrating examples of decoders that can implement some embodiments of the present disclosure.
[0021] Figures 6 to 11 A flowchart of an example method for video processing is shown. Detailed Implementation
[0022] The use of chapter headings in this document is for ease of understanding and does not limit the applicability of the techniques and embodiments disclosed in each chapter to that chapter only. Furthermore, the use of H.266 terminology in some descriptions is merely for ease of understanding and not to limit the scope of the disclosed techniques. Therefore, the techniques described herein are also applicable to other video codec protocols and designs.
[0023] 1. Introduction
[0024] This document relates to video codec technology. Specifically, it concerns defining levels and bitstream consistency for video codecs that support both single-layer and multi-layer video codecs. It can be applied to any standard or non-standard video codec that supports both single-layer and multi-layer video codecs, such as the Versatile Video Codec (VVC) currently under development.
[0025] 2. Abbreviation
[0026] APS Adaptive Parameter Set
[0027] AU Access Unit
[0028] AUD Access Unit Separator
[0029] AVC Advanced Video Codec
[0030] CLVS codec layer video sequence
[0031] CLVSS codec layer video sequence start
[0032] CPB image buffer
[0033] CRA Clean Random Access
[0034] CTU (Codec Tree Unit)
[0035] CVS codec video sequence
[0036] DCI decoding capability information
[0037] DPB Decoding Image Buffer
[0038] EOB bitstream end
[0039] EOS sequence ends
[0040] GDR Gradual Decoding and Refresh
[0041] HEVC High-Efficiency Video Encoding and Decoding
[0042] HRD Assumption Reference Decoder
[0043] IDR Instant Decoding and Refresh
[0044] Interlayer prediction in ILP
[0045] ILRP interlayer reference image
[0046] JEM Joint Exploration Model
[0047] LTRP Long-Term Reference Image
[0048] MCTS Motion Constraint Pieces
[0049] NAL Network Abstraction Layer
[0050] OLS Output Layer Set
[0051] PH image header
[0052] POC Image Sequential Counting
[0053] PPS Image Parameter Set
[0054] PTL (Level, Grade, Class)
[0055] PU Image Unit
[0056] RAP Random Access Point
[0057] RBSP raw byte sequence payload
[0058] SEI Supplemental Enhancement Information
[0059] SLI sub-image level information
[0060] SPS Sequence Parameter Set
[0061] STRP Short-Term Reference Image
[0062] SVC Scalable Video Codec
[0063] VCL (Video Codec Layer)
[0064] VPS Video Parameter Set
[0065] VTM VVC Test Model
[0066] VUI Video Availability Information
[0067] VVC Multi-Functional Video Encoding and Decoding
[0068] 3. Preliminary Discussion
[0069] Video codec standards have primarily evolved through the development of well-known ITU-T and ISO / IEC standards. ITU-T developed the H.261 and H.263 standards, while ISO / IEC developed the MPEG-1 and MPEG-4 Visual standards. The two organizations jointly developed the H.262 / MPEG-2 video standard, the H.264 / MPEG-4 Advanced Video Codec (AVC) standard, and the H.265 / HEVC standard. Starting with H.262, video codec standards are based on a hybrid video codec architecture, utilizing temporal prediction plus transform coding. To explore future video codec technologies beyond HEVC, the Joint Video Exploration Team (JVET) was jointly established by VCEG and MPEG in 2015. Since then, JVET has adopted many new methods and incorporated them into reference software called the Joint Exploration Model (JEM). JVET meetings are held quarterly, and the goal of new codec standards is to reduce the bitrate by 50% compared to HEVC. The new video codec standard was officially named Multifunctional Video Coding (VVC) at the JVET meeting in April 2018, and the first version of the VVC Test Model (VTM) was also released at that time. With ongoing efforts to standardize VVC, new codec technologies have been adopted into the VVC standard at each JVET meeting. The VVC working draft and test model VTM are updated after each meeting. The current goal of the VVC project is to achieve Technical Completion (FDIS) at the meeting in July 2020.
[0070] 3.1. Random Access and its Support in HEVC and VVC
[0071] Random access refers to accessing and decoding the bitstream starting from the first image that is not in the decoding order of the bitstream. To support tuning and channel switching in broadcast / multicast and multi-party video conferencing, searching in local playback and streaming, and stream adaptation in streaming, the bitstream needs to include closely spaced random access points, which are typically intra-frame codec images, but can also be inter-frame codec images (e.g., in the case of progressive decoding refresh).
[0072] HEVC includes signaling notification of Intra-Random Access Point (IRAP) pictures in the NAL unit header via the NAL unit type. Three types of IRAP pictures are supported: Instant Decoder Refresh (IDR), Clean Random Access (CRA), and Broken Link Access (BLA) pictures. IDR pictures constrain the inter-frame picture prediction structure to not reference any pictures preceding the current group of pictures (GOP), traditionally referred to as closed GOP random access points. CRA pictures are less restrictive by allowing specific pictures to reference pictures preceding the current GOP, where all pictures are discarded in the case of random access. CRA pictures are traditionally referred to as open GOP random access points. BLA pictures are typically derived from the concatenation of two bitstreams or a portion thereof at the CRA picture, for example, during stream switching. To better enable the system to use IRAP pictures, a total of six different NAL units are defined to signal the properties of IRAP pictures. This can be used to better match the stream access point type defined in the ISO Basic Media File Format (ISOBMFF) [7], which is used for random access support in HTTP-based Dynamic Adaptive Streaming (DASH) [8].
[0073] VVC supports three types of IRAP pictures, two types of IDR pictures (one type has an associated RADL picture, and the other does not), and one type of CRA picture. These are essentially the same as in HEVC. The BLA picture type from HEVC is not included in VVC, primarily for two reasons: i) the basic functionality of a BLA picture can be achieved by adding a sequence end NAL unit to a CRA picture, the presence of which indicates that a new CVS begins in a single-layer bitstream. ii) during the development of VVC, it was desirable to specify fewer NAL unit types than in HEVC, as indicated by using 5 bits instead of 6 bits for the NAL unit type field in the NAL unit header.
[0074] Another key difference between VVC and HEVC in random access support is that VVC supports GDR in a more canonical way. In GDR, decoding of a bitstream can begin with an inter-frame codec picture, and although not the entire picture region may be correctly decoded at the beginning, it will be correct after several pictures. Recovery point SEI messages are used to signal GDR random access points and recovery points. AVC and HEVC also support GDR. In VVC, a new NAL unit type is specified for indicating GDR pictures, and recovery points are signaled in the picture header syntax structure. CVS and bitstreams are allowed to begin with GDR pictures. This means that an entire bitstream can contain only inter-frame codec pictures, without any single intra-frame codec picture. The main benefit of specifying GDR support in this way is providing consistent behavior for GDR. GDR enables encoders to smooth the bitrate of a bitstream by distributing intra-frame encoded stripes or blocks across multiple images, as opposed to intra-frame encoding and decoding the entire image. This allows for significant end-to-end latency reduction, which is considered more important today than ever as ultra-low latency applications such as wireless displays, online gaming, and drone-based applications become more popular.
[0075] Another GDR-related feature in VVC is virtual boundary signaling notification. The boundary between the refreshed area (i.e., the correctly decoded area) and the unrefreshed area at the GDR image and its recovery point can be signaled as a virtual boundary. When signaled, loop filtering across the boundary will not be applied, thus preventing decoding mismatches at or near the boundary. This can be useful when applications determine which area is correctly decoded during the GDR process.
[0076] IRAP images and GDR images can be collectively referred to as random access point (RAP) images.
[0077] 3.2. Image resolution variations within a sequence
[0078] In AVC and HEVC, the spatial resolution of an image cannot be changed unless a new sequence with a new SPS begins with an IRAP image. VVC allows changing the image resolution within a sequence at locations where IRAP images are not encoded; IRAP images are always intra-frame encoded and decoded. This feature is sometimes called Reference Image Resampling (RPR) because it requires resampling the reference image used for inter-frame prediction when the reference image has a different resolution than the current image being decoded.
[0079] The scaling ratio is limited to greater than or equal to 1 / 2 (2x downsampling from the reference image to the current image) and less than or equal to 8 (8x upsampling). Three sets of resampling filters with different frequency cutoffs are specified to handle various scaling ratios between the reference and current images. The three sets of resampling filters are applied to scaling ratios ranging from 1 / 2 to 1 / 1.75, from 1 / 1.75 to 1 / 1.25, and from 1 / 1.25 to 8, respectively. Each set of resampling filters has 16 phases for luma and 32 phases for chroma, similar to the case for motion-compensated interpolation filters. In fact, the normal MC interpolation process is a special case of the resampling process where the scaling ratio ranges from 1 / 1.25 to 8. The horizontal and vertical scaling ratios are derived based on the image width and height, as well as the left, right, top, and bottom scaling offsets specified for the reference and current images.
[0080] Other aspects of VVC designs that support this feature, unlike HEVC, include: i) signaling the image resolution and corresponding consistency window in the PPS instead of the SPS, while in the SPS the signaling indicates the maximum image resolution. ii) for a single-layer bitstream, each image storage (the slot in the DPB used to store one decoded image) occupies the buffer size required to store the decoded image with the maximum image resolution.
[0081] 3.3. Scalable Video Codec (SVC) in General and VVC
[0082] Scalable video codec (SVC, sometimes also called scalability in video codec) refers to video codec using a base layer (BL) (sometimes called a reference layer (RL)) and one or more scalable enhancement layers (EL). In SVC, the base layer can carry video data with a base quality level. One or more enhancement layers can carry additional video data to support, for example, higher spatial, temporal, and / or signal-to-noise ratio (SNR) levels. Enhancement layers can be defined relative to previously encoded layers. For example, the bottom layer can be used as a BL, while the top layer can be used as an EL. Intermediate layers can be used as ELs or RLs, or both. For example, an intermediate layer (e.g., a layer that is neither the lowest nor the highest layer) can be an EL of a layer below the intermediate layer (e.g., a base layer or any intermediate enhancement layer) and simultaneously used as an RL of one or more enhancement layers above the intermediate layer. Similarly, in the Multiview or 3D extension of the HEVC standard, multiple views may exist, and information from one view can be used to codec (e.g., encode or decode) information from another view (e.g., motion estimation, motion vector prediction, and / or other redundancy).
[0083] In SVC, the parameters used by the encoder or decoder are grouped into parameter sets based on the codec level in which they can be used (e.g., video level, sequence level, picture level, stripe level, etc.). For example, parameters that can be used by one or more codec video sequences at different layers in a bitstream can be included in the Video Parameter Set (VPS), and parameters that can be used by one or more pictures in a codec video sequence can be included in the Sequence Parameter Set (SPS). Similarly, parameters used by one or more stripes in a picture can be included in the Picture Parameter Set (PPS), and additional parameters specific to a single strip can be included in the stripe header. Likewise, indications of which parameter set(s) a particular layer uses at a given time can be provided at various codec levels.
[0084] Because of VVC's support for Reference Picture Resampling (RPR), support for multi-layered bitstreams can be designed without requiring any additional signaling notification processing level codec tools. For example, two layers in VVC with SD and HD resolutions can be supported because the upsampling required for spatial scalability can be achieved using only RPR upsampling filters. However, supporting scalability requires a higher level of syntax changes (compared to not supporting scalability at all). Scalability support was specified in VVC version 1. Unlike scalability support in any earlier video codec standards, including extensions to AVC and HEVC, VVC scalability was designed to be as friendly as possible to single-layer decoder designs. The decoding capability of multi-layered bitstreams is specified as if there were only a single layer in the bitstream. For example, decoding capabilities such as DPB size are specified in a way that is independent of the number of layers in the bitstream to be decoded. Essentially, decoders designed for single-layered bitstreams do not require many changes to decode multi-layered bitstreams. Compared to the multi-layered extensions of AVC and HEVC, the HLS aspect is significantly simplified at the expense of some flexibility. For example, IRAP AU requires a picture of every layer present in CVS.
[0085] 3.4. Parameter Set
[0086] AVC, HEVC, and VVC specify parameter sets. Parameter set types include SPS, PPS, APS, and VPS. All AVC, HEVC, and VVC versions support SPS and PPS. VPS was introduced with HEVC and is included in both HEVC and VVC. APS is not included in AVC or HEVC, but is included in the latest VVC draft text.
[0087] SPS is designed to carry sequence-level header information, and PPS is designed to carry infrequently changing image-level header information. Using SPS and PPS eliminates the need to repeat infrequently changing information for each sequence or image, thus avoiding redundant signaling notifications. Furthermore, using SPS and PPS enables out-of-band transmission of critical header information, thereby not only avoiding redundant transmission but also improving error recovery capabilities.
[0088] The purpose of introducing a VPS is to carry sequence-level header information common to all layers in a multi-layer bitstream.
[0089] The purpose of APS is to carry such image-level or stripe-level information, which requires a considerable number of bits for encoding and decoding, can be shared by multiple images, and can have many different variations in the sequence.
[0090] 4. The technical problem solved by the disclosed technical solution
[0091] The latest design of POC, GDR, EOS, and still image formats in VVC has the following problems:
[0092] 1) When vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] equals 0 and an image exists in the current AU of the current layer's reference layer, ph_poc_msb_cycle_present_flag is required to be equal to 0. However, such images in the reference layer can be removed using the general sub-bitstream extraction procedure specified in Clause C.6. Therefore, the POC derivation would be incorrect.
[0093] 2) The value of ph_poc_msb_cycle_present_flag is used in the POC derivation process, but the flag may not exist and in this case there is no deduced value.
[0094] 3) The GDR feature is primarily used for low-end-to-end latency applications. Therefore, it makes sense to disallow its use when the bitstream is encoded in a manner unsuitable for low-end-to-end latency applications.
[0095] 4) When the EOS NAL unit of a layer exists in an AU of a multi-layer bitstream, this means that there is a search operation that jumps to that AU, or that the AU is a bitstream splicing point. In either of these cases, it is meaningless that the content is discontinuous in this layer but continuous in another layer of the same bitstream, regardless of whether there are inter-layer dependencies.
[0096] 5) The bitstream may not contain the images to be output. This should not be allowed, usually for all resolutions, or only for still image resolutions.
[0097] 5. List of Implementation Examples and Solutions
[0098] To address the aforementioned and other issues, a methodology summarized below is presented. These items should be considered as examples for explaining general concepts and should not be interpreted in a narrow sense. Furthermore, these items can be applied individually or in combination in any way.
[0099] 1) To solve problem 1, when vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] equals 0 and there is an image in the current AU of the current layer's reference layer, it is not required that ph_poc_msb_cycle_present_flag equals 0. Instead, under stricter conditions, the value of ph_poc_msb_cycle_present_flag can be required to be equal to 0.
[0100] a. In one example, the value of ph_poc_msb_cycle_present_flag is required to be 0 when vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] is equal to 0 and an ILRP entry exists in RefPicList[0] or RefPicList[1] of the current image's stripe.
[0101] b. In one example, when vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] equals 0 and there exists an image in the current AU of the reference layer of the current layer with nuh_layer_id equal to refpicLayerId and with a TemporalId less than or equal to Max(0,vps_max_tid_il_ref_pics_plus1[currLayerIdx][refLayerIdx]-1), the value of ph_poc_msb_cycle_present_flag is required to be equal to 0, where currLayerIdx and refLayerIdx are equal to GeneralLayerIdx[nuh_layer_id] and GeneralLayerIdx[refpicLayerId], respectively.
[0102] c. In one example, the value of ph_poc_msb_cycle_present_flag never needs to be equal to 0.
[0103] 2) In order to solve problem 2, instead of using "ph_poc_msb_cycle_present_flag equals 1 (0)" in the POC derivation process, use "ph_poc_msb_cycle_val exists (does not exist)".
[0104] 3) To address problem 3, assume that GDR images are only used for low-end-to-end latency applications, and that GDR images may not be allowed when the output order of the AU is different from the decoding order.
[0105] a. In one example, it is required that when `sps_gdr_enabled_flag` equals 1, the decoding order and output order of all images in CLVS must be the same. Note that this constraint also requires that the decoding order and output order of AUs be the same across multiple bitstreams, because all images within an AU need to be consecutive in the decoding order, and all images within an AU need to have the same output order.
[0106] b. In one example, it is required that when sps_gdr_enabled_flag equals 1 for the SPS referenced by the image in CVS, the decoding order and output order of all AUs in CVS should be the same.
[0107] c. In one example, it is required that when sps_gdr_enabled_flag equals 1 for the SPS referenced by the image, the decoding order and output order of all AUs in the bitstream should be the same.
[0108] d. In one example, it is required that when sps_gdr_enabled_flag equals 1 for SPS present in the bitstream, the decoding order and output order of all AUs in the bitstream should be the same.
[0109] e. In one example, it is required that when sps_gdr_enabled_flag is equal to 1 for the SPS of the bitstream (provided by the bitstream itself or by an external means), the decoding order and output order of all AUs in the bitstream should be the same.
[0110] 4) To solve problem 4, when the EOS NAL cell of a layer exists in the AU of a multi-layer bitstream, it is required that the next picture in each layer of all or a specific higher layer is a CLVSS picture.
[0111] a. In one example, it is specified that when AU auA contains EOS NAL units in layer A, for each layer B that exists in CVS and takes layer A as the reference layer, the first image in layer B in the AU following auA in the decoding order should be the CLVSS image.
[0112] b. Alternatively, in one example, it is specified that when AU auA contains EOS NAL units in layer A, for each layer B in the CVS that is higher than layer A, the first image in layer B in the AU following auA in the decoding order should be the CLVSS image.
[0113] c. Alternatively, in one example, it is stipulated that when an image in AU auA is a CLVSS image, which is a CRA or GDR image, for each layer A present in CVS, if there is an image picA of layer A in auA, then picA should be a CLVSS image; otherwise (there is no image of layer A in auA), the first image of layer A in the AU following auA in the decoding order should be a CLVSS image.
[0114] d. Alternatively, in one example, it is stipulated that when the image in layer B of AU auA is a CLVSS image, which is a CRA or GDR image, for each layer A above layer B in CVS, if there is an image picA of layer A in auA, then picA should be a CLVSS image; otherwise (there is no image of layer A in auA), the first image of layer A in the AU following auA in the decoding order should be a CLVSS image.
[0115] e. Alternatively, in one example, it is stipulated that when the image in layer B in AU auA is a CLVSS image, which is a CRA or GDR image, for each layer A existing in the CVS that uses layer B as the reference layer, if there is an image picA of layer A in auA, then picA should be a CLVSS image; otherwise (there is no image of layer A in auA), the first image of layer A in the AU following auA in the decoding order should be a CLVSS image.
[0116] f. Alternatively, in one example, it is specified that when an EOS NAL cell exists in the AU, an EOS NAL cell should exist in the AU for each layer present in the CVS.
[0117] g. Alternatively, in one example, it is specified that when there is an EOS NAL cell in layer B of the AU, there should be an EOS NAL cell in the AU for every layer above layer B in the CVS.
[0118] h. Alternatively, in one example, it is specified that when an EOS NAL cell exists in layer B of the AU, an EOS NAL cell should exist in the AU for each layer in the CVS that uses layer B as the reference layer.
[0119] i. Alternatively, in one example, it is stipulated that when the image in the AU is a CLVSS image, and that CLVSS image is a CRA or GDR image, all images in the AU should be CLVSS images.
[0120] j. Alternatively, in one example, it is stipulated that when the image in layer B of an AU is a CLVSS image, and that CLVSS image is a CRA or GDR image, the images in the AUs of all layers above layer B should be CLVSS images.
[0121] k. Alternatively, in one example, it is specified that when the image of layer B in an AU is a CLVSS image, and that CLVSS image is a CRA or GDR image, the images in the AUs of all layers with layer B as the reference layer should be CLVSS images.
[0122] l. Alternatively, in one example, it is specified that when the images in the AU are CLVSS images, which are CRA or GDR images, the AU should have images of every layer present in the CVS, and all images in the AU should be CLVSS images.
[0123] m. Alternatively, in one example, it is specified that when the image in layer B of an AU is a CLVSS image, which is a CRA or GDR image, the AU should have images of every layer present in the CVS that takes layer B as the reference layer, and all images in the AU should be CLVSS images.
[0124] n. Alternatively, in one example, it is specified that when the image in layer B of an AU is a CLVSS image, which is a CRA or GDR image, the AU should have images of each layer in the CVS that uses layer B as the reference layer, and all images in the AU should be CLVSS images.
[0125] 5) To solve problem 5, it is stipulated that the bitstream should have at least one output image.
[0126] a. In one example, it is specified that when the bitstream contains only one image, the image should have a ph_pic_output_flag equal to 1.
[0127] b. In one example, it is specified that the bitstream should have at least one image in the output layer with ph_pic_output_flag equal to 1.
[0128] c. In the example, any of the above constraints is specified as part of the definition of one or more still image grades, such as main 10 still image grade and main 4:4:4 10 still image grade.
[0129] d. In the example, none of the above constraints are specified as being part of the grade definition, therefore it applies to any grade.
[0130] 6. Example
[0131] Below are some example embodiments of the inventions summarized in Chapter 5 above, which can be applied to the VVC specification. The modified text is based on the latest VVC text in JVET-S0152-v5. Most of the relevant parts that have been added or modified are shown in bold, underlined, and italic, such as "Using A..." ", while some deleted parts are shown in italics and enclosed in bold double brackets, for example, "based on B".
[0132] 6.1. First Embodiment
[0133] This embodiment focuses on Projects 1 and 5 and some of their sub-projects.
[0134] 7.4.3.7 Image header structure and semantics ...
[0136] A flag of ph_poc_msb_cycle_present_flag equal to 1 indicates that the syntax element ph_poc_msb_cycle_val exists in the PH. A flag of ph_poc_msb_cycle_present_flag equal to 0 indicates that the syntax element ph_poc_msb_cycle_val does not exist in the PH. When vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] equals 0 and [[in the current AU of the reference layer of the current layer]]... [Image exists] The value of ph_poc_msb_cycle_present_flag should be equal to 0. ...
[0138] The `ph_pic_output_flag` affects the decoded image output and removal process specified in Appendix C. If `ph_pic_output_flag` does not exist, it is inferred to be equal to 1.
[0139]
[0140] Note 5 – There are no images in the bitstream where ph_non_ref_pic_flag is equal to 1 and ph_pic_output_flag is equal to 0. ...
[0142] 8.3.1 Decoding process of image sequential counting ...
[0144] when Furthermore, the current image is not a CLVSS image. The deduced variables prevPicOrderCntLsb and prevPicOrderCntMsb are as follows:
[0145] – Let prevTid0Pic be the previous image in the decoding order, whose nuh_layer_id is equal to the nuh_layer_id of the current image, TemporalId and ph_non_ref_pic_flag are both equal to 0, and the previous image is not a RASL or RADL image.
[0146] The variable prevPicOrderCntLsb is set to equal ph_pic_order_cnt_lsb of prevTid0Pic.
[0147] The variable prevPicOrderCntMsb is set to be equal to the PicOrderCntMsb of prevTid0Pic.
[0148] The derivation of the variable PicOrderCntMsb for the current image is as follows:
[0149] -if PicOrderCntMsb is set to equal to ph_poc_msb_cycle_val * MaxPicOrderCntLsb.
[0150] -otherwise If the current image is a CLVSS image, then PicOrderCntMsb is set to 0. ...
[0152] 7.4.3.3 Sequence Parameter Set (RBSP) Semantics ...
[0154] A value of 1 for sps_gdr_enabled_flag indicates that GDR images are enabled and may exist in CLVS. A value of 0 for sps_gdr_enabled_flag indicates that GDR images are disabled and do not exist in CLVS.
[0155] ...
[0157] 7.4.3.10 Sequence End RBSP Semantics ...
[0159] ...
[0161] Figure 1 This is a block diagram of an example video processing system 1000 that can implement the various techniques disclosed herein. Various implementations may include some or all of the components in system 1000. System 1000 may include an input 1002 for receiving video content. The video content may be received in a raw or uncompressed format (e.g., 8 or 10-bit multi-component pixel values), or in a compressed or encoded format. Input 1002 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces (such as Ethernet, Passive Optical Networking (PON), etc.) and wireless interfaces (such as Wi-Fi or cellular interfaces).
[0162] System 1000 may include an encoding / decoding component 1004 capable of implementing the various encoding / decoding or coding methods described in this document. Encoding / decoding component 1004 can reduce the average bit rate of the video from input 1002 to the output of encoding / decoding component 1004 to produce an encoded / decoded representation of the video. Therefore, encoding / decoding techniques are sometimes referred to as video compression or video transcoding techniques. The output of encoding / decoding component 1004 can be stored or transmitted via connected communication, as represented by component 1006. The stored or communicated bitstream (or encoded / decoded) representation of the video received at input 1002 can be used by component 1008 to generate pixel values or displayable video that is sent to display interface 1010. The process of generating user-visible video from the bitstream representation is sometimes referred to as video decompression. Furthermore, although some video processing operations are referred to as “encoding / decoding” operations or tools, it should be understood that the encoding / decoding tools or operations are used at the encoder, and the corresponding decoding tools or operations will be inverted by the decoder to retrieve the encoded / decoded results.
[0163] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB), High Definition Multimedia Interface (HDMI), or DisplayPort. Examples of storage interfaces include SATA (Serial Advanced Technology Accessory), PCI, IDE, etc. The technologies described in this document can be implemented in a variety of electronic devices, such as mobile phones, laptops, smartphones, or other devices capable of digital data processing and / or video display.
[0164] Figure 2 This is a block diagram of a video processing apparatus 2000. Apparatus 2000 can be used to implement one or more of the methods described herein. Apparatus 2000 can be implemented in smartphones, tablets, computers, Internet of Things (IoT) receivers, etc. Apparatus 2000 may include one or more processors 2002, one or more memories 2004, and video processing hardware 2006. The processors (multiple) 2002 can be configured to implement the methods described herein (e.g., in...). Figure 6-9 The methods described herein may be one or more. Multiple memories 2004 may be used to store data and code used to implement the methods and techniques described herein. Video processing hardware 2006 may be used to implement some of the techniques described herein in hardware circuitry. In some embodiments, hardware 2006 may be partially or wholly located in one or more processors 2002, such as graphics processors.
[0165] Figure 3 This is a block diagram illustrating an example video encoding / decoding system 100 that can utilize the techniques disclosed herein. Figure 3 As shown, the video encoding / decoding system 100 may include a source device 110 and a destination device 120. The source device 110 generates encoded video data and may be referred to as a video encoding device. The destination device 120 can decode the encoded video data generated by the source device 110 and may be referred to as a video decoding device. The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.
[0166] Video source 112 may include sources such as video capture devices, interfaces for receiving video data from video content providers, and / or computer graphics systems that generate video data, or combinations of these sources. Video data may include one or more pictures. Video encoder 114 encodes the video data from video source 112 to generate a bitstream. The bitstream may include a sequence of bits forming a codec representation of the video data. The bitstream may include codec pictures and associated data. A codec picture is a codec representation of a picture. Associated data may include sequence parameter sets, picture parameter sets, and other syntax elements. I / O interface 116 includes a modulator / demodulator (modem) and / or a transmitter. Encoded video data may be transmitted directly to destination device 120 via network 130a through I / O interface 116. Encoded video data may also be stored on storage medium / server 130b for access by destination device 120.
[0167] Destination device 120 may include I / O interface 126, video decoder 124 and display device 122.
[0168] I / O interface 126 may include a receiver and / or a modem. I / O interface 126 may acquire encoded video data from source device 110 or storage medium / server 130b. Video decoder 124 may decode the encoded video data. Display device 122 may display the decoded video data to a user. Display device 122 may be integrated with destination device 120 or may be external to destination device 120 configured to connect to an external display device.
[0169] The video encoder 114 and the video decoder 124 can operate according to video compression standards such as High Efficiency Video Codec (HEVC), Multi-Functional Video Codec (VVC), and other current and / or other standards.
[0170] Figure 4 This is a block diagram illustrating an example of a video encoder 200, which may be... Figure 3 The video encoder 114 in the system 100 shown in the figure.
[0171] The video encoder 200 can be configured to perform any or all of the techniques disclosed herein. Figure 4 In the example, the video encoder 200 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video encoder 200. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.
[0172] The functional components of the video encoder 200 may include a segmentation unit 201, a prediction unit 202 (which may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra-frame prediction unit 206), a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213, and an entropy coding unit 214.
[0173] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an intra-block copy (IBC) unit. The IBC unit may perform prediction in IBC mode, where at least one reference picture is the picture in which the current video block is located.
[0174] Furthermore, some components, such as the motion estimation unit 204 and the motion compensation unit 205, can be highly integrated, but for interpretive purposes... Figure 4 The examples are shown separately.
[0175] The segmentation unit 201 can segment an image into one or more video blocks. The video encoder 200 and the video decoder 300 can support various video block sizes.
[0176] The mode selection unit 203 can, for example, select one of the intra-frame or inter-frame encoding / decoding modes based on the error result, and provide the obtained intra-frame or inter-frame encoded / decoded blocks to the residual generation unit 207 to generate residual block data and to the reconstruction unit 212 to reconstruct the encoded blocks for use as reference images. In some examples, the mode selection unit 203 can select a combined intra-frame and inter-frame prediction (CIIP) mode, where the prediction is based on the inter-frame prediction signal and the intra-frame prediction signal. The mode selection unit 203 can also select the resolution of the motion vector (e.g., sub-pixel or integer pixel precision) for the blocks in the inter-frame prediction case.
[0177] To perform inter-frame prediction for the current video block, motion estimation unit 204 can generate motion information for the current video block by comparing one or more reference frames from buffer 213 with the current video block. Motion compensation unit 205 can determine the predicted video block for the current video block based on the motion information of the image from buffer 213 (rather than the image associated with the current video block) and decoded samples.
[0178] The motion estimation unit 204 and the motion compensation unit 205 can perform different operations on the current video block, for example, the different operations performed depend on whether the current video block is in an I-strip, a P-strip, or a B-strip.
[0179] In some examples, motion estimation unit 204 can perform unidirectional prediction of the current video block, and can search for a reference video block for the current video block in the reference images of list 0 or list 1. Motion estimation unit 204 can then generate a reference index indicating that the reference image in list 0 or list 1 contains the reference video block, and a motion vector indicating the spatial displacement between the current video block and the reference video block. Motion estimation unit 204 can output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current block based on the reference video block indicated by the motion information of the current video block.
[0180] In other examples, motion estimation unit 204 can perform bidirectional prediction of the current video block. Motion estimation unit 204 can search for a reference video block for the current video block in the reference images of list 0 and can also search for another reference video block for the current video block in the reference images of list 1. Motion estimation unit 204 can then generate a reference index indicating that the reference images in list 0 or list 1 contain the reference video block, and a motion vector indicating the spatial displacement between the reference video block and the current video block. Motion estimation unit 204 can output the reference index and the motion vector of the current video block as the motion information of the current video block. Motion compensation unit 205 can generate a predicted video block for the current video block based on the reference video block indicated by the motion information of the current video block.
[0181] In some examples, the motion estimation unit 204 can output the complete set of motion information for the decoder's decoding process.
[0182] In some examples, motion estimation unit 204 may not output the complete set of motion information for the current video. Instead, motion estimation unit 204 may signal the motion information of the current video block by referencing the motion information of another video block. For example, motion estimation unit 204 may determine that the motion information of the current video block is sufficiently similar to the motion information of adjacent video blocks.
[0183] In one example, the motion estimation unit 204 may indicate in the syntax structure associated with the current video block that the current video block has the same motion information value as another video block.
[0184] In another example, motion estimation unit 204 can identify another video block and motion vector difference (MVD) in the syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicating video block. Video decoder 300 can use the motion vector of the indicating video block and the motion vector difference to determine the motion vector of the current video block.
[0185] As discussed above, the video encoder 200 can predictively signal motion vectors. Two examples of predictive signaling notification techniques that can be implemented by the video encoder 200 include Advanced Motion Vector Prediction (AMVP) and merge pattern signaling notification.
[0186] Intra-prediction unit 206 can perform intra-prediction on the current video block. When intra-prediction unit 206 performs intra-prediction on the current video block, it can generate prediction data for the current video block based on decoded samples from other video blocks in the same frame. The prediction data for the current video block can include the predicted video block and various syntax elements.
[0187] The residual generation unit 207 can generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) multiple predicted video blocks from the current video block. The residual data for the current video block can include residual video blocks corresponding to different sample components of the samples in the current video block.
[0188] In other examples, such as in skip mode, residual data for the current video block may not exist, and the residual generation unit 207 may not perform a subtraction operation.
[0189] The transform processing unit 208 can generate one or more transform coefficient video blocks of the current video block by applying one or more transforms to the residual video block associated with the current video block.
[0190] After the transform processing unit 208 generates a transform coefficient video block associated with the current video block, the quantization unit 209 can quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0191] The inverse quantization unit 210 and the inverse transform unit 211 can apply inverse quantization and inverse transform to the transform coefficient video block respectively to reconstruct the residual video block from the transform coefficient video block. The reconstruction unit 212 can add the reconstructed residual video block to the corresponding samples of one or more predicted video blocks generated by the prediction unit 202 to produce a reconstructed video block associated with the current block for storage in the buffer 213.
[0192] After the video block is reconstructed in reconstruction unit 212, a loop filtering operation can be performed to reduce video block artifacts in the video block.
[0193] Entropy encoding unit 214 can receive data from other functional components of video encoder 200. When entropy encoding unit 214 receives data, it can perform one or more entropy encoding operations to generate entropy encoded data and output a bitstream including the entropy encoded data.
[0194] Figure 5 This is a block diagram illustrating an example of a video decoder 300, which may be... Figure 3 The video decoder 114 in the system 100 shown in the figure.
[0195] The video decoder 300 can be configured to perform any or all of the techniques disclosed herein. Figure 5 In the example, the video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video decoder 300. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.
[0196] exist Figure 5 In the example, the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra-frame prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, the video decoder 300 can perform operations related to the video encoder 200 ( Figure 4 The decoding process is the overall inversion of the encoding process described.
[0197] Entropy decoding unit 301 can retrieve the encoded bitstream. The encoded bitstream may include entropy-encoded video data (e.g., encoded blocks of video data). Entropy decoding unit 301 can decode the entropy-encoded video, and based on the entropy-encoded video data, motion compensation unit 302 can determine motion information including motion vectors, motion vector precision, reference image list index, and other motion information. Motion compensation unit 302 can determine such information, for example, by performing AMVP and merge modes.
[0198] The motion compensation unit 302 can generate motion compensation blocks, possibly based on interpolation filters. The identifier of the interpolation filter to be used at sub-pixel precision can be included in the syntax element.
[0199] The motion compensation unit 302 can use the interpolation filter used by the video encoder 200 during the encoding of the video block to calculate the interpolation values of a sub-integer number of pixels of the reference block. The motion compensation unit 302 can determine the interpolation filter used by the video encoder 200 based on the received syntax information and use the interpolation filter to generate the prediction block.
[0200] The motion compensation unit 302 can use some syntactic information to determine: the size of the blocks used to encode (multiple) frames and / or (multiple) stripes of the encoded video sequence, segmentation information describing how each macroblock of the image of the encoded video sequence is segmented, a mode indicating how each segment is encoded, one or more reference frames (and a list of reference frames) for each inter-frame coded block, and other information for decoding the encoded video sequence.
[0201] Intra-prediction unit 303 can use, for example, an intra-prediction mode received in the bitstream to form prediction blocks from spatially adjacent blocks. Inverse quantization unit 303 inverse quantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by entropy decoding unit 301. Inverse transform unit 303 applies an inverse transform.
[0202] The reconstruction unit 306 can sum the residual blocks using the corresponding prediction blocks generated by the motion compensation unit 202 or the intra-frame prediction unit 303 to form a decoded block. As desired, a deblocking filter can also be applied to filter the decoded block to remove block artifacts. The decoded video block is then stored in a buffer 307, which provides a reference block for subsequent motion compensation / intra-frame prediction and also produces the decoded video for presentation on the display device.
[0203] Figure 6-11 It shows that it can be used in, for example Figure 1-5 The embodiments shown are example methods for implementing the above-described technical solution.
[0204] Figure 6 A flowchart of an example method 600 for video processing is shown. Method 600 includes, at operation 610, performing a conversion between a video comprising one or more images and a bitstream of the video, the bitstream conforming to a format rule that specifies constraints on the value of a first syntax element, the first syntax element specifying the presence of a second syntax element in the image header syntax structure of the current image, and the second syntax element specifying the value of the most significant bit (MSB) period of the image sequence count (POC) of the current image.
[0205] Figure 7 A flowchart of an example method 700 for video processing is shown. Method 700 includes, at operation 710, performing a conversion between a video comprising one or more images and a video bitstream, the bitstream conforming to a format rule that specifies the derivation of the Picture Order Count (POC) in the absence of syntax elements, wherein the syntax elements specify the value of the most significant bit (MSB) period of the POC for the current image.
[0206] Figure 8A flowchart of an example method 800 for video processing is shown. Method 800 includes, in operation 810, performing a conversion between video and a video bitstream according to a rule, wherein the bitstream includes access units AU, and the access units AU include pictures, and the rule stipulates that the output order of the AU is different from the decoding order of the AU, and stepwise decoding refresh (GDR) of the pictures is not allowed in the bitstream.
[0207] Figure 9 A flowchart illustrating an example method 900 for video processing is provided. Method 900 includes, at operation 910, performing a conversion between video and a video bitstream according to format rules. The bitstream includes multiple layers in multiple access units (AUs), each AU including one or more pictures. The format rules specify that, in response to the presence of a first-layer End-of-Sequence (EOS) Network Abstraction Layer (NAL) unit in a first access unit (AU) of the bitstream, each subsequent picture in each of the one or more higher layers of the first layer in subsequent AUs of the bitstream is a Codec Layer Video Sequence Start (CLVSS) picture.
[0208] Figure 10 A flowchart of an example method 1000 for video processing is shown. Method 1000 includes, in operation 1010, performing a conversion between video and a video bitstream according to format rules, the bitstream comprising multiple layers in multiple access units AU, each access unit AU comprising one or more pictures, the format rules specifying that, in response to a first picture in a first access unit, a codec layer video sequence start (CLVSS) picture, a second picture is a CLVSS picture, the codec layer video sequence start (CLVSS) picture being a clean random access (CRA) picture or a progressive decode refresh (GDR) picture.
[0209] Figure 11 A flowchart of an example method 1100 for video processing is shown. Method 1100 includes, in operation 1110, performing a conversion between a video and a video bitstream comprising one or more images according to rules specifying that the bitstream includes at least a first image of output, the first image being in the output layer, the first image including a syntax element equal to one, and the syntax element affecting the decoded image output and removal process associated with a hypothetical reference decoder (HRD).
[0210] The following solutions show example implementations of the techniques discussed in the previous chapter (e.g., Items 1-5).
[0211] The following is a list of preferred solutions for some embodiments.
[0212] A1. A method for video processing, comprising: performing a conversion between a video comprising one or more images and a bitstream of the video, wherein the bitstream conforms to a format rule, wherein the format rule specifies constraints on the value of a first syntax element, the first syntax element specifying the presence of a second syntax element in an image header syntax structure of the current image, and wherein the second syntax element specifying the value of the most significant bit (MSB) period of the image sequence count (POC) of the current image.
[0213] A2. The method described in solution A1, wherein in response to the value of a flag being zero and the interlayer reference picture (ILRP) entry being in the reference picture list of the current picture's stripe, the value of the first syntax element is zero, and wherein the flag specifies whether the indexing layer uses interlayer prediction.
[0214] A3. According to the method described in solution A2, the reference image list includes a first reference image list (RefPicList[0]) or a second reference image list (RefPicList[1]).
[0215] A4. According to the method described in solution A2, the value of the first syntax element being equal to zero indicates that the second syntax element does not exist in the image header syntax structure.
[0216] A5. According to the method described in solution A2, the value of the flag being equal to zero specifies that the index layer is allowed to use inter-layer prediction.
[0217] A6. The method according to solution A1, wherein in response to a flag value equal to zero and the image having (i) a first identifier equal to a second identifier in the current access unit (AU) of the reference layer of the current layer, and (ii) a third identifier less than or equal to a threshold, wherein the flag specifies whether the index layer uses inter-layer prediction, wherein the first identifier specifies the layer to which the video codec layer (VCL) network abstraction layer (NAL) unit belongs, wherein the second identifier specifies the layer to which the reference image belongs, wherein the third identifier is a temporal identifier, and wherein the threshold is based on a second syntax element that specifies whether an image in the index layer that is neither an intra-frame random access image (IRAP) image nor a progressive decode refresh (GDR) image is used as an inter-layer reference image (IRLP) for decoding images in the index layer.
[0218] A7. According to the method described in solution A6, the first identifier is nuh_layer_id, the second identifier is refpicLayerId, and the third identifier is TemporalId, and the second syntax element is vps_max_tid_il_ref_pics_plus1.
[0219] A8. According to the method described in solution A1, the first syntax element is never required to be zero.
[0220] A9. The method according to any one of solutions A2 to A8, wherein the first syntax element is ph_poc_msb_cycle_present_flag, the flag is vps_independent_layer_flag, and wherein the second syntax element is ph_poc_msb_cycle_val.
[0221] A10. A method for video processing, comprising: performing a conversion between a video comprising one or more images and a video bitstream, wherein the bitstream conforms to a format rule, wherein the format rule specifies the derivation of a picture sequence count (POC) in the absence of a syntax element, and wherein the syntax element specifies the value of the most significant bit (MSB) period of the POC for the current image.
[0222] A11. The method described in solution A10, wherein the syntax element is ph_poc_msb_cycle_val.
[0223] A12. A video processing method comprising: performing a conversion between a video and a video bitstream according to rules, wherein the bitstream includes an access unit AU, the access unit AU including pictures, wherein the rules specify that the output order of the AU is different from the decoding order of the AU, and step-by-step decoding refresh (GDR) of the pictures is not allowed in the bitstream.
[0224] A13. The method according to solution A12, wherein, in response to a flag equal to one, the output order and decoding order of all images in the codec layer video sequence (CLVS) are the same, and wherein the flag specifies whether GDR images are enabled.
[0225] A14. The method according to solution A12, wherein in response to a flag of the sequence parameter set (SPS) referenced by the pictures in the codec video sequence (CVS) being equal to one, the output order of the AU is the same as the decoding order, and wherein the flag specifies whether GDR pictures are enabled.
[0226] A15. The method according to solution A12, wherein the flag of the sequence parameter set (SPS) referenced by the image is equal to one, the output order of the AU is the same as the decoding order, and wherein the flag specifies whether the GDR image is enabled.
[0227] A16. The method according to solution A12, wherein the flag responding to the Sequence Parameter Set (SPS) in the bitstream is equal to one, the output order of the AU is the same as the decoding order, and wherein the flag specifies whether the GDR picture is enabled.
[0228] A17. The method according to any one of solutions A13 to A16, wherein the flag is sps_gdr_enabled_flag.
[0229] The following is another list of preferred solutions for some of the embodiments.
[0230] B1. A video processing method comprising: performing a conversion between a video and a video bitstream according to a format rule, wherein the bitstream includes multiple layers in a plurality of access units (AUs), the plurality of access units (AUs) including one or more pictures, wherein the format rule specifies that, in response to the presence of a first layer End of Sequence (EOS) Network Abstraction Layer (NAL) unit in a first access unit (AU) in the bitstream, a subsequent picture of each of one or more higher layers in the first layer of AUs following the first AU in the bitstream is a codec layer Video Sequence Start (CLVSS) picture.
[0231] B2. According to the method described in solution B1, wherein the format rules further specify that for a second layer that uses the first layer as a reference layer, the first picture in the decoding order is a CLVSS picture, and the second layer exists in the codec video sequence (CVS) that includes the first layer.
[0232] B3. The method described in solution B1, wherein one or more higher layers include all or specific higher layers.
[0233] B4. According to the method described in solution B1, wherein the format rules further specify that, for a second layer that is a layer higher than the first layer, the first picture in the decoding order is a CLVSS picture, and the second layer exists in the codec video sequence (CVS) that includes the first layer.
[0234] B5. The method described in solution B1, wherein the format rules further specify that EOS NAL units exist in each layer of the codec video sequence (CVS) in the bitstream.
[0235] B6. The method according to solution B1, wherein the format rules further specify that the second layer, which is a layer higher than the first layer, includes EOS NAL units, and the second layer exists in the codec video sequence (CVS) that includes the first layer.
[0236] B7. The method described in solution B1, wherein the format rules further specify that the second layer, which uses the first layer as a reference layer, includes EOS NAL units, and the second layer exists in the codec video sequence (CVS) that includes the first layer.
[0237] B8. A video processing method comprising: performing a conversion between a video and a video bitstream according to a format rule, wherein the bitstream includes multiple layers in a plurality of access units AU, the plurality of access units AU including one or more pictures, wherein the format rule specifies that, in response to a first picture in a first access unit, a codec layer video sequence start (CLVSS) picture, a second picture is a CLVSS picture, and the codec layer video sequence start (CLVSS) picture is a clean random access (CRA) picture or a progressive decode refresh (GDR) picture.
[0238] B9. The method described in solution B8, wherein the second image is an image of the layer in the first access unit.
[0239] B10. The method according to solution B8, wherein the first layer includes a first image, and wherein the second image is an image in a second layer that is higher than the first layer.
[0240] B11. The method according to solution B8, wherein the first layer includes a first image, and wherein the second image is an image in a second layer that uses the first layer as a reference layer.
[0241] B12. The method according to solution B8, wherein the second picture is the first picture in the decoding order in the second access unit following the first access unit.
[0242] B13. The method described in solution B8, wherein the second image is any image in the first access unit.
[0243] B14. The method according to any one of solutions B1 to B13, wherein the CLVSS picture is a codec picture with a flag equal to one, the codec picture being an (IRAP) picture or a (GDR) picture, wherein the flag equal to one indicates that the decoder does not output the associated picture when determining that the associated picture includes a reference to a picture not present in the bitstream.
[0244] The following is another list of preferred solutions for some embodiments.
[0245] C1. A video processing method comprising: performing a conversion between a video and a video bitstream comprising one or more images according to rules, wherein the rules specify that the bitstream includes at least a first image of output, wherein the first image is in an output layer, wherein the first image includes a syntax element equal to one, and wherein the syntax element affects the decoded image output and removal process associated with a hypothetical reference decoder (HRD).
[0246] C2. The method described in solution C1, wherein the rules apply to all levels and the bitstream is allowed to conform to any level.
[0247] C3. The method described in solution C2, wherein the syntax element is ph_pic_output_flag.
[0248] C4. According to the method described in solution C2, the grade is either a main 10 still image grade or a main 4:4:4 10 still image grade.
[0249] The following list of solutions applies to each of the solutions listed above.
[0250] O1. The method according to any one of the foregoing solutions, wherein the conversion includes decoding video from a bitstream.
[0251] O2. The method according to any one of the foregoing solutions, wherein the conversion includes encoding the video into a bitstream.
[0252] O3. A method for storing a bitstream representing a video into a computer-readable recording medium, comprising generating a bitstream from a video according to the method described in any one or more of the foregoing solutions, and storing the bitstream in a computer-readable recording medium.
[0253] O4. A video processing apparatus comprising a processor configured to implement one or more of the aforementioned solutions.
[0254] O5. A computer-readable medium having instructions stored thereon, which, when executed, cause a processor to implement the method described in one or more of the foregoing solutions.
[0255] O6. A computer-readable medium for storing a bitstream generated according to any one or more of the foregoing solutions.
[0256] O7. A video processing apparatus for storing bitstreams, wherein the video processing apparatus is configured to implement the method described in any one or more of the foregoing solutions.
[0257] The following is another list of preferred solutions for some embodiments.
[0258] P1. A video processing method comprising performing a conversion between a video comprising one or more images and a codec representation of the video, wherein the codec representation conforms to a format rule, wherein the format rule specifies constraints on the values of syntax elements indicating the presence of the most significant bit period in the sequential counting of images in the video.
[0259] P2. According to the method described in solution P1, the formatting rules stipulate that the value of the syntax element is 0 when the independent value flag is set to zero and at least one stripe of the picture uses an interlayer reference picture in its reference list.
[0260] P3. The method according to any one of solutions P1 to P2, wherein the format rules specify that the zero value of the syntax element is indicated by omitting the syntax element in the encoding / decoding representation.
[0261] P4. A video processing method comprising performing a conversion between a video comprising one or more images and a video codec representation, wherein the conversion conforms to a prescribed rule that disallows progressively decoding and refreshing of images when the output order of an access unit differs from the decoding order of the access units.
[0262] P5. A video processing method comprising performing a conversion between a video layer comprising one or more video images and a video codec representation, wherein the codec representation conforms to a format rule that specifies that if a first Network Abstraction Layer (NAL) unit indicating the end of a video sequence exists in an access unit of the layer, then the next image of each higher layer in the codec representation must have a codec layer video sequence start type.
[0263] P6. According to the method described in solution P5, the format rules also stipulate that the first picture of the second layer using this layer as a reference layer in the decoding order should have the codec layer video sequence start type.
[0264] P7. The method according to any one of solutions P1 to P5, wherein performing the conversion includes encoding the video to generate a codec representation.
[0265] P8. The method according to any one of solutions P1 to P5, wherein performing the conversion includes parsing and decoding the codec representation to generate video.
[0266] P9. A video decoding apparatus, comprising a processor configured to implement the method described in one or more of solutions P1 to P8.
[0267] P10. A video encoding apparatus comprising a processor configured to implement the method described in one or more of solutions P1 to P8.
[0268] P11. A computer program product having computer code stored thereon, which, when executed by a processor, causes the processor to implement the method described in any one of solutions P1 to P8.
[0269] In this document, the term "video processing" can refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm can be applied during the conversion from the pixel representation of a video to its corresponding bitstream representation, and vice versa. As defined in the syntax, the bitstream representation (or simply bitstream) of the current video block can (for example) correspond to bits that are co-occurring or scattered at different locations within the bitstream. For example, a macroblock can be encoded based on the error residuals from the transformation and encoding / decoding, and also using bits in the header and other fields in the bitstream.
[0270] The disclosures and other schemes, examples, embodiments, modules, and functional operations described in this document can be implemented in digital electronic circuits or in computer software, firmware, or hardware, containing the structures disclosed in this document and their equivalents, or combinations thereof. The disclosed and other embodiments can be implemented as one or more computer program products encoded on a computer-readable medium, i.e., one or more computer program instruction modules for execution by a data processing apparatus or for controlling the operation of a data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a complex influencing machine-readable propagating signals, or combinations thereof. The term "data processing apparatus" encompasses all means, devices, and machines for processing data, including, for example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, the apparatus may also include code that creates an execution environment for the computer program in question, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, or combinations thereof. Propagating signals are artificially generated signals, such as machine-generated electrical, optical, or electromagnetic signals, which are generated to encode information for transmission to a suitable receiver device.
[0271] Computer programs (also known as programs, software, software applications, scripts, or code) can be written in any programming language, including compiled or interpreted languages, and can be deployed in any form, including standalone programs or modules, components, subroutines, or other units suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple co-located files (e.g., a file storing one or more modules, subroutines, or code portions). A computer program can be deployed to execute on one computer or on multiple computers located at a single site or distributed across multiple sites and interconnected by a communications network.
[0272] The processes and logic flows described in this document can be performed by one or more programmable processors executing one or more computer programs to perform functions by manipulating input data and generating outputs. The processes and logic flows can also be performed by special-purpose logic circuitry (e.g., field-programmable gate arrays (FPGAs) or application-specific integrated circuits (ASICs)), and the apparatus can be implemented as special-purpose logic circuitry (e.g., FPGAs or ASICs).
[0273] Processors suitable for executing computer programs include, for example, both general-purpose and special-purpose microprocessors, and any one or more processors in any type of digital computer. Typically, a processor receives instructions and data from read-only memory or random access memory, or both. The basic components of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include one or more mass storage devices (e.g., magneto-optical, magneto-optical, or optical disc) for storing data, or operatively coupled to receive data from or transfer data to a mass storage device (e.g., magneto-optical, magneto-optical, or optical disc), or both. However, a computer does not necessarily need to have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices (e.g., EPROM, EEPROM, and flash memory devices); magnetic disks (e.g., internal hard disks or removable disks); magneto-optical disks; and CD-ROM and DVD-ROM disks. Processors and memory may be supplemented by or incorporated into special-purpose logic circuitry.
[0274] While this patent document contains numerous details, these details should not be construed as limiting any subject matter or the scope of the claims, but rather as descriptions of features specific to particular embodiments of a particular art. In this patent document, certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented separately in multiple embodiments or in various suitable sub-combinations. Furthermore, although features may be described above as operating in certain combinations and even initially claimed in the same manner, in certain circumstances one or more features from the claimed combination may be removed from the combination, and the claimed combination may be for sub-combinations or variations thereof.
[0275] Similarly, although operations are depicted in a specific order in the accompanying drawings, this should not be construed as requiring such operations to be performed in the specific order or sequence shown, or to perform all the operations shown, in order to achieve the desired result. Furthermore, the separation of various system components in the embodiments described in this patent document should not be construed as requiring such separation in all embodiments.
[0276] Only a few implementations and examples are described, and other implementations, enhancements and variations can be made based on what is described and shown in this patent document.
Claims
1. A video processing method, comprising: Perform the conversion between a video containing one or more images and the bitstream of that video according to the rules. The rule stipulates that the bitstream must include at least one output first image. The first image is in the output layer. The first image includes a syntax element equal to one, and The syntax elements affect the decoded image output and removal process associated with the hypothetical reference decoder HRD.
2. The method of claim 1, wherein, The rules apply to all tiers and the bitstream is allowed to conform to any tier.
3. The method of claim 2, wherein, The syntax element is ph_pic_output_flag.
4. The method according to claim 2, wherein, The aforementioned quality level refers to either a main 10-image still image quality level or a main 4:4:4 10-image still image quality level.
5. The method of claim 1, wherein the conversion comprises decoding the video from the bitstream.
6. The method according to claim 1, wherein, The conversion includes encoding the video into the bitstream.
7. An apparatus for processing video data, comprising a processor and a non-transitory memory having instructions thereon, wherein, When the instruction is executed by the processor, the processor: Perform the conversion between a video containing one or more images and the bitstream of that video according to the rules. The rule stipulates that the bitstream must include at least one output first image. The first image is in the output layer. The first image includes a syntax element equal to one, and The syntax elements affect the decoded image output and removal process associated with the hypothetical reference decoder HRD.
8. The apparatus according to claim 7, wherein, The rules apply to all levels, and the bitstream is allowed to conform to any level.
9. The apparatus according to claim 7, wherein, The syntax element is ph_pic_output_flag.
10. The apparatus according to claim 8, wherein, The aforementioned quality level refers to either a main 10-image still image quality level or a main 4:4:4 10-image still image quality level.
11. The apparatus according to claim 7, wherein, The conversion includes decoding the video from the bitstream.
12. The apparatus according to claim 7, wherein, The conversion includes encoding the video into the bitstream.
13. A non-transitory computer-readable storage medium for storing instructions, said instructions causing a processor to: Perform the conversion between a video containing one or more images and the bitstream of that video according to the rules. in, The rule stipulates that the bitstream must include at least one output first image. The first image is in the output layer. The first image includes a syntax element equal to one, and The syntax elements affect the decoded image output and removal process associated with the hypothetical reference decoder HRD.
14. The non-transitory computer-readable storage medium according to claim 13, wherein, The rules apply to all levels, and the bitstream is allowed to conform to any level.
15. The non-transitory computer-readable storage medium according to claim 13, wherein, The syntax element is ph_pic_output_flag.
16. The non-transitory computer-readable storage medium according to claim 14, wherein, The aforementioned quality level refers to either a main 10-image still image quality level or a main 4:4:4 10-image still image quality level.
17. The non-transitory computer-readable storage medium according to claim 13, wherein, The conversion includes decoding the video from the bitstream.
18. The non-transitory computer-readable storage medium according to claim 13, wherein, The conversion includes encoding the video into the bitstream.
19. A non-transitory computer-readable recording medium having a computer program and a bit stream stored thereon, wherein, When the computer program is executed by the processor, it implements the video processing method of claim 1 to generate the bitstream.
20. The non-transitory computer-readable recording medium according to claim 19, wherein, The syntax element is ph_pic_output_flag.
21. A method for storing a bit stream, comprising: The bit stream is generated by performing the method according to any one of claims 1 to 6; as well as The bitstream is stored in a computer-readable recording medium.