Stepwise Decoding Refresh Access Unit in Scalable Video Coding
By configuring access units and managing random access points in scalable video codecs, the problems of low bandwidth usage efficiency and increased latency in multi-layer video codecs are solved, and more efficient video codecs and flexible random access are achieved.
Patent Information
- Application Number
- CN202180021965.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-03-16
- Filing Date
- 2021-03-15
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2041-03-15
AI Technical Summary
When existing video encoding and decoding technologies support multi-layer video encoding and decoding, especially in scalable video encoding and decoding, it is difficult to effectively manage and utilize random access points, resulting in inefficient bandwidth usage and increased latency, which cannot meet the needs of modern video applications.
By configuring different access units in scalable video codecs, using format rules and syntax elements to indicate the type and layer identifier of the access unit, ensuring that each access unit happens to include the picture of the video layer, supporting step-by-step decoding and intra-random access points, optimizing the management of random access points.
It improves the bandwidth usage efficiency of video encoding and codec, reduces latency, supports more flexible random access, and adapts to the needs of modern video applications.
Smart Images

Figure CN115299054B_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This is the national phase of International Patent Application No. PCT / US2021 / 022404, filed on March 15, 2021, which claims the priority and benefit of U.S. Provisional Patent Application No. 62 / 990,387, filed on March 16, 2020. For all purposes as provided by law, the entire disclosure of the above - mentioned application is incorporated by reference as part of the disclosure of this application. Technical Field
[0003] This patent document relates to image and video encoding and decoding. Background Art
[0004] In the Internet and other digital communication networks, digital video occupies the largest bandwidth usage. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth demand for digital video use is expected to continue to grow. Summary of the Invention
[0005] The techniques disclosed in this document include configuring different access units in scalable video coding and decoding, which can be used by video encoders and decoders to process the decoded representation of video using control information for decoding the decoded representation.
[0006] In one example aspect, a video processing method is disclosed. The method includes performing a conversion between a video including one or more pictures in one or more video layers and a bitstream of the video, where the bitstream includes a decoded video sequence, the decoded video sequence includes one or more access units, and where the bitstream further includes a first syntax element indicating whether the access unit includes pictures of each video layer that constitutes the decoded video sequence.
[0007] In another example aspect, another video processing method is disclosed. The method includes performing a conversion between a video including one or more pictures in one or more video layers and a bitstream of the video, where the bitstream includes a decoded video sequence, the decoded video sequence includes one or more access units, and where the bitstream conforms to format rules specifying that each picture in a given access unit carries a layer identifier that is equal to the layer identifier of the first access unit in the decoded video sequence including one or more video layers.
[0008] In yet another example aspect, a video processing method is disclosed. The method includes performing a conversion between a video including one or more pictures in one or more video layers and a bitstream of the video, where the bitstream includes a decoded video sequence, the decoded video sequence includes one or more access units, and where the bitstream further includes a first syntax element indicating an access unit in the one or more access units that starts a new decoded video sequence.
[0009] In yet another example aspect, a video processing method is disclosed. The method includes performing a conversion between a video including one or more pictures in one or more video layers and a bitstream of the video, where the bitstream includes a decoded video sequence, the decoded video sequence includes one or more access units, and where the bitstream conforms to format rules that specify that a decoded video sequence start access unit that starts a new decoded video sequence includes pictures of each video layer specified in a video parameter set.
[0010] In yet another example aspect, a video processing method is disclosed. The method includes performing a conversion between a video including one or more pictures in one or more video layers and a bitstream of the video, the bitstream includes a decoded video sequence, the decoded video sequence includes one or more access units, and where the bitstream conforms to format rules that specify that a given access unit is identified as a decoded video sequence start access unit based on whether the given access unit is the first access unit in the bitstream or an access unit preceding the given access unit includes a sequence end network abstraction layer unit.
[0011] In yet another example aspect, a video processing method is disclosed. The method includes performing a conversion between a video including one or more pictures in one or more video layers and a bitstream of the video based on a rule, where the bitstream includes a decoded video sequence, the decoded video sequence includes one or more access units, and where the rule specifies that side information is used to indicate whether an access unit in the one or more access units is a decoded video sequence start access unit.
[0012] In yet another example aspect, a video processing method is disclosed. The method includes performing a conversion between a video including one or more pictures in one or more video layers and a bitstream of the video, the bitstream includes a decoded video sequence, the decoded video sequence includes one or more access units, and where the bitstream conforms to format rules that specify that each access unit among one or more access units that are stepwise decoded refresh access units exactly includes one picture of each video layer present in the decoded video sequence.
[0013] In yet another example aspect, a video processing method is disclosed. The method includes performing a conversion between a video including one or more pictures in one or more video layers and a bitstream of the video, the bitstream including a transcoded video sequence, the transcoded video sequence including one or more access units, and wherein the bitstream conforms to format rules that specify that each access unit among the one or more access units that is an intra random access point access unit includes exactly one picture of each video layer present in the transcoded video sequence.
[0014] In yet another example aspect, a video encoder device is disclosed. The video encoder includes a processor configured to implement the above method.
[0015] In yet another example aspect, a video decoder device is disclosed. The video decoder includes a processor configured to implement the above method.
[0016] In yet another example aspect, a computer-readable medium having code stored thereon is disclosed. The code embodies one of the methods described herein in the form of processor-executable code.
[0017] These and other features are introduced throughout this document. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 is a block diagram illustrating an example video processing system in which various techniques disclosed herein may be implemented.
[0019] Figure 2 is a block diagram of an example hardware platform for video processing.
[0020] Figure 3 is a block diagram illustrating an example video codec system in which some embodiments of the present disclosure may be implemented.
[0021] Figure 4 is a block diagram illustrating an example of an encoder in which some embodiments of the present disclosure may be implemented.
[0022] Figure 5 is a block diagram illustrating an example of a decoder in which some embodiments of the present disclosure may be implemented.
[0023] Figures 6 to 13 shows a flowchart for an example video processing method. DETAILED DESCRIPTION
[0024] For ease of understanding, section headings are used in this document, and the section headings do not limit the applicability of the technologies and embodiments disclosed in each section to that section. Additionally, in some descriptions, the H.266 terminology is used only for ease of understanding and not to limit the scope of the disclosed technologies. Similarly, the technologies described herein are also applicable to other video codec protocols and designs.
[0025] 1. Preliminary Discussion
[0026] This document is related to video coding and decoding technologies. Specifically, it is about the specification and signaling of random access point access units in scalable video coding, where the video bitstream can contain more than one layer. These ideas can be applied individually or in various combinations to any video coding standard or non-standard video codec that supports multi-layer video coding (e.g., Versatile Video Coding (VVC) currently under development).
[0027] 2. Abbreviations
[0028] APS Adaptive Parameter Set
[0029] AU Access Unit
[0030] AUD Access Unit Delimiter
[0031] AVC Advanced Video Coding
[0032] CLVS Video Sequence after Coding and Decoding
[0033] CPB Coding and Decoding Picture Buffer
[0034] CRA Clean Random Access
[0035] CTU Coding Tree Unit
[0036] CVS Coding and Decoding Video Sequence
[0037] DCI Decoding Capability Information
[0038] DPB Decoded Picture Buffer
[0039] EOB End of Bitstream
[0040] EOS End of Sequence
[0041] CGR Gradual Decoding Refresh
[0042] HEVC High Efficiency Video Coding
[0043] HRD Hypothetical Reference Decoder
[0044] IDR Instantaneous Decoding Refresh
[0045] JEM Joint Exploration Model
[0046] MCTS Motion Constrained Tile Set
[0047] NAL Network Abstraction Layer
[0048] OLS Output Layer Set
[0049] PH Picture Header
[0050] PPS Picture Parameter Set
[0051] PTL Profile, Tier and Level
[0052] PU Picture Unit
[0053] RAP Random Access Point
[0054] RBSP Raw Byte Sequence Payload
[0055] SEI Supplemental Enhancement Information
[0056] SPS Sequence Parameter Set
[0057] SVC Scalable Video Coding
[0058] VCL Video Coding Layer
[0059] VPS Video Parameter Set
[0060] VTM VVC Test Model
[0061] VUI Video Usability Information
[0062] VVC Versatile Video Coding
[0063] 3. Introduction to Video Coding
[0064] Video coding standards have mainly evolved through the development of well-known ITU-T and ISO / IEC standards. ITU-T developed H.261 and H.263, ISO / IEC developed MPEG-1 and MPEG-4 Visual, and the two organizations jointly developed the H.262 / MPEG-2 video, H.264 / MPEG-4 Advanced Video Coding (AVC) standard, and H.265 / HEVC standard. Since H.262, video coding standards have been based on a hybrid video coding structure that uses temporal prediction plus transform coding. To explore future video coding technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Team (JVET) in 2015. Since then, many new methods have been adopted by JVET and incorporated into a reference software called the Joint Exploration Model (JEM). At the same time, JVET meetings are held quarterly, and the goal of the new coding standard is to reduce the bit rate by 50% compared to HEVC. The new video coding standard was officially named Versatile Video Coding (VVC) at the JVET meeting in April 2018, and the first version of the VVC Test Model (VTM) was released at that time. With continuous efforts towards VVC standardization, new coding technologies have been adopted in the VVC standard at each JVET meeting. The VVC working draft and test model VTM are updated after each meeting. The VVC project is now aiming for Feature Complete (FDIS) at the meeting in July 2020.
[0065] 3.1 Random Access and Its Support in HEVC and VVC
[0066] Random access refers to the access and decoding of a bitstream starting from a picture that is not the first picture of the bitstream in decoding order. To support tuning and channel switching in broadcast / multicast and multiparty video conferencing, seeking in local playback and streaming, and stream adaptation in streaming, the bitstream needs to include frequent random access points, which are typically intra-coded pictures, but can also be inter-coded pictures (e.g., in the case of progressive decoding refresh).
[0067] HEVC includes signaling of Intra Random Access Point (IRAP) pictures in the NAL unit header by NAL unit type. Three types of IRAP pictures are supported, namely Instantaneous Decoder Refresh (IDR), Clean Random Access (CRA), and Broken Link Access (BLA) pictures. IDR pictures constrain the picture - to - picture prediction structure not to reference any pictures before the current Group of Pictures (GOP), and are commonly referred to as closed - GOP random access points. CRA pictures are less restrictive by allowing some pictures to reference pictures before the current GOP, and in the case of random access, all such pictures are discarded. CRA pictures are commonly referred to as open - GOP random access points. For example, during stream switching, BLA pictures typically result from the splicing of two bitstreams or parts thereof at a CRA picture. To enable the system to better use IRAP pictures, a total of six different NAL units are defined to signal the attributes of IRAP pictures, which can be used to better match the stream access point types defined in the ISO Base Media File Format (ISOBMFF) for random access support in Dynamic Adaptive Streaming over HTTP (DASH).
[0068] VVC supports three types of IRAP pictures, two types of IDR pictures (one type with or the other type without an associated RADL picture), and one type of CRA picture. These are basically the same as in HEVC. The BLA picture type in HEVC is not included in VVC, mainly for two reasons: i) The basic function of BLA pictures can be achieved by CRA pictures plus a Sequence End NAL unit, the presence of which indicates that the subsequent pictures start a new CVS in a single - layer bitstream. ii) During the development of VVC, it was desired to specify fewer NAL unit types than in HEVC, such that the NAL unit type field in the NAL unit header uses 5 bits instead of 6 bits to indicate.
[0069] Another key difference between VVC and HEVC in terms of random access support is that GDR is supported in a more standardized way in VVC. In GDR, the decoding of the bitstream can start from an inter-coded picture. Although the entire picture area cannot be decoded correctly at the beginning, after several pictures, the entire picture area will be correct. AVC and HEVC also support GDR, using recovery point SEI messages for signaling of GDR random access points and recovery points. In VVC, a new NAL unit type is specified for the indication of GDR pictures, and the recovery point is signaled in the picture header syntax structure. CVS and the bitstream are allowed to start with GDR pictures. This means that the entire bitstream is allowed to contain only inter-coded pictures without a single intra-coded picture. The main benefit of specifying GDR support in this way is to provide a compliant behavior for GDR. GDR enables the encoder to smooth the bitrate of the bitstream by distributing intra-coded strips or blocks across multiple pictures instead of intra-coding the entire picture, thus allowing a significant reduction in end-to-end latency, which is considered more important today than before because ultra-low latency applications such as wireless displays, online games, and drone-based applications have become more popular.
[0070] Another feature related to GDR in VVC is virtual boundary signaling. At the pictures between a GDR picture and its recovery point, the boundary between the refreshed area (i.e., the correctly decoded area) and the non-refreshed area can be signaled as a virtual boundary, and when signaled, loop filtering across the boundary will not be applied, so decoding mismatches for some samples at or near the boundary will not occur. This can be useful when an application determines to display the correctly decoded area during the GDR process.
[0071] IRAP pictures and GDR pictures can be collectively referred to as random access point (RAP) pictures.
[0072] 3.2 General Scalable Video Coding (SVC) and Scalable Video Coding in VVC
[0073] Scalable Video Coding (SVC, sometimes also referred to simply as scalability in video coding) refers to video coding in which a Base Layer (BL) (sometimes referred to as a Reference Layer (RL)) and one or more Scalable Enhancement Layers (ELs) are used. In SVC, the Base Layer can carry video data with a basic quality level. One or more Enhancement Layers can carry additional video data to support, for example, higher spatial, temporal, and / or signal-to-noise ratio (SNR) levels. Enhancement Layers can be defined relative to previously encoded layers. For example, the bottom layer can be used as the BL, while the top layer can be used as the EL. Intermediate layers can be used as ELs or RLs or both. For example, an intermediate layer (e.g., a layer that is neither the lowest nor the highest layer) can be an EL for layers below the intermediate layer (such as the base layer or any intermediate enhancement layer) and at the same time can be used as an RL for one or more enhancement layers above the intermediate layer. Similarly, in the multi-view or 3D extension of the HEVC standard, there can be multiple views, and information from one view can be used to code (e.g., encode or decode) information from another view (such as motion estimation, motion vector prediction, and / or other redundancies).
[0074] In SVC, the parameters used by the encoder or decoder are grouped into parameter sets based on the coding levels at which they can be exploited (e.g., video level, sequence level, picture level, slice level, etc.). For example, parameters that can be used by one or more coded video sequences in different layers of a bitstream can be included in a Video Parameter Set (VPS), while parameters used by one or more pictures in a coded video sequence can be included in a Sequence Parameter Set (SPS). Similarly, parameters exploited by one or more slices in a picture can be included in a Picture Parameter Set (PPS), and other parameters specific to an individual slice can be included in the slice header. Similarly, an indication of which (which) parameter set a particular layer is using at a given time can be provided at various coding levels.
[0075] Due to the support for Reference Picture Resampling (RPR) in VVC, it is possible to design support for bitstreams containing multiple layers (e.g., two layers with SD and HD resolutions in VVC) without any additional signal processing stage codec tools, because the upsampling required for spatial scalability support can use only the RPR upsampling filter. However, for scalability support, advanced syntax changes are required (compared to non-scalability support). Scalability support is specified in VVC version 1. Different from the scalability support in any previous video codec standard (including extensions of AVC and HEVC), the scalability design of VVC is as friendly as possible to single-layer decoder design. The decoding capabilities of multi-layer bitstreams are specified in a way that is as if there were only a single layer in the bitstream. For example, decoding capabilities such as DPB size are specified in a way that is independent of the number of layers in the bitstream to be decoded. Basically, a decoder designed for single-layer bitstreams can decode multi-layer bitstreams with little change. Compared to the multi-layer extension designs of AVC and HEVC, the HLS aspect has been significantly simplified, but at the cost of some flexibility. For example, it is required that an IRAP AU contains pictures of each layer present in the CVS.
[0076] 3.3 Parameter Sets
[0077] AVC, HEVC, and VVC specify parameter sets. The types of parameter sets include SPS, PPS, APS, and VPS. All of AVC, HEVC, and VVC support SPS and PPS. VPS was introduced starting from HEVC and is included in both HEVC and VVC. APS is not included in AVC or HEVC but is included in the latest VVC draft text.
[0078] SPS is designed to carry sequence-level header information, while PPS is designed to carry picture-level header information that does not change frequently. Using SPS and PPS, there is no need to repeat the information that does not change frequently for each sequence or picture, thus avoiding redundant signaling of this information. In addition, the use of SPS and PPS enables out-of-band transmission of important header information, which not only avoids the need for redundant transmission but also improves the error resilience.
[0079] The introduction of VPS is to carry sequence-level header information that is common to all layers in a multi-layer bitstream.
[0080] The introduction of APS is to carry such picture-level or slice-level information that requires a considerable number of bits for coding and decoding, can be shared by multiple pictures, and can have quite a few different variations in a sequence.
[0081] 3.4 Related Definitions in VVC
[0082] The relevant definitions in the latest VVC text (in JVET-Q2001-Ve / v15) are as follows.
[0083] Associated IRAP picture (of a specific picture): The previous IRAP picture that has the same nuh_layer_id value as the specific picture in decoding order (when it exists).
[0084] Warped decoded video sequence (CVS): An AU sequence that consists of a CVSS AU in decoding order, followed by zero or more AUs that are not CVSS AUs (including all subsequent AUs until but not including any subsequent AU that is a CVSS AU).
[0085] Warped decoded video sequence start (CVSS) AU: An AU that has PUs for each layer in the CVS and the warped decoded pictures in each PU are CLVSS pictures.
[0086] Gradual decoding refresh (GDR) AU: An AU in which the warped decoded pictures in each present PU are GDR pictures.
[0087] Gradual decoding refresh (GDR) PU: A PU in which the warped decoded picture is a GDR picture.
[0088] Gradual decoding refresh (GDR) picture: A picture in which each VCL NAL unit has a nal_unit_type equal to GDR_NUT.
[0089] Intra random access point (IRAP) AU: An AU that has PUs for each layer in the CVS and the warped decoded pictures in each PU are IRAP pictures.
[0090] Intra random access point (IRAP) picture: A warped decoded picture in which all VCL NAL units have the same nal_unit_type value within the range from IDR_W_RADL to CRA_NUT (including the endpoints).
[0091] 3.5 VPS syntax and semantics in VVC
[0092] VVC supports scalability, also known as scalable video coding, in which multiple layers can be encoded in a warped decoded video bitstream.
[0093] In the latest VVC text (in JVET-Q2001-Ve / v15), scalability information is signaled in the VPS, and its syntax and semantics are as follows.
[0094] 7.3.2.2 Video parameter set syntax
[0095]
[0096]
[0097]
[0098]
[0099] 7.4.3.2 Video Parameter Set RBSP Semantics
[0100] The VPS RBSP shall be available for the decoding process before being referenced, including in at least one AU with TemporalId equal to 0 or provided by external means.
[0101] All VPS NAL units with a specific vps_video_parameter_set_id value in a CVS shall have the same content.
[0102] The vps_video_parameter_set_id provides an identifier for the VPS for reference by other syntax elements. The value of vps_video_parameter_set_id shall be greater than 0.
[0103] vps_max_layers_minus1 plus 1 specifies the maximum allowed number of layers in each CVS that references the VPS.
[0104] vps_max_sublayers_minus1 plus 1 specifies the maximum number of temporal sublayers that can exist in a layer in each CVS that references the VP. The value of vps_max_sublayers_minus1 shall be in the range of 0 to 6 (including the endpoints).
[0105] The Vps_all_layers_same_num_sublayers_flag equal to 1 specifies that the number of temporal sublayers is the same for all layers in each CVS that references the VPS. The Vps_all_layers_same_num_sublayers_flag equal to 0 specifies that the layers in each CVS that references the VPS may or may not have the same number of temporal sublayers. If not present, the value of vps_all_layers_same_num_sublayers_flag is inferred to be equal to 1.
[0106] A Vps_all_independent_layers_flag equal to 1 specifies that all layers in the CVS are independently coded and decoded without using inter-layer prediction. A Vps_all_independent_layers_flag equal to 0 specifies that one or more layers in the CVS may use inter-layer prediction. If not present, the value of vps_all_independent_layers_flag is inferred to be equal to 1.
[0107] Vps_layer_id[I] specifies the nuh_layer_id value of the i-th layer. For any two non-negative integer values of m and n, when m is less than n, the value of vps_layer_id[m] shall be less than the value of vps_layer_id[n].
[0108] A Vps_independent_layer_flag[I] equal to 1 specifies that the layer with index I does not use inter-layer prediction. A Vps_independent_layer_flag[I] equal to 0 specifies that the layer with index I may use inter-layer prediction and the syntax element vps_direct_ref_layer_flag[I][j] exists in the VPS for j in the range from 0 to I-1 (including the endpoints). If not present, the value of vps_independent_layer_flag[I] is inferred to be equal to 1.
[0109] A vps_direct_ref_layer_flag[I][j] equal to 0 specifies that the layer with index j is not a direct reference layer of the layer with index i. A vps_direct_ref_layer_flag[I][j] equal to 1 specifies that the layer with index j is a direct reference layer of the layer with index i. When vps_direct_ref_layer_flag[I][j] is not present for I and j in the range from 0 to vps_max_layers_minus1 (including the endpoints), it is inferred to be equal to 0. When vps_independent_layer_flag[I] is equal to 0, there is at least one j value in the range from 0 to I-1 (including the endpoints) such that the value of vps_direct_ref_layer_flag[I][j] is equal to 1.
[0110] The variables NumDirectRefLayers[I], DirectRefLayerIdx[I][d], NumRefLayers[I], RefLayerIdx[I][r] and LayerUsedAsRefLayerFlag[j] are derived as follows:
[0111]
[0112]
[0113] Specify the layer index variable GeneralLayerIdx[I] of the layer with nuh_layer_id equal to vps_layer_id[I], derived as follows:
[0114] for(I = 0; I <= vps_max_layers_minus1; i++) (38)
[0115] GeneralLayerIdx[vps_layer_id[I]] = i
[0116] For any two different values of I and j, both within the range of 0 to vps_max_layers_minus1 (including the endpoints), when dependencyFlag[I][j] is equal to 1, the requirement for bitstream consistency is that the values of chroma_format_idc and bit_depth_minus8 applied to layer i should be equal to the values of chroma_format_idc and bit_depth_minus8 applied to layer j, respectively.
[0117] Max_tid_ref_present_flag[I] equal to 1 specifies the existence of the syntax element max_tid_il_ref_pics_plus1[I]. Max_tid_ref_present_flag[I] equal to 0 specifies the non-existence of the syntax element max_tid_il_ref_pics_plus1[I].
[0118] Max_tid_il_ref_pics_plus1[I] equal to 0 specifies that inter-layer prediction is not used for non-IRAP pictures of layer i. Max_tid_il_ref_pics_plus[I] greater than 0 specifies that, for decoding pictures of layer i, no pictures with TemporalId greater than max_tid_il_ref_pics_plus1[I] - 1 are used as ILRP. If not present, the value of max_tid_il_ref_pics_plus1[I] is inferred to be equal to 7.
[0119] Each_layer_is_an_ols_flag equal to 1 specifies that each OLS contains only one layer, and each layer in the CVS of the reference VPS is itself an OLS, and the single layer it contains is the only output layer. Each_layer_is_an_ols_flag equal to 0 means that an OLS can contain more than one layer. If vps_max_layers_minus1 is equal to 0, the value of each_layer_is_an_ols_flag is inferred to be equal to 1. Otherwise, when vps_all_independent_layers_flag is equal to 0, the value of each_layer_is_an_ols_flag is inferred to be equal to 0.
[0120] Ols_mode_idc equal to 0 specifies that the total number of OLSs specified by the VPS is equal to vps_max_layers_minus1 + 1, the i-th OLS includes the layers with layer indices from 0 to I (including the endpoints), and for each OLS, only the highest layer in the OLS is output.
[0121] Ols_mode_idc equal to 1 specifies that the total number of OLSs specified by the VPS is equal to vps_max_layers_minus1 + 1, the i-th OLS includes the layers with layer indices from 0 to I (including the endpoints), and for each OLS, all layers in the OLS are output.
[0122] Ols_mode_idc equal to 2 specifies that the total number of OLSs specified by the VPS is explicitly signaled, and for each OLS, the output layer is explicitly signaled, and the other layers are direct or indirect reference layers of the output layer of the OLS.
[0123] The value of ols_mode_idc shall be in the range of 0 to 2 (including the endpoints). The value 3 of ols_mode_idc is reserved for future use by ITU-T|ISO / IEC.
[0124] When vps_all_independent_layers_flag is equal to 1 and each_layer_is_an_ols_flag is equal to 0, the value of ols_mode_idc is inferred to be equal to 2.
[0125] Num_output_layer_sets_minus1 plus 1 specifies the total number of OLSs specified by the VPS when ols_mode_idc is equal to 2.
[0126] The variable TotalNumOlss that specifies the total number of OLSs specified by the VPS is derived as follows:
[0127]
[0128] The ols_output_layer_flag[I][j] equal to 1 specifies that when ols_mode_idc is equal to 2, the layer with nuh_layer_id equal to vps_layer_id[j] is the output layer of the i-th OLS. The ols_output_layer_flag[I][j] equal to 0 specifies that when ols_mode_idc is equal to 2, the layer with nuh_layer_id equal to vps_layer_id[j] is not the output layer of the i-th OLS.
[0129] The variable NumOutputLayersInOls[I] specifying the number of output layers in the i-th OLS, the variable NumSubLayersInLayerInOLS[I][j] specifying the number of sub-layers in the j-th layer in the i-th OLS, the variable OutputLayerIdInOls[I][j] specifying the nuh_layer_id value of the j-th output layer in the i-th OLS, and the variable LayerUsedAsOutputLayerFlag[k] specifying whether the k-th layer is used as an output layer in at least one OLS are derived as follows:
[0130]
[0131]
[0132]
[0133] For each value of I in the range from 0 to vps_max_layers_minus1 (including the endpoints), the values of LayerUsedAsRefleayerFlag[I] and LayerUsedAsoutputLayerFlag[I] cannot both be equal to 0. In other words, there should not be a layer that is neither an output layer of at least one OLS nor a direct reference layer of any other layer.
[0134] For each OL, at least one layer is an output layer. In other words, for any value of I in the range from 0 to TotalNumOlss - 1 (including the endpoints), the value of NumOutputLayersInOls[I] should be greater than or equal to 1.
[0135] The variable NumLayersInOls[I] specifying the number of layers in the i-th OLS and the variable LayerIdInOls[I][j] specifying the nuh_layer_id value of the j-th layer in the i-th OLS are derived as follows:
[0136]
[0137]
[0138] Note 1 - The 0th OLS only contains the lowest layer (i.e., the layer where nuh_layer_id is equal to vps_layer_id[0]), and for the 0th OLS, the only layer contained is output.
[0139] The variable OlsLayerIdx[I][j] that specifies the OLS layer index of the layer with nuh_layer_id equal to LayerIdInOls[I][j] is derived as follows:
[0140]
[0141] The lowest layer in each OLS should be an independent layer. In other words, for each value of I in the range from 0 to TotalNumOlss - 1 (including the endpoints), the value of vps_independent_layer_flag[GeneralLayerIdx[LayerIdInOls[I][0]]] should be equal to 1.
[0142] Each layer should be included in at least one OLS specified by the VPS. In other words, for each layer with a specific value of nuh_layer_id nuhLayerId that is one of vps_layer_id[k] where k is in the range from 0 to vps_max_layers_minus1 (including the endpoints), there should be at least one pair of values of I and j, where I is in the range from 0 to TotalNumOlss - 1 (including the endpoints) and j is in the range from 0 to NumLayersInOls[I] - 1 (including the endpoints), such that the value of LayerIdInOls[I][j] is equal to nuhLayerId.
[0143] Vps_num_ptls_minus1 plus 1 specifies the number of profile_tier_level() syntax structures in the VPS. The value of vps_num_ptls_minus1 should be less than TotalNumOlss.
[0144] Pt_present_flag[I] equal to 1 specifies the presence of profile, tier, and general constraint information in the i-th profile_tier_level() syntax structure in the VPS. Pt_present_flag[I] equal to 0 specifies the absence of profile, tier, and general constraint information in the i-th profile_tier_level() syntax structure in the VPS. The value of pt_present_flag[0] is inferred to be equal to 1. When pt_present_flag[I] is equal to 0, the profile, tier, and general constraint information of the i-th profile_tier_level() syntax structure in the VPS is inferred to be the same as that of the (I-1)-th profile_tier_level() syntax structure in the VPS.
[0145] Ptl_max_temporal_id[I] specifies the TemporalId of the highest sublayer representation for which level information is present in the i-th profile_tier_level() syntax structure in the VPS. The value of ptl_max_temporal_id[I] should be in the range of 0 to vps_max_sublayers_minus1 (including the endpoints). When vps_max_sublayers_minus1 is equal to 0, the value of ptl_max_temporal_id[I] is inferred to be equal to 0. When vps_max_sublayers_minus1 is greater than 0 and vps_all_layers_same_num_sublayers_flag is equal to 1, the value of ptl_max_temporal_id[I] is inferred to be equal to vps_max_sublayers_minus1.
[0146] Vps_ptl_alignment_zero_bit should be equal to 0.
[0147] Ols_ptl_idx[I] specifies the index of the profile_tier_level() syntax structure in the VPS list of profile_tier_level() syntax structures applied to the i-th OLS. When present, the value of ols_ptl_idx[I] should be in the range of 0 to vps_num_ptls_minus1 (including the endpoints). When vps_num_ptls_minus1 is equal to 0, the value of ols_ptl_idx[I] is inferred to be equal to 0.
[0148] When NumLayersInOls[I] is equal to 1, the profile_tier_level() syntax structure applied to the i-th OLS also exists in the SPS referenced by the layer in the i-th OLS. The requirement for bitstream consistency is that when NumLayersInOls[I] is equal to 1, the profile_tier_level() syntax structure signaled in the VPS and SPS for the i-th OLS should be the same.
[0149] Vps_num_dpb_params specifies the number of dpb_parameters() syntax structures in the VPS. The value of vps_num_dpb_params shall be in the range of 0 to 16 (including the endpoints). If not present, the value of vps_num_dpb_params is inferred to be equal to 0.
[0150] Vps_sublayer_dpb_params_present_flag is used to control the presence of the max_dec_pic_buffering_minus1[], max_num_reorder_pics[], and max_latency_increase_plus1[] syntax elements in the dpb_parameters() syntax structure in the VPS. When not present, vps_sub_dpb_params_info_present_flag is inferred to be equal to 0.
[0151] Dpb_max_temporal_id[I] specifies the highest sublayer representation TemporalId in the i-th dpb_parameters() syntax structure for which DPB parameters may be present in the VPS. The value of dpb_max_temporal_id[I] should be in the range of 0 to vps_max_sublayers_minus1 (including the endpoints). When vps_max_sublayers_minus1 is equal to 0, the value of dpb_max_temporal_id[I] is inferred to be equal to 0. When vps_max_sublayers_minus1 is greater than 0 and vps_all_layers_same_num_sublayers_flag is equal to 1, the value of dpb_max_temporal_id[I] is inferred to be equal to vps_max_sublayers_minus1.
[0152] Ols_dpb_pic_width[I] specifies the width, in luma samples, of each picture storage buffer for the i-th OLS.
[0153] Ols_dpb_pic_height[I] specifies the height of each picture storage buffer of the i-th OLS in units of luma samples.
[0154] Ols_dpb_params_idx[I] specifies the index into the list of dpb_parameters() syntax structures in the VPS that applies to the dpb_parameters() syntax structure for the ith OLS when NumLayersInOls[I] is greater than 1. When present, the value of ols_dpb_params_idx[I] shall be in the range of 0 to vps_num_dpb_params-1, inclusive. When ols_dpb_params_idx[I] is not present, the value of ols_dpb_params_idx[I] is inferred to be equal to 0.
[0155] When NumLayersInOls[I] is equal to 1, the dpb_parameters() syntax structure applicable to the i-th OLS is present in the SPS referenced by the layers in the i-th OLS.
[0156] Vps_general_hrd_params_present_flag equal to 1 specifies that the VPS contains a general_hrd_parameters() syntax structure and other HRD parameters. Vps_general_hrd_params_present_flag equal to 0 specifies that the VPS does not contain a general_hrd_parameters() syntax structure or other HRD parameters. If not present, the value of vps_general_hrd_params_present_flag is inferred to be equal to 0.
[0157] When NumLayersInOls[I] is equal to 1, the general_hrd_parameters() syntax structure and the ols_hrd_parameters() syntax structure applied to the i-th OLS exist in the SPS referenced by the layers in the i-th OLS.
[0158] A Vps_sublayer_cpb_params_present_flag equal to 1 specifies that the i-th ols_hrd_parameters() syntax structure in the VPS contains HRD parameters for the sublayer representation with TemporalIds in the range 0 to hrd_max_tid[I] (including the endpoints). A Vps_sublayer_cpb_params_present_flag equal to 0 specifies that the i-th ols_hrd_parameters() syntax structure in the VPS contains HRD parameters for the sublayer representation with a TemporalId equal to only hrd_max_tid[I]. When vps_max_sublayers_minus1 is equal to 0, the value of vps_sublayer_cpb_params_present_flag is inferred to be equal to 0.
[0159] When vps_sublayer_cpb_params_present_flag is equal to 0, the HRD parameters for the sublayer representation with TemporalIds in the range 0 to hrd_max_tid[I] - 1 (including the endpoints) are inferred to be the same as the HRD parameters for the sublayer representation with a TemporalId equal to hrd_max_tid[I]. These parameters include the HRD parameters starting from the fixed_pic_rate_general_flag[I] syntax element up to the sublayer_hrd_parameters(I) syntax structure immediately following the condition "if(general_vcl_hrd_params_present_flag)" in the ols_hrd_parameters syntax structure.
[0160] Num_ols_hrd_params_minus1 plus 1 specifies the number of ols_hrd_params() syntax structures present in the VPS when vps_general_hrd_params_present_flag is equal to 1. The value of num_ols_hrd_params_minus1 should be in the range 0 to TotalNumOlss - 1 (including the endpoints).
[0161] Hrd_max_tid[I] specifies the TemporalId of the highest sublayer whose HRD parameters are included in the i-th ols_hrd_parameters() syntax structure. The value of hrd_max_tid[I] should be in the range of 0 to vps_max_sublayers_minus1 (including the endpoints). When vps_max_sublayers_minus1 is equal to 0, the value of hrd_max_tid[I] is inferred to be equal to 0. When vps_max_sublayers_minus1 is greater than 0 and vps_all_layers_same_num_sublayers_flag is equal to 1, the value of hrd_max_tid[I] is inferred to be equal to vps_max_sublayers_minus1.
[0162] Ols_hrd_idx[I] specifies the index in the list of ols_hrd_parameters() syntax structures in the VPS of the ols_hrd_parameters() syntax structure applied to the i-th OLS when NumLayersInOls[I] is greater than 1. The value of ols_hrd_idx[I] should be in the range of 0 to num_ols_hrd_params_minus1 (including the endpoints).
[0163] When NumLayersInOls[I] is equal to 1, the ols_hrd_parameters() syntax structure applied to the i-th OLS exists in the SPS referenced by the layer in the i-th OLS.
[0164] If the value of num_ols_hrd_param_minus1 + 1 is equal to TotalNumOlss, the value of ols_hrd_idx[I] is inferred to be equal to i. Otherwise, when NumLayersInOls[I] is greater than 1 and num_ols_hrd_params_minus1 is equal to 0, the value of ols_hrd_idx[[I] is inferred to be equal to 0.
[0165] A Vps_extension_flag equal to 0 specifies that the vps_extension_data_flag syntax element does not exist in the VPS RBSP syntax structure. A Vps_extension_flag equal to 1 specifies that the vps_extension_data_flag syntax element exists in the VPS RBSP syntax structure.
[0166] The Vps_extension_data_flag can have any value. Its presence and value do not affect the decoder's compliance with the profiles specified in this version of this specification. Decoders compliant with this version of this specification shall ignore all vps_extension_data_flag syntax elements.
[0167] 3.6 Access Unit Delimiter (AUD) Syntax and Semantics in VVC
[0168] In the latest VVC text (in JVET-Q2001-Ve / v15), the AUD syntax and semantics are as follows.
[0169] 7.3.2.9 AU Delimiter RBSP Syntax
[0170]
[0171] 7.4.3.9 AU Delimiter RBSP Semantics
[0172] The AU delimiter is used to indicate the start of an AU and the strip types present in the warp-decoded pictures in the AU containing the AU delimiter NAL unit. There is no canonical decoding process associated with the AU delimiter.
[0173] Pic_type indicates that the slice_type values of all strips in the warp-decoded picture in the AU containing the AU delimiter NAL unit are members of the set listed in Table 7 for the given pic_type value. The value of pic_type shall be equal to 0, 1, or 2 in the bitstream compliant with this version of this specification. Other values of pic_type are reserved for future use by ITU-T|ISO / IEC. Decoders compliant with this version of this specification shall ignore the reserved values of pic_type.
[0174] Table 7 - Interpretation of pic_type
[0175] pic_type Possible slice_type values in AU 0 I 1 P, I 2 B, P, I
[0176] 3.7 Order of Aus and PUs
[0177] In the latest VVC text (in JVET-Q2001-Ve / v15), the specification of the decoding order of Aus and PUs is as follows.
[0178] 7.4.2.4.2 Order of Aus and Their Association with CVSs
[0179] The bitstream consists of one or more CVSs.
[0180] A CVS consists of one or more Aus. The order of PUs and their association with Aus are described in Clause 7.4.2.4.3.
[0181] The first AU of the CVS is the CVSS AU, where each existing PU is a CLVSS PU, which is either an IRAP PU with NoOutputBeforeRecoveryFlag equal to 1 or a GDR PU with NoOutputBeforeRecoveryFlag equal to 1.
[0182] Each CVSS AU shall have PUs of each layer present in the CVS.
[0183] 7.4.2.4.3 Order of PUs and Their Association with Aus
[0184] An AU consists of one or more PUs in ascending order of nuh_layer_id. The order of NAL units and the warp-decoded pictures and their association with PUs are described in Clause 7.4.2.4.4.
[0185] There can be at most one AUD NAL unit in an AU. When the AUD NAL unit is present in the AU, it will be the first NAL unit of the AU, and thus it is the first NAL unit of the first PU of the AU.
[0186] There can be at most one EOB NAL unit in an AU. When the EOB NAL unit is present in the AU, it will be the last NAL unit of the AU, and thus it is the last NAL unit of the last PU of the AU.
[0187] When a VCL NAL unit is the first VCL NAL unit following a PH NAL unit and one or more of the following conditions are true, the VCL NAL unit is the first VCL NAL unit of the AU (and thus the PU containing the VCL NAL unit is the first PU of the AU):
[0188] - The value of nuh_layer_id of the VCL NAL unit is less than that of the previous picture in decoding order.
[0189] - The value of ph_pic_order_cnt_lsb of the VCL NAL unit is different from that of the previous picture in decoding order.
[0190] - The PicOrderCntVal derived for the VCL NAL unit is different from that of the previous picture in decoding order.
[0191] Let firstVclNalUnitInAu be the first VCL NAL unit of the AU. The start of a new AU is specified by the first of any of the following NAL units that comes before firstVclNalUnitInAu and after the last VCL NAL unit (if any) that comes before firstVclNalUnitInAu:
[0192] - An AUD NAL unit (when present),
[0193] - A DCI NAL unit (when present),
[0194] - A VPS NAL unit (when present),
[0195] - An SPS NAL unit (when present),
[0196] - A PPS NAL unit (when present),
[0197] - A prefix APS NAL unit (when present),
[0198] - A PH NAL unit (when present),
[0199] - A prefix SEI NAL unit (when present),
[0200] - A NAL unit with nal_unit_type equal to RSV_NVCL_26 (when present),
[0201] - A NAL unit with nal_unit_type in the range UNSPEC28..UNSPEC29 (when present).
[0202] Note — The first NAL unit that comes before firstVclNalUnitInAu and after the last VCL NAL unit that comes before firstVclNalUnitInAu, if any, can only be one of the NAL units listed above.
[0203] The bitstream conformance requirement is that when present, the next PU of a particular layer after a PU that belongs to the same layer and contains an EOS NAL unit should be a CLVSS PU, which is either an IRAP PU with NoOutputBeforeRecoveryFlag equal to 1 or a GDR PU with NoOutputBeforeRecoveryFlag equal to 1.
[0204] 4. Examples of the technical problems solved by the disclosed technical solution
[0205] The existing scalability designs in the latest VVC text (in JVET - Q2001 - Ve / v15) have the following problems:
[0206] 1) The CVSS AU requirement for starting a new CVS is complete (i.e., it is required to have pictures of each layer present in the CVS). However, according to the current design, the decoder cannot check whether the AU includes pictures of "each layer present in the CVS" until it receives the last picture of the CVS. On the other hand, it is not easy to determine even the last picture of the CVS because it is not easy to determine the start of any CVS except for the very first CVS in the bitstream. Basically, this means that the decoder can calculate the boundaries of the CVS only after receiving the entire bitstream.
[0207] 2) Currently, an IRAP AU can start a new CVS and is required to be complete (i.e., it is required to have pictures of each layer present in the CVS), while a GDR AU can also start a new CVS but is not required to be complete. This basically prohibits random access to a bitstream that never starts with a GDR AU that is complete in a compliant manner because usually the starting AU in random access will become a CVSS AU, but when such a GDR AU is incomplete, it cannot be a CVSS AU.
[0208] 5. Example Embodiments and Techniques
[0209] To solve the above - mentioned problems and other problems, the methods described below are disclosed. These inventions should be regarded as examples for explaining general concepts and should not be interpreted in a narrow way. In addition, these inventions can be applied alone or combined in any way.
[0210] Solution to the first problem
[0211] 1) To solve the first problem, an indication of whether the AU is complete can be signaled, i.e., whether the AU includes pictures of each layer present in the CVS.
[0212] a. In one example, the indication is signaled only for AUs that may start a new CVS.
[0213] i. Additionally, in one example, the indication is signaled only for AUs where each picture is an IRAP or GDR picture.
[0214] b. In one example, the indication is signaled in the AUD NAL unit.
[0215] i. In one example, the indication is signaled by a flag (e.g., irap_or_gdr_au_flag) in the AUD NAL unit.
[0216] 1. Alternatively, additionally, when irap_or_gdr_au_flag is equal to 1, a flag (i.e., named irap_au_flag) can be signaled in the AUD to specify whether the AU is an IRAP AU or a GDR AU (irap_au_flag equal to 1 specifies that the AU is an IRAP AU, and irap_au_flag equal to 0 specifies that the AU is a GDR AU). When irap_au_flag does not exist, its value is not inferred.
[0217] ii. Further, in one example, when vps_max_layers_minus1 is greater than 0, there is required to be exactly one AUD NAL unit in each IRAP or GDR AU.
[0218] 1. Alternatively, regardless of the value of vps_max_layers_minus1, there is required to be exactly one AUD NAL unit in each IRAP or GDR AU.
[0219] iii. In one example, a flag equal to 1 specifies that all slices in the AU have the same NAL unit type in the range from IDR_W_RADL to GDR_NUT (including the endpoints). Thus, if the NAL unit type is IDR_W_RADL or IDR_N_LP and the flag is equal to 1, the AU is a CVSS AU. Otherwise (NAL unit type is CRA_NUT or GDR_NUT), when the variable NoOutputBeforeRecoveryFlag for each picture in the AU is equal to 1, the AU is a CVSS AU.
[0220] c. Further, in one example, it is required that each picture in the AU in the CVS should have an nuh_layer_id equal to the nuh_layer_id of one of the pictures present in the first AU of the CVS.
[0221] d. In one example, the indication is signaled in the NAL unit header.
[0222] i. In one example, bits in the NAL unit header (e.g., nuh_reserved_zero_bit) are used to specify whether the AU is complete.
[0223] e. In one example, the indication is signaled in a new NAL unit.
[0224] i. Additionally, in one example, when vps_max_layers_minus1 is greater than 0, it is required that there be one and only one new NAL unit in each IRAP or GDR AU.
[0225] 1. Additionally, in one example, when present in an AU, the new NAL unit shall be in the AU, in decoding order, before any NAL unit other than the AUD NAL unit (when present).
[0226] 2. Alternatively, regardless of the value of vps_max_layers_minus1, it is required that there be one and only one new NAL unit in each IRAP or GDR AU.
[0227] f. In one example, this indication is signaled in an SEI message.
[0228] i. Additionally, in one example, when vps_max_layers_minus1 is greater than 0, it is required that there be an SEI message in each IRAP or GDR AU.
[0229] 1. Additionally, in one example, when present in an AU, the SEI NAL unit containing the SEI message shall be in the AU, in decoding order, before any NAL unit other than the AUD NAL unit (when present).
[0230] 2. Alternatively, regardless of the value of vps_max_layers_minus1, it is required that there be an SEI message in each IRAP or GDR AU.
[0231] 2) Alternatively, to solve the first problem, the indication of whether an AU starts a new CVS can be signaled.
[0232] a. In one example, this indication is signaled in the AUD NAL unit.
[0233] i. In one example, this indication is signaled by a flag (e.g., irap_or_gdr_au_flag) in the AUD NAL unit.
[0234] 1. Alternatively, additionally, when irap_or_gdr_au_flag is equal to 1, a flag (i.e., named irap_au_flag) can be signaled in the AUD to specify whether the AU is an IRAP AU or a GDR AU (irap_au_flag equal to 1 specifies that the AU is an IRAP AU, and irap_au_flag equal to 0 specifies that the AU is a GDR AU). When irap_au_flag is not present, its value is not inferred.
[0235] ii. Additionally, in one example, when vps_max_layers_minus1 is greater than 0, one and only one AUD NAL unit is required to be present in each CVSS AU.
[0236] 1. Alternatively, regardless of the value of vps_max_layers_minus1, one and only one AUD NAL unit is required to be present in each CVSS AU.
[0237] b. In one example, the indication is signaled in the NAL unit header.
[0238] i. In one example, a bit in the NAL unit header (e.g., nuh_reserved_zero_bit) is used to specify whether the AU is a CVSS AU.
[0239] c. In one example, the indication is signaled in a new NAL unit.
[0240] i. Additionally, in one example, when vps_max_layers_minus1 is greater than 0, one and only one new NAL unit is required to be present in each CVSS AU.
[0241] 1. Additionally, in one example, when present in the AU, the new NAL unit shall be before any NAL unit other than the AUD NAL unit (when present) in decoding order within the AU.
[0242] 2. Alternatively, regardless of the value of vps_max_layers_minus1, one and only one new NAL unit is required to be present in each CVSS AU.
[0243] ii. In one example, the presence of the new NAL unit in the AU specifies that the AU is a CVSS AU.
[0244] 1. Alternatively, a flag is included in the new NAL unit to specify whether the AU is a CVSS AU.
[0245] d. In one example, the indication is signaled in an SEI message.
[0246] i. Additionally, in one example, when vps_max_layers_minus1 is greater than 0, an SEI message is required to be present in each CVSS AU.
[0247] 1. Additionally, in one example, when present in an AU, an SEI NAL unit containing an SEI message shall be before any NAL unit other than the AUD NAL unit (when present) in decoding order within the AU.
[0248] 2. Alternatively, regardless of the value of vps_max_layers_minus1, it is required that an SEI message be present in each CVSS AU.
[0249] 3) Alternatively, to solve the first problem, a CVSS AU starting a new CVS is required to have pictures for each layer specified by the VPS. Note that a drawback of this method is that then the number of layers signaled in the VPS needs to be exact, and thus, when a layer is removed from the bitstream, the VPS needs to be modified.
[0250] 4) Alternatively, to solve the first problem, when vps_max_layers_minus1 is greater than 0, the presence of an EOS NAL unit is specified for each picture in the last AU of each CVS, and optionally the presence of an EOB NAL unit is also specified in the last AU of each bitstream. Thus, each CVSS AU will be identified by being the first AU in the bitstream or by the presence of an EOS NAL unit in the previous AU.
[0251] 5) Alternatively, to solve the first problem, variables determined by external means are specified to indicate whether an AU is a CVSS AU.
[0252] Solution to the second problem
[0253] 6) To solve the second problem, it is required that each GDR AU be complete (i.e., have pictures for each layer present in the CVS). This means that an AU consisting of GDR pictures, but if it is incomplete, then it is not a GDR AU, similar to an AU currently consisting of IRAP pictures, but if it is incomplete, then it is not an IRAP AU.
[0254] a. In one example, a GDR AU can be defined as an AU in which there is a PU for each layer in the CVS and the decoded pictures in each present PU are GDR pictures.
[0255] 6. Embodiments
[0256] The following are some example embodiments for some aspects of the present invention summarized above in Section 5, which can be applied to the VVC specification. The changed text is based on the latest VVC text in JVET-Q2001-Ve / v15. Most of the relevant parts added or modified are in Underlined, bold, and italic textIt is shown that most of the relevant parts deleted are highlighted in bold double brackets. For example, [[a]] in the table indicates that "a" has been deleted. There are also some other editorial changes that are not highlighted.
[0257] 6.1 First Embodiment
[0258] This embodiment addresses items 1, 1.a, 1.a.i, 1.b, 1.b.i, 1.b.ii, 1.b.iii, 1.c, 6, and 6a.
[0259] 3 Definitions ...
[0261] Gradual Decoding Refresh (GDR) AU: where there is a PU in each layer of the CVS and the warp-decoded picture in each existing PU is the AU of the GDR picture. ...
[0263] 7.3.2.9 AU Delimiter RBSP Syntax
[0264]
[0265] 7.4.3.9 AU Delimiter RBSP Semantics
[0266] The AU delimiter is used to indicate the start of the AU, , and the slice type present in the coded pictures in the AU containing the AU delimiter NAL unit. , and there is no canonical decoding process associated with the AU delimiter.
[0267] ...
[0269] 7.4.2.4.2 Order of AUs and Their Association with CVS
[0270] The bitstream consists of one or more CVSs.
[0271] A CVS consists of one or more AUs. The order of PUs and their association with AUs are described in Clause 7.4.2.4.3.
[0272] The first AU of the CVS is the CVSS AU, where each existing PU is a CLVSS PU, which is either an IRAP PU with NoOutputBeforeRecoveryFlag equal to 1 or a GDR PU with NoOutputBeforeRecoveryFlag equal to 1.
[0273] Each CVSS AU shall have a PU for each layer present in the CVS,
[0274] 7.4.2.4.3 Order of PUs and Its Association with Au
[0275] An AU is composed of one or more PUs in increasing order of nuh_layer_id. The ordered NAL units and the warp decoded pictures and their association with PUs are described in Clause 7.4.2.4.4.
[0276] There can be at most one AUD NAL unit in an AU.
[0277] When an AUD NAL unit is present in an AU, it will be the first NAL unit of the AU and thus it is the first NAL unit of the first PU of the AU.
[0278] There can be at most one EOB NAL unit in an AU. When an EOB NAL unit is present in an AU, it will be the last NAL unit of the AU and thus it is the last NAL unit of the last PU of the AU. ...
[0280] Figure 1 is a block diagram showing an example video processing system 1000 in which various techniques disclosed herein can be implemented. Various implementations can include some or all components of system 1000. System 1000 can include an input 1002 for receiving video content. The video content can be received in a raw or uncompressed format, e.g., 8-bit or 10-bit multi-component pixel values, or it can be received in a compressed or encoded format. Input 1002 can represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, Passive Optical Network (PON), etc., and wireless interfaces such as Wi-Fi or cellular interfaces.
[0281] System 1000 may include an encoding / decoding component 1004 that may implement various encoding / decoding or encoding methods described in this document. The encoding / decoding component 1004 may reduce the average bit rate of a video from the input 1002 to the output of the encoding / decoding 1004 to produce an encoded / decoded representation of the video. Thus, encoding / decoding techniques are sometimes referred to as video compression or video transcoding techniques. The output of the encoding / decoding component 1004 may be stored or transmitted via a connected communication, as represented by component 1006. The stored or communicated bitstream (or encoded / decoded) representation of the video received at the input 1002 may be used by component 1008 to generate pixel values or a displayable video to be sent to the display interface 1010. The process of generating a user-visible video from the bitstream representation is sometimes referred to as video decompression. Additionally, although some video processing operations are referred to as "encoding" operations or tools, it should be understood that encoding tools or operations are used at the encoder, and the corresponding decoding tools or operations that reverse the encoding result will be performed by the decoder.
[0282] Examples of a peripheral bus interface or a display interface may include a Universal Serial Bus (USB), a High-Definition Multimedia Interface (HDMI), or a DisplayPort, etc. Examples of a storage interface include SATA (Serial Advanced Technology Attachment), PCI, IDE interface, etc. The techniques described in this document may be embodied in various electronic devices, such as a mobile phone, a laptop computer, a smartphone, or other devices capable of performing digital data processing and / or video display.
[0283] Figure 2 is a block diagram of a video processing apparatus 2000. The apparatus 2000 may be used to implement one or more methods described herein. The apparatus 2000 may be embodied in a smartphone, a tablet computer, a computer, an Internet of Things (IoT) receiver, etc. The apparatus 2000 may include one or more processors 2002, one or more memories 2004, and video processing hardware 2006. The processor(s) 2002 may be configured to implement one or more methods described in this document. The memory(ies) 2004 may be used to store data and code for implementing the methods and techniques described herein. The video processing hardware 2006 may be used to implement some of the techniques described in this document in hardware circuitry. In some embodiments, the hardware 2006 may be partially or fully in one or more processors 2002 (e.g., a graphics processor).
[0284] Figure 3 is a block diagram showing an example video encoding / decoding system 100 that may utilize the techniques of the present disclosure.
[0285] As Figure 3As shown, the video encoding and decoding system 100 may include a source device 110 and a destination device 120. The source device 110 generates encoded video data, which may be referred to as a video encoding device. The destination device 120 may decode the encoded video data generated by the source device 110, and the destination device 120 may be referred to as a video decoding device.
[0286] The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.
[0287] The video source 112 may include sources such as a video capture device, an interface for receiving video data from a video content provider, and / or a computer graphics system for generating video data, or a combination of these sources. The video data may include one or more pictures. The video encoder 114 encodes the video data from the video source 112 to generate a bitstream. The bitstream may include a sequence of bits forming an encoded representation of the video data. The bitstream may include encoded pictures and associated data. The encoded pictures are encoded representations of the pictures. The associated data may include sequence parameter sets, picture parameter sets, and other syntax structures. The I / O interface 116 may include a modulator / demodulator (modem) and / or a transmitter. The encoded video data may be sent directly to the destination device 120 via the I / O interface 116 over a network 130a. The encoded video data may also be stored on a storage medium / server 130b for access by the destination device 120.
[0288] The destination device 120 may include an I / O interface 126, a video decoder 124, and a display device 122.
[0289] The I / O interface 126 may include a receiver and / or a modem. The I / O interface 126 may obtain the encoded video data from the source device 110 or the storage medium / server 130b. The video decoder 124 may decode the encoded video data. The display device 122 may display the decoded video data to a user. The display device 122 may be integrated with the destination device 120, or may be external to the destination device 120, and the destination device 120 is configured to interface with an external display device.
[0290] The video encoder 114 and the video decoder 124 may operate according to video compression standards, such as the High Efficiency Video Coding (HEVC) standard, the Versatile Video Coding (VVC) standard, and other current and / or future standards.
[0291] Figure 4 is a block diagram showing an example of a video encoder 200, and the video encoder 200 may be Figure 3 the video encoder 114 in the system 100 shown.
[0292] Video encoder 200 may be configured to perform any or all of the techniques of the present disclosure. In Figure 4 an example, video encoder 200 includes multiple functional components. The techniques described in the present disclosure may be shared among the various components of video encoder 200. In some examples, a processor may be configured to perform any or all of the techniques described in the present disclosure.
[0293] The functional components of video encoder 200 may include a splitting unit 201, a prediction unit 202, a residual generation unit 207, a transformation unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transformation unit 211, a reconstruction unit 212, a buffer 213, and an entropy codec unit 214. Prediction unit 202 may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra prediction unit 206.
[0294] In other examples, video encoder 200 may include more, fewer, or different functional components. In one example, prediction unit 202 may include an Intra Block Copy (IBC) unit. The IBC unit may perform prediction in the IBC mode, in which at least one reference picture is the picture in which the current video block is located.
[0295] Furthermore, some components such as motion estimation unit 204 and motion compensation unit 205 may be highly integrated, but are shown separately in Figure 4 the example for illustrative purposes.
[0296] Splitting unit 201 may split a picture into one or more video blocks. Video encoder 200 and video decoder 300 may support various video block sizes.
[0297] Mode selection unit 203 may select, for example based on error results, one of the intra or inter coding modes and provide the resulting intra or inter coded block to residual generation unit 207 to generate residual block data, and to reconstruction unit 212 to reconstruct the coded block for use as a reference picture. In some examples, mode selection unit 203 may select a Combined Intra Inter Prediction (CIIP) mode, in which the prediction is based on an inter prediction signal and an intra prediction signal. Mode selection unit 203 may also select the resolution of the motion vector for a block in the case of inter prediction (e.g., sub-pixel or integer pixel precision).
[0298] To perform inter prediction on a current video block, the motion estimation unit 204 may generate motion information for the current video block by comparing one or more reference frames from buffer 213 with the current video block. The motion compensation unit 205 may determine a predicted video block for the current video block based on the motion information and decoded samples of pictures from buffer 213 other than the picture associated with the current video block.
[0299] The motion estimation unit 204 and the motion compensation unit 205 may perform different operations on the current video block, for example, depending on whether the current video block is in an I-slice, a P-slice, or a B-slice.
[0300] In some examples, the motion estimation unit 204 may perform uni-directional prediction on the current video block, and the motion estimation unit 204 may search the reference pictures in list 0 or list 1 to obtain a reference video block for the current video block. Then, the motion estimation unit 204 may generate a reference index indicating the reference picture in list 0 or list 1 that contains the reference video block, and a motion vector indicating the spatial displacement between the current video block and the reference video block. The motion estimation unit 204 may output the reference index, the prediction direction indicator, and the motion vector as the motion information for the current video block. The motion compensation unit 205 may generate a predicted video block for the current block based on the reference video block indicated by the motion information of the current video block.
[0301] In other examples, the motion estimation unit 204 may perform bi-directional prediction on the current video block. The motion estimation unit 204 may search the reference pictures in list 0 for a reference video block for the current video block and may also search the reference pictures in list 1 for another reference video block for the current video block. Then, the motion estimation unit 204 may generate reference indexes indicating the reference pictures in list 0 and list 1 that contain the reference video blocks, and a motion vector indicating the spatial displacement between the reference video blocks and the current video block. The motion estimation unit 204 may output the reference indexes and the motion vector of the current video block as the motion information for the current video block. The motion compensation unit 205 may generate a predicted video block for the current video block based on the reference video blocks indicated by the motion information of the current video block.
[0302] In some examples, the motion estimation unit 204 may output the entire set of motion information for the decoding process of the decoder.
[0303] In some examples, the motion estimation unit 204 may not output the entire set of motion information for the current video. Instead, the motion estimation unit 204 may signal the motion information of the current video block by referring to the motion information of another video block. For example, the motion estimation unit 204 may determine that the motion information of the current video block is sufficiently similar to the motion information of an adjacent video block.
[0304] In one example, the motion estimation unit 204 may indicate a value in a syntax structure associated with the current video block, and this value indicates to the video decoder 300 that the current video block has the same motion information as another video block.
[0305] In another example, the motion estimation unit 204 may identify another video block and a motion vector difference (MVD) in a syntax structure associated with the current video block. This motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 300 may use the motion vector of the indicated video block and this motion vector difference to determine the motion vector of the current video block.
[0306] As described above, the video encoder 200 may predictively signal this motion vector. Two examples of predictive signaling techniques that may be implemented by the video encoder 200 include advanced motion vector prediction (AMVP) and merge mode signaling.
[0307] The intra prediction unit 206 may perform intra prediction on the current video block. When the intra prediction unit 206 performs intra prediction on the current video block, the intra prediction unit 206 may generate prediction data for the current video block based on the decoded samples of other video blocks in the same picture. The prediction data for the current video block may include the predicted video block and various syntax elements.
[0308] The residual generation unit 207 may generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) the predicted video block of the current video block from the current video block. The residual data for the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.
[0309] In other examples, such as in the skip mode, there may be no residual data for the current video block, and the residual generation unit 207 may not perform a subtraction operation.
[0310] The transform processing unit 208 may generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video block associated with the current video block.
[0311] After the transform processing unit 208 generates the transform coefficient video block associated with the current video block, the quantization unit 209 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0312] The inverse quantization unit 210 and the inverse transform unit 211 can respectively apply inverse quantization and inverse transform to the transformed coefficient video block to reconstruct the residual video block from the transformed coefficient video block. The reconstruction unit 212 can add the reconstructed residual video block to the corresponding samples of one or more predicted video blocks generated by the prediction unit 202 to generate a reconstructed video block associated with the current block for storage in the buffer 213.
[0313] After the reconstruction unit 212 reconstructs the video block, a loop filtering operation can be performed to reduce video block artifacts in the video block.
[0314] The entropy encoding unit 214 can receive data from other functional components of the video encoder 200. When the entropy encoding unit 214 receives data, the entropy encoding unit 214 can perform one or more entropy encoding operations to generate entropy-encoded data and output a bitstream including the entropy-encoded data.
[0315] Figure 5 is a block diagram showing an example of the video decoder 300, and the video decoder 300 can be Figure 3 the video decoder 114 in the system 100 shown.
[0316] The video decoder 300 can be configured to perform any or all of the techniques of the present disclosure. In Figure 5 the example, the video decoder 300 includes a plurality of functional components. The techniques described in the present disclosure can be shared among the various components of the video decoder 300. In some examples, a processor can be configured to perform any or all of the techniques described in the present disclosure.
[0317] In Figure 5 the example, the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, and a reconstruction unit 306 and a buffer 307. In some examples, the video decoder 300 can perform a decoding pass that is substantially opposite to the encoding pass ( Figure 4 ) described with respect to the video encoder 200.
[0318] The entropy decoding unit 301 can retrieve the encoded bitstream. The encoded bitstream can include entropy-coded decoded video data (e.g., decoded blocks of video data). The entropy decoding unit 301 can decode the entropy-coded decoded video data, and the motion compensation unit 302 can determine motion information from the entropy-decoded video data, which includes motion vectors, motion vector precision, reference picture list indices, and other motion information. The motion compensation unit 302 can determine such information, for example, by performing AMVP and merge modes.
[0319] The motion compensation unit 302 may generate motion-compensated blocks and may perform interpolation, possibly based on an interpolation filter. An identifier for the interpolation filter to be used with sub-pixel accuracy may be included in the syntax element.
[0320] The motion compensation unit 302 may use the interpolation filter used by the video encoder 200 during the encoding of a video block to calculate the interpolation of sub-integer pixels of a reference block. The motion compensation unit 302 may determine the interpolation filter used by the video encoder 200 according to the received syntax information, and use the interpolation filter to generate a prediction block.
[0321] The motion compensation unit 302 may use some syntax information to determine the size of the blocks used for encoding the frames and / or slices of the encoded video sequence, the partitioning information that describes how each macroblock of a picture of the encoded video sequence is partitioned, the mode indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-coded block, and other information for decoding the encoded video sequence.
[0322] The intra prediction unit 303 may use, for example, an intra prediction mode received in the bitstream to form a prediction block from spatially adjacent blocks. The inverse quantization unit 304 inverse quantizes the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 301, i.e., dequantizes. The inverse transform unit 305 applies an inverse transform.
[0323] The reconstruction unit 306 may add the residual block to the corresponding prediction block generated by the motion compensation unit 302 or the intra prediction unit 303 to form a decoded block. If necessary, a deblocking filter may also be applied to filter the decoded block in order to remove blocking artifact. Then the decoded video block is stored in the buffer 307, which provides reference blocks for subsequent motion compensation / intra prediction and also generates the decoded video for presentation on a display device.
[0324] Figures 6 - 7 An example method is shown that may implement the above technical solutions in an embodiment such as Figures 1 - 5 shown in.
[0325] Figure 6 A flowchart 600 of an example method of video processing is shown. The method 600 includes, at operation 610, performing a conversion between a video and a bitstream of the video, the bitstream including an encoded video sequence (CVS), the CVS including one or more access units (AUs), and the bitstream further including a first syntax element indicating whether an AU includes pictures of each video layer that constitutes the CVS.
[0326] Figure 7Flowchart 700 showing an example method of video processing. Method 700 includes, at operation 710, performing a conversion between a video including one or more pictures in one or more video layers and a bitstream of the video, the bitstream including a coded video sequence (CVS), the CVS including one or more access units (AUs), and the bitstream conforming to format rules specifying that each picture in a given AU carries a layer identifier that is equal to the layer identifier of the first AU in the CVS including the one or more video layers.
[0327] Figure 8 Flowchart 800 showing an example method of video processing. Method 800 includes, at operation 810, performing a conversion between a video including one or more pictures in one or more video layers and a bitstream of the video, the bitstream including a coded video sequence (CVS), the CVS including one or more access units (AUs), and the bitstream further including a first syntax element indicating an AU in the one or more AUs that starts a new CVS.
[0328] Figure 9 Flowchart 900 showing an example method of video processing. Method 900 includes, at operation 910, performing a conversion between a video including one or more pictures in one or more video layers and a bitstream of the video, the bitstream including a coded video sequence (CVS), the CVS including one or more access units (AUs), and the bitstream conforming to format rules specifying that a coded video sequence start (CVSS) AU starting a new CVS includes pictures of each video layer specified in a video parameter set (VPS).
[0329] Figure 10 Flowchart 1000 showing an example method of video processing. Method 1000 includes, at operation 1010, performing a conversion between a video including one or more pictures in one or more video layers and a bitstream of the video, the bitstream including a coded video sequence (CVS), the CVS including one or more access units (AUs), and the bitstream conforming to format rules specifying that a given AU is identified as a coded video sequence start (CVSS) AU based on whether the given AU is the first AU in the bitstream or an AU before the given AU includes a sequence end (EOS) network abstraction layer (NAL) unit.
[0330] Figure 11Flowchart 1100 shows an example method of video processing. Method 1100 includes, at operation 1110, performing a conversion between a video including one or more pictures in one or more video layers and a bitstream of the video, the bitstream including a coded video sequence (CVS), the CVS including one or more access units (AUs), and the rule specifying that side information is used to indicate whether an AU in the one or more AUs is a coded video sequence start (CVSS) AU.
[0331] Figure 12 Flowchart 1200 shows an example method of video processing. Method 1200 includes, at operation 1210, performing a conversion between a video including one or more pictures in one or more video layers and a bitstream of the video, the bitstream including a coded video sequence (CVS), the CVS including one or more access units (AUs), and the bitstream conforms to a format rule that specifies that each AU in the one or more AUs that is a progressive decoding refresh AU exactly includes one picture of each video layer present in the CVS.
[0332] Figure 13 Flowchart 1300 shows an example method of video processing. Method 1300 includes, at operation 1310, performing a conversion between a video including one or more pictures in one or more video layers and a bitstream of the video, the bitstream including a coded video sequence, the CVS including one or more access units (AUs), and the bitstream conforms to a format rule that specifies that each AU in the one or more AUs that is an intra random access point AU exactly includes one picture of each video layer present in the CVS.
[0333] Next, a list of preferred solutions for some embodiments is provided.
[0334] P1. A video processing method includes performing a conversion between a video having one or more pictures in one or more video layers and a decoded representation of an encoded version of the video; wherein the decoded representation includes one or more access units (AUs); wherein the decoded representation includes a syntax element if and only if the AU is of a type, and the syntax element indicates whether the AU includes pictures of each video layer that constitutes a decoded video sequence.
[0335] P2. The method according to solution P1, wherein the type includes all AUs.
[0336] P3. The method according to solution P1, wherein the type includes AUs that start a new CVS.
[0337] P4. A method according to solution P1, wherein the syntax element is included in an access unit delimiter network abstraction layer unit.
[0338] P5. A method according to any one of solutions P1 to P4, wherein the syntax element is included in the network abstraction layer unit header.
[0339] P6. A method according to any one of solutions P1 to P4, wherein the syntax element is included in a new network abstraction layer unit.
[0340] P7. A video processing method, comprising performing a conversion between a video having one or more pictures in one or more video layers and a transcoded representation representing an encoded version of the video; wherein the transcoded representation includes one or more access units (AUs); wherein the transcoded representation conforms to a format rule specifying that each picture in a given AU carries a layer identifier equal to the layer identifier of the first AU of the transcoded video sequence.
[0341] P8. A video processing method, comprising performing a conversion between a video having one or more pictures in one or more video layers and a transcoded representation representing an encoded version of the video; wherein the transcoded representation includes one or more access units (AUs); wherein the transcoded representation includes a syntax element if and only if the AU starts a new transcoded video sequence.
[0342] P9. A method according to solution P8, wherein the syntax element is included in an access unit delimiter network abstraction layer unit.
[0343] P10. A method according to any one of solutions P8 to P9, wherein the syntax element is included in the network abstraction layer unit header.
[0344] P11. A method according to any one of solutions P8 to P9, wherein the syntax element is included as supplementary enhancement information.
[0345] P12. A video processing method, comprising performing a conversion between a video having one or more pictures in one or more video layers and a transcoded representation of the video; wherein the transcoded representation includes one or more access units (AUs); wherein the transcoded representation conforms to a format rule specifying that the start AU of the transcoded video sequence includes video pictures of each video layer specified by a video parameter set, or the transcoded representation implicitly signals the format of the start AU based on a rule.
[0346] P13. A video processing method includes performing a conversion between a video including one or more pictures in one or more video layers and a warp - decoded representation of the video; wherein the warp - decoded representation includes one or more access units (AUs); and wherein each AU in the warp - decoded representation conforming to a specified gradual decoding refresh (GDR) type includes format rules for at least one video picture of each video layer of the warp - decoded video sequence.
[0347] P14. The method according to solution P13, wherein each AU of the GDR type further includes prediction units for each layer in the CVS, and the PU includes GDR pictures.
[0348] P15. The method according to any one of solutions 1 to 14, wherein the conversion includes encoding the video into the warp - decoded representation.
[0349] P16. The method according to any one of solutions 1 to 14, wherein the conversion includes decoding the warp - decoded representation to generate pixel values of the video.
[0350] P17. A video decoding apparatus includes a processor configured to implement one or more of the methods described in solutions P1 to P16.
[0351] P18. A video encoding apparatus includes a processor configured to implement one or more of the methods described in solutions P1 to P16.
[0352] P19. A computer program product having computer code stored thereon, which when executed by a processor causes the processor to implement the method described in any one of solutions P1 to P16.
[0353] Next, another list of preferred solutions for some embodiments is provided.
[0354] A1. A video processing method includes performing a conversion between a video including one or more pictures in one or more video layers and a bitstream of the video, wherein the bitstream includes a warp - decoded video sequence, the warp - decoded video sequence includes one or more access units, and wherein the bitstream further includes a first syntax element indicating whether an access unit includes pictures constituting each video layer of the warp - decoded video sequence.
[0355] A2. The method according to solution A1, wherein the access unit is configured to start a new warp - decoded video sequence.
[0356] A3. The method according to solution A2, wherein the pictures of each video layer are intra - random access point pictures or gradual decoding refresh pictures.
[0357] A4. A method according to Solution A1, wherein the first syntax element is included in an access unit delimiter network abstraction layer unit.
[0358] A5. A method according to Solution A4, wherein the first syntax element is a flag indicating whether the access unit containing the access unit delimiter is an intra random access point access unit or a progressive decoding refresh access unit.
[0359] A6. A method according to Solution A5, wherein the first syntax element is irap_or_gdr_au_flag.
[0360] A7. A method according to Solution A4, wherein when a second syntax element indicating the number of video layers specified by a video parameter set is greater than 1, the access unit delimiter network abstraction layer unit is the only access unit delimiter network abstraction layer unit present in each intra random access point access unit or progressive decoding refresh access unit.
[0361] A8. A method according to Solution A7, wherein the second syntax element indicates the maximum allowed number of layers of a reference video parameter set in each warp decoded video sequence.
[0362] A9. A method according to Solution A1, wherein the first syntax element being equal to 1 indicates all slices in an access unit including the same network abstraction layer unit type within the range from IDR_W_RADL to GDR_NUT (including the endpoints).
[0363] A10. A method according to Solution A9, wherein the first syntax element is equal to 1 and the network abstraction layer unit type is IDR_W_RADL or IDR_N_LP indicates that the access unit is a warp decoded video sequence start access unit.
[0364] A11. A method according to Solution A9, wherein a variable for each picture in the access unit being equal to 1 and the network abstraction layer unit type being CRA_NUT or GDR_NUT indicates that the access unit is a warp decoded video sequence start access unit.
[0365] A12. A method according to Solution A11, wherein the variable indicates whether to output a picture in the decoded picture buffer that is before the current picture in decoded order before the picture is recovered.
[0366] A13. A method according to Solution A11, wherein the variable is NoOutputBeforeRecoveryFlag.
[0367] A14. A method according to Solution A1, wherein the first syntax element is included in the network abstraction layer unit header.
[0368] A15. A method according to Solution A1, wherein the first syntax element is included in a new network abstraction layer unit.
[0369] A16. A method according to Solution A15, wherein when a second syntax element indicating the number of video layers specified by a video parameter set is greater than 1, the new network abstraction layer unit is the only network abstraction layer unit present in each intra-frame random access point access unit or progressive decoding refresh access unit.
[0370] A17. A method according to Solution A1, wherein the first syntax element is included in a supplementary enhancement information message.
[0371] A18. A method according to Solution A17, wherein when a second syntax element indicating the number of video layers specified by a video parameter set is greater than 0, the supplementary enhancement information message is the only supplementary enhancement information message present in each intra-frame random access point access unit or progressive decoding refresh access unit.
[0372] A19. A video processing method, comprising performing a conversion between a video including one or more pictures in one or more video layers and a bitstream of the video, wherein the bitstream includes a decoded video sequence, the decoded video sequence includes one or more access units, and wherein the bitstream conforms to a format rule specifying that each picture in a given access unit carries a layer identifier that is equal to the layer identifier of a first access unit in the decoded video sequence including one or more video layers.
[0373] A20. A video processing method includes performing a conversion between a video including one or more pictures in one or more video layers and a bitstream of the video, wherein the bitstream includes a decoded video sequence, the decoded video sequence includes one or more access units, and wherein the bitstream further includes a first syntax element indicating an access unit in one or more access units that starts a new decoded video sequence.
[0374] A21. A method according to Solution A20, wherein the first syntax element is included in an access unit delimiter network abstraction layer unit.
[0375] A22. A method according to Solution A20, wherein the first syntax element is included in a network abstraction layer unit header.
[0376] A23. A method according to Solution A20, wherein the first syntax element is included in a new network abstraction layer unit.
[0377] A24. A method according to Solution A20, wherein the first syntax element is included in a supplementary enhancement information message.
[0378] A25. A video processing method includes performing a conversion between a video including one or more pictures in one or more video layers and a bitstream of the video, where the bitstream includes a decoded video sequence, the decoded video sequence includes one or more access units, and where the bitstream conforms to format rules that specify that a decoded video sequence start access unit that starts a new decoded video sequence includes pictures of each video layer specified in a video parameter set.
[0379] A26. A video processing method includes performing a conversion between a video including one or more pictures in one or more video layers and a bitstream of the video, the bitstream includes a decoded video sequence, the decoded video sequence includes one or more access units, and where the bitstream conforms to format rules that specify identifying a given access unit as a decoded video sequence start access unit based on whether the given access unit is the first access unit in the bitstream or an access unit before the given access unit includes a sequence end network abstraction layer unit.
[0380] A27. A video processing method includes performing a conversion between a video including one or more pictures in one or more video layers and a bitstream of the video based on a rule, where the bitstream includes a decoded video sequence, the decoded video sequence includes one or more access units, and where the rule specifies that side information is used to indicate whether an access unit in the one or more access units is a decoded video sequence start access unit.
[0381] A28. The method according to any one of Solutions A1 to A27, where the conversion includes decoding video from the bitstream.
[0382] A29. The method according to any one of Solutions A1 to A27, where the conversion includes encoding video into the bitstream.
[0383] A30. A method of storing a bitstream representing a video in a computer-readable recording medium includes generating a bitstream from the video according to the method described in any one or more of Solutions A1 to A27; and storing the bitstream in the computer-readable recording medium.
[0384] A31. A video processing apparatus includes a processor configured to implement the method described in any one or more of Solutions A1 to A30.
[0385] A32. A computer-readable medium having instructions stored thereon that, when executed, cause a processor to implement the method described in one or more of Solutions A1 to A30.
[0386] A33. A computer-readable medium storing a bitstream generated according to any one or more of Solutions A1 to A30.
[0387] A34. A video processing apparatus for storing a bitstream, wherein the video processing apparatus is configured to implement the method described in any one or more of Solutions A1 to A30.
[0388] Next, yet another list of preferred solutions of some embodiments is provided.
[0389] B1. A video processing method, including performing a conversion between a video including one or more pictures in one or more video layers and a bitstream of the video, the bitstream including a decoded video sequence, the decoded video sequence including one or more access units, and wherein the bitstream conforms to a format rule that specifies that each of one or more access units as a progressive decoding refresh access unit exactly includes one picture of each video layer present in the decoded video sequence.
[0390] B2. The method according to Solution B1, wherein each access unit as a progressive decoding refresh access unit includes a prediction unit for each layer present in the decoded video sequence, and wherein the prediction unit for each layer includes a decoded picture as a progressive decoding refresh picture.
[0391] B3. A video processing method, including performing a conversion between a video including one or more pictures in one or more video layers and a bitstream of the video, the bitstream including a decoded video sequence, the decoded video sequence including one or more access units, and wherein the bitstream conforms to a format rule that specifies that each of one or more access units as an intra random access point access unit exactly includes one picture of each video layer present in the decoded video sequence.
[0392] B4. The method according to any one of Solutions B1 to B3, wherein the conversion includes decoding the video from the bitstream.
[0393] B5. The method according to any one of Solutions B1 to B3, wherein the conversion includes encoding the video into the bitstream.
[0394] B6. A method of storing a bitstream representing a video in a computer-readable recording medium, including generating a bitstream from the video according to the method described in any one or more of Solutions B1 to B3; and storing the bitstream in the computer-readable recording medium.
[0395] B7. A video processing apparatus, comprising a processor configured to implement the method according to any one or more of Solutions B1 to B6.
[0396] B8. A computer-readable medium having instructions stored thereon, which when executed cause a processor to implement the method described in one or more of Solutions B1 to B6.
[0397] B9. A computer-readable medium storing a bitstream generated according to any one or more of Solutions B1 to B6.
[0398] B10. A video processing apparatus for storing a bitstream, wherein the video processing apparatus is configured to implement the method according to any one or more of Solutions B1 to B6.
[0399] In this document, the term "video processing" may refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm may be applied during the conversion from the pixel representation of a video to the corresponding bitstream representation, or vice versa. For example, the bitstream representation of the current video block may correspond to bits juxtaposed or extended at different positions within the bitstream as defined by the syntax. For example, a macroblock may be encoded based on transformed and coded error residual values and also using bits in the header and other fields in the bitstream. Further, during the conversion, the decoder may parse the bitstream based on a determination as described in the above solutions, knowing that some fields may or may not be present. Similarly, the encoder may determine whether certain syntax fields are included or not included and accordingly generate an encoded representation by including or excluding the syntax fields from the encoded representation.
[0400] The disclosed and other solutions, examples, embodiments, modules, and functional operations described in this document can be implemented in digital electronic circuitry, or in computer software, firmware, or hardware, including the structures disclosed in this document and their structural equivalents, or in combinations of one or more of them. The disclosed and other embodiments can be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer-readable medium for execution by, or to control the operation of, a data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a storage device, a composition of matter affecting a machine-readable propagated signal, or a combination of one or more of them. The term "data processing apparatus" includes all apparatus, devices, and machines for processing data, including, by way of example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, the apparatus can include code that creates an execution environment for the computer program, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them. A propagated signal is an artificially generated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, generated to encode information for transmission to an appropriate receiver apparatus.
[0401] A computer program (also known as a program, software, software application, script, or code) can be written in any form of programming language, including a compiled or interpreted language, and can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored in a part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program, or in multiple coordinated files (e.g., files that store one or more modules, subroutines, or portions of code). A computer program can be deployed to execute on one computer or on multiple computers located at one site or distributed across multiple sites and interconnected by a communication network.
[0402] The processes and logical flows described in this document can be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logical flows can also be performed by, and the apparatus can also be implemented as, special-purpose logic circuitry, e.g., an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit).
[0403] Processors suitable for executing computer programs include, for example, general and special-purpose microprocessors, as well as any one or more processors of any type of digital computer. Generally, a processor will receive instructions and data from a read-only memory or a random access memory or both. The basic elements of a computer are a processor for executing instructions and one or more storage devices for storing instructions and data. Generally, a computer will also include or be operatively coupled to receive data from or transfer data to one or more mass storage devices (such as, for example, magnetic disks, magneto-optical disks, or optical disks) for storing data or both. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices (such as EPROM, EEPROM, and flash memory devices); magnetic disks (such as internal hard disks or removable disks); magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and the memory may be supplemented by, or incorporated in, special-purpose logic circuitry.
[0404] Although this patent document contains many details, these should not be construed as limiting the scope of any subject matter or what may be claimed, but rather as descriptions of features specific to particular embodiments of a particular technology. Certain features described in the context of separate embodiments in this patent document may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented separately in multiple embodiments or in any suitable sub-combination. Moreover, although the above features may be described as acting in certain combinations and even initially claimed as such, in some cases, one or more features from a claimed combination may be deleted from that combination, and the claimed combination may be directed to a sub-combination or a variation of a sub-combination.
[0405] Similarly, although operations are depicted in the drawings in a particular order, this should not be understood as requiring that the operations be performed in the particular order shown or in a sequential order, or that all of the illustrated operations be performed to obtain a desired result. Additionally, the separation of various system components in the embodiments described in this patent document should not be understood as required in all embodiments.
[0406] Only some implementations and examples have been described, and other implementations, enhancements, and variations may be made based on what is described and illustrated in this patent document.
Claims
1. A video processing method, comprising: Performing a conversion between a video including one or more pictures in one or more video layers and a bitstream of the video, wherein the bitstream includes a decoded / encoded video sequence, and the decoded / encoded video sequence includes one or more access units, wherein the bitstream complies with a first format rule, and the first format rule specifies that each of the one or more access units as progressive decoding refresh access units includes at least one picture of each video layer present in the decoded / encoded video sequence, and wherein each of the access units as the progressive decoding refresh access units includes prediction units of each layer present in the decoded / encoded video sequence, and wherein the prediction units of each layer include decoded / encoded pictures as progressive decoding refresh pictures.
2. The method according to claim 1, wherein the progressive decoding refresh access unit exactly includes one picture of each video layer present in the decoded / encoded video sequence.
3. The method according to claim 1 or 2, wherein the bitstream complies with a second format rule, and the second format rule specifies that each of the one or more access units as intra random access point access units exactly includes one picture of each video layer present in the decoded / encoded video sequence.
4. The method according to claim 1 or 2, wherein The conversion includes decoding the video from the bitstream.
5. The method according to claim 1 or 2, wherein The conversion includes encoding the video into the bitstream.
6. An apparatus for processing video data, comprising a processor and a non-transitory memory storing instructions thereon, wherein the instructions, when executed by the processor, cause the processor to: Perform a conversion between a video including one or more pictures in one or more video layers and a bitstream of the video, Among them, The bitstream includes a decoded / encoded video sequence, and the decoded / encoded video sequence includes one or more access units, wherein the bitstream complies with a first format rule, and the first format rule specifies that each of the one or more access units as progressive decoding refresh access units includes at least one picture of each video layer present in the decoded / encoded video sequence, and wherein each of the access units as the progressive decoding refresh access units includes prediction units of each layer present in the decoded / encoded video sequence, and wherein the prediction units of each layer include decoded / encoded pictures as progressive decoding refresh pictures.
7. The apparatus according to claim 6, wherein the progressive decoding refresh access unit exactly includes one picture of each video layer present in the decoded / encoded video sequence.
8. The apparatus according to claim 6 or 7, wherein the bitstream complies with a second format rule, and the second format rule specifies that each of the one or more access units as intra random access point access units exactly includes one picture of each video layer present in the decoded / encoded video sequence.
9. The apparatus according to claim 6 or 7, wherein The conversion includes decoding the video from the bitstream.
10. The device according to claim 6 or 7, wherein, The conversion includes encoding the video into the bitstream.
11. A non-transitory computer-readable storage medium storing instructions that cause a processor to: perform a conversion between a video including one or more pictures in one or more video layers and a bitstream of the video, Among them, wherein the bitstream includes a decoded video sequence, and the decoded video sequence includes one or more access units, wherein the bitstream conforms to a first format rule that specifies that each of the one or more access units that are progressive decoding refresh access units includes at least one picture of each video layer present in the decoded video sequence, and wherein each of the access units that are the progressive decoding refresh access units includes prediction units of each layer present in the decoded video sequence, and wherein the prediction units of each layer include decoded pictures that are progressive decoding refresh pictures.
12. The medium according to claim 11, wherein the progressive decoding refresh access unit exactly includes one picture of each video layer present in the decoded video sequence.
13. The medium according to claim 11 or 12, wherein the bitstream conforms to a second format rule that specifies that each of the one or more access units that are intra random access point access units exactly includes one picture of each video layer present in the decoded video sequence.
14. A non-transitory computer-readable recording medium storing a bitstream of a video generated by a method executed by a video processing device, wherein the method includes: generating a bitstream of a video including one or more pictures in one or more video layers, wherein the bitstream includes a decoded video sequence, and the decoded video sequence includes one or more access units, wherein the bitstream conforms to a first format rule that specifies that each of the one or more access units that are progressive decoding refresh access units includes at least one picture of each video layer present in the decoded video sequence, and wherein each of the access units that are the progressive decoding refresh access units includes prediction units of each layer present in the decoded video sequence, and wherein the prediction units of each layer include decoded pictures that are progressive decoding refresh pictures.
15. The medium according to claim 14, wherein the progressive decoding refresh access unit exactly includes one picture of each video layer present in the decoded video sequence.
16. The medium according to claim 14 or 15, wherein the bitstream conforms to a second format rule that specifies that each of the one or more access units that are intra random access point access units exactly includes one picture of each video layer present in the decoded video sequence.
17. A method of storing a bitstream of a video, including: generating a bitstream of a video including one or more pictures in one or more video layers; and storing the bitstream in a non-transitory computer-readable recording medium, Wherein, the bitstream includes a transcoded video sequence, and the transcoded video sequence includes one or more access units. Wherein, the bitstream conforms to a first format rule, and the first format rule specifies that each of the one or more access units that are progressive decoding refresh access units includes at least one picture of each video layer present in the transcoded video sequence, and wherein each of the access units that are the progressive decoding refresh access units includes prediction units of each layer present in the transcoded video sequence, and wherein the prediction units of each layer include transcoded pictures that are progressive decoding refresh pictures.
Citation Information
Patent Citations
Method and apparatus for video coding and decoding
CN105027567A
Inter-layer prediction for scalable video coding and decoding
CN107431819A