Random Access Point Access Unit in Scalable Video Coding
By introducing different types of access units into the video bitstream and using specific syntax elements, the problem of incomplete video layers in the prior art is solved, and a more efficient and high-quality video decoding process is achieved.
Patent Information
- Application Number
- CN202180022183.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-03-16
- Filing Date
- 2021-03-15
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2041-03-15
AI Technical Summary
The prior art is difficult to effectively process the access unit of the video in scalable video encoding and decoding, resulting in incomplete video layer problems during the decoding process, affecting video quality and decoding efficiency.
The incomplete video layer problem is solved by introducing different access unit types into the video bitstream and using specific syntax elements to indicate whether these access units contain pictures of each video layer that constitutes the codeced video sequence.
The integrity check and management of the video layer is realized, the efficiency and quality of video decoding are improved, and the stability and reliability of video streams are ensured.
Smart Images

Figure CN115299053B_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This is the national phase of International Patent Application No. PCT / US2021 / 022400, filed on March 15, 2021, which claims priority and benefit of U.S. Provisional Patent Application No. 62 / 990,387, filed on March 16, 2020. For all purposes as provided by law, the entire disclosure of the above - mentioned applications is incorporated by reference as part of the disclosure of this application. Technical Field
[0003] This patent document relates to image and video encoding and decoding. Background Art
[0004] In the Internet and other digital communication networks, digital video occupies the largest bandwidth usage. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth demand for digital video use is expected to continue to grow. Summary of the Invention
[0005] The techniques disclosed in this document include configuring different access units in scalable video coding and decoding, which can be used by video encoders and decoders to process the decoded representation of a video using control information for decoding the decoded representation.
[0006] In one example aspect, a video processing method is disclosed. The method includes performing a conversion between a video including one or more pictures in one or more video layers and a bitstream of the video, where the bitstream includes a decoded video sequence, the decoded video sequence includes one or more access units, and where the bitstream further includes a first syntax element indicating whether an access unit includes pictures of each video layer that constitutes the decoded video sequence.
[0007] In another example aspect, another video processing method is disclosed. The method includes performing a conversion between a video including one or more pictures in one or more video layers and a bitstream of the video, where the bitstream includes a decoded video sequence, the decoded video sequence includes one or more access units, and where the bitstream conforms to format rules specifying that each picture in a given access unit carries a layer identifier that is equal to the layer identifier of the first access unit in a decoded video sequence including one or more video layers.
[0008] In yet another example aspect, a video processing method is disclosed. The method includes performing a conversion between a video including one or more pictures in one or more video layers and a bitstream of the video, where the bitstream includes a decoded video sequence, the decoded video sequence includes one or more access units, and where the bitstream further includes a first syntax element indicating an access unit among the one or more access units that starts a new decoded video sequence.
[0009] In yet another example aspect, a video processing method is disclosed. The method includes performing a conversion between a video including one or more pictures in one or more video layers and a bitstream of the video, where the bitstream includes a decoded video sequence, the decoded video sequence includes one or more access units, and where the bitstream conforms to format rules that specify that a decoded video sequence start access unit starting a new decoded video sequence includes pictures of each video layer specified in a video parameter set.
[0010] In yet another example aspect, a video processing method is disclosed. The method includes performing a conversion between a video including one or more pictures in one or more video layers and a bitstream of the video, the bitstream includes a decoded video sequence, the decoded video sequence includes one or more access units, and where the bitstream conforms to format rules that specify that a given access unit is identified as a decoded video sequence start access unit based on whether the given access unit is the first access unit in the bitstream or an access unit before the given access unit includes a sequence end network abstraction layer unit.
[0011] In yet another example aspect, a video processing method is disclosed. The method includes performing a conversion between a video including one or more pictures in one or more video layers and a bitstream of the video based on a rule, where the bitstream includes a decoded video sequence, the decoded video sequence includes one or more access units, and where the rule specifies that side information is used to indicate whether an access unit among the one or more access units is a decoded video sequence start access unit.
[0012] In yet another example aspect, a video processing method is disclosed. The method includes performing a conversion between a video including one or more pictures in one or more video layers and a bitstream of the video, the bitstream includes a decoded video sequence, the decoded video sequence includes one or more access units, and where the bitstream conforms to format rules that specify that each access unit among one or more access units that are stepwise decoded refresh access units exactly includes one picture of each video layer present in the decoded video sequence.
[0013] In yet another example aspect, a video processing method is disclosed. The method includes performing a conversion between a video including one or more pictures in one or more video layers and a bitstream of the video, the bitstream including a transcoded video sequence including one or more access units, and wherein the bitstream conforms to format rules that specify that each access unit among the one or more access units that is an intra random access point access unit includes exactly one picture of each video layer present in the transcoded video sequence.
[0014] In yet another example aspect, a video encoder device is disclosed. The video encoder includes a processor configured to implement the above method.
[0015] In yet another example aspect, a video decoder device is disclosed. The video decoder includes a processor configured to implement the above method.
[0016] In yet another example aspect, a computer-readable medium having code stored thereon is disclosed. The code embodies one of the methods described herein in the form of processor-executable code.
[0017] These and other features are introduced throughout this document. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 is a block diagram illustrating an example video processing system in which various techniques disclosed herein may be implemented.
[0019] Figure 2 is a block diagram of an example hardware platform for video processing.
[0020] Figure 3 is a block diagram illustrating an example video codec system in which some embodiments of the present disclosure may be implemented.
[0021] Figure 4 is a block diagram illustrating an example of an encoder in which some embodiments of the present disclosure may be implemented.
[0022] Figure 5 is a block diagram illustrating an example of a decoder in which some embodiments of the present disclosure may be implemented.
[0023] Figures 6 to 13 shows a flowchart for an example video processing method. DETAILED DESCRIPTION
[0024] For ease of understanding, section headings are used in this document, and the section headings do not limit the applicability of the technologies and embodiments disclosed in each section to that section. Additionally, in some descriptions, the use of H.266 terms is only for ease of understanding and not to limit the scope of the disclosed technologies. Similarly, the technologies described herein are also applicable to other video codec protocols and designs.
[0025] 1. Preliminary Discussion
[0026] This document is related to video coding and decoding technologies. Specifically, it is about the specification and signaling of random access point access units in scalable video coding, where the video bitstream can contain more than one layer. These ideas can be applied individually or in various combinations to any video coding standard or non-standard video codec that supports multi-layer video coding (e.g., Versatile Video Coding (VVC) currently under development).
[0027] 2. Abbreviations
[0028] APS Adaptive Parameter Set
[0029] AU Access Unit
[0030] AUD Access Unit Delimiter
[0031] AVC Advanced Video Coding
[0032] CLVS Video Sequence after Coding and Decoding
[0033] CPB Coding and Decoding Picture Buffer
[0034] CRA Clean Random Access
[0035] CTU Coding Tree Unit
[0036] CVS Decoded Video Sequence
[0037] DCI Decoding Capability Information
[0038] DPB Decoded Picture Buffer
[0039] EOB End of Bitstream
[0040] EOS End of Sequence
[0041] CGR Gradual Decoding Refresh
[0042] HEVC High Efficiency Video Coding
[0043] HRD Hypothetical Reference Decoder
[0044] IDR Instantaneous Decoding Refresh
[0045] JEM Joint Exploration Model
[0046] MCTS Motion Constrained Tile Set
[0047] NAL Network Abstraction Layer
[0048] OLS Output Layer Set
[0049] PH Picture Header
[0050] PPS Picture Parameter Set
[0051] PTL Profile, Tier and Level
[0052] PU Picture Unit
[0053] RAP Random Access Point
[0054] RBSP Raw Byte Sequence Payload
[0055] SEI Supplemental Enhancement Information
[0056] SPS Sequence Parameter Set
[0057] SVC Scalable Video Coding
[0058] VCL Video Coding Layer
[0059] VPS Video Parameter Set
[0060] VTM VVC Test Model
[0061] VUI Video Usability Information
[0062] VVC Versatile Video Coding
[0063] 3. Introduction to Video Coding
[0064] Video coding standards have mainly evolved through the development of well-known ITU-T and ISO / IEC standards. ITU-T developed H.261 and H.263, ISO / IEC developed MPEG-1 and MPEG-4 Visual, and the two organizations jointly developed H.262 / MPEG-2 Video, H.264 / MPEG-4 Advanced Video Coding (AVC) standard, and H.265 / HEVC standard. Since H.262, video coding standards have been based on a hybrid video coding structure, which uses temporal prediction plus transform coding. To explore future video coding technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Team (JVET) in 2015. Since then, many new methods have been adopted by JVET and incorporated into a reference software called the Joint Exploration Model (JEM). At the same time, JVET meetings are held quarterly, and compared with HEVC, the goal of the new coding standard is to reduce the bitrate by 50%. The new video coding standard was officially named Versatile Video Coding (VVC) at the JVET meeting in April 2018, and the first version of the VVC Test Model (VTM) was released at that time. With continuous efforts in VVC standardization, at each JVET meeting, the VVC standard incorporates new coding technologies. The VVC working draft and the test model VTM are updated after each meeting. The VVC project now aims for Feature Complete (FDIS) at the meeting in July 2020.
[0065] 3.1 Random Access and Its Support in HEVC and VVC
[0066] Random access refers to the access and decoding of a bitstream starting from a picture that is not the first picture of the bitstream in decoding order. To support tuning and channel switching in broadcast / multicast and multiparty video conferencing, seeking in local playback and streaming, and stream adaptation in streaming, the bitstream needs to include frequent random access points, which are usually intra-coded pictures, but can also be inter-coded pictures (e.g., in the case of progressive decoding refresh).
[0067] HEVC includes signaling Intra Random Access Point (IRAP) pictures in the NAL unit header by NAL unit type. Three types of IRAP pictures are supported, namely Instantaneous Decoder Refresh (IDR), Clean Random Access (CRA), and Broken Link Access (BLA) pictures. IDR pictures constrain the inter-picture prediction structure not to reference any pictures before the current Group of Pictures (GOP), and are commonly referred to as closed GOP random access points. CRA pictures are less restrictive by allowing some pictures to reference pictures before the current GOP, and in the case of random access, all these pictures are discarded. CRA pictures are commonly referred to as open GOP random access points. For example, during stream switching, BLA pictures typically result from the splicing of two bitstreams or parts thereof at a CRA picture. To enable the system to better use IRAP pictures, a total of six different NAL units are defined to signal the attributes of IRAP pictures, which can be used to better match the stream access point types defined in the ISO Base Media File Format (ISOBMFF) for random access support in Dynamic Adaptive Streaming over HTTP (DASH).
[0068] VVC supports three types of IRAP pictures, two types of IDR pictures (one type with or the other type without associated RADL pictures), and one type of CRA picture. These are basically the same as in HEVC. The BLA picture type in HEVC is not included in VVC, mainly for two reasons: i) The basic function of BLA pictures can be achieved by CRA pictures plus a Sequence End NAL unit, the presence of which indicates that the subsequent pictures start a new CVS in a single-layer bitstream. ii) During the development of VVC, it was desired to specify fewer NAL unit types than in HEVC, such that the NAL unit type field in the NAL unit header uses 5 bits instead of 6 bits to indicate.
[0069] Another key difference between VVC and HEVC in terms of random access support is that GDR is supported in a more standardized way in VVC. In GDR, the decoding of the bitstream can start from an inter-coded picture. Although the entire picture area cannot be decoded correctly at the beginning, after several pictures, the entire picture area will be correct. AVC and HEVC also support GDR, using recovery point SEI messages for signaling GDR random access points and recovery points. In VVC, a new NAL unit type is specified for the indication of GDR pictures, and the recovery point is signaled in the picture header syntax structure. CVS and the bitstream are allowed to start with GDR pictures. This means that the entire bitstream is allowed to contain only inter-coded pictures without a single intra-coded picture. The main benefit of specifying GDR support in this way is to provide a compliant behavior for GDR. GDR enables the encoder to smooth the bitrate of the bitstream by distributing intra-coded strips or blocks across multiple pictures instead of intra-coding the entire picture, thus allowing a significant reduction in end-to-end latency, which is considered more important today than before because ultra-low latency applications such as wireless displays, online games, and drone-based applications have become more popular.
[0070] Another feature related to GDR in VVC is virtual boundary signaling. At the pictures between a GDR picture and its recovery point, the boundary between the refreshed area (i.e., the correctly decoded area) and the non-refreshed area can be signaled as a virtual boundary, and when signaled, loop filtering across the boundary will not be applied, so decoding mismatches for some samples at or near the boundary will not occur. This can be useful when an application determines to display the correctly decoded area during the GDR process.
[0071] IRAP pictures and GDR pictures can be collectively referred to as random access point (RAP) pictures.
[0072] 3.2 General Scalable Video Coding (SVC) and Scalable Video Coding in VVC
[0073] Scalable Video Coding (SVC, sometimes also referred to simply as scalability in video coding) refers to video coding in which a base layer (BL) (sometimes referred to as a reference layer (RL)) and one or more scalable enhancement layers (El) are used. In SVC, the base layer can carry video data with a basic quality level. One or more enhancement layers can carry additional video data to support, for example, higher spatial, temporal, and / or signal-to-noise ratio (SNR) levels. Enhancement layers can be defined relative to previously encoded layers. For example, the bottom layer can be used as the BL, and the top layer can be used as the EL. Intermediate layers can be used as El or RL or both. For example, an intermediate layer (e.g., a layer that is neither the lowest nor the highest layer) can be an EL for layers below the intermediate layer (such as the base layer or any intermediate enhancement layer) and at the same time be used as an RL for one or more enhancement layers above the intermediate layer. Similarly, in the multi-view or 3D extension of the HEVC standard, there can be multiple views, and the information of one view can be used to code (e.g., encode or decode) the information of another view (such as motion estimation, motion vector prediction, and / or other redundancies).
[0074] In SVC, the parameters used by the encoder or decoder are grouped into parameter sets based on the coding levels at which they can be exploited (e.g., video level, sequence level, picture level, slice level, etc.). For example, the parameters that can be used by one or more coded video sequences of different layers in the bitstream can be included in a Video Parameter Set (VPS), while the parameters used by one or more pictures in the coded video sequence can be included in a Sequence Parameter Set (SPS). Similarly, the parameters exploited by one or more slices in a picture can be included in a Picture Parameter Set (PPS), and other parameters specific to a single slice can be included in the slice header. Similarly, an indication of which (which) parameter set a particular layer is using at a given time can be provided at various coding levels.
[0075] Due to the support for reference picture resampling (RPR) in VVC, it is possible to design support for bitstreams containing multiple layers (e.g., two layers with SD and HD resolutions in VVC) without any additional signal processing stage codec tools, because the upsampling required for spatial scalability support can use only the RPR upsampling filter. However, for scalability support, advanced syntax changes are required (compared to non-scalability support). Scalability support is specified in VVC version 1. Different from the scalability support in any previous video codec standard (including extensions of AVC and HEVC), the scalability design of VVC is as friendly as possible to single-layer decoder design. The decoding capabilities of the multi-layer bitstream are specified in a way that is as if there were only a single layer in the bitstream. For example, decoding capabilities such as DPB size are specified in a way that is independent of the number of layers in the bitstream to be decoded. Basically, a decoder designed for a single-layer bitstream can decode a multi-layer bitstream with little change. Compared with the multi-layer extension design of AVC and HEVC, the HLS aspect has been significantly simplified, but at the cost of some flexibility. For example, it is required that the IRAP AU contains pictures of each layer present in the CVS.
[0076] 3.3 Parameter Sets
[0077] AVC, HEVC, and VVC specify parameter sets. The types of parameter sets include SPS, PPS, APS, and VPS. All of AVC, HEVC, and VVC support SPS and PPS. VPS was introduced starting from HEVC and is included in both HEVC and VVC. APS is not included in AVC or HEVC but is included in the latest VVC draft text.
[0078] SPS is designed to carry sequence-level header information, while PPS is designed to carry picture-level header information that does not change frequently. Using SPS and PPS, there is no need to repeat the information that does not change frequently for each sequence or picture, thus avoiding redundant signaling of this information. In addition, the use of SPS and PPS enables out-of-band transmission of important header information, which not only avoids the need for redundant transmission but also improves the error resilience.
[0079] The introduction of VPS is to carry sequence-level header information that is common to all layers in a multi-layer bitstream.
[0080] The introduction of APS is to carry such picture-level or slice-level information that requires a considerable number of bits for encoding and decoding, can be shared by multiple pictures, and can have quite a lot of different variations in a sequence.
[0081] 3.4 Related Definitions in VVC
[0082] The relevant definitions in the latest VVC text (in JVET-Q2001-Ve / v15) are as follows.
[0083] The previous IRAP picture that has the same nuh_layer_id value as a specific picture in decoding order (when present).
[0084] An AU sequence that consists, in decoding order, of a CVSS AU followed by zero or more AUs that are not CVSS AUs (including all subsequent AUs up to but not including any subsequent AU that is a CVSS AU).
[0085] An AU that has PUs for each layer in CVS and the warp-decoded picture in each PU is a CLVSS picture.
[0086] An AU in which the warp-decoded picture in each present PU is a GDR picture.
[0087] A PU in which the warp-decoded picture is a GDR picture.
[0088] Each VCL NAL unit has a picture with a nal_unit_type equal to GDR_NUT.
[0089] An AU that has PUs for each layer in CVS and the warp-decoded picture in each PU is an IRAP picture.
[0090] All VCL NAL units have warp-decoded pictures with the same nal_unit_type value within the range from IDR_W_RADL to CRA_NUT (including the endpoints).
[0091] 3.5 VPS Syntax and Semantics in VVC
[0092] VVC supports scalability, also known as scalable video coding, in which multiple layers can be encoded in a single warp-decoded video bitstream.
[0093] In the latest VVC text (in JVET-Q2001-Ve / v15), scalability information is signaled in the VPS, and its syntax and semantics are as follows.
[0094] 7.3.2.2 Video Parameter Set Syntax
[0095]
[0096]
[0097]
[0098]
[0099] 7.4.3.2 Video Parameter Set RBSP Semantics
[0100] The VPS RBSP shall be available for the decoding process before being referenced, including in at least one AU with TemporalId equal to 0 or provided by external means.
[0101] All VPS NAL units with a specific vps_video_parameter_set_id value in a CVS shall have the same content.
[0102] Provide an identifier for the VPS for reference by other syntax elements. The value of vps_video_parameter_set_id shall be greater than 0.
[0103] Adding 1 specifies the maximum allowed number of layers in each CVS that references the VPS.
[0104] Adding 1 specifies the maximum number of temporal sublayers that can exist in the layers of each CVS that references the VP. The value of vps_max_sublayers_minus1 shall be in the range of 0 to 6 (including the endpoints).
[0105] Equal to 1 Specifies that the number of temporal sublayers is the same for all layers in each CVS that references the VPS. A Vps_all_layers_same_num_sublayers_flag equal to 0 specifies that the layers in each CVS that references the VPS may or may not have the same number of temporal sublayers. If not present, the value of vps_all_layers_same_num_sublayers_flag is inferred to be equal to 1.
[0106] Equal to 1 Specifies that all layers in the CVS are independently coded and decoded without using inter-layer prediction. A Vps_all_independent_layers_flag equal to 0 specifies that one or more layers in the CVS may use inter-layer prediction. If not present, the value of vps_all_independent_layers_flag is inferred to be equal to 1.
[0107] Specify the nuh_layer_id value for the i-th layer. For any two non-negative integer values of m and n, when m is less than n, the value of vps_layer_id[m] should be less than the value of vps_layer_id[n].
[0108] equal to 1 Specify that the layer with index I does not use inter-layer prediction. Vps_independent_layer_flag[I] equal to 0 specifies that the layer with index I can use inter-layer prediction, and for the j syntax elements vps_direct_ref_layer_flag[I][j] in the range from 0 to I-1 (including the endpoints), they exist in the VPS. If not, the value of vps_independent_layer_flag[I] is inferred to be equal to 1.
[0109] equal to 0 Specify that the layer with index j is not a direct reference layer of the layer with index i. vps_direct_ref_layer_flag[I][j] equal to 1 specifies that the layer with index j is a direct reference layer of the layer with index i. When vps_direct_ref_layer_flag[I][j] does not exist for I and j in the range from 0 to vps_max_layers_minus1 (including the endpoints), it is inferred to be equal to 0. When vps_independent_layer_flag[I] is equal to 0, there is at least one j value in the range from 0 to I-1 (including the endpoints) such that the value of vps_direct_ref_layer_flag[I][j] is equal to 1.
[0110] The variables NumDirectRefLayers[I], DirectRefLayerIdx[I][d], NumRefLayers[I], RefLayerIdx[I][r], and LayerUsedAsRefLayerFlag[j] are derived as follows:
[0111]
[0112]
[0113] Specify the layer index variable GeneralLayerIdx[I] for the layer with nuh_layer_id equal to vps_layer_id[I], which is derived as follows:
[0114] for(I = 0; I <= vps_max_layers_minus1; i++) (38)
[0115] GeneralLayerIdx[vps_layer_id[I]] = i
[0116] For any two different values of I and j, both within the range of 0 to vps_max_layers_minus1 (including the endpoints), when dependencyFlag[I][j] is equal to 1, the requirement for bitstream consistency is that the values of chroma_format_idc and bit_depth_minus8 applied to layer i should be equal to the values of chroma_format_idc and bit_depth_minus8 applied to layer j, respectively.
[0117] equal to 1 Specifies the existence of the syntax element max_tid_il_ref_pics_plus1[I]. Max_tid_ref_present_flag[I] equal to 0 specifies the non-existence of the syntax element max_tid_il_ref_pics_plus1[I].
[0118] Specifies that non-IRAP pictures of layer i do not use inter-layer prediction. Max_tid_il_ref_pics_plus[I] greater than 0 specifies that, for decoding pictures of layer i, no pictures with TemporalId greater than max_tid_il_ref_pics_plus1[I] - 1 are used as ILRP. If it does not exist, the value of max_tid_il_ref_pics_plus1[I] is inferred to be equal to 7.
[0119] equal to 1 Specifies that each OLS contains only one layer, and each layer in the CVS of the reference VPS is itself an OLS, and the single layer it contains is the only output layer. Each_layer_is_an_ols_flag equal to 0 means that an OLS can contain more than one layer. If vps_max_layers_minus1 is equal to 0, the value of each_layer_is_an_ols_flag is inferred to be equal to 1. Otherwise, when vps_all_independent_layers_flag is equal to 0, the value of each_layer_is_an_ols_flag is inferred to be equal to 0.
[0120] equal to 0 Specify that the total number of OLSs specified by the VPS is equal to vps_max_layers_minus1 + 1. The i-th OLS includes the layers with layer indices from 0 to i (including the endpoints), and for each OLS, only the highest layer in the OLS is output.
[0121] An ols_mode_idc equal to 1 specifies that the total number of OLSs specified by the VPS is equal to vps_max_layers_minus1 + 1. The i-th OLS includes the layers with layer indices from 0 to i (including the endpoints), and for each OLS, all layers in the OLS are output.
[0122] An ols_mode_idc equal to 2 specifies that the total number of OLSs specified by the VPS is signaled explicitly, and for each OLS, the output layers are signaled explicitly, while the other layers are direct or indirect reference layers of the output layers of the OLS.
[0123] The value of ols_mode_idc shall be in the range of 0 to 2 (including the endpoints). The value 3 of ols_mode_idc is reserved for future use by ITU-T|ISO / IEC.
[0124] When vps_all_independent_layers_flag is equal to 1 and each_layer_is_an_ols_flag is equal to 0, the value of ols_mode_idc is inferred to be equal to 2.
[0125] plus 1 specifies the total number of OLSs specified by the VPS when ols_mode_idc is equal to 2.
[0126] The variable TotalNumOlss that specifies the total number of OLSs specified by the VPS is derived as follows:
[0127]
[0128] equal to 1 specifies that when ols_mode_idc is equal to 2, the layer with nuh_layer_id equal to vps_layer_id[j] is the output layer of the i-th OLS. An Ols_output_layer_flag[I][j] equal to 0 specifies that when ols_mode_idc is equal to 2, the layer with nuh_layer_id equal to vps_layer_id[j] is not the output layer of the i-th OLS.
[0129] The variable NumOutputLayersInOls[I] that specifies the number of output layers in the i-th OLS, the variable NumSubLayersInLayerInOLS[I][j] that specifies the number of sub-layers in the j-th layer in the i-th OLS, the variable OutputLayerIdInOls[I][j] that specifies the nuh_layer_id value of the j-th output layer in the I-th OLS, and the variable LayerUsedAsOutputLayerFlag[k] that specifies whether the k-th layer is used as an output layer in at least one OLS are derived as follows:
[0130]
[0131]
[0132]
[0133] For each value of I in the range from 0 to vps_max_layers_minus1 (including the endpoints), the values of LayerUsedAsRefleayerFlag[I] and LayerUsedAsoutputLayerFlag[I] cannot both be equal to 0. In other words, there should not be a layer that is neither an output layer of at least one OLS nor a direct reference layer of any other layer.
[0134] For each OL, at least one layer is an output layer. In other words, for any value of I in the range from 0 to TotalNumOlss - 1 (including the endpoints), the value of NumOutputLayersInOls[I] should be greater than or equal to 1.
[0135] The variable NumLayersInOls[I] that specifies the number of layers in the i-th OLS and the variable LayerIdInOls[I][j] that specifies the nuh_layer_id value of the j-th layer in the i-th OLS are derived as follows:
[0136]
[0137]
[0138] Note 1 - The 0-th OLS only contains the lowest layer (i.e., the layer with nuh_layer_id equal to vps_layer_id[0]), and for the 0-th OLS, the only layer contained is output.
[0139] The variable OlsLayerIdx[I][j] that specifies the OLS layer index of the layer with nuh_layer_id equal to LayerIdInOls[I][j] is derived as follows:
[0140]
[0141] The lowest layer in each OLS should be an independent layer. In other words, for each I value in the range from 0 to TotalNumOlss - 1 (including the endpoints), the value of vps_independent_layer_flag[GeneralLayerIdx[LayerIdInOls[I[0]]] should be equal to 1.
[0142] Each layer should be included in at least one OLS specified by the VPS. In other words, for each layer with a specific value of nuh_layer_id nuhLayerId that is one of vps_layer_id[k] where k is in the range from 0 to vps_max_layers_minus1 (including the endpoints), there should be at least one pair of values of I and j, where I is in the range from 0 to TotalNumOlss - 1 (including the endpoints) and j is in the range from NumLayersInOls[I] - 1 (including the endpoints), such that the value of LayerIdInOls[I][j] is equal to nuhLayerId.
[0143] Adding 1 specifies the number of profile_tier_level() syntax structures in the VPS. The value of vps_num_ptls_minus1 should be less than TotalNumOlss.
[0144] equal to 1 Specifies the presence of profile, tier, and general constraint information in the i-th profile_tier_level() syntax structure in the VPS. Pt_present_flag[I] equal to 0 specifies the absence of profile, tier, and general constraint information in the i-th profile_tier_level() syntax structure in the VPS. The value of pt_present_flag[0] is inferred to be equal to 1. When pt_present_flag[I] is equal to 0, the profile, tier, and general constraint information of the i-th profile_tier_level() syntax structure in the VPS is inferred to be the same as that of the (I - 1)-th profile_tier_level() syntax structure in the VPS.
[0145] Specify the TemporalId representing the highest sublayer with level information in the i-th profile_tier_level() syntax structure in the VPS. The value of ptl_max_temporal_id[I] should be in the range of 0 to vps_max_sublayers_minus1 (including the endpoints). When vps_max_sublayers_minus1 is equal to 0, the value of ptl_max_temporal_id[I] is inferred to be equal to 0. When vps_max_sublayers_minus1 is greater than 0 and vps_all_layers_same_num_sublayers_flag is equal to 1, the value of ptl_max_temporal_id[I] is inferred to be equal to vps_max_sublayers_minus1.
[0146] Should be equal to 0.
[0147] Specify the index of the profile_tier_level() syntax structure list in the VPS that is applied to the i-th OLS. When present, the value of ols_ptl_idx[I] should be in the range of 0 to vps_num_ptls_minus1 (including the endpoints). When vps_num_ptls_minus1 is equal to 0, the value of ols_ptl_idx[I] is inferred to be equal to 0.
[0148] When NumLayersInOls[I] is equal to 1, the profile_tier_level() syntax structure applied to the i-th OLS also exists in the SPS referenced by the layer in the i-th OLS. The requirement for bitstream consistency is that when NumLayersInOls[I] is equal to 1, the profile_tier_level() syntax structures signaled in the VPS and SPS for the i-th OLS should be the same.
[0149] Specify the number of dpb_parameters() syntax structures in the VPS. The value of vps_num_dpb_params should be in the range of 0 to 16 (including the endpoints). If not present, the value of vps_num_dpb_params is inferred to be equal to 0.
[0150] Used to control the presence of the max_dec_pic_buffering_minus1[], max_num_reorder_pics[], and max_latency_increase_plus1[] syntax elements in the dpb_parameters() syntax structure in the VPS. When not present, vps_sub_dpb_params_info_present_flag is inferred to be equal to 0.
[0151] Specifies the TemporalId of the highest sublayer representation in the i-th dpb_parameters() syntax structure for which DPB parameters may be present in the VPS. The value of dpb_max_temporal_id[I] should be in the range of 0 to vps_max_sublayers_minus1, inclusive. When vps_max_sublayers_minus1 is equal to 0, the value of dpb_max_temporal_id[I] is inferred to be equal to 0. When vps_max_sublayers_minus1 is greater than 0 and vps_all_layers_same_num_sublayers_flag is equal to 1, the value of dpb_max_temporal_id[I] is inferred to be equal to vps_max_sublayers_minus1.
[0152] Specifies the width of each picture storage buffer of the ith OLS in units of luma samples.
[0153] Specifies the height of each picture storage buffer of the ith OLS in units of luma samples.
[0154] Specifies the index into the list of dpb_parameters() syntax structures in the VPS that applies to the dpb_parameters() syntax structure for the ith OLS when NumLayersInOls[I] is greater than 1. When present, the value of ols_dpb_params_idx[I] shall be in the range of 0 to vps_num_dpb_params-1, inclusive. When ols_dpb_params_idx[I] is not present, the value of ols_dpb_params_idx[I] is inferred to be equal to 0.
[0155] When NumLayersInOls[I] is equal to 1, the dpb_parameters() syntax structure applied to the i-th OLS exists in the SPS referenced by the layer in the i-th OLS.
[0156] equal to 1 Specifies that the VPS contains the general_hrd_parameters() syntax structure and other HRD parameters. The Vps_general_hrd_params_present_flag equal to 0 specifies that the VPS does not contain the general_hrd_parameters() syntax structure or other HRD parameters. If not present, the value of vps_general_hrd_params_present_flag is inferred to be equal to 0.
[0157] When NumLayersInOls[I] is equal to 1, the general_hrd_parameters() syntax structure and the ols_hrd_parameters() syntax structure applied to the i-th OLS exist in the SPS referenced by the layer in the i-th OLS.
[0158] equal to 1 Specifies that the i-th ols_hrd_parameters() syntax structure in the VPS contains the HRD parameters of the sublayer representation with TemporalId in the range of 0 to hrd_max_tid[I] (including the endpoints). The Vps_sublayer_cpb_params_present_flag equal to 0 specifies that the i-th ols_hrd_parameters() syntax structure in the VPS contains the HRD parameters of the sublayer representation with TemporalId equal to only hrd_max_tid[I]. When vps_max_sublayers_minus1 is equal to 0, the value of vps_sublayer_cpb_params_present_flag is inferred to be equal to 0.
[0159] When vps_sublayer_cpb_params_present_flag is equal to 0, the HRD parameters represented by the sublayer with a TemporalId in the range of 0 to hrd_max_tid[I] - 1 (including the endpoints) are inferred to be the same as the HRD parameters represented by the sublayer with a TemporalId equal to hrd_max_tid[I]. These parameters include the HRD parameters starting from the fixed_pic_rate_general_flag[I] syntax element until the sublayer_hrd_parameters(I) syntax structure immediately following the condition "if(general_vcl_hrd_params_present_flag)" in the ols_hrd_parameters syntax structure.
[0160] plus1 specifies the number of ols_hrd_params() syntax structures present in the VPS when vps_general_hrd_params_present_flag is equal to 1. The value of num_ols_hrd_params_minus1 shall be in the range of 0 to TotalNumOlss - 1 (including the endpoints).
[0161] Specifies the TemporalId of the highest sublayer representation whose HRD parameters are included in the i-th ols_hrd_parameters() syntax structure. The value of hrd_max_tid[I] should be in the range of 0 to vps_max_sublayers_minus1 (including the endpoints). When vps_max_sublayers_minus1 is equal to 0, the value of hrd_max_tid[I] is inferred to be equal to 0. When vps_max_sublayers_minus1 is greater than 0 and vps_all_layers_same_num_sublayers_flag is equal to 1, the value of hrd_max_tid[I] is inferred to be equal to vps_max_sublayers_minus1.
[0162] Specifies the index in the list of ols_hrd_parameters() syntax structures in the VPS for the ols_hrd_parameters() syntax structure applied to the i-th OLS when NumLayersInOls[I] is greater than 1. The value of ols_hrd_idx[I] should be in the range of 0 to num_ols_hrd_params_minus1 (including the endpoints).
[0163] When NumLayersInOls[I] is equal to 1, the ols_hrd_parameters() syntax structure applied to the i-th OLS exists in the SPS referenced by the layer in the i-th OLS.
[0164] If the value of num_ols_hrd_param_minus1 + 1 is equal to TotalNumOlss, the value of ols_hrd_idx[I] is inferred to be equal to i. Otherwise, when NumLayersInOls[I] is greater than 1 and num_ols_hrd_params_minus1 is equal to 0, the value of ols_hrd_idx[[I] is inferred to be equal to 0.
[0165] Equal to 0 Specifies that the vps_extension_data_flag syntax element does not exist in the VPS RBSP syntax structure. A Vps_extension_flag equal to 1 specifies that the vps_extension_data_flag syntax element exists in the VPS RBSP syntax structure.
[0166] It can have any value. Its presence and value do not affect the profiles specified in this version of the decoder conforming to this specification. Decoders conforming to this version of the specification shall ignore all vps_extension_data_flag syntax elements.
[0167] 3.6 Access Unit Delimiter (AUD) Syntax and Semantics in VVC
[0168] In the latest VVC text (in JVET-Q2001-Ve / v15), the AUD syntax and semantics are as follows.
[0169] 7.3.2.9 AU Delimiter RBSP Syntax
[0170]
[0171] 7.4.3.9 AU Delimiter RBSP Semantics
[0172] The AU delimiter is used to indicate the start of an AU and the slice types present in the warp-decoded pictures in the AU containing the AU delimiter NAL unit. There is no specified decoding process associated with the AU delimiter.
[0173] Indicate that the slice_type values of all slices of the warp-decoded picture in the AU containing the AU separator NAL unit are members of the set listed in Table 7 for the given pic_type value. The value of pic_type shall be equal to 0, 1, or 2 in the bitstream of this version that complies with this specification. Other values of pic_type are reserved for future use by ITU-T|ISO / IEC. The decoder that complies with this version of this specification shall ignore the reserved values of pic_type.
[0174] Table 7 - Interpretation of pic_type
[0175]
[0176] 3.7 Order of Au and PU
[0177] In the latest VVC text (in JVET-Q2001-Ve / v15), the specification of the decoding order of Au and PU is as follows.
[0178] 7.4.2.4.2 Order of Au and Its Association with CVS
[0179] The bitstream consists of one or more CVSs.
[0180] A CVS consists of one or more Aus. The order of PUs and their association with Au are described in Clause 7.4.2.4.3.
[0181] The first AU of a CVS is the CVSS AU, where each existing PU is a CLVSS PU, which is either an IRAP PU with NoOutputBeforeRecoveryFlag equal to 1 or a GDR PU with NoOutputBeforeRecoveryFlag equal to 1.
[0182] Each CVSS AU shall have PUs of each layer present in the CVS.
[0183] 7.4.2.4.3 Order of PUs and Their Association with Au
[0184] An AU consists of one or more PUs in ascending order of nuh_layer_id. The order of NAL units and the warp-decoded picture and their association with PUs are described in Clause 7.4.2.4.4.
[0185] There can be at most one AUD NAL unit in an AU. When the AUD NAL unit exists in the AU, it will be the first NAL unit of the AU, and thus it is the first NAL unit of the first PU of the AU.
[0186] There can be at most one EOB NAL unit in an AU. When an EOB NAL unit exists in an AU, it will be the last NAL unit of the AU, and thus it is the last NAL unit of the last PU of the AU.
[0187] When a VCL NAL unit is the first VCL NAL unit following a PH NAL unit and one or more of the following conditions are true, the VCL NAL unit is the first VCL NAL unit of the AU (and thus the PU containing the VCL NAL unit is the first PU of the AU):
[0188] - The value of nuh_layer_id of the VCL NAL unit is less than that of the previous picture in decoding order.
[0189] - The value of ph_pic_order_cnt_lsb of the VCL NAL unit is different from that of the previous picture in decoding order.
[0190] - The PicOrderCntVal derived for the VCL NAL unit is different from that of the previous picture in decoding order.
[0191] Let firstVclNalUnitInAu be the first VCL NAL unit of the AU. The first of any of the following NAL units before firstVclNalUnitInAu and after the last VCL NAL unit (if any) before firstVclNalUnitInAu specifies the start of a new AU:
[0192] - AUD NAL unit (when present),
[0193] - DCI NAL unit (when present),
[0194] - VPS NAL unit (when present),
[0195] - SPS NAL unit (when present),
[0196] - PPS NAL unit (when present),
[0197] - Prefix APS NAL unit (when present),
[0198] - PH NAL unit (when present),
[0199] - Prefix SEI NAL unit (when present),
[0200] - The NAL unit with - nal_unit_type equal to RSV_NVCL_26 (when present),
[0201] - The NAL unit with - nal_unit_type in the range UNSPEC28..UNSPEC29 (when present).
[0202] Note - The first NAL unit before firstVclNalUnitInAu and after the last VCL NAL unit before firstVclNalUnitInAu, if any, can only be one of the NAL units listed above.
[0203] The requirement for bitstream conformance is that when present, the next PU of a particular layer after a PU belonging to the same layer and containing an EOS NAL unit should be a CLVSS PU, which is either an IRAP PU with NoOutputBeforeRecoveryFlag equal to 1 or a GDR PU with NoOutputBeforeRecoveryFlag equal to 1.
[0204] 4. Examples of Technical Problems Solved by the Disclosed Technical Solutions
[0205] The existing scalability design in the latest VVC text (in JVET - Q2001 - Ve / v15) has the following problems:
[0206] 1) The requirement for a CVSS AU starting a new CVS is that it is complete (i.e., it is required to have pictures of each layer present in the CVS). However, according to the current design, the decoder cannot check whether an AU includes pictures of "each layer present in the CVS" before receiving the last picture of the CVS. On the other hand, it is not easy to determine even the last picture of the CVS because it is not easy to determine the start of any CVS except for the very first CVS in the bitstream. Basically, this means that the decoder can calculate the boundaries of the CVS only after receiving the entire bitstream.
[0207] 2) Currently, an IRAP AU can start a new CVS and is required to be complete (i.e., it is required to have pictures of each layer present in the CVS), while a GDR AU can also start a new CVS but is not required to be complete. This basically prohibits random access to a bitstream that never starts with a GDR AU that is complete in a compliant manner because usually the starting AU in random access will become a CVSS AU, but when such a GDR AU is incomplete, it cannot be a CVSS AU.
[0208] 5. Example Embodiments and Techniques
[0209] To address the above and other issues, the following methods are disclosed. These inventions should be regarded as examples for explaining general concepts and should not be construed in a narrow manner. Additionally, these inventions can be applied individually or combined in any way.
[0210] Solution to the first problem
[0211] 1) To address the first issue, an indication of whether the AU is complete can be signaled, i.e., whether the AU includes pictures of each layer present in the CVS.
[0212] a. In one example, the indication is signaled only for AUs that may start a new CVS.
[0213] i. Additionally, in one example, the indication is signaled only for AUs where each picture is an IRAP or GDR picture.
[0214] b. In one example, the indication is signaled in the AUD NAL unit.
[0215] i. In one example, the indication is signaled by a flag (e.g., irap_or_gdr_au_flag) in the AUD NAL unit.
[0216] 1. Alternatively, additionally, when irap_or_gdr_au_flag is equal to 1, a flag (i.e., named irap_au_flag) can be signaled in the AUD to specify whether the AU is an IRAP AU or a GDR AU (irap_au_flag equal to 1 specifies that the AU is an IRAP AU, and irap_au_flag equal to 0 specifies that the AU is a GDR AU). When the irap_au_flag does not exist, its value is not inferred.
[0217] ii. Additionally, in one example, when vps_max_layers_minus1 is greater than 0, there is required to be exactly one AUD NAL unit in each IRAP or GDR AU.
[0218] 1. Alternatively, regardless of the value of vps_max_layers_minus1, there is required to be exactly one AUD NAL unit in each IRAP or GDR AU.
[0219] iii. In one example, a flag equal to 1 specifies that all strips in the AU have the same NAL unit type in the range from IDR_W_RADL to GDR_NUT (including the endpoints). Thus, if the NAL unit type is IDR_W_RADL or IDR_N_LP and the flag is equal to 1, the AU is a CVSS AU. Otherwise (the NAL unit type is CRA_NUT or GDR_NUT), when the variable NoOutputBeforeRecoveryFlag for each picture in the AU is equal to 1, the AU is a CVSS AU.
[0220] c. Additionally, in one example, it is required that each picture in the AU in CVS should have a nuh_layer_id equal to that of one of the pictures present in the first AU of CVS.
[0221] d. In one example, the indication is signaled in the NAL unit header.
[0222] i. In one example, bits in the NAL unit header (e.g., nuh_reserved_zero_bit) are used to specify whether the AU is complete.
[0223] e. In one example, the indication is signaled in a new NAL unit.
[0224] i. Additionally, in one example, when vps_max_layers_minus1 is greater than 0, it is required that there is one and only one new NAL unit present in each IRAP or GDR AU.
[0225] 1. Additionally, in one example, when present in the AU, the new NAL unit should be before any NAL unit other than the AUD NAL unit (when present) in decoding order in the AU.
[0226] 2. Alternatively, regardless of the value of vps_max_layers_minus1, it is required that there is one and only one new NAL unit present in each IRAP or GDR AU.
[0227] f. In one example, the indication is signaled in an SEI message.
[0228] i. Additionally, in one example, when vps_max_layers_minus1 is greater than 0, it is required that an SEI message is present in each IRAP or GDR AU.
[0229] 1. Additionally, in one example, when present in an AU, the SEINAL unit containing the SEI message shall be before any NAL unit other than the AUD NAL unit (when present) in decoding order within the AU.
[0230] 2. Alternatively, regardless of the value of vps_max_layers_minus1, it is required that an SEI message be present in each IRAP or GDR AU.
[0231] 2) Alternatively, to address the first issue, an indication of whether an AU starts a new CVS can be signaled.
[0232] a. In one example, this indication is signaled in the AUD NAL unit.
[0233] i. In one example, this indication is signaled by a flag (e.g., irap_or_gdr_au_flag) in the AUD NAL unit.
[0234] 1. Alternatively, additionally, when irap_or_gdr_au_flag is equal to 1, a flag (i.e., named irap_au_flag) can be signaled in the AUD to specify whether the AU is an IRAP AU or a GDR AU (irap_au_flag equal to 1 specifies the AU is an IRAP AU, and irap_au_flag equal to 0 specifies the AU is a GDR AU). When irap_au_flag is not present, its value is not inferred.
[0235] ii. Additionally, in one example, when vps_max_layers_minus1 is greater than 0, one and only one AUD NAL unit is required to be present in each CVSS AU.
[0236] 1. Alternatively, regardless of the value of vps_max_layers_minus1, one and only one AUD NAL unit is required to be present in each CVSS AU.
[0237] b. In one example, this indication is signaled in the NAL unit header.
[0238] i. In one example, a bit (e.g., nuh_reserved_zero_bit) in the NAL unit header is used to specify whether the AU is a CVSS AU.
[0239] c. In one example, this indication is signaled in a new NAL unit.
[0240] i. Additionally, in one example, when vps_max_layers_minus1 is greater than 0, it is required that there is one and only one new NAL unit in each CVSS AU.
[0241] 1. Additionally, in one example, when present in the AU, the new NAL unit shall be before any NAL unit other than the AUD NAL unit (when present) in decoding order within the AU.
[0242] 2. Alternatively, regardless of the value of vps_max_layers_minus1, it is required that there is one and only one new NAL unit in each CVSS AU.
[0243] ii. In one example, the presence of the new NAL unit in the AU specifies that the AU is a CVSS AU.
[0244] 1. Alternatively, a flag is included in the new NAL unit to specify whether the AU is a CVSS AU.
[0245] d. In one example, the indication is signaled in an SEI message.
[0246] i. Additionally, in one example, when vps_max_layers_minus1 is greater than 0, it is required that an SEI message is present in each CVSS AU.
[0247] 1. Additionally, in one example, when present in the AU, the SEI NAL unit containing the SEI message shall be before any NAL unit other than the AUD NAL unit (when present) in decoding order within the AU.
[0248] 2. Alternatively, regardless of the value of vps_max_layers_minus1, it is required that an SEI message is present in each CVSS AU.
[0249] 3) Alternatively, to solve the first problem, the CVSS AUs starting a new CVS are required to have pictures for each layer specified by the VPS. Note that the drawback of this method is that then the number of layers signaled in the VPS needs to be exact and thus, when one layer is removed from the bitstream, the VPS needs to be modified.
[0250] 4) Alternatively, to solve the first problem, when vps_max_layers_minus1 is greater than 0, the presence of an EOS NAL unit is specified for each picture in the last AU of each CVS, and optionally the presence of an EOB NAL unit is also specified in the last AU of each bitstream. Thus, each CVSS AU will be identified by being the first AU in the bitstream or by the presence of an EOS NAL unit in the previous AU.
[0251] 5) Alternatively, to solve the first problem, a variable determined by external means is specified to specify whether an AU is a CVSS AU.
[0252] Solution to the second problem
[0253] 6) To solve the second problem, each GDR AU is required to be complete (i.e., having pictures of each layer present in the CVS). This means that an AU consisting of GDR pictures, but if it is incomplete, then it is not a GDR AU, similar to an AU currently consisting of IRAP pictures, but if it is incomplete, then it is not an IRAP AU.
[0254] a. In one example, a GDR AU can be defined as an AU in which there is a PU of each layer in the CVS and the warp-decoded pictures in each present PU are GDR pictures.
[0255] 6. Embodiment
[0256]
[0257] 6.1 First Embodiment
[0258] This embodiment addresses items 1, 1.a, 1.a.i, 1.b, 1.b.i, 1.b.ii, 1.b.iii, 1.c, 6, and 6a.
[0259] 3 Definitions ...
[0261] Where In each present PU Is an AU of GDR pictures. ...
[0263] 7.3.2.9 AU Delimiter RBSP Syntax
[0264]
[0265] 7.4.3.9 AU Delimiter RBSP Semantics
[0266] The AU delimiter is used to indicate the start of an AU, and the slice types present in the coded pictures in the AU containing the AU delimiter NAL unit. There is no canonical decoding process associated with the AU delimiter.
[0267] ...
[0269] 7.4.2.4.2 Order of AUs and their association with CVSs
[0270] The bitstream consists of one or more CVSs.
[0271] A CVS consists of one or more AUs. The order of PUs and their association with AUs are described in Clause 7.4.2.4.3.
[0272] The first AU of a CVS is the CVSS AU, where each present PU is a CLVSS PU, which is either an IRAP PU with NoOutputBeforeRecoveryFlag equal to 1 or a GDR PU with NoOutputBeforeRecoveryFlag equal to 1.
[0273] Each CVSS AU shall have PUs for each layer present in the CVS,
[0274] 7.4.2.4.3 Order of PUs and their association with AUs
[0275] An AU consists of one or more PUs in increasing order of nuh_layer_id. The order NAL units and the decoded pictures and their association with PUs are described in Clause 7.4.2.4.4.
[0276] There can be at most one AUD NAL unit in an AU,
[0277] When the AUD NAL unit is present in the AU, it will be the first NAL unit of the AU and thus it is the first NAL unit of the first PU of the AU.
[0278] There can be at most one EOB NAL unit in an AU. When the EOB NAL unit is present in the AU, it will be the last NAL unit of the AU and thus it is the last NAL unit of the last PU of the AU. ...
[0280] Figure 1 is a block diagram showing an example video processing system 1000 in which various techniques disclosed herein may be implemented. Various implementations may include some or all components of system 1000. System 1000 may include an input 1002 for receiving video content. The video content may be received in a raw or uncompressed format, e.g., 8-bit or 10-bit multi-component pixel values, or may be received in a compressed or encoded format. Input 1002 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, Passive Optical Network (PON), etc., and wireless interfaces such as Wi-Fi or cellular interfaces.
[0281] System 1000 may include a codec component 1004, which may implement various codec or encoding methods described in this document. The codec component 1004 may reduce the average bit rate of the video from the input 1002 to the output of the codec 1004 to produce a codec representation of the video. Thus, codec techniques are sometimes referred to as video compression or video transcoding techniques. The output of the codec component 1004 may be stored or transmitted via a connected communication, as represented by component 1006. The stored or communicated bitstream (or codec) representation of the video received at input 1002 may be used by component 1008 to generate pixel values or a displayable video to be sent to the display interface 1010. The process of generating a user-visible video from the bitstream representation is sometimes referred to as video decompression. Additionally, although some video processing operations are referred to as "encoding" operations or tools, it should be understood that encoding tools or operations are used at the encoder, and the corresponding decoding tools or operations that reverse the encoding result will be performed by the decoder.
[0282] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB), High-Definition Multimedia Interface (HDMI), or DisplayPort, etc. Examples of storage interfaces include SATA (Serial Advanced Technology Attachment), PCI, IDE interfaces, etc. The techniques described in this document may be embodied in various electronic devices, such as mobile phones, laptop computers, smart phones, or other devices capable of performing digital data processing and / or video display.
[0283] Figure 2is a block diagram of a video processing apparatus 2000. The apparatus 2000 can be used to implement one or more methods described herein. The apparatus 2000 can be embodied in a smart phone, a tablet computer, a computer, an Internet of Things (IoT) receiver, etc. The apparatus 2000 can include one or more processors 2002, one or more memories 2004, and video processing hardware 2006. The processor(s) 2002 can be configured to implement one or more methods described in this document. The memory(ies) 2004 can be used to store data and code for implementing the methods and techniques described herein. The video processing hardware 2006 can be used to implement some of the techniques described in this document in hardware circuitry. In some embodiments, the hardware 2006 can be partially or fully in one or more processors 2002 (e.g., a graphics processor).
[0284] Figure 3 is a block diagram showing an example video codec system 100 that can utilize the techniques of the present disclosure.
[0285] As Figure 3 shown, the video codec system 100 can include a source device 110 and a destination device 120. The source device 110 generates encoded video data, which can be referred to as a video encoding device. The destination device 120 can decode the encoded video data generated by the source device 110, and the destination device 120 can be referred to as a video decoding device.
[0286] The source device 110 can include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.
[0287] The video source 112 can include sources such as a video capture device, an interface for receiving video data from a video content provider, and / or a computer graphics system for generating video data, or a combination of these sources. The video data can include one or more pictures. The video encoder 114 encodes the video data from the video source 112 to generate a bitstream. The bitstream can include a sequence of bits that form an encoded representation of the video data. The bitstream can include encoded pictures and associated data. The encoded pictures are encoded representations of pictures. The associated data can include a sequence parameter set, a picture parameter set, and other syntax structures. The I / O interface 116 can include a modulator / demodulator (modem) and / or a transmitter. The encoded video data can be sent directly to the destination device 120 via the I / O interface 116 over a network 130a. The encoded video data can also be stored on a storage medium / server 130b for access by the destination device 120.
[0288] The destination device 120 can include an I / O interface 126, a video decoder 124, and a display device 122.
[0289] The I / O interface 126 may include a receiver and / or a modem. The I / O interface 126 may obtain the encoded video data from the source device 110 or the storage medium / server 130b. The video decoder 124 may decode the encoded video data. The display device 122 may display the decoded video data to the user. The display device 122 may be integrated with the destination device 120 or may be external to the destination device 120, and the destination device 120 is configured to interface with an external display device.
[0290] The video encoder 114 and the video decoder 124 may operate according to video compression standards, such as the High Efficiency Video Coding (HEVC) standard, the Versatile Video Coding (VVC) standard, and other current and / or future standards.
[0291] Figure 4 is a block diagram showing an example of the video encoder 200, and the video encoder 200 may be Figure 3 the video encoder 114 in the system 100 shown.
[0292] The video encoder 200 may be configured to perform any or all of the techniques of the present disclosure. In Figure 4 an example, the video encoder 200 includes a plurality of functional components. The techniques described in the present disclosure may be shared among the various components of the video encoder 200. In some examples, a processor may be configured to perform any or all of the techniques described in the present disclosure.
[0293] The functional components of the video encoder 200 may include a splitting unit 201, a prediction unit 202, a residual generation unit 207, a transformation unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transformation unit 211, a reconstruction unit 212, a buffer 213, and an entropy coding / decoding unit 214. The prediction unit 202 may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra prediction unit 206.
[0294] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an Intra Block Copy (IBC) unit. The IBC unit may perform prediction in the IBC mode, in which at least one reference picture is the picture in which the current video block is located.
[0295] In addition, some components, such as the motion estimation unit 204 and the motion compensation unit 205, may be highly integrated, but are shown separately in Figure 4 an example for illustrative purposes.
[0296] The splitting unit 201 can split an image into one or more video blocks. The video encoder 200 and the video decoder 300 can support various video block sizes.
[0297] The mode selection unit 203 can select, for example based on an error result, one of the intra or inter coding modes, and provide the resulting intra or inter coded block to the residual generation unit 207 to generate residual block data, and to the reconstruction unit 212 to reconstruct the coded block for use as a reference picture. In some examples, the mode selection unit 203 can select a combined intra-inter prediction (CIIP) mode, where the prediction is based on an inter prediction signal and an intra prediction signal. The mode selection unit 203 can also select the resolution of the motion vector for a block in the case of inter prediction (e.g., sub-pixel or integer pixel accuracy).
[0298] To perform inter prediction on a current video block, the motion estimation unit 204 can generate motion information for the current video block by comparing one or more reference frames from the buffer 213 with the current video block. The motion compensation unit 205 can determine a predicted video block for the current video block based on the motion information and decoded samples of a picture from the buffer 213 other than the picture associated with the current video block.
[0299] The motion estimation unit 204 and the motion compensation unit 205 can perform different operations on the current video block, for example depending on whether the current video block is in an I slice, a P slice, or a B slice.
[0300] In some examples, the motion estimation unit 204 can perform uni-directional prediction on the current video block, and the motion estimation unit 204 can search the reference pictures in list 0 or list 1 to obtain a reference video block for the current video block. Then, the motion estimation unit 204 can generate a reference index indicating the reference picture in list 0 or list 1 that contains the reference video block, and a motion vector indicating the spatial displacement between the current video block and the reference video block. The motion estimation unit 204 can output the reference index, the prediction direction indicator, and the motion vector as the motion information for the current video block. The motion compensation unit 205 can generate a predicted video block for the current block based on the reference video block indicated by the motion information of the current video block.
[0301] In other examples, the motion estimation unit 204 may perform bidirectional prediction on a current video block. The motion estimation unit 204 may search for a reference video block for the current video block in the reference pictures in list 0, and may also search for another reference video block for the current video block in the reference pictures in list 1. Then, the motion estimation unit 204 may generate a reference index indicating the reference pictures in list 0 and list 1 that contain the reference video blocks, and a motion vector indicating the spatial displacement between the reference video block and the current video block. The motion estimation unit 204 may output the reference index and the motion vector of the current video block as the motion information of the current video block. The motion compensation unit 205 may generate a predicted video block of the current video block based on the reference video block indicated by the motion information of the current video block.
[0302] In some examples, the motion estimation unit 204 may output the entire set of motion information for the decoding process of the decoder.
[0303] In some examples, the motion estimation unit 204 may not output the entire set of motion information for the current video. Instead, the motion estimation unit 204 may signal the motion information of the current video block with reference to the motion information of another video block. For example, the motion estimation unit 204 may determine that the motion information of the current video block is sufficiently similar to the motion information of an adjacent video block.
[0304] In one example, the motion estimation unit 204 may indicate a value in a syntax structure associated with the current video block, and this value indicates to the video decoder 300 that the current video block has the same motion information as another video block.
[0305] In another example, the motion estimation unit 204 may identify another video block and a motion vector difference (MVD) in a syntax structure associated with the current video block. This motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 300 may use the motion vector of the indicated video block and this motion vector difference to determine the motion vector of the current video block.
[0306] As described above, the video encoder 200 may predictively signal the motion vector. Two examples of predictive signaling techniques that may be implemented by the video encoder 200 include advanced motion vector prediction (AMVP) and merge mode signaling.
[0307] The intra prediction unit 206 may perform intra prediction on a current video block. When the intra prediction unit 206 performs intra prediction on the current video block, the intra prediction unit 206 may generate prediction data for the current video block based on the decoded samples of other video blocks in the same picture. The prediction data for the current video block may include the predicted video block and various syntax elements.
[0308] The residual generation unit 207 may generate residual data for a current video block by subtracting a predicted video block of the current video block (e.g., indicated by a minus sign) from the current video block. The residual data for the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.
[0309] In other examples, such as in the skip mode, there may be no residual data for the current video block, and the residual generation unit 207 may not perform a subtraction operation.
[0310] The transform processing unit 208 may generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video block associated with the current video block.
[0311] After the transform processing unit 208 generates the transform coefficient video block associated with the current video block, the quantization unit 209 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0312] The inverse quantization unit 210 and the inverse transform unit 211 may apply inverse quantization and inverse transform to the transform coefficient video block, respectively, to reconstruct the residual video block from the transform coefficient video block. The reconstruction unit 212 may add the reconstructed residual video block to the corresponding samples of one or more predicted video blocks generated by the prediction unit 202 to generate a reconstructed video block associated with the current block for storage in the buffer 213.
[0313] After the reconstruction unit 212 reconstructs the video block, a loop filter operation may be performed to reduce video block artifacts in the video block.
[0314] The entropy coding unit 214 may receive data from other functional components of the video encoder 200. When the entropy coding unit 214 receives data, the entropy coding unit 214 may perform one or more entropy coding operations to generate entropy-coded data and output a bitstream including the entropy-coded data.
[0315] Figure 5 is a block diagram showing an example of a video decoder 300, and the video decoder 300 may be Figure 3 the video decoder 114 in the system 100 shown.
[0316] The video decoder 300 may be configured to perform any or all of the techniques of the present disclosure. In Figure 5 the example, the video decoder 300 includes a plurality of functional components. The techniques described in the present disclosure may be shared among the various components of the video decoder 300. In some examples, a processor may be configured to perform any or all of the techniques described in the present disclosure.
[0317] In Figure 5 example, video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, and a reconstruction unit 306 and a buffer 307. In some examples, video decoder 300 may perform a decoding pass that is substantially the reverse of the encoding pass ( Figure 4 ) described with respect to video encoder 200.
[0318] Entropy decoding unit 301 may retrieve the encoded bitstream. The encoded bitstream may include entropy-coded video data (e.g., coded blocks of video data). Entropy decoding unit 301 may decode the entropy-coded video data, and motion compensation unit 302 may determine motion information from the entropy-decoded video data, which includes motion vectors, motion vector precision, reference picture list indices, and other motion information. Motion compensation unit 302 may determine such information, for example, by performing AMVP and merge modes.
[0319] Motion compensation unit 302 may generate a motion-compensated block, possibly performing interpolation based on an interpolation filter. An identifier for the interpolation filter to be used with sub-pixel precision may be included in the syntax element.
[0320] Motion compensation unit 302 may use the interpolation filter used by video encoder 200 during the encoding of video blocks to compute the interpolation of sub-integer pixels of a reference block. Motion compensation unit 302 may determine the interpolation filter used by video encoder 200 according to the received syntax information, and generate a prediction block using the interpolation filter.
[0321] Motion compensation unit 302 may use some syntax information to determine the size of the blocks used to encode frames and / or slices of the encoded video sequence, the partitioning information that describes how each macroblock of a picture of the encoded video sequence is partitioned, the mode indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-coded block, and other information for decoding the encoded video sequence.
[0322] Intra prediction unit 303 may use, for example, the intra prediction mode received in the bitstream to form a prediction block from spatially adjacent blocks. Inverse quantization unit 304 inverse quantizes the quantized video block coefficients provided in the bitstream and decoded by entropy decoding unit 301, i.e., de-quantizes. Inverse transform unit 305 applies an inverse transform.
[0323] The reconstruction unit 306 can add the residual block to the corresponding prediction block generated by the motion compensation unit 302 or the intra prediction unit 303 to form a decoded block. If necessary, a deblocking filter can also be applied to filter the decoded block to remove block effect artifacts. Then the decoded video block is stored in the buffer 307, which provides reference blocks for subsequent motion compensation / intra prediction and also generates the decoded video for presentation on a display device.
[0324] Figures 6 - 7 An example method is shown that can implement the above technical solutions in an embodiment such as Figures 1 - 5 shown in.
[0325] Figure 6 A flowchart 600 of an example method of video processing is shown. Method 600 includes, at operation 610, performing a conversion between a video including one or more pictures in one or more video layers and a bitstream of the video, the bitstream including a coded video sequence (CVS), the CVS including one or more access units (AUs), and the bitstream further including a first syntax element indicating whether an AU includes pictures of each video layer constituting the CVS.
[0326] Figure 7 A flowchart 700 of an example method of video processing is shown. Method 700 includes, at operation 710, performing a conversion between a video including one or more pictures in one or more video layers and a bitstream of the video, the bitstream including a coded video sequence (CVS), the CVS including one or more access units (AUs), and the bitstream conforms to a format rule specifying a layer identifier carried by each picture in a given AU, the layer identifier being equal to the layer identifier of the first AU in the CVS including one or more video layers.
[0327] Figure 8 A flowchart 800 of an example method of video processing is shown. Method 800 includes, at operation 810, performing a conversion between a video including one or more pictures in one or more video layers and a bitstream of the video, the bitstream including a coded video sequence (CVS), the CVS including one or more access units (AUs), and the bitstream further including a first syntax element indicating an AU among one or more AUs that starts a new CVS.
[0328] Figure 9Flowchart 900 illustrates an example method of video processing. Method 900 includes, at operation 910, performing a conversion between a video including one or more pictures in one or more video layers and a bitstream of the video, the bitstream including a coded video sequence (CVS), the CVS including one or more access units (AUs), and the bitstream conforming to format rules that specify that a coded video sequence start (CVSS) AU that starts a new CVS includes pictures of each video layer specified in a video parameter set (VPS).
[0329] Figure 10 Flowchart 1000 illustrates an example method of video processing. Method 1000 includes, at operation 1010, performing a conversion between a video including one or more pictures in one or more video layers and a bitstream of the video, the bitstream including a coded video sequence (CVS), the CVS including one or more access units (AUs), and the bitstream conforming to format rules that specify that a given AU is identified as a coded video sequence start (CVSS) AU based on whether the given AU is the first AU in the bitstream or an AU preceding the given AU includes an end-of-sequence (EOS) network abstraction layer (NAL) unit.
[0330] Figure 11 Flowchart 1100 illustrates an example method of video processing. Method 1100 includes, at operation 1110, performing a conversion between a video including one or more pictures in one or more video layers and a bitstream of the video, the bitstream including a coded video sequence (CVS), the CVS including one or more access units (AUs), and the rule specifying that side information is used to indicate whether an AU among one or more AUs is a coded video sequence start (CVSS) AU.
[0331] Figure 12 Flowchart 1200 illustrates an example method of video processing. Method 1200 includes, at operation 1210, performing a conversion between a video including one or more pictures in one or more video layers and a bitstream of the video, the bitstream including a coded video sequence (CVS), the CVS including one or more access units (AUs), and the bitstream conforming to format rules that specify that each AU among one or more AUs that are stepwise decoded refresh AUs exactly includes one picture of each video layer present in the CVS.
[0332] Figure 13FIG. 1300 is a flow chart showing an example method of video processing. Method 1300 includes, at operation 1310, performing a conversion between a video including one or more pictures in one or more video layers and a bitstream of the video, the bitstream including a coded video sequence (CVS), the CVS including one or more access units (AUs), and the bitstream conforming to format rules that specify that each AU of one or more AUs that is an intra random access point AU exactly includes one picture of each video layer present in the CVS.
[0333] Next, a list of preferred solutions for some embodiments is provided.
[0334] P1. A video processing method includes performing a conversion between a video including one or more pictures in one or more video layers and a coded representation of an encoded version of the video; wherein the coded representation includes one or more access units (AUs); wherein the coded representation includes a syntax element if and only if the AU is of a type, and wherein the syntax element indicates whether the AU includes pictures of each video layer that makes up the coded video sequence.
[0335] P2. The method according to solution P1, wherein the type includes all AUs.
[0336] P3. The method according to solution P1, wherein the type includes AUs that start a new CVS.
[0337] P4. The method according to solution P1, wherein the syntax element is included in an access unit delimiter network abstraction layer unit.
[0338] P5. The method according to any one of solutions P1 to P4, wherein the syntax element is included in a network abstraction layer unit header.
[0339] P6. The method according to any one of solutions P1 to P4, wherein the syntax element is included in a new network abstraction layer unit.
[0340] P7. A video processing method includes performing a conversion between a video including one or more pictures in one or more video layers and a coded representation of an encoded version of the video; wherein the coded representation includes one or more access units (AUs); wherein the coded representation conforms to format rules that specify that each picture in a given AU carries a layer identifier equal to the layer identifier of the first AU of the coded video sequence.
[0341] P8. A video processing method includes performing a conversion between a video having one or more pictures in one or more video layers and a transcoded representation representing an encoded version of the video; wherein the transcoded representation includes one or more access units (AUs); and wherein the transcoded representation includes syntax elements if and only if an AU starts a new transcoded video sequence.
[0342] P9. The method according to solution P8, wherein the syntax elements are included in an access unit delimiter network abstraction layer unit.
[0343] P10. The method according to any one of solutions P8 to P9, wherein the syntax elements are included in a network abstraction layer unit header.
[0344] P11. The method according to any one of solutions P8 to P9, wherein the syntax elements are included as supplementary enhancement information.
[0345] P12. A video processing method includes performing a conversion between a video having one or more pictures in one or more video layers and a transcoded representation of the video; wherein the transcoded representation includes one or more access units (AUs); and wherein the transcoded representation conforming to the start AU specifying a transcoded video sequence includes video pictures of each video layer specified by a video parameter set, or the transcoded representation implicitly signals the format rule of the start AU based on a rule.
[0346] P13. A video processing method includes performing a conversion between a video having one or more pictures in one or more video layers and a transcoded representation of the video; wherein the transcoded representation includes one or more access units (AUs); and wherein the transcoded representation conforming to each AU specifying a gradual decoding refresh (GDR) type includes a format rule of at least one video picture of each video layer of a transcoded video sequence.
[0347] P14. The method according to solution P13, wherein each AU of the GDR type further includes prediction units for each layer in a CVS, and a PU includes a GDR picture.
[0348] P15. The method according to any one of solutions 1 to 14, wherein the conversion includes encoding the video into the transcoded representation.
[0349] P16. The method according to any one of solutions 1 to 14, wherein the conversion includes decoding the transcoded representation to generate pixel values of the video.
[0350] P17. A video decoding apparatus includes a processor configured to implement one or more of the methods described in solutions P1 to P16.
[0351] P18. A video encoding device, comprising a processor configured to implement one or more of the methods described in Solutions P1 to P16.
[0352] P19. A computer program product having computer code stored thereon, which when executed by a processor causes the processor to implement the method described in any one of Solutions P1 to P16.
[0353] Next, another list of preferred solutions of some embodiments is provided.
[0354] A1. A video processing method, comprising performing a conversion between a video including one or more pictures in one or more video layers and a bitstream of the video, wherein the bitstream includes a decoded video sequence, the decoded video sequence includes one or more access units, and wherein the bitstream further includes a first syntax element indicating whether the access unit includes pictures of each video layer constituting the decoded video sequence.
[0355] A2. The method according to Solution A1, wherein the access unit is configured to start a new decoded video sequence.
[0356] A3. The method according to Solution A2, wherein the pictures of each video layer are intra random access point pictures or progressive decoded refresh pictures.
[0357] A4. The method according to Solution A1, wherein the first syntax element is included in an access unit delimiter network abstraction layer unit.
[0358] A5. The method according to Solution A4, wherein the first syntax element is a flag indicating whether the access unit including the access unit delimiter is an intra random access point access unit or a progressive decoded refresh access unit.
[0359] A6. The method according to Solution A5, wherein the first syntax element is irap_or_gdr_au_flag.
[0360] A7. The method according to Solution A4, wherein when a second syntax element indicating the number of video layers specified by a video parameter set is greater than 1, the access unit delimiter network abstraction layer unit is the only access unit delimiter network abstraction layer unit present in each intra random access point access unit or progressive decoded refresh access unit.
[0361] A8. The method according to Solution A7, wherein the second syntax element indicates the maximum allowed number of layers of a reference video parameter set in each decoded video sequence.
[0362] A9. A method according to Solution A1, wherein a first syntax element being equal to 1 indicates all slices in access units of the same network abstraction layer unit type within the range from IDR_W_RADL to GDR_NUT (including endpoints).
[0363] A10. A method according to Solution A9, wherein the first syntax element is equal to 1 and the network abstraction layer unit type is IDR_W_RADL or IDR_N_LP indicates that the access unit is a warp decoded video sequence start access unit.
[0364] A11. A method according to Solution A9, wherein a variable for each picture in the access unit being equal to 1 and the network abstraction layer unit type being CRA_NUT or GDR_NUT indicates that the access unit is a warp decoded video sequence start access unit.
[0365] A12. A method according to Solution A11, wherein the variable indicates whether to output a picture in the decoded picture buffer that is before the current picture in decoded order before the picture is recovered.
[0366] A13. A method according to Solution A11, wherein the variable is the NoOutputBeforeRecoveryFlag.
[0367] A14. A method according to Solution A1, wherein the first syntax element is included in the network abstraction layer unit header.
[0368] A15. A method according to Solution A1, wherein the first syntax element is included in a new network abstraction layer unit.
[0369] A16. A method according to Solution A15, wherein when a second syntax element indicating the number of video layers specified by the video parameter set is greater than 1, the new network abstraction layer unit is the only network abstraction layer unit present in each intra random access point access unit or progressive decoded refresh access unit.
[0370] A17. A method according to Solution A1, wherein the first syntax element is included in the supplementary enhancement information message.
[0371] A18. A method according to Solution A17, wherein when the second syntax element indicating the number of video layers specified by the video parameter set is greater than 0, the supplementary enhancement information message is the only supplementary enhancement information message present in each intra random access point access unit or progressive decoded refresh access unit.
[0372] A19. A video processing method includes performing a conversion between a video including one or more pictures in one or more video layers and a bitstream of the video, where the bitstream includes a decoded video sequence, the decoded video sequence includes one or more access units, and where the bitstream conforms to format rules specifying a layer identifier carried by each picture in a given access unit, the layer identifier being equal to the layer identifier of a first access unit in the decoded video sequence including one or more video layers.
[0373] A20. A video processing method includes performing a conversion between a video including one or more pictures in one or more video layers and a bitstream of the video, where the bitstream includes a decoded video sequence, the decoded video sequence includes one or more access units, and where the bitstream further includes a first syntax element indicating an access unit in one or more access units that starts a new decoded video sequence.
[0374] A21. The method according to solution A20, where the first syntax element is included in an access unit delimiter network abstraction layer unit.
[0375] A22. The method according to solution A20, where the first syntax element is included in a network abstraction layer unit header.
[0376] A23. The method according to solution A20, where the first syntax element is included in a new network abstraction layer unit.
[0377] A24. The method according to solution A20, where the first syntax element is included in a supplementary enhancement information message.
[0378] A25. A video processing method includes performing a conversion between a video including one or more pictures in one or more video layers and a bitstream of the video, where the bitstream includes a decoded video sequence, the decoded video sequence includes one or more access units, and where the bitstream conforms to format rules specifying that a decoded video sequence start access unit starting a new decoded video sequence includes pictures of each video layer specified in a video parameter set.
[0379] A26. A video processing method includes performing a conversion between a video including one or more pictures in one or more video layers and a bitstream of the video, the bitstream includes a decoded video sequence, the decoded video sequence includes one or more access units, and where the bitstream conforms to format rules specifying that a given access unit is identified as a decoded video sequence start access unit based on whether the given access unit is the first access unit in the bitstream or an access unit before the given access unit includes a sequence end network abstraction layer unit.
[0380] A27. A video processing method, comprising performing a conversion between a video and a bitstream of the video including one or more pictures in one or more video layers based on a rule, wherein the bitstream includes a decoded / encoded video sequence, the decoded / encoded video sequence includes one or more access units, and wherein the rule specifies that side information is used to indicate whether an access unit in the one or more access units is a start access unit of the decoded / encoded video sequence.
[0381] A28. The method according to any one of Solutions A1 to A27, wherein the conversion includes decoding the video from the bitstream.
[0382] A29. The method according to any one of Solutions A1 to A27, wherein the conversion includes encoding the video into the bitstream.
[0383] A30. A method of storing a bitstream representing a video in a computer-readable recording medium, comprising generating a bitstream from the video according to the method described in any one or more of Solutions A1 to A27; and storing the bitstream in the computer-readable recording medium.
[0384] A31. A video processing apparatus, comprising a processor configured to implement the method described in any one or more of Solutions A1 to A30.
[0385] A32. A computer-readable medium having instructions stored thereon, which when executed cause the processor to implement the method described in one or more of Solutions A1 to A30.
[0386] A33. A computer-readable medium storing a bitstream generated according to any one or more of Solutions A1 to A30.
[0387] A34. A video processing apparatus for storing a bitstream, wherein the video processing apparatus is configured to implement the method described in any one or more of Solutions A1 to A30.
[0388] Next, yet another list of preferred solutions of some embodiments is provided.
[0389] B1. A video processing method includes performing a conversion between a video including one or more pictures in one or more video layers and a bitstream of the video, the bitstream including a decoded video sequence, the decoded video sequence including one or more access units, and wherein the bitstream conforms to format rules that specify that each access unit among one or more access units that are progressive decoding refresh access units exactly includes one picture of each video layer present in the decoded video sequence.
[0390] B2. The method according to solution B1, wherein each access unit that is a progressive decoding refresh access unit includes prediction units for each layer present in the decoded video sequence, and wherein the prediction units for each layer include decoded pictures that are progressive decoding refresh pictures.
[0391] B3. A video processing method includes performing a conversion between a video including one or more pictures in one or more video layers and a bitstream of the video, the bitstream including a decoded video sequence, the decoded video sequence including one or more access units, and wherein the bitstream conforms to format rules that specify that each access unit among one or more access units that are intra random access point access units exactly includes one picture of each video layer present in the decoded video sequence.
[0392] B4. The method according to any one of solutions B1 to B3, wherein the conversion includes decoding the video from the bitstream.
[0393] B5. The method according to any one of solutions B1 to B3, wherein the conversion includes encoding the video into the bitstream.
[0394] B6. A method of storing a bitstream representing a video in a computer-readable recording medium includes generating a bitstream from the video according to the method described in any one or more of solutions B1 to B3; and storing the bitstream in the computer-readable recording medium.
[0395] B7. A video processing apparatus includes a processor configured to implement the method described in any one or more of solutions B1 to B6.
[0396] B8. A computer-readable medium having instructions stored thereon that, when executed, cause a processor to implement the method described in one or more of solutions B1 to B6.
[0397] B9. A computer-readable medium stores a bitstream generated according to any one or more of solutions B1 to B6.
[0398] B10. A video processing apparatus for storing a bitstream, wherein the video processing apparatus is configured to implement the method according to any one or more of Solutions B1 to B6.
[0399] In this document, the term "video processing" may refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm may be applied during the conversion from the pixel representation of a video to the corresponding bitstream representation, or vice versa. For example, the bitstream representation of a current video block may correspond to bits juxtaposed or extended at different positions within the bitstream as defined by the syntax. For example, a macroblock may be encoded based on transformed and decoded error residual values, and also using bits in the header and other fields in the bitstream. Additionally, during the conversion, the decoder may parse the bitstream based on a determination, as described in the above solutions, knowing that some fields may or may not be present. Similarly, the encoder may determine that certain syntax fields are included or not included and accordingly generate an encoded representation by including or excluding the syntax fields from the encoded representation.
[0400] The disclosed and other solutions, examples, embodiments, modules, and functional operations described in this document may be implemented in digital electronic circuits, or in computer software, firmware, or hardware, including the structures disclosed in this document and their structural equivalents, or in a combination of one or more of them. The disclosed and other embodiments may be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer-readable medium for execution or control of the operation by a data processing apparatus. The computer-readable medium may be a machine-readable storage device, a machine-readable storage motherboard, a storage device, a substance composition affecting a machine-readable propagated signal, or a combination of one or more of them. The term "data processing apparatus" includes all apparatuses, devices, and machines for processing data, including, for example, programmable processors, computers, or multiple processors or computers. In addition to the hardware, the apparatus may include code for creating an execution environment for the said computer program, e.g., code constituting processor firmware, protocol stacks, database management systems, operating systems, or a combination of one or more of them. A propagated signal is an artificially generated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal generated to encode information for transmission to a suitable receiver device.
[0401] A computer program (also known as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, and can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. The program can be stored in a part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the said program, or in multiple coordinated files (e.g., files that store one or more modules, subroutines, or portions of code). A computer program can be deployed to execute on one computer or on multiple computers located at one site or distributed across multiple sites and interconnected by a communication network.
[0402] The processes and logical flows described in this document can be executed by one or more programmable processors that execute one or more computer programs to perform functions by operating on input data and generating output. The processes and logical flows can also be executed by dedicated logic circuitry, and the apparatus can also be implemented as dedicated logic circuitry, such as an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit).
[0403] Processors suitable for executing a computer program include, by way of example, both general and special purpose microprocessors, and any one or more processors of any type of digital computer. In general, a processor will receive instructions and data from a read only memory or a random access memory or both. The essential elements of a computer are a processor for executing the instructions and one or more storage devices for storing the instructions and data. In general, a computer will also include or be operatively coupled to receive data from or transfer data to one or more mass storage devices for storing data (e.g., magnetic disks, magneto-optical disks, or optical disks). However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including by way of example semiconductor memory devices (e.g., EPROM, EEPROM, and flash memory devices); magnetic disks (e.g., internal hard disks or removable disks); magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, dedicated logic circuitry.
[0404] Although this patent document contains many details, these should not be construed as limiting the scope of any subject matter or what can be claimed, but rather as descriptions of features specific to particular embodiments of a particular technology. Certain features described in the context of separate embodiments in this patent document may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments. Moreover, although the above features may be described as acting in certain combinations and even initially claimed as such, in some cases, one or more features from a claimed combination may be deleted from that combination, and the claimed combination may be directed to a sub-combination or a variation of a sub-combination.
[0405] Similarly, although operations are depicted in the drawings in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in a sequential order, or that all of the operations shown be performed to obtain a desired result. Additionally, the separation of various system components in the embodiments described in this patent document should not be construed as requiring such separation in all embodiments.
[0406] Only some implementations and examples have been described, and other implementations, enhancements, and variations can be made based on what is described and illustrated in this patent document.
Claims
1. A video processing method, comprising: Perform a conversion between a video including one or more pictures in one or more video layers and a bitstream of the video, wherein the bitstream includes a coded video sequence, the coded video sequence includes access units, and wherein the bitstream further includes a first syntax element indicating whether the access unit includes a first type of picture that constitutes each video layer of the coded video sequence; wherein the first syntax element is included in an access unit delimiter network abstraction layer unit; When a second syntax element indicating the maximum allowed number of layers in each coded video sequence referring to a video parameter set is greater than 1, the access unit delimiter network abstraction layer unit is the only access unit delimiter network abstraction layer unit present in each intra-random access point access unit or progressive decoding refresh access unit.
2. The method according to claim 1, wherein, The access unit includes the first type of picture of each video layer; and the access unit is configured to start a new coded video sequence, or each picture of the access unit is the first type of picture.
3. The method according to claim 1, wherein, The first type of picture of each video layer is an intra-random access point picture or a progressive decoding refresh picture.
4. The method according to claim 1, wherein, The first syntax element is a flag indicating whether the access unit including the access unit delimiter is an intra-random access point access unit or a progressive decoding refresh access unit.
5. The method according to claim 4, wherein, The first syntax element is irap_or_gdr_au_flag.
6. The method according to claim 1, wherein, The second syntax element indicates the maximum allowed number of layers by indicating the maximum allowed number minus 1.
7. The method according to claim 1, wherein the bitstream further conforms to a format rule that specifies that each picture in an access unit carries a layer identifier, and the layer identifier is equal to the layer identifier of one of the pictures in the first access unit of the coded video sequence.
8. The method according to claim 7, wherein the layer identifier is nuh_layer_id.
9. The method according to claim 1, wherein, The first syntax element being equal to 1 indicates that all slices in the access unit include the same network abstraction layer unit type within the range from IDR_W_RADL to GDR_NUT, and the range includes the endpoints IDR_W_RADL and GDR_NUT.
10. The method according to claim 9, wherein, The first syntax element being equal to 1 and the network abstraction layer unit type being IDR_W_RADL or IDR_N_LP indicates that the access unit is a coded video sequence start access unit.
11. The method according to claim 9, wherein, Each picture in the access unit having a variable equal to 1 and the network abstraction layer unit type being CRA_NUT or GDR_NUT indicates that the access unit is a coded video sequence start access unit.
12. The method according to claim 11, wherein, The variable indicates whether to output pictures in the decoded picture buffer that are in decoded order before the current picture before the current picture is recovered.
13. The method according to claim 12, wherein, The variable is NoOutputBeforeRecoveryFlag.
14. The method according to any one of claims 1 to 13, wherein, The conversion includes decoding the video from the bitstream.
15. The method according to any one of claims 1 to 13, wherein, The conversion includes encoding the video into the bitstream.
16. An apparatus for processing video data, comprising a processor and a non-transitory memory storing instructions thereon, wherein the instructions, when executed by the processor, cause the processor to: perform a conversion between a video including one or more pictures in one or more video layers and a bitstream of the video, Among them, The bitstream includes a coded video sequence, the coded video sequence includes access units, and wherein the bitstream further includes a first syntax element indicating whether the access unit includes a first type of picture that constitutes each video layer of the coded video sequence; wherein the first syntax element is included in an access unit delimiter network abstraction layer unit; When the second syntax element indicating the maximum allowed number of layers in each coded video sequence of the reference video parameter set is greater than 1, the access unit delimiter network abstraction layer unit is the only access unit delimiter network abstraction layer unit present in each intra-random access point access unit or progressive decoding refresh access unit.
17. A non - transitory computer - readable storage medium storing instructions, the instructions causing a processor to: Perform a conversion between a video including one or more pictures in one or more video layers and a bitstream of the video, Among them, The bitstream includes a coded video sequence, the coded video sequence includes access units, and wherein, the bitstream further includes a first syntax element indicating whether the access unit includes a first type of picture constituting each video layer of the coded video sequence; wherein, the first syntax element is included in the access unit delimiter network abstraction layer unit; When the second syntax element indicating the maximum allowed number of layers in each coded video sequence of the reference video parameter set is greater than 1, the access unit delimiter network abstraction layer unit is the only access unit delimiter network abstraction layer unit present in each intra-random access point access unit or progressive decoding refresh access unit.
18. A non - transitory computer - readable recording medium storing a bitstream of a video generated by a method executed by a video processing device, where the method includes: Generate a bitstream of a video including one or more pictures in one or more video layers, wherein, the bitstream includes a coded video sequence, the coded video sequence includes access units, and wherein, the bitstream further includes a first syntax element indicating whether the access unit includes a first type of picture constituting each video layer of the coded video sequence; wherein, the first syntax element is included in the access unit delimiter network abstraction layer unit; When the second syntax element indicating the maximum allowed number of layers in each coded video sequence of the reference video parameter set is greater than 1, the access unit delimiter network abstraction layer unit is the only access unit delimiter network abstraction layer unit present in each intra-random access point access unit or progressive decoding refresh access unit.
19. A method of storing a bitstream of a video, including: Generate a bitstream of a video including one or more pictures in one or more video layers; and Store the bitstream in a non-transitory computer-readable recording medium, wherein, the bitstream includes a coded video sequence, the coded video sequence includes access units, and wherein, the bitstream further includes a first syntax element indicating whether the access unit includes a first type of picture constituting each video layer of the coded video sequence; wherein, the first syntax element is included in the access unit delimiter network abstraction layer unit; When the second syntax element indicating the maximum allowed number of layers in each coded video sequence of the reference video parameter set is greater than 1, the access unit delimiter network abstraction layer unit is the only access unit delimiter network abstraction layer unit present in each intra-random access point access unit or progressive decoding refresh access unit.
Citation Information
Patent Citations
Parameter set updates in video coding
CN104380747A
An apparatus, a method and a computer program for omnidirectional video
WO2019038473A1