HRD parameters for layer-based conformance tests
By encoding lower layer HRD parameters to match the top layer and optimizing compliance testing, the solution addresses inefficiencies in multi-layer bitstream compression, achieving reduced bitstream size and resource utilization.
Patent Information
- Application Number
- JP2025076670
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-09-24
- Filing Date
- 2025-05-02
- Publication Date
- 2025-07-30
- Estimated Expiration
- 2040-09-08
AI Technical Summary
Existing video coding systems face challenges in efficiently compressing multi-layer bitstreams due to redundant HRD parameters, complex compliance testing, and increased resource utilization, leading to larger bitstream sizes and inefficient use of processor, memory, and network resources.
The proposed solution involves encoding HRD parameters of lower layers to match those of the top layer by setting the sublayer_cpb_params_present_flag to 0, reducing redundant parameter inclusion, and optimizing compliance testing mechanisms to minimize bitstream size and resource usage.
This approach reduces the bitstream size and resource utilization by eliminating redundant HRD parameters and simplifying compliance testing, thereby enhancing coding efficiency and reducing processor, memory, and network resource consumption.
Smart Images

Figure 2025111786000001_ABST
Abstract
Description
Technical Field
[0001] Cross - reference to Related Applications This patent application claims the benefit of U.S. Provisional Patent Application No. 62 / 905,244, filed on September 24, 2019, entitled "Hypothetical Reference Decoder (HRD) for Multi - Layer Video Bitstreams" by Ye - Kui Wang, which is hereby incorporated by reference in its entirety.
[0002] This disclosure generally relates to video coding, and more particularly to virtual reference decoder (HRD) parameter modification to support efficient encoding and / or compliance testing of multi - layer bitstreams.
Background Art
[0003] Even for relatively short videos, the amount of video data required for their depiction can be enormous. Therefore, when streaming or otherwise communicating data over a communication network with limited bandwidth capacity, difficulties may arise. Thus, video data is generally compressed before being communicated over modern telecommunications networks. Also, when storing a video on a storage device, the size of the video can be a problem because memory resources may be limited. Video compression devices often use software and / or hardware at the source to encode the video data before transmission or storage, thereby reducing the amount of data required to represent the digital video image. Subsequently, the compressed data is received at the destination by a video decompression device that decodes the video data. In the context of limited network resources and an increasing demand for high - quality video, improved compression and decompression techniques that improve the compression ratio without sacrificing much or any of the image quality are desired.
Summary of the Invention
[0004] In one embodiment, the present disclosure includes a method implemented by a decoder, the method comprising receiving, by a receiver of the decoder, a bitstream including a plurality of sublayer representations, virtual reference decoder (HRD) parameters, and a sublayer coding picture buffer (CPB) parameter presence flag (sublayer_cpb_params_present_flag); inferring, by a processor of the decoder, that HRD parameters of all sublayer representations whose temporal identifier (TemporalId) is less than the maximum TemporalId are equal to the HRD parameters of the maximum sublayer representation having the maximum TemporalId when the sublayer_cpb_params_present_flag is set to 0; and decoding, by the processor, an image from the sublayer representation.
[0005] Video coding systems employ various compliance tests to ensure that the bitstream is decodable by the decoder. For example, the compliance check can include testing the entire bitstream for compliance, then testing each layer of the bitstream for compliance, and finally checking the potential decodable output for compliance. To perform the compliance check, corresponding parameters are included in the bitstream. The HRD can read the parameters and execute the tests. A video can include multiple layers and multiple different output layer sets (OLSs). In response to a request, the encoder transmits one or more layers of a selected OLS. For example, the encoder can transmit the best layer from the OLSs that can be supported by the current network bandwidth. Problems can occur when the video is separated into multiple layers and / or sublayers. The encoder can encode these layers into the bitstream. Additionally, the encoder can use the HRD that executes a compliance test to check the bitstream for compliance with the standard. The encoder can be configured to include layer-specific HRD parameters in the bitstream to support such compliance tests. The layer-specific HRD parameters may be encoded layer by layer in some video coding systems. In some cases, the layer-specific HRD parameters are the same for each layer, resulting in redundant information that unnecessarily increases the size of the video coding. This embodiment includes a mechanism for reducing the HRD parameter redundancy of a video that employs multiple layers. The encoder can encode the HRD parameters of the top layer. The encoder can also encode the sublayer_cpb_params_present_flag. Setting the sublayer_cpb_params_present_flag to 0 can indicate that all lower layers should use the same HRD parameters as the top layer.In this context, the top layer has the largest layer identifier (ID), and the lower layers are any layers having a layer ID smaller than the layer ID of the top layer. In this way, the HRD parameters of the lower layers can be omitted from the bitstream. This reduces the bitstream size and thus reduces the utilization of processor, memory, and / or network resources in both the encoder and the decoder.
[0006] Optionally, in any of the foregoing aspects, another implementation of the aspect is provided where the sublayer_cpb_params_present_flag is included in the video parameter set (VPS) in the bitstream.
[0007] Optionally, in any of the foregoing aspects, another implementation of the aspect is provided where the maximum TemporalId of the maximum sublayer representation is represented as the HRD maximum TemporalId (hrd_max_tid[i]), and i indicates the i-th HRD parameter syntax structure.
[0008] Optionally, in any of the foregoing aspects, another implementation of the aspect is provided where the TemporalId smaller than the maximum TemporalId is in the range from 0 to hrd_max_tid[i] - 1.
[0009] Optionally, in any of the foregoing aspects, another implementation of the aspect is provided where the HRD parameters include a fixed picture rate general flag (fixed_pic_rate_general_flag[i]) indicating whether the temporal distance between the HRD output times of consecutive pictures in the output order is constrained.
[0010] Optionally, in any of the foregoing aspects, another implementation of the aspect is provided where the HRD parameters include a sublayer HRD parameters (sublayer_hrd_parameters(i)) syntax structure that includes the HRD parameters for one or more sublayers.
[0011] Optionally, in any of the foregoing aspects, another embodiment of the aspect is provided, wherein the HRD parameter includes a general video coding layer (VCL) HRD parameter presence flag (general_vcl_hrd_params_present_flag) indicating whether the VCL HRD parameter related to the adaptation point is present in the general HRD parameter syntax structure.
[0012] In one embodiment, the present disclosure includes a method implemented by an encoder, the method comprising: encoding, by a processor of the encoder, a plurality of sublayer representations into a bitstream; encoding, by the processor, HRD parameters and a sublayer_cpb_params_present_flag into the bitstream; inferring, by the processor, that the HRD parameters of all sublayer representations whose TemporalId is less than the maximum TemporalId are equal to the HRD parameters of the maximum sublayer representation having the maximum TemporalId when the sublayer_cpb_params_present_flag is set to 0; and executing, by the processor, a set of bitstream compliance tests on the bitstream based on the HRD parameters.
[0013] Video coding systems employ various compliance tests to ensure that the bitstream is decodable by the decoder. For example, the compliance check can include testing the entire bitstream for compliance, then testing each layer of the bitstream for compliance, and finally checking the potential decodable output for compliance. To perform the compliance check, corresponding parameters are included in the bitstream. The HRD can read the parameters and execute the tests. The video can include multiple layers and multiple OLSs. In response to a request, the encoder transmits one or more layers of the selected OLS. For example, the encoder can transmit the best layer from the OLSs supported by the current network bandwidth. The problem can occur when the video is separated into multiple layers and / or sublayers. The encoder can encode these layers into the bitstream. Further, the encoder can use the HRD that performs a compliance test to check the bitstream for compliance with the standard. The encoder may be configured to include layer-specific HRD parameters in the bitstream to support such compliance tests. The layer-specific HRD parameters may be encoded layer by layer in some video coding systems. In some cases, the layer-specific HRD parameters are the same for each layer, resulting in redundant information that unnecessarily increases the size of the video coding. This embodiment includes a mechanism for reducing the HRD parameter redundancy of a video that employs multiple layers. The encoder can encode the HRD parameters of the top layer. The encoder can also encode the sublayer_cpb_params_present_flag. Setting the sublayer_cpb_params_present_flag to 0 can indicate that all lower layers should use the same HRD parameters as the top layer.In this context, the top layer has the largest layer ID, and a lower layer is any layer having a layer ID smaller than the layer ID of the top layer. In this way, the HRD parameters of the lower layer can be omitted from the bitstream. Thereby, the bitstream size is reduced, and thus the utilization of processor, memory, and / or network resources in both the encoder and the decoder is reduced.
[0014] Optionally, in any of the foregoing aspects, another embodiment of the aspect is provided in which the sublayer_cpb_params_present_flag is encoded in the VPS in the bitstream.
[0015] Optionally, in any of the foregoing aspects, another embodiment of the aspect is provided in which the maximum TemporalId of the maximum sublayer representation is represented as hrd_max_tid[i], where i indicates the i-th HRD parameter syntax structure.
[0016] Optionally, in any of the foregoing aspects, another embodiment of the aspect is provided in which the TemporalId smaller than the maximum TemporalId is in the range from 0 to hrd_max_tid[i] - 1.
[0017] Optionally, in any of the foregoing aspects, another embodiment of the aspect is provided in which the HRD parameter includes a fixed_pic_rate_general_flag[i] indicating whether the temporal distance between the HRD output times of consecutive pictures in the output order is restricted.
[0018] Optionally, in any of the foregoing aspects, another embodiment of the aspect is provided in which the HRD parameter includes a sublayer_hrd_parameters(i) syntax structure that includes the HRD parameters for one or more sublayers.
[0019] Optionally, in any of the foregoing aspects, another embodiment of the aspect is provided, where the HRD parameter includes a general_vcl_hrd_params_present_flag indicating whether the VCL HRD parameter related to the adaptation point exists in the general HRD parameter syntax structure.
[0020] In one embodiment, the present disclosure includes a video coding device including a processor, a receiver coupled to the processor, a memory coupled to the processor, and a transmitter coupled to the processor, wherein the processor, the receiver, the memory, and the transmitter are configured to execute any of the methods of the foregoing aspects.
[0021] In one embodiment, the present disclosure includes a non-transitory computer-readable medium including a computer program product for use by a video coding device, the computer program product including computer-executable instructions stored on the non-transitory computer-readable medium that, when executed by a processor, cause the video coding device to execute any of the methods of the foregoing aspects.
[0022] In one embodiment, the present disclosure includes a decoder including receiving means for receiving a bitstream including a plurality of sublayer representations, HRD parameters, and a sublayer_cpb_params_present_flag, inferring means for inferring that the HRD parameters of all sublayer representations whose TemporalId is less than the maximum TemporalId are equal to the HRD parameters of the maximum sublayer representation having the maximum TemporalId when the sublayer_cpb_params_present_flag is set to 0, decoding means for decoding an image from the sublayer representation, and transferring means for transferring the image for display as part of the decoded video sequence.
[0023] Optionally, in any of the foregoing aspects, another embodiment of the aspect is provided, where the decoder is further configured to execute any of the methods of the foregoing aspects.
[0024] In one embodiment, the present disclosure includes an encoder comprising: encoding means for encoding a plurality of sublayer representations into a bitstream; encoding means for encoding HRD parameters and a sublayer_cpb_params_present_flag into the bitstream; inferring means for inferring that the HRD parameters of all sublayer representations whose TemporalId is less than the maximum TemporalId are equal to the HRD parameters of the maximum sublayer representation having the maximum TemporalId when the sublayer_cpb_params_present_flag is set to 0; HRD means for performing a set of bitstream compliance tests on the bitstream based on the HRD parameters; and storage means for storing the bitstream for communication to a decoder.
[0025] Optionally, in any of the foregoing aspects, another embodiment of the aspect is provided, wherein the encoder is further configured to execute the method of any of the foregoing aspects.
[0026] For clarity, any one of the foregoing embodiments may be combined with any one or more of the other foregoing embodiments to create a new embodiment within the scope of the present disclosure.
[0027] These and other features will be more clearly understood from the following detailed description in conjunction with the accompanying drawings and the claims.
Brief Description of the Drawings
[0028] To understand the present disclosure more fully, reference is now made to the following brief description, taken in conjunction with the accompanying drawings and detailed description, in which like reference numerals represent like parts.
[0029]
Figure 1
[0030]
Figure 2
[0031]
Figure 3
[0032]
Figure 4
[0033]
Figure 5
[0034]
Figure 6
[0035]
Figure 7
[0036]
Figure 8
[0037]
Figure 9
[0038]
Figure 10
[0039]
Figure 11
[0040]
Figure 12
[0041] Exemplary embodiments of one or more embodiments are provided below, but it should first be understood that the disclosed system and / or method may be implemented using any number of techniques, whether currently known or existing. The present disclosure should in no way be limited to the exemplary embodiments, drawings, and techniques shown below, which include the exemplary designs and embodiments illustrated and described herein, but may be modified within the scope of the appended claims, together with the full scope of their equivalents.
[0042] The following terms are defined as follows, unless used in a contrary context herein. Specifically, the following definitions are intended to provide further clarity to the present disclosure. However, terms may be described differently in different contexts. Accordingly, the following definitions should be considered supplementary and should not be considered to limit any other definitions provided herein for such terms.
[0043] A bitstream is a sequence of bits that contains video data compressed for transmission between an encoder and a decoder. An encoder is a device configured to employ an encoding process to compress video data into a bitstream. A decoder is a device configured to employ a decoding process to reconstruct video data from the bitstream for display. An image is an array of luma samples and / or an array of chroma samples that constructs a frame or a field thereof. An encoded or decoded image may be referred to as the current picture for clarity of explanation. A Network Abstraction Layer (NAL) unit is a syntax structure that contains a Raw Byte Sequence Payload (RBSP), a representation of the type of data, and data in the form of emulation prevention bytes interspersed as necessary. A Video Coding Layer (VCL) NAL unit is a NAL unit encoded to contain video data, such as an encoded slice of an image. A non-VCL NAL unit is a NAL unit that contains non-video data, such as syntax and / or parameters that support decoding of video data, performing compliance checks, or other operations. An Access Unit (AU) is a set of NAL units associated with each other according to a specified classification rule and related to one specific output time. A Decoding Unit (DU) is an AU or a subset of an AU and associated non-VCL NAL units. For example, an AU includes a VCL NAL unit and any non-VCL NAL unit associated with the VCL NAL unit within the AU. Further, a DU includes a set or a subset of VCL NAL units from an AU, as well as any non-VCL NAL unit associated with the VCL NAL unit within the DU. A layer is a set of VCL NAL units sharing specified characteristics (e.g., common resolution, frame rate, image size, etc.) and associated non-VCL NAL units. The decoding order is the order in which syntax elements are processed by the decoding process.A Video Parameter Set (VPS) is a data unit that contains parameters related to an entire video.
[0044] A temporally scalable bitstream is a bitstream encoded in multiple layers that provide various temporal resolutions / frame rates (e.g., each layer is encoded to support a different frame rate). A sublayer is a temporally scalable layer of a temporally scalable bitstream that includes VCL NAL units having a specific temporal identifier value and associated non-VCL NAL units. For example, a temporal sublayer is a layer that contains video data associated with a specified frame rate. A sublayer representation is a subset of the bitstream that includes NAL units of a specific sublayer and lower sublayers. Thus, it is possible to achieve a sublayer representation that can be decoded into a video sequence having a specified frame rate by combining one or more temporal sublayers. An Output Layer Set (OLS) is a set of layers in which one or more layers are designated as output layers. An output layer is a layer designated for output (e.g., to a display). An OLS index is an index that uniquely identifies the corresponding OLS. The zero-th (0-th) OLS includes only the lowest layer (the layer having the lowest layer identifier) and thus is an OLS that includes only the output layer. A temporal identifier (ID) is a data element that indicates that the data corresponds to a temporal position within a video sequence. A sub-bitstream extraction process is a process of removing NAL units that do not belong to a target set determined by a target OLS index and a target maximum temporal ID from the bitstream. The sub-bitstream extraction process outputs a sub-bitstream that includes NAL units that are part of the target set from the bitstream.
[0045] HRD is a decoder model that operates on an encoder, checks the variability of the bitstream generated by the encoding process, and verifies its compliance with specified constraints. The bitstream compliance test is a test for determining whether the encoded bitstream conforms to a standard such as VVC (Versatile Video Coding). HRD parameters are syntax elements that initialize and / or define the operating conditions of HRD. Sequence-level HRD parameters are HRD parameters applied to the entire coded video sequence. The maximum HRD time ID (hrd_max_tid[i]) specifies the time ID of the highest sublayer representation for which the HRD parameters are included in the i-th OLS HRD parameter set. The general HRD parameters syntax structure is a syntax structure that includes sequence-level HRD parameters. An operation point (OP) is a time subset of OLS and is identified by the OLS index and the highest time ID. The test target OP (targetOp) is the OP selected for the compliance test in HRD. The target OLS is the OLS selected for extraction from the bitstream. The decoding unit HRD parameter presence flag (decoding_unit_hrd_params_present_flag) is a flag indicating whether the corresponding HRD parameters operate at the DU level or the AU level. The coded picture buffer (CPB) is a first-in-first-out buffer in HRD that contains coded pictures in decoding order for use during bitstream compliance verification. The decoded picture buffer (DPB) is a buffer for holding decoded pictures for reference, output reordering, and / or output delay.
[0046] The supplementary enhancement information (SEI) message is a syntax structure with specified semantics that conveys information not required by the decoding process to determine the sample values of the decoded image. The scalable nesting SEI message is a message that contains one or more OLSs or multiple SEI messages corresponding to one or more layers. The non-scalable nesting SEI message is a non-nested message and thus contains a single SEI message. The buffering period (BP) SEI message is an SEI message that contains HRD parameters for initializing the HRD to manage the CPB. The picture timing (PT) SEI message is an SEI message that contains HRD parameters for managing the delivery information of AUs in the CPB and / or DPB. The decoding unit information (DUI) SEI message is an SEI message that contains HRD parameters for managing the delivery information for DUs in the CPB and / or DPB.
[0047] The CPB removal delay is the period during which the corresponding current AU can stay in the CPB before being removed and output to the DPB. The initial CPB removal delay is the default CPB removal delay for each picture, AU, and / or DU in the bitstream, OLS, and / or layer. The CPB removal offset is the location within the CPB that is used to determine the boundary of the corresponding AU in the CPB. The initial CPB removal offset is the default CPB removal offset associated with each picture, AU, and / or DU in the bitstream, OLS, and / or layer. The decoded picture buffer (DPB) output delay information is the period during which the corresponding AU can stay in the DPB before output. The CPB removal delay information is the information related to the removal of the corresponding DU from the CPB. The delivery schedule specifies the timing for the delivery of video data to and / or from memory locations such as the CPB and / or DPB. The VPS layer ID (vps_layer_id) is a syntax element that indicates the layer ID of the i-th layer indicated in the VPS. The number of output layer sets minus 1 (num_output_layer_sets_minus1) is a syntax element that specifies the total number of OLSs specified by the VPS. The HRD coded picture buffer count (hrd_cpb_cnt_minus1) is a syntax element that specifies the number of alternative CPB delivery schedules. The sublayer CPB parameters present flag (sublayer_cpb_params_present_flag) is a syntax element that specifies whether the set of OLS HRD parameters includes the HRD parameters for the specified sublayer representation. The schedule index (ScIdx) is an index that identifies the delivery schedule. The BP CPB count minus 1 (bp_cpb_cnt_minus1) is a syntax element that specifies the number of pairs of initial CPB removal delay and offset, and thus the number of delivery schedules available for the temporal sublayer. The NAL unit header layer identifier (nuh_layer_id) is a syntax element that specifies the identifier of the layer containing the NAL unit.The fixed_pic_rate_general_flag syntax element is a syntax element that specifies whether to constrain the temporal distance between the HRD output times of consecutive pictures in output order. The sublayer_hrd_parameters syntax structure is a syntax structure that contains the HRD parameters of the corresponding sublayer. The general_vcl_hrd_params_present_flag is a flag that specifies whether VCL HRD parameters are present in the syntax structure of the general HRD parameters. The bp_max_sublayers_minus1 syntax element is a syntax element that specifies the maximum number of temporal sublayers in which the CPB removal delay and CPB removal offset are indicated in the BP SEI message. The vps_max_sublayers_minus1 syntax element is a syntax element that specifies the maximum number of temporal sublayers that can exist in the layer specified by the VPS. The scalable_nesting_OLS_flag is a flag that specifies whether the scalable nesting SEI message is applied to a specific OLS or to a specific layer. The num_olss_minus1 is a syntax element that specifies the number of OLSs to which the scalable nesting SEI message is applied. The NestingOlsIdx is a syntax element that specifies the OLS index of the OLS to which the scalable nesting SEI message is applied. The targetOlsIdx is a variable that identifies the OLS index of the OLS to be decoded. The OLS-1 is a syntax element that specifies the total number of OLSs specified in the VPS.
[0048] The following acronyms, access unit (AU), coding tree block (CTB), coding tree unit (CTU), coding unit (CU), coded layer video sequence (CLVS), coded layer video sequence start (CLVSS), coded video sequence (CVS), coded video sequence start (CVSS), Joint Video Experts Team (JVET), hypothetical reference decoder (HRD), motion constrained tile set (MCTS), maximum transfer unit (MTU), network abstraction layer (NAL), output layer set (OLS), picture order count (POC), random access point (RAP), raw byte sequence payload (RBSP), sequence parameter set (SPS), video parameter set (VPS), Versatile Video Coding (VVC) are used in this specification.
[0049] To minimize data loss and reduce the size of video files, many video compression techniques can be employed. For example, video compression techniques can include performing spatial (e.g., within an image) prediction and / or temporal (e.g., between images) prediction to reduce or remove data redundancy in a video sequence. In the case of block-based video coding, a video slice (e.g., a video image or a portion of a video image) may be divided into video blocks that may also be referred to as tree blocks, coded tree blocks (CTBs), coded tree units (CTUs), coding units (CUs), and / or coding nodes. Video blocks within an intra-coded (I) slice of an image are coded using spatial prediction against reference samples of adjacent blocks within the same image. Video blocks within an inter-coded unidirectional prediction (P) slice or bidirectional prediction (B) slice of an image may be coded by employing spatial prediction against reference samples of adjacent blocks within the same image or temporal prediction against reference samples of other reference images. A picture may sometimes be referred to as a frame and / or an image, and a reference picture may sometimes be referred to as a reference frame and / or a reference image. Spatial or temporal prediction results in a predicted block representing the image block. Residual data represents the difference in pixels between the original image block and the predicted block. Thus, an inter-coded block is coded according to a motion vector indicating a block of reference samples forming the predicted block and residual data indicating the difference between the coded block and the predicted block. An intra-coded block is coded according to an intra coding mode and residual data. For further compression, the residual data may be transformed from a pixel domain to a transform domain. As a result, residual transform coefficients that can be quantized are obtained. The quantized transform coefficients may first be arranged in a two-dimensional array. The quantized transform coefficients may be scanned to generate a one-dimensional vector of the transform coefficients. Entropy coding may be applied to achieve further compression. Such video compression techniques will be described in more detail below.
[0050] To ensure that the encoded video can be accurately decoded, the video is encoded and decoded according to the corresponding video coding standard. Video coding standards include International Telecommunication Union (ITU) Telecommunication Standardization Sector (ITU-T) H.261, International Organization for Standardization / International Electrotechnical Commission (ISO / IEC) Motion Picture Experts Group (MPEG)-1 Part2, ITU-T H.262 or ISO / IEC MPEG-2 Part2, ITU-T H.263, ISO / IEC MPEG-4 Part2, ITU-T H.264 or Advanced Video Coding (AVC) also known as ISO / IEC MPEG-4 Part10, and ITU-T H.265 or High Efficiency Video Coding (HEVC) also known as MPEG-H Part2. AVC includes extensions such as Scalable Video Coding (SVC), Multiview Video Coding (MVC), and Multiview Video Coding plus Depth (MVC+D), as well as three-dimensional (3D) AVC (3D-AVC). HEVC includes extensions such as Scalable HEVC (SHVC), Multiview HEVC (MV-HEVC), and 3D HEVC (3D-HEVC). The joint video experts team (JVET) of ITU-T and ISO / IEC has started developing a video coding standard called Versatile Video Coding (VVC). VVC is included in the Working Draft (WD) including JVET-O2001-v14.
[0051] A video coding system employs various compliance tests to ensure that a bitstream is decodable by a decoder. For example, the compliance check can include testing the entire bitstream for compliance, then testing each layer of the bitstream for compliance, and finally checking the potential decodable output for compliance. To perform the compliance check, corresponding parameters are included in the bitstream. A Hypothetical Reference Decoder (HRD) can read the parameters and execute the tests. A video can include multiple layers and multiple different Output Layer Sets (OLSs). In response to a request, the encoder transmits one or more layers of a selected OLS. For example, the encoder can transmit the best layer from the OLSs supported by the current network bandwidth. The first problem with this approach is that a significant number of layers are tested but not actually transmitted to the decoder. However, the parameters for supporting such tests may still be included in the bitstream, unnecessarily increasing the size of the bitstream.
[0052] In a first embodiment, a mechanism for applying bitstream compliance tests only to each OLS is disclosed herein. In this way, when testing the corresponding OLS, the entire bitstream, each layer, and the decodable output are tested together. Thus, the number of compliance tests is reduced, thereby reducing the usage of processor and memory resources at the encoder. Further, by reducing the number of compliance tests, the number of associated parameters included in the bitstream can be reduced. This results in a smaller bitstream size, and thus, reduced utilization of processor, memory, and / or network resources at both the encoder and the decoder.
[0053] The second problem is that the signaling process of HRD parameters used in the HRD compliance test in some video coding systems can become complex in a multi-layer context. For example, a set of HRD parameters can be signaled for each layer of each OLS. Such HRD parameters can be signaled at different locations within the bitstream depending on the intended range of the parameters. As a result, the scheme becomes more complex as more layers and / or OLSs are added. Furthermore, the HRD parameters of different layers and / or OLSs may contain redundant information.
[0054] In a second embodiment, a mechanism for signaling a global set of HRD parameters for an OLS and corresponding layers is disclosed herein. For example, all sequence-level HRD parameters applied to all OLSs and all layers included in the OLSs are signaled in a video parameter set (VPS). The VPS is signaled once in the bitstream, and thus the sequence-level HRD parameters are signaled once. Furthermore, the sequence-level HRD parameters may be constrained to be the same for all OLSs. In this way, redundant signaling is reduced and coding efficiency is improved. Also, by this approach, the HRD process is simplified. As a result, the usage of signaling resources of the processor, memory, and / or network is reduced in both the encoder and the decoder.
[0055] The third problem may occur when a video coding system performs a bitstream compliance check. The video may be coded into multiple layers and / or sub-layers, which can then be assembled into an OLS. Each layer and / or sub-layer of each OLS is checked for compliance according to a delivery schedule. Each delivery schedule is associated with a different coded picture buffer (CPB) size and CPB delay in order to take into account different transmission bandwidths and system capabilities. Some video coding systems allow each sub-layer to define any number of delivery schedules. This can result in a large amount of signaling to support the compliance check, and as a result, the coding efficiency of the bitstream will decrease.
[0056] In a third embodiment, a mechanism for enhancing the coding efficiency of a video including multiple layers is disclosed herein. Specifically, all layers and / or sub-layers are constrained to include the same number of CPB delivery schedules. For example, the encoder can determine the maximum number of CPB delivery schedules used for any one layer and set the number of CPB delivery schedules for all layers to this maximum number. Then, the number of delivery schedules may be signaled once, for example, as part of the HRD parameters in the VPS. This eliminates the need to signal several schedules for each layer / sublayer. In some examples, all layers / sublayers of the OLS can also share the same delivery schedule index. These changes reduce the amount of data used to signal data related to the compliance check. This reduces the bitstream size and thus reduces the utilization of processor, memory, and / or network resources in both the encoder and decoder.
[0057] The fourth problem may occur when the video is coded into multiple layers and / or sub-layers, and then they are arranged in the OLS. The OLS may include a zero-th (0-th) OLS that contains only the output layer. A supplementary enhancement information (SEI) message may be included in the bitstream to notify the HRD of layer / OLS-specific parameters used to test multiple layers of the bitstream for compliance with the standard. Specifically, when the OLS is included in the bitstream, a scalable nesting SEI message is adopted. The scalable nesting SEI message includes a group of nested SEI messages applied to one or more OLSs and / or one or more layers of the OLS. Each of the nested SEI messages may include an indicator for indicating the association with the corresponding OLS and / or layer. The nested SEI messages are configured for use in multiple layers, and when applied to the zero-th OLS containing a single layer, may include irrelevant information.
[0058] In a fourth embodiment, a mechanism for enhancing the coding efficiency of a video including a zero-th OLS is disclosed herein. A non-scalable nesting SEI message is adopted for the zero-th OLS. The non-scalable nesting SEI message is restricted to be applied only to the zero-th OLS, and thus only to the output layer included in the zero-th OLS. In this way, irrelevant information such as nesting relationships and layer indications can be omitted from the SEI message. The non-scalable nesting SEI message may be used as a buffering period (BP) SEI message, a picture timing (PT) SEI message, a decoding unit (DU) SEI message, or a combination thereof. These changes reduce the amount of data used to signal compliance check-related information for the zero-th OLS. This reduces the bitstream size, and thus reduces the utilization of processor, memory, and / or network resources in both the encoder and decoder.
[0059] The fifth problem may also occur when the video is separated into multiple layers and / or sublayers. The encoder can encode these layers into a bitstream. Further, the encoder can use an HRD that performs a compliance test to check the bitstream for compliance with the standard. The encoder may be configured to include layer-specific HRD parameters in the bitstream to support such a compliance test. The layer-specific HRD parameters may be encoded for each layer in some video coding systems. In some cases, the layer-specific HRD parameters are the same for each layer, resulting in redundant information that unnecessarily increases the size of the video encoding.
[0060] In a fifth embodiment, a mechanism for reducing HRD parameter redundancy is disclosed herein for videos employing multiple layers. The encoder can encode the HRD parameters of the top layer. The encoder can also encode a sublayer CPB parameters present flag (sublayer_cpb_params_present_flag). Setting the sublayer_cpb_params_present_flag to 0 can indicate that all lower layers should use the same HRD parameters as the top layer. In this context, the top layer has the largest layer identifier (ID), and a lower layer is any layer having a layer ID smaller than the layer ID of the top layer. In this way, the HRD parameters of the lower layers can be omitted from the bitstream. This reduces the bitstream size and thus reduces the utilization of processor, memory, and / or network resources in both the encoder and decoder.
[0061] The sixth problem relates to the usage of the sequence parameter set (SPS) that contains syntax elements associated with each video sequence in the video. The video coding system can encode the video in layers and / or sub - layers. The video sequences may behave differently in different layers and / or sub - layers. Thus, different layers may refer to different SPSs. The BP SEI message can indicate the layer / sublayer for which compliance to the standard is checked. In some video coding systems, the BP SEI message may indicate that it applies to the layer / sublayer indicated by the SPS. This can cause a problem of unexpected errors because when different layers refer to different SPSs, such SPSs may contain conflicting information.
[0062] In a sixth embodiment, a mechanism for handling errors related to compliance checking when multiple layers are employed in a video sequence is disclosed herein. Specifically, the BP SEI message is modified to indicate that it can check the compliance of any number of layers / sublayers described in the VPS. For example, the BP SEI message can include a BP maximum sublayers minus 1 (bp_max_sublayers_minus1) syntax element that indicates the number of layers / sublayers associated with the data in the BP SEI message. On the other hand, the VPS maximum sublayers minus 1 (vps_max_sublayers_minus1) syntax element in the VPS indicates the number of sublayers in the entire video. The bp_max_sublayers_minus1 syntax element can be set to any value from 0 to the value of the vps_max_sublayers_minus1 syntax element. In this way, the compliance of any number of layers / sublayers within the video can be checked while avoiding layer-based sequence problems related to SPS inconstancy. Thus, the present disclosure avoids layer-based coding errors and, therefore, improves the functionality of the encoder and / or decoder. Further, this embodiment supports layer-based coding that can improve coding efficiency. Thus, this embodiment supports a reduction in the usage of processor, memory, and / or network resources in the encoder and / or decoder.
[0063] The seventh problem relates to the layers included in the OLS. Each OLS includes at least one output layer configured as displayed by the decoder. The HRD of the encoder can check each OLS for compliance with the standard. A compliant OLS can always be decoded and displayed by a compliant decoder. The HRD process may be partially managed by the SEI message. For example, a scalable nesting SEI message can include scalable nested SEI messages. Each scalable nested SEI message can include data related to the corresponding layer. When performing the compliance check, the HRD can perform a bitstream extraction process on the target OLS. Data not related to the layers within the OLS is generally removed before the compliance test (e.g., before transmission) so that each OLS can be checked separately. Depending on the video coding system, since the scalable nesting SEI message is related to multiple layers, there are some that do not remove such a message during the sub-bitstream extraction process. For this reason, even when the scalable nesting SEI message is not related to any layer of the target OLS (the OLS being extracted), the scalable nesting SEI message may remain in the bitstream after the sub-bitstream extraction. This can increase the size of the final bitstream without providing any additional functionality.
[0064] In a seventh embodiment, a mechanism for reducing the size of a multi-layer bitstream is disclosed herein. During sub-bitstream extraction, scalable nested SEI messages may be considered for removal from the bitstream. If a scalable nested SEI message is associated with one or more OLSs, the scalable nested SEI messages within the scalable nested SEI message are checked. If a scalable nested SEI message is not associated with any layer of the target OLS, the entire scalable nested SEI message can be removed from the bitstream. As a result, the size of the bitstream sent to the decoder is reduced. Thus, this embodiment increases coding efficiency and reduces the usage of processor, memory, and / or network resources in both the encoder and the decoder.
[0065] FIG. 1 is a flowchart of an exemplary operational method 100 for coding a video signal. Specifically, a video signal is encoded by an encoder. The encoding process compresses the video signal by employing various mechanisms to reduce the video file size. By reducing the file size, the compressed video file can be transmitted to the user while reducing the associated bandwidth overhead. A decoder decodes the compressed video file and reconstructs the original video signal for display to the end user. The decoding process generally mirrors the encoding process so that the decoder can consistently reconstruct the video signal.
[0066] In step 101, a video signal is input into an encoder. For example, the video signal may be an uncompressed video file stored in a memory. As another example, the video file may be captured by a video capture device such as a video camera and encoded to support live streaming of the video. The video file can include both an audio component and a video component. The video component includes a series of image frames, which, when viewed continuously, give a visually moving impression. The frames include pixels represented with respect to light, called the luma component (or luma samples) in this specification, and color, called the chroma component (or color samples). In some examples, the frames can also include depth values to support a 3D view.
[0067] In step 103, the video is divided into blocks. The division includes re - dividing the pixels within each frame into square and / or rectangular blocks for compression. For example, in High Efficiency Video Coding (HEVC), also known as H.265 and MPEG - H Part 2, a frame can first be divided into Coding Tree Units (CTUs), which are blocks of a predefined size (e.g., 64 pixels × 64 pixels). A CTU includes both luma samples and chroma samples. By adopting a coding tree, the CTU can be divided into blocks, and the blocks can be recursively re - divided until a configuration that supports further encoding is achieved. For example, the luma component of a frame may be re - divided until the individual blocks contain relatively homogeneous lighting values. Further, the chroma component of a frame may be re - divided until the individual blocks contain relatively uniform color values. Thus, the division mechanism varies according to the content of the video frame.
[0068] In step 105, various compression mechanisms are employed to compress the image blocks divided in step 103. For example, inter prediction and / or intra prediction can be adopted. Inter prediction is designed to utilize the fact that objects tend to appear in consecutive frames within a common scene. Therefore, it is not necessary to repeatedly describe the blocks indicating the objects in the reference frame in adjacent frames. Specifically, an object such as a table may remain at a fixed position over multiple frames. Thus, the table is described once, and adjacent frames can refer to the reference frame. The object can be matched over multiple frames using a pattern matching mechanism. Furthermore, a moving object may be represented over multiple frames, for example, due to the movement of the object or the movement of the camera. As a specific example, a video may show a car moving across the screen over multiple frames. Such movement can be described using motion vectors. A motion vector is a two-dimensional vector that provides an offset from the coordinates of an object in a frame to the coordinates of the object in the reference frame. Therefore, inter prediction can encode an image block in the current frame as a set of motion vectors indicating the offsets from the corresponding blocks in the reference frame.
[0069] Intra prediction encodes a block within a common frame. Intra prediction utilizes the fact that the luma and chroma components tend to cluster within a frame. For example, a green patch of a part of a tree tends to be placed adjacent to similar green patches. Intra prediction employs multiple directional prediction modes (e.g., 33 in HEVC), a planar mode, and a direct current (DC) mode. The directional mode indicates that the current block is similar / identical to the samples of adjacent blocks in the corresponding direction. The planar mode indicates that a series of blocks along a row / column (e.g., a plane) can be interpolated based on adjacent blocks at the edge of the row. The planar mode effectively indicates a smooth transition of light / color across the row / column by using a relatively constant slope when changing values. The DC mode is employed for boundary smoothing and indicates that the block is similar / identical to the average value associated with the samples of all adjacent blocks associated with the angular direction of the directional prediction mode. Thus, an intra prediction block can represent an image block as various relational prediction mode values rather than actual values. Furthermore, an inter prediction block can represent an image block as motion vector values rather than actual values. In either case, the prediction block may not accurately represent the image block in some cases. The difference is stored in a residual block. A transform can be applied to the residual block to further compress the file.
[0070] In step 107, various filtering techniques can be applied. In HEVC, the filter is applied according to the in-loop filtering method. In the block-based prediction described above, a block-shaped image may be generated in the decoder. Further, in the block-based prediction method, after encoding a block, the encoded block may be reconstructed for later use as a reference block. The in-loop filtering method repeatedly applies a noise reduction filter, a deblocking filter, an adaptive loop filter, and a sample adaptive offset (SAO) filter to blocks / frames. These filters reduce such blocking artifacts, and as a result, the encoded file can be accurately reconstructed. Further, these filters reduce the artifacts of the reconstructed reference block, and as a result, the possibility that the artifacts generate additional artifacts in subsequent blocks encoded based on the reconstructed reference block is reduced.
[0071] When the video signal is divided, compressed, and filtered, in step 109, the obtained data is encoded into a bitstream. The bitstream includes the data described above, as well as any signaling data desired to support proper video signal reconstruction in the decoder. For example, such data may include split data, prediction data, residual blocks, and various flags that give coding instructions to the decoder. The bitstream may be stored in memory for transmission to the decoder in response to a request. Also, the bitstream may be broadcast and / or multicast to multiple decoders. The creation of the bitstream is an iterative process. Therefore, steps 101, 103, 105, 107, and 109 may be performed continuously and / or simultaneously over many frames and blocks. The order shown in FIG. 1 is presented to clarify and facilitate the discussion and is not intended to limit the video coding process to a specific order.
[0072] The decoder receives a bitstream and starts the decoding process at step 111. Specifically, the decoder adopts an entropy decoding method to convert the bitstream into corresponding syntax and video data. At step 111, the decoder determines the frame partitioning using the syntax data from the bitstream. This partitioning should match the result of the block partitioning at step 103. Next, the entropy coding / decoding employed at step 111 will be described. The encoder makes many choices during the compression process, such as selecting a block partitioning method from several possible options based on the spatial arrangement of values in the input image. Signaling the exact choice may require using a large number of bins. As used herein, a bin is a binary value treated as a variable (e.g., a bit value that may vary depending on the context). Through entropy coding, the encoder can discard any option that is clearly infeasible for a particular case and retain a set of acceptable options. Each acceptable option is assigned a codeword. The length of the codeword is based on the number of acceptable options (e.g., 1 bin for 2 options, 2 bins for 3 - 4 options, etc.). Then, the encoder encodes the codeword of the selected option. This method reduces the size of the codeword because the codeword is of the desired size to uniquely indicate a selection from a small subset of acceptable options rather than from a potentially large set of all possible options. Then, the decoder decodes this selection by determining the set of acceptable options in the same way as the encoder. The decoder can read the codeword and determine the selection made by the encoder by determining the set of acceptable options.
[0073] In step 113, the decoder performs block decoding. Specifically, the decoder uses inverse transformation to generate a residual block. Next, the decoder uses the residual block and the corresponding prediction block to reconstruct an image block according to the partition. The prediction block can include both the intra-prediction block and the inter-prediction block generated by the encoder in step 105. Then, according to the partition data determined in step 111, the reconstructed image block is arranged in the frame of the reconstructed video signal. The syntax of step 113 may also be signaled in the bitstream via entropy coding as described above.
[0074] In step 115, filtering is performed on the frame of the reconstructed video signal in the same way as in step 107 in the encoder. For example, a noise suppression filter, a deblocking filter, an adaptive loop filter, and an SAO filter can be applied to the frame to remove blocking artifacts. When the frame is filtered, the video signal can be output to a display for the end user to view in step 117.
[0075] Figure 2 is a schematic diagram of an exemplary coding / decoding (codec) system 200 for video coding. Specifically, the codec system 200 provides functions that support the implementation of the operation method 100. The codec system 200 is generalized to show components used in both the encoder and the decoder. As described with respect to steps 101 and 103 of the operation method 100, the codec system 200 receives and splits a video signal, and as a result, the split video signal 201 is obtained. Next, as described with respect to steps 105, 107, and 109 of method 100, when operating as an encoder, the codec system 200 compresses the split video signal 201 into a coded bitstream. When operating as a decoder, the codec system 200 generates an output video signal from the bitstream, as described with respect to steps 111, 113, 115, and 117 of the operation method 100. The codec system 200 includes a general codec control component 211, a transform scaling and quantization component 213, an intra picture estimation component 215, an intra picture prediction component 217, a motion compensation component 219, a motion estimation component 221, a scaling and inverse transform component 229, a filter control analysis component 227, an in-loop filter component 225, a decoded picture buffer component 223, and a header formatting and context adaptive binary arithmetic coding (CABAC) component 231. Such components are coupled as shown. In Figure 2, the black lines indicate the movement of data to be encoded / decoded, and the dashed lines indicate the movement of control data that controls the operation of other components. All components of the codec system 200 may be present within the encoder. The decoder may include a subset of the components of the codec system 200.For example, the decoder can include an intra picture prediction component 217, a motion compensation component 219, a scaling and inverse transform component 229, an in-loop filter component 225, and a decoded picture buffer component 223. Here, these components will be described.
[0076] The segmented video signal 201 is a captured video sequence that has been segmented into blocks of pixels by a coding tree. The coding tree uses various splitting modes to further split the blocks of pixels into smaller blocks of pixels. These blocks can then be further split into even smaller blocks. The blocks are sometimes referred to as nodes on the coding tree. Larger parent nodes are split into smaller child nodes. The number of times a node is split is called the depth of the node / coding tree. The segmented blocks may sometimes be included in a coding unit (CU). For example, a CU can be a lower part of a CTU that includes a luma block, a chroma red difference (Cr) block, and a chroma blue difference (Cb) block, along with the corresponding syntax instruction for the CU. The splitting modes can include a binary tree (BT), a ternary tree (TT), and a quaternary tree (QT) that are employed to split a node into two, three, or four child nodes of various shapes depending on the splitting mode used. The segmented video signal 201 is transferred to a general coder control component 211, a transform scaling and quantization component 213, an intra picture estimation component 215, a filter control analysis component 227, and a motion estimation component 221 for compression.
[0077] The general coder control component 211 is configured to make decisions related to coding the images of a video sequence into a bitstream according to application constraints. For example, the general coder control component 211 manages the optimization of the bitrate / bitstream size with respect to the reconstructed quality. Such decisions can be made based on memory space / bandwidth availability and image resolution requests. The general coder control component 211 also manages the buffer utilization considering the transmission speed and reduces the problems of buffer underrun and overrun. To manage these problems, the general coder control component 211 manages the splitting, prediction, and filtering by other components. For example, the general coder control component 211 can dynamically increase the compression complexity to improve the resolution and increase the bandwidth usage, or decrease the compression complexity to reduce the resolution and bandwidth usage. Therefore, the general coder control component 211 controls other components of the codec system 200 to balance the relationship between the video signal reconstructed quality and the bitrate. The general coder control component 211 creates control data for controlling the operations of other components. The control data is also transferred to the header formatting and CABAC component 231, encoded into the bitstream, and signals the parameters for decoding at the decoder.
[0078] The split video signal 201 is also sent to the motion estimation component 221 and the motion compensation component 219 for inter prediction. The frames or slices of the split video signal 201 may be split into a plurality of video blocks. The motion estimation component 221 and the motion compensation component 219 perform relative inter prediction coding of the received video block with respect to one or more blocks in one or more reference frames and provide temporal prediction. The codec system 200 can execute a plurality of coding paths, for example, to select an appropriate coding mode for each block of video data.
[0079] The motion estimation component 221 and the motion compensation component 219 may be highly integrated, but are shown separately for conceptual purposes. The motion estimation performed by the motion estimation component 221 is a process of generating a motion vector that estimates the motion of a video block. The motion vector can indicate, for example, the relative displacement of the coded object with respect to the prediction block. The prediction block is a block that is found to closely match the block to be coded from the perspective of pixel differences. The prediction block may also be referred to as a reference block. Such pixel differences may be determined by the sum of absolute differences (SAD), the sum of squared differences (SSD), or other difference metrics. HEVC employs several coding objects including a coding tree unit (CTU), a coding tree block (CTB), and a coding unit (CU). For example, a CTU can be divided into CTBs, and then the CTBs can be divided into CUs for inclusion in the CUs. A CU can be coded as a prediction unit (PU) containing prediction data and / or a transform unit (TU) containing residual data transformed for the CU. The motion estimation component 221 generates motion vectors, PUs, and TUs by using rate-distortion analysis as part of the rate-distortion optimization process. For example, the motion estimation component 221 can determine a plurality of reference blocks, a plurality of motion vectors, etc. for the current block / frame, and can select a reference block, a motion vector, etc. having the best rate-distortion characteristics. The best rate-distortion characteristics balance both the quality of video reconstruction (e.g., the amount of data loss due to compression) and the coding efficiency (e.g., the size of the final coding).
[0080] In some examples, the codec system 200 can calculate values for sub-integer pixel positions of a reference image stored in the decoded image buffer component 223. For example, the video codec system 200 can interpolate values for 1 / 4 pixel positions, 1 / 8 pixel positions, or other fractional pixel positions of the reference image. Thus, the motion estimation component 221 can perform a relative motion search between full pixel positions and fractional pixel positions and output motion vectors with fractional pixel accuracy. The motion estimation component 221 calculates the motion vector of the prediction unit (PU) of the reference image of a video block in an inter-coded slice by comparing the position of the PU with the position of the prediction block. The motion estimation component 221 outputs the calculated motion vector as motion data to the header formatting and CABAC component 231 for encoding and outputs the motion to the motion compensation component 219.
[0081] Motion compensation performed by the motion compensation component 219 can include fetching or generating a prediction block based on the motion vector determined by the motion estimation component 221. Also in this case, the motion estimation component 221 and the motion compensation component 219 may be functionally integrated in some examples. Upon receiving the motion vector for the current video block's PU, the motion compensation component 219 can identify the position of the prediction block indicated by the motion vector. Then, by subtracting the pixel values of the prediction block from the pixel values of the currently encoded video block to form a pixel difference value, a residual video block is formed. Generally, the motion estimation component 221 performs motion estimation on the luma component, and the motion compensation component 219 uses the motion vector calculated based on the luma component for both the chroma component and the luma component. The prediction block and the residual block are transferred to the transform scaling and quantization component 213.
[0082] The segmented video signal 201 is also sent to the intra picture estimation component 215 and the intra picture prediction component 217. Similar to the motion estimation component 221 and the motion compensation component 219, the intra picture estimation component 215 and the intra picture prediction component 217 may be highly integrated, but are shown separately for conceptual purposes. The intra picture estimation component 215 and the intra picture prediction component 217 perform intra prediction of the current block for a block within the current frame as an alternative to inter prediction performed by the motion estimation component 221 and the motion compensation component 219 between frames as described above. In particular, the intra picture estimation component 215 determines the intra prediction mode to be used to encode the current block. In some examples, the intra picture estimation component 215 selects an intra prediction mode appropriate for encoding the current block from a plurality of tested intra prediction modes. The selected intra prediction mode is then transferred to the header formatting and CABAC component 231 for encoding.
[0083] For example, the intra picture estimation component 215 calculates rate distortion values using rate distortion analysis for various tested intra prediction modes and selects the intra prediction mode having the best rate distortion characteristics among the tested modes. Rate distortion analysis generally determines the amount of distortion (or error) between the encoded block and the original unencoded block that was encoded to produce the encoded block, as well as the bit rate (e.g., number of bits) used to produce the encoded block. The intra picture estimation component 215 calculates a ratio from the distortion and rate of various encoded blocks and determines which intra prediction mode exhibits the best rate distortion value for the block. Additionally, the intra picture estimation component 215 may be configured to encode depth blocks of a depth map using a depth modeling mode (DMM) based on rate distortion optimization (RDO).
[0084] When the intra picture prediction component 217 is implemented on the encoder, it can generate a residual block from a prediction block based on the selected intra prediction mode determined by the intra picture estimation component 215, or when implemented on the decoder, it can read the residual block from the bitstream. The residual block includes the difference in values between the prediction block and the original block and is represented as a matrix. The residual block is then transferred to the transform scaling and quantization component 213. The intra picture estimation component 215 and the intra picture prediction component 217 can operate on both the luma component and the chroma component.
[0085] The transform scaling and quantization component 213 is configured to further compress the residual block. The transform scaling and quantization component 213 applies a transform, such as a discrete cosine transform (DCT), a discrete sine transform (DST), or a conceptually similar transform, to the residual block to generate a video block containing residual transform coefficient values. A wavelet transform, an integer transform, a subband transform, or other types of transforms can also be used. The transform can transform the residual information from the pixel value domain to a transform domain such as the frequency domain. The transform scaling and quantization component 213 is also configured to scale the transformed residual information, for example, based on frequency. Such scaling includes applying a scale factor to the residual information so that different frequency information is quantized at different granularities, which can affect the final visual quality of the reconstructed video. The transform scaling and quantization component 213 is also configured to quantize the transform coefficients to further reduce the bitrate. The quantization process can reduce the bit depth associated with some or all of the coefficients. The degree of quantization can be adjusted by adjusting the quantization parameter. In some examples, the transform scaling and quantization component 213 can then perform a scan of the matrix containing the quantized transform coefficients. The quantized transform coefficients are transferred to the header formatting and CABAC component 231 and encoded into the bitstream.
[0086] The scaling and inverse transform component 229 applies the inverse operations of the transform scaling and quantization component 213 to support motion estimation. The scaling and inverse transform component 229 applies inverse scaling, transform, and / or quantization, for example, to reconstruct the residual block in the pixel domain for later use as a reference block that can become a predicted block for another current block. The motion estimation component 221 and / or the motion compensation component 219 can calculate the reference block by adding the residual block to the corresponding predicted block for use in motion estimation of subsequent blocks / frames. A filter is applied to the reconstructed reference block to reduce artifacts generated during scaling, quantization, and transform. Such artifacts can typically cause inaccurate predictions (and generate additional artifacts) when predicting subsequent blocks.
[0087] The filter control analysis component 227 and the in-loop filter component 225 apply filters to residual blocks and / or reconstructed image blocks. For example, a transformed residual block from the scaling and inverse transform component 229 can be combined with a corresponding prediction block from the intra-image prediction component 217 and / or the motion compensation component 219 to reconstruct an original image block. A filter can then be applied to the reconstructed image block. In some examples, a filter can be applied to the residual block instead. Like the other components in FIG. 2, the filter control analysis component 227 and the in-loop filter component 225 may be highly integrated and implemented together, but are shown separately for conceptual purposes. The filters applied to reconstructed reference blocks are applied to specific spatial regions and include multiple parameters for adjusting how such filters are applied. The filter control analysis component 227 analyzes the reconstructed reference blocks to determine where such filters should be applied and set the corresponding parameters. Such data is forwarded to the header formatting and CABAC component 231 as filter control data for encoding. The in-loop filter component 225 applies such filters based on the filter control data. The filters may include deblocking filters, noise suppression filters, SAO filters, and adaptive loop filters. Such filters may be applied in the spatial / pixel domain (e.g., to reconstructed pixel blocks) or in the frequency domain, depending on the implementation.
[0088] When operating as an encoder, the filtered and reconstructed image blocks, residual blocks, and / or prediction blocks are stored in the decoded image buffer component 223 for later use in the motion estimation described above. When operating as a decoder, the decoded image buffer component 223 stores the reconstructed and filtered blocks and transfers them to the display as part of the output video signal. The decoded image buffer component 223 may be any memory device capable of storing prediction blocks, residual blocks, and / or reconstructed image blocks.
[0089] The header formatting and CABAC component 231 receives data from various components of the codec system 200 and encodes such data into a coded bitstream for transmission towards a decoder. Specifically, the header formatting and CABAC component 231 generates various headers for encoding control data such as general control data and filter control data. Further, all prediction data including intra prediction and motion data, as well as residual data in the form of quantized transform coefficient data, are also encoded into the bitstream. The final bitstream contains all the information desired by the decoder to reconstruct the original segmented video signal 201. Such information may include an intra prediction mode index table (also called a codeword mapping table), the definition of coding contexts for various blocks, an indication of the most accurate intra prediction mode, an indication of segmentation information, etc. Such data can be encoded by using entropy coding. For example, the information can be encoded by employing context-adaptive variable-length coding (CAVLC), CABAC, syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or another entropy coding technique. Following entropy coding, the encoded bitstream may be transmitted to another device (e.g., a video decoder) or archived for later transmission or retrieval.
[0090] FIG. 3 is a block diagram showing an exemplary video encoder 300. The video encoder 300 can be employed to implement the encoding function of the codec system 200 and / or to implement steps 101, 103, 105, 107, and / or 109 of the operation method 100. The encoder 300 splits the input video signal, resulting in a split video signal 301 that is substantially similar to the split video signal 201. The split video signal 301 is then compressed by components of the encoder 300 and encoded into a bitstream.
[0091] Specifically, the split video signal 301 is transferred to an intra-picture prediction component 317 for intra prediction. The intra-picture prediction component 317 may be substantially similar to the intra-picture estimation component 215 and the intra-picture prediction component 217. The split video signal 301 is also transferred to a motion compensation component 321 for inter prediction based on reference blocks in a decoded picture buffer component 323. The motion compensation component 321 may be substantially similar to the motion estimation component 221 and the motion compensation component 219. The prediction blocks and residual blocks from the intra-picture prediction component 317 and the motion compensation component 321 are transferred to a transform and quantization component 313 for transformation and quantization of the residual blocks. The transform and quantization component 313 may be substantially similar to the transform scaling and quantization component 213. The transformed and quantized residual blocks and the corresponding prediction blocks (along with associated control data) are transferred to an entropy coding component 331 for encoding into a bitstream. The entropy coding component 331 may be substantially similar to the header formatting and CABAC component 231.
[0092] The transformed and quantized residual blocks and / or corresponding prediction blocks are also transferred from the transform and quantization component 313 to the inverse transform and quantization component 329 to be reconstructed into reference blocks for use by the motion compensation component 321. The inverse transform and quantization component 329 may be substantially similar to the scaling and inverse transform component 229. Depending on the embodiment, the in-loop filter of the in-loop filter component 325 is also applied to the residual blocks and / or the reconstructed reference blocks. The in-loop filter component 325 may be substantially similar to the filter control analysis component 227 and the in-loop filter component 225. The in-loop filter component 325 may include multiple filters as described for the in-loop filter component 225. The filtered blocks are then stored in the decoded picture buffer component 323 for use as reference blocks by the motion compensation component 321. The decoded picture buffer component 323 may be substantially similar to the decoded picture buffer component 223.
[0093] FIG. 4 is a block diagram showing an exemplary video decoder 400. The video decoder 400 can be employed to implement the decoding function of the codec system 200 and / or to implement steps 111, 113, 115, and / or 117 of the operation method 100. The decoder 400 receives, for example, a bitstream from the encoder 300 and generates an output video signal reconstructed based on the bitstream for display to an end user.
[0094] The bitstream is received by the entropy decoding component 433. The entropy decoding component 433 is configured to implement an entropy decoding method such as CAVLC, CABAC, SBAC, PIPE coding, or other entropy coding techniques. For example, the entropy decoding component 433 can adopt header information to provide a context for interpreting additional data encoded as codewords in the bitstream. The decoded information includes any desired information for decoding the video signal, such as general control data, filter control data, segmentation information, motion data, prediction data, and quantized transform coefficients from residual blocks. The quantized transform coefficients are transferred to the inverse transform and quantization component 429 for reconstruction into residual blocks. The inverse transform and quantization component 429 may be similar to the inverse transform and quantization component 329.
[0095] The reconstructed residual block and / or prediction block is transferred to the intra-picture prediction component 417 to reconstruct the picture block based on the intra-prediction operation. The intra-picture prediction component 417 may be similar to the intra-picture estimation component 215 and the intra-picture prediction component 217. Specifically, the intra-picture prediction component 417 identifies the position of the reference block within the frame using the prediction mode, and applies the residual block to the result to reconstruct the intra-predicted picture block. The reconstructed intra-predicted picture block and / or residual block and the corresponding inter-prediction data are transferred to the decoded picture buffer component 423 via the in-loop filter component 425, which may be substantially similar to the decoded picture buffer component 223 and the in-loop filter component 225, respectively. The in-loop filter component 425 filters the reconstructed picture block, residual block, and / or prediction block, and such information is stored in the decoded picture buffer component 423. The reconstructed picture block from the decoded picture buffer component 423 is transferred to the motion compensation component 421 for inter-prediction. The motion compensation component 421 may be substantially similar to the motion estimation component 221 and / or the motion compensation component 219. Specifically, the motion compensation component 421 generates a prediction block using the motion vector from the reference block, and applies the residual block to the result to reconstruct the picture block. The resulting reconstructed block may also be transferred to the decoded picture buffer component 423 via the in-loop filter component 425. The decoded picture buffer component 423 continues to store additional reconstructed picture blocks that can be reconstructed into the frame via the segmentation information. Such a frame can also be placed in the sequence. The sequence is output to the display as the reconstructed output video signal.
[0096] FIG. 5 is a schematic diagram showing an exemplary HRD500. The HRD500 can be employed in an encoder such as the codec system 200 and / or the encoder 300. The HRD500 can check the bitstream created in step 109 of method 100 before the bitstream is transferred to a decoder such as the decoder 400. In some examples, the bitstream may be continuously transferred through the HRD500 when the bitstream is encoded. If a portion of the bitstream does not conform to the associated constraints, the HRD500 can indicate such a failure to the encoder and cause the encoder to re-encode the corresponding section of the bitstream with a different mechanism.
[0097] The HRD500 includes a virtual stream scheduler (HSS) 541. The HSS541 is a component configured to execute a virtual delivery mechanism. The virtual delivery mechanism is used to check the bitstream or decoder compliance with respect to the timing and data flow of the bitstream 551 input to the HRD500. For example, the HSS541 can receive the bitstream 551 output from the encoder and manage the compliance test process for the bitstream 551. In a particular example, the HSS541 can control the rate at which the encoded image moves through the HRD500 and verify that the bitstream 551 does not contain non-conforming data.
[0098] HSS541 can transfer the bitstream 551 to the CPB543 at a predetermined rate. The HRD500 can manage data in the decoding unit (DU) 553. The DU553 is an access unit (AU) or a subset of AUs, and the associated non-video coding layer (VCL) network abstraction layer (NAL) unit. Specifically, an AU includes one or more pictures associated with the output time. For example, an AU can include a single picture in a single-layer bitstream and can include pictures of each layer in a multi-layer bitstream. Each picture of an AU can be divided into slices each contained in a corresponding VCL NAL unit. Thus, the DU553 can include one or more pictures, one or more slices of a picture, or a combination thereof. Also, the parameters used to decode an AU, a picture, and / or a slice can be included in a non-VCL NAL unit. Thus, the DU553 includes non-VCL NAL units containing the data necessary to support the decoding of the VCL NAL units in the DU553. The CPB543 is the first-in first-out buffer in the HRD500. The CPB543 contains the DU553 including video data in decoding order. The CPB543 stores video data for use during bitstream conformity verification.
[0099] The CPB543 transfers the DU553 to the decoding process component 545. The decoding process component 545 is a component compliant with the VVC standard. For example, the decoding process component 545 can emulate the decoder 400 used by an end user. The decoding process component 545 decodes the DU553 at a rate achievable by an exemplary end-user decoder. If the decoding process component 545 cannot decode the DU553 fast enough to prevent an overflow of the CPB543, the bitstream 551 does not conform to the standard and needs to be re-encoded.
[0100] The decoding process component 545 decodes the DU 553 and generates the decoded DU 555. The decoded DU 555 includes the decoded image. The decoded DU 555 is transferred to the DPB 547. The DPB 547 may be substantially similar to the decoded image buffer components 223, 323, and / or 423. To support inter prediction, the image marked to be used as the reference image 556 obtained from the decoded DU 555 is returned to the decoding process component 545 to support further decoding. The DPB 547 outputs the decoded video sequence as a series of images 557. The images 557 are reconstructed images that entirely mirror the images encoded in the bitstream 551 by the encoder.
[0101] The images 557 are transferred to the output clipping component 549. The output clipping component 549 is configured to apply a conforming clipping window to the images 557. Thereby, the cropped output image 559 is obtained. The cropped output image 559 is a fully reconstructed image. Thus, the cropped output image 559 mimics what the end user would expect to see when decoding the bitstream 551. In this way, the encoder can review the cropped output image 559 to ensure that the encoding is satisfactory.
[0102] HRD500 is initialized based on the HRD parameters within bitstream 551. For example, HRD500 can read the HRD parameters from a VPS, SPS, and / or SEI message. Then, HRD500 can perform a compliance test operation on bitstream 551 based on the information of such HRD parameters. As a specific example, HRD500 can determine one or more CPB delivery schedules 561 from the HRD parameters. The delivery schedule specifies the timing for the delivery of video data to and / or from memory locations such as CPB and / or DPB. Thus, CPB delivery schedule 561 specifies the delivery timing of AUs, DUs 553, and / or images to / from CPB543. For example, CPB delivery schedule 561 can describe the bitrate and buffer size of CPB543, and such bitrate and buffer size correspond to a specific class of decoder and / or network conditions. Thus, CPB delivery schedule 561 can indicate how long data can stay in CPB543 before backoff. The inability to maintain CPB delivery schedule 561 in HRD500 during the compliance test indicates that the decoder corresponding to CPB delivery schedule 561 cannot decode the corresponding bitstream. Note that HRD500 can use a DPB delivery schedule similar to CPB delivery schedule 561 for DPB547.
[0103] The video may be coded into different layers and / or OLSs for use by decoders having various levels of hardware capabilities and for various network conditions. The CPB delivery schedule 561 is selected to reflect these issues. Thus, the upper layer sub-bitstream is specified for optimal hardware and network conditions, and thus the upper layer can receive one or more CPB delivery schedules 561 that use a large amount of memory within the CPB 543 and a short delay for the transfer of the DU 553 to the DPB 547. Similarly, the lower layer sub-bitstream is specified for limited decoder hardware capabilities and / or poor network conditions. Thus, the lower layer can receive one or more CPB delivery schedules 561 that use a small amount of memory within the CPB 543 and a longer delay for the transfer of the DU 553 to the DPB 547. Then, the OLS, layer, sub-layer, or combinations thereof can be tested according to the corresponding delivery schedule 561 to ensure that the resulting sub-bitstream can be correctly decoded under the conditions expected for the sub-bitstream. Each of the CPB delivery schedules 561 is associated with a schedule index (ScIdx) 563. The ScIdx 563 is an index that identifies the delivery schedule. Thus, the HRD parameters within the bitstream 551 can not only indicate the CPB delivery schedule 561 by the ScIdx 563, but also the HRD 500 can contain sufficient data to determine the CPB delivery schedule 561 and correlate the CPB delivery schedule 561 to the corresponding OLS, layer, and / or sub-layer.
[0104] FIG. 6 is a schematic diagram showing an exemplary multi-layer video sequence 600 configured for inter-layer prediction 621. The multi-layer video sequence 600 may be encoded by an encoder such as codec system 200 and / or encoder 300, for example, according to method 100, and decoded by a decoder such as codec system 200 and / or decoder 400. Further, the multi-layer video sequence 600 can be checked for compliance by an HRD such as HRD 500. The multi-layer video sequence 600 is included to show an exemplary application of layers in an encoded video sequence. The multi-layer video sequence 600 is any video sequence that uses multiple layers such as layer N 631 and layer N+1 632.
[0105] In one example, the multi-layer video sequence 600 can employ inter-layer prediction 621. Inter-layer prediction 621 is applied between images 611, 612, 613, and 614 of different layers and images 615, 616, 617, and 618. In the illustrated example, images 611, 612, 613, and 614 are part of layer N+1 632, and images 615, 616, 617, and 618 are part of layer N 631. Layers such as layer N 631 and / or layer N+1 632 are groups of images all associated with the same characteristic values such as similar size, quality, resolution, signal-to-noise ratio, capabilities, etc. A layer can be formally defined as a set of VCL NAL units and associated non-VCL NAL units. A VCL NAL unit is a NAL unit encoded to contain video data, such as an encoded slice of an image. A non-VCL NAL unit is a NAL unit that contains non-video data such as syntax and / or parameters that support decoding of video data, performing compliance checks, or other operations.
[0106] In the illustrated example, layer N+1 632 is associated with a larger image size than layer N 631. Thus, in this example, the images 611, 612, 613, and 614 of layer N+1 632 have a larger image size (e.g., greater height and width, and thus more samples) than the images 615, 616, 617, and 618 of layer N 631. However, such images can be separated between layer N+1 632 and layer N 631 by other characteristics. Only two layers, layer N+1 632 and layer N 631, are shown, but a set of images can be separated into any number of layers based on the associated characteristics. Layer N+1 632 and layer N 631 may also be denoted by layer IDs. A layer ID is an item of data associated with an image that indicates that the image is part of a specified layer. Thus, each of the images 611-618 can be associated with a corresponding layer ID to indicate which of layer N+1 632 or layer N 631 includes the corresponding image. For example, the layer ID can include a NAL unit header layer identifier (nuh_layer_id) that is a syntax element specifying an identifier of a layer that includes a NAL unit (e.g., including slices and / or parameters of images within the layer). A layer associated with a lower quality / bitstream size, such as layer N 631, is generally assigned a lower layer ID and is called a lower layer. Further, a layer associated with a higher quality / bitstream size, such as layer N+1 632, is generally assigned a higher layer ID and is called an upper layer.
[0107] The images 611 - 618 of different layers 631 - 632 are configured to be selectively displayed. Thus, the images of different layers 631 - 632 can share the time ID 622 as long as these images are included in the same AU. The time ID 622 is a data element indicating that the data corresponds to a time position within the video sequence. An AU is a set of NAL units associated with each other according to a specified classification rule and related to a specific output time. For example, an AU can include one or more images in different layers, such as image 611 and image 615, when such images are associated with the same time ID 622. As a specific example, when a smaller image is desired, the decoder can decode and display image 615 at the current display time, or when a larger image is desired, the decoder can decode and display image 611 at the current display time. Thus, the images 611 - 614 of the upper layer N+1 632 include substantially the same image data as the corresponding images 615 - 618 of the lower layer N631 (despite the difference in image size). Specifically, image 611 includes substantially the same image data as image 615, image 612 includes substantially the same image data as image 616, and so on.
[0108] Images 611 to 618 can be coded by referring to other images 611 to 618 of the same layer N631 or N+1 632. Coding an image by referring to another image of the same layer results in an inter prediction 623. The inter prediction 623 is depicted by the solid arrows. For example, image 613 may be coded by adopting an inter prediction 623 that uses one or two of images 611, 612, and / or 614 of layer N+1 632 as references. One image is referred to for uni-directional inter prediction and / or two images are referred to for bi-directional inter prediction. Further, image 617 may be coded by adopting an inter prediction 623 that uses one or two of images 615, 616, and / or 618 of layer N531 as references. One image is referred to for uni-directional inter prediction and / or two images are referred to for bi-directional inter prediction. When performing the inter prediction 623, if an image is used as a reference to another image of the same layer, that image may be called a reference image. For example, image 612 may be a reference image used to code image 613 according to the inter prediction 623. The inter prediction 623 may also be called an intra-layer prediction in a multi-layer context. Therefore, the inter prediction 623 is a mechanism for coding samples of a current image by referring to the indicated samples in a reference image different from the current image, and the reference image and the current image are in the same layer.
[0109] The images 611 to 618 can also be coded by referring to other images 611 to 618 of different layers. This process is known as inter-layer prediction 621 and is depicted by the dashed arrows. The inter-layer prediction 621 is a mechanism for coding the samples of the current image by referring to the indicated samples in the reference image, where the current image and the reference image are in different layers and thus have different layer IDs. For example, the image of the lower layer N631 can be used as a reference image for coding the corresponding image of the upper layer N+1 632. As a specific example, the image 611 can be coded by referring to the image 615 according to the inter-layer prediction 621. In such a case, the image 615 is used as an inter-layer reference image. The inter-layer reference image is the reference image used for the inter-layer prediction 621. In most cases, the inter-layer prediction 621 is restricted so that the current image such as the image 611 can only use the inter-layer reference images in the lower layer such as the image 615 that are included in the same AU. If multiple layers (for example, three or more) are available, the inter-layer prediction 621 can encode / decode the current image based on a plurality of inter-layer reference images at a level lower than the current image.
[0110] The video encoder can encode images 611-618 through many different combinations and / or permutations of inter prediction 623 and inter-layer prediction 621 using the multi-layer video sequence 600. For example, image 615 may be encoded according to intra prediction. Then, images 616-618 can be encoded according to inter prediction 623 by using image 615 as a reference image. Further, image 611 may be encoded according to inter-layer prediction 621 by using image 615 as an inter-layer reference image. Then, images 612-614 can be encoded according to inter prediction 623 by using image 611 as a reference image. Thus, the reference image can serve both as a single-layer reference image and an inter-layer reference image for different coding mechanisms. By coding the image of the upper layer N+1 632 based on the image of the lower layer N 631, the upper layer N+1 632 can avoid adopting intra prediction with much lower coding efficiency than inter prediction 623 and inter-layer prediction 621. Thus, the low coding efficiency of intra prediction can be limited to the minimum / lowest quality images, and thus can be limited to the coding of the minimum amount of video data. The image used as a reference image and / or an inter-layer reference image can be indicated by an entry in the reference image list included in the reference image list structure.
[0111] To perform such operations, layers such as layer N631 and layer N+1 632 may be included in one or more OLSs 625 and 626. Specifically, images 611-618 are encoded as layers 631-632 of bitstream 600, and then each layer 631-632 of the image is assigned to one or more of OLSs 625 and 626. Then, OLS 625 and / or 626 can be selected, and depending on the capabilities in the decoder and / or network conditions, the corresponding layers 631 and / or 632 can be sent to the decoder. OLS 625 is a set of layers in which one or more layers are designated as output layers. The output layer is the layer designated for output (e.g., to a display). For example, layer N631 may be included only to support inter-layer prediction 621 and may never be output. In such a case, layer N+1 632 is decoded and output based on layer N631. In such a case, OLS 625 includes layer N+1 632 as the output layer. If the OLS includes only the output layer, that OLS is called the 0th OLS 626. The 0th OLS 626 is an OLS that includes only the lowest layer (the layer with the lowest layer identifier) and thus an OLS that includes only the output layer. In other cases, OLS 625 may include many layers in different combinations. For example, the output layer of OLS 625 can be encoded according to inter-layer prediction 621 based on one, two, or many lower layers. Further, OLS 625 can include two or more output layers. Thus, OLS 625 can include one or more output layers and any support layers necessary to reconstruct the output layers. Only two OLSs 625 and 626 are shown, but the multi-layer video sequence 600 may be encoded by employing many different OLSs 625 and / or 626, each employing a different combination of layers. OLSs 625 and 626 are each associated with an OLS index 629, which is an index that uniquely identifies the corresponding OLSs 625 and 626.
[0112] Checking the compliance of the multi-layer video sequence 600 with HRD500 may become complicated depending on layers 631 - 632 and OLSs 625 and 626. HRD500 can separate the multi-layer video sequence 600 into a series of operation points 627 for testing. OLSs 625 and / or 626 are identified by the OLS index 629. The operation point 627 is a temporal subset of OLS 625 / 626. The operation point 627 can be identified by both the OLS index 629 of the corresponding OLS 625 / 626 and the highest temporal ID 622. As a specific example, the first operation point 627 can include all the images within the first OLS 625 from temporal ID 0 to temporal ID 200, the second operation point 627 can include all the images within the first OLS 625 from temporal ID 201 to temporal ID 400, and so on. In such a case, the first operation point 627 is described by the OLS index 629 of the first OLS 625 and the temporal ID of 200. Further, the second operation point 627 is described by the OLS index 629 of the first OLS 625 and the temporal ID of 400. The operation point 627 selected to be tested at a specified instant is called the test target OP (targetOp). Therefore, targetOp is the operation point 627 selected for the compliance test in HRD500.
[0113] FIG. 7 is a schematic diagram showing an exemplary multi-layer video sequence 700 configured for temporal scalability. The multi-layer video sequence 700 can be encoded by an encoder such as codec system 200 and / or encoder 300, for example, according to method 100, and decoded by a decoder such as codec system 200 and / or decoder 400. Further, the multi-layer video sequence 700 can be checked for compliance by an HRD such as HRD 500. The multi-layer video sequence 700 is included to show another exemplary application for layers in an encoded video sequence. For example, the multi-layer video sequence 700 may be employed as a separate embodiment or may be combined with the techniques described with respect to the multi-layer video sequence 600.
[0114] The multi-layer video sequence 700 includes sub-layers 710, 720, and 730. The sub-layers are temporal scalable layers of a temporally scalable bitstream that include VCL NAL units (e.g., pictures) having specific time identifier values as well as associated non-VCL NAL units (e.g., support parameters). For example, layers such as layer N631 and / or layer N+1 632 can be further divided into sub-layers 710, 720, and 730 to support temporal scalability. Sub-layer 710 may be referred to as the base layer, and sub-layers 720 and 730 may be referred to as enhancement layers. As shown, sub-layer 710 includes pictures 711 at a first frame rate, such as 30 frames per second. Since sub-layer 710 includes the base / lowest frame rate, sub-layer 710 is the base layer. Sub-layer 720 includes pictures 721 that are temporally offset from pictures 711 of sub-layer 710. As a result, sub-layer 710 and sub-layer 720 can be combined, resulting in a higher overall frame rate than the frame rate of sub-layer 710 alone. For example, sub-layers 710 and 720 can together have a frame rate of 60 frames per second. Thus, sub-layer 720 improves the frame rate of sub-layer 710. Further, sub-layer 730 also includes pictures 731 that are temporally offset from pictures 721 and 711 of sub-layers 720 and 710. Thus, sub-layer 730 can be combined with sub-layers 720 and 710 to further enhance sub-layer 710. For example, sub-layers 710, 720, and 730 can together have a frame rate of 90 frames per second.
[0115] The sublayer representation 740 can be dynamically created by combining sublayers 710, 720, and / or 730. The sublayer representation 740 is a subset of the bitstream that includes NAL units of specific sublayers and lower sublayers. In the illustrated example, the sublayer representation 740 includes an image 741 that is the composite images 711, 721, and 731 of sublayers 710, 720, and 730. Thus, the multi-layer video sequence 700 can be temporally scaled to a desired frame rate by selecting a sublayer representation 740 that includes a desired set of sublayers 710, 720, and / or 730. The sublayer representation 740 may be created by adopting an OLS that includes sublayers 710, 720, and / or 730 as layers. In such a case, the sublayer representation 740 is selected as the output layer. Thus, temporal scalability is one of several mechanisms that can be achieved using a multi-layer mechanism.
[0116] FIG. 8 is a schematic diagram showing an exemplary bitstream 800. For example, the bitstream 800 can be generated by an encoder 300 and / or a codec system 200 for decoding by a decoder 400 and / or a codec system 200 according to method 100. Further, the bitstream 800 can include a multi-layer video sequence 600 and / or 700. Further, the bitstream 800 can include various parameters for controlling the operation of an HRD such as HRD 500. Based on such parameters, the HRD can check the bitstream 800 for compliance with the standard before transmitting it to the decoder for decoding.
[0117] The bitstream 800 includes a VPS 811, one or more SPSs 813, a plurality of Picture Parameter Sets (PPSs) 815, a plurality of slice headers 817, image data 820, and SEI messages 819. The VPS 811 includes data related to the entire bitstream 800. For example, the VPS 811 can include data related to the OLS, layer, and / or sublayer used in the bitstream 800. The SPS 813 includes sequence data common to all images within the coded video sequence included in the bitstream 800. For example, each layer can include one or more coded video sequences, and each coded video sequence can reference the SPS 813 for corresponding parameters. The parameters of the SPS 813 can include image sizing, bit depth, coding tool parameters, bitrate limits, and the like. Each sequence references the SPS 813, but note that in some examples, a single SPS 813 can include data related to multiple sequences. The PPS 815 includes parameters applied to the entire picture. Thus, each image within the video sequence can reference the PPS 815. Each image references the PPS 815, but note that in some examples, a single PPS 815 can include data for multiple images. For example, multiple similar images may be coded according to similar parameters. In such cases, a single PPS 815 can include data for such similar images. The PPS 815 can indicate coding tools, quantization parameters, offsets, etc. available for the slices of the corresponding image.
[0118] Slice header 817 contains parameters specific to each slice within an image. Thus, there may be one slice header 817 for each slice within a video sequence. The slice header 817 can include slice type information, POC, reference picture list, prediction weights, tile entry point, deblocking parameters, etc. Note that in some examples, the bitstream 800 can also include an image header, which is a syntax structure containing parameters applicable to all slices within a single image. For this reason, the image header and the slice header 817 may be used interchangeably in some contexts. For example, certain parameters can be moved between the slice header 817 and the image header depending on whether such parameters are common to all slices within the image.
[0119] The image data 820 includes video data encoded according to inter-prediction and / or intra-prediction, as well as corresponding transformed and quantized residual data. For example, the image data 820 can include AU821, DU822, and / or image 823. AU821 is a set of NAL units associated with each other according to specified classification rules and related to one specific output time. DU822 is an AU or a subset of AUs, and related non-VCL NAL units. Image 823 is an array of luma samples and / or an array of chroma samples that create a frame or its fields. Put simply, AU821 includes various video data that may be displayed at a specified instant within a video sequence, as well as supporting syntax data. Thus, AU821 can include a single image 823 in a single-layer bitstream, or multiple images from multiple layers all associated with the same instant in a multi-layer bitstream. On the other hand, image 823 is an encoded image that can be output for display or used to support the encoding of other image 823s for output. DU822 can include one or more image 823s and any supporting syntax data necessary for decoding. For example, DU822 and AU821 may be interchangeably used in a simple bitstream (e.g., when the AU includes a single image). However, in a more complex multi-layer bitstream, DU822 can include only a portion of the video data from AU821. For example, AU821 can include image 823s in several layers and / or sub-layers, and some of the image 823s are associated with different OLSs. In such a case, DU822 can include only the image 823s from the specified OLS and / or the specified layer / sublayer.
[0120] Image 823 includes one or more slices 825. The slice 825 may be defined as an integer number of complete tiles (e.g., within a tile) or an integer number of consecutive complete coding tree units (CTUs) of the image 823, and the tile or CTU row is exclusively included in a single NAL unit 829. Thus, the slice 825 is also included in a single NAL unit 829. The slice 825 is further divided into CTUs and / or coding tree blocks (CTBs). A CTU is a group of samples of a predefined size that can be divided by a coding tree. A CTB is a subset of a CTU and includes the luma component or chroma component of the CTU. The CTU / CTB is further divided into coded blocks based on a coding tree. Thereafter, the coded blocks can be encoded / decoded according to a prediction mechanism.
[0121] The bitstream 800 is a sequence of NAL units 829. The NAL unit 829 is a container for video data and / or support syntax. The NAL unit 829 can be a VCL NAL unit or a non-VCL NAL unit. A VCL NAL unit is a NAL unit 829 encoded to include video data such as an encoded slice 825 and an associated slice header 817. A non-VCL NAL unit is a NAL unit 829 that includes non-video data such as syntax and / or parameters that support decoding of video data, performing compliance checks, or other operations. For example, the non-VCL NAL unit can include a VPS 811, an SPS 813, a PPS 815, an SEI message 819, or other support syntax.
[0122] SEI message 819 is a syntax structure with specified semantics that conveys information not required by the decoding process to determine the sample values of the decoded image. For example, the SEI message can include data for supporting the HRD process or other support data not directly related to decoding the bitstream 800 in the decoder. SEI message 819 can include scalable nesting SEI messages and / or non-scalable nested SEI messages. A scalable nesting SEI message is a message that includes one or more OLSs or multiple SEI messages corresponding to one or more layers. A non-scalable nested SEI message is a non-nested message and thus includes a single SEI message. SEI message 819 can include a BP SEI message that contains HRD parameters for initializing HRD to manage the CPB. SEI message 819 can also include a PT SEI message that contains HRD parameters for managing the delivery information about AU821 in the CPB and / or DPB. SEI message 819 can also include a DUI SEI message that contains HRD parameters for managing the delivery information about DU822 in the CPB and / or DPB.
[0123] The bitstream 800 includes a set of integer (i) HRD parameters 833, which are syntax elements that initialize and / or define the operating conditions of an HRD such as the HRD 500. In some examples, the general_hrd_parameters syntax structure can include the HRD parameters 833 that apply to all OLSs specified by the VPS 811. In one example, an encoder can encode a video sequence into layers. The encoder can then encode the HRD parameters 833 into the bitstream to properly configure the HRD to perform compliance checks. The HRD parameters 833 can also indicate to the decoder that the decoder can decode the bitstream according to a delivery schedule. The HRD parameters 833 can be included in the VPS 811 and / or the SPS 813. Additional parameters used to configure the HRD may also be included in the SEI message 819.
[0124] As described above, a video stream can include many OLSs and many layers, such as OLS625, layer N631, layer N+1 632, sublayer 710, sublayer 720, and / or sublayer 730. Further, some layers may be included in multiple OLSs. Therefore, a multi-layer video sequence such as multi-layer video sequence 600 and / or 700 can be quite complex. As a result, the bitstream compliance check process in HRD can become complex. Some video coding systems use layer-specific HRD parameters 833 for each layer / sublayer. HRD reads the layer-specific HRD parameters 833 from the bitstream 800 and then performs a bitstream compliance test for each layer based on the HRD parameters 833. In some cases, some of the various layers / sublayers use the same HRD parameters 833. As a result, redundant HRD parameters 833 are encoded into the bitstream 800, reducing coding efficiency. Further, with this approach, redundant information is repeatedly retrieved from the bitstream 800 by HRD, wasting the encoder's memory and / or processor resources. Therefore, redundant HRD parameters 833 can waste the processor, memory, and / or network resources of the encoder and / or decoder.
[0125] The present disclosure includes a mechanism for reducing the redundancy of HRD parameters 833 for videos using multiple layers. If the HRD parameters 833 are the same for all sublayers in the sublayer representation and / or OLS, the encoder can encode the HRD parameters 833 for the top layer. The encoder can also encode the sublayer_cpb_params_present_flag 831. The sublayer_cpb_params_present_flag 831 is a syntax element that specifies whether a set of HRD parameters (e.g., for OLS) includes the HRD parameters for the specified sublayer / sublayer representation. Setting the sublayer_cpb_params_present_flag 831 to 0 can indicate that all lower layers should use the same HRD parameters as the top layer. Setting the sublayer_cpb_params_present_flag 831 to 1 can also indicate that each layer includes separate (e.g., different) HRD parameters 833. Thus, setting the sublayer_cpb_params_present_flag 831 to 0 allows inferring that the HRD parameters 833 of the lower layers are equal to the HRD parameters 833 of the top layer. Thus, the HRD parameters 833 of the lower sublayers can be omitted from the bitstream 800 if they are the same as the HRD parameters 833 of the top sublayer to avoid redundant signaling. This mechanism reduces the size of the bitstream 800. Thus, this mechanism reduces the utilization of processor, memory, and / or network resources in both the encoder and the decoder. Further, by reducing the number of HRD parameters 833, the amount of resource usage during the HRD process in the encoder can be reduced because the set of signaled HRD parameters 833 can be read and adopted for the complete set of sublayers.
[0126] The top layer / sublayer is the layer with the highest value of the corresponding layer ID in the OLS and / or sublayer representation. In one example, VPS811 can include hrd_max_tid[i]832. hrd_max_tid[i]832 specifies the time ID of the highest sublayer representation in which the HRD parameter 833 is included in the i-th set of OLS HRD parameters 833. Thus, HRD can read the sublayer_cpb_params_present_flag831. Then, HRD can determine that the HRD parameter 833 is applied to the top layer / sublayer as indicated by hrd_max_tid[i]832 within VPS811. HRD can also infer that the same HRD parameter 833 is applied to all lower layers / sublayers with an ID smaller than hrd_max_tid[i]832.
[0127] With the above method, various redundant HRD parameters 833 can be omitted from the bitstream 800 of the lower sublayer. The omission of redundancy can be applied to some of the HRD parameters 833. In a specific example, the HRD parameters 833 can include fixed_pic_rate_general_flag 835, sublayer_hrd_parameters 837, and general_vcl_hrd_params_present_flag 839, and it can be inferred that each of these is applied to the topmost sublayer and can be applied to the lower sublayers in an equivalent manner. The fixed_pic_rate_general_flag 835 is a syntax element that specifies whether the temporal distance between the HRD output times of consecutive pictures in the output order is constrained by other HRD parameters 833. For example, the fixed_pic_rate_general_flag 835 can be set to 1 to indicate that such a constraint is applied, or set to 0 to indicate that such a constraint is not applied. The sublayer_hrd_parameters 837 is a syntax structure that includes the HRD parameters of the corresponding sublayer indicated by the sublayer ID. The general_vcl_hrd_params_present_flag 839 is a flag that specifies whether the VCL HRD parameters are present in the general HRD parameter syntax structure. For example, the general_vcl_hrd_params_present_flag 839 can be set to 1 to indicate that the VCL HRD parameters related to the first type of compliance point are present in the general HRD parameter syntax structure, or set to 0 to indicate that such VCL HRD parameters do not exist (e.g., the second type of compliance point is adopted).
[0128] Here, the above-mentioned information will be described in more detail below in this specification. Hierarchical video coding is also referred to as scalable video coding or video coding with scalability. Scalability in video coding can be supported by using multi-layer coding techniques. A multi-layer bitstream includes a base layer (BL) and one or more enhancement layers (EL). Examples of scalability include spatial scalability, quality / signal-to-noise ratio (SNR) scalability, multi-view scalability, frame rate scalability, and the like. When using multi-layer coding techniques, an image or a part thereof may be coded without using a reference image (intra prediction), may be coded by referring to a reference image in the same layer (inter prediction), and / or may be coded by referring to a reference image in another layer (inter-layer prediction). The reference image used for inter-layer prediction of the current image is called an inter-layer reference image (ILRP). FIG. 6 shows an example of multi-layer coding for spatial scalability where images of different layers have different resolutions.
[0129] Some video coding families provide support for scalability in profiles that are separated from the profiles for single-layer coding. Scalable video coding (SVC) is a scalable extension of advanced video coding (AVC) that supports spatial scalability, temporal scalability, and quality scalability. In the case of SVC, a flag indicating whether each macroblock (MB) of the EL image is predicted using a collocation block from the lower layer is signaled. Prediction from the collocation block can include texture, motion vectors, and / or coding modes. The implementation of SVC cannot directly reuse the unchanged implementation of AVC in the design. The syntax and decoding process of SVC EL macroblocks are different from those of AVC.
[0130] Scalable High Efficiency Video Coding (SHVC) is an extension of HEVC that supports spatial scalability and quality scalability. Multi-View HEVC (MV-HEVC) is an extension of HEVC that supports multi-view scalability. 3D HEVC (3D-HEVC) is an extension of HEVC that supports 3D video coding that is more advanced and efficient than MV-HEVC. Temporal scalability may be included as an integral part of a single-layer HEVC codec. In the multi-layer extension of HEVC, the decoded pictures used for inter-layer prediction are only from the same access unit and are treated as long-term reference pictures (LTRPs). Such pictures are assigned reference indexes in the reference picture list together with other temporal reference pictures of the current layer. Inter-layer prediction (ILP) is achieved at the prediction unit (PU) level by setting the value of the reference index to refer to the inter-layer reference pictures in the reference picture list. Spatial scalability resamples a reference picture or a part thereof when the ILRP has a different spatial resolution from the current picture being encoded or decoded. Resampling of the reference picture can be realized either at the picture level or at the coding block level.
[0131] VVC can also support hierarchical video coding. The VVC bitstream can contain multiple layers. All layers may be independent of each other. For example, each layer can be coded without using inter-layer prediction. In this case, the layer is also called a simulcast layer. In some cases, some of the layers are coded using ILP. The flags in the VPS can indicate whether a layer is a simulcast layer or whether some layers use ILP. When some layers use ILP, the layer dependencies between the layers are also signaled in the VPS. Different from SHVC and MV-HEVC, VVC may not specify an OLS. The OLS includes a specified set of layers, and one or more layers in the set of layers are specified as output layers. The output layer is the layer of the OLS that is output. In some embodiments of VVC, when a layer is a simulcast layer, only one layer may be selected for decoding and output. In some embodiments of VVC, the entire bitstream containing all layers is specified to be decoded when any layer uses ILP. Further, a particular one of the layers is specified as the output layer. The output layer may be indicated to be only the top layer, all layers, or a set of the top layer + the indicated lower layers.
[0132] A video coding standard can specify an HRD for verifying the compliance of a bitstream through a specified HRD compliance test. In SHVC and MV-HEVC, three sets of bitstream compliance tests are adopted to check the compliance of the bitstream. The bitstream is referred to as the entire bitstream and denoted as entireBitstream. The first set of bitstream compliance tests is for testing the compliance of the entire bitstream and the corresponding temporal subsets. Such tests are adopted regardless of whether there is a layer set specified by the active VPS that includes all nuh_layer_id values of the VCL NAL units present in the entire bitstream. Thus, even if one or more layers are not included in the output set, the entire bitstream is always checked for compliance. The second set of bitstream compliance tests is adopted for testing the compliance of the layer set specified by the active VPS and the associated temporal subsets. For all these tests, only the base layer image (e.g., the image with nuh_layer_id equal to 0) is decoded and output. Other images are ignored by the decoder when the decoding process is invoked. The third set of bitstream compliance tests is adopted for testing the compliance of the OLS and the associated temporal subsets specified by the VPS extension part of the active VPS based on OLS and bitstream partitioning. Bitstream partitioning includes one or more layers of the OLS of a multi-layer bitstream.
[0133] The foregoing aspects include certain problems. For example, the first two sets of compliance tests may be applied to layers that are not decoded and output. For example, layers other than the lowest layer are not decoded and output. In an actual application, the decoder can only receive data that is decoded. Therefore, using the first two sets of compliance tests can both complicate the codec design and waste bits for carrying both sequence level parameters and picture level parameters used to support the compliance tests. The third set of compliance tests includes bitstream splitting. Such splitting may be related to one or more layers of the OLS of the multi-layer bitstream. Instead, if the compliance tests always operate separately for each layer, the HRD may be significantly simplified.
[0134] The signaling of sequence-level HRD parameters can be complex. For example, sequence-level HRD parameters may be signaled in multiple places, such as both in the SPS and the VPS. Furthermore, sequence-level HRD parameter signaling may include redundancy. For example, information that may generally be the same for the entire bitstream may be repeated at each layer of each OLS. Additionally, an exemplary HRD method allows for selecting different delivery schedules for each layer. Such a delivery schedule may be selected from a list of schedules signaled for each layer for each operation point, where the operation point is an OLS or a time subset of an OLS. Such a system is complex. Moreover, an exemplary HRD method allows for associating incomplete AUs with buffering period SEI messages. An incomplete AU is an AU that does not have images for all layers present in the CVS. However, there may be problems with HRD initialization with such AUs. For example, HRD may not be properly initialized for layers that have layer access units that do not exist in the incomplete AU. In addition, the demultiplexing process for deriving the layer bitstream may not be able to efficiently remove nested SEI messages that are not applicable to the target layer. The layer bitstream occurs when only one layer is included in the bitstream splitting. Furthermore, non-scalable nested buffering periods, picture timings, and applicable OLSs for the decode unit information SEI messages may be specified for the entire bitstream. However, the non-scalable nested buffering period should instead be applicable to the 0th OLS.
[0135] Furthermore, depending on the implementation of VVC, when sub_layer_cpb_params_present_flag is equal to 0, it may fail to infer HDR parameters. Such inferences may enable proper HRD operation. In addition, the values of bp_max_sub_layers_minus1 and pt_max_sub_layers_minus1 may need to be equal to the value of sps_max_sub_layers_minus1. However, the buffering period and picture timing SEI messages may be nested and applicable to multiple OLSs and multiple layers of each of the multiple OLSs. In such a context, the layers involved may refer to multiple SPSs. Therefore, it may be difficult for the system to track which SPS is the SPS corresponding to each layer. Therefore, the values of these two syntax elements should instead be constrained based on the value of vps_max_sub_layers_minus1. Furthermore, since different layers can have different numbers of sub-layers, the values of these two syntax elements are not always equal to a specific value in all buffering periods and picture timing SEI messages.
[0136] Also, the following issues are associated with HRD design in both SHVC / MV-HEVC and VVC. The sub-bitstream extraction process may not remove SEI NAL units containing nested SEI messages that are not required for the target OLS.
[0137] Generally, the present disclosure describes a technique for scalable nesting of SEI messages for an output layer set in a multi-layer video bitstream. The description of this technique is based on VVC. However, this technique is also applicable to hierarchical video coding based on other video codec specifications.
[0138] One or more of the above problems can be solved as follows. Specifically, the present disclosure includes methods and related aspects for HRD design that enable efficient signaling of HRD parameters using a much simpler HRD operation compared to SHVC and MV-HEVC. Each of the solutions described below addresses the above problems. For example, instead of requiring three sets of compliance tests, the present disclosure may employ only one set of compliance tests to test the compliance of the OLS specified by the VPS. Further, instead of a design based on bitstream segmentation, the disclosed HRD mechanism can always operate separately for each layer of the OLS. Further, sequence level HRD parameters that are global for all layers and sub-layers of all OLSs may be signaled, for example, only once in the VPS. In addition, a single number of delivery schedules may be signaled for all layers and sub-layers of all OLSs. The same delivery schedule index can also be applied to all layers of the OLS. Further, an incomplete AU may not be associated with the buffering period SEI message. An incomplete AU is an AU that does not include an image for all layers present in the CVS. This enables reliable initialization of HRD at all times for all layers of the OLS. Also disclosed is a mechanism for efficiently removing nested SEI messages that are not applied to the target layer in the OLS. This supports the demultiplexing process for deriving the layer bitstream. In addition, applicable OLSs for non-scalable nested buffering periods, picture timing, and decoder unit information SEI messages may be designated as the 0th OLS. Further, when sub_layer_cpb_params_present_flag is equal to 0, HDR parameters can be inferred, which may enable appropriate HRD operation. The values of bp_max_sub_layers_minus1 and pt_max_sub_layers_minus1 may need to be in the range from 0 to vps_max_sub_layers_minus1.Thus, such parameters do not need to be specific values for all buffering periods and picture timing SEI messages. Also, the sub-bitstream extraction process can remove SEI NAL units containing nested SEI messages that are not applicable to the target OLS.
[0139] Exemplary embodiments of the foregoing mechanisms are as follows. The output layer is the layer that is output among the output layer set. The OLS is a set of layers including a specified set of layers, and one or more layers within the set of layers are designated as output layers. The OLS layer index is the index of the layer within the OLS with respect to the list of layers within the OLS. The sub-bitstream extraction process is a specified process in which NAL units in the bitstream that do not belong to the target set, determined by the target OLS index and the target maximum TemporalId, are removed from the bitstream, and the output sub-bitstream includes NAL units in the bitstream that belong to the target set.
[0140] The syntax of an exemplary video parameter set is as follows. [Table 1]
[0141] The syntax of an exemplary sequence parameter set RBSP is as follows. [Table 2]
[0142] The syntax of exemplary DPB parameters is as follows. [Table 3]
[0143] The syntax of exemplary general HRD parameters is as follows. [Table 4]
[0144] The semantics of an exemplary video parameter set RBSP are as follows. By setting each_layer_is_an_ols_flag equal to 1, it is specified that each output layer set contains only one layer, each layer itself in the bitstream is the output layer set, and the single layer contained is the only output layer. By setting each_layer_is_an_ols_flag equal to 0, it is specified that the output layer set can contain two or more layers. When vps_max_layers_minus1 is equal to 0, the value of each_layer_is_an_ols_flag is presumed to be equal to 1. Otherwise, when vps_all_independent_layers_flag is equal to 0, the value of each_layer_is_an_ols_flag is presumed to be equal to 0.
[0145] By setting ols_mode_idc equal to 0, it is specified that the total number of OLSs specified by the VPS is equal to vps_max_layers_minus1 + 1. The i-th OLS includes the layers with layer indices from 0 to i, including the end values. For each OLS, only the top layer of the OLS is output. By setting ols_mode_idc equal to 1, it is specified that the total number of OLSs specified by the VPS is equal to vps_max_layers_minus1 + 1. The i-th OLS includes the layers with layer indices from 0 to i, including the end values. For each OLS, all layers of the OLS are output. By setting ols_mode_idc equal to 2, it is specified that the total number of OLSs specified by the VPS is explicitly signaled. For each OLS, a set of the top layer of the OLS and the explicitly signaled lower layers is output. The value of ols_mode_idc is in the range from 0 to 2, including the end values. The value 3 of ols_mode_idc is reserved. When vps_all_independent_layers_flag is equal to 1 and each_layer_is_an_ols_flag is equal to 0, the value of ols_mode_idc is presumed to be equal to 2. num_output_layer_sets_minus1 + 1 specifies the total number of OLSs specified by the VPS when ols_mode_idc is equal to 2.
[0146] The variable TotalNumOlss that specifies the total number of OLSs specified by the VPS is derived as follows.
Number
[0147] When ols_mode_idc is equal to 2, layer_included_flag[i][j] specifies whether to include the j-th layer (the layer with nuh_layer_id equal to vps_layer_id[j]) in the i-th OLS. By setting layer_included_flag[i][j] equal to 1, it is specified to include the j-th layer in the i-th OLS. By setting layer_included_flag[i][j] equal to 0, it is specified not to include the j-th layer in the i-th OLS.
[0148] The variable NumLayersInOls[i] that specifies the number of layers in the i-th OLS and the variable LayerIdInOls[i][j] that specifies the nuh_layer_id value of the j-th layer in the i-th OLS are derived as follows.
Number
[0149] The variable LayerIdInOls[i][j] that specifies the OLS layer index of the layer with nuh_layer_id equal to OlsLayeIdx[i][j] is derived as follows.
Number
[0150] The bottom layer in each OLS is assumed to be an independent layer. In other words, for each i in the range from 0 to TotalNumOlss - 1, including both end values, the value of vps_independent_layer_flag[GeneralLayerIdx[LayerIdInOls[i][0]]] shall be equal to 1. Each layer shall be included in at least one OLS specified by the VPS. In other words, for each layer for which a specific value of nuh_layer_id nuhLayerId is equal to one of vps_layer_id[k] for k in the range from 0 to vps_max_layers_minus1, including both end values, there may be at least one pair of values of i and j, where i is in the range from 0 to TotalNumOlss - 1, including both end values, j is in the range from 0 to NumLayersInOls[i] - 1, including both end values, and the value of LayerIdInOls[i][j] is equal to nuhLayerId. Any layer of the OLS shall be the output layer of the OLS or a (direct or indirect) reference layer of the output layer of the OLS.
[0151] vps_output_layer_flag[i][j] specifies whether the j-th layer of the i-th OLS is output when ols_mode_idc is equal to 2. vps_output_layer_flag[i] equal to 1 specifies that the j-th layer of the i-th OLS is output. Setting vps_output_layer_flag[i] equal to 0 specifies that the j-th layer of the i-th OLS is not output. When vps_all_independent_layers_flag is equal to 1 and each_layer_is_an_ols_flag is equal to 0, the value of vps_output_layer_flag[i] is assumed to be equal to 1. The variable OutputLayerFlag[i][j], where the value 1 specifies that the j-th layer of the i-th OLS is output and the value 0 specifies that the j-th layer of the i-th OLS is not output, is derived as follows.
Number
[0152] Setting vps_extension_flag equal to 0 specifies that the vps_extension_data_flag syntax element does not exist in the VPS RBSP syntax structure. Setting vps_extension_flag equal to 1 specifies that the vps_extension_data_flag syntax element exists in the VPS RBSP syntax structure. vps_extension_data_flag can have any value. The presence and value of vps_extension_data_flag do not affect the decoder's compliance with the specified profile. The decoder shall ignore all vps_extension_data_flag syntax elements.
[0153] The semantics of exemplary DPB parameters are as follows. The dpb_parameters() syntax structure provides DPB size information and, optionally, maximum picture rearrangement number and maximum waiting time (MRML) information. Each SPS contains one or the dpb_parameters() syntax structure. The first dpb_parameters() syntax structure of an SPS contains both DPB size information and MRML information. If present, the second dpb_parameters() syntax structure of an SPS contains only DPB size information. The MRML information of the first dpb_parameters() syntax structure of an SPS is applied to the layer that references the SPS, regardless of whether that layer is the output layer of the OLS. The DPB size information of the first dpb_parameters() syntax structure of an SPS is applied to the layer that references the SPS if that layer is the output layer of the OLS. The DPB size information contained in the second dpb_parameters() syntax structure of an SPS, if present, is applied to the layer that references the SPS if that layer is a non-output layer of the OLS. If an SPS contains only one dpb_parameters() syntax structure, the DPB size information for a layer as a non-output layer is presumed to be the same as that for the layer as an output layer.
[0154] The semantics of the exemplary general HRD parameters are as follows. The general_hrd_parameters() syntax structure provides the HRD parameters used in the HRD operation. When sub_layer_cpb_params_present_flag is set equal to 1, it specifies that the i-th layer_level_hrd_parameters() syntax structure contains the HRD parameters for the sub-layer representation where the TemporalId ranges from 0 to hrd_max_temporal_id[i] inclusive of both end values. When sub_layer_cpb_params_present_flag is set equal to 0, it specifies that the i-th layer_level_hrd_parameters() syntax structure contains the HRD parameters for the sub-layer representation where the TemporalId is equal to hrd_max_temporal_id[i] only. When vps_max_sub_layers_minus1 is equal to 0, the value of sub_layer_cpb_params_present_flag is assumed to be 0. When sub_layer_cpb_params_present_flag is equal to 0, the HRD parameters for the sub-layer representation where the TemporalId ranges from 0 to hrd_max_temporal_id[i]-1 inclusive of both end values are assumed to be the same as those for the sub-layer representation where the TemporalId is equal to hrd_max_temporal_id[i]. These include the HRD parameters starting from the fixed_pic_rate_general_flag[i] syntax element in the layer_level_hrd_parameters syntax structure up to the sub_layer_hrd_parameters(i) syntax structure immediately below the conditional statement if(general_vcl_hrd_params_present_flag). num_layer_hrd_params_minus1+1 specifies the number of layer_level_hrd_parameters() syntax structures present in the general_hrd_parameters() syntax structure.The value of num_layer_hrd_params_minus1 shall be in the range of 0 to 63, inclusive of both end values. hrd_cpb_cnt_minus1 + 1 specifies the number of alternative CPB specifications in the CVS bitstream. The value of hrd_cpb_cnt_minus1 shall be in the range of 0 to 31, inclusive of both end values. hrd_max_temporal_id[i] specifies the TemporalId of the top - level sub - layer representation in which the HRD parameter is included in the i - th layer_level_hrd_parameters() syntax structure. The value of hrd_max_temporal_id[i] shall be in the range of 0 to vps_max_sub_layers_minus1, inclusive of both end values. When vps_max_sub_layers_minus1 is equal to 0, the value of hrd_max_temporal_id[i] is assumed to be 0. layer_level_hrd_idx[i][j] specifies the index of the layer_level_hrd_parameters() syntax structure applied to the j - th layer of the i - th OLS. The value of layer_level_hrd_idx[[i][j] shall be in the range of 0 to num_layer_hrd_params_minus1, inclusive of both end values. If it does not exist, the value of layer_level_hrd_idx[[0][0] is assumed to be 0.
[0155] An exemplary sub-bitstream extraction process is as follows. The inputs to this process are the bitstream inBitstream, the target OLS index targetOlsIdx, and the target maximum TemporalId value tIdTarget. The output of this process is the sub-bitstream outBitstream. The requirements for bitstream compliance with respect to the input bitstream are the output of the process specified in this section with respect to the bitstream with targetOlsIdx equal to the index of the list of OLSs specified by the VPS, and tIdTarget equal to any value in the range of 0 to 6, inclusive, as inputs, and any output sub-bitstream that satisfies the following conditions is a compliant bitstream. The output sub-bitstream should contain at least one VCL NAL unit whose nuh_layer_id is equal to each of the nuh_layer_id values in LayerIdInOls[targetOlsIdx]. The output sub-bitstream should contain at least one VCL NAL unit whose TemporalId is equal to tIdTarget. A compliant bitstream contains one or more coded slice NAL units whose TemporalId is equal to 0, but need not contain a coded slice NAL unit whose nuh_layer_id is equal to 0.
[0156] The output sub-bitstream OutBitstream is derived as follows. The bitstream outBitstream is set to be the same as the bitstream inBitstream. Remove all NAL units from outBitstream for which the TemporalId is greater than tIdTarget. Remove all NAL units from outBitstream for which the nuh_layer_id is not included in the list LayerIdInOls[targetOlsIdx]. Remove all SEI NAL units from outBitstream that contain a scalable nesting SEI message for which the nesting_ols_flag is equal to 1 and the value of i is not in the range from 0 to nesting_num_olss_minus1 inclusive, where NestingOlsIdx[i] is equal to targetOlsIdx. If targetOlsIdx is greater than 0, remove all SEI NAL units from outBitstream that contain a non-scalable nesting SEI message for which the payloadType is equal to 0 (buffering period), 1 (picture timing), or 130 (decoding unit information).
[0157] Exemplary general aspects of HRD are as follows. This section specifies HRD and its usage for checking bitstream and decoder compliance. A set of bitstream compliance tests is used to check the compliance of a bitstream called the entire bitstream, denoted as entireBitstream. A set of bitstream compliance tests is for testing the compliance of each OLS specified by the VPS and each temporal subset of each OLS. For each test, the following ordered steps are applied in the order listed.
[0158] The operation point to be tested, denoted as targetOp, is selected by selecting the target OLS having the OLS index opOlsIdx and the highest TemporalId value opTid. The value of opOlsIdx ranges from 0 to TotalNumOlss - 1, inclusive. The value of opTid ranges from 0 to vps_max_sub_layers_minus1, inclusive. The values of opOlsIdx and opTid are such that the sub-bitstream BitstreamToDecode, which is the output by calling the sub-bitstream extraction process with entireBitstream, opOlsIdx, and opTid as inputs, satisfies the following conditions. There exists at least one VCL NAL unit for which nuh_layer_id is equal to each of the nuh_layer_id values of LayerIdInOls[opOlsIdx] of BitstreamToDecode. There exists at least one VCL NAL unit for which TemporalId is equal to opTid of BitstreamToDecode.
[0159] The values of TargetOlsIdx and Htid are set equal to opOlsIdx and opTid of targetOp, respectively. The value of ScIdx is selected. The selected ScIdx shall range from 0 to hrd_cpb_cnt_minus1, inclusive. The access unit within BitstreamToDecode associated with the buffering period SEI message applicable to TargetOlsIdx (present in TargetLayerBitstream or available through an external mechanism not specified herein) is selected as the HRD initialization point and is called access unit 0 for each layer of the target OLS.
[0160] The subsequent steps are applied to each layer having an OLS layer index TargetOlsLayerIdx within the target OLS. If there is only one layer in the target OLS, the test target layer bitstream TargetLayerBitstream is set to be the same as BitstreamToDecode. Otherwise, TargetLayerBitstream is derived by calling a demultiplexing process for deriving the layer bitstream with BitstreamToDecode, TargetOlsIdx, and TargetOlsLayerIdx as inputs, and the output is assigned to TargetLayerBitstream.
[0161] The layer_level_hrd_parameters() syntax structure and sub_layer_hrd_parameters() syntax structure applicable to the TargetLayerBitstream are selected as follows. The layer_level_hrd_parameters() syntax structure at layer_level_hrd_idx[TargetOlsIdx][TargetOlsLayerIdx] within the VPS (or provided through an external mechanism such as user input) is selected. In the selected layer_level_hrd_parameters() syntax structure, if BitstreamToDecode is a type I bitstream, the sub_layer_hrd_parameters(Htid) syntax structure immediately following the condition if(general_vcl_hrd_params_present_flag) is selected and the variable NalHrdModeFlag is set equal to 0. Otherwise (if BitstreamToDecode is a type II bitstream), the sub_layer_hrd_parameters(Htid) syntax structure immediately following either the condition if(general_vcl_hrd_params_present_flag) (in which case the variable NalHrdModeFlag is set equal to 0) or the condition if(general_nal_hrd_params_present_flag) (in which case the variable NalHrdModeFlag is set equal to 1) is selected. If BitstreamToDecode is a type II bitstream and NalHrdModeFlag is equal to 0, all non-VCL NAL units except filler data NAL units, and all leading_0_8bits, 0_byte, start_code_prefix_one_3bytes, and trailing_0_8 bit syntax elements that form a byte stream from the NAL unit stream, if present, are discarded from the TargetLayerBitstream and the remaining bitstream is assigned to the TargetLayerBitstream.
[0162] If the decoding_unit_hrd_params_present_flag is equal to 1, the CPB is scheduled to operate either at the access unit level (in which case the variable DecodingUnitHrdFlag is set equal to 0) or at the decoding unit level (in which case the variable DecodingUnitHrdFlag is set equal to 1). Otherwise, DecodingUnitHrdFlag is set equal to 0 and the CPB is scheduled to operate at the access unit level. For each access unit in the TargetLayerBitstream starting from access unit 0, a buffering period SEI message (present in the TargetLayerBitstream or available through an external mechanism) associated with the access unit and applied to TargetOlsIdx and TargetOlsLayerIdx is selected, an image timing SEI message (present in the TargetLayerBitstream or available through an external mechanism) associated with the access unit and applied to TargetOlsIdx and TargetOlsLayerIdx is selected, and if DecodingUnitHrdFlag is equal to 1 and decoding_unit_cpb_params_in_pic_timing_sei_flag is equal to 0, a decoding unit information SEI message (present in the TargetLayerBitstream or available through an external mechanism) associated with the decoding units in the access unit and applied to TargetOlsIdx and TargetOlsLayerIdx is selected.
[0163] Each conformance test includes a combination of one option in each of the above steps. If there are two or more options for a step, only one option is selected for any particular conformance test. All possible combinations of all steps form the complete set of conformance tests. For each operation point under test, the number of bitstream conformance tests to be executed is equal to n0*n1*n2*n3, where the values of n0, n1, n2, and n3 are specified as follows. n1 is equal to hrd_cpb_cnt_minus1 + 1, and n1 is the number of access units in BitstreamToDecode associated with the buffering period SEI message. n2 is derived as follows. If BitstreamToDecode is a type I bitstream, n0 is equal to 1. Otherwise (if BitstreamToDecode is a type II bitstream), n0 is equal to 2. n3 is derived as follows. If decoding_unit_hrd_params_present_flag is equal to 0, n3 is equal to 1. Otherwise, n3 is equal to 2.
[0164] HRD includes a bitstream demultiplexer (present optionally), coded picture buffers (CPBs) for each layer, an instantaneous decoding process for each layer, a decoded picture buffer (DPB) including sub-DPBs for each layer, and output cropping.
[0165] In one example, the HRD operates as follows. The HRD is initialized to zero in the decoding unit, and each CPB of the DPB and each sub-DPB are set to be empty. The sub-DPB fullness of each sub-DPB is set to be equal to zero. After initialization, the HRD is not re-initialized by subsequent SEI messages during the buffering period. The data associated with the decoding units flowing into each CPB according to the specified arrival schedule is delivered by the HSS. The data associated with each decoding unit is removed by the instantaneous decoding process at the CPB removal time of the decoding unit and is decoded instantaneously. Each decoded image is placed in the DPB. The decoded image is removed from the DPB when it is no longer needed for inter-prediction reference and is no longer needed for output.
[0166] In one example, the demultiplexing process for deriving the layer bitstream is as follows. The inputs to this process are the bitstream inBitstream, the target OLS index targetOlsIdx, and the target OLS layer index targetOlsLayerIdx. The output of this process is the layer bitstream outBitstream. The output layer bitstream outBitstream is derived as follows. The bitstream outBitstream is set to be the same as the bitstream inBitstream. Remove all NAL units from outBitstream for which nuh_layer_id is not equal to LayerIdInOls[targetOlsIdx][targetOlsLayerIdx]. For which nesting_ols_flag is equal to 1 and there are no values of i and j in the range from 0 to nesting_num_ols_minus1, inclusive, and from 0 to nesting_num_olss_layers_minus1[i], inclusive, such that NestingOlsLayerIdx[i][j] is equal to targetOlsLayerIdx. Remove all SEI NAL units containing scalable nesting SEI messages from outBitstream. Remove all SEI NAL units containing scalable nesting SEI messages from outBitstream for which nesting_ols_flag is equal to 1 and there are values of i and j in the range from 0 to nesting_num_olss_minus1, inclusive, and from 0 to nesting_num_ols_layers_minus1[i], inclusive, such that NestingOlsLayerIdx[i][j] is less than targetOlsLayerIdx.Remove all SEI NAL units from outBitstream that contain scalable nesting SEI messages such that nesting_ols_flag is equal to 0 and, including both end values, there is no value of i in the range from 0 to LayerIdInOls - 1, so that NestingLayerId[i] is equal to NestingNumLayers[targetOlsIdx][targetOlsLayerIdx]. Remove all SEI NAL units from outBitstream that contain scalable nesting SEI messages such that nesting_ols_flag is equal to 0 and, including both end values, there is at least one value of i in the range from 0 to LayerIdInOls - 1, so that NestingLayerId[i] is less than NestingNumLayers[targetOlsIdx][targetOlsLayerIdx].
[0167] An exemplary buffering period SEI message syntax is as follows.
Table 5
[0168] An exemplary scalable nesting SEI message syntax is as follows.
Table 6
[0169] Exemplary general SEI payload semantics are as follows. The following applies (in the context of OLS or generally) to applicable layers of the non-scalable nested SEI message. In the case of a non-scalable nested SEI message, if the payloadType is equal to 0 (buffering period), 1 (picture timing), or 130 (decoding unit information), the non-scalable nested SEI message is applied only to the lowest layer in the context of the 0-th OLS. In the case of a non-scalable nested SEI message, if the payloadType is equal to any value in the VclAssociatedSeiList, the non-scalable nested SEI message is applied only to the layer where the nuh_layer_id of the VCL NAL unit is equal to the nuh_layer_id of the SEI NAL unit containing the SEI message.
[0170] Exemplary semantics of the buffering period SEI message are as follows. The buffering period SEI message provides the initial CPB removal delay and initial CPB removal delay offset information for initializing the HRD at the position of the access unit associated in decoding order. If a buffering period SEI message exists and the TemporalId is equal to 0 and the picture is not a RASL or RADL (random access decodable leading) picture, the picture is said to be a notDiscardablePic picture. If the current picture is not the first picture in the bitstream in decoding order, prevNonDiscardablePic is set to the picture preceding in decoding order where the TemporalId is equal to 0 and the picture is not a RASL or RADL picture.
[0171] The presence of the buffering period SEI message is specified as follows. When NalHrdBpPresentFlag is equal to 1 or VclHrdBpPresentFlag is equal to 1, the following applies to each access unit in the CVS. If the access unit is an IRAP or Gradual Decoder Refresh (GDR) access unit, the buffering period SEI message applicable to the operation point shall be associated with the access unit. Otherwise, if the access unit contains notDiscardablePic, the buffering period SEI message applicable to the operation point may or may not be associated with the access unit. Otherwise, the access unit shall not be associated with the buffering period SEI message applicable to the operation point. Otherwise (when both NalHrdBpPresentFlag and VclHrdBpPresentFlag are equal to 0), none of the access units in the CVS shall be associated with the buffering period SEI message. In some applications, it may be desirable for the buffering period SEI message to be present frequently (for example, in the case of random access in an IRAP image or a non-IRAP image, or in the case of bitstream splicing). If the image of an access unit is associated with the buffering period SEI message, the access unit shall be assumed to have each image of the layer present in the CVS, and each image of the access unit shall be associated with the buffering period SEI message.
[0172] bp_max_sub_layers_minus1 + 1 specifies the maximum number of temporal sub-layers for which the CPB removal delay and CBP removal offset are indicated in the buffering period SEI message. The value of bp_max_sub_layers_minus1 shall be in the range from 0 to vps_max_sub_layers_minus1, inclusive. bp_cpb_cnt_minus1 + 1 specifies the number of syntax element pairs of nal_initial_cpb_removal_delay[i][j] and nal_initial_cpb_removal_offset[i][j] for the i-th temporal sub-layer when bp_nal_hrd_params_present_flag is equal to 1, and specifies the number of syntax element pairs of vcl_initial_cpb_removal_delay[i][j] and vcl_initial_cpb_removal_offset[i][j] for the i-th temporal sub-layer when bp_vcl_hrd_params_present_flag is equal to 1. The value of bp_cpb_cnt_minus1 shall be in the range from 0 to 31, inclusive. The value of bp_cpb_cnt_minus1 shall be equal to the value of hrd_cpb_cnt_minus1.
[0173] The semantics of an exemplary picture timing SEI message are as follows. The picture timing SEI message provides CPB removal delay and DPB output delay information of the access unit associated with the SEI message. When the bp_nal_hrd_params_present_flag or bp_vcl_hrd_params_present_flag of the buffering period SEI message applicable to the current access unit is equal to 1, the variable CpbDpbDelaysPresentFlag is set equal to 1. Otherwise, CpbDpbDelaysPresentFlag is set equal to 0. The presence of the picture timing SEI message is specified as follows. When CpbDpbDelaysPresentFlag is equal to 1, the picture timing SEI message is considered to be associated with the current access unit. Otherwise (when CpbDpbDelaysPresentFlag is equal to 0), it is considered that there is no picture timing SEI message associated with the current access unit. The TemporalId in the picture timing SEI message syntax is the TemporalId of the SEI NAL unit containing the picture timing SEI message. pt_max_sub_layers_minus1 + 1 specifies the TemporalId of the topmost sublayer representation in which the CPB removal delay information is included in the picture timing SEI message. The value of pt_max_sub_layers_minus1 is assumed to be in the range from 0 to vps_max_sub_layers_minus1 inclusive.
[0174] The semantics of an exemplary scalable nesting SEI message are as follows. The scalable nesting SEI message provides a mechanism for associating an SEI message with a specific layer within the context of a particular OLS, or a specific layer not within the context of an OLS. The scalable nesting SEI message contains one or more SEI messages. The SEI messages contained in the scalable nesting SEI message are also referred to as scalable nested SEI messages. The bitstream compliance requirement is that the following restrictions apply to including SEI messages in a scalable nesting SEI message. SEI messages with a payloadType equal to 132 (decoded picture hash) or 133 (scalable nesting) shall not be included in a scalable nesting SEI message. If a scalable nesting SEI message contains a buffering period, picture timing, or decoded unit information SEI message, the scalable nesting SEI message shall not contain any other SEI messages with a payloadType not equal to 0 (buffering period), 1 (picture timing), or 130 (decoded unit information).
[0175] The requirements for bitstream compliance are that the following restrictions apply to the value of nal_unit_type of the SEI NAL unit containing the scalable nesting SEI message. When the scalable nesting SEI message contains an SEI message whose payloadType is equal to 0 (buffering period), 1 (picture timing), 130 (decoding unit information), 145 (dependent RAP indication), or 168 (frame field information), the SEI NAL unit containing the scalable nesting SEI message shall have nal_unit_type equal to PREFIX_SEI_NUT. When the scalable nesting SEI message contains an SEI message whose payloadType is equal to 132 (decoded picture hash), the SEI NAL unit containing the scalable nesting SEI message shall have nal_unit_type equal to SUFFIX_SEI_NUT.
[0176] Set the nesting_ols_flag to 1 to specify that the scalable nesting SEI message is applied to a specific layer in the context of a specific OLS. Set the nesting_ols_flag to 0 to specify that the scalable nesting SEI message is applied to a specific layer generally (not in the context of an OLS). The bitstream compliance requirement is that the following restrictions apply to the value of the nesting_ols_flag. If the scalable nesting SEI message includes an SEI message where the payloadType is equal to 0 (buffering period), 1 (picture timing), or 130 (decoding unit information), the value of the nesting_ols_flag shall be equal to 1. If the scalable nesting SEI message includes an SEI message where the payloadType is equal to a value in the VclAssociatedSeiList, the value of the nesting_ols_flag shall be equal to 0. nesting_num_olss_minus1 + 1 specifies the number of OLSs to which the scalable nesting SEI message is applied. The value of nesting_num_olss_minus1 shall be in the range from 0 to TotalNumOlss - 1, inclusive. nesting_ols_idx_delta_minus1[i] is used to derive the variable NestingOlsIdx[i] that specifies the OLS index of the i-th OLS to which the scalable nesting SEI message is applied when the nesting_ols_flag is equal to 1. The value of nesting_ols_idx_delta_minus1[i] shall be in the range from 0 to TotalNumOlss - 2, inclusive. The variable NestingOlsIdx[i] is derived as follows.
Number
[0177] nesting_num_ols_layers_minus1[ i ]+1 specifies the number of layers to which the scalable nesting SEI message is applied in the context of the NestingOlsIdx[i]-th OLS. The value of nesting_num_ols_layers_minus1[i] shall be in the range from 0 to NumLayersInOls[NestingOlsIdx[i]]-1, inclusive. nesting_ols_layer_idx_delta_minus1[i][j] is used to derive the variable NestingOlsLayerIdx[i][j] that specifies the OLS layer index of the j-th layer to which the scalable nesting SEI message is applied in the context of the NestingOlsIdx[i]-th OLS when nesting_ols_flag is equal to 1. The value of nesting_ols_layer_idx_delta_minus1[i] shall be in the range from 0 to NumLayersInOls[nestingOlsIdx[i]]-2, inclusive. The variable NestingOlsLayerIdx[i][j] is derived as follows.
Number
[0178] For i in the range from 0 to nesting_num_olss_minus1, inclusive, the smallest value among all values of LayerIdInOls[NestingOlsIdx[i]][NestingOlsLayerIdx[i][0]] shall be equal to the nuh_layer_id of the current SEI NAL unit (the SEI NAL unit including the scalable nesting SEI message). The nesting_all_layers_flag is set to 1 to specify that the scalable nesting SEI message is generally applicable to all layers whose nuh_layer_id is greater than or equal to the nuh_layer_id of the current SEI NAL unit. The nesting_all_layers_flag is set to 0 to specify that the scalable nesting SEI message may or may not be generally applicable to all layers whose nuh_layer_id is greater than or equal to the nuh_layer_id of the current SEI NAL unit. nesting_num_layers_minus1 + 1 specifies the number of layers to which the scalable nesting SEI message is generally applicable. The value of nesting_num_layers_minus1 shall be in the range from 0 to vps_max_layers_minus1 - GeneralLayerIdx[nuh_layer_id], inclusive, where nuh_layer_id is the nuh_layer_id of the current SEI NAL unit. The nesting_layer_id[i] specifies the nuh_layer_id value of the i-th layer to which the scalable nesting SEI message is generally applicable when the nesting_all_layers_flag is equal to 0. The value of nesting_layer_id[i] shall be greater than nuh_layer_id, where nuh_layer_id is the nuh_layer_id of the current SEI NAL unit.When nesting_ols_flag is equal to 0, the variable NestingNumLayers that specifies the number of layers to which the scalable nesting SEI message is generally applied, and the list NestingLayerId[i] for i in the range from 0 to NestingNumLayers - 1, inclusive, that specifies the list of nuh_layer_id values of the layers to which the scalable nesting SEI message is generally applied, are derived as follows, where nuh_layer_id is the nuh_layer_id of the current SEI NAL unit.
Number
[0179] nesting_num_seis_minus1 + 1 specifies the number of scalable nesting SEI messages. The value of nesting_num_seis_minus1 is assumed to be in the range from 0 to 63, inclusive. nesting_0_bit is assumed to be equal to 0.
[0180] FIG. 9 is a schematic diagram of an exemplary video coding device 900. The video coding device 900 is suitable for implementing the disclosed embodiments / implementations described herein. The video coding device 900 includes a transceiver unit (Tx / Rx) 910 that includes a downstream port 920, an upstream port 950, and / or a transmitter and / or receiver for communicating data upstream and / or downstream via a network. The video coding device 900 also includes a processor 930 that includes a logic unit and / or a central processing unit (CPU) for processing data, and a memory 932 for storing data. The video coding device 900 can also include electrical, optical, or wireless communication components coupled to the upstream port 950 and / or the downstream port 920 for communication of data via an electrical, optical-electrical (OE) network, an electrical-optical (EO) network, and / or a wireless communication network. The video coding device 900 can also include an input and / or output (I / O) device 960 for communicating data with a user. The I / O device 960 can include output devices such as a display for displaying video data, a speaker for outputting audio data, etc. The I / O device 960 can also include input devices such as a keyboard, a mouse, a trackball, etc., and / or corresponding interfaces for interacting with such output devices.
[0181] Processor 930 is implemented by hardware and software. Processor 930 may be implemented as one or more CPU chips, cores (e.g., as a multi-core processor), a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), and a digital signal processor (DSP). Processor 930 communicates with downstream port 920, Tx / Rx 910, upstream port 950, and memory 932. Processor 930 includes a coding module 914. Coding module 914 implements the disclosed embodiments described herein, such as methods 100, 1000, and 1100 that can employ multi-layer video sequence 600, multi-layer video sequence 700, and / or bitstream 800. Coding module 914 can also implement any other method / mechanism described herein. Further, coding module 914 can implement codec system 200, encoder 300, decoder 400, and / or HRD 500. For example, coding module 914 may be employed to implement HRD. Further, coding module 914 is employed to encode parameters into a bitstream and can support an HRD compliance check process. Accordingly, coding module 914 may be configured to execute a mechanism for addressing one or more of the above problems. Accordingly, coding module 914 provides additional functionality and / or coding efficiency to video encoding device 900 when encoding video data. Accordingly, coding module 914 improves the functionality of video encoding device 900 and also addresses problems specific to video encoding techniques. Further, coding module 914 results in a transition of video coding device 900 to different states. Alternatively, coding module 914 can be implemented as instructions executed by processor 930 stored in memory 932 (e.g., as a computer program product stored on a non-transitory medium).
[0182] The memory 932 comprises one or more memory types such as a disk, a tape drive, a solid state drive, a read-only memory (ROM), a random access memory (RAM), a flash memory, a ternary content addressable memory (TCAM), a static random access memory (SRAM). The memory 932 is used as an overflow data storage device, can store such a program when a program is selected for execution, and can store instructions and data read during program execution.
[0183] Figure 10 is a flowchart of an exemplary method 1000 for encoding a video sequence into a bitstream by including the inferred HRD parameters to support the bitstream compliance test by the HRD. The method 1000 may be employed by an encoder such as the codec system 200, the encoder 300, and / or the video coding device 900 when executing the method 100. Further, the method 1000 can operate on the HRD 500, and thus can perform a compliance test on the multi-layer video sequence 600, the multi-layer video sequence 700, and / or the bitstream 800.
[0184] Method 1000 can start when an encoder receives a video sequence and decides to encode the video sequence into a multi-layer bitstream, for example, based on user input. In step 1001, the encoder encodes a plurality of sub-layers / sublayer representations into the bitstream. The encoder determines the HRD parameters of the sub-layers. In this example, the HRD parameters are the same for all of the plurality of sub-layers / sublayer representations. The encoder encodes a set of HRD parameters into the bitstream for the maximum sub-layer / sublayer representation. Further, the encoder encodes sublayer_cpb_params_present_flag into the bitstream. Setting sublayer_cpb_params_present_flag to 0 can indicate that the HRD parameters for the topmost sub-layer / sublayer representation apply to all sub-layers / sublayer representations. sublayer_cpb_params_present_flag may be encoded in the VPS in the bitstream.
[0185] In step 1003, the HRD reads the HRD parameters and the sublayer_cpb_params_present_flag. The HRD can then infer that when the sublayer_cpb_params_present_flag is set to 0, the HRD parameters of all lower sublayers / sublayer representations with a TemporalId less than the maximum TemporalId are equal to the HRD parameters of the maximum sublayer / sublayer representation having the maximum TemporalId. For example, multiple sublayers / sublayer representations may be associated with a TemporalId such as TemporalId 622. The TemporalId of the maximum sublayer representation can be represented as the HRD maximum TemporalId (hrd_max_tid[i]), where i indicates the i-th HRD parameter syntax structure. Thus, the TemporalId of the lower sublayers / sublayer representations may range from 0 to hrd_max_tid[i] - 1. hrd_max_tid[i] may be encoded in the VPS. Inferable HRD parameters can include, for example, the fixed_pic_rate_general_flag[i] syntax element, the sublayer_hrd_parameters(i) syntax structure, and / or the general_vcl_hrd_params_present_flag. The fixed_pic_rate_general_flag[i] is a syntax element indicating whether the temporal distance between the HRD output times of consecutive pictures in the output order is constrained. The sublayer_hrd_parameters(i) is a syntax structure containing the HRD parameters of one or more sublayers. The general_vcl_hrd_params_present_flag is a flag indicating whether the VCL HRD parameters related to the conformity point are present in the general HRD parameter syntax structure.
[0186] In step 1005, the HRD can perform a set of bitstream compliance tests on the bitstream by using the HRD parameters. Specifically, the HRD can perform compliance tests on all sublayer / sublayer representations (including lower sublayers / representations) by using the HRD parameters from the maximum sublayer / sublayer representation.
[0187] In step 1007, the encoder can store the bitstream for communication to the decoder.
[0188] FIG. 11 is a flowchart of an exemplary method 1100 for decoding a video sequence from a bitstream including inferred HRD parameters for use in a bitstream compliance test by an HRD such as HRD500. Method 1100 may be employed by a decoder such as codec system 200, decoder 400, and / or video coding device 900 when executing method 100. Further, method 1100 can operate on a bitstream such as bitstream 800 including multi-layer video sequence 600 and / or multi-layer video sequence 700.
[0189] Method 1100 can start when the decoder begins to receive a bitstream of coded data representing a multi-layer video sequence, for example, as a result of method 1000. In step 1101, the receiver receives a bitstream including a plurality of sublayer / sublayer representations. The bitstream also includes HRD parameters and sublayer_cpb_params_present_flag. Setting sublayer_cpb_params_present_flag to 0 can indicate that the HRD parameters for the topmost sublayer / sublayer representation are applied to all sublayer / sublayer representations. sublayer_cpb_params_present_flag may be encoded in the VPS in the bitstream.
[0190] In step 1103, when the sublayer_cpb_params_present_flag is set to 0, the decoder infers that the HRD parameters of all the lower sublayer representations whose TemporalId is less than the maximum TemporalId are equal to the HRD parameters of the maximum sublayer representation having the maximum TemporalId. For example, a plurality of sublayers / sublayer representations may be associated with a TemporalId such as TemporalId 622. The TemporalId of the maximum sublayer representation can be represented as the HRD maximum TemporalId (hrd_max_tid[i]), where i indicates the i-th HRD parameter syntax structure. Thus, the TemporalId of the lower sublayer / sublayer representation may range from 0 to hrd_max_tid[i] - 1. hrd_max_tid[i] may be encoded in the VPS. Inferable HRD parameters can include, for example, the fixed_pic_rate_general_flag[i] syntax element, the sublayer_hrd_parameters(i) syntax structure, and / or the general_vcl_hrd_params_present_flag. The fixed_pic_rate_general_flag[i] is a syntax element indicating whether the temporal distance between the HRD output times of consecutive pictures in the output order is constrained. The sublayer_hrd_parameters(i) is a syntax structure including the HRD parameters of one or more sublayers. The general_vcl_hrd_params_present_flag is a flag indicating whether the VCL HRD parameters related to the compliance points are present in the general HRD parameter syntax structure.
[0191] In step 1105, the decoder decodes an image from the sublayer / sublayer representation. In step 1107, the decoder transfers the decoded image for display as part of the decoded video sequence.
[0192] FIG. 12 is a schematic diagram of an exemplary system 1200 for coding a video sequence into a bitstream by including the inferred HRD parameters. System 1200 may be implemented by an encoder and a decoder such as codec system 200, encoder 300, decoder 400, and / or video coding device 900. Further, system 1200 may employ HRD 500 to perform compliance tests on multi-layer video sequence 600, multi-layer video sequence 700, and / or bitstream 800. Additionally, system 1200 may be employed when implementing methods 100, 1000, and / or 1100.
[0193] System 1200 includes a video encoder 1202. The video encoder 1202 includes an encoding module 1203 for encoding a plurality of sublayer representations into a bitstream. The encoding module 1203 is further for encoding the HRD parameters and the sublayer_cpb_params_present_flag into the bitstream. The video encoder 1202 further includes an inference module 1204 for inferring that the HRD parameters of all sublayer representations with a TemporalId less than the maximum TemporalId are equal to the HRD parameters of the maximum sublayer representation having the maximum TemporalId when the sublayer_cpb_params_present_flag is set to 0. The video encoder 1202 further includes an HRD module 1205 for performing a set of bitstream compliance tests on the bitstream based on the HRD parameters. The video encoder 1202 further includes a storage module 1206 for storing a bitstream for communicating to the decoder. The video encoder 1202 further includes a transmission module 1207 for transmitting the bitstream towards video decoder 1210. The video encoder 1202 may be further configured to perform any of the steps of method 1000.
[0194] System 1200 also includes a video decoder 1210. The video decoder 1210 includes a receiving module 1211 for receiving a bitstream including a plurality of sublayer representations, HRD parameters, and a sublayer_cpb_params_present_flag. The video decoder 1210 further includes an inferring module 1213 for inferring that when the sublayer_cpb_params_present_flag is set to 0, the HRD parameters of all sublayer representations whose TemporalId is less than the maximum TemporalId are equal to the HRD parameters of the maximum sublayer representation having the maximum TemporalId. The video decoder 1210 further includes a decoding module 1215 for decoding an image from the sublayer representation. The video decoder 1210 further includes a transferring module 1217 for transferring the image for display as part of the decoded video sequence. The video decoder 1210 may be further configured to perform any of the steps of method 1100.
[0195] If there are no intervening components between a first component and a second component other than a line, trace, or another medium, the first component is directly coupled to the second component. If there are intervening components between a first component and a second component other than a line, trace, or another medium, the first component is indirectly coupled to the second component. The terms "coupled" and its variations include both directly coupled and indirectly coupled. The use of the term "about" means a range including ±10% of the subsequent number unless otherwise specified.
[0196] It should also be understood that the steps of the exemplary methods described herein need not necessarily be executed in the order described, and that the order of such method steps is to be understood as merely exemplary. Similarly, in methods consistent with various embodiments of the present disclosure, additional steps may be included in such methods, and specific steps may be omitted or combined.
[0197] While several embodiments are provided in the present disclosure, it can be understood that the disclosed systems and methods can be embodied in many other specific forms without departing from the spirit or scope of the present disclosure. This example should be considered illustrative and not restrictive, and the present invention should not be limited to the details provided herein. For example, various elements or components may be combined or integrated in another system, or certain features may be omitted or not implemented.
[0198] In addition, the technologies, systems, subsystems, and methods described and illustrated in various embodiments as discrete or separate may be combined or integrated with other systems, components, technologies, or methods without departing from the scope of the present disclosure. Other examples of changes, substitutions, and alternatives will be recognizable to those skilled in the art and can be made without departing from the spirit and scope disclosed herein. Other possible items [Item 1] A method implemented by a decoder, receiving, by a receiver of the decoder, a bitstream including a plurality of sublayer representations, virtual reference decoder (HRD) parameters, and a sublayer coding picture buffer (CPB) parameter presence flag (sublayer_cpb_params_present_flag); inferring, by a processor of the decoder, that when the sublayer_cpb_params_present_flag is set to 0, the HRD parameters of all sublayer representations whose temporal identifier (TemporalId) is less than the maximum TemporalId are equal to the HRD parameters of the maximum sublayer representation having the maximum TemporalId; decoding, by the processor, an image from the plurality of sublayer representations; comprising a method. [Item 2] The method according to item 1, wherein the sublayer_cpb_params_present_flag is included in a video parameter set (VPS) in the bitstream. [Item 3] The method according to item 1 or 2, wherein the maximum TemporalId of the maximum sublayer representation is represented as the HRD maximum TemporalId (hrd_max_tid[i]), and i indicates the i-th HRD parameter syntax structure. [Item 4] The method according to any one of items 1 to 3, wherein the TemporalId smaller than the maximum TemporalId is in the range from 0 to hrd_max_tid[i] - 1. [Item 5] The method according to any one of items 1 to 4, wherein the HRD parameter includes a fixed picture rate general flag (fixed_pic_rate_general_flag[i]) indicating whether the time distance between the HRD output times of consecutive pictures in the output order is restricted. [Item 6] The method according to any one of items 1 to 4, wherein the HRD parameter includes a sublayer HRD parameter (sublayer_hrd_parameters(i)) syntax structure including HRD parameters for one or more sublayers. [Item 7] The method according to any one of items 1 to 4, wherein the HRD parameter includes a general video coding layer (VCL) HRD parameter presence flag (general_vcl_hrd_params_present_flag) indicating whether VCL HRD parameters regarding the compliance point exist in the general HRD parameter syntax structure. [Item 8] A method performed by an encoder, encoding, by a processor of the encoder, a plurality of sublayer representations into a bitstream; Encoding, by the processor, a virtual reference decoder (HRD) parameter and a sublayer coding picture buffer (CPB) parameter presence flag (sublayer_cpb_params_present_flag) into the bitstream; Inferring, by the processor, that when the sublayer_cpb_params_present_flag is set to 0, the HRD parameters of all sublayer representations where the temporal identifier (TemporalId) is less than the maximum TemporalId are equal to the HRD parameters of the maximum sublayer representation having the maximum TemporalId; Executing, by the processor, a set of bitstream compliance tests on the bitstream based on the HRD parameters; A method comprising the above. [Item 9] The method according to item 8, wherein the sublayer_cpb_params_present_flag is encoded in a video parameter set (VPS) in the bitstream. [Item 10] The method according to item 8 or 9, wherein the maximum TemporalId of the maximum sublayer representation is represented as an HRD maximum TemporalId (hrd_max_tid[i]), and i indicates the i-th HRD parameter syntax structure. [Item 11] The method according to any one of items 8 to 10, wherein the TemporalId less than the maximum TemporalId is in the range from 0 to hrd_max_tid[i] - 1. [Item 12] The method according to any one of items 8 to 11, wherein the HRD parameters include a fixed picture rate general flag (fixed_pic_rate_general_flag[i]) indicating whether the temporal distance between the HRD output times of consecutive pictures in the output order is restricted. [Item 13] The method according to any one of items 8 to 12, wherein the HRD parameter includes a syntax structure of sublayer HRD parameters (sublayer_hrd_parameters(i)) for one or more sublayers. [Item 14] The method according to any one of items 8 to 13, wherein the HRD parameter includes a general video coding layer (VCL) HRD parameter presence flag (general_vcl_hrd_params_present_flag) indicating whether the VCL HRD parameter regarding the compliance point exists in the general HRD parameter syntax structure. [Item 15] A video coding device, comprising a processor, a receiver coupled to the processor, a memory coupled to the processor, and a transmitter coupled to the processor, wherein the processor, the receiver, the memory, and the transmitter are configured to execute the method according to any one of items 1 to 14. Video encoding device. [Item 16] A non-transitory computer-readable medium comprising a computer program product for use by a video coding device, the computer program product comprising computer-executable instructions stored on the non-transitory computer-readable medium such that when executed by a processor, cause the video encoding device to execute the method according to any one of items 1 to 14. [Item 17] A decoder, receiving means for receiving a bitstream including a plurality of sublayer representations, virtual reference decoder (HRD) parameters, and a sublayer coding picture buffer (CPB) parameter presence flag (sublayer_cpb_params_present_flag); Inference means for inferring that when the sublayer_cpb_params_present_flag is set to 0, the HRD parameters of all sublayer representations with a temporal identifier (TemporalId) smaller than the maximum TemporalId are equal to the HRD parameters of the maximum sublayer representation having the maximum TemporalId, Decoding means for decoding an image from the plurality of sublayer representations, Transfer means for transferring the image for display as part of the decoded video sequence, A decoder comprising. [Item 18] The decoder according to item 17, further configured to execute the method according to any one of items 1 to 7. [Item 19] An encoder, Encoding a plurality of sublayer representations into a bitstream, Encoding a virtual reference decoder (HRD) parameter and a sublayer coding picture buffer (CPB) parameter presence flag (sublayer_cpb_params_present_flag) into the bitstream, Encoding means for, Inference means for inferring that when the sublayer_cpb_params_present_flag is set to 0, the HRD parameters of all sublayer representations with a temporal identifier (TemporalId) smaller than the maximum TemporalId are equal to the HRD parameters of the maximum sublayer representation having the maximum TemporalId, HRD means for performing a set of bitstream compliance tests on the bitstream based on the HRD parameters, Storage means for storing the bitstream for communication to a decoder, An encoder comprising. [Item 20] The encoder according to item 19, further configured to execute the method according to any one of items 8 to 14.
Claims
Claim 1 A method for decoding a bitstream, the method comprising: receiving a bitstream having a video parameter set (VPS), wherein the VPS has a vps_extension_flag, and the vps_extension_flag equal to 0 specifies that a vps_extension_data_flag syntax element does not exist in the VPS Raw Byte Sequence Payload (RBSP) syntax structure, and the vps_extension_flag equal to 1 specifies that a vps_extension_data_flag syntax element exists in the VPS RBSP syntax structure; the VPS further includes an ols_mode_idc, and the ols_mode_idc equal to 0 specifies that the total number of output layer sets (OLSs) specified by the VPS is equal to vps_max_layers_minus1 + 1, and the i-th OLS includes layers having layer indices from 0 to i, inclusive, and for each OLS, only the topmost layer within the OLS is output; the OLS is a set of layers in which one or more layers are specified as output layers; parsing the VPS to obtain the vps_extension_flag and the ols_mode_idc; decoding one or more of the layers based on the vps_extension_flag and the ols_mode_idc A method comprising the steps of: Claim 2 The method of claim 1, wherein the bitstream further has a sublayer_cpb_params_present_flag, and when the sublayer_cpb_params_present_flag is equal to 0, the virtual reference decoder (HRD) parameters for a sublayer representation having a TemporalId in the range from 0 to hrd_max_temporal_id[i] - 1, inclusive, are inferred to be the same as the HRD parameters for a sublayer representation having a TemporalId equal to hrd_max_temporal_id[i]. Claim 3 When the sublayer_cpb_params_present_flag is equal to 1, the i-th hrd_parameters() syntax structure includes, including both end values, HRD parameters for the sublayer representation having a TemporalId within the range from 0 to the hrd_max_temporal_id[i], the method according to claim 2.
4. When the sublayer_cpb_params_present_flag is equal to 0, the i-th hrd_parameters() syntax structure includes HRD parameters for the sublayer representation having a TemporalId equal to only the hrd_max_temporal_id[i], the method according to claim 2 or 3.
5. The HRD parameters include a sublayer HRD parameters (sublayer_hrd_parameters(i)) syntax structure including HRD parameters for one or more sublayers, the method according to any one of claims 2 to 4.
6. A method for encoding a bitstream, the method comprising: encoding a video parameter set (VPS) into the bitstream comprising the VPS has a vps_extension_flag, the vps_extension_flag equal to 0 specifies that the vps_extension_data_flag syntax element does not exist in the VPS Raw Byte Sequence Payload (RBSP) syntax structure, the vps_extension_flag equal to 1 specifies that the vps_extension_data_flag syntax element exists in the VPS RBSP syntax structure; the VPS further includes an ols_mode_idc, the ols_mode_idc equal to 0 specifies that the total number of output layer sets (OLSs) specified by the VPS is equal to vps_max_layers_minus1 + 1, the i-th OLS includes layers having layer indices from 0 to i, including both end values, and for each OLS, only the topmost layer within the OLS is output; the OLS is a set of layers in which one or more layers are specified as output layers, method.
7. The method according to claim 6, further comprising encoding a sublayer_cpb_params_present_flag into the bitstream, wherein when the sublayer_cpb_params_present_flag is equal to 0, the virtual reference decoder (HRD) parameters for the sublayer representation having a TemporalId in the range from 0 to hrd_max_temporal_id[i] - 1, including both end values, are inferred to be the same as the HRD parameters for the sublayer representation having a TemporalId equal to hrd_max_temporal_id[i].
8. The method according to claim 7, wherein when the sublayer_cpb_params_present_flag is equal to 1, the i-th hrd_parameters() syntax structure includes HRD parameters for the sublayer representation having a TemporalId in the range from 0 to the hrd_max_temporal_id[i], including both end values. [[ID= A computer program for causing a video coding device to execute the method according to any one of claims 1 to 5 or any one of claims 6 to 10.
14. An apparatus for storing a bitstream, comprising one or more storage media and a receiver, the receiver being configured to receive one or more bitstreams, the one or more bitstreams having a video parameter set (VPS); the VPS having a vps_extension_flag, and when the vps_extension_flag is equal to 0, the apparatus identifies that the vps_extension_data_flag syntax element does not exist in the VPS Raw Byte Sequence Payload (RBSP) syntax structure, and when the vps_extension_flag is equal to 1, the apparatus identifies that the vps_extension_data_flag syntax element exists in the VPS RBSP syntax structure; the VPS further including an ols_mode_idc, and when the ols_mode_idc is equal to 0, the apparatus identifies that the total number of output layer sets (OLSs) specified by the VPS is equal to vps_max_layers_minus1 + 1, and the i-th OLS includes layers having layer indices from 0 to i, inclusive, and for each OLS, only the topmost layer within the OLS is output; the OLS being a set of layers in which one or more layers are specified as output layers; The apparatus, wherein the one or more storage media are configured to store the one or more bitstreams.
15. The apparatus according to claim 14, further comprising a processor configured to retrieve the bitstream from the one or more storage media and transmit the bitstream to another device.
16. A method for storing a bitstream, comprising: receiving a bitstream through a receiver; Storing the bitstream in one or more storage media, the bitstream having a video parameter set (VPS); the VPS having a vps_extension_flag, the vps_extension_flag equal to 0 specifying that a vps_extension_data_flag syntax element does not exist in the VPS Raw Byte Sequence Payload (RBSP) syntax structure, and the vps_extension_flag equal to 1 specifying that a vps_extension_data_flag syntax element exists in the VPS RBSP syntax structure; the VPS further including an ols_mode_idc, the ols_mode_idc equal to 0 specifying that the total number of output layer sets (OLSs) specified by the VPS is equal to vps_max_layers_minus1 + 1, the i-th OLS including layers having layer indices from 0 to i, inclusive, and for each OLS, only the topmost layer within the OLS is output; the OLS being a set of layers in which one or more layers are specified as output layers, and A method comprising. Claim 17 A device for transmitting a bitstream, the device comprising At least one storage medium configured to store at least one bitstream, wherein the bitstream has a video parameter set (VPS); the VPS has a vps_extension_flag, and when the vps_extension_flag is set to 0 by the device, the vps_extension_flag specifies that the vps_extension_data_flag syntax element does not exist in the VPS Raw Byte Sequence Payload (RBSP) syntax structure, and when the vps_extension_flag is set to 1 by the device, the vps_extension_flag specifies that the vps_extension_data_flag syntax element exists in the VPS RBSP syntax structure; the VPS further includes an ols_mode_idc, and when the ols_mode_idc is set to 0 by the device, the ols_mode_idc specifies that the total number of output layer sets (OLSs) specified by the VPS is equal to vps_max_layers_minus1 + 1, and the i-th OLS includes layers having layer indices from 0 to i, including both end values, and for each OLS, only the topmost layer within the OLS is output; the OLS is a set of layers in which one or more layers are specified as output layers, at least one storage medium; At least one processor configured to obtain one or more bitstreams from one of the at least one storage medium, At least one processor configured to transmit the one or more bitstreams to another device A device comprising. [
18. ] A method for transmitting a bitstream, the method comprising: Obtaining one or more bitstreams from at least one storage medium, wherein the at least one storage medium stores at least one bitstream, and the bitstream has a video parameter set (VPS); the VPS has a vps_extension_flag, and the vps_extension_flag equal to 0 specifies that the vps_extension_data_flag syntax element does not exist in the VPS Raw Byte Sequence Payload (RBSP) syntax structure, and the vps_extension_flag equal to 1 specifies that the vps_extension_data_flag syntax element exists in the VPS RBSP syntax structure; the VPS further includes an ols_mode_idc, and the ols_mode_idc equal to 0 specifies that the total number of output layer sets (OLSs) specified by the VPS is equal to vps_max_layers_minus1 + 1, and the i-th OLS includes layers having layer indices from 0 to i including both end values, and for each OLS, only the topmost layer within the OLS is output; the OLS is a set of layers in which one or more layers are specified as output layers, and transmitting the one or more bitstreams to another device A method comprising. Claim 19 A method for generating a bitstream, comprising The encoder comprises a stage of generating a bitstream having a video parameter set (VPS), wherein the VPS has a vps_extension_flag, and the vps_extension_flag equal to 0 specifies that the vps_extension_data_flag syntax element does not exist in the VPS Raw Byte Sequence Payload (RBSP) syntax structure, and the vps_extension_flag equal to 1 specifies that the vps_extension_data_flag syntax element exists in the VPS RBSP syntax structure; the VPS further includes an ols_mode_idc, and the ols_mode_idc equal to 0 specifies that the total number of output layer sets (OLSs) specified by the VPS is equal to vps_max_layers_minus1 + 1, the i-th OLS includes layers having layer indices from 0 to i including both end values, and for each OLS, only the topmost layer within the OLS is output; the OLS is a set of layers in which one or more layers are specified as output layers, a generation method.