Scalable Nesting SEI Message Management

KR103023389B1Active Publication Date: 2026-09-21HUAWEI TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
KR1020227013760
Authority / Receiving Office
KR · KR
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-09-24
Filing Date
2020-09-08
Publication Date
2026-09-21
Estimated Expiration
2040-09-08

Smart Images

  • Figure 112022043811500-PCT00017_ABST
    Figure 112022043811500-PCT00017_ABST
Patent Text Reader

Abstract

A video coding mechanism is initiated. This mechanism includes encoding a bitstream containing one or more sets of output layers (OLS). A sub-bitstream extraction process is performed by a Hypothetical Reference Decoder (HRD) to extract a target OLS from the OLS. If no scalable nested SEI message within a scalable nesting SEI message references the target OLS, the SEI network abstraction layer (NAL) unit containing the scalable nesting SEI message is removed from the bitstream. A bitstream conformance test set is performed on the target OLS. The bitstream is stored for communication with the decoder.
Need to check novelty before this filing date? Find Prior Art

Description

Technology Field

[0001] Cross-reference regarding related applications

[0002] This patent application claims the benefit of U.S. provisional patent application No. 62 / 905,244, filed by Ye-Kui Wang on September 24, 2019, titled “Virtual reference decoder for multilayer video bitstream (HRD)”, which is incorporated herein by reference.

[0003] The present disclosure generally relates to video coding, and in particular to changing virtual reference decoder (HRD) parameters to support efficient encoding and / or conformance testing of multi-layer bitstreams. Background Technology

[0004] Even for relatively short videos, the amount of video data required to depict them can be substantial, which can cause difficulties when data is streamed or transmitted over communication networks with limited bandwidth. Therefore, video data is typically compressed before being transmitted over modern communication networks. Since memory resources may be limited, the size of the video can also be an issue when it is stored on a storage device. Video compression devices often use software and / or hardware at the source to code video data before transmission or storage, thereby reducing the amount of data required to represent a digital video image. The compressed data is then received at the destination by a video decompression device that decodes the video data. Due to limited network resources and the ever-increasing demand for higher video quality, improved compression and decompression technologies that improve compression ratios with little to no sacrifice in image quality are desirable.

[0005] In one embodiment, the present disclosure comprises a method implemented by a decoder, the method comprising: receiving a bitstream containing a target output layer set (OLS) by a receiver of the decoder—in which no scalable nested SEI message within a scalable nesting SEI (supplemental enhancement information) message refers to the target OLS and, when the scalable nesting SEI message is applied to a specific OLS, the SEI NAL (network abstraction layer) unit containing the scalable nesting SEI message is removed from the bitstream as part of a sub-bitstream extraction process—; and decoding a picture from the target OLS by the processor.

[0006] Video coding systems use various conformance tests to verify whether a bitstream can be decoded by a decoder. For example, conformance testing may involve testing the entire bitstream for conformance, then testing each layer of the bitstream for conformance, and finally checking the potential decodingable output for conformance. To implement conformance testing, the relevant parameters are included in the bitstream. A virtual reference decoder (HRD) can read the parameters and perform the tests. Video can contain many layers and various OLSs. Upon request, the encoder transmits one or more layers of a selected OLS. For example, the encoder can transmit the best layer(s) from an OLS that can be supported by the current network bandwidth. The issue relates to the layers included in the OLS. Each OLS contains one or more output layers configured to be displayed to the decoder. The encoder's HRD can check whether each OLS complies with a standard. A standard-compliant OLS can always be decoded and displayed by a standard-compliant decoder. The HRD process can be partially managed by SEI messages. For example, a scalable nesting SEI message may contain a scalable nested SEI message. Each scalable nested SEI message may contain data related to that layer. When performing conformance checks, HRD may perform a bitstream extraction process for the target OLS. Data unrelated to the OLS layer is typically removed before conformance testing so that each OLS can be inspected individually (e.g., before transmission). Some video coding systems do not remove scalable nesting SEI messages during the sub-bitstream extraction process because these messages are related to multiple layers.This can result in scalable nested SEI messages remaining in the bitstream after sub-bitstream extraction, even if the scalable nested SEI messages are not associated with the layer of the target OLS (the OLS being extracted). This can increase the size of the final bitstream without providing additional functionality. This example includes a mechanism to reduce the size of a multi-layer bitstream. It can be considered as removing scalable nested SEI messages from the bitstream during sub-bitstream extraction. If the scalable nested SEI messages are associated with one or more OLSs, the scalable nested SEI messages are examined. If the scalable nested SEI messages are not associated with the layer of the target OLS, the entire scalable nested SEI message can be removed from the bitstream. This results in a reduction in the size of the bitstream to be sent to the decoder. Therefore, this example increases coding efficiency and reduces processor, memory, and / or network resource usage in both the encoder and the decoder.

[0007] Optionally, in any of the preceding modes, another implementation of this mode is provided, wherein the sub-bitstream extraction process is performed by a virtual reference decoder (HRD) on the encoder.

[0008] Optionally, in any of the preceding modes, another implementation of this mode is provided, wherein the scalable nesting SEI message is applied to a specific OLS when the scalable nesting OLS flag is set to 1.

[0009] Optionally, in any of the preceding modalities, another implementation of this modal is provided, and if the scalable nesting SEI message does not contain any index(i) value in the range from 0 to 'scalable nesting OLS number minus 1 (scalable nesting num_olss_minus1)', no scalable nested SEI message refers to a target OLS, so that the i-th nesting OLS index (NestingOlsIdx[ i ]) is identical to the target OLS index (targetOlsIdx) associated with the target OLS.

[0010] Optionally, in any of the preceding modes, another implementation of this mode is provided, wherein scalable nesting num_olss_minus1 specifies the number of OLS to which the scalable nesting SEI message applies, and the value of scalable nesting num_olss_minus1 is limited to a range of 0 or greater and 1 or less than the total number of OLS minus 1 (TotalNumOlss-1).

[0011] Optionally, in any of the preceding modes, another implementation of this mode is provided, where targetOlsIdx identifies the OLS index of the target OLS.

[0012] Optionally, in any of the preceding modalities, another implementation of this modal is provided, where NestingOlsIdx[ i ] specifies the OLS index of the i-th OLS to which the scalable nesting SEI message is applied when the scalable nesting OLS flag is set to 1.

[0013] In one embodiment, the present disclosure comprises a method implemented by an encoder, the method comprising: encoding a bitstream comprising one or more sets of output layers (OLS) by a processor of the encoder; performing a sub-bitstream extraction process for extracting a target OLS from the OLS by a Hypothetical Reference Decoder (HRD) operating in the processor; removing a network abstraction layer (NAL) unit containing the scalable nesting SEI message from the bitstream when the scalable nesting SEI message is applied to a specific OLS, such that no scalable nested SEI message within the scalable nesting SEI message refers to the target OLS by the HRD operating in the processor; and performing a bitstream conformance test set for the target OLS by the HRD operating in the processor.

[0014] Video coding systems use various conformance tests to verify whether a bitstream can be decoded by a decoder. For example, conformance testing may involve testing the entire bitstream for conformance, then testing each layer of the bitstream for conformance, and finally examining the potential decodingable output for conformance. To implement conformance testing, the relevant parameters are included in the bitstream. A virtual reference decoder (HRD) can read these parameters and perform the tests. Video can contain many layers and various OLSs. Upon request, the encoder transmits one or more layers of a selected OLS. For example, the encoder may transmit the best layer(s) from an OLS that can be supported by the current network bandwidth. The issue lies in the layers included in the OLS. Each OLS contains one or more output layers configured to be presented to the decoder. The encoder's HRD can verify whether each OLS complies with the standard. Standard-compliant OLS can always be decoded and presented by a standard-compliant decoder.

[0015] The HRD process can be partially managed by SEI messages. For example, a scalable nesting SEI message may contain a scalable nested SEI message. Each scalable nested SEI message may contain data related to its corresponding layer. When performing a conformance check, HRD may perform a bitstream extraction process for the target OLS. Data unrelated to the OLS layer is typically removed before the conformance test so that each OLS can be inspected individually (e.g., before transmission). Some video coding systems do not remove scalable nesting SEI messages during the sub-bitstream extraction process because these messages are related to multiple layers. This can result in scalable nesting SEI messages remaining in the bitstream after sub-bitstream extraction, even if the scalable nesting SEI message is unrelated to the layer of the target OLS (the OLS being extracted). This can increase the size of the final bitstream without providing additional functionality. This example includes a mechanism to reduce the size of a multi-layer bitstream. During sub-bitstream extraction, scalable nested SEI messages can be considered to be removed from the bitstream. If a scalable nested SEI message is associated with one or more OLSs, the scalable nested SEI message is examined. If a scalable nested SEI message is not associated with a layer of the target OLS, the entire scalable nested SEI message can be removed from the bitstream. This results in a reduction in the size of the bitstream sent to the decoder. Therefore, this example increases coding efficiency and reduces processor, memory, and / or network resource usage in both the encoder and the decoder.

[0016] Optionally, in any of the preceding modes, another implementation of this mode is provided, wherein the scalable nesting SEI message is applied to a specific OLS when the scalable nesting OLS flag is set to 1.

[0017] Optionally, in any of the preceding modalities, another implementation of this modal is provided, and if the scalable nesting SEI message does not contain any index(i) value in the range of 0 or greater and scalable nesting num_olss_minus1 or less, no scalable nested SEI message refers to a target OLS, so that the i-th nesting OLS index (NestingOlsIdx[ i ]) is identical to the target OLS index (targetOlsIdx) associated with the target OLS.

[0018] Optionally, in any of the preceding modalities, another implementation of this modal is provided, where scalable nesting num_olss_minus1 specifies the number of OLS to which scalable nesting SEI messages are applied.

[0019] Optionally, in any of the preceding modes, another implementation of this mode is provided, wherein the value of the scalable nesting num_olss_minus1 is limited to a range of 0 or greater and TotalNumOlss-1 or less.

[0020] Optionally, in any of the preceding modes, another implementation of this mode is provided, where targetOlsIdx identifies the OLS index of the target OLS.

[0021] Optionally, in any of the preceding modalities, another implementation of this modal is provided, where NestingOlsIdx[ i ] specifies the OLS index of the i-th OLS to which the scalable nesting SEI message is applied when the scalable nesting OLS flag is set to 1.

[0022] In one embodiment, the present disclosure comprises a video coding device including a processor, a receiver coupled to the processor, a memory coupled to the processor, and a transmitter coupled to the processor, wherein the processor, the receiver, the memory, and the transmitter perform a method of one of the aforementioned embodiments.

[0023] In one embodiment, the present disclosure comprises a non-transient computer-readable medium comprising a computer program product for use by a video coding device, the computer program product comprising computer-executable instructions stored in the non-transient computer-readable medium, wherein the computer-executable instructions, when executed by a processor, cause the video coding device to perform a method of any of the aforementioned embodiments.

[0024] In one embodiment, the present disclosure comprises: receiving means for receiving a bitstream including a target output layer set (OLS)—in which no scalable nested SEI message within a scalable nesting SEI (supplemental enhancement information) message refers to the target OLS and, when the scalable nesting SEI message is applied to a specific OLS, the SEI NAL (network abstraction layer) unit including the scalable nesting SEI message is removed from the bitstream as part of a sub-bitstream extraction process—; decoding means for decoding a picture from the target OLS; and delivering means for delivering the picture for display as part of a decoded video sequence.

[0025] Optionally, in any of the preceding modes, another implementation of this mode is provided, wherein the decoder is further configured to perform the method of any of the aforementioned modes.

[0026] In one embodiment, the present disclosure comprises: encoding means for encoding a bitstream comprising one or more OLSs; HRD means: performing a sub-bitstream extraction process to extract a target OLS from an OLS; removing an SEI NAL unit containing a scalable nesting SEI message from a bitstream when the scalable nesting SEI message within the scalable nesting SEI message does not refer to the target OLS and the scalable nesting SEI message is applied to a specific OLS; and performing a bitstream conformance test set for the target OLS; and storage means for storing a bitstream for communication to a decoder.

[0027] Optionally, in any of the preceding modes, another implementation of this mode is provided, wherein the encoder is additionally configured to perform the method of any of the preceding modes.

[0028] For clarity, any one of the aforementioned embodiments may be combined with any one or more other aforementioned embodiments to create a new embodiment within the scope of the present disclosure.

[0029] These and other features will be more clearly understood from the following detailed description taken in relation to the attached drawings and claims. Brief explanation of the drawing

[0030] For a more complete understanding of the present disclosure, the following brief description taken in connection with the accompanying drawings and detailed description, in which similar reference numbers indicate similar parts, is now referred to. Figure 1 is a flowchart of an exemplary method for coding a video signal. Figure 2 is a schematic diagram of an exemplary coding and decoding (codec) system for video coding. Figure 3 is a schematic diagram illustrating an exemplary video encoder. Figure 4 is a schematic diagram illustrating an exemplary video decoder. Figure 5 is a schematic diagram illustrating an exemplary virtual reference decoder (HRD). Figure 6 is a schematic diagram illustrating an exemplary multi-layer video sequence configured for inter-layer prediction. Figure 7 is a schematic diagram illustrating an exemplary multi-layer video sequence configured for temporal scalability. Figure 8 is a schematic diagram illustrating an exemplary bitstream. Figure 9 is a schematic diagram of an exemplary video coding device. FIG. 10 is a flowchart of an exemplary method for encoding a video sequence into a bitstream by removing a scalable nested SEI message when no scalable nested SEI message within a scalable nested SEI (Supplemental enhancement information) message refers to a target OLS. FIG. 11 is a flowchart of an exemplary method for decoding a video sequence from a bitstream from which scalable nesting SEI messages have been removed, when no scalable nesting SEI message within the scalable nesting SEI message refers to a target OLS. FIG. 12 is a schematic diagram of an exemplary system for coding a video sequence in a bitstream by removing a scalable nested SEI message when no scalable nested SEI message within the scalable nested SEI message refers to a target OLS. Specific details for implementing the invention

[0031] While exemplary implementations of one or more embodiments are provided below, it should be understood from the outset that the disclosed system and / or method may be implemented using any number of technologies currently known or existing. The present disclosure shall by no means be limited to the exemplary implementations, drawings, and technologies illustrated below, including the exemplary designs and implementations illustrated and described herein, and may be modified within the scope of the appended claims, together with the full scope of equivalents.

[0032] The following terms are defined as follows unless otherwise used herein. Specifically, the following definitions are intended to provide additional clarity to the present disclosure. However, terms may be described differently in different contexts. Accordingly, the following definitions should be considered supplementary and should not be construed as limiting any other definitions of the description provided herein for such terms.

[0033] A bitstream is a sequence of bits containing compressed video data for transmission between an encoder and a decoder. An encoder is a device configured to use an encoding process to compress video data into a bitstream. A decoder is a device configured to use a decoding process to reconstruct video data from a bitstream for display. A picture is an array of luminance samples and / or chroma samples that generate a frame or its fields. The picture being encoded or decoded may be referred to as the current picture for clarity of discussion. A Network Abstraction Layer (NAL) unit is a syntax structure containing data in the form of a Raw Byte Sequence Payload (RBSP), data type indicators, and anti-emulation bytes interspersed as desired. A Video Coding Layer (VCL) NAL unit is a NAL unit coded to contain video data, such as a coded slice of a picture. A Non-VCL NAL unit is a NAL unit containing non-video data, such as syntax and / or parameters, that supports video data decoding, the performance of conformance checks, or other operations. An access unit (AU) is a set of NAL units associated with one another according to specified classification rules and belonging to a single specific output time. A decoding unit (DU) is an AU or a subset of the AU and associated non-VCL NAL units. For example, an AU includes VCL NAL units and any non-VCL NAL units associated with the VCL NAL units within the AU. Additionally, a DU includes a set of VCL NAL units from the AU or a subset thereof, as well as any non-VCL NAL units associated with the VCL NAL units of the DU. A layer is a set of VCL NAL units that share specified characteristics (e.g., common resolution, frame rate, image size, etc.) and associated non-VCL NAL units. The decoding order is the order in which syntax elements are processed in the decoding process.A Video Parameter Set (VPS) is a data unit containing parameters for the entire video.

[0034] A temporally scalable bitstream is a bitstream coded across multiple layers providing various temporal resolutions / frame rates (e.g., each layer is coded to support different frame rates). A sublayer is a temporally scalable layer of the temporally scalable bitstream containing VCL NAL units having specific temporal identifier values ​​and associated non-VCL NAL units. For example, a temporal sublayer is a layer containing video data associated with a specified frame rate. A sublayer representation is a subset of bitstreams containing NAL units of a specific sublayer and a lower sublayer. Thus, one or more temporal sublayers can be combined to achieve a sublayer representation that can be decoded to produce a video sequence having a specified frame rate. An Output Layer Set (OLS) is a set of layers in which one or more layers are designated as output layers. An output layer is a layer designated for output (e.g., display). An OLS index is an index that uniquely identifies the corresponding OLS. The 0th OLS is an OLS that includes only the lowest layer (the layer with the lowest layer identifier) ​​and therefore only the output layer. The temporal identifier (ID) is a data element that indicates that the data corresponds to a time position in a video sequence. The sub-bitstream extraction process is a process that removes NAL units from bitstreams that do not belong to the target set determined by the target OLS index and the target highest time ID. The sub-bitstream extraction process generates an output sub-bitstream containing NAL units from bitstreams that are part of the target set.

[0035] HRD is a decoder model that operates in an encoder to examine the variability of a bitstream generated by an encoding process to verify compliance with specified constraints. Bitstream compliance testing is a test that determines whether the encoded bitstream complies with standards such as VVC (Versatile Video Coding). HRD parameters are syntax elements that initialize and / or define the operating conditions of HRD. HRD parameters can be accommodated within an HRD parameter syntax structure. A syntax structure is a data object configured to contain multiple different parameters. A syntax element is a data object that accommodates one or more parameters of the same type. Therefore, a syntax structure can accommodate multiple syntax elements. Sequence-level HRD parameters are HRD parameters applied to the entire coded video sequence. The maximum HRD time ID (hrd_max_tid[i]) specifies the time ID of the top-level sublayer representation in which the HRD parameter is included in the i-th set of OLS HRD parameters. The general_hrd_parameters syntax structure is a syntax structure containing sequence-level HRD parameters. An Operation Point (OP) is a temporal subset of OLS identified by an OLS index and the highest time ID. The targetOp (targetOp) is the OP selected for conformance testing in HRD. The target OLS is the OLS selected for extraction from the bitstream. The decoding_unit_hrd_params_present_flag is a flag indicating whether the corresponding HRD parameter operates at the DU level or the AU level. The coded picture buffer (CPB) is a first-in, first-out buffer of HRD containing pictures coded in decoding order for use during bitstream conformance verification.The decoded picture buffer (DPB) is a buffer for holding the decoded picture for reference, output reordering, and / or output delay.

[0036] Supplemental Enhancement Information (SEI) messages are syntax structures with explicit semantics that convey information not required in the decoding process to determine sample values ​​of a decoded picture. Scalable-nesting SEI messages are messages containing multiple SEI messages corresponding to one or more OLS or one or more layers. Non-scalable-nested SEI messages are messages containing a single SEI message that is not nested. Buffering Period (BP) SEI messages are SEI messages containing HRD parameters to initialize HRD for managing the CPB. Picture Timing (PT) SEI messages are SEI messages containing HRD parameters to manage delivery information for AUs in the CPB and / or DPB. Decoding Unit Information (DUI) SEI messages are SEI messages containing HRD parameters to manage delivery information for DUs in the CPB and / or DPB.

[0037] The CPB elimination delay is the period during which the corresponding current AU may remain in the CPB before being eliminated and output to the DPB. The initial CPB elimination delay is the default CPB elimination delay for each picture AU and / or DU in the bitstream, OLS, and / or layer. The CPB elimination offset is the position within the CPB used to determine the boundary of the corresponding AU in the CPB. The initial CPB elimination offset is the default CPB elimination offset associated with each picture, AU, and / or DU in the bitstream, OLS, and / or layer. The DPB (decoded picture buffer) output delay information is the period during which the corresponding AU may remain in the DPB before being output. The CPB elimination delay information is information related to eliminating the corresponding DU from the CPB. The delivery schedule specifies the timing for delivering video data to and from memory locations such as the CPB and / or DPB. The VPS layer ID (vps_layer_id) is a syntax element representing the layer ID of the i-th layer indicated by the VPS. The number of output layer sets minus 1 (num_output_layer_sets_minus1) is a syntax element specifying the total number of OLS specified by the VPS. The HRD-coded picture buffer count (hrd_cpb_cnt_minus1) is a syntax element specifying the number of alternative CPB delivery schedules. The sublayer CPB parameter present flag (sublayer_cpb_params_present_flag) is a syntax element specifying whether the OLS HRD parameter set contains HRD parameters for the specified sublayer representation. The schedule index (ScIdx) is an index identifying the delivery schedule. The BP CPB count minus 1 (bp_cpb_cnt_minus1) is a syntax element specifying the number of initial CPB elimination delay and offset pairs, and thus the number of delivery schedules available for the temporal sublayer.The NAL unit header layer identifier (nuh_layer_id) is a syntax element that specifies the identifier of the layer containing the NAL unit. The fixed_pic_rate_general_flag syntax element specifies whether the temporal distance between HRD output times of consecutive pictures in the output order is limited. The sublayer_hrd_parameters syntax structure is a syntax structure containing HRD parameters for the corresponding sublayer. The general_vcl_hrd_params_present_flag is a flag that specifies whether VCL HRD parameters exist in the general HRD parameter syntax structure. The bp_max_sublayers_minus1 syntax element specifies the maximum number of temporal sublayers where the CPB removal delay and CPB removal offset are indicated in the BP SEI message. The syntax element vps_max_sublayers_minus1 specifies the maximum number of temporal sublayers that can exist in the layer specified by the VPS. The Scalable Nesting OLS flag specifies whether a scalable nested SEI message applies to a specific OLS or a specific layer. The Scalable Nesting Count of OLS minus1 is a syntax element specifying the number of OLSs to which a scalable nested SEI message applies. Nesting OLS Index is a syntax element specifying the OLS index of the OLS to which the scalable nested SEI message applies. Target OLS Index is a variable that identifies the OLS index of the target OLS to be decoded.TotalNumOlss-1 is a syntax element that specifies the total number of OLS specified in the VPS.

[0038] The following abbreviations are used in this specification: Access Unit (AU), Coding Tree Block (CTB), Coding Tree Unit (CTU), Coding Unit (CU), Coded Layer Video Sequence (CLVS), Start of Coded Layer Video Sequence (CLVSS), Coded Layer Video Sequence (CVS), Start of Coded Layer Video Sequence (CVSS), Joint Video Expert Team (JVET), Virtual Reference Decoder (HRD), Motion Constrained Tile Set (MCTS), Max Transport Unit (MTU), Network Abstraction Layer (NAL), Output Layer Set (OLS), Picture Order Count (POC), Random Access Point (RAP), Raw Byte Sequence Payload (RBSP), Sequence Parameter Set (SPS), Video Parameter Set (VPS), Versatile Video Coding (VVC).

[0039] Many video compression techniques can be used to reduce the size of video files while minimizing data loss. For example, video compression techniques may include performing spatial (e.g., intra-picture) prediction and / or temporal (e.g., inter-picture) prediction to reduce or eliminate data redundancy in video sequences. For block-based video coding, a video slice (e.g., a video picture or a part of a video picture) may be divided into video blocks, also called tree blocks, coding tree blocks (CTB), coding tree units (CTU), coding units (CU), and / or coding nodes. Video blocks of an intra-coded (I) slice of a picture are coded using spatial prediction for reference samples in neighboring blocks of the same picture. Video blocks of an inter-coded unidirectional prediction (P) or bidirectional prediction (B) slice of a picture may be coded using spatial prediction for reference samples in neighboring blocks of the same picture or temporal prediction for reference samples in other reference pictures. A picture can be referred to as a frame and / or image, and a reference picture can be referred to as a reference frame and / or reference image. Spatial or temporal prediction generates prediction blocks representing image blocks. Residual data represents the pixel differences between the original image blocks and the prediction blocks. Therefore, inter-coded blocks are encoded according to motion vectors pointing to blocks of reference samples forming the prediction blocks, and residual data representing the differences between the coded blocks and the prediction blocks. Intra-coded blocks are encoded according to the intra-coding mode and residual data. For further compression, residual data can be transformed from the pixel domain to the transform domain. This generates residual transform coefficients that can be quantized. The quantized transform coefficients can initially be arranged in a two-dimensional array. The quantized transform coefficients can be scanned to generate a one-dimensional vector of transform coefficients. Entropy coding can be applied to achieve greater compression.These video compression technologies are discussed in more detail below.

[0040] In order to ensure that encoded video can be accurately decoded, the video is encoded and decoded according to the corresponding video coding standards. Video coding standards include Advanced Video Coding (AVC), also known as ITU-T H.261, ISO / IEC Motion Picture Experts Group (MPEG)-1 Part 2, ITU-T H.262 or ISO / IEC MPEG-2 Part 2, ITU-T H.263, ISO / IEC MPEG-4 Part 2, ITU-T H.264 or ISO / IEC MPEG-4 Part 10, and High Efficiency Video Coding (HEVC), also known as ITU-T H.265 or MPEG-H Part 2. AVC includes extensions such as SVC (Scalable Video Coding), MVC (Multiview Video Coding), MVC+D (Multiview Video Coding plus Depth), and 3D AVC (3D-AVC). HEVC includes extensions such as Scalable HEVC (SHVC), Multiview HEVC (MV-HEVC), and 3D HEVC (3D-HEVC). The Joint Video Expert Team (JVET) of ITU-T and ISO / IEC began developing a video coding standard called VVC (Versatile Video Coding). VVC is included in the Draft Specification (WD) containing JVET-O2001-v14.

[0041] Video coding systems use various conformance tests to verify whether a bitstream can be decoded by a decoder. For example, conformance testing may involve testing the entire bitstream for conformance, then testing each layer of the bitstream for conformance, and finally checking the potential decodingable output for conformance. To implement conformance testing, the corresponding parameters are included in the bitstream. A virtual reference decoder (HRD) can read the parameters and perform the tests. Video can contain many layers and various sets of output layers (OLS). Upon request, the encoder transmits one or more layers of the selected OLS. For example, the encoder can transmit the best layer(s) from the OLS that can be supported by the current network bandwidth. The first problem with this approach is that a significant number of layers are tested but are not actually transmitted to the decoder. However, parameters supporting these tests may still be included in the bitstream, which unnecessarily increases the bitstream size.

[0042] In the first example, a mechanism for applying bitstream conformance tests to each OLS only is disclosed herein. In this way, when the corresponding OLS is tested, the entire bitstream, each layer, and the decodingable output are tested collectively. Consequently, the number of conformance tests is reduced, thereby reducing the use of processor and memory resources in the encoder. Furthermore, reducing the number of conformance tests can reduce the number of relevant parameters included in the bitstream. This reduces the bitstream size and thus reduces the use of processor, memory, and / or network resources in both the encoder and the decoder.

[0043] The second issue is that the HRD parameter signaling process used for HRD conformance testing in some video coding systems can become complex in a multi-layer context. For example, a set of HRD parameters can be signaled for each layer of each OLS. These HRD parameters may be signaled at different locations in the bitstream depending on the intended range of the parameters. This results in a situation where complexity increases as more layers and / or OLS are added. Furthermore, HRD parameters for different layers and / or OLS may contain redundant information.

[0044] In the second example, a mechanism for signaling a global set of HRD parameters for an OLS and its corresponding layers is disclosed herein. For example, all sequence-level HRD parameters applicable to all OLS and all layers included in the OLS are signaled to a Video Parameter Setter (VPS). Since the VPS is signaled once in the bitstream, the sequence-level HRD parameters are signaled once. Additionally, the sequence-level HRD parameters may be restricted to be the same for all OLS. In this way, redundant signaling is reduced, thereby increasing coding efficiency. Furthermore, this approach simplifies the HRD process. As a result, the use of processor, memory, and / or network signaling resources in both the encoder and the decoder is reduced.

[0045] A third problem can arise when a video coding system performs conformance checks on a bitstream. Video can be coded with multiple layers and / or sublayers and then organized into an OLS. Each layer and / or sublayer of each OLS performs conformance checks according to a delivery schedule. Each delivery schedule is associated with different coded picture buffer (CPB) sizes and CPB delays to account for different transmit bandwidths and system capabilities. In some video coding systems, each sublayer can define an arbitrary number of delivery schedules. This can generate a large amount of signaling to support conformance checks, which reduces coding efficiency for the bitstream.

[0046] In the third example, a mechanism for increasing coding efficiency for video containing multiple layers is disclosed herein. Specifically, all layers and / or sublayers are limited to containing the same number of CPB forward schedules. For example, the encoder may determine the maximum number of CPB forward schedules used for any one layer and set the number of CPB forward schedules for all layers to the maximum. Then, the number of forward schedules may be signaled once, for example, as part of the HRD parameters in the VPS. This avoids the need to signal multiple schedules for each layer / sublayer. In some examples, all layers / sublayers of the OLS may share the same forward schedule index. This change reduces the amount of data used for signal data related to conformance checks. This reduces the bitstream size and thus reduces processor, memory, and / or network resource usage in both the encoder and the decoder.

[0047] The fourth problem can occur when video is coded with multiple layers and / or sublayers and then configured into OLS. The OLS may include a zeroth OLS containing only the output layer. Supplemental enhancement information (SEI) messages may be included in the bitstream to inform the HRD of layer / OLS-specific parameters used to test the layers of the bitstream for compliance with the standard. Specifically, scalable nesting SEI messages are used when OLS is included in the bitstream. Scalable nesting SEI messages contain a group of nested SEI messages applied to one or more OLSs and / or one or more OLS layers. Each nested SEI message may include an indicator representing an association with the corresponding OLS and / or layer. Nested SEI messages are configured for use with multiple layers and may contain irrelevant information when applied to a zeroth OLS containing a single layer.

[0048] In the fourth example, a mechanism for increasing coding efficiency for a video containing the 0th OLS is disclosed herein. A non-scalable-nested SEI message is used for the 0th OLS. The non-scalable-nested SEI message is restricted to apply only to the 0th OLS and thus only to the output layer contained within the 0th OLS. In this way, irrelevant information, such as nesting relationships and layer indications, can be omitted from the SEI message. The non-scalable-nested SEI message may be used as a buffering cycle (BP) SEI message, a picture timing (PT) SEI message, a decoding unit (DU) SEI message, or a combination thereof. This change reduces the amount of data used to signal conformance-related information for the 0th OLS. This reduces the bitstream size and thus reduces processor, memory, and / or network resource usage in both the encoder and the decoder.

[0049] The fifth problem can also occur when the video is separated into multiple layers and / or sublayers. The encoder can encode these layers into a bitstream. Additionally, the encoder can use HRD to perform conformity tests to inspect the bitstream for compliance with standards. To support these conformity tests, the encoder can be configured to include layer-specific HRD parameters in the bitstream. In some video coding systems, layer-specific HRD parameters can be encoded for each layer. In some cases, since the layer-specific HRD parameters are the same for each layer, redundant information is generated, unnecessarily increasing the size of the video encoding.

[0050] In the fifth example, a mechanism for reducing HRD parameter redundancy for video using multiple layers is disclosed herein. The encoder may encode HRD parameters for the top layer. The encoder may also encode a sublayer CPB parameter present flag (sublayer_cpb_params_present_flag). The sublayer_cpb_params_present_flag may be set to 0 to indicate that all lower layers must use the same HRD parameters as the top layer. In this context, the top layer has the largest layer ID (ID), and the lower layers are all layers with a layer ID smaller than the layer ID of the top layer. In this way, HRD parameters for the lower layers may be omitted from the bitstream. This reduces the bitstream size and thus reduces processor, memory, and / or network resource usage in both the encoder and the decoder.

[0051] The sixth issue concerns the use of Sequence Parameter Sets (SPS) to include syntax elements associated with each video sequence. Video coding systems may code video at layers and / or sublayers. Video sequences may behave differently at different layers and / or sublayers. Therefore, different layers may reference different SPSs. BP SEI messages may indicate the layers / sublayers to be checked for compliance with standards. Some video coding systems may indicate that BP SEI messages apply to the layers / sublayers specified in the SPS. This can cause problems when different layers reference different SPSs, as these SPSs may contain contradictory information, potentially leading to unexpected errors.

[0052] In the sixth example, a mechanism for handling errors related to conformity checks when multiple layers are used in a video sequence is disclosed herein. Specifically, the BP SEI message is modified to indicate that any number of layers / sublayers described in the VPS can be checked for conformity. For example, the BP SEI message may include a BP max_sublayers_minus1 (bp_max_sublayers_minus1) syntax element indicating the number of layers / sublayers associated with the data in the BP SEI message. Meanwhile, in the VPS, the VPS max_sublayers_minus1 (vps_max_sublayers_minus1) syntax element indicates the number of sublayers in the entire video. The bp_max_sublayers_minus1 syntax element can be set to any value from 0 to the value of the vps_max_sublayers_minus1 syntax element. In this way, conformity can be checked on any number of layers / sublayers in the video while preventing layer-based sequence problems related to SPS mismatches. Accordingly, the present disclosure avoids layer-based coding errors and thus enhances the capabilities of the encoder and / or decoder. Additionally, since the present example supports layer-based coding, coding efficiency can be improved. As such, the present example supports reduced processor, memory, and / or network resource usage in the encoder and / or decoder.

[0053] The seventh issue concerns the layers included in an OLS. Each OLS includes at least one output layer configured to be displayed to a decoder. The encoder's HRD can verify that each OLS complies with the standard. A matching OLS can always be decoded and displayed by a matching decoder. The HRD process can be partially managed by SEI messages. For example, a scalable nesting SEI message may contain a scalable nested SEI message. Each scalable nested SEI message may contain data associated with that layer. The HRD can perform a bitstream extraction process for the target OLS when conducting a conformance check. Data unrelated to the OLS's layers is typically removed before the conformance test so that each OLS can be verified individually (e.g., before transmission). Some video coding systems do not remove scalable nesting SEI messages during the sub-bitstream extraction process because these messages are associated with multiple layers. This can result in scalable nesting SEI messages remaining in the bitstream after sub-bitstream extraction, even if the scalable nesting SEI messages are not related to the layer of the target OLS (the OLS being extracted). This can increase the size of the final bitstream without providing additional functionality.

[0054] In the seventh example, a mechanism for reducing the size of a multilayer bitstream is disclosed herein. This may be considered as removing scalable nested SEI messages from the bitstream during sub-bitstream extraction. If the scalable nested SEI message is associated with at least one OLS, the scalable nested SEI message is identified from the scalable nested SEI message. If the scalable nested SEI message is not associated with a layer of the target OLS, the entire scalable nested SEI message may be removed from the bitstream. This results in a reduction in the size of the bitstream to be transmitted to the decoder. Thus, this example increases coding efficiency and reduces the use of processor, memory, and / or network resources in both the encoder and the decoder.

[0055] FIG. 1 is a flowchart of an exemplary operation method (100) for encoding a video signal. Specifically, the video signal is encoded in an encoder. The encoding process compresses the video signal using various mechanisms to reduce the video file size. The smaller the file size, the more the compressed video file can be transmitted to the user, and the associated bandwidth overhead is reduced. Then, a decoder decodes the compressed video file to reconstruct the original video signal to be displayed to the end user. The decoding process generally mirrors the encoding process so that the decoder can consistently reconstruct the video signal.

[0056] In step 101, the video signal is input into the encoder. For example, the video signal may be an uncompressed video file stored in memory. As another example, the video file may be captured by a video capture device, such as a video camera, and encoded to support live streaming of the video. The video file may contain both audio and video components. The video component contains a series of image frames that provide a visual impression of motion when viewed as a sequence. The frames contain pixels represented by brightness, called the luminance component (or luminance sample), and color, called the chroma component (or color sample). In some examples, the frames may also include depth values ​​to support a three-dimensional view.

[0057] In step 103, the video is divided into blocks. Dividing involves subdividing the pixels of each frame into square and / or rectangular blocks for compression. For example, in HEVC (also known as High Efficiency Video Coding, H.265, and MPEG-H Part 2), a frame can first be divided into Coding Tree Units (CTUs), which are blocks of a predefined size (e.g., 64 pixels x 64 pixels). CTUs contain both luminance and chroma samples. Using a coding tree, CTUs can be divided into blocks, and then the blocks can be recursively subdivided until a configuration supporting additional encoding is achieved. For example, the luminance component of a frame can be subdivided until individual blocks contain relatively uniform brightness values. Additionally, the chroma component of a frame can be subdivided until individual blocks contain relatively uniform color values. Thus, the division mechanism varies depending on the content of the video frame.

[0058] In step 105, various compression mechanisms are used to compress the image blocks segmented in step 103. For example, inter-prediction and / or intra-prediction may be used. Inter-prediction is designed to take advantage of the fact that objects in a common scene tend to appear in consecutive frames. Therefore, blocks representing objects in a reference frame do not need to be described repeatedly in adjacent frames. Specifically, objects such as tables may remain in a constant position across multiple frames. Thus, a table is described once, and adjacent frames may refer back to the reference frame. Objects can be matched across multiple frames using a pattern matching mechanism. Additionally, moving objects may be represented across multiple frames due to, for example, the movement of the object or the movement of the camera. As a specific example, a video may show a car moving across the screen over multiple frames. Motion vectors can be used to describe such movement. A motion vector is a two-dimensional vector that provides an offset from the coordinates of an object in a frame to the coordinates of an object in a reference frame. In this way, inter-prediction can encode an image block of the current frame into a set of motion vectors representing an offset from the corresponding block of the reference frame.

[0059] Intra prediction encodes blocks of a common frame. Intra prediction utilizes the fact that lumina and chroma components tend to cluster within a frame. For example, green patches within a part of a tree tend to be placed adjacent to similar green patches. Intra prediction uses multiple directional prediction modes (e.g., 33 in HEVC), planar modes, and direct current (DC) modes. Directional modes indicate that the current block is similar to / identical to samples from neighboring blocks in that direction. Planar modes indicate that a series of blocks along a row / column (e.g., a plane) can be interpolated based on adjacent blocks at the row edges. In practice, planar modes represent smooth transitions of light / color across rows / columns using a relatively constant gradient when changing values. DC modes are used for boundary smoothing and indicate blocks similar to / identical to the average value associated with samples from all adjacent blocks related to the angular direction of the directional prediction modes. Therefore, intra prediction blocks can represent image blocks using various relational prediction mode values ​​rather than actual values. Additionally, intra prediction blocks can represent image blocks using motion vector values ​​rather than actual values. In any case, the prediction block may not accurately represent the image block depending on the circumstances. All differences are stored in the residual block. Transformations can be applied to the residual block to further compress the file.

[0060] In step 107, various filtering techniques can be applied. In HEVC, filters are applied according to the in-loop filtering method. The block-based prediction discussed above can result in the generation of block images in the decoder. Additionally, the block-based prediction method can encode blocks and then reconstruct the encoded blocks to be used later as reference blocks. The in-loop filtering method repeatedly applies noise suppression filters, deblocking filters, adaptive loop filters, and sample adaptive offset (SAO) filters to blocks / frames. These filters mitigate these blocking artifacts so that the encoded file can be accurately reconstructed. Furthermore, these filters mitigate artifacts in the reconstructed reference blocks, making it less likely that artifacts will generate additional artifacts in subsequent blocks encoded based on the reconstructed reference blocks.

[0061] Once the video signal is split, compressed, and filtered, the resulting data is encoded into a bitstream in step 109. The bitstream contains the data discussed above, as well as the desired signaling data, to support the appropriate reconstruction of the video signal in the decoder. For example, this data may include partition data, prediction data, residual blocks, and various flags that provide coding instructions to the decoder. The bitstream may be stored in memory for transmission to the decoder upon request. The bitstream may also be broadcast and / or multicast to multiple decoders. Bitstream generation is an iterative process. Therefore, steps 101, 103, 105, 107, and 109 may occur sequentially and / or simultaneously across many frames and blocks. The order shown in FIG. 1 is presented for clarity and ease of discussion and is not intended to restrict the video coding process to a specific order.

[0062] The decoder receives the bitstream and starts the decoding process in step 111. Specifically, the decoder uses an entropy decoding method to convert the bitstream into corresponding syntax and video data. In step 111, the decoder uses the syntax data of the bitstream to determine the partition for the frame. The partition must match the block partitioning result of step 103. Now, the entropy encoding / decoding used in step 111 is described. The encoder makes many choices during the compression process, such as selecting a block partitioning method from among several possible choices based on the spatial location of values ​​in the input image. A large number of bins can be used as signals for the correct choice. As used herein, a bin is a binary value treated as a variable (e.g., a bit value that can change depending on the context). Using entropy coding allows the encoder to discard options that are not clearly feasible in a specific case, leaving a set of acceptable options. Then, a code word is assigned to each acceptable option. The length of the code word is based on the number of acceptable options (e.g., one bin for two options, two bins for three to four options, etc.). The encoder then encodes the code word for the selected option. This method reduces the size of the code word, as the code word is as large as desired to uniquely represent a choice within a small subset of acceptable options, in contrast to uniquely representing a choice within a potentially large set of all possible options. The decoder then decodes the choice by determining the set of acceptable options in a manner similar to the encoder. By determining the set of acceptable options, the decoder can read the code word and determine the choice made by the encoder.

[0063] In step 113, the decoder performs block decoding. Specifically, the decoder generates residual blocks using an inverse transform. The decoder then reconstructs image blocks according to the segmentation using the residual blocks and the corresponding prediction blocks. The prediction blocks may include both intra-prediction blocks and inter-prediction blocks, as generated by the encoder in step 105. The reconstructed image blocks are then placed in frames of the reconstructed video signal according to the segmentation data determined in step 111. The syntax for step 113 can also be signaled in the bitstream via entropy coding as discussed above.

[0064] In step 115, filtering is performed on frames of the video signal reconstructed in the encoder in a manner similar to step 107. For example, a noise suppression filter, a deblocking filter, an adaptive loop filter, and an SAO filter can be applied to the frames to remove blocking artifacts. Once the frames are filtered, the video signal can be output to a display in step 117 for viewing by an end user.

[0065] FIG. 2 is a schematic diagram of an exemplary coding and decoding (codec) system (200) for video coding. Specifically, the codec system (200) provides functions that support the implementation of the operation method (100). The codec system (200) is generalized to describe components used in both the encoder and the decoder. The codec system (200) receives a video signal and divides it to generate a divided video signal (201), as discussed in relation to steps 101 and 103 of the operation method (100). The codec system (200) compresses the divided video signal (201) into a coded bitstream when acting as an encoder, as discussed in relation to steps 105, 107, and 109 of the method (100). When acting as a decoder, the codec system (200) generates an output video signal from a bitstream as discussed in relation to steps 111, 113, 115, and 117 in the method of operation (100). The codec system (200) includes a general coder control component (211), a transform scaling and quantization component (213), an intra-picture estimation component (215), an intra-picture prediction component (217), a motion compensation component (219), a motion estimation component (221), a scaling and inverse transform component (229), a filter control analysis component (227), an in-loop filter component (225), a decoded picture buffer component (223), and a header formatting and CABAC (context adaptive binary arithmetic coding) component (231). These components are combined as illustrated. In FIG. 2, the black line indicates the movement of data to be encoded / decoded, and the dotted line indicates the movement of control data that controls the operation of other components. All components of the codec system (200) may exist in the encoder. The decoder may include a subset of the components of the codec system (200).For example, the decoder may include an intra-picture prediction component (217), a motion compensation component (219), a scaling and inverse transformation component (229), an in-loop filter component (225), and a decoded picture buffer component (223). These components will now be described.

[0066] The segmented video signal (201) is a captured video sequence that has been segmented into pixel blocks by a coding tree. The coding tree uses various segmentation modes to subdivide the pixel blocks into smaller pixel blocks. Then, these blocks can be subdivided into smaller blocks. Blocks can be called nodes on the coding tree. Larger parent nodes are segmented into smaller child nodes. The number of times a node is segmented is called the depth of the node / coding tree. The segmented blocks may be included in Coding Units (CUs) in some cases. For example, a CU may be a sub-part of a CTU containing a Luma block, a Red Difference Chroma (Cr) block, and a Blue Difference Chroma (Cb) block, along with corresponding syntax instructions for the CU. Depending on the segmentation mode used, the segmentation modes may include binary trees (BT), triad trees (TT), and quad trees (QT) used to segment nodes of various shapes into 2, 3, or 4 child nodes, respectively. The segmented video signal (201) is passed to a general coder control component (211), a transform scaling and quantization component (213), an intra-picture estimation component (215), a filter control analysis component (227), and a motion estimation component (221) for compression.

[0067] The general coder control component (211) is configured to make decisions regarding coding images of a video sequence into a bitstream according to application limitations. For example, the general coder control component (211) manages the optimization of bitrate / bitstream size versus reconstruction quality. These decisions may be based on storage space / bandwidth availability and image resolution requests. The general coder control component (211) also manages buffer utilization in relation to the transmission rate to mitigate buffer underrun and overrun issues. To manage these issues, the general coder control component (211) manages splitting, prediction, and filtering by other components. For example, the general coder control component (211) can dynamically increase compression complexity to increase resolution and bandwidth usage, or decrease compression complexity to decrease resolution and bandwidth usage. Thus, the general coder control component (211) controls other components of the codec system (200) to balance video signal reconstruction quality and bitrate issues. The general coder control component (211) generates control data that controls the operation of other components. The control data is also passed to the header formatting and CABAC component (231) to be encoded in the bitstream as signal parameters for decoding in the decoder.

[0068] The segmented video signal (201) is also transmitted to a motion estimation component (221) and a motion compensation component (219) for inter-prediction. A frame or slice of the segmented video signal (201) may be divided into multiple video blocks. The motion estimation component (221) and the motion compensation component (219) perform inter-prediction coding of the received video blocks for one or more blocks of one or more reference frames to provide time prediction. The codec system (200) may perform multiple coding passes, for example, to select an appropriate coding mode for each block of video data.

[0069] The motion estimation component (221) and the motion compensation component (219) may be highly integrated, but are exemplified separately for conceptual purposes. Motion estimation performed by the motion estimation component (221) is a process of generating motion vectors that estimate motion for a video block. For example, motion vectors may represent the displacement of a coded object relative to a prediction block. A prediction block is a block found to closely match the block to be coded in terms of pixel differences. A prediction block may also be called a reference block. These pixel differences may be determined by the sum of absolute differences (SAD), the sum of squared differences (SSD), or other difference metrics. HEVC uses several coded objects, including CTUs, coding tree blocks (CTBs), and CUs. For example, a CTU may be divided into CTBs and then divided into CBs to be included in a CU. A CU may be encoded as a prediction unit (PU) containing prediction data and / or a transformation unit (TU) containing transformed residual data for the CU. The motion estimation component (221) generates motion vectors, PUs, and TUs by using rate distortion analysis as part of the rate distortion optimization process. For example, the motion estimation component (221) can determine multiple reference blocks, multiple motion vectors, etc. for the current block / frame, and can select the reference blocks, motion vectors, etc. that have the best rate distortion characteristics. The best rate distortion characteristics balance the quality of video reconstruction (e.g., amount of data loss due to compression) and coding efficiency (e.g., size of the final encoding).

[0070] In some examples, the codec system (200) may calculate values ​​for integer-level pixel locations of reference pictures stored in the decoded picture buffer component (223). For example, the video codec system (200) may interpolate values ​​for 1 / 4 pixel locations, 1 / 8 pixel locations, or other fractional pixel locations of the reference picture. Thus, the motion estimation component (221) may perform motion search for whole pixel locations and partial pixel locations and output a motion vector with partial pixel precision. The motion estimation component (221) calculates a motion vector for the PU of the video block in the intercoded slice by comparing the position of the PU with the position of the prediction block of the reference picture. The motion estimation component (221) outputs the calculated motion vector to the header formatting as motion data and outputs the CABAC component (231) for encoding motion to the motion compensation component (219).

[0071] Motion compensation performed by the motion compensation component (219) may involve fetching or generating a prediction block based on a motion vector determined by the motion estimation component (221). Again, the motion estimation component (221) and the motion compensation component (219) may be functionally integrated in some examples. Upon receiving a motion vector for the PU of the current video block, the motion compensation component (219) may find the prediction block indicated by the motion vector. Then, a residual video block is formed by subtracting the pixel value of the prediction block from the pixel value of the current video block being coded to form a pixel difference value. Generally, the motion estimation component (221) performs motion estimation for the luminance component, and the motion compensation component (219) uses a motion vector calculated based on the luminance component for both the chroma component and the luminance component. The prediction block and the residual block are passed to the transform scaling and quantization component (213).

[0072] The segmented video signal (201) is also transmitted to the intra-picture estimation component (215) and the intra-picture prediction component (217). As with the motion estimation component (221) and the motion compensation component (219), the intra-picture estimation component (215) and the intra-picture prediction component (217) can be highly integrated, but are exemplified separately for conceptual purposes. As described above, the intra-picture estimation component (215) and the intra-picture prediction component (217) intra-predict the current block for the current frame as an alternative to the inter-prediction performed by the motion estimation component (221) and the motion compensation component (219) between frames. In particular, the intra-picture estimation component (215) determines the intra-prediction mode to use to encode the current block. In some examples, the intra-picture estimation component (215) selects an appropriate intra-prediction mode to encode the current block from a number of tested intra-prediction modes. The selected intra prediction mode is passed to the header formatting and CABAC component (231) for encoding.

[0073] For example, the intra-picture estimation component (215) calculates rate distortion values ​​using rate distortion analysis for various tested intra-prediction modes and selects the intra-prediction mode having the best rate distortion characteristics among the tested modes. Ratio-distortion analysis generally determines the amount of distortion (or error) between the encoded block and the original unencoded block that was encoded to generate the encoded block, as well as the bitrate (e.g., number of bits) used to generate the encoded block. The intra-picture estimation component (215) calculates the ratio from the distortion and rate for various encoded blocks to determine the intra-prediction mode that represents the best rate distortion value for the block. Additionally, the intra-picture estimation component (215) may be configured to code the depth blocks of the depth map using a depth modeling mode (DMM) based on rate distortion optimization (RDO).

[0074] The intra-picture prediction component (217) can generate a residual block from a prediction block based on a selected intra-prediction mode determined by the intra-picture estimation component (215) when implemented on an encoder, or read a residual block from a bitstream when implemented on a decoder. The residual block contains the difference in values ​​between the prediction block and the original block and is represented as a matrix. The residual block is then passed to the transform scaling and quantization component (213). The intra-picture estimation component (215) and the intra-picture prediction component (217) can operate on both the luminance and chroma components.

[0075] The transform scaling and quantization component (213) is configured to further compress the residual block. The transform scaling and quantization component (213) applies a transform, such as the Discrete Cosine Transform (DCT), Discrete Sine Transform (DST), or a conceptually similar transform, to the residual block to generate a video block containing residual transform coefficient values. Wavelet transform, integer transform, subband transform, or other types of transform may also be used. The transform can transform residual information from a pixel value domain to a transform domain, such as a frequency domain. The transform scaling and quantization component (213) is also configured to scale the transformed residual information, for example, based on frequency. This scaling involves applying a scale factor to the residual information so that different frequency information is quantized at different granularities, which may affect the final visual quality of the reconstructed video. The transform scaling and quantization component (213) is also configured to quantize the transform coefficients to further reduce the bit rate. The quantization process may reduce the bit depth associated with some or all of the coefficients. The degree of quantization can be modified by adjusting the quantization parameters. In some examples, the transform scaling and quantization component (213) may then perform a scan of a matrix containing the quantized transform coefficients. The quantized transform coefficients are passed to the header formatting and CABAC component (231) and encoded in the bitstream.

[0076] The scaling and inverse transform component (229) applies the inverse operation of the transform scaling and quantization component (213) to support motion estimation. The scaling and inverse transform component (229) applies inverse scaling, transform, and / or quantization to reconstruct a residual block in the pixel domain for later use as a reference block, which may be a prediction block for another current block, for example. The motion estimation component (221) and / or motion compensation component (219) may also compute a reference block by adding the residual block back to the corresponding prediction block for use in motion estimation of a later block / frame. Filters are applied to the reconstructed reference block to mitigate artifacts generated during scaling, quantization, and transform. Otherwise, when a subsequent block is predicted, these artifacts may cause inaccurate predictions and generate additional artifacts.

[0077] The filter control analysis component (227) and the in-loop filter component (225) apply filters to residual blocks and / or reconstructed image blocks. For example, the transformed residual blocks from the scaling and inverse transformation component (229) can be combined with the corresponding prediction blocks from the intra-picture prediction component (217) and / or motion compensation component (219) to reconstruct the original image blocks. Then, filters can be applied to the reconstructed image blocks. In some examples, filters can be applied to residual blocks instead. Like other components in FIG. 2, the filter control analysis component (227) and the in-loop filter component (225) can be implemented together in a highly integrated manner, but are shown separately for conceptual purposes. Filters applied to the reconstructed reference blocks are applied to specific spatial regions and include a number of parameters to adjust how these filters are applied. The filter control analysis component (227) analyzes the reconstructed reference blocks to determine where these filters should be applied and sets the corresponding parameters. This data is passed to the header formatting and CABAC component (231) as filter control data for encoding. The in-loop filter component (225) applies these filters based on the filter control data. The filters may include deblocking filters, noise suppression filters, SAO filters, and adaptive loop filters. These filters may be applied in the spatial / pixel domain (e.g., in a reconstructed pixel block) or the frequency domain, depending on the example.

[0078] When operating as an encoder, the filtered reconstructed image blocks, residual blocks, and / or prediction blocks are stored in a decoded picture buffer component (223) for later use in motion estimation as previously discussed. When operating as a decoder, the decoded picture buffer component (223) stores the reconstructed and filtered blocks and forwards them toward a display as part of the output video signal. The decoded picture buffer component (223) may be any memory device capable of storing the prediction blocks, residual blocks, and / or reconstructed image blocks.

[0079] The header formatting and CABAC component (231) receives data from various components of the codec system (200) and encodes this data into a coded bitstream for transmission to the decoder. Specifically, the header formatting and CABAC component (231) generates various headers for encoding control data, such as general control data and filter control data. Additionally, prediction data including intra prediction and motion data, and residual data in the form of quantized transform coefficient data, are all encoded in the bitstream. The final bitstream contains all information desired by the decoder to reconstruct the original segmented video signal (201). This information may also include an intra prediction mode index table (also called a codeword mapping table), definitions of encoding contexts for various blocks, indications of the most likely intra prediction mode, indications of partition information, etc. This data may be encoded using entropy coding. For example, information may be encoded using context-adaptive variable length coding (CAVLC), CABAC, syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or other entropy coding techniques. Following entropy coding, the encoded bitstream may be transmitted to another device (e.g., a video decoder) or stored for later transmission or retrieval.

[0080] FIG. 3 is a block diagram illustrating an exemplary video encoder (300). The video encoder (300) may be used to implement the encoding function of the codec system (200) and / or to implement steps 101, 103, 105, 107 and / or 109 of the operation method (100). The encoder (300) divides the input video signal to generate a divided video signal (301) substantially similar to the divided video signal (201). Then, the divided video signal (301) is compressed by a component of the encoder (300) and encoded into a bitstream.

[0081] Specifically, the segmented video signal (301) is forwarded to the intra-picture prediction component (317) for intra-prediction. The intra-picture prediction component (317) may be substantially similar to the intra-picture estimation component (215) and the intra-picture prediction component (217). The segmented video signal (301) is also forwarded to the motion compensation component (321) for inter-prediction based on the reference block of the decoded picture buffer component (323). The motion compensation component (321) may be substantially similar to the motion estimation component (221) and the motion compensation component (219). The prediction block and residual block from the intra-picture prediction component (317) and the motion compensation component (321) are forwarded to the transformation and quantization component (313) for transformation and quantization of the residual block. The transformation and quantization component (313) may be substantially similar to the transformation scaling and quantization component (213). The transformed and quantized residual blocks and the corresponding prediction blocks (along with the associated control data) are passed to the entropy coding component (331) for coding into a bitstream. The entropy coding component (331) may be substantially similar to the header formatting and CABAC component (231).

[0082] The transformed and quantized residual block and / or corresponding prediction block is also transferred from the transformed and quantized component (313) to the inverse transformed and quantized component (329) for reconstruction into a reference block to be used by the motion compensation component (321). The inverse transformed and quantized component (329) may be substantially similar to the scaling and inverse transformed component (229). The in-loop filter of the in-loop filter component (325) is applied to the residual block and / or the reconstructed reference block, depending on the example. The in-loop filter component (325) may be substantially similar to the filter control analysis component (227) and the in-loop filter component (225). The in-loop filter component (325) may include a number of filters as discussed in relation to the in-loop filter component (225). The filtered block is then stored in the decoded picture buffer component (323) to be used as a reference block by the motion compensation component (321). The decoded picture buffer component (323) may be substantially similar to the decoded picture buffer component (223).

[0083] FIG. 4 is a block diagram illustrating an exemplary video decoder (400). The video decoder (400) may be used to implement the decoding function of the codec system (200) and / or to implement steps 111, 113, 115 and / or 117 of the operation method (100). The decoder (400) receives a bitstream, for example, from the encoder (300) and generates an output video signal reconstructed based on the bitstream for display to an end user.

[0084] The bitstream is received by the entropy decoding component (433). The entropy decoding component (433) is configured to implement an entropy decoding method such as CAVLC, CABAC, SBAC, PIPE coding, or other entropy coding techniques. For example, the entropy decoding component (433) may use header information to provide context for interpreting additional data encoded as codewords in the bitstream. The decoded information contains information desired for decoding the video signal, such as general control data, filter control data, partition information, motion data, prediction data, and quantized transform coefficients from the residual block. The quantized transform coefficients are passed to the inverse transform and quantization component (429) for reconstruction into the residual block. The inverse transform and quantization component (429) may be similar to the inverse transform and quantization component (329).

[0085] The reconstructed residual block and / or prediction block is passed to the intra-picture prediction component (417) for reconstruction into an image block based on the intra-prediction operation. The intra-picture prediction component (417) may be similar to the intra-picture estimation component (215) and the intra-picture prediction component (217). Specifically, the intra-picture prediction component (417) uses a prediction mode to find a reference block in the frame and applies the residual block to the result to reconstruct the intra-predicted image block. The reconstructed intra-predicted image block and / or residual block and the corresponding intra-prediction data are passed to the decoded picture buffer component (423) via the in-loop filter component (425), which may be substantially similar to the decoded picture buffer component (223) and the in-loop filter component (225), respectively. The in-loop filter component (425) filters the reconstructed image blocks, residual blocks, and / or prediction blocks, and this information is stored in the decoded picture buffer component (423). The reconstructed image blocks from the decoded picture buffer component (423) are passed to the motion compensation component (421) for inter-prediction. The motion compensation component (421) may be substantially similar to the motion estimation component (221) and / or the motion compensation component (219). Specifically, the motion compensation component (421) generates prediction blocks using motion vectors from reference blocks and reconstructs image blocks by applying residual blocks to the result. The resulting reconstructed blocks may also be passed to the decoded picture buffer component (423) via the in-loop filter component (425). The decoded picture buffer component (423) continues to store additional reconstructed image blocks that can be reconstructed into frames via partition information. These frames may also be placed in a sequence. The sequence is output toward the display as a reconstructed output video signal.

[0086] FIG. 5 is a schematic diagram illustrating an exemplary HRD (500). The HRD (500) may be used in an encoder, such as a codec system (200) and / or an encoder (300). The HRD (500) may verify the bitstream generated in step 109 of method 100 before the bitstream is forwarded to a decoder, such as a decoder (400). In some examples, the bitstream may be forwarded continuously through the HRD (500) as the bitstream is encoded. If a portion of the bitstream fails to comply with an associated limit, the HRD (500) may indicate such failure to the encoder, causing the encoder to re-encode the corresponding portion of the bitstream by a different mechanism.

[0087] The HRD (500) includes a virtual stream scheduler (HSS, 541). The HSS (541) is a component configured to perform a virtual delivery mechanism. The virtual delivery mechanism is used to check the suitability of the bitstream or decoder with respect to the timing and data flow of the bitstream (551) input to the HRD (500). For example, the HSS (541) can receive the bitstream (551) output from the encoder and manage the suitability test process for the bitstream (551). In a specific example, the HSS (541) can control the rate at which the coded picture travels through the HRD (500) and verify that the bitstream (551) does not contain unsuitable data.

[0088] The HSS (541) can transmit the bitstream (551) to the CPB (543) at a predefined rate. The HRD (500) can manage the data of the DU (553). The DU (553) is an AU or a subset of the AU and the associated non-video coding layer (VCL) network abstraction layer (NAL) units. Specifically, the AU contains one or more pictures related to the output time. For example, the AU may contain a single picture in a single-layer bitstream and pictures for each layer in a multi-layer bitstream. Each picture in the AU may be divided into slices, each contained in a corresponding VCL NAL unit. Thus, the DU (553) may contain one or more pictures, one or more slices of pictures, or a combination thereof. Additionally, parameters used to decode the AU, pictures, and / or slices may be contained in the non-VCL NAL units. Thus, the DU (553) contains non-VCL NAL units containing data necessary to support the decoding of VCL NAL units in the DU (553). The CPB (543) is a first-in, first-out buffer of the HRD (500). The CPB (543) contains the DU (553) containing video data in decoding order. The CPB (543) stores video data for use during bitstream conformance verification.

[0089] CPB (543) passes DU (553) to the decoding process component (545). The decoding process component (545) is a component that follows the VVC standard. For example, the decoding process component (545) can emulate the decoder (400) used by the end user. The decoding process component (545) decodes the DU (553) at a rate achievable by an exemplary end-user decoder. If the decoding process component (545) cannot decode the DU (553) fast enough to prevent an overflow of CPB (543), the bitstream (551) must be re-encoded without following the standard.

[0090] The decoding process component (545) decodes the DU (553) to generate the decoded DU (555). The decoded DU (555) contains the decoded picture. The decoded DU (555) is passed to the DPB (547). The DPB (547) may be substantially similar to the decoded picture buffer component (223, 323 and / or 423). To support inter-prediction, the picture marked for use as a reference picture (556) obtained from the decoded DU (555) is returned to the decoding process component (545) to support further decoding. The DPB (547) outputs the decoded video sequence as a series of pictures (557). The pictures (557) are reconstructed pictures that generally mirror the picture encoded into the bitstream (551) by the encoder.

[0091] The picture (557) is passed to the output cropping component (549). The output cropping component (549) is configured to apply a suitability cropping window to the picture (557). As a result, a cropped picture (559) is output. The output cropped picture (559) is a completely reconstructed picture. Thus, the output cropped picture (559) mimics what the end user will see when decoding the bitstream (551). In this way, the encoder can review the output cropped picture (559) to check if the encoding is satisfactory.

[0092] The HRD (500) is initialized based on the HRD parameters of the bitstream (551). For example, the HRD (500) may read HRD parameters from VPS, SPS, and / or SEI messages. Then, the HRD (500) may perform a conformity test operation on the bitstream (551) based on the information of these HRD parameters. As a specific example, the HRD (500) may determine one or more CPB delivery schedules (561) from the HRD parameters. The delivery schedule specifies the timing for delivering video data to and / or from memory locations such as CPB and / or DPB. Thus, the CPB delivery schedule (561) specifies the timing for delivering AU, DU (553), and / or pictures to and from CPB (543). For example, the CPB forwarding schedule (561) may describe the bit rate and buffer size for the CPB (543), where these bit rate and buffer size correspond to a specific class of decoder and / or network conditions. Thus, the CPB forwarding schedule (561) may indicate how long data can remain in the CPB (543) before evicting. Failure to maintain the CPB forwarding schedule (561) in the HRD (500) during the conformity test is an indication that the decoder corresponding to the CPB forwarding schedule (561) will not be able to decode the corresponding bitstream. The HRD (500) may use a DPB forwarding schedule for the DPB (547) similar to the CPB forwarding schedule (561).

[0093] Video may be coded into different layers and / or OLS for use in decoders with varying levels of hardware capabilities and various network conditions. The CPB forwarding schedule (561) is selected to reflect these issues. Thus, since the upper layer sub-bitstream is specified for optimal hardware and network conditions, the upper layer may receive one or more CPB forwarding schedules (561) that use a large amount of memory at the CPB (543) and a short delay for the transmission of the DU (553) toward the DPB (547). Likewise, the lower layer sub-bitstream is specified for limited decoder hardware capabilities and / or poor network conditions. Thus, the lower layer may receive one or more CPB forwarding schedules (561) that use a small amount of memory at the CPB (543) and a longer delay for the transmission of the DU (553) toward the DPB (547). Next, the OLS, layer, sublayer, or combination thereof may be tested according to the corresponding forwarding schedule (561) to ensure that the resulting sub-bitstream can be correctly decoded under expected conditions for the sub-bitstream. Each CPB forwarding schedule (561) is associated with a schedule index (ScIdx, 563). ScIdx (563) is an index that identifies the forwarding schedule. Thus, the HRD parameter of the bitstream (551) may not only represent the CPB forwarding schedule (561) by ScIdx (563), but may also contain sufficient data for the HRD (500) to determine the CPB forwarding schedule (561) and correlate the CPB forwarding schedule (561) to the corresponding OLS, layer, and / or sublayer.

[0094] FIG. 6 is a schematic diagram illustrating an exemplary multilayer video sequence configured for interlayer prediction (621). The multilayer video sequence (600) may be encoded by an encoder such as a codec system (200) and / or an encoder (300) and decoded by a decoder such as a codec system (200) and / or a decoder (400), for example, according to method (100). Additionally, the multilayer video sequence (600) may be checked for standard compliance by an HRD such as an HRD (500). The multilayer video sequence (600) is included to illustrate an exemplary application to the layers of a coded video sequence. The multiple-layer video sequence (600) is any video sequence using multiple layers, such as layer N (631) and layer N+1 (632).

[0095] In the example, a multilayer video sequence (600) may use interlayer prediction (621). Interlayer prediction (621) is applied between pictures (611, 612, 613, 614) and pictures (615, 616, 617, 618) of different layers. In the illustrated example, pictures 611, 612, 613, and 614 are part of layer N+1 (632), and pictures 615, 616, 617, and 618 are part of layer N (631). A layer is a group of pictures all associated with similar values ​​of characteristics such as similar size, quality, resolution, signal-to-noise ratio, capability, etc. A layer can be formally defined as a set of VCL NAL units and associated non-VCL NAL units. A VCL NAL unit is a NAL unit coded to contain video data, such as a coded slice of a picture. A non-VCL NAL unit is an NAL unit that contains non-video data, such as syntax and / or parameters, that supports video data decoding, the performance of conformance checks, or other operations.

[0096] In the illustrated example, layer N+1 (632) is associated with a larger image size than layer N (631). Thus, the pictures (611, 612, 613, 614) of layer N+1 (632) have larger picture sizes (e.g., greater height and width, and therefore more samples) than the pictures (615, 616, 617, 618) of layer N (631) in this example. However, these pictures may be separated between layer N+1 (632) and layer N (631) by other characteristics. Although only two layers, layer N+1 (632) and layer N (631), are illustrated, a set of pictures may be separated into any number of layers based on associated characteristics. Layer N+1 (632) and layer N (631) may also be indicated by layer IDs. A layer ID is a data item associated with a picture and indicating that the picture is part of the layer in which it is displayed. Thus, each picture (611-618) may be associated with a corresponding layer ID to indicate layer N+1 (632) or layer N (631) containing the corresponding picture. For example, a layer ID may include a NAL unit header layer identifier (nuh_layer_id), which is a syntax element specifying the identifier of the layer containing the NAL unit (e.g., including slices and / or parameters of the picture in the layer). A layer associated with a lower quality / bitstream size, such as layer N (631), is typically assigned a lower layer ID and is referred to as the lower layer. Additionally, a layer associated with a higher quality / bitstream size, such as layer N+1 (632), is typically assigned a higher layer ID and is referred to as the upper layer.

[0097] Pictures (611-618) of other layers (631-632) are configured to be displayed in alternatives. As such, pictures of different layers (631-632) may share a temporal ID (622) as long as the pictures are included in the same AU. The temporal ID (622) is a data element indicating that the data corresponds to a time position in a video sequence. An AU is a set of NAL units associated with one another according to specified classification rules and belonging to a specific output time. For example, an AU may include one or more pictures, such as picture (611) and picture (615), in different layers if these pictures are associated with the same temporal ID (622). As a specific example, the decoder may decode and display picture (615) at the current display time if a smaller picture is required, or the decoder may decode and display picture (611) at the current display time if a larger picture is required. Thus, the pictures (611-614) of the upper layer N+1 (632) contain substantially the same image data as the corresponding pictures (615-618) of the sub-layer N (631) (despite differences in picture size). Specifically, picture (611) contains substantially the same image data as picture (615), and picture (612) contains substantially the same image data as picture (616).

[0098] Pictures (611-618) can be coded by referencing other pictures (611-618) of the same layer N (631) or N+1 (632). When a picture is coded by referencing another picture of the same layer, an inter prediction (623) occurs. The inter prediction (623) is illustrated by a solid arrow. For example, a picture (613) can be coded by using the inter prediction (623) with one or two of the pictures (611, 612 and / or 614) of layer N+1 (632) as references, where one picture is referenced for a unidirectional inter prediction or two pictures are referenced for a bidirectional inter prediction. Additionally, a picture (617) may be coded by using inter-prediction (623) with one or two of the pictures (615, 616 and / or 618) of layer N (531) as references, where one picture is referenced for unidirectional inter-prediction or two pictures are referenced for bidirectional inter-prediction. When performing inter-prediction (623), if a picture is used as a reference to another picture of the same layer, the picture may be called a reference picture. For example, picture (612) may be a reference picture used to code picture (613) according to inter-prediction (623). Inter-prediction (623) may also be called intra-layer prediction in a multi-layer context. As such, inter-prediction (623) is a mechanism for coding samples of the current picture by referencing samples directed from a reference picture that is different from the current picture, where the reference picture and the current picture are in the same layer.

[0099] Pictures (611-618) can also be coded by referencing other pictures (611-618) of different layers. This process is known as inter-layer prediction (621) and is indicated by a dashed arrow. Inter-layer prediction (621) is a mechanism for coding samples of the current picture by referencing samples indicated in a reference picture, where the current picture and the reference picture are on different layers and therefore have different layer IDs. For example, a picture of a lower layer N (631) can be used as a reference picture to code a corresponding picture of an upper layer N+1 (632). As a specific example, picture (611) can be coded by referencing picture (615) according to inter-layer prediction (621). In this case, picture (615) is used as an inter-layer reference picture. An inter-layer reference picture is a reference picture used in inter-layer prediction (621). In most cases, the interlayer prediction (621) is restricted to using only the interlayer reference picture(s) that are in the same AU as the current picture (611), such as picture (615), and are in the lower layer. If multiple layers (e.g., two or more) are available, the interlayer prediction (621) can encode / decode the current picture based on multiple interlayer reference picture(s) at a lower level than the current picture.

[0100] A video encoder may use multiple layer video sequences (600) to encode pictures (611-618) through many different combinations and / or permutations of inter-prediction (623) and inter-layer prediction (621). For example, picture (615) may be coded according to intra-prediction. Then, pictures (616-618) may be coded according to inter-prediction (623) by using picture (615) as a reference picture. Additionally, picture (611) may be coded according to inter-layer prediction (621) by using picture (615) as an inter-layer reference picture. Then, pictures (612-614) may be coded according to inter-prediction (623) by using picture (611) as a reference picture. In this way, the reference picture can serve as both a single-layer reference picture and an inter-layer reference picture for different coding mechanisms. By coding the upper layer N+1 (632) based on the lower layer N (631) picture, the upper layer N+1 (632) can avoid using intra prediction, which has a much lower coding efficiency than inter prediction (623) and inter layer prediction (621). As such, the poor coding efficiency of intra prediction can be limited to the minimum / lowest quality picture, and thus limited to coding the smallest amount of video data. The picture used as a reference picture and / or inter layer reference picture can be indicated in the entry of the reference picture list(s) included in the reference picture list structure.

[0101] To perform this operation, one or more OLSs (625, 626) may include layers such as layer N (631) and layer N+1 (632). Specifically, pictures (611-618) are encoded as layers (631-632) in the bitstream (600), and then each layer (631-632) of the image is assigned to one or more OLSs (625 and 626). Then, the OLSs (625 and / or 626) may be selected, and the corresponding layers (631 and / or 632) may be transmitted to the decoder depending on the capabilities of the decoder and / or network conditions. OLS (625) is a set of layers in which one or more layers are designated as output layers. An output layer is a layer designated for output (e.g., display). For example, layer N (631) may be included alone to support inter-layer prediction (621) and may never be output. In this case, layer N+1 (632) is decoded and output based on layer N (631). In this case, the OLS (625) includes layer N+1 (632) as an output layer. When the OLS includes only an output layer, the OLS is called the 0th OLS (626). The 0th OLS (626) is an OLS that includes only the lowest layer (the layer with the lowest layer identifier) ​​and thus includes only an output layer. In other cases, the OLS (625) may include many layers in different combinations. For example, the output layer of the OLS (625) may be coded according to inter-layer prediction (621) based on one, two, or many lower layers. Additionally, the OLS (625) may include one or more output layers. Thus, the OLS (625) may include one or more output layers and any support layer necessary to reconstruct the output layer.Although only two OLSs (625 and 626) are shown, a multilayer video sequence (600) can be coded by using many different OLSs (625 and / or 626), each using a different combination of layers. Each OLS (625, 626) is associated with an OLS index (629), which is an index that uniquely identifies the corresponding OLS (625, 626).

[0102] Inspecting a multilayer video sequence (600) for standard compliance in HRD (500) can be complicated depending on the number of layers (631-632) and OLS (625, 626). HRD (500) can separate the multilayer video sequence (600) into a sequence of action points (627) for testing. OLS (625 and / or 626) are identified by OLS indices (629). An action point (627) is a temporal subset of OLS (625 / 626). An action point (627) can be identified by both the OLS index (629) of the corresponding OLS (625 / 626) and the highest temporal ID (622). As a specific example, the first operation point (627) may include all pictures of the first OLS (625) from temporal ID 0 to temporal ID 200, and the second operation point (627) may include all pictures of the first OLS (625) from temporal ID 201 to temporal ID 400. In this case, the first operation point (627) is described by the OLS index (629) of the first OLS (625) and the temporal ID 200. Additionally, the second operation point (627) is described by the OLS index (629) of the first OLS (625) and the temporal ID 400. The operation point (627) selected for testing at a specified moment is called the targetOp. Thus, the targetOp is the operation point (627) selected for conformity testing in the HRD (500).

[0103] FIG. 7 is a schematic diagram illustrating an exemplary multilayer video sequence (700) configured for temporal scalability. The multilayer video sequence (700) may be encoded by an encoder such as a codec system (200) and / or an encoder (300) and decoded by a decoder such as a codec system (200) and / or a decoder (400), for example, according to method (100). Additionally, the multilayer video sequence (700) may be checked for standard compliance by an HRD such as an HRD (500). The multilayer video sequence (700) is included to illustrate other exemplary applications for layers of a coded video sequence. For example, the multilayer video sequence (700) may be used as a separate embodiment or may be combined with the techniques described for the multilayer video sequence (600).

[0104] A multilayer video sequence (700) includes sublayers (710, 720, 730). A sublayer is a temporally scalable layer of a temporally scalable bitstream that includes VCL NAL units (e.g., pictures) having specific temporal identifier values ​​and associated non-VCL NAL units (e.g., support parameters). For example, a layer such as layer N (631) and / or layer N+1 (632) may be further divided into sublayers (710, 720, 730) to support temporal scalability. Sublayer (710) may be referred to as a base layer, and sublayers (720 and 730) may be referred to as enhancement layers. As illustrated, sublayer (710) includes a picture (711) of a first frame rate, such as 30 frames per second. Since the sublayer (710) contains a base / lowest frame rate, the sublayer (710) is a base layer. The sublayer (720) contains a picture (721) that is time-offset from the picture (711) of the sublayer (710). As a result, the sublayer (710) and the sublayer (720) can be combined, which results in a collectively higher frame rate than the frame rate of the sublayer (710) alone. For example, the sublayers (710 and 720) can have a combined frame rate of 60 frames per second. Thus, the sublayer (720) enhances the frame rate of the sublayer (710). Additionally, the sublayer (730) contains a picture (731) that is time-offset from the pictures (721 and 711) of the sublayers (720 and 710). In this way, the sublayer (730) can be combined with the sublayers (720 and 710) to further enhance the sublayer (710). For example, the sublayers (710, 720, 730) can have a composite frame rate of 90 frames per second.

[0105] A sublayer representation (740) can be dynamically generated by combining sublayers (710, 720 and / or 730). A sublayer representation (740) is a subset of bitstreams containing NAL units of specific sublayers and lower sublayers. In the illustrated example, the sublayer representation (740) includes a picture (741) which is a combined picture (711, 721, 731) of sublayers (710, 720, 730). Thus, a multilayer video sequence (700) can be temporally scaled to a desired frame rate by selecting a sublayer representation (740) containing a desired set of sublayers (710, 720 and / or 730). A sublayer representation (740) can be generated by using an OLS that includes sublayers (710, 720 and / or 730) as layers. In this case, the sublayer representation (740) is selected as the output layer. Thus, temporal scalability is one of several mechanisms that can be achieved using a multilayer mechanism.

[0106] FIG. 8 is a schematic diagram illustrating an exemplary bitstream (800). For example, the bitstream (800) may be generated by a codec system (200) and / or an encoder (300) for decoding by a codec system (200) and / or a decoder (400) according to method (100). Additionally, the bitstream (800) may include a multi-layer video sequence (600 and / or 700). Additionally, the bitstream (800) may include various parameters for controlling the operation of an HRD, such as an HRD (500). Based on these parameters, the HRD may inspect the bitstream (800) for compliance with a standard before transmission to a decoder for decoding.

[0107] The bitstream (800) includes a VPS (811), one or more SPS (813), a plurality of picture parameter sets (PPS, 815), a plurality of slice headers (817), image data (820), and SEI messages (819). The VPS (811) includes data related to the entire bitstream (800). For example, the VPS (811) may include data related to the OLS, layers, and / or sublayers used in the bitstream (800). The SPS (813) includes sequence data common to all pictures of the coded video sequences included in the bitstream (800). For example, each layer may include one or more coded video sequences, and each coded video sequence may refer to the SPS (813) for corresponding parameters. The parameters of the SPS (813) may include picture scaling, bit depth, coding tool parameters, bit rate limits, etc. While each sequence refers to the SPS (813), a single SPS (813) may contain data for multiple sequences in some examples. The PPS (815) contains parameters applied to the entire picture. Thus, each picture in a video sequence may refer to the PPS (815). Although each picture refers to the PPS (815), in some examples, a single PPS (815) may contain data for multiple pictures. For example, multiple similar pictures may be coded according to similar parameters. In such cases, a single PPS (815) may contain data for these similar pictures. The PPS (815) may represent coding tools available for slices, quantization parameters, offsets, etc., of the corresponding picture.

[0108] A slice header (817) contains parameters specific to each slice of a picture. Thus, there may be one slice header (817) per slice of a video sequence. A slice header (817) may include slice type information, POC, reference picture list, prediction weights, tile entry points, deblocking parameters, etc. In some examples, the bitstream (800) may also include a picture header, which is a syntax structure containing parameters that apply to all slices in a single picture. For this reason, picture headers and slice headers (817) may be used interchangeably in some contexts. For example, specific parameters may be moved between the slice header (817) and the picture header depending on whether these parameters are common to all slices of the picture.

[0109] Image data (820) includes video data encoded according to inter-prediction and / or intra-prediction, and corresponding transformed and quantized residual data. For example, image data (820) may include AU (821), DU (822), and / or picture (823). AU (821) is a set of NAL units associated with one another according to a specified classification rule and belonging to one specific output time. DU (822) is a subset of AU or AU and associated non-VCL NAL units. Picture (823) is an array of luminance samples and / or chroma samples that generate a frame or its field. In plain language, AU (821) includes various video data that may be displayed at a specified moment in a video sequence, as well as supporting syntax data. Thus, AU (821) may include a single picture (823) of a single-layer bitstream or multiple pictures from multiple layers associated with the same moment in a multi-layer bitstream. Meanwhile, the picture (823) is a coded image that can be output for display or used to support the coding of other picture(s) (823) for output. The DU (822) may contain one or more pictures (823) and any supporting syntax data required for decoding. For example, the DU (822) and the AU (821) may be used interchangeably in a simple bitstream (e.g., where the AU contains a single picture). However, in a more complex multi-layer bitstream, the DU (822) may contain only a portion of the video data from the AU (821). For example, the AU (821) may contain the picture (823) in multiple layers and / or sublayers where a portion of the picture (823) is associated with a different OLS. In such cases, the DU (822) may contain only the picture(s) (823) from the specified OLS and / or specified layer / sublayer.

[0110] A picture (823) includes one or more slices (825). A slice (825) may be defined as an integer of a complete tile of the picture (823) or an integer of a consecutive complete coding tree unit (CTU) row (e.g., within a tile), wherein the tile or CTU row is exclusively contained within a single NAL unit (829). Thus, a slice (825) is also contained within a single NAL unit (829). A slice (825) is further divided into CTUs and / or coding tree blocks (CTBs). A CTU is a group of samples of a predefined size that can be divided into coding trees. A CTB is a subset of CTUs and contains the luminance or chroma components of the CTUs. The CTU / CTB is further divided into coding blocks based on the coding trees. Then, the coding blocks can be encoded / decoded according to a prediction mechanism.

[0111] The bitstream (800) is a sequence of NAL units (829). A NAL unit (829) is a container for video data and / or supporting syntax. A NAL unit (829) may be a VCL NAL unit or a non-VCL NAL unit. A VCL NAL unit is a NAL unit (829) coded to contain video data, such as a coded slice (825) and an associated slice header (817). A non-VCL NAL unit is a NAL unit (829) containing non-video data, such as syntax and / or parameters, that support decoding of video data, performing conformance checks, or other operations. For example, a non-VCL NAL unit may contain a VPS (811), SPS (813), PPS (815), SEI message (819), or other supporting syntax.

[0112] A SEI message (819) is a syntax structure having a specified meaning that conveys information not required in the decoding process to determine the sample values ​​of the decoded picture. For example, the SEI message (819) may contain data to support the HRD process or other supporting data that is not directly related to decoding the bitstream (800) in the decoder. The SEI message (819) used for the bitstream (800) may include a scalable nesting SEI message and / or a scalable nested SEI message. A scalable nesting SEI message is a message containing multiple SEI messages corresponding to one or more OLS or one or more layers. A scalable nested SEI message is an SEI message contained within a scalable nesting SEI message. A non-scalable-nested SEI message is a message that is not nested and contains a single SEI message. The SEI message (819) may include a BP SEI message containing HRD parameters for initializing the HRD to manage the CPB. The SEI message (819) may also include a PT SEI message containing HRD parameters for managing delivery information for the AU (821) in the CPB and / or DPB. The SEI message (819) may also include a DUI SEI message containing HRD parameters for managing delivery information for the DU (822) in the CPB and / or DPB. The parameters, PT SEI messages, and / or DUI SEI messages included within the BP SEI message may be employed to determine the CPB delivery schedule in the HRD. For example, a scalable nesting SEI message may include a set of BP SEI messages, a set of PT SEI messages, or a set of DUI SEI messages.

[0113] As mentioned above, a video stream may contain many OLSs and many layers, such as OLS (625), Layer N (631), Layer N+1 (632), Sublayer (710), Sublayer (720) and / or Sublayer (730). Additionally, some layers may be included in multiple OLSs. As such, a multilayer video sequence, such as a multilayer video sequence (600 and / or 700), can be quite complex. For example, a scalable nested SEI message may contain a scalable nested SEI message applied to many OLSs, layers and / or sublayers. When the decoder requests a target OLS, the encoder / HRD may perform a subbitstream extraction process for the bitstream (800). The encoder / HRD extracts image data (820) that supports parameters and forms the target OLS from, for example, the VPS (811), SPS (813), PPS (815), slice header (817), SEI message (819), etc. This extraction is performed by the NAL unit (829). The result is a subbitstream of the bitstream (800) containing sufficient information to decode the target OLS. The unextracted information is not part of the requested target OLS and is not transmitted to the decoder. The scalable nesting SEI message may contain data related to many layers, sublayers, and / or OLS. Therefore, some video coding systems may include all scalable nesting SEI messages in the extracted subbitstream to ensure that the target OLS is checked for standard compliance and can be properly decoded by the decoder. The HRD then performs a series of conformity tests before transmitting to the decoder. This approach can be inefficient because it is overly comprehensive. For example, some of the scalable nesting SEI messages may be completely irrelevant to the target OLS.These irrelevant scalable nesting SEI messages remain in the extracted subbitstream and are sent to the decoder. This can increase the size of the final subbitstream without providing additional functionality.

[0114] The present disclosure includes a mechanism for reducing the size of a subbitstream extracted from a bitstream (800) encoding a multilayer bitstream, such as a multilayer bitstream (600 and / or 700). During subbitstream extraction, scalable nested SEI messages may be considered to be removed from the bitstream. For example, scalable nested SEI messages may be associated with an OLS and / or layer. If a scalable nested SEI message is associated with a specific OLS, scalable nested SEI messages (e.g., SEI message (819)) within the scalable nested SEI message are examined. If a scalable nested SEI message is not associated with a layer of the target OLS, the entire scalable nested SEI message may be removed from the subbitstream. This results in a reduction in the size of the subbitstream to be sent to the decoder. Thus, the present example increases coding efficiency and reduces the use of processor, memory, and / or network resources in both the encoder and the decoder.

[0115] It should be noted that a subbitstream is a type of bitstream (800). Specifically, a subbitstream is a bitstream (800) extracted from a larger bitstream (800). Thus, the term bitstream (800) may refer to the initially encoded bitstream (800), the extracted subbitstream, or both, depending on the context.

[0116] For example, the VPS (811) may include a TotalNumOlss (833). TotalNumOlss (833) is a syntax element that specifies the total number of OLS assigned to the VPS (811). This may include all OLS of the bitstream (800) containing the target OLS, which will be extracted along with any other OLS that may not be associated with a specific user request (e.g., other OLS associated with different screen resolutions, frame rates, etc.).

[0117] Additionally, the SEI message (819) may include a scalable nesting OLS flag (831). The scalable nesting OLS flag (831) is a flag that specifies whether the scalable nesting SEI message applies to a specific OLS or to a specific layer. The scalable nesting OLS flag (831) may be set to 1 to specify that the scalable nested SEI message within the scalable nesting SEI message applies to a specific OLS, or set to 0 to specify that the scalable nested SEI message applies to a specific layer. The SEI message (819) may also include other syntax elements such as the number of OLS minus 1 (num_olss_minus1) (835) and the difference in OLS index minus 1 (ols_idx_delta_minus1) (837). num_olss_minus1 (835) is a syntax element that specifies the number of OLSs to which the corresponding scalable nested SEI message is applied in the scalable nesting SEI message. The value of num_olss_minus1 (835) is restricted to a range of 0 to TotalNumOlss (833) minus 1. ols_idx_delta_minus1 (837) contains a value that can be used to derive a variable of nesting OLS index (NestingOlsIdx[ i ]) which specifies the OLS index of the i-th OLS to which the scalable nested SEI message is applied when the scalable nesting OLS flag (831) is 1. The value of ols_idx_delta_minus1 (837) for the i-th OLS is restricted to a range of 0 to TotalNumOlss (833) minus 2. The minus paradigm can be used to reduce the number of bits used to represent a value. For example, the actual value of a minus 1 syntax element can be determined by adding 1.

[0118] Preceding data can be used to determine whether to include the scalable nested SEI message in the subbitstream or to remove the scalable nested SEI message before conformity testing and / or transmission. Specifically, when the scalable nested OLS flag (831) is set to indicate that the scalable nested SEI message applies to one or more specific OLSs, for example, the scalable nested OLS flag is set to 1. Then, the encoder / HRD checks the scalable nested SEI message to determine whether the included scalable nested SEI message is associated with a layer of the target OLS. When neither the scalable nested SEI message nor the scalable nested SEI message refers to the target OLS, the scalable nested SEI message may be removed from the subbitstream. For example, the encoder / HRD may check the value of num_olss_minus1 (835) to determine how many OLSs are associated with the scalable nested SEI message. Then, the encoder can check each OLS between 0 and num_olss_minus1 (835) for relevance to the target OLS. For each current OLS, the encoder / HRD can determine the OLS index of the current (i-th) OLS to which the scalable nested SEI message is applied by examining the value of ols_idx_delta_minus1 (837) to derive the value of NestingOlsIdx[ i ]. If the value of NestingOlsIdx[ i ]. derived from ols_idx_delta_minus1 (837) is not the same as the target OLS index (targetOlsIdx) associated with the target OLS, none of the scalable nested SEI messages within the scalable nested SEI message are applied to the target OLS. In such cases, the scalable nested SEI message can be removed from the extracted sub-bitstream without negatively affecting HRD suitability testing or decoding.This mechanism allows some SEI messages (819) to be removed during sub-bitstream extraction from the bitstream (800). This increases coding efficiency for the resulting sub-bitstream, which reduces the use of processor, memory, and / or network resources in both the encoder and the decoder. Additionally, removing irrelevant SEI messages (819) can reduce the complexity of the HRD conformity testing process, which reduces the use of processor and / or memory resources in the encoder and / or HRD.

[0119] The previous information is now explained in more detail below. Layered video coding is also referred to as scalable video coding or video coding with scalability. Scalability in video coding can be supported by using multi-layer coding techniques. A multi-layer bitstream includes a base layer (BL) and one or more enhancement layers (EL). Examples of scalability include spatial scalability, quality / SNR (signal-to-noise ratio) scalability, multi-view scalability, and frame rate scalability. When using multi-layer coding techniques, a picture or part thereof can be coded without using a reference picture (intra-prediction), coded by referencing a reference picture on the same layer (inter-prediction), or coded by referencing a reference picture on a different layer(s) (inter-layer prediction). The reference picture used for inter-layer prediction of the current picture is called the Inter-Layer Reference Picture (ILRP). Figure 6 illustrates an example of multi-layer coding for spatial scalability where pictures of different layers have different resolutions.

[0120] Some video coding families support profiles for single-layer coding and the scalability of separate profiles. Scalable Video Coding (SVC) is a scalable extension of Advanced Video Coding (AVC) that supports spatial, temporal, and quality scalability. In the case of SVC, a flag is signaled at each macroblock (MB) of an EL picture to indicate whether the EL MB is predicted using a colocated block from the lower layer. Predictions from the colocated block may include textures, motion vectors, and / or coding modes. An implementation of SVC cannot directly reuse an AVC implementation without modifications in the design. The SVC EL macroblock syntax and decoding process differs from the AVC syntax and decoding process.

[0121] Scalable HEVC (SHVC) is an extension of HEVC that supports spatial and quality scalability. Multi-view HEVC (MV-HEVC) is an extension of HEVC that supports multi-view scalability. 3D HEVC (3D-HEVC) is an extension of HEVC that provides support for 3D video coding that is more advanced and efficient than MV-HEVC. Temporal scalability can be included as an essential part of a single-layer HEVC codec. In multi-layer extensions of HEVC, the decoded picture used for inter-layer prediction is provided only within the same AU and is treated as a long-term reference picture (LTRP). These pictures are assigned a reference index from the reference picture list(s) along with other temporal reference pictures of the current layer. Inter-layer prediction (ILP) is achieved at the prediction unit (PU) level by setting the value of the reference index to reference the inter-layer reference picture(s) in the reference picture list(s). Spatial scalability involves resampling the reference picture or a portion thereof when the ILRP has a different spatial resolution than the current picture being encoded or decoded. Reference picture resampling can be implemented at the picture level or the coding block level.

[0122] VVC can also support layered video coding. A VVC bitstream can contain multiple layers. The layers can all be independent of each other. For example, each layer can be coded without using inter-layer prediction. In this case, the layer is also called a simulcast layer. In some cases, some layers are coded using ILP. Flags in the VPS can indicate whether a layer is a simulcast layer or whether some layers use ILP. If some layers use ILP, layer dependencies between layers are also signaled in the VPS. Unlike SHVC and MV-HEVC, VVC may not specify an OLS. An OLS contains a specified set of layers, and one or more layers of the layer set are designated as output layers. An output layer is the layer of the OLS to be output. In some implementations of VVC, if a layer is a simulcast layer, only one layer may be selected for decoding and output. In some implementations of VVC, the entire bitstream containing all layers is designated to be decoded when any layer uses ILP. Additionally, certain layers among the layers are designated as output layers. Output layers can be designated as the top layer, all layers, or only the top layer plus the designated set of sub-layers.

[0123] Video coding standards may specify HRD to verify the conformity of a bitstream through designated HRD conformity tests. In SHVC and MV-HEVC, three sets of bitstream conformity tests are used to verify the conformity of a bitstream. The bitstream is referred to as the entire bitstream and is denoted as entireBitstream. The first set of bitstream conformity tests is intended to test the conformity of the entire bitstream and its corresponding temporal subset. These tests are used regardless of whether there is a set of layers specified by the active VPS that includes all nuh_layer_id values ​​of VCL NAL units present in the entire bitstream. Therefore, the conformity of the entire bitstream is always checked even if one or more layers are not included in the output set. The second set of bitstream conformity tests is used to test the conformity of the set of layers specified by the active VPS and its associated temporal subset. In all these tests, only the base layer picture (e.g., the picture with nuh_layer_id 0) is decoded and output. Other pictures are ignored by the decoder when the decoding process is invoked. The third set of bitstream conformance tests is used to test the conformance of OLS specified by the VPS extension portion of the active VPS and the relevant temporal subset based on OLS and bitstream partitions. Bitstream partitions include one or more layers of OLS of a multilayer bitstream.

[0124] The previous approach has specific problems. For example, the first two sets of conformance tests can be applied to layers that are not decoded and are not output. For instance, layers other than the lowest layer may not be decoded and are not output. In a real-world application, a decoder may only receive data to be decoded. Therefore, using both sets of the first conformance tests complicates the codec design and can waste bits by passing both the sequence-level and picture-level parameters used to support the conformance tests. The third set of conformance tests includes bitstream partitions. These partitions can relate to one or more layers of OLS for a multi-layer bitstream. Instead, HRD can be significantly simplified if the conformance tests always operate separately for each layer.

[0125] The signaling of sequence-level HRD parameters can be complex. For example, sequence-level HRD parameters may be signaled at multiple locations, such as both SPS and VPS. Additionally, the signaling of sequence-level HRD parameters may involve redundancy. For instance, information that may generally be the same for the entire bitstream may be repeated at each layer of each OLS. Furthermore, examples of HRD schemes allow for the selection of different delivery schedules for each layer. These delivery schedules may be selected from a list of schedules signaled for each layer for each operation point, where the operation point is an OLS or a temporal subset of the OLS. Such a system is complex. Additionally, exemplary HRD schemes allow incomplete AUs to be associated with buffering period SEI messages. An incomplete AU is an AU that lacks a picture for all layers in the CVS. However, HRD initialization in such AUs can be problematic. For example, HRD may not be properly initialized for layers that have layer access units that are not present in the incomplete AU. Additionally, the demultiplexing process for deriving the layer bitstream may not sufficiently and efficiently remove nested SEI messages that are not applied to the target layer. The layer bitstream occurs when the bitstream partition contains only one layer. Furthermore, applicable OLS for scalable unnested buffering periods, picture timing, and decoding unit information SEI messages can be specified for the entire bitstream. However, scalable unnested buffering periods should instead be applied to the 0th OLS.

[0126] Additionally, some VVC implementations may fail to infer HDR parameters when sub_layer_cpb_params_present_flag is equal to 0. This inference can enable proper HRD operations. Furthermore, the values ​​of bp_max_sub_layers_minus1 and pt_max_sub_layers_minus1 may need to be equal to the value of sps_max_sub_layers_minus1. However, buffering cycle and picture timing SEI messages can be nested and applied to multiple OLSs and multiple layers of each OLS. In this context, the relevant layers may reference multiple SPSs. Therefore, it may be difficult for the system to track which SPS corresponds to each layer. Consequently, the values ​​of these two syntax elements should instead be constrained based on the value of vps_max_sub_layers_minus1. Furthermore, since different layers may have different numbers of sub-layers, the values ​​of these two syntax elements may not always be the same as specific values ​​in all buffering cycles and picture timing SEI messages.

[0127] In addition, there is the following issue related to HRD design in both SHVC / MV-HEVC and VVC. The sub-bitstream extraction process may not remove SEI NAL units containing nested SEI messages that are not required by the target OLS.

[0128] Generally, the present disclosure describes an approach for scalable nesting of SEI messages for a set of output layers in a multi-layer video bitstream. The description of the technique is based on VVC. However, this technique also applies to layered video coding based on other video codec specifications.

[0129] One or more of the problems mentioned above can be solved as follows. Specifically, the present disclosure includes methods for HRD designs and related aspects that allow for efficient signaling of HRD parameters with much simpler HRD behavior compared to SHVC and MV-HEVC. Each solution described below corresponds to the problem described above. For example, instead of requiring three sets of conformance tests, the present disclosure may use only one set of conformance tests to test the conformance of the OLS specified by the VPS. Furthermore, instead of a design based on bitstream partitioning, the disclosed HRD mechanism may always operate separately for each layer of the OLS. Additionally, global sequence-level HRD parameters for all layers and all sublayers of the OLS may be signaled only once, for example, by the VPS. Additionally, a single number of forwarding schedules may be signaled for all layers and sublayers of the OLS. The same forwarding schedule index may also be applied to all layers of the OLS. Additionally, incomplete AUs may not be associated with buffering cycle SEI messages. An incomplete AU is an AU that does not contain pictures for all layers in the CVS. This ensures that HRD can always be properly initialized for all layers of the OLS. Additionally, a mechanism is disclosed for efficiently removing nested SEI messages that are not applied to the target layer in the OLS. This supports a demultiplexing process for deriving the layer bitstream. Furthermore, the applicable OLS for scalable unnested buffering cycles, picture timing, and decoding unit information SEI messages can be specified as the 0th OLS. Additionally, HDR parameters can be inferred when sub_layer_cpb_params_present_flag is equal to 0, which enables proper HRD operation.The values ​​of bp_max_sub_layers_minus1 and pt_max_sub_layers_minus1 may be in the range of 0 or greater and vps_max_sub_layers_minus1 or less. In this way, these parameters do not need to be specific values ​​for all buffering cycles and picture timing SEI messages. Additionally, the sub-bitstream extraction process can remove SEI NAL units containing nested SEI messages that are not applied to the target OLS.

[0130] An example of the implementation of the previous mechanism is as follows. An output layer is a layer of the output layer set being output. An OLS is a layer set containing a specified layer set, where one or more layers of the layer set are specified as output layers. An OLS layer index is a layer index of the OLS for the layer list of the OLS. The sub-bitstream extraction process is a specified process in which NAL units of a bitstream not belonging to the target set determined by the target OLS index and the target max TemporalId are removed from the bitstream along with an output sub-bitstream containing NAL units of a bitstream belonging to the target set.

[0131] An example of video parameter set syntax is as follows.

[0132] video_parameter_set_rbsp( ) { Descriptor ... general_hrd_params_present_flag u(1) if( general_hrd_params_present_flag ) { num_units_in_tick u(32) time_scale u(32) general_hrd_parameters( ) } vps_extension_flag u(1) if( vps_extension_flag ) while( more_rbsp_data( ) ) vps_extension_data_flag u(1) rbsp_trailing_bits( ) }

[0133] An exemplary sequence parameter set RBSP syntax is as follows.

[0134] seq_parameter_set_rbsp( ) { Descriptor sps_decoding_parameter_set_id u(4) sps_video_parameter_set_id u(4) sps_max_sub_layers_minus1 u(3) sps_reserved_zero_4bits u(4) same_nonoutput_level_and_dpb_size_flag u(1) profile_tier_level( 1, sps_max_sub_layers_minus1 ) if( !same_nonoutput_level_and_dpb_size_flag ) profile_tier_level( 0, sps_max_sub_layers_minus1 ) ... if( sps_max_sub_layers_minus1 > 0 ) sps_sub_layer_ordering_info_present_flag u(1) dpb_parameters( 1 ) if( !same_nonoutput_level_and_dpb_size_flag ) dpb_parameters( 0 ) long_term_ref_pics_flag u(1) ... sps_scaling_list_enabled_flag u(1) vui_parameters_present_flag u(1) if( vui_parameters_present_flag ) vui_parameters( ) sps_extension_flag u(1) if( sps_extension_flag ) while( more_rbsp_data( ) ) sps_extension_data_flag u(1) rbsp_trailing_bits( ) }

[0135] An example of DPB parameter syntax is as follows.

[0136] dpb_parameters( reorderMaxLatencyPresentFlag ) { Descriptor for( i = ( sps_sub_layer_ordering_info_present_flag ? 0 : sps_max_sub_layers_minus1 ); i <= sps_max_sub_layers_minus1; i++ ) { sps_max_dec_pic_buffering_minus1[ i ] ue(v) if( reorderMaxLatencyPresentFlag ) { sps_max_num_reorder_pics[ i ] ue(v) sps_max_latency_increase_plus1[ i ] ue(v) } } }

[0137] An example of a typical HRD parameter syntax is as follows.

[0138] general_hrd_parameters( ) { Descriptor general_nal_hrd_params_present_flag u(1) general_vcl_hrd_params_present_flag u(1) if( general_nal_hrd_params_present_flag | | general_vcl_hrd_params_present_flag ) { decoding_unit_hrd_params_present_flag u(1) if( decoding_unit_hrd_params_present_flag ) { tick_divisor_minus2 u(8) decoding_unit_cpb_params_in_pic_timing_sei_flag u(1) } bit_rate_scale u(4) cpb_size_scale u(4) if( decoding_unit_hrd_params_present_flag ) cpb_size_du_scale u(4) } if( vps_max_sub_layers_minus1 > 0 ) sub_layer_cpb_params_present_flag u(1) if( TotalNumOlss > 1 ) num_layer_hrd_params_minus1 ue(v) hrd_cpb_cnt_minus1 ue(v) for( i = 0; i <= num_layer_hrd_params_minus1; i++ ) { if( vps_max_sub_layers_minus1 > 0 ) hrd_max_temporal_id[ i ] u(3) layer_level_hrd_parameters( hrd_max_temporal_id[ i ] ) } if( num_layer_hrd_params_minus1 > 0 ) for( i = 1; i < TotalNumOlss; i++ ) for( j = 0; j < NumLayersInOls[ i ]; j++ ) layer_level_hrd_idx[ i ][ j ] ue(v) }

[0139] The semantics of the exemplary video parameter set RBSP are as follows: each_layer_is_an_ols_flag is set to 1 to specify that each output layer set is a single-include layer set containing only one layer, where each layer of the bitstream itself is the only output layer. each_layer_is_an_ols_flag is set to 0 to specify that the output layer set may contain one or more layers. If vps_max_layers_minus1 is equal to 0, the value of each_layer_is_an_ols_flag is inferred to be equal to 1. Otherwise, when vps_all_independent_layers_flag is equal to 0, the value of each_layer_is_an_ols_flag is inferred to be equal to 0. ols_mode_idc is set to 0 to specify that the total number of OLS specified by the VPS is equal to vps_max_layers_minus1 + 1, the i-th OLS contains layers with layer indices from 0 to I, and for each OLS, only the highest layer index in the OLS is output. ols_mode_idc is set to 1 to specify that the total number of OLS specified by the VPS is equal to vps_max_layers_minus1 + 1, the i-th OLS contains layers with layer indices from 0 to I, and for each OLS, all layers in the OLS are output. ols_mode_idc is set to 2 to specify that the total number of OLS specified by the VPS is explicitly signaled, and for each OLS, the highest layer in the OLS and the set of lower layers explicitly signaled in the OLS are output. The value of ols_mode_idc must be in the range of 0 to 2. The value of ols_mode_idc, 3, is reserved.When vps_all_independent_layers_flag is equal to 1 and each_layer_is_an_ols_flag is 0, the value of ols_mode_idc is inferred to be equal to 2. num_output_layer_sets_minus1 plus 1 specifies the total number of OLS specified by VPS when ols_mode_idc is 2.

[0140] The variable TotalNumOlss, which specifies the total number of OLS specified by VPS, is derived as follows.

[0141]

[0142] layer_included_flag[ i ][ j ] specifies whether the j-th layer (the layer with the same nuh_layer_id as vps_layer_id[ j ]) is included in the i-th OLS when ols_mode_idc is 2. layer_included_flag[ i ][ j ] is set to 1 to specify that the j-th layer is included in the i-th OLS. layer_included_flag[ i ][ j ] is set to 0 to specify that the j-th layer is not included in the i-th OLS.

[0143] The variable NumLayerInOls[ i ], which specifies the number of layers in the i-th OLS, and the variable LayerIdInOls[ i ][ j ], which specifies the nuh_layer_id value of the j-th layer in the i-th OLS, are derived as follows.

[0144]

[0145] The variable OlsLayeIdx[ i ][ j ], which specifies the OLS layer index of the layer where nuh_layer_id is the same as LayerIdInOls[ i ][ j ], is derived as follows.

[0146]

[0147] The lowest layer of each OLS must be an independent layer. In other words, for each i in the range from 0 to TotalNumOlss - 1, the value of vps_independent_layer_flag[ GeneralLayerIdx[ LayerIdInOls[ i ]

[0000] ] ] must be equal to 1. Each layer must be included in at least one OLS specified by the VPS. In other words, for each layer having a specific value of nuh_layer_id nuhLayerId that is equal to one of vps_layer_id[ k ] for k in the range from 0 to vps_max_layers_minus1, there must be at least one pair of values ​​i and j, and since i is in the range from 0 to TotalNumOlss - 1 and j is in the range from 0 to NumLayerInOls[ i ] - 1, the value of LayerIdInOls[ i ][ j ] is equal to nuhLayerId. All layers of OLS must be either OLS output layers or (direct or indirect) reference layers of OLS output layers.

[0148] vps_output_layer_flag[ i ][ j ] specifies whether the j-th layer of the i-th OLS is output when ols_mode_idc is 2. vps_output_layer_flag[ i ], which is equal to 1, specifies that the j-th layer of the i-th OLS is output. vps_output_layer_flag[ i ] is set to equal 0 to specify that the j-th layer of the i-th OLS is not output. When vps_all_independent_layers_flag is 1 and each_layer_is_an_ols_flag is 0, the value of vps_output_layer_flag[ i ] is inferred to be equal to 1. The variable OutputLayerFlag[ i ][ j ], where a value of 1 specifies that the j-th layer of the i-th OLS is output and a value of 0 specifies that the j-th layer of the i-th OLS is not output, is derived as follows.

[0149]

[0150] The 0th OLS contains only the lowest layer (the layer where nuh_layer_id is vps_layer_id

[0000] ), and for the 0th OLS, the only included layer is output.

[0151] vps_extension_flag is set to 0 to indicate that the vps_extension_data_flag syntax element does not exist in the VPS RBSP syntax structure. vps_extension_flag is set to 1 to indicate that the vps_extension_data_flag syntax element exists in the VPS RBSP syntax structure. vps_extension_data_flag can have an arbitrary value. The presence and value of vps_extension_data_flag do not affect decoder suitability for the specified profile. The decoder must ignore all vps_extension_data_flag syntax elements.

[0152] Examples of DPB parameter semantics are as follows. The dpb_parameters() syntax structure provides DPB size information and optionally provides information on the maximum picture reorder number and maximum latency (MRML). Each SPS contains one or more dpb_parameters() syntax structures. The first dpb_parameters() syntax structure in an SPS contains both DPB size information and MRML information. If present, the second dpb_parameters() syntax structure in an SPS contains only DPB size information. The MRML information in the first dpb_parameters() syntax structure in an SPS applies to layers referencing the SPS regardless of whether the corresponding layer is an output layer of the OLS. The DPB size information in the first dpb_parameters() syntax structure in an SPS applies to layers referencing the SPS if the layer is an output layer of the OLS. If present, the DPB size information contained in the second dpb_parameters() syntax structure in the SPS is applied to the layer referencing the SPS when the layer is a non-output layer of OLS. If the SPS contains only one dpb_parameters() syntax structure, the DPB size information for a non-output layer is inferred to be the same as the DPB size information for a layer that is an output layer.

[0153] Examples of general HRD parameter semantics are as follows. The general_hrd_parameters() syntax structure provides HRD parameters used for HRD operations. sub_layer_cpb_params_present_flag is set to 1 to specify that the i-th layer_level_hrd_parameters() syntax structure contains HRD parameters for sublayer representations with TemporalIds in the range of 0 to hrd_max_temporal_id[ i ]. sub_layer_cpb_params_present_flag is set to 0 to specify that the i-th layer_level_hrd_parameters() syntax structure contains only HRD parameters for sublayer representations with TemporalIds equal to hrd_max_temporal_id[ i ]. When vps_max_sub_layers_minus1 is equal to 0, the value of sub_layer_cpb_params_present_flag is inferred to be equal to 0. When sub_layer_cpb_params_present_flag is equal to 0, HRD parameters for sublayer representations with TemporalId in the range of 0 to hrd_max_temporal_id[ i ] - 1 are presumed to be identical to HRD parameters for sublayer representations with a TemporalId equal to hrd_max_temporal_id[ i ]. This includes HRD parameters from the fixed_pic_rate_general_flag[ i ] syntax element immediately below the condition if( general_vcl_hrd_params_present_flag ) in the layer_level_hrd_parameters syntax structure up to the sub_layer_hrd_parameters( i ) syntax structure.num_layer_hrd_params_minus1 + 1 specifies the number of layer_level_hrd_parameters() syntax structures present in the general_hrd_parameters() syntax structure. The value of num_layer_hrd_params_minus1 must be in the range of 0 to 63. hrd_cpb_cnt_minus1 + 1 specifies the number of alternate CPB specifications in the CVS bitstream. The value of hrd_cpb_cnt_minus1 must be in the range of 0 to 31. hrd_max_temporal_id[ i ] specifies the TemporalId of the top-level sublayer representation in which the HRD parameter is included in the i-th layer_level_hrd_parameters() syntax structure. The value of hrd_max_temporal_id[ i ] must be in the range of 0 to vps_max_sub_layers_minus1. When vps_max_sub_layers_minus1 is equal to 0, the value of hrd_max_temporal_id[ i ] is inferred to be equal to 0. layer_level_hrd_idx[ i ][ j ] specifies the index of the layer_level_hrd_parameters() syntax structure applied to the j-th layer in the i-th OLS. The value of layer_level_hrd_idx[[ i ][ j ] must be in the range from 0 to num_layer_hrd_params_minus1. If it does not exist, the value of layer_level_hrd_idx[

[0000]

[0000] is inferred to be equal to 0.

[0154] An exemplary sub-bitstream extraction process is as follows. The inputs to this process are the bitstream inBitstream, the target OLS index targetOlsIdx, and the target highest TemporalId value tIdTarget. The output of this process is the sub-bitstream outBitstream. The bitstream conformity requirement for the input bitstream is that any output sub-bitstream, which is the output of the process specified in this section, must be a bitstream that satisfies the following conditions as input, and that the targetOlsIdx is the same as the index for the OLS list specified by the VPS, and that the tIdTarget is the same as any value in the range from 0 to 6. The output sub-bitstream must contain at least one VCL NAL unit having a nuh_layer_id that is the same as each nuh_layer_id value of LayerIdInOls[ targetOlsIdx ]. The output sub-bitstream must contain at least one VCL NAL unit whose TemporalId is equal to tIdTarget. The matching bitstream contains one or more coded slice NAL units with TemporalId 0, but does not need to contain coded slice NAL units with nuh_layer_id 0.

[0155] The output sub-bitstream OutBitstream is derived as follows. The bitstream outBitstream is set to be the same as the bitstream inBitstream. All NAL units with TemporalId greater than tIdTarget are removed from outBitstream. All NAL devices with nuh_layer_id not included in the LayerIdInOls[ targetOlsIdx ] list are removed from outBitstream. All SEI NAL units containing scalable nesting SEI messages with no i value in the range from 0 to nesting_num_olss_minus1, such that nesting_ols_flag is 1 and NestingOlsIdx[ i ] is equal to targetOlsIdx, are removed from outBitstream. If targetOlsIdx is greater than 0, remove all SEI NAL units from outBitstream that contain scalable non-nested SEI messages with payloadType 0 (buffering period), 1 (picture timing), or 130 (decoding unit information).

[0156] Examples of general aspects of HRD are as follows. This section specifies HRD and its use to check bitstream and decoder conformity. The bitstream conformity test set is referred to as the entire bitstream, is used to check the conformity of the bitstream, and is denoted as entireBitstream. The bitstream conformity test set is intended to test the conformity of each OLS specified by the VPS and a temporal subset of each OLS. For each test, the following sequential steps are applied in the order listed.

[0157] Select the target OLS with the OLS index opOlsIdx and the highest TemporalId value opTid, and select the operation point under test, denoted as targetOp. The value of opOlsIdx is in the range from 0 to TotalNumOlss - 1. The value of opTid is in the range from 0 to vps_max_sub_layers_minus1. The values ​​of opOlsIdx and opTid ensure that the sub-bitstream BitstreamToDecode output by calling the sub-bitstream extraction process with entireBitstream, opOlsIdx, and opTid as input satisfies the following conditions: There is at least one VCL NAL unit having a nuh_layer_id that is identical to each of the nuh_layer_id values ​​in LayerIdInOls[ opOlsIdx ]. BitstreamToDecode has at least one VCL NAL unit whose TemporalId is equal to opTid.

[0158] The values ​​of TargetOlsIdx and Htid are set to be equal to the opOlsIdx and opTid of targetOp, respectively. A ScIdx value is selected. The selected ScIdx must be in the range of 0 or greater and hrd_cpb_cnt_minus1 or less. An access unit of BitstreamToDecode associated with a buffering period SEI message applicable to TargetOlsIdx (which is in TargetLayerBitstream or available through an external mechanism not specified herein) is selected as the HRD initialization point and is called access unit 0 for each layer of the target OLS.

[0159] The subsequent step is applied to each layer in the target OLS with the OLS layer index TargetOlsLayerIdx. If the target OLS has only one layer, the bitstream of the layer being tested, TargetLayerBitstream, is set to be the same as BitstreamToDecode. Otherwise, the demultiplexing process is called to derive the layer bitstream using BitstreamToDecode, TargetOlsIdx, and TargetOlsLayerIdx as inputs to derive TargetLayerBitstream, and the output is assigned to TargetLayerBitstream.

[0160] The layer_level_hrd_parameters() and sub_layer_hrd_parameters() syntax structures applicable to TargetLayerBitstream are selected as follows. The layer_level_hrd_parameters() syntax structure at the layer_level_hrd_idx[ TargetOlsIdx ][ TargetOlsLayerIdx ] of the VPS (or provided through an external mechanism such as user input) is selected. Within the selected layer_level_hrd_parameters() syntax structure, if BitstreamToDecode is a bitstream of type I, the sub_layer_hrd_parameters( Htid ) syntax structure immediately following the condition if( general_vcl_hrd_params_present_flag ) is selected, and the variable NalHrdModeFlag is set to 0. Otherwise (BitstreamToDecode is a type II bitstream), the condition if( general_vcl_hrd_params_present_flag )(if the variable NalHrdModeFlag is set to 0) or the condition if( general_nal_hrd_params_present_flag )(if the variable NalHrdModeFlag is set to 1) is selected. When BitstreamToDecode is a type II bitstream and NalHrdModeFlag is equal to 0, all non-VCL NAL units except the filler data NAL unit, and if present, all leading_zero_8bits, zero_byte, start_code_prefix_one_3bytes, and trailing_zero_8bits syntax elements forming a byte stream from the NAL unit stream are removed from TargetLayerBitstream and the remaining bitstream is assigned to TargetLayerBitstream.

[0161] When Decoding_unit_hrd_params_present_flag is equal to 1, CPB is scheduled to operate at the access unit level (in which case the variable DecodingUnitHrdFlag is set to 0) or the decoding unit level (in which case the variable DecodingUnitHrdFlag is set to 1). Otherwise, DecodingUnitHrdFlag is set to 0 and CPB is scheduled to operate at the access unit level. For each access unit of TargetLayerBitstream starting from access unit 0, a buffering duration SEI message (present in TargetLayerBitstream or available through an external mechanism) associated with the access unit and applicable to TargetOlsIdx and TargetOlsLayerIdx is selected, a picture timing SEI message (present in TargetLayerBitstream or available through an external mechanism) associated with the access unit and applicable to TargetOlsIdx and TargetOlsLayerIdx is selected, and when DecodingUnitHrdFlag is 1 and decoding_unit_cpb_params_in_pic_timing_sei_flag is 0, a decoding unit information SEI message (present in TargetLayerBitstream or available through an external mechanism) associated with the decoding unit of the access unit and applicable to TargetOlsIdx and TargetOlsLayerIdx is selected.

[0162] Each conformance test includes one combination of options from each of the steps above. If there are more than one option for a step, only one option is selected for a specific conformance test. All possible combinations of all steps form the entire set of conformance tests. For each operation point under test, the number of bitstream conformance tests to be performed is equal to n0 * n1 * n2 * n3, where the values ​​of n0, n1, n2, and n3 are specified as follows: n1 is equal to hrd_cpb_cnt_minus1 + 1. n1 is the number of access units of BitstreamToDecode associated with the buffering period SEI message. n2 is derived as follows: If BitstreamToDecode is a type I bitstream, n0 is equal to 1. Otherwise (BitstreamToDecode is a type II bitstream), n0 is equal to 2. n3 is derived as follows: If Decoding_unit_hrd_params_present_flag is equal to 0, n3 is equal to 1. Otherwise, n3 is equal to 2.

[0163] HRD includes a bitstream demultiplexer (optional), a coded picture buffer (CPB) for each layer, an immediate decoding process for each layer, a decoded picture buffer (DPB) containing a sub-DPB for each layer, and output cropping.

[0164] In the example, the HRD operates as follows. The HRD is initialized at Decoding Unit 0, and each sub-DPB and each CPB of the DPB is set to empty. The sub-DPB fullness for each sub-DPB is set to 0. After initialization, the HRD is not re-initialized by subsequent buffering cycle SEI messages. Data associated with the Decoding Unit flowing to each CPB according to the specified arrival schedule is delivered by the HSS. The data associated with each Decoding Unit is immediately removed and decoded by the instantaneous decoding process at the time of removal of the CPB of the Decoding Unit. Each decoded picture is placed in the DPB. The decoded picture is removed from the DPB when it is no longer needed for inter-predictive reference and no longer needed for the output.

[0165] In one example, the demultiplexing process for deriving a layer bitstream is as follows. The inputs to this process are the bitstream inBitstream, the target OLS index targetOlsIdx, and the target OLS layer index targetOlsLayerIdx. The output of this process is the layer bitstream outBitstream. The output layer bitstream outBitstream is derived as follows. The bitstream outBitstream is set to be the same as the bitstream inBitstream. All NAL units whose nuh_layer_id is not equal to LayerIdInOls[ targetOlsIdx ][ targetOlsLayerIdx ] are removed from outBitstream. Remove all SEI NAL units from outBitstream containing scalable nesting SEI messages where nesting_ols_flag is 1 and i and j values ​​are not in the ranges from 0 to nesting_num_olss_minus1 and from 0 to nesting_num_ols_layers_minus1[ i ], respectively, so that NestingOlsLayerIdx[ i ][ j ] is equal to targetOlsLayerIdx. Remove all SEI NAL units from outBitstream containing scalable nesting SEI messages where nesting_ols_flag is 1 and i and j values ​​are in the ranges from 0 to nesting_num_olss_minus1 and from 0 to nesting_num_ols_layers_minus1[ i ], respectively, so that NestingOlsLayerIdx[ i ][ j ] is less than targetOlsLayerIdx.Remove all SEI NAL units from outBitstream containing scalable nesting SEI messages where nesting_ols_flag is 0 and i is not in the range of 0 to NestingNumLayers - 1, so that NestingLayerId[ i ] is equal to LayerIdInOls[ targetOlsIdx ][ targetOlsLayerIdx ].

[0166] An example of buffering period SEI message syntax is as follows.

[0167] buffering_period( payloadSize ) { Descriptor ... bp_max_sub_layers_minus1 u(3) bp_cpb_cnt_minus1 ue(v) ... }

[0168] An example of scalable nesting SEI message syntax is as follows.

[0169] scalable_nesting( payloadSize ) { Descriptor nesting_ols_flag u(1) if( nesting_ols_flag ) { nesting_num_olss_minus1 ue(v) for( i = 0; i <= nesting_num_olss_minus1; i++ ) { nesting_ols_idx_delta_minus1[ i ] ue(v) if( NumLayersInOls[ NestingOlsIdx[ i ] ] > 1 ) { nesting_num_ols_layers_minus1[ i ] ue(v) for( j = 0; j <= nesting_num_ols_layers_minus1[ i ]; j++ ) nesting_ols_layer_idx_delta_minus1[ i ][ j ] ue(v) } } } else { nesting_all_layers_flag u(1) if( !nesting_all_layers_flag ) { nesting_num_layers_minus1 ue(v) for( i = 1; i <= nesting_num_layers_minus1; i++ ) nesting_layer_id[ i ] u(6) } } nesting_num_seis_minus1 ue(v) while( !byte_aligned( ) ) nesting_zero_bit / * equal to 0 * / u(1) for( i = 0; i <= nesting_num_seis_minus1; i++ ) sei_message( ) }

[0170] Examples of general SEI payload semantics are as follows. The following apply to the applicable layers (in the context of OLS or generally) of a scalable non-nested SEI message. For a scalable non-nested SEI message, when payloadType is 0 (buffering cycle), 1 (picture timing), or 130 (decoding unit information), the scalable non-nested SEI message applies only to the lowest layer in the context of the 0th OLC. For a scalable non-nested SEI message, when payloadType is equal to any value in VclAssociatedSeiList, the scalable non-nested SEI message applies only to the layer where the VCL NAL unit has a nuh_layer_id identical to the nuh_layer_id of the SEI NAL unit containing the SEI message. An exemplary buffering cycle SEI message semantic is as follows. The buffering cycle SEI message provides the initial CPB elimination delay and initial CPB elimination delay offset information for HRD initialization at the position of the associated access unit in the decoding order. When a buffering period SEI message is present, if a picture has a TemporalId equal to 0 and is not a RASL or RADL (Random Access Decodeable Leading) picture, the picture is called a notDiscardablePic picture. If the current picture is not the first picture in the bitstream in the decoding order, prevNonDiscardablePic is made to be the previous picture in the decoding order with a TemporalId equal to 0 that is not a RASL or RADL picture. The presence of a buffering period SEI message is specified as follows. If NalHrdBpPresentFlag is equal to 1 or VclHrdBpPresentFlag is equal to 1, the following applies to each access unit of the CVS.If the access unit is an IRAP or GDR (Gradual Decoder Refresh) access unit, the buffering period SEI message applicable to the operation point must be associated with the access unit. Otherwise, if the access unit includes notDiscardablePic, the buffering period SEI message applicable to the operation point may or may not be associated with the access unit. Otherwise, the access unit is not associated with the buffering period SEI message applicable to the operation point. Otherwise (when both NalHrdBpPresentFlag and VclHrdBpPresentFlag are 0), the access unit of the CVS is not associated with the buffering period SEI message. For some applications, the frequent presence of buffering period SEI messages may be desirable (e.g., for random access or bitstream splicing on IRAP or non-IRAP pictures). When a picture of an access unit is associated with a buffering cycle SEI message, the access unit must have a picture of each layer existing in the CVS, and each picture of the access unit must be accompanied by a buffering cycle SEI message.

[0171] bp_max_sub_layers_minus1 + 1 specifies the maximum number of temporal sublayers where the CPB removal delay and CBP removal offset are indicated in the buffering cycle SEI message. The value of bp_max_sub_layers_minus1 must be in the range of 0 or greater and vps_max_sub_layers_minus1 or less. bp_cpb_cnt_minus1 + 1 specifies the number of syntax element pairs of nal_initial_cpb_removal_delay[ i ][ j ] and nal_initial_cpb_removal_offset[ i ][ j ] in the i-th temporal sublayer when bp_nal_hrd_params_present_flag is equal to 1, and specifies the number of syntax element pairs of vcl_initial_cpb_removal_delay[ i ][ j ] and vcl_initial_cpb_removal_offset[ i ][ j ] in the i-th temporal sublayer when bp_vcl_hrd_params_present_flag is equal to 1. The value of bp_cpb_cnt_minus1 must be in the range of 0 to 31. The value of bp_cpb_cnt_minus1 must be the same as the value of hrd_cpb_cnt_minus1.

[0172] The meaning of an exemplary picture timing SEI message is as follows. A picture timing SEI message provides CPB removal delay and DPB output delay information for the access unit associated with the SEI message. If bp_nal_hrd_params_present_flag or bp_vcl_hrd_params_present_flag of the buffering cycle SEI message applicable to the current access unit is 1, the variable CpbDpbDelaysPresentFlag is set to 1. Otherwise, CpbDpbDelaysPresentFlag is set to 0. The existence of a picture timing SEI message is specified as follows: If CpbDpbDelaysPresentFlag is equal to 1, the picture timing SEI message must be associated with the current access unit. Otherwise (if CpbDpbDelaysPresentFlag is equal to 0), there must be no picture timing SEI message associated with the current access unit. The TemporalId within the picture timing SEI message syntax is the TemporalId of the SEI NAL unit containing the picture timing SEI message. pt_max_sub_layers_minus1 + 1 specifies the TemporalId of the top-most sublayer representation in which CPB removal delay information is included in the picture timing SEI message. The value of pt_max_sub_layers_minus1 must be in the range of 0 or greater and vps_max_sub_layers_minus1 or less.

[0173] Examples of the semantics of a Scalable Nesting SEI message are as follows. Scalable Nesting SEI messages provide a mechanism to associate an SEI message with a specific layer within the context of a specific OLS, or with a specific layer outside the context of an OLS. A Scalable Nesting SEI message contains one or more SEI messages. SEI messages included in a Scalable Nesting SEI message are also referred to as Scalable Nested SEI messages. The following restrictions apply when including SEI messages in a Scalable Nesting SEI message as requirements of bitstream conformity. SEI messages with payloadTypes such as 132 (Decoded Picture Hash) or 133 (Scalable Nesting) are not included in a Scalable Nesting SEI message. When a scalable nesting SEI message contains a buffering period, picture timing, or decoding unit information SEI message, the scalable nesting SEI message does not contain other SI messages having a payloadType that is not equal to 0 (buffering period), 1 (picture timing), or 130 (decoding unit information).

[0174] It is a requirement of bitstream conformity that the following restrictions apply to the nal_unit_type value of an SEI NAL unit containing a scalable nesting SEI message. If the scalable nesting SEI message contains an SEI message with a payloadType such as 0 (buffering period), 1 (picture timing), 130 (decoding unit information), 145 (dependent RAP indication), or 168 (frame field information), the SEI NAL unit containing the scalable nesting SEI message must have a nal_unit_type equal to PREFIX_SEI_NUT. If the scalable nesting SEI message contains an SEI message with a payloadType such as 132 (decoded picture hash), the SEI NAL unit containing the scalable nesting SEI message must have a nal_unit_type equal to SUFFIX_SEI_NUT.

[0175] nesting_ols_flag is set to 1 to specify that scalable nested SEI messages are applied to a specific layer within the context of a specific OLS. nesting_ols_flag is set to 0 to specify that scalable nested SEI messages are generally applied to a specific layer (not within the context of an OLS). The following restrictions apply to the nesting_ols_flag value as requirements for bitstream conformity. When a scalable nested SEI message contains an SEI message with a payloadType such as 0 (buffering cycle), 1 (picture timing), or 130 (decoding unit information), the value of nesting_ols_flag must be equal to 1. When a scalable nested SEI message contains an SEI message with a payloadType identical to the value of VclAssociatedSeiList, the value of nesting_ols_flag must be equal to 0. nesting_num_olss_minus1 + 1 specifies the number of OLSs to which scalable nested SEI messages apply. The value of nesting_num_olss_minus1 must be in the range of 0 to TotalNumOlss - 1. nesting_ols_idx_delta_minus1[ i ] is used to derive the variable NestingOlsIdx[ i ], which specifies the OLS index of the i-th OLS to which scalable nested SEI messages apply when nesting_ols_flag is 1. The value of nesting_ols_idx_delta_minus1[ i ] must be in the range of 0 to TotalNumOlss - 2. The variable NestingOlsIdx[ i ] is derived as follows.

[0176]

[0177] nesting_num_ols_layers_minus1[ i ] + 1 specifies the number of layers to which scalable nested SEI messages are applied in the context of the NestingOlsIdx[ i ]-1 OLS. The value of nesting_num_ols_layers_minus1[ i ] must be in the range of 0 or greater and NumLayersInOls[ NestingOlsIdx[ i ] ] - 1 or less. nesting_ols_layer_idx_delta_minus1[ i ][ j ] is used to derive the variable NestingOlsLayerIdx[ i ][ j ], which specifies the OLS layer index of the j-th layer to which scalable nested SEI messages are applied in the context of the NestingOlsIdx[ i ]-1 OLS when nesting_ols_flag is equal to 1. The value of nesting_ols_layer_idx_delta_minus1[ i ] must be in the range of 0 or greater, or the binary NumLayersInOls[ nestingOlsIdx[ i ] ] - 2. The variable NestingOlsLayerIdx[ i ][ j ] is derived as follows.

[0178]

[0179] The lowest value among all values ​​of LayerIdInOls[ NestingOlsIdx[ i ] ] [ NestingOlsLayerIdx[ i ]

[0000] ] for i in the range from 0 to nesting_num_olss_minus1 must be equal to the nuh_layer_id of the current SEI NAL unit (the SEI NAL unit containing scalable nesting SEI messages). nesting_all_layers_flag is set to 1 to specify that scalable nested SEI messages are generally applied to all layers with a nuh_layer_id greater than or equal to the nuh_layer_id of the current SEI NAL unit. nesting_all_layers_flag is set to 0 to specify that scalable nested SEI messages may or may not be applied to all layers with a nuh_layer_id greater than or equal to the nuh_layer_id of the current SEI NAL unit. nesting_num_layers_minus1 + 1 specifies the number of layers to which scalable nested SEI messages are generally applied. The value of nesting_num_layers_minus1 must be in the range of 0 or greater, and vps_max_layers_minus1 - GeneralLayerIdx[ nuh_layer_id ] or less, where nuh_layer_id is the nuh_layer_id of the current SEI NAL unit. nesting_layer_id[ i ] specifies the nuh_layer_id value of the i-th layer to which scalable nested SEI messages are generally applied when nesting_all_layers_flag is 0. The value of nesting_layer_id[ i ] must be greater than nuh_layer_id, where nuh_layer_id is the nuh_layer_id of the current SEI NAL unit.When nesting_ols_flag is equal to 0, the variable NestingNumLayers, which specifies the number of layers to which scalable nested SEI messages are typically applied, and the list NestingLayerId[ i ] for i in the range of 0 to NestingNumLayers - 1, which specifies a list of nuh_layer_id values ​​of layers to which scalable nested SEI messages are typically applied, are derived as follows, where nuh_layer_id is the nuh_layer_id of the current SEI NAL unit.

[0180]

[0181] nesting_num_seis_minus1 + 1 specifies the number of scalable nested SEI messages. The value of nesting_num_seis_minus1 must be in the range of 0 to 63. nesting_zero_bit must be equal to 0.

[0182] FIG. 9 is a schematic diagram of an exemplary video coding device (900). The video coding device (900) is suitable for implementing the disclosed examples / executions as described herein. The video coding device (900) includes a downstream port (920) comprising a transmitter and / or receiver for communicating data upstream and / or downstream over a network, an upstream port (950), and / or a transceiver unit (Tx / Rx, 910). The video coding device (900) also includes a processor (930) comprising a logic unit and / or a central processing unit (CPU) for processing data and a memory (932) for storing data. The video coding device (900) may also include an electric, optical-electric (OE) component, an electric-optical (EO) component, and / or a wireless communication component coupled to the upstream port (950) and / or downstream port (920) for communicating data over an electric, optical, or wireless communication network. The video coding device (900) may also include an input and / or output (I / O) device (960) for communicating data with a user. The input / output device (960) may include an output device such as a display for displaying video data, a speaker for outputting audio data, etc. The I / O device (960) may also include an input device such as a keyboard, mouse, trackball, etc., and / or a corresponding interface for interacting with such output devices.

[0183] The processor (930) is implemented in hardware and software. The processor (930) may be implemented as one or more CPU chips, cores (e.g., as multi-core processors), field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), and digital signal processors (DSPs). The processor (930) communicates with a downstream port (920), a Tx / Rx (910), an upstream port (950), and memory (932). The processor (930) includes a coding module (914). The coding module (914) implements the disclosed embodiments described herein, such as methods (100, 1000, 1100) that can use a multi-layer video sequence (600), a multi-layer video sequence (700), and / or a bitstream (800). The coding module (914) may also implement any other method / mechanism described herein. Additionally, the coding module (914) can implement a codec system (200), an encoder (300), a decoder (400), and / or HRD (500). For example, the coding module (914) can be used to implement HRD. Additionally, the coding module (914) can be used to encode parameters into a bitstream to support the HRD conformity check process. Thus, the coding module (914) can be configured to perform a mechanism to solve one or more of the problems discussed above. Thus, the coding module (914) enables the video coding device (900) to provide additional functionality and / or coding efficiency when coding video data. In this way, the coding module (914) not only improves the functionality of the video coding device (900) but also solves problems specific to video coding technology. Furthermore, the coding module (914) transforms the video coding device (900) into different states.Alternatively, the coding module (914) may be implemented as an instruction stored in memory (932) and executed by the processor (930) (e.g., as a computer program product stored on a non-transient medium).

[0184] Memory (932) includes one or more types of memory such as disk, tape drive, solid-state drive, read-only memory (ROM), random access memory (RAM), flash memory, ternary content addressable memory (TCAM), static random access memory (SRAM), etc. Memory (932) may be used as an overflow data storage device and stores a program when such program is selected for execution, and stores instructions and data read during program execution.

[0185] FIG. 10 is a flowchart of an exemplary method (1000) for encoding a video sequence into a bitstream by removing a scalable nested SEI message when no scalable nested SEI message within the scalable nested SEI message refers to a target OLS. This method (1000) may be utilized by an encoder, such as a codec system (200), an encoder (300), and / or a video coding device (900), when performing this method (100). Additionally, this method (1000) may be operated on an HRD (500) and thus perform a conformance test on a multilayer video sequence (600), a multilayer video sequence (700), and / or a bitstream (800).

[0186] This method (1000) may begin when an encoder receives a video sequence and decides to encode the video sequence into a multi-layer bitstream, for example, based on user input. In step 1001, the encoder encodes a bitstream containing one or more OLSs, such as an OLS (625) containing a layer (631), a layer (632), a sublayer (710), a sublayer (720), and / or a sublayer (730). In step 1003, the encoder and / or HRD may perform a sub-bitstream extraction process to extract a target OLS from the OLS.

[0187] In step 1005, the encoder / HRD may remove SEI NAL units from the bitstream. Specifically, the SEI NAL units contain scalable nesting SEI messages. For example, the scalable nesting SEI messages are removed when the scalable nesting SEI message is applied to a specific OLS without any scalable nested SEI messages within the scalable nesting SEI message referencing a target OLS. In a specific example, the scalable nesting SEI message may be applied to specific OLS(s) (e.g., not a specific layer) when the scalable nesting OLS flag is set to 1. Thus, the encoder / HRD may resolve the scalable nesting SEI message for removal from the sub-bitstream when the scalable nesting OLS flag is set to indicate that the scalable nesting SEI message is applied to one or more specific OLSs.

[0188] The encoder / HRD can then examine the scalable nesting SEI messages to determine which of the included scalable nesting SEI messages is associated with which layer of the target OLS. When neither the scalable nesting SEI message nor the scalable nested SEI message refers to the target OLS, the scalable nesting SEI message can be removed from the sub-bitstream. In a specific example, if the scalable nesting SEI message does not contain an index(i) value within the range of 0 or greater and scalable nesting num_olss_minus1 or less, the scalable nested SEI message within the scalable nesting SEI message does not refer to the target OLS, and thus NestingOlsIdx[i] is equal to targetOlsIdx associated with the target OLS. Scalable nesting num_olss_minus1 can specify the number of OLS to which the scalable nesting SEI message applies. Scalable nesting num_olss_minus1 can be restricted to a range of 0 or greater and 'TotalNumOlss minus 1 (TotalNumOlss-1)'. Thus, the encoder / HRD can determine how many OLS are associated with the scalable nesting SEI message by checking the value of scalable nesting num_olss_minus1. The encoder can then check the relevance to the target OLS for each OLS between 0 and scalable nesting num_olss_minus1. The value of scalable nesting num_olss_minus1 can be restricted to 0 or greater and 'Total number of OLS minus 1 (TotalNumOlss-1)'. TotalNumOlss-1 can be included in the VPS specifying OLS in the bitstream / sub-bitstream. NestingOlsIdx[ i ] can specify the OLS index of the i-th OLS to which scalable nested SEI messages are applied when the scalable nesting OLS flag is set to 1.For each current OLS, the encoder / HRD can derive the NestingOlsIdx[i] value to determine the OLS index of the current / i-th OLS to which the scalable nesting SEI message applies by checking the value of scalable nesting ols_idx_delta_minus1. If none of the values ​​of NestingOlsIdx[i] derived from ols_idx_delta_minus1 are identical to the targetOlsIdx associated with the target OLS, none of the scalable nested SEI messages within the scalable nesting SEI message apply to the target OLS. In such cases, the scalable nesting SEI message can be removed from the extracted sub-bitstream without negatively affecting HRD conformance testing or decoding. It should be noted that targetOlsIdx can identify the OLS index of the target OLS, for example, as requested by the decoder.

[0189] In step 1007, the encoder / HRD performs a bitstream conformance test set for the target OLS based on SEI messages that are included in or not removed from the extracted sub-bitstream associated with, for example, the target OLS. In step 1009, the encoder stores the bitstream for communication toward the decoder. The preceding mechanism removes scalable nested SEI messages from the sub-bitstream that do not contain scalable nested SEI messages associated with any layer of the target OLS. This results in a reduction in the size of the sub-bitstream to be sent to the decoder. Thus, this example increases coding efficiency and reduces processor, memory, and / or network resource usage in both the encoder and the decoder.

[0190] FIG. 11 is a flowchart of an exemplary method (1100) for decoding a video sequence from a bitstream from which scalable nesting SEI messages have been removed, when no scalable nesting SEI messages within the scalable nesting SEI messages refer to a target OLS. The remaining SEI messages can be used for bitstream conformance testing in an HRD, such as an HRD (500). This method (1100) may be utilized by a decoder, such as a codec system (200), a decoder (400), and / or a video coding device (900), when performing this method (100). Additionally, this method (1100) may be operated on a bitstream, such as a bitstream (800) containing a multilayer video sequence (600) and / or a multilayer video sequence (700).

[0191] This method (1100) may begin when the decoder begins to receive a bitstream of coded data representing a multilayer video sequence, for example, as a result of method (1000). In step 1101, the decoder may receive a bitstream containing a target OLS. The target OLS may be an OLS such as an OLS (625) containing layer (631), layer (632), sublayer (710), sublayer (720), and / or sublayer (730). Additionally, the decoder may request a target OLS. When no scalable nested SEI message within a scalable nested SEI message refers to the target OLS and when a scalable nested SEI message is applied to a specific OLS, the scalable nested SEI NAL unit containing the scalable nested SEI message is removed from the bitstream before being received by the decoder as part of a sub-bitstream extraction process. The sub-bitstream extraction process can be performed by the HRD on the encoder. For example, a scalable nesting SEI message is applied to a specific OLS (e.g., not a layer) when the scalable nesting OLS flag within the scalable nesting SEI message is set to 1. Therefore, when the scalable nesting OLS flag is set to indicate that the scalable nesting SEI message is applied to one or more specific OLSs, the scalable nesting SEI message can be removed from the sub-bitstream.

[0192] Additionally, if a scalable nesting SEI message or any scalable nesting SEI message does not reference a target OLS, the scalable nesting SEI message may be removed from the sub-bitstream. In a specific example, if a scalable nesting SEI message does not contain any index(i) value within the range of 0 to scalable nesting num_olss_minus1, then no scalable nested SEI message within the scalable nesting SEI message references a target OLS. Thus, NestingOlsIdx[ i ] is equal to targetOlsIdx associated with the target OLS. Scalable nesting num_olss_minus1 may specify the number of OLS to which the scalable nesting SEI message applies. Scalable nesting num_olss_minus1 may be restricted to the range of 0 to 'TotalNumOlss minus 1 (TotalNumOlss-1)'. Therefore, the value of Scalable Nesting num_olss_minus1 can indicate how many OLSs are associated with the Scalable Nesting SEI message. Each OLS between 0 and Scalable Nesting num_olss_minus1 can be checked for relevance to the target OLS. The value of Scalable Nesting num_olss_minus1 can be restricted to a range of 0 or greater and TotalNumOlss-1 or less. TotalNumOlss-1 can be included in the VPS specifying OLSs in the bitstream / sub-bitstream. NestingOlsIdx[ i ] can specify the OLS index of the i-th OLS to which the Scalable Nested SEI message applies when the Scalable Nesting OLS flag is set to 1.For each current OLS, the value of NestingOlsIdx[i] can be derived by examining the value of Scalable Nesting ols_idx_delta_minus1 to determine the OLS index of the current / i-th OLS to which the Scalable Nesting SEI message applies. If the value of NestingOlsIdx[i] derived from ols_idx_delta_minus1 is not the same as targetOlsIdx associated with the target OLS, none of the Scalable Nested SEI messages within the Scalable Nesting SEI message apply to the target OLS. In such cases, the Scalable Nesting SEI message is removed from the extracted sub-bitstream before being received by the decoder. It should be noted that targetOlsIdx can identify the OLS index of the target OLS, for example, as requested by the decoder.

[0193] In step 1103, the decoder can decode a picture from the target OLS. The decoder can also pass the picture for display as part of the video sequence decoded in step 1107.

[0194] FIG. 12 is a schematic diagram of an exemplary system (1200) for coding a video sequence in a bitstream by removing a scalable nested SEI message when no scalable nested SEI message within the scalable nested SEI message refers to a target OLS. This system (1200) may be implemented by an encoder and a decoder such as a codec system (200), an encoder (300), a decoder (400), and / or a video coding device (900). Additionally, this system (1200) may use an HRD (500) to perform a conformity test on a multilayer video sequence (600), a multilayer video sequence (700), and / or a bitstream (800). Additionally, this system (1200) may be used when implementing methods (100, 1000, and / or 1100).

[0195] The system (1200) includes a video encoder (1202). The video encoder (1202) includes an encoding module (1203) for encoding a bitstream containing one or more OLSs. The video encoder (1202) further includes an HRD module (1205) for performing a sub-bitstream extraction process to extract a target OLS from the OLSs. The HRD module (1205) is intended to remove SEI NAL units containing scalable nesting SEI messages from the bitstream when a scalable nesting SEI message is applied to a specific OLS and no scalable nested SEI message within the scalable nesting SEI message refers to the target OLS. The HRD module (1205) is also intended to perform a bitstream conformance test set for the target OLS. The video encoder (1202) further includes a storage module (1206) for storing the bitstream for communication toward a decoder. The video encoder (1202) further includes a transmission module (1207) for transmitting a bitstream toward the video decoder (1210). The video encoder (1202) may be further configured to perform any step of the method (1000).

[0196] The system (1200) also includes a video decoder (1210). The video decoder (1210) includes a receiving module (1211) for receiving a bitstream containing a target OLS, and when no scalable nested SEI message within a scalable nested SEI message refers to the target OLS and when a scalable nested SEI message is applied to a specific OLS, the scalable nested SEI NAL unit containing the scalable nested SEI message is removed from the bitstream before receiving it in the decoder as part of a sub-bitstream extraction process.

[0197] The video decoder (1210) further includes a decoding module (1213) for decoding a picture from a target OLS. The video decoder (1210) further includes a transfer module (1215) for transferring the picture for display as part of a decoded video sequence. The video decoder (1210) may be further configured to perform any step of the method (1100).

[0198] The first component is directly coupled to the second component when there is no intervening component between the first component and the second component, except for a line, trajectory, or other medium. The first component is indirectly coupled to the second component when there is an intervening component other than a line, trajectory, or other medium between the first component and the second component. The term "coupled" and its variations include both directly coupled and indirectly coupled. The use of the term "approximately" implies a range including ±10% of the subsequent number, unless otherwise specified.

[0199] Furthermore, the steps of the exemplary methods described herein do not need to be performed in the order described, and the order of the steps of such methods should be understood as merely exemplary. Likewise, such methods may include additional steps, and specific steps in the methods according to various embodiments of the present disclosure may be omitted or combined.

[0200] While some embodiments have been provided in this disclosure, it will be understood that the disclosed systems and methods may be implemented in many other specific forms without departing from the spirit or scope of this disclosure.

[0201] This example should be considered exemplary and not limiting, and its intent is not limited to the details given herein.

[0202] For example, various components or components may be combined or integrated into other systems, or specific functions may be omitted or not implemented.

[0203] Additionally, the technologies, systems, subsystems, and methods described and illustrated in various embodiments as separate or distinct may be combined with or integrated with other systems, components, technologies, or methods without departing from the scope of this disclosure.

[0204] Changes, substitutions, and other examples of modifications are verifiable by those skilled in the art and may be made without departing from the spirit and scope disclosed herein.

Claims

Claim 1 A method implemented by a decoder, wherein a bitstream containing a target output layer set (OLS) is received by a receiver of the decoder; wherein no scalable nested SEI message within a scalable nesting SEI (supplemental enhancement information) message refers to the target OLS, and when said scalable nesting SEI message is applied to a specific OLS, a scalable nesting SEI NAL (network abstraction layer) unit containing said scalable nesting SEI message is removed from said bitstream as part of a sub-bitstream extraction process, and said scalable nesting SEI message does not contain any index (i) value in the range of 0 or greater and 'scalable nesting OLS number minus 1 (scalable nesting num_olss_minus1)' or less, then no scalable nested SEI message refers to said target OLS, and thus the i-th nesting OLS The index (NestingOlsIdx[ i ]) is identical to the target OLS index (targetOlsIdx) associated with the target OLS, the sub-bitstream extraction process is performed by a hypothetical reference decoder (HRD), the bitstream additionally includes a decoding unit HRD parameter presence flag, the decoding unit HRD parameter presence flag is set to 1 when the HRD operates at the access unit (AU) level or the decoding unit (DU) level, and the decoding unit HRD parameter presence flag is set to 0 when the HRD operates at the AU level, and the operation point (OP) (targetOp) to be tested is,A method comprising the step of selecting a target OLS based on an OP OLS index (opOlsIdx) and the highest OP temporal identifier (opTid), wherein opOlsIdx is a value limited to a range of 0 or greater minus 1 of the total number of OLS (TotalNumOlss); and decoding a picture from the target OLS by a processor of the decoder. Claim 2 A method according to claim 1, wherein the scalable nesting SEI message is applied to a specific OLS when the scalable nesting OLS flag is set to 1. Claim 3 A method according to claim 1, wherein the scalable nesting num_olss_minus1 specifies the number of OLS to which the scalable nesting SEI message is applied, and the value of the scalable nesting num_olss_minus1 is limited to a range of 0 or greater and 1 or less than or equal to the total number of OLS minus 1 (TotalNumOlss-1). Claim 4 In claim 1, the method wherein the targetOlsIdx identifies the OLS index of the target OLS. Claim 5 A method according to claim 1, wherein NestingOlsIdx[ i ] specifies the OLS index of the i-th OLS to which the scalable nested SEI message is applied when the scalable nesting OLS flag is set to 1. Claim 6 A method implemented by an encoder, wherein the processor of the encoder encodes a bitstream comprising one or more sets of output layers (OLS); wherein a Hypothetical Reference Decoder (HRD) operating in the processor performs a sub-bitstream extraction process for extracting a target OLS from the OLS, wherein the bitstream comprises a decoding unit HRD parameter presence flag, the decoding unit HRD parameter presence flag is set to 1 when the HRD operates at the access unit (AU) level or the decoding unit (DU) level, and the decoding unit HRD parameter presence flag is set to 0 when the HRD operates at the AU level; wherein, by the HRD operating in the processor, when no scalable nested SEI message within the scalable nesting SEI message references the target OLS and the scalable nesting SEI message is applied to a specific OLS, a SEI NAL (network abstraction layer) unit containing the scalable nesting SEI message is extracted from the bitstream. Removal step - If the above scalable nesting SEI message does not contain any index(i) value within the range of 0 or greater minus 1 of the scalable nesting OLS number, no scalable nested SEI message refers to the target OLS, and thus the i-th nesting OLS index (NestingOlsIdx[ i ]) is identical to the target OLS index (targetOlsIdx) associated with the target OLS -;A method comprising the step of selecting an operation point (OP) as a target OP (targetOp) based on the target OLS having the OP OLS index (opOlsIdx) and the highest OP temporal identifier (opTid), wherein opOlsIdx is a value limited to a range of 0 or greater minus 1 of the total number of OLS (TotalNumOlss); and the step of performing a bitstream conformance test set on the target OLS by an HRD operating on the processor. Claim 7 In paragraph 6, the scalable nesting SEI message is applied to a specific OLS when the scalable nesting OLS flag is set to 1, a method. Claim 8 In claim 6, the scalable nesting num_olss_minus1 specifies the number of OLS to which the scalable nesting SEI message is applied. Claim 9 A method according to claim 6, wherein the value of the scalable nesting num_olss_minus1 is limited to a range of 0 or greater minus 1 (TotalNumOlss-1) of the total number of OLS. Claim 10 In claim 6, the method wherein the targetOlsIdx identifies the OLS index of the target OLS. Claim 11 In claim 6, the above NestingOlsIdx[ i ] specifies the OLS index of the i-th OLS to which the scalable nested SEI message is applied when the scalable nesting OLS flag is set to 1. Claim 12 A video coding device comprising a processor, a receiver connected to the processor, a memory connected to the processor, and a transmitter connected to the processor, wherein the processor, the receiver, the memory, and the transmitter are configured to perform the method of any one of claims 1 to 11. Claim 13 A non-transient computer-readable medium storing a computer program used by a video coding device, wherein the computer program includes computer-executable instructions stored in the non-transient computer-readable medium, and the computer-executable instructions cause the video coding device to perform the method of any one of claims 1 to 11 when executed by a processor. Claim 14 As a decoder, a receiving means for receiving a bitstream containing a target output layer set (OLS)—in which no scalable nested SEI message within a scalable nesting SEI (supplemental enhancement information) message refers to a target OLS, and when said scalable nesting SEI message is applied to a specific OLS, a scalable nesting SEI NAL (network abstraction layer) unit containing said scalable nesting SEI message is removed from said bitstream as part of a sub-bitstream extraction process, and said scalable nesting SEI message does not contain any index (i) value in the range of 0 or greater and 'scalable nesting OLS number minus 1 (scalable nesting num_olss_minus1)' or less, in which case no scalable nested SEI message refers to a target OLS, and thus the i-th nesting OLS index (NestingOlsIdx[ i ]) is associated with said target OLS It is identical to the target OLS index (targetOlsIdx), and the sub-bitstream extraction process is performed by a hypothetical reference decoder (HRD), and the bitstream additionally includes a decoding unit HRD parameter presence flag, the decoding unit HRD parameter presence flag is set to 1 when the HRD operates at the access unit (AU) level or the decoding unit (DU) level, and the decoding unit HRD parameter presence flag is set to 0 when the HRD operates at the AU level, and the operation point (OP) (targetOp) to be tested is selected based on the target OLS having the OP OLS index (opOlsIdx) and the highest OP temporal identifier (opTid), andopOlsIdx is a value limited to a range of 0 or greater minus 1 of the total number of OLS (TotalNumOlss); a decoding means for decoding a picture from the target OLS; and a decoder comprising a delivery means for delivering the picture for display as part of a decoded video sequence. Claim 15 In paragraph 14, the decoder is further configured to perform the method of any one of paragraphs 2 through 5. Claim 16 As an encoder, an encoding means for encoding a bitstream comprising one or more sets of output layers (OLS); To extract a target OLS from the above OLS, a sub-bitstream extraction process is performed—the bitstream further includes a decoding unit HRD parameter presence flag, which is set to 1 when the HRD operates at the access unit (AU) level or the decoding unit (DU) level, and which is set to 0 when the HRD operates at the AU level—, and when no scalable nested SEI message within the scalable nesting SEI message refers to the target OLS, and when the scalable nesting SEI message is applied to a specific OLS, the SEI NAL (network abstraction layer) unit containing the scalable nesting SEI message is removed from the bitstream—and if the scalable nesting SEI message does not contain any index (i) value within the range of 0 or greater and 'scalable nesting OLS number minus 1 (scalable nesting num_olss_minus1)' or less, any scalable nested The SEI message also does not refer to the target OLS, so that the i-th nesting OLS index (NestingOlsIdx[ i ]) is identical to the target OLS index (targetOlsIdx) associated with the target OLS - , select an operation point (OP) as the target OP (targetOp) based on the target OLS having the OP OLS index (opOlsIdx) and the highest OP temporal identifier (opTid) - opOlsIdx is a value limited to a range of 0 or greater minus 1 of the total number of OLS (TotalNumOlss) - , and a virtual reference decoder (HRD) means for performing a bitstream conformance test set on the target OLS;An encoder comprising a storage means for storing the bitstream for communication with a decoder. Claim 17 In paragraph 16, the encoder is further configured to perform the method of any one of paragraphs 7 through 11. Claim 18 A computer-readable medium comprising a bitstream encoded or decoded by any one of the methods of paragraphs 6 through 11. Claim 19 delete Claim 20 delete Claim 21 delete

Citation Information

Patent Citations

  • Method and apparatus for video coding and decoding

    US20150264404A1