Scalable nesting SEI messaging for OLS
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- HUAWEI TECH CO LTD
- Filing Date
- 2025-11-05
- Publication Date
- 2026-08-04
AI Technical Summary
【0005】 いくつかのビデオコーディングシステムは、SEIメッセージを使用する。SEIメッセージは、デコードされたピクチャ内のサンプルの値を決定するためにデコード処理によって必要とされない情報を含む。例えば、SEIメッセージは、標準に適合しているかどうかについてビットストリームをチェックするために使用されるパラメータを含み得る。仮想参照デコーダ(HRD)は、標準適合性についてビットストリームをチェックする方法を決定するためにSEIメッセージを読み取り得る。このようなシステムは、レイヤに関連するデータと、レイヤを含むOLSに関連するデータとのために、別々のタイプのSEIメッセージを使用し得る。これは、複雑かつ冗長なシステムをもたらす場合がある。本例は、レイヤまたはOLSのいずれかに関連するパラメータを含むように構成されたスケーラブルネスティングSEIメッセージを含む。例えば、スケーラブルネスティングSEIメッセージは、スケーラブルネスティングOLSフラグを含み得、これは、スケーラブルネスティングSEIメッセージがレイヤに関連するパラメータを含む、またはOLSに関連するパラメータを含むかどうかを示すように設定され得る。スケーラブルネスティングSEIメッセージはまた、レイヤまたはOLSに関連する1つまたは複数のスケーラブルネスティングされたSEIメッセージを含み得る。スケーラブルネスティングSEIメッセージがOLSに関する場合、スケーラブルネスティングSEIメッセージはまた、スケーラブルネスティングSEIメッセージに関連付けられたOLSの数を指示するフラグと、OLSをスケーラブルネスティングされたSEIメッセージに関連付けるためのOLSインデックスを指示するフラグとを含む。スケーラブルネスティングSEIメッセージがレイヤに関する場合、スケーラブルネスティングSEIメッセージはまた、スケーラブルネスティングSEIメッセージに関連付けられたレイヤの数を指示するフラグと、レイヤをスケーラブルネスティングされたSEIメッセージに関連付けるためのレイヤ識別子(ID)を指示するフラグとを含む。このようにして、SEIメッセージタイプの数を低減することができ、これは複雑度を減少させ、メッセージタイプの総数を減少させる。これにより、各タイプのメッセージを識別するために使用されるメッセージIDデータの長さが低減される。その結果、コーディング効率が向上し、エンコーダおよびデコーダの両方でのプロセッサ、メモリ、および/またはネットワークシグナリングリソースの使用は低減される。
Smart Images

Figure 0007899509000016 
Figure 0007899509000017 
Figure 0007899509000018
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications This patent application claims the benefit of U.S. Provisional Patent Application No. 62 / 905,143, filed Sep. 24, 2019, by Ye - Kui Wang, titled "Scalable Nesting of SEI Messages for Output Layer Sets", which is incorporated herein by reference.
[0002] This disclosure generally relates to video coding, and more particularly to scalable nesting supplementary enhancement information (SEI) messages used to support encoding layers for output layer sets (OLS) in a multi - layer bitstream.
Background Art
[0003] Even for relatively short videos, the amount of video data required to depict them can be quite large, which can pose difficulties when the data is streamed or communicated over a communication network with limited bandwidth capacity. Thus, video data is generally compressed before being communicated over today's telecommunications networks. The size of the video can also be a problem when the video is stored on a storage device due to potentially limited memory resources. Video compression devices often use software and / or hardware at the source to code the video data before transmission or storage, thereby reducing the amount of data required to represent the digital video image. The compressed data is then received at the destination by a video decompression device that decodes the video data. Given limited network resources and the increasing demand for high video quality, improved compression and decompression techniques that improve the compression ratio without sacrificing much or any of the picture quality are desirable.
Summary of the Invention
Means for Solving the Problems
[0004] In one embodiment, the present disclosure includes a method implemented by a decoder, the receiving step of the decoder receiving a bitstream comprising one or more layers and Scalable Nesting Supplementary Extension Information (SEI) messages, wherein the Scalable Nesting SEI message comprises one or more Scalable Nesting SEI messages and a Scalable Nesting Output Layer Set (OLS) flag, the Scalable Nesting OLS flag being set to specify whether the Scalable Nesting SEI message applies to a particular OLS or a particular layer; and the decoder's processor decoding coded pictures from one or more layers based on the Scalable Nesting SEI messages to produce a decoded picture.
[0005] Some video coding systems use SEI messages. SEI messages contain information that is not required by the decoding process to determine the values of samples in a decoded picture. For example, an SEI message may contain parameters used to check whether a bitstream conforms to a standard. A virtual reference decoder (HRD) may read SEI messages to determine how to check the bitstream for standard conformance. Such systems may use separate types of SEI messages for data related to layers and data related to OLSs that contain layers. This can result in a complex and redundant system. This example includes a scalable nesting SEI message configured to contain parameters related to either layers or OLSs. For example, a scalable nesting SEI message may include a scalable nesting OLS flag, which may be set to indicate whether the scalable nesting SEI message contains parameters related to layers or parameters related to OLSs. A scalable nesting SEI message may also contain one or more scalable nested SEI messages related to layers or OLSs. If a scalable nesting SEI message relates to an OLS, the message also includes a flag indicating the number of OLS associated with the message and a flag indicating the OLS index for associating the OLS with the scalable nested SEI message. If a scalable nesting SEI message relates to a layer, the message also includes a flag indicating the number of layers associated with the message and a flag indicating the layer identifier (ID) for associating the layer with the scalable nested SEI message. In this way, the number of SEI message types can be reduced, which reduces complexity and decreases the total number of message types.This reduces the length of the message ID data used to identify each type of message. As a result, coding efficiency is improved, and the use of processor, memory, and / or network signaling resources in both the encoder and decoder is reduced.
[0006] Optionally, in any of the embodiments described above, another implementation of the embodiment provides that the scalable nesting OLS flag is set to 1 if it specifies that a scalable nesting SEI message applies to a particular OLS, and that the scalable nesting OLS flag is set to 0 if it specifies that a scalable nesting SEI message applies to a particular layer.
[0007] Optionally, in any of the embodiments described above, another implementation of the embodiment provides that the scalable nesting OLS flag is set to 1 if the scalable nesting SEI message includes an SEI message whose payload type is buffering period, picture timing, or decode unit information.
[0008] Optionally, in any of the embodiments described above, another implementation of the embodiment provides that a scalable nesting SEI message includes a syntax element of num_olss_minus1, where the scalable nesting OLS flag is set to 1, the scalable nesting num_olss_minus1 syntax element specifies the number of OLS to which the scalable nesting SEI message applies, and the value of the scalable nesting num_olss_minus1 syntax element is within the range of 0 to the total number of OLS (TotalNumOlss)-1 (including both ends).
[0009] Optionally, in any of the embodiments described above, another implementation of the embodiment provides that a scalable nesting SEI message includes a syntax element of ols_idx_delta_minus1[i], which is the scalable nesting OLS delta minus 1, used to derive a nesting OLS index (NestingOlsIdx[i]) that specifies the OLS index of the i-th OLS to which the scalable nesting SEI message applies, if the scalable nesting OLS flag is equal to 1, and the value of the scalable nesting ols_idx_delta_minus1[i] syntax element is in the range of 0 to TotalNumOlss-2 (inclusive).
[0010] Optionally, in any of the embodiments described above, another implementation of the embodiment further includes the step of deriving NestingOlsIdx[i] as follows:
number
[0011] Optionally, in any of the embodiments described above, another implementation of the embodiment provides that a scalable nesting SEI message includes a syntax element of num_layers_minus1, where num_layers_minus1 is the number of layers to which the scalable nesting SEI message applies, if the scalable nesting OLS flag is set to 0, and the num_layers_minus1 syntax element specifies the number of layers to which the scalable nesting SEI message applies.
[0012] In one embodiment, the present disclosure includes a method implemented by an encoder, comprising: the steps of: encoding a bitstream including one or more layers by the encoder's processor; encoding a scalable nesting SEI message into a bitstream by the processor, the scalable nesting SEI message including one or more scalable nesting SEI messages and a scalable nesting OLS flag, the scalable nesting OLS flag being set to specify whether the scalable nesting SEI message applies to a particular OLS or a particular layer; the steps of the processor performing a set of bitstream conformance tests based on the scalable nesting SEI message; and storing the bitstream for communication to a decoder in a memory coupled to the processor.
[0013] Some video coding systems use SEI messages. SEI messages contain information that is not required by the decoding process to determine the values of samples in a decoded picture. For example, an SEI message may contain parameters used to check a bitstream for conformance to a standard. The HRD may read the SEI message to determine how to check the bitstream for standard conformance. Such systems may use separate types of SEI messages for data related to layers and data related to OLSs containing layers. This can result in a complex and redundant system. This example includes a scalable nesting SEI message configured to contain parameters related to either layers or OLSs. For example, a scalable nesting SEI message may include a scalable nesting OLS flag, which may be set to indicate whether the scalable nesting SEI message contains parameters related to layers or parameters related to OLSs. A scalable nesting SEI message may also contain one or more scalable nested SEI messages related to layers or OLSs. If a scalable nesting SEI message relates to an OLS, the message also includes a flag indicating the number of OLS associated with the message and a flag indicating the OLS index for associating the OLS with the scalable nested SEI message. If a scalable nesting SEI message relates to a layer, the message also includes a flag indicating the number of layers associated with the message and a flag indicating the layer ID for associating the layer with the scalable nested SEI message. In this way, the number of SEI message types can be reduced, which reduces complexity and decreases the total number of message types. This reduces the length of the message ID data used to identify each type of message.As a result, coding efficiency is improved, and the use of processor, memory, and / or network signaling resources in both the encoder and decoder is reduced.
[0014] Optionally, in any of the embodiments described above, another implementation of the embodiment provides that the scalable nesting OLS flag is set to 1 if it specifies that a scalable nesting SEI message applies to a particular OLS, and that the scalable nesting OLS flag is set to 0 if it specifies that a scalable nesting SEI message applies to a particular layer.
[0015] Optionally, in any of the embodiments described above, another implementation of the embodiment provides that the scalable nesting OLS flag is set to 1 if the scalable nesting SEI message includes an SEI message whose payload type is buffering period, picture timing, or decode unit information.
[0016] Optionally, in any of the embodiments described above, another implementation of the embodiment provides that a scalable nesting SEI message includes a scalable nesting num_olss_minus1 syntax element when the scalable nesting OLS flag is set to 1, the scalable nesting num_olss_minus1 syntax element specifies the number of OLS to which the scalable nesting SEI message applies, and the value of the scalable nesting num_olss_minus1 syntax element is within the range of 0 to TotalNumOlss-1 (inclusive).
[0017] Optionally, in any of the embodiments described above, another implementation of the embodiment provides that a scalable nesting SEI message includes a syntax element of scalable nesting ols_idx_delta_minus1[i] used to derive NestingOlsIdx[i] which specifies the OLS index of the i-th OLS to which the scalable nesting SEI message applies, if the scalable nesting OLS flag is equal to 1, and that the value of the scalable nesting ols_idx_delta_minus1[i] syntax element is in the range of 0 to TotalNumOlss-2 (inclusive).
[0018] Optionally, in any of the embodiments described above, another implementation of the embodiment further includes the step of deriving NestingOlsIdx[i] as follows:
number
[0019] Optionally, in any of the embodiments described above, another implementation of the embodiment provides that a scalable nesting SEI message includes a scalable nesting num_layers_minus1 syntax element when the scalable nesting OLS flag is set to 0, and the scalable nesting num_layers_minus1 syntax element specifies the number of layers to which the scalable nesting SEI message applies.
[0020] In one embodiment, the disclosure includes a video coding device comprising a processor, a receiver coupled to the processor, a memory coupled to the processor, and a transmitter coupled to the processor, wherein the processor, receiver, memory, and transmitter are configured to perform the method described in any of the above embodiments.
[0021] In one embodiment, the Disclosure relates to a non-temporary computer-readable medium including a computer program product for use by a video coding device, wherein the computer program product includes computer-executable instructions stored on the non-temporary computer-readable medium, which, when executed by a processor, causes the video coding device to perform the method described in any of the preceding embodiments.
[0022] In one embodiment, the Disclosure provides a receiving means for receiving a bitstream comprising one or more layers and a scalable nesting SEI message, wherein the scalable nesting SEI message comprises one or more scalable nesting SEI messages and a scalable nesting OLS flag, the scalable nesting OLS flag being set to specify whether the scalable nesting SEI message applies to a particular OLS or a particular layer; and a decoder comprising a decoding means for decoding coded pictures from one or more layers based on the scalable nesting SEI message to produce a decoded picture, and a transfer means for transferring the decoded picture for display as part of a decoded video sequence.
[0023] Some video coding systems use SEI messages. SEI messages contain information that is not required by the decoding process to determine the values of samples in a decoded picture. For example, an SEI message may contain parameters used to check a bitstream for conformance to a standard. The HRD may read the SEI message to determine how to check the bitstream for standard conformance. Such systems may use separate types of SEI messages for data related to layers and data related to OLSs containing layers. This can result in a complex and redundant system. This example includes a scalable nesting SEI message configured to contain parameters related to either layers or OLSs. For example, a scalable nesting SEI message may include a scalable nesting OLS flag, which may be set to indicate whether the scalable nesting SEI message contains parameters related to layers or parameters related to OLSs. A scalable nesting SEI message may also contain one or more scalable nested SEI messages related to layers or OLSs. If a scalable nesting SEI message relates to an OLS, the message also includes a flag indicating the number of OLS associated with the message and a flag indicating the OLS index for associating the OLS with the scalable nested SEI message. If a scalable nesting SEI message relates to a layer, the message also includes a flag indicating the number of layers associated with the message and a flag indicating the layer ID for associating the layer with the scalable nested SEI message. In this way, the number of SEI message types can be reduced, which reduces complexity and decreases the total number of message types. This reduces the length of the message ID data used to identify each type of message.As a result, the coding efficiency is improved, and the use of processor, memory, and / or network signaling resources in both the encoder and the decoder is reduced.
[0024] Optionally, in any of the above-described aspects, another implementation of the aspect provides that the decoder is further configured to perform any of the methods of the above-described aspects.
[0025] In one embodiment, the present disclosure provides encoding means for encoding a bitstream including one or more layers, and encoding a scalable nesting SEI message into the bitstream, the scalable nesting SEI message including one or more scalable nested SEI messages and a scalable nesting OLS flag, the scalable nesting OLS flag being set to specify whether the scalable nested SEI message is applied to a specific OLS or a specific layer; HRD means for performing a set of bitstream compliance tests based on the scalable nesting SEI message; and storage means for storing a bitstream for communicating to a decoder, including an encoder.
[0026] Some video coding systems use SEI messages. SEI messages contain information that is not required by the decoding process to determine the values of samples in a decoded picture. For example, an SEI message may contain parameters used to check a bitstream for conformance to a standard. The HRD may read the SEI message to determine how to check the bitstream for standard conformance. Such systems may use separate types of SEI messages for data related to layers and data related to OLSs containing layers. This can result in a complex and redundant system. This example includes a scalable nesting SEI message configured to contain parameters related to either layers or OLSs. For example, a scalable nesting SEI message may include a scalable nesting OLS flag, which may be set to indicate whether the scalable nesting SEI message contains parameters related to layers or parameters related to OLSs. A scalable nesting SEI message may also contain one or more scalable nested SEI messages related to layers or OLSs. If a scalable nesting SEI message relates to an OLS, the message also includes a flag indicating the number of OLS associated with the message and a flag indicating the OLS index for associating the OLS with the scalable nested SEI message. If a scalable nesting SEI message relates to a layer, the message also includes a flag indicating the number of layers associated with the message and a flag indicating the layer ID for associating the layer with the scalable nested SEI message. In this way, the number of SEI message types can be reduced, which reduces complexity and decreases the total number of message types. This reduces the length of the message ID data used to identify each type of message.As a result, the coding efficiency is improved, and the use of processor, memory, and / or network signaling resources in both the encoder and the decoder is reduced.
[0027] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the encoder is further configured to perform the method of any of the foregoing aspects.
[0028] For clarity, any one of the foregoing embodiments can be combined with any one or more of the other foregoing embodiments to create new embodiments within the scope of the present disclosure.
[0029] These and other features will be more clearly understood from the following detailed description in conjunction with the accompanying drawings and the claims.
Brief Description of the Drawings
[0030] To more fully understand the present disclosure, reference is made to the following brief description in connection with the accompanying drawings and detailed description, in which like reference numerals represent like parts.
[0031] [Figure 1] It is a flowchart of an exemplary method for coding a video signal.
[0032] [Figure 2] It is a schematic diagram of an exemplary coding and decoding (codec) system for video coding.
[0033] [Figure 3] It is a schematic diagram showing an exemplary video encoder.
[0034] [Figure 4] It is a schematic diagram showing an exemplary video decoder.
[0035] [Figure 5]This is a schematic diagram illustrating an exemplary virtual reference decoder (HRD).
[0036] [Figure 6] This is a schematic diagram showing an exemplary multilayer video sequence configured for interlayer prediction.
[0037] [Figure 7] This is a schematic diagram showing an example bitstream.
[0038] [Figure 8] This is a schematic diagram illustrating an exemplary video coding device.
[0039] [Figure 9] This is a flowchart illustrating an exemplary method for encoding a video sequence into a bitstream containing scalable nesting SEI messages.
[0040] [Figure 10] This is a flowchart illustrating an exemplary method for decoding a video sequence from a bitstream containing scalable nesting SEI messages.
[0041] [Figure 11] This is a schematic diagram of an exemplary system that codes a video sequence using a bitstream containing scalable nesting SEI messages. [Modes for carrying out the invention]
[0042] First, while exemplary implementations of one or more embodiments are provided below, it should be understood that the disclosed systems and / or methods may be implemented using any number of techniques, whether currently known or existing. This disclosure should not be limited in any way to the exemplary implementations, drawings, and techniques shown below, including the exemplary designs and implementations shown and described herein, and may be modified in the entirety of their equivalents and within the scope of the appended claims.
[0043] The following terms are defined as follows, unless used in the opposite context herein. Specifically, the following definitions are intended to further clarify this disclosure. However, terms may be described differently in different contexts. Therefore, the following definitions should be considered supplementary and not intended to limit any other definitions of such terms provided herein.
[0044] A bitstream is a sequence of bits containing compressed video data for transmission between an encoder and a decoder. An encoder is a device configured to compress video data into a bitstream using an encoding process. A decoder is a device configured to reconstruct video data from the bitstream for display using a decoding process. A picture is an array of luminance samples and / or saturation samples that make up a frame or its fields. For clarity, the picture being encoded or decoded can be referred to as the current picture. A coded picture is a coded representation of a picture, comprising a Video Coding Layer (VCL) Network Abstraction Layer (NAL) unit with a specific value for the NAL unit header layer identifier (nuh_layer_id) within an Access Unit (AU), and containing all coding tree units (CTUs) of the picture. A decoded picture is a picture produced by applying a decoding process to a coded picture. A NAL unit is a syntax structure containing a raw byte sequence payload (RBSP), data in the form of an indication of the data type, scattered as needed with emulation-preventing bytes. A VCL NAL unit is a NAL unit coded to contain video data, such as coded slices of a picture. A non-VCL NAL unit is a NAL unit that contains non-video data, such as syntax and / or parameters that support decoding video data, performing conformance checks, or other operations. A layer is a set of VCL NAL units that share specified characteristics (e.g., common resolution, frame rate, image size, etc.) and associated non-VCL NAL units, as indicated by a layer identifier (ID). The NAL unit header layer identifier (nuh_layer_id) is a syntax element that specifies the identifier of the layer containing the NAL unit. A video parameter set (VPS) is a data unit that contains parameters for the entire video.A coded video sequence is a set of one or more coded pictures. A decoded video sequence is a set of one or more decoded pictures.
[0045] An Output Layer Set (OLS) is a set of layers in which one or more layers are designated as output layers. An output layer is a layer designated for output (e.g., to a display). An OLS index is an index that uniquely identifies a corresponding OLS. A Virtual Reference Decoder (HRD) is a decoder model that operates on an encoder to check the variability of the bitstream produced by the encoding process and verify its conformance to specified constraints. A bitstream conformance test is a test to determine whether the encoded bitstream conforms to a standard such as Versatile Video Coding (VVC). HRD parameters are syntactic elements that initialize and / or define the operating conditions of the HRD. HRD parameters may be included in Supplemental Extension Information (SEI) messages. An SEI message is a syntactic structure with specified semantics that conveys information not required by the decoding process to determine the value of a sample in a decoded picture. A scalable nesting SEI message is a message containing multiple SEI messages corresponding to one or more OLSs or one or more layers. The Buffering Period (BP) SEI message is an SEI message that contains HRD parameters for initializing the HRD to manage coded picture buffers (CPBs). The Picture Timing (PT) SEI message is an SEI message that contains HRD parameters for managing delivery information for access units (AUs) in the CPB and / or decoded picture buffers (DPBs). The Decoded Unit Information (DUI) SEI message is an SEI message that contains HRD parameters for managing delivery information for DUs in the CPB and / or DPBs.
[0046] A scalable nesting SEI message contains a set of scalable nested SEI messages. A scalable nested SEI message is an SEI message nested within a scalable nesting SEI message. A flag is a variable or single-bit syntax element that can take one of two possible values: 0 and 1. The scalable nesting OLS flag is a flag that specifies whether a scalable nested SEI message applies to a particular OLS or a particular layer. The number of scalable nesting OLS minus 1 (num_olss_minus1) is a syntax element that specifies the number of OLS to which a scalable nested SEI message applies. The number of OLS total minus 1 (TotalNumOlss-1) is a syntax element that specifies the total number of OLS specified in the VPS. The scalable nesting OLS delta minus 1 (ols_idx_delta_minus1[i]) is a syntax element containing enough data to derive the nesting OLS index. The nesting OLS index (NestingOlsIdx) is a syntax element that specifies the OLS index of the OLS to which the scalable nested SEI message applies. The number of scalable nesting layers minus 1 (num_layers_minus1) is a syntax element that specifies the number of layers to which the scalable nested SEI message applies. The scalable nesting layer id (layer_id[i]) is a syntax element that specifies the nuh_layer_id value of the i-th layer to which the scalable nested SEI message applies.
[0047] In this specification, the following acronyms are used: Access Unit (AU), Coding Tree Block (CTB), Coding Tree Unit (CTU), Coding Unit (CU), Coding Layer Video Sequence (CLVS), Coding Layer Video Sequence Start (CLVSS), Coded Video Sequence (CVS), Coded Video Sequence Start (CVSS), Joint Video Expert Team (JVET), Motion Constraint Tile Set (MCTS), Maximum Transmission Unit (MTU), Network Abstraction Layer (NAL), Output Layer Set (OLS), Picture Order Count (POC), Random Access Point (RAP), Raw Byte Sequence Payload (RBSP), Sequence Parameter Set (SPS), Video Parameter Set (VPS), and Multipurpose Video Coding (VVC).
[0048] Many video compression techniques can be used to reduce the size of video files while minimizing data loss. For example, video compression techniques may include performing spatial (e.g., intra-picture) prediction and / or temporal (e.g., inter-picture) prediction to reduce or eliminate data redundancy in a video sequence. In block-based video coding, a video slice (e.g., a video picture or part of a video picture) may be divided into video blocks, which may also be called tree blocks, coding tree blocks (CTBs), coding tree units (CTUs), coding units (CUs), and / or coding nodes. Video blocks in an intra-coded (I) slice of a picture are coded using spatial prediction for reference samples in adjacent blocks within the same picture. Video blocks in an intercoded (P) or bidirectional (B) slice of a picture may be coded by using spatial prediction for reference samples in adjacent blocks within the same picture or temporal prediction for reference samples in other reference pictures. A picture may be called a frame and / or image, and a reference picture may be called a reference frame and / or reference image. Spatial or temporal prediction yields predicted blocks representing image blocks. Residual data represents the pixel difference between the original image blocks and the predicted blocks. Thus, the intercoded blocks are encoded according to a motion vector pointing to the reference sample blocks that form the predicted blocks and residual data showing the difference between the coded blocks and the predicted blocks. Intracoded blocks are encoded according to the intracoded mode and residual data. For further compression, the residual data can be transformed from the pixel domain to the transformation domain. These yield residual transformation coefficients that can be quantized. The quantized transformation coefficients may first be placed in a two-dimensional array. The quantized transformation coefficients can be scanned to generate a one-dimensional vector of transformation coefficients. Entropy coding can be applied to achieve even more compression.Such video compression techniques are described in more detail below.
[0049] To ensure that encoded video can be accurately decoded, video is encoded and decoded according to the corresponding video coding standard. Video coding standards include International Telecommunication Union (ITU) Standardization Sector (ITU-T) H.261, International Organization for Standardization / International Electrotechnical Commission (ISO / IEC) Motion Picture Experts Group (MPEG)-1 Part 2, Advanced Video Coding (AVC), also known as ITU-T H.262 or ISO / IEC MPEG-2 Part 2, ITU-T H.263, ISO / IEC MPEG-4 Part 2, ITU-T H.264 or ISO / IEC MPEG-4 Part 10, and High Efficiency Video Coding (HEVC), also known as ITU-T H.265 or MPEG-H Part 2. AVC includes extensions such as Scalable Video Coding (SVC), Multiview Video Coding (MVC), Multiview Video Coding plus Depth (MVC+D), and 3D AVC (3D-AVC). HEVC includes extensions such as Scalable HEVC (SHVC), Multiview HEVC (MV-HEVC), and 3D HEVC (3D-HEVC). The ITU-T and ISO / IEC joint video experts team (JVET) has begun development of a video coding standard called Multipurpose Video Coding (VVC). VVC is included in the Working Draft (WD), which includes JVET-O2001-v14.
[0050] Some video coding systems use Supplemental Expansion Information (SEI) messages. SEI messages contain information that is not required by the decoding process to determine the values of samples in the decoded picture. For example, SEI messages may contain parameters used to check whether a bitstream conforms to a standard. A virtual reference decoder (HRD) may read SEI messages to determine how to check the bitstream for standard conformance. Such systems may use separate types of SEI messages for data related to layers and data related to the output layer set (OLS) containing the layers. This can result in a complex and redundant system.
[0051] This specification discloses scalable nesting SEI messages configured to include parameters related to either layers or OLSs. For example, a scalable nesting SEI message may include a scalable nesting OLS flag, which may be set to indicate whether the scalable nesting SEI message includes parameters related to layers or parameters related to OLSs. A scalable nesting SEI message may also include one or more scalable nested SEI messages related to layers or OLSs. As used herein, one or more indicates any positive number of corresponding items, including one or more such items. If a scalable nesting SEI message relates to OLSs, the scalable nesting SEI message may also include a flag indicating the number of OLSs associated with the scalable nesting SEI message and a flag indicating an OLS index for associating the OLSs with the scalable nested SEI messages. When a scalable nesting SEI message relates to layers, the message also includes a flag indicating the number of layers associated with the message and a flag indicating a layer identifier (ID) for associating the layer with the scalable nesting SEI message. In this way, the number of SEI message types can be reduced, which reduces complexity and decreases the total number of message types. This reduces the length of the message ID data used to identify each type of message. As a result, coding efficiency is improved and the use of processor, memory, and / or network signaling resources in both the encoder and decoder is reduced.
[0052] Figure 1 is a flowchart of an exemplary operation method 100 for coding a video signal. Specifically, the video signal is encoded by an encoder. The encoding process compresses the video signal and reduces the video file size by using various mechanisms. The smaller file size allows the compressed video file to be sent to the user while reducing the associated bandwidth overhead. The decoder then decodes the compressed video file to reconstruct the original video signal for display to the end user. The decoding process generally mirrors the encoding process to allow the decoder to consistently reconstruct the video signal.
[0053] In step 101, a video signal is input to the encoder. For example, the video signal may be an uncompressed video file stored in memory. As another example, the video file may be captured by a video capture device such as a video camera and encoded to support live streaming of the video. The video file may contain both an audio component and a video component. The video component contains a series of image frames that give a visual impression of motion when viewed in sequence. Each frame contains pixels that are represented with respect to light, which is referred to herein as the luminance component (or luminance sample), and color, which is referred to herein as the saturation component (or color sample). In some examples, the frame may also contain depth values to support three-dimensional display.
[0054] In step 103, the video is divided into blocks. This division involves subdividing the pixels within each frame into square and / or rectangular blocks for compression. For example, in High Efficiency Video Coding (HEVC) (also known as H.265 and MPEG-H Part 2), a frame may first be divided into coding tree units (CTUs), which are blocks of a predetermined size (e.g., 64 pixels × 64 pixels). A CTU contains both luminance and saturation samples. Using the coding tree, the CTUs can be divided into blocks, and then the blocks can be recursively subdivided until a configuration supporting further encoding is achieved. For example, the luminance component of a frame may be subdivided until each individual block contains relatively uniform illumination values. Furthermore, the saturation component of a frame may be subdivided until each individual block contains relatively uniform color values. Thus, the division mechanism varies depending on the content of the video frame.
[0055] In stage 105, various compression mechanisms are used to compress the image blocks divided in stage 103. For example, inter-prediction and / or intra-prediction may be used. Inter-prediction is designed to take advantage of the fact that objects in a common scene tend to appear in consecutive frames. Therefore, a block depicting an object in a reference frame does not need to be described repeatedly in adjacent frames. Specifically, an object such as a table may remain in the same position across multiple frames. Thus, a table is described once, and adjacent frames can refer to the reference frame. An object can be matched across multiple frames using a pattern matching mechanism. Furthermore, moving objects can be represented across multiple frames, for example, due to the movement of the object or the movement of the camera. As a particular example, a video may show a car moving across the screen across multiple frames. Such motion can be described using motion vectors. A motion vector is a two-dimensional vector that provides an offset from the coordinates of an object in a frame to the coordinates of an object in a reference frame. Thus, inter-prediction can encode an image block in the current frame as a set of motion vectors indicating the offset from the corresponding block in the reference frame.
[0056] Intra-prediction encodes blocks within a common frame. Intra-prediction leverages the fact that luminance and saturation components tend to cluster within a frame. For example, some green patches in a tree tend to be adjacent to similar green patches. Intra-prediction uses multiple directional prediction modes (e.g., 33 in HEVC), planar mode, and DC mode. Directional modes indicate that the current block is similar to / identical to samples in adjacent blocks in the corresponding direction. Planar mode indicates that a series of blocks along a row / column (e.g., a plane) can be interpolated based on adjacent blocks at the ends of the row. Planar mode actually shows smooth light / color transitions across rows / columns by employing a relatively constant slope to the changing values. DC mode is used for boundary smoothing, indicating that a block is similar to / identical to the mean associated with samples in all adjacent blocks associated with the angular direction of the directional prediction mode. Thus, intra-prediction blocks can represent image blocks as various relational prediction mode values rather than actual values. Furthermore, intra-prediction blocks can represent image blocks as motion vector values rather than actual values. In all cases, the predicted block may not accurately represent the image block in some instances. Any differences are stored in the residual block. Transformations may be applied to the residual block to further compress the file.
[0057] In step 107, various filtering techniques can be applied. In HEVC, filters are applied according to an in-loop filtering scheme. The block-based prediction described above can result in the generation of blocky images in the decoder. Furthermore, the block-based prediction scheme can encode blocks and then reconstruct the encoded blocks for later use as reference blocks. The in-loop filtering scheme repeatedly applies noise suppression filters, deblocking filters, adaptive loop filters, and sample adaptive offset (SAO) filters to blocks / frames. These filters mitigate such blocking artifacts so that the encoded file can be accurately reconstructed. In addition, these filters mitigate artifacts in the reconstructed reference blocks, and as a result, the artifacts are less likely to generate additional artifacts in subsequent blocks encoded based on the reconstructed reference blocks.
[0058] Once the video signal has been split, compressed, and filtered, in step 109 the resulting data is encoded into a bitstream. The bitstream contains the data described above, as well as any signaling data that is desired to support proper video signal reconstruction in the decoder. For example, such data may include split data, prediction data, residual blocks, and various flags that provide coding instructions to the decoder. The bitstream may be stored in memory for transmission to the decoder, as needed. The bitstream may also be broadcast and / or multicast to multiple decoders. The creation of the bitstream is an iterative process. Therefore, steps 101, 103, 105, 107, and 109 may be performed sequentially and / or simultaneously over many frames and blocks. The order shown in Figure 1 is presented for clarity and ease of explanation and is not intended to limit the video coding process to a specific order.
[0059] The decoder receives the bitstream and begins the decoding process in step 111. Specifically, the decoder uses an entropy decoding scheme to convert the bitstream into corresponding syntax and video data. In step 111, the decoder uses the syntax data from the bitstream to determine the division of the frame. The division should match the result of the block division in step 103. Herein lies the entropy encoding / decoding used in step 111. During the compression process, the encoder makes many choices, including selecting a block division scheme from several possible choices based on the spatial location of values in the input image. A number of bins may be used to signal the precise selection. As used herein, a bin is a binary value treated as a variable (e.g., a bit value that can vary depending on the context). Entropy coding allows the encoder to discard any options that are clearly not viable for a particular case, leaving a set of acceptable options. Each acceptable option is then assigned a codeword. The length of the codeword is based on the number of acceptable options (e.g., one bin for two options, two bins for three or four options, etc.). The encoder then encodes the codeword for the selected option. This scheme reduces the size of the codeword so that it is large enough to uniquely indicate a selection from a small subset of acceptable options, as opposed to uniquely indicating a selection from a potentially large set of all possible options. The decoder then decodes the selection by determining the set of acceptable options in a similar manner to the encoder. By determining the set of acceptable options, the decoder can read the codeword and determine the selection made by the encoder.
[0060] In step 113, the decoder performs the decoding of the blocks. Specifically, the decoder uses the inverse transform to generate residual blocks. The decoder then uses the residual blocks and the corresponding prediction blocks to reconstruct the image blocks according to the partitioning. The prediction blocks may include both intra-prediction blocks and inter-prediction blocks, as generated by the encoder in step 105. The reconstructed image blocks are then placed in frames of the video signal reconstructed according to the partitioning data determined in step 111. The syntax of step 113 may also be signaled within the bitstream via entropy coding, as described above.
[0061] In step 115, filtering is performed on the frames of the reconstructed video signal in a manner similar to that in step 107 of the encoder. For example, noise suppression filters, deblocking filters, adaptive loop filters, and SAO filters can be applied to the frames to remove blocking artifacts. Once the frames are filtered, the video signal can be output to a display in step 117 for viewing by the end user.
[0062] Figure 2 is a schematic diagram of an exemplary coding and decoding (codec) system 200 for video coding. Specifically, the codec system 200 provides functions to support the implementation of operation method 100. The codec system 200 is generalized to show the components used in both the encoder and the decoder. The codec system 200 receives and divides the video signal, as described in relation to steps 101 and 103 of operation method 100, resulting in the divided video signal 201. The codec system 200 then compresses the divided video signal 201 into a coded bitstream, as described in relation to steps 105, 107, and 109 of method 100. When operating as a decoder, the codec system 200 generates an output video signal from the bitstream, as described in relation to steps 111, 113, 115, and 117 of operation method 100. The codec system 200 includes a general-purpose coder control component 211, a transform scaling and quantization component 213, an intra-picture estimation component 215, an intra-picture prediction component 217, a motion compensation component 219, a motion estimation component 221, a scaling and inverse transform component 229, a filter control analysis component 227, an in-loop filter component 225, a decoded picture buffer component 223, and a header formatting and context-adaptive binary arithmetic coding (CABAC) component 231. Such components are combined as shown in the figure. In Figure 2, black lines indicate the movement of data to be encoded / decoded, and dashed lines indicate the movement of control data that controls the operation of other components. All components of the codec system 200 may also be present in the encoder. The decoder may contain a subset of the components of the codec system 200.For example, the decoder may include an intrapicture prediction component 217, a motion compensation component 219, a scaling and inverse transform component 229, an in-loop filter component 225, and a decoded picture buffer component 223. These components will now be described.
[0063] The split video signal 201 is a captured video sequence divided into blocks of pixels by a coding tree. The coding tree uses various split modes to subdivide blocks of pixels into smaller blocks of pixels. These blocks can then be further subdivided into even smaller blocks. Blocks are sometimes called nodes on the coding tree. Larger parent nodes are divided into smaller child nodes. The number of times a node is subdivided is called the node / coding tree depth. The divided blocks may, in some cases, be contained within a coding unit (CU). For example, a CU may be a sub-part of a CTU, including luminance blocks, red difference saturation (Cr) blocks, and blue difference saturation (Cb) blocks, along with the corresponding syntax instructions for the CU. Split modes may include binary trees (BT), triple trees (TT), and quad trees (QT), which are used to divide nodes of various shapes into two, three, or four child nodes, depending on the split mode used. The divided video signal 201 is transferred for compression to a general-purpose coder control component 211, a transformation scaling and quantization component 213, an intrapicture estimation component 215, a filter control analysis component 227, and a motion estimation component 221.
[0064] The general-purpose coder control component 211 is configured to make decisions regarding the coding of images in a video sequence into a bitstream, according to application constraints. For example, the general-purpose coder control component 211 manages the optimization of bitrate / bitstream size versus reconstruction quality. Such decisions can be made based on memory space / bandwidth availability and image resolution requirements. The general-purpose coder control component 211 also manages buffer utilization in relation to the transmission rate to mitigate buffer underrun and overrun issues. To manage these issues, the general-purpose coder control component 211 manages partitioning, prediction, and filtering by other components. For example, the general-purpose coder control component 211 may dynamically increase compression complexity to increase resolution and bandwidth usage, or decrease compression complexity to decrease resolution and bandwidth usage. Thus, the general-purpose coder control component 211 controls other components of the codec system 200 to balance video signal reconstruction quality with bitrate concerns. The general-purpose coder control component 211 creates control data that controls the operation of other components. The control data is also transferred to the header formatting and CABAC component 231 so that it is encoded in a bitstream into signal parameters for decoding by the decoder.
[0065] The divided video signal 201 is also sent to the motion estimation component 221 and the motion compensation component 219 for interpretation. A frame or slice of the divided video signal 201 may be divided into multiple video blocks. The motion estimation component 221 and the motion compensation component 219 perform interpretation coding of the received video blocks for one or more blocks in one or more reference frames in order to provide time prediction. The codec system 200 may perform multiple coding passes, for example, to select an appropriate coding mode for each block of video data.
[0066] The motion estimation component 221 and the motion compensation component 219 may be highly integrated, but are shown separately for conceptual purposes. The motion estimation performed by the motion estimation component 221 is the process of generating motion vectors that estimate the motion of video blocks. The motion vectors may, for example, represent the displacement of a coded object relative to a predicted block. A predicted block is a block that is found to be in exact agreement with the block to be coded, with respect to pixel difference. A predicted block is also sometimes called a reference block. Such pixel difference may be determined by the sum of absolute difference (SAD), the sum of square difference (SSD), or other difference metrics. HEVC uses several coded objects, including CTUs, coding tree blocks (CTBs), and CUs. For example, a CTU can be split into CTBs, which can then be split into CBs to be included in CUs. A CU can be encoded as a prediction unit (PU) containing prediction data and / or a transform unit (TU) containing transformed residual data for the CU. The motion estimation component 221 generates motion vectors, PUs, and TUs by using rate distortion analysis as part of the rate distortion optimization process. For example, the motion estimation component 221 can determine multiple reference blocks, multiple motion vectors, etc. for the current block / frame and select the reference blocks, motion vectors, etc. with the best rate distortion characteristics. The best rate distortion characteristics balance both the quality of video reconstruction (e.g., the amount of data loss due to compression) and coding efficiency (e.g., the size of the final encode).
[0067] In some cases, the codec system 200 can calculate the values of sub-integer pixel positions of the reference picture stored in the decoded picture buffer component 223. For example, the video codec system 200 can interpolate the values of quarter-pixel, eighth-pixel, or other fractional pixel positions of the reference picture. Thus, the motion estimation component 221 can perform motion searches on all pixel positions and fractional pixel positions and output motion vectors with fractional pixel precision. The motion estimation component 221 calculates motion vectors for the PU of video blocks in the intercoding slice by comparing the PU positions with the predicted block positions of the reference picture. The motion estimation component 221 outputs the calculated motion vectors as motion data to the header formatting and outputs the CABAC component 231 for encoding and motion to the motion compensation component 219.
[0068] The motion compensation performed by the motion compensation component 219 may include fetching or generating a predicted block based on the motion vector determined by the motion estimation component 221. Again, in some examples, the motion estimation component 221 and the motion compensation component 219 may be functionally integrated. Upon receiving the motion vector of the current video block's PU, the motion compensation component 219 may position itself in the predicted block pointed to by the motion vector. The residual video block is then formed by subtracting the pixel values of the predicted block from the pixel values of the current video block being encoded to form a pixel difference value. Generally, the motion estimation component 221 performs motion estimation for the luminance component, and the motion compensation component 219 uses a motion vector calculated based on the luminance component for both the saturation and luminance components. The predicted and residual blocks are then transferred to the transformation scaling and quantization component 213.
[0069] The split video signal 201 is also sent to the intra-picture estimation component 215 and the intra-picture prediction component 217. Similar to the motion estimation component 221 and the motion compensation component 219, the intra-picture estimation component 215 and the intra-picture prediction component 217 may be highly integrated, but are shown separately for conceptual purposes. The intra-picture estimation component 215 and the intra-picture prediction component 217 intra-predict the current block for the block in the current frame, as an alternative to the inter-prediction performed by the motion estimation component 221 and the motion compensation component 219 between frames, as described above. In particular, the intra-picture estimation component 215 determines the intra-prediction mode to use for encoding the current block. In some examples, the intra-picture estimation component 215 selects an appropriate intra-prediction mode from several tested intra-prediction modes to encode the current block. The selected intra-prediction mode is then forwarded to the header formatting and CABAC component 231 for encoding.
[0070] For example, the intra-picture estimation component 215 calculates rate distortion values using rate distortion analysis of various tested intra-prediction modes and selects the intra-prediction mode with the best rate distortion characteristics among the tested modes. Rate distortion analysis generally determines the amount of distortion (or error) between the encoded block and the original unencoded block encoded to produce the encoded block, and the bit rate (e.g., number of bits) used to produce the encoded block. The intra-picture estimation component 215 calculates ratios from the distortion and rates of various encoded blocks to determine which intra-prediction mode exhibits the best rate distortion value for the block. In addition, the intra-picture estimation component 215 may be configured to code depth blocks of a depth map using a depth modeling mode (DMM) based on rate-distortion optimization (RDO).
[0071] The intra-picture prediction component 217, when implemented on an encoder, can generate residual blocks from prediction blocks based on a selected intra-prediction mode determined by the intra-picture estimation component 215, or, when implemented on a decoder, can read residual blocks from a bitstream. The residual blocks contain the difference in values between the prediction blocks and the original blocks, represented as a matrix. The residual blocks are then transferred to the transformation scaling and quantization component 213. The intra-picture estimation component 215 and the intra-picture prediction component 217 can operate on both luminance and saturation components.
[0072] The transform scaling and quantization component 213 is configured to further compress the residual block. The transform scaling and quantization component 213 applies a transform, such as a discrete cosine transform (DCT), discrete sine transform (DST), or a conceptually similar transform, to the residual block to generate a video block containing residual transform coefficient values. Wavelet transforms, integer transforms, subband transforms, or other types of transforms can also be used. Transforms can convert residual information from a pixel value domain to a transformation domain, such as a frequency domain. The transform scaling and quantization component 213 is also configured to scale the transformed residual information, for example, based on frequency. Such scaling involves applying a scale factor to the residual information so that different frequency information is quantized at different granularities, which can affect the final visual quality of the reconstructed video. The transform scaling and quantization component 213 is also configured to quantize the transform coefficients to further reduce the bitrate. Quantization can reduce the bit depth associated with some or all of the coefficients. The degree of quantization can be changed by adjusting the quantization parameters. In some examples, the transformation scaling and quantization component 213 can then perform a scan of a matrix containing the quantized transformation coefficients. The quantized transformation coefficients are then transferred to the header formatting and CABAC component 231 to be encoded within the bitstream.
[0073] The scaling and inverse transform component 229 applies the inverse operations of the transform scaling and quantization component 213 to support motion estimation. The scaling and inverse transform component 229 reconstructs residual blocks in the pixel domain by applying inverse scaling, transform, and / or quantization for later use as reference blocks that may become predicted blocks for another current block, for example. The motion estimation component 221 and / or motion compensation component 219 can compute the reference blocks by adding the residual blocks back to the corresponding predicted blocks for use in motion estimation of later blocks / frames. Filters are applied to the reconstructed reference blocks to mitigate artifacts generated during scaling, quantization, and transform. Otherwise, such artifacts may cause inaccurate predictions (and generate additional artifacts) when subsequent blocks are predicted.
[0074] The filter-controlled analysis component 227 and the in-loop filter component 225 apply filters to residual blocks and / or reconstructed image blocks. For example, a transformed residual block from the scaling and inverse transform component 229 can be combined with the corresponding predicted block from the intra-picture prediction component 217 and / or the motion compensation component 219 to reconstruct the original image block. Filters can then be applied to the reconstructed image block. In some examples, filters may be applied to residual blocks instead. Like the other components in Figure 2, the filter-controlled analysis component 227 and the in-loop filter component 225 are highly integrated and can be implemented together, but are shown separately for conceptual purposes. Filters applied to reconstructed reference blocks are applied to a specific spatial region and include several parameters to adjust how such filters are applied. The filter-controlled analysis component 227 analyzes the reconstructed reference block to determine where such filters should be applied and sets the corresponding parameters. Such data is transferred to the header formatting and CABAC component 231 as filter-controlled data for encoding. The in-loop filter component 225 applies such filters based on filter control data. These filters may include deblocking filters, noise suppression filters, SAO filters, and adaptive loop filters. Such filters can be applied, as an example, to the spatial / pixel domain (e.g., on a reconstructed pixel block) or the frequency domain.
[0075] When operating as an encoder, filtered and reconstructed image blocks, residual blocks, and / or prediction blocks are stored in the decoded picture buffer component 223 for later use in motion estimation, as described above. When operating as a decoder, the decoded picture buffer component 223 stores the reconstructed and filtered blocks and transfers them to the display as part of the output video signal. The decoded picture buffer component 223 can be any memory device capable of storing prediction blocks, residual blocks, and / or reconstructed image blocks.
[0076] The header formatting and CABAC component 231 receives data from various components of the codec system 200 and encodes such data into a coding bitstream for transmission to the decoder. Specifically, the header formatting and CABAC component 231 generates various headers for encoding control data, such as general control data and filter control data. Furthermore, prediction data, including intra-prediction and motion data, as well as residual data in the form of quantized transformation coefficient data, are all encoded within the bitstream. The final bitstream contains all the information desired by the decoder to reconstruct the original segmented video signal 201. Such information may also include an intra-prediction mode index table (also called a codeword mapping table), definitions of encoding contexts for various blocks, indications of the most likely intra-prediction mode, indications of segmentation information, etc. Such data can be encoded by applying entropy coding. For example, information can be encoded using context-adaptive variable length coding (CAVLC), CABAC, syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or other entropy coding techniques. Following entropy coding, the coded bitstream can be transmitted to another device (e.g., a video decoder) or archived for later transmission or retrieval.
[0077] Figure 3 is a block diagram of an exemplary video encoder 300. The video encoder 300 may be used to implement the encoding function of the codec system 200 and / or to implement stages 101, 103, 105, 107, and / or 109 of the operation method 100. The encoder 300 splits the input video signal, resulting in a split video signal 301 substantially similar to the split video signal 201. The split video signal 301 is then compressed by the components of the encoder 300 and encoded into a bitstream.
[0078] Specifically, the divided video signal 301 is transferred to the intra-picture prediction component 317 for intra-prediction. The intra-picture prediction component 317 may be substantially the same as the intra-picture estimation component 215 and the intra-picture prediction component 217. The divided video signal 301 is also transferred to the motion compensation component 321 for inter-prediction based on the reference block in the decoded picture buffer component 323. The motion compensation component 321 may be substantially the same as the motion estimation component 221 and the motion compensation component 219. The prediction block and residual block from the intra-picture prediction component 317 and the motion compensation component 321 are transferred to the transformation and quantization component 313 for transformation and quantization of the residual block. The transformation and quantization component 313 may be substantially the same as the transformation scaling and quantization component 213. The transformed and quantized residual block and the corresponding prediction block (along with the associated control data) are transferred to the entropy coding component 331 for coding into a bitstream. The entropy coding component 331 may be substantially the same as the header formatting and CABAC component 231.
[0079] The transformed and quantized residual blocks and / or corresponding predicted blocks are also transferred from the transform and quantization component 313 to the inverse transform and quantization component 329 for reconstruction into reference blocks for use by the motion compensation component 321. The inverse transform and quantization component 329 may be substantially the same as the scaling and inverse transform component 229. The in-loop filter in the in-loop filter component 325 is also applied, as an example, to the residual blocks and / or the reconstructed reference blocks. The in-loop filter component 325 may be substantially the same as the filter-controlled analysis component 227 and the in-loop filter component 225. The in-loop filter component 325 may contain multiple filters, as described with respect to the in-loop filter component 225. The filtered blocks are then stored in the decoded picture buffer component 323 for use as reference blocks by the motion compensation component 321. The decoded picture buffer component 323 may be substantially the same as the decoded picture buffer component 223.
[0080] Figure 4 is a block diagram showing an exemplary video decoder 400. The video decoder 400 may be used to implement the decoding function of the codec system 200 and / or to implement steps 111, 113, 115, and / or 117 of the operation method 100. The decoder 400, for example, receives a bitstream from the encoder 300 and generates an output video signal reconstructed based on the bitstream for display to the end user.
[0081] The bitstream is received by the entropy decoding component 433. The entropy decoding component 433 is configured to implement an entropy decoding scheme such as CAVLC, CABAC, SBAC, PIPE coding, or other entropy coding techniques. For example, the entropy decoding component 433 may use header information to provide context for interpreting additional data encoded as codewords in the bitstream. The decoded information includes any desired information for decoding the video signal, such as general control data, filter control data, division information, motion data, prediction data, and quantized transformation coefficients from residual blocks. The quantized transformation coefficients are transferred to the inverse transform and quantization component 429 for reconstruction into residual blocks. The inverse transform and quantization component 429 may be similar to the inverse transform and quantization component 329.
[0082] The reconstructed residual blocks and / or predicted blocks are transferred to the intra-picture prediction component 417 for reconstruction into image blocks based on intra-prediction calculations. The intra-picture prediction component 417 may be similar to the intra-picture estimation component 215 and the intra-picture prediction component 217. Specifically, the intra-picture prediction component 417 uses a prediction mode to identify reference blocks in the frame and applies the residual blocks to reconstruct the resulting intra-predicted image blocks. The reconstructed intra-predicted image blocks and / or residual blocks and the corresponding intra-prediction data are transferred to the decoded picture buffer component 423 via the in-loop filter component 425, which may be substantially similar to the decoded picture buffer component 223 and the in-loop filter component 225, respectively. The in-loop filter component 425 filters the reconstructed image blocks, residual blocks, and / or predicted blocks, and such information is stored in the decoded picture buffer component 423. The reconstructed image blocks from the decoded picture buffer component 423 are transferred to the motion compensation component 421 for interpretation. The motion compensation component 421 may be substantially the same as the motion estimation component 221 and / or motion compensation component 219. Specifically, the motion compensation component 421 generates a prediction block using the motion vector from the reference block and reconstructs the image block by applying the residual block to the result. The resulting reconstructed block may also be transferred to the decoded picture buffer component 423 via the in-loop filter component 425. The decoded picture buffer component 423 continues to store additional reconstructed image blocks that can be reconstructed into frames via the division information. Such frames may also be placed in a sequence. The sequence is output to a display as a reconstructed output video signal.
[0083] Figure 5 is a schematic diagram showing an exemplary HRD500. The HRD500 may be applied in an encoder, for example, a codec system 200 and / or encoder 300. The HRD500 can check the bitstream created in step 109 of method 100 before it is transferred to a decoder, such as decoder 400. In some examples, the bitstream may be transferred sequentially through the HRD500 as the bitstream is encoded. If a portion of the bitstream does not conform to the relevant constraints, the HRD500 can indicate such failure to the encoder, causing the encoder to re-encode the corresponding section of the bitstream using a different mechanism.
[0084] The HRD500 includes a hypothetical stream scheduler (HSS) 541. The HSS 541 is a component configured to run a virtual distribution mechanism. The virtual distribution mechanism is used to check the conformity of a bitstream or decoder with respect to the timing and data flow of the bitstream 551 input to the HRD500. For example, the HSS 541 can receive the bitstream 551 output from an encoder and manage the conformity testing process for the bitstream 551. In a specific example, the HSS 541 can control the rate at which coded pictures move through the HRD500 and verify that the bitstream 551 does not contain non-conforming data.
[0085] HSS541 can transfer bitstream 551 to CPB543 at a predetermined rate. HRD500 can manage the data in the Decode Unit (DU) 553. DU553 is an Access Unit (AU), or a subset of AUs and associated Non-Video Coding Layer (VCL) Network Abstraction Layer (NAL) units. Specifically, an AU contains one or more pictures associated with output time. For example, an AU may contain a single picture in a single-layer bitstream, or per-layer pictures in a multi-layer bitstream. Each picture in an AU may be divided into slices, each contained in a corresponding VCL NAL unit. Thus, DU553 may contain one or more pictures, one or more slices of pictures, or a combination thereof. Parameters used to decode the AUs, pictures, and / or slices may also be contained in non-VCL NAL units. In this way, DU553 contains non-VCL NAL units containing the data necessary to support the decoding of VCL NAL units within DU553. CPB543 is a first-in, first-out buffer in the HRD500. CPB543 contains DU553, which contains video data in the decode order. CPB543 stores video data for use during bitstream conformance verification.
[0086] CPB543 forwards DU553 to decoding component 545. Decoding component 545 is a VVC standard compliant component. For example, decoding component 545 can emulate a decoder 400 used by an end user. Decoding component 545 decodes DU553 at a rate achievable by an exemplary end-user decoder. If decoding component 545 cannot decode DU553 fast enough to prevent overflow of CPB543, bitstream 551 is non-standard and should be re-encoded.
[0087] The decoding processing component 545 decodes DU553 and creates the decoded DU555. The decoded DU555 contains the decoded picture. The decoded DU555 is transferred to DPB547. DPB547 may be substantially the same as the decoded picture buffer components 223, 323, and / or 423. To support interpretation, the picture marked to be used as a reference picture 556 obtained from the decoded DU555 is returned to the decoding processing component 545 to support further decoding. DPB547 outputs the decoded video sequence as a series of pictures 557. Picture 557 is a reconstructed picture that generally mirrors the picture encoded into the bitstream 551 by the encoder.
[0088] Picture 557 is transferred to the output cropping component 549. The output cropping component 549 is configured to apply a conformance cropping window to picture 557. This yields the output trimmed picture 559. The output trimmed picture 559 is a fully reconstructed picture. Thus, the output trimmed picture 559 mimics what an end user would see when decoding the bitstream 551. In this way, the encoder can review the output trimmed picture 559 to ensure that the encoding is satisfactory.
[0089] The HRD500 is initialized based on HRD parameters in bitstream 551. For example, the HRD500 can read HRD parameters from VPS, SPS, and / or SEI messages. The HRD500 can then perform conformity testing operations on bitstream 551 based on the information in such HRD parameters. Specifically, the HRD500 may determine one or more CPB delivery schedules from the HRD parameters. The delivery schedules specify the timing of video data delivery to and from memory locations such as the CPB and / or DPB. Thus, the CPB delivery schedule specifies the timing of AU, DU553, and / or picture delivery to and from the CPB543. Note that the HRD500 can use a DPB delivery schedule for DPB547 similar to the CPB delivery schedule.
[0090] Video can be coded into different layers and / or OLS for use by decoders with varying levels of hardware capabilities and for various network conditions. CPB distribution schedules are selected to reflect these issues. Thus, upper-layer subbitstreams are specified for optimal hardware and network conditions, and therefore, the upper layers may receive one or more CPB distribution schedules that use a large amount of memory in CPB543 and a short delay for the transfer of DU553 toward DPB547. Similarly, lower-layer subbitstreams are specified for limited decoder hardware capabilities and / or poor network conditions. Thus, the lower layers may receive one or more CPB distribution schedules that use a small amount of memory in CPB543 and a longer delay for the transfer of DU553 toward DPB547. The OLS, layers, sublayers, or combinations thereof can then be tested according to the corresponding distribution schedules to ensure that the resulting subbitstreams can be correctly decoded under the conditions expected for those subbitstreams. Therefore, the HRD parameters in bitstream 551 may indicate the CPB delivery schedule and may contain enough data for HRD500 to determine the CPB delivery schedule and correlate the CPB delivery schedule to the corresponding OLS, layer, and / or sublayer.
[0091] Figure 6 is a schematic diagram showing an exemplary multilayer video sequence 600 configured for interlayer prediction 621. The multilayer video sequence 600 may be encoded by an encoder such as codec system 200 and / or encoder 300, for example, according to method 100, and decoded by a decoder such as codec system 200 and / or decoder 400. Furthermore, the multilayer video sequence 600 may be checked for standards compliance by an HRD such as HRD 500. The multilayer video sequence 600 is included to illustrate an exemplary application of layers in a coded video sequence. The multilayer video sequence 600 is any video sequence that uses multiple layers, such as layer N 631 and layer N+1 632.
[0092] In one example, a multilayer video sequence 600 may use interlayer prediction 621. Interlayer prediction 621 is applied between pictures 611, 612, 613, and 614 and pictures 615, 616, 617, and 618 of a different layer. In the example shown, pictures 611, 612, 613, and 614 are part of layer N+1 632, and pictures 615, 616, 617, and 618 are part of layer N 631. Layers such as layer N 631 and / or layer N+1 632 are all groups of pictures associated with similar values of characteristics such as similar size, quality, resolution, signal-to-noise ratio, and capability. A layer can be formally defined as a set of VCL NAL units and associated non-VCL NAL units. A VCL NAL unit is a NAL unit coded to contain video data, such as coded slices of a picture. A non-VCL NAL unit is a NAL unit that contains non-video data such as syntax and / or parameters that support decoding of video data, performing conformance checks, or other operations.
[0093] In the example shown, layer N+1 632 is associated with a larger image size than layer N 631. Therefore, in this example, the picture sizes of pictures 611, 612, 613, and 614 in layer N+1 632 are larger than the picture sizes of pictures 615, 616, 617, and 618 in layer N 631 (e.g., larger height and width, and therefore more samples). However, such pictures may be separated between layer N+1 632 and layer N 631 by other characteristics. Although only two layers, layer N+1 632 and layer N 631, are shown, a set of pictures can be separated into any number of layers based on the relevant characteristics. Layers N+1 632 and N 631 may also be indicated by layer IDs. A layer ID is a data item associated with a picture, indicating that the picture is part of the indicated layer. Therefore, each picture 611-618 can be associated with a corresponding layer ID to indicate which layer N+1 632 or layer N 631 contains the corresponding picture. For example, a layer ID may include a NAL unit header layer identifier (nuh_layer_id), which is a syntax element that specifies the identifier of the layer containing the NAL unit (e.g., slices and / or parameters of the picture within the layer). Layers associated with low quality / bitstream size, such as layer N 631, are generally assigned a lower layer ID and are referred to as lower layers. Furthermore, layers associated with high quality / bitstream size, such as layer N+1 632, are generally assigned a higher layer ID and are referred to as higher layers.
[0094] Pictures 611-618 in different layers 631-632 are configured to be displayed in alternate forms. Thus, pictures in different layers 631-632 may share a time ID and be included in the same AU. A time ID is a data element that indicates the time location of the data within the video sequence. An AU is a set of NAL units related to one particular output time, associated with each other according to a specified classification rule. For example, an AU may contain one or more pictures in different layers, such as picture 611 and picture 615, if they are associated with the same time ID. Specifically, the decoder may decode and display picture 615 at the current display time if a smaller picture is desired, or the decoder may decode and display picture 611 at the current display time if a larger picture is desired. Thus, pictures 611-614 in the upper layer N+1 632 contain substantially the same image data as their corresponding pictures 615-618 in the lower layer N 631 (regardless of differences in picture size). Specifically, picture 611 contains substantially the same image data as picture 615, picture 612 contains substantially the same image data as picture 616, and so on.
[0095] Pictures 611-618 can be coded by referencing other pictures 611-618 in the same layer N 631 or N+1 632. Coding a picture by referencing another picture in the same layer yields an interprediction 623. Interpredictions 623 are indicated by solid arrows. For example, picture 613 can be coded using an interprediction 623 that references one or two of pictures 611, 612, and / or 614 in layer N+1 632, with one picture referenced for a one-way interprediction and / or two pictures referenced for a two-way interprediction. Furthermore, picture 617 can be coded using an interprediction 623 that references one or two of pictures 615, 616, and / or 618 in layer N 631, with one picture referenced for a one-way interprediction and / or two pictures referenced for a two-way interprediction. When performing interprediction 623, if one picture is used as a reference to another picture in the same layer, that picture may be called a reference picture. For example, picture 612 may be a reference picture used to code picture 613 according to interprediction 623. Interprediction 623 may also be called intralayer prediction in a multilayer context. Thus, interprediction 623 is a mechanism for coding a sample of the current picture by referencing an indicated sample in a reference picture that is different from the current picture and is in the same layer as the reference picture and the current picture.
[0096] Pictures 611-618 can also be coded by referencing other pictures 611-618 in different layers. This process is known as inter-layer prediction 621 and is indicated by a dashed arrow. Inter-layer prediction 621 is a mechanism for coding a sample of the current picture by referencing an indicated sample in a reference picture, where the current picture and the reference picture are in different layers and therefore have different layer IDs. For example, a picture in a lower layer N 631 can be used as the reference picture to code the corresponding picture in a higher layer N+1 632. Specifically, picture 611 can be coded by referencing picture 615 according to inter-layer prediction 621. In such a case, picture 615 is used as the inter-layer reference picture. The inter-layer reference picture is the reference picture used in inter-layer prediction 621. In most cases, inter-layer prediction 621 is constrained so that the current picture, such as picture 611, can only use inter-layer reference pictures that are in the same AU and in lower layers, such as picture 615. If multiple layers (e.g., more than two) are available, the inter-layer prediction 621 can encode / decode the current picture based on multiple inter-layer reference pictures that are at a lower level than the current picture.
[0097] The video encoder can encode pictures 611-618 using a multi-layer video sequence 600 via many different combinations and / or permutations of inter-prediction 623 and inter-layer prediction 621. For example, picture 615 may be coded according to intra-prediction. Then, pictures 616-618 may be coded according to inter-prediction 623 by using picture 615 as a reference picture. Furthermore, picture 611 may be coded according to inter-layer prediction 621 by using picture 615 as an inter-layer reference picture. Then, pictures 612-614 may be coded according to inter-prediction 623 by using picture 611 as a reference picture. Thus, a reference picture can function as both a single-layer reference picture and an inter-layer reference picture for different encoding mechanisms. By coding the upper layer N+1 632 picture based on the lower layer N 631 picture, the upper layer N+1 632 can avoid using intra-prediction, which has much lower coding efficiency than inter-prediction 623 and inter-layer prediction 621. Thus, the poor coding efficiency of intra-prediction may be limited to the smallest / lowest quality picture and therefore limited to coding the smallest amount of video data. Pictures used as reference pictures and / or inter-layer reference pictures may be indicated by entries in the reference picture list contained in the reference picture list structure.
[0098] To perform such operations, layers such as layer N 631 and layer N+1 632 may be included in OLS625. OLS625 is a set of layers in which one or more layers are designated as output layers. An output layer is a layer designated for output (e.g., to display). For example, layer N 631 may be included only to support inter-layer prediction 621 and may never be output. In such a case, layer N+1 632 is decoded based on layer N 631 and output. In such a case, OLS625 includes layer N+1 632 as an output layer. In some cases, OLS625 includes only an output layer called a simulcast layer. In other cases, OLS625 may include many layers in different combinations. For example, an output layer in OLS625 may be coded according to an inter-layer prediction 621 based on one, two, or many lower layers. Furthermore, OLS625 may include two or more output layers. Therefore, an OLS625 may include one or more output layers and any supporting layers necessary to reconstruct the output layers. A multilayer video sequence 600 may be coded by using many different OLS625s, each using a different combination of layers. Each OLS625 is associated with an OLS index, which is an index that uniquely identifies the corresponding OLS625.
[0099] Checking multilayer video sequence 600 for standards compliance in HRD500 can be complex depending on the number of layers 631-632 and OLS625. Scalable nesting SEI messages can be used to indicate the parameters required to check layers 631-632 and OLS625 for standards compliance.
[0100] Figure 7 is a schematic diagram showing an exemplary bitstream 700. For example, bitstream 700 may be generated by codec system 200 and / or encoder 300 for decoding by codec system 200 and / or decoder 400 according to method 100. Furthermore, bitstream 700 may include a multilayer video sequence 600. In addition, bitstream 700 may include various parameters for controlling the operation of an HRD such as HRD500. Based on such parameters, the HRD can check bitstream 700 for conformance to a standard before sending it to the decoder for decoding.
[0101] Bitstream 700 includes a VPS 711, one or more SPS 713s, multiple picture parameter sets (PPS) 715s, multiple slice headers 717s, image data 720, and SEI messages 719. The VPS 711 contains data relating to the entire bitstream 700. For example, the VPS 711 may contain data-related OLSs, layers, and / or sublayers used in the bitstream 700. The SPS 713 contains sequence data common to all pictures within the coded video sequences contained in the bitstream 700. For example, each layer may contain one or more coded video sequences, and each coded video sequence may reference the SPS 713 for its corresponding parameters. Parameters within the SPS 713 may include picture sizing, bit depth, coding tool parameters, bitrate limits, etc. While each sequence points to an SPS 713, note that in some examples, a single SPS 713 may contain data for multiple sequences. The PPS 715 contains parameters applicable to the entire picture. Therefore, each picture in a video sequence can be referred to as a PPS715. While each picture refers to a PPS715, it should be noted that in some examples, a single PPS715 can contain data for multiple pictures. For example, multiple similar pictures may be coded according to similar parameters. In such cases, a single PPS715 may contain data for such similar pictures. A PPS715 may indicate the coding tools, quantization parameters, offsets, etc., available for the slices within the corresponding picture.
[0102] The slice header 717 contains parameters specific to each slice within the picture. Therefore, there may be one slice header 717 for each slice in a video sequence. The slice header 717 may include slice type information, POC, reference picture list, prediction weights, tile entry points, deblocking parameters, etc. Note that in some examples, the bitstream 700 may also include a picture header, which is a syntactic structure containing parameters applicable to all slices within a single picture. For this reason, picture headers and slice headers 717 may be used interchangeably in some contexts. For example, certain parameters may be moved between the slice header 717 and the picture header depending on whether such parameters are common to all slices within the picture.
[0103] Image data 720 includes video data encoded according to inter-prediction and / or intra-prediction, as well as corresponding transformed and quantized residual data. For example, image data 720 may include OLS721, layer 723, picture 725, and / or slice 727. OLS721 is a set of layer 723, one or more layers of which are designated as output layers. OLS721 may be substantially the same as OLS625. Layer 723 is a set of VCL NAL units that share a specified characteristic (e.g., common resolution, frame rate, image size, etc.) and associated non-VCL NAL units, as indicated by a layer ID such as nuh_layer_id. For example, layer 723 may include a set of picture 725 that share the same nuh_layer_id. Layer 723 may be substantially the same as layer 631 and / or 632. Picture 725 is an array of luminance samples and / or chroma samples that generate a frame or its field. For example, picture 725 is a coded image that can be output for display or used to support the coding of other picture 725s for output. Picture 725 contains one or more slices 727. A slice 727 may be defined as a sequence of complete coding tree unit (CTU) rows (e.g., within a tile) of picture 725s exclusively contained in an integer number of complete tiles or an integer number of single NAL units. A slice 727 is further divided into CTUs and / or coding tree blocks (CTBs). A CTU is a group of samples of a given size that can be divided by a coding tree. A CTB is a subset of a CTU and contains the luminance or chrominance components of the CTU. A CTU / CTB is further divided into coding blocks based on the coding tree. The coding blocks can then be encoded / decoded according to a prediction mechanism.
[0104] Bitstream 700 may be coded as a series of NAL units. A NAL unit is a container for video data and / or supporting syntax. A NAL unit can be a VCL NAL unit or a non-VCL NAL unit. A VCL NAL unit is a NAL unit coded to contain video data, such as image data 720 and an associated slice header 717. A non-VCL NAL unit is a NAL unit that contains non-video data, such as syntax and / or parameters that support decoding video data, performing conformance checks, or other operations. For example, a non-VCL NAL unit may contain VPS 711, SPS 713, PPS 715, SEI message 719, or other supporting syntax.
[0105] SEI message 719 is a syntactic structure with specified semantics that conveys information not required by the decoding process to determine the values of samples in a decoded picture. For example, SEI message 719 may contain data to support the HRD process, or other supporting data not directly related to the decoding of bitstream 700 in the decoder. SEI message 719 can be a scalable nesting SEI message. A scalable nesting SEI message is a message that contains multiple scalable nested SEI messages corresponding to one or more OLS 721 or one or more layers 723. Thus, a scalable nesting SEI message is an SEI message 719 that contains a set of scalable nested SEI messages of the same type. SEI message 719 may contain a BP SEI message containing HRD parameters for initializing HRD to manage CPB. SEI message 719 may also contain a PT SEI message containing HRD parameters for managing delivery information for AUs in CPB and / or DPB. SEI message 719 may also include a DUI SEI message containing HRD parameters for managing delivery information for DU in CPB and / or DPB.
[0106] Bitstream 700 contains various flags for signaling the structure of SEI message 719. For example, if SEI message 719 is a scalable nesting SEI message, it may contain the scalable nesting (SN)OLS flag 731, the number of scalable nesting OLS minus 1 (num_olss_minus1) 733, the number of scalable nesting OLS delta minus 1 (ols_idx_delta_minus1[i]) 735, the number of scalable nesting layers minus 1 (num_layers_minus1) 737, and / or the scalable nesting layer ID (layer_id[i]) 739.
[0107] The scalable nesting OLS flag 731 is a syntax element that specifies whether a scalable nested SEI message within a scalable nesting SEI message applies to a specific OLS 721 or a specific layer 723. For example, the scalable nesting OLS flag 731 may be set to 1 if a scalable nested SEI message applies to a specific OLS 721 (not a layer). Furthermore, the scalable nesting OLS flag 731 may be set to 0 if a scalable nested SEI message applies to a specific layer 723 (not an OLS). Thus, HRD can read the scalable nesting OLS flag 731 within an SEI message 719 and determine whether all scalable nested SEI messages contained within it describe an OLS 721 or a layer 723.
[0108] The scalable nesting num_olss_minus1 733 is used when an SEI message 719 relates to an OLS721, as indicated by the scalable nesting OLS flag 731. The scalable nesting num_olss_minus1 733 is a syntax element that specifies the number of OLS721 to which the scalable nested SEI messages within the scalable nesting SEI message apply. The scalable nesting num_olss_minus1 733 uses the format minus 1, and therefore includes values one less than the actual value. For example, if a scalable nesting SEI message contains scalable nested SEI messages relating to five OLS721s, the scalable nesting num_olss_minus1 733 would be set to a value of 4.
[0109] The scalable nesting ols_idx_delta_minus1[i]735 is used when an SEI message 719 is associated with OLS721, as indicated by the scalable nesting OLS flag 731. The scalable nesting ols_idx_delta_minus1[i]735 is a syntax element that contains enough data to derive a nesting OLS index. Specifically, the scalable nesting ols_idx_delta_minus1[i]735 contains the OLS index for each scalable nested SEI message within the scalable nesting SEI message. Thus, the scalable nesting ols_idx_delta_minus1[i]735 can be used to correlate scalable nested SEI messages with OLS721. In a concrete example, the nesting OLS index (NestingOlsIdx) for each scalable nested SEI message can be determined using ols_idx_delta_minus1[i]735. NestingOlsIdx is a syntax element that specifies the OLS index of OLS721 to which the corresponding scalable nested SEI message applies. In one example, the variable NestingOlsIdx[i] is derived as follows:
number
[0110] The scalable nesting num_layers_minus1 737 is used when an SEI message 719 relates to layer 723, as indicated by the scalable nesting OLS flag 731. The scalable nesting num_layers_minus1 737 is a syntax element that specifies the number of layers 723 to which the scalable nested SEI message applies within the scalable nesting SEI message. The scalable nesting num_layers_minus1 737 uses the minus-1 format and therefore includes values one less than the actual value. For example, if a scalable nesting SEI message contains scalable nested SEI messages relating to five layers 723, the scalable nesting num_layers_minus1 737 would be set to a value of 4.
[0111] The layer_id[i]739 is used when SEI message 719 is associated with layer 723, as indicated by the scalable nesting OLS flag 731. layer_id[i]739 is a syntax element that specifies the nuh_layer_id value of the i-th layer to which the scalable nested SEI message applies. Thus, layer_id[i]739 can be used to associate each scalable nested SEI message with its corresponding layer 723.
[0112] Therefore, the flags described in bitstream 700 allow the HRD and / or decoder to quickly determine the structure of the SEI message 719. The HRD / decoder may use the scalable nesting OLS flag 731 to determine whether the set of scalable nested messages relates to OLS 721 or layer 723. The HRD / decoder can then determine how to apply the scalable nested message if it relates to OLS 721, by using the scalable nesting num_olss_minus1 733 to determine the number of corresponding OLS 721s and the index of each corresponding OLS 721 using the scalable nesting ols_idx_delta_minus1[i] 735. Furthermore, the HRD / decoder can determine how to apply the scalable nested message if the scalable nested message is related to a layer 723 by using scalable nesting num_layers_minus1 737 to determine the number of corresponding layers 723 and layer_id[i] 739 to determine the index of each corresponding layer 723. This technique reduces the number of SEI message types 719, which reduces complexity and the total number of message types. This reduces the length of the message ID data used to identify each type of message. As a result, coding efficiency is improved and the use of processor, memory, and / or network signaling resources in both the encoder and decoder is reduced.
[0113] The information mentioned above will be explained in more detail below. Layered video coding is also called scalable video coding or video coding with scalability. Scalability in video coding can be supported by using multi-layer coding techniques. A multi-layer bitstream comprises a base layer (BL) and one or more enhancement layers (EL). Examples of scalability include spatial scalability, signal-to-noise ratio (SNR) scalability, multi-view scalability, and frame rate scalability. When multi-layer coding techniques are used, a picture or part of a picture may be coded without using a reference picture (intra-prediction), by referencing a reference picture in the same layer (inter-prediction), and / or by referencing a reference picture in another layer (inter-layer prediction). A reference picture used for inter-layer prediction of the current picture is called an inter-layer reference picture (ILRP). Figure 6 shows an example of multi-layer coding for spatial scalability, where pictures on different layers have different resolutions.
[0114] Several video coding families offer support for scalability in profiles separated from profiles for single-layer coding. Scalable Video Coding (SVC) is a scalable extension of Advanced Video Coding (AVC) that provides support for spatial, temporal, and quality scalability. In SVC, each macroblock (MB) in an EL picture is signaled to indicate whether an EL MB is predicted using collated blocks from lower layers. Predictions from collated blocks may include textures, motion vectors, and / or coding modes. Implementations of SVC may not directly reuse AVC implementations that are not modified in their design. SVC EL macroblock syntax and decoding processes differ from AVC syntax and decoding processes.
[0115] Scalable HEVC (SHVC) is an extension of HEVC that provides support for spatial and quality scalability. Multi-view HEVC (MV-HEVC) is an extension of HEVC that provides support for multi-view scalability. 3D HEVC (3D-HEVC) is an extension of HEVC that provides support for more advanced and efficient 3D video coding than MV-HEVC. Temporal scalability can be included as an integral part of the single-layer HEVC codec. In the multi-layer extension of HEVC, decoded pictures used for inter-layer prediction come from only the same AU and are treated as long-term reference pictures (LTRPs). Such pictures are assigned a reference index in the reference picture list, along with other temporal reference pictures in the current layer. Inter-layer prediction (ILP) is achieved at the prediction unit (PU) level by setting the value of the reference index to refer to the inter-layer reference picture in the reference picture list. Spatial scalability involves resampling a reference picture or a portion of it if the ILRP has a different spatial resolution than the current picture being encoded or decoded. Reference picture resampling can be achieved at either the picture level or the coding block level.
[0116] VVC can also support layered video coding. A VVC bitstream can contain multiple layers. The layers may all be independent of each other. For example, each layer may be coded without using inter-layer prediction. In this case, the layers are also called simulcast layers. In some cases, some of the layers are coded using ILP. Flags in the VPS may indicate whether a layer is a simulcast layer or whether some layers use ILP. If some layers use ILP, the layer dependencies between layers are also signaled in the VPS. Unlike SHVC and MV-HEVC, VVC may not specify an OLS. The OLS contains a specified set of layers, and one or more layers in the set of layers are specified to be output layers. The output layers are the layers of the output OLS. In some implementations of VVC, if a layer is a simulcast layer, only one layer may be selected for decoding and output. In some implementations of VVC, when any layer uses ILP, the entire bitstream containing all layers is specified to be decoded. Furthermore, certain layers are specified as output layers. The output layers may be indicated as only the highest layer, all layers, or the highest layer plus a set of lower layers indicated by the highest layer.
[0117] The aforementioned embodiments involve specific problems. HEVC, including Scalable Extended SHVC and MV-HEVC, may use scalable nesting SEI messages to associate SEI messages with bitstream subsets or specific layers or sublayers corresponding to various operating points. HEVC may also use bitstream partitioning nesting to associate SEI messages with bitstream partitioning in OLS. Bitstream partitioning involves one or more layers of a multilayer bitstream. Each bitstream partitioning nesting SEI message may be contained within a scalable nesting SEI message. This two-level nesting scheme for SEI messages in OLS is complex.
[0118] Generally, this disclosure describes a technique for scalable nesting of SEI messages for output layer sets in a multilayer video bitstream. The description of the technique is based on VVC. However, the technique is also applicable to layered video coding based on other video codec specifications.
[0119] One or more of the above problems can be solved as follows. Specifically, this disclosure includes a method for simple and efficient scalable nesting of SEI messages for OLS in a multilayer video bitstream. Instead of using a two-level nesting scheme, only one nesting SEI message is defined to directly include nesting SEI messages applied to one or more layers in the OLS.
[0120] An exemplary implementation of the aforementioned mechanism is as follows. An exemplary scalable nesting SEI message syntax is as follows. [Table 1-1] [Table 1-2]
[0121] In an alternative example, a flag may be added when nesting_ols_flag is equal to 1. This flag may be set to equal to 1 to indicate that scalable nested SEI messages apply to all OLSs and are applicable to all layers within each OLS. If this flag is set to equal to 1, all syntax elements from after this flag up to nesting_num_seis_minus1 will not be signaled. In another alternative example, a flag may be used and set to equal to 1 to indicate that scalable nested SEI messages apply to all OLSs. If this flag is equal to 1, the syntax elements nesting_num_olss_minus1 and the list of syntax elements nesting_ols_idx_delta_minus1[i] will not be signaled. In yet another alternative example, nesting OLS index values signaled by syntax element nesting_ols_idx_delta_minus1[i] are directly coded instead of delta coded. In another alternative, a flag may be used and set to equal to 1 to indicate that scalable nested SEI messages apply to all layers of the OLS. If this flag is equal to 1, the lists of syntax elements nesting_num_ols_layers_minus1[i] and nesting_ols_layer_idx_delta_minus1[i][j] are not signaled. In yet another alternative, nesting OLS layer index values signaled by syntax elements nesting_ols_layer_idx_delta_minus1[i][j] are directly coded instead of delta coded.
[0122] An exemplary scalable nesting SEI message semantics is as follows:
[0123] Scalable nesting SEI messages provide a mechanism for associating SEI messages with specific layers in a particular OLS context, or with specific layers that are not in an OLS context. A scalable nesting SEI message contains one or more SEI messages. SEI messages contained within a scalable nesting SEI message are also called scalable nested SEI messages. Bitstream compliance may require the following restrictions to apply when including SEI messages within a scalable nesting SEI message.
[0124] SEI messages with a payloadType equal to 132 (decoded picture hash) or 133 (scalable nesting) should not be included in a scalable nesting SEI message. If a scalable nesting SEI message includes a buffering period, picture timing, or decode unit information SEI message, the scalable nesting SEI message should not include any other SEI messages whose payloadType is not equal to 0 (buffering period), 1 (picture timing), or 130 (decode unit information).
[0125] Bitstream compliance may also require the following restrictions to be applied to the nal_unit_type value of an SEI NAL unit containing a scalable nesting SEI message: If a scalable nesting SEI message contains an SEI message with payloadType equal to 0 (buffering period), 1 (picture timing), 130 (decode unit information), 145 (dependent RAP instruction), or 168 (frame field information), the SEI NAL unit containing the scalable nesting SEI message should have a nal_unit_type set equal to PREFIX_SEI_NUT. If a scalable nesting SEI message contains an SEI message with payloadType equal to 132 (decoded picture hash), the SEI NAL unit containing the scalable nesting SEI message should have a nal_unit_type set equal to SUFFIX_SEI_NUT.
[0126] The nesting_ols_flag can be set to equal to 1 to specify that scalable nested SEI messages apply to a specific layer within the context of a particular OLS. The nesting_ols_flag can also be set to equal to 0 to specify that scalable nested SEI messages generally apply to a specific layer (e.g., not within the context of an OLS).
[0127] Bitstream compliance may require the following restrictions to be applied to the value of nesting_ols_flag: If a scalable nesting SEI message contains an SEI message whose payloadType is equal to 0 (buffering period), 1 (picture timing), or 130 (decode unit information), the value of nesting_ols_flag should be equal to 1. If a scalable nesting SEI message contains an SEI message whose payloadType is equal to a value in VclAssociatedSeiList, the value of nesting_ols_flag should be equal to 0.
[0128] The number nesting_num_olss_minus1 plus 1 specifies the number of OLS to which the scalable nested SEI message applies. The value of nesting_num_olss_minus1 should be in the range of 0 to TotalNumOlss-1 (inclusive). nesting_ols_idx_delta_minus1[i] is used to derive the variable NestingOlsIdx[i], which specifies the OLS index of the i-th OLS to which the scalable nested SEI message applies, when nesting_ols_flag is equal to 1. The value of nesting_ols_idx_delta_minus1[i] should be in the range of 0 to TotalNumOlss-2 (inclusive). The variable NestingOlsIdx[i] can be derived as follows:
number
[0129] The number nesting_num_ols_layers_minus1[i] plus 1 specifies the number of layers to which scalable nested SEI messages are applied in the context of the NestingOlsIdx[i]th OLS. The value of nesting_num_ols_layers_minus1[i] should be within the range of 0 to NumLayersInOls[NestingOlsIdx[i]]-1 (including both ends).
[0130] nesting_ols_layer_idx_delta_minus1[i][j] is used to derive the variable NestingOlsLayerIdx[i][j], which specifies the OLS layer index of the j-th layer to which a scalable nested SEI message applies in the context of the NestingOlsIdx[i]-th OLS, when nesting_ols_flag is equal to 1. The value of nesting_ols_layer_idx_delta_minus1[i] should be in the range of 0 to NumLayersInOls[nestingOlsIdx[i]]-2 (inclusive).
[0131] The variable NestingOlsLayerIdx[i][j] can be derived as follows:
number
[0132] The lowest value among all values of LayerIdInOls[NestingOlsIdx[i]][NestingOlsLayerIdx[i][0]] for i within the range 0 to nesting_num_olss_minus1 (including both ends) should be equal to the nuh_layer_id of the current SEI NAL unit (e.g., the SEI NAL unit containing the scalable nesting SEI message). nesting_all_layers_flag may be set to equal to 1 to specify that the scalable nesting SEI message generally applies to all layers that have a nuh_layer_id greater than or equal to the nuh_layer_id of the current SEI NAL unit. nesting_all_layers_flag may be set to equal to 0 to specify that the scalable nesting SEI message generally applies to or does not apply to all layers that have a nuh_layer_id greater than or equal to the nuh_layer_id of the current SEI NAL unit.
[0133] The number nesting_num_layers_minus1 plus 1 specifies the number of layers to which scalable nested SEI messages generally apply. The value of nesting_num_layers_minus1 should be within the range of 0 to vps_max_layers_minus1 - GeneralLayerIdx[nuh_layer_id] (inclusive), where nuh_layer_id is the nuh_layer_id of the current SEI NAL unit. nesting_layer_id[i] specifies the nuh_layer_id value of the i-th layer to which scalable nested SEI messages generally apply, if nesting_all_layers_flag is equal to 0. The value of nesting_layer_id[i] should be greater than nuh_layer_id, where nuh_layer_id is the nuh_layer_id of the current SEI NAL unit.
[0134] If nesting_ols_flag is equal to 1, the variable NestingNumLayers, which specifies the number of layers to which scalable nested SEI messages generally apply, and NestingLayerId[i], a list of i values in the range 0 to NestingNumLayers - 1 (including both ends), which specifies a list of nuh_layer_id values for the layers to which scalable nested SEI messages generally apply, are derived as follows, where nuh_layer_id is the nuh_layer_id of the current SEI NAL unit.
number
[0135] The value of nesting_num_seis_minus1 plus 1 specifies the number of scalable nested SEI messages. The value of nesting_num_seis_minus1 should be in the range of 0 to 63 (inclusive). nesting_zero_bit should be set to equal to 0.
[0136] Figure 8 is a schematic diagram showing an exemplary video coding device 800. The video coding device 800 is suitable for implementing the examples / embodiments disclosed herein. The video coding device 800 comprises a downstream port 820, an upstream port 850, and / or a transceiver unit (Tx / Rx) 810, including a transmitter and / or receiver for communicating data upstream and / or downstream over a network. The video coding device 800 also includes a processor 830, including a logic unit and / or a central processing unit (CPU) for processing data, and a memory 832 for storing data. The video coding device 800 may also include electrical, optical-electrical (OE) components, electrical-optical (EO) components, and / or wireless communication components coupled to the upstream port 850 and / or the downstream port 820 for communicating data over an electrical, optical, or wireless communication network. The video coding device 800 may also include input and / or output (I / O) devices 860 for communicating data with a user. I / O device 860 may include output devices such as a display for showing video data and speakers for outputting audio data. I / O device 860 may also include input devices such as a keyboard, mouse, and trackball, and / or corresponding interfaces for interacting with such output devices.
[0137] The processor 830 is implemented by hardware and software. The processor 830 may be implemented as one or more CPU chips, cores (e.g., as a multi-core processor), field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), and digital signal processors (DSPs). The processor 830 communicates with downstream ports 820, Tx / Rx 810, upstream port 850, and memory 832. The processor 830 includes a coding module 814. The coding module 814 implements the disclosed embodiments described herein, such as methods 100, 900, and 1000, which may use a multi-layer video sequence 600 and / or bitstream 700. The coding module 814 may also implement any other methods / mechanisms described herein. Furthermore, the coding module 814 may implement a codec system 200, an encoder 300, a decoder 400, and / or an HRD 500. For example, the coding module 814 may be used to implement an HRD. Furthermore, the coding module 814 can be used to encode scalable nesting SEI messages having corresponding flags to support clear and concise signaling of scalable nesting SEI messages within scalable nesting SEI messages. Thus, the coding module 814 can be configured to perform a mechanism to address one or more of the problems described above. Thus, the coding module 814 causes the video coding device 800 to provide additional functionality and / or coding efficiency when coding video data. In this way, the coding module 814 improves the functionality of the video coding device 800 and addresses problems specific to video coding techniques. Furthermore, the coding module 814 transforms the video coding device 800 into different states. Alternatively, the coding module 814 can be implemented as an instruction stored in memory 832 (for example, as a computer program product stored on a non-temporary medium) and executed by the processor 830.
[0138] Memory 832 includes one or more memory types, such as disks, tape drives, solid-state drives, read-only memory (ROM), random-access memory (RAM), flash memory, ternary associative memory (TCAM), and static random-access memory (SRAM). Memory 832 is used as an overflow data storage device and can store programs when such programs are selected for execution and can store instructions and data that are read during program execution.
[0139] Figure 9 is a flowchart of an exemplary method 900 for encoding a video sequence into a bitstream, such as bitstream 700 containing scalable nesting SEI messages. Method 900 can be used by an encoder, such as a codec system 200, encoder 300, and / or video coding device 800, when performing Method 100. Furthermore, Method 900 can operate on HRD 500 and therefore can perform conformance testing against a multilayer video sequence 600.
[0140] Method 900 can be initiated when an encoder receives a video sequence and decides to encode the video sequence into a multilayer bitstream, for example, based on user input. In step 901, the encoder encodes the video sequence into one or more layers and encodes the layers into a multilayer bitstream. A layer may include a set of VCL NAL units having the same layer ID and associated non-VCL NAL units. For example, a layer may include a set of VCL NAL units containing video data of an encoded picture and an arbitrary set of parameters used to code such a picture. A layer may be contained in an OLS. For example, an OLS may include an output layer and an arbitrary support layer that can be used to decode the output layer according to interlayer predictions. Thus, an OLS may contain enough data to decode a representation of the video sequence, for example, corresponding image size, SNR, frame rate, etc. Since a video sequence can be coded into several representations, a video sequence may include several layers that are organized into several OLSs as needed. In this way, the encoder may, as needed, select an OLS having the corresponding layer to send to the decoder.
[0141] In step 903, the encoder encodes the SEI message into a bitstream. The SEI message is a syntactic structure that includes data not used for decoding. For example, the SEI message may include data to support conformance testing to ensure that the bitstream conforms to a standard. To support simplified signaling when used with multilayer bitstreams, the SEI message is encoded as a scalable nesting SEI message. A scalable nesting SEI message includes a set of one or more scalable nested SEI messages. Each scalable nested SEI message may apply to one or more OLSs and / or one or more layers. To support simplified signaling, the scalable nesting SEI message includes a scalable nesting OLS flag. The scalable nesting OLS flag may be set to specify whether the scalable nested SEI messages within the scalable nesting SEI message apply to a particular OLS or to a particular layer. For example, the scalable nesting OLS flag may be set to 1 to specify that a scalable nested SEI message applies to a specific / corresponding OLS (not, for example, a layer). As another example, the scalable nesting OLS flag may be set to 0 to specify that a scalable nested SEI message applies to a specific / corresponding layer (not, for example, an OLS). Scalable nesting SEI messages can include several types of scalable nested SEI messages. Specifically, scalable nested SEI messages may include buffering period SEI messages, picture timing SEI messages, and / or decode unit information SEI messages.The Scalable Nesting OLS flag may be set to 1 to indicate that a scalable nesting SEI message applies to a specific OLS (not, for example, a layer) if the scalable nesting SEI message contains any SEI message whose payload type is buffering period, picture timing, or decode unit information.
[0142] A scalable nesting SEI message may include other data indicating how the corresponding scalable nested SEI message should be used by the HRD in the encoder. For example, a scalable nesting SEI message may include a scalable nesting num_olss_minus1 syntax element that specifies the number of OLS to which the corresponding scalable nested SEI message applies. The scalable nesting num_olss_minus1 syntax element may be used when the scalable nesting OLS flag is set to 1 to indicate that the scalable nested SEI message applies to an OLS. The value of the scalable nesting num_olss_minus1 syntax element may be constrained to remain within the range of 0 to TotalNumOlss-1 (inclusive). Similarly, a scalable nesting SEI message may include a scalable nesting num_layers_minus1 specifying the number of layers to which the corresponding scalable nesting SEI message applies, if the scalable nesting OLS flag is set to 0 to indicate that the scalable nesting SEI message applies to a layer.
[0143] A scalable nesting SEI message may also contain a scalable nesting ols_idx_delta_minus1[i] syntax element, which is used to derive a nesting OLS index (NestingOlsIdx[i]) that specifies the OLS index of the i-th OLS to which the scalable nesting SEI message applies, if the scalable nesting OLS flag is equal to 1, indicating that the scalable nesting SEI message applies to an OLS. Specifically, the scalable nesting ols_idx_delta_minus1[i] syntax element can be used to specify the OLS corresponding to each scalable nesting SEI message. In this way, the scalable nesting num_olss_minus1 can be used to determine the number of OLS referenced by a scalable nesting SEI message, and the scalable nesting ols_idx_delta_minus1[i] can be used to correlate each scalable nesting SEI message to its corresponding OLS. The value of the scalable nesting ols_idx_delta_minus1[i] syntax element can be constrained to remain within the range of 0 to TotalNumOlss-2 (including both ends). In a concrete example, NestingOlsIdx[i] is derived as follows:
number
[0144] Similarly, a scalable nesting SEI message may include a scalable nesting layer_id[i] if the scalable nesting OLS flag is equal to 0, indicating that the scalable nesting SEI message applies to a layer. The scalable nesting layer_id[i] specifies the layer ID (e.g., nuh_layer_id) value of the i-th layer to which the scalable nesting SEI message applies.
[0145] In step 905, the HRD operating in the encoder can perform a set of bitstream conformance tests based on the scalable nesting SEI messages. For example, the HRD may read flags within the scalable nesting SEI messages to determine how the scalable nested SEI messages contained within the scalable nesting SEI messages should be interpreted. The HRD may then read the scalable nested SEI messages to determine how to check whether the OLS and / or layers conform to the standards. The HRD can then perform conformance tests on the OLS and / or layers based on the scalable nested SEI messages and / or the corresponding flags within the scalable nesting SEI messages. In step 907, the encoder may, upon request, store the bitstream for communication to the decoder.
[0146] Figure 10 is a flowchart of an exemplary method 1000 for decoding a video sequence from a bitstream, such as bitstream 700 containing scalable nesting SEI messages. Method 1000 can be used by a decoder, such as a codec system 200, a decoder 400, and / or a video coding device 800, when performing Method 100. Furthermore, Method 1000 can be used on a multilayer video sequence 600 whose conformance has been checked by an HRD, such as an HRD 500.
[0147] Method 1000 can begin when the decoder begins receiving a bitstream of coded data representing a multilayer video sequence, for example, as a result of Method 900. In step 1001, the decoder receives a bitstream containing one or more layers. A layer may contain a set of VCL NAL units having the same layer ID and associated non-VCL NAL units. For example, a layer may contain a set of VCL NAL units containing video data of an encoded picture and an arbitrary set of parameters used to code such a picture. A layer may be contained in an OLS. For example, an OLS may contain an output layer and any supporting layers that can be used to decode the output layer according to inter-layer predictions. Thus, an OLS may contain enough data to decode a representation of the video sequence, for example, corresponding image size, SNR, frame rate, etc. Since a video sequence can be coded into several representations, a video sequence may contain several layers that are organized into several OLSs as needed. In this way, the decoder may request and receive a specified OLS having the corresponding layers as needed to decode and display a particular representation of the video sequence.
[0148] A bitstream also contains one or more scalable nesting SEI messages. An SEI message is a syntactic structure that contains data not used for decoding. For example, an SEI message may contain data to support conformance testing to ensure that the bitstream conforms to a standard. To support simplified signaling when used with multilayer bitstreams, SEI messages are coded within scalable nesting SEI messages. A scalable nesting SEI message contains a set of one or more scalable nested SEI messages. Each scalable nested SEI message may apply to one or more OLSs and / or one or more layers, respectively. To support simplified signaling, a scalable nesting SEI message includes a scalable nesting OLS flag. The scalable nesting OLS flag may be set to specify whether the scalable nested SEI messages within the scalable nesting SEI message apply to a particular OLS or to a particular layer. For example, the scalable nesting OLS flag may be set to 1 to specify that a scalable nested SEI message applies to a specific / corresponding OLS (not, for example, a layer). As another example, the scalable nesting OLS flag may be set to 0 to specify that a scalable nested SEI message applies to a specific / corresponding layer (not, for example, an OLS). Scalable nesting SEI messages can include several types of scalable nested SEI messages. Specifically, scalable nested SEI messages may include buffering period SEI messages, picture timing SEI messages, and / or decode unit information SEI messages.The Scalable Nesting OLS flag may be set to 1 to indicate that a scalable nesting SEI message applies to a specific OLS (not, for example, a layer) if the scalable nesting SEI message contains any SEI message whose payload type is buffering period, picture timing, or decode unit information.
[0149] A scalable nesting SEI message may include other data indicating how the corresponding scalable nested SEI message should be used by the HRD in the encoder. For example, a scalable nesting SEI message may include a scalable nesting num_olss_minus1 syntax element that specifies the number of OLS to which the corresponding scalable nested SEI message applies. The scalable nesting num_olss_minus1 syntax element may be used when the scalable nesting OLS flag is set to 1 to indicate that the scalable nested SEI message applies to an OLS. The value of the scalable nesting num_olss_minus1 syntax element may be constrained to remain within the range of 0 to TotalNumOlss-1 (inclusive). Similarly, a scalable nesting SEI message may include a scalable nesting num_layers_minus1 specifying the number of layers to which the corresponding scalable nesting SEI message applies, if the scalable nesting OLS flag is set to 0 to indicate that the scalable nesting SEI message applies to a layer.
[0150] A scalable nesting SEI message may also contain a scalable nesting ols_idx_delta_minus1[i] syntax element, which is used to derive a nesting OLS index (NestingOlsIdx[i]) that specifies the OLS index of the i-th OLS to which the scalable nesting SEI message applies, if the scalable nesting OLS flag is equal to 1, indicating that the scalable nesting SEI message applies to an OLS. Specifically, the scalable nesting ols_idx_delta_minus1[i] syntax element can be used to specify the OLS corresponding to each scalable nesting SEI message. In this way, the scalable nesting num_olss_minus1 can be used to determine the number of OLS referenced by a scalable nesting SEI message, and the scalable nesting ols_idx_delta_minus1[i] can be used to correlate each scalable nesting SEI message to its corresponding OLS. The value of the scalable nesting ols_idx_delta_minus1[i] syntax element can be constrained to remain within the range of 0 to TotalNumOlss-2 (including both ends). In a concrete example, NestingOlsIdx[i] is derived as follows:
number
[0151] Similarly, a scalable nesting SEI message may include a scalable nesting layer_id[i] if the scalable nesting OLS flag is equal to 0, indicating that the scalable nesting SEI message applies to a layer. The scalable nesting layer_id[i] specifies the layer ID (e.g., nuh_layer_id) value of the i-th layer to which the scalable nesting SEI message applies.
[0152] In step 1003, the decoder may decode coded pictures from one or more layers based on scalable nesting SEI messages to produce a decoded picture. For example, the presence of scalable nesting SEI messages may indicate that the bitstream has been checked by HRD in the encoder and is therefore compliant with the standard. Thus, the presence of scalable nesting SEI messages indicates that the bitstream can be decoded. In step 1005, the decoder may transfer the decoded picture for display as part of the decoded video sequence.
[0153] Figure 11 is a schematic diagram of an exemplary system 1100 that codes a video sequence using a bitstream containing scalable nesting SEI messages. System 1100 may be implemented by encoders and decoders such as a codec system 200, encoder 300, decoder 400, and / or video coding device 800. Furthermore, system 1100 can perform conformance testing against a multilayer video sequence 600 and / or bitstream 700 using HRD 500. In addition, system 1100 may be used when implementing methods 100, 900, and / or 1000.
[0154] System 1100 includes a video encoder 1102. The video encoder 1102 includes an encoding module 1103 for encoding a bitstream having one or more layers. The encoding module 1103 further encodes Scalable Nesting Supplemental Extension Information (SEI) messages into the bitstream, each Scalable Nesting SEI message including one or more Scalable Nesting SEI messages and a Scalable Nesting Output Layer Set (OLS) flag, the Scalable Nesting OLS flag being set to specify whether the Scalable Nesting SEI message applies to a particular OLS or a particular layer. The video encoder 1102 further includes an HRD module 1105 for performing a set of bitstream conformance tests based on the Scalable Nesting SEI messages. The video encoder 1102 further includes a storage module 1106 for storing the bitstream for communication toward the decoder. The video encoder 1102 further includes a transmission module 1107 for transmitting the bitstream toward the video decoder 1110. The video encoder 1102 may further be configured to perform any of the steps of method 900.
[0155] System 1100 also includes a video decoder 1110. The video decoder 1110 includes a receive module 1111 that receives a bitstream containing one or more layers and Scalable Nesting Supplemental Extension Information (SEI) messages, the Scalable Nesting SEI messages containing one or more Scalable Nesting SEI messages and a Scalable Nesting Output Layer Set (OLS) flag, the Scalable Nesting OLS flag being set to specify whether the Scalable Nesting SEI message applies to a particular OLS or a particular layer. The video decoder 1110 further includes a decode module 1113 that decodes coded pictures from one or more layers based on the Scalable Nesting SEI messages to produce a decoded picture. The video decoder 1110 further includes a transfer module 1115 for transferring the decoded picture for display as part of the decoded video sequence. The video decoder 1110 may further be configured to perform any of the steps of Method 1000.
[0156] If there are no intermediary components between the first and second components, except for lines, traces, or other media, the first component is directly joined to the second component. If there are intermediary components other than lines, traces, or other media between the first and second components, the first component is indirectly joined to the second component. The term "joined" and its variations include both directly and indirectly joined components. The use of the term "about" means a range including ±10% of the subsequent number unless otherwise specified.
[0157] It should also be understood that the steps of the exemplary methods described herein do not necessarily have to be performed in the order described, and that the order of the steps of such methods is merely illustrative. Similarly, methods consistent with various embodiments of this disclosure may include additional steps, and certain steps may be omitted or combined.
[0158] While several embodiments are provided in this disclosure, it will be understood that the disclosed systems and methods can be embodied in many other specific forms without departing from the spirit or scope of this disclosure. These embodiments should be considered illustrative and non-limiting, and their intent should not be limited to the details given herein. For example, various elements or components may be combined or integrated into other systems, or certain features may be omitted or not implemented.
[0159] In addition, without departing from the scope of this disclosure, the technologies, systems, subsystems, and methods described and illustrated individually or separately in various embodiments may be combined with or integrated with other systems, components, technologies, or methods. Other examples of modifications, substitutions, and changes are readily apparent to those skilled in the art and can be made without departing from the spirit and scope disclosed herein. [Other adjacent items] (Item 1) A method implemented by the decoder, The receiving step of the decoder's receiver receiving a bitstream comprising one or more layers and Scalable Nesting Supplementary Extension Information (SEI) messages, wherein the Scalable Nesting SEI message comprises one or more Scalable Nesting SEI messages and a Scalable Nesting Output Layer Set (OLS) flag, and the Scalable Nesting OLS flag is set to specify whether the Scalable Nesting SEI message applies to a particular OLS or a particular layer; The decoder's processor decodes the coded pictures from one or more layers based on the scalable nested SEI messages to generate a decoded picture. A method for providing this. (Item 2) The method according to item 1, wherein the scalable nesting OLS flag is set to 1 if it specifies that the scalable nesting SEI message applies to a particular OLS, and the scalable nesting OLS flag is set to 0 if it specifies that the scalable nesting SEI message applies to a particular layer. (Item 3) If the scalable nesting SEI message includes an SEI message whose payload type is buffering period, picture timing, or decode unit information, the scalable nesting OLS flag is set to 1, as described in item 1 or 2. (Item 4) The scalable nesting SEI message includes a syntax element of the number obtained by subtracting 1 from the scalable nesting number of OLS (num_olss_minus1) when the scalable nesting OLS flag is set to 1, the scalable nesting num_olss_minus1 syntax element specifies the number of OLS to which the scalable nesting SEI message applies, and the value of the scalable nesting num_olss_minus1 syntax element is in the range of 0 to the total number of OLS (TotalNumOlss) - 1 (including both ends), as described in any of items 1 to 3. (Item 5) The scalable nesting SEI message includes, if the scalable nesting OLS flag is equal to 1, a syntax element of ols_idx_delta_minus1[i] which is the scalable nesting OLS delta minus 1 used to derive a nesting OLS index (NestingOlsIdx[i]) that specifies the OLS index of the i-th OLS to which the scalable nesting SEI message applies, and the value of the scalable nesting ols_idx_delta_minus1[i] syntax element is in the range of 0 to TotalNumOlss-2 (including both ends), as described in any of items 1 to 4. (Item 6) The method described in any of items 1 to 5, further including the step of deriving NestingOlsIdx[i] as follows:
number
number
Claims
1. A method implemented by a decoder, wherein the method is Steps include receiving a bitstream containing one or more layers, Scalable Nesting Supplemental Extension Information (SEI) messages and Scalable Nesting Layer IDs (layer_id[i]), wherein the Scalable Nesting SEI message contains one or more Scalable Nested SEI messages and a Scalable Nesting Output Layer Set (OLS) flag, the Scalable Nesting OLS flag is set to specify whether the Scalable Nested SEI message applies to a particular OLS or a particular layer, and the Scalable Nesting SEI message, when the Scalable Nesting OLS flag is set to 1, the Scalable Nesting OLS number The step includes a `nuh_olss_minus1` syntax element, where the `nuh_olss_minus1` syntax element specifies the number of OLS to which the scalable nesting SEI message applies, the value of the `nuh_olss_minus1` syntax element is in the range of 0 to the total number of OLS (TotalNumOlss) - 1 (including both ends), and the `nuh_layer_id` syntax element specifies the `nuh_layer_id` value of the i-th layer to which the scalable nesting SEI message applies, where the `nuh_layer_id` syntax element specifies the identifier of the layer containing the network abstraction layer unit. The step of deriving a nesting OLS index (NestingOlsIdx[i]) based on the scalable nesting OLS index delta minus 1 (ols_idx_delta_minus1[i]) syntax element included in the scalable nesting SEI message, wherein NestingOlsIdx[i] specifies the OLS index of the i-th OLS to which the scalable nesting SEI message applies when the scalable nesting OLS flag is equal to 1, the value of the scalable nesting ols_idx_delta_minus1[i] syntax element is in the range of 0 to TotalNumOlss-2 (including both ends), and NestingOlsIdx[i] is 【Number 11】 The steps are derived as follows, The steps include: applying the scalable nested SEI message to the OLS specified by the nesting OLS index to decode the coded picture from one or more layers and generate a decoded picture; A method for providing this.
2. If the scalable nesting SEI message is to be applied to a specific OLS, the scalable nesting OLS flag is set to 1; if the scalable nesting SEI message is to be applied to a specific layer, the scalable nesting OLS flag is set to 0. The method according to claim 1.
3. If the scalable nesting SEI message includes an SEI message with a payload type of buffering period, picture timing, or decode unit information, the scalable nesting OLS flag is set to 1. The method according to claim 1 or 2.
4. A method implemented by an encoder, wherein the method is The step of encoding a bitstream that includes one or more layers, The step of encoding a Scalable Nesting Supplemental Extension Information (SEI) message into the bitstream, wherein the Scalable Nesting SEI message includes one or more Scalable Nesting SEI messages and a Scalable Nesting Output Layer Set (OLS) flag, the Scalable Nesting OLS flag is set to specify whether the Scalable Nesting SEI message applies to a particular OLS or a particular layer, and the Scalable Nesting SEI message When the Scalable Nesting OLS flag is set to 1, the Sage includes a Scalable Nesting OLS number minus 1 (num_olss_minus1) syntax element, where the Scalable Nesting num_olss_minus1 syntax element specifies the number of OLS to which the scalable nesting SEI message applies, and the value of the Scalable Nesting num_olss_minus1 syntax element is within the range of 0 to the total number of OLS (TotalNumOlss) minus 1 (including both ends), in steps. A step of encoding a scalable nesting layer ID (layer_id[i]) into the bitstream, wherein the scalable nesting layer ID is a syntax element that specifies the nuh_layer_id value of the i-th layer to which the scalable nested SEI message is applied, and the nuh_layer_id is a syntax element that specifies the identifier of the layer containing the network abstraction layer unit, The step of deriving a nesting OLS index (NestingOlsIdx[i]) based on the scalable nesting OLS index delta minus 1 (ols_idx_delta_minus1[i]) syntax element included in the scalable nesting SEI message, wherein NestingOlsIdx[i] specifies the OLS index of the i-th OLS to which the scalable nesting SEI message applies when the scalable nesting OLS flag is equal to 1, the value of the scalable nesting ols_idx_delta_minus1[i] syntax element is in the range of 0 to TotalNumOlss-2 (including both ends), and NestingOlsIdx[i] is [Math 12] The steps are derived as follows, The steps include applying the scalable nested SEI message to the OLS specified by the nesting OLS index, A step of storing the bitstream in order to communicate it to the decoder. A method for providing this.
5. If the scalable nesting SEI message is to be applied to a specific OLS, the scalable nesting OLS flag is set to 1; if the scalable nesting SEI message is to be applied to a specific layer, the scalable nesting OLS flag is set to 0. The method according to claim 4.
6. If the scalable nesting SEI message includes an SEI message with a payload type of buffering period, picture timing, or decode unit information, the scalable nesting OLS flag is set to 1. The method according to claim 4 or 5.
7. The system comprises a processor, a receiver coupled to the processor, a memory coupled to the processor, and a transmitter coupled to the processor, wherein the processor, the receiver, the memory, and the transmitter are configured to perform the method according to any one of claims 1 to 3. Video coding device.
8. Encoding means for performing the method described in any one of claims 4 to 6, A storage means for storing the bitstream in order to communicate it to the decoder. An encoder equipped with the following features.
9. A method for storing a bitstream, Steps include receiving a bitstream containing one or more layers, Scalable Nesting Supplemental Extension Information (SEI) messages and Scalable Nesting Layer IDs (layer_id[i]), wherein the Scalable Nesting SEI message contains one or more Scalable Nested SEI messages and a Scalable Nesting Output Layer Set (OLS) flag, the Scalable Nesting OLS flag is set to specify whether the Scalable Nested SEI message applies to a particular OLS or a particular layer, and the Scalable Nesting SEI message, when the Scalable Nesting OLS flag is set to 1, the Scalable Nesting OLS number The step includes a `nuh_olss_minus1` syntax element, where the `nuh_olss_minus1` syntax element specifies the number of OLS to which the scalable nesting SEI message applies, the value of the `nuh_olss_minus1` syntax element is in the range of 0 to the total number of OLS (TotalNumOlss) - 1 (including both ends), and the `nuh_layer_id` syntax element specifies the `nuh_layer_id` value of the i-th layer to which the scalable nesting SEI message applies, where the `nuh_layer_id` syntax element specifies the identifier of the layer containing the network abstraction layer unit. The step of storing the bitstream Equipped with, The bitstream is sent to the coding device, The nesting OLS index (NestingOlsIdx[i]) is derived based on the scalable nesting OLS index delta minus 1 (ols_idx_delta_minus1[i]) syntax element included in the scalable nesting SEI message, where NestingOlsIdx[i] specifies the OLS index of the i-th OLS to which the scalable nesting SEI message applies when the scalable nesting OLS flag is equal to 1, the value of the scalable nesting ols_idx_delta_minus1[i] syntax element is within the range of 0 to TotalNumOlss-2 (including both ends), and NestingOlsIdx[i] is [Number 13] It is derived as follows; and, The scalable nested SEI message is applied to the OLS specified by the nesting OLS index; A method of configuration.