Scalable nesting SEI message for ols

The introduction of a scalable nesting SEI message in video coding systems addresses the complexity of existing systems by consolidating parameters for layers and OLS within a single message, enhancing coding efficiency and reducing resource usage.

JP2025087788AActive Publication Date: 2025-06-10HUAWEI TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2025033357
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-09-24
Filing Date
2025-03-04
Publication Date
2025-06-10
Estimated Expiration
2040-09-11

AI Technical Summary

Technical Problem

Existing video coding systems use multiple types of SEI messages for data related to layers and output layer sets (OLS), leading to a complex and redundant system.

Method used

A scalable nesting SEI message is introduced, which includes a flag to specify whether it applies to a specific OLS or layer, reducing the number of SEI message types by including parameters related to either layers or OLS within the same message.

Benefits of technology

This approach simplifies the system, reduces message complexity, and improves coding efficiency by minimizing processor, memory, and network resource usage in both encoders and decoders.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025087788000001_ABST
    Figure 2025087788000001_ABST
Patent Text Reader

Abstract

To provide a decoder and an encoder which improve an SEI message in a multilayer bit stream.SOLUTION: A method for decoding a video sequence includes: receiving a bitstream including one or more layers and a scalable nesting supplemental enhancement information (SEI) message; the scalable nesting SEI message including one or more scalable-nested SEI messages and a scalable nesting output layer set (OLS) flag, which is a set to specify whether the scalable-nested SEI message is to be applied to specific OLSs or specific layers; and the scalable nesting OLS flag decoding a coded picture from the one or more layers to produce a decoded picture.SELECTED DRAWING: Figure 10
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - Reference to Related Applications This patent application claims the benefit of U.S. Provisional Patent Application No. 62 / 905,143, filed Sep. 24, 2019, by Ye - Kui Wang, which is incorporated herein by reference.

[0002] This disclosure generally relates to video coding, and more specifically to scalable nesting supplementary enhancement information (SEI) messages used to support encoding layers for output layer sets (OLS) in a multi - layer bitstream.

Background Art

[0003] Even for relatively short videos, the amount of video data required to depict them can be quite large, which can pose difficulties when the data is streamed or communicated over a communication network with limited bandwidth capacity. Thus, video data is generally compressed before being communicated over today's telecommunications networks. The size of the video can also be a problem when the video is stored on a storage device because memory resources may be limited. Video compression devices often use software and / or hardware at the source to code the video data before transmission or storage, thereby reducing the amount of data required to represent the digital video image. The compressed data is then received at the destination by a video decompression device that decodes the video data. Improved compression and decompression techniques that improve the compression ratio without sacrificing much or any of the picture quality are desirable because network resources are limited and the demand for high video quality is increasing.

Summary of the Invention

Means for Solving the Problems

[0004] In one embodiment, the present disclosure is a method implemented by a decoder, the method comprising: receiving, by a receiver of the decoder, a bitstream including one or more layers and scalable nesting supplementary enhancement information (SEI) messages, wherein the scalable nesting SEI message includes one or more scalable nested SEI messages and a scalable nesting output layer set (OLS) flag, and the scalable nesting OLS flag is set to specify whether the scalable nested SEI message is applicable to a specific OLS or a specific layer; and decoding, by a processor of the decoder, coded pictures from one or more layers based on the scalable nested SEI messages to generate decoded pictures.

[0005] Some video coding systems use SEI messages. SEI messages contain information that is not required by the decoding process to determine the values of samples within the decoded picture. For example, an SEI message may contain parameters used to check the bitstream for compliance with the standard. A Hypothetical Reference Decoder (HRD) may read the SEI message to determine how to check the bitstream for compliance. Such systems may use different types of SEI messages for data related to a layer and data related to an OLS that includes the layer. This can result in a complex and redundant system. This example includes a scalable nesting SEI message configured to include parameters related to either a layer or an OLS. For example, a scalable nesting SEI message may include a scalable nesting OLS flag that can be set to indicate whether the scalable nesting SEI message includes parameters related to a layer or parameters related to an OLS. A scalable nesting SEI message may also include one or more scalable nested SEI messages related to a layer or an OLS. When the scalable nesting SEI message relates to an OLS, the scalable nesting SEI message also includes a flag indicating the number of OLSs associated with the scalable nesting SEI message and a flag indicating an OLS index for associating the OLS with the scalable nested SEI message. When the scalable nesting SEI message relates to a layer, the scalable nesting SEI message also includes a flag indicating the number of layers associated with the scalable nesting SEI message and a flag indicating a layer identifier (ID) for associating the layer with the scalable nested SEI message. In this way, the number of SEI message types can be reduced, which reduces complexity and the total number of message types.This reduces the length of the message ID data used to identify each type of message. As a result, the coding efficiency is improved, and the use of processor, memory, and / or network signaling resources in both the encoder and decoder is reduced.

[0006] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the scalable nesting OLS flag is set to 1 when the scalable nesting SEI message specifies that it is applied to a particular OLS, and the scalable nesting OLS flag is set to 0 when the scalable nesting SEI message specifies that it is applied to a particular layer.

[0007] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that when the scalable nesting SEI message includes an SEI message whose payload type is buffering period, picture timing, or decode unit information, the scalable nesting OLS flag is set to 1.

[0008] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that when the scalable nesting SEI message has the scalable nesting OLS flag set to 1, it includes a syntax element of the number (num_olss_minus1) obtained by subtracting 1 from the scalable nesting number of OLSs, and the scalable nesting num_olss_minus1 syntax element specifies the number of OLSs to which the scalable nesting SEI message is applied, and the value of the scalable nesting num_olss_minus1 syntax element is within the range of 0 to the total number of OLSs (TotalNumOlss) - 1 (including both ends).

[0009] Optionally, in any of the foregoing aspects, another implementation of the aspect includes a syntax element of a number (ols_idx_delta_minus1[i]) obtained by subtracting 1 from a scalable nesting OLS delta used to derive a nesting OLS index (NestingOlsIdx[i]) that specifies the OLS index of the i-th OLS to which a scalable nested SEI message is applied when the scalable nesting OLS flag is equal to 1, and provides that the value of the scalable nesting ols_idx_delta_minus1[i] syntax element is within the range from 0 to TotalNumOlss - 2 (inclusive).

[0010] Optionally, in any of the foregoing aspects, another implementation of the aspect further provides a step of deriving NestingOlsIdx[i] as follows.

Number

[0011] Optionally, in any of the foregoing aspects, another implementation of the aspect includes a syntax element of a number (num_layers_minus1) obtained by subtracting 1 from the number of scalable nestings of the layer when the scalable nesting SEI message has the scalable nesting OLS flag set to 0, and provides that the scalable nesting num_layers_minus1 syntax element specifies the number of layers to which the scalable nested SEI message is applied.

[0012] In one embodiment, the present disclosure is a method implemented by an encoder, the method including: encoding, by a processor of the encoder, a bitstream including one or more layers; encoding, by the processor, a scalable nesting SEI message into the bitstream, the scalable nesting SEI message including one or more scalable nested SEI messages and a scalable nesting OLS flag, the scalable nesting OLS flag being set to specify whether the scalable nested SEI message is applicable to a specific OLS or a specific layer; executing, by the processor, a set of bitstream compliance tests based on the scalable nesting SEI message; and storing, by a memory coupled to the processor, a bitstream for communication to a decoder.

[0013] Some video coding systems use SEI messages. An SEI message contains information that is not required by the decoding process to determine the values of samples within a decoded picture. For example, an SEI message may contain parameters used to check a bitstream for compliance with a standard. HRD may read an SEI message to determine how to check a bitstream for standard compliance. Such systems may use different types of SEI messages for data related to a layer and data related to an OLS that includes the layer. This can result in a complex and redundant system. This example includes a scalable nesting SEI message configured to include parameters related to either a layer or an OLS. For example, a scalable nesting SEI message may include a scalable nesting OLS flag, which can be set to indicate whether the scalable nesting SEI message includes parameters related to a layer or parameters related to an OLS. A scalable nesting SEI message may also include one or more scalable nested SEI messages related to a layer or an OLS. When a scalable nesting SEI message relates to an OLS, the scalable nesting SEI message also includes a flag indicating the number of OLSs associated with the scalable nesting SEI message and a flag indicating an OLS index for associating the OLS with the scalable nested SEI message. When a scalable nesting SEI message relates to a layer, the scalable nesting SEI message also includes a flag indicating the number of layers associated with the scalable nesting SEI message and a flag indicating a layer ID for associating the layer with the scalable nested SEI message. In this way, the number of SEI message types can be reduced, which reduces complexity and the total number of message types. This reduces the length of the message ID data used to identify each type of message.As a result, the coding efficiency is improved, and the use of processor, memory, and / or network signaling resources in both the encoder and the decoder is reduced.

[0014] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the scalable nesting OLS flag is set to 1 when the scalable nesting SEI message specifies that it is applied to a specific OLS, and the scalable nesting OLS flag is set to 0 when the scalable nesting SEI message specifies that it is applied to a specific layer.

[0015] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the scalable nesting OLS flag is set to 1 when the scalable nesting SEI message includes an SEI message whose payload type is buffering period, picture timing, or decode unit information.

[0016] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the scalable nesting SEI message includes a scalable nesting num_olss_minus1 syntax element when the scalable nesting OLS flag is set to 1, the scalable nesting num_olss_minus1 syntax element specifies the number of OLSs to which the scalable nesting SEI message is applied, and the value of the scalable nesting num_olss_minus1 syntax element is within the range of 0 to TotalNumOlss - 1 (inclusive).

[0017] Optionally, in any of the foregoing aspects, another implementation of the aspect includes a syntax element of scalable nesting ols_idx_delta_minus1[i] used to derive NestingOlsIdx[i] that specifies the OLS index of the i-th OLS to which the scalable nested SEI message is applied when the scalable nesting OLS flag is equal to 1, and provides that the value of the scalable nesting ols_idx_delta_minus1[i] syntax element is within the range from 0 to TotalNumOlss - 2 (including both ends).

[0018] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that it further includes the step of deriving NestingOlsIdx[i] as follows.

Number

[0019] Optionally, in any of the foregoing aspects, another implementation of the aspect includes a syntax element of scalable nesting num_layers_minus1 when the scalable nesting SEI message and the scalable nesting OLS flag are set to 0, and provides that the scalable nesting num_layers_minus1 syntax element specifies the number of layers to which the scalable nested SEI message is applied.

[0020] In one embodiment, the present disclosure includes a video coding device comprising a processor, a receiver coupled to the processor, a memory coupled to the processor, and a transmitter coupled to the processor, wherein the processor, the receiver, the memory, and the transmitter are configured to execute the method described in any of the foregoing aspects.

[0021] In one embodiment, the present disclosure is a non-transitory computer-readable medium including a computer program product for use by a video coding device, the computer program product including computer-executable instructions stored in the non-transitory computer-readable medium such that, when executed by a processor, cause the video coding device to execute the method according to any of the foregoing aspects.

[0022] In one embodiment, the present disclosure includes a receiving means for receiving a bitstream including one or more layers and a scalable nesting SEI message, the scalable nesting SEI message including one or more scalable nested SEI messages and a scalable nesting OLS flag, the scalable nesting OLS flag being set to specify whether the scalable nested SEI message is applied to a specific OLS or a specific layer, a decoding means for decoding coded pictures from one or more layers based on the scalable nested SEI message to generate decoded pictures, and a transferring means for transferring the decoded pictures for display as part of the decoded video sequence.

[0023] Some video coding systems use SEI messages. An SEI message contains information that is not required by the decoding process to determine the values of samples within the decoded picture. For example, an SEI message may contain parameters used to check the bitstream for compliance with the standard. HRD may read the SEI message to determine how to check the bitstream for compliance. Such systems may use different types of SEI messages for data related to a layer and data related to the OLS containing the layer. This can result in a complex and redundant system. This example includes a scalable nesting SEI message configured to include parameters related to either a layer or an OLS. For example, a scalable nesting SEI message may include a scalable nesting OLS flag, which can be set to indicate whether the scalable nesting SEI message contains parameters related to a layer or parameters related to an OLS. A scalable nesting SEI message may also include one or more scalable nested SEI messages related to a layer or an OLS. When the scalable nesting SEI message relates to an OLS, the scalable nesting SEI message also includes a flag indicating the number of OLSs associated with the scalable nesting SEI message and a flag indicating an OLS index for associating the OLS with the scalable nested SEI message. When the scalable nesting SEI message relates to a layer, the scalable nesting SEI message also includes a flag indicating the number of layers associated with the scalable nesting SEI message and a flag indicating a layer ID for associating the layer with the scalable nested SEI message. In this way, the number of SEI message types can be reduced, which reduces complexity and the total number of message types. This reduces the length of the message ID data used to identify each type of message.As a result, the coding efficiency is improved, and the use of processor, memory, and / or network signaling resources in both the encoder and the decoder is reduced.

[0024] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the decoder is further configured to perform any of the methods of the foregoing aspects.

[0025] In one embodiment, the present disclosure provides encoding means for encoding a bitstream including one or more layers, and encoding a scalable nesting SEI message into the bitstream, the scalable nesting SEI message including one or more scalable nested SEI messages and a scalable nesting OLS flag, the scalable nesting OLS flag being set to specify whether the scalable nested SEI message is applied to a specific OLS or a specific layer; HRD means for performing a set of bitstream compliance tests based on the scalable nesting SEI message; and storage means for storing a bitstream for communication to a decoder, including an encoder.

[0026] Some video coding systems use SEI messages. SEI messages contain information that is not required by the decoding process to determine the values of samples within a decoded picture. For example, an SEI message may contain parameters used to check a bitstream for compliance with a standard. HRD may read an SEI message to determine how to check a bitstream for standard compliance. Such systems may use different types of SEI messages for data related to a layer and data related to an OLS that includes the layer. This can result in a complex and redundant system. This example includes a scalable nesting SEI message configured to include parameters related to either a layer or an OLS. For example, a scalable nesting SEI message may include a scalable nesting OLS flag, which can be set to indicate whether the scalable nesting SEI message includes parameters related to a layer or parameters related to an OLS. A scalable nesting SEI message may also include one or more scalable nested SEI messages related to a layer or an OLS. When a scalable nesting SEI message relates to an OLS, the scalable nesting SEI message also includes a flag indicating the number of OLSs associated with the scalable nesting SEI message and a flag indicating an OLS index for associating the OLS with the scalable nested SEI message. When a scalable nesting SEI message relates to a layer, the scalable nesting SEI message also includes a flag indicating the number of layers associated with the scalable nesting SEI message and a flag indicating a layer ID for associating the layer with the scalable nested SEI message. In this way, the number of SEI message types can be reduced, which reduces complexity and the total number of message types. This reduces the length of the message ID data used to identify each type of message.As a result, the coding efficiency is improved, and the use of processor, memory, and / or network signaling resources in both the encoder and the decoder is reduced.

[0027] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the encoder is further configured to perform the method of any of the foregoing aspects.

[0028] For clarity, any one of the foregoing embodiments can be combined with any one or more of the other foregoing embodiments to create new embodiments within the scope of the present disclosure.

[0029] These and other features will be more clearly understood from the following detailed description in conjunction with the accompanying drawings and the claims.

Brief Description of the Drawings

[0030] To understand the present disclosure more fully, reference is made to the following brief description in connection with the accompanying drawings and detailed description, in which like reference numerals represent like parts.

[0031]

Figure 1

[0032]

Figure 2

[0033]

Figure 3

[0034]

Figure 4

[0035]

Figure 5

[0036]

Figure 6

[0037]

Figure 7

[0038]

Figure 8

[0039]

Figure 9

[0040]

Figure 10

[0041]

Figure 11

Best Mode for Carrying Out the Invention

[0042] Initially, an exemplary implementation of one or more embodiments is provided below, but it should be understood that the disclosed system and / or method may be implemented using any number of techniques, whether currently known or existing. The present disclosure should in no way be limited to the exemplary implementations, drawings, and techniques shown below, including the exemplary designs and implementations shown and described herein, but may be modified within the scope of the appended claims together with the full scope of their equivalents.

[0043] The following terms are defined as follows, unless used in the opposite context in this specification. Specifically, the following definitions are intended to further clarify the present disclosure. However, the terms may be described differently in different contexts. Accordingly, the following definitions should be regarded as supplementary and should not be regarded as limiting any other definitions provided herein for such terms.

[0044] A bitstream is a sequence of bits that contains compressed video data for transmission between an encoder and a decoder. An encoder is a device configured to compress video data into a bitstream using an encoding process. A decoder is a device configured to reconstruct video data from a bitstream for display using a decoding process. A picture is an array of luminance samples and / or an array of chrominance samples that generates a frame or a field thereof. For clarity of explanation, a picture being encoded or decoded can be referred to as the current picture. A coded picture is a coded representation of a picture that comprises a video coding layer (VCL) network abstraction layer (NAL) unit having a specific value of a NAL unit header layer identifier (nuh_layer_id) within an access unit (AU) and includes all coding tree units (CTUs) of the picture. A decoded picture is a picture generated by applying a decoding process to a coded picture. A NAL unit is a syntax structure that contains a raw byte sequence payload (RBSP) and data in the form of an indication of the type of data, and is interspersed as necessary using emulation prevention bytes. A VCL NAL unit is a NAL unit coded to contain video data such as a coded slice of a picture. A non-VCL NAL unit is a NAL unit that contains non-video data such as syntax and / or parameters that support decoding of video data, performing compliance checks, or other operations. A layer is a set of VCL NAL units that share specified characteristics (e.g., common resolution, frame rate, image size, etc.) and associated non-VCL NAL units as indicated by a layer identifier (ID). The NAL unit header layer identifier (nuh_layer_id) is a syntax element that specifies the identifier of the layer containing the NAL unit. A video parameter set (VPS) is a data unit that contains parameters relating to the entire video.A coded video sequence is a set of one or more coded pictures. A decoded video sequence is a set of one or more decoded pictures.

[0045] An output layer set (OLS) is a set of layers in which one or more layers are designated as output layers. An output layer is a layer designated for output (e.g., to a display). An OLS index is an index that uniquely identifies the corresponding OLS. A hypothetical reference decoder (HRD) is a decoder model that operates on an encoder to check the variability of a bitstream generated by an encoding process and verify its compliance with specified constraints. A bitstream compliance test is a test to determine whether an encoded bitstream conforms to a standard such as Versatile Video Coding (VVC). HRD parameters are syntax elements that initialize and / or define the operating conditions of the HRD. HRD parameters may be included in a supplementary enhancement information (SEI) message. An SEI message is a syntax structure with specified semantics that conveys information not required by the decoding process to determine the values of samples within a decoded picture. A scalable nesting SEI message is a message that includes one or more SEI messages corresponding to one or more OLSs or one or more layers. A buffering period (BP) SEI message is an SEI message that includes HRD parameters for initializing the HRD to manage a coded picture buffer (CPB). A picture timing (PT) SEI message is an SEI message that includes HRD parameters for managing delivery information for an access unit (AU) in the CPB and / or a decoded picture buffer (DPB). A decode unit information (DUI) SEI message is an SEI message that includes HRD parameters for managing delivery information for a DU in the CPB and / or DPB.

[0046] The scalable nesting SEI message contains a set of scalable nested SEI messages. A scalable nested SEI message is an SEI message nested within a scalable nesting SEI message. A flag is a variable or single-bit syntax element that can take one of two possible values: 0 and 1. The scalable nesting OLS flag is a flag that specifies whether a scalable nested SEI message is applied to a specific OLS or a specific layer. The number obtained by subtracting 1 from the scalable nesting number of OLSs (num_olss_minus1) is a syntax element that specifies the number of OLSs to which the scalable nested SEI message is applied. The number obtained by subtracting 1 from the total number of OLSs (TotalNumOlss - 1) is a syntax element that specifies the total number of OLSs specified in the VPS. The number obtained by subtracting 1 from the scalable nesting OLS delta (ols_idx_delta_minus1[i]) is a syntax element that contains sufficient data to derive the nesting OLS index. The nesting OLS index (NestingOlsIdx) is a syntax element that specifies the OLS index of the OLS to which the scalable nested SEI message is applied. The number obtained by subtracting 1 from the scalable nesting number of layers (num_layers_minus1) is a syntax element that specifies the number of layers to which the scalable nested SEI message is applied. The scalable nesting layer id (layer_id[i]) is a syntax element that specifies the nuh_layer_id value of the i-th layer to which the scalable nested SEI message is applied.

[0047] In this specification, the following acronyms are used: Access Unit (AU), Coding Tree Block (CTB), Coding Tree Unit (CTU), Coding Unit (CU), Coding Layer Video Sequence (CLVS), Coding Layer Video Sequence Start (CLVSS), Coded Video Sequence (CVS), Coded Video Sequence Start (CVSS), Joint Video Expert Team (JVET), Motion Constraint Tile Set (MCTS), Maximum Transfer Unit (MTU), Network Abstraction Layer (NAL), Output Layer Set (OLS), Picture Order Count (POC), Random Access Point (RAP), Raw Byte Sequence Payload (RBSP), Sequence Parameter Set (SPS), Video Parameter Set (VPS), Versatile Video Coding (VVC).

[0048] To reduce the size of video files while minimizing data loss, many video compression techniques can be used. For example, video compression techniques may include performing spatial (e.g., intra-picture) prediction and / or temporal (e.g., inter-picture) prediction to reduce or remove data redundancy in a video sequence. In the case of block-based video coding, a video slice (e.g., a video picture or a part of a video picture) may be divided into video blocks, which may also be referred to as tree blocks, coding tree blocks (CTB), coding tree units (CTU), coding units (CU), and / or coding nodes. Video blocks within an intra-coded (I) slice of a picture are coded using spatial prediction with respect to reference samples in adjacent blocks within the same picture. Video blocks within an inter-coded unidirectional prediction (P) or bidirectional prediction (B) slice of a picture may be coded by using spatial prediction with respect to reference samples in adjacent blocks within the same picture or temporal prediction with respect to reference samples in other reference pictures. A picture may sometimes be referred to as a frame and / or an image, and a reference picture may sometimes be referred to as a reference frame and / or a reference image. Spatial or temporal prediction results in a prediction block representing the image block. Residual data represents the pixel difference between the original image block and the prediction block. Thus, an inter-coded block is encoded according to a motion vector that points to a block of reference samples forming the prediction block and residual data indicating the difference between the coded block and the prediction block. An intra-coded block is encoded according to an intra-coding mode and residual data. For further compression, the residual data may be transformed from the pixel domain to the transform domain. These result in residual transform coefficients that may be quantized. The quantized transform coefficients may first be arranged in a two-dimensional array. The quantized transform coefficients may be scanned to generate a one-dimensional vector of transform coefficients. Entropy coding may be applied to achieve further compression.Such video compression techniques are described in more detail below.

[0049] To be able to accurately decode an encoded video, the video is encoded and decoded according to the corresponding video coding standard. Video coding standards include International Telecommunication Union (ITU) Telecommunication Standardization Sector (ITU-T) H.261, International Organization for Standardization / International Electrotechnical Commission (ISO / IEC) Motion Picture Experts Group (MPEG)-1 Part 2, ITU-T H.262 or ISO / IEC MPEG-2 Part 2, ITU-T H.263, ISO / IEC MPEG-4 Part 2, ITU-T H.264 or Advanced Video Coding (AVC) also known as ISO / IEC MPEG-4 Part 10, and High Efficiency Video Coding (HEVC) also known as ITU-T H.265 or MPEG-H Part 2. AVC includes extensions such as Scalable Video Coding (SVC), Multiview Video Coding (MVC) and Multiview Video Coding plus Depth (MVC+D), and 3D AVC (3D-AVC). HEVC includes extensions such as Scalable HEVC (SHVC), Multiview HEVC (MV-HEVC), and 3D HEVC (3D-HEVC). The joint video experts team (JVET) of ITU-T and ISO / IEC has started the development of a video coding standard called Versatile Video Coding (VVC). VVC is included in the Working Draft (WD) including JVET-O2001-v14.

[0050] Some video coding systems use Supplemental Enhancement Information (SEI) messages. SEI messages contain information that is not required by the decoding process to determine the values of samples within the decoded picture. For example, an SEI message may contain parameters that are used to check the bitstream for compliance with the standard. A Hypothetical Reference Decoder (HRD) may read an SEI message to determine how to check the bitstream for compliance. Such systems may use different types of SEI messages for data related to a layer and data related to an Output Layer Set (OLS) that includes the layer. This may result in a complex and redundant system.

[0051] Disclosed herein is a scalable nested SEI message configured to include parameters related to either a layer or OLS. For example, the scalable nested SEI message may include a scalable nested OLS flag, which may be set to indicate whether the scalable nested SEI message includes parameters related to a layer or parameters related to OLS. The scalable nested SEI message may also include one or more scalable nested SEI messages related to a layer or OLS. As used herein, one or more refers to any positive number of corresponding items, including one or more such items. When the scalable nested SEI message relates to OLS, the scalable nested SEI message also includes a flag indicating the number of OLSs associated with the scalable nested SEI message and a flag indicating an OLS index for associating the OLS with the scalable nested SEI message. When the scalable nested SEI message relates to a layer, the scalable nested SEI message also includes a flag indicating the number of layers associated with the scalable nested SEI message and a flag indicating a layer identifier (ID) for associating the layer with the scalable nested SEI message. In this way, the number of SEI message types can be reduced, which reduces complexity and the total number of message types. As a result, the length of the message ID data used to identify each type of message is reduced. The coding efficiency is thereby improved, and the use of processor, memory, and / or network signaling resources in both the encoder and decoder is reduced.

[0052] FIG. 1 is a flowchart of an exemplary operating method 100 for coding a video signal. Specifically, the video signal is encoded by an encoder. The encoding process compresses the video signal and reduces the video file size by using various mechanisms. The smaller file size enables the compressed video file to be transmitted to the user while reducing the associated bandwidth overhead. The decoder then decodes the compressed video file and reconstructs the original video signal for display to the end user. The decoding process generally mirrors the encoding process to enable the decoder to consistently reconstruct the video signal.

[0053] In step 101, a video signal is input to the encoder. For example, the video signal may be an uncompressed video file stored in memory. As another example, the video file may be captured by a video capture device such as a video camera and encoded to support live streaming of the video. The video file may include both an audio component and a video component. The video component includes a series of image frames that give a visual impression of movement when viewed in sequence. The frames include pixels that are represented with respect to light, referred to herein as the luminance component (or luminance samples), and color, referred to as the chrominance component (or color samples). In some examples, the frames may also include depth values to support 3D displays.

[0054] In stage 103, the video is divided into blocks. The division includes subdividing the pixels within each frame into square and / or rectangular blocks for compression. For example, in High Efficiency Video Coding (HEVC), also known as H.265 and MPEG-H Part 2, a frame can first be divided into Coding Tree Units (CTUs), which are blocks of a predetermined size (e.g., 64 pixels × 64 pixels). A CTU contains both luminance samples and chrominance samples. Using a coding tree, the CTU can be divided into blocks, and the blocks can be recursively subdivided until a configuration that supports further encoding is achieved. For example, the luminance component of a frame can be subdivided until the individual blocks contain relatively uniform illumination values. Further, the chrominance component of a frame can be subdivided until the individual blocks contain relatively uniform color values. Thus, the division mechanism varies according to the content of the video frame.

[0055] In stage 105, various compression mechanisms are used to compress the image blocks divided in stage 103. For example, inter prediction and / or intra prediction may be used. Inter prediction is designed to utilize the fact that objects within a common scene tend to appear in consecutive frames. Thus, blocks depicting objects in a reference frame need not be repeatedly described in adjacent frames. Specifically, an object such as a table may remain in a fixed position over multiple frames. Thus, the table is described once and adjacent frames can refer to the reference frame. A pattern matching mechanism can be used to match objects over multiple frames. Further, a moving object may be represented over multiple frames, for example, due to the movement of the object or the camera. As a specific example, a video may show a car moving across the screen over multiple frames. Such movement can be described using motion vectors. A motion vector is a two-dimensional vector that provides an offset from the coordinates of an object in a frame to the coordinates of the object in a reference frame. Thus, inter prediction can encode an image block in the current frame as a set of motion vectors indicating the offsets from corresponding blocks in the reference frame.

[0056] Intra prediction encodes blocks within a common frame. Intra prediction takes advantage of the fact that luminance and chrominance components tend to cluster within a frame. For example, some green patches of a tree tend to be located adjacent to similar green patches. Intra prediction uses multiple directional prediction modes (e.g., 33 in HEVC), a planar mode, and a direct current (DC) mode. The directional mode indicates that the current block is similar / the same as the samples of adjacent blocks in the corresponding direction. The planar mode indicates that a series of blocks along a row / column (e.g., a plane) can be interpolated based on adjacent blocks at the end of the row. The planar mode actually indicates a smooth transition of light / color across the row / column by adopting a relatively constant slope for varying values. The DC mode is used for boundary smoothing and indicates that the block is similar / the same as the average value associated with the samples of all adjacent blocks associated with the angular direction of the directional prediction mode. Thus, an intra prediction block can represent an image block as various relational prediction mode values rather than actual values. Further, an inter prediction block can represent an image block as motion vector values rather than actual values. In either case, the prediction block may not accurately represent the image block in some cases. Any difference is stored in a residual block. A transform can be applied to the residual block to further compress the file.

[0057] In stage 107, various filtering techniques can be applied. In HEVC, the filters are applied according to an in-loop filtering scheme. The block-based prediction described above can result in the generation of a block-shaped image at the decoder. Further, the block-based prediction scheme can encode a block and then reconstruct it for later use as a reference block. The in-loop filtering scheme repeatedly applies a noise reduction filter, a deblocking filter, an adaptive loop filter, and a sample adaptive offset (SAO) filter to the blocks / frames. These filters reduce such blocking artifacts so that the encoded file can be accurately reconstructed. Further, these filters mitigate the artifacts in the reconstructed reference blocks, so that the artifacts are less likely to generate additional artifacts in subsequent blocks encoded based on the reconstructed reference blocks.

[0058] When the video signal is segmented, compressed, and filtered, in stage 109, the resulting data is encoded into a bitstream. The bitstream includes the data described above and any signaling data that is desirably supported for proper video signal reconstruction at the decoder. For example, such data can include the segmentation data, prediction data, residual blocks, and various flags that provide coding instructions to the decoder. The bitstream may be stored in memory for transmission to the decoder upon request. The bitstream may also be broadcast and / or multicast to multiple decoders. The creation of the bitstream is an iterative process. Thus, stages 101, 103, 105, 107, and 109 can be performed continuously and / or simultaneously over many frames and blocks. The order shown in FIG. 1 is presented for clarity and ease of explanation and is not intended to limit the video coding process to a particular order.

[0059] The decoder receives a bitstream and starts the decoding process at stage 111. Specifically, the decoder uses an entropy decoding scheme to convert the bitstream into the corresponding syntax and video data. At stage 111, the decoder determines the frame partitioning using the syntax data from the bitstream. The partitioning should match the result of the block partitioning at stage 103. Here, the entropy encoding / decoding used at stage 111 will be described. The encoder makes many choices during the compression process, such as selecting a block partitioning scheme from several possible choices based on the spatial position of the values in the input image. To signal the exact choice, a large number of bins may be used. As used herein, a bin is a binary value treated as a variable (e.g., a bit value that can change depending on the context). Entropy coding allows the encoder to discard any options that are clearly not executable for a particular case and leave a set of acceptable options. Then, a codeword is assigned to each acceptable option. The length of the codeword is based on the number of acceptable options (e.g., one bin for two options, two bins for three to four options, etc.). Then, the encoder encodes the codeword for the selected option. This scheme reduces the size of the codeword because the codeword is as large as desired to uniquely indicate a selection from a small subset of acceptable options, as opposed to uniquely indicating a selection from a potentially large set of all possible options. Then, the decoder decodes the selection by determining the set of acceptable options in a manner similar to the encoder. By determining the set of acceptable options, the decoder can read the codeword and determine the selection made by the encoder.

[0060] In stage 113, the decoder performs the decoding of the block. Specifically, the decoder uses inverse transformation to generate the residual block. Then, the decoder uses the residual block and the corresponding prediction block to reconstruct the image block according to the segmentation. The prediction block can include both intra-prediction blocks and inter-prediction blocks such as those generated by the encoder in stage 105. Then, the reconstructed image block is arranged in the frame of the reconstructed video signal according to the segmentation data determined in stage 111. The syntax of stage 113 can also be signaled in the bitstream via entropy coding as described above.

[0061] In stage 115, filtering is performed on the frame of the reconstructed video signal in a manner similar to stage 107 in the encoder. For example, a noise suppression filter, a deblocking filter, an adaptive loop filter, and an SAO filter can be applied to the frame to remove blocking artifacts. Once the frame is filtered, the video signal can be output to the display in stage 117 for the end user to view.

[0062] Figure 2 is a schematic diagram of an exemplary coding and decoding (codec) system 200 for video coding. Specifically, the codec system 200 provides functions that support the implementation of the operation method 100. The codec system 200 is generalized to show components used in both the encoder and the decoder. As described with respect to steps 101 and 103 in the operation method 100, the codec system 200 receives and splits a video signal, and as a result, a split video signal 201 is obtained. Next, as described with respect to steps 105, 107, and 109 in method 100, when functioning as an encoder, the codec system 200 compresses the split video signal 201 into a coded bitstream. When operating as a decoder, the codec system 200 generates an output video signal from the bitstream as described with respect to steps 111, 113, 115, and 117 in the operation method 100. The codec system 200 includes a general-purpose codec control component 211, a transform scaling and quantization component 213, an intra-picture estimation component 215, an intra-picture prediction component 217, a motion compensation component 219, a motion estimation component 221, a scaling and inverse transform component 229, a filter control analysis component 227, an in-loop filter component 225, a decoded picture buffer component 223, and a header formatting and context-adaptive binary arithmetic coding (CABAC) component 231. Such components are coupled as shown. In Figure 2, the solid lines indicate the movement of data to be encoded / decoded, and the dashed lines indicate the movement of control data that controls the operation of other components. All components of the codec system 200 may be present in the encoder. The decoder may include a subset of the components of the codec system 200.For example, the decoder may include an intra-picture prediction component 217, a motion compensation component 219, a scaling and inverse transform component 229, an in-loop filter component 225, and a decoded picture buffer component 223. Next, these components will be described.

[0063] The segmented video signal 201 is a captured video sequence segmented into blocks of pixels by a coding tree. The coding tree uses various split modes to subdivide blocks of pixels into smaller blocks of pixels. These blocks can then be further subdivided into even smaller blocks. A block may sometimes be referred to as a node on the coding tree. A larger parent node is split into smaller child nodes. The number of times a node is subdivided is called the depth of the node / coding tree. The segmented blocks may, in some cases, be included in a coding unit (CU). For example, a CU can be a lower part of a CTU that includes a luminance block, a red differential chrominance (Cr) block, and a blue differential chrominance (Cb) block, along with the corresponding syntax instruction for the CU. Split modes can include a binary tree (BT), a triple tree (TT), and a quad tree (QT) used to split nodes of various shapes into two, three, or four child nodes respectively, depending on the split mode used. The segmented video signal 201 is transferred for compression to a general-purpose coder control component 211, a transform scaling and quantization component 213, an intra-picture estimation component 215, a filter control analysis component 227, and a motion estimation component 221.

[0064] The general-purpose coder control component 211 is configured to make decisions regarding the coding of the images of the video sequence into a bitstream according to the application constraints. For example, the general-purpose coder control component 211 manages the optimization of the bitrate / bitstream size versus the reconstructed quality. Such decisions can be made based on the availability of memory space / bandwidth and the image resolution requirements. The general-purpose coder control component 211 also manages the buffer utilization in light of the transmission speed in order to mitigate the problems of buffer underrun and overrun. To manage these problems, the general-purpose coder control component 211 manages the splitting, prediction, and filtering by other components. For example, the general-purpose coder control component 211 can dynamically increase the compression complexity to increase the resolution and the bandwidth usage, or decrease the compression complexity to decrease the resolution and the bandwidth usage. Thus, the general-purpose coder control component 211 controls the other components of the codec system 200 in order to balance the concerns of the video signal reconstructed quality and the bitrate. The general-purpose coder control component 211 creates control data for controlling the operation of other components. The control data is also transferred to the header formatting and the CABAC component 231 so as to be encoded in the bitstream with the signal parameters for decoding at the decoder.

[0065] The split video signal 201 is also transmitted to the motion estimation component 221 and the motion compensation component 219 for inter prediction. The frames or slices of the split video signal 201 can be split into a plurality of video blocks. The motion estimation component 221 and the motion compensation component 219 perform inter prediction coding of the received video blocks with respect to one or more blocks in one or more reference frames to provide temporal prediction. The codec system 200 can execute a plurality of coding paths, for example, to select an appropriate coding mode for each block of the video data.

[0066] The motion estimation component 221 and the motion compensation component 219 may be highly integrated but are shown separately for conceptual purposes. The motion estimation performed by the motion estimation component 221 is a process of generating a motion vector that estimates the motion of a video block. The motion vector can indicate, for example, the displacement of a coded object with respect to a prediction block. The prediction block is a block that is found to exactly match the block to be coded in terms of pixel difference. The prediction block may also be referred to as a reference block. Such pixel difference can be determined by the sum of absolute difference (SAD), the sum of square difference (SSD), or other difference metrics. HEVC uses several coded objects including CTUs, coding tree blocks (CTBs), and CUs. For example, a CTU can be divided into CTBs, which can then be divided into CUs for inclusion in CUs. A CU can be encoded as a prediction unit (PU) containing prediction data and / or a transform unit (TU) containing transformed residual data for the CU. The motion estimation component 221 generates motion vectors, PUs, and TUs by using rate-distortion analysis as part of the rate-distortion optimization process. For example, the motion estimation component 221 can determine multiple reference blocks, multiple motion vectors, etc. for the current block / frame and can select the reference block, motion vector, etc. with the best rate-distortion characteristics. The best rate-distortion characteristics balance both the quality of video reconstruction (e.g., the amount of data loss due to compression) and the coding efficiency (e.g., the size of the final encoding).

[0067] In some examples, codec system 200 can calculate the values of sub - integer pixel positions of reference pictures stored in decoded picture buffer component 223. For example, video codec system 200 can interpolate the values of 1 / 4 pixel positions, 1 / 8 pixel positions, or other fractional pixel positions of the reference picture. Thus, motion estimation component 221 can perform motion search for both full - pixel positions and fractional pixel positions and output motion vectors with fractional - pixel accuracy. Motion estimation component 221 calculates the motion vector for a prediction unit (PU) of a video block in an inter - coded slice by comparing the position of the PU with the position of the prediction block of the reference picture. Motion estimation component 221 outputs the calculated motion vector as motion data to header formatting and outputs it to motion compensation component 219 for encoding and motion - related context - adaptive binary arithmetic coding (CABAC) component 231.

[0068] Motion compensation performed by motion compensation component 219 may include fetching or generating a prediction block based on the motion vector determined by motion estimation component 221. Again, in some examples, motion estimation component 221 and motion compensation component 219 may be functionally integrated. Upon receiving the motion vector of the current video block's PU, motion compensation component 219 may be located at the prediction block indicated by the motion vector. Then, a residual video block is formed by subtracting the pixel values of the prediction block from the pixel values of the currently encoded video block to form a pixel difference value. Generally, motion estimation component 221 performs motion estimation for the luminance component, and motion compensation component 219 uses the motion vector calculated based on the luminance component for both the chrominance components and the luminance component. The prediction block and the residual block are transferred to transform, scaling, and quantization component 213.

[0069] The split video signal 201 is also sent to the intra picture estimation component 215 and the intra picture prediction component 217. Similar to the motion estimation component 221 and the motion compensation component 219, the intra picture estimation component 215 and the intra picture prediction component 217 may be highly integrated but are shown separately for conceptual purposes. The intra picture estimation component 215 and the intra picture prediction component 217 perform intra prediction of the current block for blocks within the current frame as an alternative to the inter prediction performed by the motion estimation component 221 and the motion compensation component 219 between frames, as described above. In particular, the intra picture estimation component 215 determines the intra prediction mode to be used to encode the current block. In some examples, the intra picture estimation component 215 selects an appropriate intra prediction mode from a plurality of tested intra prediction modes to encode the current block. The selected intra prediction mode is then transferred to the header formatting and CABAC component 231 for encoding.

[0070] For example, the intra picture estimation component 215 calculates rate-distortion values using rate-distortion analysis of various tested intra prediction modes and selects an intra prediction mode having the best rate-distortion characteristics among the tested modes. Rate-distortion analysis generally determines the amount of distortion (or error) between the encoded block and the original unencoded block that was encoded to generate the encoded block, and the bit rate (e.g., number of bits) used to generate the encoded block. The intra picture estimation component 215 calculates a ratio from the distortion and rate of various encoded blocks to determine which intra prediction mode exhibits the best rate-distortion value for the block. Additionally, the intra picture estimation component 215 may be configured to code depth blocks of a depth map using a depth modeling mode (DMM) based on rate-distortion optimization (RDO).

[0071] The intra picture prediction component 217 may generate a residual block from a prediction block based on the selected intra prediction mode determined by the intra picture estimation component 215 when implemented on an encoder, or may read a residual block from a bitstream when implemented on a decoder. The residual block includes the difference in values between the prediction block and the original block, represented as a matrix. The residual block is then transferred to the transform scaling and quantization component 213. The intra picture estimation component 215 and the intra picture prediction component 217 may operate on both the luminance component and the chrominance component.

[0072] The transform scaling and quantization component 213 is configured to further compress the residual block. The transform scaling and quantization component 213 applies a transform, such as a discrete cosine transform (DCT), a discrete sine transform (DST), or a conceptually similar transform, to the residual block to generate a video block including residual transform coefficient values. A wavelet transform, an integer transform, a subband transform, or other types of transforms can also be used. The transform can transform the residual information from the pixel value domain to a transform domain such as the frequency domain. The transform scaling and quantization component 213 is also configured to scale the transformed residual information, for example, based on frequency. Such scaling includes applying a scale factor to the residual information such that different frequency information is quantized at different granularities, which can affect the final visual quality of the reconstructed video. The transform scaling and quantization component 213 is also configured to quantize the transform coefficients to further reduce the bitrate. The quantization process can reduce the bit depth associated with some or all of the coefficients. The degree of quantization can be changed by adjusting the quantization parameter. In some examples, the transform scaling and quantization component 213 can then perform a scan of the matrix including the quantized transform coefficients. The quantized transform coefficients are transferred to the header formatting and CABAC component 231 to be encoded within the bitstream.

[0073] The scaling and inverse transform component 229 applies the inverse operations of the transform scaling and quantization component 213 to support motion estimation. The scaling and inverse transform component 229 applies inverse scaling, transform, and / or quantization to reconstruct the residual block in the pixel domain, for example, for later use as a reference block that can be a predicted block of another current block. The motion estimation component 221 and / or the motion compensation component 219 can calculate the reference block by adding the residual block back to the corresponding predicted block for use in motion estimation of subsequent blocks / frames. A filter is applied to the reconstructed reference block to reduce artifacts generated during scaling, quantization, and transform. Otherwise, such artifacts may cause inaccurate predictions (and generate additional artifacts) when subsequent blocks are predicted.

[0074] The filter control analysis component 227 and the in-loop filter component 225 apply a filter to the residual block and / or the reconstructed image block. For example, the transformed residual block from the scaling and inverse transform component 229 can be combined with the corresponding prediction block from the intra-picture prediction component 217 and / or the motion compensation component 219 to reconstruct the original image block. Then, the filter can be applied to the reconstructed image block. In some examples, the filter may instead be applied to the residual block. Similar to the other components in FIG. 2, the filter control analysis component 227 and the in-loop filter component 225 are highly integrated and can be implemented together, but are shown separately for conceptual purposes. The filter applied to the reconstructed reference block is applied to a specific spatial region and includes a plurality of parameters for adjusting how such a filter is applied. The filter control analysis component 227 analyzes the reconstructed reference block to determine where such a filter should be applied and sets the corresponding parameters. Such data is transferred to the header formatting and CABAC component 231 as filter control data for encoding. The in-loop filter component 225 applies such a filter based on the filter control data. The filter can include a deblocking filter, a noise suppression filter, an SAO filter, and an adaptive loop filter. Such a filter can be applied in the spatial / pixel domain (e.g., on the reconstructed pixel block) or the frequency domain depending on the example.

[0075] When operating as an encoder, the filtered and reconstructed image blocks, residual blocks, and / or prediction blocks are stored in the decoded picture buffer component 223 as described above for later use in motion estimation. When operating as a decoder, the decoded picture buffer component 223 stores the reconstructed and filtered blocks and transfers them towards the display as part of the output video signal. The decoded picture buffer component 223 can be any memory device capable of storing prediction blocks, residual blocks, and / or reconstructed image blocks.

[0076] The header formatting and CABAC component 231 receives data from various components of the codec system 200 and encodes such data into a coding bitstream for transmission to the decoder. Specifically, the header formatting and CABAC component 231 generates various headers for encoding control data such as general control data and filter control data. Further, prediction data including intra prediction and motion data, and residual data in the form of quantized transform coefficient data are all encoded within the bitstream. The final bitstream contains all the information desired by the decoder to reconstruct the original segmented video signal 201. Such information may also include an intra prediction mode index table (also called a codeword mapping table), the definition of encoding contexts for various blocks, an indication of the most likely intra prediction mode, an indication of the segmentation information, etc. Such data can be encoded by applying entropy coding. For example, the information can be encoded by using context adaptive variable length coding (CAVLC), CABAC, syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or another entropy coding technique. Following entropy coding, the coded bitstream can be transmitted to another device (e.g., a video decoder), or can be archived for later transmission or retrieval.

[0077] FIG. 3 is a block diagram showing an exemplary video encoder 300. The video encoder 300 can be used to implement the encoding function of the codec system 200 and / or to implement steps 101, 103, 105, 107, and / or 109 of the operation method 100. The encoder 300 splits the input video signal, resulting in a split video signal 301 that is substantially the same as the split video signal 201. The split video signal 301 is then compressed by components of the encoder 300 and encoded into a bitstream.

[0078] Specifically, the split video signal 301 is transferred to an intra-picture prediction component 317 for intra prediction. The intra-picture prediction component 317 may be substantially the same as the intra-picture estimation component 215 and the intra-picture prediction component 217. The split video signal 301 is also transferred to a motion compensation component 321 for inter prediction based on reference blocks in a decoded picture buffer component 323. The motion compensation component 321 may be substantially the same as the motion estimation component 221 and the motion compensation component 219. The prediction blocks and residual blocks from the intra-picture prediction component 317 and the motion compensation component 321 are transferred to a transform and quantization component 313 for transformation and quantization of the residual blocks. The transform and quantization component 313 may be substantially the same as the transform scaling and quantization component 213. The transformed and quantized residual blocks and the corresponding prediction blocks (along with associated control data) are transferred to an entropy coding component 331 for coding into a bitstream. The entropy coding component 331 may be substantially the same as the header formatting and CABAC component 231.

[0079] The transformed and quantized residual blocks and / or the corresponding prediction blocks are also transferred from the transform and quantization component 313 to the inverse transform and quantization component 329 for reconstruction into reference blocks for use by the motion compensation component 321. The inverse transform and quantization component 329 may be substantially similar to the scaling and inverse transform component 229. The in-loop filter within the in-loop filter component 325 is also applied to the residual blocks and / or the reconstructed reference blocks, depending on the example. The in-loop filter component 325 may be substantially similar to the filter control analysis component 227 and the in-loop filter component 225. The in-loop filter component 325 may include multiple filters as described for the in-loop filter component 225. The filtered blocks are then stored in the decoded picture buffer component 323 for use as reference blocks by the motion compensation component 321. The decoded picture buffer component 323 may be substantially similar to the decoded picture buffer component 223.

[0080] Figure 4 is a block diagram showing an exemplary video decoder 400. The video decoder 400 may be used to implement the decoding function of the codec system 200 and / or to implement steps 111, 113, 115, and / or 117 of the method of operation 100. The decoder 400 receives, for example, a bitstream from the encoder 300 and generates an output video signal reconstructed based on the bitstream for display to an end user.

[0081] The bitstream is received by the entropy decoding component 433. The entropy decoding component 433 is configured to implement an entropy decoding scheme such as CAVLC, CABAC, SBAC, PIPE coding, or other entropy coding techniques. For example, the entropy decoding component 433 can use the header information to provide a context for interpreting additional data encoded as codewords within the bitstream. The decoded information includes any desired information for decoding the video signal, such as general control data, filter control data, partition information, motion data, prediction data, and quantized transform coefficients from residual blocks. The quantized transform coefficients are transferred to the inverse transform and quantization component 429 for reconstruction into the residual blocks. The inverse transform and quantization component 429 may be similar to the inverse transform and quantization component 329.

[0082] The reconstructed residual block and / or prediction block is transferred to the intra-picture prediction component 417 to reconstruct the image block based on the intra prediction operation. The intra-picture prediction component 417 may be similar to the intra-picture estimation component 215 and the intra-picture prediction component 217. Specifically, the intra-picture prediction component 417 identifies a reference block within the frame using a prediction mode and applies the residual block, thereby reconstructing the intra prediction image block. The reconstructed intra prediction image block and / or residual block and the corresponding inter prediction data are transferred via the in-loop filter component 425 to the decoded picture buffer component 423, which may be substantially similar to the decoded picture buffer component 223 and the in-loop filter component 225 respectively. The in-loop filter component 425 filters the reconstructed image block, residual block, and / or prediction block, and such information is stored in the decoded picture buffer component 423. The reconstructed image block from the decoded picture buffer component 423 is transferred to the motion compensation component 421 for inter prediction. The motion compensation component 421 may be substantially similar to the motion estimation component 221 and / or the motion compensation component 219. Specifically, the motion compensation component 421 generates a prediction block using the motion vector from the reference block and applies the residual block to the result to reconstruct the image block. The resulting reconstructed block may also be transferred to the decoded picture buffer component 423 via the in-loop filter component 425. The decoded picture buffer component 423 continues to store additional reconstructed image blocks that can be reconstructed into the frame via the segmentation information. Such a frame may also be placed within the sequence. The sequence is output to the display as the reconstructed output video signal.

[0083] FIG. 5 is a schematic diagram showing an exemplary HRD500. HRD500 can be applied in an encoder such as, for example, codec system 200 and / or encoder 300. HRD500 can check the bitstream created at stage 109 of method 100 before the bitstream is transferred to a decoder such as decoder 400. In some examples, the bitstream can be continuously transferred through HRD500 when the bitstream is encoded. If a portion of the bitstream does not conform to the relevant constraints, HRD500 can indicate such a failure to the encoder and cause the encoder to re-encode the corresponding section of the bitstream by a different mechanism.

[0084] HRD500 includes a hypothetical stream scheduler (HSS) 541. HSS541 is a component configured to execute a virtual delivery mechanism. The virtual delivery mechanism is used to check the compliance of the bitstream or decoder with respect to the timing and data flow of bitstream 551 input to HRD500. For example, HSS541 can receive bitstream 551 output from the encoder and manage the compliance test process for bitstream 551. In a particular example, HSS541 can control the rate at which coded pictures move through HRD500 and verify that bitstream 551 does not contain non-conforming data.

[0085] HSS541 can transfer the bitstream 551 to the CPB543 at a predetermined rate. The HRD500 can manage data in the decode unit (DU) 553. The DU553 is a subset of the access unit (AU), or the AU and the related non-video coding layer (VCL) network abstraction layer (NAL) units. Specifically, the AU includes one or more pictures associated with the output time. For example, the AU can include a single picture in a single-layer bitstream, or can include pictures for each layer in a multi-layer bitstream. Each picture of the AU can be divided into slices each included in the corresponding VCL NAL unit. Therefore, the DU553 can include one or more pictures, one or more slices of a picture, or a combination thereof. Also, the parameters used to decode the AU, picture, and / or slice can be included in the non-VCL NAL unit. Thus, the DU553 includes non-VCL NAL units containing the data necessary to support the decoding of the VCL NAL units within the DU553. The CPB543 is the first-in first-out buffer in the HRD500. The CPB543 includes the DU553 containing video data in decode order. The CPB543 stores video data for use during bitstream compliance verification.

[0086] The CPB543 transfers the DU553 to the decode processing component 545. The decode processing component 545 is a component compliant with the VVC standard. For example, the decode processing component 545 can emulate the decoder 400 used by the end user. The decode processing component 545 decodes the DU553 at a rate that can be achieved by an exemplary end-user decoder. If the decode processing component 545 cannot decode the DU553 fast enough to prevent an overflow of the CPB543, the bitstream 551 is non-compliant and should be re-encoded.

[0087] The decoding processing component 545 decodes the DU 553 and creates the decoded DU 555. The decoded DU 555 contains the decoded picture. The decoded DU 555 is transferred to the DPB 547. The DPB 547 may be substantially similar to the decoded picture buffer components 223, 323, and / or 423. To support inter prediction, the pictures marked as reference pictures 556 obtained from the decoded DU 555 for use are returned to the decoding processing component 545 to support further decoding. The DPB 547 outputs the decoded video sequence as a series of pictures 557. The pictures 557 are reconstructed pictures that generally mirror the pictures encoded in the bitstream 551 by the encoder.

[0088] The pictures 557 are transferred to the output clipping component 549. The output clipping component 549 is configured to apply the compliance clipping window to the pictures 557. Thereby, the output trimmed picture 559 is obtained. The output trimmed picture 559 is a completely reconstructed picture. Thus, the output trimmed picture 559 mimics what the end user would see when decoding the bitstream 551. In this way, the encoder can review the output trimmed picture 559 to ensure that the encoding is satisfactory.

[0089] HRD500 is initialized based on the HRD parameters within bitstream 551. For example, HRD500 can read the HRD parameters from the VPS, SPS, and / or SEI messages. Then, HRD500 can perform a compliance test operation on bitstream 551 based on the information within such HRD parameters. As a specific example, HRD500 may determine one or more CPB delivery schedules from the HRD parameters. The delivery schedule specifies the timing of the delivery of video data between memory locations such as CPB and / or DPB. Thus, the CPB delivery schedule specifies the timing of the delivery of AUs, DUs 553, and / or pictures to / from CPB543. It should be noted that HRD500 can use a DPB delivery schedule for DPB547 similar to the CPB delivery schedule.

[0090] The video can be coded into different layers and / or OLSs for use by decoders having various levels of hardware capabilities and for various network conditions. The CPB delivery schedule is selected to reflect these issues. Thus, the upper layer sub-bitstream is specified for optimal hardware and network conditions, and thus the upper layer may receive one or more CPB delivery schedules that use a short delay for the transfer of DU553 to the large amount of memory in CPB543 and DPB547. Similarly, the lower layer sub-bitstream is specified for limited decoder hardware capabilities and / or poor network conditions. Thus, the lower layer may receive one or more CPB delivery schedules that use a longer delay for the transfer of DU553 to the small amount of memory in CPB543 and DPB547. Then, the OLS, layer, sub-layer, or combinations thereof can be tested according to the corresponding delivery schedule to ensure that the resulting sub-bitstream can be correctly decoded under the conditions expected for the sub-bitstream. Thus, the HRD parameters in bitstream 551 may indicate the CPB delivery schedule, and the HRD500 may contain sufficient data to determine the CPB delivery schedule and correlate the CPB delivery schedule to the corresponding OLS, layer, and / or sub-layer.

[0091] FIG. 6 is a schematic diagram showing an exemplary multi-layer video sequence 600 configured for inter-layer prediction 621. The multi-layer video sequence 600 may be encoded by an encoder such as codec system 200 and / or encoder 300, for example, according to method 100, and may be decoded by a decoder such as codec system 200 and / or decoder 400. Further, the multi-layer video sequence 600 may be checked for compliance by an HRD such as HRD 500. The multi-layer video sequence 600 is included to show an exemplary application of layers within a coded video sequence. The multi-layer video sequence 600 is any video sequence that uses multiple layers such as layer N 631 and layer N+1 632.

[0092] In one example, the multi-layer video sequence 600 may use inter-layer prediction 621. Inter-layer prediction 621 is applied between pictures 611, 612, 613, and 614 and pictures 615, 616, 617, and 618 of a different layer. In the example shown, pictures 611, 612, 613, and 614 are part of layer N+1 632, and pictures 615, 616, 617, and 618 are part of layer N 631. Layers such as layer N 631 and / or layer N+1 632 are all groups of pictures associated with similar values of characteristics such as similar size, quality, resolution, signal-to-noise ratio, capabilities, etc. A layer may be formally defined as a set of VCL NAL units and associated non-VCL NAL units. A VCL NAL unit is a NAL unit coded to contain video data such as a coded slice of a picture. A non-VCL NAL unit is a NAL unit that contains non-video data such as syntax and / or parameters that support decoding of video data, performing compliance checks, or other operations.

[0093] In the example shown, layer N+1 632 is associated with a larger image size than layer N 631. Thus, in this example, the picture sizes of pictures 611, 612, 613, and 614 in layer N+1 632 are larger than the picture sizes of pictures 615, 616, 617, and 618 in layer N 631 (e.g., greater height and width, and thus more samples). However, such pictures can be separated between layer N+1 632 and layer N 631 by other characteristics. Only two layers, layer N+1 632 and layer N 631, are shown, but a set of pictures can be separated into any number of layers based on the relevant characteristics. Layer N+1 632 and layer N 631 can also be indicated by a layer ID. The layer ID is an item of data associated with the picture and indicates that the picture is part of the indicated layer. Thus, each of pictures 611 to 618 can be associated with a corresponding layer ID to indicate which layer N+1 632 or layer N 631 contains the corresponding picture. For example, the layer ID can include a NAL unit header layer identifier (nuh_layer_id) which is a syntax element that specifies an identifier of a layer containing a NAL unit (including, for example, slices and / or parameters of pictures within the layer). Layers associated with a low quality / bitstream size, such as layer N 631, are generally assigned a lower layer ID and are called lower layers. Further, layers associated with a high quality / bitstream size, such as layer N+1 632, are generally assigned a higher layer ID and are called upper layers.

[0094] Pictures 611-618 within different layers 631-632 are configured to be displayed in alternative forms. Thus, pictures of different layers 631-632 can share a time ID and can be included in the same AU. The time ID is a data element indicating that the data corresponds to a temporal location within the video sequence. An AU is a set of NAL units related to each other according to a specified classification rule and regarding a specific output time. For example, if picture 611 and picture 615 are associated with the same time ID, an AU can include one or more pictures within different layers such as those pictures. As a specific example, if a smaller picture is desired, the decoder may decode and display picture 615 at the current display time, or if a larger picture is desired, the decoder may decode and display picture 611 at the current display time. Thus, pictures 611-614 in the upper layer N+1 632 include substantially the same image data as the corresponding pictures 615-618 in the lower layer N 631 (regardless of the difference in picture size). Specifically, picture 611 includes substantially the same image data as picture 615, picture 612 includes substantially the same image data as picture 616, and so on.

[0095] Pictures 611 to 618 can be coded by referring to other pictures 611 to 618 within the same layer N 631 or N+1 632. When a picture is coded by referring to another picture within the same layer, an inter prediction 623 is obtained. The inter prediction 623 is indicated by a solid arrow. For example, picture 613 can be coded by using an inter prediction 623 that uses one or two of pictures 611, 612, and / or 614 within layer N+1 632 as references. One picture is referred to for uni-directional inter prediction and / or two pictures are referred to for bi-directional inter prediction. Further, picture 617 can be coded by using an inter prediction 623 that uses one or two of pictures 615, 616, and / or 618 within layer N 631 as references. One picture is referred to for uni-directional inter prediction and / or two pictures are referred to for bi-directional inter prediction. When performing the inter prediction 623, when a picture is used as a reference to another picture within the same layer, that picture can be called a reference picture. For example, picture 612 can be a reference picture used to code picture 613 according to the inter prediction 623. The inter prediction 623 can also be called an intra-layer prediction in a multi-layer context. Thus, the inter prediction 623 is a mechanism for coding samples of a current picture by referring to indicated samples in a reference picture different from the current picture where the reference picture and the current picture are within the same layer.

[0096] Pictures 611 to 618 can also be coded by referring to other pictures 611 to 618 in different layers. This process is known as inter-layer prediction 621 and is indicated by the dashed arrow. Inter-layer prediction 621 is a mechanism for coding samples of the current picture by referring to the indicated samples in a reference picture within a different layer where the current picture and the reference picture are in different layers and thus have different layer IDs. For example, a picture in the lower layer N 631 can be used as a reference picture to code the corresponding picture in the upper layer N+1 632. As a specific example, picture 611 can be coded by referring to picture 615 according to inter-layer prediction 621. In such a case, picture 615 is used as an inter-layer reference picture. An inter-layer reference picture is a reference picture used for inter-layer prediction 621. In most cases, inter-layer prediction 621 is restricted such that the current picture, such as picture 611, can only use inter-layer reference pictures in a lower layer, such as picture 615, that are included in the same AU. If multiple layers (e.g., more than 2) are available, inter-layer prediction 621 can encode / decode the current picture based on multiple inter-layer reference pictures at a level lower than the current picture.

[0097] The video encoder can encode pictures 611 to 618 through many different combinations and / or permutations of inter prediction 623 and inter-layer prediction 621 using the multi-layer video sequence 600. For example, picture 615 can be coded according to intra prediction. Then, pictures 616 to 618 can be coded according to inter prediction 623 by using picture 615 as a reference picture. Further, picture 611 can be coded according to inter-layer prediction 621 by using picture 615 as an inter-layer reference picture. Then, pictures 612 to 614 can be coded according to inter prediction 623 by using picture 611 as a reference picture. Thus, the reference picture can function as both a single-layer reference picture and an inter-layer reference picture for different encoding mechanisms. By coding the upper-layer N+1 picture 632 based on the picture of the lower-layer N 631, the upper-layer N+1 632 can avoid using intra prediction, which has a much lower coding efficiency than inter prediction 623 and inter-layer prediction 621. Thus, the poor coding efficiency of intra prediction can be limited to the minimum / lowest quality pictures and, therefore, can be limited to coding the minimum amount of video data. The pictures used as reference pictures and / or inter-layer reference pictures can be indicated by the entries of the reference picture list included in the reference picture list structure.

[0098] To perform such operations, layers such as layer N 631 and layer N+1 632 may be included in OLS625. OLS625 is a set of layers in which one or more layers are designated as output layers. An output layer is a layer designated for output (e.g., to a display). For example, layer N 631 may be included only to support inter-layer prediction 621 and may never be output. In such a case, layer N+1 632 is decoded and output based on layer N 631. In such a case, OLS625 includes layer N+1 632 as the output layer. In some cases, OLS625 includes only output layers called simulcast layers. In other cases, OLS625 may include many layers in different combinations. For example, the output layers within OLS625 may be coded according to inter-layer prediction 621 based on one, two, or a number of lower layers. Further, OLS625 may include two or more output layers. Thus, OLS625 may include one or more output layers and any support layers necessary to reconstruct the output layers. The multi-layer video sequence 600 may be coded by using many different OLS625s, each using a different combination of layers. OLS625 is associated with an OLS index, which is an index that uniquely identifies the corresponding OLS625.

[0099] The check of the multi-layer video sequence 600 for compliance in HRD500 may become complex depending on the number of layers 631-632 and OLS625. Scalable nesting SEI messages may be used to indicate the parameters necessary to check layers 631-632 and OLS625 for compliance.

[0100] FIG. 7 is a schematic diagram showing an exemplary bitstream 700. For example, the bitstream 700 may be generated by the codec system 200 and / or the encoder 300 for decoding by the codec system 200 and / or the decoder 400 according to method 100. Further, the bitstream 700 may include a multi-layer video sequence 600. Additionally, the bitstream 700 may include various parameters for controlling the operation of an HRD such as the HRD 500. Based on such parameters, the HRD can check the bitstream 700 for compliance with the standard before sending it to the decoder for decoding.

[0101] The bitstream 700 includes a VPS 711, one or more SPSs 713, a plurality of picture parameter sets (PPSs) 715, a plurality of slice headers 717, picture data 720, and SEI messages 719. The VPS 711 includes data related to the entire bitstream 700. For example, the VPS 711 may include data-related OLSs, layers, and / or sublayers used in the bitstream 700. The SPS 713 includes sequence data common to all pictures within the coded video sequence included in the bitstream 700. For example, each layer may include one or more coded video sequences, and each coded video sequence may refer to the SPS 713 for corresponding parameters. The parameters within the SPS 713 can include picture sizing, bit depth, coding tool parameters, bitrate limits, and the like. Note that each sequence refers to the SPS 713, but in some examples, a single SPS 713 can include data for multiple sequences. The PPS 715 includes parameters applied to the entire picture. Thus, each picture in the video sequence may refer to the PPS 715. Note that each picture refers to the PPS 715, but in some examples, a single PPS 715 can include data for multiple pictures. For example, multiple similar pictures may be coded according to similar parameters. In such cases, a single PPS 715 may include data for such similar pictures. The PPS 715 may indicate coding tools, quantization parameters, offsets, etc. available for slices within the corresponding picture.

[0102] Slice header 717 contains parameters specific to each slice within a picture. Thus, there may be one slice header 717 for each slice within a video sequence. The slice header 717 may include slice type information, POC, reference picture list, prediction weight, tile entry point, deblocking parameters, etc. Note that in some examples, the bitstream 700 may also include a picture header, which is a syntax structure containing parameters applicable to all slices within a single picture. For this reason, the picture header and the slice header 717 may be used interchangeably in some contexts. For example, certain parameters may be moved between the slice header 717 and the picture header depending on whether such parameters are common to all slices within a picture.

[0103] The picture data 720 includes video data encoded according to inter-prediction and / or intra-prediction, as well as corresponding transformed and quantized residual data. For example, the picture data 720 may include OLS 721, layer 723, picture 725, and / or slice 727. OLS 721 is a set of layers 723 where one or more layers are designated as output layers. OLS 721 may be substantially the same as OLS 625. The layer 723 is a set of VCL NAL units that share specified characteristics (such as common resolution, frame rate, picture size, etc.) and associated non-VCL NAL units, as indicated by a layer ID such as nuh_layer_id. For example, the layer 723 may include a set of pictures 725 that share the same nuh_layer_id. The layer 723 may be substantially the same as layer 631 and / or 632. The picture 725 is an array of luminance samples and / or an array of chrominance samples that generate a frame or a field thereof. For example, the picture 725 is a coded image that can be output for display or used to support the coding of other pictures 725 for output. The picture 725 includes one or more slices 727. The slice 727 may be defined as a continuous complete coding tree unit (CTU) row (e.g., within a tile) of the picture 725 that is exclusively included in an integer number of complete tiles or an integer number of single NAL units. The slice 727 is further divided into CTUs and / or coding tree blocks (CTBs). The CTU is a group of samples of a predetermined size that can be divided by a coding tree. The CTB is a subset of the CTU and includes the luminance component or chrominance component of the CTU. The CTU / CTB is further divided into coding blocks based on a coding tree. Then, the coding blocks can be encoded / decoded according to a prediction mechanism.

[0104] The bitstream 700 can be coded as a series of NAL units. A NAL unit is a container for video data and / or supporting syntax. A NAL unit can be a VCL NAL unit or a non-VCL NAL unit. A VCL NAL unit is a NAL unit coded to contain video data such as picture data 720 and an associated slice header 717. A non-VCL NAL unit is a NAL unit containing non-video data such as syntax and / or parameters to support the decoding of video data, the execution of compliance checks, or other operations. For example, a non-VCL NAL unit can include a VPS 711, an SPS 713, a PPS 715, a SEI message 719, or other supporting syntax.

[0105] The SEI message 719 is a syntax structure having a specified semantics that conveys information not required by the decoding process to determine the values of samples within the decoded picture. For example, the SEI message 719 may include data for supporting HRD processing or other support data not directly related to the decoding of the bitstream 700 in the decoder. The SEI message 719 may be a scalable nesting SEI message. A scalable nesting SEI message is a message that includes a plurality of scalable nested SEI messages corresponding to one or more OLSs 721 or one or more layers 723. Thus, a scalable nesting SEI message is the SEI message 719 that includes a set of scalable nested SEI messages of the same type. The SEI message 719 may include a BP SEI message that includes HRD parameters for initializing HRD to manage the CPB. The SEI message 719 may also include a PT SEI message that includes HRD parameters for managing the delivery information for AUs in the CPB and / or DPB. The SEI message 719 may also include a DUI SEI message that includes HRD parameters for managing the delivery information for DUs in the CPB and / or DPB.

[0106] The bitstream 700 includes various flags for signaling the composition of the SEI message 719. For example, the SEI message 719 may include a scalable nesting (SN) OLS flag 731, a number obtained by subtracting 1 from the number of scalable nestings of the OLS (num_olss_minus1) 733, a number obtained by subtracting 1 from the scalable nesting OLS delta (ols_idx_delta_minus1[i]) 735, a number obtained by subtracting 1 from the number of scalable nestings of the layer (num_layers_minus1) 737, and / or a scalable nesting layer ID (layer_id[i]) 739 when the SEI message 719 is a scalable nesting SEI message.

[0107] The scalable nesting OLS flag 731 is a syntax element that specifies whether a scalable nested SEI message within a scalable nesting SEI message applies to a specific OLS 721 or to a specific layer 723. For example, the scalable nesting OLS flag 731 can be set to 1 when a scalable nested SEI message applies to a specific OLS 721 (rather than a layer). Further, the scalable nesting OLS flag 731 can be set to 0 when a scalable nested SEI message applies to a specific layer 723 (rather than an OLS). Thus, the HRD can read the scalable nesting OLS flag 731 within the SEI message 719 and determine whether all of the scalable nested SEI messages contained therein describe an OLS 721 or a layer 723.

[0108] The scalable nesting num_olss_minus1 733 is used when the SEI message 719 is related to an OLS 721 as indicated by the scalable nesting OLS flag 731. The scalable nesting num_olss_minus1 733 is a syntax element that specifies the number of OLSs 721 to which the scalable nested SEI message within the scalable nesting SEI message applies. The scalable nesting num_olss_minus1 733 uses a format of minus 1 and thus includes a value that is 1 less than the actual value. For example, if the scalable nesting SEI message contains a scalable nested SEI message related to 5 OLSs 721, the scalable nesting num_olss_minus1 733 is set to a value of 4.

[0109] Scalable nesting ols_idx_delta_minus1[i] 735 is used when the SEI message 719 is related to OLS 721, as indicated by the scalable nesting OLS flag 731. Scalable nesting ols_idx_delta_minus1[i] 735 is a syntactic element that contains sufficient data to derive the nesting OLS index. Specifically, scalable nesting ols_idx_delta_minus1[i] 735 contains the OLS index of each scalable nested SEI message within the scalable nesting SEI message. Thus, scalable nesting ols_idx_delta_minus1[i] 735 can be used to correlate the scalable nested SEI message with OLS 721. In a specific example, ols_idx_delta_minus1[i] 735 can be used to determine the NestingOlsIdx (Nesting OLS index) of each scalable nested SEI message. NestingOlsIdx is a syntactic element that specifies the OLS index of the OLS 721 to which the corresponding scalable nested SEI message applies. In one example, the variable NestingOlsIdx[i] is derived as follows.

Number

[0110] Scalable nesting num_layers_minus1 737 is used when the SEI message 719 is related to layer 723, as indicated by the Scalable nesting OLS flag 731. Scalable nesting num_layers_minus1 737 is a syntax element that specifies the number of layers 723 to which the scalable nested SEI messages within the Scalable nesting SEI message apply. Scalable nesting num_layers_minus1 737 uses a format of minus 1 and thus includes a value that is 1 less than the actual value. For example, if the Scalable nesting SEI message includes scalable nested SEI messages related to 5 layers 723, the Scalable nesting num_layers_minus1 737 is set to a value of 4.

[0111] layer_id[i] 739 is used when the SEI message 719 is related to layer 723, as indicated by the Scalable nesting OLS flag 731. layer_id[i] 739 is a syntax element that specifies the nuh_layer_id value of the i-th layer to which the scalable nested SEI message applies. In this way, layer_id[i] 739 can be used to associate each of the scalable nested SEI messages with the corresponding layer 723.

[0112] Thus, the flags described in the bitstream 700 enable the HRD and / or decoder to quickly determine the composition of the SEI message 719. The HRD / decoder may use the scalable nesting OLS flag 731 to determine whether a set of scalable nested messages is related to the OLS 721 or layer 723. The HRD / decoder then uses the scalable nesting num_olss_minus1 733 to determine the number of corresponding OLSs 721 and the scalable nesting ols_idx_delta_minus1[i] 735 to determine the index of each corresponding OLS 721, so that when the scalable nested message is related to the OLS 721, it can determine how to apply the scalable nested message. Further, the HRD / decoder uses the scalable nesting num_layers_minus1 737 to determine the number of corresponding layers 723 and the layer_id[i] 739 to determine the index of each corresponding layer 723, so that when the scalable nested message is related to the layer 723, it can determine how to apply the scalable nested message. This approach reduces the number of SEI message 719 types. This reduces complexity and the total number of message types. As a result, the length of the message ID data used to identify each type of message is reduced. Consequently, coding efficiency is improved, and the use of processor, memory, and / or network signaling resources in both the encoder and decoder is reduced.

[0113] Here, the aforementioned information will be described in more detail below. Layered video coding is also referred to as scalable video coding or video coding with scalability. Scalability in video coding can be supported by using multi-layer coding techniques. A multi-layer bitstream comprises a base layer (BL) and one or more enhancement layers (EL). Examples of scalability include spatial scalability, quality / signal-to-noise ratio (SNR) scalability, multi-view scalability, frame rate scalability, and the like. When multi-layer coding techniques are used, a picture or a part thereof may be coded without using a reference picture (intra prediction), may be coded by referring to a reference picture within the same layer (inter prediction), and / or may be coded by referring to a reference picture within another layer (inter-layer prediction). A reference picture used for inter-layer prediction of the current picture is called an inter-layer reference picture (ILRP). FIG. 6 shows an example of multi-layer coding for spatial scalability where pictures of different layers have different resolutions.

[0114] Some video coding families provide support for scalability in profiles that are separate from the profiles for single-layer coding. Scalable Video Coding (SVC) is a scalable extension of Advanced Video Coding (AVC) that provides support for spatial, temporal, and quality scalability. In the case of SVC, a flag is signaled at each macroblock (MB) in the EL picture to indicate whether the EL MB is predicted using a collocated block from the lower layer. Prediction from a collocated block may include texture, motion vectors, and / or coding mode. Implementations of SVC may not directly reuse AVC implementations that have not been modified in their design. The SVC EL macroblock syntax and decode processing are different from the AVC syntax and decode processing.

[0115] Scalable High Efficiency Video Coding (SHVC) is an extension of HEVC that provides support for spatial and quality scalability. Multi-View HEVC (MV-HEVC) is an extension of HEVC that provides support for multi-view scalability. 3D HEVC (3D-HEVC) is an extension of HEVC that provides support for 3D video coding that is more advanced and efficient than MV-HEVC. Temporal scalability can be included as an essential part of a single-layer HEVC codec. In the multi-layer extension of HEVC, the decoded pictures used for inter-layer prediction come only from the same access unit and are treated as long-term reference pictures (LTRPs). Such pictures are assigned reference indexes in the reference picture list together with other temporal reference pictures within the current layer. Inter-layer prediction (ILP) is achieved at the prediction unit (PU) level by setting the value of the reference index for referring to the inter-layer reference pictures within the reference picture list. Spatial scalability resamples the reference picture or a part thereof when the ILRP has a different spatial resolution from the current picture being encoded or decoded. Resampling of the reference picture can be realized either at the picture level or at the coding block level.

[0116] VVC can also support layered video coding. The VVC bitstream may contain multiple layers. The layers may all be independent of each other. For example, each layer can be coded without using inter-layer prediction. In this case, the layers are also called simulcast layers. In some cases, a part of the layers is coded using ILP. The flag in the VPS can indicate whether the layer is a simulcast layer or which layers use ILP. When some layers use ILP, the layer dependencies between the layers are also signaled in the VPS. Different from SHVC and MV-HEVC, VVC may not specify an OLS. The OLS includes a specified set of layers, and one or more layers in the set of layers are specified to be output layers. The output layer is the layer of the OLS to be output. In some implementations of VVC, when the layer is a simulcast layer, only one layer can be selected for decoding and output. In some implementations of VVC, when any layer uses ILP, the entire bitstream including all layers is specified to be decoded. Further, a specific layer among the layers is designated as the output layer. The output layer can be shown to be only the highest layer, all layers, or the set of lower layers indicated by the highest layer.

[0117] The foregoing aspects include certain problems. HEVC, including scalable extended SHVC and MV-HEVC, may use scalable nesting SEI messages to associate SEI messages with bitstream subsets corresponding to various operating points or specific layers or sub-layers. HEVC may also use bitstream segmentation nesting to associate SEI messages with bitstream segmentation in OLS. Bitstream segmentation includes one or more layers of a multi-layer bitstream. Each bitstream segmentation nesting SEI message may be included within a scalable nesting SEI message. This two-level nesting scheme for SEI messages for OLS is complex.

[0118] Generally, the present disclosure describes a technique for scalable nesting of SEI messages for an output layer set in a multi-layer video bitstream. The description of the technique is based on VVC. However, the technique is also applicable to layered video coding based on other video codec specifications.

[0119] One or more of the above problems can be solved as follows. Specifically, the present disclosure includes a method for simple and efficient scalable nesting of SEI messages for OLS in a multi-layer video bitstream. Instead of using a two-level nesting scheme, only one nesting SEI message is defined to directly include the nesting SEI message applied to one or more layers within OLS.

[0120] Exemplary implementations of the foregoing mechanisms are as follows. An exemplary scalable nesting SEI message syntax is as follows.

Table 1-1

Table 1-2

[0121] In an alternative example, a flag may be added when nesting_ols_flag is equal to 1. This flag may be set equal to 1 to indicate that the scalable nested SEI message applies to all OLSs and is applicable to all layers within each OLS. If this flag is set equal to 1, all syntax elements from this flag up to nesting_num_seis_minus1 are not signaled. In another alternative example, a flag is used and may be set equal to 1 to indicate that the scalable nested SEI message applies to all OLSs. If this flag is equal to 1, the syntax element nesting_num_olss_minus1 and the list of syntax elements nesting_ols_idx_delta_minus1[i] are not signaled. In another alternative example, the nested OLS index value signaled by the syntax element nesting_ols_idx_delta_minus1[i] is directly coded instead of being delta-coded. In another alternative example, a flag is used and may be set equal to 1 to indicate that the scalable nested SEI message applies to all layers of the OLS. If this flag is equal to 1, the list of syntax elements nesting_num_ols_layers_minus1[i] and nesting_ols_layer_idx_delta_minus1[i][j] are not signaled. In another alternative example, the nested OLS layer index value signaled by the syntax element nesting_ols_layer_idx_delta_minus1[i][j] is directly coded instead of being delta-coded.

[0122] Exemplary scalable nesting SEI message semantics are as follows.

[0123] The scalable nesting SEI message provides a mechanism for associating SEI messages with a specific layer in the context of a particular OLS or with a specific layer not in the context of the OLS. The scalable nesting SEI message contains one or more SEI messages. The SEI messages contained in the scalable nesting SEI message are also referred to as scalable nested SEI messages. Bitstream compliance may require applying the following restrictions when including SEI messages within a scalable nesting SEI message.

[0124] SEI messages having a payloadType equal to 132 (decoded picture hash) or 133 (scalable nesting) should not be included in a scalable nesting SEI message. If a scalable nesting SEI message contains a buffering period, picture timing, or decode unit information SEI message, the scalable nesting SEI message should not contain any other SEI message whose payloadType is not equal to 0 (buffering period), 1 (picture timing), or 130 (decode unit information).

[0125] Bitstream compliance may also require applying the following restrictions to the value of nal_unit_type of an SEI NAL unit containing a scalable nesting SEI message. When the scalable nesting SEI message contains an SEI message whose payloadType is equal to 0 (buffering period), 1 (picture timing), 130 (decode unit information), 145 (dependent RAP indication), or 168 (frame field information), the SEI NAL unit containing the scalable nesting SEI message should have a nal_unit_type set equal to PREFIX_SEI_NUT. When the scalable nesting SEI message contains an SEI message having a payloadType equal to 132 (decoded picture hash), the SEI NAL unit containing the scalable nesting SEI message should have a nal_unit_type set equal to SUFFIX_SEI_NUT.

[0126] The nesting_ols_flag may be set equal to 1 to specify that a scalable nested SEI message applies to a particular layer in the context of a particular OLS. The nesting_ols_flag may be set equal to 0 to specify that a scalable nested SEI message generally applies to a particular layer (e.g., not in the context of an OLS).

[0127] Bitstream compliance may require applying the following restrictions to the value of nesting_ols_flag. For scalable nesting SEI messages, when the SEI message contains an SEI message with payloadType equal to 0 (buffering period), 1 (picture timing), or 130 (decode unit information), the value of nesting_ols_flag should be equal to 1. For scalable nesting SEI messages, when the SEI message contains an SEI message with payloadType equal to a value within VclAssociatedSeiList, the value of nesting_ols_flag should be equal to 0.

[0128] The number obtained by adding 1 to nesting_num_olss_minus1 specifies the number of OLSs to which the scalable nested SEI message is applied. The value of nesting_num_olss_minus1 should be in the range of 0 to TotalNumOlss - 1 (both ends included). nesting_ols_idx_delta_minus1[i] is used to derive the variable NestingOlsIdx[i] that specifies the OLS index of the i-th OLS to which the scalable nested SEI message is applied when nesting_ols_flag is equal to 1. The value of nesting_ols_idx_delta_minus1[i] should be in the range of 0 to TotalNumOlss - 2 (both ends included). The variable NestingOlsIdx[i] can be derived as follows.

Number

[0129] The number obtained by adding 1 to nesting_num_ols_layers_minus1[i] specifies the number of layers to which the scalable nested SEI message is applied in the context of the NestingOlsIdx[i]-th OLS. The value of nesting_num_ols_layers_minus1[i] should be within the range of 0 to NumLayersInOls[NestingOlsIdx[i]] - 1 (inclusive of both ends).

[0130] nesting_ols_layer_idx_delta_minus1[i][j] is used to derive the variable NestingOlsLayerIdx[i][j] that specifies the OLS layer index of the j-th layer to which the scalable nested SEI message is applied in the context of the NestingOlsIdx[i]-th OLS when nesting_ols_flag is equal to 1. The value of nesting_ols_layer_idx_delta_minus1[i] should be within the range of 0 to NumLayersInOls[nestingOlsIdx[i]] - 2 (inclusive of both ends).

[0131] The variable NestingOlsLayerIdx[i][j] can be derived as follows. [Number]

[0132] For i in the range of 0 to nesting_num_olss_minus1 (including both ends), the minimum value among all values of LayerIdInOls[NestingOlsIdx[i]][NestingOlsLayerIdx[i][0]] should be equal to the nuh_layer_id of the current SEI NAL unit (e.g., the SEI NAL unit including the scalable nesting SEI message). The nesting_all_layers_flag can be set equal to 1 to specify that the scalable nesting SEI message is generally applied to all layers having a nuh_layer_id greater than or equal to the nuh_layer_id of the current SEI NAL unit. The nesting_all_layers_flag can be set equal to 0 to specify that the scalable nesting SEI message may or may not be applied to all layers having a nuh_layer_id greater than or equal to the nuh_layer_id of the current SEI NAL unit.

[0133] The number obtained by adding 1 to nesting_num_layers_minus1 specifies the number of layers to which the scalable nesting SEI message is generally applied. The value of nesting_num_layers_minus1 should be in the range of 0 to vps_max_layers_minus1 - GeneralLayerIdx[nuh_layer_id] (including both ends), where nuh_layer_id is the nuh_layer_id of the current SEI NAL unit. When nesting_all_layers_flag is equal to 0, nesting_layer_id[i] specifies the nuh_layer_id value of the i-th layer to which the scalable nesting SEI message is generally applied. The value of nesting_layer_id[i] should be greater than nuh_layer_id, where nuh_layer_id is the nuh_layer_id of the current SEI NAL unit.

[0134] When nesting_ols_flag is equal to 1, the variable NestingNumLayers that specifies the number of layers to which the scalable nested SEI message is generally applicable, and the list NestingLayerId[i] for i in the range of 0 to NestingNumLayers - 1 (inclusive) that specifies the list of nuh_layer_id values of the layers to which the scalable nested SEI message is generally applicable are derived as follows, where nuh_layer_id is the nuh_layer_id of the current SEI NAL unit.

Number

[0135] The number obtained by adding 1 to nesting_num_seis_minus1 specifies the number of scalable nested SEI messages. The value of nesting_num_seis_minus1 should be in the range of 0 to 63 (inclusive). nesting_zero_bit should be set equal to 0.

[0136] FIG. 8 is a schematic diagram showing an exemplary video coding device 800. The video coding device 800 is suitable for implementing the disclosed examples / embodiments described herein. The video coding device 800 includes a downstream port 820, an upstream port 850, and / or a transceiver unit (Tx / Rx) 810 including a transmitter and / or a receiver for communicating data upstream and / or downstream via a network. The video coding device 800 also includes a processor 830 including a logic unit and / or a central processing unit (CPU) for processing data, and a memory 832 for storing data. The video coding device 800 may also include electrical, optical, or radio communication components coupled to the upstream port 850 and / or the downstream port 820 for communication of data via an electrical, optical, or radio communication network. The video coding device 800 may also include an input and / or output (I / O) device 860 for communicating data with a user. The I / O device 860 may include output devices such as a display for displaying video data, a speaker for outputting audio data, etc. The I / O device 860 may also include input devices such as a keyboard, a mouse, a trackball, etc., and / or corresponding interfaces for interacting with such output devices.

[0137] Processor 830 is implemented by hardware and software. Processor 830 can be implemented as one or more CPU chips, cores (e.g., as a multi-core processor), field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), and digital signal processors (DSPs). Processor 830 communicates with downstream port 820, Tx / Rx 810, upstream port 850, and memory 832. Processor 830 includes a coding module 814. Coding module 814 implements the disclosed embodiments described herein, such as methods 100, 900, and 1000 that may use multi-layer video sequence 600 and / or bitstream 700. Coding module 814 may also implement any other method / mechanism described herein. Further, coding module 814 may implement codec system 200, encoder 300, decoder 400, and / or HRD 500. For example, coding module 814 may be used to implement HRD. Further, using coding module 814, a scalable nesting SEI message with corresponding flags may be encoded to support clear and concise signaling of scalable nested SEI messages within a scalable nesting SEI message. Thus, coding module 814 may be configured to execute a mechanism for addressing one or more of the problems described above. Thus, coding module 814 provides additional functionality and / or coding efficiency to video coding device 800 when coding video data. In this way, coding module 814 improves the functionality of video coding device 800 and addresses problems specific to video coding techniques. Further, coding module 814 converts video coding device 800 to a different state. Alternatively, coding module 814 can be implemented as instructions stored in memory 832 and executed by processor 830 (e.g., as a computer program product stored on a non-transitory medium).

[0138] Memory 832 includes one or more memory types such as a disk, a tape drive, a solid state drive, a read only memory (ROM), a random access memory (RAM), a flash memory, a ternary content addressable memory (TCAM), a static random access memory (SRAM), etc. Memory 832 is used as an overflow data storage device, can store a program when such a program is selected for execution, and can store instructions and data read during program execution.

[0139] FIG. 9 is a flowchart of an exemplary method 900 for encoding a video sequence into a bitstream, such as bitstream 700 including scalable nesting SEI messages. Method 900 can be used by an encoder such as codec system 200, encoder 300, and / or video coding device 800 when executing method 100. Further, method 900 can operate on HRD 500, and thus, can perform a compliance test on the multi-layer video sequence 600.

[0140] Method 900 can start when an encoder receives a video sequence and decides to encode that video sequence into a multi-layer bitstream, for example, based on user input. At step 901, the encoder encodes the video sequence into one or more layers and encodes the layers into a multi-layer bitstream. A layer can include a set of VCL NAL units having the same layer ID and associated non-VCL NAL units. For example, a layer can include a set of VCL NAL units containing the video data of an encoded picture and any parameter set used to code such a picture. A layer can be included in an OLS. For example, an OLS can include an output layer and any support layers that can be used to decode the output layer according to inter-layer prediction. Thus, an OLS can include sufficient data to decode the representation of the video sequence, for example, with a corresponding image size, SNR, frame rate, etc. Since a video sequence can be coded into several representations, the video sequence can include several layers organized into several OLSs as needed. In this way, the encoder can select an OLS with corresponding layers for transmission to the decoder as required.

[0141] In stage 903, the encoder encodes the SEI message into the bitstream. The SEI message is a syntax structure that contains data not used for decoding. For example, the SEI message may contain data for supporting compliance testing to ensure that the bitstream conforms to the standard. To support simplified signaling when used with a multi-layer bitstream, the SEI message is encoded as a scalable nesting SEI message. The scalable nesting SEI message includes a set of one or more scalable nested SEI messages. The scalable nested SEI messages can be applied to one or more of the OLSs and / or one or more of the layers respectively. To support simplified signaling, the scalable nesting SEI message includes a scalable nesting OLS flag. The scalable nesting OLS flag can be set to specify whether the scalable nested SEI message within the scalable nesting SEI message is applied to a specific OLS or to a specific layer. For example, the scalable nesting OLS flag can be set to 1 when the scalable nested SEI message is specified to be applied to a specific / corresponding OLS (e.g., rather than a layer). As another example, the scalable nesting OLS flag can be set to 0 when the scalable nested SEI message is specified to be applied to a specific / corresponding layer (e.g., rather than an OLS). The scalable nesting SEI message may include several types of scalable nested SEI messages. As a specific example, the scalable nested SEI message may include a buffering period SEI message, a picture timing SEI message, and / or a decode unit information SEI message.The scalable nesting OLS flag can be set to 1 to indicate that the scalable nesting SEI message is applied to a specific OLS (e.g., rather than a layer) when the scalable nesting SEI message contains any SEI message whose payload type is buffering period, picture timing, or decode unit information.

[0142] The scalable nesting SEI message can contain other data to indicate how the corresponding scalable nesting SEI message should be used by the HRD in the encoder. For example, the scalable nesting SEI message can contain a scalable nesting num_olss_minus1 syntax element that specifies the number of OLs to which the corresponding scalable nesting SEI message is applied. The scalable nesting num_olss_minus1 syntax element can be used when the scalable nesting OLS flag is set to 1 to indicate that the scalable nesting SEI message is applied to an OLS. The value of the scalable nesting num_olss_minus1 syntax element can be constrained to stay within the range of 0 to TotalNumOlss - 1 (inclusive). In a similar manner, the scalable nesting SEI message can contain a scalable nesting num_layers_minus1 that specifies the number of layers to which the corresponding scalable nesting SEI message is applied when the scalable nesting OLS flag is set to 0 to indicate that the scalable nesting SEI message is applied to a layer.

[0143] The scalable nesting SEI message may also include the scalable nesting ols_idx_delta_minus1[i] syntax element, which is used to derive the nesting Ols index (NestingOlsIdx[i]) that specifies the Ols index of the i-th OLS to which the scalable nested SEI message is applied when the scalable nesting OLS flag is equal to 1 to indicate that the scalable nested SEI message is applied to OLS. Specifically, the scalable nesting ols_idx_delta_minus1[i] syntax element can be used to specify the OLS corresponding to each scalable nested SEI message. In this way, the scalable nesting num_olss_minus1 can be used to determine the number of OLSs referred to by the scalable nesting SEI message, and the scalable nesting ols_idx_delta_minus1[i] can be used to correlate each scalable nested SEI message with the corresponding OLS. The value of the scalable nesting ols_idx_delta_minus1[i] syntax element can be restricted to stay within the range of 0 to TotalNumOlss - 2 (including both ends). In a specific example, NestingOlsIdx[i] is derived as follows. [Number]

[0144] In a similar manner, the scalable nesting SEI message may include the scalable nesting layer_id[i] when the scalable nesting OLS flag is equal to 0 to indicate that the scalable nested SEI message is applied to the layer. The scalable nesting layer_id[i] specifies the layer ID (e.g., nuh_layer_id) value of the i-th layer to which the scalable nested SEI message is applied.

[0145] In stage 905, the HRD operating in the encoder can perform a set of bitstream compliance tests based on the scalable nesting SEI message. For example, the HRD can read a flag in the scalable nesting SEI message and determine how to interpret the scalable nested SEI messages contained in the scalable nesting SEI message. The HRD can then read the scalable nested SEI messages and determine how to check whether the OLS and / or layers comply with the standard. The HRD can then perform compliance tests on the OLS and / or layers based on the scalable nested SEI messages and / or the corresponding flags in the scalable nesting SEI message. In stage 907, the encoder can store the bitstream for communication to the decoder upon request.

[0146] FIG. 10 is a flowchart of an exemplary method 1000 for decoding a video sequence from a bitstream such as bitstream 700 that includes a scalable nesting SEI message. Method 1000 can be used by a decoder such as codec system 200, decoder 400, and / or video coding device 800 when performing method 100. Further, method 1000 can be used on a multi-layer video sequence 600 whose compliance has been checked by an HRD such as HRD 500.

[0147] Method 1000 may start when a decoder begins to receive a bitstream of coded data representing a multi-layer video sequence, for example, as a result of Method 900. In stage 1001, the decoder receives a bitstream including one or more layers. A layer may include a set of VCL NAL units having the same layer ID and associated non-VCL NAL units. For example, a layer may include a set of VCL NAL units including video data of an encoded picture and any parameter set used to code such a picture. A layer may be included in an OLS. For example, an OLS may include an output layer and any support layers that may be used to decode the output layer according to inter-layer prediction. Thus, an OLS may include sufficient data to decode the representation of the video sequence, for example, with a corresponding image size, SNR, frame rate, etc. Since a video sequence may be coded into several representations, the video sequence may include several layers organized into several OLSs as needed. In this way, the decoder may request and receive a specified OLS with corresponding layers as needed to decode and display a particular representation of the video sequence.

[0148] The bitstream also includes one or more scalable nesting SEI messages. An SEI message is a syntax structure that contains data not used for decoding. For example, an SEI message may contain data for supporting compliance testing to ensure that the bitstream conforms to the standard. To support simplified signaling when used with a multi-layer bitstream, the SEI message is coded in a scalable nesting SEI message. A scalable nesting SEI message includes a set of one or more scalable nested SEI messages. A scalable nested SEI message can be applied to one or more of the OLSs and / or one or more of the layers respectively. To support simplified signaling, the scalable nesting SEI message includes a scalable nesting OLS flag. The scalable nesting OLS flag can be set to specify whether the scalable nested SEI message within the scalable nesting SEI message is applied to a specific OLS or to a specific layer. For example, the scalable nesting OLS flag can be set to 1 when the scalable nested SEI message is specified to be applied to a specific / corresponding OLS (e.g., rather than a layer). As another example, the scalable nesting OLS flag can be set to 0 when the scalable nested SEI message is specified to be applied to a specific / corresponding layer (e.g., rather than an OLS). The scalable nesting SEI message can include several types of scalable nested SEI messages. As a specific example, the scalable nested SEI message can include a buffering period SEI message, a picture timing SEI message, and / or a decode unit information SEI message.The scalable nesting OLS flag can be set to 1 to indicate that a scalable nesting SEI message is applied to a specific OLS (e.g., rather than a layer) when the scalable nesting SEI message contains any SEI message whose payload type is buffering period, picture timing, or decoder unit information.

[0149] The scalable nesting SEI message may contain other data to indicate how the corresponding scalable nesting SEI message should be used by the HRD in the encoder. For example, the scalable nesting SEI message may contain a scalable nesting num_olss_minus1 syntax element that specifies the number of OLSs to which the corresponding scalable nesting SEI message is applied. The scalable nesting num_olss_minus1 syntax element can be used when the scalable nesting OLS flag is set to 1 to indicate that the scalable nesting SEI message is applied to an OLS. The value of the scalable nesting num_olss_minus1 syntax element can be constrained to stay within the range of 0 to TotalNumOlss - 1 (inclusive). In a similar manner, the scalable nesting SEI message may contain a scalable nesting num_layers_minus1 that specifies the number of layers to which the corresponding scalable nesting SEI message is applied when the scalable nesting OLS flag is set to 0 to indicate that the scalable nesting SEI message is applied to a layer.

[0150] The scalable nesting SEI message may also include a scalable nesting ols_idx_delta_minus1[i] syntax element, which is used to derive the nesting Ols index (NestingOlsIdx[i]) that specifies the Ols index of the i-th Ols to which the scalable nested SEI message is applied when the scalable nesting Ols flag is equal to 1 to indicate that the scalable nested SEI message is applied to OLS. Specifically, the scalable nesting ols_idx_delta_minus1[i] syntax element can be used to specify the OLS corresponding to each scalable nested SEI message. In this way, the number of OLSs referred to by the scalable nesting SEI message can be determined using scalable nesting num_olss_minus1, and each scalable nested SEI message can be correlated with the corresponding OLS using scalable nesting ols_idx_delta_minus1[i]. The value of the scalable nesting ols_idx_delta_minus1[i] syntax element can be constrained to stay within the range of 0 to TotalNumOlss-2 (including both ends). In a specific example, NestingOlsIdx[i] is derived as follows.

Number

[0151] In a similar manner, the scalable nesting SEI message may include scalable nesting layer_id[i] when the scalable nesting Ols flag is equal to 0 to indicate that the scalable nested SEI message is applied to the layer. The scalable nesting layer_id[i] specifies the layer ID (e.g., nuh_layer_id) value of the i-th layer to which the scalable nested SEI message is applied.

[0152] In stage 1003, the decoder may decode coded pictures from one or more layers based on scalable nested SEI messages to generate decoded pictures. For example, the presence of a scalable nested SEI message may indicate that the bitstream has been checked by the HRD in the encoder and thus conforms to the standard. Thus, the presence of a scalable nested SEI message indicates that the bitstream can be decoded. In stage 1005, the decoder may transfer the decoded pictures for display as part of the decoded video sequence.

[0153] FIG. 11 is a schematic diagram of an exemplary system 1100 for coding a video sequence using a bitstream that includes scalable nested SEI messages. System 1100 may be implemented by an encoder and decoder such as codec system 200, encoder 300, decoder 400, and / or video coding device 800. Further, system 1100 may use HRD 500 to perform a compliance test on multi-layer video sequence 600 and / or bitstream 700. Additionally, system 1100 may be used when implementing methods 100, 900, and / or 1000.

[0154] System 1100 includes a video encoder 1102. The video encoder 1102 includes an encoding module 1103 for encoding a bitstream having one or more layers. The encoding module 1103 is further for encoding a scalable nesting supplementary enhancement information (SEI) message into the bitstream, and the scalable nesting SEI message includes one or more scalable nested SEI messages and a scalable nesting output layer set (OLS) flag, and the scalable nesting OLS flag is set to specify whether the scalable nested SEI message is applied to a specific OLS or a specific layer. The video encoder 1102 further includes an HRD module 1105 that performs a set of bitstream compliance tests based on the scalable nesting SEI message. The video encoder 1102 further includes a storage module 1106 for storing the bitstream for communication to a decoder. The video encoder 1102 further includes a transmission module 1107 for transmitting the bitstream towards a video decoder 1110. The video encoder 1102 can further be configured to perform any of the steps of method 900.

[0155] System 1100 also includes a video decoder 1110. The video decoder 1110 includes a receiving module 1111 that receives a bitstream including one or more layers and scalable nesting supplementary enhancement information (SEI) messages, where the scalable nesting SEI messages include one or more scalable nested SEI messages and a scalable nesting output layer set (OLS) flag, and the scalable nesting OLS flag is set to specify whether the scalable nested SEI messages are applicable to a specific OLS or a specific layer. The video decoder 1110 further includes a decoding module 1113 that decodes coded pictures from one or more layers based on the scalable nested SEI messages to generate decoded pictures. The video decoder 1110 further includes a transfer module 1115 for transferring the decoded pictures for display as part of the decoded video sequence. The video decoder 1110 can further be configured to perform any of the steps of method 1000.

[0156] If there are no intervening components between a first component and a second component other than a line, trace, or another medium, the first component is directly coupled to the second component. If there are intervening components between the first component and the second component other than a line, trace, or another medium, the first component is indirectly coupled to the second component. The term "coupled" and its variations include both those that are directly coupled and those that are indirectly coupled. The use of the term "about" means a range that includes ±10% of the subsequent number unless otherwise specified.

[0157] The steps of the exemplary methods described herein need not be performed in the order described, and it should also be understood that the order of such method steps is merely exemplary. Similarly, in methods consistent with various embodiments of the present disclosure, additional steps can be included in such methods, and specific steps can be omitted or combined.

[0158] Although several embodiments have been provided in the present disclosure, it will be understood that the disclosed systems and methods may be embodied in many other specific forms without departing from the spirit or scope of the present disclosure. The present examples are to be considered as illustrative and not restrictive, and the intention should not be limited to the details given herein. For example, various elements or components may be combined or integrated into another system, or certain features may be omitted or not implemented.

[0159] In addition, without departing from the scope of the present disclosure, the techniques, systems, subsystems, and methods described and illustrated as separate or discrete in various embodiments may be combined with or integrated into other systems, components, techniques, or methods. Other examples of changes, substitutions, and alterations can be ascertained by those skilled in the art and can be made without departing from the spirit and scope disclosed herein. [Other possible items] (Item 1) A method implemented by a decoder, Receiving, by a receiver of the decoder, a bitstream including one or more layers and scalable nesting supplementary enhancement information (SEI) messages, wherein the scalable nesting SEI message includes one or more scalable nested SEI messages and a scalable nesting output layer set (OLS) flag, and the scalable nesting OLS flag is set to specify whether the scalable nested SEI message is applied to a specific OLS or a specific layer; Decoding, by a processor of the decoder, coded pictures from the one or more layers based on the scalable nested SEI messages to generate decoded pictures; A method comprising the above steps. (Item 2) The method according to Item 1, wherein the scalable nesting OLS flag is set to 1 when the scalable nested SEI message is specified to be applied to a specific OLS, and the scalable nesting OLS flag is set to 0 when the scalable nested SEI message is specified to be applied to a specific layer. (Item 3) The method according to Item 1 or 2, wherein the scalable nesting OLS flag is set to 1 when the scalable nesting SEI message includes an SEI message whose payload type is buffering period, picture timing, or decode unit information. (Item 4) When the scalable nesting SEI message has the scalable nesting OLS flag set to 1, it includes syntax elements of a number (num_olss_minus1) obtained by subtracting 1 from the number of scalable nestings of OLS. The scalable nesting num_olss_minus1 syntax element specifies the number of OLSs to which the scalable nested SEI message is applied, and the value of the scalable nesting num_olss_minus1 syntax element is within the range of 0 to the total number of OLSs (TotalNumOlss) - 1 (including both ends). The method according to any one of items 1 to 3. (Item 5) When the scalable nesting SEI message is equal to 1 for the scalable nesting OLS flag, it includes syntax elements of a number (ols_idx_delta_minus1[i]) obtained by subtracting 1 from the scalable nesting OLS delta used to derive the nesting OLS index (NestingOlsIdx[i]) of the i-th OLS to which the scalable nested SEI message is applied. The value of the scalable nesting ols_idx_delta_minus1[i] syntax element is within the range of 0 to the TotalNumOlss - 2 (including both ends). The method according to any one of items 1 to 4. (Item 6) The method according to any one of items 1 to 5, further including the step of deriving NestingOlsIdx[i] as follows: [Number] (Item 7) When the scalable nesting SEI message has the scalable nesting OLS flag set to 0, it includes syntax elements of the number obtained by subtracting 1 from the scalable nesting number of layers (num_layers_minus1), and the scalable nesting num_layers_minus1 syntax element specifies the number of layers to which the scalable nested SEI message is applied, according to the method described in any of Items 1 to 6. (Item 8) A method implemented by an encoder, encoding, by a processor of the encoder, a bitstream including one or more layers; encoding, by the processor, a scalable nesting supplementary enhancement information (SEI) message into the bitstream, where the scalable nesting SEI message includes one or more scalable nested SEI messages and a scalable nesting output layer set (OLS) flag, and the scalable nesting OLS flag is set to specify whether the scalable nested SEI message is applied to a specific OLS or a specific layer; executing, by the processor, a set of bitstream conformity tests based on the scalable nesting SEI message; storing, by a memory coupled to the processor, the bitstream for communication to a decoder; A method comprising the above steps. (Item 9) The scalable nesting OLS flag is set to 1 when it specifies that the scalable nested SEI message is applied to a specific OLS, and the scalable nesting OLS flag is set to 0 when it specifies that the scalable nested SEI message is applied to a specific layer, according to the method described in Item 8. (Item 10) If the scalable nesting SEI message includes an SEI message whose payload type is buffering period, picture timing, or decode unit information, the scalable nesting OLS flag is set to 1, as described in item 8 or 9. (Item 11) When the scalable nesting SEI message has the scalable nesting OLS flag set to 1, it includes syntax elements of num_olss_minus1, which is the number obtained by subtracting 1 from the number of scalable nestings of OLS. The scalable nesting num_olss_minus1 syntax element specifies the number of OLSs to which the scalable nested SEI message is applied, and the value of the scalable nesting num_olss_minus1 syntax element is within the range of 0 to TotalNumOlss - 1 (both ends included), as described in any of items 8 to 10. (Item 12) When the scalable nesting SEI message has the scalable nesting OLS flag equal to 1, it includes syntax elements of ols_idx_delta_minus1[i], which is the number obtained by subtracting 1 from the scalable nesting OLS delta used to derive the nesting OLS index (NestingOlsIdx[i]) of the i-th OLS to which the scalable nested SEI message is applied. The value of the scalable nesting ols_idx_delta_minus1[i] syntax element is within the range of 0 to TotalNumOlss - 2 (both ends included), as described in any of items 8 to 11. (Item 13) The method described in any of items 8 to 12 further includes the step of deriving NestingOlsIdx[i] as follows:

Number

Claims

1. 1. A method implemented by a decoder, the method comprising: receiving a bitstream including one or more layers and a scalable nesting supplemental enhancement information (SEI) message, the scalable nesting SEI message including one or more scalable nested SEI messages and a scalable nesting output layer set (OLS) flag, the scalable nesting OLS flag being set to specify whether the scalable nested SEI message applies to a particular OLS or a particular layer, the scalable nesting OLS flag being set to specify whether the scalable nested SEI message applies to a particular OLS or a particular layer, If the scalable nesting OLS flag is set to 1, the SEI message includes a scalable nesting OLS number minus 1 (num_olss_minus1) syntax element, which specifies the number of OLSs to which the scalable nested SEI message applies, and the value of the scalable nesting num_olss_minus1 syntax element is in the range of 0 to a total number of OLSs (TotalNumOlss) minus 1 (inclusive); deriving a nesting OLS index (NestingOlsIdx[i]) based on a scalable nesting OLS index delta minus one (ols_idx_delta_minus1[i]) syntax element included in the scalable nesting SEI message, where the NestingOlsIdx[i] specifies an OLS index of an ith OLS to which the scalable nested SEI message applies if the scalable nesting OLS flag is equal to 1, a value of the scalable nesting ols_idx_delta_minus1[i] syntax element is in the range of 0 to the TotalNumOlss-2, inclusive, and NestingOlsIdx[i] is ##EQU00011## The steps are derived as follows: decoding a coded picture from the one or more layers by applying the scalable nested SEI message to the OLS specified by the nesting OLS index to generate a decoded picture; A method for providing the above.

2. When the scalable nested SEI message specifies that it applies to a specific OLS, the scalable nesting OLS flag is set to 1, and when the scalable nested SEI message specifies that it applies to a specific layer, the scalable nesting OLS flag is set to 0. The method of claim 1.

3. If the scalable nesting SEI message includes an SEI message with a payload type of buffering period, picture timing, or decode unit information, the scalable nesting OLS flag is set to 1.

3. The method according to claim 1 .

4. If the scalable nesting OLS flag is set to 0, the scalable nesting SEI message includes a scalable nesting layer number minus 1 (num_layers_minus1) syntax element, which specifies the number of layers to which the scalable nested SEI message applies, and the scalable nested SEI message applies to the layer.

4. The method according to any one of claims 1 to 3.

5. 1. A method implemented by an encoder, the method comprising: encoding a bitstream comprising one or more layers; encoding a scalable nesting supplemental enhancement information (SEI) message into the bitstream, the scalable nesting SEI message including one or more scalable nested SEI messages and a scalable nesting output layer set (OLS) flag, the scalable nesting OLS flag being set to specify whether the scalable nested SEI message applies to a particular OLS or a particular layer, If the scalable nesting OLS flag is set to 1, the message includes a scalable nesting OLS number minus 1 (num_olss_minus1) syntax element, which specifies the number of OLSs to which the scalable nested SEI message applies, and a value of the scalable nesting num_olss_minus1 syntax element is in the range of 0 to a total number of OLSs (TotalNumOlss)-1, inclusive; deriving a nesting OLS index (NestingOlsIdx[i]) based on a scalable nesting OLS index delta minus one (ols_idx_delta_minus1[i]) syntax element included in the scalable nesting SEI message, where the NestingOlsIdx[i] specifies an OLS index of an ith OLS to which the scalable nested SEI message applies if the scalable nesting OLS flag is equal to 1, a value of the scalable nesting ols_idx_delta_minus1[i] syntax element is in the range of 0 to the TotalNumOlss-2, inclusive, and NestingOlsIdx[i] is ##EQU00012## The steps are derived as follows: applying the scalable nested SEI message to the OLS specified by the nesting OLS index; storing the bitstream for communication to a decoder; A method for providing the above.

6. When the scalable nested SEI message specifies that it applies to a specific OLS, the scalable nesting OLS flag is set to 1, and when the scalable nested SEI message specifies that it applies to a specific layer, the scalable nesting OLS flag is set to 0. The method according to claim 5.

7. If the scalable nesting SEI message includes an SEI message with a payload type of buffering period, picture timing, or decode unit information, the scalable nesting OLS flag is set to 1. The method according to claim 5 or 6.

8. If the scalable nesting OLS flag is set to 0, the scalable nesting SEI message includes a scalable nesting layer number minus 1 (num_layers_minus1) syntax element, which specifies the number of layers to which the scalable nested SEI message applies, and the scalable nested SEI message applies to the layer.

8. The method according to any one of claims 5 to 7.

9. A method for implementing a method according to any one of claims 1 to 4, comprising: a processor; a receiver coupled to the processor; a memory coupled to the processor; and a transmitter coupled to the processor, the processor, the receiver, the memory, and the transmitter being configured to perform the method according to any one of claims 1 to 4. Video coding device.

10. encoding means for carrying out the method according to any one of claims 5 to 8; storage means for storing said bitstream for communication to a decoder; An encoder comprising:

11. 1. A method for storing a bitstream, comprising the steps of: receiving a bitstream including one or more layers and a scalable nesting supplemental enhancement information (SEI) message, the scalable nesting SEI message including one or more scalable nested SEI messages and a scalable nesting output layer set (OLS) flag, the scalable nesting OLS flag being set to specify whether the scalable nested SEI message applies to a particular OLS or a particular layer; the scalable nesting SEI message includes a scalable nesting OLS number minus 1 (num_olss_minus1) syntax element when the scalable nesting OLS flag is set to 1, the scalable nesting num_olss_minus1 syntax element specifying the number of OLSs to which the scalable nested SEI message applies, and a value of the scalable nesting num_olss_minus1 syntax element is in the range of 0 to a total number of OLSs (TotalNumOlss)-1 (inclusive); storing the bitstream; Equipped with The bitstream is transmitted to a coding device. derive a nesting OLS index (NestingOlsIdx[i]) based on a scalable nesting OLS index delta minus one (ols_idx_delta_minus1[i]) syntax element included in the scalable nesting SEI message, where the NestingOlsIdx[i] specifies an OLS index of the ith OLS to which the scalable nested SEI message applies if the scalable nesting OLS flag is equal to 1, a value of the scalable nesting ols_idx_delta_minus1[i] syntax element is in the range from 0 to the TotalNumOlss-2, inclusive, and NestingOlsIdx[i] is ##EQU00013## and apply the scalable nested SEI message to the OLS specified by the nesting OLS index; How it is configured.

Citation Information

Patent Citations

  • Identification of operation points applicable to nested SEI message in video coding

    US20140098894A1

  • Generic use of HEVC SEI messages for multi-layer codecs

    US20150271529A1

  • Method and apparatus for video coding and decoding

    WO2015104451A1