Scalable Nesting for Suffix SEI Messages

By allowing suffix SEI messages to be used in scalable nesting SEI messages, the solution improves coding efficiency and reduces resource usage in video coding systems, addressing the restriction of certain SEI messages in existing systems.

JP7721511B2Active Publication Date: 2025-08-12HUAWEI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2022518798
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-09-24
Filing Date
2020-09-17
Publication Date
2025-08-12
Estimated Expiration
2040-09-17

Smart Images

  • Figure 0007721511000003
    Figure 0007721511000003
  • Figure 0007721511000004
    Figure 0007721511000004
  • Figure 0007721511000005
    Figure 0007721511000005
Patent Text Reader

Abstract

A video coding mechanism is disclosed. The mechanism includes encoding one or more coded pictures into a bitstream. Also encoded into the bitstream is a supplemental enhancement information (SEI) network abstraction layer (NAL) unit having a NAL unit type (nal_unit_type) equal to a suffix SEI NAL unit type (SUFFIX_SEI_NUT). The SEI NAL unit includes a scalable nesting SEI message. A set of bitstream conformance tests is performed on the bitstream based on the scalable nesting SEI message. The bitstream is stored for communication to a decoder.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This patent application claims the benefit of U.S. Provisional Patent Application No. 62 / 905,236, entitled "Video Coding Improvements," filed September 24, 2019 by Ye-Kui Wang, which is incorporated herein by reference.

[0002] FIELD This disclosure relates generally to video coding, and more particularly to improved signaling parameters to support coding of multiple layer bitstreams. [Background technology]

[0003] The significant amount of video data required to represent even a relatively short video can present challenges when streaming data or transmitting data across communication networks with limited bandwidth capacity. For this reason, video data is typically compressed before transmission over today's communication networks. When video is stored on a storage device, the size of the video can also be an issue because memory resources may be scarce. Video compression devices often use software and / or hardware to code video data at the source prior to transmission or storage, thereby reducing the amount of data required to represent a digital video image. The compressed data is then received by a video decompressor, which decodes the video data at the destination. Due to limited network resources and increasing demand for higher video quality, improved compression and decompression techniques that increase compression ratios with little or no sacrifice in image quality are desirable. Summary of the Invention [Means for solving the problem]

[0004] In one embodiment, the present disclosure includes a method implemented by a decoder, the method including: receiving, by a receiver of the decoder, a bitstream including a coded picture and a supplemental enhancement information (SEI) network abstraction layer (NAL) unit having a NAL unit type (nal_unit_type) equal to a suffix SEI NAL unit type (SUFFIX_SEI_NUT) and including a scalable nesting SEI message; and decoding, by a processor of the decoder, the coded picture to generate a decoded picture.

[0005] Some video coding systems use SEI messages. SEI messages contain information that is not required by the decoding process to determine the values of samples in a decoded picture. For example, the SEI message may contain parameters used to check the bitstream for standard compliance. Furthermore, video coding systems can encode pictures into multiple layers and / or output layer sets (OLSs). Scalable nesting SEI messages can be used to correlate prefix SEI messages to layers and / or OLSs. Some video coding systems also use suffix SEI messages, but scalable nesting SEI messages cannot be used in suffix SEI messages. Various types of SEI messages can be included as either prefix SEI messages or suffix SEI messages. However, certain types of SEI messages, such as the decoded picture hash SEI message, are restricted to use as suffix SEI messages. Therefore, certain SEI messages, such as the decoded picture hash SEI message, cannot be coded in scalable nesting SEI messages in such systems. This example includes a mechanism for encoding certain SEI messages, such as a decoded picture hash SEI message, in a scalable nesting SEI message. Specifically, a suffix SEI message is enabled to be used in conjunction with a scalable nesting SEI message. In other words, an SEI NAL unit with a nal_unit_type equal to SUFFIX_SEI_NUT may include a scalable nesting SEI message. In this way, a suffix SEI message, such as a decoded picture hash SEI message, can be used as a scalable nesting SEI message and / or as a scalably nested SEI message included in a scalable nesting SEI message. This allows the suffix SEI message to be used in conjunction with a scalable nesting SEI message while applying to a specified layer and / or OLS. Messages can be nested, which improves the performance of the encoder and decoder. Furthermore, coding efficiency can be improved, which reduces the use of processor, memory, and / or network signaling resources in both the encoder and decoder.

[0006] Optionally, in any of the above aspects, another embodiment of the above aspect is provided, wherein the scalable nesting SEI message includes one or more scalably nested SEI messages.

[0007] Optionally, in any of the aforementioned aspects, another embodiment of the aspect is provided, wherein the one or more scalably nested SEI messages include a decoded picture hash SEI message.

[0008] Optionally, in any of the above aspects, another embodiment of the aspect is provided, wherein the scalable nesting SEI message associates the SEI message with a particular OLS.

[0009] Optionally, in any of the above aspects, another embodiment of the aspect is provided, wherein the scalable nesting SEI message associates the SEI message with a particular layer.

[0010] Optionally, in any of the above aspects, another embodiment of the aspect is provided, wherein the scalable nesting SEI message includes a payloadType set to 133.

[0011] Optionally, in any of the above aspects, another implementation of the aspect is provided, wherein the scalable nesting SEI message includes a scalable nesting layer identifier (layer_id[i]) that specifies the NAL unit header layer identifier (nuh_layer_id) value of the i-th layer to which the scalably nested SEI message applies.

[0012] In one embodiment, the present disclosure includes a method implemented by an encoder, the method including: encoding, by a processor of the encoder, one or more coded pictures into a bitstream; encoding, by the processor, an SEI NAL unit having a nal_unit_type equal to SUFFIX_SEI_NUT and including a scalable nesting SEI message into the bitstream; performing, by the processor, a set of bitstream conformance tests on the bitstream based on the scalable nesting SEI message; and storing, by a memory coupled to the processor, the bitstream for communication to a decoder.

[0013] Some video coding systems use SEI messages. SEI messages contain information that is not required by the decoding process to determine the values of samples in a decoded picture. For example, the SEI message may contain parameters used to check the bitstream for standard compliance. Furthermore, video coding systems can encode pictures into multiple layers and / or output layer sets (OLSs). Scalable nesting SEI messages can be used to correlate prefix SEI messages to layers and / or OLSs. Some video coding systems also use suffix SEI messages, but scalable nesting SEI messages cannot be used in suffix SEI messages. Various types of SEI messages can be included as either prefix SEI messages or suffix SEI messages. However, certain types of SEI messages, such as the decoded picture hash SEI message, are restricted to use as suffix SEI messages. Therefore, certain SEI messages, such as the decoded picture hash SEI message, cannot be coded in scalable nesting SEI messages in such systems. This example includes a mechanism for encoding certain SEI messages, such as a decoded picture hash SEI message, in a scalable nesting SEI message. Specifically, a suffix SEI message is enabled to be used in conjunction with a scalable nesting SEI message. In other words, an SEI NAL unit with a nal_unit_type equal to SUFFIX_SEI_NUT may include a scalable nesting SEI message. In this way, a suffix SEI message, such as a decoded picture hash SEI message, can be used as a scalable nesting SEI message and / or as a scalably nested SEI message included in a scalable nesting SEI message. This allows the suffix SEI message to be used in conjunction with a scalable nesting SEI message while applying to a specified layer and / or OLS. Messages can be nested, which improves the performance of the encoder and decoder. Furthermore, coding efficiency can be improved, which reduces the use of processor, memory, and / or network signaling resources in both the encoder and decoder.

[0014] Optionally, in any of the above aspects, another embodiment of the above aspect is provided, wherein the scalable nesting SEI message includes one or more scalably nested SEI messages.

[0015] Optionally, in any of the aforementioned aspects, another embodiment of the aspect is provided, wherein the one or more scalably nested SEI messages include a decoded picture hash SEI message.

[0016] Optionally, in any of the above aspects, another embodiment of the aspect is provided, wherein the scalable nesting SEI message associates the SEI message with a particular OLS.

[0017] Optionally, in any of the above aspects, another embodiment of the aspect is provided, wherein the scalable nesting SEI message associates the SEI message with a particular layer.

[0018] Optionally, in any of the aforementioned aspects, another embodiment of the aspect is provided, wherein the scalable nesting SEI message includes a payloadType set to 133.

[0019] Optionally, in any of the aforementioned aspects, another embodiment of the aspect is provided, wherein the scalable nesting SEI message includes a scalable nesting layer_id[i] that specifies the nuh_layer_id value of the i-th layer to which the scalably nested SEI message applies.

[0020] In one embodiment, the present disclosure includes a video coding apparatus comprising a processor, a receiver coupled to the processor, a memory coupled to the processor, and a transmitter coupled to the processor, wherein the processor, receiver, memory, and transmitter are configured to perform the method of the aforementioned aspect.

[0021] In one embodiment, the present disclosure includes a non-transitory computer-readable medium including a computer program product for use by a video coding device, the computer program product including computer-executable instructions stored on the non-transitory computer-readable medium that, when executed by a processor, cause the video coding device to perform the method of any of the aforementioned aspects.

[0022] In one embodiment, the present disclosure includes a decoder comprising: receiving means for receiving a bitstream including coded pictures and SEI NAL units having a nal_unit_type equal to SUFFIX_SEI_NUT and including a scalable nesting SEI message; decoding means for decoding the coded pictures to generate decoded pictures; and forwarding means for forwarding the decoded pictures for display as part of a decoded video sequence.

[0023] Some video coding systems use SEI messages. SEI messages contain information that is not required by the decoding process to determine the values of samples in a decoded picture. For example, the SEI message may contain parameters used to check the bitstream for standard compliance. Furthermore, video coding systems can encode pictures into multiple layers and / or output layer sets (OLSs). Scalable nesting SEI messages can be used to correlate prefix SEI messages to layers and / or OLSs. Some video coding systems also use suffix SEI messages, but scalable nesting SEI messages cannot be used in suffix SEI messages. Various types of SEI messages can be included as either prefix SEI messages or suffix SEI messages. However, certain types of SEI messages, such as the decoded picture hash SEI message, are restricted to use as suffix SEI messages. Therefore, certain SEI messages, such as the decoded picture hash SEI message, cannot be coded in scalable nesting SEI messages in such systems. This example includes a mechanism for encoding certain SEI messages, such as a decoded picture hash SEI message, in a scalable nesting SEI message. Specifically, a suffix SEI message is enabled to be used in conjunction with a scalable nesting SEI message. In other words, an SEI NAL unit with a nal_unit_type equal to SUFFIX_SEI_NUT may include a scalable nesting SEI message. In this way, a suffix SEI message, such as a decoded picture hash SEI message, can be used as a scalable nesting SEI message and / or as a scalably nested SEI message included in a scalable nesting SEI message. This allows the suffix SEI message to be used in conjunction with a scalable nesting SEI message while applying to a specified layer and / or OLS. Messages can be nested, which improves the performance of the encoder and decoder. Furthermore, coding efficiency can be improved, which reduces the use of processor, memory, and / or network signaling resources in both the encoder and decoder.

[0024] Optionally, in any of the preceding aspects, another implementation of the aspect is provided, wherein the decoder is further configured to perform the method of any of the preceding aspects.

[0025] In one embodiment, the present disclosure includes an encoder, the encoder including encoding means for encoding one or more coded pictures into a bitstream and encoding an SEI NAL unit having a nal_unit_type equal to SUFFIX_SEI_NUT and including a scalable nesting SEI message into the bitstream; HRD means for performing a set of bitstream conformance tests on the bitstream based on the scalable nesting SEI message; and storage means for storing the bitstream for communication to a decoder.

[0026] Some video coding systems use SEI messages. SEI messages contain information that is not required by the decoding process to determine the values of samples in a decoded picture. For example, the SEI message may contain parameters used to check the bitstream for standard compliance. Furthermore, video coding systems can encode pictures into multiple layers and / or output layer sets (OLSs). Scalable nesting SEI messages can be used to correlate prefix SEI messages to layers and / or OLSs. Some video coding systems also use suffix SEI messages, but scalable nesting SEI messages cannot be used in suffix SEI messages. Various types of SEI messages can be included as either prefix SEI messages or suffix SEI messages. However, certain types of SEI messages, such as the decoded picture hash SEI message, are restricted to use as suffix SEI messages. Therefore, certain SEI messages, such as the decoded picture hash SEI message, cannot be coded in scalable nesting SEI messages in such systems. This example includes a mechanism for encoding certain SEI messages, such as a decoded picture hash SEI message, in a scalable nesting SEI message. Specifically, a suffix SEI message is enabled to be used in conjunction with a scalable nesting SEI message. In other words, an SEI NAL unit with a nal_unit_type equal to SUFFIX_SEI_NUT may include a scalable nesting SEI message. In this way, a suffix SEI message, such as a decoded picture hash SEI message, can be used as a scalable nesting SEI message and / or as a scalably nested SEI message included in a scalable nesting SEI message. This allows the suffix SEI message to be used in conjunction with a scalable nesting SEI message while applying to a specified layer and / or OLS. Messages can be nested, which improves the performance of the encoder and decoder. Furthermore, coding efficiency can be improved, which reduces the use of processor, memory, and / or network signaling resources in both the encoder and decoder.

[0027] Optionally, in any of the preceding aspects, another embodiment of the aspect is provided, wherein the encoder is further configured to perform the method of any of the preceding aspects.

[0028] For clarity, any one of the above-described embodiments may be combined with any one or more of the other above-described embodiments to create new embodiments within the scope of the present disclosure.

[0029] These and other features will be more clearly understood from the following detailed description taken in conjunction with the accompanying drawings and claims.

[0030] For a more complete understanding of the present disclosure, reference is now made to the following brief description taken in conjunction with the accompanying drawings and detailed description, in which like reference numerals represent like parts. [Brief explanation of the drawings]

[0031] [Figure 1] 1 is a flowchart of an exemplary method for coding a video signal. [Figure 2] 1 is a schematic diagram of an example coding-decoding (codec) system for video coding. [Figure 3] FIG. 1 is a schematic diagram illustrating an exemplary video encoder. [Figure 4] FIG. 1 is a schematic diagram illustrating an exemplary video decoder. [Figure 5] 1 is a schematic diagram illustrating an exemplary hypothetical reference decoder (HRD). [Figure 6]FIG. 1 is a schematic diagram illustrating an exemplary multi-layer video sequence. [Figure 7] FIG. 2 is a schematic diagram illustrating an exemplary bitstream. [Figure 8] 1 is a schematic diagram of an exemplary video coding device; [Figure 9] 1 is a flowchart of an example method for encoding a video sequence into a bitstream by applying a suffix supplemental enhancement information (SEI) message that includes a scalable nesting SEI message. [Figure 10] 10 is a flowchart of an example method for decoding a video sequence from a bitstream that uses a suffix SEI message that includes a scalable nesting SEI message. [Figure 11] FIG. 1 is a schematic diagram of an example system for coding a video sequence using a bitstream that uses a suffix SEI message that includes a scalable nesting SEI message. DETAILED DESCRIPTION OF THE INVENTION

[0032] Initially, exemplary implementations of one or more embodiments are provided below, but it should be understood that the disclosed systems and / or methods may be implemented using any number of technologies, whether currently known or in existence. The present disclosure should in no way be limited to the exemplary implementations, drawings, and techniques shown below, including the exemplary designs and implementations shown or described herein, but may be modified within the scope of the appended claims, along with their full range of equivalents.

[0033] The following terms are defined as follows, unless used in the context to the contrary herein. Specifically, the following definitions are intended to further clarify the present disclosure. However, terms may be described differently in different contexts. Therefore, the following definitions should be considered supplementary and should not be considered limiting of other definitions given to such terms herein.

[0034] A bitstream is a sequence of bits containing compressed video data for transmission between an encoder and a decoder. An encoder is a device configured to use an encoding process to compress video data into a bitstream. A decoder is a device configured to use a decoding process to recover video data from a bitstream for display. A picture is an array of luma samples and / or chroma samples that generate a frame or its fields. A slice is an integer number of complete tiles contained exclusively in a single Network Abstraction Layer (NAL) unit, or an integer number of consecutive complete coding tree unit (CTU) rows (e.g., within a tile) of a picture. For clarity, an encoded or decoded picture can be referred to as the current picture. A coded picture is a coded representation of a picture comprising a video coding layer (VCL) NAL unit with a particular value of the NAL unit header layer identifier (nuh_layer_id) within an access unit (AU), including all coding tree units (CTUs) of the picture. A decoded picture is a picture produced by applying the decoding process to a coded picture.

[0035] An AU is a set of coded pictures contained in different layers and associated with the same time for output from the decoded picture buffer (DPB). A NAL unit is a syntax structure containing data in the form of a raw byte sequence payload (RBSP), an indication of the type of data, and optionally interspersed with emulation prevention bytes. A VCL NAL unit is a NAL unit coded to contain video data, such as a coded slice of a picture. A non-VCL NAL unit is a NAL unit containing non-video data, such as syntax and / or parameters that support decoding the video data, performing conformance checks, or other operations. A NAL unit type (nal_unit_type) is a syntax element contained in a NAL unit that indicates the type of data contained in the NAL unit. A layer is a set of VCL NAL units that share specified characteristics (e.g., a common resolution, frame rate, picture size, etc.), as indicated by the layer ID and the associated non-VCL NAL unit. nuh_layer_id is a syntax element that specifies the identifier of the layer containing the NAL unit. An output layer set (OLS) is a collection of layers where one or more layers are designated as output layers.

[0036] A hypothetical reference decoder (HRD) is a decoder model that runs on an encoder and checks the variability of the bitstream produced by the encoding process to verify conformance with specified constraints. Bitstream conformance tests are tests to determine whether the encoded bitstream conforms to standards such as Versatile Video Coding (VVC). HRD parameters are syntax elements that initialize and / or define the operating conditions of the HRD. HRD parameters may be included in supplemental enhancement information (SEI) messages and / or video parameter sets (VPS).

[0037] An SEI message is a syntax structure with specific semantics that conveys information that is not required by the decoding process to determine the values of samples in the decoded picture. An SEI NAL unit is a NAL unit that contains one or more SEI messages. A particular SEI NAL unit may be referred to as the current SEI NAL unit. A scalable nesting SEI message is a message that contains multiple SEI messages corresponding to one or more output layer sets (OLSs) or one or more layers. A prefix SEI message is an SEI message that applies to one or more subsequent NAL units. The prefix SEI NAL unit type (PREFIX_SEI_NUT) indicates that the corresponding SEI message is a prefix SEI message. A suffix SEI message is an SEI message that applies to one or more preceding NAL units. The suffix SEI NAL unit type (SUFFIX_SEI_NUT) indicates that the corresponding SEI message is a suffix SEI message. The payload type (payloadType) is a syntax element that represents the type of data contained in an SEI message and indicates the type of SEI message contained in the SEI NAL unit. The buffering period (BP) SEI message is a type of SEI message that contains HRD parameters for initializing the HRD to manage the coded picture buffer (CPB). The picture timing (PT) SEI message is a type of SEI message that contains HRD parameters for managing delivery information for AUs in the CPB and / or decoded picture buffer (DPB). The decoded unit information (DUI) SEI message is a type of SEI message that contains HRD parameters for managing delivery information for DUs in the CPB and / or DPB. The decoded picture hash SEI message is a type of SEI message that contains a checksum derived from sample values of the decoded picture.The decoded picture hash SEI message can be used to detect whether a picture has been correctly received and decoded at a decoder. A scalable nesting SEI message is a set of scalably nested SEI messages. A scalably nested SEI message is an SEI message nested within a scalable nesting SEI message. The scalable nesting layer id (layer_id[i]) is a syntax element in a scalable nesting SEI message that specifies the nuh_layer_id value of the ith layer to which the scalably nested SEI message applies.

[0038] A picture parameter set (PPS) is a syntax structure containing syntax elements that apply to the entire coded picture, determined by the syntax elements found in each picture header. A picture header is a syntax structure containing syntax elements that apply to all slices of a coded picture. A slice header is part of a coded slice that contains data elements for all tiles or CTU rows within the tiles represented in the slice. A coded video sequence is a set of one or more coded pictures. A decoded video sequence is a set of one or more decoded pictures.

[0039] The following acronyms are used herein: Access Unit (AU), Coding Tree Block (CTB), Coding Tree Unit (CTU), Coding Unit (CU), Coded Layer Video Sequence (CLVS), Coded Video Sequence Start (CLVSS), Coded Video Sequence (CVS), Coded Video Sequence Start (CVSS), Joint Video Experts Team (JVET), Hypothetical Reference Decoder (HRD), Motion Constrained Tile Set (MCTS), Maximum Transmission Unit (MTU), Network Abstraction Layer (NAL), Output Layer Set (OLS), Picture Order Count (POC), Random Access Point (RAP), Raw Byte Sequence Payload (RBSP), Sequence Parameter Set (SPS), Video Parameter Set (VPS), and Versatile Video Coding (VVC).

[0040] Many video compression techniques can be used to reduce the size of video files while minimizing data loss. For example, video compression techniques may include performing spatial (e.g., intra-picture) prediction and / or temporal (e.g., inter-picture) prediction to reduce or remove data redundancy in a video sequence. In block-based video coding, a video slice (e.g., a video picture or a portion of a video picture) may be partitioned into video blocks, which may also be referred to as tree blocks, coding tree blocks (CTBs), coding tree units (CTUs), coding units (CUs), and / or coding nodes. Video blocks in an intra-coded (I) slice of a picture are coded using spatial prediction with respect to reference samples in neighboring blocks within the same picture. Video blocks in an inter-coded unidirectionally predicted (P) or bidirectionally predicted (B) slice of a picture may be coded using spatial prediction with respect to reference samples in neighboring blocks within the same picture or temporal prediction with respect to reference samples in other reference pictures. A picture may be referred to as a frame and / or an image, and a reference picture may be referred to as a frame and / or a reference image. Spatial or temporal prediction results in a prediction block that represents an image block. Residual data represents pixel differences between the original image block and the prediction block. Thus, inter-coded blocks are encoded according to a motion vector that points to a block of reference samples that form the prediction block, and residual data that indicates the difference between the coded block and the prediction block. Intra-coded blocks are encoded according to an intra-coding mode and residual data. For further compression, the residual data may be transformed from the pixel domain to a transform domain, which results in residual transform coefficients and may be quantized. The quantized transform coefficients may initially be arranged in a two-dimensional array. The quantized transform coefficients may be scanned to generate a one-dimensional vector of transform coefficients. To achieve further compression, entropy coding may be applied. Such video compression techniques are described in more detail below.

[0041] To ensure accurate decoding of the encoded video, the video is encoded and decoded according to a corresponding video coding standard. Video coding standards include: Advanced Video Coding (AVC), also known as International Telecommunication Union (ITU) Standardization Sector (ITU-T) H.261, International Organization for Standardization / International Electrotechnical Commission (ISO / IEC) Moving Picture Experts Group (MPEG)-1 Part 2, ITU-T H.262 or ISO / IEC MPEG-2 Part 2, ITU-T H.263, ISO / IEC MPEG-4 Part 2, ITU-T H.264 or ISO / IEC MPEG-4 Part 10, and High Efficiency Video Coding (HEVC), also known as ITU-T H.265 or MPEG-H Part 2. AVC includes extensions such as Scalable Video Coding (SVC), Multiview Video Coding (MVC), Multiview Video Coding plus Depth (MVC+D), and Three-Dimensional (3D) AVC (3D-AVC). HEVC includes extensions such as Scalable HEVC (SHVC), Multiview HEVC (MV-HEVC), and 3D HEVC (3D-HEVC). The Joint Video Experts Team (JVET) of ITU-T and ISO / IEC has initiated the development of a video coding standard called Versatile Video Coding (VVC). VVC is contained in working drafts (WDs) including JVET-O 2001-v14.

[0042] Some video coding systems use supplemental enhancement information (SEI) messages. SEI messages contain information not required by the decoding process to determine the values of samples in a decoded picture. For example, SEI messages may contain parameters used to check the bitstream for standard compliance. Furthermore, video coding systems can encode pictures into multiple layers and / or output layer sets (OLSs). Scalable nesting SEI messages can be used to correlate prefix SEI messages to layers and / or OLSs. Some video coding systems also use suffix SEI messages, but scalable nesting SEI messages cannot be used in suffix SEI messages. Various types of SEI messages can be included as either prefix SEI messages or suffix SEI messages. However, certain types of SEI messages, such as the decoded picture hash SEI message, are restricted to use as suffix SEI messages. Therefore, certain SEI messages, such as the decoded picture hash SEI message, cannot be coded in scalable nesting SEI messages in such systems.

[0043] Disclosed herein is a mechanism for coding certain SEI messages, such as a decoded picture hash SEI message, in a scalable nesting SEI message. Specifically, a suffix SEI message is enabled to be used in conjunction with a scalable nesting SEI message. In other words, an SEI NAL unit having a NAL unit type (nal_unit_type) equal to a suffix SEI NAL unit type (SUFFIX_SEI_NUT) may include a scalable nesting SEI message. In this manner, a suffix SEI message, such as a decoded picture hash SEI message, can be used as a scalable nesting SEI message and / or as a scalably nested SEI message included in a scalable nesting SEI message. This allows the suffix SEI message to be nested while applying it to a specified layer and / or OLS. This results in improved encoder and decoder functionality. Furthermore, coding efficiency can be increased, thereby reducing the use of processor, memory, and / or network signaling resources in both the encoder and decoder.

[0044] 1 is a flowchart of an exemplary operational method 100 for coding a video signal. Specifically, a video signal is encoded by an encoder. The encoding process compresses the video signal using various techniques to reduce the size of the video file. The smaller file size allows the compressed video file to be transmitted to a user, reducing the associated bandwidth overhead. A decoder then decodes the compressed video file to recover the original video signal for display to the end user. To enable the decoder to consistently recover the video signal, the decoding process typically mirrors the encoding process.

[0045] In step 101, a video signal is input to an encoder. For example, the video signal may be an uncompressed video file stored in memory. As another example, the video file may be captured by a video capture device, such as a video camera, and encoded to support live streaming of the video. The video file may include both an audio component and a video component. The video component includes a series of image frames that, when viewed in succession, create the visual impression of movement. The frames include pixels, which are represented by light, referred to herein as luma components (or luma samples), and color, referred to herein as chroma components (or color samples). In some examples, the frames may also include depth values to support three-dimensional viewing.

[0046] In step 103, the video is partitioned into blocks. Partitioning involves subdividing pixels within each frame into square and / or rectangular blocks for compression. For example, in High Efficiency Video Coding (HEVC) (also known as H.265 and MPEG-H Part 2), a frame can first be subdivided into coding tree units (CTUs), which are blocks of a predetermined size (e.g., 64 pixels by 64 pixels). CTUs contain both luma and chroma samples. A coding tree is used to divide the CTUs into blocks, which can then be recursively subdivided until a configuration is achieved that aids further encoding. For example, the luma component of a frame can be subdivided until each block contains relatively uniform illumination values. Furthermore, the chroma component of a frame can be subdivided until each block contains relatively uniform color values. Thus, the partitioning technique varies depending on the content of the video frame.

[0047] In step 105, the image blocks partitioned in step 103 are compressed using various compression techniques. For example, inter-prediction and / or intra-prediction may be used. Inter-prediction exploits the fact that objects in a typical scene tend to appear in consecutive frames. Therefore, a block representing an object in a reference frame need not be repeatedly described in adjacent frames. Specifically, an object such as a desk may remain in a constant position across multiple frames. Therefore, the desk may be described once, and adjacent frames may reference the reference frame. Pattern matching techniques can be used to match objects across multiple frames. Furthermore, an object may be depicted as moving across multiple frames, for example, due to object motion or camera motion. As a specific example, a video may show a car moving across the screen over multiple frames. Such motion can be described using motion vectors. A motion vector is a two-dimensional vector that provides an offset from the coordinates of an object in a frame to the coordinates of that object in a reference frame. Thus, inter-prediction allows image blocks in a current frame to be encoded as a set of motion vectors indicating their offsets from corresponding blocks in a reference frame.

[0048] Intra prediction encodes blocks within a given frame. It takes advantage of the fact that luma and chroma components tend to cluster together within a frame. For example, green patches in a tree tend to be located near similar green patches. Intra prediction uses several directional prediction modes (e.g., 33 in HEVC), planar mode, and direct current (DC) mode. Directional mode indicates that the current block is similar to samples from neighboring blocks in the corresponding direction. Planar mode indicates that a series of blocks along a row / column (e.g., a plane) can be interpolated based on neighboring blocks at the edge of the row. Planar mode effectively indicates a smooth transition of light / color across a row / column by using a relatively constant slope of the changing values. DC mode is used for boundary smoothing and indicates that the block is similar to the average value associated with samples from all neighboring blocks associated with the angular direction of the directional prediction mode. Therefore, intra-predicted blocks can represent image blocks as various related prediction mode values instead of actual values. Furthermore, inter-predicted blocks can represent image blocks as motion vector values instead of actual values. In either case, the predicted block may not exactly represent the image block. The difference is stored in a residual block, to which a transformation can be applied to further compress the file.

[0049] In step 107, various filtering techniques are applied. In HEVC, filters are applied according to an in-loop filtering scheme. The block-based prediction described above can produce blocky images at the decoder. Furthermore, block-based prediction schemes encode blocks and allow the encoded blocks to be reconstructed for later use as reference blocks. In-loop filtering schemes iteratively apply noise suppression filters, deblocking filters, adaptive loop filters, and sample adaptive offset (SAO) filters to blocks / frames. These filters reduce such blocking artifacts so that the encoded file can be accurately reconstructed. Furthermore, these filters reduce artifacts in the reconstructed reference blocks, reducing the likelihood of further artifacts in subsequent blocks that are encoded based on the reconstructed reference blocks.

[0050] Once the video signal has been segmented, compressed, and filtered, the resulting data is encoded into a bitstream at step 109. This bitstream includes the data described above, as well as any signaling data desired to assist a decoder in properly recovering the video signal. For example, such data may include segmentation data, prediction data, residual blocks, and various flags that provide coding instructions to the decoder. The bitstream may be stored in memory for transmission to the decoder upon request. The bitstream may also be broadcast and / or multicast to multiple decoders. Creating the bitstream is an iterative process. Therefore, steps 101, 103, 105, 107, and 109 may be performed sequentially and / or simultaneously across multiple frames and blocks. The order depicted in FIG. 1 is presented for clarity and simplicity of explanation and is not intended to limit the video coding process to any particular order.

[0051] The decoder receives the bitstream and begins the decoding process in step 111. Specifically, the decoder uses an entropy decoding scheme to convert the bitstream into corresponding syntax data and video data. In step 111, the decoder uses the syntax data from the bitstream to determine the frame partitioning. The partitioning must match the block partitioning results from step 103. This section describes the entropy encoding / decoding used in step 111. The encoder makes many choices during the compression process, such as selecting a block partitioning scheme from several options based on the spatial arrangement of values in the input image. To convey the best possible options, it may use multiple bins. A bin, as used here, is a binary value treated as a variable (e.g., a bit value that can change depending on the situation). Entropy coding allows the encoder to discard options that are clearly invalid in a particular situation, leaving a set of acceptable options. Each acceptable option is then assigned a codeword. The length of the codeword is based on the number of acceptable options (e.g., one bin for two options, two bins for three or four options, etc.). The encoder then encodes a codeword for the selected option. This scheme reduces the size of the codeword because it is desirable if the codeword uniquely indicates a choice from a small subset of allowable options, as opposed to a unique selection from a large set of all possible options. The decoder then decodes the selection by determining the set of allowable options in the same way as the encoder. By determining the set of allowable options, the decoder can read the codeword and determine the selection made by the encoder.

[0052] In step 113, the decoder performs block decoding. Specifically, the decoder generates a residual block using an inverse transform. The decoder then uses the residual block and its corresponding prediction block to reconstruct an image block based on the partitioning. The prediction block may include both intra-predicted blocks and inter-predicted blocks generated by the encoder in step 105. The reconstructed image block is then placed within a frame of the reconstructed video signal based on the partitioning data determined in step 111. The syntax of step 113 may also be conveyed in the bitstream using entropy coding as described above.

[0053] In step 115, filtering is performed on the frames of the reconstructed video signal, similar to step 107 in the encoder. For example, noise suppression filters, deblocking filters, adaptive loop filters, and SAO filters can be applied to the frames to remove blocking artifacts. Once the frames have been filtered, in step 117 the video signal can be output to a display for viewing by an end user.

[0054] 2 is a schematic diagram of an exemplary coding and decoding (codec) system 200 for video coding. Specifically, codec system 200 provides functionality to support the implementation of operational method 100. Codec system 200 is generalized to depict components used in both the encoder and decoder. Codec system 200 receives and segments a video signal as described with respect to steps 101 and 103 of operational method 100, resulting in segmented video signal 201. Codec system 200 then compresses segmented video signal 201 into a coded bitstream when functioning as an encoder as described with respect to steps 105, 107, and 109 of method 100. When functioning as a decoder, codec system 200 generates an output video signal from the bitstream as described with respect to steps 111, 113, 115, and 117 of operational method 100. Codec system 200 includes an overall coder control component 211, a transform scaling and quantization component 213, an intra-picture estimation component 215, an intra-picture prediction component 217, a motion compensation component 219, a motion estimation component 221, a scaling inverse transform component 229, a filter control analysis component 227, an in-loop filter component 225, a decoded picture buffer component 223, and a header formatting and context-adaptive binary arithmetic coding (CABAC) component 231. These components are coupled as shown. In FIG. 2, black lines indicate the movement of data to be encoded / decoded, and dashed lines indicate the movement of control data that controls the operation of other components. Any of the components of codec system 200 may reside within an encoder. A decoder may include a subset of the components of codec system 200. For example, a decoder may include an intra-picture prediction component 217, a motion compensation component 219, a scaling inverse transform component 229, an in-loop filter component 225, and a decoded picture buffer component 223. These components are now explained.

[0055] The partitioned video signal 201 is a captured video sequence that has been partitioned into blocks of pixels by a coding tree. The coding tree uses various split modes to subdivide blocks of pixels into smaller blocks of pixels. These blocks can then be further subdivided into smaller blocks. Blocks are sometimes called nodes on the coding tree. Large parent nodes are split into smaller child nodes. The number of times a node is subdivided is called the depth of the node / coding tree. The partitioned blocks may be included in coding units (CUs). For example, a CU may be a subpart of a CTU that includes a luma block, one or more red-difference chroma (Cr) blocks, and one or more blue-difference chroma (Cb) blocks, along with the corresponding CU syntax instructions. Split modes include binary tree (BT), triple tree (TT), and quad tree (QT), which are used to partition a node into two, three, or four child nodes, each with a different shape depending on the split mode used. The segmented video signal 201 is forwarded to an overall coder control component 211, a transform scaling and quantization component 213, an intra-picture estimation component 215, a filter control analysis component 227, and a motion estimation component 221 for compression.

[0056] The overall coder control component 211 is configured to make decisions related to coding images of a video sequence into a bitstream according to application constraints. For example, the overall coder control component 211 manages the optimization of bitrate / bitstream size and reconstruction quality. Such decisions may be made based on available storage space / bandwidth and image resolution requirements. The overall coder control component 211 also manages buffer usage in relation to transmission rate to mitigate buffer underrun and overrun issues. To manage these issues, the overall coder control component 211 manages segmentation, prediction, and filtering by other components. For example, the overall coder control component 211 can dynamically increase compression complexity to increase resolution and bandwidth usage, or decrease compression complexity to decrease resolution and bandwidth usage. Thus, the overall coder control component 211 controls other components of the codec system 200 to balance the reconstruction quality and bitrate issues of the video signal. The overall coder control component 211 generates control data that controls the operation of the other components. Control data is also forwarded to the header formatting CABAC component 231 and encoded in the bitstream to signal parameters for decoding at the decoder.

[0057] The partitioned video signal 201 is also sent to a motion estimation component 221 and a motion compensation component 219 for inter-prediction. A frame or slice of the partitioned video signal 201 can be divided into multiple video blocks. The motion estimation component 221 and the motion compensation component 219 perform inter-predictive coding of the received video blocks relative to one or more blocks in one or more reference frames to provide temporal prediction. The codec system 200 can perform multiple coding steps, for example, to select an appropriate coding mode for each block of video data.

[0058] The motion estimation component 221 and the motion compensation component 219 may be highly integrated but are shown separately for conceptual purposes. Motion estimation, performed by the motion estimation component 221, is the process of generating motion vectors that estimate the motion of video blocks. A motion vector may indicate, for example, the movement of a coded object relative to a predictive block. A predictive block is a block known to closely match a coded block in terms of pixel differences. A predictive block may also be referred to as a reference block. Such pixel differences may be determined by sum of absolute differences (SAD), sum of squared differences (SSD), or other difference metrics. HEVC uses several coded objects, including CTUs, coding tree blocks (CTBs), and CUs. For example, a CTU can be divided into CTBs, which can then be divided into CBs for inclusion in CUs. A CU can be encoded as a prediction unit containing prediction data and / or as a transform unit (TU) containing transformed residual data of the CU. The motion estimation component 221 generates motion vectors, prediction units, and TUs by using rate-distortion analysis as part of a rate-distortion optimization process. For example, the motion estimation component 221 can determine multiple reference blocks, multiple motion vectors, etc. for a current block / frame and select the reference block, motion vector, etc. with the best rate-distortion performance, which balances the quality of video reconstruction (e.g., the amount of data lost due to compression) and coding efficiency (e.g., the size of the final encoding).

[0059] In some examples, the codec system 200 can calculate values for sub-integer pixel positions of reference pictures stored in the decoded picture buffer component 223. For example, the video codec system 200 can interpolate values for quarter-pixel positions, eighth-pixel positions, or other fractional pixel positions of the reference pictures. Therefore, the motion estimation component 221 can perform motion searches based on whole-pixel and fractional pixel positions and output fractional pixel motion vectors. The motion estimation component 221 calculates motion vectors for prediction units of video blocks in inter-coded slices by comparing the positions of the prediction units to the positions of the prediction blocks in the reference pictures. The motion estimation component 221 outputs the calculated motion vectors as motion data to the header formatting CABAC component 231 for encoding and outputs the motion to the motion compensation component 219.

[0060] The motion compensation performed by the motion compensation component 219 can obtain or generate a prediction block based on the motion vector determined by the motion estimation component 221. Again, in some examples, the motion estimation component 221 and the motion compensation component 219 may be functionally integrated. After receiving the motion vector of the prediction unit of the current video block, the motion compensation component 219 can determine the location of the prediction block to which the motion vector points. A residual video block is then formed by subtracting pixel values of the prediction block from pixel values of the current video block being coded to form pixel difference values. Typically, the motion estimation component 221 performs motion estimation on the luma component, and the motion compensation component 219 uses the motion vector calculated based on the luma component for both the chroma and luma components. The prediction block and the residual block are forwarded to the transform, scaling, and quantization component 213.

[0061] The segmented video signal 201 is also sent to an intra-picture estimation component 215 and an intra-picture prediction component 217. Like the motion estimation component 221 and motion compensation component 219, the intra-picture estimation component 215 and the intra-picture prediction component 217 may be highly integrated but are shown separately for conceptual purposes. The intra-picture estimation component 215 and the intra-picture prediction component 217 intra-predict the current block based on blocks within the current frame, instead of the inter-prediction performed from frame to frame by the motion estimation component 221 and the motion compensation component 219, as described above. Specifically, the intra-picture estimation component 215 determines the intra-prediction mode to use to encode the current block. In some examples, the intra-picture estimation component 215 selects an appropriate intra-prediction mode to encode the current block from multiple validated intra-prediction modes. The selected intra-prediction mode is then forwarded to the header formatting CABAC component 231 for encoding.

[0062] For example, the intra picture estimation component 215 may use rate-distortion analysis to calculate rate-distortion values for various validated intra prediction modes and select the intra prediction mode with the best rate-distortion characteristics from among the validated modes. Rate-distortion analysis typically determines the amount of distortion (or error) between an encoded block and the original unencoded block that was encoded to generate the encoded block, as well as the bitrate (e.g., number of bits) used to generate the encoded block. The intra picture estimation component 215 may calculate a ratio between the distortion and rate of the various encoded blocks to determine which intra prediction mode exhibits the best rate-distortion value for the block. Additionally, the intra picture estimation component 215 may be configured to code the depth blocks of the depth map using a rate-distortion optimization (RDO)-based depth modeling mode (DMM).

[0063] The intra-picture prediction component 217, when performed on an encoder, can generate a residual block from the prediction block based on the selected intra-prediction mode determined by the intra-picture estimation component 215, or, when performed on a decoder, reads the residual block from the bitstream. The residual block contains the difference in values between the prediction block and the original block, expressed as a matrix. The residual block is then forwarded to the transform, scaling, and quantization component 213. The intra-picture estimation component 215 and the intra-picture prediction component 217 can operate on both the luma and chroma components.

[0064] The transform scaling quantization component 213 is configured to further compress the residual block. The transform scaling quantization component 213 applies a transform, such as a discrete cosine transform (DCT), a discrete sine transform (DST), or a conceptually similar transform, to the residual block to generate a video block containing residual transform coefficient values. Wavelet transforms, integer transforms, subband transforms, or other types of transforms may also be used. The transform may convert the residual information from the pixel value domain to a transform domain, such as the frequency domain. The transform scaling quantization component 213 is also configured to scale the transformed residual information, for example, based on frequency. Such scaling applies a scale factor to the residual information, which causes different frequency information to be quantized with different granularities, which may affect the final visual quality of the reconstructed video. The transform scaling quantization component 213 is also configured to quantize the transform coefficients to further reduce the bitrate. The quantization process may reduce the bit depth associated with some or all of the coefficients. The degree of quantization can be changed by adjusting a quantization parameter. In some examples, the transform scaling and quantization component 213 can then perform a scan of the matrix containing the quantized transform coefficients, which are forwarded to the header formatting CABAC component 231 for encoding in the bitstream.

[0065] The inverse scaling component 229 applies the inverse operation of the transform scaling quantization component 213 to aid in motion estimation. The inverse scaling component 229 applies inverse scaling, transformation, and / or quantization to reconstruct a residual block in the pixel domain for later use as a reference block, which may become a prediction block for another current block, for example. The motion estimation component 221 and / or motion compensation component 219 can calculate a reference block by adding the residual block to a corresponding prediction block used in motion estimation for a later block / frame. To mitigate artifacts introduced during scaling, quantization, and transformation, a filter is applied to the reconstructed reference block. Otherwise, such artifacts may lead to inaccurate predictions (and further artifacts) when subsequent blocks are predicted.

[0066] The filter control analysis component 227 and the in-loop filter component 225 apply filters to residual blocks and / or reconstructed image blocks. For example, to reconstruct the original image block, a transformed residual block from the inverse scaling transform component 229 can be combined with a corresponding prediction block from the intra-picture prediction component 217 and / or motion compensation component 219. The filter can then be applied to the reconstructed image block. In some examples, the filter can be applied to the residual block instead. Like the other components in Figure 2, the filter control analysis component 227 and the in-loop filter component 225 may be highly integrated and performed together, but are depicted separately for conceptual purposes. The filters applied to reconstructed reference blocks are applied to specific spatial regions and include multiple parameters that adjust how such filters are applied. The filter control analysis component 227 analyzes the reconstructed reference blocks to determine when such filters should be applied and sets the appropriate parameters. This data is forwarded to the header formatting CABAC component 231 as filter control data for encoding. The in-loop filter component 225 applies such filters based on the filter control data. The filters may include deblocking filters, noise suppression filters, SAO filters, and adaptive loop filters. Such filters may be applied in the spatial / pixel domain (e.g., on reconstructed pixel blocks) or in the frequency domain, depending on the example.

[0067] When operating as an encoder, the filtered reconstructed image blocks, residual blocks, and / or prediction blocks are stored in the decoded picture buffer component 223 for later use in motion estimation as described above. When operating as a decoder, the decoded picture buffer component 223 stores the reconstructed and filtered blocks and forwards them to the display as part of the output video signal. The decoded picture buffer component 223 may be any memory device capable of storing prediction blocks, residual blocks, and / or reconstructed image blocks.

[0068] The header formatting CABAC component 231 receives data from various components of the codec system 200 and encodes such data into a coded bitstream for transmission to a decoder. Specifically, the header formatting CABAC component 231 generates various headers for encoding control data, such as general control data and filter control data. Additionally, prediction data, including intra-prediction and motion data, and residual data in the form of quantized transform coefficient data are all encoded in the bitstream. The final bitstream contains all the information a decoder needs to reconstruct the original segmented video signal 201. Such information may also include an intra-prediction mode index table (also called a codeword mapping table), definitions of the encoding contexts of various blocks, indications of the most likely intra-prediction modes, indications of segmentation information, and so on. Such data may be encoded using entropy coding. For example, the information can be encoded using context-adaptive variable length coding (CAVLC), CABAC, syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval partitioned entropy (PIPE) coding, or other entropy coding techniques. After entropy coding, the coded bitstream can be transmitted to another device (e.g., a video decoder) or stored for later transmission or retrieval.

[0069] 3 is a block diagram illustrating an example video encoder 300. Video encoder 300 may be used to perform the encoding functionality of codec system 200 and / or to perform steps 101, 103, 105, 107, and / or 109 of method of operation 100. Encoder 300 segments an input video signal, resulting in a segmented video signal 301 that is substantially similar to segmented video signal 201. Segmented video signal 301 is then compressed and encoded into a bitstream by components of encoder 300.

[0070] Specifically, the partitioned video signal 301 is forwarded to an intra-picture prediction component 317 for intra prediction. The intra-picture prediction component 317 may be substantially similar to the intra-picture estimation component 215 and the intra-picture prediction component 217. The partitioned video signal 301 is also forwarded to a motion compensation component 321 for inter prediction based on reference blocks in a decoded picture buffer component 323. The motion compensation component 321 may be substantially similar to the motion estimation component 221 and the motion compensation component 219. The prediction block and residual block from the intra-picture prediction component 317 and the motion compensation component 321 are forwarded to a transform quantization component 313 for transforming and quantizing the residual block. The transform quantization component 313 may be substantially similar to the transform scaling quantization component 213. The transformed and quantized residual block and the corresponding prediction block (along with associated control data) are forwarded to an entropy coding component 331 for coding into a bitstream. The entropy coding component 331 may be substantially similar to the header formatting CABAC component 231 .

[0071] The transformed and quantized residual block and / or the corresponding prediction block are also forwarded from the transform quantization component 313 to an inverse transform quantization component 329 for reconstruction into a reference block used by the motion compensation component 321. The inverse transform quantization component 329 may be substantially similar to the inverse scaling transform component 229. In some examples, an in-loop filter in an in-loop filter component 325 is also applied to the residual block and / or the reconstructed reference block. The in-loop filter component 325 may be substantially similar to the filter control analysis component 227 and the in-loop filter component 225. The in-loop filter component 325 may include multiple filters as described with respect to the in-loop filter component 225. The filtered block is then stored in a decoded picture buffer component 323 for use as a reference block by the motion compensation component 321. The decoded picture buffer component 323 may be substantially similar to the decoded picture buffer component 223.

[0072] 4 is a block diagram illustrating an exemplary video decoder 400. Video decoder 400 may be used to perform the decoding functionality of codec system 200 and / or to perform steps 111, 113, 115, and / or 117 of method of operation 100. Decoder 400 receives a bitstream, for example, from encoder 300, and generates a reconstructed output video signal based on the bitstream for display to an end user.

[0073] The bitstream is received by the entropy decoding component 433. The entropy decoding component 433 is configured to perform an entropy decoding scheme, such as CAVLC, CABAC, SBAC, PIPE coding, or other entropy coding techniques. For example, the entropy decoding component 433 can use header information to provide context for interpreting additional data encoded as codewords in the bitstream. The decoded information includes information desirable for decoding the video signal, such as general control data, filter control data, partition information, motion data, prediction data, and quantized transform coefficients from residual blocks. The quantized transform coefficients are forwarded to the inverse transform and quantization component 429 for reconstruction into residual blocks. The inverse transform and quantization component 429 can be similar to the inverse transform and quantization component 329.

[0074] The reconstructed residual blocks and / or predictive blocks are forwarded to the intra-picture prediction component 417 for reconstruction into image blocks based on intra-prediction operations. The intra-picture prediction component 417 may be similar to the intra-picture estimation component 215 and the intra-picture prediction component 217. Specifically, the intra-picture prediction component 417 uses the prediction mode to locate reference blocks within a frame and applies the residual blocks to the result to reconstruct intra-predicted image blocks. The reconstructed intra-predicted image blocks and / or residual blocks and corresponding inter-prediction data are forwarded to the decoded picture buffer component 423 via the in-loop filter component 425, which may be substantially similar to the decoded picture buffer component 223 and the in-loop filter component 225, respectively. The in-loop filter component 425 filters the reconstructed image blocks, residual blocks, and / or predictive blocks, and such information is stored in the decoded picture buffer component 423. The reconstructed image blocks are transferred from the decoded picture buffer component 423 to the motion compensation component 421 for inter prediction. The motion compensation component 421 may be substantially similar to the motion estimation component 221 and / or the motion compensation component 219. Specifically, the motion compensation component 421 generates a prediction block using a motion vector from a reference block and applies a residual block to the result to reconstruct the image block. The resulting reconstructed block may also be transferred to the decoded picture buffer component 423 via an in-loop filter component 425. The decoded picture buffer component 423 continues to store additional reconstructed image blocks, which can be reconstructed into frames according to the segmentation information. Such frames may also be arranged in a sequence. This sequence is output to a display as a reconstructed output video signal.

[0075] 5 is a schematic diagram illustrating an example HRD 500. The HRD 500 may be used in the codec system 200 and / or an encoder, such as the encoder 300. The HRD 500 may inspect the bitstream generated in step 109 of the method 100 before the bitstream is forwarded to a decoder, such as the decoder 400. In some examples, the bitstream may be continuously forwarded through the HRD 500 as it is encoded. If a portion of the bitstream does not conform to an associated constraint, the HRD 500 may indicate such anomaly to the encoder so that the encoder re-encodes the corresponding portion of the bitstream with another mechanism.

[0076] The HRD 500 includes a virtual stream scheduler (HSS) 541. The HSS 541 is a component configured to implement a virtual distribution mechanism. The virtual distribution mechanism is used to check the conformance of a bitstream or decoder with respect to the timing and data flow of a bitstream 551 input to the HRD 500. For example, the HSS 541 may receive the bitstream 551 output from an encoder and manage the process of conformance testing of the bitstream 551. In one particular example, the HSS 541 may control the rate at which coded pictures pass through the HRD 500 and verify that the bitstream 551 does not contain non-conforming data.

[0077] The HSS 541 may transfer the bitstream 551 to the CPB 543 at a predetermined rate. The HRD 500 may manage data in decode units (DUs) 553. A DU 553 is a subset of an access unit (AU) or an AU and associated non-video coding layer (VCL) network abstraction layer (NAL) units. Specifically, an AU includes one or more pictures associated with an output time. For example, an AU may include a single picture in a single-layer bitstream or pictures per layer in a multi-layer bitstream. Each picture in an AU may be divided into slices, each of which is included in a corresponding VCL NAL unit. Thus, a DU 553 may include one or more pictures, one or more slices of a picture, or a combination thereof. Additionally, parameters used to decode the AU, picture, and / or slice may be included in the non-VCL NAL units. Thus, DU 553 contains non-VCL NAL units, which contain data necessary to support decoding of VCL NAL units within DU 553. CPB 543 is a first-in, first-out buffer for HRD 500. CPB 543 contains DUs 553 containing video data in decode order. CPB 543 stores video data for use during bitstream conformance verification.

[0078] The CPB 543 forwards the DUs 553 to a decode process component 545. The decode process component 545 is a component that conforms to the VVC standard. For example, the decode process component 545 may emulate the decoder 400 used by an end user. The decode process component 545 decodes the DUs 553 at a rate achievable by an exemplary end-user decoder. If the decode process component 545 cannot decode the DUs 553 fast enough to prevent the CPB 543 from overflowing, the bitstream 551 does not conform to the standard and must be re-encoded.

[0079] The decode process component 545 decodes the DU 553 to generate a decoded DU 555. The decoded DU 555 includes a decoded picture. The decoded DU 555 is forwarded to a DPB 547. The DPB 547 may be substantially similar to the decoded picture buffer components 223, 323, and / or 423. Pictures obtained from the decoded DU 555 and marked for use as reference pictures 556 to support inter-prediction are returned to the decode process component 545 to support further decoding. The DPB 547 outputs the decoded video sequence as a series of pictures 557. The pictures 557 are reconstructed pictures that typically reflect the pictures encoded into the bitstream 551 by an encoder.

[0080] Picture 557 is forwarded to output cropping component 549, which is configured to apply an adaptive cropping window to picture 557. This results in output cropped picture 559. Output cropped picture 559 is a fully reconstructed picture. Thus, output cropped picture 559 mimics what an end user will see when decoding bitstream 551. In this way, the encoder can review output cropped picture 559 to ensure the encoding is good.

[0081] The HRD 500 is initialized based on HRD parameters in the bitstream 551. For example, the HRD 500 can read the HRD parameters from a VPS, SPS, and / or SEI message. The HRD 500 can then perform conformance testing operations on the bitstream 551 based on the information in such HRD parameters. As a specific example, the HRD 500 can determine one or more CPB delivery schedules from the HRD parameters. The delivery schedules specify the timing of delivery of video data to and from memory locations such as the CPB and / or DPB. Thus, the CPB delivery schedules specify the timing of delivery of AUs, DUs 553, and / or pictures to and from the CPB 543. It should be noted that the HRD 500 can use a DPB delivery schedule for the DPB 547 that is similar to the CPB delivery schedule.

[0082] Video may be coded into different layers and / or OLSs for use by decoders with varying levels of hardware capabilities and for varying network conditions. CPB delivery schedules are selected to reflect these considerations. Thus, higher layer sub-bitstreams are designated for optimal hardware and network conditions, and therefore the higher layers may receive one or more CPB delivery schedules that use large amounts of memory in the CPB 543 and short delays for the transfer of DUs 553 towards the DPB 547. Similarly, lower layer sub-bitstreams are designated for limited decoder hardware capabilities and / or poor network conditions. Thus, the lower layers may receive one or more CPB delivery schedules that use small amounts of memory in the CPB 543 and longer delays for the transfer of DUs 553 towards the DPB 547. The OLSs, layers, sub-layers, or combinations thereof may then be tested according to the corresponding delivery schedules to ensure that the resulting sub-bitstreams can be correctly decoded under the conditions expected for the sub-bitstreams. Thus, the HRD parameters in the bitstream 551 may indicate the CPB delivery schedule and include sufficient data to enable the HRD 500 to determine the CPB delivery schedule and correlate the CPB delivery schedule to the corresponding OLS, tier, and / or sub-tier.

[0083] 6 is a schematic diagram illustrating an exemplary multi-layer video sequence 600. The multi-layer video sequence 600 may be encoded by an encoder, such as codec system 200 and / or encoder 300, and decoded by a decoder, such as codec system 200 and / or decoder 400, according to, for example, method 100. Additionally, the multi-layer video sequence 600 may be checked for conformance by an HRD, such as HRD 500. The multi-layer video sequence 600 is included to illustrate an exemplary application of layers within a coded video sequence. The multi-layer video sequence 600 is any video sequence that uses multiple layers, such as layer N 631 and layer N+1 632.

[0084] In one example, the multiple-layer video sequence 600 can use inter-layer prediction 621. Inter-layer prediction 621 is applied between pictures 611, 612, 613, 614 and pictures 615, 616, 617, 618 of different layers. In the example shown, pictures 611, 612, 613, and 614 are part of layer N+1 632, and pictures 615, 616, 617, and 618 are part of layer N 631. Layers, such as layer N 631 and / or layer N+1 632, are groups of pictures that are all associated with similar values of characteristics such as similar size, quality, resolution, signal-to-noise ratio, capacity, etc. A layer may be formally defined as a set of VCL NAL units that share the same layer ID and associated non-VCL NAL units. A VCL NAL unit is a NAL unit coded to contain video data, such as a coded slice of a picture. A non-VCL NAL unit is a NAL unit that contains non-video data, such as syntax and / or parameters that support decoding video data, performing conformance checks, or other operations.

[0085] In the illustrated example, layer N+1 632 is associated with a larger image size than layer N 631. Thus, in this example, the picture sizes (e.g., larger heights and widths, and therefore more samples) of pictures 611, 612, 613, and 614 of layer N+1 632 are larger than the picture sizes of pictures 615, 616, 617, and 618 of layer N 631. However, such pictures may be separated between layer N+1 632 and layer N 631 by other characteristics. Although only two layers, layer N+1 632 and layer N 631, are shown, a set of pictures may be separated into any number of layers based on associated characteristics. Layer N+1 632 and layer N 631 may also be indicated by a layer ID. A layer ID is an item of data associated with a picture and indicates that the picture is part of the indicated layer. Thus, each picture 611-618 may be associated with a corresponding layer ID to indicate which layer N+1 632 or layer N 631 contains the corresponding picture. For example, the layer ID may include nuh_layer_id 635, a syntax element that specifies the identifier of the layer that contains the NAL unit (e.g., containing slices and / or parameters of the picture within the layer). A layer associated with lower quality / smaller image size / smaller bitstream size, such as layer N 631, is generally assigned a lower layer ID and is referred to as a lower layer. Furthermore, a layer associated with higher quality / larger image size / larger bitstream size, such as layer N+1 632, is generally assigned a higher layer ID and is referred to as a higher layer.

[0086] Pictures 611-618 in different layers 631-632 are configured to be displayed in alternative manners. As a specific example, a decoder may decode and display picture 615 at the current display time if a smaller picture is desired, or the decoder may decode and display picture 611 at the current display time if a larger picture is desired. Thus, pictures 611-614 in upper layer N+1 632 contain substantially the same image data as corresponding pictures 615-618 in lower layer N 631 (despite differences in picture size). Specifically, picture 611 contains substantially the same image data as picture 615, picture 612 contains substantially the same image data as picture 616, and so on.

[0087] Pictures 611-618 may be coded by referencing other pictures 611-618 in the same layer N 631 or N+1 632. Coding a picture with reference to another picture in the same layer results in inter-prediction 623. Inter-prediction 623 is indicated by a solid arrow. For example, picture 613 may be coded using inter-prediction 623 using one or two of pictures 611, 612, and / or 614 in layer N+1 632 as references, with one picture referenced for unidirectional inter-prediction and / or two pictures referenced for bidirectional inter-prediction. Furthermore, picture 617 may be coded using inter-prediction 623 using one or two of pictures 615, 616, and / or 618 in layer N 631 as references, with one picture referenced for unidirectional inter-prediction and / or two pictures referenced for bidirectional inter-prediction. When performing inter prediction 623, a picture may be called a reference picture when it is used as a reference for another picture in the same layer. For example, picture 612 may be a reference picture used to code picture 613 according to inter prediction 623. Inter prediction 623 may also be called intra-layer prediction in a multi-layer context. Thus, inter prediction 623 is a mechanism for coding samples of a current picture by referencing indicated samples in a reference picture different from the current picture, where the reference picture and the current picture are in the same layer.

[0088] Pictures 611-618 can also be coded by referencing other pictures 611-618 in different layers. This process is known as inter-layer prediction 621 and is indicated by the dashed arrows. Inter-layer prediction 621 is a mechanism for coding samples of a current picture by referencing indicated samples in reference pictures where the current picture and the reference picture are in different layers and therefore have different values for nuh_layer_id 635. For example, a picture in a lower layer N 631 can be used as a reference picture to code a corresponding picture in an upper layer N+1 632. As a specific example, picture 611 can be coded by referencing picture 615 according to inter-layer prediction 621. In such a case, picture 615 is used as the inter-layer reference picture. An inter-layer reference picture is a reference picture used for inter-layer prediction 621. In most cases, inter-layer prediction 621 is constrained so that a current picture, such as picture 611, can only use inter-layer reference pictures that are contained in the same AU 627 and that are in a lower layer, such as picture 615. When multiple layers (e.g., more than two) are available, inter-layer prediction 621 can code / decode the current picture based on multiple inter-layer reference pictures that are in a lower level than the current picture.

[0089] A video encoder can use the multiple-layer video sequence 600 to code pictures 611-618 via many different combinations and / or permutations of inter-prediction 623 and inter-layer prediction 621. For example, picture 615 may be coded according to intra-prediction. Thereafter, pictures 616-618 may be coded according to inter-prediction 623 by using picture 615 as a reference picture. Furthermore, picture 611 may be coded according to inter-layer prediction 621 by using picture 615 as an inter-layer reference picture. Thereafter, pictures 612-614 may be coded according to inter-prediction 623 by using picture 611 as a reference picture. Thus, reference pictures can function as both single-layer reference pictures and inter-layer reference pictures for different coding mechanisms. By coding the upper layer N+1 631 pictures based on the lower layer N 632 pictures, the upper layer N+1 632 can avoid using intra prediction, which has much lower coding efficiency than inter prediction 623 and inter-layer prediction 621. In this way, the poor coding efficiency of intra prediction may be limited to pictures of the smallest / lowest quality and therefore limited to coding a minimum amount of video data. Pictures used as reference pictures and / or inter-layer reference pictures may be indicated by entries in a reference picture list included in a reference picture list structure.

[0090] To perform such operations, layers such as layer N 631 and layer N+1 632 may be included in OLS 628. OLS 628 is a collection of layers, one or more of which are designated as output layers. An output layer is a layer designated for output (e.g., to a display). For example, layer N 631 may be included only to support inter-layer prediction 621 and may never be output. In this case, layer N+1 631 is decoded and output based on layer N 632. In this case, OLS 628 has layer N+1 632 as an output layer. In some cases, OLS 628 includes only an output layer, called a simulcast layer. In other cases, OLS 628 may include many layers in different combinations. For example, an output layer in OLS 628 can be coded according to inter-layer prediction 621 based on one, two, or many lower layers. OLS 628 may also include multiple output layers. Thus, an OLS 628 may include one or more output layers and any support layers necessary to reconstruct the output layers. A multi-layer video sequence 600 can be coded by using many different OLSs 628, each using a different combination of layers. An OLS 628 has an associated OLS index, which is an index that uniquely identifies the OLS 628.

[0091] Pictures 611-618 may also be included in an access unit (AU) 627. AU 627 is a collection of coded pictures included in different layers and having the same output time during decoding. Therefore, coded pictures in the same AU 627 are scheduled to be output from the DPB at the same time in the decoder. For example, pictures 614 and 618 are in the same AU 627. Pictures 613 and 617 are in a different AU 627 from pictures 614 and 618. In an alternative example, pictures 614 and 618 in the same AU 627 may be displayed. For example, picture 618 may be displayed when a small picture size is desired, and picture 614 may be displayed when a large picture size is desired. When a large picture size is desired, picture 614 is output, and picture 618 is used only for inter-layer prediction 621. In this case, picture 618 is discarded without being output once inter-layer prediction 621 is completed.

[0092] 7 is a schematic diagram illustrating an exemplary bitstream 700. For example, the bitstream 700 may be generated by the codec system 200 and / or the encoder 300 for decoding by the codec system 200 and / or the decoder 400 according to the method 100. Furthermore, the bitstream 700 may include the multi-layer video sequence 600. Furthermore, the bitstream 700 may include various parameters for controlling the operation of an HRD, such as the HRD 500. Based on such parameters, the HRD may check the bitstream 700 for compliance with the standard before sending it to the decoder for decoding.

[0093] The bitstream 700 includes a VPS 711, one or more SPSs 713, multiple picture parameter sets (PPSs) 715, multiple slice headers 717, image data 720, a prefix SEI message 718, and a suffix SEI message 719. The VPS 711 includes data related to the entire bitstream 700. For example, the VPS 711 may include data related to image sequences, layers, and / or sublayers used in the bitstream 700. The SPS 713 includes sequence data common to all pictures in a coded video sequence included in the bitstream 700. For example, each layer may include one or more coded video sequences, and each coded video sequence may reference the SPS 713 for corresponding parameters. The parameters in the SPS 713 may include picture sizing, bit depth, coding tool parameters, bit rate limits, etc. Note that while each sequence points to an SPS 713, in some examples, a single SPS 713 may contain data for multiple sequences. The PPS 715 contains parameters that apply to an entire picture. Thus, each picture in a video sequence may reference a PPS 715. Note that, in some examples, while each picture references a PPS 715, a single PPS 715 can contain data for multiple pictures. For example, multiple similar pictures may be coded according to similar parameters. In such cases, a single PPS 715 may contain data for such similar pictures. The PPS 715 may indicate coding tools, quantization parameters, offsets, etc. available for slices within the corresponding picture.

[0094] The slice header 717 contains parameters specific to each slice in a picture. Thus, there may be one slice header 717 per slice in a video sequence. The slice header 717 may include slice type information, filtering information, prediction weights, tile entry points, deblocking parameters, etc. Note that in some examples, the bitstream 700 may also include a picture header, which is a syntax structure that contains parameters that apply to all slices in a single picture. For this reason, in some contexts, the picture header and the slice header 717 may be used interchangeably. For example, some parameters may move between the slice header 717 and the picture header depending on whether such parameters are common to all slices in a picture.

[0095] Image data 720 includes video data encoded according to inter-prediction and / or intra-prediction, as well as corresponding transformed and quantized residual data. For example, image data 720 may include layers 723, pictures 725, and / or slices 727. Layer 723 is a set of VCL NAL units 745 and associated non-VCL NAL units 741 that share specified characteristics (e.g., a common resolution, frame rate, picture size, etc.) as indicated by a layer ID such as nuh_layer_id. For example, layer 723 may include a set of pictures 725 that share the same nuh_layer_id. Layer 723 may be substantially similar to layers 631 and / or 632. nuh_layer_id is a syntax element that specifies an identifier for layer 723 that includes at least one NAL unit. For example, the lowest quality layer 723, known as the base layer, may include the lowest value of nuh_layer_id, with nuh_layer_id values increasing for higher quality layers 723. Therefore, the lower layer is the layer 723 with the smaller value of nuh_layer_id, and the higher layer is the layer 723 with the larger value of nuh_layer_id.

[0096] A picture 725 is an array of luma samples and / or chroma samples that generate a frame or a field thereof. For example, a picture 725 is a coded image that can be output for display or used to support coding of other pictures 725 for output. A picture 725 includes one or more slices 727. A slice 727 may be defined as an integer number of complete tiles or an integer number of contiguous complete coding tree unit (CTU) rows (e.g., within a tile) of a picture 725 that are exclusively contained in a single NAL unit. A slice 727 is further divided into CTUs and / or coding tree blocks (CTBs). A CTU is a group of samples of a predetermined size that can be partitioned in the coding tree. A CTB is a subset of a CTU and contains the luma or chroma component of the CTU. The CTUs / CTBs are further divided into coding blocks based on the coding tree. The coding blocks can then be encoded / decoded according to a prediction mechanism.

[0097] An SEI message is a syntax structure with specific semantics that conveys information that is not required by the decoding process to determine the values of samples in a decoded picture. For example, an SEI message may contain data to support the HRD process or other support data that is not directly related to decoding of the bitstream 700 at the decoder. An SEI message can be structured as a prefix SEI message 718 or a suffix SEI message 719. A prefix SEI message 718 is an SEI message that applies to one or more subsequent NAL units. A suffix SEI message 719 is an SEI message that applies to one or more preceding NAL units. A prefix SEI message 718 can contain an SEI message of a specified type, while a suffix SEI message 719 can contain an SEI message of another type.

[0098] The prefix SEI message 718 may include a buffering period (BP) SEI message including HRD parameters for initializing an HRD to manage a CPB for testing the corresponding OLS and / or layer 723. The prefix SEI message 718 may also include a picture timing (PT) SEI message including HRD parameters for managing distribution information for AUs in the CPB and / or DPB for testing the corresponding OLS and / or layer 723. The prefix SEI message 718 may also include a decode unit information (DUI) SEI message including HRD parameters for managing distribution information for decode units (DUs) in the CPB and / or DPB for testing the corresponding OLS and / or layer 723.

[0099] The suffix SEI message 719 may include a decoded picture hash SEI message 748. The decoded picture hash SEI message 748 is a type of SEI message that includes a checksum derived from sample values of the decoded picture. The decoded picture hash SEI message 748 may be used by a decoder to detect whether a picture was received and decoded correctly. In this manner, the decoded picture hash SEI message 748 can be used to detect transmission and decoding errors associated with the picture 725 that precedes the decoded picture hash SEI message 748.

[0100] A set of SEI messages may be implemented as a scalable nesting SEI message 746. The scalable nesting SEI message 746 provides a mechanism for associating an SEI message with a particular layer 723. A scalable nesting SEI message 746 is a message that includes multiple scalably nested SEI messages 747. A scalably nested SEI message 747 is an SEI message that corresponds to one or more OLSs or one or more layers 723. An OLS is a collection of layers 723 where at least one of the layers 723 is an output layer. Thus, depending on the context, a scalable nesting SEI message 746 may be said to include a set of scalably nested SEI messages 747 or a set of SEI messages. Furthermore, a scalable nesting SEI message 746 includes a set of scalably nested SEI messages 747 of the same type.

[0101] Some video coding systems do not allow the scalable nesting SEI message 746 to be used in the suffix SEI message 719. Various types of SEI messages can be included as either prefix SEI messages 718 or suffix SEI messages 719. However, certain types of SEI messages, such as the decoded picture hash SEI message 748, are restricted to use as suffix SEI messages 719. Thus, certain SEI messages, such as the decoded picture hash SEI message 748, cannot be coded into the scalable nesting SEI message 746 in such systems.

[0102] The bitstream 700 is modified to address the aforementioned deficiencies. Specifically, the suffix SEI message 719 is configured to include a scalable nesting SEI message 746. Accordingly, the suffix SEI message 719 may also include a scalably nested SEI message 747 corresponding to the layer 743 and / or the OLS. This allows certain SEI messages, such as a decoded picture hash SEI message 748, to be used as the scalably nested SEI message 747 within the scalable nesting SEI message 746.

[0103] Note that bitstream 700 can be coded as a sequence of NAL units. NAL units are containers for video data and / or supporting syntax. The NAL units can be VCL NAL units 745 or non-VCL NAL units 741. VCL NAL units 745 are NAL units coded to contain video data. Specifically, VCL NAL units 745 include slices 727 and associated slice headers 717. Non-VCL NAL units 741 are NAL units that contain non-video data, such as syntax and / or parameters that support decoding the video data, performing conformance checks, or other operations. Non-VCL NAL units 741 can include VPS NAL units, SPS NAL units, PPS NAL units, and SEI NAL units 744, which include VPS 711, SPS 713, PPS 715, and prefix SEI message 718 or suffix SEI message 719, respectively. It should be noted that the above list of NAL units is exemplary and not exhaustive. Each NAL unit has a NAL unit type (nal_unit_type) 731. The nal_unit_type 731 is a syntax element contained in a NAL unit that indicates the type of data contained in the NAL unit.

[0104] An SEI NAL unit 744 is a NAL unit that contains an SEI message. An SEI NAL unit 744 can have nal_unit_type 731 set to indicate that the SEI NAL unit 744 is a prefix SEI NAL unit type (PREFIX_SEI_NUT) 742. An SEI NAL unit 744 can also have nal_unit_type 731 set to indicate that the SEI NAL unit 744 is a SUFFIX_SEI_NUT 743. Thus, an SEI NAL unit 744 with nal_unit_type 731 equal to SUFFIX_SEI_NUT 743 can contain a scalable nesting SEI message 746 and / or one or more scalably nested SEIs 747. In this manner, a suffix SEI message 719, such as a decoded picture hash SEI message 748, can be used as a scalable nesting SEI message 746 and / or as a scalably nested SEI message 747 included in the scalable nesting SEI message 746. This allows the suffix SEI message 719 to be nested while applying to a specified layer 723 and / or OLS (e.g., OLS 628 as shown in FIG. 6). Additionally, other SEI messages that can be used within the suffix SEI message 719, such as BP, PT, and / or DUI SEI messages, can also be used as scalable nesting SEI messages 746 and / or scalably nested SEI messages 747 when placed within the suffix SEI message 719 (e.g., in an SEI NAL unit 744 having nal_unit_type 731 set to indicate that the SEI NAL unit 744 is SUFFIX_SEI_NUT 743). This results in improved encoder and decoder performance, and can increase coding efficiency, thereby reducing the use of processor, memory, and / or network signaling resources in both the encoder and decoder.

[0105] The scalable nesting SEI message 746 also includes syntax elements for associating the scalably nested SEI message 747 with a layer 723 and / or an OLS. For example, the scalable nesting SEI message 746 may include a payload type (payloadType) 733 and a scalable nesting layer identifier (layer_id[i]) 735. These syntax elements represent the type of data included in the SEI message and therefore represent the type of SEI message (e.g., suffix SEI message 719) included in the SEI NAL unit 744. For example, payloadType 733 can be set to indicate that suffix SEI message 719 includes the scalable nesting SEI message 746. In a particular example, payloadType 733 can be set to a value 133 to indicate that suffix SEI message 719 includes the scalable nesting SEI message 746. In other words, payloadType 733 may be set to a value of 133, indicating that the SEI NAL unit 744 has a nal_unit_type 731 of SUFFIX_SEI_NUT 743 and contains a scalable nesting SEI message 746. scalable nesting layer_id[i] 735 is a syntax element in the scalable nesting SEI message 746 that specifies the nuh_layer_id value of the ith layer 723 to which the scalably nested SEI message 747 applies. Thus, scalable nesting layer_id[i] 735 may associate each of the scalable nested SEI messages 747 in the scalable nesting SEI message 746 with a corresponding layer 723. Thus, the HRD and / or decoder may read payloadType 733 and determine that a scalable nesting SEI message 746 is present in the suffix SEI message 719.The HRD and / or decoder can then determine the layer 723 for each scalably nested SEI message 747 within the scalable nesting SEI message 746 based on the layer ID value in the scalable nesting layer_id[i] 735.

[0106] The aforementioned information will now be explained in more detail below. Layered video coding is also referred to as scalable coding or scalable video coding. Scalability in video coding can be supported by using multi-layer coding techniques. A multi-layer bitstream comprises a base layer (BL) and one or more enhancement layers (EL). Examples of scalability include spatial scalability, quality / signal-to-noise ratio (SNR) scalability, multiview scalability, frame rate scalability, etc. When a multi-layer coding technique is used, a picture or a portion thereof may be coded without using a reference picture (intra-prediction), coded by referencing a reference picture within the same layer (inter-prediction), and / or coded by referencing a reference picture in another layer (inter-layer prediction). A reference picture used for inter-layer prediction of a current picture is called an inter-layer reference picture (ILRP). Figure 6 shows an example of multi-layer coding for spatial scalability in which pictures in different layers have different resolutions.

[0107] Some video coding families provide scalability support in profiles separate from profiles for single-layer coding. Scalable Video Coding (SVC) is a scalable extension of Advanced Video Coding (AVC) that supports spatial, temporal, and quality scalability. For SVC, a flag is signaled in each macroblock (MB) in an EL picture to indicate whether the EL MB is predicted using co-located blocks from lower layers. Predictions from co-located blocks may include texture, motion vectors, and / or coding modes. SVC implementations may not directly reuse unmodified AVC implementations in their design. The SVC EL macroblock syntax and decoding process differ from the AVC syntax and decoding process.

[0108] Scalable HEVC (SHVC) is an extension of HEVC that supports spatial and quality scalability. Multiview HEVC (MV-HEVC) is an extension of HEVC that provides support for multiview scalability. 3D HEVC (3D-HEVC) is an extension of HEVC that provides support for more advanced and efficient 3D video coding than MV-HEVC. Temporal scalability can be included as an integral part of a single-layer HEVC codec. In multi-layer extensions of HEVC, decoded pictures used for inter-layer prediction come only from the same AU and are treated as long-term reference pictures (LTRPs). Such pictures are assigned reference indices in a reference picture list along with other temporal reference pictures in the current layer. Inter-layer prediction (ILP) is achieved at the prediction unit level by setting the value of a reference index to reference an inter-layer reference picture in the reference picture list. Spatial scalability involves resampling a reference picture or part of it if the ILRP has a different spatial resolution than the current picture being coded or decoded. Reference picture resampling can be achieved either at the picture level or at the coding block level.

[0109] VVC also supports layered video coding. A VVC bitstream may contain multiple layers. The layers may all be independent of each other. For example, each layer may be coded without inter-layer prediction. In this case, each layer is also referred to as a simulcast layer. In some cases, some layers are coded using ILP. A flag in the VPS may indicate whether a layer is a simulcast layer or whether some layers use ILP. If some layers use ILP, the layer dependency between layers is also signaled in the VPS. Unlike SHVC and MV-HEVC, VVC may not specify an OLS. An OLS contains a specified set of layers, and one or more layers in the set of layers are designated as output layers. An output layer is a layer of the OLS that is output. In some implementations of VVC, if a layer is a simulcast layer, only one layer may be selected for decoding and output. In some implementations of VVC, when any layer uses ILP, the entire bitstream, including all layers, is designated to be decoded. Furthermore, one of these layers is designated as the output layer. The output layer may be designated as the top layer only, all layers, or the top layer plus the set of designated lower layers.

[0110] The above-described aspect has several problems. For example, the nuh_layer_id values of SPS, PPS, and APS NAL units may not be properly constrained. Furthermore, the TemporalId value of SEI NAL units may not be properly constrained. Also, when reference picture resampling is enabled and pictures in a CLVS have different spatial resolutions, the setting of NoOutputOfPriorPicsFlag may not be properly specified. Also, in some video coding systems, suffix SEI messages cannot be included in scalable nesting SEI messages. As another example, buffering period, picture timing, and decode unit information SEI messages may contain parsing dependencies on VPS and / or SPS.

[0111] Generally, this disclosure describes video coding improvement techniques. The techniques are based on VVC. However, these techniques also apply to layered video coding based on other video codec specifications.

[0112] One or more of the problems described above may be solved as follows: The nuh_layer_id values of SPS, PPS, and APS NAL units are appropriately constrained herein. The TemporalId values of SEI NAL units are appropriately constrained herein. The setting of NoOutputOfPriorPicsFlag is appropriately specified when reference picture resampling is enabled and pictures in a CLVS have different spatial resolutions. A suffix SEI message can be included in a scalable nesting SEI message. The parsing dependency of BP, PT, and DUI SEI messages on VPS or SPS can be removed by repeating the syntax element decoding_unit_hrd_params_present_flag in the BP SEI message syntax, the syntax elements decoding_unit_hrd_params_present_flag and decoding_unit_cpb_params_in_pic_timing_sei_flag in the PT SEI message syntax, and the syntax element decoding_unit_cpb_params_in_pic_timing_sei_flag in the DUI SEI message.

[0113] An exemplary implementation of the aforementioned mechanism is as follows: An example of general NAL unit semantics is given below:

[0114] nuh_temporal_id_plus1-1 specifies the temporal identifier of the NAL unit. The value of nuh_temporal_id_plus1 should not be equal to 0. The variable TemporalId can be derived as follows: TemporalId=nuh_temporal_id_plus1-1 When nal_unit_type is in the range IDR_W_RADL to RSV_IRAP_13 (inclusive), TemporalId must be equal to 0. When nal_unit_type is equal to STSA_NUT, TemporalId must not be equal to 0.

[0115] The value of TemporalId MUST be the same for all VCL NAL units of an access unit. The value of TemporalId of a coded picture, layer access unit, or access unit MAY be the value of TemporalId of the VCL NAL units of the coded picture, layer access unit, or access unit. The value of TemporalId of a sub-layer representation MAY be the maximum value of TemporalId of all VCL NAL units in the sub-layer representation.

[0116] The value of TemporalId for non-VCL NAL units is constrained as follows: If nal_unit_type is equal to DPS_NUT, VPS_NUT, or SPS_NUT, then TemporalId shall be equal to 0, and the TemporalId of the access unit that contains the NAL unit shall be equal to 0. Otherwise, if nal_unit_type is equal to EOS_NUT or EOB_NUT, then TemporalId shall be equal to 0. Otherwise, if nal_unit_type is equal to AUD_NUT, FD_NUT, PREFIX_SEI_NUT, or SUFFIX_SEI_NUT, then TemporalId shall be equal to the TemporalId of the access unit that contains the NAL unit. Otherwise, if nal_unit_type is equal to PPS_NUT or APS_NUT, then TemporalId shall be greater than or equal to the TemporalId of the access unit that contains the NAL unit. If the NAL unit is a non-VCL NAL unit, the value of TemporalId MUST be equal to the minimum of the TemporalId values of all access units to which the non-VCL NAL unit applies. If nal_unit_type is equal to PPS_NUT or APS_NUT, TemporalId MAY be greater than or equal to the TemporalId of the containing access unit. This is because all PPSs and APSs may be included at the beginning of the bitstream. Furthermore, the first coded picture has TemporalId equal to 0.

[0117] Examples of sequence parameter set RBSP semantics are as follows: An SPS RBSP MUST be available to the decoding process before it is referenced. An SPS MAY be included in at least one access unit with TemporalId equal to 0, or MAY be provided via an external mechanism. An SPS NAL unit containing an SPS MAY be constrained to have nuh_layer_id equal to the lowest nuh_layer_id value of the PPS NAL unit that references the SPS.

[0118] Exemplary picture parameter set RBSP semantics are as follows: The PPS RBSP must be available to the decoding process before it is referenced. The PPS should be included in at least one access unit with a TemporalId less than or equal to the TemporalId of the PPS NAL unit, or provided via an external mechanism. The PPS NAL unit containing the PPS RBSP should have a nuh_layer_id equal to the lowest nuh_layer_id value of the coded slice NAL units that reference the PPS.

[0119] The semantics of an exemplary adaptation parameter set are as follows: Each APS RBSP MUST be available to the decoding process before it is referenced. The APS SHOULD also be included in at least one access unit with a TemporalId less than or equal to the TemporalId of the coded slice NAL unit that references the APS or is provided via an external mechanism. An APS NAL unit can be shared by pictures / slices of multiple layers. The nuh_layer_id of an APS NAL unit MUST be equal to the lowest nuh_layer_id value of the coded slice NAL unit that references the APS NAL unit. Alternatively, an APS NAL unit may not be shared by pictures / slices of multiple layers. The nuh_layer_id of an APS NAL unit MUST be equal to the nuh_layer_id of the slice that references the APS.

[0120] In one example, removing a picture from the DPB before decoding the current picture is described as follows: Removing a picture from the DPB before decoding the current picture (but after parsing the slice header of the first slice of the current picture) may be done at the CPB removal time of the first decoding unit of access unit n (containing the current picture). This proceeds as follows: A decoding process for reference picture list construction is invoked, and a decoding process for reference picture marking is invoked.

[0121] If the current picture is a Coding Layer Video Sequence Start (CLVSS) picture that is not picture 0, the following ordered steps are applied: For the decoder under test, derive the variable NoOutputOfPriorPicsFlag as follows: If the values of pic_width_in_luma_samples, pic_height_max_in_luma_samples, chroma_format_idc, separate_colour_plane_flag, bit_depth_luma_minus8, bit_depth_chroma_minus8, or sps_max_dec_pic_buffering_minus1[Htid] derived from the SPS differ from the values of pic_width_in_luma_samples, pic_height_in_luma_samples, chroma_format_idc, separate_colour_plane_flag, bit_depth_luma_minus8, bit_depth_chroma_minus8, or sps_max_dec_pic_buffering_minus1[Htid] respectively derived from the preceding SPS, NoOutputOfPriorPicsFlag may be set to 1 by the decoder under test regardless of the value of no_output_of_prior_pics_flag. Note that under these conditions it may be preferable to set NoOutputOfPriorPicsFlag equal to no_output_of_prior_pics_flag, in which case the decoder under test may set NoOutputOfPriorPicsFlag to 1. Otherwise, NoOutputOfPriorPicsFlag may be set equal to no_output_of_prior_pics_flag.

[0122] The derived NoOutputOfPriorPicsFlag value for the decoder under test is applied to the HRD, and when the resulting value of NoOutputOfPriorPicsFlag is equal to 1, all picture storage buffers in the DPB are emptied without outputting the pictures they contain, and DPB fullness is set to 0. If both of the following conditions apply to any picture k in the DPB, all such pictures k in the DPB are removed from the DPB: Picture k is marked as unused for reference, and picture k has a PictureOutputFlag equal to 0, or the corresponding DPB output time is less than or equal to the CPB removal time of the first decode unit (denoted as decode unit m) of the current picture n. This may occur if DpbOutputTime[k] is less than or equal to DuCpbRemovalTime[m]. For each picture removed from the DPB, DPB fullness is decremented by 1.

[0123] In one example, outputting and removing a picture from the DPB is described as follows: Outputting and removing a picture from the DPB before decoding the current picture (but after parsing the slice header of the first slice of the current picture) may occur when the first decoding unit of the access unit containing the current picture is removed from the CPB and proceeds as follows: The decoding process for reference picture list construction and the decoding process for reference picture marking are invoked.

[0124] If the current picture is a CLVSS picture that is not picture 0, the following ordered steps are applied: For the decoder under test, the variable NoOutputOfPriorPicsFlag is obtained as follows: If the values of pic_width_in_luma_samples, pic_height_max_in_luma_samples, chroma_format_idc, separate_colour_plane_flag, bit_depth_luma_minus8, bit_depth_chroma_minus8 or sps_max_dec_pic_reconstruction_buffering_minus1[Htid] derived from the SPS differ from the values of pic_width_in_luma_samples, pic_height_in_luma_samples, chroma_format_idc, separate_colour_plane_flag, bit_depth_luma_minus8, bit_depth_chroma_minus8 or sps_max_dec_pic_buffering_minus1[Htid], respectively, derived from the SPS referenced by the previous picture, NoOutputOfPriorPicsFlag may be set to 1 by the decoder under test regardless of the value of no_output_of_prior_pics_flag. Note that under these conditions it is preferable to set NoOutputOfPriorPicsFlag equal to no_output_of_prior_pics_flag, but in this case the decoder under test may set NoOutputOfPriorPicsFlag to 1. Otherwise, it may set NoOutputOfPriorPicsFlag equal to no_output_of_prior_pics_flag.

[0125] The value of NoOutputOfPriorPicsFlag derived for the decoder under test can be applied to the HRD as follows: If NoOutputOfPriorPicsFlag is equal to 1, all picture storage buffers in the DPB are emptied without output of the pictures they contain and DPB fullness is set to 0. Otherwise (NoOutputOfPriorPicsFlag is equal to 0), all picture storage buffers containing pictures marked as not needed for output and not used for reference are emptied (no output), and all non-empty picture storage buffers in the DPB are emptied by repeatedly invoking the bumping process and DPB fullness is set to 0.

[0126] Otherwise (the current picture is not a CLVSS picture), all picture storage buffers containing pictures marked as not needed for output and unused for reference are emptied (no output). For each picture storage buffer emptied, DPB fullness is decremented by 1. When one or more of the following conditions are true, the bumping process is invoked repeatedly, further decrementing DPB fullness by 1 for each additional picture storage buffer emptied until none of the following conditions are true: A condition is that the number of pictures in the DPB marked as needed for output is greater than sps_max_num_reorder_pics[Htid]. Another condition is that there is at least one picture in the DPB marked as needed for output whose associated variable PicLatencyCount is greater than or equal to SpsMaxLatencyPictures[Htid] and whose associated variable PicLatencyCount is greater than or equal to SpsMaxLatencyPictures[Htid]. Another condition is that the number of pictures in the DPB is equal to or greater than SubDpbSize[Htid].

[0127] An exemplary general SEI message syntax is as follows:

[0128] [Table 1]

[0129] An example of scalable nesting SEI message syntax is below:

[0130] [Table 2]

[0131] Examples of scalable nesting SEI message semantics are as follows: A scalable nesting SEI message provides a mechanism for associating an SEI message with a particular layer in the context of a particular OLS or with a particular layer outside the context of an OLS. A scalable nesting SEI message contains one or more SEI messages. An SEI message contained in a scalable nesting SEI message is also referred to as a scalably nested SEI message. Bitstream conformance may require that the following restrictions apply when an SEI message is included in a scalable nesting SEI message:

[0132] An SEI message with payloadType of 132 (decoded picture hash) or 133 (scalable nesting) should not be included in a scalable nesting SEI message. If a scalable nesting SEI message contains a buffering period, picture timing, or decode unit information SEI message, the scalable nesting SEI message should not contain any other SEI messages with payloadType other than 0 (buffering period), 1 (picture timing), or 130 (decode unit information).

[0133] Bitstream conformance may also require that the following restrictions apply to the value of nal_unit_type of SEI NAL units that contain scalable nesting SEI messages: If the scalable nesting SEI message contains an SEI message with payloadType equal to 0 (buffering duration), 1 (picture timing), 130 (decoded unit information), 145 (dependent RAP indication), or 168 (frame field information), the SEI NAL unit that contains the scalable nesting SEI message SHOULD have nal_unit_type equal to PREFIX_SEI_NUT. If the scalable nesting SEI message contains an SEI message with payloadType equal to 132 (decoded picture hash), the SEI NAL unit that contains the scalable nesting SEI message SHOULD have nal_unit_type set equal to SUFFIX_SEI_NUT.

[0134] nesting_ols_flag may be set equal to 1 to specify that the scalably nested SEI message applies to a particular layer in the context of a particular OLS. nesting_ols_flag may be set equal to 0 to specify that the scalably nested SEI message applies to a particular layer in general (e.g., not in the context of an OLS).

[0135] Bitstream conformance may require that the following restrictions apply to the value of nesting_ols_flag: If the scalable nesting SEI message contains an SEI message with payloadType equal to 0 (buffering duration), 1 (picture timing), or 130 (decode unit information), the value of nesting_ols_flag shall be equal to 1. When the scalable nesting SEI message contains an SEI message with payloadType equal to a value in VclAssociatedSeiList, the value of nesting_ols_flag shall be equal to 0.

[0136] nesting_num_olss_minus1 plus 1 specifies the number of OLSs to which the scalably nested SEI message applies. The value of nesting_num_olss_minus1 must be in the range from 0 to TotalNumOlss-1, inclusive. nesting_ols_idx_delta_minus1[i] is used to derive the variable NestingOlsIdx[i], which specifies the OLS index of the ith OLS to which the scalably nested SEI message applies, when nesting_ols_flag is 1. The value of nesting_ols_idx_delta_minus1[i] must be in the range from 0 to TotalNumOlss-2, inclusive. The variable NestingOlsIdx[i] can be derived as follows: if(i==0) NestingOlsIdx[i]=nesting_ols_idx_delta_minus1[i] else NestingOlsIdx[i]=NestingOlsIdx[i-1]+nesting_ols_idx_delta_minus1[i]+1

[0137] nesting_num_ols_layers_minus1[i]+1 specifies the number of layers to which the scalably nested SEI message applies in the context of the NestingOlsIdx[i]th OLS. The value of nesting_num_ols_layers_minus1[i] must be in the range from 0 to NumLayersInOls[NestingOlsIdx[i]]-1, inclusive.

[0138] nesting_ols_layer_idx_delta_minus1[i][j] is used to derive the variable NestingOlsLayerIdx[i][j], which specifies the OLS layer index of the jth layer to which the scalably nested SEI message applies, in the context of the NestingOlsIdx[i]th OLS, if nesting_ols_flag is 1. The value of nesting_ols_layer_idx_delta_minus1[i] must be in the range from 0 to NumLayersInOls[nestingOlsIdx[i]]-2, inclusive.

[0139] The variable NestingOlsLayerIdx[i][j] can be derived as follows: if(j==0) NestingOlsLayerIdx[i][j]=nesting_ols_layer_idx_delta_minus1[i][j] else NestingOlsLayerIdx[i][j]=NestingOlsLayerIdx[i][j-1]+ nesting_ols_layer_idx_delta_minus1[i][j]+1

[0140] The lowest value of all values of LayerIdInOls[NestingOlsIdx[i]][NestingOlsLayerIdx[i][0]], for i in the range of 0 to nesting_num_olss_minus1, inclusive, MUST be equal to the nuh_layer_id of the current SEI NAL unit (e.g., the SEI NAL unit that contains the scalable nesting SEI message). nesting_all_layers_flag may be set equal to 1 to specify that the scalably nested SEI message generally applies to all layers with a nuh_layer_id greater than or equal to the nuh_layer_id of the current SEI NAL unit. nesting_all_layers_flag may be set equal to 0 to specify that the scalably nested SEI message generally may or may not apply to all layers with a nuh_layer_id greater than or equal to the nuh_layer_id of the current SEI NAL unit.

[0141] nesting_num_layers_minus1 plus 1 specifies the number of layers to which a scalably nested SEI message generally applies. If nuh_layer_id is the nuh_layer_id of the current SEI NAL unit, the value of nesting_num_layers_minus1 MUST be in the range of 0 to vps_max_layers_minus1-GeneralLayerIdx[nuh_layer_id]. nesting_layer_id[i] specifies the nuh_layer_id value of the ith layer to which a scalably nested SEI message generally applies if nesting_all_layers_flag is 0. The value of nesting_layer_id[i] MUST be greater than nuh_layer_id, where nuh_layer_id is the nuh_layer_id of the current SEI NAL unit.

[0142] If nesting_ols_flag is equal to 1, the variable NestingNumLayers, which specifies the number of layers to which a scalably nested SEI message generally applies, and the list NestingLayerId[i], for i in the range 0 to NestingNumLayers-1, which specifies a list of nuh_layer_id values of the layers to which a scalably nested SEI message generally applies, are derived as follows, where nuh_layer_id is the nuh_layer_id of the current SEI NAL unit: if(nesting_all_layers_flag){ NestingNumLayers=vps_max_layers_minus1+1-GeneralLayerIdx[nuh_layer_id] for(i=0;i <NestingNumLayers;i++) NestingLayerId[i]=vps_layer_id[GeneralLayerIdx[nuh_layer_id]+i](D-2) }else{ NestingNumLayers=nesting_num_layers_minus1+1 for(i=0;i <NestingNumLayers;i++) NestingLayerId[i]=(i==0)?nuh_layer_id:nesting_layer_id[i] }

[0143] nesting_num_seis_minus1 plus 1 specifies the number of scalably nested SEI messages. The value of nesting_num_seis_minus1 can be any value in the range 0 to 63, inclusive. nesting_0_bit should be set equal to 0.

[0144] FIG. 8 is a schematic diagram of an exemplary video coding device 800. The video coding device 800 is suitable for implementing the disclosed examples / embodiments as described herein. The video coding device 800 includes a downstream port 820, an upstream port 850, and / or a transceiver unit (Tx / Rx) 810 including a transmitter and / or receiver for communicating data upstream and / or downstream across a network. The video coding device 800 also includes a processor 830 including a logic unit and / or central processing unit (CPU) for processing data and a memory 832 for storing data. The video coding device 800 may also include electrical, optical-electrical (OE) components, electrical-optical (EO) components, and / or wireless communication components coupled to the upstream port 850 and / or downstream port 820 for data communication over an electrical, optical, or wireless communication network. The video coding device 800 may also include input and / or output (I / O) devices 860 for exchanging data with a user. The I / O devices 860 may include output devices such as a display for displaying video data, speakers for outputting audio data, etc. The I / O devices 860 may also include input devices such as a keyboard, mouse, trackball, etc., and / or corresponding interfaces for interacting with such output devices.

[0145] The processor 830 is implemented by hardware and software. The processor 830 may be implemented as one or more CPU chips, cores (e.g., as a multi-core processor), field programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), and digital signal processors (DSPs). The processor 830 communicates with the downstream port 820, the Tx / Rx 810, the upstream port 850, and the memory 832. The processor 830 includes a coding module 814. The coding module 814 implements the disclosed embodiments described herein, such as methods 100, 900, and 1000, which may use the multi-layer video sequence 600 and / or the bitstream 700. The coding module 814 may also implement any other method / mechanism described herein. Additionally, the coding module 814 may implement the codec system 200, the encoder 300, the decoder 400, and / or the HRD 500. For example, coding module 814 may signal and / or read various parameters as described herein. Furthermore, coding module 814 may be used to encode and / or decode video sequences based on such parameters. Accordingly, the signaling modifications described herein may increase efficiency and / or avoid errors in coding module 814. Accordingly, coding module 814 may be configured to implement mechanisms to address one or more of the above-mentioned problems. Thus, coding module 814 allows video coding device 800 to provide additional functionality and / or coding efficiency when coding video data. In this manner, coding module 814 improves the functionality of video coding device 800 while addressing problems inherent in video coding techniques. Furthermore, coding module 814 achieves transformation of video coding device 800 to a different state. Alternatively, coding module 814 may be implemented as instructions stored in memory 832 (e.g., computer program code stored on a non-transitory medium). The program may be executed by the processor 830 (as a computer program product).

[0146] Memory 832 includes one or more memory types such as a disk, a tape drive, a solid-state drive, read-only memory (ROM), random access memory (RAM), flash memory, ternary content addressable memory (TCAM), static random access memory (SRAM), etc. Memory 832 may be used as an overflow data storage device for storing such programs when such programs are selected for execution, and for storing instructions and data read during the execution of the programs.

[0147] 9 is a flowchart of an example method 900 for encoding a video sequence into a bitstream, such as bitstream 700, by using a suffix SEI message that includes a scalable nesting SEI message. Method 900 may be used in an encoder, such as codec system 200, encoder 300, and / or video coding device 800, when performing method 100. Furthermore, method 900 may operate on HRD 500 and, therefore, may perform conformance testing on multi-layer video sequence 600.

[0148] Method 900 may begin when an encoder receives a video sequence and determines, for example, based on user input, to encode the video sequence into a multi-layer bitstream. In step 901, the encoder encodes one or more coded pictures in one or more VCL NAL units in the bitstream. For example, the coded pictures may be included in an AU within a layer. Furthermore, the encoder may encode one or more layers containing coded pictures into a multi-layer bitstream. A layer may include a set of VCL NAL units with the same layer ID and associated non-VCL NAL units. For example, a set of VCL NAL units is part of a layer if the set of VCL NAL units all have a particular value of nuh_layer_id. A layer may include a set of VCL NAL units containing video data of an encoded picture and any parameter set used to code such a picture. These layers may be included in one or more OLSs. One or more layers may be output layers (e.g., each OLS includes at least one output layer). Layers that are not output layers are coded to support the reconstruction of the output layers, but such support layers are not intended for output at the decoder. In this way, the encoder can code various combinations of layers to send to the decoder on demand. Layers can be sent as needed to allow the decoder to obtain different representations of the video sequence depending on network conditions, hardware capabilities, and / or user preferences.

[0149] In step 903, the encoder may code one or more non-VCL NAL units into the bitstream. For example, a layer and / or a set of layers also includes various non-VCL NAL units. The non-VCL NAL units are associated with a set of VCL NAL units that all have a particular value of nuh_layer_id. Specifically, the encoder may code an SEI NAL unit with nal_unit_type equal to SUFFIX_SEI_NUT, where the SEI NAL unit includes a scalable nesting SEI message. In other words, the encoder may encode a scalable nesting SEI message into a suffix SEI message. The scalable nesting SEI message includes one or more scalably nested SEI messages. The scalably nested SEI message may be any type of SEI message that can be included in a suffix SEI message. For example, a scalable nested SEI message may include a decoded picture hash SEI message (e.g., or a BP SEI, PT SEI, and / or DUI SEI message). Thus, the scalable nesting SEI message associates the suffix SEI message with a particular OLS and / or a particular layer, depending on the example. Note that the scalable nesting SEI message (within the suffix SEI message) may include a scalable nesting layer_id[i] that specifies the nuh_layer_id value of the ith layer to which the scalably nested SEI message applies. Thus, the scalable nesting layer_id[i] can associate the scalably nested SEI message within the scalable nesting SEI message with a layer. In some examples, the scalable nesting SEI message within the suffix SEI message may include a payloadType set to 133. PayloadType is a syntax element that indicates the type of data included in the SEI message, and therefore indicates the type of SEI message included in the SEI NAL unit.Thus, the payloadType may indicate that the suffix SEI message includes a scalable nesting SEI message and / or one or more scalably nested SEI messages, which are thus encoded to apply to NAL units, such as VCL NAL units, that precede the scalable nesting SEI message and / or one or more scalably nested SEI messages.

[0150] In step 905, the encoder can use the HRD to perform a set of bitstream conformance tests on the bitstream based on the scalable nesting SEI message. The set of bitstream conformance tests may include one or more tests. For example, the HRD can use the nal_unit_type to determine the suffix SEI message that contains the scalable nesting SEI message and / or one or more scalably nested SEI messages. Furthermore, the HRD can use the scalable nesting layer_id[i] and / nuh_layer_id values to correlate the scalably nested SEI messages in the suffix SEI message to coded pictures, layers, and / or OLSs. Thus, the HRD can use parameters from the SEI message to perform one or more conformance tests on coded pictures, layers, and / or OLSs starting from the VCL NAL unit immediately preceding the suffix SEI message. The encoder can then store the bitstream for communication to the decoder in step 907. The encoder can also send the bitstream to a decoder as needed.

[0151] 10 is a flowchart of an example method 1000 of decoding a video sequence from a bitstream, such as bitstream 700, using a suffix SEI message that includes a scalable nesting SEI message. Method 1000 may be used in a decoder, such as codec system 200, decoder 400, and / or video coding device 800, when performing method 100. Furthermore, method 1000 may be used on a multi-layer video sequence 600 that has been checked for conformance by an HRD, such as HRD 500.

[0152] Method 1000 may begin when a decoder begins receiving a bitstream of coded data representing a multi-layer video sequence, e.g., as a result of method 900 and / or in response to a request by the decoder. In step 1001, the decoder receives a bitstream comprising coded pictures in one or more VCL NAL units. Furthermore, the bitstream may include one or more layers containing the coded pictures. A layer may include a set of VCL NAL units with the same layer ID and associated non-VCL NAL units. For example, a set of VCL NAL units is part of a layer if the set of VCL NAL units all have a particular value of nuh_layer_id. A layer may include a set of VCL NAL units containing video data of a coded picture and any parameter set used to code such a picture. One or more layers may be output layers. Layers that are not output layers are coded to support the reconstruction of an output layer, but such support layers are not intended for output. In this way, the decoder can obtain different representations of a video sequence depending on network conditions, hardware capabilities, and / or user settings. A layer also contains various non-VCL NAL units, which are associated with a set of VCL NAL units that all have a particular value of nuh_layer_id.

[0153] Specifically, a bitstream may include an SEI NAL unit with a nal_unit_type equal to SUFFIX_SEI_NUT, indicating that the SEI NAL unit includes a suffix SEI message. The SEI NAL unit may also include a scalable nesting SEI message. In other words, the suffix SEI message includes a scalable nesting SEI message. The scalable nesting SEI message includes one or more scalably nested SEI messages. The scalably nested SEI message may be any type of SEI message that can be included in a suffix SEI message. For example, the scalably nested SEI message may include a decoded picture hash SEI message (e.g., or a BP SEI, PT SEI, and / or DUI SEI message). Thus, the scalable nesting SEI message associates the suffix SEI message with a specific OLS and / or a specific layer, depending on the example. Note that the scalable nesting SEI message (within the suffix SEI message) may include a scalable nesting layer_id[i] that specifies the nuh_layer_id value of the ith layer to which the scalably nested SEI message applies. Thus, the scalable nesting layer_id[i] may associate a scalable nested SEI message within the scalable nesting SEI message with a layer. In some examples, the scalable nesting SEI message within the suffix SEI message may include a payloadType set to 133. The payloadType is a syntax element that indicates the type of data contained in the SEI message, and thus indicates the type of SEI message contained in the SEI NAL unit. Thus, the payloadType may indicate that the suffix SEI message includes a scalable nesting SEI message and / or one or more scalably nested SEI messages.Thus, the received scalable nesting SEI message and / or one or more scalably nested SEI messages are applied to a NAL unit, such as a VCL NAL unit, that precedes the scalable nesting SEI message and / or the scalably nested SEI message.

[0154] In step 1003, the decoder may decode the coded picture from the VCL NAL unit to generate a decoded picture. For example, the decoder may use the nal_unit_type to determine the suffix SEI message that includes the scalable nesting SEI message and / or one or more scalably nested SEI messages. Furthermore, the decoder may use the scalable nesting layer_id[i] and / nuh_layer_id values to correlate the scalably nested SEI messages in the suffix SEI message to the coded picture, layer, and / or OLS. The decoder may then use the scalable nesting SEI message and / or the scalably nested SEI messages from the suffix SEI message, as desired, when decoding the coded picture. For example, the decoder may use the decoded picture hash SEI message to verify that one or more pictures in one or more layers (e.g., including the coded picture) were correctly received and decoded without error. In step 1005, the decoder can transfer the decoded pictures for display as part of the decoded video sequence.

[0155] 11 is a schematic diagram of an exemplary system 1100 for coding a video sequence using a bitstream that uses a suffix SEI message that includes a scalable nesting SEI message. System 1100 may be implemented by an encoder and a decoder, such as codec system 200, encoder 300, decoder 400, and / or video coding device 800. Furthermore, system 1100 may perform conformance testing on multi-layer video sequence 600 and / or bitstream 700 using HRD 500. Furthermore, system 1100 may be used when implementing methods 100, 900, and / or 1000.

[0156] The system 1100 includes a video encoder 1102. The video encoder 1102 comprises an encoding module 1103 for encoding the coded one or more pictures into a bitstream. Further, the encoding module 1103 is for encoding an SEI NAL unit having a nal_unit_type equal to SUFFIX_SEI_NUT and including a scalable nesting SEI message into the bitstream. The video encoder 1102 further comprises an HRD module 1105 for performing a set of bitstream conformance tests on the bitstream based on the scalable nesting SEI message. The video encoder 1102 further comprises a storage module 1106 for storing the bitstream for communication to a decoder. The video encoder 1102 further comprises a transmission module 1107 for transmitting the bitstream toward the video decoder 1110. The video encoder 1102 may be further configured to perform any of the steps of method 900.

[0157] System 1100 also includes a video decoder 1110. The video decoder 1110 comprises a receiving module 1111 for receiving a bitstream comprising one or more coded pictures and an SEI NAL unit having a nal_unit_type equal to SUFFIX_SEI_NUT and including a scalable nesting SEI message. The video decoder 1110 further comprises a decoding module 1113 for decoding the coded pictures to generate decoded pictures. The video decoder 1110 further comprises a transport module 1115 for transporting the decoded pictures for display as part of a decoded video sequence. The video decoder 1110 may be further configured to perform any of the steps of method 1000.

[0158] A first component is directly coupled to a second component when there are no intervening components other than wires, traces, or other intermediaries between the first and second components. A first component is indirectly coupled to a second component when there are intervening components other than wires, traces, or other intermediaries between the first and second components. The term "coupled" and variations thereof include both direct and indirect coupling. The use of the term "about," unless otherwise stated, means a range that includes ±10% of the subsequent numerical value.

[0159] It should also be understood that the steps of the exemplary methods described herein do not necessarily have to be performed in the order described, and the order of steps in such methods should be understood as merely exemplary. Similarly, such methods may include additional steps, and some steps may be omitted or combined, in methods consistent with various embodiments of the present disclosure.

[0160] While several embodiments are provided in this disclosure, it will be understood that the disclosed systems and methods may be embodied in many other specific forms without departing from the spirit or scope of the disclosure. The examples of the disclosure should be considered illustrative rather than limiting, and the intention should not be limited to the details given herein. For example, various elements or components may be combined or integrated in another system, or certain features may be omitted or not implemented.

[0161] Additionally, techniques, systems, subsystems, and methods described and illustrated as separate or distinct in various embodiments may be combined or integrated with other systems, components, techniques, or methods without departing from the scope of the present disclosure. Other examples of changes, substitutions, and alterations will be apparent to those skilled in the art and may be made without departing from the spirit and scope disclosed herein. [Explanation of symbols]

[0162] 100 How it works 200 Codec System 201 segmented video signal 211 Integrated Coder Control Component 213 Transform Scaling Quantization Component 215 Intra-picture Estimation Component 217 Intra-picture Prediction Component 219 Motion Compensation Component 221 Motion Estimation Component 223 Decoded Picture Buffer Component 225 In-Loop Filter Components 227 Filter Control Analysis Component 229 Scaling Inverse Transformation Component 231 Header Formatting CABAC Component 300 Encoder 301 segmented video signal 313 Transform Quantization Component 317 Intra-picture Prediction Component 321 Motion Compensation Component 323 Decoded Picture Buffer Component 325 In-Loop Filter Components 329 Inverse Transform Quantization Component 331 Entropy Coding Component 400 decoder 417 Intra-picture Prediction Component 421 Motion Compensation Component 423 Decoded Picture Buffer Component 425 In-Loop Filter Components 429 Inverse Transform Quantization Component 433 Entropy Decoding Component 500 HRD 600 multi-layer video sequences 611, 612, 613, 614, 615, 616, 617, 618 Pictures 621 Interlayer Prediction 623 Inter Prediction 627 Access Units (AU) 628 OLS 631 Layer N 632 layer N+1 635 nuh_layer_id 700 bitstream 711 VPS 713 SPS 715 PPS 717 Slice Header 718 prefix SEI message 719 Suffix SEI Message 720 image data 723 layers 725 Pictures 727 slices 731 nal_unit_type 733 Payload Type 735 Scalable nesting layer_id[i] 741 Non-VCL NAL Units 742 PREFIX_SEI_NUT 743 SUFFIX_SEI_NUT 744 SEI NAL unit 745 VCL NAL Unit 746 Scalable Nesting SEI Message 747 Scalable Nested SEI Messages 748 Decoded Picture Hash SEI Message 800 Video Coding Device 810 Transceiver Unit (Tx / Rx) 814 Coding Module 820 downstream ports 830 processor 832 memory 850 upstream ports 860 Input and / or Output (I / O) Devices 1100 System 1102 Video Encoder 1103 Encoding Module 1105 HRD module 1106 Storage Module 1107 Transmitting Module 1110 Video Decoder 1111 Receiver Module 1113 Decoding Module 1115 Transfer Module

Claims

1. 1. A method implemented by a decoder, said method comprising: receiving, by a receiver of the decoder, a bitstream including coded pictures and supplemental enhancement information (SEI) network abstraction layer (NAL) units having a NAL unit type (nal_unit_type) equal to a suffix SEI NAL unit type (SUFFIX_SEI_NUT) and including a scalable nesting SEI message, where the scalable nesting SEI message applies to one or more preceding NAL units; decoding, by a processor of the decoder, the coded picture to generate a decoded picture, the processor correlates one or more scalably nested SEI messages included in the scalable nesting SEI message to coded pictures, layers, and / or output layer sets (OLSs) using values of scalable nesting layer identifiers (layer_id[i]); using, by the processor, the one or more scalably nested SEI messages included in the scalable nesting SEI message when decoding the coded picture; A method comprising:

2. The method of claim 1 , wherein the one or more scalably nested SEI messages include a decoded picture hash SEI message.

3. The method of claim 1 or 2, wherein the scalable nesting SEI message includes a payload type set to 133.

4. 4. The method of claim 1, wherein the scalable nesting SEI message includes a scalable nesting layer identifier (layer_id[i]) that specifies the NAL unit header layer identifier (nuh_layer_id) value of the i-th layer to which the scalably nested SEI message applies.

5. A computer-readable storage medium having a program recorded thereon, the program causing a computer to execute the method of any one of claims 1 to 4.

6. A decoding device, a receiver configured to receive a bitstream consisting of a coded picture and supplemental enhancement information (SEI) network abstraction layer (NAL) units having a NAL unit type (nal_unit_type) equal to a suffix SEI NAL unit type (SUFFIX_SEI_NUT) and including a scalable nesting SEI message, the scalable nesting SEI message applying to one or more preceding NAL units; a memory coupled to the receiver, the memory storing instructions; a processor coupled to the memory, the processor configured to execute instructions to cause the decoding device to perform the method of any one of claims 1 to 4; A decoding device comprising:

7. A program that causes a computer to execute the method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • How to encode video data in a scalable way

    JP2010531554A

  • Generic use of hevc SEI messages for multi-layer codecs

    JP2017510198A