Simplifying SEI Message Dependencies in Video Coding
By integrating DU HRD parameters and CPB removal delay parameters in SEI messages, the dependency on the VPS is removed, enhancing coding efficiency and reducing resource usage in video coding systems.
Patent Information
- Application Number
- JP2022519005
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-09-24
- Filing Date
- 2020-09-17
- Publication Date
- 2025-12-15
- Estimated Expiration
- 2040-09-17
AI Technical Summary
Video coding systems face challenges due to dependencies between the video parameter set (VPS) and supplemental enhancement information (SEI) messages, leading to errors when the VPS is omitted, particularly in single-layer transmissions, which affects coding efficiency and resource usage.
Incorporating decoding unit (DU) HRD parameters and DU-level CPB removal delay parameters in the SEI message, allowing it to be parsed independently of the VPS, thereby eliminating dependencies and enabling error-free parsing even when the VPS is omitted.
This approach improves coding efficiency by reducing processor, memory, and network signaling resource usage, while ensuring seamless parsing and decoding of video sequences.
Smart Images

Figure 0007785665000003 
Figure 0007785665000004 
Figure 0007785665000005
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This patent application claims the benefit of U.S. Provisional Patent Application No. 62 / 905,236, entitled "Video Coding Improvements," filed September 24, 2019, by Ye-Kui Wang, which is incorporated herein by reference.
[0002] FIELD This disclosure relates generally to video coding, and more particularly to improvements in signaling parameters to support coding of multi-layer bitstreams. [Background technology]
[0003] The significant amount of video data required to represent even a relatively short video can present challenges when streaming data or transmitting data across communication networks with limited bandwidth capacity. For this reason, video data is typically compressed before transmission over today's communication networks. When video is stored on a storage device, the size of the video can also be an issue because memory resources may be scarce. Video compression devices often use software and / or hardware to encode video data at the source prior to transmission or storage, thereby reducing the amount of data required to represent a digital video image. The compressed data is then received by a video decompression device, which decodes the video data at the destination. Due to limited network resources and increasing demand for higher video quality, improved compression and decompression techniques that increase compression ratios with little or no sacrifice in image quality are desirable. Summary of the Invention [Means for solving the problem]
[0004] In one embodiment, the present disclosure includes a method implemented by a decoder, the method including receiving, by a receiver of the decoder, a bitstream including an encoded picture and a current supplemental enhancement information (SEI) message including a decoding unit (DU) hypothetical reference decoder (HRD) parameters present flag (du_hrd_params_present_flag) specifying whether DU-level HRD parameters are present in the bitstream, and decoding, by a processor of the decoder, the encoded picture to generate a decoded picture.
[0005] A video coding system can encode a video sequence as a series of coded pictures into a bitstream. Various parameters can also be coded to support the decoding of the video sequence. For example, a video parameter set (VPS) can contain parameters related to the configuration of layers, sublayers, and / or output layer sets (OLSs) within the video sequence. Furthermore, a video sequence can be checked for conformance to a standard by an HRD. To support such conformance testing, the VPS and / or SPS can contain HRD parameters. HRD-related parameters can also be included in an SEI message. The SEI message contains information not required by the decoding process to determine the values of samples in a decoded picture. For example, the SEI message can contain HRD parameters that further describe the HRD process in light of the HRD parameters included in the VPS. In some video coding systems, the SEI message can contain parameters that directly reference the VPS. This dependency poses certain challenges. For example, the VPS can be removed from the bitstream when transmitting an OLS containing a single layer. This approach can be beneficial in some cases because the VPS does not contain useful information when a decoder receives only one layer. However, omitting the VPS may prevent the SEI message from being properly parsed due to its dependency on the VPS. Specifically, omitting the VPS may cause the SEI message to return an error because the decoder does not receive the data on which the SEI message depends in the VPS.
[0006] This example includes a mechanism for removing the dependency between the VPS and the SEI message. For example, the du_hrd_params_present_flag can be coded in the current SEI message. The du_hrd_params_present_flag specifies whether the HRD operates at the access unit (AU) level or the DU level. Furthermore, the current SEI message can include a DU coded picture buffer (CPB) parameter-in-picture timing (PT) SEI flag (du_cpb_params_in_pic_timing_sei_flag), which specifies whether the DU-level CPB removal delay parameter is present in a PT SEI message or a decoding unit information (DUI) SEI message. Including these flags in the current SEI message makes the current SEI message independent of the VPS. Therefore, the current SEI message can be parsed even if the VPS is omitted from the bitstream. Therefore, various errors can be avoided. This improves the performance of encoders and decoders. Furthermore, removing the dependency between the SEI message and the VPS supports the removal of the VPS in certain cases, which improves coding efficiency and therefore reduces processor, memory, and / or network signaling resource usage in both the encoder and decoder.
[0007] Optionally, in any of the above-mentioned aspects, in another implementation of the aspect, du_hrd_params_present_flag further specifies whether the HRD operates at an access unit (AU) level or a DU level.
[0008] Optionally, in any of the above-described aspects, in another implementation of the aspect, du_hrd_params_present_flag is set to 1 when a DU-level HRD parameter is present and specifies that the HRD can operate at the AU level or the DU level, and is set to 0 when a DU-level HRD parameter is not present and specifies that the HRD operates at the AU level.
[0009] Optionally, in any of the aforementioned aspects, in another implementation of the aspect, the current SEI message further includes a DU coded picture buffer (CPB) parameters-in-picture timing (PT) SEI flag (du_cpb_params_in_pic_timing_sei_flag) that specifies whether DU-level CPB removal delay parameters are present in the PT SEI message.
[0010] Optionally, in any of the above-mentioned aspects, in another implementation form of the aspect, du_cpb_params_in_pic_timing_sei_flag further specifies whether DU-level CPB removal delay parameters are present in the decoding unit information (DUI) SEI message.
[0011] Optionally, in any of the above-mentioned aspects, in another implementation of the aspect, du_cpb_params_in_pic_timing_sei_flag is set to 1 when a DU-level CPB removal delay parameter is present in the PT SEI message specifying that the DUI SEI message is not available, and du_cpb_params_in_pic_timing_sei_flag is set to 0 when a DU-level CPB removal delay parameter is present in the DUI SEI message specifying that the PT SEI message does not include a DU-level CPB removal delay parameter.
[0012] Optionally, in any of the aforementioned aspects, in another implementation of the aspect, the current SEI message is a buffering period (BP) SEI message, a PT SEI message, or a DUI SEI message.
[0013] In one embodiment, the present disclosure includes a method implemented by an encoder, the method including the steps of: encoding, by a processor of the encoder, an encoded picture into a bitstream; encoding, by the processor, a current SEI message into the bitstream, the current SEI message including du_hrd_params_present_flag specifying whether DU level HRD parameters are present in the bitstream; performing, by the processor, a set of bitstream conformance tests on the bitstream based on the current SEI message; and storing, by a memory coupled to the processor, the bitstream for communication to a decoder.
[0014] A video coding system can encode a video sequence as a series of coded pictures into a bitstream. Various parameters can also be coded to support the decoding of the video sequence. For example, a VPS can include parameters related to the configuration of layers, sublayers, and / or an OLS within the video sequence. Furthermore, a video sequence can be checked for conformance to a standard by an HRD. To support such conformance testing, the VPS and / or SPS can include HRD parameters. HRD-related parameters can also be included in an SEI message. The SEI message contains information not required by the decoding process to determine the values of samples in a decoded picture. For example, the SEI message can include HRD parameters that further describe the HRD process in light of the HRD parameters included in the VPS. In some video coding systems, the SEI message can include parameters that directly reference the VPS. This dependency poses certain challenges. For example, the VPS can be removed from the bitstream when transmitting an OLS containing a single layer. This approach can be beneficial in some cases because the VPS does not contain useful information when a decoder receives only one layer. However, omitting the VPS may prevent the SEI message from being properly parsed due to its dependency on the VPS. Specifically, omitting the VPS may cause the SEI message to return an error because the decoder does not receive the data on which the SEI message depends in the VPS.
[0015] This example includes a mechanism for removing the dependency between the VPS and the SEI message. For example, du_hrd_params_present_flag can be encoded in the current SEI message. du_hrd_params_present_flag specifies whether the HRD operates at the AU level or the DU level. Furthermore, the current SEI message can include du_cpb_params_in_pic_timing_sei_flag, which specifies whether the DU-level CPB removal delay parameter is present in the PT SEI message or the DUI SEI message. By including these flags in the current SEI message, the current SEI message does not depend on the VPS. Therefore, the current SEI message can be parsed even if the VPS is omitted from the bitstream. Therefore, various errors can be avoided. As a result, the functionality of the encoder and decoder is improved. Furthermore, removing the dependency between the SEI message and the VPS supports the removal of the VPS in certain cases, which improves coding efficiency and therefore reduces processor, memory, and / or network signaling resource usage in both the encoder and the decoder.
[0016] Optionally, in any of the aforementioned aspects, in another implementation of the aspect, du_hrd_params_present_flag further specifies whether the HRD operates at the AU level or the DU level.
[0017] Optionally, in any of the above-described aspects, in another implementation of the aspect, du_hrd_params_present_flag is set to 1 when a DU-level HRD parameter is present and specifies that the HRD can operate at the AU level or the DU level, and is set to 0 when a DU-level HRD parameter is not present and specifies that the HRD operates at the AU level.
[0018] Optionally, in any of the aforementioned aspects, in another implementation form of the aspect, the current SEI message further includes du_cpb_params_in_pic_timing_sei_flag that specifies whether a DU-level CPB removal delay parameter is present in the PT SEI message.
[0019] Optionally, in any of the above-mentioned aspects, in another implementation of the aspect, du_cpb_params_in_pic_timing_sei_flag further specifies whether a DU-level CPB removal delay parameter is present in the DUI SEI message.
[0020] Optionally, in any of the above-mentioned aspects, in another implementation of the aspect, du_cpb_params_in_pic_timing_sei_flag is set to 1 when a DU-level CPB removal delay parameter is present in the PT SEI message specifying that the DUI SEI message is not available, and du_cpb_params_in_pic_timing_sei_flag is set to 0 when a DU-level CPB removal delay parameter is present in the DUI SEI message specifying that the PT SEI message does not include a DU-level CPB removal delay parameter.
[0021] Optionally, in any of the aforementioned aspects, in another implementation of the aspect, the current SEI message is a BP SEI message, a PT SEI message, or a DUI SEI message.
[0022] In one embodiment, the present disclosure includes a video encoding apparatus comprising a processor, a receiver coupled to the processor, a memory coupled to the processor, and a transmitter coupled to the processor, wherein the processor, receiver, memory, and transmitter are configured to perform the method of any of the aforementioned aspects.
[0023] In one embodiment, the present disclosure includes a non-transitory computer-readable medium including a computer program product for use by a video encoding device, the computer program product including computer-executable instructions stored on the non-transitory computer-readable medium, the computer-executable instructions, when executed by a processor, causing the video encoding device to perform the method of any of the aforementioned aspects.
[0024] In one embodiment, the present disclosure includes a decoder comprising receiving means for receiving a bitstream including an encoded picture and a current SEI message including a du_hrd_params_present_flag specifying whether DU level HRD parameters are present in the bitstream, decoding means for decoding the encoded picture to generate a decoded picture, and forwarding means for forwarding the decoded picture for display as part of a decoded video sequence.
[0025] A video coding system can encode a video sequence as a series of coded pictures into a bitstream. Various parameters can also be coded to support the decoding of the video sequence. For example, a VPS can include parameters related to the configuration of layers, sublayers, and / or an OLS within the video sequence. Furthermore, a video sequence can be checked for conformance to a standard by an HRD. To support such conformance testing, the VPS and / or SPS can include HRD parameters. HRD-related parameters can also be included in an SEI message. The SEI message contains information not required by the decoding process to determine the values of samples in a decoded picture. For example, the SEI message can include HRD parameters that further describe the HRD process in light of the HRD parameters included in the VPS. In some video coding systems, the SEI message can include parameters that directly reference the VPS. This dependency poses certain challenges. For example, the VPS can be removed from the bitstream when transmitting an OLS containing a single layer. This approach can be beneficial in some cases because the VPS does not contain useful information when a decoder receives only one layer. However, omitting the VPS may prevent the SEI message from being properly parsed due to its dependency on the VPS. Specifically, omitting the VPS may cause the SEI message to return an error because the decoder does not receive the data on which the SEI message depends in the VPS.
[0026] This example includes a mechanism for removing the dependency between the VPS and the SEI message. For example, du_hrd_params_present_flag can be encoded in the current SEI message. du_hrd_params_present_flag specifies whether the HRD operates at the AU level or the DU level. Furthermore, the current SEI message can include du_cpb_params_in_pic_timing_sei_flag, which specifies whether the DU-level CPB removal delay parameter is present in the PT SEI message or the DUI SEI message. By including these flags in the current SEI message, the current SEI message does not depend on the VPS. Therefore, the current SEI message can be parsed even if the VPS is omitted from the bitstream. Therefore, various errors can be avoided. As a result, the functionality of the encoder and decoder is improved. Furthermore, removing the dependency between the SEI message and the VPS supports the removal of the VPS in certain cases, which improves coding efficiency and therefore reduces processor, memory, and / or network signaling resource usage in both the encoder and the decoder.
[0027] Optionally, in any of the aforementioned aspects, in another implementation of the aspect, the decoder is further configured to perform the method of any of the aforementioned aspects.
[0028] In one embodiment, the present disclosure includes an encoder comprising: encoding means for encoding an encoded picture into a bitstream and encoding into the bitstream a current SEI message including du_hrd_params_present_flag specifying whether DU level HRD parameters are present in the bitstream; HRD means for performing a set of bitstream conformance tests on the bitstream based on the current SEI message; and storage means for storing the bitstream for communication to a decoder.
[0029] A video coding system can encode a video sequence as a series of coded pictures into a bitstream. Various parameters can also be coded to support the decoding of the video sequence. For example, a VPS can include parameters related to the configuration of layers, sublayers, and / or an OLS within the video sequence. Furthermore, a video sequence can be checked for conformance to a standard by an HRD. To support such conformance testing, the VPS and / or SPS can include HRD parameters. HRD-related parameters can also be included in an SEI message. The SEI message contains information not required by the decoding process to determine the values of samples in a decoded picture. For example, the SEI message can include HRD parameters that further describe the HRD process in light of the HRD parameters included in the VPS. In some video coding systems, the SEI message can include parameters that directly reference the VPS. This dependency poses certain challenges. For example, the VPS can be removed from the bitstream when transmitting an OLS containing a single layer. This approach can be beneficial in some cases because the VPS does not contain useful information when a decoder receives only one layer. However, omitting the VPS may prevent the SEI message from being properly parsed due to its dependency on the VPS. Specifically, omitting the VPS may cause the SEI message to return an error because the decoder does not receive the data on which the SEI message depends in the VPS.
[0030] This example includes a mechanism for removing the dependency between the VPS and the SEI message. For example, du_hrd_params_present_flag can be encoded in the current SEI message. du_hrd_params_present_flag specifies whether the HRD operates at the AU level or the DU level. Furthermore, the current SEI message can include du_cpb_params_in_pic_timing_sei_flag, which specifies whether the DU-level CPB removal delay parameter is present in the PT SEI message or the DUI SEI message. By including these flags in the current SEI message, the current SEI message does not depend on the VPS. Therefore, the current SEI message can be parsed even if the VPS is omitted from the bitstream. Therefore, various errors can be avoided. As a result, the functionality of the encoder and decoder is improved. Furthermore, removing the dependency between the SEI message and the VPS supports the removal of the VPS in certain cases, which improves coding efficiency and therefore reduces processor, memory, and / or network signaling resource usage in both the encoder and the decoder.
[0031] Optionally, in any of the aforementioned aspects, the encoder is further configured to perform the method of any of the aforementioned aspects.
[0032] For clarity, any one of the above-described embodiments may be combined with any one or more of the other above-described embodiments to create new embodiments within the scope of the present disclosure.
[0033] These and other features will be more clearly understood from the following detailed description taken in conjunction with the accompanying drawings and claims.
[0034] For a more complete understanding of the present disclosure, reference is now made to the following brief description taken in conjunction with the accompanying drawings and detailed description, in which like reference numerals represent like parts. [Brief explanation of the drawings]
[0035] [Figure 1] 1 is a flowchart of an exemplary method for encoding a video signal. [Figure 2] 1 is a schematic diagram of an exemplary encoding-decoding (codec) system for video encoding. [Figure 3] FIG. 1 is a schematic diagram illustrating an exemplary video encoder. [Figure 4] FIG. 1 is a schematic diagram illustrating an exemplary video decoder. [Figure 5] 1 is a schematic diagram illustrating an exemplary hypothetical reference decoder (HRD). [Figure 6] FIG. 1 is a schematic diagram illustrating an exemplary multi-layer video sequence. [Figure 7] FIG. 2 is a schematic diagram illustrating an exemplary bitstream. [Figure 8] 1 is a schematic diagram of an exemplary video encoding device; [Figure 9] 1 is a flowchart of an example method for encoding a video sequence into a bitstream by using supplemental enhancement information (SEI) messages that may be independent of a video parameter set (VPS). [Figure 10] 1 is a flowchart of an example method for decoding a video sequence from a bitstream using SEI messages that may not rely on a VPS. [Figure 11] 1 is a schematic diagram of an example system for encoding a video sequence using a bitstream that uses SEI messages that may not rely on a VPS. DETAILED DESCRIPTION OF THE INVENTION
[0036] Initially, while exemplary implementations of one or more embodiments are provided below, it should be understood that the disclosed systems and / or methods may be implemented using any number of technologies, whether currently known or in existence. The present disclosure should in no way be limited to the exemplary implementations, drawings, and techniques shown below, including the exemplary designs and implementations shown or described herein, but may be modified within the scope of the appended claims along with their full range of equivalents.
[0037] The following terms are defined as follows, unless used herein in a contrary context. Specifically, the following definitions are intended to further clarify the present disclosure. However, terms may be described differently in different contexts. Therefore, the following definitions should be considered supplementary and should not be considered limiting of other definitions given to such terms herein.
[0038] A bitstream is a sequence of bits containing compressed video data for transmission between an encoder and a decoder. An encoder is a device configured to use an encoding process to compress video data into a bitstream. A decoder is a device configured to use a decoding process to recover video data from the bitstream for display. A picture is an array of luma samples and / or chroma samples that generate a frame or a field thereof. A slice is an integer number of complete tiles or an integer number of contiguous complete coding tree unit (CTU) rows of a picture (e.g., within a tile) contained exclusively in a single network abstraction layer (NAL) unit. For clarity, the picture being coded or decoded can be referred to as the current picture. A coded picture is a coded representation of a picture comprising video coding layer (VCL) NAL units with a particular value of the NAL unit header layer identifier (nuh_layer_id) in the access unit (AU) and including all coding tree units (CTUs) of the picture. A decoded picture is a picture produced by applying a decoding process to a coded picture.
[0039] An AU is a set of coded pictures contained in different layers and associated at the same time for output from the decoded picture buffer (DPB). A decoding unit (DU) is an AU or a subset of an AU that includes one or more VCL NAL units and associated non-VCL NAL units within the AU. A NAL unit is a syntax structure that contains data in the form of a Raw Byte Sequence Payload (RBSP), an indication of the type of data, and is optionally interspersed with emulation prevention bytes. A VCL NAL unit is a NAL unit coded to contain video data, such as a coded slice of a picture. A non-VCL NAL unit is a NAL unit that contains non-video data, such as syntax and / or parameters that support decoding video data, performing conformance checks, or other operations. A layer is a set of VCL NAL units that share specified characteristics (e.g., a common resolution, frame rate, picture size, etc.), as indicated by a layer ID and associated non-VCL NAL units.
[0040] A hypothetical reference decoder (HRD) is a decoder model that operates in an encoder and checks the variability of the bitstream generated by the encoding process to verify conformance to specified constraints. Bitstream conformance tests determine whether the encoded bitstream conforms to a standard such as Versatile Video Coding (VVC). HRD parameters are syntax elements that initialize and / or define the operating conditions of the HRD. HRD parameters may be included in a supplemental enhancement information (SEI) message, a sequence parameter set (SPS), and / or a video parameter set (VPS). The DU HRD parameters present flag (du_hrd_params_present_flag) is a syntax element that specifies whether DU-level HRD parameters are present in the bitstream. The AU level is a description of operations that apply to one or more AUs (e.g., to one or more picture groups that share the same output time). The DU level is a description of operations that apply to one or more DUs (e.g., to one or more pictures).
[0041] An SEI message is a syntax structure with specified semantics that conveys information that is not required in the decoding process to determine the values of samples in a decoded picture. An SEI NAL unit is a NAL unit that contains one or more SEI messages. A particular SEI NAL unit may be referred to as the current SEI NAL unit. A buffering period (BP) SEI message is an SEI message that contains HRD parameters for initializing an HRD to manage the coded picture buffer (CPB). A picture timing (PT) SEI message is a type of SEI message that contains HRD parameters for managing delivery information for AUs in the CPB and / or decoded picture buffer (DPB). A decoding unit information (DUI) SEI message is a type of SEI message that contains HRD parameters for managing delivery information for DUs in the CPB and / or DPB. The DU CPB parameter flag in PT SEI (du_cpb_params_in_pic_timing_sei_flag) is a syntax element that specifies whether the DU-level CPB removal delay parameter is present in the PT SEI message and / or the DUI SEI message. The CPB is a first-in, first-out buffer in the HRD that contains DUs in decoding order. The CPB removal delay is the time that one or more pictures can remain in the CPB before being transferred to the DPB in the HRD. The VPS is a syntax structure that contains data about the entire bitstream. The SPS is a syntax structure that contains syntax elements that apply to zero or more entire coded layer video sequences. A coded video sequence is a set of one or more coded pictures. A decoded video sequence is a set of one or more decoded pictures.
[0042] In this specification, the following acronyms are used: Access Unit (AU), Coding Tree Block (CTB), Coding Tree Unit (CTU), Coding Unit (CU), Coded Layer Video Sequence (CLVS), Coded Layer Video Sequence Start (CLVSS), Coded Video Sequence (CVS), Coded Video Sequence Start (CVSS), Joint Video Experts Team (JVET), Hypothetical Reference Decoder HRD, Motion Constrained Tile Set (MCTS), Maximum Transfer Unit (MTU), Network Abstraction Layer (NAL), Output Layer Set (OLS), Picture Order Count (POC), Random Access Point (RAP), and The following are used: Point, Raw Byte Sequence Payload (RBSP), Sequence Parameter Set (SPS), Video Parameter Set (VPS), and Versatile Video Coding (VVC).
[0043] Many video compression techniques can be used to reduce the size of video files while minimizing data loss. For example, video compression techniques may include performing spatial (e.g., intra-picture) prediction and / or temporal (e.g., inter-picture) prediction to reduce or remove data redundancy in a video sequence. In block-based video coding, a video slice (e.g., a video picture or a portion of a video picture) may be divided into video blocks, which may also be referred to as tree blocks, coding tree blocks (CTBs), coding tree units (CTUs), coding units (CUs), and / or coding nodes. Video blocks in an intra-coded (I) slice of a picture are coded using spatial prediction with respect to reference samples in neighboring blocks in the same picture. Video blocks in an inter-coded unidirectionally predicted (P) or bidirectionally predicted (B) slice of a picture may be coded using spatial prediction with respect to reference samples in neighboring blocks in the same picture or temporal prediction with respect to reference samples in other reference pictures. A picture may be referred to as a frame and / or an image, and a reference picture may be referred to as a reference frame and / or a reference image. Spatial or temporal prediction results in a predicted block representing an image block. Residual data represents pixel differences between the original image block and the predicted block. Thus, inter-coded blocks are coded according to a motion vector pointing to a block of reference samples forming the predicted block and residual data indicating the difference between the coded block and the predicted block. Intra-coded blocks are coded according to an intra-coding mode and residual data. For further compression, the residual data may be transformed from the pixel domain to a transform domain, resulting in residual transform coefficients, which may be quantized. The quantized transform coefficients may initially be arranged in a two-dimensional array. The quantized transform coefficients may be scanned to generate a one-dimensional vector of transform coefficients. To achieve further compression, entropy coding may be applied. Such video compression techniques are described in more detail below.
[0044] To enable accurate decoding of the encoded video, the video is encoded and decoded according to a corresponding video coding standard, including International Telecommunication Union (ITU) Standardization Sector (ITU-T) H.261, International Organization for Standardization / International Electrotechnical Commission (ISO / IEC) Moving Picture Experts Group (MPEG)-1 Part 2, Advanced Video Coding (AVC), also known as ITU-T H.262 or ISO / IEC MPEG-2 Part 2, ITU-T H.263, ISO / IEC MPEG-4 Part 2, ITU-T H.264 or ISO / IEC MPEG-4 Part 10, and High Efficiency Video Coding (HEVC), also known as ITU-T H.265 or MPEG-H Part 2. AVC includes extensions such as Scalable Video Coding (SVC), Multiview Video Coding (MVC), Multiview Video Coding plus Depth (MVC+D), and three-dimensional (3D) AVC (3D-AVC). HEVC includes extensions such as Scalable HEVC (SHVC), Multiview HEVC (MV-HEVC), and 3D HEVC (3D-HEVC). The Joint Video Experts Team (JVET) of ITU-T and ISO / IEC has begun development of a video coding standard called Versatile Video Coding (VVC). VVC is included in working drafts (WDs), including JVET-O2001-v14.
[0045] A video coding system can encode a video sequence as a series of coded pictures into a bitstream. Various parameters can also be coded to support decoding of the video sequence. For example, a video parameter set (VPS) can include parameters related to the configuration of layers, sublayers, and / or output layer sets (OLSs) within the video sequence. Furthermore, a video sequence can be checked for conformance to a standard by a hypothetical reference decoder (HRD). To support such conformance testing, the VPS and / or SPS can include HRD parameters. HRD-related parameters can also be included in supplemental enhancement information (SEI) messages. SEI messages contain information not required by the decoding process to determine the values of samples in decoded pictures. For example, the SEI message can include HRD parameters that further describe the HRD process in light of the HRD parameters included in the VPS. In some video coding systems, the SEI message can include parameters that directly reference the VPS. This dependency poses certain challenges. For example, the VPS can be removed from the bitstream when transmitting an OLS containing a single layer. This approach can be beneficial in some cases because the VPS does not contain useful information when the decoder receives only one layer. However, omitting the VPS can prevent SEI messages from being properly analyzed due to their dependency on the VPS. Specifically, omitting the VPS can cause SEI messages to return errors because the decoder does not receive the data in the VPS on which the SEI messages depend.
[0046] This specification discloses a mechanism for removing the dependency between the VPS and an SEI message. For example, a decoding unit (DU) hypothetical reference decoder (HRD) parameters present flag (du_hrd_params_present_flag) can be coded in the current SEI message. The du_hrd_params_present_flag specifies whether the HRD operates at the access unit (AU) level or the decoding unit (DU) level. Furthermore, the current SEI message can include a DU coded picture buffer (CPB) parameters-in-picture timing (PT) SEI flag (du_cpb_params_in_pic_timing_sei_flag) that specifies whether the DU-level CPB removal delay parameter is present in a PT SEI message or a decoding unit information (DUI) SEI message. By including these flags in the current SEI message, the current SEI message is independent of the VPS. Therefore, the current SEI message can be parsed even if the VPS is omitted from the bitstream. This avoids various errors. As a result, the functionality of the encoder and decoder is improved. Furthermore, removing the dependency between the SEI message and the VPS supports the removal of the VPS in certain cases, which improves coding efficiency and therefore reduces processor, memory, and / or network signaling resource usage in both the encoder and decoder.
[0047] 1 is a flowchart of an exemplary operational method 100 for encoding a video signal. Specifically, a video signal is encoded by an encoder. The encoding process compresses the video signal using various mechanisms to reduce the size of the video file. The smaller file size allows the compressed video file to be transmitted to a user while reducing the associated bandwidth overhead. A decoder then decodes the compressed video file to recover the original video signal for display to the end user. To enable the decoder to reconstruct the video signal consistently, the decoding process typically mirrors the encoding process.
[0048] In step 101, a video signal is input to an encoder. For example, the video signal may be an uncompressed video file stored in memory. As another example, the video file may be captured by a video capture device such as a video camera and encoded to support live streaming of the video. The video file may include both an audio component and a video component. The video component includes a series of image frames that, when viewed in succession, create the visual impression of movement. The frames include pixels, which are represented by light, referred to herein as luma components (or luma samples), and color, referred to herein as chroma components (or color samples). In some examples, the frames may also include depth values to support three-dimensional viewing.
[0049] In step 103, the video is divided into blocks. Partitioning involves subdividing pixels within each frame into square and / or rectangular blocks for compression. For example, in High Efficiency Video Coding (HEVC) (also known as H.265 and MPEG-H Part 2), a frame can first be subdivided into coding tree units (CTUs), which are blocks of a predetermined size (e.g., 64 pixels by 64 pixels). CTUs contain both luma and chroma samples. A coding tree is used to subdivide the CTUs into blocks, which can then be recursively subdivided until a configuration that supports further encoding is achieved. For example, the luma component of a frame may be subdivided until each block contains relatively uniform illumination values. Furthermore, the chroma component of a frame may be subdivided until each block contains relatively uniform color values. Thus, the partitioning mechanism varies depending on the content of the video frame.
[0050] In step 105, the image blocks partitioned in step 103 are compressed using various compression mechanisms. For example, inter-prediction and / or intra-prediction may be used. Inter-prediction is designed to take advantage of the fact that objects in a typical scene tend to appear in consecutive frames. Thus, a block representing an object in a reference frame need not be repeatedly described in adjacent frames. Specifically, an object such as a desk may remain in a constant position across multiple frames. Thus, the desk may be described once, and adjacent frames may reference the reference frame. A pattern matching mechanism may be used to match objects across multiple frames. Furthermore, an object may be depicted as moving across multiple frames, for example, due to object movement or camera movement. As a specific example, a video may show a car moving across the screen across multiple frames. Such movement can be described using a motion vector. A motion vector is a two-dimensional vector that provides an offset from the coordinates of an object in one frame to the coordinates of that object in the reference frame. Thus, inter-prediction allows image blocks in a current frame to be coded as a set of motion vectors indicating their offsets from corresponding blocks in the reference frame.
[0051] Intra prediction encodes blocks within a given frame. It takes advantage of the fact that luma and chroma components tend to cluster together within a frame. For example, green patches in a section of a tree tend to be located near other similar green patches. Intra prediction uses several directional prediction modes (e.g., 33 in HEVC), planar mode, and direct current (DC) mode. Directional mode indicates that the current block is similar to samples from neighboring blocks in the corresponding direction. Planar mode indicates that a series of blocks along a row / column (e.g., a plane) can be interpolated based on neighboring blocks at the edge of the row. Planar mode effectively indicates a smooth transition of light / color across a row / column by using a relatively constant slope of changing values. DC mode is used for boundary smoothing and indicates that the block is similar to the average value associated with samples from all neighboring blocks associated with the angular direction of the directional prediction mode. Therefore, intra-predicted blocks can represent image blocks as various related prediction mode values instead of actual values. Furthermore, inter-predicted blocks can represent image blocks as motion vector values instead of actual values. In either case, the predicted block may not exactly represent the image block. The difference is stored in a residual block, to which a transform can be applied to further compress the file.
[0052] Various filtering techniques may be applied in step 107. In HEVC, filters are applied according to an in-loop filtering scheme. The block-based prediction described above may produce blocky images at the decoder. Furthermore, block-based prediction schemes may encode blocks and reconstruct the encoded blocks for later use as reference blocks. In-loop filtering schemes iteratively apply noise suppression filters, deblocking filters, adaptive loop filters, and sample adaptive offset (SAO) filters to blocks / frames. These filters reduce such blocking artifacts so that the encoded file can be accurately reconstructed. Furthermore, these filters reduce artifacts in the reconstructed reference blocks, reducing the likelihood of further artifacts in subsequent blocks that are coded based on the reconstructed reference blocks.
[0053] After the video signal is segmented, compressed, and filtered, the resulting data is encoded into a bitstream at step 109. This bitstream includes the data described above, as well as any signaling data desired to support proper video signal reconstruction at the decoder. For example, such data may include segmentation data, prediction data, residual blocks, and various flags that provide coding instructions to the decoder. The bitstream may be stored in memory so that it can be transmitted to the decoder upon request. The bitstream may also be broadcast and / or multicast to multiple decoders. Creating the bitstream is an iterative process. Thus, steps 101, 103, 105, 107, and 109 may occur sequentially and / or simultaneously across multiple frames and blocks. The order depicted in FIG. 1 is presented for clarity and simplicity of explanation and is not intended to limit the video encoding process to any particular order.
[0054] The decoder receives the bitstream and begins the decoding process in step 111. Specifically, the decoder uses an entropy decoding scheme to convert the bitstream into corresponding syntax data and video data. In step 111, the decoder uses the syntax data from the bitstream to determine the frame partitioning. The partitioning must match the block partitioning results from step 103. The entropy encoding / decoding used in step 111 is described below. The encoder makes many choices during the compression process, such as selecting a block partitioning scheme from several possible options based on the spatial arrangement of values in the input image. To convey the exact options, it may use multiple bins. As used herein, a bin is a binary value treated as a variable (e.g., a bit value that can change depending on the situation). Entropy coding allows the encoder to discard options that are clearly invalid in a particular situation, leaving a set of acceptable options. Each acceptable option is then assigned a codeword. The length of the codeword is based on the number of acceptable options (e.g., one bin for two options, two bins for three or four options, etc.). The encoder then encodes a codeword of the selected option. This scheme reduces the size of the codeword because it is desirable if the codeword uniquely indicates a choice from a small subset of acceptable options, rather than a choice from a potentially large set of all possible options. The decoder then decodes the selection by determining the set of acceptable options in the same way as the encoder. The decoder can read the codeword and determine the selection made by the encoder by determining the set of acceptable options.
[0055] In step 113, the decoder performs block decoding. Specifically, the decoder generates residual blocks using an inverse transform. The decoder then uses the residual blocks and corresponding prediction blocks to reconstruct image blocks based on the partitioning. The prediction blocks may include both intra-predicted blocks and inter-predicted blocks generated by the encoder in step 105. The reconstructed image blocks are then placed within frames of the reconstructed video signal based on the partitioning data determined in step 111. The syntax of step 113 may also be conveyed in the bitstream using entropy coding as described above.
[0056] In step 115, filtering is performed on the frames of the reconstructed video signal, similar to step 107 in the encoder. For example, noise suppression filters, deblocking filters, adaptive loop filters, and SAO filters can be applied to the frames to remove blocking artifacts. After the frames are filtered, in step 117 the video signal can be output to a display for viewing by an end user.
[0057] 2 is a schematic diagram of an exemplary encoding and decoding (codec) system 200 for video encoding. Specifically, codec system 200 provides functionality supporting implementation of operational method 100. Codec system 200 is generalized to represent components used in both encoders and decoders. Codec system 200 receives and segments a video signal as described with respect to steps 101 and 103 of operational method 100, resulting in segmented video signal 201. Codec system 200 then compresses segmented video signal 201 into an encoded bitstream when functioning as an encoder as described with respect to steps 105, 107, and 109 of method 100. When functioning as a decoder, codec system 200 generates an output video signal from the bitstream as described with respect to steps 111, 113, 115, and 117 of operational method 100. Codec system 200 includes a general coder control component 211, a transform scaling and quantization component 213, an intra-picture estimation component 215, an intra-picture prediction component 217, a motion compensation component 219, a motion estimation component 221, a scaling inverse transform component 229, a filter control analysis component 227, an in-loop filter component 225, a decoded picture buffer component 223, and a header formatting and context adaptive binary arithmetic coding (CABAC) component 231. These components are coupled as shown. In FIG. 2, black lines indicate the movement of data to be coded / decoded, and dashed lines indicate the movement of control data that controls the operation of other components. All of the components of codec system 200 may reside within an encoder. A decoder may include a subset of the components of codec system 200.For example, the decoder may include an intra-picture prediction component 217, a motion compensation component 219, a scaling inverse transform component 229, an in-loop filter component 225, and a decoded picture buffer component 223. These components are described below.
[0058] The segmented video signal 201 is a captured video sequence that has been segmented into blocks of pixels by a coding tree. The coding tree uses various split modes to subdivide blocks of pixels into smaller blocks of pixels. These blocks can then be further subdivided into smaller blocks. Blocks are sometimes referred to as nodes on the coding tree. Larger parent nodes are split into smaller child nodes. The number of times a node is subdivided is referred to as the depth of the node / coding tree. The segmented blocks may be included in coding units (CUs). For example, a CU may be a subpart of a CTU that includes a luma block, one or more red differential chroma (Cr) blocks, and one or more blue differential chroma (Cb) blocks, along with the corresponding CU syntax instructions. Split modes include binary tree (BT), triple tree (TT), and quad tree (QT), which are used to split a node into two, three, or four child nodes, each with a different shape depending on the split mode used. The segmented video signal 201 is forwarded to a general coder control component 211, a transform scaling and quantization component 213, an intra-picture estimation component 215, a filter control analysis component 227, and a motion estimation component 221 for compression.
[0059] The general coder control component 211 is configured to make decisions related to encoding images of a video sequence into a bitstream according to application constraints. For example, the general coder control component 211 manages the optimization of bitrate / bitstream size and reconstruction quality. Such decisions may be made based on storage space / bandwidth availability and image resolution requirements. The general coder control component 211 also manages buffer usage in light of transmission rate to mitigate buffer underrun and overrun issues. To manage these issues, the general coder control component 211 manages segmentation, prediction, and filtering by other components. For example, the general coder control component 211 can dynamically increase compression complexity to increase resolution and bandwidth usage, or decrease compression complexity to decrease resolution and bandwidth usage. Thus, the general coder control component 211 balances the reconstruction quality and bitrate issues of the video signal by controlling other components of the codec system 200. The general coder control component 211 generates control data that controls the operation of the other components. The control data is also forwarded to the header formatting CABAC component 231 and encoded into the bitstream to signal parameters for decoding at the decoder.
[0060] The segmented video signal 201 is also sent to a motion estimation component 221 and a motion compensation component 219 for inter prediction. A frame or slice of the segmented video signal 201 can be subdivided into multiple video blocks. The motion estimation component 221 and the motion compensation component 219 perform inter predictive coding of the received video blocks with reference to one or more blocks in one or more reference frames to provide temporal prediction. The codec system 200 can perform multiple coding passes, for example, to select an appropriate coding mode for each block of video data.
[0061] The motion estimation component 221 and the motion compensation component 219 may be highly integrated but are shown separately for conceptual purposes. Motion estimation, performed by the motion estimation component 221, is the process of generating motion vectors that estimate the motion of video blocks. A motion vector may indicate, for example, the displacement of a coded object relative to a predictive block. A predictive block is a block known to closely match a block to be coded in terms of pixel differences. A predictive block may also be referred to as a reference block. Such pixel differences may be determined by sum of absolute differences (SAD), sum of square differences (SSD), or other difference metrics. HEVC uses several coded objects, including CTUs, coding tree blocks (CTBs), and CUs. For example, a CTU can be divided into CTBs, which can then be subdivided into CBs for inclusion in CUs. A CU can be coded as a prediction unit containing prediction data and / or a transform unit (TU) containing transform residual data for the CU. The motion estimation component 221 generates motion vectors, prediction units, and TUs by using rate-distortion analysis as part of a rate-distortion optimization process. For example, the motion estimation component 221 can determine multiple reference blocks, multiple motion vectors, etc. for a current block / frame and select the reference block, motion vector, etc. with the best rate-distortion characteristics. The best rate-distortion characteristics balance the quality of the video reconstruction (e.g., the amount of data lost due to compression) and the coding efficiency (e.g., the size of the final encoding).
[0062] In some examples, the codec system 200 can calculate values for sub-integer pixel positions of reference pictures stored in the decoded picture buffer component 223. For example, the video codec system 200 can interpolate values for quarter-pixel positions, eighth-pixel positions, or other fractional-pixel positions of the reference picture. Accordingly, the motion estimation component 221 can perform motion searches for full-pixel and fractional-pixel positions and output fractional-pixel-accuracy motion vectors. The motion estimation component 221 calculates motion vectors for prediction units of video blocks in inter-coded slices by comparing the positions of the prediction units with the positions of prediction blocks in the reference pictures. The motion estimation component 221 outputs the calculated motion vectors as motion data to the header formatting CABAC component 231 for encoding and outputs the motion to the motion compensation component 219.
[0063] The motion compensation performed by the motion compensation component 219 may include obtaining or generating a prediction block based on the motion vector determined by the motion estimation component 221. Again, in some examples, the motion estimation component 221 and the motion compensation component 219 may be functionally integrated. After receiving the motion vector of the prediction unit of the current video block, the motion compensation component 219 can identify the location of the prediction block to which the motion vector points. A residual video block is then formed by subtracting pixel values of the prediction block from pixel values of the current video block being coded to form pixel difference values. Generally, the motion estimation component 221 performs motion estimation on the luma component, and the motion compensation component 219 uses the motion vector calculated based on the luma component for both the chroma and luma components. The prediction block and the residual block are forwarded to the transform, scaling, and quantization component 213.
[0064] The segmented video signal 201 is also sent to an intra-picture estimation component 215 and an intra-picture prediction component 217. Similar to the motion estimation component 221 and the motion compensation component 219, the intra-picture estimation component 215 and the intra-picture prediction component 217 may be highly integrated but are shown separately for conceptual purposes. The intra-picture estimation component 215 and the intra-picture prediction component 217 intra-predict the current block relative to blocks within the current frame, instead of the inter-prediction performed between frames by the motion estimation component 221 and the motion compensation component 219, as described above. Specifically, the intra-picture estimation component 215 determines the intra-prediction mode to use to encode the current block. In some examples, the intra-picture estimation component 215 selects an appropriate intra-prediction mode to encode the current block from multiple tested intra-prediction modes. The selected intra-prediction mode is then forwarded to the header formatting CABAC component 231 for encoding.
[0065] For example, the intra-picture estimation component 215 may calculate rate-distortion values for various tested intra-prediction modes using a rate-distortion analysis and select an intra-prediction mode with the best rate-distortion characteristics from among the tested modes. The rate-distortion analysis typically determines the amount of distortion (or error) between an encoded block and the original uncoded block that was coded to generate the encoded block, as well as the bitrate (e.g., number of bits) used to generate the encoded block. The intra-picture estimation component 215 may calculate a ratio from the distortion and rate of the various coded blocks to determine which intra-prediction mode exhibits the best rate-distortion value for the block. Additionally, the intra-picture estimation component 215 may be configured to code depth blocks of a depth map using a rate-distortion optimization (RDO)-based depth modeling mode (DMM).
[0066] The intra picture prediction component 217, when implemented on an encoder, can generate a residual block from the prediction block based on a selected intra prediction mode determined by the intra picture estimation component 215, or, when implemented on a decoder, can read the residual block from the bitstream. The residual block contains the value difference between the prediction block and the original block and is represented as a matrix. The residual block is then forwarded to the transform scaling quantization component 213. The intra picture estimation component 215 and the intra picture prediction component 217 can operate on both the luma and chroma components.
[0067] The transform scaling quantization component 213 is configured to further compress the residual block. The transform scaling quantization component 213 applies a transform, such as a discrete cosine transform (DCT), a discrete sine transform (DST), or a conceptually similar transform, to the residual block to generate a video block containing residual transform coefficient values. A wavelet transform, an integer transform, a subband transform, or other types of transforms may also be used. The transform may convert the residual information from a pixel value domain to a transform domain, such as the frequency domain. The transform scaling quantization component 213 is also configured to scale the transformed residual information, for example, based on frequency. Such scaling involves applying a scale factor to the residual information, which causes different frequency information to be quantized with different granularities, which may affect the final visual quality of the restored video. The transform scaling quantization component 213 is also configured to quantize the transform coefficients to further reduce the bit rate. The quantization process may reduce the bit depth associated with some or all of the coefficients. The degree of quantization can be changed by adjusting a quantization parameter. In some examples, the transform scaling and quantization component 213 can then perform a scan of a matrix containing the quantized transform coefficients, which are forwarded to the header formatting CABAC component 231 for encoding into the bitstream.
[0068] The inverse scaling component 229 applies the inverse operation of the transform scaling quantization component 213 to support motion estimation. The inverse scaling component 229 applies inverse scaling, transformation, and / or quantization to reconstruct a residual block in the pixel domain for later use as a reference block, which may become a prediction block for another current block, for example. The motion estimation component 221 and / or motion compensation component 219 can calculate a reference block by adding the residual block to a corresponding prediction block for use in motion estimation of a later block / frame. A filter is applied to the reconstructed reference block to mitigate artifacts introduced during scaling, quantization, and transformation. Otherwise, such artifacts may cause inaccurate predictions (and further artifacts) when subsequent blocks are predicted.
[0069] The filter control analysis component 227 and the in-loop filter component 225 apply filters to residual blocks and / or reconstructed image blocks. For example, to reconstruct the original image block, a transformed residual block from the inverse scaling transform component 229 can be combined with a corresponding predicted block from the intra-picture prediction component 217 and / or motion compensation component 219. A filter can then be applied to the reconstructed image block. In some examples, a filter can be applied to the residual block instead. Like the other components in FIG. 2, the filter control analysis component 227 and the in-loop filter component 225 may be highly integrated and implemented together, but are depicted separately for conceptual purposes. The filters applied to reconstructed reference blocks are applied to specific spatial regions and include multiple parameters that adjust how such filters are applied. The filter control analysis component 227 analyzes the reconstructed reference blocks to determine when such filters should be applied and sets the corresponding parameters. Such data is forwarded to the header formatting CABAC component 231 as filter control data for encoding. The in-loop filter component 225 applies such filters based on the filter control data. The filters may include a deblocking filter, a noise suppression filter, an SAO filter, and an adaptive loop filter. Such filters may be applied in the spatial / pixel domain (e.g., on the reconstructed pixel blocks) or in the frequency domain, depending on the example.
[0070] When operating as an encoder, the filtered reconstructed image blocks, residual blocks, and / or prediction blocks are stored in the decoded picture buffer component 223 for later use in motion estimation as described above. When operating as a decoder, the decoded picture buffer component 223 stores the reconstructed and filtered blocks and forwards them to a display as part of the output video signal. The decoded picture buffer component 223 may be any memory device capable of storing prediction blocks, residual blocks, and / or reconstructed image blocks.
[0071] The header formatting CABAC component 231 receives data from various components of the codec system 200 and encodes such data into an encoded bitstream for transmission to a decoder. Specifically, the header formatting CABAC component 231 generates various headers to encode control data, such as general control data and filter control data. Additionally, prediction data, including intra-prediction and motion data, and residual data in the form of quantized transform coefficient data are all encoded into the bitstream. The final bitstream contains all the information a decoder needs to reconstruct the original segmented video signal 201. Such information may also include an intra-prediction mode index table (also called a codeword mapping table), definitions of the coding contexts of various blocks, indications of the most probable intra-prediction modes, indications of segmentation information, and so on. Such data may be encoded using entropy coding. For example, the information can be coded using context adaptive variable length coding (CAVLC), CABAC, syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or other entropy coding techniques. After entropy coding, the coded bitstream can be transmitted to another device (e.g., a video decoder) or stored for later transmission or retrieval.
[0072] 3 is a block diagram illustrating an example video encoder 300. Video encoder 300 may be used to perform the encoding functions of codec system 200 and / or to perform steps 101, 103, 105, 107, and / or 109 of method of operation 100. Encoder 300 segments an input video signal, and the resulting segmented video signal 301 is substantially similar to segmented video signal 201. Segmented video signal 301 is then compressed and encoded into a bitstream by components of encoder 300.
[0073] Specifically, the segmented video signal 301 is forwarded to an intra-picture prediction component 317 for intra-prediction. The intra-picture prediction component 317 may be substantially similar to the intra-picture estimation component 215 and the intra-picture prediction component 217. The segmented video signal 301 is also forwarded to a motion compensation component 321 for inter-prediction based on reference blocks in a decoded picture buffer component 323. The motion compensation component 321 may be substantially similar to the motion estimation component 221 and the motion compensation component 219. The prediction blocks and residual blocks from the intra-picture prediction component 317 and the motion compensation component 321 are forwarded to a transform and quantization component 313 for transforming and quantizing the residual blocks. The transform and quantization component 313 may be substantially similar to the transform scaling and quantization component 213. The transformed and quantized residual blocks and corresponding prediction blocks (along with associated control data) are forwarded to an entropy coding component 331 for encoding into a bitstream. The entropy coding component 331 may be substantially similar to the header formatting CABAC component 231 .
[0074] The transformed and quantized residual block and / or the corresponding prediction block are also forwarded from the transform quantization component 313 to the inverse transform quantization component 329 for reconstruction into a reference block used by the motion compensation component 321. The inverse transform quantization component 329 may be substantially similar to the scaling inverse transform component 229. In some examples, an in-loop filter in an in-loop filter component 325 is also applied to the residual block and / or the reconstructed reference block. The in-loop filter component 325 may be substantially similar to the filter control analysis component 227 and the in-loop filter component 225. The in-loop filter component 325 may include multiple filters as described with respect to the in-loop filter component 225. The filtered block is then stored in the decoded picture buffer component 323 for use as a reference block by the motion compensation component 321. The decoded picture buffer component 323 may be substantially similar to the decoded picture buffer component 223.
[0075] 4 is a block diagram illustrating an exemplary video decoder 400. Video decoder 400 may be used to perform the decoding functions of codec system 200 and / or to perform steps 111, 113, 115, and / or 117 of method of operation 100. Decoder 400 receives a bitstream, for example, from encoder 300, and generates a reconstructed output video signal based on the bitstream for display to an end user.
[0076] The bitstream is received by the entropy decoding component 433. The entropy decoding component 433 is configured to perform an entropy decoding scheme, such as CAVLC, CABAC, SBAC, PIPE coding, or other entropy coding techniques. For example, the entropy decoding component 433 can use header information to provide context for interpreting additional data encoded as codewords in the bitstream. The decoded information includes desired information for decoding the video signal, such as general control data, filter control data, segmentation information, motion data, prediction data, and quantized transform coefficients from residual blocks. The quantized transform coefficients are forwarded to the inverse transform and quantization component 429 for reconstruction into residual blocks. The inverse transform and quantization component 429 can be similar to the inverse transform and quantization component 329.
[0077] The reconstructed residual block and / or predictive block are forwarded to an intra-picture prediction component 417 for reconstruction into an image block based on an intra-prediction operation. The intra-picture prediction component 417 may be similar to the intra-picture estimation component 215 and the intra-picture prediction component 217. Specifically, the intra-picture prediction component 417 uses the prediction mode to locate a reference block within a frame and applies the residual block to the result to reconstruct an intra-predicted image block. The reconstructed intra-predicted image block and / or residual block and corresponding inter-prediction data are forwarded to the decoded picture buffer component 423 via an in-loop filter component 425, which may be substantially similar to the decoded picture buffer component 223 and the in-loop filter component 225, respectively. The in-loop filter component 425 filters the reconstructed image block, residual block, and / or predictive block, and such information is stored in the decoded picture buffer component 423. The reconstructed image blocks are transferred from the decoded picture buffer component 423 to the motion compensation component 421 for inter prediction. The motion compensation component 421 may be substantially similar to the motion estimation component 221 and / or the motion compensation component 219. Specifically, the motion compensation component 421 generates a prediction block using a motion vector from a reference block and applies a residual block to the result to reconstruct the image block. The resulting reconstructed block may also be transferred to the decoded picture buffer component 423 via the in-loop filter component 425. The decoded picture buffer component 423 continues to store additional reconstructed image blocks, which can be reconstructed into frames using the partition information. Such frames may also be arranged in a sequence. This sequence is output to a display as a reconstructed output video signal.
[0078] 5 is a schematic diagram illustrating an exemplary HRD 500. The HRD 500 may be used in codec system 200 and / or an encoder, such as encoder 300. The HRD 500 may inspect the bitstream generated in step 109 of method 100 before the bitstream is forwarded to a decoder, such as decoder 400. In some examples, the bitstream may be continuously forwarded through the HRD 500 as it is encoded. If a portion of the bitstream does not conform to an associated constraint, the HRD 500 may indicate such anomaly to the encoder so that the encoder re-encodes the corresponding portion of the bitstream with another mechanism.
[0079] The HRD 500 includes a hypothetical stream scheduler (HSS) 541. The HSS 541 is a component configured to execute a virtual distribution mechanism. The virtual distribution mechanism is used to check the conformance of a bitstream or a decoder with respect to the timing and data flow of a bitstream 551 input to the HRD 500. For example, the HSS 541 may receive the bitstream 551 output from an encoder and manage the process of conformance testing of the bitstream 551. In a particular example, the HSS 541 may control the rate at which coded pictures pass through the HRD 500 and verify that the bitstream 551 does not contain non-conforming data.
[0080] The HSS 541 may transfer the bitstream 551 to the CPB 543 at a predetermined rate. The HRD 500 may manage data in decoding units (DUs) 553. A DU 553 is a subset of an access unit (AU) or an AU and associated non-video coding layer (VCL) network abstraction layer (NAL) units. Specifically, an AU includes one or more pictures associated with an output time. For example, an AU may include a single picture in a single-layer bitstream or a picture of each layer in a multi-layer bitstream. Each picture in an AU may be subdivided into slices, each of which is included in a corresponding VCL NAL unit. Thus, a DU 553 may include one or more pictures, one or more slices of a picture, or a combination thereof. Additionally, parameters used to decode the AU, picture, and / or slice may be included in the non-VCL NAL units. Thus, DU 553 contains non-VCL NAL units that contain data necessary to support decoding of VCL NAL units within DU 553. CPB 543 is a first-in, first-out buffer for HRD 500. CPB 543 contains DUs 553 that contain video data in decoding order. CPB 543 stores video data for use during bitstream conformance verification.
[0081] The CPB 543 forwards the DU 553 to a decoding process component 545. The decoding process component 545 is a component that conforms to the VVC standard. For example, the decoding process component 545 may emulate the decoder 400 used by an end user. The decoding process component 545 decodes the DU 553 at a rate achievable by an exemplary end-user decoder. If the decoding process component 545 cannot decode the DU 553 fast enough to prevent overflow of the CPB 543, then the bitstream 551 does not conform to the standard and should be re-encoded.
[0082] The decoding process component 545 decodes the DU 553 to generate a decoded DU 555. The decoded DU 555 includes a decoded picture. The decoded DU 555 is forwarded to a DPB 547. The DPB 547 may be substantially similar to the decoded picture buffer components 223, 323, and / or 423. To support inter-prediction, pictures obtained from the decoded DU 555 and marked for use as reference pictures 556 are returned to the decoding process component 545 to support further decoding. The DPB 547 outputs the decoded video sequence as a series of pictures 557. The pictures 557 are generally reconstructed pictures that reflect the pictures coded into the bitstream 551 by the encoder.
[0083] Picture 557 is forwarded to output cropping component 549, which is configured to apply an adaptive cropping window to picture 557. This results in output cropped picture 559. Output cropped picture 559 is a perfectly reconstructed picture. Thus, output cropped picture 559 mimics what an end user would see when decoding bitstream 551. In this way, the encoder can review output cropped picture 559 to ensure a good encoding.
[0084] The HRD 500 is initialized based on HRD parameters in the bitstream 551. For example, the HRD 500 can read the HRD parameters from a VPS, SPS, and / or SEI message. The HRD 500 can then perform conformance testing operations on the bitstream 551 based on the information in such HRD parameters. As a specific example, the HRD 500 can determine one or more CPB delivery schedules from the HRD parameters. The delivery schedules specify the timing of delivery of video data to and from memory locations such as the CPB and / or DPB. Thus, the CPB delivery schedules specify the timing of delivery of AUs, DUs 553, and / or pictures to and from the CPB 543. It should be noted that the HRD 500 can use a DPB delivery schedule for the DPB 547 that is similar to the CPB delivery schedule.
[0085] Video may be encoded into different layers and / or OLSs for use by decoders with varying levels of hardware capabilities and for varying network conditions. CPB delivery schedules are selected to reflect these considerations. Thus, upper layer sub-bitstreams are designated for optimal hardware and network conditions, and therefore the upper layers may receive one or more CPB delivery schedules that use large amounts of memory in the CPB 543 and short delays for the transfer of DUs 553 to the DPB 547. Similarly, lower layer sub-bitstreams are designated for limited decoder hardware capabilities and / or poor network conditions. Thus, the lower layers may receive one or more CPB delivery schedules that use small amounts of memory in the CPB 543 and longer delays for the transfer of DUs 553 to the DPB 547. The OLSs, layers, sub-layers, or combinations thereof may then be tested according to the corresponding delivery schedules to ensure that the resulting sub-bitstreams can be correctly decoded under the conditions expected for the sub-bitstreams. Thus, the HRD parameters in the bitstream 551 may indicate the CPB delivery schedule and include sufficient data to enable the HRD 500 to determine the CPB delivery schedule and correlate the CPB delivery schedule to the corresponding OLS, layer, and / or sublayer.
[0086] 6 is a schematic diagram illustrating an exemplary multi-layer video sequence 600. The multi-layer video sequence 600 may be encoded by an encoder, such as codec system 200 and / or encoder 300, and decoded by a decoder, such as codec system 200 and / or decoder 400, according to, for example, method 100. Additionally, the multi-layer video sequence 600 may be checked for standards conformance by an HRD, such as HRD 500. The multi-layer video sequence 600 is included to illustrate an exemplary application of layers within an encoded video sequence. The multi-layer video sequence 600 is any video sequence that uses multiple layers, such as layer N 631 and layer N+1 632.
[0087] In one example, the multi-layer video sequence 600 may use inter-layer prediction 621. Inter-layer prediction 621 is applied between pictures 611, 612, 613, and 614 and pictures 615, 616, 617, and 618 of different layers. In the illustrated example, pictures 611, 612, 613, and 614 are part of layer N+1 632, and pictures 615, 616, 617, and 618 are part of layer N 631. Layers, such as layer N 631 and / or layer N+1 632, are groups of pictures that are all associated with similar values of characteristics such as similar size, quality, resolution, signal-to-noise ratio, capacity, etc. A layer may be formally defined as a set of VCL NAL units that share the same layer ID and associated non-VCL NAL units. A VCL NAL unit is a NAL unit coded to contain video data, such as a coded slice of a picture. A non-VCL NAL unit is a NAL unit that contains non-video data, such as syntax and / or parameters that support decoding video data, performing conformance checks, or other operations.
[0088] In the illustrated example, layer N+1 632 is associated with a larger image size than layer N 631. Thus, in this example, the picture size (e.g., larger height and width, and therefore more samples) of pictures 611, 612, 613, and 614 of layer N+1 632 is larger than the picture size of pictures 615, 616, 617, and 618 of layer N 631. However, such pictures may be separated between layer N+1 632 and layer N 631 by other characteristics. Although only two layers, layer N+1 632 and layer N 631, are shown, a set of pictures may be separated into any number of layers based on associated characteristics. Layer N+1 632 and layer N 631 may also be indicated by a layer ID. A layer ID is an item of data associated with a picture and indicates that the picture is part of the indicated layer. Thus, each picture 611-618 may be associated with a corresponding layer ID to indicate which layer N+1 632 or layer N 631 contains the corresponding picture. For example, the layer ID may include a NAL unit header layer ID (nuh_layer_id), which is a syntax element that specifies an identifier of a layer that contains an NAL unit (e.g., containing slices and / or parameters of a picture within a layer). A layer associated with a lower quality / smaller picture size / smaller bitstream size, such as layer N 631, is generally assigned a lower layer ID and is referred to as a lower layer. Furthermore, a layer associated with a higher quality / larger picture size / larger bitstream size, such as layer N+1 632, is generally assigned a higher layer ID and is referred to as a higher layer.
[0089] Pictures 611-618 in different layers 631-632 are configured to be alternatively displayed. As a specific example, a decoder may decode and display picture 615 at the current display time if a smaller picture is desired, or the decoder may decode and display picture 611 at the current display time if a larger picture is desired. Thus, pictures 611-614 in higher layer N+1 632 contain substantially the same image data as corresponding pictures 615-618 in lower layer N 631 (despite differences in picture size). Specifically, picture 611 contains substantially the same image data as picture 615, picture 612 contains substantially the same image data as picture 616, and so on.
[0090] Pictures 611-618 may be coded by referencing other pictures 611-618 in the same layer N 631 or N+1 632. Coding a picture with reference to another picture in the same layer results in inter-prediction 623. Inter-prediction 623 is indicated by a solid arrow. For example, picture 613 may be coded using inter-prediction 623 with one or two of pictures 611, 612, and / or 614 in layer N+1 632 as references, with one picture referenced for unidirectional inter-prediction and / or two pictures referenced for bidirectional inter-prediction. Furthermore, picture 617 may be coded using inter-prediction 623 with one or two of pictures 615, 616, and / or 618 in layer N 631 as references, with one picture referenced for unidirectional inter-prediction and / or two pictures referenced for bidirectional inter-prediction. When performing inter prediction 623, if a picture is used as a reference for another picture in the same layer, the picture may be called a reference picture. For example, picture 612 may be a reference picture used to encode picture 613 according to inter prediction 623. Inter prediction 623 may also be called intra-layer prediction in a multi-layer context. Thus, inter prediction 623 is a mechanism for encoding samples of a current picture by referencing indicated samples in a reference picture different from the current picture, where the reference picture and the current picture are in the same layer.
[0091] Pictures 611-618 can also be coded by referencing other pictures 611-618 in different layers. This process is known as inter-layer prediction 621 and is indicated by the dashed arrows. Inter-layer prediction 621 is a mechanism for coding samples of a current picture by referencing indicated samples in a reference picture where the current picture and the reference picture are in different layers and therefore have different values for nuh_layer_id. For example, a picture in a lower layer N 631 can be used as a reference picture to code a corresponding picture in an upper layer N+1 632. As a specific example, picture 611 can be coded by referencing picture 615 according to inter-layer prediction 621. In such a case, picture 615 is used as the inter-layer reference picture. An inter-layer reference picture is a reference picture used for inter-layer prediction 621. In most cases, inter-layer prediction 621 is constrained so that a current picture, such as picture 611, can only use inter-layer reference pictures that are included in the same AU 627 and that are in a lower layer, such as picture 615. When multiple layers (e.g., more than two) are available, inter-layer prediction 621 can encode / decode the current picture based on multiple inter-layer reference pictures that are in a lower level than the current picture.
[0092] A video encoder can use the multi-layer video sequence 600 to encode pictures 611-618 according to many different combinations and / or permutations of inter-prediction 623 and inter-layer prediction 621. For example, picture 615 may be encoded according to intra-prediction. Thereafter, pictures 616-618 may be encoded according to inter-prediction 623 by using picture 615 as a reference picture. Furthermore, picture 611 may be encoded according to inter-layer prediction 621 by using picture 615 as an inter-layer reference picture. Thereafter, pictures 612-614 may be encoded according to inter-prediction 623 by using picture 611 as a reference picture. Thus, reference pictures can serve as both single-layer reference pictures and inter-layer reference pictures for different encoding mechanisms. By encoding the upper layer N+1 632 pictures based on the lower layer N 631 pictures, the upper layer N+1 632 can avoid using intra prediction, which has much lower coding efficiency than inter prediction 623 and inter-layer prediction 621. In this way, the poor coding efficiency of intra prediction may be limited to pictures of the smallest / lowest quality and therefore limited to encoding a minimum amount of video data. Pictures used as reference pictures and / or inter-layer reference pictures may be indicated in entries of a reference picture list included in a reference picture list structure.
[0093] Pictures 611-618 may also be included in an access unit (AU) 627. AU 627 is a set of coded pictures included in different layers and having the same output time during decoding. Therefore, coded pictures in the same AU 627 are scheduled to be output from the DPB at the same time in the decoder. For example, pictures 614 and 618 are in the same AU 627. Pictures 613 and 617 are in a different AU 627 from pictures 614 and 618. Pictures 614 and 618 in the same AU 627 may be displayed alternatively. For example, picture 618 may be displayed when a small picture size is desired, and picture 614 may be displayed when a large picture size is desired. When a large picture size is desired, picture 614 is output, and picture 618 is used only for inter-layer prediction 621. In this case, picture 618 is discarded without being output once inter-layer prediction 621 is completed.
[0094] For the purpose of conformance testing, AUs 627 can be further subdivided into DUs 628. A DU 628 may be defined as an AU 627 or a subset of an AU 627 that includes one or more VCL NAL units and associated non-VCL NAL units within the AU 627. In other words, a DU 628 may include a single coded picture along with syntax elements to support decoding of the picture. In a single-layer bitstream, a DU 628 is an AU 627. In a multi-layer bitstream, a DU 628 is a subset of an AU 627. The distinction between AUs 627 and DUs 628 can be used when performing conformance testing in the HRD. For example, some conformance tests are configured to apply to each AU 627, and other conformance tests are configured to apply to each DU 628 within each AU 627. Conformance tests applied to one or more complete AUs 627 can be referred to as AU-level operations. Conformance tests applied to one or more DUs 628 can be referred to as DU-level operations. Thus, an AU level is a description of operations that apply to one or more complete AUs 627, and therefore to one or more complete groups of pictures that share the same output time. Also, a DU level is a description of operations that apply to one or more complete DUs 628, and therefore to one or more pictures.
[0095] 7 is a schematic diagram illustrating an exemplary bitstream 700. For example, the bitstream 700 may be generated by the codec system 200 and / or the encoder 300 for decoding by the codec system 200 and / or the decoder 400 according to the method 100. Furthermore, the bitstream 700 may include the multi-layer video sequence 600. Furthermore, the bitstream 700 may include various parameters for controlling the operation of an HRD, such as the HRD 500. Based on such parameters, the HRD may check the bitstream 700 for conformance to a standard before sending it to the decoder for decoding.
[0096] The bitstream 700 includes a VPS 711, one or more SPSs 713, multiple picture parameter sets (PPSs) 715, multiple slice headers 717, image data 720, a buffering period (BP) SEI message 716, a PT SEI message 718, and / or a DUI SEI message 719. The VPS 711 includes data related to the entire bitstream 700. For example, the VPS 711 may include data related to the OLS, layers, and / or sublayers used in the bitstream 700. The SPS 713 includes sequence data common to all pictures in a coded video sequence included in the bitstream 700. For example, each layer may include one or more coded video sequences, and each coded video sequence may reference the SPS 713 for corresponding parameters. The parameters in the SPS 713 may include picture sizing, bit depth, coding tool parameters, bit rate limits, etc. Note that while each sequence references an SPS 713, in some examples, a single SPS 713 may contain data for multiple sequences. The PPS 715 contains parameters that apply to an entire picture. Thus, each picture in a video sequence may reference a PPS 715. Note that while each picture references a PPS 715, in some examples, a single PPS 715 may contain data for multiple pictures. For example, multiple similar pictures may be coded according to similar parameters. In such cases, a single PPS 715 may contain data for such similar pictures. The PPS 715 may indicate coding tools, quantization parameters, offsets, etc. available for slices within the corresponding picture.
[0097] The slice header 717 contains parameters specific to each slice in a picture. Thus, there may be one slice header 717 for each slice 727 in a video sequence. The slice header 717 may include slice type information, filtering information, prediction weights, tile entry points, deblocking parameters, etc. Note that in some examples, the bitstream 700 may also include a picture header, which is a syntax structure that contains parameters that apply to all slices 727 in a single picture 725. For this reason, the terms picture header and slice header 717 may be used interchangeably in some contexts. For example, certain parameters may be moved between the slice header 717 and the picture header depending on whether such parameters are common to all slices 727 in the picture 725.
[0098] The image data 720 includes video data encoded according to inter-prediction and / or intra-prediction, as well as corresponding transformed and quantized residual data. For example, the image data 720 may include a layer 723, a picture 725, and / or a slice 727. A layer 723 is a set of VCL NAL units and associated non-VCL NAL units that share specified characteristics (e.g., a common resolution, frame rate, picture size, etc.) as indicated by a layer ID such as nuh_layer_id. For example, a layer 723 may include a set of pictures 725 that share the same nuh_layer_id and associated parameter sets and / or SEI messages. A layer 723 may be substantially similar to layers 631 and / or 632. nuh_layer_id is a syntax element that specifies an identifier of a layer 723 that includes at least one NAL unit. For example, the lowest quality layer 723, known as the base layer, may include the lowest value of nuh_layer_id, with the values of nuh_layer_id increasing for higher quality layers 723. Thus, lower layers are layers 723 with lower values of nuh_layer_id, and higher layers are layers 723 with higher values of nuh_layer_id. Layers 723 can also include an OLS. An OLS is a set of layers 723 where one or more layers 723 are designated as output layers. An output layer is any layer 723 designated for output and display at a decoder. Layers 723 that are not output layers can be included in an OLS to support decoding of the output layer, for example, via inter-layer prediction.
[0099] A picture 725 is an array of luma samples and / or chroma samples that generate a frame or a field thereof. For example, a picture 725 is a coded image that can be output for display or used to support coding of other pictures 725 for output. A picture 725 includes one or more slices 727. A slice 727 may be defined as an integer number of complete tiles or an integer number of contiguous complete coding tree unit (CTU) rows (e.g., within a tile) of a picture 725 that are exclusively contained in a single NAL unit. A slice 727 is further subdivided into CTUs and / or coding tree blocks (CTBs). A CTU is a group of samples of a predetermined size that can be divided in a coding tree. A CTB is a subset of a CTU and contains the luma or chroma component of the CTU. The CTUs / CTBs are further subdivided into coding blocks based on the coding tree. The coding blocks can then be coded / decoded according to a prediction mechanism.
[0100] An SEI message is a syntax structure with specified semantics that conveys information not required by the decoding process to determine values of samples in a decoded picture. For example, an SEI message may include data to support the HRD process or other support data not directly related to decoding of the bitstream 700 at the decoder. The BP SEI message 716 is an SEI message that includes HRD parameters for initializing the HRD to manage the CPB for testing the corresponding OLS and / or layer 723. The PT SEI message 718 is an SEI message that includes HRD parameters for managing the distribution information of AUs in the CPB and / or DPB for testing the corresponding OLS and / or layer 723. The DUI SEI message 719 is an SEI message that includes HRD parameters for managing the distribution information of DUs in the CPB and / or DPB for testing the corresponding OLS and / or layer 723.
[0101] Note that bitstream 700 may be encoded as a sequence of NAL units. NAL units are containers for video data and / or supporting syntax. NAL units may be VCL NAL units or non-VCL NAL units. VCL NAL units are NAL units coded to contain video data. Specifically, VCL NAL units include slices 727 and associated slice headers 717. Non-VCL NAL units are NAL units containing non-video data, such as syntax and / or parameters that support decoding the video data, performing conformance checks, or other operations. Non-VCL NAL units may include VPS NAL units, SPS NAL units, and PPS NAL units, which include VPS 711, SPS 713, and PPS 715, respectively. Non-VCL NAL units may also include SEI NAL units, which may contain BP SEI messages 716, PT SEI messages 718, and / or DUI SEI messages 719. Thus, an SEI NAL unit is a NAL unit that contains an SEI message. Note that the foregoing list of NAL units is exemplary and not exhaustive.
[0102] An HRD, such as HRD 500, can be used to check bitstream 700 for conformance to a standard. The HRD can perform conformance tests on bitstream 700 using HRD parameters. HRD parameters 735 can be stored in syntax structures within VPS 711 and / or SPS 713. HRD parameters 735 are syntax elements that initialize and / or define the operating conditions of the HRD. BP SEI message 716, PT SEI message 718, and / or DUI SEI message 719 contain parameters that further define the behavior of the HRD for particular sequences, AUs, and / or DUs based on the HRD parameters 735 within VPS 711 and / or SPS 713.
[0103] In some video coding systems, the BP SEI message 716, the PT SEI message 718, and / or the DUI SEI message 719 may contain parameters that directly reference the VPS 711. This dependency poses certain challenges. For example, the bitstream 700 may be encoded with various layers 723 and / or OLSs. When a decoder requests an OLS, the encoder, slicer, and / or intermediate storage server may send the OLS for the layers 723 to the decoder based on the decoder capabilities and / or current network conditions. Specifically, the encoder, slicer, and / or storage server uses a bitstream extraction process to remove layers 723 outside the OLS from the bitstream 700 and send the remaining layers 723 toward the decoder. This process allows many different decoders to obtain different representations of the bitstream 700 based on the decoder-side conditions. As described above, the VPS 711 contains data related to the OLS and / or layers 723. However, some OLSs include a single layer 723. When an OLS with a single layer 723 is transmitted to a decoder, the decoder may not need the data in the VPS 711 because the data in the SPS 713, PPS 715, and slice header 717 may be sufficient to decode the bitstream for the single layer 723. To avoid transmitting unnecessary data, the encoder, slicer, and / or storage server can remove the VPS 711 as part of the bitstream extraction process. This approach can be beneficial because it can increase the coding efficiency of the sub-bitstreams extracted and transmitted to the decoder. However, dependencies between the BP SEI message 716, the PT SEI message 718, and / or the DUI SEI message 719 may result in errors when the VPS 711 is removed. Specifically, omitting the VPS 711 may cause the SEI messages to return errors because the data on which the SEI messages depend within the VPS 711 is not received at the decoder when the VPS 711 is removed.Additionally, the HRD checks the bitstream for conformance by mimicking a decoder. Thus, the HRD may check an OLS with a single layer 723 without parsing the VPS 711. Therefore, the HRD of some systems may not be able to resolve the parameters of the BP SEI message 716, the PT SEI message 718, and / or the DUI SEI message 719, and therefore the HRD may not be able to check such an OLS for conformance.
[0104] The bitstream 700 is improved to correct the above problem. Specifically, parameters are added to the BP SEI message 716, the PT SEI message 718, and / or the DUI SEI message 719 to remove the dependency on the VPS 711. Thus, the BP SEI message 716, the PT SEI message 718, and / or the DUI SEI message 719 can be fully parsed and resolved by the HRD and / or decoder even when the VPS 711 is omitted due to an OLS that includes a single layer 723. The dependency can be removed by including du_hrd_params_present_flag 731 and du_cpb_params_in_pic_timing_sei_flag 733 in one or more of the SEI messages. In the illustrated example, du_hrd_params_present_flag 731 and du_cpb_params_in_pic_timing_sei_flag 733 are included in the BP SEI message 716.
[0105] du_hrd_params_present_flag 731 is a syntax element that specifies whether the HRD should operate at the AU level or the DU level. The AU level is a description of an operation that applies to one or more entire AUs (e.g., applied to one or more entire groups of pictures that share the same output time). The DU level is a description of an operation that applies to one or more entire DUs (e.g., applied to one or more pictures). In a specific example, du_hrd_params_present_flag 731 is set to 1 if a DU-level HRD parameter is present and specifies that the HRD can operate at the AU level or the DU level, and is set to 0 if a DU-level HRD parameter is absent and specifies that the HRD operates at the AU level.
[0106] When the HRD operates at the DU level (e.g., when du_hrd_params_present_flag 731 is set to 1), the HRD should refer to DU parameters such as CPB removal delay 737. CPB removal delay 737 is a syntax element that specifies the CPB removal delay of a DU, which is the time that one or more DUs (pictures 725) can remain in the CPB before being forwarded to the HRD's DPB. However, depending on the example, the associated CPB removal delay 737 may be included in the PT SEI message 718 or the DUI SEI message 719. du_cpb_params_in_pic_timing_sei_flag 733 is a syntax element that specifies whether the DU-level CPB removal delay 737 parameter is present in the PT SEI message 718 or the DUI SEI message 719. In a particular example, du_cpb_params_in_pic_timing_sei_flag 733 is set to 1 when a DU-level CPB removal delay 737 parameter is present in the PT SEI message 718 and specifies that the DUI SEI message 719 is not available. Additionally, du_cpb_params_in_pic_timing_sei_flag 733 can be set to 0 when a DU-level CPB removal delay 737 parameter is present in the DUI SEI message 719 and specifies that the PT SEI message 718 does not include a DU-level CPB removal delay 737 parameter.
[0107] By including du_hrd_params_present_flag 731 and du_cpb_params_in_pic_timing_sei_flag 733 in the SEI message, the SEI message does not depend on the VPS 711. Therefore, even if the VPS 711 is omitted from the bitstream 700, the BP SEI message 716, the PT SEI message 718, and / or the DUI SEI message 719 can be parsed. Therefore, the HRD can properly parse the BP SEI message 716, the PT SEI message 718, and / or the DUI SEI message 719 and perform a conformance test for the OLS with a single layer. Furthermore, the decoder can parse and use syntax elements in the BP SEI message 716, the PT SEI message 718, and / or the DUI SEI message 719 as needed to support the decoding process. As a result, the encoder and decoder can improve their functionality and avoid errors. Additionally, removing the dependency between the SEI message and the VPS 711 supports the removal of the VPS 711 in certain cases, which improves coding efficiency and therefore reduces processor, memory, and / or network signaling resource usage in both the encoder and decoder in such cases.
[0108] The aforementioned information will now be explained in more detail below. Layered video coding is also referred to as scalable coding or scalable video coding. Scalability in video coding can be supported by using multi-layer coding techniques. A multi-layer bitstream includes a base layer (BL) and one or more enhancement layers (EL). Examples of scalability include spatial scalability, quality / signal-to-noise ratio (SNR) scalability, multiview scalability, frame rate scalability, etc. When multi-layer coding techniques are used, a picture or a portion thereof may be coded without using a reference picture (intra-prediction), coded by referencing a reference picture in the same layer (inter-prediction), and / or coded by referencing a reference picture in another layer (inter-layer prediction). A reference picture used for inter-layer prediction of a current picture is called an inter-layer reference picture (ILRP). FIG. 6 shows an example of multi-layer coding for spatial scalability, where pictures in different layers have different resolutions.
[0109] Some video coding families provide scalability support in profiles separate from profiles for single-layer coding. Scalable Video Coding (SVC) is a scalable extension of Advanced Video Coding (AVC) that supports spatial, temporal, and quality scalability. For SVC, a flag is signaled in each macroblock (MB) in an EL picture to indicate whether the EL MB is predicted using a co-located block from a lower layer. Predictions from the co-located block may include texture, motion vectors, and / or coding mode. SVC implementations may not directly reuse unmodified AVC implementations in their design. The syntax and decoding process of SVC EL macroblocks differ from that of AVC.
[0110] Scalable HEVC (SHVC) is an extension of HEVC that supports spatial and quality scalability. Multiview HEVC (MV-HEVC) is an extension of HEVC that provides support for multiview scalability. 3D HEVC (3D-HEVC) is an extension of HEVC that provides support for more advanced and efficient 3D video coding than MV-HEVC. Temporal scalability can be included as an integral part of a single-layer HEVC codec. In multi-layer extensions of HEVC, decoded pictures used for inter-layer prediction originate only from the same AU and are treated as long-term reference pictures (LTRPs). Such pictures are assigned reference indices in a reference picture list along with other temporal reference pictures in the current layer. Inter-layer prediction (ILP) is achieved at the prediction unit level by setting the values of reference indices to refer to inter-layer reference pictures in the reference picture list. Spatial scalability involves resampling a reference picture or part of it when the ILRP has a different spatial resolution than the current picture being coded or decoded. Resampling of reference pictures can be achieved either at the picture level or at the coding block level.
[0111] VVC also supports layered video coding. A VVC bitstream may contain multiple layers. The layers may all be independent of each other. For example, each layer may be coded without inter-layer prediction. In this case, each layer is also referred to as a simulcast layer. In some cases, some of the layers are coded using ILP. A flag in the VPS may indicate whether a layer is a simulcast layer or whether some layers use ILP. If some layers use ILP, layer dependencies between layers are also signaled in the VPS. Unlike SHVC and MV-HEVC, VVC may not specify an OLS. An OLS contains a specified set of layers, and one or more layers in the set of layers are designated as output layers. An output layer is a layer in the OLS that is output. In some implementations of VVC, if a layer is a simulcast layer, only one layer may be selected for decoding and output. In some implementations of VVC, when any layer uses ILP, the entire bitstream, including all layers, is designated to be decoded. Furthermore, one of these layers is designated as an output layer. The output layer may be indicated as the top layer only, all layers, or the top layer plus a designated set of lower layers.
[0112] The above-described aspect has several problems. For example, the nuh_layer_id values of SPS, PPS, and APS NAL units may not be properly constrained. Furthermore, the TemporalId value of SEI NAL units may not be properly constrained. Also, when reference picture resampling is enabled and pictures in a CLVS have different spatial resolutions, the setting of NoOutputOfPriorPicsFlag may not be properly specified. Also, in some video coding systems, suffix SEI messages cannot be included in scalable nesting SEI messages. As another example, buffering period, picture timing, and decoding unit information SEI messages may include analysis dependencies on VPS and / or SPS.
[0113] Generally, this disclosure describes video coding improvement techniques. The techniques are based on VVC. However, these techniques also apply to layered video coding based on other video codec specifications.
[0114] One or more of the problems mentioned above may be solved as follows: The nuh_layer_id values of SPS, PPS, and APS NAL units are appropriately constrained herein. The TemporalId values of SEI NAL units are appropriately constrained herein. The setting of NoOutputOfPriorPicsFlag is appropriately specified when reference picture resampling is enabled and pictures in a CLVS have different spatial resolutions. A suffix SEI message can be included in a scalable nesting SEI message. The parsing dependency of BP, PT, and DUI SEI messages on VPS or SPS can be removed by repeating the syntax element decoding_unit_hrd_params_present_flag in the BP SEI message syntax, the syntax elements decoding_unit_hrd_params_present_flag and decoding_unit_cpb_params_in_pic_timing_sei_flag in the PT SEI message syntax, and the syntax element decoding_unit_cpb_params_in_pic_timing_sei_flag in the DUI SEI message.
[0115] An exemplary implementation of the aforementioned mechanism is as follows: Exemplary general NAL unit semantics are as follows:
[0116] nuh_temporal_id_plus1 minus 1 specifies the temporal identifier of the NAL unit. The value of nuh_temporal_id_plus1 should not be equal to 0. The variable TemporalId can be derived as follows: TemporalId=nuh_temporal_id_plus1-1 When nal_unit_type is in the range from IDR_W_RADL to RSV_IRAP_13 (inclusive), TemporalId must be equal to 0. When nal_unit_type is equal to STSA_NUT, TemporalId must not be equal to 0.
[0117] The value of TemporalId MUST be the same for all VCL NAL units of an access unit. The value of TemporalId of a coded picture, layer access unit, or access unit MAY be the value of TemporalId of the VCL NAL units of the coded picture, layer access unit, or access unit. The value of TemporalId of a sub-layer representation MAY be the maximum value of TemporalId of all VCL NAL units in the sub-layer representation.
[0118] The value of TemporalId for non-VCL NAL units is constrained as follows: If nal_unit_type is equal to DPS_NUT, VPS_NUT, or SPS_NUT, then TemporalId shall be equal to 0, and the TemporalId of the access unit that contains the NAL unit shall be equal to 0. Otherwise, if nal_unit_type is equal to EOS_NUT or EOB_NUT, then TemporalId shall be equal to 0. Otherwise, if nal_unit_type is equal to AUD_NUT, FD_NUT, PREFIX_SEI_NUT, or SUFFIX_SEI_NUT, then TemporalId shall be equal to the TemporalId of the access unit that contains the NAL unit. Otherwise, if nal_unit_type is equal to PPS_NUT or APS_NUT, then TemporalId shall be greater than or equal to the TemporalId of the access unit that contains the NAL unit. If the NAL unit is a non-VCL NAL unit, the value of TemporalId MUST be equal to the minimum of the TemporalId values of all access units to which the non-VCL NAL unit applies. If nal_unit_type is equal to PPS_NUT or APS_NUT, TemporalId MAY be greater than or equal to the TemporalId of the containing access unit. This is because all PPSs and APSs may be included at the beginning of the bitstream. Furthermore, the first coded picture has TemporalId equal to 0.
[0119] Exemplary sequence parameter set RBSP semantics are as follows: An SPS RBSP must be available to the decoding process before it can be referenced. An SPS may be included in at least one access unit with TemporalId equal to 0, or may be provided via an external mechanism. An SPS NAL unit that contains an SPS may be constrained to have a nuh_layer_id equal to the lowest nuh_layer_id value of the PPS NAL units that reference the SPS.
[0120] Exemplary picture parameter set RBSP semantics are as follows: The PPS RBSP must be available to the decoding process before it is referenced. The PPS should be included in at least one access unit with a TemporalId less than or equal to the TemporalId of the PPS NAL unit, or provided via an external mechanism. The PPS NAL unit containing the PPS RBSP should have a nuh_layer_id equal to the lowest nuh_layer_id value of the coded slice NAL units that reference the PPS.
[0121] The semantics of an exemplary adaptation parameter set are as follows: Each APS RBSP must be available to the decoding process before it is referenced. The APS should also be included in at least one access unit with a TemporalId less than or equal to the TemporalId of the coded slice NAL unit that references the APS or is provided via an external mechanism. An APS NAL unit can be shared by pictures / slices of multiple layers. The nuh_layer_id of an APS NAL unit must be equal to the lowest nuh_layer_id value of the coded slice NAL units that reference the APS NAL unit. Alternatively, an APS NAL unit may not be shared by pictures / slices of multiple layers. The nuh_layer_id of an APS NAL unit must be equal to the nuh_layer_id of the slice that references the APS.
[0122] In one example, removing a picture from the DPB before decoding the current picture is described as follows: Removing a picture from the DPB before decoding the current picture (but after parsing the slice header of the first slice of the current picture) may be done at the CPB removal time of the first decoding unit of access unit n (containing the current picture). This proceeds as follows: A decoding process for reference picture list construction is invoked, and a decoding process for reference picture marking is invoked.
[0123] If the current picture is a Coded Layer Video Sequence Start (CLVSS) picture that is not picture 0, the following ordered steps are applied: For the decoder under test, the variable NoOutputOfPriorPicsFlag may be derived as follows: If the values of pic_width_in_luma_samples, pic_height_max_in_luma_samples, chroma_format_idc, separate_colour_plane_flag, bit_depth_luma_minus8, bit_depth_chroma_minus8, or sps_max_dec_pic_buffering_minus1[Htid] derived from the SPS differ from the values of pic_width_in_luma_samples, pic_height_in_luma_samples, chroma_format_idc, separate_colour_plane_flag, bit_depth_luma_minus8, bit_depth_chroma_minus8, or sps_max_dec_pic_buffering_minus1[Htid], respectively, derived from the SPS referenced by the previous picture, then NoOutputOfPriorPicsFlag may be set to 1 by the decoder under test regardless of the value of no_output_of_prior_pics_flag. Under these conditions, it may be preferable to set NoOutputOfPriorPicsFlag equal to no_output_of_prior_pics_flag, provided that the decoder under test is capable of setting NoOutputOfPriorPicsFlag to 1. Otherwise, NoOutputOfPriorPicsFlag may be set equal to no_output_of_prior_pics_flag.
[0124] The value of NoOutputOfPriorPicsFlag obtained for the decoder under test is applied to the HRD. If the resulting value of NoOutputOfPriorPicsFlag is 1, all picture storage buffers in the DPB are emptied without outputting the pictures they contain, and the DPB fullness is set to zero. If both of the following conditions apply to any picture k in the DPB, all such pictures k in the DPB are removed from the DPB: Picture k is marked as unused for reference, and picture k has a PictureOutputFlag equal to 0, or the corresponding DPB output time is less than or equal to the CPB removal time of the first decoding unit (denoted as decoding unit m) of the current picture n. This may occur if DpbOutputTime[k] is less than or equal to DuCpbRemovalTime[m]. For each picture removed from the DPB, the DPB fullness is decremented by 1.
[0125] An exemplary output or removal of a picture from the DPB is as follows: Outputting and removing a picture from the DPB before decoding the current picture (but after parsing the slice header of the first slice of the current picture) may occur when the first decoding unit of the access unit containing the current picture is removed from the CPB and proceeds as follows: The decoding process for reference picture list construction and the decoding process for reference picture marking are invoked.
[0126] If the current picture is a CLVSS picture that is not picture 0, the following ordered steps may be applied: For the decoder under test, the variable NoOutputOfPriorPicsFlag may be derived as follows: If the values of pic_width_in_luma_samples, pic_height_max_in_luma_samples, chroma_format_idc, separate_colour_plane_flag, bit_depth_luma_minus8, bit_depth_chroma_minus8 or sps_max_dec_pic_buffering_minus1[Htid] derived from the SPS are different from the values of pic_width_in_luma_samples, pic_height_in_luma_samples, chroma_format_idc, separate_colour_plane_flag, bit_depth_luma_minus8, bit_depth_chroma_minus8 or sps_max_dec_pic_buffering_minus1[Htid] respectively derived from the SPS referenced by the previous picture, NoOutputOfPriorPicsFlag may be set to 1 by the decoder under test regardless of the value of no_output_of_prior_pics_flag. Under these conditions, it is preferable to set NoOutputOfPriorPicsFlag equal to no_output_of_prior_pics_flag, in which case the decoder under test may set NoOutputOfPriorPicsFlag to 1. Otherwise, NoOutputOfPriorPicsFlag may be set equal to no_output_of_prior_pics_flag.
[0127] The value of NoOutputOfPriorPicsFlag derived for the decoder under test may be applied to the HRD as follows: If the value of NoOutputOfPriorPicsFlag is 1, all picture storage buffers in the DPB may be emptied without outputting the pictures they contain, and DPB fullness may be set equal to 0. Otherwise (when the value of NoOutputOfPriorPicsFlag is equal to 0), all picture storage buffers that do not need to be output and contain pictures marked as not to be used for reference may be emptied (without outputting), and by repeatedly invoking the bumping process, all non-empty picture storage buffers in the DPB may be emptied and DPB fullness may be set equal to 0.
[0128] Otherwise (the current picture is not a CLVSS picture), all picture storage buffers containing pictures marked as not needed for output and unused for reference are emptied (without output). For each picture storage buffer emptied, DPB fullness is decremented by 1. When one or more of the following conditions are true, the bumping process is repeatedly invoked, further decrementing DPB fullness by 1 for each additional picture storage buffer emptied until none of the following conditions are true: A condition is that the number of pictures in the DPB marked as needed for output is greater than sps_max_num_reorder_pics[Htid]. Another condition is that sps_max_latency_increase_plus1[Htid] is not equal to 0 and there is at least one picture in the DPB marked as needed for output whose associated variable PicLatencyCount is greater than or equal to SpsMaxLatencyPictures[Htid]. Another condition is that the number of pictures in the DPB is equal to or greater than SubDpbSize[Htid].
[0129] An exemplary general SEI message syntax is as follows:
[0130] [Table 1]
[0131] An exemplary scalable nesting SEI message syntax is as follows:
[0132] [Table 2]
[0133] Exemplary scalable nesting SEI message semantics are as follows: A scalable nesting SEI message provides a mechanism for associating an SEI message with a particular layer in the context of a particular OLS or with a particular layer outside the context of an OLS. A scalable nesting SEI message contains one or more SEI messages. An SEI message included in a scalable nesting SEI message is also referred to as a scalably nested SEI message. Bitstream conformance may require that the following restrictions apply when an SEI message is included in a scalable nesting SEI message:
[0134] An SEI message with payloadType equal to 132 (decoded picture hash) or 133 (scalable nesting) should not be included in a scalable nesting SEI message. If a scalable nesting SEI message contains a buffering period, picture timing, or decoding unit information SEI message, the scalable nesting SEI message should not contain any other SEI messages with payloadType not equal to 0 (buffering period), 1 (picture timing), or 130 (decoding unit information).
[0135] Bitstream conformance may also require that the following restrictions apply to the value of nal_unit_type of SEI NAL units that contain scalable nesting SEI messages: If the scalable nesting SEI message contains an SEI message with payloadType equal to 0 (buffering duration), 1 (picture timing), 130 (decoded unit information), 145 (dependent RAP indication), or 168 (frame field information), the SEI NAL unit that contains the scalable nesting SEI message SHOULD have nal_unit_type equal to PREFIX_SEI_NUT. If the scalable nesting SEI message contains an SEI message with payloadType (decoded picture hash) equal to 132, the SEI NAL unit that contains the scalable nesting SEI message SHOULD have nal_unit_type set equal to SUFFIX_SEI_NUT.
[0136] nesting_ols_flag may be set equal to 1 to specify that the scalably nested SEI message applies to a particular layer in the context of a particular OLS. nesting_ols_flag may be set equal to 0 to specify that the scalably nested SEI message applies to a particular layer in general (e.g., not within the context of an OLS).
[0137] Bitstream conformance may require that the following restrictions apply to the value of nesting_ols_flag: When a scalable nesting SEI message contains an SEI message with payloadType equal to 0 (buffering duration), 1 (picture timing), or 130 (decoding unit information), the value of nesting_ols_flag must be equal to 1. When a scalable nesting SEI message contains an SEI message with payloadType equal to a value in VclAssociatedSeiList, the value of nesting_ols_flag must be equal to 0.
[0138] nesting_num_olss_minus1 plus 1 specifies the number of OLSs to which the scalably nested SEI message applies. The value of nesting_num_olss_minus1 must be in the range from 0 to TotalNumOlss-1, inclusive. nesting_ols_idx_delta_minus1[i] is used to derive the variable NestingOlsIdx[i], which specifies the OLS index of the ith OLS to which the scalably nested SEI message applies, when nesting_ols_flag is equal to 1. The value of nesting_ols_idx_delta_minus1[i] must be in the range from 0 to TotalNumOlss-2, inclusive. The variable NestingOlsIdx[i] can be derived as follows: if(i==0) NestingOlsIdx[ i ]=nesting_ols_idx_delta_minus1[ i ] else NestingOlsIdx[ i ]=NestingOlsIdx[ i-1 ]+nesting_ols_idx_delta_minus1[ i ]+1
[0139] nesting_num_ols_layers_minus1[i] plus 1 specifies the number of layers to which the scalably nested SEI message applies in the context of the NestingOlsIdx[i]th OLS. The value of nesting_num_ols_layers_minus1[i] must be in the range from 0 to NumLayersInOls[NestingOlsIdx[i]]-1, inclusive.
[0140] nesting_ols_layer_idx_delta_minus1[i][j] is used to derive the variable NestingOlsLayerIdx[i][j] that specifies the OLS layer index of the jth layer to which the scalably nested SEI message applies, in the context of the NestingOlsIdx[i]th OLS, when nesting_ols_flag is equal to 1. The value of nesting_ols_layer_idx_delta_minus1[i] must be in the range from 0 to NumLayersInOls[nestingOlsIdx[i]]-2, inclusive.
[0141] The variable NestingOlsLayerIdx[i][j] can be derived as follows: if(j==0) NestingOlsLayerIdx[ i ][ j ]=nesting_ols_layer_idx_delta_minus1[ i ][ j ] else NestingOlsLayerIdx[ i ][ j ]=NestingOlsLayerIdx[ i ][ j-1 ]+ nesting_ols_layer_idx_delta_minus1[ i ][ j ]+1
[0142] The lowest of all values of LayerIdInOls[NestingOlsIdx[i]][NestingOlsLayerIdx[i][0]] for i in the range from 0 to nesting_num_olss_minus1 (inclusive) MUST be equal to the nuh_layer_id of the current SEI NAL unit (e.g., the SEI NAL unit that contains the scalable nesting SEI message). nesting_all_layers_flag may be set equal to 1 to specify that the scalably nested SEI message generally applies to all layers with a nuh_layer_id greater than or equal to the nuh_layer_id of the current SEI NAL unit. nesting_all_layers_flag may be set equal to 0 to specify that the scalably nested SEI message generally may or may not apply to all layers with a nuh_layer_id greater than or equal to the nuh_layer_id of the current SEI NAL unit.
[0143] nesting_num_layers_minus1 plus 1 specifies the number of layers to which a scalably nested SEI message generally applies. If nuh_layer_id is the nuh_layer_id of the current SEI NAL unit, the value of nesting_num_layers_minus1 MUST be in the range from 0 to vps_max_layers_minus1-GeneralLayerIdx[nuh_layer_id], inclusive. nesting_layer_id[i] specifies the nuh_layer_id value of the ith layer to which a scalably nested SEI message generally applies when nesting_all_layers_flag is 0. The value of nesting_layer_id[i] MUST be greater than nuh_layer_id, where nuh_layer_id is the nuh_layer_id of the current SEI NAL unit.
[0144] When nesting_ols_flag is equal to 1, the variable NestingNumLayers, which specifies the number of layers to which a scalably nested SEI message generally applies, and the list NestingLayerId[i], for i in the range 0 to NestingNumLayers-1 (inclusive), which specifies a list of nuh_layer_id values of layers to which a scalably nested SEI message generally applies, are derived as follows, where nuh_layer_id is the nuh_layer_id of the current SEI NAL unit: if(nesting_all_layers_flag){ NestingNumLayers= vps_max_layers_minus1+1-GeneralLayerIdx[ nuh_layer_id ] for(i=0;i < NestingNumLayers;i++) NestingLayerId[ i ]=vps_layer_id[ GeneralLayerIdx[ nuh_layer_id ]+i ](D-2) } else { NestingNumLayers=nesting_num_layers_minus1+1 for(i=0;i < NestingNumLayers;i++) NestingLayerId[ i ]=(i==0)? nuh_layer_id:nesting_layer_id[ i ] }
[0145] nesting_num_seis_minus1 plus 1 specifies the number of scalably nested SEI messages. The value of nesting_num_seis_minus1 must be in the range 0 to 63 (inclusive). nesting_0_bit should be set equal to 0.
[0146] FIG. 8 is a schematic diagram of an exemplary video encoding device 800. The video encoding device 800 is suitable for implementing the disclosed examples / embodiments described herein. The video encoding device 800 includes a downstream port 820, an upstream port 850, and / or a transceiver unit (Tx / Rx) 810 including a transmitter and / or a receiver for communicating data upstream and / or downstream over a network. The video encoding device 800 also includes a processor 830 including a logic unit and / or central processing unit (CPU) for processing data and a memory 832 for storing data. The video encoding device 800 may also include electrical, optical-electrical (OE) components, electrical-optical (EO) components, and / or wireless communication components coupled to the upstream port 850 and / or the downstream port 820 for communication of data over an electrical, optical, or wireless communication network. The video encoding device 800 may also include input and / or output (I / O) devices 860 for communicating data to and from a user. The I / O devices 860 may include output devices such as a display for displaying video data, speakers for outputting audio data, etc. The I / O devices 860 may also include input devices such as a keyboard, mouse, trackball, etc., and / or corresponding interfaces for interacting with such output devices.
[0147] The processor 830 is implemented by hardware and software. The processor 830 may be implemented as one or more CPU chips, cores (e.g., as a multi-core processor), field programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), and digital signal processors (DSPs). The processor 830 communicates with the downstream port 820, the Tx / Rx 810, the upstream port 850, and the memory 832. The processor 830 includes an encoding module 814. The encoding module 814 implements the disclosed embodiments described herein, such as the methods 100, 900, and 1000, which may use the multi-layer video sequence 600 and / or the bitstream 700. The encoding module 814 may also implement any other method / mechanism described herein. Additionally, the encoding module 814 may implement the codec system 200, the encoder 300, the decoder 400, and / or the HRD 500. For example, the encoding module 814 may be used to signal and / or read various parameters as described herein. Furthermore, the encoding module may be used to encode and / or decode video sequences based on such parameters. Accordingly, the signaling modifications described herein can increase efficiency and / or avoid errors in the encoding module 814. Accordingly, the encoding module 814 may be configured to implement mechanisms to address one or more of the above-mentioned problems. Thus, the encoding module 814 causes the video encoding device 800 to provide additional functionality and / or coding efficiency when encoding video data. Thus, the encoding module 814 improves the functionality of the video encoding device 800 and addresses problems specific to video encoding techniques. Furthermore, the encoding module 814 transforms the video encoding device 800 into a different state. Alternatively, the encoding module 814 may be implemented as instructions stored in the memory 832 and executed by the processor 830 (e.g., as a computer program product stored on a non-transitory medium).
[0148] Memory 832 includes one or more memory types such as a disk, tape drive, solid state drive, read-only memory (ROM), random access memory (RAM), flash memory, ternary content addressable memory (TCAM), static random access memory (SRAM), etc. Memory 832 may be used as an overflow data storage device to store programs when such programs are selected for execution and to store instructions and data retrieved during program execution.
[0149] 9 is a flowchart of an example method 900 for encoding a video sequence into a bitstream, such as bitstream 700, by using SEI messages that may not rely on a VPS. Method 900 may be used by an encoder, such as codec system 200, encoder 300, and / or video encoding device 800, when performing method 100. Furthermore, method 900 may operate on HRD 500 and, therefore, may perform conformance testing on multi-layer video sequence 600.
[0150] Method 900 may start when an encoder receives a video sequence and determines, for example, based on user input, to encode the video sequence into a multi-layer bitstream. In step 901, the encoder encodes multiple coded pictures into a bitstream. For example, the coded pictures may be organized into layers to create a multi-layer bitstream. Furthermore, each coded picture may be included in a DU. The coded pictures may also be included in an AU, which includes a set of pictures from different layers with the same output time. A layer may include a set of VCL NAL units with the same layer ID and associated non-VCL NAL units. For example, a set of VCL NAL units is part of a layer if the set of VCL NAL units all have a particular value of nuh_layer_id. A layer may include a set of VCL NAL units, each of which includes a slice of a coded picture. A layer may also include any parameter sets used to encode such pictures, which are included in non-VCL NAL units. These layers may be included in one or more OLSs. One or more layers may be output layers (e.g., each OLS includes at least one output layer). Layers that are not output layers are coded to support the reconstruction of the output layers, but such support layers are not intended for output at a decoder. In this way, the encoder can code various combinations of layers for transmission to the decoder upon request. Layers can be transmitted as needed to allow the decoder to obtain different representations of the video sequence depending on network conditions, hardware capabilities, and / or user preferences.
[0151] In step 903, the encoder may encode the current SEI message into the bitstream. The current SEI message may be a BP SEI message, a PT SEI message, or a DUI SEI message, depending on the example. The current SEI message includes du_hrd_params_present_flag, which specifies whether DU-level HRD parameters are present in the bitstream. du_hrd_params_present_flag may further specify whether the HRD operates at the AU level or the DU level. The AU level indicates that the HRD process is applied to the entire AU, and the DU level indicates that the HRD process is applied to each individual DU. Thus, du_hrd_params_present_flag may specify the granularity of the conformance test (e.g., AU granularity or DU granularity). As a specific example, du_hrd_params_present_flag may be set to 1 if DU-level HRD parameters are present and specify that the HRD can operate at the AU level or the DU level. Also, du_hrd_params_present_flag can be set to 0 to specify that there are no DU level HRD parameters and the HRD operates at the AU level.
[0152] The current SEI message may also include du_cpb_params_in_pic_timing_sei_flag, which specifies whether DU-level CPB removal delay parameters are present in the PT SEI message. du_cpb_params_in_pic_timing_sei_flag may further specify whether DU-level CPB removal delay parameters are present in the DUI SEI message. In a particular example, when DU-level CPB removal delay parameters are present in the PT SEI message and the DUI SEI message specifies that they are not available, du_cpb_params_in_pic_timing_sei_flag may be set to 1. Furthermore, when DU-level CPB removal delay parameters are present in the DUI SEI message and the PT SEI message specifies that the DU-level CPB removal delay parameters are not included, du_cpb_params_in_pic_timing_sei_flag may be set to 0. The foregoing constraints and / or requirements ensure that the bitstream complies with, for example, VVC or any other standard as modified as set forth herein. However, the encoder may also be able to operate in other less constrained modes, such as when operating under different standards or different versions of the same standard.
[0153] In step 905, the HRD can perform a set of bitstream conformance tests on the bitstream based on the current SEI message. For example, the HRD can read the du_hrd_params_present_flag to determine whether to test the bitstream at the AU level, or whether a DU parameter is present that allows testing at the DU level as well. Furthermore, the HRD can read the du_cpb_params_in_pic_timing_sei_flag to determine whether the DU parameter, if present, is found in the DUI SEI message or the PT SEI message. The HRD can then obtain the desired parameters from the indicated SEI message and perform conformance tests based on those parameters. Based on the aforementioned flags, the current SEI message is VPS-independent. In this way, the current SEI message can be fully analyzed and resolved even if the VPS is unavailable. Therefore, the conformance tests work properly even when tests are performed on an OLS with a single layer, and therefore the VPS is not available to the HRD because it is not configured to transmit as part of the OLS.
[0154] In step 907, the encoder can store the bitstream for communication to the decoder upon request. The encoder can also transmit the bitstream to the decoder as needed.
[0155] 10 is a flowchart of an example method 1000 of decoding a video sequence from a bitstream, such as bitstream 700, that applies SEI messages that may not rely on a VPS. Method 1000 may be used by a decoder, such as codec system 200, decoder 400, and / or video encoding device 800, when performing method 100. Additionally, method 1000 may be used on a multi-layer video sequence 600 that has been checked for conformance by an HRD, such as HRD 500.
[0156] Method 1000 may begin when a decoder begins receiving a bitstream of coded data representing a multi-layer video sequence, e.g., as a result of method 900 and / or in response to a request by the decoder. In step 1001, the decoder receives a bitstream including coded pictures in one or more VCL NAL units. Furthermore, the bitstream may include one or more layers including coded pictures. Furthermore, each coded picture may be included in a DU. A coded picture may also be included in an AU, which includes a set of pictures from different layers having the same output time. A layer may include a set of VCL NAL units with the same layer ID and associated non-VCL NAL units. For example, a set of VCL NAL units is part of a layer if the set of VCL NAL units all have a particular value of nuh_layer_id. A layer may include a set of VCL NAL units, each of which includes a slice of a coded picture. A layer may also include any parameter set used to code such pictures, such parameters being included in the non-VCL NAL units. Layers may be included in an OLS. One or more layers may be output layers. Layers that are not output layers are coded to support the reconstruction of output layers, but such support layers are not intended for output. In this way, a decoder can obtain different representations of a video sequence depending on network conditions, hardware capabilities, and / or user preferences.
[0157] The bitstream also includes a current SEI message. The current SEI message may be a BP SEI message, a PT SEI message, or a DUI SEI message, depending on the example. The current SEI message includes du_hrd_params_present_flag, which specifies whether DU-level HRD parameters are present in the bitstream. du_hrd_params_present_flag can further specify whether the HRD operates at the AU level or the DU level. The AU level indicates that the HRD process applies to the entire AU, and the DU level indicates that the HRD process applies to individual DUs. Therefore, du_hrd_params_present_flag can specify the granularity of the conformance test (e.g., AU granularity or DU granularity). As a specific example, du_hrd_params_present_flag can be set to 1 if DU-level HRD parameters are present and specify that the HRD can operate at the AU level or the DU level. Also, du_hrd_params_present_flag can be set to 0 to specify that there are no DU level HRD parameters and the HRD operates at the AU level.
[0158] The current SEI message may also include du_cpb_params_in_pic_timing_sei_flag, which specifies whether DU-level CPB removal delay parameters are present in the PT SEI message. du_cpb_params_in_pic_timing_sei_flag may further specify whether DU-level CPB removal delay parameters are present in the DUI SEI message. In a particular example, du_cpb_params_in_pic_timing_sei_flag may be set to 1 when DU-level CPB removal delay parameters are present in the PT SEI message and the DUI SEI message specifies that they are not available. Furthermore, du_cpb_params_in_pic_timing_sei_flag may be set to 0 when DU-level CPB removal delay parameters are present in the DUI SEI message and the PT SEI message specifies that the DU-level CPB removal delay parameters are not included.
[0159] In one embodiment, a video decoder expects du_hrd_params_present_flag and du_cpb_params_in_pic_timing_sei_flag to indicate the presence and / or location of DU level parameters as described above under VVC or some other standard. However, if the decoder determines that this condition is not true, the decoder may detect an error, signal the error, request that a corrected bitstream (or part thereof) be retransmitted, or take some other corrective measures to ensure that a conforming bitstream is received.
[0160] In step 1003, the decoder may decode the coded picture from the VCL NAL unit to generate a decoded picture. For example, the decoder may determine that the bitstream has been checked for conformance with the standard based on the presence of the current SEI message. Thus, the decoder may determine that the bitstream is decodable based on the presence of the current SEI message. Note that in some cases, the OLS received at the decoder may include a single layer. In such cases, the layer / OLS may not include a VPS. Due to the presence of du_hrd_params_present_flag and du_cpb_params_in_pic_timing_sei_flag, the current SEI message does not depend on the VPS. Thus, the absence of a VPS does not cause an error when the current SEI message is parsed. In step 1005, the decoder may forward the decoded picture for display as part of the decoded video sequence.
[0161] 11 is a schematic diagram of an example system 1100 for encoding a video sequence using a bitstream that employs VPS-independent SEI messages. System 1100 may be implemented by an encoder and decoder, such as codec system 200, encoder 300, decoder 400, and / or video encoding device 800. Furthermore, system 1100 may use HRD 500 to perform conformance testing on multi-layer video sequence 600 and / or bitstream 700. Furthermore, system 1100 may be used when implementing methods 100, 900, and / or 1000.
[0162] The system 1100 includes a video encoder 1102. The video encoder 1102 comprises an encoding module 1103 for encoding an encoded picture into a bitstream. Further, the encoding module 1103 is for encoding into the bitstream a current SEI message including a du_hrd_params_present_flag that specifies whether DU-level HRD parameters are present in the bitstream. The video encoder 1102 further comprises an HRD module 1105 for performing a set of bitstream conformance tests on the bitstream based on the current SEI message. The video encoder 1102 further comprises a storage module 1106 for storing the bitstream for communication to a decoder. The video encoder 1102 further comprises a transmission module 1107 for transmitting the bitstream toward a video decoder 1110. The video encoder 1102 may be further configured to perform any of the steps of the method 900.
[0163] The system 1100 also includes a video decoder 1110. The video decoder 1110 comprises a receiving module 1111 for receiving a bitstream including coded pictures and a current SEI message including a du_hrd_params_present_flag that specifies whether DU-level HRD parameters are present in the bitstream. The video decoder 1110 further comprises a decoding module 1113 for decoding the coded pictures to generate decoded pictures. The video decoder 1110 further comprises a forwarding module 1115 for forwarding the decoded pictures for display as part of a decoded video sequence. The video decoder 1110 may be further configured to perform any of the steps of the method 1000.
[0164] A first component is directly coupled to a second component when there are no intervening components other than wires, traces, or other intermediaries between the first and second components. A first component is indirectly coupled to a second component when there are intervening components other than wires, traces, or other intermediaries between the first and second components. The term "coupled" and variations thereof include both direct and indirect coupling. The use of the term "about," unless otherwise stated, means a range that includes ±10% of the subsequent numerical value.
[0165] It should also be understood that the steps of the exemplary methods described herein do not necessarily have to be performed in the order described, and the order of steps in such methods should be understood as merely exemplary. Similarly, such methods may include additional steps, and some steps may be omitted or combined, in methods consistent with various embodiments of the present disclosure.
[0166] While several embodiments are provided in this disclosure, it will be understood that the disclosed systems and methods may be embodied in many other specific forms without departing from the spirit or scope of the disclosure. The examples of the disclosure should be considered illustrative rather than limiting, and the intention should not be limited to the details given herein. For example, various elements or components may be combined or integrated in another system, or certain features may be omitted or not implemented.
[0167] Additionally, techniques, systems, subsystems, and methods described and illustrated as separate or distinct in various embodiments may be combined or integrated with other systems, components, techniques, or methods without departing from the scope of the present disclosure. Other examples of changes, substitutions, and alterations will be apparent to those skilled in the art and may be made without departing from the spirit and scope disclosed herein. [Explanation of symbols]
[0168] 100 How it works 200 Codec System 201 split video signal 211 General Coder Control Component 213 Transform Scaling Quantization Component 215 In-Picture Estimation Component 217 Intra-Picture Prediction Component 219 Motion Compensation Component 221 Motion Estimation Component 223 Decoded Picture Buffer Component 225 In-Loop Filter Components 227 Filter Control Analysis Component 229 Scaling Inverse Transformation Component 231 Header Formatting Context-Adaptive Binary Arithmetic Coding (CABAC) Component 300 Video Encoder 301 Split Video Signal 313 Transform Quantization Component 317 Intra-Picture Prediction Component 321 Motion Compensation Component 323 Decoded Picture Buffer Component 325 In-Loop Filter Components 329 Inverse Transform Quantization Component 331 Entropy Coding Components 400 Video Decoder 417 Intra-Picture Prediction Component 421 Motion Compensation Component 423 Decoded Picture Buffer Component 425 In-Loop Filter Components 429 Inverse Transform Quantization Component 433 Entropy Decoding Component 500 HRD 541 Virtual Stream Scheduler (HSS) 543 CPB 545 Decryption Process Components 547 DPB 549 Output Cropping Component 551 bitstream 553 Decoding Unit (DU) 555 Decrypted DUs 556 Reference Pictures 557 Pictures 559 Output Cropped Picture 600 multi-layer video sequences 611 Pictures 612 Pictures 613 Pictures 614 Pictures 615 Pictures 616 Pictures 617 Pictures 618 Pictures 621 Inter-layer Prediction 623 Inter Prediction 627 Access Units (AU) 628 DU 631 Layer N 632 Layer N+1 700 bitstream 711 VPS 713 SPS 715 Picture Parameter Set (PPS) 716 Buffering Period (BP) SEI Message 717 Slice Header 718 PT SEI Message 719 DUI SEI Message 720 image data 723 Layer 725 Pictures 727 slices 731 du_hrd_params_present_flag 733 du_cpb_params_in_pic_timing_sei_flag 737 CPB removal delay 800 Video Encoder 810 Transceiver Unit (Tx / Rx) 814 Encoding Module 820 downstream ports 830 processor 832 memory 850 upstream ports 860 Input and / or Output (I / O) Devices 1100 System 1102 Video Encoder 1103 Encoding Module 1105 HRD module 1106 Storage Module 1107 Transmitting Module 1110 Video Decoder 1111 Receiver Module 1113 Decryption Module 1115 Transfer Module
Claims
1. A method implemented by a decoder, comprising: receiving, by a receiver of the decoder, a bitstream including a coded picture and a current supplemental enhancement information (SEI) message including a decoding unit (DU) hypothetical reference decoder (HRD) parameters present flag (du_hrd_params_present_flag) specifying whether DU-level HRD parameters are present in the bitstream; decoding, by a processor of the decoder, the coded picture to generate a decoded picture; Including, the SEI message is a buffering period (BP) SEI message, the coded picture is included in one or more video coding layer (VCL) network abstraction layer (NAL) units, the BP SEI message is included in an SEI NAL unit, the SEI NAL unit is a non-VCL NAL unit; the BP SEI message is included in a scalable nesting SEI message, the BP SEI message has payloadType equal to 0, and the SEI NAL unit containing the scalable nesting SEI message has nal_unit_type equal to PREFIX_SEI_NUT; The method of claim 1, wherein the TemporalId of the non-VCL NAL unit when the nal_unit_type is equal to the PREFIX_SEI_NUT is equal to the TemporalId of an access unit that contains the non-VCL NAL unit.
2. The method of claim 1 , wherein the du_hrd_params_present_flag further specifies whether HRD operates at an access unit (AU) level or a DU level.
3. 3. The method of claim 2, wherein the du_hrd_params_present_flag is set to 1 when the DU level HRD parameter is present and specifies that the HRD can operate at the AU level or the DU level, and the du_hrd_params_present_flag is set to 0 when the DU level HRD parameter is not present and specifies that the HRD operates at the AU level.
4. 4. The method of claim 1, wherein the current SEI message further includes a DU coded picture buffer (CPB) parameters in picture timing (PT) SEI flag (du_cpb_params_in_pic_timing_sei_flag) that specifies whether DU-level CPB removal delay parameters are present in a PT SEI message.
5. The method of claim 4 , wherein the du_cpb_params_in_pic_timing_sei_flag further specifies whether the DU-level CPB removal delay parameters are present in a Decoding Unit Information (DUI) SEI message.
6. 6. The method of claim 5, wherein the du_cpb_params_in_pic_timing_sei_flag is set to 1 when the DU-level CPB removal delay parameter is present in the PT SEI message and specifies that the DUI SEI message is not available, and the du_cpb_params_in_pic_timing_sei_flag is set to 0 when the DU-level CPB removal delay parameter is present in the DUI SEI message and specifies that the PT SEI message does not include the DU-level CPB removal delay parameter.
7. 7. A video coding apparatus comprising a processor, a receiver coupled to the processor, a memory coupled to the processor, and a transmitter coupled to the processor, wherein the processor, receiver, memory, and transmitter are configured to perform the method of any one of claims 1 to 6.
8. 10. A non-transitory computer-readable medium including a computer program product for use by a video coding device, the computer program product including computer-executable instructions stored on the non-transitory computer-readable medium, the computer-executable instructions, when executed by a processor, causing the video coding device to perform the method of any one of claims 1 to 6.
9. receiving means for receiving a bitstream including a coded picture and a current Supplemental Enhancement Information (SEI) message including a Decoding Unit (DU) Hypothetical Reference Decoder (HRD) Parameters Present Flag (du_hrd_params_present_flag) specifying whether DU-level HRD parameters are present in the bitstream; decoding means for decoding the coded picture to generate a decoded picture; transfer means for transferring the decoded pictures for display as part of a decoded video sequence; Equipped with the SEI message is a buffering period (BP) SEI message, the coded picture is included in one or more video coding layer (VCL) network abstraction layer (NAL) units, the BP SEI message is included in an SEI NAL unit, the SEI NAL unit is a non-VCL NAL unit; the BP SEI message is included in a scalable nesting SEI message, the BP SEI message has payloadType equal to 0, and the SEI NAL unit containing the scalable nesting SEI message has nal_unit_type equal to PREFIX_SEI_NUT; a decoder, wherein the TemporalId of the non-VCL NAL unit when the nal_unit_type is equal to the PREFIX_SEI_NUT is equal to the TemporalId of an access unit that contains the non-VCL NAL unit.
10. The decoder of claim 9 , wherein the du_hrd_params_present_flag further specifies whether HRD operates at an access unit (AU) level or a DU level.
11. 11. The decoder of claim 10, wherein the du_hrd_params_present_flag is set to 1 when the DU level HRD parameter is present and specifies that the HRD can operate at the AU level or the DU level, and the du_hrd_params_present_flag is set to 0 when the DU level HRD parameter is not present and specifies that the HRD operates at the AU level.
12. 12. The decoder of claim 9, wherein the current SEI message further includes a DU coded picture buffer (CPB) parameters in picture timing (PT) SEI flag (du_cpb_params_in_pic_timing_sei_flag) that specifies whether DU-level CPB removal delay parameters are present in a PT SEI message.
13. The decoder of claim 12 , wherein the du_cpb_params_in_pic_timing_sei_flag further specifies whether the DU-level CPB removal delay parameters are present in a Decoding Unit Information (DUI) SEI message.
14. 14. The decoder of claim 13, wherein the du_cpb_params_in_pic_timing_sei_flag is set to 1 when the DU-level CPB removal delay parameter is present in the PT SEI message specifying that the DUI SEI message is not available, and the du_cpb_params_in_pic_timing_sei_flag is set to 0 when the DU-level CPB removal delay parameter is present in the DUI SEI message specifying that the PT SEI message does not include the DU-level CPB removal delay parameter.
Citation Information
Patent Citations
Low-latency video encoding system and its operation method
JP2015526970A
Instructions and activation of parameter sets for video coding.
JP2015529437A
Buffering period and recovery point supplemental enhancement information message
JP2015529438A
Virtual reference decoder parameter syntax structure
JP2015532551A
Signaling indications and restrictions
JP2016531451A