Simulcast layer for multi-view in video coding
The mechanism addresses the issue of improper decoding of multi-view videos by setting specific flags in the bitstream to ensure correct decoding and efficient resource utilization in simulcast scenarios, enhancing video coding systems.
Patent Information
- Application Number
- JP2025061282
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-09-24
- Filing Date
- 2025-04-02
- Publication Date
- 2025-07-03
- Estimated Expiration
- 2040-09-08
AI Technical Summary
Existing video coding systems fail to properly decode multi-view videos when all layers are simulcast and inter-layer prediction is not used, leading to errors and improper rendering.
A mechanism is introduced where the 'vps_all_independent_layers_flag' is set to 1 when all layers are simulcast, and 'each_layer_is_an_ols_flag' is signaled to specify whether each OLS contains a single layer or multiple layers, with 'ols_mode_idc' set to 2 to explicitly signal the number of OLSs and associated layers, enabling correct decoding of multi-view videos.
This approach supports coding efficiency while correcting errors, reducing bitstream size, and minimizing processor, memory, and network resource utilization in both encoders and decoders.
Smart Images

Figure 2025100571000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to video coding, and more particularly to configuring an output layer set (OLS) in a multi-layer bitstream for use in a multi-view application.
Background Art
[0002] The video data required to depict even relatively short videos can be quite large and can pose difficulties when the data is streamed or otherwise communicated over a communication network having a limited bandwidth capacity. Thus, video data is generally compressed before being communicated over modern telecommunications networks. The size of the video can also be a problem when the video is stored on a storage device, as memory resources may be limited. Video compression devices often use software and / or hardware at the source to encode the video data before transmission or storage, thereby reducing the amount of data required to represent the digital video image. The compressed data is then received at the destination by a video decompression device that decodes the video data. Improved compression and decompression techniques are desirable that improve the compression ratio without sacrificing much of the image quality, given limited network resources and the ever-increasing demands for higher video quality.
Summary of the Invention
Means for Solving the Problems
[0003] In one implementation, the present disclosure includes a method executed in a decoder, the method comprising: receiving, by a receiver of the decoder, a bitstream including coded pictures of one or more layers and a video parameter set (VPS), wherein when all layers specified by the VPS are independently coded without inter-layer prediction, the VPS includes an "each layer is an output layer set (OLS)" flag (each_layer_is_an_ols_flag), and the each_layer_is_an_ols_flag specifies whether each OLS includes only one layer; decoding, by a processor of the decoder, the coded pictures from the output layers of the OLS based on the each_layer_is_an_ols_flag in the VPS to generate decoded pictures; and transferring, by the processor, the decoded pictures for display as part of the decoded video sequence.
[0004] To support scalability, layers of pictures can be used. For example, a video can be encoded into multiple layers. The layers can be encoded without referring to other layers. Such layers are called simulcast layers. Thus, a simulcast layer can be decoded without referring to other layers. As another example, a layer can be encoded using inter-layer prediction. This enables the current layer to be encoded by including only the difference between the current layer and a reference layer. A layer can be organized into an OLS. An OLS is a set of layers that includes at least one output layer and any layer that supports decoding of the output layer. As a specific example, the first OLS may include a base layer, and the second OLS may include the base layer and, further, an enhancement layer with increased characteristics. In one example, the first OLS can be sent to the decoder to enable the video to be decoded at a base resolution, or the second OLS can be sent to enable the video to be decoded at a higher enhanced resolution. Thus, the video can be scaled based on user requirements. In some cases, scalability is not used and each layer is encoded as a simulcast layer. Some systems infer that each OLS should include a single layer (since no reference layer is used) if all layers are simulcast. This inference can improve coding efficiency because signaling can be omitted from the encoded bitstream. However, such an inference does not support multi-view. Multi-view is also known as stereoscopic video. In multi-view, two video sequences of the same scene are recorded by spatially offset cameras. The two video sequences are presented to the user through different lenses in a headset. In this way, a three-dimensional (3D) video and / or an impression of visual depth can be created by presenting different spatially offset sequences for each eye.Therefore, the OLS that implements multi-view contains two layers (e.g., one for each eye). However, when all layers are simulcast, the video decoder may use inference to infer that each OLS contains only one layer. This can lead to errors because the decoder may display only one layer of the multi-view or may not be able to continue displaying any layer. Therefore, when all layers are simulcast, the inference that each OLS contains a single layer may prevent the multi-view application from rendering properly in the decoder.
[0005] This example includes a mechanism that enables a video coding system to properly decode a multi-view video when all layers within the video are simulcast and inter-layer prediction is not used. The "vps_all_independent_layers_flag" can be set to 1 when none of the layers within the VPS use inter-layer prediction (e.g., all are simulcast), which is included in the bitstream within the VPS. When this flag is set to 1, the each_layer_is_an_ols_flag is signaled within the VPS. The each_layer_is_an_ols_flag can be set to specify whether each OLS contains a single layer or at least one OLS contains multiple layers (e.g., to support multi-views). Thus, the vps_all_independent_layers_flag and the each_layer_is_an_ols_flag can be used to support multi-view applications. Further, when this occurs, the ols_mode_idc can be set to 2 within the VPS. This explicitly signals the number of OLSs and the layers associated with the OLSs. The decoder can then use this information to correctly decode the OLSs including the multi-view video. This approach supports coding efficiency while correcting errors. Thus, the disclosed mechanism improves the functionality of the encoder and / or decoder. Further, the disclosed mechanism can reduce the bitstream size and thus reduce the utilization of processor, memory, and / or network resources in both the encoder and the decoder.
[0006] Optionally, in any of the foregoing aspects, other implementations of the aspect are provided, and the each_layer_is_an_ols_flag is set to 1 when it is specified that each OLS contains only one layer and each layer is the only output layer within each OLS.
[0007] Optionally, in any of the foregoing aspects, when other implementation forms of the aspect are provided and it is specified that at least one OLS includes a plurality of layers, each_layer_is_an_ols_flag is set to 0.
[0008] Optionally, in any of the foregoing aspects, when other implementation forms of the aspect are provided and the OLS mode identification code (ols_mode_idc) is equal to 2, the total number of OLSs is explicitly signaled, the layers associated with the OLSs are explicitly signaled, the "all independent layers of the VPS" flag (vps_all_independent_layers_flag) is set to 1, and when each_layer_is_an_ols_flag is set to 0, it is inferred that ols_mode_idc is equal to 2.
[0009] Optionally, in any of the foregoing aspects, when other implementation forms of the aspect are provided, the VPS includes the vps_all_independent_layers_flag set to 1 to specify that all layers specified by the VPS are coded independently without inter-layer prediction.
[0010] Optionally, in any of the foregoing aspects, when other implementation forms of the aspect are provided, the VPS includes the "VPS maximum layers minus 1" (vps_max_layers_minus1) syntax element that specifies the number of layers specified by the VPS, and when vps_max_layers_minus1 is greater than 0, the vps_all_independent_layers_flag is signaled.
[0011] Optionally, in any of the foregoing aspects, when other implementation forms of the aspect are provided, the VPS includes the "number of output layer sets minus 1" (num_output_layer_sets_minus1) that specifies the total number of OLSs specified by the VPS when ols_mode_idc is equal to 2.
[0012] In one embodiment, the present disclosure includes a method executed in an encoder, the method comprising: encoding, by a processor of the encoder, a bitstream including coded pictures of one or more layers; encoding, by the processor, a VPS into the bitstream, the VPS including an each_layer_is_an_ols_flag when all layers specified by the VPS are independently encoded without inter-layer prediction, the each_layer_is_an_ols_flag specifying whether each OLS includes only one layer; and storing, by a memory coupled to the processor, the bitstream for communication to a decoder.
[0013] To support scalability, layers of pictures can be used. For example, a video can be encoded into multiple layers. The layers can be encoded without referring to other layers. Such layers are called simulcast layers. Therefore, a simulcast layer can be decoded without referring to other layers. As another example, layers can be encoded using inter-layer prediction. This enables the current layer to be encoded by including only the difference between the current layer and a reference layer. Layers can be organized into OLSs. An OLS is a set of layers that includes at least one output layer and any layer that supports the decoding of the output layer. As a specific example, the first OLS may include a base layer, and the second OLS may include the base layer and, further, an enhancement layer with increased characteristics. In one example, the first OLS can be sent to the decoder to enable the video to be decoded at a base resolution, or the second OLS can be sent to enable the video to be decoded at a higher enhanced resolution. Thus, the video can be scaled based on user requirements. In some cases, scalability is not used and each layer is encoded as a simulcast layer. Some systems infer that each OLS should include a single layer (since no reference layer is used) when all layers are simulcast. This inference can improve coding efficiency because signaling can be omitted from the encoded bitstream. However, such an inference does not support multi-view. Multi-view is also known as stereoscopic video. In multi-view, two video sequences of the same scene are recorded by spatially offset cameras. The two video sequences are presented to the user through different lenses in a headset. In this way, a 3D video and / or an impression of visual depth can be created by presenting different spatially offset sequences for each eye.Accordingly, OLS that implements multi-view includes two layers (e.g., one for each eye). However, when all layers are simulcast, the video decoder may use inference to infer that each OLS includes only one layer. This may result in an error because the decoder may display only one layer of the multi-view or may not be able to continue displaying any layer. Therefore, when all layers are simulcast, the inference that each OLS includes a single layer may prevent the multi-view application from rendering properly in the decoder.
[0014] This example includes a mechanism that enables a video coding system to properly decode a multi-view video when all layers within the video are simulcast and layer - to - layer prediction is not used. The vps_all_independent_layers_flag is included in the bitstream within the VPS and can be set to 1 when none of the layers use layer - to - layer prediction (e.g., all are simulcast). When this flag is set to 1, the each_layer_is_an_ols_flag is signaled in the VPS. The each_layer_is_an_ols_flag can be set to specify whether each OLS contains a single layer or at least one OLS contains multiple layers (e.g., to support multi - view). Thus, the vps_all_independent_layers_flag and the each_layer_is_an_ols_flag can be used to support multi - view applications. Further, when this occurs, the ols_mode_idc can be set to 2 in the VPS. This explicitly signals the number of OLSs and the layers associated with the OLSs. The decoder can then use this information to correctly decode the OLSs including the multi - view video. This approach supports coding efficiency while correcting errors. Thus, the disclosed mechanism improves the functionality of the encoder and / or decoder. Further, the disclosed mechanism can reduce the bitstream size and thus reduce the utilization of processor, memory, and / or network resources in both the encoder and the decoder.
[0015] Optionally, in any of the foregoing aspects, other implementations of the aspect are provided, and when specifying that each OLS contains only one layer and each layer is the only output layer within each OLS, the each_layer_is_an_ols_flag is set to 1.
[0016] Optionally, in any of the foregoing aspects, when other implementation forms of the aspect are provided and it is specified that at least one OLS includes a plurality of layers, each_layer_is_an_ols_flag is set to 0.
[0017] Optionally, in any of the foregoing aspects, when other implementation forms of the aspect are provided and ols_mode_idc is equal to 2, the total number of OLSs is explicitly signaled, the layers associated with the OLSs are explicitly signaled, vps_all_independent_layers_flag is set to 1, and when each_layer_is_an_ols_flag is set to 0, ols_mode_idc is inferred to be equal to 2.
[0018] Optionally, in any of the foregoing aspects, when other implementation forms of the aspect are provided, the VPS includes vps_all_independent_layers_flag set to 1 to specify that all layers specified by the VPS are independently coded without inter-layer prediction.
[0019] Optionally, in any of the foregoing aspects, when other implementation forms of the aspect are provided, the VPS includes the vps_max_layers_minus1 syntax element that specifies the number of layers specified by the VPS, and when vps_max_layers_minus1 is greater than 0, vps_all_independent_layers_flag is signaled.
[0020] Optionally, in any of the foregoing aspects, when other implementation forms of the aspect are provided, the VPS includes num_output_layer_sets_minus1 that specifies the total number of OLSs specified by the VPS when ols_mode_idc is equal to 2.
[0021] In one embodiment, the present disclosure includes a video coding device including a processor, a receiver coupled to the processor, a memory coupled to the processor, and a transmitter coupled to the processor, wherein the processor, the receiver, the memory, and the transmitter are configured to execute the method according to any of the foregoing aspects.
[0022] In one embodiment, the present disclosure includes a non-transitory computer-readable medium including a computer program product for use by a video coding device, the computer program product including computer-executable instructions stored in the non-transitory computer-readable medium that, when executed by a processor, cause the video coding device to execute the method according to any of the foregoing aspects.
[0023] In one embodiment, the present disclosure includes a decoder, the decoder including receiving means for receiving a bitstream including coded pictures of one or more layers and a VPS, wherein when all layers specified by the VPS are independently coded without inter-layer prediction, the VPS includes an each_layer_is_an_ols_flag, the each_layer_is_an_ols_flag specifying whether each OLS includes only one layer; decoding means for decoding a coded picture from an output layer of the OLS based on the each_layer_is_an_ols_flag in the VPS to generate a decoded picture; and transfer means for transferring the decoded picture for display as part of the decoded video sequence.
[0024] To support scalability, layers of a picture can be used. For example, a video can be encoded into multiple layers. Layers can be encoded without referring to other layers. Such layers are called simulcast layers. Thus, a simulcast layer can be decoded without referring to other layers. As another example, layers can be encoded using inter-layer prediction. This enables the current layer to be encoded by including only the difference between the current layer and a reference layer. Layers can be organized into OLSs. An OLS is a set of layers that includes at least one output layer and any layer that supports decoding of the output layer. As a specific example, a first OLS may include a base layer, and a second OLS may include the base layer and, additionally, an enhancement layer with increased characteristics. In one example, the first OLS can be sent to a decoder to enable the video to be decoded at a base resolution, or the second OLS can be sent to enable the video to be decoded at a higher enhanced resolution. Thus, the video can be scaled based on user requirements. In some cases, scalability is not used and each layer is encoded as a simulcast layer. Some systems infer that each OLS should contain a single layer (since no reference layer is used) when all layers are simulcast. This inference can increase coding efficiency because signaling can be omitted from the encoded bitstream. However, such an inference does not support multi-view. Multi-view is also known as stereoscopic video. In multi-view, two video sequences of the same scene are recorded by spatially offset cameras. The two video sequences are presented to the user through different lenses in a headset. In this way, by presenting different spatially offset sequences for each eye, an impression of 3D video and / or visual depth can be created.Accordingly, OLS that implements multi-view includes two layers (e.g., one for each eye). However, when all layers are simulcast, the video decoder may use inference to infer that each OLS includes only one layer. This can result in errors because the decoder may display only one layer of the multi-view or may not be able to continue displaying any of the layers. Thus, when all layers are simulcast, the inference that each OLS includes a single layer may prevent the multi-view application from rendering properly in the decoder.
[0025] This example includes a mechanism that enables a video coding system to properly decode a multi-view video when all layers within the video are simulcast and inter-layer prediction is not used. The vps_all_independent_layers_flag is included in the bitstream within the VPS and can be set to 1 when none of the layers use inter-layer prediction (e.g., all are simulcast). When this flag is set to 1, the each_layer_is_an_ols_flag is signaled in the VPS. The each_layer_is_an_ols_flag can be set to specify whether each OLS contains a single layer or at least one OLS contains multiple layers (e.g., to support multi-view). Thus, the vps_all_independent_layers_flag and the each_layer_is_an_ols_flag can be used to support multi-view applications. Further, when this occurs, the ols_mode_idc can be set to 2 in the VPS. This explicitly signals the number of OLSs and the layers associated with the OLSs. The decoder can then use this information to correctly decode the OLSs including the multi-view video. This approach supports coding efficiency while correcting errors. Thus, the disclosed mechanism improves the functionality of the encoder and / or decoder. Further, the disclosed mechanism can reduce the bitstream size and thus reduce the utilization of processor, memory, and / or network resources in both the encoder and decoder.
[0026] Optionally, in any of the foregoing aspects, other implementations of the aspect are provided, and the decoder is further configured to execute the method described in any of the foregoing aspects.
[0027] In one embodiment, the present disclosure includes an encoder that encodes a bitstream including coded pictures of one or more layers, and encoding means for encoding a VPS including an each_layer_is_an_ols_flag into the bitstream when all layers specified by the VPS are independently encoded without inter-layer prediction, where the each_layer_is_an_ols_flag specifies whether each OLS includes only one layer, and storage means for storing the bitstream for communication to a decoder.
[0028] To support scalability, layers of pictures can be used. For example, a video can be encoded into multiple layers. The layers can be encoded without referring to other layers. Such layers are called simulcast layers. Therefore, a simulcast layer can be decoded without referring to other layers. As another example, layers can be encoded using inter-layer prediction. This enables the current layer to be encoded by including only the difference between the current layer and a reference layer. Layers can be organized into OLSs. An OLS is a set of layers that includes at least one output layer and any layer that supports decoding of the output layer. As a specific example, the first OLS may include a base layer, and the second OLS may include the base layer and, further, an enhancement layer with increased characteristics. In one example, the first OLS can be sent to a decoder to enable the video to be decoded at a base resolution, or the second OLS can be sent to enable the video to be decoded at a higher enhanced resolution. Thus, the video can be scaled based on user requirements. In some cases, scalability is not used and each layer is encoded as a simulcast layer. Some systems infer that each OLS should include a single layer (since no reference layer is used) when all layers are simulcast. This inference can improve coding efficiency because signaling can be omitted from the encoded bitstream. However, such an inference does not support multi-view. Multi-view is also known as stereoscopic video. In multi-view, two video sequences of the same scene are recorded by spatially offset cameras. The two video sequences are presented to the user through different lenses in a headset. In this way, by presenting different spatially offset sequences for each eye, an impression of 3D video and / or visual depth can be created.Therefore, OLS that implements multi-view includes two layers (for example, one for each eye). However, when all layers are simulcast, the video decoder may use inference to infer that each OLS includes only one layer. This can result in errors because the decoder may display only one layer of the multi-view or may not be able to continue displaying any of the layers. Therefore, when all layers are simulcast, the inference that each OLS includes a single layer may prevent the multi-view application from rendering properly in the decoder.
[0029] This example includes a mechanism that enables a video coding system to properly decode a multi-view video when all layers within the video are simulcast and layer - to - layer prediction is not used. The vps_all_independent_layers_flag is included in the bitstream within the VPS and can be set to 1 when none of the layers use layer - to - layer prediction (e.g., all are simulcast). When this flag is set to 1, the each_layer_is_an_ols_flag is signaled in the VPS. The each_layer_is_an_ols_flag can be set to specify whether each OLS contains a single layer or at least one OLS contains multiple layers (e.g., to support multi - views). Thus, the vps_all_independent_layers_flag and the each_layer_is_an_ols_flag can be used to support multi - view applications. Further, when this occurs, the ols_mode_idc can be set to 2 in the VPS. This explicitly signals the number of OLSs and the layers associated with the OLSs. The decoder can then use this information to correctly decode the OLSs that include multi - view video. This approach supports coding efficiency while correcting errors. Thus, the disclosed mechanism improves the functionality of the encoder and / or decoder. Further, the disclosed mechanism can reduce the bitstream size and thus reduce the utilization of processor, memory, and / or network resources in both the encoder and the decoder.
[0030] Optionally, in any of the foregoing aspects, other implementations of the aspect are provided and the encoder is further configured to execute the method described in any of the foregoing aspects.
[0031] For the sake of clarity speed, any one of the above embodiments may be combined with any one or more of the other above embodiments to create new embodiments within the scope of the present disclosure.
[0032] These and other features will be more clearly understood from the following detailed description together with the accompanying drawings and the claims.
[0033] For a more complete understanding of the present disclosure, reference is now made to the following brief description taken in connection with the accompanying drawings and detailed description in which like reference numerals represent like parts.
Brief Description of the Drawings
[0034]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
DETAILED DESCRIPTION OF THE INVENTION
[0035] Exemplary implementations of one or more embodiments are provided below, but it should be first understood that the disclosed systems and / or methods can be implemented using any number of techniques, whether currently known or existing. The present disclosure is not limited to the exemplary implementations, drawings, and techniques shown and described herein, including the exemplary designs and implementations shown herein, but can be varied within the scope of the appended claims, together with the full scope of their equivalents.
[0036] The following terms are defined as follows unless used in the opposite context herein. Specifically, the following definitions are for the purpose of further clarifying the present disclosure. However, the terms may be described differently in different contexts. Therefore, the following definitions are to be regarded as supplementary and not as limiting any other definitions provided herein for such terms.
[0037] A bitstream is a sequence of bits that contains video data compressed for transmission between an encoder and a decoder. An encoder is a device configured to compress video data into a bitstream using an encoding process. A decoder is a device configured to reconstruct video data from a bitstream for display using a decoding process. A picture is an array of luma samples and / or an array of chroma samples that makes up a frame or a field thereof. For clarity of explanation, a picture that is encoded or decoded may be referred to as the current picture.
[0038] A Network Abstraction Layer (NAL) unit is a syntax structure that contains data in the form of a Raw Byte Sequence Payload (RBSP), an indication of the type of data, and, optionally, emulation prevention bytes interspersed. A Video Coding Layer (VCL) NAL unit is a NAL unit encoded to contain video data, such as a coded slice of a picture. A non-VCL NAL unit is a NAL unit that contains non-video data, such as syntax and / or parameters to support decoding of video data, performing compliance checks, or other operations. A layer is a set of VCL NAL units that share certain specified characteristics (e.g., common resolution, frame rate, image size, etc.) and the associated non-VCL NAL units. The VCL NAL units of a layer may share a particular value of the NAL unit header layer identifier (nuh_layer_id). A coded picture is a coded representation of a picture that includes VCL NAL units having a particular value of the NAL unit header layer identifier (nuh_layer_id) within an Access Unit (AU) and includes all Coding Tree Units (CTUs) of the picture. A decoded picture is a picture generated by applying a decoding process to a coded picture.
[0039] An output layer set (OLS) is a set of layers where one or more layers are designated as output layers. An output layer is the layer designated for output (e.g., to a display). The 0th OLS contains only the bottom layer (the layer with the bottom layer identifier), and thus is an OLS that contains only an output layer. A video parameter set (VPS) is a data unit that contains parameters related to the entire video. Inter-layer prediction is a mechanism for encoding a current picture in a current layer by referring to a reference picture in a reference layer, where the current picture and the reference picture are included in the same access unit, and the reference layer contains a lower nuh_layer_id than the current layer.
[0040] The "each layer is an OLS" flag (each_layer_is_an_ols_flag) is a syntax element that signals whether each OLS in the bitstream contains a single layer. The OLS mode identification code (ols_mode_idc) is a syntax element that indicates information about the number of OLSs, the layers of the OLSs, and the output layers of the OLSs. The "all independent layers of the VPS" flag (vps_all_independent_layers_flag) is a syntax element that signals whether inter-layer prediction is used to encode any of the layers in the bitstream. "VPS maximum layers minus 1" (vps_max_layers_minus1) is a syntax element that signals the number of layers specified by the VPS, and thus the maximum number of layers permitted in the corresponding coded video sequence. "Number of output layer sets minus 1" (num_output_layer_sets_minus1) is a syntax element that specifies the total number of OLSs specified by the VPS.
[0041] In this specification, the following initials, Coding Tree Block (CTB), Coding Tree Unit (CTU), Coding Unit (CU), Coded Video Sequence (CVS), Joint Video Experts Team (JVET), Motion Constrained Tile Set (MCTS), Maximum Transmission Unit (MTU), Network Abstraction Layer (NAL), Output Layer Set (OLS), Picture Order Count (POC), Raw Byte Sequence Payload (RBSP), Sequence Parameter Set (SPS), Video Parameter Set (VPS), Versatile Video Coding (VVC), and Working Draft (WD) are used.
[0042] To reduce the size of video files while minimizing data loss, many video compression techniques can be used. For example, video compression techniques can include performing spatial (e.g., intra-picture) prediction and / or temporal (e.g., inter-picture) prediction to reduce or remove data redundancy in a video sequence. In the case of block-based video coding, a video slice (e.g., a video picture, or a part of a video picture) can be partitioned into video blocks, also sometimes called tree blocks, coding tree blocks (CTBs), coding tree units (CTUs), coding units (CUs), and / or coding nodes. Video blocks within an intra-coded (I) slice of a picture are coded using spatial prediction against reference samples in neighboring blocks within the same picture. Video blocks within an inter-coded unidirectional prediction (P) slice or bidirectional prediction (B) slice of a picture can be coded by using spatial prediction against reference samples in neighboring blocks within the same picture or temporal prediction against reference samples in other reference pictures. A picture may sometimes be called a frame and / or an image, and a reference picture may sometimes be called a reference frame and / or a reference image. Spatial or temporal prediction results in a predicted block representing the image block. Residual data represents the pixel difference between the original image block and the predicted block. Thus, an inter-coded block is coded according to a motion vector indicating a block of reference samples forming the predicted block and residual data indicating the difference between the coded block and the predicted block. An intra-coded block is coded according to an intra-coding mode and residual data. For further compression, the residual data can be transformed from a pixel domain to a transform domain. As a result, residual transform coefficients are obtained, which can be quantized. The quantized transform coefficients may first be arranged in a two-dimensional array. The quantized transform coefficients can be scanned to generate a one-dimensional vector of transform coefficients. To achieve further compression, entropy coding may be applied.Such video compression techniques will be described in more detail below.
[0043] To ensure that the encoded video can be accurately decoded, the video is encoded and decoded according to the corresponding video coding standard. Video coding standards include International Telecommunication Union (ITU) Telecommunication Standardization Sector (ITU-T) H.261, International Organization for Standardization / International Electrotechnical Commission (ISO / IEC) Moving Picture Experts Group (MPEG)-1 Part 2, ITU-T H.262 or ISO / IEC MPEG-2 Part 2, ITU-T H.263, ISO / IEC MPEG-4 Part 2, ITU-T H.264 or Advanced Video Coding (AVC) also known as ISO / IEC MPEG-4 Part 10, and High Efficiency Video Coding (HEVC) also known as ITU-T H.265 or MPEG-H Part 2. AVC includes extensions such as Scalable Video Coding (SVC), Multi-View Video Coding (MVC) and Multi-View Video Coding Plus Depth (MVC+D), and 3D AVC (3D-AVC). HEVC includes extensions such as Scalable HEVC (SHVC), Multi-View HEVC (MV-HEVC), and 3D HEVC (3D-HEVC). The joint video experts team (JVET) of ITU-T and ISO / IEC has started the development of a video coding standard called Versatile Video Coding (VVC). VVC is included in the Working Draft (WD) including JVET-O2001-v14.
[0044] The layers of a picture can be used to support scalability. For example, a video can be encoded into multiple layers. A layer can be encoded without referring to other layers. Such a layer is called a simulcast layer. Thus, a simulcast layer can be decoded without referring to other layers. As another example, a layer can be encoded using inter-layer prediction. This enables the current layer to be encoded by including only the difference between the current layer and a reference layer. For example, the current layer and the reference layer can include the same video sequence encoded by varying characteristics such as signal-to-noise ratio (SNR), picture size, frame rate, etc. Layers can be organized into an output layer set (OLS). An OLS is a set of layers that includes at least one output layer and any layer that supports the decoding of the output layer. As a specific example, the first OLS can include a base layer, and the second OLS can include the base layer and, further, an enhancement layer with increased characteristics. In an example where the characteristic is picture resolution, the first OLS can be sent to the decoder to enable the video to be decoded at the base resolution, or the second OLS can be sent to enable the video to be decoded at a higher improved resolution. Thus, the video can be scaled based on user requirements.
[0045] In some cases, scalability is not used and each layer is coded as a simulcast layer. Some systems infer that each OLS should contain a single layer (since no reference layer is used) when all layers are simulcast. This inference can increase coding efficiency since signaling can be omitted from the encoded bitstream. However, such an inference does not support multi-view, which is also known as stereoscopic video. In multi-view, two video sequences of the same scene are recorded by spatially offset cameras. The two video sequences are presented to the user through different lenses in a headset. Thus, by presenting different spatially offset sequences for each eye, a three-dimensional (3D) video and / or an impression of visual depth can be created. Therefore, an OLS implementing multi-view contains two layers (e.g., one for each eye). However, when all layers are simulcast, the video decoder may use the inference to infer that each OLS contains only one layer. This can result in errors since the decoder may display only one layer of the multi-view or may not be able to continue displaying any layer. Therefore, the inference that each OLS contains a single layer when all layers are simulcast can prevent multi-view applications from being properly rendered in the decoder.
[0046] In this specification, a mechanism is disclosed that enables a video coding system to properly decode a multi-view video when all layers within the video are simulcast and inter-layer prediction is not used. The "vps_all_independent_layers_flag" can be set to 1 when none of the layers within the VPS use inter-layer prediction (e.g., all are simulcast), which is included in the bitstream within the VPS. When this flag is set to 1, the "each_layer_is_an_ols_flag" is signaled in the VPS. The each_layer_is_an_ols_flag can be set to specify whether each OLS contains a single layer or at least one OLS contains multiple layers (e.g., to support multi-view). Thus, the vps_all_independent_layers_flag and each_layer_is_an_ols_flag can be used to support multi-view applications. Further, when this occurs, the OLS mode identification code (ols_mode_idc) can be set to 2 in the VPS. This explicitly signals the number of OLSs and the layers associated with the OLSs. The decoder can then use this information to correctly decode the OLSs including the multi-view video. This approach supports coding efficiency while correcting errors. Thus, the disclosed mechanism improves the functionality of the encoder and / or decoder. Further, the disclosed mechanism can reduce the bitstream size and thus reduce the utilization of processor, memory, and / or network resources in both the encoder and decoder.
[0047] FIG. 1 is a flowchart showing an exemplary operational method 100 for coding a video signal. Specifically, the video signal is encoded in an encoder. The encoding process compresses the video signal using various mechanisms to reduce the video file size. The smaller file size enables the compressed video file to be transmitted to the user while reducing the associated bandwidth overhead. The decoder then decodes the compressed video file and reconstructs the original video signal for display to the end user. The decoding process generally mirrors the encoding process to enable the decoder to consistently reconstruct the video signal.
[0048] In step 101, a video signal is input to the encoder. For example, the video signal may be an uncompressed video file stored in memory. As another example, the video file may be captured by a video capture device such as a video camera and encoded to support live streaming of the video. The video file may include both an audio component and a video component. The video component includes a series of image frames that give a visual impression of movement when viewed in sequence. The frames include pixels represented by light, referred to herein as the luma component (or luma samples), and color, referred to as the chroma component (or color samples). In some examples, the frames may also include depth values to support three-dimensional viewing.
[0049] In step 103, the video is partitioned into blocks. Partitioning includes re - dividing the pixels within each frame into square and / or rectangular blocks for compression. For example, in High Efficiency Video Coding (HEVC), also known as H.265 and MPEG - H Part 2, a frame can first be divided into Coding Tree Units (CTUs) which are blocks of a predefined size (e.g., 64 pixels × 64 pixels). A CTU contains both luma samples and chroma samples. Using a coding tree, the CTU is divided into blocks, and the blocks may be recursively re - divided until a configuration that supports further encoding is achieved. For example, the luma component of a frame may be re - divided until the individual blocks contain relatively uniform illumination values. Further, the chroma component of a frame may be re - divided until the individual blocks contain relatively uniform color values. Thus, the partitioning mechanism varies according to the content of the video frame.
[0050] In step 105, various compression mechanisms are used to compress the image blocks partitioned in step 103. For example, inter prediction and / or intra prediction may be used. Inter prediction is designed to utilize the fact that objects within a common scene tend to appear in successive frames. Thus, blocks representing objects in a reference frame need not be repeatedly described in adjacent frames. Specifically, an object such as a table may remain in a fixed position over multiple frames. Thus, the table is described once, and adjacent frames can refer to the reference frame. A pattern matching mechanism may be used for matching objects over multiple frames. Further, a moving object can be represented over multiple frames, for example, by the movement of the object or the movement of the camera. As a specific example, a video may show a car moving across the screen over multiple frames. To describe such movement, motion vectors can be used. A motion vector is a two-dimensional vector that provides an offset from the coordinates of an object in a frame to the coordinates of the object in a reference frame. Thus, inter prediction can encode an image block in the current frame as a set of motion vectors indicating the offsets from corresponding blocks in the reference frame.
[0051] Intra prediction encodes blocks within a common frame. Intra prediction exploits the fact that luma and chroma components tend to cluster within a frame. For example, a green patch of a tree part tends to be placed adjacent to similar green patches. Intra prediction uses multiple directional prediction modes (e.g., 33 in HEVC), a planar mode, and a direct current (DC) mode. The directional mode indicates that the current block is similar / same as samples of neighboring blocks in the corresponding direction. The planar mode indicates that a series of blocks along a row / column (e.g., a plane) can be interpolated based on neighboring blocks at the edge of the row. The planar mode effectively indicates a smooth transition of light / color across the row / column by using a relatively constant slope in the change of values. The DC mode is used for boundary smoothing and indicates that the block is similar / same as the average value associated with samples of all neighboring blocks associated with the angular direction of the directional prediction mode. Thus, an intra prediction block can represent an image block as various relational prediction mode values instead of actual values. Further, an inter prediction block can represent an image block as motion vector values instead of actual values. In either case, the prediction block may not accurately represent the image block in some cases. Any difference is stored in the residual block. To further compress the file, a transform may be applied to the residual block.
[0052] In step 107, various filtering techniques can be applied. In HEVC, the filter is applied according to the in-loop filtering method. In the block-based prediction described above, a blocky image may be created at the decoder. Further, the block-based prediction method may encode a block and then reconstruct the encoded block for later use as a reference block. The in-loop filtering method repeatedly applies a noise suppression filter, a deblocking filter, an adaptive loop filter, and a sample adaptive offset (SAO) filter to blocks / frames. These filters reduce such blocking artifacts so that the encoded file can be accurately reconstructed. Further, these filters reduce the artifacts in the reconstructed reference block so that the artifacts are less likely to create additional artifacts in subsequent blocks encoded based on the reconstructed reference block.
[0053] When the video signal is partitioned, compressed, and filtered, in step 109, the resulting data is encoded into a bitstream. The bitstream includes the data described above, as well as any signaling data desired to support proper video signal reconstruction at the decoder. For example, such data may include partition data, prediction data, residual blocks, and various flags that provide coding instructions to the decoder. The bitstream can be stored in memory for transmission to a decoder on demand. The bitstream may be broadcast and / or multicast to multiple decoders. The creation of the bitstream is an iterative process. Thus, steps 101, 103, 105, 107, and 109 can be performed continuously and / or simultaneously over a number of frames and blocks. The order shown in FIG. 1 is presented for clarity and ease of explanation and is not intended to limit the video coding process to a particular order.
[0054] The decoder receives a bitstream and starts the decoding process at step 111. Specifically, the decoder uses an entropy decoding method to convert the bitstream into corresponding syntax and video data. At step 111, the decoder determines the partitioning of the frame using the syntax data from the bitstream. The partitioning should match the result of the block partitioning at step 103. Next, the entropy coding / decoding used at step 111 will be described. The encoder makes many choices during the compression process, such as selecting a block partitioning method from several possible options based on the spatial positioning of the values within the input image. Signaling the exact choice can use a large number of bins. As used herein, a bin is a binary value treated as a variable (e.g., a bit value that can vary depending on the context). Entropy coding makes it possible to discard any options that are clearly not executable in a particular case by the encoder and leave a set of acceptable options. Then, a codeword is assigned to each acceptable option. The length of the codeword is based on the number of acceptable options (e.g., one bin for two options, two bins for three to four options, etc.). Then, the encoder encodes the codeword for the selected option. This method reduces the size of the codeword because, in contrast to uniquely indicating a choice from a potentially large set of all possible options, the codeword only needs to be as large as necessary to uniquely indicate a choice from a small subset of acceptable options. Then, the decoder decodes the choice by determining the set of acceptable options in the same way as the encoder. By determining the set of acceptable options, the decoder can read the codeword and determine the choice made by the encoder.
[0055] In step 113, the decoder performs block decoding. Specifically, the decoder generates a residual block using the inverse transform. Next, the decoder reconstructs the image block according to the partitioning using the residual block and the corresponding prediction block. The prediction block may include both an intra prediction block and an inter prediction block as generated by the encoder in step 105. Then, the reconstructed image block is placed into the frame of the reconstructed video signal according to the partitioning data determined in step 111. The syntax in step 113 may also be signaled in the bitstream via entropy coding as described above.
[0056] In step 115, filtering is performed on the frame of the reconstructed video signal in a manner similar to step 107 in the encoder. For example, a noise suppression filter, a deblocking filter, an adaptive loop filter, and an SAO filter may be applied to the frame to remove blocking artifacts. Once the frame is filtered, in step 117, the video signal can be output to a display for the end user to view.
[0057] Figure 2 is a schematic diagram of an exemplary coding and decoding (codec) system 200 for video coding. Specifically, the codec system 200 provides functions that support the implementation of the operating method 100. The codec system 200 is generalized to show components used in both the encoder and the decoder. The codec system 200 receives and partitions a video signal as described for steps 101 and 103 of the operating method 100, resulting in a partitioned video signal 201. Next, when operating as an encoder, the codec system 200 compresses the partitioned video signal 201 into a coded bitstream as described for steps 105, 107, and 109 of method 100. When operating as a decoder, the codec system 200 generates an output video signal from the bitstream as described for steps 111, 113, 115, and 117 of the operating method 100. The codec system 200 includes a general codec control component 211, a transform scaling and quantization component 213, an intra-picture estimation component 215, an intra-picture prediction component 217, a motion compensation component 219, a motion estimation component 221, a scaling and inverse transform component 229, a filter control analysis component 227, an in-loop filter component 225, a decoded picture buffer component 223, and a header formatting and context-adaptive binary arithmetic coding (CABAC) component 231. Such components are coupled as shown. In FIG. 2, solid lines indicate the movement of data to be encoded / decoded, and dashed lines indicate the movement of control data that controls the operation of other components. All components of the codec system 200 may be present within the encoder. The decoder may include a subset of the components of the codec system 200.For example, the decoder may include an intra-picture prediction component 217, a motion compensation component 219, a scaling and inverse transform component 229, an in-loop filter component 225, and a decoded picture buffer component 223. Next, these components will be described.
[0058] The partitioned video signal 201 is a captured video sequence partitioned into blocks of pixels by a coding tree. The coding tree uses various partitioning modes to re-partition a block of pixels into smaller blocks of pixels. These blocks can then be further re-partitioned into even smaller blocks. A block may sometimes be referred to as a node on the coding tree. A larger parent node is divided into smaller child nodes. The number of times a node is re-partitioned is called the depth of the node / coding tree. In some cases, the partitioned blocks can be included in a coding unit (CU). For example, a CU can be a sub-part of a CTU that includes a luminance block, a chrominance red (Cr) block, and a chrominance blue (Cb) block, along with the corresponding syntax instruction for the CU. The partitioning modes can include a binary tree (BT), a ternary tree (TT), and a quadtree (QT) that are used to partition a node into two, three, or four child nodes of various shapes depending on the partitioning mode used. The partitioned video signal 201 is transferred to a general coder control component 211, a transform scaling and quantization component 213, an intra-picture estimation component 215, a filter control analysis component 227, and a motion estimation component 221 for compression.
[0059] The general coder control component 211 is configured to make decisions related to encoding the images of a video sequence into a bitstream according to application constraints. For example, the general coder control component 211 manages the optimization of bitrate / bitstream size versus reconstructed quality. Such decisions can be made based on the availability of memory space / bandwidth and the image resolution requirements. The general coder control component 211 also manages the utilization of the buffer in light of the transmission speed to mitigate buffer underrun and overrun problems. To manage these problems, the general coder control component 211 manages the partitioning, prediction, and filtering by other components. For example, the general coder control component 211 may dynamically increase the compression complexity to increase the resolution and increase the bandwidth usage, or decrease the compression complexity to decrease the resolution and bandwidth usage. Thus, the general coder control component 211 controls other components of the codec system 200 to balance the video signal reconstructed quality with respect to the bitrate. The general coder control component 211 creates control data for controlling the operation of other components. The control data is also transferred to the header formatting and CABAC component 231 and encoded into the bitstream to signal the parameters for decoding at the decoder.
[0060] The partitioned video signal 201 is also sent to the motion estimation component 221 and the motion compensation component 219 for inter prediction. A frame or slice of the partitioned video signal 201 may be divided into a plurality of video blocks. The motion estimation component 221 and the motion compensation component 219 perform inter prediction coding of the received video block with respect to one or more blocks in one or more reference frames to provide temporal prediction. The codec system 200 may execute a plurality of coding paths, for example, to select an appropriate coding mode for each block of video data.
[0061] The motion estimation component 221 and the motion compensation component 219 can be highly integrated but are shown separately for conceptual purposes. The motion estimation performed by the motion estimation component 221 is a process of generating motion vectors that estimate the motion of video blocks. The motion vectors can indicate, for example, the displacement of the coded object with respect to the prediction block. The prediction block is a block that is found to closely match the block to be coded in terms of pixel difference. The prediction block may also be referred to as a reference block. Such pixel differences can be determined by the sum of absolute differences (SAD), the sum of squared differences (SSD), or other difference metrics. HEVC uses several coding objects including coding tree units (CTUs), coding tree blocks (CTBs), and coding units (CUs). For example, a CTU can be divided into CTBs, and a CTB can be divided into coding blocks (CBs) for inclusion in CUs. A CU can be coded as a prediction unit (PU) containing prediction data and / or a transform unit (TU) containing transformed residual data for the CU. The motion estimation component 221 generates motion vectors, PUs, and TUs by using rate-distortion analysis as part of the rate-distortion optimization process. For example, the motion estimation component 221 may determine multiple reference blocks, multiple motion vectors, etc. for the current block / frame, and may select the reference block, motion vector, etc. with the best rate-distortion characteristics. The best rate-distortion characteristics balance both the quality of video reconstruction (e.g., the amount of data loss due to compression) and the coding efficiency (e.g., the size of the final encoded data).
[0062] In some examples, the codec system 200 may calculate values of sub-integer pixel positions of reference pictures stored in the decoded picture buffer component 223. For example, the video codec system 200 may interpolate values of 1 / 4 pixel positions, 1 / 8 pixel positions, or other fractional pixel positions of the reference picture. Thus, the motion estimation component 221 may perform motion searches for full pixel positions and fractional pixel positions and output motion vectors with fractional pixel accuracy. The motion estimation component 221 calculates the motion vector of the PU of the video block in the inter-coded slice by comparing the position of the PU with the position of the predicted block of the reference picture. The motion estimation component 221 outputs the calculated motion vector as motion data to the header formatting and CABAC component 231 for encoding and outputs the motion to the motion compensation component 219.
[0063] Motion compensation performed by the motion compensation component 219 may involve fetching or generating a predicted block based on the motion vector determined by the motion estimation component 221. Also in this case, in some examples, the motion estimation component 221 and the motion compensation component 219 may be functionally integrated. Upon receiving the motion vector for the current video block's PU, the motion compensation component 219 may find the predicted block indicated by the motion vector. Then, by subtracting the pixel values of the predicted block from the pixel values of the currently encoded video block, a residual video block is formed and pixel difference values are formed. Generally, the motion estimation component 221 performs motion estimation for the luma component, and the motion compensation component 219 uses the motion vector calculated based on the luma component for both the chroma component and the luma component. The predicted block and the residual block are transferred to the transform scaling and quantization component 213.
[0064] The partitioned video signal 201 is also sent to the in-picture estimation component 215 and the in-picture prediction component 217. Similar to the motion estimation component 221 and the motion compensation component 219, the in-picture estimation component 215 and the in-picture prediction component 217 can be highly integrated, but are shown separately for conceptual purposes. The in-picture estimation component 215 and the in-picture prediction component 217 perform intra prediction of the current block for a block within the current frame as an alternative to inter prediction performed by the inter-frame motion estimation component 221 and the motion compensation component 219 as described above. In particular, the in-picture estimation component 215 determines the intra prediction mode to use for encoding the current block. In some examples, the in-picture estimation component 215 selects the intra prediction mode appropriate for encoding the current block from a plurality of tested intra prediction modes. The selected intra prediction mode is then transferred to the header formatting and CABAC component 231 for encoding.
[0065] For example, the in-picture estimation component 215 calculates a rate-distortion value using rate-distortion analysis of various tested intra prediction modes and selects the intra prediction mode having the best rate-distortion characteristics among the tested modes. Rate-distortion analysis generally determines the amount of distortion (or error) between the encoded block and the original unencoded block encoded to produce the encoded block, as well as the bit rate (e.g., number of bits) used to produce the encoded block. The in-picture estimation component 215 calculates a ratio from the distortion and rate of various encoded blocks to determine which intra prediction mode exhibits the best rate-distortion value for the block. Additionally, the in-picture estimation component 215 can be configured to encode depth blocks of the depth map using a depth modeling mode (DMM) based on rate-distortion optimization (RDO).
[0066] When the picture-in-prediction component 217 is implemented in the encoder, it may generate a residual block from the prediction block based on the selected intra prediction mode determined by the picture-in-estimation component 215, or when implemented in the decoder, it may read the residual block from the bitstream. The residual block includes the difference in values between the prediction block and the original block, represented as a matrix. The residual block is then transferred to the transform scaling and quantization component 213. The picture-in-estimation component 215 and the picture-in-prediction component 217 can operate on both the luma component and the chroma component.
[0067] The transform scaling and quantization component 213 is configured to further compress the residual block. The transform scaling and quantization component 213 applies a transform, such as a discrete cosine transform (DCT), a discrete sine transform (DST), or a conceptually similar transform, to the residual block to generate a video block containing residual transform coefficient values. A wavelet transform, an integer transform, a subband transform, or other types of transforms can also be used. The transform can transform the residual information from the pixel value domain to a transform domain such as the frequency domain. The transform scaling and quantization component 213 is also configured to scale the transformed residual information, for example, based on frequency. Such scaling involves applying a scale factor to the residual information so that different frequency information is quantized at different granularities, which can affect the final visual quality of the reconstructed video. The transform scaling and quantization component 213 is also configured to quantize the transform coefficients to further reduce the bitrate. The quantization process can reduce the bit depth associated with some or all of the coefficients. The degree of quantization may be adjusted by adjusting the quantization parameter. In some examples, the transform scaling and quantization component 213 may then perform a scan of the matrix containing the quantized transform coefficients. The quantized transform coefficients are transferred to the header formatting and CABAC component 231 and encoded into the bitstream.
[0068] The scaling and inverse transform component 229 applies the opposite operations of the transform scaling and quantization component 213 to support motion estimation. The scaling and inverse transform component 229 applies inverse scaling, transform, and / or quantization to reconstruct the residual block within the pixel region for later use, for example, as a reference block that can be a prediction block for other current blocks. The motion estimation component 221 and / or the motion compensation component 219 may calculate the reference block by returning and adding the residual block to the corresponding prediction block for use in motion estimation of subsequent blocks / frames. A filter is applied to the reconstructed reference block to reduce artifacts created during scaling, quantization, and transformation. Otherwise, such artifacts can cause inaccurate predictions (and create additional artifacts) when subsequent blocks are predicted.
[0069] The filter control analysis component 227 and the in-loop filter component 225 apply a filter to the residual block and / or the reconstructed image block. For example, the transformed residual block from the scaling and inverse transform component 229 may be combined with the corresponding prediction block from the intra-picture prediction component 217 and / or the motion compensation component 219 to reconstruct the original image block. Then, the filter may be applied to the reconstructed image block. In some examples, the filter may alternatively be applied to the residual block. Similar to the other components in FIG. 2, the filter control analysis component 227 and the in-loop filter component 225 are highly integrated and may be implemented together, but are shown separately for conceptual purposes. The filter applied to the reconstructed reference block is applied to a specific spatial region and includes a plurality of parameters for adjusting how such a filter is applied. The filter control analysis component 227 analyzes the reconstructed reference block to determine where such a filter should be applied and sets the corresponding parameters. Such data is transferred as filter control data for encoding to the header formatting and CABAC component 231. The in-loop filter component 225 applies such a filter based on the filter control data. The filter may include a deblocking filter, a noise suppression filter, a SAO filter, and an adaptive loop filter. Such a filter may be applied in the spatial / pixel domain (e.g., the reconstructed pixel block) or the frequency domain depending on the example.
[0070] When operating as an encoder, the filtered reconstructed image block, residual block, and / or prediction block are stored in the decoded picture buffer component 223 for later use in motion estimation, as described above. When operating as a decoder, the decoded picture buffer component 223 stores the reconstructed and filtered blocks and transfers them towards the display as part of the output video signal. The decoded picture buffer component 223 may be any memory device capable of storing prediction blocks, residual blocks, and / or reconstructed image blocks.
[0071] The header formatting and CABAC component 231 receives data from various components of the codec system 200 and encodes such data into a coded bitstream for transmission towards a decoder. Specifically, the header formatting and CABAC component 231 generates various headers for encoding control data such as general control data and filter control data. Further, prediction data including intra prediction and motion data, as well as residual data in the form of quantized transform coefficient data are all encoded into the bitstream. The final bitstream contains all the information desired by the decoder for reconstructing the original partitioned video signal 201. Such information may also include an intra prediction mode index table (also referred to as a codeword mapping table), definitions of coding contexts for various blocks, indication of the most accurate intra prediction mode, indication of partition information, etc. Such data may be encoded by using entropy coding. For example, the information may be encoded by using context adaptive variable length coding (CAVLC), CABAC, syntax based context adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or other entropy coding techniques. Following entropy coding, the coded bitstream may be transmitted to other devices (e.g., a video decoder), or may be archived for later transmission or retrieval.
[0072] FIG. 3 is a block diagram showing an exemplary video encoder 300. The video encoder 300 can be used to implement the encoding function of the codec system 200 and / or to implement steps 101, 103, 105, 107, and / or 109 of the operation method 100. The encoder 300 partitions the input video signal to obtain a partitioned video signal 301 that is substantially the same as the partitioned video signal 201. The partitioned video signal 301 is then compressed by components of the encoder 300 and encoded into a bitstream.
[0073] Specifically, the partitioned video signal 301 is transferred to the intra-prediction component 317 for intra prediction. The intra-prediction component 317 may be substantially the same as the intra-estimation component 215 and the intra-prediction component 217. The partitioned video signal 301 is also transferred to the motion compensation component 321 for inter prediction based on reference blocks in the decoded picture buffer component 323. The motion compensation component 321 may be substantially the same as the motion estimation component 221 and the motion compensation component 219. The prediction blocks and residual blocks from the intra-prediction component 317 and the motion compensation component 321 are transferred to the transform and quantization component 313 for transformation and quantization of the residual blocks. The transform and quantization component 313 may be substantially the same as the transform scaling and quantization component 213. The transformed and quantized residual blocks and the corresponding prediction blocks (along with associated control data) are transferred to the entropy coding component 331 for coding into the bitstream. The entropy coding component 331 may be substantially the same as the header formatting and CABAC component 231.
[0074] The transformed and quantized residual blocks and / or corresponding prediction blocks are also transferred from the transform and quantization component 313 to the inverse transform and quantization component 329 and are reconstructed into reference blocks for use by the motion compensation component 321. The inverse transform and quantization component 329 may be substantially the same as the scaling and inverse transform component 229. The in-loop filter within the in-loop filter component 325 is also applied to the residual blocks and / or the reconstructed reference blocks, depending on the example. The in-loop filter component 325 may be substantially the same as the filter control analysis component 227 and the in-loop filter component 225. The in-loop filter component 325 may include multiple filters as described for the in-loop filter component 225. The filtered blocks are then stored in the decoded picture buffer component 323 for use as reference blocks by the motion compensation component 321. The decoded picture buffer component 323 may be substantially the same as the decoded picture buffer component 223.
[0075] Figure 4 is a block diagram showing an exemplary video decoder 400. The video decoder 400 can be used to implement the decoding function of the codec system 200 and / or to implement steps 111, 113, 115, and / or 117 of the operation method 100. The decoder 400 receives, for example, a bitstream from the encoder 300 and generates an output video signal reconstructed based on the bitstream for display to an end user.
[0076] The bitstream is received by the entropy decoding component 433. The entropy decoding component 433 is configured to implement an entropy decoding method such as CAVLC, CABAC, SBAC, PIPE coding, or other entropy coding techniques. For example, the entropy decoding component 433 may use the header information to provide a context for interpreting additional data encoded as codewords within the bitstream. The decoded information includes any desired information for decoding the video signal, such as general control data, filter control data, partition information, motion data, prediction data, and quantized transform coefficients from residual blocks. The quantized transform coefficients are transferred to the inverse transform and quantization component 429 for reconstruction into the residual blocks. The inverse transform and quantization component 429 may be similar to the inverse transform and quantization component 329.
[0077] The reconfigured residual block and / or prediction block is transferred to the intra-prediction component 417 to reconstruct the image block based on the intra-prediction operation. The intra-prediction component 417 may be similar to the intra-estimation component 215 and the intra-prediction component 217. Specifically, the intra-prediction component 417 uses the prediction mode to find the position of the reference block within the frame and applies the residual block to the result to reconstruct the intra-predicted image block. The reconfigured intra-predicted image block and / or residual block and the corresponding inter-prediction data are transferred to the decoded picture buffer component 423 via the in-loop filter component 425, which may be substantially similar to the decoded picture buffer component 223 and the in-loop filter component 225 respectively. The in-loop filter component 425 filters the reconfigured image block, residual block, and / or prediction block, and such information is stored in the decoded picture buffer component 423. The reconfigured image block from the decoded picture buffer component 423 is transferred to the motion compensation component 421 for inter-prediction. The motion compensation component 421 may be substantially similar to the motion estimation component 221 and / or the motion compensation component 219. Specifically, the motion compensation component 421 uses the motion vector from the reference block to generate a prediction block and applies the residual block to the result to reconstruct the image block. The resulting reconfigured block may be transferred to the decoded picture buffer component 423 via the in-loop filter component 425. The decoded picture buffer component 423 continues to store additional reconfigured image blocks that can be reconstructed into a frame via the partition information. Such a frame may be placed within the sequence. This sequence is output to the display as the reconfigured output video signal.
[0078] FIG. 5 is a schematic diagram showing an exemplary multi-layer video sequence 500 configured for inter-layer prediction 521. The multi-layer video sequence 500 is encoded by an encoder such as codec system 200 and / or encoder 300 and can be decoded by a decoder such as codec system 200 and / or decoder 400, for example, according to method 100. The multi-layer video sequence 500 is included to illustrate an exemplary application for layers in an encoded video sequence. The multi-layer video sequence 500 is any video sequence that uses multiple layers, such as layer N 531 and layer N+1 532.
[0079] In one example, the multi-layer video sequence 500 may use inter-layer prediction 521. Inter-layer prediction 521 is applied between pictures 511, 512, 513, and 514 and pictures 515, 516, 517, and 518 of different layers. In the illustrated example, pictures 511, 512, 513, and 514 are part of layer N+1 532, and pictures 515, 516, 517, and 518 are part of layer N 531. Layers such as layer N 531 and / or layer N+1 532 are groups of pictures all associated with similar values of characteristics such as similar size, quality, resolution, signal-to-noise ratio, capabilities, etc. A layer can be formally defined as a set of VCL NAL units and associated non-VCL NAL units that share the same nuh_layer_id. A VCL NAL unit is a NAL unit encoded to contain video data, such as an encoded slice of a picture. A non-VCL NAL unit is a NAL unit that contains non-video data, such as syntax and / or parameters that support decoding of video data, performing compliance checks, or other operations.
[0080] In the illustrated example, layer N+1 532 is associated with a larger image size than layer N 531. Thus, pictures 511, 512, 513, and 514 within layer N+1 532 have, in this example, a larger picture size (e.g., greater height and width, and thus more samples) than pictures 515, 516, 517, and 518 within layer N 531. However, such pictures can be separated between layer N+1 532 and layer N 531 by other characteristics. Only two layers, layer N+1 532 and layer N 531, are shown, but a set of pictures can be divided into any number of layers based on the associated characteristics. Layer N+1 532 and layer N 531 may also be indicated by a layer identifier (ID). The layer ID is an item of data associated with a picture and indicates that the picture is part of the indicated layer. Thus, each picture 511 - 518 can be associated with a corresponding layer ID to indicate which layer N+1 532 or layer N 531 contains the corresponding picture. For example, the layer ID may include the NAL unit header layer identifier (nuh_layer_id), which is a syntax element that specifies an identifier for a layer that contains a NAL unit (e.g., including slices and / or parameters of pictures within the layer). A layer associated with a lower quality / bitstream size, such as layer N 531, is generally assigned a lower layer ID and is called a lower layer. Further, a layer associated with a higher quality / bitstream size, such as layer N+1 532, is generally assigned a higher layer ID and is called an upper layer.
[0081] Pictures 511 - 518 within different layers 531 - 532 are configured to be displayed as alternatives. As a specific example, if a smaller picture is desired, the decoder may decode and display picture 515 at the current display time, or if a larger picture is desired, the decoder may decode and display picture 511 at the current display time. Thus, pictures 511 - 514 of the upper layer N + 1 532 contain substantially the same image data as the corresponding pictures 515 - 518 of the lower layer N 531 (despite the difference in picture size). Specifically, picture 511 contains substantially the same image data as picture 515, picture 512 contains substantially the same image data as picture 516, and so on.
[0082] Pictures 511 to 518 can be coded by referring to other pictures 511 to 518 within the same layer N531 or N+1 532. When a picture is coded by referring to other pictures within the same layer, an inter prediction 523 occurs. The inter prediction 523 is indicated by a solid arrow. For example, picture 513 may be coded by using an inter prediction 523 that uses one or two of pictures 511, 512, and / or 514 within layer N+1 532 as references. One picture is referred to for uni-directional inter prediction and / or two pictures are referred to for bi-directional inter prediction. Further, picture 517 may be coded by using an inter prediction 523 that uses one or two of pictures 515, 516, and / or 518 within layer N531 as references. One picture is referred to for uni-directional inter prediction and / or two pictures are referred to for bi-directional inter prediction. When performing the inter prediction 523, when a picture is used as a reference to other pictures within the same layer, that picture may be called a reference picture. For example, picture 512 may be a reference picture used to code picture 513 according to the inter prediction 523. The inter prediction 523 can also be called an intra-layer prediction in a multi-layer context. Thus, the inter prediction 523 is a mechanism for coding samples of a current picture by referring to the indicated samples in a reference picture different from the current picture when the reference picture and the current picture are in the same layer.
[0083] Pictures 511 to 518 can also be coded by referring to other pictures 511 to 518 in different layers. This process is known as inter-layer prediction 521 and is indicated by the dashed arrows. Inter-layer prediction 521 is a mechanism for coding samples of the current picture by referring to the indicated samples in the reference picture when the current picture and the reference picture are in different layers and thus have different layer IDs. For example, a picture in the lower layer N531 can be used as a reference picture for coding the corresponding picture in the upper layer N+1 532. As a specific example, picture 511 can be coded by referring to picture 515 according to inter-layer prediction 521. In such a case, picture 515 is used as an inter-layer reference picture. An inter-layer reference picture is a reference picture used for inter-layer prediction 521. In most cases, inter-layer prediction 521 is restricted such that the current picture, such as picture 511, can only use inter-layer reference pictures that are in the same AU and in a lower layer, such as picture 515. An AU is a set of pictures associated with a specific output time in a video sequence, and thus, an AU can contain one picture per layer. If multiple layers (e.g., three or more) are available, inter-layer prediction 521 can encode / decode the current picture based on multiple inter-layer reference pictures at a level lower than the current picture.
[0084] The video encoder can use the multi-layer video sequence 500 to encode pictures 511-518 through many different combinations and / or substitutions of inter prediction 523 and inter-layer prediction 521. For example, picture 515 can be encoded according to intra prediction. Then, by using picture 515 as a reference picture, pictures 516-518 can be encoded according to inter prediction 523. Further, by using picture 515 as an inter-layer reference picture, picture 511 may be encoded according to inter-layer prediction 521. Then, by using picture 511 as a reference picture, pictures 512-514 can be encoded according to inter prediction 523. Therefore, the reference picture can function as both a single-layer reference picture and an inter-layer reference picture of different coding mechanisms. By encoding the upper-layer N+1 532 picture based on the lower-layer N 531 picture, the upper-layer N+1 532 can avoid using intra prediction that has a much lower coding efficiency than inter prediction 523 and inter-layer prediction 521. Therefore, the poor coding efficiency of intra prediction can be limited to the minimum / lowest quality pictures, and thus, the coding of the minimum amount of video data can be limited. The pictures used as reference pictures and / or inter-layer reference pictures can be shown in the entries of the reference picture list included in the reference picture list structure.
[0085] To perform such an operation, layers such as layer N531 and layer N+1 532 may be included in OLS525. OLS525 is a set of layers in which one or more layers are designated as output layers. An output layer is a layer designated for output (e.g., to a display). For example, layer N531 may be included only to support inter-layer prediction 521 and may never be output. In such a case, layer N+1 532 is decoded and output based on layer N531. In such a case, OLS525 includes layer N+1 532 as the output layer. OLS525 may include many layers in different combinations. For example, the output layer within OLS525 can be coded according to inter-layer prediction 521 based on one, two, or many lower layers. Further, OLS525 may include more than two output layers. Thus, OLS525 may include one or more output layers and any support layers necessary to reconstruct the output layers. Multi-layer video sequence 500 can be coded by using many different OLS525s, each using a different combination of layers.
[0086] As a specific example, inter-layer prediction 521 can be used to support scalability. For example, a video can be encoded into a base layer such as layer N 531 and several enhancement layers such as layer N+1 532, layer N+2, layer N+3, etc., which are encoded according to the inter-layer prediction 521. Due to several scalable characteristics such as resolution, frame rate, picture size, etc., a video sequence can be encoded. Then, an OLS 525 can be created for each acceptable characteristic. For example, the OLS 525 for the first resolution may only include layer N 531, the OLS 525 for the second resolution may include layer N 531 and layer N+1 532, and the OLS for the third resolution may include layer N 531, layer N+1 532, layer N+2, etc. In this way, the OLS 525 can be transmitted to enable the decoder to decode which version of the multi-layer video sequence 500 is desired based on network conditions, hardware constraints, etc.
[0087] FIG. 6 is a schematic diagram showing an exemplary multi-view sequence 600 including simulcast layers 631, 632, 633, and 634 for use in multi-view. The multi-view sequence 600 is a type of multi-layer video sequence 500. Therefore, the multi-view sequence 600 can be encoded by an encoder such as the codec system 200 and / or the encoder 300, and can be decoded by a decoder such as the codec system 200 and / or the decoder 400, for example, according to method 100.
[0088] Multi-view video is also called stereoscopic video. In multi-view, the video sequence is simultaneously captured from multiple camera angles into a single video stream. For example, a video can be captured using a pair of cameras that are spatially offset. Each camera captures the video from a different angle. As a result, a pair of views of the same subject are obtained. The first view can be presented to the user's right eye and the second view can be presented to the user's left eye. For example, this can be achieved by using a head-mounted display (HMD) that includes a separate left-eye display and a separate right-eye display. Displaying a pair of streams of the same subject from different angles creates an impression of visual depth and thus creates a 3D viewing experience.
[0089] To implement multi-view, a video can be encoded into multiple OLSs such as OLS625 and OLS626 similar to OLS525. Each view is encoded into layers such as layers 631, 632, 633, and 634 which can be similar to layer N531. As a specific example, the right-eye view can be encoded into layer 631 and the left-eye view can be encoded into layer 632. Then, layers 631 and 632 can be included in OLS625. In this way, OLS625 can be transmitted to the decoder with layers 631 and 632 marked as output layers. Then, the decoder can decode and display both layers 631 and 632. Therefore, OLS625 provides sufficient data to enable the representation of multi-view video. Similar to other types of video, multi-view video may be encoded into several representations to enable different display devices, different network conditions, etc. Therefore, OLS626 includes video encoded to achieve different characteristics but is substantially the same as OLS625. For example, layer 633 may be substantially the same as layer 631 and layer 634 may be substantially the same as layer 632. However, layers 633 and 634 may have different characteristics from layers 631 and 632. As a specific example, layers 633 and 634 may be encoded with a different resolution, frame rate, screen size, etc. from layers 631 and 632. As a specific example, if a first picture resolution is desired, OLS625 can be transmitted to the decoder, and if a second picture resolution is desired, OLS626 can be transmitted to the decoder.
[0090] In some cases, scalability is not used. Layers that do not use inter-layer prediction are called simulcast layers. A simulcast layer can be fully decoded without referring to other layers. For example, the shown layers 631 - 634 do not depend on any reference layers and are thus all simulcast layers. This configuration can cause errors in some video coding systems.
[0091] For example, some video coding systems can be configured to infer that each OLS contains a single layer when all layers are simulcast. In some cases, such an inference is valid. For example, when scalability is not used for standard video, the system can assume that each simulcast layer can be displayed without any other layer, and thus, the OLS is assumed to contain only one layer. This inference can prevent multi-view from operating properly. As shown, OLSs 625 and 626 each contain two layers 631 and 632, and layers 633 and 634, respectively. In such a case, the decoder is unsure which layer to decode, and since only one layer is expected, it may decode both layers and not display them.
[0092] The present disclosure addresses this problem by using each_layer_is_an_ols_flag in a bitstream. Specifically, when all layers 631 - 634 are simulcast, as indicated by vps_all_independent_layers_flag, each_layer_is_an_ols_flag is signaled. each_layer_is_an_ols_flag indicates whether each OLS contains a single layer or whether any OLS, such as OLSs 625 and 626, contains two or more layers. This enables proper decoding of the multi-view sequence 600. Further, ols_mode_idc can be set to indicate that information regarding the number of OLSs 625 - 626 and layers 631 - 634 should be explicitly signaled (e.g., an indication of which of layers 631 - 634 are output layers). These flags provide sufficient information for a decoder to correctly decode and display OLSs 625 and / or 626 using multi-view. Note that each_layer_is_an_ols_flag, vps_all_independent_layers_flag, and ols_mode_idc are named based on the nomenclature used in the VVC standardization. For the sake of consistency and clarity of the description, such names are included here. However, such syntax elements may be referred to by other names without departing from the scope of the present disclosure.
[0093] FIG. 7 is a schematic diagram showing an exemplary bitstream 700 including an OLS having simulcast layers for use in multi-view. For example, bitstream 700 can be generated by codec system 200 and / or encoder 300 for decoding by codec system 200 and / or decoder 400 according to method 100. Further, bitstream 700 can include an encoded multi-layer video sequence 500 and / or multi-view sequence 600.
[0094] The bitstream 700 includes a VPS 711, one or more sequence parameter sets (SPSs) 713, a plurality of picture parameter sets (PPSs) 715, a plurality of slice headers 717, and picture data 720. The VPS 711 includes data regarding the entire bitstream 700. For example, the VPS 711 may include data-related OLSs, layers, and / or sub-layers used in the bitstream 700. The SPS 713 includes sequence data common to all pictures within the coded video sequence included in the bitstream 700. For example, each layer may include one or more coded video sequences, and each coded video sequence may refer to the SPS 713 for corresponding parameters. The parameters within the SPS 713 can include picture sizing, bit depth, coding tool parameters, bitrate limits, and the like. Each sequence refers to the SPS 713, but note that in some examples, a single SPS 713 can include data for multiple sequences. The PPS 715 includes parameters applied to the entire picture. Thus, each picture within the video sequence may refer to the PPS 715. Each picture refers to the PPS 715, but note that in some examples, a single PPS 715 can include data for multiple pictures. For example, multiple similar pictures may be coded according to similar parameters. In such cases, a single PPS 715 may include data for such similar pictures. The PPS 715 can indicate coding tools available for slices, quantization parameters, offsets, etc. within the corresponding picture.
[0095] Slice header 717 contains parameters specific to each slice 727 within picture 725. Thus, there may be one slice header 717 for each slice 727 in a video sequence. The slice header 717 may include slice type information, POC, reference picture list, prediction weight, tile entry point, deblocking parameters, etc. Note that in some examples, bitstream 700 may also include a picture header, which is a syntax structure containing parameters applicable to all slices 727 within a single picture. For this reason, picture header and slice header 717 may be used interchangeably in some contexts. For example, some parameters may be moved between slice header 717 and picture header depending on whether such parameters are common to all slices 727 within picture 725.
[0096] Image data 720 includes video data encoded according to inter prediction and / or intra prediction, as well as corresponding transformed and quantized residual data. For example, image data 720 may include layer 723 of picture 725. Layer 723 may be compiled into OLS721. OLS721 may be substantially similar to OLS525, 625, and / or 626. Specifically, OLS721 is a set of layer 723s in which one or more layer 723s are designated as output layers. For example, bitstream 700 may be encoded to include several OLS721s with video encoded at different resolutions, frame rates, picture 725 sizes, etc. Upon request by the decoder, the sub-bitstream extraction process can remove all but the requested OLS721 from bitstream 700. The encoder can then send to the decoder a bitstream 700 containing only the requested OLS721, and thus only the video that meets the requested criteria.
[0097] Layer 723 may be substantially similar to layer N531, layer N+1 532, and / or layers 631, 632, 633, and / or 634. Layer 723 is generally a set of encoded pictures 725. Layer 723, when decoded, may be formally defined as a set of VCL NAL units that share specified characteristics (e.g., common resolution, frame rate, image size, etc.). Layer 723 also includes associated non-VCL NAL units to support the decoding of VCL NAL units. The VCL NAL units of layer 723 may share a particular value of nuh_layer_id. Layer 723 may be a simulcast layer encoded without inter-layer prediction, or layer 723 encoded according to inter-layer prediction as described for FIGS. 6 and 5 respectively.
[0098] Picture 725 is an array of luma samples and / or an array of chroma samples that create a frame or fields thereof. For example, picture 725 may be an encoded image that can be output for display or used to support the encoding of other pictures 725 for output. Picture 725 may include a set of VCL NAL units. Picture 725 includes one or more slices 727. Slice 727 may be defined as a single NAL unit, in particular an integral number of complete tiles of picture 725 exclusively included in a VCL NAL unit, or an integral number of consecutive complete coding tree unit (CTU) rows (e.g., within a tile). Slice 727 is further divided into CTUs and / or coding tree blocks (CTBs). A CTU is a predefined size group of samples that can be partitioned by a coding tree. A CTB is a subset of a CTU and includes the luma or chroma component of the CTU. CTU / CTB is further divided into coding blocks based on a coding tree. The coding blocks can then be encoded / decoded according to a prediction mechanism.
[0099] The present disclosure includes a mechanism that enables a video coding system to properly decode a multi-view video, such as a multi-view sequence 600, when all layers 723 within a video are simulcast and inter-layer prediction is not used. For example, the VPS 711 can include various data for indicating to the decoder that all layers 723 are simulcast and that the OLS 721 includes two or more layers 723. The vps_all_independent_layers_flag 731 can be included in the bitstream 700 within the VPS 711. The vps_all_independent_layers_flag 731 is a syntax element that signals whether inter-layer prediction is used to code any of the layers 723 within the bitstream 700. For example, the vps_all_independent_layers_flag 731 can be set to 1 when none of the layers 723 use inter-layer prediction and thus all are simulcast. In another example, the vps_all_independent_layers_flag 731 can be set to 0 to indicate that at least one of the layers 723 uses inter-layer prediction. When the vps_all_independent_layers_flag 731 is set to 1 to indicate that all layers 723 are simulcast, the each_layer_is_an_ols_flag 733 is signaled in the VPS 711. The each_layer_is_an_ols_flag 733 is a syntax element that signals whether each OLS 721 within the bitstream 700 includes a single layer 723. For example, each OLS 721 can typically include a single simulcast layer. However, when a multi-view video is encoded into the bitstream 700, one or more OLS 721s can include two simulcast layers.Therefore, each_layer_is_an_ols_flag733 can be set (e.g., to 1) to specify that each OLS721 includes a single layer 723, or can be set (e.g., to 0) to specify that at least one OLS721 includes two or more layers 723 to support multi-views. Therefore, vps_all_independent_layers_flag731 and each_layer_is_an_ols_flag733 can be used to support multi-view applications.
[0100] Furthermore, VPS711 may include ols_mode_idc735. ols_mode_idc735 is a syntax element indicating information about the number of OLSs721, the layers 723 of the OLSs721, and the output layers of the OLSs721. The output layer 723 is any layer specified for decoder output, in contrast to being used only for reference-based coding. ols_mode_idc735 can be set to 0 or 1 for coding other types of video. ols_mode_idc735 can be set to 2 to support multi-views. For example, when vps_all_independent_layers_flag731 is set to 1 (indicating simulcast layers), each_layer_is_an_ols_flag733 is set to 0, and at least one OLS721 includes two or more layers 723, ols_mode_idc735 can be set to 2. When ols_mode_idc735 is set to 2, information about the number of OLSs721 and the number of layers 723 and / or output layers included in each OLS721 is explicitly signaled.
[0101] VPS711 may also include vps_max_layers_minus1 737. vps_max_layers_minus1 737 is a syntax element that signals the number of layers 723 specified by VPS711, and thus the maximum number of layers 723 permitted in the corresponding coded video sequence within bitstream 700. VPS711 may also include num_output_layer_sets_minus1 739. num_output_layer_sets_minus1 739 is a syntax element that specifies the total number of OLSs 721 specified by VPS711. In one example, vps_max_layers_minus1 737 and num_output_layer_sets_minus1 739 may be signaled in VPS711 when ols_mode_idc735 is set to 2. Thereby, when the video includes multiple views, the number of OLSs 721 and the number of layers 723 are signaled. Specifically, vps_max_layers_minus1 737 and num_output_layer_sets_minus1 739 may be signaled when vps_all_independent_layers_flag731 is set to 1 (indicating simulcast layers), each_layer_is_an_ols_flag733 is set to 0, and at least one OLS 721 includes two or more layers 723. The decoder can then use this information to correctly decode the OLSs 721 that include multi-view video. This approach supports coding efficiency while correcting errors. In particular, multiple views are supported. However, when multiple views are not used, the number of OLSs 721 and / or the number of layers 723 can still be inferred from bitstream 700 and may be omitted. Thus, the disclosed mechanism improves the functionality of the encoder and / or decoder by enabling such devices to properly code multi-view video.Furthermore, the disclosed mechanism can maintain a reduced bitstream size and thus reduce the utilization of processor, memory, and / or network resources in both the encoder and the decoder.
[0102] Here, the above-mentioned information will be described in more detail below. Hierarchical video coding is also referred to as scalable video coding or video coding with scalability. Scalability in video coding can be supported by using multi-layer coding techniques. A multi-layer bitstream includes a base layer (BL) and one or more enhancement layers (EL). Examples of scalability include spatial scalability, quality / signal-to-noise ratio (SNR) scalability, multi-view scalability, frame rate scalability, and the like. When multi-layer coding techniques are used, a picture or a part thereof may be coded (intra prediction) without using a reference picture, may be coded (inter prediction) by referring to a reference picture in the same layer, and / or may be coded (inter-layer prediction) by referring to a reference picture in another layer. The reference picture used for inter-layer prediction of the current picture is called an inter-layer reference picture (ILRP). FIG. 5 shows an example of multi-layer coding for spatial scalability where pictures of different layers have different resolutions.
[0103] Some video coding families provide support for scalability in profiles separated from the profile for single layer coding. Scalable Video Coding (SVC) is a scalable extension of Advanced Video Coding (AVC) that provides support for spatial, temporal, and quality scalability. In the case of SVC, for each macroblock (MB) in an EL picture, a flag is signaled indicating whether the EL MB is predicted using a collocated block from the lower layer. Prediction from collocated blocks can include texture, motion vectors, and / or coding modes. Implementations of SVC cannot directly reuse unmodified AVC implementations in their designs. The SVC EL macroblock syntax and decoding process are different from the AVC syntax and decoding process.
[0104] Scalable High Efficiency Video Coding (SHVC) is an extension of HEVC that supports spatial scalability and quality scalability. Multi-View HEVC (MV-HEVC) is an extension of HEVC that supports multi-view scalability. 3D HEVC (3D-HEVC) is an extension of HEVC that provides support for 3D video coding that is more advanced and efficient than MV-HEVC. Temporal scalability may be included as an integral part of a single-layer HEVC codec. In the multi-layer extension of HEVC, the decoded pictures used for inter-layer prediction are obtained only from the same access unit and are treated as long-term reference pictures (LTRPs). Such pictures have reference indexes within the reference picture list assigned along with other temporal reference pictures within the current layer. Inter-layer prediction (ILP) is achieved at the prediction unit (PU) level by setting the value of the reference index to reference the inter-layer reference pictures within the reference picture list. Spatial scalability resamples a reference picture or a portion thereof when the ILRP has a different spatial resolution than the current picture being encoded or decoded. Resampling of the reference picture can be achieved either at the picture level or at the coding block level.
[0105] VVC can also support hierarchical video coding. The VVC bitstream can include multiple layers. All layers can be independent of each other. For example, each layer can be coded without using inter-layer prediction. In this case, the layer is also called a simulcast layer. In some cases, some of the layers are coded using ILP. The flags in the VPS can indicate whether a layer is a simulcast layer or whether some layers use ILP. When some layers use ILP, the layer dependencies between the layers are also signaled in the VPS. Different from SHVC and MV-HEVC, VVC does not need to specify an OLS. The OLS includes a specified set of layers, and one or more layers within the set of layers are specified to be output layers. The output layer is the layer of the OLS to be output. In some implementations of VVC, when a layer is a simulcast layer, only one layer can be selected for decoding and output. In some implementations of VVC, the entire bitstream including all layers is specified to be decoded when any layer uses ILP. Further, some of the layers are specified as output layers. The output layer may be indicated as only the top layer, all layers, or the top layer plus a set of the indicated lower layers.
[0106] The foregoing aspects include some problems. For example, when a layer is a simulcast layer, only one layer may be selected for decoding and output. However, this approach does not support cases where two or more layers can be decoded and output, such as in a multi-view application.
[0107] Generally, the present disclosure describes an approach for supporting operating points having multiple output layers of simulcast layers. The description of the technique is based on VVC by ITU-T and JVET of ISO / IEC. However, this technique is also applicable to hierarchical video coding based on other video codec specifications.
[0108] One or more of the above problems can be solved as follows. Specifically, the present disclosure includes a simple and efficient method for supporting decoding and output of multiple layers of a bitstream including simulcast layers, as summarized below. The VPS may include an indication of whether each layer is an OLS. When each layer is an OLS, only one layer can be decoded and output. In this case, it is inferred that the number of OLSs is equal to the number of layers. Further, each OLS includes one layer, and that layer is the output layer. Otherwise, the number of OLSs is explicitly signaled. For each OLS except the 0th OLS, the layers included in the OLS may be explicitly signaled. Further, each layer within each OLS can be inferred to be the output layer. The 0th OLS includes only the lowest layer that is the output layer.
[0109] An exemplary implementation of the foregoing mechanism is as follows. An exemplary video parameter set syntax is as follows.
[0110] [Table 1A] [Table 1B]
[0111] An exemplary video parameter set semantics is as follows. A VPS RBSP is available for the decoding process before being referenced and is included in at least one access unit having a TemporalId equal to 0 or provided via an external mechanism. The VPS NAL unit containing the VPS RBSP should have a nuh_layer_id equal to vps_layer_id[0]. All VPS NAL units having a specific value of vps_video_parameter_set_id within a CVS should have the same content. vps_video_parameter_set_id provides an identifier for the VPS for reference by other syntax elements. vps_max_layers_minus1 plus 1 specifies the maximum allowable number of layers in each CVS that references the VPS. vps_max_sub_layers_minus1 plus 1 specifies the maximum number of temporal sub-layers that may exist in each CVS that references the VPS. The value of vps_max_sub_layers_minus1 should be in the range from 0 to 6 (inclusive).
[0112] The vps_all_independent_layers_flag can be set equal to 1 to specify that all layers within the CVS are coded independently without using inter-layer prediction. The vps_all_independent_layers_flag can be set equal to 0 to specify that one or more of the layers within the CVS can use inter-layer prediction. When it does not exist, the value of the vps_all_independent_layers_flag is inferred to be equal to 0. When the vps_all_independent_layers_flag is equal to 1, the value of vps_independent_layer_flag[i] is inferred to be equal to 1. When the vps_all_independent_layers_flag is equal to 0, the value of vps_independent_layer_flag[0] is inferred to be equal to 0. vps_layer_id[i] specifies the nuh_layer_id value of the i-th layer. For any two non-negative integer values of m and n, when m is less than n, the value of vps_layer_id[m] should be less than vps_layer_id[n]. The vps_independent_layer_flag[i] can be set equal to 1 to specify that the layer with index i does not use inter-layer prediction. The vps_independent_layer_flag[i] can be set equal to 0 to specify that the layer with index i uses inter-layer prediction and vps_layer_dependency_flag[i] exists in the VPS. When it does not exist, the value of vps_independent_layer_flag[i] is inferred to be equal to 1.
[0113] vps_direct_dependency_flag[i][j] can be set equal to 0 to specify that the layer with index j is not a direct reference layer of the layer with index i. vps_direct_dependency_flag[i][j] can be set equal to 1 to specify that the layer with index j is a direct reference layer of the layer with index i. When vps_direct_dependency_flag[i][j] does not exist for i and j in the range from 0 to vps_max_layers_minus1 (inclusive), vps_direct_dependency_flag[i][j] is inferred to be equal to 0. The variable DirectDependentLayerIdx[i][j] that specifies the j-th direct dependent layer of the i-th layer is derived as follows. for( i = 1; i < vps_max_layers_minus1; i++ ) if( !vps_independent_layer_flag[ i ] ) for( j = i, k = 0; j >= 0; j-- ) if( vps_direct_dependency_flag[ i ][ j ] ) DirectDependentLayerIdx[ i ][ k++ ] = j
[0114] The variable GeneralLayerIdx[i] that specifies the layer index of the layer having nuh_layer_id equal to vps_layer_id[i] is derived as follows. for( i = 0; i <= vps_max_layers_minus1; i++ ) GeneralLayerIdx[ vps_layer_id[ i ] ] = i
[0115] each_layer_is_an_ols_flag can be set equal to 1 to specify that each output layer set contains only one layer and each layer itself within the bitstream is an output layer set where the single included layer is the only output layer. each_layer_is_an_ols_flag can be set equal to 0 to specify that the output layer set can contain two or more layers. When vps_max_layers_minus1 is equal to 0, the value of each_layer_is_an_ols_flag is inferred to be equal to 1. Otherwise, when vps_all_independent_layers_flag is equal to 0, the value of each_layer_is_an_ols_flag is inferred to be equal to 0.
[0116] ols_mode_idc can be set equal to 0 to specify that the total number of OLSs specified by the VPS is equal to vps_max_layers_minus1 + 1, the i-th OLS contains the layers with layer indices from 0 to i (inclusive), and for each OLS, only the topmost layer within the OLS is output. ols_mode_idc can be set equal to 1 to specify that the total number of OLSs specified by the VPS is equal to vps_max_layers_minus1 + 1, the i-th OLS contains the layers with layer indices from 0 to i (inclusive), and for each OLS, all layers within the OLS are output. ols_mode_idc can be set equal to 2 to specify that the total number of OLSs specified by the VPS is explicitly signaled and for each OLS, the set of the topmost layer and the explicitly signaled lower layers within the OLS are output. The value of ols_mode_idc should be in the range from 0 to 2 (inclusive). The value 3 of ols_mode_idc is reserved. When vps_all_independent_layers_flag is equal to 1 and each_layer_is_an_ols_flag is equal to 0, the value of ols_mode_idc is inferred to be equal to 2.
[0117] num_output_layer_sets_minus1 plus 1 specifies the total number of OLSs specified by the VPS when ols_mode_idc is equal to 2. The variable TotalNumOlss that specifies the total number of OLSs specified by the VPS is derived as follows. if( vps_max_layers_minus1 == 0 ) TotalNumOlss = 1 else if( each_layer_is_an_ols_flag || ols_mode_idc == 0 || ols_mode_idc == 1 ) TotalNumOlss = vps_max_layers_minus1 + 1 else if( ols_mode_idc == 2 ) TotalNumOlss = num_output_layer_sets_minus1 + 1
[0118] layer_included_flag[i][j] specifies whether the j-th layer (for example, the layer with nuh_layer_id equal to vps_layer_id[j]) is included in the i-th OLS when ols_mode_idc is equal to 2. layer_included_flag[i][j] can be set equal to 1 to specify that the j-th layer is included in the i-th OLS. layer_included_flag[i][j] can be set equal to 0 to specify that the j-th layer is not included in the i-th OLS.
[0119] The variable NumLayersInOls[i] that specifies the number of layers in the i-th OLS and the variable NumLayersInOls[i] that specifies the nuh_layer_id value of the j-th layer in the i-th OLS can be derived as follows. NumLayersInOls
[0000] = 1 LayerIdInOls
[0000]
[0000] = vps_layer_id
[0000] for( i = 1; i < TotalNumOlss; i++ ) { if( each_layer_is_an_ols_flag ) { NumLayersInOls[ i ] = 1 LayerIdInOls[ i ]
[0000] = vps_layer_id[ i ] } else if( ols_mode_idc == 0 || ols_mode_idc == 1 ) { NumLayersInOls[ i ] = i + 1 for( j = 0; j < NumLayersInOls[ i ]; j++ ) LayerIdInOls[ i ][ j ] = vps_layer_id[ j ] } else if( ols_mode_idc == 2 ) { for( k = 0, j = 0; k <= vps_max_layers_minus1; k++ ) if( layer_included_flag[ i ][ k ] ) LayerIdInOls[ i ][ j++ ] = vps_layer_id[ k ] NumLayersInOls[ i ] = j } }
[0120] The variable OlsLayeIdx[i][j] that specifies the OLS layer index of the layer having the nuh_layer_id equal to LayerIdInOls[i][j] is derived as follows. for( i = 0; i < TotalNumOlss; i++ ) for j = 0; j < NumLayersInOls[ i ]; j++ ) OlsLayeIdx[i][LayerIdInOls[i][j]] = j
[0121] The lowest layer within each OLS should be an independent layer. In other words, for each i within the range from 0 to TotalNumOlss - 1 (inclusive), the value of vps_independent_layer_flag[GeneralLayerIdx[LayerIdInOls[i][0]]] should be equal to 1. The top layer, for example, the layer with nuh_layer_id equal to vps_layer_id[vps_max_layers_minus1], should be included in at least one OLS specified by the VPS. In other words, for at least one i within the range from 0 to TotalNumOlss - 1 (inclusive), the value of LayerIdInOls[i][NumLayersInOls[i] - 1] should be equal to vps_layer_id[vps_max_layers_minus1].
[0122] vps_output_layer_flag[i][j] specifies whether the j-th layer of the i-th OLS is output when ols_mode_idc is 2. vps_output_layer_flag[i] can be set equal to 1 to specify that the j-th layer of the i-th OLS is output. vps_output_layer_flag[i] can be set equal to 0 to specify that the j-th layer of the i-th OLS is not output. When vps_all_independent_layers_flag is equal to 1 and each_layer_is_an_ols_flag is equal to 0, it can be inferred that the value of vps_output_layer_flag[i] is equal to 1.
[0123] The variable OutputLayerFlag[i][j] has a value of 1 to specify that the j-th layer of the i-th OLS is output and a value of 0 to specify that the j-th layer of the i-th OLS is not output, and can be derived as follows. for( i = 0; i < TotalNumOlss; i++ ) { OutputLayerFlag[ i ][ NumLayersInOls[ i ] - 1 ] = 1 for( j = 0; j < NumLayersInOls[ i ] - 1; j++ ) if( ols_mode_idc[ i ] == 0 ) OutputLayerFlag[ i ][ j ] = 0 else if( ols_mode_idc[ i ] == 1 ) OutputLayerFlag[ i ][ j ] = 1 else if( ols_mode_idc[ i ] == 2 ) OutputLayerFlag[ i ][ j ] = vps_output_layer_flag[ i ][ j ] }
[0124] Any layer within the OLS should be the output layer of the OLS or a (direct or indirect) reference layer to the output layer of the OLS. The 0th OLS contains only the lowest layer (e.g., the layer with nuh_layer_id equal to vps_layer_id[0]), and for the 0th OLS, only the included layer is output. vps_constraint_info_present_flag can be set equal to 1 to specify that the general_constraint_info() syntax structure is present in the VPS. vps_constraint_info_present_flag can be set equal to 0 to specify that the general_constraint_info() syntax structure is not present in the VPS. In a compliant bitstream, vps_reserved_zero_7bits should be equal to 0. Other values of vps_reserved_zero_7bits are reserved. The decoder should ignore the value of vps_reserved_zero_7bits.
[0125] The general_hrd_params_present_flag can be set equal to 1 to indicate that the syntax elements num_units_in_tick and time_scale, and the syntax structure general_hrd_parameters() are present within the SPS RBSP syntax structure. The general_hrd_params_present_flag can be set equal to 0 to indicate that the syntax elements num_units_in_tick and time_scale, and the syntax structure general_hrd_parameters() are not present within the SPS RBSP syntax structure. num_units_in_tick is the number of time units of a clock that operates at a frequency of time_scale Hertz (Hz) corresponding to a 1 increment of the clock tick counter (referred to as a clock tick). num_units_in_tick should be greater than 0. The clock tick in seconds is equal to the quotient of num_units_in_tick divided by time_scale. For example, when the picture rate of a video signal is 25 Hz, time_scale is equal to 27000000, num_units_in_tick is equal to 1080000, and as a result, the clock tick can be equal to 0.04 seconds.
[0126] The time_scale is the number of time units that pass in one second. For example, a time coordinate system that measures time using a 27 megahertz (MHz) clock has a time_scale of 27000000. The value of the time_scale should be greater than 0. The vps_extension_flag can be set equal to 0 to indicate that the vps_extension_data_flag syntax element does not exist in the VPS RBSP syntax structure. The vps_extension_flag can be set equal to 1 to indicate that there is a vps_extension_data_flag syntax element within the VPS RBSP syntax structure. The vps_extension_data_flag can have any value. The presence and value of the vps_extension_data_flag do not affect the decoder's compliance with the profile. A compliant decoder should ignore all vps_extension_data_flag syntax elements.
[0127] FIG. 8 is a schematic diagram showing an exemplary video coding device 800. The video coding device 800 is suitable for implementing the disclosed examples / embodiments described herein. The video coding device 800 includes a transceiver unit (Tx / Rx) 810 that includes a transmitter and / or receiver for communicating data upstream and / or downstream via a downstream port 820, an upstream port 850, and / or a network. The video coding device 800 also includes a processor 830 that includes a logic unit and / or a central processing unit (CPU) for processing data and a memory 832 for storing data. The video coding device 800 may also include wireless communication components coupled to the upstream port 850 and / or the downstream port 820 for communicating data via an electrical, electro-optical (OE) component, an electro-optical (EO) component, and / or an electrical, optical, or wireless communication network. The video coding device 800 may also include an input and / or output (I / O) device 860 for communicating data with a user. The I / O device 860 may include output devices such as a display for displaying video data, a speaker for outputting audio data, etc. The I / O device 860 may also include input devices such as a keyboard, a mouse, a trackball, etc., and / or corresponding interfaces for interacting with such output devices.
[0128] Processor 830 is implemented by hardware and software. Processor 830 can be implemented as one or more CPU chips, cores (e.g., as a multi-core processor), field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), and digital signal processors (DSPs). Processor 830 communicates with downstream port 820, Tx / Rx 810, upstream port 850, and memory 832. Processor 830 includes coding module 814. Coding module 814 implements the disclosed embodiments described herein, such as methods 100, 900, and 1000, which may use multi-layer video sequence 500, multi-view sequence 600, and / or bitstream 700. Coding module 814 may also implement any other method / mechanism described herein. Further, coding module 814 may implement codec system 200, encoder 300, and / or decoder 400. For example, coding module 814 can be used to code each_layer_is_an_ols_flag to indicate whether each OLS includes a single layer or at least one OLS includes multiple layers, and support multi-view when simulcast layers are used. Thus, coding module 814 causes video coding device 800 to provide additional functionality and / or coding efficiency when coding video data. Thus, coding module 814 not only improves the functionality of video coding device 800, but also addresses problems specific to the video coding field. Further, coding module 814 converts video coding device 800 to a different state. Alternatively, coding module 814 can be implemented as instructions stored in memory 832 and executed by processor 830 (e.g., as a computer program product stored on a non-transitory medium).
[0129] Memory 832 includes one or more memory types such as a disk, a tape drive, a solid state drive, a read-only memory (ROM), a random access memory (RAM), a flash memory, a ternary content addressable memory (TCAM), a static random access memory (SRAM), etc. Memory 832 may be used as an overflow data storage device, store a program when such a program is selected for execution, and store instructions and data read during program execution.
[0130] FIG. 9 is a flowchart showing an exemplary method 900 for encoding a video sequence in a bitstream 700, such as having an OLS of a simulcast layer for use in a multi-view such as a multi-view sequence 600. Method 900 may be used by an encoder such as codec system 200, encoder 300, and / or video coding device 800 when executing method 100.
[0131] Method 900 may begin when an encoder receives a video sequence and determines to encode the video sequence into a set of simulcast layers for use in a multi-view, for example, based on user input. In step 901, the encoder encodes a bitstream including coded pictures of one or more layers. For example, the layer may be a simulcast layer and may not be encoded according to inter-layer prediction. Further, the layer may be encoded to support multi-view video. Thus, the layer may be organized into an OLS in which one or more OLSs include two layers (for example, one layer for display to each eye of an end user).
[0132] In step 903, the encoder can encode the VPS into a bitstream. The VPS may include various syntax elements for indicating the layer / OLS configuration to the decoder for proper multi-view decoding and display. For example, the VPS may include a vps_all_independent_layers_flag that can be set to 1 to specify that all layers specified by the VPS are encoded independently without inter-layer prediction. Thus, when the vps_all_independent_layers_flag is set to 1, and thus all layers specified by the VPS are encoded independently without inter-layer prediction, the VPS may also include an each_layer_is_an_ols_flag. The each_layer_is_an_ols_flag can specify whether each OLS contains only one layer or at least one OLS contains multiple layers. For example, the each_layer_is_an_ols_flag can be set to 1 when each OLS contains only one layer and / or each layer is an OLS where the single included layer is the only output layer. Thus, the each_layer_is_an_ols_flag can be set to 1 when multi-view is not used. As another example, the each_layer_is_an_ols_flag can be set to 0 when at least one OLS contains two or more layers and thus specifies that the bitstream encoded in step 901 includes a multi-view video.
[0133] The VPS may also include an ols_mode_idc syntax element. For example, when each_layer_is_an_ols_flag is set to 0 and vps_all_independent_layers_flag is set to 1, ols_mode_idc may be set to 2. When ols_mode_idc is set to / equal to 2, the total number of OLSs is explicitly signaled in the VPS. Further, when ols_mode_idc is set to / equal to 2, the number of layers associated with each OLS and / or the number of output layers is explicitly signaled in the VPS. In a particular example, the vps_max_layers_minus1 syntax element may be included in the VPS to explicitly specify the number of layers specified by the VPS, and thus, the number of layers that can be included in the OLS may also be specified. In some examples, vps_all_independent_layers_flag may be signaled when vps_max_layers_minus1 is greater than 0. In other particular examples, when ols_mode_idc is equal to 2, num_output_layer_sets_minus1 may be included in the VPS. num_output_layer_sets_minus1 may specify the total number of OLSs specified by the VPS. Thus, vps_max_layers_minus1 and num_output_layer_sets_minus1 may be signaled in the VPS to indicate the number of layers and the number of OLSs, respectively, when such data is explicitly signaled (e.g., when each_layer_is_an_ols_flag is set to 0, vps_all_independent_layers_flag is set to 1, ols_mode_idc is set to 2, and / or is inferred to be equal to 2). As a specific example, when vps_all_independent_layers_flag is set to 1 and each_layer_is_an_ols_flag is set to 0, it can be inferred that ols_mode_idc is equal to 2.
[0134] In step 905, the bitstream is stored for communication to the decoder.
[0135] FIG. 10 is a flowchart showing an exemplary method 1000 for decoding a video sequence from a bitstream 700, such as a multi-view bitstream 600, including OLSs of simulcast layers for use in multi-view. Method 1000 may be used by a decoder, such as codec system 200, decoder 400, and / or video coding device 800 when executing method 100.
[0136] Method 1000 may begin when the decoder begins to receive a bitstream including OLSs of a simulcast multi-view layer, for example, as a result of method 900. In step 1001, the decoder receives the bitstream. The bitstream may include one or more OLSs and one or more layers. For example, the layer may be a simulcast layer and may not be coded according to inter-layer prediction. Further, the layer may be coded to support multi-view video. Thus, the layer may be organized into an OLS in which one or more OLSs include two layers (e.g., one layer for display to each eye of an end user).
[0137] The bitstream may also include a VPS. The VPS may include various syntax elements for indicating the layer / OLS configuration to the decoder for proper multi-view decoding and presentation. For example, the VPS may include a vps_all_independent_layers_flag that can be set to 1 to specify that all layers specified by the VPS are coded independently without inter-layer prediction. When the vps_all_independent_layers_flag is set to 1, thus, when all layers specified by the VPS are coded independently without inter-layer prediction, the VPS may also include an each_layer_is_an_ols_flag. The each_layer_is_an_ols_flag can specify whether an OLS includes multiple layers. For example, the each_layer_is_an_ols_flag can be set to 1 when each OLS includes only one layer and / or when each layer is an OLS where the single included layer is the only output layer. Thus, the each_layer_is_an_ols_flag can be set to 1 when multi-view is not used. As another example, when at least one OLS includes two or more layers and thus it is specified that the bitstream includes multi-view video, the each_layer_is_an_ols_flag can be set to 0.
[0138] The VPS may also include an ols_mode_idc syntax element. For example, when each_layer_is_an_ols_flag is set to 0 and vps_all_independent_layers_flag is set to 1, ols_mode_idc may be set equal to 2. When ols_mode_idc is set equal to 2, the total number of OLSs is signaled explicitly in the VPS. Further, when ols_mode_idc is set / equal to 2, the number of layers associated with each OLS and / or the number of output layers is signaled explicitly in the VPS. In a particular example, the vps_max_layers_minus1 syntax element may be included in the VPS to explicitly specify the number of layers specified by the VPS, and thus also the number of layers that can be included in the OLS. In some examples, vps_all_independent_layers_flag may be signaled when vps_max_layers_minus1 is greater than 0. In other particular examples, when ols_mode_idc is equal to 2, num_output_layer_sets_minus1 may be included in the VPS. num_output_layer_sets_minus1 may specify the total number of OLSs specified by the VPS. Thus, vps_max_layers_minus1 and num_output_layer_sets_minus1 may be signaled in the VPS to indicate the number of layers and the number of OLSs, respectively, when such data is signaled explicitly (e.g., when each_layer_is_an_ols_flag is set to 0, vps_all_independent_layers_flag is set to 1, ols_mode_idc is set to 2, and / or is inferred to be equal to 2). As a specific example, when vps_all_independent_layers_flag is set to 1 and each_layer_is_an_ols_flag is set to 0, ols_mode_idc can be inferred to be equal to 2.
[0139] In step 1003, based on each_layer_is_an_ols_flag in the VPS, the coded picture is decoded from the output layer of the OLS to generate a decoded picture. For example, the decoder may read vps_all_independent_layers_flag to determine that all layers are simulcast. The decoder may also read each_layer_is_an_ols_flag to determine that at least one OLS contains two or more layers. The decoder may also read ols_mode_idc to determine that the number of OLSs and the number of layers are explicitly signaled. Then, the decoder can determine the number of OLSs and the number of layers by reading num_output_layer_sets_minus1 and vps_max_layers_minus1 respectively. Then, the decoder can use this information to identify the correct multi-view layer position in the bitstream. The decoder can also identify the position of the correct coded picture from the layer. Then, the decoder can decode the picture to generate a decoded picture.
[0140] In step 1005, the decoder can transfer the decoded picture for display as part of the decoded video sequence.
[0141] FIG. 11 is a schematic diagram showing an exemplary system 1100 for encoding a video sequence, such as in bitstream 700, having OLSs of simulcast layers for use in multi-views such as multi-view sequence 600. System 1100 may be implemented by an encoder and decoder such as codec system 200, encoder 300, decoder 400, and / or video coding device 800. Further, system 1100 may use multi-layer video sequence 500. Additionally, system 1100 may be used when implementing methods 100, 900, and / or 1000.
[0142] System 1100 includes a video encoder 1102. The video encoder 1102 includes an encoding module 1105 for encoding a bitstream including one or more layers of coded pictures. The encoding module 1105 is further for encoding a VPS including each_layer_is_an_ols_flag into the bitstream when all layers specified by the VPS are independently encoded without inter-layer prediction, and each_layer_is_an_ols_flag specifies whether each OLS includes only one layer. The video encoder 1102 further includes a storage module 1106 for storing the bitstream for communication to a decoder. The video encoder 1102 further includes a transmission module 1107 for transmitting the bitstream towards a video decoder 1110. The video encoder 1102 may be further configured to execute any of the steps of method 900.
[0143] System 1100 also includes a video decoder 1110. The video decoder 1110 is a receiving module 1111 for receiving a bitstream including coded pictures of one or more layers and a VPS, and when all layers specified by the VPS are independently encoded without inter-layer prediction, the VPS includes each_layer_is_an_ols_flag, and each_layer_is_an_ols_flag specifies whether each OLS includes only one layer, and includes the receiving module 1111. The video decoder 1110 further includes a decoding module 1113 for decoding the coded pictures from the output layer of the OLS based on each_layer_is_an_ols_flag in the VPS to generate decoded pictures. The video decoder 1110 further includes a transfer module 1115 for transferring the decoded pictures for display as part of the decoded video sequence. The video decoder 1110 may be further configured to execute any of the steps of method 1000.
[0144] When there are no intervening components between the first component and the second component except for a line, trace, or other medium, the first component is directly coupled to the second component. When there are intervening components other than a line, trace, or other medium between the first component and the second component, the first component is indirectly coupled to the second component. The term "coupled" and its variants include both those that are directly coupled and those that are indirectly coupled. The use of the term "about" means a range that includes ±10% of the subsequent number, unless otherwise specified.
[0145] The steps of the exemplary methods described herein need not necessarily be performed in the order described, and it should also be understood that the order of such method steps is merely exemplary. Similarly, in methods consistent with various embodiments of the present disclosure, additional steps may be included in such methods, or specific steps may be omitted or combined.
[0146] Although several embodiments are provided in the present disclosure, it will be understood that the disclosed systems and methods may be implemented in many other specific forms without departing from the spirit or scope of the present disclosure. This example is to be considered illustrative and not restrictive, and the invention is not limited to the details shown herein. For example, various elements or components may be combined or integrated into other systems, or some functions may be omitted or not implemented.
[0147] In addition, the techniques, systems, subsystems, and methods described individually or separately in various embodiments may be combined with or integrated into other systems, components, techniques, or methods without departing from the scope of the present disclosure. Other examples of changes, substitutions, and modifications are ascertainable by those skilled in the art and may be made without departing from the spirit and scope disclosed herein.
Description of Reference Numerals
[0148] 100 Method 200 Codec System 201 Partitioned Video Signal 211 General Coder Control Component 213 Transformation Scaling and Quantization Component 215 Intra-Picture Estimation Component 217 Intra-Picture Prediction Component 219 Motion Compensation Component 221 Motion Estimation Component 223 Decoded Picture Buffer Component 225 In-Loop Filter Component 227 Filter Control Analysis Component 229 Scaling and Inverse Transformation Component 231 Header Formatting and Context-Adaptive Binary Arithmetic Coding (CABAC) Component 300 Video Encoder 301 Partitioned Video Signal 313 Transformation and Quantization Component 317 Intra-Picture Prediction Component 321 Motion Compensation Component 323 Decoded Picture Buffer Component 325 In-Loop Filter Component 329 Inverse Transformation and Quantization Component 331 Entropy Coding Component 400 Video Decoder 417 Intra-Picture Prediction Component 421 Motion Compensation Component 423 Decoded Picture Buffer Component 425 In-Loop Filter Component 429 Inverse Transformation and Quantization Component 433 Entropy Decoding Component 500 Multilayer Video Sequence 511 Picture 512 Picture 513 Picture 514 Picture 515 Picture 516 Picture 517 Picture 518 Picture 521 Inter-layer Prediction 523 Inter Prediction 525 OLS 531 Layer N 532 Layer N+1 600 Multiview Sequence 625 OLS 626 OLS 631 Simulcast Layer 632 Simulcast Layer 633 Simulcast Layer 634 Simulcast Layer 700 Bitstream 711 VPS 713 Sequence Parameter Set (SPS) 715 Picture Parameter Set (PPS) 717 Slice Header 720 Image Data 721 OLS 723 Layer 725 Picture 727 Slice 731 vps_all_independent_layers_flag 733 each_layer_is_an_ols_flag 735 ols_mode_idc 737 vps_max_layers_minus1 739 num_output_layer_sets_minus1 800 Video Coding Device 810 Transceiver Unit (Tx / Rx) 814 Coding Module 820 Downstream Port 830 Processor 832 Memory 850 Upstream Port 860 I / O Device 900 Method 1000 Method 1100 System 1102 Video Encoder 1105 Encoding Module 1106 Memory Module 1107 Transmission Module 1110 Video Decoder 1111 Reception Module 1113 Decoding Module 1115 Transfer Module
Claims
Claim 1 A method executed in a decoder, the method comprising: receiving, by a receiver of the decoder, a bitstream including coded pictures of one or more layers and a video parameter set (VPS), wherein when all layers specified by the VPS are coded independently without inter-layer prediction, the VPS includes an "each_layer_is_an_ols_flag" flag, and the each_layer_is_an_ols_flag specifies whether each OLS includes only one layer; decoding, by a processor of the decoder, the coded pictures from the output layers of the OLS based on the each_layer_is_an_ols_flag in the VPS to generate decoded pictures; transferring, by the processor, the decoded pictures for display as part of the decoded video sequence and a method including this. Claim 2 The method according to claim 1, wherein the each_layer_is_an_ols_flag is set to 1 when each OLS includes only one layer and each layer is specified as the only output layer within each OLS. Claim 3 The method according to any one of claims 1 to 2, wherein the each_layer_is_an_ols_flag is set to 0 when it is specified that at least one OLS includes a plurality of layers. Claim 4 When the OLS mode identification code (ols_mode_idc) is equal to 2, the total number of OLSs is explicitly signaled, the layers associated with the OLSs are explicitly signaled, the "vps_all_independent_layers_flag" flag is set to 1, and when the each_layer_is_an_ols_flag is set to 0, it is inferred that the ols_mode_idc is equal to 2. The method according to any one of claims 1 to 3. Claim 5 The method according to any one of claims 1 to 4, wherein the VPS includes a vps_all_independent_layers_flag set to 1 to specify that all layers specified by the VPS are independently coded without inter-layer prediction.
6. The method according to any one of claims 1 to 5, wherein the VPS includes a VPS maximum layers minus 1 (vps_max_layers_minus1) syntax element that specifies the number of layers specified by the VPS, and when vps_max_layers_minus1 is greater than 0, the vps_all_independent_layers_flag is signaled.
7. The method according to any one of claims 1 to 6, wherein the VPS includes a "number of output layer sets minus 1" (num_output_layer_sets_minus1) that specifies the total number of OLSs specified by the VPS when ols_mode_idc is equal to 2.
8. A method executed in an encoder, the method comprising: encoding, by a processor of the encoder, a bitstream including coded pictures of one or more layers; encoding, by the processor, a video parameter set (VPS) into the bitstream, wherein the VPS includes an "each layer is an output layer set (OLS) flag" (each_layer_is_an_ols_flag) when all layers specified by the VPS are independently coded without inter-layer prediction, and the each_layer_is_an_ols_flag specifies whether each OLS includes only one layer; storing, by a memory coupled to the processor, the bitstream for communication to a decoder and the method includes.
9. The method according to claim 8, wherein the each_layer_is_an_ols_flag is set to 1 when each OLS includes only one layer and each layer is the only output layer within each OLS.
10. The method according to any one of claims 8 to 10, wherein when specifying that at least one OLS includes a plurality of layers, the each_layer_is_an_ols_flag is set to 0.
11. The method according to any one of claims 8 to 10, wherein when the OLS mode identification code (ols_mode_idc) is equal to 2, the total number of OLSs is explicitly signaled, the layers associated with the OLSs are explicitly signaled, the "all independent layers of the VPS" flag (vps_all_independent_layers_flag) is set to 1, and when the each_layer_is_an_ols_flag is set to 0, it is inferred that the ols_mode_idc is equal to 2.
12. The method according to any one of claims 8 to 11, wherein the VPS includes the vps_all_independent_layers_flag set to 1 to specify that all layers specified by the VPS are coded independently without inter-layer prediction.
13. The method according to any one of claims 8 to 12, wherein the VPS includes a VPS maximum layers minus 1 (vps_max_layers_minus1) syntax element that specifies the number of layers specified by the VPS, and when vps_max_layers_minus1 is greater than 0, the vps_all_independent_layers_flag is signaled.
14. The method according to any one of claims 8 to 13, wherein the VPS includes a "number of output layer sets minus 1" (num_output_layer_sets_minus1) that specifies the total number of OLSs specified by the VPS when the ols_mode_idc is equal to 2.
15. A video coding device, comprising a processor, a receiver coupled to the processor, a memory coupled to the processor, and a transmitter coupled to the processor, wherein the processor, receiver, memory, and transmitter are configured to execute the method according to any one of claims 1 to 14. Video coding device.
16. A non-transitory computer-readable medium including a computer program product for use by a video coding device, wherein the computer program product, when executed by a processor, causes the video coding device to execute the method according to any one of claims 1 to 14, and the non-transitory computer-readable medium includes computer-executable instructions stored in the non-transitory computer-readable medium.
17. A decoder, receiving means for receiving a bitstream including one or more layers of coded pictures and a video parameter set (VPS), wherein when all layers specified by the VPS are independently coded without inter-layer prediction, the VPS includes a "each_layer_is_an_ols_flag" flag, and the each_layer_is_an_ols_flag specifies whether each OLS includes only one layer; decoding means for decoding a coded picture from an output layer of the OLS based on the each_layer_is_an_ols_flag in the VPS to generate a decoded picture; transfer means for transferring the decoded picture for display as part of the decoded video sequence and including a decoder.
18. The decoder according to claim 17, wherein the decoder is further configured to execute the method according to any one of claims 1 to 7.
19. An encoder, encoding a bitstream including one or more layers of coded pictures, encoding means for encoding the VPS into the bitstream, including a "each_layer_is_an_ols_flag" flag when all layers specified by the video parameter set (VPS) are independently coded without inter-layer prediction, and the each_layer_is_an_ols_flag specifies whether each OLS includes only one layer; storage means for storing the bitstream for communication with a decoder and including an encoder.
20. The encoder according to claim 19, further configured to execute the method according to any one of claims 8 to 14.
Citation Information
Patent Citations
Image decoder and image encoder
JP2015195543A
Method and an apparatus and a computer program for encoding media content
US20170347026A1
Image decoding device and image decoding method
WO2015137432A1