Disallowing unnecessary layers in multi-layer video bitstreams
By ensuring that only relevant layers are included in the output layer set within the video bitstream, the method addresses the inefficiencies in existing video coding technologies, resulting in improved coding efficiency and user experience.
Patent Information
- Application Number
- JP2025024968
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-09-24
- Filing Date
- 2025-02-19
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2040-09-15
Smart Images

Figure 2025084807000001_ABST
Abstract
Description
Technical Field
[0001]
[0002] [Technical Field] Generally, the present disclosure describes techniques for multi-layer video bitstreams in video coding. More specifically, the present disclosure ensures that unwanted and / or unused layers are prohibited within the multi-layer video bitstream in video coding.
Background Art
[0003] The amount of video data required to depict even relatively short videos can be substantial. This can pose difficulties when the data is streamed across a communication network having limited bandwidth capabilities or otherwise communicated. Thus, video data is typically compressed before being communicated across today's telecommunications networks. When videos are stored on storage devices, the size of the videos can also be an issue since memory resources may be limited. Video compression devices often use software and / or hardware to code the video data at the source prior to transmission or storage, thereby reducing the amount of data required to represent the digital video image. The compressed data is then received at the destination by a video decompression device that decodes the video data. With limited network resources and the ever-increasing demands for higher video quality, improved compression and decompression techniques that improve the compression ratio without sacrificing much or any of the image quality are desirable.
Summary of the Invention
[0004] A first aspect is a method of decoding performed by a video decoder, Receiving, by the video decoder, a video bitstream including a video parameter set (VPS) and a plurality of layers, wherein none of the layers is an output layer of at least one output layer set (OLS) or a direct reference layer of any other layer; Decoding, by the video decoder, a picture from one of the plurality of layers; relates to a method including the above.
[0005] The method provides a technique for prohibiting unnecessary layers in a multi-layer video bitstream. That is, all layers included in an output layer set (OLS) are either output layers or used as direct or indirect reference layers of output layers. This avoids having irrelevant information in the coding process and improves coding efficiency. Therefore, a coder / decoder (also known as a "codec") in video coding is improved over the current codec. Practically, the improved video coding process provides a good user experience to users when the video is transmitted, received, and / or viewed.
[0006] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the picture is included in the output layer of the at least one OLS.
[0007] Optionally, in any of the foregoing aspects, another implementation of the aspect provides selecting an output layer from the at least one OLS before the decoding step.
[0008] Optionally, in any of the foregoing aspects, another implementation of the aspect provides selecting the picture from the output layer after the output picture is selected.
[0009] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that each layer among the plurality of layers includes a set of a video coding layer (VCL) network abstraction layer (NAL) unit all having a specific value of a layer identifier (ID) and an associated non-VCL NAL unit.
[0010] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that each layer of the at least one OLS includes the output layer or the direct reference layer of any other layer.
[0011] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the at least one OLS includes one or more output layers.
[0012] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the VPS includes a layer used as a reference flag and a layer used as an output layer flag, and values of the layer used as the reference flag and the layer used as the output layer flag are both not zero.
[0013] Optionally, in any of the above aspects, another implementation of this aspect provides to display a decoded picture on a display of an electronic device.
[0014] Optionally, in any of the foregoing aspects, another implementation of the aspect is a step of receiving, by the video decoder, a second video bitstream including a second video parameter set (VPS) and a second plurality of layers, wherein at least one layer is neither an output layer of at least one output layer set (OLS) nor a direct reference layer of any other layer. In response to the receiving step, taking several other corrective measures to ensure that a compliant bitstream corresponding to the second video bitstream is received before decoding a picture from one of the second plurality of layers. To provide.
[0015] A second aspect is a method of encoding implemented by a video encoder, the method comprising: Generating, by the video encoder, a plurality of layers and a video parameter set (VPS) that specifies at least one output layer set (OLS), wherein the video encoder is constrained such that no layer is an output layer of at least one OLS or a direct reference layer of any other layer. Encoding, by the video encoder, the plurality of layers and the VPS into a video bitstream. Storing, by the video encoder, the video bitstream for communication to a video decoder. Relates to a method including.
[0016] The method provides a technique for prohibiting unnecessary layers in a multi-layer video bitstream. That is, all layers included in an output layer set (OLS) are either output layers or used as direct or indirect reference layers of output layers. This avoids having irrelevant information in the coding process and improves coding efficiency. Accordingly, a coder / decoder (alias: "codec") in video coding is improved over the current codec. As a practical matter, the improved video coding process provides a good user experience to the user when the video is transmitted, received, and / or viewed.
[0017] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that each of the one or more OLSs includes one or more output layers, and each of the output layers includes one or more pictures.
[0018] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that each layer among the plurality of layers includes a set of a video coding layer (VCL) network abstraction layer (NAL) unit all having a specific value of a layer identifier (ID) and an associated non-VCL NAL unit.
[0019] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that one of the OLSs includes two output layers, and one of the two output layers references the other of the two output layers.
[0020] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the VPS includes a layer used as a reference flag and a layer used as an output layer flag, and the values of the layer used as the reference flag and the layer used as the output layer flag are both non-zero.
[0021] A third aspect is a decoding apparatus, a receiver configured to receive a video bitstream including a video parameter set (VPS) and a plurality of layers, where none of the layers is an output layer of at least one OLS or a direct reference layer of any other layer, the receiver, a memory coupled to the receiver, where the memory stores instructions, the memory, a processor coupled to the memory, where the processor executes the instructions to cause the decoding apparatus to decode a picture from one of the plurality of layers to obtain a decoded picture, the processor, relates to a decoding apparatus including the above.
[0022] The decoding device provides a technique for prohibiting unnecessary layers in a multi-layer video bitstream. That is, all layers included in an output layer set (OLS) are either output layers or are used as direct or indirect reference layers of output layers. This avoids having irrelevant information in the coding process and improves coding efficiency. Therefore, a coder / decoder (also known as a "codec") in video coding is improved with respect to the current codec. As a practical matter, the improved video coding process provides a good user experience to users when the video is transmitted, received, and / or viewed.
[0023] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the picture is included in an output layer of the at least one OLS.
[0024] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the processor is further configured to select an output layer from the at least one OLS before the picture is decoded.
[0025] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the processor is further configured to select the picture from the output layer after the output picture is selected.
[0026] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that each layer in the plurality of layers includes a set of a video coding layer (VCL) network abstraction layer (NAL) unit all having a specific value of a layer identifier (ID) and an associated non-VCL NAL unit.
[0027] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that each layer of the at least one OLS includes the output layer or the direct reference layer of any other layer.
[0028] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the at least one OLS includes one or more output layers.
[0029] Optionally, in any of the above aspects, another implementation of this aspect provides that the display is configured to display the decoded picture.
[0030] Optionally, in any of the above aspects, another implementation of this aspect is that the processor executes the instructions to further cause the decoding device to receive a second video parameter set (VPS) and a second video bitstream including a second plurality of layers, where at least one layer is neither an output layer of at least one output layer set (OLS) nor a direct reference layer of any other layer, and in response to receiving the second video bitstream, take some other corrective measures to ensure that a compliant bitstream corresponding to the second video bitstream is received before decoding a picture from one of the second plurality of layers.
[0031] A fourth aspect is an encoding device, including a memory containing instructions, and a processor coupled to the memory, where the processor executes the instructions to cause the encoding device to generate a plurality of layers and a video parameter set (VPS) specifying at least one output layer set (OLS), and the encoding device is restricted such that none of the layers is an output layer of at least one OLS or a direct reference layer of any other layer, and a processor that encodes the plurality of layers and the VPS into a video bitstream. A transmitter coupled to the processor, the transmitter being configured to transmit the video bitstream towards a video decoder, a transmitter; Relates to an encoding device including.
[0032] The encoding device provides a technique for prohibiting unnecessary layers in a multi-layer video bitstream. That is, all layers included in an output layer set (OLS) are either output layers or are used as direct or indirect reference layers of an output layer. This avoids having irrelevant information in the coding process and improves coding efficiency. Accordingly, a coder / decoder (alias: "codec") in video coding is improved over the current codec. As a practical matter, the improved video coding process provides a good user experience to the user when the video is transmitted, received, and / or viewed.
[0033] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that each of the one or more OLSs includes one or more output layers, and each of the output layers includes one or more pictures.
[0034] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that each layer among the plurality of layers includes a set of a video coding layer (VCL) network abstraction layer (NAL) unit all having a specific value of a layer identifier (ID) and an associated non-VCL NAL unit.
[0035] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that each layer of the at least one OLS includes the output layer or a direct reference layer of any other layer.
[0036] Optionally, in any of the foregoing aspects, another implementation of the aspect includes a layer used as a reference flag by the VPS and a layer used as an output layer flag, and provides that the values of the layer used as a reference flag and the layer used as an output layer flag are both non-zero.
[0037] A fifth aspect relates to a coding device. The coding device a receiver configured to receive and encode a picture or receive and decode a bitstream, and a transmitter coupled to the receiver, the transmitter being configured to transmit the bitstream to a decoder or transmit a decoded image to a display, and a memory coupled to at least one of the receiver or the transmitter, the memory being configured to store instructions, and a processor coupled to the memory, the processor being configured to execute the instructions stored in the memory to execute any of the methods disclosed in the present specification, and including.
[0038] The coding device provides a technique for prohibiting unnecessary layers in a multi-layer video bitstream. That is, all layers included in an output layer set (OLS) are either output layers or used as direct or indirect reference layers of output layers. This avoids having irrelevant information in the coding process and improves coding efficiency. Therefore, a coder / decoder (also known as a "codec") in video coding is improved over the current codec. As a practical matter, the improved video coding process provides a good user experience to the user when the video is transmitted, received, and / or viewed.
[0039] Optionally, in any of the above aspects, another implementation of this aspect provides a display configured to display a decoded picture.
[0040] The sixth aspect relates to a system. The system is an encoder, a decoder communicating with the encoder, wherein the encoder or the decoder includes a decoding device, an encoding device, or a coding device disclosed in the present specification, and includes.
[0041] The system provides a technique for prohibiting unnecessary layers in a multi-layer video bitstream. That is, all layers included in the output layer set (OLS) are either output layers or used as direct or indirect reference layers of the output layer. This avoids having irrelevant information in the coding process and improves coding efficiency. Therefore, the coder / decoder (also known as "codec") in video coding is improved compared to the current codec. As a practical matter, the improved video coding process provides a good user experience to the user when the video is transmitted, received, and / or viewed.
[0042] The seventh aspect relates to means for coding. The means for coding is receiving means configured to receive and encode a picture or receive and decode a bitstream, transmitting means coupled to the receiving means, wherein the transmitting means is configured to transmit the bitstream to decoding means or transmit a decoded image to display means, storage means coupled to at least one of the receiving means or the transmitting means, wherein the storage means is configured to store instructions, A processing means coupled to the memory means, the processing means being configured to execute the instructions stored in the memory means in order to execute any of the methods described in the present specification, the processing means; including.
[0043] The coding means provides a technique for prohibiting unnecessary layers in a multi-layer video bitstream. That is, all layers included in the output layer set (OLS) are either output layers or are used as direct or indirect reference layers of the output layer. This avoids having irrelevant information in the coding process and improves coding efficiency. Therefore, a coder / decoder (alias: "codec") in video coding is improved over the current codec. As a practical matter, the improved video coding process provides a good user experience to the user when the video is transmitted, received, and / or viewed.
[0044] For the purpose of clarity, any one of the foregoing embodiments may be combined with any one or more of the other foregoing embodiments to generate new embodiments within the scope of the present disclosure.
[0045] The above and other features will be more clearly understood from the following detailed description taken in conjunction with the accompanying drawings and the claims.
Brief Description of the Drawings
[0046] For a more complete understanding of the present disclosure, reference is now made to the following brief description taken in conjunction with the accompanying drawings and detailed description. Here, like reference numerals represent like parts.
[0047]
Figure 1
[0048]
Figure 2
[0049]
Figure 3
[0050]
Figure 4
[0051]
Figure 5
[0052]
Figure 6
[0053]
Figure 7
[0054]
Figure 8
[0055]
Figure 9
[0056]
Figure 10
[0057]
Figure 11
DETAILED DESCRIPTION OF THE INVENTION
[0058] It should be understood first that illustrative implementations of one or more embodiments are provided below, but the disclosed system and / or method may be implemented using any number of techniques, whether currently known or existing. The present disclosure should in no way be limited to the illustrative implementations, drawings, and techniques described below, including the exemplary designs and implementations illustrated and described herein, but may be modified within the scope of the appended claims, along with the full scope of their equivalents.
[0059] The following terms are defined as follows, unless used in a context inconsistent herein. In particular, the following definitions are intended to provide further clarity to the present disclosure. However, terms may be described differently in different contexts. Accordingly, the following definitions should be considered as supplementary and should not be considered to limit any other definition provided herein for such terms.
[0060] A bitstream is a sequence of bits that contains video data compressed for transmission between an encoder and a decoder. An encoder is a device configured to utilize an encoding process to compress video data into a bitstream. A decoder is a device configured to utilize a decoding process to reconstruct video data from a bitstream for display. A picture is an array of chroma samples and / or an array of luma samples that generates a frame or a field thereof. A picture that has been encoded or decoded can be referred to as the current picture for clarity of discussion. A reference picture is a picture that contains reference samples that can be used when coding other pictures by reference according to inter prediction and / or inter-layer prediction. A reference picture list is a list of reference pictures used for inter prediction and / or inter-layer prediction. Some video coding systems utilize two reference picture lists that can be denoted as reference picture list 1 and reference picture list 0. A reference picture list structure is an addressable syntax structure that includes a plurality of reference picture lists. Inter prediction is a mechanism that codes samples of the current picture by reference to indicated samples in a reference picture different from the current picture. Here, the reference picture and the current picture are in the same layer. A reference picture list structure entry is an addressable position within the reference picture list structure that indicates a reference picture associated with the reference picture list. A slice header is a part of a coding slice that contains data elements related to all the video data within the tiles represented within the slice. A picture parameter set (PPS) is a parameter set that contains data related to an entire picture. More specifically, the PPS is a syntax structure that includes syntax elements applicable to zero or more coding pictures as determined by syntax elements found within each picture header.A sequence parameter set (SPS) is a parameter set that contains data related to a sequence of pictures. An access unit (AU) is a set of one or more coded pictures associated with the same presentation time (e.g., the same picture order count) for output from a decoded picture buffer (DPB) (e.g., for display to a user). An access unit delimiter (AUD) is an indicator or data structure used to indicate the start of an AU or the boundary between AUs. A decoded video sequence is a sequence of pictures reconstructed by a decoder when preparing for display to a user.
[0061] A network abstraction layer (NAL) unit is a syntax structure that contains data in the form of a raw byte sequence payload (RBSP), an indication of the type of data, and, optionally, emulation prevention bytes scattered therein. A video coding layer (VCL) NAL unit is a NAL unit coded to contain video data such as a coding slice of a picture. A non-VCL NAL unit is a NAL unit that contains non-video data such as syntax and / or parameters that support decoding of video data, compliance checking performance, or other operations. A layer is a set of VCL NAL units and associated non-VCL NAL units that share certain characteristics (e.g., common resolution, frame rate, picture size, etc.). The VCL NAL units of a layer may share a particular value of the NAL unit header layer identifier (nuh_layer_id). A coded picture is a coded representation of a picture that has a particular value of the NAL unit header layer identifier (nuh_layer_id) within an access unit (AU) and contains VCL NAL units that include all coding tree units (CTUs) of the picture. A decoded picture is a picture generated by applying decoding processing to a coded picture.
[0062] An output layer set (OLS) is a set of layers in which one or more layers are designated as output layers. An output layer is a layer designated for output (e.g., to a display). The zero-th OLS contains only the bottom layer (the layer with the bottom layer identifier), and thus is an OLS that contains only output layers. A video parameter set (VPS) is a data unit that contains parameters related to the entire video. Inter-layer prediction is a mechanism for coding a current picture in a current layer by referring to a reference picture in a reference layer. Here, the current picture and the reference picture are included in the same AU, and the reference layer contains a lower nuh_layer_id than the current layer.
[0063] The following abbreviations are used in this specification. Coding tree block (CTB), coding tree unit (CTU), coding unit (CU), coded video sequence (CVS), Joint Video Experts Team (JVET), network abstraction layer (NAL), picture order count (POC), Picture Parameter Set (PPS), raw byte sequence payload (RBSP), sequence parameter set (SPS), versatile video coding (VVC), and working draft (WD).
[0064] FIG. 1 is a flowchart of an exemplary operating method 100 for coding a video signal. Specifically, the video signal is encoded by an encoder. The encoding process compresses the video signal by utilizing various mechanisms to reduce the video file size. A smaller file size enables the transmission of a compressed video file to the user while reducing the associated bandwidth overhead. The decoder then decodes the compressed video file to reconstruct the original video signal for display to the end user. The decoding process is typically a mirror of the encoding process, enabling the decoder to reconstruct the video signal without contradictions.
[0065] In step 101, the video signal is input to the encoder. For example, the video signal may be an uncompressed video file stored in memory. As another example, the video file may be captured by a video capture device such as a video camera and encoded to support live streaming of the video. The video file may include both an audio component and a video component. The video component includes a series of image frames that give a visual impression of movement when viewed within a sequence. The frames include pixels represented in terms of light, herein referred to as the luma component (or luma samples), and color, herein referred to as the chroma component (or chroma samples). In some examples, the frames may also include depth values to support 3D display.
[0066] In step 103, the video is partitioned into blocks. Partitioning includes subdividing the pixels within each frame into square and / or rectangular blocks for compression. For example, in High Efficiency Video Coding (HEVC) (also known as H.265 and MPEG-H Part 2), a frame can first be divided into coding tree units (CTUs) that are blocks of a predetermined size (e.g., 64 pixels × 64 pixels). A CTU includes both luma and chroma samples. Coding trees may be used to divide a CTU into blocks and then repeatedly subdivide the blocks until a configuration that supports further encoding is achieved. For example, the luma component of a frame may be subdivided until the individual blocks contain relatively homogeneous light values. Further, the chroma component of a frame may be subdivided until the individual blocks contain relatively homogeneous color values. Thus, the partitioning mechanism varies depending on the content of the video frame.
[0067] In step 105, various compression mechanisms are utilized to compress the image blocks partitioned in step 103. For example, inter prediction and / or intra prediction may be utilized. Inter prediction is designed to utilize the fact that objects in a common scene tend to appear in consecutive frames. Therefore, blocks depicting objects within a reference frame do not need to be repeatedly shown in adjacent frames. Specifically, an object such as a table may remain in a fixed position across multiple frames. Thus, the table is shown once, and adjacent frames can return to the reference frame for reference. A pattern matching mechanism may be utilized to match objects across multiple frames. Further, for example, due to the movement of an object or the movement of a camera, a moving object may be displayed across multiple frames. As a specific example, a video may show a car moving horizontally across the screen across multiple frames. To indicate such movement, motion vectors can be utilized. A motion vector is a two-dimensional vector that provides an offset from the coordinates of an object within a frame to the coordinates of the object within a reference frame. Therefore, inter prediction can encode an image block within the current frame as a set of motion vectors indicating the offsets from the corresponding blocks within the reference frame.
[0068] Intra prediction encodes blocks within a common frame. Intra prediction exploits the fact that luma and chroma components tend to cluster within a frame. For example, a green patch of a part of a tree tends to be located adjacent to a similar green patch. Intra prediction utilizes multiple directional prediction modes (e.g., 33 in HEVC), a planar mode, and a direct current (DC) mode. The directional mode indicates that the current block is similar / same as the samples of neighboring blocks in the corresponding direction. The planar mode indicates that a series of blocks along a row / column (e.g., a plane) can be interpolated based on neighboring blocks at the end of the row. The planar mode effectively indicates a smooth transition of light / color across the row / column by utilizing a relatively constant gradient of changing values. The DC mode is utilized for boundary smoothing and indicates that the block is similar / same as the average value related to the samples of all neighboring blocks related to the angular direction of the directional prediction mode. Thus, an intra prediction block can represent an image block as various relational prediction modes instead of the actual values. Further, an inter prediction block can represent an image block as a motion vector value instead of the actual values. In any case, the prediction block may not accurately represent the image in some cases. Any difference is stored in the residual block. A transform may be applied to the residual block to further compress the file.
[0069] In step 107, various filtering techniques may be applied. In HEVC, the filters are applied according to the in-loop filtering method. The prediction based on the above-mentioned blocks may result in the generation of an image with uneven shading in the decoder. Further, the block-based prediction method may encode a block and then reconstruct the encoded block for later use as a reference block. The in-loop filtering method repeatedly applies a noise reduction filter, a deblocking filter, an adaptive loop filter, and a sample adaptive offset (SAO) filter to the blocks / filters. These filters mitigate such artifacts of uneven shading, and as a result, the encoded file can be accurately reconstructed. Further, these filters mitigate the artifacts in the reconstructed reference blocks, and as a result, there is a low possibility of generating additional artifacts in the subsequent blocks encoded based on the reconstructed reference blocks.
[0070] When the video signal is partitioned, compressed, and filtered, the resulting data is encoded into a bitstream in step 109. The bitstream includes the above-mentioned data and any desired signaling data for supporting proper video signal reconstruction in the decoder. For example, such data may include partition data, prediction data, residual blocks, and various flags that provide coding instructions to the decoder. The bitstream may be stored in memory for transmission to the decoder upon request. The bitstream may be broadcast and / or multicast to a plurality of decoders. The generation of the bitstream is an iterative process. Accordingly, steps 101, 103, 105, 107, and 109 may occur continuously and / or simultaneously over a number of frames and blocks. The order shown in FIG. 1 is presented for clarity and ease of discussion and is not intended to limit the video coding process to a specific order.
[0071] In step 111, the decoder receives the bitstream and starts the decoding process. Specifically, the decoder uses an entropy decoding method to convert the bitstream into the corresponding syntax and video data. In step 111, the decoder determines the frame partition using the syntax from the bitstream. The partition should match the result of the block partition in step 103. Entropy coding / decoding as used in step 111 is described below. The encoder generates a number of options such that during the compression process, it selects a block partitioning method from several possible options based on the spatial position of the values in the input image. Signaling the exact option could use a huge number of bins. As used here, a bin is a binary value treated as a variable (e.g., a bit value that can vary depending on the context). Entropy coding makes it possible to discard any option that is clearly not executable by the encoder in a particular case, leaving a set of acceptable options. Each acceptable option is then assigned a codeword. The length of the codeword is based on the number of acceptable options (e.g., 1 bin for 2 options, 2 bins for 3 - 4 options, etc.). The encoder then encodes the codeword for the selected option. This method is desirable in order to reduce the size of the codeword because it uniquely represents a selection from a small subset of possible options, as opposed to uniquely representing a selection from a potentially large set of all possible options. The decoder then decodes the selection by determining the set of acceptable options in the same way as the encoder. By determining the set of acceptable options, the decoder can read the codeword and determine the selection made by the encoder.
[0072] In step 113, the decoder performs block decoding. Specifically, the decoder generates a residual block using inverse transformation. Next, the decoder reconstructs an image block according to a partition using the residual block and the corresponding prediction block. The prediction block may include both an intra prediction block and an inter prediction block generated in step 105 in the encoder. The reconstructed image block is then positioned into a frame of the reconstructed video signal according to the partition data determined in step 111. The syntax of step 113 may also be signaled in the bitstream by entropy coding as described above.
[0073] In step 115, filtering is performed on the frame of the reconstructed video signal in a manner similar to step 107 in the encoder. For example, a noise suppression filter, a deblocking filter, an adaptive loop filter, and an SAO filter may be applied to the frame to remove blocking artifacts. Once the frame is filtered, the video signal can be output to a display in step 117 for viewing by an end user.
[0074] Figure 2 is a schematic diagram of an exemplary coding and decoding (codec) system 200 for video coding. Specifically, the codec system 200 provides functions to support the implementation of the operation method 100. The codec system 200 is generalized to show components that are used in both the encoder and the decoder. The codec system 200 receives and partitions the video signal described above with respect to steps 101 and 103 in the operation method 100 to produce a partitioned video signal 201. The codec system 200 then compresses the partitioned video signal 201 into a coding bitstream when operating as the encoder described above with respect to steps 105, 107, and 109 of method 100. When operating as a decoder, the codec system 200 generates an output video signal from the bitstream as described above with respect to steps 111, 113, 115, and 117 in the operation method 100. The codec system 200 includes a general-purpose codec control component 211, a transform scaling and quantization component 213, an intra-picture estimation component 215, an intra-picture prediction component 217, a motion compensation component 219, a motion estimation component 221, a scaling and inverse transform component 229, a filter control analysis component 227, an in-loop filter component 225, a decoded picture buffer component 223, and a header format and context adaptive binary arithmetic coding (CABAC) component 231. Such components are coupled as shown. In Figure 2, the solid lines indicate the movement of data to be encoded / decoded, while the dashed lines indicate the movement of control data that controls the operation of other components. The components of the codec system 200 may all be present within the encoder. The decoder may include some of the components of the codec system 200.For example, the decoder may include an intra-picture prediction component 217, a motion compensation component 219, a scaling and inverse transform component 229, an in-loop filter component 225, and a decoded picture buffer component 223. These components are described herein.
[0075] The partitioned video signal 201 is a captured video sequence that has been partitioned into blocks of pixels by a coding tree. The coding tree uses various partitioning modes to subdivide blocks of pixels into smaller blocks of pixels. These blocks can then be further subdivided into even smaller blocks. The blocks may be referred to as nodes on the coding tree. Larger parent nodes are divided into smaller child nodes. The number of times a node is subdivided is called the depth of the node / coding tree. The partitioned blocks may, in some cases, be included in a coding unit (CU). For example, a CU may be a sub-part of a CTU that includes a luma block, a red difference chroma (Cr) block, and a blue difference chroma (Cb) block, as well as the corresponding syntax instruction for the CU. The partitioning modes may include a binary tree (BT), a triple tree (TT), and a quad tree (QT) that are used to partition a node into two, three, or four child nodes of various shapes depending on the partitioning mode used. The partitioned video signal 201 is transferred for compression to a general-purpose coder control component 211, a transform scaling and quantization component 213, an intra-picture estimation component 215, a filter control analysis component 227, and a motion estimation component 221.
[0076] The general coder control component 211 is configured to make decisions related to the coding of images of a video sequence into a bitstream according to application constraints. For example, the general coder control component 211 manages the optimization of the bit rate / bitstream size with respect to the reconstructed quality. Such decisions may be made based on the availability of memory space / bandwidth and the image resolution requirements. The general coder control component 211 also manages the buffer utilization in terms of the conversion speed in order to mitigate buffer underrun and overrun problems. To address these problems, the general coder control component 211 manages the partitioning, prediction, and filtering by other components. For example, the general coder control component 211 may dynamically increase the compression complexity to increase the resolution, and increase the bandwidth usage or decrease the compression complexity to reduce the resolution and bandwidth usage. Thus, the general coder control component 211 controls other components of the codec system 200 to balance the video signal reconstructed quality and the bit rate concerns. The general coder control component 211 generates control data for controlling the operation of other components. The control data is also transferred to the header format and the CABAC component 231 so as to be encoded in the bitstream for signaling the parameters for decoding in the decoder.
[0077] The partitioned video signal 201 is also sent to the motion estimation component 221 and the motion compensation component 219 for inter prediction. The frame or slice of the partitioned video signal 201 may be divided into a plurality of video blocks. The motion estimation component 221 and the motion compensation component 219 perform the inter prediction coding of the received video blocks in relation to one or more blocks in one or more reference frames and provide temporal prediction. The codec system 200 may execute a plurality of coding paths, for example, to select an appropriate coding mode for each block of the video data.
[0078] The motion estimation component 221 and the motion compensation component 219 may be highly integrated but are shown separately for conceptual purposes. The motion estimation performed by the motion estimation component 221 is a process of generating a motion vector that estimates the motion for a video block. The motion vector may indicate, for example, the placement of a coding object in relation to a prediction block. The prediction block is a block that has been found to exactly match the block to be coded in terms of pixel differences. The prediction block may also be referred to as a reference block. Such pixel differences may be determined by the sum of absolute difference (SAD), the sum of square difference (SSD), or other difference metrics. HEVC utilizes several coding objects including CTUs, coding tree blocks (CTBs), and CUs. For example, a CTU can be divided into CTBs, and a CTB can then be divided into CUs for inclusion in a CB. A CU can be coded as a prediction unit (PU) containing prediction data and / or a transform unit (TU) containing the transformed residual data of the CU. The motion estimation component 221 uses rate distortion analysis as part of the rate distortion optimization process to generate motion vectors, PUs, and TUs. For example, the motion estimation component 221 may determine multiple reference blocks, multiple motion vectors, etc. for the current block / frame and may select the reference blocks, motion vectors, etc. having optimal rate distortion characteristics. The optimal rate distortion characteristics balance both the quality of video reconstruction (e.g., the amount of data loss due to compression) and the coding efficiency (e.g., the size of the final coding).
[0079] In some examples, the codec system 200 may calculate the values of the sub-integer picture positions of the reference pictures stored in the decoded picture buffer component 223. For example, the video codec system 200 may interpolate the values of the quarter-pixel positions, the eighth-pixel positions, or other fractional-pixel positions of the reference pictures. Thus, the motion estimation component 221 may perform motion search in relation to both the full-pixel positions and the fractional-pixel positions and output motion vectors with fractional-pixel accuracy. The motion estimation component 221 calculates the motion vectors for the PUs of the video blocks in the inter-coded slice by comparing the position of the PU with the position of the predicted block of the reference picture. The motion estimation component 221 outputs the calculated motion vectors as motion data to the header format and the CABAC component 231 for encoding and outputs the motion to the motion compensation component 219.
[0080] The motion compensation performed by the motion compensation component 219 may include fetching or generating a predicted block based on the motion vector determined by the motion estimation component 221. Again, the motion estimation component 221 and the motion compensation component 219 may be functionally integrated in some examples. Upon receiving the motion vector of the current video block's PU, the motion compensation component 219 may identify the position of the predicted block pointed to by the motion vector. Next, a residual video block is formed by subtracting the pixel values of the predicted block from the pixel values of the currently coded video block to form a pixel difference value. Generally, the motion estimation component 221 performs motion estimation in relation to the luma component, and the motion compensation component 219 uses the motion vectors calculated based on the luma component for both the chroma components and the luma component. The predicted block and the residual block are transferred to the transform scaling and quantization component 213.
[0081] The partitioned video signal 201 is also sent to the intra picture estimation component 215 and the intra picture prediction component 217. Similar to the motion estimation component 221 and the motion compensation component 219, the intra picture estimation component 215 and the intra picture prediction component 217 may be highly integrated, but are shown separately for conceptual purposes. Instead of the inter prediction performed by the inter-frame motion estimation component 221 and the motion compensation component 219 as described above, the intra picture estimation component 215 and the intra picture prediction component 217 perform intra prediction of the current block in relation to blocks within the current frame. In particular, the intra picture estimation component 215 determines the intra prediction mode to be used to encode the current block. In some examples, the intra picture estimation component 215 selects an appropriate intra prediction mode for encoding the current block from a plurality of tested intra prediction modes. The selected intra prediction mode is then transferred to the header format and the CABAC component 231 for encoding.
[0082] For example, the intra picture estimation component 215 calculates rate-distortion values using rate-distortion analysis for various tested intra prediction modes, and selects an intra prediction mode having optimal rate-distortion characteristics among the tested modes. Rate-distortion analysis generally determines the amount of distortion (or error) between an encoded block and the original unencoded block from which the encoded block was generated, as well as the bit rate (e.g., the number of bits) used to generate the encoded block. The intra picture estimation component 215 calculates a ratio from the distortion and rate for various encoded blocks to determine, for a block, which intra prediction mode exhibits an optimal rate-distortion value. Further, the intra picture estimation component 215 may be configured to code depth blocks of a depth map using a depth modeling mode (DMM) based on rate-distortion optimization (RDO).
[0083] When implemented in an encoder, the intra picture prediction component 217 may generate a residual block from a prediction block based on the selected intra prediction mode determined by the intra picture estimation component 215, or when implemented in a decoder, may read a residual block from a bit stream. The residual block includes the difference in values between the prediction block and the original block, represented as a matrix. The residual block is then transferred to the transform scaling and quantization component 213. The intra picture estimation component 215 and the intra picture prediction component 217 may operate on both luma and chroma components.
[0084] The transform scaling and quantization component 213 is configured to further compress the residual block. The transform scaling and quantization component 213 applies a transform, such as a discrete cosine transform (DCT), a discrete sine transform (DST), or a conceptually similar transform, to the residual block to generate a video block including residual transform coefficient values. A wavelet transform, an integer transform, a subband transform, or other types of transforms may also be used. The transform may transform the residual information from the pixel value domain to a transform domain, such as the frequency domain. The transform scaling and quantization component 213 is also configured to scale the transformed residual information, for example, based on frequency. Such scaling includes applying a magnification factor to the residual information. As a result, different frequency information is quantized at different granularities, which may affect the final visual quality of the reconstructed video. The transform scaling and quantization component 213 is also configured to quantize the transform coefficients to further reduce the bitrate. The quantization process may reduce the bit depth associated with some or all of the coefficients. The degree of quantization may be changed by adjusting the quantization parameter. In some examples, the transform scaling and quantization component 213 may then perform a scan of the matrix including the quantized transform coefficients. The quantized transform coefficients are transferred to the header format and the CABAC component 231 for encoding in the bitstream.
[0085] The scaling and inverse transform component 229 applies the inverse operations of the transform scaling and quantization component 213 to support motion estimation. The scaling and inverse transform component 229 applies inverse scaling, transform, and / or quantization to reconstruct the residual block in the pixel domain for later use, for example, as a reference block that can become a prediction block for another current block. The motion estimation component 221 and / or the motion compensation component 219 may calculate the reference block by adding the residual block back to the corresponding prediction block for use in motion estimation of subsequent blocks / frames. A filter is applied to the reconstructed reference block to reduce artifacts generated during scaling, quantization, and transform. Such artifacts could otherwise cause inaccurate predictions (and generate additional artifacts) when subsequent blocks are predicted.
[0086] The filter control analysis component 227 and the in-loop filter component 225 apply filters to the residual block and / or the reconstructed image block. For example, the transformed residual block from the scaling and inverse transform component 229 may be combined with the corresponding prediction block from the intra-picture prediction component 217 and / or the motion compensation component 219 to reconstruct the original image block. The filter may then be applied to the reconstructed image block. In some examples, the filter may instead be applied to the residual block. Similar to the other components in FIG. 2, the filter control analysis component 227 and the in-loop filter component 225 may be highly integrated and implemented together, but are shown separately for conceptual purposes. The filter applied to the reconstructed reference block is applied to a specific spatial region and includes a plurality of parameters for adjusting how such a filter is applied. The filter control analysis component 227 analyzes the reconstructed reference block to determine when such a filter should be applied and sets the corresponding parameters. Such data is transferred as filter control data for encoding to the header format and the CABAC component 231. The in-loop filter component 225 applies such a filter based on the filter control data. The filter may include a deblocking filter, a noise suppression filter, an SAO filter, and an adaptive loop filter. Such a filter may be applied in the spatial / pixel domain (e.g., on the reconstructed pixel block) or in the frequency domain, depending on the example.
[0087] When operating as an encoder, the filtered reconstructed image block, residual block, and / or prediction block are stored in the decoded picture buffer component 223 for later use in motion estimation as described above. When operating as a decoder, the decoded picture buffer component 223 stores the reconstructed and filtered blocks as part of the output video signal and transfers them towards the display. The decoded picture buffer component 223 may be any memory device capable of storing prediction blocks, residual blocks, and / or reconstructed image blocks.
[0088] The header format and CABAC component 231 receive data from various components of the codec system 200 and encode such data into a coding bitstream for transmission to the decoder. Specifically, the header format and CABAC component 231 generate various headers for encoding control data such as general control data and filter control data. Further, prediction data including intra prediction and motion data, and residual data in the form of quantized transform coefficient data are all encoded within the bitstream. The final bitstream contains all the information desired by the decoder to reconstruct the original partitioned video signal 201. Such information may also include an intra prediction mode index table (also referred to as a codeword mapping table), the definition of the coding context for various blocks, an indication of the most promising intra prediction mode, an indication of partition information, etc. Such data may be encoded by utilizing entropy coding. For example, the information may be encoded by utilizing context adaptive variable length coding (CAVLC), syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or another entropy coding technique. In accordance with the entropy coding, the encoded bitstream may be transmitted to another device (e.g., a video decoder) or stored for later transmission or reading.
[0089] FIG. 3 is a block diagram showing an exemplary video encoder 300. The video encoder 300 may be utilized to implement the encoding function of the codec system 200 and / or to implement steps 101, 103, 105, 107 and / or 109 of the operation method 100. The encoder 300 partitions the input video signal to yield a partitioned video signal 301 that is substantially similar to the partitioned video signal 201. The partitioned video signal 301 is then compressed and encoded into a bitstream by components of the encoder 300.
[0090] Specifically, the partitioned video signal 301 is transferred to an intra-picture prediction component 317 for intra prediction. The intra-picture prediction component 317 may be substantially similar to the intra-picture estimation component 215 and the intra-picture prediction component 217. The partitioned video signal 301 is also transferred to a motion compensation component 321 for inter prediction based on reference blocks in the decoded picture buffer component 323. The motion compensation component 321 may be substantially similar to the motion estimation component 221 and the motion compensation component 219. The prediction blocks and residual blocks from the intra-picture prediction component 317 and the motion compensation component 321 are transferred to a transform and quantization component 313 for transformation and quantization of the residual blocks. The transform and quantization component 313 may be substantially similar to the transform scaling and quantization component 213. The transformed and quantized residual blocks and the corresponding prediction blocks (along with associated control data) are transferred to an entropy coding component 313 for coding into the bitstream. The entropy coding component 331 may be substantially similar to the header format and CABAC component 231.
[0091] The transformed and quantized residual block and / or the corresponding prediction block are also transferred from the transform and quantization component 313 to the inverse transform and quantization component 329 in order to be reconstructed into a reference block for use by the motion compensation component 321. The inverse transform and quantization component 329 may be substantially the same as the scaling and inverse transform component 229. The in-loop filter in the in-loop filter component 325 is also applied to the residual block and / or the reconstructed reference block, depending on the example. The in-loop filter component 325 may be substantially the same as the filter control analysis component 227 and the in-loop filter component 225. The in-loop filter component 325 may include multiple filters as discussed with respect to the in-loop filter component 225. The filtered block is then stored in the decoded picture buffer component 323 for use as a reference block by the motion compensation component 321. The decoded picture buffer component 323 may be substantially the same as the decoded picture buffer component 223.
[0092] Figure 4 is a block diagram showing an exemplary video decoder 400. The video encoder 400 may be utilized to implement the decoding function of the codec system 200 and / or to implement steps 111, 113, 115, and / or 117 of the method of operation 100. The decoder 400 receives, for example, a bitstream from the encoder 300 and generates an output video signal reconstructed based on the bitstream for display to an end user.
[0093] The bitstream is received by the entropy decoding component 433. The entropy decoding component 433 is configured to implement an entropy decoding method such as CAVLC, CABAC, SBAC, PIPE coding, or other entropy coding techniques. For example, the entropy decoding component 433 may utilize header information to provide context for interpreting additional data encoded as codewords within the bitstream. The decoded information includes any desired information for decoding a video signal, such as general control data, filter control data, partition information, motion data, prediction data, and quantized transform coefficients from residual blocks. The quantized transform coefficients are transferred to the inverse transform and quantization component 429 for reconstruction into residual blocks. The inverse transform and quantization component 429 may be similar to the inverse transform and quantization component 329.
[0094] The reconstructed residual block and / or prediction block is transferred to the intra-picture prediction component 417 to reconstruct into an image block based on the intra prediction operation. The intra-picture prediction component 417 may be similar to the intra-picture estimation component 215 and the intra-picture prediction component 217. Specifically, the intra-picture prediction component 417 utilizes a prediction mode to identify the position of a reference block within a frame, applies a residual block to the result, and reconstructs the intra-predicted image block. The reconstructed intra-predicted image block and / or residual block, and the corresponding inter-prediction data, are transferred to the decoded picture buffer component 423 via the in-loop filter component 425 which may be substantially similar to the decoded picture buffer component 223 and the in-loop filter component 225 respectively. The in-loop filter component 425 filters the reconstructed image block, residual block, and / or prediction block, and such information is stored in the decoded picture buffer component 423. The reconstructed image block from the decoded picture buffer component 423 is transferred to the motion compensation component 421 for inter prediction. The motion compensation component 421 may be substantially similar to the motion estimation component 221 and / or the motion compensation component 219. Specifically, the motion compensation component 421 utilizes a motion vector from a reference block to generate a prediction block, provides a residual block to the result, and reconstructs the image block. The resulting reconstructed block may be transferred to the decoded picture buffer component 423 via the in-loop filter component 425. The decoded picture buffer component 423 may continue to store additional reconstructed image blocks that can be reconstructed into a frame according to partition information. Such a frame may be arranged within a sequence. The sequence is output towards a display as a reconstructed output video signal.
[0095] Keeping the above in mind, video compression techniques perform spatial (intrapicture) prediction and / or temporal (interpicture) prediction to reduce or remove redundancy inherent in a video sequence. In block-based video coding, a video slice (i.e., a video picture or a portion of a video picture) may be partitioned into video blocks that may also be referred to as tree blocks, coding tree blocks (CTBs), coding tree units (CTUs), coding units (CUs), and / or coding nodes. Video blocks within an intra-coding (I) slice of a picture are coded using spatial prediction with respect to reference samples in neighboring blocks within the same picture. Video blocks within an inter-coding (P or B) slice of a picture may utilize spatial prediction with respect to reference samples in neighboring blocks within the same picture or temporal prediction with respect to reference samples in other reference pictures. A picture may be referred to as a frame, and a reference picture may be referred to as a reference frame. POC is a variable associated with each picture that uniquely identifies the associated picture among all pictures within a coded layer video sequence (CLVS), indicates when the associated picture is output from the DPB, and indicates the position of the associated picture in the output order relative to the output order positions of other pictures within the same CLVS that should be output from the DPB. A flag is a variable or single-bit syntax element that can take on one of two possible values: 0 and 1.
[0096] Spatial or temporal prediction results in a predicted block of the block to be coded. Residual data represents the pixel difference between the original block to be coded and the predicted block. An inter-coded block is coded according to a motion vector indicating a block of reference samples forming the predicted block, and residual data indicating the difference between the coded block and the predicted block. An intra-coded block is coded according to an intra-coding mode and residual data. For further compression, the residual data may be transformed from the pixel domain to a transform domain, resulting in residual transform coefficients, which may then be quantized. The quantized transform coefficients are first arranged in a two-dimensional array and may be scanned to generate a one-dimensional vector of the transform coefficients, and entropy coding may be applied to achieve further compression.
[0097] Image and video compression have experienced rapid growth and have led to various coding standards. Such video coding standards include ITU-T H.261, International Organization for Standardization / International Electrotechnical Commission (ISO / IEC) MPEG-1 Part 2, ITU-T H.262 or ISO / IEC MPEG-2 Part 2, ITU-T H.263, ISO / IEC MPEG-4 Part 2, ITU-T H.264 or Advanced Video Coding (AVC) also known as ISO / IEC MPEG-4 Part 10, ITU-T H.265 or High Efficiency Video Coding (HEVC) also known as MPEG-H Part 2. AVC includes extensions such as Scalable Video Coding (SVC), Multiview Video Coding (MVC), and Multiview Video Coding plus Depth (MVC+D), as well as 3D-AVC. HEVC includes extensions such as Scalable HEVC (SHVC), Multiview HEVC (MV-HEVC), and 3D HEVC (3D-HEVC).
[0098] There is also a new coding standard named Versatile Video Coding (VVC) by the joint video experts team (JVET) of ITU-T and ISO / IEC. The VVC standard has several working drafts, and in particular, one working draft (WD) of VVC, namely B. Bross, J. Chen, and S. Liu, "Versatile Video Coding (Draft5)", JVET-N1001-v3, 13th JVET Meeting, March 27, 2019 (VVC Draft5) is referred to here.
[0099] Scalability in video coding is supported by using multi-layer coding techniques. A multi-layer bitstream includes a base layer (BL) and one or more enhancement layers (EL). Examples of scalability include spatial scalability, quality / signal-to-noise ratio (SNR) scalability, multi-view scalability, etc. When multi-layer coding techniques are used, a picture or a part thereof can be coded by (1) using intra prediction without using a reference picture, i.e., by using intra prediction, (2) referring to a reference picture within the same layer, i.e., by using inter prediction, or (3) referring to a reference picture in another layer, i.e., by using inter-layer prediction. The reference picture used for inter-layer prediction of the current picture is called an inter-layer reference picture (ILRP).
[0100] FIG. 5 is a schematic diagram showing an example of layer - based prediction 500, such as is executed to determine an MV in, for example, block compression step 105, block decoding step 113, motion estimation component 221, motion compensation component 219, motion compensation component 321, and / or motion compensation component 421. Layer - based prediction 500 is compatible with one - direction inter - prediction and / or bi - direction inter - prediction, but is also executed between pictures of different layers.
[0101] Layer - based prediction 500 is applied between pictures 511, 512, 513, and 514 and pictures 515, 516, 517, and 518 in different layers. In the illustrated example, pictures 511, 512, 513, and 514 are part of layer N + 1 532, and pictures 515, 516, 517, and 518 are part of layer N 531. Layers such as layer N 531 and / or layer N + 1 532 are groups of pictures all related to characteristics of similar values, such as similar size, quality, resolution, signal - to - noise ratio, capabilities, etc. Thus, pictures 511, 512, 513, and 514 within layer N + 1 532 have a larger picture size (e.g., greater height and width, and thus more samples) than pictures 515, 516, 517, and 518 within layer N 531 of this example. However, such pictures can be separated between layer N + 1 532 and layer N 531 by other characteristics. Only two layers, layer N + 1 532 and layer N 531, are shown, but a set of pictures can be separated into any number of layers based on related characteristics. Layer N + 1 532 and layer N 531 may be indicated by a layer ID. The layer ID is an item of data associated with a picture and indicates that the picture is part of the layer in which it is shown. Thus, each picture 511 - 518 can be associated with a corresponding layer ID to indicate which layer N + 1 532 or layer N 531 contains the corresponding figure.
[0102] Pictures 511 - 518 within different layers 531 - 532 are configured to be displayed as alternatives. Thus, pictures 511 - 518 within different layers 531 - 532 can share the same time identifier (ID) and can be included within the same AU. As used herein, an AU is a set of one or more coded pictures related to the same display time for output from the DPB. For example, if a smaller picture is desired, the decoder can decode and display picture 515 at the current display time, and if a larger picture is desired, the decoder can decode and display picture 511 at the current display time. Thus, pictures 511 - 514 in the upper layer N + 1 532 contain substantially the same image data as the corresponding pictures 515 - 518 in the lower layer N 531 (despite the difference in picture size). Specifically, picture 511 contains substantially the same image data as picture 515, picture 512 contains substantially the same image data as picture 516, and so on.
[0103] Pictures 511 to 518 can be coded by referring to other pictures 511 to 518 within the same layer N 531 or N+1 532. When a picture is coded by referring to another picture within the same layer, an inter prediction 523, which is a compatible unidirectional inter prediction and / or bidirectional inter prediction, is obtained. The inter prediction 523 is indicated by solid arrows. For example, picture 513 may be coded by adopting an inter prediction 523 that uses one or two of pictures 511, 512, and / or 514 within layer N+1 532 as references, where one picture is referred to for unidirectional inter prediction and / or two pictures are referred to for bidirectional inter prediction. Further, picture 517 may be coded by adopting an inter prediction 523 that uses one or two of pictures 515, 516, and / or 518 within layer N 531 as references, where one picture is referred to for unidirectional inter prediction and / or two pictures are referred to for bidirectional inter prediction. When a picture is used as a reference for another picture within the same layer when performing the inter prediction 523, the picture may be called a reference picture. For example, picture 512 may be a reference picture used to code picture 513 according to the inter prediction 523. The inter prediction 523 may also be called an intra-layer prediction in a multi-layer context. Thus, the inter prediction 523 is a mechanism that codes samples of the current picture, by reference, to samples shown in a reference picture different from the current picture. Here, the reference picture and the current picture are in the same layer.
[0104] Pictures 511 to 518 can also be coded by referring to other pictures 511 to 518 in different layers. This process is known as inter-layer prediction 521 and is indicated by the dashed arrows. Inter-layer prediction 521 is a mechanism for coding samples of the current picture by referring to the indicated samples in the reference picture when the current picture and the reference picture are in different layers and thus have different layer IDs. For example, a picture in the lower layer N 531 can be used as a reference picture to code the corresponding picture in the upper layer N+1 532. As a specific example, picture 511 can be coded by referring to picture 515 according to inter-layer prediction 521. In such a case, picture 515 is used as an inter-layer reference picture. An inter-layer reference picture is a reference picture used for inter-layer prediction 521. In most cases, inter-layer prediction 521 is restricted such that the current picture, such as picture 511, can only use inter-layer reference pictures that are in the same AU and in a lower layer, such as picture 515. If multiple layers (e.g., two or more) are available, inter-layer prediction 521 can encode / decrypt the current picture based on multiple inter-layer reference pictures at a level lower than the current picture.
[0105] The video encoder can encode pictures 511 - 518 through many different combinations and / or permutations of inter prediction 523 and inter-layer prediction 521 using layer-based prediction 500. For example, picture 515 may be coded according to intra prediction. Then, by using picture 515 as a reference picture, pictures 516 - 518 can be coded according to inter prediction 523. Further, picture 511 may be coded according to inter-layer prediction 521 by using picture 515 as an inter-layer reference picture. Then, by using picture 511 as a reference picture, pictures 512 - 514 can be coded according to inter prediction 523. Thus, the reference picture can function as both a single-layer reference picture and an inter-layer reference picture for different coding mechanisms. By coding the upper-layer N+1 picture 532 based on the lower-layer N picture 531, the upper-layer N+1 picture 532 can avoid using intra prediction which has a much lower coding efficiency than inter prediction 523 and inter-layer prediction 521. Thus, the poor coding efficiency of intra prediction can be limited to the minimum / lowest quality pictures and thus can be limited to coding the minimum amount of video data. The pictures used as reference pictures and / or inter-layer reference pictures can be indicated among the entries of the reference picture list included in the reference picture list structure.
[0106] Each AU506 in FIG. 5 can include several pictures. For example, one AU506 can include pictures 511 and 515. Another AU506 can include pictures 512 and 516. In fact, each AU506 is a set of one or more coding pictures associated with the same display time (e.g., the same time ID) for output from a decoded picture buffer (DPB) (e.g., for display to a user). Each AUD508 is an indicator or data structure used to indicate the start of an AU (e.g., AU508) or the boundary between AUs.
[0107] Previous H.26x video coding families have provided support for scalability in a profile separate from the profile for single-layer coding. Scalable video coding (SVC) is a scalable extension of AVC / H.264 that provides support for spatial, temporal, and quality scalability. In SVC, a flag is signaled within each macroblock (MB) in an EL picture to indicate whether the EL MB is predicted using the same-position block from the lower layer. Prediction from the same-position block may include texture, motion vectors, and / or coding modes. Implementations of SVC cannot directly reuse an unmodified H.264 / AVC implementation in the design. The syntax and decoding process of SVC EL macroblocks are different from those of H.264 / AVC.
[0108] Scalable HEVC (SHVC) is an extension of the HEVC / H.265 standard that provides support for spatial and quality scalability, multiview HEVC (MV-HEVC) is an extension of HEVC / H.265 that provides support for multiview scalability, and 3D HEVC (3D-HEVC) is an extension of HEVC / H.264 that provides support for 3D video coding that is more advanced and efficient than MV-HEVC. Note that temporal scalability is included as an essential part of the single-layer HEVC codec. The design of the multi-layer extension of HEVC utilizes the idea that the decoded pictures used for inter-layer prediction come only from the same access unit (AU), are treated as long-term reference pictures (LTRP), and are assigned reference indices in the reference picture list together with other temporal reference pictures of the current layer. Inter-layer prediction (ILP) is achieved at the prediction unit (PU) level by setting the value of the reference index to refer to the inter-layer reference pictures in the reference picture list.
[0109] It should be noted that both the resampling of reference pictures and the features of spatial scalability require the resampling of the reference picture or a part thereof. Reference picture resampling (RPR) can be implemented either at the picture level or at the coding block level. However, if RPR is called a coding feature, it is a feature for single-layer coding. Even so, it is possible or desirable from the perspective of codec design to use the same resampling filter for both the RPR feature of single-layer coding and the spatial scalability feature of multi-layer coding.
[0110] FIG. 6 shows an example of layer-based prediction 600 that utilizes an output layer set (OLS) such as to determine an MV in, for example, block compression step 105, block decoding step 113, motion estimation component 221, motion compensation component 219, motion compensation component 321, and / or motion compensation component 421. The layer-based prediction 500 is compatible with one-way and / or bidirectional inter prediction, but is also executed between pictures of different layers. The layer-based prediction in FIG. 6 is similar to that in FIG. 5. Therefore, for simplicity, a complete description of the layer-based prediction is not repeated.
[0111] Some of the layers in the coded video sequence (CVS) 690 of FIG. 6 are included in the OLS. The OLS is a set of layers in which one or more layers are designated as output layers. The output layer is the layer of the OLS that is output. FIG. 6 shows three different OLSs, namely OLS1, OLS2, and OLS3. As shown, OLS1 includes layer N631 and layer N+1 632. OLS2 includes layer N631, layer N+1 632, layer N+2 633, and layer N+1 634. OLS3 includes layer N631, layer N+1 632, and layer N+2 633. Although three OLSs are shown, a different number of OLSs may be used in actual applications.
[0112] Each of the different OLSs may include any number of layers. The different OLSs are generated to attempt to correspond to the coding capabilities of various different devices having various coding capabilities. For example, OLS1 may be generated to correspond to a mobile phone that includes only two layers and has relatively limited coding capabilities. On the other hand, OLS2 may be generated to correspond to a large-screen television that includes four layers and can decode more layers than a mobile phone. OLS3 may be generated to correspond to a personal computer, laptop computer, or tablet computer that includes three layers and can decode more layers than a mobile phone but cannot decode as many layers as a large-screen television.
[0113] The layers in FIG. 6 can all be independent of each other. That is, each layer can be coded without using inter-layer prediction (ILP). In this case, the layers are called peer layers. One or more of the layers in FIG. 6 may be coded using ILP. Whether a layer is a peer layer or whether some of the layers are coded using ILP is signaled by a flag within a video parameter set (VPS). This is discussed more fully below. When some of the layers use ILP, the layer dependencies between the layers are also signaled within the VPS.
[0114] In an embodiment, when a layer is a feedback layer, only one layer is selected for decoding and output. In an embodiment, when several layers use ILP, all layers (e.g., the entire bitstream) are specified to be decoded, and a specific layer among the layers is specified as the output layer. One or more output layers may be, for example, 1) only the top layer, 2) all layers, or 3) a set of the top layer and layers lower than those indicated. For example, when a set of layers lower than the indicated top layer is specified for output by a flag in the VPS, layer N+3 634 from OLS2 (which is the top layer) and layers N631 and N+1 632 (which are lower layers) are output.
[0115] Referring further to FIG. 6, several layers are included in the OLS but are not output or are not used for reference by the output layer. For example, layer N+2 633 is included in OLS2 but is not output. Rather, layer N+3 634 is the top layer in OLS2 and is thus output as the output layer. Further, layer N+2 633 in OLS2 is not used for reference by the output layer layer N+3 634. In fact, layer N+3 634 in OLS2 depends on layer N+2 633, layer N+1 632, and / or layer N631 in OLS2. Thus, layer N+2 633 in OLS2 is called an unnecessary layer. Unfortunately, SHVC and MV-HEVC allow such unnecessary layers to be included in the multi-layer video bitstream. This unnecessarily burdens the coding resources and reduces the coding efficiency.
[0116] The present specification discloses a technique for prohibiting unnecessary layers in a multi-layer video bitstream. That is, all layers included in an output layer set (OLS) are either output layers or used as direct or indirect reference layers of output layers. This avoids having irrelevant information in the coding process and improves coding efficiency. Therefore, a coder / decoder (also known as a "codec") in video coding is improved over the current codec. As a practical matter, the improved video coding process provides a good user experience to users when the video is transmitted, received, and / or viewed.
[0117] FIG. 7 shows an embodiment of a video bitstream 700. As used herein, the video bitstream 700 can also represent a coded video bitstream, a bitstream, or variations thereof. As shown in FIG. 7, the bitstream 700 includes at least one picture unit (PU) 701. Three PUs 701 are shown in FIG. 7, but in actual applications, different numbers of PUs 701 may be present in the bitstream 700. Each PU 701 is a set of NAL units that includes exactly one coded picture (e.g., picture 714) that is consecutive in decoding order and associated with each other according to a specified classification rule.
[0118] In an embodiment, each PU 701 includes one or more of decoding capability information (DCI) 702, a video parameter set (VPS) 704, a sequence parameter set (SPS) 706, a picture parameter set (PPS) 708, a picture header (PH) 712, and a picture 714. Each of DCI 702, VPS 704, SPS 706, and PPS 708 may be collectively referred to as a parameter set. In an embodiment, other parameter sets not shown in FIG. 7, for example, an adaption parameter set (APS), may be included in the bitstream 700, which is a syntax structure including syntax elements applied to zero or more slices determined by zero or more syntax elements found in a slice header.
[0119] DCI702 may also be referred to as a decoding parameter set (DPS) or a decoder parameter set and is a syntax structure that includes syntax elements applied to the entire bitstream. DCI702 includes parameters that remain constant over the lifetime of a video bitstream (e.g., bitstream 700), which can be converted over the lifetime of a session. DCI702 can include profile, level, and sub-profile information to determine an interoperability point of maximum complexity that is guaranteed never to be exceeded, even if splicing of the video sequence occurs within a session. It can further optionally include constraint flags that indicate that the video bitstream is restricted in the use of certain features as indicated by the values of those flags. This allows the bitstream to be labeled as not using certain tools, which in particular enables resource allocation in decoder implementations. Like all parameter sets, DCI702 exists when first referenced, is referenced by the first picture of the video sequence, and must be transmitted between the first NAL units of the bitstream. Multiple DCI702s can exist within a bitstream, but the values of the syntax elements therein cannot conflict when referenced.
[0120] VPS704 includes decoding dependencies or information for the reference picture set configuration of an enhancement layer. VPS704 provides an overall perspective or view of a scalable sequence that includes which type of operation point is provided, the profile, tier, and level of the operation point, as well as several other high-level characteristics of the bitstream that can be used as a basis for session negotiation and content selection.
[0121] In an embodiment, when it is shown that some of the layers use ILP, VPS704 indicates that the total number of OLSs specified by the VPS is equal to the number of layers, that the i-th OLS includes the layers having layer indices from 0 to i including both ends, and that for each OLS, only the topmost layer within the OLS is output.
[0122] In an embodiment, VPS704 includes the syntax and semantics corresponding to the OLSs in the CLVS and / or the video bitstream. The following syntax and semantics corresponding to VPS704 may be used to implement the embodiments disclosed in this specification.
[0123] The syntax for VPS704 may be as follows. [Table 1-1] [Table 1-2]
[0124] The semantics for VPS704 may be as follows. In an embodiment, VPS704 includes one or more of the flags and parameters described below.
[0125] The VPS unprocessed byte sequence payload (RBSP) shall be available for decoding before it is referenced and shall be included in at least one access unit provided through external means or having a TemporalId equal to 0. The VPS NAL unit containing the VPS RBSP shall have a nuh_layer_id equal to vps_layer_id[0]. The nuh_layer_id specifies the identifier of the layer to which a VCL NAL unit belongs or the layer to which a non-VCL NAL unit applies. The value of the nuh_layer_id shall be in the range of 0 to 55, inclusive. The TemporalId is an ID used to uniquely identify a particular NAL unit with respect to other NAL units. The value of the TemporalId shall be the same for all VCL NAL units of an AU. The value of the TemporalId of a coding picture, PU, or AU shall be the value of the TemporalId of the VCL NAL units of the coding picture, PU, or AU.
[0126] All VPS NAL units having a particular value of vps_video_parameter_set_id within a CVS shall have the same content. The vps_video_parameter_set_id provides an identifier for the VPS for reference by other syntax elements. One plus vps_max_layers_minus1 specifies the maximum number of layers allowed within each CVS that references the VPS. One plus vps_max_sub_layers_minus1 specifies the maximum number of temporal sub-layers that may exist within each CVS that references the VPS. The value of vps_max_sub_layers_minus1 shall be in the range of 0 to 6, inclusive.
[0127] The fact that vps_all_independent_layers_flag is equal to 1 specifies that all layers within the CVS are coded independently without using inter-layer prediction. The fact that vps_all_independent_layers_flag is equal to 0 specifies that one or more of the layers within the CVS may use inter-layer prediction. When it does not exist, the value of vps_all_independent_layers_flag is presumed to be equal to 1. When vps_all_independent_layers_flag is equal to 1, the value of vps_independent_layer_flag[i] is presumed to be equal to 1. When vps_all_independent_layers_flag is equal to 0, the value of vps_independent_layer_flag[0] is presumed to be equal to 1.
[0128] vps_layer_id[i] specifies the nuh_layer_id value of the i-th layer. For any two non-negative integer values of m and n, when m is less than n, the value of vps_layer_id[m] should be less than the value of vps_layer_id[n]. The fact that vps_independent_layer_flag[i] is equal to 1 specifies that the layer with index i does not use inter-layer prediction. The fact that vps_independent_layer_flag[i] is equal to 0 specifies that the layer with index i may use inter-layer prediction and that vps_layer_dependency_flag[i] exists within the VPS. When it does not exist, the value of vps_independent_layer_flag[i] is presumed to be equal to 1.
[0129] That vps_direct_dependency_flag[i][j] is equal to 0 specifies that the layer with index j is not a direct reference layer of the layer with index i. That vps_direct_dependency_flag[i][j] is equal to 1 specifies that the layer with index j is a direct reference layer of the layer with index i. When vps_direct_dependency_flag[i][j] does not exist for i and j within the range of 0 to vps_max_layers_minus1 including both ends, it is assumed to be equal to 0.
[0130] The variable DirectDependentLayerIdx[i][j] that specifies the j-th direct dependent layer of the i-th layer is derived as follows.
Number
[0131] The variable GeneralLayerIdx[i] that specifies the layer index of the layer having nuh_layer_id equal to vps_layer_id[i] is derived as follows.
Number
[0132] That each_layer_is_an_ols_flag is equal to 1 specifies that each output layer set contains only one layer and each layer itself within the bitstream is an output layer set where the single layer contained is the only output layer. That each_layer_is_an_ols_flag is equal to 0 means that the output layer may contain more than one layer. When vps_max_layers_minus1 is equal to 0, the value of each_layer_is_an_ols_flag is presumed to be 1. In other cases, when vps_all_independent_layers_flag is equal to 0, the value of each_layer_is_an_ols_flag is presumed to be 0.
[0133] That ols_mode_idc is equal to 0 specifies that the total number of OLSs specified by the VPS is equal to vps_max_layers_minus1 + 1, the i-th OLS contains the layers with layer indices from 0 to i including both ends, and for each OLS, only the topmost layer within the OLS is output. That ols_mode_idc is equal to 1 specifies that the total number of OLSs specified by the VPS is equal to vps_max_layers_minus1 + 1, the i-th OLS contains the layers with layer indices from 0 to i including both ends, and for each OLS, all the layers within the OLS are output. That ols_mode_idc is equal to 2 specifies that the total number of OLSs specified by the VPS is explicitly signaled and for each OLS, the set of the topmost layer and the lower layers within the explicitly signaled OLS is output. The value of ols_mode_idc should be in the range from 0 to 2 inclusive. An ols_mode_idc value of 3 is reserved for future use by ITU-T / ISO / IEC. When vps_all_independent_layers_flag is equal to 1 and each_layer_is_an_ols_flag is equal to 0, the value of ols_mode_idc is presumed to be 2.
[0134] When 1 is added to num_output_layer_sets_minus1, it specifies the total number of OLSs specified by the VPS when ols_mode_idc is equal to 2.
[0135] The variable TotalNumOlss that specifies the total number of OLSs specified by the VPS is derived as follows.
Number
[0136] When ols_mode_idc is equal to 2, layer_included_flag[i][j] specifies whether the j-th layer (i.e., the layer with the nuh_layer_id equal to vps_layer_id[j]) is included in the i-th OLS. That layer_included_flag[i][j] is equal to 1 specifies that the j-th layer is included in the i-th OLS. That layer_included_flag[i][j] is equal to 0 specifies that the j-th layer is not included in the i-th OLS.
[0137] The variable NumLayersInOls[i] that specifies the number of layers in the i-th OLS and the variable LayerIdInOls[i][j] that specifies the nuh_layer_id value of the j-th layer in the i-th OLS are derived as follows.
Number
[0138] The variable OlsLayerIdx[i][j] that specifies the OLS layer index of the layer with the nuh_layer_id equal to LayerIdInOls[i][j] is derived as follows.
Number
[0139] The lowest layer within each OLS should be an independent layer. In other words, for each i in the range of 0 to TotalNumOlss - 1 including both ends, the value of vps_independent_layer_flag[GeneralLayerIdx[LayerIdInOls[i][0]]] should be equal to 1.
[0140] Each layer should be included in at least one OLS specified by the VPS. That is, for each layer having a specific value of nuh_layer_id (nuhLayerId) equal to one of vps_layer_id[k] in the range of 0 to vps_max_layers_minus1 including both ends, there may exist at least one pair of values of i and j. Here, i is in the range of 0 to TotalNumOlss - 1 including both ends, j is in the range of 0 to NumLayersInOls[i] - 1 including the ends, and as a result, the value of LayerIdInOls[i][j] is equal to nuhLayerId.
[0141] Any layer within an OLS should be the output layer of the OLS or a (direct or indirect) reference layer of the output layer of the OLS.
[0142] vps_output_layer_flag[i][j] specifies whether the j-th layer in the i-th OLS is output when ols_mode_idc is equal to 2. That vps_output_layer_flag[i] is equal to 1 specifies that the j-th layer in the i-th OLS is output. That vps_output_layer_flag[i] is equal to 0 specifies that the j-th layer in the i-th OLS is not output. When vps_all_independent_layers_flag is equal to 1 and each_layer_is_an_ols_flag is equal to 0, the value of vps_output_layer_flag[i] is presumed to be equal to 1.
[0143] For the variable OutputLayerFlag[i][j], the value 1 specifies that the j-th layer in the i-th OLS is output, the value 0 specifies that the j-th layer in the i-th OLS is not output, and it is derived as follows.
Number
[0144] Note: The 0-th OLS includes only the lowest layer (i.e., the layer with nuh_layer_id equal to vps_layer_id[0]), and for the 0-th OLS, only the included layers are output.
[0145] That vps_constraint_info_present_flag is equal to 1 specifies that the general_constraint_info( ) syntax structure exists within the VPS. That vps_constraint_info_present_flag is equal to 0 specifies that the general_constraint_info() syntax structure does not exist within the VPS.
[0146] vps_reserved_zero_7bits should be equal to 0 in a bitstream compliant with this version of this VVC draft. Other values of vps_reserved_zero_7bits are reserved for future use by ITU-T / ISO / IEC. In an embodiment, the decoder should ignore the value of vps_reserved_zero_7bits.
[0147] The fact that general_hrd_params_present_flag is equal to 1 specifies that the syntax elements num_units_in_tick and time_scale and the syntax structure general_hrd_parameters() are present within the SPS RBSP syntax structure. The fact that general_hrd_params_present_flag is equal to 0 specifies that the syntax elements num_units_in_tick and time_scale and the syntax structure general_hrd_parameters() are not present within the SPS RBSP syntax structure.
[0148] num_units_in_tick is the number of time units of a clock operating at a frequency of time_scale Hz corresponding to 1 increment of a clock tick counter (referred to as a clock tick). num_units_in_tick should be greater than 0. A clock tick is in seconds and is equal to the quotient of num_units_in_tick divided by time_scale. For example, when the picture rate of a video signal is 25 Hz, time_scale may be equal to 27,000,000 and num_units_in_tick may be equal to 1,080,000, and thus a clock tick may be equal to 0.04 seconds.
[0149] time_scale is the number of time units elapsed in 1 second. For example, a time coordinate system that measures time using a 27 MHz clock has a time_scale of 27,000,000. time_s The value of cale should be greater than 0.
[0150] That vps_extension_flag is equal to 0 specifies that vps_extension_data_flag does not exist within the VPS RBSP syntax structure. That vps_extension_flag is equal to 1 specifies that the vps_extension_data_flag syntax element exists within the VPS RBSP syntax structure.
[0151] vps_extension_data_flag may have any value. Its presence and value do not affect decoder compliance to the profiles specified within this version of this specification. Decoders compliant with this version of this specification should ignore all vps_extension_data_flag syntax elements.
[0152] SPS706 contains data common to all pictures in a sequence of pictures (SOP). SPS706 is a syntax structure that contains zero or more syntax elements applicable to an entire CLVS, as determined by the content of the syntax elements found within the PPS, which is referred to by syntax elements found within each picture header. In contrast, PPS708 contains data common to an entire picture. PPS708 is a syntax structure that contains zero or more syntax elements applicable to an entire coded picture, as determined by the syntax elements found within each picture header (e.g., PH712).
[0153] DCI 702, VPS 704, SPS 706, and PPS 708 are included in different types of Network Abstraction Layer (NAL) units. An NAL unit is a syntax structure that contains an indication of the type of data to follow (e.g., coded video data). NAL units are classified into video coding layer (VCL) NAL units and non-VCL NAL units. A VCL NAL unit contains data representing the values of samples within a video picture, and a non-VCL NAL unit contains any relevant additional information such as parameter sets (important data applicable to a number of VCL NAL units) and supplementary enhancement information (other supplementary data that is not necessary to decode the values of samples within a video picture but may enhance the usefulness of the decoded video signal).
[0154] In an embodiment, DCI 702 is included in a non-VCL NAL unit designated as a DCI NAL unit or a DPS NAL unit. That is, a DCI NAL unit has a DCI NAL unit type (NAL unit type (NUT)), and a DPS NAL unit has a DPS NUT. In an embodiment, VPS 704 is included in a non-VCL NAL unit designated as a DPS NAL unit. Accordingly, a VPS NAL unit has a VPS NUT. In an embodiment, SPS 706 is a non-VCL NAL unit designated as an SPS NAL unit. Accordingly, an SPS NAL unit has an SPS NUT. In an embodiment, PPS 708 is included in a non-VCL NAL unit designated as a PPS NAL unit. Accordingly, a PPS NAL unit has a PPS NUT.
[0155] PH712 is a syntax structure that includes syntax elements applicable to all slices (e.g., slice 718) of a coding picture (e.g., picture 714). In an embodiment, PH712 is a new type of non-VCL NAL unit designated as a PH NAL unit. Thus, the PH NAL unit has a PH NUT (e.g., PH_NUT). In an embodiment, each PU701 includes one and only one PH712. That is, PU701 includes a single or isolated PH712. In an embodiment, exactly one PH NAL unit exists for each picture 701 within bitstream 700.
[0156] In an embodiment, the PH NAL unit associated with PH712 has a temporal ID and a layer ID. The temporal ID identifier indicates the temporal position of the PH NAL unit relative to other PH NAL units within a bitstream (e.g., bitstream 701). The layer ID indicates the layer (e.g., layer 531 or layer 532) that includes the PH NAL unit. In an embodiment, the temporal ID is similar to, but different from, the POC. The POC uniquely identifies each picture in sequence. In a single-layer bitstream, the temporal ID and the POC are the same. In a multi-layer bitstream (e.g., see Figure 5), pictures within the same AU have different POCs but the same temporal ID.
[0157] In an embodiment, the PH NAL unit precedes the VCL NAL unit that includes the first slice 718 of the associated picture 714. This is signaled within the PH 712 and establishes the association between the PH 712 and the slice 718 of the picture 714 associated with the PH 712 without having to have a picture header ID that is referenced from the slice header 720. Thus, all VCL NAL units between two PH 712s belong to the same picture 714, and the picture 714 can be presumed to be associated with the first PH 712 between the two PH 712s. In an embodiment, the first VCL NAL unit following the PH 712 includes the first slice 718 of the picture 714 associated with the PH 712.
[0158] In an embodiment, the PH NAL unit follows a picture level parameter set (e.g., PPS), or a higher level parameter set, e.g., DCI (also known as DPS), VPS, SPS, PPS, etc., and has a temporal ID and a layer ID that are both smaller than the temporal ID and the layer ID of the PH NAL unit, respectively. As a result, those parameter sets are not repeated within a picture or access unit. With this order, the PH 712 can be resolved immediately. That is, the parameter set that includes parameters related to the entire picture is placed before the PH NAL unit in the bitstream. Those that include parameters for a part of the picture are placed after the PH NAL unit.
[0159] As an alternative, the PH NAL unit follows a picture level parameter set and a supplemental enhancement information (SEI) message, or a higher level parameter set such as DCI (also known as DPS), VPS, SPS, PPS, APS, SEI message, etc.
[0160] Picture 714 is an array of luma samples in monochrome format, or an array of luma samples and two corresponding arrays of chroma samples in 4:2:0, 4:2:2, and 4:4:4 color formats. In an embodiment, each PU701 includes one and only one picture 714. Thus, within each PU701, there is only one PH712 and only one picture 714 corresponding to that PH712. That is, PU701 includes a single or isolated picture 714.
[0161] Picture 714 may be a frame or a field. However, in one CVS716, all pictures 714 are frames or all pictures 714 are fields. CVS716 is a coded video sequence for each coded layer video sequence (CLVS) within video bitstream 700. Notably, when video bitstream 700 includes a single layer, CVS716 and CLVS are the same. CVS716 and CLVS are different only when video bitstream 700 includes multiple layers (as shown in FIGS. 5 and 6, for example).
[0162] Each picture 714 includes one or more slices 718. A slice 718 is an integral number of complete tiles or an integral number of consecutive complete CTU rows within a tile of a picture (e.g., picture 714). Each slice 718 is exclusively included in a single NAL unit (e.g., a VCL NAL unit). A tile (not shown) is a rectangular region of CTUs within a particular tile column and a particular tile row within a picture (e.g., picture 714). A CTU (not shown) is a CTB of luma samples, two corresponding CTBs of chroma samples of a picture having three sample arrays, or a CTB of samples of a picture coded using a syntax structure used to code a monochrome picture or three separate color planes and samples. A CTB (not shown) may be an N×N block of samples for some value of N. As a result, the partitioning of components into CTBs is a partition. A block (not shown) is an M×N (M columns × N rows) array of samples (e.g., pixels) or an M×N array of transform coefficients.
[0163] In an embodiment, each slice 718 includes a slice header 720. The slice header 720 is part of the coding slice 718 that includes data elements related to all the tiles or CTU rows within the tiles represented within the slice 718. That is, the slice header 720 includes information regarding the slice 718, such as, for example, the slice type, which reference pictures are used, etc.
[0164] Pictures 714, and their slices 718, contain data related to an image or video to be encoded or decoded. Thus, pictures 714, and their slices 718, may simply be referred to as the payload or data carried within the bitstream 700.
[0165] One skilled in the art will understand that the bitstream 700 may include other parameters and information in actual applications.
[0166] FIG. 8 is an embodiment of a decoding method 800 implemented by a video decoder (e.g., video decoder 400). The method 800 may be executed after a bitstream is received directly or indirectly from a video encoder (e.g., video encoder 300). The method 800 improves the decoding process by suppressing unnecessary layers in a multi-layer video bitstream. That is, all layers included in an output layer set (OLS) are either output layers or are used as direct or indirect reference layers of an output layer. This avoids having information irrelevant in the coding process and improves coding efficiency. Accordingly, a coder / decoder (aka: “codec”) in video coding is improved over the current codec. In practical terms, the improved video coding process provides a good user experience to a user when the video is transmitted, received, and / or viewed.
[0167] In block 802, the video decoder receives a video bitstream including a VPS (e.g., VPS 704) and a plurality of layers (e.g., layer N 631, layer N+1 632, etc.). In an embodiment, none of the layers are output layers of at least one OLS (e.g., OLS1, OLS2, etc.) or direct reference layers of any other layer. That is, each layer of at least one OLS may be an output layer (e.g., layer N+2 633 in OLS3) or a direct reference layer of any other layer (e.g., layer N+1 632 and layer N 631 in OLS3). Any layer that is not an output layer or a layer directly referenced by another layer is unnecessary and is thus excluded from the CVS (e.g., 690) and / or the video bitstream.
[0168] In an embodiment, the video decoder expects that a layer is neither an output layer in at least one output layer set (OLS) nor a direct reference layer of any other layer based on VVC or some other standard as described above. However, if the decoder determines that this condition is not true, the decoder may detect an error, signal the error, request that the received bitstream (or a portion thereof) be retransmitted, or take some other corrective means to ensure that a compliant bitstream is received.
[0169] In an embodiment, the VPS includes a layer used as a reference flag (e.g., LayerUsedAsRefLayerFlag) and a layer used as an output layer flag (e.g., LayerUsedAsOutputLayerFlag), and the value of the layer used as a reference flag (e.g., [i]) and the value of the layer used as an output layer flag (e.g., [i]) are both non-zero.
[0170] In an embodiment, each layer among a plurality of layers includes a set of a video coding layer (VCL) network abstraction layer (NAL) unit all having a specific value of a layer identifier (ID) and an associated non-VCL NAL unit. In an embodiment, at least one OLS includes one or more output layers. In an embodiment, for each of the plurality of layers having a layer ID of a specific value specified within the VPS, one of the layers within at least one OLS should also have a layer ID of the specific value.
[0171] At block 804, the video decoder decodes a picture (e.g., picture 615) from one of the multiple layers. In an embodiment, the picture is included in the output layer of at least one OLS. In an embodiment, method 800 further includes, prior to the decoding step, a step of selecting an output layer from at least one OLS. In an embodiment, method 800 further includes, after the output layer is selected, a step of selecting a picture from the output layer.
[0172] Once the picture is decoded, the picture may be used to produce or generate an image or video sequence for display to a user on a display or screen of an electronic device (e.g., smartphone, tablet, laptop, personal computer, etc.).
[0173] FIG. 9 is an embodiment of a method 900 for encoding a video bitstream implemented by a video encoder (e.g., video encoder 300). Method 900 may be executed when a picture (e.g., from a video) is encoded into a video bitstream and transmitted to a video decoder (e.g., video decoder 400). Method 900 improves the decoding process by suppressing unnecessary layers in a multi-layer video bitstream. That is, all layers included in an output layer set (OLS) are either output layers or are used as direct or indirect reference layers of output layers. This avoids having irrelevant information in the coding process and improves coding efficiency. Accordingly, the coder / decoder (alias: "codec") in video coding is improved over the current codec. As a practical matter, the improved video coding process provides a good user experience to the user when the video is transmitted, received, and / or viewed.
[0174] At block 902, the video encoder generates a VPS (e.g., VPS 704) that specifies a plurality of layers (e.g., N631, layer N+1 632, etc.) and one or more OLSs (e.g., OLS1, OLS2, etc.). In an embodiment, each layer from the plurality of layers is included in at least one of the OLSs specified by the VPS. In an embodiment, no layer is an output layer of at least one OLS (e.g., OLS1, OLS2, etc.) or a direct reference layer of any other layer. That is, each layer of at least one OLS may be an output layer (e.g., layer N+2 633 in OLS3) or a direct reference layer of any other layer (e.g., layer N+1 632 and layer N 631 in OLS3). Any layer that is not an output layer or a layer directly referenced by another layer is unnecessary and is thus excluded from the CVS (e.g., 690) and / or the video bitstream. In an embodiment, the video encoder is constrained such that no layer is an output layer of at least one output layer set (OLS) or a direct reference layer of any other layer. That is, the video encoder is required to encode such that no layer is an output layer of at least one output layer set (OLS) or a direct reference layer of any other layer. Such a constraint or requirement ensures that the bitstream conforms to, for example, VVC or some other standard.
[0175] In an embodiment, each of the one or more OLSs includes one or more output layers, and each of the output layers includes one or more pictures. In an embodiment, each layer among the plurality of layers includes a set of a video coding layer (VCL) network abstraction layer (NAL) unit having a specific value of a layer identifier (ID) and an associated non-VCL NAL unit.
[0176] In an embodiment, one of the OLSs includes two output layers, and one of the two output layers refers to the other of the two output layers. In an embodiment, the VPS includes a layer used as a reference flag and a layer used as an output layer flag, and the values of both the layer used as a reference flag and the layer used as an output layer flag are not zero.
[0177] In an embodiment, a hypothetical reference decoder (HRD) disposed within the encoder checks to ensure that all layers included in an output layer set (OLS) are either output layers or are used as direct or indirect reference layers for output layers. When the HRD finds a layer that is unnecessary as described in this specification, the HRD returns a compliance test error. That is, the HRD compliance test ensures that there are no unnecessary layers. Thus, the encoder encodes in accordance with the requirement of no unused layers, while the HRD enforces this requirement.
[0178] At block 904, the video encoder encodes a plurality of layers and a VPS into a video bitstream. At block 906, the video encoder stores the video bitstream for communication to a video decoder. The video bitstream may be stored in memory until the video bitstream is transmitted towards the video decoder. When received by the video decoder, the encoded video bitstream may be decoded to produce or generate an image or video sequence for display to a user on a display or screen of an electronic device (such as, for example, a smartphone, tablet, laptop, personal computer, etc.).
[0179] FIG. 10 is a schematic diagram of a video coding apparatus 1000 (e.g., video encoder 300 or video decoder 400) according to an embodiment of the present disclosure. The video coding apparatus 1000 is suitable for implementing the embodiments of the disclosure as described herein. The video coding apparatus 1000 includes an ingress port 1010 and receiver units (Rx) 1020 for receiving data, a processor, logic unit, or central processing unit (CPU) 1030 for processing data, transmitter units (Tx) 1040 and an egress port 1050 for transmitting data, and a memory 1060 for storing data. The video coding apparatus 1000 may also include optical-to-electrical (OE) components and electrical-to-optical (EO) components for ingress or egress of optical or electrical signals connected to the ingress port 1010, the receiver units 1020, the transmitter units 1040, and the egress port 1050.
[0180] Processor 1030 is implemented by hardware and software. Processor 1030 may be implemented as one or more CPU chips, cores (e.g., multi-core processors), field-programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), and digital signal processors (DSPs). Processor 1030 communicates with ingress port 1010, receiver unit 1020, transmitter unit 1040, ingress port 1050, and memory 1060. Processor 1030 includes coding module 1070. Coding module 1070 implements the embodiments of the above disclosure. For example, coding module 1070 implements, processes, prepares, or provides various codec functions. What is included in coding module 1070 thus provides a substantial improvement to the functionality of video coding device 1000 and results in a conversion of video coding device 1000 to different states. Alternatively, coding module 1070 is implemented as instructions stored in memory 1060 and executed by processor 1030.
[0181] Video coding device 1000 may also include an input and / or output (I / O) device 1080 that communicates data to and from a user. I / O device 1080 may include output devices such as a display for displaying video data, a speaker for outputting audio data, etc. I / O device 1080 may also include input devices such as a keyboard, a mouse, a trackball, etc., and / or corresponding interfaces that interface with such output devices.
[0182] Memory 1060 may include one or more disks, tape drives, and solid state drives, and may be used as an overflow data storage device for storing a program when the program is selected for execution and for storing instructions and data read during execution of the program. Memory 1060 may be volatile and / or non-volatile, and may be read-only memory (ROM), random access memory (RAM), ternary content-addressable memory (TCAM), and / or static random-access memory (SRAM).
[0183] FIG. 11 is a schematic diagram of an embodiment of coding means 1100. In the embodiment, coding means 1100 is implemented within a video coding device 1102 (e.g., video encoder 300 or video decoder 400). Video coding device 1102 includes receiving means 1101. Receiving means 1101 is configured to receive a picture to be coded or to receive a bitstream to be decoded. Video coding device 1102 includes transmitting means 1107 coupled to receiving means 1101. Transmitting means 1107 is configured to transmit a bitstream to a decoder or to transmit a decoded picture to a display means (e.g., one of I / O devices 1080).
[0184] Video coding device 1102 includes storage means 1103. Storage means 1103 is coupled to at least one of receiving means 1101 or transmitting means 1107. Storage means 1103 is configured to store instructions. Video coding device 1102 further includes processing means 1105. Processing means 1105 is coupled to storage means 1103. Processing means 1105 is configured to execute instructions stored in storage means 1103 in order to execute the methods disclosed herein.
[0185] It should be further understood that the steps of the exemplary methods described herein need not necessarily be performed in the order described, and the order of such method steps should be understood to be merely illustrative. Similarly, additional steps may be included in such methods, and certain steps may be omitted or combined in methods according to various embodiments of the present disclosure.
[0186] Although several embodiments have been provided in the present disclosure, it should be understood that the disclosed systems and methods may be implemented in many other specific forms without departing from the spirit or scope of the present disclosure. The examples of the present invention are for illustrative purposes and should be considered non-limiting, and are not intended to be limited to the details given herein. For example, various elements or components may be combined or integrated into another system, or certain functions may be omitted or not implemented.
[0187] Furthermore, the techniques, systems, subsystems, and methods described and shown in various embodiments may be combined or integrated with other systems, modules, techniques, or methods without departing from the scope of the present disclosure. Other items shown or discussed as being coupled to, directly coupled to, or communicating with each other may be indirectly coupled or communicate through some interface, device, or intermediate component, whether electrically, mechanically, or otherwise. Other examples of changes, substitutions, and modifications can be ascertained by those skilled in the art and can be made without departing from the spirit and scope disclosed herein.
Claims
1. 1. A method of decoding implemented by a video decoder, the method comprising: receiving, at the video decoder, a video bitstream including a video parameter set (VPS) and a plurality of layers, wherein none of the layers is an output layer of at least one output layer set (OLS) or a direct reference layer of any other layer, the VPS including an ols_mode_idc and a vps_max_layers_minus1, the ols_mode_idc equal to 0 specifying that the total number of OLSs specified by the VPS is equal to vps_max_layers_minus1+1, the i-th OLS including layers with layer indices from 0 to i (inclusive), and for each OLS, only the top layer of the OLS is output; and the video decoder decoding a picture from one of the plurality of layers. method.
2. 1. A method of encoding implemented by a video encoder, the method comprising: generating a video parameter set (VPS) specifying a number of layers and at least one output layer set (OLS), wherein the video encoder is constrained such that no layer is an output layer of the at least one OLS or a direct reference layer of any other layer, the VPS including ols_mode_idc and vps_max_layers_minus1, wherein the ols_mode_idc equal to 0 specifies that the total number of OLSs specified by the VPS is equal to vps_max_layers_minus1+1, the i-th OLS includes layers with layer indices from 0 to i (inclusive), and for each OLS, only the top layer of the OLS is output; and the video encoder encoding the plurality of layers and the VPS into a video bitstream. method.
3. A decoding device, the decoding device comprising: a receiver configured to receive a video bitstream including a video parameter set (VPS) and a number of layers, none of which is an output layer of at least one OLS or a direct reference layer of any other layer, the VPS including an ols_mode_idc and a vps_max_layers_minus1, the ols_mode_idc equal to 0 specifying that a total number of OLSs specified by the VPS is equal to vps_max_layers_minus1+1, the i-th OLS including layers with layer indices from 0 to i inclusive, and for each OLS, only the top layer of the OLS is output; a memory coupled to the receiver, the memory storing instructions; a processor coupled to the memory and configured to execute the instructions to cause the decoding device to decode a picture from one of the plurality of layers to obtain a decoded picture. Decoding device.
4. An encoding device, the encoding device comprising: a memory containing instructions; a processor coupled to the memory; The processor executes the instructions to cause the encoding device to: generating a video parameter set (VPS) specifying a number of layers and one or more output layer sets (OLS), wherein the encoding device is constrained such that no layer is an output layer of at least one OLS or a direct reference layer of any other layer, the VPS including ols_mode_idc and vps_max_layers_minus1, wherein the ols_mode_idc equal to 0 specifies that the total number of OLSs specified by the VPS is equal to vps_max_layers_minus1+1, the i-th OLS includes layers with layer indices from 0 to i (inclusive), and for each OLS, only the top layer of the OLS is output; encoding the plurality of layers and the VPS into a video bitstream. Encoding device.
5. 1. A coding device, the coding device comprising: a receiver configured to receive and encode a picture or to receive and decode a bitstream; a transmitter coupled to the receiver, the transmitter configured to transmit the bitstream to a decoder or to transmit a decoded image to a display; a memory coupled to at least one of the receiver or the transmitter, the memory configured to store instructions; a processor coupled to the memory and configured to execute the instructions stored in the memory to perform the method of claims 1 and 2. Coding device.
6. 1. A system comprising: An encoder; and a decoder in communication with said encoder, said encoder or decoder comprising a decoding device, an encoding device or a coding device according to any one of claims 3 to 5. system.
7. An encoder comprising processing circuitry for carrying out the method of claim 2.
8. A decoder comprising processing circuitry for performing the method of claim 1.
9. A computer program comprising a program code for carrying out the method according to claim 1 or 2.
10. 1. A storage medium storing an encoded bitstream for a video signal, the encoded bitstream including a video parameter set (VPS) and a number of layers, none of which is an output layer of at least one output layer set (OLS) or a direct reference layer of any other layer, the VPS including an ols_mode_idc and a vps_max_layers_minus1, the ols_mode_idc equal to 0 specifying that a total number of OLSs specified by the VPS is equal to vps_max_layers_minus1+1, the i-th OLS including layers with layer indices from 0 to i inclusive, and for each OLS, only the top layer of the OLS is output. storage medium.
11. A terminal, the terminal comprising one or more processors, a memory and a communication interface, the memory and the communication interface being connected to the one or more processors, the terminal communicating with another device via the communication interface, the memory being configured to store computer program code, the computer program code comprising instructions, the one or more processors executing the instructions causing the terminal to perform the method of claim 1 or 2. Terminal.
12. a data structure for use by a decoder, the data structure including an encoded bitstream, the bitstream including a video parameter set (VPS) and a number of layers, none of which is an output layer of at least one output layer set (OLS) or a direct reference layer of any other layer, the VPS including an ols_mode_idc and a vps_max_layers_minus1, the ols_mode_idc equal to 0 specifying that a total number of OLSs specified by the VPS is equal to vps_max_layers_minus1+1, the i-th OLS including layers with layer indices from 0 to i inclusive, and for each OLS, only a top layer of the OLS is output, and a video decoder decodes a picture from one of the layers. Data structure.
13. An apparatus for storing a bitstream, the apparatus including at least one storage medium and at least one communication interface; the at least one communication interface is configured to receive or transmit the bitstream, and the at least one storage medium is configured to store the bitstream; the bitstream includes a video parameter set (VPS) and a number of layers, none of which is an output layer of at least one output layer set (OLS) or a direct reference layer of any other layer, the VPS includes an ols_mode_idc and a vps_max_layers_minus1, the ols_mode_idc equal to 0 specifying that the total number of OLSs specified by the VPS is equal to vps_max_layers_minus1+1, the i-th OLS includes layers with layer indices from 0 to i (inclusive), and for each OLS, only the top layer of the OLS is output; Device.
14. 1. A method for storing a bitstream, the method comprising: receiving or transmitting a bitstream via a communication interface; storing the bitstream in one or more storage media, the bitstream including a video parameter set (VPS) and a number of layers, none of which is an output layer of at least one output layer set (OLS) or a direct reference layer of any other layer, the VPS including an ols_mode_idc and a vps_max_layers_minus1, the ols_mode_idc equal to 0 specifying that the total number of OLSs specified by the VPS is equal to vps_max_layers_minus1+1, the i-th OLS including layers with layer indices from 0 to i inclusive, and for each OLS, only the top layer of the OLS is output. method.
15. 1. An apparatus for transmitting a bitstream, the apparatus comprising: at least one storage medium configured to store at least one bitstream, the bitstream including a video parameter set (VPS) and a number of layers, none of which is an output layer of at least one output layer set (OLS) or a direct reference layer of any other layer, the VPS including an ols_mode_idc and a vps_max_layers_minus1, the ols_mode_idc equal to 0 specifying that a total number of OLSs specified by the VPS is equal to vps_max_layers_minus1+1, the i-th OLS including layers with layer indices from 0 to i inclusive, and for each OLS, only the top layer of the OLS is output; at least one processor configured to obtain one or more bitstreams from one of the at least one storage medium and transmit the one or more bitstreams to a destination device; Device.
16. 1. A method for transmitting a bitstream, the method comprising: storing at least one bitstream on at least one storage medium, the bitstream including a video parameter set (VPS) and a number of layers, none of which is an output layer of at least one output layer set (OLS) or a direct reference layer of any other layer, the VPS including an ols_mode_idc and a vps_max_layers_minus1, the ols_mode_idc equal to 0 specifying that the total number of OLSs specified by the VPS is equal to vps_max_layers_minus1+1, the i-th OLS including layers with layer indices from 0 to i (inclusive), and for each OLS, only the top layer of the OLS is output; obtaining one or more bitstreams from one of said at least one storage medium; transmitting the one or more bitstreams to a destination device. method.
17. 1. A system for processing a bitstream, the system including an encoding device, one or more storage devices, and a decoding device; the encoding device is configured to obtain a video signal and encode the video signal to obtain one or more bitstreams, the bitstream including a video parameter set (VPS) and a number of layers, none of which is an output layer of at least one output layer set (OLS) or a direct reference layer of any other layer, the VPS including ols_mode_idc and vps_max_layers_minus1, ols_mode_idc equal to 0 specifies that a total number of OLSs specified by the VPS is equal to vps_max_layers_minus1+1, the i-th OLS includes layers with layer indices from 0 to i (inclusive), and for each OLS, only the top layer of the OLS is output; the one or more storage devices are used to store the one or more bitstreams; the decoding device is used for decoding the one or more bitstreams; system.
Citation Information
Patent Citations
Image decoder and image encoder
JP2015195543A
Method and an apparatus and a computer program for encoding media content
US20170347026A1
Image decoding device and image decoding method
WO2015137432A1