Same picture order count (POC) numbering for scalability support
By adopting weighted prediction technology and zero-value picture sequential counting offset in the video processing system, the problems of heavy burden of video data processing and poor scalability in the prior art are solved, and efficient and flexible video data encoding and decoding are achieved.
Patent Information
- Application Number
- CN202080043476.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-06-01
- Filing Date
- 2020-06-02
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2040-06-02
AI Technical Summary
When processing and storing high-quality video data, existing video decoding technologies face the problems of large amount of data and heavy processing burden, and it is difficult to effectively realize the scalability of video decoding.
By introducing weighted prediction technology in the video processing system, the reference picture is identified using zero-value picture sequential count offset, and the encoding and decoding process of video data is optimized to realize the scalability of the system.
It improves the efficiency and flexibility of the video processing system, reduces the bit rate of data transmission and storage, and enhances the fidelity and scalability of video quality.
Smart Images

Figure CN113950835B_ABST
Abstract
Description
Technical Field
[0001] The present application is related to video coding. More specifically, the present application relates to systems, methods, and computer-readable media that provide scalability support for video coding systems. Background Art
[0002] Many devices and systems allow video data to be processed and output for consumption. Digital video data includes large amounts of data to meet the needs of consumers and video providers. For example, consumers of video data expect the highest quality video with high fidelity, resolution, frame rate, etc. As a result, the large amounts of video data required to meet these needs place a burden on the communication networks and devices that process and store video data.
[0003] Various video coding techniques can be used to compress video data. Video coding is performed according to one or more video coding standards. For example, video coding standards include general video coding (VVC), high efficiency video coding (HEVC), advanced video coding (AVC), moving picture experts group (MPEG) coding and others. Video coding can utilize prediction methods (e.g., inter-frame prediction, intra-frame prediction, etc.), which utilize redundancy present in video images or sequences. An important goal of video coding technology is to compress video data into a form using a lower bit rate while avoiding or minimizing degradation of video quality. As evolving video services become available, coding techniques with better decoding efficiency are needed. Summary of the invention
[0004] This article describes systems and methods for improved video processing. Digital video data includes a large amount of data to meet the needs of consumers and video providers, which places a burden on communication networks and devices that process and store video data. Some examples of video processing use video compression techniques that utilize predictions to efficiently encode and decode video data. For example, the prediction can be determined as the difference between the pixel values in the block being encoded and the predicted block. The prediction error can also be referred to as a residual. The video encoder can also apply a transform to the prediction error (e.g., a discrete cosine transform (DCT) or other appropriate transform) to generate transform coefficients. After the transform, the video encoder can quantize the transform coefficients. The quantized transform coefficients and motion vectors can be represented using syntax elements, and together with control information, form a decoded representation of the video sequence. In some instances, the video encoder can entropy decode the syntax elements to further reduce the number of bits required for their representation.
[0005] In some examples, the prediction may use a reference frame in the same layer as the frame being analyzed. Such a prediction may include copies of the same frame with different sizes or resolutions. In some such examples, the reference frame may be identified within a reference list and referenced using a picture sequence count offset value. For weighted prediction, an example reference frame within the same layer as the frame being encoded or decoded may be referenced using a zero-valued picture sequence count offset. In some such examples, when weighted prediction is not used, the encoder or decoder may interpret the picture sequence count offset value as plus or minus 1 from the value transmitted with the signal. Such operations improve the efficiency of the network and encoding or decoding devices by reducing signaling and providing efficient processing for operations for identifying and using reference frames.
[0006] In one illustrative example, an apparatus for decoding video data is provided. The apparatus includes: a memory; and a processor implemented in a circuit. The processor is configured to: obtain at least a portion of a picture included in a bitstream. The processor is also configured to: determine, based on the bitstream, that weighted prediction is enabled for at least a portion of the picture; and based on determining that weighted prediction is enabled for at least a portion of the picture, identify a zero-valued picture order count offset for indicating a reference picture from a reference picture list. The processor is also configured to: reconstruct at least a portion of the picture using at least a portion of the reference picture identified by the zero-valued picture order count offset.
[0007] In another example, a method for processing video data is provided. The method includes obtaining at least a portion of a picture from a bitstream. The method also includes determining, based on the bitstream, that weighted prediction is enabled for a portion of the picture. The method also includes identifying a zero-valued picture order count offset for indicating a reference picture from a reference picture list based on determining that weighted prediction is enabled for a portion of the picture. The method also includes reconstructing at least a portion of the picture using at least a portion of the reference picture identified by the zero-valued picture order count offset.
[0008] In another example, a computer-readable storage medium stores instructions that, when executed, cause one or more processors of a device for decoding video data to perform the following operations: obtain at least a portion of a picture included in a bitstream; determine, based on the bitstream, that weighted prediction is enabled for at least a portion of the picture; based on determining that weighted prediction is enabled for at least a portion of the picture, identify a zero-valued picture order count offset for indicating a reference picture from a reference picture list; and reconstruct at least a portion of the picture using at least a portion of the reference picture identified by the zero-valued picture order count offset.
[0009] In another example, an apparatus for decoding video data is provided. The apparatus includes: a unit for obtaining at least a portion of a picture included in a bitstream; a unit for determining, from the bitstream, that weighted prediction is enabled for at least a portion of the picture; a unit for identifying a zero-valued picture order count offset for indicating a reference picture from a reference picture list based on the determination that weighted prediction is enabled for at least a portion of the picture; and a unit for reconstructing at least a portion of the picture using at least a portion of the reference picture identified by the zero-valued picture order count offset.
[0010] In some cases, the methods, apparatus, and computer-readable storage media described above include: obtaining multiple portions of a picture from a bitstream, the multiple portions including at least a portion of the picture reconstructed using a reference picture identified by a zero-valued picture order count offset; identifying multiple corresponding picture order count offsets for the multiple portions of the picture, wherein the zero-valued picture order count offset is a corresponding picture order count offset associated with the reference picture among the multiple corresponding picture order count offsets; and reconstructing the picture using the multiple reference pictures identified by the multiple corresponding picture order count offsets.
[0011] In some cases, the methods, apparatus, and computer-readable storage media described above include: determining to enable weighted prediction by parsing a bitstream to identify one or more weighted prediction flags for a picture. In some cases, the one or more weighted prediction flags for a portion of a picture include a sequence parameter set weighted prediction flag and a picture parameter set weighted prediction flag. In some cases, the sequence parameter set weighted prediction flag and the picture parameter set weighted prediction flag are flags for unidirectional prediction frames. In some cases, the sequence parameter set weighted prediction flag and the picture parameter set weighted prediction flag are flags for bidirectional prediction frames. In some cases, the picture parameter set weighted prediction flag is constrained by the sequence parameter set weighted prediction flag.
[0012] In some cases, the methods, devices, and computer-readable storage media described above include: obtaining at least a portion of a second picture included in a bitstream; determining, based on the bitstream, that weighted prediction is disabled for a portion of the second picture; and based on determining that weighted prediction is disabled for at least a portion of the second picture, determining not to allow a second zero-valued picture order count offset to indicate a second reference picture for a portion of the second picture from a reference picture list.
[0013] In some cases, the methods, apparatus, and computer-readable storage media described above include: based on determining that a second zero-valued picture order count offset is not allowed, parsing a syntax element by reconstructing a value of a syntax element of a bitstream indicating a picture order count offset for a second reference picture to a signaled value plus 1 based on disabling weighted prediction for the second picture; and reconstructing the second picture using the reconstructed value of the syntax element.
[0014] In some cases, when weighted prediction is disabled, the value of the reference picture syntax element specifies the absolute difference between the picture order count value of the second picture and the previous reference picture entry in the reference picture list for the second reference picture as a value plus one.
[0015] In some cases, the reference picture is associated with a first set of weights, and the picture is associated with a second set of weights that is different from the first set of weights. In some cases, the reference picture has a different size than the picture. In some cases, the reference picture is signaled as a short-term reference picture with a picture order count least significant bit value equal to zero. In some cases, the picture is a non-instantaneous decoding refresh (non-IDR) picture. In some cases, at least a portion of the picture is a slice. A slice may include multiple blocks of a picture. In some cases, at least a portion of the picture is a block of the picture (e.g., a coding tree unit (CTU), a macroblock, a coding unit or block, a prediction unit or block, or other types of blocks of a picture).
[0016] In some cases, the methods, apparatus, and computer-readable storage media described above include determining, from a bitstream, a layer identifier for a reference picture, the layer identifier indicating a layer of a layer index for inter-layer prediction. In some examples, when the layer identifier is different from the current layer identifier, a zero-valued picture order count offset is identified by inferring a zero value based on that a picture order count for the reference picture is not signaled.
[0017] In another illustrative example, an apparatus for encoding video data is provided. The apparatus includes: a memory; and a processor implemented in circuitry and configured to: identify at least a portion of a picture. The processor is further configured to: select weighted prediction as enabled for a portion of the picture. The processor is further configured to: identify a reference picture for a portion of the picture; and generate a zero-valued picture order count offset for indicating a reference picture from a reference picture list. The processor is further configured to: generate a bitstream, the bitstream including a portion of the picture and a zero-valued picture order count offset associated with the portion of the picture.
[0018] In another example, a method for encoding video data is provided. The method includes: identifying at least a portion of a picture. The method includes: determining that weighted prediction is selected to be enabled for the portion of the picture. The method includes: identifying a reference picture for the portion of the picture; and generating a zero-valued picture order count offset for indicating the reference picture from a reference picture list. The method also includes: generating a bitstream, wherein the bitstream includes the portion of the picture and the zero-valued picture order count offset as associated with the portion of the picture.
[0019] In another illustrative example, a computer-readable storage medium stores instructions that, when executed, cause one or more processors of a device for encoding video data to: identify at least a portion of a picture; select weighted prediction as enabled for the portion of the picture; identify a reference picture for the portion of the picture; generate a zero-valued picture order count offset for indicating a reference picture from a reference picture list; and generate a bitstream, the bitstream including the portion of the picture and the zero-valued picture order count offset as associated with the portion of the picture.
[0020] In another illustrative example, an apparatus for encoding video data is provided. The apparatus includes: means for identifying at least a portion of a picture; means for selecting weighted prediction as enabled for the portion of the picture; means for identifying a reference picture for the portion of the picture; means for generating a zero-valued picture order count offset for indicating a reference picture from a reference picture list; and means for generating a bitstream, the bitstream including the portion of the picture and the zero-valued picture order count offset as associated with the portion of the picture.
[0021] In another example, a method for encoding video data is provided. The method includes: identifying at least a portion of a picture. The method includes: determining that weighted prediction is selected to be enabled for the portion of the picture. The method includes: identifying a reference picture for the portion of the picture; and generating a zero-valued picture order count offset for indicating the reference picture from a reference picture list. The method also includes: generating a bitstream, wherein the bitstream includes the portion of the picture and the zero-valued picture order count offset as associated with the portion of the picture.
[0022] In some aspects, the apparatus for decoding video data and / or the apparatus for encoding video data includes a camera, a mobile device (e.g., a mobile phone or so-called "smart phone" or other mobile device), a wearable device, an extended reality device (e.g., a virtual reality (VR) device, an augmented reality (AR) device, or a mixed reality (MR) device), a personal computer, a laptop computer, a server computer, or other device. In some aspects, the apparatus for decoding video data and / or the apparatus for encoding video data includes a camera or multiple cameras for capturing one or more images. In some aspects, the apparatus for decoding video data and / or the apparatus for encoding video data includes a display for displaying one or more images, notifications, and / or other displayable data.
[0023] The aspects described above relating to any of the method, apparatus, and computer-readable medium may be used alone or in any suitable combination.
[0024] This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used alone to determine the scope of the claimed subject matter. The subject matter should be understood by reference to appropriate portions of the entire specification of this patent, any or all drawings, and each claim.
[0025] The foregoing, together with other features and embodiments, will become more apparent after reference to the following description, claims, and accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Illustrative examples of the present application are described in detail below with reference to the following drawings:
[0027] Figure 1 is a block diagram illustrating examples of encoding devices and decoding devices according to some examples.
[0028] Figure 2 is a schematic diagram illustrating an example implementation of a filter unit for performing the techniques of this disclosure.
[0029] Figure 3 is a diagram illustrating examples of multiple access units (AUs) with different picture order counts (POCs) according to some examples.
[0030] Figure 4 is a flow chart illustrating an example method according to various examples described herein.
[0031] Figure 5 is a flow chart illustrating an example method according to various examples described herein.
[0032] Figure 6 is a block diagram illustrating an example video encoding device according to some examples.
[0033] Figure 7 is a block diagram illustrating an example video decoding device according to some examples. DETAILED DESCRIPTION
[0034] Some aspects and examples of the present disclosure are provided below. As will be apparent to those skilled in the art, some of these aspects and examples may be applied independently, and some of them may be applied in combination. In the following description, for the purpose of explanation, specific details are set forth to provide a thorough understanding of the examples of the present application. However, it will be apparent that each example may be implemented without these specific details. The accompanying drawings and description are not intended to be restrictive.
[0035] The description that follows provides only exemplary examples and is not intended to limit the scope, applicability or configuration of the present disclosure. Specifically, the description that follows the exemplary examples will provide those skilled in the art with a description that enables the implementation of the exemplary examples. It should be understood that various changes may be made to the function and arrangement of elements without departing from the spirit and scope of the present application as set forth in the appended claims.
[0036] Video decoding devices implement video compression techniques to efficiently encode and decode video data. Video compression techniques may include applying different prediction modes, including spatial prediction (e.g., intra-frame prediction or intra-prediction), temporal prediction (e.g., inter-frame prediction or inter-prediction), inter-layer prediction (across different layers of video data), and / or other prediction techniques for reducing or removing redundancy inherent in video sequences. The video encoder may partition each picture of the original video sequence into a plurality of rectangular areas, which are referred to as video blocks or coding units (described in more detail below). These video blocks may be encoded using a specific prediction mode.
[0037] Video blocks may be divided into one or more groups of smaller blocks in one or more ways. Blocks may include coding tree blocks, prediction blocks, transform blocks, and / or other appropriate blocks. Unless otherwise specified, references to "blocks" may generally refer to such video blocks (e.g., coding tree blocks, coding blocks, prediction blocks, transform blocks, or other appropriate blocks or sub-blocks, as will be understood by those of ordinary skill in the art). Further, each of these blocks may also be referred to interchangeably herein as a "unit" (e.g., coding tree unit (CTU), decoding unit, prediction unit (PU), transform unit (TU), etc.). In some cases, a unit may indicate a decoding logic unit encoded in a bitstream, and a block may indicate a portion of a video frame buffer to which the process is directed.
[0038] For inter-prediction mode, the video encoder may search for a block similar to the block being encoded in a frame (or picture) located at another temporal position, which is called a reference frame or reference picture. The video encoder may limit the search to a certain spatial displacement from the block to be encoded. In some systems, the best match is located using a two-dimensional (2D) motion vector that includes a horizontal displacement component and a vertical displacement component. For intra-prediction mode, the video encoder may use spatial prediction techniques to form a prediction block based on data from previously encoded neighboring blocks within the same picture.
[0039] The video encoder can determine the prediction error. For example, the prediction can be determined as the difference between the pixel values in the block being encoded and the pixel values in the prediction block. The prediction error can also be referred to as a residual. The video encoder can also apply a transform to the prediction error (e.g., a discrete cosine transform (DCT) or other appropriate transform) to generate transform coefficients. After the transform, the video encoder can quantize the transform coefficients. In some examples, the quantized transform coefficients and motion vectors can be represented using syntax elements, and together with the control information, a decoded representation of the video sequence is formed. In some instances, the video encoder can entropy decode the syntax elements to further reduce the number of bits used for its representation.
[0040] The video decoder can use the syntax elements and control information discussed above to construct prediction data (e.g., a prediction block) for decoding the current frame. For example, the video decoder can add the prediction block and the compressed prediction error. The video decoder can determine the compressed prediction error by weighting the transform basis function using the quantization coefficients. The difference between the reconstructed frame and the original frame is called the reconstruction error.
[0041] As mentioned above, the reference picture may be one that is used when performing inter-frame prediction. To identify a reference picture, a picture order count (POC) offset value or a "delta POC" value may be used. The POC value identifies the selected picture based on the difference (e.g., offset or delta) between the picture order count value for the selected picture and the picture order count value for the previous picture or the original picture. Some reference pictures may have an incremental POC value equal to zero, such as when multiple reference pictures have the same POC value (e.g., the same picture is used as a reference multiple times). For example, a zero incremental POC value may be used with weighted prediction. However, in some examples, when weighted prediction is disabled, a zero incremental POC value may not be used. Signaling a non-zero incremental POC value uses additional bits, which may waste resources when the incremental POC values are typically the same non-zero value.
[0042] Described herein are systems and techniques that can reduce the bit rate for signaling encoded video data by determining when a related mode (e.g., weighted prediction) is enabled or disabled and interpreting a signaled incremental POC value based on the determination. In some examples, a flag is used to determine the current mode, such as a weighted prediction flag to indicate that weighted prediction is enabled. In such examples, when weighted prediction is enabled, a signaled incremental POC value of zero can be used as the actual incremental POC for reconstructing a picture. When weighted prediction is disabled, the signaled incremental POC of zero can be modified to determine an actual (non-zero) incremental POC value. As described above, such an implementation can improve the operation of the system by reducing the bit rate used for image signaling without reducing image quality.
[0043] The techniques described herein can be applied to any video codec in an existing video codec (e.g., High Efficiency Video Coding (HEVC), Advanced Video Coding (AVC), or other suitable existing video codecs), and / or can be an efficient decoding tool for any video decoding standard being developed and / or future video decoding standards, such as, for example, Versatile Video Coding (VVC), Joint Exploration Model (JEM), and / or other video decoding standards being developed or to be developed. Although the examples described herein provide examples using the JEM model, VVC, HEVC standards, and / or their extensions, the techniques and systems described herein may also be applicable to other decoding standards. Therefore, although the techniques and systems described herein may be described with reference to a specific video decoding standard, it will be understood by those of ordinary skill in the art that the description should not be interpreted as applying only to that specific standard.
[0044] Figure 11 is a block diagram showing an example of a system 100 including an encoding device 104 and a decoding device 112. The encoding device 104 may be part of a source device, and the decoding device 112 may be part of a receiving device. The source device and / or the receiving device may include an electronic device, such as a mobile or stationary telephone handset (e.g., a smart phone, a cellular phone, etc.), a desktop computer, a laptop or notebook computer, a tablet computer, a set-top box, a television, a camera, a display device, a digital media player, a video game console, a video streaming device, an Internet Protocol (IP) camera, or any other suitable electronic device. In some examples, the source device and the receiving device may include one or more wireless transceivers for wireless communication. The decoding techniques described herein are applicable to video decoding in various multimedia applications, including streaming video transmission (e.g., via the Internet), television broadcasting or transmission, encoding of digital video for storage on a data storage medium, decoding of digital video stored on a data storage medium, or other applications. In some examples, the system 100 may support one-way or two-way video transmission to support applications such as video conferencing, video streaming, video playback, video broadcasting, games, and / or video calls.
[0045] The encoding device 104 (or encoder) can be used to encode video data using a video coding standard or protocol to generate a coded video bitstream. Examples of video coding standards include ITU-T H.261, ISO / IEC MPEG-1 Vision, ITU-T H.262, or ISO / IEC MPEG-2 Vision, ITU-T H.263, ISO / IEC MPEG-4 Vision, ITU-TH.264 (also known as ISO / IEC MPEG-4 AVC) (including its scalable video coding (SVC) and multi-view video coding (MVC) extensions), and HEVC or ITU-T H.265. There are various extensions to HEVC for processing multi-layer video coding, including range and screen content coding extensions, 3D video coding (3D-HEVC) and multi-view extensions (MV-HEVC) and scalable extensions (SHVC). The new video coding standard being developed by the Joint Exploration Video Team (JVET) is called Universal Video Coding (VVC).
[0046] refer to Figure 1, video source 102 may provide video data to encoding device 104. Video source 102 may be part of a source device, or may be part of a device other than a source device. Video source 102 may include a video capture device (e.g., a video camera, a camera phone, a video phone, etc.), a video archive containing stored video, a video server or content provider for providing video data, a video feed interface for receiving video from a video server or content provider, a computer graphics system for generating computer graphics video data, a combination of such sources, or any other suitable video source.
[0047] The video data from video source 102 may include one or more input pictures. A picture may also be referred to as a "frame". A picture or frame is a still image, which in some cases is part of a video. In some examples, the data from video source 102 may be a still image that is not part of a video. In HEVC, VVC, and other video coding specifications, a video sequence may include a series of pictures. A picture may include three sample arrays, denoted S L , S Cb and S Cr . S L is a two-dimensional array of brightness samples, S Cb is a two-dimensional array of Cb chrominance samples, and S Cr is a two-dimensional array of Cr chrominance samples. Chrominance samples may also be referred to herein as "chroma" samples. A pixel may refer to a location within a picture that includes a luma component, a Cb component, and a Cr component. In other instances, a picture may be monochrome and may include only an array of luma samples.
[0048] The encoder engine 106 (or encoder) of the encoding device 104 encodes the video data to generate a coded video bitstream. In some examples, the coded video bitstream (or "video bitstream" or "bitstream") is a series of one or more decoded video sequences. The decoded video sequence (CVS) includes a series of access units (AUs), which start from an AU having a random access point picture and certain attributes in the base layer, until the next AU having a random access point picture and certain attributes in the base layer and does not include the next AU. The AU includes one or more decoded pictures and control information corresponding to the decoded pictures sharing the same output time. The decoded slices of the picture are encapsulated as data units at the bitstream level, which are called network abstraction layer (NAL) units. For example, a HEVC video bitstream may include one or more CVSs, which include NAL units. Each NAL unit in the NAL unit has a NAL unit header.
[0049] The encoder engine 106 generates a coded representation of a picture by partitioning each picture into multiple slices. Slices are independent of other slices so that the information in a slice is decoded without relying on the data of other slices in the same picture. The slice is then partitioned into coding tree blocks (CTBs) of luma samples and chroma samples. The CTB of luma samples and one or more CTBs of chroma samples together with the syntax for the samples are called a coding tree unit (CTU). CTU is the basic processing unit for HEVC encoding. CTU can be split into multiple coding units (CUs) of different sizes. CU contains arrays of luma and chroma samples called coding blocks (CBs).
[0050] Luma and chroma CBs can be further split into prediction blocks (PBs). A PB is a block of samples of a luminance component or a chroma component that uses the same motion parameters for inter-frame prediction or intra-block copy prediction (when available or enabled for use). A luminance PB and one or more chroma PBs together with associated syntax form a prediction unit (PU). For inter-frame prediction, a set of motion parameters (e.g., one or more motion vectors, reference indices, etc.) are signaled in the bitstream for each PU, as well as for inter-frame prediction of the luminance PB and one or more chroma PBs. Motion parameters may also be referred to as motion information. A CB may also be partitioned into one or more transform blocks (TBs). A TB represents a square block of samples of a color component to which the same two-dimensional transform is applied for decoding a prediction residual signal. A transform unit (TU) represents a TB of luminance and chroma samples and corresponding syntax elements.
[0051] The size of the CU corresponds to the size of the decoding mode, and can be square in shape. For example, the size of the CU can be 8x 8 samples, 16x 16 samples, 32x 32 samples, 64x 64 samples, or any other appropriate size up to the size of the corresponding CTU. The phrase "N x N" is used herein to refer to the pixel size of the video block in the vertical and horizontal dimensions (e.g., 8 pixels x 8 pixels). The pixels in the block can be arranged in rows and columns. In some examples, the block may not have the same number of pixels in the horizontal direction as in the vertical direction. Syntax data associated with the CU can describe, for example, the partitioning of the CU into one or more PUs. The partitioning mode can be different between whether the CU is encoded in an intra-frame prediction mode or an inter-frame prediction mode. The PU can be partitioned into a non-square shape.
[0052] According to the HEVC standard, the transform may be performed using TUs, as described above. TUs may be different for different CUs. TUs may be sized based on the size of PUs within a given CU. TUs may be the same size as PUs or smaller than PUs. In some examples, the residual samples corresponding to a CU may be subdivided into smaller units using a quadtree structure called a residual quadtree (RQT). The leaf nodes of the RQT may correspond to TUs. The pixel difference values associated with the TUs may be transformed to produce transform coefficients. The transform coefficients may then be quantized by the encoder engine 106.
[0053] Once the picture of the video data is partitioned into CUs, the encoder engine 106 predicts each PU using a prediction mode. The prediction unit or prediction block is then subtracted from the original video data to obtain a residual (described below). For each CU, the prediction mode may be signaled within the bitstream using syntax data. The prediction mode may include intra-frame prediction (or intra-picture prediction) or inter-frame prediction (or inter-picture prediction). The decision on whether to use inter-picture prediction or intra-picture prediction to decode a picture region may be made, for example, at the CU level.
[0054] Intra-picture prediction exploits the correlation between spatially adjacent samples within a picture. For example, using intra-frame prediction, each PU is predicted from adjacent image data in the same picture using, for example, DC prediction to find the average for the PU, planar prediction to fit a planar surface to the PU, directional prediction to infer from adjacent data, or any other suitable type of prediction.
[0055] Inter-picture prediction uses temporal correlation between pictures in order to derive motion compensated prediction for blocks of image samples. Using a translational motion model, the position of a block in a previously decoded picture (reference picture) is represented by a motion vector (Δx, Δy), where Δx specifies the horizontal displacement of the reference block relative to the position of the current block, and Δy specifies the vertical displacement of the reference block relative to the position of the current block. In some cases, the motion vector (Δx, Δy) may be integer sample precision (also referred to as integer precision), in which case the motion vector points to an integer pixel grid (or integer pixel sampling grid) of a reference frame. In some cases, the motion vector (Δx, Δy) may have fractional sample precision (also referred to as fractional pixel precision or non-integer precision) to more accurately capture the motion of the underlying object without being limited to the integer pixel grid of the reference frame. The precision of the motion vector may be expressed by the quantization level of the motion vector. For example, the quantization level may be integer precision (e.g., 1 pixel) or fractional pixel precision (e.g., 1 / 4 pixel, 1 / 2 pixel, or other sub-pixel value). When the corresponding motion vector has fractional sample precision, interpolation is applied on the reference picture to derive the prediction signal. For example, samples available at integer positions can be filtered (e.g., using one or more interpolation filters) to estimate values at fractional positions. A previously decoded reference picture is indicated by a reference index (refIdx) to a reference picture list. The motion vector and the reference index may be referred to as motion parameters. Two types of inter-picture prediction may be performed, including unidirectional prediction and bidirectional prediction.
[0056] In the case of inter-frame prediction using bidirectional prediction, two motion parameter sets (Δx0, Δy0, refIdx0 and Δx1, Δy1, refIdx1) are used to generate two motion compensated predictions (from the same reference picture or possibly from different reference pictures). For example, in the case of bidirectional prediction, each prediction block uses two motion compensated prediction signals, and B prediction units are generated. The two motion compensated predictions are then combined to obtain the final motion compensated prediction. For example, the two motion compensated predictions can be combined by averaging. In another example, weighted prediction can be used, in which case different weights can be applied to each motion compensated prediction. The reference pictures that can be used in bidirectional prediction are stored in two separate lists, denoted as List 0 and List 1. The motion parameters can be derived using a motion estimation process at the encoder.
[0057] In the case of inter prediction using unidirectional prediction, a motion parameter set (Δx0, Δy0, refIdx0) is used to generate motion compensated prediction from the reference picture. For example, in the case of unidirectional prediction, each prediction block uses at most one motion compensated prediction signal and generates P prediction units.
[0058] The PU may include data related to the prediction process (e.g., motion parameters or other appropriate data). For example, when the PU is encoded using intra prediction, the PU may include data describing the intra prediction mode for the PU. As another example, when the PU is encoded using inter prediction, the PU may include data defining a motion vector for the PU. The data defining the motion vector for the PU may describe, for example, a horizontal component of the motion vector (Δx), a vertical component of the motion vector (Δy), a resolution for the motion vector (e.g., integer precision, quarter-pixel precision, or eighth-pixel precision), a reference picture to which the motion vector points, a reference index, a reference picture list for the motion vector (e.g., list 0, list 1, or list C), or any combination thereof.
[0059] The encoding device 104 may then perform a transform and quantization. For example, after prediction, the encoder engine 106 may calculate a residual value corresponding to the PU. The residual value may include a pixel difference value between the current pixel block (PU) being decoded and a prediction block (e.g., a predicted version of the current block) used to predict the current block. For example, after generating a prediction block (e.g., using inter-frame prediction or intra-frame prediction), the encoder engine 106 may generate a residual block by subtracting the prediction block generated by the prediction unit from the current block. The residual block includes a set of pixel difference values that quantizes the difference between the pixel values of the current block and the pixel values of the prediction block. In some examples, the residual block may be represented in a two-dimensional block format (e.g., a two-dimensional matrix or array of pixel values). In such an example, the residual block is a two-dimensional representation of pixel values.
[0060] Any residual data that may remain after performing the prediction is transformed using a block transform, which may be based on a discrete cosine transform, a discrete sine transform, an integer transform, a wavelet transform, other appropriate transform functions, or any combination thereof. In some cases, one or more block transforms (e.g., sizes 32x 32, 16x 16, 8x 8, 4x 4, or other appropriate sizes) may be applied to the residual data in each CU. In some examples, a TU may be used for a transform and quantization process implemented by the encoder engine 106. A given CU with one or more PUs may also include one or more TUs. As described in further detail below, the residual values may be transformed into transform coefficients using a block transform, and then may be quantized and scanned using the TU to produce serialized transform coefficients for entropy coding.
[0061] In some examples, after intra prediction or inter prediction decoding using the PU of the CU, the encoder engine 106 can calculate residual data for the TU of the CU. The PU may include pixel data in the spatial domain (or pixel domain). After applying the block transform, the TU may include coefficients in the transform domain. As mentioned above, the residual data may correspond to the pixel difference between the pixel of the unencoded picture and the predicted value corresponding to the PU. The encoder engine 106 may form a TU including the residual data for the CU, and then the TU may be transformed to generate transform coefficients for the CU.
[0062] The encoder engine 106 may perform quantization of the transform coefficients. Quantization provides further compression by quantizing the transform coefficients to reduce the amount of data used to represent the coefficients. For example, quantization may reduce the bit depth associated with some or all of the coefficients. In one example, a coefficient having an n-bit value may be rounded down to an m-bit value during quantization, where n is greater than m.
[0063] Once quantization is performed, the decoded video bitstream includes quantized transform coefficients, prediction information (e.g., prediction mode, motion vector, block vector, etc.), segmentation information, and any other appropriate data, such as other syntax data. The different elements of the decoded video bitstream can then be entropy encoded by the encoder engine 106. In some examples, the encoder engine 106 can scan the quantized transform coefficients using a predefined scanning order to generate a serialized vector that can be entropy encoded. In some examples, the encoder engine 106 can perform adaptive scanning. After scanning the quantized transform coefficients to form a vector (e.g., a one-dimensional vector), the encoder engine 106 can entropy encode the vector. For example, the encoder engine 106 can use context adaptive variable length coding, context adaptive binary arithmetic coding, context adaptive binary arithmetic coding based on syntax, probability interval segmentation entropy coding, or another appropriate entropy coding technology.
[0064] The output 110 of the encoding device 104 may send the NAL units constituting the encoded video bitstream data to the decoding device 112 of the receiving device over the communication link 120. The input 114 of the decoding device 112 may receive the NAL units. The communication link 120 may include a channel provided by a wireless network, a wired network, or a combination of a wired network and a wireless network. The wireless network may include any wireless interface or combination of wireless interfaces, and may include any suitable wireless network (e.g., the Internet or other wide area network, a packet-based network, a WiFi network, or a wireless network). TM , Radio Frequency (RF), UWB, WiFi Direct, Cellular, Long Term Evolution (LTE), WiMax TMThe wired network may include any wired interface (e.g., optical fiber, Ethernet, power line Ethernet, Ethernet over coaxial cable, digital signal line (DSL), etc.). The wired and / or wireless network may be implemented using various devices (e.g., base stations, routers, access points, bridges, gateways, switches, etc.). The encoded video bitstream data may be modulated according to a communication standard such as a wireless communication protocol and sent to a receiving device.
[0065] In some examples, the encoding device 104 may store the encoded video bitstream data in the storage space 108. The output 110 may retrieve the encoded video bitstream data from the encoder engine 106 or from the storage space 108. The storage space 108 may include any of a variety of distributed or locally accessed data storage media. For example, the storage space 108 may include a hard drive, a storage disk, a flash memory, a volatile or non-volatile memory, or any other suitable digital storage medium for storing the encoded video data. The storage space 108 may also include a decoded picture buffer (DPB) for storing reference pictures used in inter-frame prediction.
[0066] The encoder engine 106 and the decoder engine 116 (described in more detail below) can be configured to operate according to VVC. According to VVC, a video decoder (such as the encoder engine 106 and / or the decoder engine 116) partitions a picture into multiple coding tree units (CTUs). The video decoder can partition the CTU according to a tree structure (such as a quadtree-binary tree (QTBT) structure or a multi-type tree (MTT) structure). The QTBT structure removes the concept of multiple partition types, such as the distinction between CU, PU, and TU in HEVC. The QTBT structure includes two levels, including a first level partitioned according to quadtree partitioning, and a second level partitioned according to binary tree partitioning. The root node of the QTBT structure corresponds to the CTU. The leaf nodes of the binary tree correspond to the coding units (CUs).
[0067] In the MTT segmentation structure, blocks can be segmented using quadtree segmentation, binary tree segmentation, and one or more types of ternary tree segmentation. Ternary tree segmentation is a segmentation in which a block is divided into three sub-blocks. In some examples, ternary tree segmentation divides a block into three sub-blocks without dividing the original block through the center. The segmentation types in MTT (e.g., quadtree, binary tree, and ternary tree) can be symmetric or asymmetric.
[0068] In some examples, the video coder may use a single QTBT or MTT structure to represent each of the luma and chroma components, while in other examples, the video coder may use two or more QTBT or MTT structures, such as one QTBT or MTT structure for the luma component and another QTBT or MTT structure for the two chroma components (or two QTBT and / or MTT structures for respective chroma components).
[0069] The video decoder may be configured to use quadtree segmentation, QTBT segmentation, MTT segmentation, or other segmentation structures according to HEVC. For illustrative purposes, the description herein may refer to QTBT segmentation. However, it should be understood that the technology of the present disclosure may also be applied to a video decoder configured to use quadtree segmentation or also use other types of segmentation.
[0070] In VVC, pictures can be divided into slices, tiles, and bricks. Typically, a brick can be a rectangular area of a CTU row within a specific tile in a picture. A tile can be a rectangular area of a CTU within a specific tile column and a specific tile row in a picture. A tile column is a rectangular area of a CTU having a height equal to the height of the picture and a width specified by a syntax element in a picture parameter set. A tile row is a rectangular area of a CTU having a height specified by a syntax element in a picture parameter set and a width equal to the width of the picture. In some cases, a tile can be divided into multiple bricks, each of which can include one or more CTU rows within a tile. Tiles that are not divided into multiple bricks are also called bricks. However, bricks that are true subsets of tiles are not called tiles. A slice can be an integer number of bricks of a picture that are uniquely contained in a single NAL unit. In some cases, a slice can include a number of complete tiles or only a continuous sequence of complete bricks of a tile.
[0071] The input 114 of the decoding device 112 receives the encoded video bitstream data, and the video bitstream data may be provided to the decoder engine 116, or to the storage space 118 for later use by the decoder engine 116. The input 114 of the decoding device 112 receives the encoded video bitstream data, and the video bitstream data may be provided to the decoder engine 116, or to the storage space 118 for later use by the decoder engine 116. For example, the storage space 118 may include a DPB for storing reference pictures used in inter-frame prediction. A receiving device including the decoding device 112 may receive the encoded video data to be decoded via the storage space 108. The encoded video data may be modulated according to a communication standard such as a wireless communication protocol, and sent to the receiving device. The decoder engine 116 may decode the encoded video bitstream data by entropy decoding (e.g., using an entropy decoder) and extracting elements of one or more decoded video sequences constituting the encoded video data. The decoder engine 116 may then rescale the encoded video bitstream data and perform an inverse transform on it. The residual data is then passed to the prediction stage of the decoder engine 116. The decoder engine 116 then predicts a block of pixels (eg, a PU). In some examples, the prediction is added to the output of the inverse transform (the residual data).
[0072] The decoding device 112 may output the decoded video to a video destination device 122, which may include a display or other output device for displaying the decoded video data to a consumer of the content. In some aspects, the video destination device 122 may be part of a receiving device that includes the decoding device 112. In some aspects, the video destination device 122 may be part of a separate device distinct from the receiving device.
[0073] Examples of specific details of the encoding device 104 are provided below with reference to Figure 6 Examples of specific details of the decoding device 112 are described below with reference to Figure 7 To describe.
[0074] As previously described, the HEVC bitstream includes a set of NAL units, including VCL NAL units and non-VCL NAL units. The VCL NAL unit includes decoded picture data that forms a decoded video bitstream. For example, a bit sequence that forms a decoded video bitstream is present in a VCL NAL unit. In addition to other information, a non-VCL NAL unit may contain a parameter set with high-level information related to the encoded video bitstream. For example, a parameter set may include a video parameter set (VPS), a sequence parameter set (SPS), and a picture parameter set (PPS). Examples of the goals of a parameter set include bit rate efficiency, error resilience, and providing a system layer interface. Each slice references a single valid PPS, SPS, and VPS to access information that a decoding device 112 can use to decode the slice. An identifier (ID) may be decoded for each parameter set (including a VPS ID, an SPS ID, and a PPS ID). The SPS includes an SPS ID and a VPS ID. The PPS includes a PPS ID and an SPS ID. Each slice header includes a PPS ID. Using the ID, an active parameter set may be identified for a given slice.
[0075] The PPS includes information applied to all slices in a given picture. Because of this, all slices in the picture refer to the same PPS. Slices in different pictures may also refer to the same PPS. The SPS includes information applied to all pictures in the same decoded video sequence (CVS) or bitstream. The information in the SPS may not change from picture to picture within the decoded video sequence. The pictures in the decoded video sequence may use the same SPS. The VPS includes information applied to all layers within the decoded video sequence or bitstream. The VPS includes a syntax structure having syntax elements applied to the entire decoded video sequence. In some examples, the VPS, SPS, or PPS may be sent in-band with the encoded bitstream. In some examples, the VPS, SPS, or PPS may be sent out-of-band in a transmission separate from the NAL unit containing the decoded video data.
[0076] The video bitstream may also include supplemental enhancement information (SEI) messages. For example, a SEI NAL unit may be part of the video bitstream. In some cases, the SEI message may contain information that is not used by the decoding process. For example, the information in the SEI message may not be necessary for the decoder to decode the video pictures of the bitstream, but the decoder may use the information to improve the display or processing of the pictures (e.g., decoded output).
[0077] As described herein, for each block, a set of motion information (also referred to herein as motion parameters) may be available. The motion information set contains motion information for a forward prediction direction and a backward prediction direction. The forward prediction direction and the backward prediction direction are two prediction directions of a bidirectional prediction mode, in which case the terms "forward" and "backward" do not necessarily have a geometric meaning. Instead, "forward" and "backward" correspond to reference picture list 0 (RefPicList0 or L0) and reference picture list 1 (RefPicList1 or L1) of the current picture. In some examples, when only one reference picture list is available for a picture or slice, only RefPicList0 is available, and the motion information for each block of the slice is always forward.
[0078] In some cases, the motion vector together with its reference index is used in the decoding process (e.g., motion compensation). Such a motion vector with an associated reference index is represented as a unidirectional prediction set of motion information. For each prediction direction, the motion information may contain a reference index and a motion vector. In some cases, for simplicity, the motion vector itself may be referenced in a manner assuming that it has an associated reference index. The reference index is used to identify the reference picture in the current reference picture list (RefPicList0 or RefPicList1). The motion vector has a horizontal component and a vertical component, which provides an offset from the coordinate position in the current picture to the coordinate in the reference picture identified by the reference index. For example, the reference index may indicate a specific reference picture that should be used for a block in the current picture, and the motion vector may indicate where the best matching block in the reference picture (the block that best matches the current block) is located in the reference picture.
[0079] The picture order count (POC) can be used in the video coding standard to identify the display order of pictures. Although there are cases where two pictures in a decoded video sequence may have the same POC value, this does not usually happen in a decoded video sequence. When there are multiple decoded video sequences in the bitstream, pictures with the same POC value may be closer to each other in decoding order. The POC value of a picture can be used for reference picture list construction, such as reference picture set derivation in HEVC, and motion vector scaling.
[0080] For motion prediction in HEVC, there are two inter prediction modes for prediction units (PUs), including merge mode and advanced motion vector prediction (AMVP) mode. Skipping is considered a special case of merging. In AMVP mode or merge mode, a motion vector (MV) candidate list is maintained for multiple motion vector predictors. The motion vector of the current PU and the reference index in merge mode are generated by getting a candidate from the reference list.
[0081] In an example where the MV candidate list is used for motion prediction of a block, the reference list may be constructed separately by an encoding device and a decoding device. For example, the candidate list may be generated by an encoding device when encoding a block (e.g., a CTU, a CU, or other block of a picture), and may be generated by a decoding device when decoding a block. Information related to the motion information candidates in the candidate list may be transmitted by signal between the encoding device and the decoding device. For example, in merge mode, the index value of the stored motion information candidate may be transmitted by signal from the encoding device to the decoding device (e.g., in a syntax structure such as a PPS, SPS, VPS, a slice header, a SEI message sent in a video bitstream or separately from a video bitstream, and / or other signaling). The decoding device may construct a candidate list, and use a reference or index transmitted by signal to obtain one or more motion information candidates from the constructed candidate list for motion compensation prediction. For example, the decoding device 112 may construct an MV candidate list, and use a motion vector from an index position to perform motion prediction on a block. In the case of AMVP mode, in addition to the reference or index, the difference or residual value may also be transmitted as an increment signal. For example, for AMVP mode, the decoding device may construct one or more MV candidate lists, and apply the incremental value to one or more motion information candidates obtained using the signaled index value when performing motion compensated prediction of the block.
[0082] In some examples, the MV candidate list contains up to five candidates for merge mode and two candidates for AMVP mode. In other examples, different numbers of candidates may be included in the candidate lists for merge mode and / or AMVP mode. The merge candidate may include a motion information set. For example, a motion information set may include motion vectors corresponding to two reference picture lists (list 0 and list 1) and a reference index. If the merge candidate is identified by a merge index, the reference picture is used to predict the current block and to determine the associated motion vector. However, under AMVP mode, for each potential prediction direction from list 0 or list 1, the reference index needs to be explicitly signaled along with the index to the candidate list, because the AMVP candidate only contains motion vectors. In AMVP mode, the predicted motion vector can be further refined.
[0083] As seen above, the merge candidate corresponds to a complete motion information set, while the AMVP candidate contains only one motion vector and reference index for a specific prediction direction. Candidates for both modes can be derived similarly from the same spatial and temporal neighboring blocks. In some examples, the merge mode allows an inter-predicted PU to inherit the same motion vector or multiple motion vectors, prediction direction, and reference picture index or multiple reference picture indexes from an inter-predicted PU that includes a motion data position selected from a set of spatially adjacent motion data positions and one of two temporally co-located motion data positions. For the AMVP mode, the motion vector or multiple motion vectors of the PU can be predictively decoded relative to one or more motion vector predictors (MVPs) from an AMVP candidate list constructed by the encoder and / or decoder. In some instances, for unidirectional inter prediction of a PU, the encoder and / or decoder can generate a single AMVP candidate list. In some instances, for bidirectional prediction of a PU, the encoder and / or decoder may generate two AMVP candidate lists, one using motion data of spatially and temporally neighboring PUs from a forward prediction direction, and one using motion data of spatially and temporally neighboring PUs from a backward prediction direction.
[0084] In VVC, there is a reference picture resampling (RPR) tool under consideration, which is described in the following document: S. Wenger, BD. Choi, SCHan, X. Li, S. Liu, "AHG8: Spatial Scalability using Reference Picture Resampling", JVET-00045, the entire contents of which are hereby incorporated by reference and for all purposes. The RPR tool allows the use of reference pictures with a picture size different from the current picture size. In this case, a picture resampling process is called to provide an upsampled or downsampled version of a picture (e.g., a reference picture) that matches the current picture size. The tool is similar to the spatial scalability present in, for example, the scalable extension (SHVC) of the H.265 / HEVC standard.
[0085] This document describes systems, methods, apparatus, and computer-readable media that provide several aspects to increase support for spatial scalability using RPR in VVC. In some cases, pictures with the same content but in different representations (e.g., with different resolutions) may be assumed to have the same picture order count (POC) value (similar to the approach in, for example, HEVC and SHVC).
[0086] Using the techniques described herein, no access unit (AU) definition needs to be changed in VVC, because the AU starts with a VCL NAL unit (or a non-VCL NAL unit associated with a VCL NAL unit with a new POC value), which occurs when the next picture VCL NAL unit is encountered, such as Figure 3 As shown, across two different layers labeled as layer 0 and layer 1 (e.g., layer 0 has lower resolution pictures compared to layer 1), the first AU has a POC value of n-1, the second AU has a POC value of n, and the third AU has a POC value of n+1.
[0087] However, other layer pictures may need to be included in the reference picture structure (RPS) and reference picture list (RPL) to indicate which pictures can be used for inter-layer prediction and from which layer. In current VVC, referencing pictures from other layers is not allowed.
[0088] The following provides various examples of reference draft VVC documents, such as "Versatile Video Coding (Draft 5)", Version 7, JVET-N1001-v7, the entire contents of which are hereby incorporated by reference and for all purposes. As mentioned above, the techniques described herein can be applied to any existing video codec (such as High Efficiency Video Coding (HEVC) and / or Advanced Video Coding (AVC)), or can be proposed as a promising coding tool for currently developing standards (such as Versatile Video Coding (VVC)) and for other future video coding standards.
[0089] Figure 2 is a schematic diagram showing an example implementation of the filter unit 91, which may be as described below with respect to Figure 6 and Figure 7 2, including placing the reference picture in a reference picture storage space 208 (e.g., picture memory 92) that can be identified by a reference picture table 209. Additional aspects of the reference picture table 209 are discussed below. Filter unit 63 may be implemented in the same manner. For scalability support as described herein, filter units 63 and 91 may use POC numbers and reference picture offset values to implement the implementations described herein and other implementations, possibly in conjunction with other components of the video encoding device 104 or the video decoding device 112.
[0090] The filter unit 91 as shown includes a deblocking filter 202, a sample adaptive offset (SAO) filter 204 and a conventional filter 206, but is different from Figure 2The filters shown in FIG. 1 may include fewer filters and / or may include additional filters. Figure 2 The specific filters shown in the figure may be implemented in a different order. Other loop filters (either in the decoding loop or after the decoding loop) may also be used to smooth pixel transitions or otherwise improve video quality. When in the decoding loop, the decoded video blocks in a given frame or picture are then stored in a decoded picture buffer (DPB), which stores reference pictures as part of the reference picture storage space 208. The DPB (e.g., the reference picture storage space 208) may be a storage space for storing decoded video for later display on a display device (e.g., a video buffer). Figure 1 The memory may be part of or separate from the additional memory presented on a display of the video destination device 122.
[0091] Weighted prediction (WP) is a form of prediction using a reference picture, in which case a scaling factor (denoted by a), a shift number (denoted by s), and an offset (denoted by b) are used in motion compensation. For example, a linear model may be used in WP, as shown in the following equation (1):
[0092] p(i,j)=a*r(i+dv x ,j+dv y )+b, where (i,j)∈PU c Equation (1)
[0093] In equation (1), PU c is the current PU, (i,j) is the PU c The coordinates of the samples (or pixels in some cases) in x ,dv y ) is PU c The difference vector of , p(i,j) is PU c , r is the reference picture of the PU, and a and b are model parameters (where a is the scaling factor and b is the offset, as mentioned above).
[0094] When WP is enabled, for each reference picture of the current slice, a flag is signaled to indicate whether WP is applied to the reference picture. If WP is applied to a reference picture, the WP parameter set (i.e., a, s, and b) is sent to the decoder and used for motion compensation from the reference picture. In some examples, in order to provide flexibility when turning WP on or off for luma and chroma samples, the WP flag and WP parameters can be signaled separately for the luma component and chroma component of the pixel. In some cases, in WP, the same WP parameter set is used for all samples in a reference picture. In some examples, variables specify the width and height of the current decoding block, and an array with the same width and height is used as the prediction sample. The prediction sample can be derived using a weighted sample prediction process and is used when reconstructing the picture.
[0095] The picture may include one or more offsets (e.g., offset b), one or more weights or scaling factors (e.g., scaling factor a), a shift number, and / or other appropriate weighted prediction parameters. For bidirectional inter prediction, the one or more weights may include a first weight for a first reference picture and a second weight for a second reference picture. As described herein, in some examples, the reference block may be the same block as the current block, but with different associated weight parameters (e.g., for weighted prediction). In other examples, the reference block may be a block of the same picture but of a different size (e.g., a different resolution), such that the reference block is part of an inter-layer prediction process, such as with respect to Figure 3 Described in more detail. In some examples, both weighted prediction and inter-layer prediction may be performed on a single current block with reference blocks having different weights and being associated with pictures of different sizes.
[0096] Figure 3 is a schematic diagram showing a series of access units 310, 320, and 330. Each access unit has an associated POC value, shown as POC values 316, 326, and 336 corresponding to access units 310, 320, and 330. The access units each have two layers for picture frames having different sizes (e.g., resolutions). Figure 3Examples include layer 0 unit 312 and layer 1 unit 314 for access unit 310, layer 0 unit 322 and layer 1 unit 324 for access unit 320, and layer 0 unit 332 and layer 1 unit 334 for access unit 330. Layer 0 unit 312 (e.g., a VCL NAL unit) will have a different size than layer 1 unit 314 (e.g., also a VCL NAL unit). As mentioned above, using the techniques described herein, the VVC AU definition does not need to change because the AU starts with a VCL NAL unit with a new POC value, or with a non-VCL NAL unit associated with a VCL NAL unit with a new POC. For example, when a next picture VCL NAL unit is encountered, such as Figure 3 As shown (wherein across two different layers labeled as layer 0 and layer 1 (e.g., layer 0 has lower resolution pictures compared to layer 1), the first AU has a POC value of n-1, the second AU has a POC value of n, and the third AU has a POC value of n+1), new POC values may be used. Figure 3 The example includes two layers, but in various examples, other numbers of layers may be used.
[0097] The examples described herein utilize syntax and operational structures for efficiently allowing the use of data from other layers to be used as reference pictures by pictures in different layers, improving the operation of VVC devices and networks. As described herein, such improvements allow efficient signaling for identifying reference pictures in another layer of a shared access unit when weighted prediction is enabled, and efficient signaling for allowing reference pictures with zero-valued offsets to be used in other frames when weighted prediction is not enabled.
[0098] VVC operates with both short-term and long-term reference pictures. Some examples described herein operate with signaling of reference pictures for short-term reference pictures. In various systems, short-term reference pictures are signaled relative to previous short-term reference pictures as a POC delta to the previous reference picture POC. The initial POC value can be initialized to be equal to the current picture POC. The incremental POC is allowed to be equal to 0 (e.g., there may be multiple reference pictures with the same POC value, such as where there are different resolutions of the same picture), but the POC value cannot be equal to the current picture POC (e.g., the current picture cannot be inserted as a reference picture in a reference picture structure or list). In various examples, the long-term reference picture POC is signaled as the most significant bit (MSB) and the least significant bit (LSB), and the LSB portion can be equal to 0.
[0099] In some examples, such as Figure 2 The table 209 of reference pictures shown is stored in the DPB (e.g., using a memory or such as Figure 2209 is a table of picture sequence count offset values from a current picture or a previous picture that identifies a picture from the reference picture storage space 208 (e.g., as part of a picture memory such as picture memory 92). In some examples, such a table 209 may be generated in a loop, where an initial table value is equal to an initial picture POC value. A base POC value (e.g., a picture sequence count offset value) is assigned to identify the POC value of the reference picture for the initial picture. The POC value of the first reference picture for the initial picture associated with the initial POC value is assigned as the base POC value for the table. All subsequent reference pictures are then identified as delta values from the base POC value (e.g., the POC value of the previous reference picture), with a delta POC and a new base POC value for each entry in the table. The process loops until the table 209 is complete, with all reference pictures for each picture associated with the table. As described above, the initial incremental POC is added to the first picture POC, and the subsequent incremental POC is added to the previous reference picture POC. The bitstream signal is then constructed using the incremental POC and the base POC to limit the data usage in the bitstream while allowing the decoder to reconstruct the table 209. The repeated reference picture from the table 209 described above can then be used for prediction. In some examples using weighted prediction, when the same picture is being used for prediction but has different weights, the incremental POC for the reference picture and the picture being processed can be zero. In some examples, certain types of predictions can use modified tables or modified incremental POC values. For example, two pictures with the same POC can have different weights. In an illustrative example, one of the pictures can have a weight of 1.5 and a reference picture with the same weight of 2.5. Therefore, the same picture with different weighting parameters can be used as a reference with a zero-value incremental POC (also called a "zero-value picture sequence count offset" or a "zero-value POC offset"), which identifies that the reference picture is being used for weighted prediction. In another example, the modified delta POC value is used when weighted prediction is disabled. For example, in some cases, weighted prediction may be turned off by a flag such as weighted_pred_flag, which may be signaled in a picture parameter set (PPS). weighted_pred_flag may be referred to as a PPS weighted prediction flag. In some examples, when weighted prediction is disabled, the same value POC signaling for more than one reference picture is not allowed.To implement such restrictions or constraints, weighted prediction flags (e.g., sps_weighted_pred_flag and sps_weighted_bipred_flag) may be signaled in a parameter set (e.g., in an SPS, as RPS information may be signaled in an SPS or other parameter set). sps_weighted_pred_flag and / or sps_weighted_bipred_flag may be referred to as sequence parameter set (SPS) weighted prediction flags.
[0100] Allowing zero incremental POC values when weighted prediction is disabled results in inefficiencies in incremental POC signaling because there is no reason to insert the same reference picture into the reference picture list multiple times. In such a case, the incremental POC value will be at least equal to 1, and the incremental POC value of 0 is not used. However, in some examples, an incremental POC value of zero may be associated with a reserved codeword. The use of PPS and SPS flags allows different treatments of incremental POC signaling in the weighted prediction and non-weighted prediction cases described herein. Thus, a flag may be used to allow a zero-value picture order count offset (e.g., a zero incremental POC value) based on a determination that weighted prediction is enabled for a picture slice (or other portion of a picture, such as a CTU, CU, or other block), and to enable the interpretation of a signaled zero incremental POC as a different value (e.g., a signaled value plus 1) when weighted prediction is disabled. Thus, for weighted prediction, the same picture with different weights is used for weighted prediction. For inter-layer prediction, different layer pictures from the same access unit may be used as reference pictures. Conditional use of zero delta POC signaling values provides efficiency in reference picture signaling and improved system and device operation, while limiting additional overhead for signaling and improving system and device performance when weighted prediction is disabled.
[0101] For a zero-valued picture order count offset (where the reference picture is in the same access unit as the current picture being processed), a reference picture sampling tool may be used to generate the necessary reference data for processing the current picture. In some examples, the reference picture sampling tool may be part of a filter unit (e.g., filter unit 91), or in other examples may be part of any aspect of an apparatus for encoding or decoding as described herein.
[0102] An illustrative example of a modification to a VVC draft (e.g., Universal Video Coding (Draft 5), Version 7) for such marking is shown below. Additional examples of syntax tables, syntax terms, syntax logic and values, and other aspects of example implementations are shown below. Where an addition is shown, in the “ <insert>"and" <insertend>" Use underscores and text between symbols to describe them (for example, " <insert>Added text <insertend>”). An example of a weighted prediction tag is as follows:
[0103] <insert>sps_weighted_pred_flag equal to 0 specifies that weighted prediction is not applied to P slices. weighted_pred_flag equal to 1 specifies that weighted prediction can be applied to P slices.
[0104] sps_weighted_bipred_flag equal to 0 specifies that default weighted prediction is applied to B slices. weighted_bipred_flag equal to 1 specifies that weighted prediction can be applied to B slices. <insertend>
[0105] In such an example, the flag signaled in the PPS may be constrained by the SPS weighted prediction flag (e.g., the PPS flag may be equal to 1 only when the SPS flag is equal to 1). In another illustrative example:
[0106] weighted_pred_flag equal to 0 specifies that weighted prediction is not applied to P slices. weighted_pred_flag equal to 1 specifies that weighted prediction can be applied to P slices; <insert>weighted_pred_flag can be equal to 1 only when the corresponding sps_weighted_pred_flag is equal to 1. <insertend>.
[0107] weighted_bipred_flag equal to 0 specifies that the default weighted prediction is applied to B slices. weighted_bipred_flag equal to 1 specifies that weighted prediction can be applied to B slices; <insert>weighted_bipred_flag can be equal to 1 only when the corresponding sps_weighted_bipred_flag is equal to 1. <insertend>.
[0108] In some examples, delta POC equal to zero and POC LSB equal to zero may be allowed only when weighted prediction is enabled (e.g., as indicated by a flag in the SPS). In another example, zero delta POC and zero POC LSB are required when inter-layer prediction is enabled. The enabling of the inter-layer prediction flag may be signaled in a parameter set such as an SPS. Putting the above two cases together, delta POC equal to zero and POC LSB equal to zero may be allowed only when weighted prediction or inter-layer prediction is enabled.
[0109] In VVC, the POC value of the reference picture is signaled in the ref_pic_list_struct (RPS) as follows:
[0110]
[0111] Table 1
[0112] Examples of syntax include:
[0113] abs_delta_poc_st[listIdx][rplsIdx][i], when the i-th entry is the first STRP entry in the ref_pic_list_struct(listIdx,rplsIdx) syntax structure, specifies the absolute difference between the picture order count value of the current picture and the picture referenced by the i-th entry, or, when the i-th entry is a STRP entry but not the first STRP entry in the ref_pic_list_struct(listIdx,rplsIdx) syntax structure, specifies the absolute difference between the picture order count value of the picture referenced by the i-th entry and the picture order count value referenced by the previous STRP entry in the ref_pic_list_struct(listIdx,rplsIdx) syntax structure.
[0114] The value of abs_delta_poc_st[listIdx][rplsIdx][i] shall be between 0 and 2. 15 The range is -1 (inclusive).
[0115] rpls_poc_lsb_lt[listIdx][rplsIdx][i] specifies the value of the picture order count modulo MaxPicOrderCntLsb of the picture referenced by the i-th entry in the ref_pic_list_struct(listIdx, rplsIdx) syntax structure. The length of the rpls_poc_lsb_lt[listIdx][rplsIdx][i] syntax element is log2_max_pic_order_cnt_lsb_minus4+4 bits.
[0116] In one illustrative implementation example, when both weighted prediction and inter-layer prediction are disabled, syntax elements that may indicate zero delta POC (e.g., using a zero-valued picture order count offset to identify a reference picture) or the same value POC for a reference picture (such as abs_delta_poc_st and rpls_poc_lsb_lt) may not be equal to 0 or may not indicate the same POC value as already used for other reference pictures. In some examples, zero delta POC is not allowed when weighted prediction is disabled. In some examples, zero delta POC is not allowed when inter-layer prediction is disabled.
[0117] In another illustrative implementation example, when both weighted prediction and inter-layer prediction are disabled, the values of these syntax elements (e.g., the values used to identify reference pictures signaled in the bitstream, such as picture order count offsets) can be treated as syntax elements minus 1. In some such examples, the value of the syntax element is reconstructed as the signaled value plus 1. In such examples, it is not possible for a value of 0 to be signaled. An implementation example is shown below:
[0118] abs_delta_poc_st[listIdx][rplsIdx][i], when the i-th entry is the first STRP entry in the ref_pic_list_struct(listIdx,rplsIdx) syntax structure, specifies the absolute difference between the picture order count value of the current picture and the picture referenced by the i-th entry, or, when the i-th entry is a STRP entry but not the first STRP entry in the ref_pic_list_struct(listIdx,rplsIdx) syntax structure, specifies the absolute difference between the picture order count value of the picture referenced by the i-th entry and the picture order count value of the picture referenced by the previous STRP entry in the ref_pic_list_struct(listIdx,rplsIdx) syntax structure <insert>Or when weighted prediction and inter-layer prediction are disabled, the absolute difference is increased by 1 <insertend>.
[0119] The value of abs_delta_poc_st[listIdx][rplsIdx][i] shall be between 0 and 2. 15 The range is -1 (inclusive).
[0120] rpls_poc_lsb_lt[listIdx][rplsIdx][i] specifies the picture order count modulo the value of MaxPicOrderCntLsb for the picture referenced by the i-th entry in the ref_pic_list_struct(listIdx, rplsIdx) syntax structure, <insert>Or when weighted prediction and inter-layer prediction are disabled, the picture order count modulo the value of MaxPicOrderCntLsb is increased by 1. <insertend>The length of the rpls_poc_lsb_lt[listIdx][rplsIdx][i] syntax element is log2_max_pic_order_cnt_lsb_minus4+4 bits.
[0121] In some examples, it is proposed to add other layer pictures to the RPS, but since the RPS only has POC values to identify pictures, the layer ID is additionally signaled for each reference picture.
[0122] In another example, to avoid signaling a layer ID for each reference picture, the layer ID is signaled only when the POC value of the reference picture is equal to the POC value of the current picture. If the delta POC relative to the current POC value is equal to 0 (e.g., the reference picture POC is equal to the current picture POC), the layer ID is signaled to identify the layer. In this example, a picture with a layer ID different from the layer ID of the current picture can be used for prediction only when the picture has the same POC as the current picture (or belongs to the same access unit as the current picture). In addition, inter-layer reference pictures can be considered as long-term reference pictures (LTRPs) only, and can be marked as LTRPs to avoid motion vector scaling. In this case, the layer ID can be signaled only for LTRP types. In another example, the layer ID is signaled only for LTRPs and when the delta POC is equal to 0 or the reference picture has the same POC value as the current picture. In another example, before the POC value of the reference picture is signaled, the layer ID can be signaled for the reference picture or only the LTRP. In this case, the POC value may be conditionally signaled based on the signaled layer ID and the current layer ID. For example, the POC is not signaled when the layer ID is not equal to the current layer ID (e.g., an inter-layer reference picture), and then the POC value is inferred to be equal to the current picture POC. In some examples, the layer ID may be an indicator for identifying the layer used for prediction, and may be a layer number, a layer index, or another indicator within a list of layers.
[0123]
[0124]
[0125] Table 2
[0126] According to the above syntax table for ref_pic_list_struct shown by Table 2, one example may include:
[0127] num_ref_entries may include the total number of both conventional and inter-layer reference pictures. In an alternative approach, the number of inter-layer reference pictures, num_inter_layer_ref_entries[listIdx][rplsIdx], is signaled separately. For example, an advanced flag signaled in a parameter set (PS), such as sps_interlayer_ref_pics_present_flag signaled in an SPS, may indicate whether inter-layer reference pictures are present.
[0128] <insert>sps_inter_layer_ref_pics_present_flag equal to 1 specifies that inter-layer prediction may be used in decoding of pictures that reference the SPS. sps_inter_layer_ref_pics_present_flag equal to 0 specifies that inter-layer prediction is not used in decoding of pictures that reference the SPS. <insertend>
[0129] In some such examples, instead of directly signaling nuh_layer_id, a layer index may be signaled. For example, layer 10 may use layers 0 and 9 as dependent layers. Instead of directly signaling the values 0 and 9, layer indices 0 (corresponding to layer ID 0) and 1 (corresponding to layer ID 9) may be signaled instead. The mapping between layer IDs and layer indices may be inferred or derived at the decoder.
[0130] In some examples, a separate flag, il_ref_pic_flag, may be signaled as shown below in Table 3 to indicate that a picture is used for inter-layer prediction. This flag may be signaled for each picture or only for LTRPs (e.g., when st_ref_pic_flag is equal to 0). In some such examples, a layer ID or layer index may be signaled only when il_ref_pic_flag is equal to 1. When a layer index is signaled, such a flag may be needed because the current layer ID is not included as a dependent layer and may not have an associated layer index.
[0131]
[0132] Table 3
[0133] Then, Table 4 below shows an example of a syntax table for ref_pic_list_struct, with the following details:
[0134] <insert>layer_dependency_idc[listIdx][rplsIdx][i] specifies the dependency layer index of the i-th entry in the ref_pic_list_struct(listIdx, rplsIdx) syntax structure. The value of layer_dependency_idc[listIdx][rplsIdx][i] shall be in the range of 0 to the current layer idc LayerIdc[nuh_layer_id] minus 1 minus 1, inclusive.
[0135] il_ref_pic_flag[listIdx][rplsIdx][i] equal to 1 specifies that the i-th entry in the ref_pic_list_struct(listIdx, rplsIdx) syntax structure is an inter-layer reference picture. il_ref_pic_flag[listIdx][rplsIdx][i] equal to 0 specifies that the i-th entry in the ref_pic_list_struct(listIdx, rplsIdx) syntax structure is not an inter-layer reference picture entry. When not present, the value of il_ref_pic_flag[listIdx][rplsIdx][i] is inferred to be equal to 0. num_inter_layer_ref_entries[listIdx][rplsIdx] specifies the number of direct inter-layer entries in the ref_pic_list_struct(listIdx, rplsIdx) syntax structure. The value of num_inter_layer_ref_entries[listIdx][rplsIdx] shall be in the range of 0 to sps_max_dec_pic_buffering_minus1+14–num_ref_entries[listIdx][rplsIdx], inclusive. <insertend>
[0136] In some such examples, the number of inter-layer reference pictures, num_inter_layer_ref_entries, may be constrained to not exceed the maximum number of reference pictures minus the number of conventional inter-reference pictures, e.g., the value of num_inter_layer_ref_entries[listIdx][rplsIdx] shall be in the range of 0 to sps_max_dec_pic_buffering_minus 1+14–num_ref_entries[listIdx][rplsIdx], inclusive, added to the semantics of num_inter_layer_ref_entries. The num_inter_layer_ref_entries syntax element may be referred to as an inter-layer reference picture entry syntax element. The ref_pic_list_struct(listIdx, rplsIdx) syntax structure may be referred to as a reference picture list syntax structure. In some examples, when inter-layer reference pictures are allowed, il_ref_pic_flag[][] is signaled for all long-term reference pictures.
[0137]
[0138]
[0139] Table 4
[0140] Then, Tables 5 and 6 show example syntax tables for ref_pic_list_struct and associated slice headers. In such additional examples, layer information may be signaled in the slice header. This signaling may be in addition to the signaling allowed in the ref_pic_list_struct() syntax structure. Such a definition facilitates the case where layer information is signaled only when the LSB of the reference picture is equal to the LSB of the current picture.
[0141]
[0142]
[0143] Table 5
[0144]
[0145]
[0146] Table 6
[0147] In other embodiments described by Table 7, il_ref_pic_flag[][] in the slice header is signaled only when the POC LSB of the long-term reference picture is equal to the POC LSB of the current picture. In some examples, rpls_poc_lsb_lt may also be used.
[0148]
[0149] Table 7
[0150] In some examples, Table 8 below may operate as an alternative to the examples of Tables 6 and 7 above:
[0151]
[0152] Table 8
[0153] In some examples, the layer_depdendency_idc syntax element may not be sent in cases where an indication of layer_depdendency_idc is not necessary to identify an inter-layer reference picture. In some examples, if there is only one inter-layer reference picture that can be used for reference and il_ref_pic_flag[][][] indicates that the reference picture can be used for reference that satisfies condition A (e.g., as defined below), then layer_dependency_idc is not signaled and is inferred to be the inter-layer reference picture that can be used for reference. In some cases, when there are more than one inter-layer reference pictures that can be used for reference that satisfies condition A (e.g., n in number), it may be sufficient to signal no more than n-1 layer_dependency_idc values to indicate a specific inter-layer reference picture for position i in the list. In such examples, condition A may be any condition that allows inter-layer prediction for the current picture. For example, in some cases, condition A checks whether the inter-layer reference picture shares the same POC LSB as the current picture. In some cases, condition A checks whether the inter-layer reference picture shares the same POC as the current picture. In some cases, condition A checks that the difference between the POC of the inter-layer reference picture and the POC of the current picture does not exceed a threshold. In some cases, condition A may be restricted to be applied within an access unit. In some examples, when there is no picture satisfying condition A for the current picture, an indication of the inter-layer reference picture (e.g., a related syntax element) may not be signaled.
[0154] In some examples, when reference pictures are indicated as inter-layer reference pictures (e.g., using il_ref_pic_flag[] or similar syntax), the picture flag may be modified to indicate those pictures as "used for inter-layer reference pictures". In some cases, this flag may be combined with other flags (e.g., used for short-term reference or used for long-term reference), or may not include other flags (e.g., a picture may be marked as only two of the following: used for short-term reference or used for long-term reference or used for inter-layer reference pictures). In some cases, the flag of a picture may be modified as follows:
[0155] - After a picture is decoded, it is marked as "used for inter-layer reference".
[0156] - After an access unit is completed, or when starting the next access unit, all pictures in the current access unit are marked for short-term reference.
[0157] In some cases, a discardable_flag is signaled to indicate that a picture may not be used for inter-layer reference. In this case, some modifications may be applied to some of the methods disclosed above. In one example, a picture with a discardable_flag equal to 1 may be marked as "not used for reference" after decoding of the picture. In another example, a picture with a discardable_flag equal to 1 may never be marked as anything other than "not used for reference". In another example, a picture with a discardable_flag equal to 1 may not be considered for any prediction; therefore, it may not be considered when determining certain conditions that may otherwise apply to an inter-layer reference picture or reference picture. For example, if there are two pictures A and B in a layer different from the current picture in the current access unit, and only one of them (e.g., picture A) has a discardable_flag equal to 1, when an inter-layer reference picture is indicated, the system may infer that the inter-layer reference picture is picture B, in which case picture A is not added to any reference picture list in the reference picture list.
[0158] The mapping table mentioned above (eg, vps_direct_dependency_flag[i][j]) may be signaled, for example, in a parameter set (such as VPS as shown in example Table 10).
[0159]
[0160] Table 10
[0161] Example Table 10 may operate with the following syntax elements:
[0162] <insert>The variable LayerIdc[i][j]is derived as follows:
[0163] for(i=0,j=0;i<=vps_max_layers_minus1;i++,j++){
[0164] LayerIdc[vps_included_layer_id[i]][j]=j
[0165] }
[0166] ( <insert>The variable LayerIdc[i][j] is derived as follows:
[0167] for(i=0,j=0;i<=vps_max_layers_minus1;i++,j++){
[0168] LayerIdc[vps_included_layer_id[i]][j]=j
[0169] })
[0170] vps_independent_layer_flag[i] equal to 1 specifies that the layer with index i does not use inter-layer prediction. vps_independent_layer_flag[i] equal to 1 specifies that the layer with index i may use inter-layer prediction and vps_layer_dependency_flag is present in the VPS.
[0171] vps_direct_dependency_flag[i][j] equal to 0 specifies that the layer with index j is not a direct reference layer for the layer with index i. vps_direct_dependency_flag[i][j] equal to 1 specifies that the layer with index j is a direct reference layer for the layer with index i. When vps_direct_dependency_flag[i][j] is not present for i and j in the range of 0 to MaxLayersMinus1, it is inferred to be equal to 0.
[0172] The variable DependencyIdc[i][j] is derived as follows:
[0173]
[0174] In a reference picture list according to some examples, when the POC value is not signaled (e.g., when the layer ID is not equal to the current layer ID, such as for an inter-layer reference picture), the POC value is inferred to be equal to the current picture POC. The layer ID is an indicator used to identify the layer, and can be the layer index used for prediction within the list of layers. In other examples, other indicators can be used. Num_ref_entries may include the total number of both conventional and inter-layer reference pictures. In some examples, a separate reference flag may be signaled to indicate that the picture is used for inter-layer prediction in various examples, and the flag may be signaled for each reference picture.
[0175] In some examples, the inter-layer reference picture is not indicated by a POC value. In such examples, the inter-layer reference picture may be indicated by a flag (e.g., inter_layer_ref_pic_flag) and a layer index. The POC value of the inter-layer reference picture may be derived to be equal to the current picture POC according to the following condition: there is a reference picture in the DPB with a nuh_layer_id equal to the reference picture layer ID and the same picture sequence count value as the current picture. Therefore, as described above, in some examples, the reference picture list structure may be constructed using the layer ID and an inferred POC value of 0 for signaling inter-layer prediction. In other examples, other reference picture list structures may be used.
[0176] In some such examples, reference picture lists RefPicList[0] and RefPicList[1] are constructed as follows:
[0177]
[0178]
[0179] In some cases, if nuh_layer_id[i][RplsIdx[i]][j] is not signaled directly, it can be derived from LayerIdc and DependencyIdc variables using layer_dependency_idc. For example, nuh_layer_id[listIdx][rplsIdx][j] = DependencyIdc[LayerIdc[nuh_layer_id]][layer_dependency_idc[listIdx][rplsIdx][j]], where nuh_layer_id is the layer ID of the current picture. This value can be used as the layer ID in reference picture list derivation.
[0180] In some cases, the following constraint can be added, that is, when an inter-layer reference picture is signaled in the RPS, the POC value of the inter-layer reference picture derived from the RPS should be equal to the POC value of the current picture. When inter-layer prediction is used, a zero motion vector can be assumed, in which case there is no displacement of content in pictures of different representations (resolutions). Motion information can be signaled for inter-frame prediction mode. To save overhead, motion information can be signaled when (in some cases, only when) it is not used for inter-layer prediction. Otherwise (if inter-layer prediction) the motion vector is inferred to be equal to 0 in the absence of signaling.
[0181] Table 11 is an implementation example based on a VVC draft (e.g., Universal Video Coding (Draft 5), Version 7), which has changes in coding_unit (x0, y0, cbWidth, cbHeight, treeType) as follows:
[0182]
[0183]
[0184] Table 11
[0185] As shown in example Table 12, in the slice header, information about the number of active inter-layer pictures may be signaled using the syntax elements described below:
[0186]
[0187]
[0188] Table 12
[0189] <insert>num_inter_layer_active_override_flag equal to 1 specifies that the syntax element num_inter_layer_ref_pics_minus1[0] is present for P and B slices and the syntax element num_inter_layer_ref_pics_minus1[1] is present for B slices. num_inter_layer_active_override_flag equal to 0 specifies that the syntax elements num_inter_layer_ref_pics_minus1[0] and num_inter_layer_ref_pics_minus1[1] are not present. When not present, the value of num_inter_layer_active_override_flag is inferred to be equal to 0.
[0190] num_inter_layer_ref_pics_minus1[i] is used in the derivation of the variable NumInterLayerActive[i]. The value of num_ref_idx_active_minus1[i] should be in the range of 0 to num_inter_layer_ref_entries[i][RplsIdx[i]].
[0191] The variable NumInterLayerActive[i] is derived as follows:
[0192]
[0193] In some examples, a spatial ID similar to the temporal ID may be defined that indicates a spatial representation of a picture that may not be captured by the nuh_layer_id. One of the purposes of signaling the nuh_layer_id in the NAL unit header is to easily identify the NAL units belonging to a specific layer, which may be removed for sub-bitstream extraction purposes. Easy identification of NAL units belonging to a specific layer provides benefits to the operation of the system, and therefore the spatial_id is also considered to be signaled in the NAL unit header. To this end, the nuh_reserved_zero_bit may be renamed to layer_id_interpret_flag, and the semantics may be modified as follows, or in some cases, a new syntax element spatial_id may also be sent.
[0194] - When layer_id_interpret_flag is equal to 0, the syntax element nuh_layer_id_plus1 is currently interpreted as specifying the layer ID of the current picture.
[0195] - When layer_id_interpret_flag is equal to 1, the syntax element nuh_layer_id_plus1 is interpreted as containing some bits for indicating the layer ID and some bits for specifying the spatial ID. For example, the first four bits may indicate the layer ID and the last three bits may indicate the spatial ID. Typically, in this case, the actual layer ID and spatial ID may be derived from the syntax element nuh_layer_id_plus1 using a predetermined method.
[0196] The distinction between spatial_id and layer ID is that the consistency of a "single-layer" decoder will also be concerned when decoding a bitstream containing multiple spatial_ids. The value of spatial_id can be used to identify reference pictures in a "single-layer" decoder without reusing the layer concept, and in this way the concept of layer will be clearly distinguished from the temporal and spatial layers. For simplicity, the consistency can be further simplified so that a simple decoder does not have to consider bitstreams with non-zero spatial_id: that is, one type of single-layer bitstream with only a single layer and one spatial ID (e.g., main-simple) and a single-layer bitstream with a single layer but which may contain multiple spatial IDs (e.g., main-representation).
[0197] Within a single layer, pictures with multiple spatial ID values may be considered to be part of different spatial access units: pictures with POC 0 with spatial_id equal to 0 belong to one spatial access unit, while pictures with POC 0 with spatial ID equal to 1 (in the same access unit) belong to another spatial access unit. The concept of spatial access units may not necessarily have a different relationship to output time (similar to layer access units), but they are important in different representations within the CPB and DPB. The output order may also be affected by the spatial access units when one or more spatial access units in a layer access unit are to be output. Pictures belonging to spatial_id may depend on other pictures with lower spatial_id; pictures belonging to spatial ID S may not reference pictures belonging to the same layer with spatial ID greater than S.
[0198] The maximum number of spatial layers may be specified for the bitstream, and the POC value for each picture may be derived from the picture order count LSB (syntax element) and the spatial ID. For example, the POC of pictures across layers may be constrained to be the same; whereas the POC of pictures in the same access unit within a layer but with different spatial IDs may not be the same. In some cases, this constraint may not be explicitly signaled, but derived in the decoder based on the pic_order_cnt_lsb and spatial_id values.
[0199] Figure 4 4 is a flowchart illustrating an example process 400 according to some examples. In some examples, process 400 is performed by a decoding device (e.g., decoding device 112). In some examples, process 400 may be performed by an encoding device (e.g., encoding device 104), such as when performing a decoding process for storing one or more reference pictures in a decoded picture buffer (DPB). In other examples, process 400 may be implemented as instructions in a non-transitory storage medium, and when a processor of the device executes the instructions, the instructions cause the device to perform process 400. In some cases, when process 400 is performed by a video decoder, the video data may include a decoded picture or a portion of a decoded picture (e.g., one or more blocks) included in an encoded video bitstream, or may include multiple decoded pictures included in an encoded video bitstream.
[0200] At block 402, process 400 obtains at least a portion of a picture from a bitstream. In some examples, the portion of the picture is a slice of the picture. As described herein, this can be a bitstream obtained by a mobile device such as a smart phone, a computer, any device with a screen, a television, or any device with decoding hardware.
[0201] At block 404, process 400 determines from the bitstream that weighted prediction is enabled for the portion of the picture. Process 400 may determine that weighted prediction is enabled by parsing the bitstream for a flag indicating weighted prediction. As described herein, parsing may be performed using an SPS flag and / or a PPS flag. In some cases, the PPS flag is constrained by an SPS flag. In some examples, a flag (e.g., an SPS flag and / or a PPS flag) may specify a unidirectional predicted frame (e.g., a P frame) or a bidirectional predicted frame (e.g., a B frame).
[0202] At block 406, based on determining that weighted prediction is enabled for the portion of the picture, process 400 identifies a zero-valued picture order count offset for indicating a reference picture from a reference picture list. As described herein, a zero-valued picture order count offset can identify a reference picture from a table, and the reference picture can then be used for weighted prediction. The reference picture for weighted prediction referenced by the zero-valued picture order count offset at block 404 can be a version of the picture that includes the portion of the picture obtained at block 402, wherein the picture from block 402 has different weighting parameters than the reference picture indicated by the zero-valued picture order count offset (e.g., a copy of the picture from block 402 with different weights).
[0203] At block 408, process 400 reconstructs at least a portion of a picture using at least a portion of a reference picture identified by a zero-valued picture order count offset. The reconstruction of the picture may be performed as part of reconstruction of the picture according to a bitstream (as part of a decoding operation compliant with VVC), or may be part of any similar decoding operation. The process 400 may be repeated for additional portions of the picture or other pictures obtained from the bitstream. As part of the process, different corresponding picture order count offsets for multiple portions of the picture are identified. The process may also be repeated for any number of pictures, some of which use weighted prediction and other pictures have weighted prediction disabled. When weighted prediction is enabled, some reference pictures of the video will have an associated picture order count offset value of zero, and other reference pictures will have a non-zero value, where frames from other access units are used for weighted prediction of the current frame. Then, as the process is repeated for additional pictures, the associated pictures are reconstructed.
[0204] In some examples, the additional frame identification is used to indicate a flag to disable weighted prediction. In some such embodiments, the additional pictures or picture portions or picture slices may be processed using the following operation: parsing a syntax element of the bitstream indicating a plurality of corresponding picture order count offsets by reconstructing the value of the syntax element to be signaled plus 1. In some such examples, when weighted prediction is disabled, the abs_delta_poc_st short-term reference picture syntax value specifies the absolute difference between the picture order count value of the second picture and the previous short-term reference picture entry in the reference picture list for the second reference picture as the abs_delta_poc_st short-term reference picture syntax value plus one.
[0205] In some examples, both weighted prediction and inter-layer prediction may be applied to the portion of the picture such that the reference picture has a different size than the picture, the reference picture being associated with a set of weights that is different from a second different set of weights associated with the picture (e.g., from block 402). In some such examples, the reference picture is signaled as a short-term reference picture having a picture order count least significant bit value equal to zero.
[0206] In some examples, the picture is a non-instantaneous decoding refresh (non-IDR) picture (eg, the picture may include another type of picture or a random access picture), as described above.
[0207] Figure 5 is a flow chart illustrating an example of a process 500 according to some examples. In some examples, the process 500 is performed by an encoding device (e.g., encoding device 104). In other examples, the process 500 may be implemented as instructions in a non-transitory storage medium, and when a processor of the device executes the instructions, the instructions cause the device to perform the process 500. In some cases, when the process 500 is performed by a video encoder, the video data may include a picture or a portion of a picture (e.g., one or more blocks) to be encoded in an encoded video bitstream, or may include multiple pictures to be encoded in an encoded video bitstream.
[0208] At block 502, process 500 includes identifying at least a portion of a picture. At block 504, process 500 includes selecting weighted prediction to be enabled for the portion of the picture. At block 506, process 500 includes identifying a reference picture for the portion of the picture. Identifying the reference picture may include selecting weights for each picture and associating the picture with a copy having a different weight as a reference picture for the first picture.
[0209] At block 508, process 500 includes generating a zero-valued picture order count offset for indicating a reference picture from a reference picture list. At block 510, process 500 includes generating a bitstream. The bitstream includes the portion of the picture and the zero-valued picture order count offset as associated with the portion of the picture.
[0210] As above for process 400, process 500 can be implemented in a variety of devices. In some implementations, process 500 is performed by processing circuitry of a mobile device having a camera coupled to a processor and a memory for capturing and storing at least one picture. In other examples, process 500 is performed by any device having a display coupled to a processor, wherein the display is configured to display at least one picture before the picture is processed and sent in a bitstream using process 400.
[0211] In some implementations, the processes (or methods) described herein (including processes 400 and 500) may be performed by a computing device or apparatus (such as on a Figure 1 For example, the process may be performed by the system 100 shown in Figure 1 and Figure 6 The encoding device 104 shown in FIG. 1 , by another video source side device or video transmission device, by Figure 1 and Figure 7 , and / or by another client-side device (such as a player device, a display, or any other client-side device). In some cases, the computing device or apparatus may include one or more input devices, one or more processors, one or more microprocessors, one or more microcomputers, and / or other components of a device configured to perform the steps of the processes described herein.
[0212] In some examples, the computing device or apparatus may include or may be a mobile device, a desktop computer, a server computer and / or a server system, or other types of computing devices. The components of the computing device (e.g., one or more input devices, one or more output devices, one or more processors, one or more microprocessors, one or more microcomputers, and / or other components) may be implemented in a circuit system. For example, the components may include electronic circuits or other electronic hardware and / or may be implemented using electronic circuits or other electronic hardware, which may include one or more programmable circuits (e.g., microprocessors, graphics processing units (GPUs), digital signal processors (DSPs), central processing units (CPUs), and / or other appropriate electronic circuits), and / or may include computer software, firmware, or any combination thereof and / or may be implemented using computer software, firmware, or any combination thereof to perform the various operations described herein. In some examples, the computing device or apparatus may include a camera configured to capture video data (e.g., a video sequence) including video frames. In some examples, the camera or other capture device that captures video data is separate from the computing device, in which case the computing device receives or obtains the captured video data. The computing device may include a network interface configured to transmit video data. The network interface may be configured to transmit Internet Protocol (IP) based data or other types of data.In some examples, the computing device or apparatus may include a display for displaying output video content (such as samples of pictures of a video bitstream).
[0213] The process may be described with respect to a logical flow diagram, the operations of which represent a series of operations that can be implemented using hardware, computer instructions, or a combination thereof. In the context of computer instructions, the operations represent computer-executable instructions stored on one or more computer-readable storage media that, when executed by one or more processors, implement the described operations. Typically, computer-executable instructions include routines, programs, objects, components, data structures, etc. that perform specific functions or implement specific data types. The order in which the operations are described is not intended to be construed as a limitation, and any number of the described operations may be combined in any order and / or in parallel to implement the process.
[0214] Additionally, the process may be executed under the control of one or more computer systems configured with executable instructions, and may be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that is executed together on one or more processors, implemented by hardware, or a combination thereof. As previously described, the code may be stored on a computer-readable or machine-readable storage medium, for example, in the form of a computer program that includes multiple instructions that can be executed by one or more processors. The computer-readable storage medium or machine-readable storage medium may be non-transitory.
[0215] The decoding techniques discussed herein may be implemented in an example video encoding and decoding system (e.g., system 100). In some examples, the system includes a source device that provides encoded video data to be decoded later by a destination device. Specifically, the source device provides the video data to the destination device via a computer-readable medium. The source device and the destination device may include any of a variety of devices, including desktop computers, notebook (i.e., laptop) computers, tablet computers, set-top boxes, telephone handsets (such as so-called "smart" phones), so-called "smart" boards, televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, etc. In some cases, the source device and the destination device may be equipped for wireless communication.
[0216] The destination device may receive the encoded video data to be decoded via a computer-readable medium. The computer-readable medium may include any type of medium or device capable of moving the encoded video data from the source device to the destination device. In one example, the computer-readable medium may include a communication medium for enabling the source device to directly send the encoded video data to the destination device in real time. The encoded video data may be modulated according to a communication standard such as a wireless communication protocol, and sent to the destination device. The communication medium may include any wireless or wired communication medium, such as a radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium may form a part of a packet-based network, such as a local area network, a wide area network, or a global network such as the Internet. The communication medium may include a router, a switch, a base station, or any other device that may be useful for facilitating communication from the source device to the destination device.
[0217] In some examples, the encoded data can be output to a storage device from an output interface. Similarly, the encoded data can be accessed from a storage device through an input interface. The storage device may include any of a variety of distributed or locally accessed data storage media, such as a hard drive, a Blu-ray disc, a DVD, a CD-ROM, a flash memory, a volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded video data. In a further example, the storage device may correspond to a file server or another intermediate storage device, which may store the encoded video generated by the source device. The destination device may access the stored video data from the storage device via streaming or downloading. The file server may be any type of server capable of storing encoded video data and sending the encoded video data to the destination device. Example file servers include web servers (e.g., for websites), FTP servers, network attached storage (NAS) devices, or local disk drives. The destination device may access the encoded video data through any standard data connection (including an Internet connection). This may include a wireless channel (e.g., a Wi-Fi connection), a wired connection (e.g., DSL, cable modem, etc.), or a combination of both suitable for accessing the encoded video data stored on the file server. The transmission of the encoded video data from the storage device may be a streaming transmission, a download transmission, or a combination thereof.
[0218] The technology of the present disclosure is not necessarily limited to wireless applications or settings. The technology can be applied to video decoding to support any of a variety of multimedia applications, such as over-the-air television broadcasting, cable television transmission, satellite television transmission, Internet streaming video transmission (such as dynamic adaptive streaming (DASH) via HTTP), digital video encoded on a data storage medium, decoding of digital video stored on a data storage medium, or other applications. In some examples, the system can be configured to support one-way or two-way video transmission to support applications such as video streaming, video playback, video broadcasting, and / or video telephony.
[0219] In one example, the source device includes a video source, a video encoder, and an output interface. The destination device may include an input interface, a video decoder, and a display device. The video encoder of the source device may be configured to apply the technology disclosed herein. In other examples, the source device and the destination device may include other components or arrangements. For example, the source device may receive video data from an external video source such as an external camera. Similarly, the destination device may interface with an external display device rather than including an integrated display device.
[0220] The example system above is only an example. The technology for processing video data in parallel can be performed by any digital video encoding and / or decoding device. Although in general, the technology of the present disclosure is performed by a video encoding device, the technology can also be performed by a video encoder / decoder commonly referred to as a "CODEC". In addition, the technology of the present disclosure can also be performed by a video preprocessor. The source device and the destination device are only examples of encoding devices such as follows: in which the source device generates encoded video data for transmission to the destination device. In some examples, the source device and the destination device can operate in a substantially symmetrical manner so that each device in the device includes a video encoding and decoding component. Therefore, the example system can support one-way or two-way video transmission between video devices, for example, for video streaming, video playback, video broadcasting or video telephony.
[0221] The video source can include a video capture device, such as a video camera, a video archive comprising previously captured video, and / or a video feed interface for receiving video from a video content provider. As a further alternative, the video source can generate computer graphics-based data as the source video, or generate a combination of real-time video, archived video, and computer-generated video. In some cases, if the video source is a video camera, the source device and the destination device can form a so-called camera phone or video phone. However, as described above, the technology described in this disclosure can be generally applicable to video encoding, and can be applied to wireless and / or wired applications. In each case, the captured, pre-captured or computer-generated video can be encoded by a video encoder. Then, the encoded video information can be output to the computer-readable medium through an output interface.
[0222] As mentioned, computer-readable media may include transient media such as wireless broadcast or wired network transmissions, or storage media such as hard disks, flash drives, compact discs, digital versatile discs, Blu-ray discs (i.e., non-transitory storage media), or other computer-readable media. In some examples, a network server (not shown) may receive encoded video data from a source device, for example, via a network transmission, and provide the encoded video data to a destination device. Similarly, a computing device of a media production facility such as a disc stamping facility may receive encoded video data from a source device, and manufacture a disc containing the encoded video data. Therefore, in various examples, a computer-readable medium may be understood to include one or more computer-readable media in various forms.
[0223] An input interface of the destination device receives information from a computer-readable medium. The information of the computer-readable medium may include syntax information defined by a video encoder (which is also used by a video decoder), the syntax information including syntax elements that describe characteristics and / or processing of blocks and other decoding units (e.g., groups of pictures (GOPs)). A display device displays the decoded video data to a user and may include any of a variety of display devices, such as a cathode ray tube (CRT), a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or another type of display device. Various examples of the present application have been described. In Figure 6 and 7 Specific details of the encoding device 104 and the decoding device 112 are shown separately.
[0224] The specific details of the encoding device 104 and the decoding device 112 are in Figure 6 and Figure 7 are shown separately in . Figure 6 1 is a block diagram illustrating an example encoding device 104 that can implement one or more of the techniques described in the present disclosure. The encoding device 104 can, for example, generate the syntax structures described herein (e.g., syntax structures of VPS, SPS, PPS or other syntax elements). The encoding device 104 can perform intra-frame prediction and inter-frame prediction decoding of video blocks within a video. As previously described, intra-frame decoding relies at least in part on spatial prediction to reduce or remove spatial redundancy within a given video frame or picture. Inter-frame decoding relies at least in part on temporal prediction to reduce or remove temporal redundancy within adjacent or surrounding frames of a video sequence. Intra-frame mode (I mode) can refer to any of several spatial-based compression modes. Inter-frame modes such as unidirectional prediction (P mode) or bidirectional prediction (B mode) can refer to any of several temporal-based compression modes.
[0225] The encoding device 104 includes a segmentation unit 35, a prediction processing unit 41, a filter unit 63, a picture memory 64, an adder 50, a transform processing unit 52, a quantization unit 54, and an entropy encoding unit 56. The prediction processing unit 41 includes a motion estimation unit 42, a motion compensation unit 44, and an intra-frame prediction processing unit 46. For video block reconstruction, the encoding device 104 also includes an inverse quantization unit 58, an inverse transform processing unit 60, and an adder 62. The filter unit 63 is intended to represent one or more loop filters, such as a deblocking filter, an adaptive loop filter (ALF), and a sample adaptive offset (SAO) filter. Although in Figure 6 Filter unit 63 is shown as an in-loop filter in FIG. 1 , but in other configurations, filter unit 63 may be implemented as a post-loop filter. Post-processing device 57 may perform additional processing on the encoded video data generated by encoding device 104. In some examples, the techniques of the present disclosure may be implemented by encoding device 104. However, in other examples, one or more of the techniques of the present disclosure may be implemented by post-processing device 57.
[0226] like Figure 6 As shown, the encoding device 104 receives video data, and the segmentation unit 35 segments the data into video blocks. Segmentation may also include, for example, segmentation into slices, slice segments, tiles or other larger units according to the quadtree structure of LCU and CU, as well as video block segmentation. The encoding device 104 generally shows a component for encoding a video block within a video slice to be encoded. A slice may be divided into a plurality of video blocks (and may be divided into a set of video blocks called tiles). The prediction processing unit 41 may select a decoding mode in a plurality of possible decoding modes for the current video block based on error results (e.g., coding rate and distortion level, etc.), such as an intra-frame prediction decoding mode in a plurality of intra-frame prediction decoding modes or an inter-frame prediction decoding mode in a plurality of inter-frame prediction decoding modes. The prediction processing unit 41 may provide the resulting intra-frame or inter-frame decoded block to the adder 50 to generate residual block data, and to the adder 62 to reconstruct the encoded block for use as a reference picture.
[0227] Intra-prediction processing unit 46 within prediction processing unit 41 may perform intra-prediction decoding of the current video block relative to one or more neighboring blocks in the same frame or slice as the current video block to be decoded to provide spatial compression. Motion estimation unit 42 and motion compensation unit 44 within prediction processing unit 41 may perform inter-prediction decoding of the current video block relative to one or more prediction blocks in one or more reference pictures to provide temporal compression.
[0228] Motion estimation unit 42 may be configured to determine an inter-prediction mode for a video slice according to a predetermined pattern for a video sequence. The predetermined pattern may designate a video slice in a sequence as a P slice, a B slice, or a GPB slice. Motion estimation unit 42 and motion compensation unit 44 may be highly integrated, but are shown separately for conceptual purposes. Motion estimation performed by motion estimation unit 42 is the process of generating a motion vector that estimates motion for a video block. A motion vector may, for example, indicate the displacement of a prediction unit (PU) of a video block within a current video frame or picture relative to a prediction block within a reference picture.
[0229] A prediction block is a block that is found to closely match a PU of a video block to be decoded in terms of pixel difference, which may be determined by sum of absolute difference (SAD), sum of squared difference (SSD), or other difference metrics. In some examples, encoding device 104 may calculate values for sub-integer pixel positions for a reference picture stored in picture memory 64. For example, encoding device 104 may interpolate values for quarter-pixel positions, eighth-pixel positions, or other fractional pixel positions of a reference picture. Thus, motion estimation unit 42 may perform motion searches relative to full pixel positions and fractional pixel positions, and output motion vectors with fractional pixel precision.
[0230] Motion estimation unit 42 calculates a motion vector for a PU by comparing the position of the PU of the video block in the inter-coded slice with the position of a prediction block of a reference picture. The reference picture may be selected from a first reference picture list (list 0) or a second reference picture list (list 1), each of which identifies one or more reference pictures stored in picture memory 64. Motion estimation unit 42 sends the calculated motion vector to entropy encoding unit 56 and motion compensation unit 44.
[0231] Motion compensation performed by the motion compensation unit 44 may involve extracting or generating a prediction block based on a motion vector determined by motion estimation, possibly performing interpolation to sub-pixel precision. Upon receiving a motion vector for a PU of a current video block, the motion compensation unit 44 may locate the prediction block pointed to by the motion vector in a reference picture list. The encoding device 104 forms a residual video block by subtracting the pixel values of the prediction block from the pixel values of the current video block being decoded, forming pixel difference values. The pixel difference values form residual data for the block, and may include both a luma difference component and a chroma difference component. The adder 50 represents a component or components that perform such a subtraction operation. The motion compensation unit 44 may also generate syntax elements associated with video blocks and video slices for use by the decoding device 112 when decoding video blocks of video slices.
[0232] As described above, the intra-prediction processing unit 46 may perform intra-prediction on the current block as an alternative to the inter-prediction performed by the motion estimation unit 42 and the motion compensation unit 44. Specifically, the intra-prediction processing unit 46 may determine an intra-prediction mode to be used to encode the current block. In some examples, the intra-prediction processing unit 46 may use various intra-prediction modes to encode the current block (e.g., during a separate encoding pass), and the intra-prediction processing unit 46 may select a suitable intra-prediction mode from the tested modes to use. For example, the intra-prediction processing unit 46 may calculate a rate-distortion value using a rate-distortion analysis for various tested intra-prediction modes, and may select an intra-prediction mode having the best rate-distortion characteristic among the tested modes. The rate-distortion analysis typically determines the amount of distortion (or error) between the coded block and the original uncoded block that is coded to generate the coded block, and the bit rate (i.e., the number of bits) used to generate the coded block. Intra-prediction processing unit 46 may calculate ratios from the distortions and rates for the various encoded blocks to determine which intra-prediction mode exhibits the best rate-distortion value for the block.
[0233] In any case, after selecting an intra-prediction mode for a block, the intra-prediction processing unit 46 may provide information indicating the intra-prediction mode selected for the block to the entropy encoding unit 56. The entropy encoding unit 56 may encode the information indicating the selected intra-prediction mode. The encoding device 104 may include in the transmitted bitstream configuration data the definition of the coding contexts for the various blocks, as well as an indication of the most probable intra-prediction mode to be used for each of the contexts, an intra-prediction mode index table, and a modified intra-prediction mode index table. The bitstream configuration data may include a plurality of intra-prediction mode index tables and a plurality of modified intra-prediction mode index tables (also referred to as codeword mapping tables).
[0234] After the prediction processing unit 41 generates a prediction block for the current video block via inter-frame prediction or intra-frame prediction, the encoding device 104 forms a residual video block by subtracting the prediction block from the current video block. The residual video data in the residual block may be included in one or more TUs and applied to the transform processing unit 52. The transform processing unit 52 uses a transform (such as a discrete cosine transform (DCT) or a conceptually similar transform) to transform the residual video data into residual transform coefficients. The transform processing unit 52 may convert the residual video data from a pixel domain to a transform domain (such as a frequency domain).
[0235] Transform processing unit 52 may send the resulting transform coefficients to quantization unit 54. Quantization unit 54 quantizes the transform coefficients to further reduce the bit rate. The quantization process may reduce the bit depth associated with some or all of the coefficients. The degree of quantization may be modified by adjusting a quantization parameter. In some examples, quantization unit 54 may then perform a scan of the matrix including the quantized transform coefficients. Alternatively, entropy encoding unit 56 may perform the scan.
[0236] After quantization, entropy encoding unit 56 performs entropy encoding on the quantized transform coefficients. For example, entropy encoding unit 56 may perform context adaptive variable length coding (CAVLC), context adaptive binary arithmetic coding (CABAC), syntax-based context adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or another entropy encoding technique. After entropy encoding by entropy encoding unit 56, the encoded bitstream may be sent to decoding device 112, or archived for later transmission or retrieved by decoding device 112. Entropy encoding unit 56 may also entropy encode motion vectors and other syntax elements for the current video slice being decoded.
[0237] Inverse quantization unit 58 and inverse transform processing unit 60 apply inverse quantization and inverse transform, respectively, to reconstruct the residual block in the pixel domain for later use as a reference block of a reference picture. Motion compensation unit 44 may calculate a reference block by adding the residual block to a prediction block of one of the reference pictures in the reference picture list. Motion compensation unit 44 may also apply one or more interpolation filters to the reconstructed residual block to calculate sub-integer pixel values for use in motion estimation. Adder 62 adds the reconstructed residual block to the motion compensated prediction block produced by motion compensation unit 44 to produce a reference block for storage in picture memory 64. The reference block may be used by motion estimation unit 42 and motion compensation unit 44 as a reference block to inter-predict a block in a subsequent video frame or picture.
[0238] In this way, Figure 6 The encoding device 104 is configured to perform one or more of the techniques described herein (including the above-mentioned Figure 4 and Figure 5 In some cases, some of the techniques of the present disclosure may also be implemented by post-processing device 57.
[0239] Figure 7 1 is a block diagram illustrating an example decoding device 112. The decoding device 112 includes an entropy decoding unit 80, a prediction processing unit 81, an inverse quantization unit 86, an inverse transform processing unit 88, an adder 90, a filter unit 91, and a picture memory 92. The prediction processing unit 81 includes a motion compensation unit 82 and an intra-frame prediction processing unit 84. In some examples, the decoding device 112 may perform operations generally associated with the processing of the image from Figure 6 The encoding device 104 described herein is a decoding path that is the inverse of the encoding path.
[0240] During the decoding process, the decoding device 112 receives a coded video bitstream sent by the encoding device 104, the coded video bitstream representing video blocks and associated syntax elements of the coded video slice. In some examples, the decoding device 112 may receive the coded video bitstream from the encoding device 104. In some examples, the decoding device 112 may receive the coded video bitstream from the network entity 79 (such as a server, a media-aware network element (MANE), a video editor / splicer, or other such devices configured to implement one or more of the techniques described above). The network entity 79 may or may not include the encoding device 104. Some of the techniques described in the present disclosure may be implemented by the network entity 79 before the network entity 79 sends the coded video bitstream to the decoding device 112. In some video decoding systems, the network entity 79 and the decoding device 112 may be parts of separate devices, while in other instances, the functions described with respect to the network entity 79 may be performed by the same device including the decoding device 112.
[0241] The entropy decoding unit 80 of the decoding device 112 entropy decodes the bitstream to generate quantized coefficients, motion vectors and other syntax elements. The entropy decoding unit 80 forwards the motion vectors and other syntax elements to the prediction processing unit 81. The decoding device 112 can receive syntax elements at the video slice level and / or the video block level. The entropy decoding unit 80 can process and parse both fixed length syntax elements and variable length syntax elements in more parameter sets such as VPS, SPS and PPS.
[0242] When the video slice is encoded as an intra-coded (I) slice, the intra-prediction processing unit 84 of the prediction processing unit 81 can generate prediction data for the video block of the current video slice based on the intra-prediction mode transmitted by the signal and the data of the previously decoded block from the current frame or picture. When the video frame is decoded as an inter-coded (i.e., B, P or GPB) slice, the motion compensation unit 82 of the prediction processing unit 81 generates a prediction block for the video block of the current video slice based on the motion vector and other syntax elements received from the entropy decoding unit 80. The prediction block can be generated from one of the reference pictures in the reference picture list. The decoding device 112 can construct the reference frame lists, List 0 and List 1, using the construction technique based on the reference pictures stored in the picture memory 92.
[0243] Motion compensation unit 82 determines prediction information for a video block of the current video slice by parsing the motion vectors and other syntax elements, and uses the prediction information to generate a prediction block for the current video block being decoded. For example, motion compensation unit 82 may use one or more syntax elements in a parameter set to determine a prediction mode (e.g., intra or inter prediction) for coding a video block of a video slice, an inter-prediction slice type (e.g., a B slice, a P slice, or a GPB slice), construction information for one or more reference picture lists for the slice, motion vectors for each inter-coded video block of the slice, inter-prediction status for each inter-coded video block of the slice, and other information for decoding video blocks in the current video slice.
[0244] The motion compensation unit 82 may also perform interpolation based on an interpolation filter. The motion compensation unit 82 may calculate interpolated values for pixels below an integer of a reference block using an interpolation filter as used by the encoding device 104 during encoding of the video block. In this case, the motion compensation unit 82 may determine the interpolation filter used by the encoding device 104 based on the received syntax elements, and may use the interpolation filter to generate a prediction block.
[0245] The inverse quantization unit 86 inverse quantizes or dequantizes the quantized transform coefficients provided in the bitstream and decoded by the entropy decoding unit 80. The inverse quantization process may include using the quantization parameters calculated by the encoding device 104 for each video block in the video slice to determine the degree of quantization, and likewise the degree of inverse quantization that should be applied. The inverse transform processing unit 88 applies an inverse transform (e.g., an inverse DCT or other suitable inverse transform), an inverse integer transform, or a conceptually similar inverse transform process to the transform coefficients to produce residual blocks in the pixel domain.
[0246] After the motion compensation unit 82 generates a prediction block for the current video block based on the motion vector and other syntax elements, the decoding device 112 forms a decoded video block by adding the residual block from the inverse transform processing unit 88 to the corresponding prediction block generated by the motion compensation unit 82. Adder 90 represents a component or multiple components that perform this summation operation. If desired, a loop filter (in the encoding loop or after the encoding loop) can also be used to smooth pixel transitions or otherwise improve video quality. Filter unit 91 is intended to represent one or more loop filters, such as a deblocking filter, an adaptive loop filter (ALF), and a sample adaptive offset (SAO) filter. Although in Figure 7 Filter unit 91 is shown as an in-loop filter in FIG. 1 , but in other configurations, filter unit 91 may be implemented as a post-loop filter. The decoded video blocks in a given frame or picture are then stored in picture memory 92, which stores reference pictures for subsequent motion compensation. Picture memory 92 also stores the decoded video for later display on a display device (e.g., Figure 1 122).
[0247] In this way, Figure 7 The decoding device 112 represents a device configured to perform one or more of the techniques described herein (including the above-mentioned Figure 3 and Figure 4 An example of a video decoder of the process described above.
[0248] The filter unit 91 filters the reconstructed block (e.g., the output of the adder 90), and stores the filtered reconstructed block in the DPB 94 for use as a reference block and / or outputs the filtered reconstructed block (decoded video). The reference block can be used as a reference block by the motion compensation unit 82 to perform inter-frame prediction on blocks in subsequent video frames or pictures. The filter unit 91 can perform any type of filtering, such as deblocking filtering, SAO filtering, peak SAO filtering, ALF and / or GALF, and / or other types of loop filters. The deblocking filter can, for example, apply deblocking filtering to filter block boundaries to remove blocky artifacts from the reconstructed video. The peak SAO filter can apply an offset to the reconstructed pixel value to improve the overall decoding quality. Additional loop filters (in the loop or after the loop) can also be used.
[0249] In addition, the filter unit 91 may be configured to perform any of the techniques related to adaptive loop filtering in the present disclosure. For example, as described above, the filter unit 91 may be configured to determine the parameters for filtering the current block based on the parameters for filtering the previous block included in the same APS as the current block, a different APS, or a predefined filter.
[0250] As used herein, the term "computer-readable medium" includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other media capable of storing, containing or carrying instructions and / or data. Computer-readable media may include non-transient media in which data may be stored and does not include carrier waves and / or temporary electronic signals that are propagated wirelessly or on a wired connection. Examples of non-transient media may include, but are not limited to, disks or tapes, optical storage media such as compact discs (CDs) or digital versatile discs (DVDs), flash memory, memory, or memory devices. Computer-readable media may have codes and / or machine-executable instructions stored thereon, which may represent any combination of processes, functions, subroutines, programs, routines, subroutines, modules, software packages, classes, or instructions, data structures, or program statements. Code segments may be coupled to another code segment or hardware circuit by transmitting and / or receiving information, data, independent variables, parameters, or memory contents. Information, independent variables, parameters, data, etc. may be transmitted, forwarded, or sent via any appropriate means including memory sharing, message passing, token passing, network transmission, etc.
[0251] In some examples, computer-readable storage devices, media, and memories may include cables or wireless signals containing bit streams, etc. However, when referred to, non-transitory computer-readable storage media explicitly excludes media such as energy, carrier signals, electromagnetic waves, and signals themselves.
[0252] Specific details are provided in the above description to provide a thorough understanding and examples of the examples provided herein. However, it will be appreciated by those of ordinary skill in the art that examples may be implemented without these specific details. For clarity of explanation, in some instances, the technology herein may be presented as a separate functional block including the following functional block, the functional block including equipment, device components, steps or routines in the method embodied in software, or a combination of hardware and software. In addition to the components shown in the figures and / or described herein, additional components may be used. For example, circuits, systems, networks, processes, and other components may be shown as components in block diagram form, so as not to make the examples difficult to understand in unnecessary details. In other instances, circuits, processes, algorithms, structures, and techniques are shown as known without unnecessary details, so as to avoid making the examples difficult to understand.
[0253] Individual examples may be described as processes or methods above, and the process or method is depicted as a flow chart, a schematic flow chart, a data flow chart, a structure diagram or a block diagram. Although a flow chart can describe an operation as a sequential process, many operations in the operation can be performed in parallel or simultaneously. In addition, the order of the operation can be rearranged. A process terminates when its operation is completed, but may have additional steps that are not included in the figure. A process may correspond to a method, a function, a procedure, a subroutine, a subprogram, etc. When a process corresponds to a function, its termination may correspond to the function returning to a calling function or a main function.
[0254] The process and method according to the example described above can be implemented using computer executable instructions, which are stored in a computer readable medium or can be obtained from a computer readable medium in other ways. Such instructions may include, for example, instructions or data, which enable or otherwise configure a general-purpose computer, a special-purpose computer or a processing device to perform a certain function or group of functions. The part of the computer resources used may be accessible through a network. Computer executable instructions may be, for example, binary, such as intermediate format instructions of assembly language, firmware, source code, etc. Examples of computer readable media that can be used to store instructions, information used, and / or information created during the method according to the described examples include disks or optical disks, flash memory, USB devices with non-volatile memory, storage devices of networks, etc.
[0255] The equipment implementing the process and method according to these disclosures may include hardware, software, firmware, middleware, microcode, hardware description language or any combination thereof, and may adopt any of various form factors. When implemented with software, firmware, middleware or microcode, the program code or code segment (e.g., computer program product) for performing the necessary tasks may be stored in a computer-readable or machine-readable medium. The processor may perform the necessary tasks. Typical examples of form factors include personal computers, personal digital assistants, rack-mounted devices, stand-alone devices, etc., of laptop computers, smart phones, mobile phones, tablet devices or other small form factors. The functions described herein may also be embodied in peripheral devices or plug-in cards. By further example, such functions may also be implemented on circuit boards among different chips or different processes executed in a single device.
[0256] Instructions, media for communicating such instructions, computing resources for executing them, and other structures for supporting such computing resources are example means for providing the functionality described in this disclosure.
[0257] In the foregoing description, various aspects of the present application are described with reference to the specific examples of the present application, but those skilled in the art will recognize that the present application is not limited thereto. Therefore, although the illustrative examples of the present application have been described in detail herein, it should be understood that the inventive concept may be embodied and adopted differently in other ways, and the appended claims are intended to be interpreted as including such variants, except for the variants limited by the prior art. The various features and aspects of the application described above may be used individually or collectively. Further, without departing from the broader spirit and scope of this specification, the examples may be utilized in any number of environments and applications other than the environments and applications described herein. Therefore, the description and the accompanying drawings are considered to be illustrative rather than restrictive. For the purpose of illustration, the method is described in a specific order. It should be understood that in an alternative example, the method may be performed in an order different from the described order.
[0258] It will be understood by those of ordinary skill in the art that the less than ("<") and greater than (">") symbols or terms used herein may be replaced with less than or equal to ("≤") and greater than or equal to ("≥") symbols, respectively, without departing from the scope of the present specification.
[0259] Where a component is described as being “configured to” perform certain operations, such configuration may be accomplished, for example, by designing an electronic circuit or other hardware to perform the operation, by programming a programmable electronic circuit (e.g., a microprocessor or other appropriate circuit) to perform the operation, or any combination thereof.
[0260] The phrase "coupled to" refers to any component that is directly or indirectly physically connected to another component, and / or any component that directly or indirectly communicates with another component (e.g., connected to another component via a wired or wireless connection and / or other appropriate communication interface).
[0261] Claim language or other language that recites "at least one of" a set and / or "one or more of" a set indicates that one member of the set or multiple members of the set (in any combination) satisfies the claim. For example, claim language that recites "at least one of A and B" means A, B, or A and B. In another example, claim language that recites "at least one of A, B, and C" means A, B, C, or A and B, or A and C, or B and C, or A and B and C. The language "at least one of" a set and / or "one or more of" a set does not limit the set to the items listed in the set. For example, claim language that recites "at least one of A and B" may mean A, B, or A and B, and may additionally include items not listed in the set of A and B.
[0262] Various illustrative logical blocks, modules, circuits and algorithmic steps described in conjunction with the examples disclosed herein can be implemented as electronic hardware, computer software, firmware or a combination thereof. In order to clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits and steps have been generally described above around their functions. Whether such functional implementation is hardware or software depends on specific applications and the design constraints imposed on the entire system. Those skilled in the art can implement the described functions in an alternative manner for each specific application, but such implementation decision should not be interpreted as causing deviations from the scope of the present application.
[0263] The techniques described herein may also be implemented in electronic hardware, computer software, firmware, or any combination thereof. Such techniques may be implemented in any of a variety of devices, such as general-purpose computers, wireless communication device phones, or integrated circuit devices with multiple uses, including applications in wireless communication device phones and other devices. Any features described as modules or components may be implemented together in an integrated logic device, or separately as discrete but interoperable logic devices. If implemented in software, the techniques may be implemented at least in part by a computer-readable data storage medium, which includes a program code, which includes instructions for implementing one or more of the methods described above when executed. The computer-readable data storage medium may form part of a computer program product, which may include packaging materials. The computer-readable medium may include a memory or data storage medium, such as a random access memory (RAM) (such as a synchronous dynamic random access memory (SDRAM)), a read-only memory (ROM), a non-volatile random access memory (NVRAM), an electrically erasable programmable read-only memory (EEPROM), a flash memory, a magnetic or optical data storage medium, and the like. Additionally or alternatively, the techniques may be implemented at least in part by a computer-readable communication medium (such as a propagated signal or wave) that carries or communicates program code in the form of instructions or data structures and can be accessed, read, and / or executed by a computer.
[0264] The program code can be executed by a processor, which may include one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Such a processor can be configured to perform any of the technologies described in the present disclosure. A general-purpose processor can be a microprocessor, but in an alternative, the processor can be any conventional processor, controller, microcontroller, or state machine. The processor can also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, a plurality of microprocessors, a combination of one or more microprocessors and a DSP core, or any other such configuration. Therefore, the term "processor" as used herein may refer to any of the aforementioned structures, any combination of the aforementioned structures, or any other structure or device suitable for implementing the technology described herein. In addition, in some aspects, the functions described herein may be provided in a dedicated software module or hardware module configured for encoding and decoding, or incorporated in a combined video encoder-decoder (CODEC).
[0265] Illustrative examples of the present disclosure include:
[0266] Example 1. A method for processing video data, the method comprising: obtaining video data including one or more pictures; and determining that weighted prediction is disabled for at least a portion of the one or more pictures; and in response to determining that weighted prediction is disabled for at least a portion of the one or more pictures, determining that the same picture order count (POC) value is not allowed for more than one reference picture.
[0267] Example 2. The method of Example 1, wherein at least a portion of the one or more pictures comprises a slice of the picture.
[0268] Example 3. The method according to any one of Examples 1 to 2 further includes: obtaining a weighted prediction flag; and determining to disable weighted prediction for one or more pictures based on the weighted prediction flag.
[0269] Example 4. The method according to Example 3, wherein the weighted prediction flag indicates whether weighted prediction is applied to a unidirectionally predicted slice (P slice).
[0270] Example 5. The method according to Example 3, wherein the weighted prediction flag indicates whether weighted prediction is applied to a bidirectionally predicted slice (B slice).
[0271] Example 6. An apparatus comprising: a memory configured to store video data; and a processor configured to process the video data according to any one of Examples 1 to 5.
[0272] Example 7. An apparatus according to Example 6, wherein the apparatus includes an encoder.
[0273] Example 8. The apparatus of Example 6, wherein the apparatus comprises a decoder.
[0274] Example 9. An apparatus according to any one of Examples 6 to 8, wherein the apparatus is a mobile device.
[0275] Example 10. The apparatus according to any one of Examples 6 to 9, further comprising: a display configured to display video data.
[0276] Example 11. The apparatus according to any one of Examples 6 to 10 further includes: a camera configured to capture one or more pictures.
[0277] Example 12. A computer-readable medium having instructions stored thereon, which when executed by a processor implement the method according to any one of Examples 1 to 5.
[0278] Example 13. A method for processing video data, the method comprising: obtaining video data including multiple reference pictures in multiple layers; and generating a reference picture structure (RPS), the RPS comprising a picture order count (POC) value and a layer identifier (ID) for at least one reference picture among the multiple reference pictures, wherein the layer ID for the reference picture identifies the layer of the reference picture.
[0279] Example 14. The method of Example 13, further comprising: generating the RPS to include a layer ID for each reference picture of the plurality of reference pictures.
[0280] Example 15. The method according to Example 13 further includes: when the POC value of the reference picture is equal to the POC value of the current picture, generating the RPS to include a layer ID for the reference picture.
[0281] Example 16. An apparatus comprising: a memory configured to store video data; and a processor configured to process the video data according to any of Examples 13 to 15.
[0282] Example 17. An apparatus according to Example 16, wherein the apparatus includes an encoder.
[0283] Example 18. An apparatus according to any one of Examples 16 to 17, wherein the apparatus is a mobile device.
[0284] Example 19. The apparatus according to any one of Examples 16 to 18, further comprising: a display configured to display video data.
[0285] Example 20. The apparatus of any one of Examples 16 to 19, further comprising: a camera configured to capture one or more pictures.
[0286] Example 21. An apparatus according to any one of Examples 16 to 20, further comprising: a communication circuit configured to send processed video data.
[0287] Example 22. A computer-readable medium having instructions stored thereon, which when executed by a processor implement the method according to any one of Examples 13 to 15.
[0288] Example 23. A method for processing video data, the method comprising: obtaining video data including multiple reference pictures in multiple layers; and processing a reference picture structure (RPS) based on the video data, the RPS comprising a picture order count (POC) value and a layer identifier (ID) for at least one reference picture among the multiple reference pictures, wherein the layer ID for the reference picture identifies the layer of the reference picture.
[0289] Example 24. The method of Example 23, wherein the RPS includes a layer ID for each reference picture in the plurality of reference pictures.
[0290] Example 25. The method of Example 23, wherein when the POC value of the reference picture is equal to the POC value of the current picture, the RPS includes a layer ID for the reference picture.
[0291] Example 26. An apparatus comprising: a memory configured to store video data; and a processor configured to process the video data according to any of Examples 23 to 25.
[0292] Example 27. The apparatus of Example 26, wherein the apparatus comprises a decoder.
[0293] Example 28. An apparatus according to any one of Examples 26 to 27, wherein the apparatus is a mobile device.
[0294] Example 29. The apparatus according to any one of Examples 26 to 28, further comprising: a display configured to display video data.
[0295] Example 30. The apparatus of any one of Examples 26 to 29 further includes a camera configured to capture one or more pictures.
[0296] Example 31. A computer-readable medium having instructions stored thereon, which when executed by a processor implement the method according to any one of Examples 23 to 25.
[0297] Example 32. A method for processing video data, the method comprising: obtaining video data including multiple pictures in multiple layers, the multiple layers corresponding to multiple representations of video content; determining to enable inter-layer prediction for the multiple pictures; and inferring a zero motion vector based on determining to enable inter-layer prediction for the multiple pictures, the zero motion vector indicating that there is no content displacement in the multiple pictures of the multiple representations.
[0298] Example 33. A method according to Example 33, wherein when the motion information is not used for inter-layer prediction, the motion information is transmitted by a signal in an encoded video bitstream including multiple pictures.
[0299] Example 34. An apparatus comprising: a memory configured to store video data; and a processor configured to process the video data according to any of Examples 32 to 33.
[0300] Example 35. An apparatus according to Example 34, wherein the apparatus includes an encoder.
[0301] Example 36. The apparatus of Example 34, wherein the apparatus comprises a decoder.
[0302] Example 37. An apparatus according to any one of Examples 34 to 36, wherein the apparatus is a mobile device.
[0303] Example 38. The apparatus according to any one of Examples 34 to 37 further includes: a display configured to display video data.
[0304] Example 39. The apparatus of any one of Examples 34 to 38 further includes a camera configured to capture one or more pictures.
[0305] Example 40. A computer-readable medium having instructions stored thereon, the instructions, when executed by a processor, implementing the method according to any one of Examples 32 to 33.
[0306] Example 41. A method for processing video data, the method comprising: obtaining video data; and generating an encoded video bitstream based on the video data, the encoded video bitstream comprising one or more syntax elements and multiple pictures in multiple layers, the multiple layers corresponding to multiple representations of video content, wherein the number of entries in the inter-layer reference picture entry syntax element is limited to not more than the maximum number of reference pictures minus the number of inter-frame prediction reference pictures.
[0307] Example 42. The method of Example 41, wherein the inter-layer reference picture entry syntax element specifies the number of direct inter-layer reference picture entries in the reference picture list syntax structure.
[0308] Example 43. An apparatus comprising: a memory configured to store video data; and a processor configured to process the video data according to any of Examples 41 to 42.
[0309] Example 44. An apparatus according to Example 43, wherein the apparatus includes an encoder.
[0310] Example 45. The apparatus of Example 43, wherein the apparatus is a mobile device.
[0311] Example 46. The apparatus according to any one of Examples 43 to 45 further includes: a display configured to display video data.
[0312] Example 47. The apparatus of any one of Examples 43 to 46 further includes a camera configured to capture one or more pictures.
[0313] Example 48. A computer-readable medium having instructions stored thereon, which when executed by a processor implement the method according to any one of Examples 41 to 42.
[0314] Example 49. A method for processing video data, the method comprising: obtaining an encoded video bitstream, the encoded video bitstream comprising one or more syntax elements and multiple pictures in multiple layers, the multiple layers corresponding to multiple representations of video content; and processing an inter-layer reference picture entry syntax element based on the encoded video bitstream, wherein the number of entries in the inter-layer reference picture entry syntax element is limited to not more than the maximum number of reference pictures minus the number of inter-frame prediction reference pictures.
[0315] Example 50. The method of Example 33, wherein the inter-layer reference picture entry syntax element specifies the number of direct inter-layer reference picture entries in the reference picture list syntax structure.
[0316] Example 51. An apparatus comprising: a memory configured to store video data; and a processor configured to process the video data according to any of Examples 49 to 50.
[0317] Example 52. The apparatus of Example 51, wherein the apparatus comprises a decoder.
[0318] Example 53. An apparatus according to Example 51, wherein the apparatus is a mobile device.
[0319] Example 54. An example apparatus according to any one of Examples 51 to 53, further comprising: a display configured to display video data.
[0320] Example 55. An example apparatus according to any one of Examples 51 to 54, further comprising: a camera configured to capture one or more pictures.
[0321] Example 56. An example computer-readable medium having instructions stored thereon, which when executed by a processor implements the method of any one of Examples 49 to 50.
[0322] Example 57. A method for decoding video data, the method comprising: obtaining at least a portion of a picture from a bitstream; determining, based on the bitstream, to enable weighted prediction for a portion of the picture; based on determining to enable weighted prediction for a portion of the picture, identifying a zero-valued picture order count offset for indicating a reference picture from a reference picture list; and reconstructing at least a portion of the picture using at least a portion of the reference picture identified by the zero-valued picture order count offset.
[0323] Example 58. The method according to Example 57 further includes: obtaining multiple parts of a picture from a bitstream, the multiple parts including at least a part of the picture reconstructed using a reference picture identified by a zero-valued picture sequence count offset; identifying multiple corresponding picture sequence count offsets for the multiple parts of the picture, wherein the zero-valued picture sequence count offset is a corresponding picture sequence count offset for the picture among the multiple corresponding picture sequence count offsets; and reconstructing the picture using multiple reference pictures identified by the multiple corresponding picture sequence count offsets.
[0324] Example 59. The method according to Example 57 further includes: determining to enable weighted prediction by parsing the bitstream to identify one or more weighted prediction flags for the picture.
[0325] Example 60. A method according to Example 59, wherein the one or more weighted prediction tags for a picture include a sequence parameter set weighted prediction tag and a picture parameter set weighted prediction tag.
[0326] Example 61. A method according to Example 60, wherein the sequence parameter set weighted prediction flag and the picture parameter set weighted prediction flag are flags for unidirectional prediction frames.
[0327] Example 62. A method according to Example 60, wherein the sequence parameter set weighted prediction flag and the picture parameter set weighted prediction flag are flags for bidirectional prediction frames.
[0328] Example 63. A method according to Example 60, wherein the picture parameter set weighted prediction marking is constrained by the sequence parameter set weighted prediction marking.
[0329] Example 64. The method according to Example 57 further includes: obtaining at least a portion of a second picture included in a bitstream; determining, based on the bitstream, to disable weighted prediction for the second picture; based on disabling weighted prediction for the second picture, parsing the syntax element by reconstructing the value of a syntax element of the bitstream indicating a picture sequence count offset for the second reference picture to a value transmitted by a signal plus 1; and reconstructing the second picture using the reconstructed value of the syntax element.
[0330] Example 65. A method according to Example 64, wherein when weighted prediction is disabled, the value of the reference picture syntax element specifies the absolute difference between the picture order count value of the second picture and the previous reference picture entry for the second reference picture in the reference picture list as a value plus one.
[0331] Example 66. A method according to Example 57, wherein the reference picture is associated with a first set of weights and the picture is associated with a second set of weights different from the first set of weights; wherein the reference picture has a different size than the picture; and wherein the reference picture is signaled as a short-term reference picture having a picture order count least significant bit value equal to zero.
[0332] Example 67. The method of Example 57, wherein the picture is a non-instantaneous decoding refresh (non-IDR) picture.
[0333] Example 68. The method of Example 57, wherein at least a portion of the picture is a slice.
[0334] Example 69. The method according to Example 57 further includes: determining a layer identifier for a reference picture based on a bitstream, the layer identifier indicating a layer of a layer index used for inter-layer prediction; wherein, when the layer identifier is different from the current layer identifier, the zero-value picture sequence count offset is identified by inferring a zero value based on that the picture sequence count for the reference picture is not transmitted with a signal.
[0335] Example 70. An apparatus comprising: a memory configured to store video data; and a processor configured to process the video data according to any of Examples 57 to 69.
[0336] Example 71. An apparatus for encoding video data, the apparatus comprising: a memory; and a processor implemented in a circuit and configured to: identify at least a portion of a picture; select weighted prediction as enabled for the portion of the picture; identify a reference picture for the portion of the picture; generate a zero-valued picture order count offset for indicating a reference picture from a reference picture list; and generate a bitstream, the bitstream comprising the portion of the picture and the zero-valued picture order count offset as associated with the portion of the picture.
[0337] Example 72. A method for encoding video data, the method comprising: identifying at least a portion of a picture; determining that weighted prediction is selected to be enabled for the portion of the picture; identifying a reference picture for the portion of the picture; generating a zero-valued picture order count offset for indicating a reference picture from a reference picture list; and generating a bitstream, the bitstream comprising the portion of the picture and the zero-valued picture order count offset as associated with the portion of the picture.< / insert> < / insert> < / insert> < / insertend> < / insert> < / insertend> < / insert> < / insertend> < / insert> < / insertend> < / insert> < / insertend> < / insert> < / insertend> < / insert> < / insertend> < / insert> < / insertend> < / insert> < / insertend> < / insert>
Claims
1. A device for decoding video data, the device comprising: Memory; as well as A processor implemented in circuitry and configured to: obtaining at least a portion of a first picture included in a bitstream; determining, based on the bitstream, to enable weighted prediction for at least a portion of the first picture; determining, from the bitstream, a picture order count offset associated with a portion of the first picture, wherein a zero-valued picture order count offset indicates that the same reference picture is being used for prediction with different weights; enabling weighted prediction for at least a portion of the first picture based on the determining, reconstructing at least a portion of the first picture using at least a portion of the reference picture identified by the picture order count offset; obtaining at least a portion of a second picture included in the bitstream; determining, based on the bitstream, to disable weighted prediction for a portion of the second picture; parsing the second syntax element by generating a reconstructed value of a second syntax element of the bitstream indicating a second picture order count offset for a second reference picture as a signaled value plus 1 based on disabling the weighted prediction for a portion of the second picture; and A portion of the second picture is reconstructed using the reconstructed value of the second syntax element.
2. The device according to claim 1, wherein: The processor is further configured to: obtaining, from the bitstream, a plurality of portions of the first picture, the plurality of portions comprising at least a portion of the first picture reconstructed using the reference picture identified by the picture order count offset; identifying a plurality of corresponding picture order count offsets for the plurality of portions of the first picture, wherein the picture order count offset identified for the first portion is a corresponding picture order count offset associated with the reference picture among the plurality of corresponding picture order count offsets; as well as The picture is reconstructed using a plurality of reference pictures identified by the plurality of corresponding picture order count offsets.
3. The device according to claim 1, wherein: The processor is further configured to determine that weighted prediction can be enabled by parsing the bitstream to identify one or more weighted prediction flags for the picture.
4. The device according to claim 3, wherein: The one or more weighted prediction flags for the picture include a sequence parameter set weighted prediction flag and a picture parameter set weighted prediction flag.
5. The device according to claim 4, wherein: The sequence parameter set weighted prediction flag and the picture parameter set weighted prediction flag are flags for unidirectional prediction frames.
6. The device according to claim 4, wherein: The sequence parameter set weighted prediction flag and the picture parameter set weighted prediction flag are flags for bidirectional prediction frames.
7. The device according to claim 4, wherein: The picture parameter set weighted prediction flag is constrained by the sequence parameter set weighted prediction flag.
8. The device according to claim 1, wherein: When weighted prediction is disabled, the value of the reference picture syntax element specifies the absolute difference between the picture order count value of the second picture and the previous reference picture entry in the reference picture list for the second reference picture as the signaled value plus one.
9. The device according to claim 1, wherein: The reference picture is associated with a first set of weights, and the picture is associated with a second set of weights different from the first set of weights; wherein the reference picture has a different size than the picture; and Therein, the reference picture is signaled as a short-term reference picture having a picture order count least significant bit value equal to zero.
10. The device according to claim 1, wherein: The picture is a non-instantaneous decoding refresh (non-IDR) picture.
11. The device according to claim 1, wherein: The processor is configured to: determining, from the bitstream, a layer identifier for the reference picture, the layer identifier indicating a layer of a layer index used for inter-layer prediction; and Wherein, when the layer identifier is different from a current layer identifier, the zero-valued picture order count offset is identified by inferring a zero value based on that a picture order count for the reference picture is not signaled.
12. The device according to claim 1, wherein: The apparatus includes a mobile device having a camera coupled to the processor and the memory for capturing and storing the picture.
13. The apparatus of claim 1, further comprising a display.
14. A method for decoding video data, the method comprising: obtaining at least a portion of a first picture from a bitstream; determining, based on the bitstream, to enable weighted prediction for a portion of the first picture; determining, from the bitstream, a picture order count offset associated with a portion of the first picture, wherein a zero-valued picture order count offset indicates that the same reference picture is being used for prediction with different weights; reconstructing at least a portion of the first picture using at least a portion of the reference picture identified by the picture order count offset based on the determination to enable weighted prediction for a portion of the first picture; obtaining at least a portion of a second picture included in the bitstream; determining, based on the bitstream, to disable weighted prediction for a portion of the second picture; parsing the second syntax element by generating a reconstructed value of a second syntax element of the bitstream indicating a second picture order count offset for a second reference picture as a signaled value plus 1 based on disabling the weighted prediction for a portion of the second picture; and A portion of the second picture is reconstructed using the reconstructed value of the second syntax element.
15. The method according to claim 14, further comprising: obtaining, from the bitstream, a plurality of portions of the picture, the plurality of portions comprising at least a portion of the picture reconstructed using the reference picture identified by the zero-valued picture order count offset; identifying a plurality of corresponding picture order count offsets for the plurality of portions of the picture, wherein the zero-valued picture order count offset is a corresponding picture order count offset for the picture of the plurality of corresponding picture order count offsets; as well as The picture is reconstructed using a plurality of reference pictures identified by the plurality of corresponding picture order count offsets.
16. The method according to claim 14, further comprising: Enabling weighted prediction is determined by parsing the bitstream to identify one or more weighted prediction flags for the picture.
17. The method according to claim 16, wherein: The one or more weighted prediction flags for the picture include a sequence parameter set weighted prediction flag and a picture parameter set weighted prediction flag.
18. The method according to claim 17, wherein: The sequence parameter set weighted prediction flag and the picture parameter set weighted prediction flag are flags for unidirectional prediction frames.
19. The method according to claim 17, wherein: The sequence parameter set weighted prediction flag and the picture parameter set weighted prediction flag are flags for bidirectional prediction frames.
20. The method according to claim 17, wherein: The picture parameter set weighted prediction flag is constrained by the sequence parameter set weighted prediction flag.
21. The method according to claim 14, wherein: When weighted prediction is disabled, the value of the reference picture syntax element specifies an absolute difference between a picture order count value of the second picture and a previous reference picture entry in a reference picture list for the second reference picture as the signaled value plus one.
22. The method according to claim 14, wherein: The reference picture is associated with a first set of weights, and the picture is associated with a second set of weights different from the first set of weights; wherein the reference picture has a different size from the picture; and Therein, the reference picture is signaled as a short-term reference picture having a picture order count least significant bit value equal to zero.
23. The method according to claim 14, wherein: The picture is a non-instantaneous decoding refresh (non-IDR) picture.
24. The method according to claim 14, wherein: At least a portion of the picture is a slice.
25. The method of claim 14, further comprising: determining, from the bitstream, a layer identifier for the reference picture, the layer identifier indicating a layer of a layer index used for inter-layer prediction; as well as Wherein, when the layer identifier is different from a current layer identifier, the zero-valued picture order count offset is identified by inferring a zero value based on that a picture order count for the reference picture is not signaled.
26. An apparatus for encoding video data, the apparatus comprising: Memory; as well as A processor implemented in circuitry and configured to: identifying at least a portion of a first image; selecting weighted prediction to be enabled for a portion of the first picture; identifying a reference picture for a portion of the first picture; generating a picture order count offset indicating the reference picture from a reference picture list based on determining that weighted prediction is enabled for a portion of the first picture, wherein a zero-valued picture order count offset indicates that the same reference picture is being used for prediction with different weights; identifying at least a portion of a second image; selecting weighted prediction to be disabled for a portion of the second picture; identifying a second reference picture for a portion of the second picture; generating a picture order count offset for indicating the second reference picture from a reference picture list; generating a second syntax element indicating a picture order count offset minus one for the second reference picture based on disabling the weighted prediction for the second picture; and A bitstream is generated, the bitstream including a portion of the first picture and the picture order count offset associated with the portion of the first picture, a portion of the second picture, and the second syntax element.
27. A method for encoding video data, the method comprising: identifying at least a portion of a first image; selecting weighted prediction to be enabled for a portion of the first picture; identifying a reference picture for a portion of the first picture; generating a picture order count offset indicating the reference picture from a reference picture list based on determining that weighted prediction is enabled for a portion of the first picture, wherein a zero-valued picture order count offset indicates that the same reference picture is being used for prediction with different weights; identifying at least a portion of a second image; selecting weighted prediction to be disabled for a portion of the second picture; identifying a second reference picture for a portion of the second picture; generating a picture order count offset for indicating the second reference picture from a reference picture list; generating a second syntax element indicating a picture order count offset minus one for the second reference picture based on disabling the weighted prediction for the second picture; and A bitstream is generated, the bitstream including a portion of the first picture and the picture order count offset associated with the portion of the first picture, a portion of the second picture, and the second syntax element.