Same picture order count (POC) numbers for scalability support

By introducing weighted prediction technology and zero-value picture sequential counting offset into the video decoding system, the encoding and decoding process of video data is optimized, and the problem of inefficient video decoding in the prior art is solved, and more efficient video data transmission and decoding are achieved.

CN120281900APending Publication Date: 2025-07-08QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510519747.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2020-06-01
Filing Date
2020-06-02
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The existing video decoding technology causes excessive burden on communication networks and devices when processing video data of high fidelity, resolution and frame rate, and the existing video encoder has problems of wasted resources and inefficiency when signaling reference frames.

Method used

By introducing weighted prediction technology in the video decoding system, the reference picture is identified using zero-value picture sequential count offset, the encoding and decoding process of video data is optimized, signaling overhead is reduced, and the efficiency of encoding equipment and network is improved.

Benefits of technology

It effectively reduces the bit rate during the video decoding process, improves the encoding efficiency of video data and the performance of decoding equipment, reduces resource waste, and improves the transmission efficiency of video quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120281900A_ABST
    Figure CN120281900A_ABST
Patent Text Reader

Abstract

Techniques are described for video encoding and decoding for scalability support with the same picture or count number. One example relates to obtaining a portion (e.g., a slice, block, or other portion) of a picture, and determining whether weighted prediction is enabled for the portion of the picture. When weighted prediction is enabled, although different portions of a picture (e.g., different slices of the picture) may have different picture order count offset values, a zero-valued picture order count offset indicative of a reference picture from the reference picture may be used. The portion of the picture may then be reconstructed using the reference picture identified by the zero-valued picture order count offset. Additional embodiments may use weighted prediction flags and different offset values to determine reference pictures that support scalability or have a different size than the picture being reconstructed in weighted prediction.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of an application with an application date of June 2, 2020, an application number of 202080043476.7, and a title of "Same Picture Order Count (POC) Number for Scalability Support". Technical Field

[0002] This application is related to video coding. More specifically, this application relates to systems, methods, and computer-readable media for providing scalability support for a video coding system. Background Art

[0003] Many devices and systems allow video data to be processed and output for consumption. Digital video data includes a large amount of data to meet the needs of consumers and video providers. For example, consumers of video data expect the highest quality video with high fidelity, resolution, frame rate, etc. As a result, the large amount of video data required to meet these needs burdens the communication networks and devices that process and store the video data.

[0004] Various video coding techniques can be used to compress video data. Video coding is performed according to one or more video coding standards. For example, video coding standards include Versatile Video Coding (VVC), High Efficiency Video Coding (HEVC), Advanced Video Coding (AVC), Moving Picture Experts Group (MPEG) coding, and others. Video coding can utilize prediction methods (e.g., inter-frame prediction, intra-frame prediction, etc.), which utilize the redundancy present in video images or sequences. An important goal of video coding techniques is to compress video data into a form that uses a lower bit rate while avoiding or minimizing degradation of the video quality. As evolving video services become available, there is a need for coding techniques with better coding efficiency. Summary of the Invention

[0005] This document describes systems and methods for improved video processing. Digital video data includes a large amount of data to meet the needs of consumers and video providers, which places a burden on communication networks and devices for processing and storing video data. Some examples of video processing use video compression techniques that utilize prediction to efficiently encode and decode video data. For example, prediction can be determined as the difference between the pixel values in the block being encoded and a predicted block. The prediction error can also be referred to as a residual. The video encoder can also apply a transform to the prediction error (e.g., a discrete cosine transform (DCT) or other suitable transform) to generate transform coefficients. After the transform, the video encoder can quantize the transform coefficients. The quantized transform coefficients and motion vectors can be represented using syntax elements and, together with control information, form a decoded representation of the video sequence. In some instances, the video encoder can entropy code the syntax elements, thereby further reducing the number of bits required for their representation.

[0006] In some examples, the prediction can use a reference frame in the same layer as the frame being analyzed. Such prediction can include copies of the same frame with different sizes or resolutions. In some such examples, the reference frame can be identified within a reference list and referenced using a picture order count offset value. For weighted prediction, an example reference frame within the same layer as the frame being encoded or decoded can be referenced using a zero value picture order count offset. In some such examples, when weighted prediction is not used, the encoder or decoder can interpret the picture order count offset value as plus or minus 1 from the value signaled. Such operations improve the efficiency of the network and the encoding or decoding device by reducing signaling and providing efficient processing for operations to identify and use reference frames.

[0007] In an illustrative example, a device for decoding video data is provided. The device includes: a memory; and a processor implemented in circuitry. The processor is configured to: obtain at least a portion of a picture included in a bitstream. The processor is further configured to: determine, based on the bitstream, that weighted prediction is enabled for at least a portion of the picture; and based on determining that weighted prediction is enabled for at least a portion of the picture, identify a zero value picture order count offset for indicating a reference picture from a reference picture list. The processor is further configured to: use at least a portion of the reference picture identified by the zero value picture order count offset to reconstruct at least a portion of the picture.

[0008] In another example, a method for processing video data is provided. The method includes: obtaining at least a portion of a picture from a bitstream. The method further includes: determining, based on the bitstream, that weighted prediction is enabled for a portion of the picture. The method further includes: based on determining that weighted prediction is enabled for a portion of the picture, identifying a zero picture order count offset for indicating a reference picture from a reference picture list. The method further includes: using at least a portion of the reference picture identified by the zero picture order count offset to reconstruct at least a portion of the picture.

[0009] In another example, a computer-readable storage medium stores instructions that, when executed, cause one or more processors of a device for decoding video data to perform the following operations: obtaining at least a portion of a picture included in a bitstream; determining, based on the bitstream, that weighted prediction is enabled for at least a portion of the picture; based on determining that weighted prediction is enabled for at least a portion of the picture, identifying a zero picture order count offset for indicating a reference picture from a reference picture list; and using at least a portion of the reference picture identified by the zero picture order count offset to reconstruct at least a portion of the picture.

[0010] In another example, a device for decoding video data is provided. The device includes: a unit for obtaining at least a portion of a picture included in a bitstream; a unit for determining, based on the bitstream, that weighted prediction is enabled for at least a portion of the picture; a unit for, based on determining that weighted prediction is enabled for at least a portion of the picture, identifying a zero picture order count offset for indicating a reference picture from a reference picture list; and a unit for using at least a portion of the reference picture identified by the zero picture order count offset to reconstruct at least a portion of the picture.

[0011] In some cases, the methods, devices, and computer-readable storage media described above include: obtaining multiple portions of a picture from a bitstream, the multiple portions including at least a portion of the picture reconstructed using the reference picture identified by the zero picture order count offset; identifying multiple respective picture order count offsets for the multiple portions of the picture, wherein the zero picture order count offset is the respective picture order count offset associated with the reference picture among the multiple respective picture order count offsets; and reconstructing the picture using the multiple reference pictures identified by the multiple respective picture order count offsets.

[0012] In some cases, the methods, apparatuses, and computer-readable storage media described above include: determining weighted prediction enablement by parsing a bitstream to identify one or more weighted prediction flags for a picture. In some cases, the one or more weighted prediction flags for a portion of a picture include a sequence parameter set weighted prediction flag and a picture parameter set weighted prediction flag. In some cases, the sequence parameter set weighted prediction flag and the picture parameter set weighted prediction flag are flags for unidirectional prediction frames. In some cases, the sequence parameter set weighted prediction flag and the picture parameter set weighted prediction flag are flags for bidirectional prediction frames. In some cases, the picture parameter set weighted prediction flag is constrained by the sequence parameter set weighted prediction flag.

[0013] In some cases, the methods, apparatuses, and computer-readable storage media described above include: obtaining at least a portion of a second picture included in a bitstream; determining weighted prediction disablement for a portion of the second picture according to the bitstream; and determining that a second zero-valued picture order count offset is not allowed to indicate a second reference picture for a portion of the second picture from a reference picture list based on determining weighted prediction disablement for at least a portion of the second picture.

[0014] In some cases, the methods, apparatuses, and computer-readable storage media described above include: parsing a syntax element by reconstructing a value of the syntax element indicating a picture order count offset for the second reference picture in the bitstream to a signaled value plus 1 based on determining that the second zero-valued picture order count offset is not allowed and based on weighted prediction disablement for the second picture; and reconstructing the second picture using the reconstructed value of the syntax element.

[0015] In some cases, when weighted prediction is disabled, the value of a reference picture syntax element specifies the absolute difference between the picture order count value of the second picture and a previous reference picture entry for the second reference picture in a reference picture list as a value plus one.

[0016] In some cases, a reference picture is associated with a first weight set, and a picture is associated with a second weight set different from the first weight set. In some cases, the reference picture has a different size from the picture. In some cases, the reference picture is signaled as a short-term reference picture having a picture order count least significant bit value equal to zero. In some cases, the picture is a non-instantaneous decoding refresh (non-IDR) picture. In some cases, at least a portion of the picture is a slice. The slice may include multiple blocks of the picture. In some cases, at least a portion of the picture is a block of the picture (e.g., a coding tree unit (CTU), a macroblock, a coding unit or block, a prediction unit or block, or other types of blocks of the picture).

[0017] In some cases, the methods, apparatuses, and computer-readable storage media described above include: determining a layer identifier for a reference picture according to a bitstream, the layer identifier indicating a layer of a layer index used for inter-layer prediction. In some examples, when the layer identifier is different from the current layer identifier, a zero picture order count offset is identified by inferring a zero value according to that the picture order count for the reference picture is not signaled.

[0018] In another illustrative example, there is provided an apparatus for encoding video data. The apparatus includes: a memory; and a processor implemented in circuitry and configured to: identify at least a portion of a picture. The processor is further configured to: select weighted prediction to be enabled for a portion of the picture. The processor is further configured to: identify a reference picture for a portion of the picture; and generate a zero picture order count offset for indicating a reference picture from a reference picture list. The processor is further configured to: generate a bitstream, the bitstream including a portion of the picture and the zero picture order count offset associated with the portion of the picture.

[0019] In another example, there is provided a method for encoding video data. The method includes: identifying at least a portion of a picture. The method includes: determining that weighted prediction is selected to be enabled for a portion of the picture. The method includes: identifying a reference picture for a portion of the picture; and generating a zero picture order count offset for indicating a reference picture from a reference picture list. The method further includes: generating a bitstream, where the bitstream includes a portion of the picture and the zero picture order count offset associated with the portion of the picture.

[0020] In another illustrative example, a computer-readable storage medium stores instructions that, when executed, cause one or more processors of a device for encoding video data to perform the following operations: identify at least a portion of a picture; select weighted prediction to be enabled for a portion of the picture; identify a reference picture for a portion of the picture; generate a zero picture order count offset for indicating a reference picture from a reference picture list; and generate a bitstream, the bitstream including a portion of the picture and the zero picture order count offset associated with the portion of the picture.

[0021] In another illustrative example, an apparatus for encoding video data is provided. The apparatus includes: a unit for identifying at least a portion of a picture; a unit for selecting weighted prediction to be enabled for a portion of a picture; a unit for identifying a reference picture for a portion of a picture; a unit for generating a zero picture order count offset for indicating a reference picture from a reference picture list; and a unit for generating a bitstream, the bitstream including a portion of a picture and a zero picture order count offset associated with the portion of the picture.

[0022] In another example, a method for encoding video data is provided. The method includes: identifying at least a portion of a picture. The method includes: determining that weighted prediction is selected to be enabled for a portion of a picture. The method includes: identifying a reference picture for a portion of a picture; and generating a zero picture order count offset for indicating a reference picture from a reference picture list. The method further includes: generating a bitstream, where the bitstream includes a portion of a picture and a zero picture order count offset associated with the portion of the picture.

[0023] In some aspects, the apparatus for decoding video data and / or the apparatus for encoding video data includes a camera, a mobile device (e.g., a mobile phone or a so-called "smartphone" or other mobile device), a wearable device, an extended reality device (e.g., a virtual reality (VR) device, an augmented reality (AR) device, or a mixed reality (MR) device), a personal computer, a laptop computer, a server computer, or other device. In some aspects, the apparatus for decoding video data and / or the apparatus for encoding video data includes a camera or multiple cameras for capturing one or more images. In some aspects, the apparatus for decoding video data and / or the apparatus for encoding video data includes a display for displaying one or more images, notifications, and / or other displayable data.

[0024] Aspects related to any one of the methods, apparatuses, and computer-readable media described above may be used alone or in any suitable combination.

[0025] This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used alone to determine the scope of the claimed subject matter. The subject matter should be understood by reference to the appropriate portions of the entire specification of this patent, any or all of the drawings, and each claim.

[0026] The foregoing, together with other features and embodiments, will become more apparent after referring to the following specification, claims, and drawings. Description of the Drawings

[0027] Exemplary examples of the present application are described in detail below with reference to the following drawings:

[0028] Figure 1 is a block diagram showing examples of an encoding device and a decoding device according to some examples.

[0029] Figure 2 is a schematic diagram showing an example implementation of a filter unit for performing the techniques of the present disclosure.

[0030] Figure 3 is a schematic diagram showing examples of multiple access units (AUs) having different picture order counts (POCs) according to some examples.

[0031] Figure 4 is a flowchart showing an example method according to various examples described herein.

[0032] Figure 5 is a flowchart showing an example method according to various examples described herein.

[0033] Figure 6 is a block diagram showing an example video encoding device according to some examples.

[0034] Figure 7 is a block diagram showing an example video decoding device according to some examples. Detailed Description

[0035] Certain aspects and examples of the present disclosure are provided below. As will be apparent to those skilled in the art, some of these aspects and examples may be applied independently, and some of them may be applied in combination. In the following description, for the purpose of explanation, specific details are set forth to provide a thorough understanding of the examples of the present application. However, it will be apparent that the various examples may be practiced without these specific details. The drawings and the description are not intended to be restrictive.

[0036] The following description only provides exemplary examples and is not intended to limit the scope, applicability, or configuration of the present disclosure. Rather, the following description of the exemplary examples will provide those skilled in the art with a description that enables the implementation of the exemplary examples. It should be understood that various changes may be made to the functions and arrangements of the elements without departing from the spirit and scope of the present application as set forth in the appended claims.

[0037] Video decoding devices implement video compression techniques to effectively encode and decode video data. Video compression techniques can include applying different prediction modes, including spatial prediction (e.g., intra-frame prediction or intra-prediction), temporal prediction (e.g., inter-frame prediction or inter-prediction), inter-layer prediction (across different layers of video data), and / or other prediction techniques for reducing or removing redundancy inherent in a video sequence. A video encoder can divide each picture of an original video sequence into a plurality of rectangular regions, which are referred to as video blocks or coding units (described in more detail below). These video blocks can be encoded using a specific prediction mode.

[0038] Video blocks can be divided into one or more groups of smaller blocks in one or more ways. Blocks can include coding tree blocks, prediction blocks, transform blocks, and / or other suitable blocks. Unless otherwise specified, a reference to "block" generally can refer to such video blocks (e.g., coding tree blocks, coding blocks, prediction blocks, transform blocks, or other suitable blocks or sub-blocks, as would be understood by a person of ordinary skill in the art). Further, each of these blocks can also be interchangeably referred to herein as a "unit" (e.g., coding tree unit (CTU), coding unit, prediction unit (PU), transform unit (TU), etc.). In some cases, a unit can indicate a coding logic unit encoded in a bitstream, while a block can indicate a portion of a video frame buffer that the process is directed to.

[0039] For inter-frame prediction modes, a video encoder can search for a block similar to the block being encoded in a frame (or picture) located at another temporal position, which is referred to as a reference frame or reference picture. The video encoder can limit the search to a certain spatial displacement from the block to be encoded. In some systems, the best match is located using a two-dimensional (2D) motion vector that includes a horizontal displacement component and a vertical displacement component. For intra-frame prediction modes, a video encoder can use spatial prediction techniques to form a prediction block based on data from previously encoded neighboring blocks within the same picture.

[0040] A video encoder can determine a prediction error. For example, the prediction can be determined as the difference between the pixel values in the block being encoded and the pixel values in a prediction block. The prediction error can also be referred to as a residual. The video encoder can also apply a transform to the prediction error (e.g., a discrete cosine transform (DCT) or other suitable transform) to generate transform coefficients. After the transform, the video encoder can quantize the transform coefficients. In some examples, the quantized transform coefficients and motion vectors can be represented using syntax elements and, together with control information, form a decoded representation of the video sequence. In some instances, the video encoder can entropy code the syntax elements, thereby further reducing the number of bits used for their representation.

[0041] A video decoder can use the syntax elements and control information discussed above to construct prediction data (e.g., a prediction block) for decoding the current frame. For example, the video decoder can add the prediction block and the compressed prediction error. The video decoder can determine the compressed prediction error by weighting the transform basis functions using the quantization coefficients. The difference between the reconstructed frame and the original frame is referred to as the reconstruction error.

[0042] As mentioned above, reference pictures can be used during inter prediction. To identify a reference picture, a picture order count (POC) offset value or “delta POC” value can be used. The POC value identifies the selected picture based on the difference (e.g., offset or delta) between the picture order count value for the selected picture and the picture order count value for a previous picture or the original picture. Some reference pictures can have a delta POC value equal to zero, such as when multiple reference pictures have the same POC value (e.g., the same picture is used as a reference multiple times). For example, a zero delta POC value can be used with weighted prediction. However, in some examples, when weighted prediction is disabled, the zero delta POC value may not be used. Signaling a non - zero delta POC value uses additional bits, which can be wasteful when the delta POC value is typically the same non - zero value.

[0043] The present document describes systems and techniques for reducing the bitrate used to signal encoded video data by determining when to enable or disable a related mode (e.g., weighted prediction) and interpreting the signaled delta POC value based on that determination. In some examples, a flag is used to determine the current mode, such as a weighted prediction flag that indicates whether weighted prediction is enabled. In such examples, when weighted prediction is enabled, a signaled delta POC value of zero can be used as the actual delta POC for picture reconstruction. When weighted prediction is disabled, the signaled delta POC value of zero can be modified to determine the actual (non-zero) delta POC value. As described above, such an implementation can improve the operation of a system by reducing the bitrate used for picture signaling without degrading picture quality.

[0044] The techniques described herein can be applied to any video codec in existing video codecs (e.g., High Efficiency Video Coding (HEVC), Advanced Video Coding (AVC), or other suitable existing video codecs), and / or can be efficient coding tools for any video coding standard being developed and / or future video coding standards, such as, for example, Versatile Video Coding (VVC), Joint Exploration Model (JEM), and / or other video coding standards being developed or to be developed. Although the examples described herein provide examples using the JEM model, VVC, HEVC standards, and / or extensions thereof, the techniques and systems described herein can also be applicable to other coding standards. Thus, although the techniques and systems described herein may be described with reference to a particular video coding standard, those of ordinary skill in the art will understand that the description should not be construed as being applicable only to that particular standard.

[0045] Figure 1FIG. 0 is a block diagram illustrating an example of a system 100 that includes an encoding device 104 and a decoding device 112. The encoding device 104 may be part of a source device, and the decoding device 112 may be part of a receiving device. The source device and / or the receiving device may include an electronic device such as a mobile or stationary telephone handset (e.g., a smart phone, a cellular phone, etc.), a desktop computer, a laptop or notebook computer, a tablet computer, a set-top box, a television, a camera, a display device, a digital media player, a video game console, a video streaming device, an Internet Protocol (IP) camera, or any other suitable electronic device. In some examples, the source device and the receiving device may include one or more wireless transceivers for wireless communication. The decoding techniques described herein are applicable to video decoding in a variety of multimedia applications, including streaming video transmission (e.g., over the Internet), television broadcast or transmission, encoding of digital video for storage on a data storage medium, decoding of digital video stored on a data storage medium, or other applications. In some examples, the system 100 may support unidirectional or bidirectional video transmission to support applications such as video conferencing, video streaming, video playback, video broadcast, gaming, and / or video telephony.

[0046] The encoding device 104 (or encoder) may be used to encode video data using a video coding standard or protocol to generate an encoded video bitstream. Examples of video coding standards include ITU-T H.261, ISO / IEC MPEG-1 Visual, ITU-T H.262, or ISO / IEC MPEG-2 Visual, ITU-T H.263, ISO / IEC MPEG-4 Visual, ITU-T H.264 (also known as ISO / IEC MPEG-4 AVC) (including its scalable video coding (SVC) and multi-view video coding (MVC) extensions), and HEVC or ITU-T H.265. There are various extensions to HEVC for handling multi-layer video coding, including range and screen content coding extensions, 3D video coding (3D-HEVC) and multi-view extensions (MV-HEVC) and scalable extensions (SHVC). A new video coding standard being developed by the Joint Video Exploration Team (JVET) is called Versatile Video Coding (VVC).

[0047] Reference Figure 1, the video source 102 can provide video data to the encoding device 104. The video source 102 can be part of the source device or can be part of a device other than the source device. The video source 102 can include a video capture device (e.g., a camera, a camera phone, a video phone, etc.), a video archive containing stored video, a video server or content provider for providing video data, a video feed interface for receiving video from a video server or content provider, a computer graphics system for generating computer graphics video data, a combination of such sources, or any other suitable video source.

[0048] The video data from the video source 102 can include one or more input pictures. A picture can also be referred to as a "frame". A picture or frame is a still image, which in some cases is part of a video. In some examples, the data from the video source 102 can be a still image that is not part of a video. In HEVC, VVC, and other video coding specifications, a video sequence can include a series of pictures. A picture can include three sample arrays, denoted as S L 、S Cb and S Cr . S L is a two-dimensional array of luminance samples, S Cb is a two-dimensional array of Cb chrominance samples, and S Cr is a two-dimensional array of Cr chrominance samples. Chrominance samples can also be referred to as "chroma" samples herein. A pixel can refer to a position within a picture that includes a luminance component, a Cb component, and a Cr component. In other instances, a picture can be monochromatic and can include only an array of luminance samples.

[0049] The encoder engine 106 (or encoder) of the encoding device 104 encodes the video data to generate an encoded video bitstream. In some examples, the encoded video bitstream (or "video bitstream" or "bitstream") is a series of one or more decoded video sequences. A decoded video sequence (CVS) includes a series of access units (AUs), the series of access units starting from an access unit with a random access point picture in the base layer and having certain attributes until the next access unit with a random access point picture in the base layer and having certain attributes and not including that next access unit. An AU includes one or more decoded pictures and control information corresponding to the decoded pictures sharing the same output time. The decoded slices of a picture are encapsulated as data units at the bitstream level, and this data unit is called a network abstraction layer (NAL) unit. For example, an HEVC video bitstream can include one or more CVSs, and each CVS includes NAL units. Each NAL unit in the NAL units has a NAL unit header.

[0050] The encoder engine 106 generates a decoded representation of a picture by dividing each picture into multiple slices. The slices are independent of other slices such that the information in a slice is decoded without relying on data from other slices within the same picture. The slices are then divided into decoded tree blocks (CTBs) of luma samples and chroma samples. A CTB of luma samples and one or more CTBs of chroma samples, together with the syntax for the samples, are referred to as a coding tree unit (CTU). The CTU is the basic processing unit for HEVC encoding. A CTU can be split into multiple coding units (CUs) of different sizes. A CU contains an array of luma and chroma samples called a coding block (CB).

[0051] The luma and chroma CBs can be further split into prediction blocks (PBs). A PB is a block of samples of a luma component or a chroma component that uses the same motion parameters for inter - frame prediction or intra - block copy prediction (when available or enabled for use). A luma PB and one or more chroma PBs, together with the associated syntax, form a prediction unit (PU). For inter - frame prediction, a set of motion parameters (e.g., one or more motion vectors, reference indices, etc.) is signaled in the bitstream for each PU and is used for inter - frame prediction of the luma PB and one or more chroma PBs. The motion parameters can also be referred to as motion information. A CB can also be split into one or more transform blocks (TBs). A TB represents a square block of samples of a color component to which the same two - dimensional transform is applied for decoding the prediction residual signal. A transform unit (TU) represents the TBs of luma and chroma samples and the corresponding syntax elements.

[0052] The size of a CU corresponds to the size of the coding mode and can be square in shape. For example, the size of a CU can be 8x8 samples, 16x16 samples, 32x32 samples, 64x64 samples, or any other suitable size up to the size of the corresponding CTU. The phrase “N x N” is used herein to refer to the pixel dimensions of a video block in the vertical and horizontal dimensions (e.g., 8 pixels x 8 pixels). The pixels in a block can be arranged in rows and columns. In some examples, the block may not have the same number of pixels in the horizontal direction as in the vertical direction. The syntax data associated with a CU can describe, for example, splitting the CU into one or more PUs. The splitting pattern can differ between whether the CU is encoded in an intra - prediction mode or an inter - prediction mode. A PU can be split into a non - square shape.

[0053] According to the HEVC standard, the transform can be performed using TUs, as described above. For different CUs, the TUs can be different. The TUs can be sized based on the size of the PUs within a given CU. The TUs can be the same size as the PUs or smaller than the PUs. In some examples, the residual samples corresponding to a CU can be subdivided into smaller units using a quadtree structure called a Residual Quadtree (RQT). The leaf nodes of the RQT can correspond to the TUs. The pixel differences associated with the TUs can be transformed to produce transform coefficients. The transform coefficients can then be quantized by the encoder engine 106.

[0054] Once the picture of the video data is segmented into CUs, the encoder engine 106 uses a prediction mode to predict each PU. The prediction unit or prediction block is then subtracted from the original video data to obtain a residual (described below). For each CU, the prediction mode can be signaled in the bitstream using syntax data. The prediction mode can include intra prediction (or picture-in prediction) or inter prediction (or picture - to - picture prediction). The decision as to whether to use inter prediction or intra prediction to decode a picture region can be made, for example, at the CU level.

[0055] Intra prediction exploits the correlation between spatially adjacent samples within a picture. For example, using intra prediction, each PU is predicted from adjacent picture data in the same picture using, for example: DC prediction to find the average value for the PU, planar prediction to fit a planar surface to the PU, directional prediction to infer from adjacent data, or any other suitable type of prediction.

[0056] Inter-picture prediction uses the temporal correlation between pictures to derive motion-compensated predictions for image sample blocks. Using a translational motion model, the position of a block in a previously decoded picture (reference picture) is represented by a motion vector (Δx, Δy), where Δx specifies the horizontal displacement of the reference block relative to the position of the current block, and Δy specifies the vertical displacement of the reference block relative to the position of the current block. In some cases, the motion vector (Δx, Δy) can be at integer sample precision (also referred to as integer precision), in which case the motion vector points to the integer pixel grid (or integer pixel sampling grid) of the reference frame. In some cases, the motion vector (Δx, Δy) can have fractional sample precision (also referred to as fractional pixel precision or non-integer precision) to more accurately capture the motion of the underlying object, without being limited to the integer pixel grid of the reference frame. The precision of the motion vector can be expressed by the quantization level of the motion vector. For example, the quantization level can be integer precision (e.g., 1 pixel) or fractional pixel precision (e.g., 1 / 4 pixel, 1 / 2 pixel, or other sub-pixel values). When the corresponding motion vector has fractional sample precision, interpolation is applied to the reference picture to derive the prediction signal. For example, samples available at integer positions can be filtered (e.g., using one or more interpolation filters) to estimate the values at fractional positions. The previously decoded reference picture is indicated by a reference index (refIdx) for the reference picture list. The motion vector and the reference index can be referred to as motion parameters. Two types of inter-picture prediction can be performed, including uni-directional prediction and bi-directional prediction.

[0057] In the case of performing inter-frame prediction using bi-directional prediction, two sets of motion parameters (Δx0, Δy0, refIdx0 and Δx1, Δy1, refIdx1) are used to generate two motion-compensated predictions (from the same reference picture or possibly from different reference pictures). For example, in the case of bi-directional prediction, each prediction block uses two motion-compensated prediction signals, and B prediction units are generated. Then, the two motion-compensated predictions are combined to obtain the final motion-compensated prediction. For example, the two motion-compensated predictions can be combined by averaging. In another example, weighted prediction can be used, in which case different weights can be applied to each motion-compensated prediction. The reference pictures that can be used in bi-directional prediction are stored in two separate lists, denoted as list 0 and list 1. The motion parameters can be derived at the encoder using a motion estimation process.

[0058] In the case of performing inter-frame prediction using uni-directional prediction, one set of motion parameters (Δx0, Δy0, refIdx0) is used to generate a motion-compensated prediction from the reference picture. For example, in the case of uni-directional prediction, each prediction block uses at most one motion-compensated prediction signal, and P prediction units are generated.

[0059] The PU may include data related to the prediction process (e.g., motion parameters or other suitable data). For example, when the PU is encoded using intra prediction, the PU may include data for describing the intra prediction mode for the PU. As another example, when the PU is encoded using inter prediction, the PU may include data for defining the motion vector for the PU. The data for defining the motion vector for the PU may describe, for example, the horizontal component (Δx) of the motion vector, the vertical component (Δy) of the motion vector, the resolution for the motion vector (e.g., integer precision, quarter-pixel precision, or eighth-pixel precision), the reference picture to which the motion vector points, the reference index, the reference picture list for the motion vector (e.g., list 0, list 1, or list C), or any combination thereof.

[0060] The encoding device 104 may then perform transformation and quantization. For example, after prediction, the encoder engine 106 may calculate the residual value corresponding to the PU. The residual value may include the pixel difference between the current pixel block (PU) being decoded and the predicted block used to predict the current block (e.g., the predicted version of the current block). For example, after generating the predicted block (e.g., using inter prediction or intra prediction), the encoder engine 106 may generate a residual block by subtracting the predicted block generated by the prediction unit from the current block. The residual block includes a set of pixel differences that quantize the difference between the pixel values in the current block and the pixel values in the predicted block. In some examples, the residual block may be represented in a two-dimensional block format (e.g., a two-dimensional matrix or array of pixel values). In such examples, the residual block is a two-dimensional representation of pixel values.

[0061] Any residual data that may remain after performing prediction is transformed using a block transform, which may be based on a discrete cosine transform, a discrete sine transform, an integer transform, a wavelet transform, other suitable transform functions, or any combination thereof. In some cases, one or more block transforms (e.g., of size 32x32, 16x16, 8x8, 4x4, or other suitable sizes) may be applied to the residual data in each CU. In some examples, TUs may be used for the transformation and quantization processes implemented by the encoder engine 106. A given CU having one or more PUs may also include one or more TUs. As described in further detail below, the residual values may be transformed into transform coefficients using a block transform, and then may be quantized and scanned using TUs to produce serialized transform coefficients for entropy coding.

[0062] In some examples, after performing intra prediction or inter prediction decoding on a PU of a CU, the encoder engine 106 may compute residual data for a TU of the CU. The PU may include pixel data in the spatial domain (or pixel domain). After applying a block transform, the TU may include coefficients in the transform domain. As previously mentioned, the residual data may correspond to the pixel differences between the pixels in the unencoded picture and the predicted values corresponding to the PU. The encoder engine 106 may form a TU including the residual data for the CU, and then may transform the TU to produce transform coefficients for the CU.

[0063] The encoder engine 106 may perform quantization of the transform coefficients. Quantization provides further compression by quantizing the transform coefficients to reduce the amount of data used to represent the coefficients. For example, quantization may reduce the bit depth associated with some or all of the coefficients. In one example, a coefficient with an n-bit value may be rounded down to an m-bit value during quantization, where n is greater than m.

[0064] Once quantization is performed, the encoded video bitstream includes the quantized transform coefficients, prediction information (e.g., prediction mode, motion vectors, block vectors, etc.), segmentation information, and any other appropriate data, such as other syntax data. Different elements of the encoded video bitstream may then be entropy encoded by the encoder engine 106. In some examples, the encoder engine 106 may use a predefined scan order to scan the quantized transform coefficients to produce a serialized vector that can be entropy encoded. In some examples, the encoder engine 106 may perform adaptive scanning. After scanning the quantized transform coefficients to form a vector (e.g., a one-dimensional vector), the encoder engine 106 may entropy encode the vector. For example, the encoder engine 106 may use context-adaptive variable-length decoding, context-adaptive binary arithmetic decoding, syntax-based context-adaptive binary arithmetic decoding, probability interval segmentation entropy decoding, or another appropriate entropy coding technique.

[0065] The output 110 of the encoding device 104 may send NAL units constituting the encoded video bitstream data over the communication link 120 to the decoding device 112 of the receiving device. The input 114 of the decoding device 112 may receive the NAL units. The communication link 120 may include a channel provided by a wireless network, a wired network, or a combination of a wired network and a wireless network. The wireless network may include any wireless interface or combination of wireless interfaces, and may include any appropriate wireless network (e.g., the Internet or other wide area network, packet-based network, WiFi TM , radio frequency (RF), UWB, WiFi Direct, cellular, Long Term Evolution (LTE), WiMax TMetc.). The wired network may include any wired interface (e.g., optical fiber, Ethernet, power line Ethernet, Ethernet over coaxial cable, digital subscriber line (DSL), etc.). The wired and / or wireless network may be implemented using various devices such as base stations, routers, access points, bridges, gateways, switches, etc. The encoded video bitstream data may be modulated according to a communication standard such as a wireless communication protocol and sent to a receiving device.

[0066] In some examples, the encoding device 104 may store the encoded video bitstream data in the storage space 108. The output 110 may retrieve the encoded video bitstream data from the encoder engine 106 or from the storage space 108. The storage space 108 may include any one of various distributed or local access data storage media. For example, the storage space 108 may include a hard disk drive, a storage disk, a flash memory, a volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded video data. The storage space 108 may also include a decoded picture buffer (DPB) for storing reference pictures used in inter prediction.

[0067] The encoder engine 106 and the decoder engine 116 (described in more detail below) may be configured to operate according to VVC. According to VVC, a video decoder (such as the encoder engine 106 and / or the decoder engine 116) divides a picture into multiple coding tree units (CTUs). The video decoder may divide a CTU according to a tree structure such as a quadtree-binary tree (QTBT) structure or a multi-type tree (MTT) structure. The QTBT structure eliminates the concept of multiple partition types, such as the distinction between CUs, PUs, and TUs in HEVC. The QTBT structure includes two levels, including a first level that is divided according to quadtree partitioning and a second level that is divided according to binary tree partitioning. The root node of the QTBT structure corresponds to the CTU. The leaf nodes of the binary tree correspond to coding units (CUs).

[0068] In the MTT partitioning structure, a block may be partitioned using quadtree partitioning, binary tree partitioning, and one or more types of ternary tree partitioning. Ternary tree partitioning is a partitioning in which a block is divided into three sub-blocks. In some examples, the ternary tree partitioning divides a block into three sub-blocks without dividing the original block through the center. The partitioning types in the MTT (e.g., quadtree, binary tree, and ternary tree) may be symmetric or asymmetric.

[0069] In some examples, a video decoder may use a single QTBT or MTT structure to represent each of the luminance and chrominance components, while in other examples, the video decoder may use two or more QTBT or MTT structures, such as one QTBT or MTT structure for the luminance component and another QTBT or MTT structure for the two chrominance components (or two QTBTs and / or MTTs for each respective chrominance component).

[0070] The video decoder may be configured to use quadtree partitioning, QTBT partitioning, MTT partitioning, or other partitioning structures according to HEVC. For illustrative purposes, the description herein may refer to QTBT partitioning. However, it should be understood that the techniques of the present disclosure may also be applied to video decoders configured to use quadtree partitioning or also use other types of partitioning.

[0071] In VVC, a picture may be partitioned into slices, tiles, and bricks. Generally, a brick may be a rectangular region of CTU rows within a particular tile in the picture. A tile may be a rectangular region of CTUs within a particular tile column and a particular tile row in the picture. A tile column is a rectangular region of CTUs that has a height equal to the height of the picture and a width specified by a syntax element in the picture parameter set. A tile row is a rectangular region of CTUs that has a height specified by a syntax element in the picture parameter set and a width equal to the width of the picture. In some cases, a tile may be partitioned into multiple bricks, and each brick may include one or more CTU rows within the tile. A tile that is not partitioned into multiple bricks is also referred to as a brick. However, a brick that is a proper subset of a tile is not referred to as a tile. A slice may be an integer number of bricks of a picture that are uniquely contained within a single NAL unit. In some cases, a slice may include a number of complete tiles or only a contiguous sequence of complete bricks of one tile.

[0072] Input 114 of decoding device 112 receives the encoded video bitstream data and can provide the video bitstream data to decoder engine 116 or to storage space 118 for later use by decoder engine 116. Input 114 of decoding device 112 receives the encoded video bitstream data and can provide the video bitstream data to decoder engine 116 or to storage space 118 for later use by decoder engine 116. For example, storage space 118 may include a DPB for storing reference pictures used during inter prediction. A receiving device including decoding device 112 may receive the encoded video data to be decoded via storage space 108. The encoded video data may be modulated according to a communication standard such as a wireless communication protocol and sent to the receiving device. Decoder engine 116 can decode the encoded video bitstream data by entropy decoding (e.g., using an entropy decoder) and extracting elements of one or more decoded video sequences that make up the encoded video data. Decoder engine 116 can then rescale the encoded video bitstream data and perform an inverse transform on it. The residual data is then passed to the prediction stage of decoder engine 116. Decoder engine 116 then predicts pixel blocks (e.g., PUs). In some examples, the prediction is added to the output of the inverse transform (residual data).

[0073] Decoding device 112 can output the decoded video to video destination device 122, which may include a display or other output device for displaying the decoded video data to a consumer of the content. In some aspects, video destination device 122 may be part of the receiving device that includes decoding device 112. In some aspects, video destination device 122 may be part of a separate device that is different from the receiving device.

[0074] Examples of the specific details of encoding device 104 are described below with reference to Figure 6 Examples of the specific details of decoding device 112 are described below with reference to Figure 7 to describe.

[0075] As previously described, the HEVC bitstream includes a set of NAL units, including VCL NAL units and non-VCL NAL units. The VCL NAL units include the decoded picture data that forms the decoded video bitstream. For example, the bit sequence that forms the decoded video bitstream is present in the VCL NAL units. Among other information, the non-VCL NAL units can contain parameter sets having high-level information related to the encoded video bitstream. For example, the parameter sets can include a Video Parameter Set (VPS), a Sequence Parameter Set (SPS), and a Picture Parameter Set (PPS). Examples of the objectives of the parameter sets include bitrate efficiency, error resilience, and providing a system layer interface. Each slice references a single valid PPS, SPS, and VPS to access the information that the decoding device 112 can use to decode the slice. An identifier (ID) can be decoded for each parameter set (including VPSID, SPSID, and PPSID). The SPS includes the SPS ID and the VPS ID. The PPS includes the PPS ID and the SPS ID. Each slice header includes the PPS ID. Using the ID, the active parameter set can be identified for a given slice.

[0076] The PPS includes information that applies to all slices in a given picture. As such, all slices in a picture reference the same PPS. Slices in different pictures can also reference the same PPS. The SPS includes information that applies to all pictures in the same decoded video sequence (CVS) or bitstream. The information in the SPS may not change from picture to picture within the decoded video sequence. Pictures in the decoded video sequence can use the same SPS. The VPS includes information that applies to all layers within the decoded video sequence or bitstream. The VPS includes a syntax structure that has syntax elements that apply to the entire decoded video sequence. In some examples, the VPS, SPS, or PPS can be sent in-band with the encoded bitstream. In some examples, the VPS, SPS, or PPS can be sent out-of-band in a separate transmission from the NAL units that contain the decoded video data.

[0077] The video bitstream can also include Supplemental Enhancement Information (SEI) messages. For example, SEI NAL units can be part of the video bitstream. In some cases, the SEI messages may contain information that is not used by the decoding process. For example, the information in the SEI messages may not be necessary for the decoder to decode the video pictures of the bitstream, but the decoder can use the information to improve the display or processing of the pictures (e.g., the decoded output).

[0078] As described herein, for each block, a set of motion information (also referred to herein as motion parameters) may be available. The set of motion information contains motion information for both the forward prediction direction and the backward prediction direction. The forward prediction direction and the backward prediction direction are the two prediction directions of the bi - directional prediction mode, in which case the terms "forward" and "backward" do not necessarily have a geometric meaning. Instead, "forward" and "backward" correspond to Reference Picture List 0 (RefPicList0 or L0) and Reference Picture List 1 (RefPicList1 or L1) of the current picture. In some examples, when only one reference picture list is available for a picture or slice, only RefPicList0 is available, and the motion information for each block of the slice is always forward.

[0079] In some cases, the motion vector together with its reference index is used in the decoding process (e.g., motion compensation). Such a motion vector with an associated reference index is represented as a unidirectional prediction set of motion information. For each prediction direction, the motion information may contain a reference index and a motion vector. In some cases, for simplicity, the motion vector itself may be referred to in a way that assumes it has an associated reference index. The reference index is used to identify the reference picture in the current reference picture list (RefPicList0 or RefPicList1). The motion vector has a horizontal component and a vertical component, which provide an offset from the coordinate position in the current picture to the coordinates in the reference picture identified by the reference index. For example, the reference index may indicate the specific reference picture that should be used for a block in the current picture, and the motion vector may indicate where the block that best matches the current block (the best - matching block) is located in the reference picture.

[0080] Picture Order Count (POC) may be used in video coding standards to identify the display order of pictures. Although there are cases where two pictures within a coded video sequence may have the same POC value, this generally does not occur within a coded video sequence. When there are multiple coded video sequences in a bitstream, pictures with the same POC value may be closer to each other in terms of decoding order. The POC value of a picture can be used for reference picture list construction, such as the derivation of the reference picture set in HEVC, and motion vector scaling.

[0081] For motion prediction in HEVC, there are two inter - prediction modes for Prediction Units (PUs), including the Merge mode and the Advanced Motion Vector Prediction (AMVP) mode. Skip is considered a special case of Merge. In the AMVP mode or the Merge mode, a motion vector (MV) candidate list is maintained for multiple motion predictors. The motion vector of the current PU and the reference index in the Merge mode are generated by obtaining a candidate from the reference list.

[0082] In an example where an MV candidate list is used for motion prediction of a block, the reference list may be constructed by an encoding device and a decoding device separately. For example, the candidate list may be generated by the encoding device when encoding a block (e.g., a CTU, a CU, or other blocks of a picture), and may be generated by the decoding device when decoding a block. Information related to motion information candidates in the candidate list may be signaled between the encoding device and the decoding device. For example, in the merge mode, an index value of the stored motion information candidate may be signaled from the encoding device to the decoding device (e.g., in a syntax structure such as PPS, SPS, VPS, slice header, SEI message sent in the video bitstream or sent separately from the video bitstream, and / or other signaling). The decoding device may construct the candidate list and use the signaled reference or index to obtain one or more motion information candidates from the constructed candidate list for motion compensation prediction. For example, the decoding device 112 may construct an MV candidate list and use the motion vector from the index position to perform motion prediction on a block. In the case of the AMVP mode, in addition to the reference or index, a difference or residual value may also be signaled as an increment. For example, for the AMVP mode, the decoding device may construct one or more MV candidate lists and apply the increment value to one or more motion information candidates obtained by using the signaled index value when performing motion compensation prediction on a block.

[0083] In some examples, the MV candidate list contains up to five candidates for the merge mode and two candidates for the AMVP mode. In other examples, different numbers of candidates may be included in the candidate list for the merge mode and / or the AMVP mode. The merge candidates may contain a set of motion information. For example, the set of motion information may include motion vectors corresponding to two reference picture lists (list 0 and list 1) and reference indices. If the merge candidate is identified by a merge index, the reference picture is used for prediction of the current block and the associated motion vector is determined. However, in the AMVP mode, for each potential prediction direction from list 0 or list 1, the reference index needs to be explicitly signaled together with the index to the candidate list because the AMVP candidate contains only the motion vector. In the AMVP mode, the predicted motion vector may be further refined.

[0084] As seen above, the merge candidate corresponds to a complete motion information set, while the AMVP candidate contains only one motion vector and reference index for a specific prediction direction. Candidates for both modes can be derived similarly from the same spatial and temporal neighboring blocks. In some examples, the merge mode allows an inter-predicted PU to inherit the same motion vector or multiple motion vectors, prediction direction, and reference picture index or multiple reference picture indexes from an inter-predicted PU that includes a motion data position selected from a set of spatially adjacent motion data positions and one of two temporally co-located motion data positions. For the AMVP mode, the motion vector or multiple motion vectors of the PU can be predictively decoded relative to one or more motion vector predictors (MVPs) from an AMVP candidate list constructed by the encoder and / or decoder. In some instances, for unidirectional inter prediction of a PU, the encoder and / or decoder can generate a single AMVP candidate list. In some instances, for bidirectional prediction of a PU, the encoder and / or decoder may generate two AMVP candidate lists, one using motion data of spatially and temporally neighboring PUs from a forward prediction direction, and one using motion data of spatially and temporally neighboring PUs from a backward prediction direction.

[0085] In VVC, there is a reference picture resampling (RPR) tool under consideration, which is described in the following document: S. Wenger, BD. Choi, SCHan, X. Li, S. Liu, "AHG8: Spatial Scalability using Reference Picture Resampling", JVET-00045, the entire contents of which are hereby incorporated by reference and for all purposes. The RPR tool allows the use of reference pictures with a picture size different from the current picture size. In this case, a picture resampling process is called to provide an upsampled or downsampled version of a picture (e.g., a reference picture) that matches the current picture size. The tool is similar to the spatial scalability present in, for example, the scalable extension (SHVC) of the H.265 / HEVC standard.

[0086] This document describes systems, methods, apparatus, and computer-readable media that provide several aspects to increase support for spatial scalability using RPR in VVC. In some cases, pictures with the same content but in different representations (e.g., with different resolutions) may be assumed to have the same picture order count (POC) value (similar to the approach in, for example, HEVC and SHVC).

[0087] Using the techniques described herein, there is no need to change the access unit (AU) definition in VVC because an AU starts from a VCL NAL unit with a new POC value (or a non-VCL NAL unit associated with a VCL NAL unit with a new POC), which occurs when the next picture VCL NAL unit is encountered, as Figure 3 shown, where two different layers labeled as layer 0 and layer 1 are spanned (e.g., layer 0 has a picture with a lower resolution compared to layer 1), the first AU has a POC value of n - 1, the second AU has a POC value of n, and the third AU has a POC value of n + 1.

[0088] However, other layer pictures may need to be included in the reference picture structure (RPS) and the reference picture list (RPL) to indicate which pictures can be used for inter-layer prediction and from which layer. In the current VVC, it is not allowed to reference pictures from other layers.

[0089] Examples from the reference draft VVC document are provided below, such as "Versatile Video Coding (Draft5)" ("Versatile Video Coding (Draft 5)"), version 7, JVET-N1001-v7, the entire content of which is hereby incorporated by reference and for all purposes. As mentioned above, the techniques described herein can be applied to any existing video codec (such as High Efficiency Video Coding (HEVC) and / or Advanced Video Coding (AVC)), or can be proposed as a promising coding tool for current standards under development (such as Versatile Video Coding (VVC)) and for other future video coding standards.

[0090] Figure 2 is a schematic diagram showing an example implementation of the filter unit 91, which can be used as described below with respect to Figure 6 and Figure 7 described, including placing the reference picture in a reference picture storage space 208 (e.g., picture memory 92) that can be identified by the reference picture table 209. Additional aspects of the reference picture table 209 are discussed below. The filter unit 63 can be implemented in the same way. For the scalability support described herein, the filter units 63 and 91 can use POC numbers and reference picture offset values to implement the implementations described herein and other implementations, possibly in combination with other components of the video encoding device 104 or the video decoding device 112.

[0091] As shown, the filter unit 91 includes a deblocking filter 202, a sample adaptive offset (SAO) filter 204, and a conventional filter 206, but compared to Figure 2The filter shown in can include fewer filters and / or can include additional filters. Additionally, in Figure 2 the specific filters shown can be implemented in a different order. Other loop filters (in or after the decoding loop) can also be used to smooth pixel transitions or otherwise improve video quality. When in the decoding loop, the decoded video blocks in a given frame or picture are then stored in the decoded picture buffer (DPB), which stores reference pictures as part of the reference picture storage space 208. The DPB (e.g., the reference picture storage space 208) can be part of additional memory that stores decoded video for later presentation on a display device (such as Figure 1 the display of the video destination device 122) or can be separate from such additional memory.

[0092] Weighted prediction (WP) is a form of prediction that uses reference pictures, in which case a scaling factor (represented by a), a shift amount (represented by s), and an offset (represented by b) are used in motion compensation. For example, a linear model can be used in WP, as shown by the following equation (1):

[0093] p(i,j) = a*r(i + dv x ,j + dv y ) + b, where (i,j) ∈ PU c Equation (1)

[0094] In equation (1), PU c is the current PU, (i,j) are the coordinates of the samples (or pixels in some cases) in the PU c , (dv x ,dv y ) is the difference vector of the PU c , p(i,j) is the prediction of the PU c , r is the reference picture of the PU, and a and b are model parameters (where a is the scaling factor and b is the offset, as mentioned above).

[0095] When WP is enabled, for each reference picture of the current slice, signaling flags are used to indicate whether WP is applied to that reference picture. If WP is applied to a reference picture, the WP parameter set (i.e., a, s, and b) is sent to the decoder and used for motion compensation from the reference picture. In some examples, to provide flexibility in turning WP on or off for luminance and chrominance samples, the WP flag and WP parameters may be signaled separately for the luminance and chrominance components of the pixels. In some cases, in WP, the same WP parameter set is used for all samples in a reference picture. In some examples, variables specify the width and height of the current decoding block, and an array with the same width and height is used as the prediction samples. The prediction samples may be derived using a weighted sample prediction process and are used when reconstructing the picture.

[0096] Pictures may include one or more offsets (e.g., offset b), one or more weights or scaling factors (e.g., scaling factor a), shift amounts, and / or other suitable weighted prediction parameters. For bi-directional inter prediction, one or more weights may include a first weight for a first reference picture and a second weight for a second reference picture. As described herein, in some examples, a reference block may be the same block as the current block but with different associated weight parameters (e.g., for weighted prediction). In other examples, the reference block may be a block of the same picture but with a different size (e.g., different resolution), such that the reference block is part of an inter-layer prediction process, as described in more detail with respect to Figure 3 In some examples, both weighted prediction and inter-layer prediction may be performed on a single current block with reference blocks having different weights and associated with pictures of different sizes.

[0097] Figure 3 is a schematic diagram showing a series of access units 310, 320, and 330. Each access unit has an associated POC value, shown as POC values 316, 326, and 336 corresponding to access units 310, 320, and 330. Each access unit has two layers for picture frames of different sizes (e.g., resolutions). Figure 3Examples include layer 0 unit 312 and layer 1 unit 314 for access unit 310, layer 0 unit 322 and layer 1 unit 324 for access unit 320, and layer 0 unit 332 and layer 1 unit 334 for access unit 330. Layer 0 unit 312 (e.g., VCL NAL unit) will have a different size from layer 1 unit 314 (e.g., also a VCL NAL unit). As mentioned above, using the techniques described herein, the VVC AU definition does not need to be changed because the AU starts from a VCL NAL unit with a new POC value or from a non-VCL NAL unit associated with a VCL NAL unit with a new POC. For example, when the next picture VCL NAL unit is encountered, as Figure 3 shown (where there are two different layers labeled layer 0 and layer 1 (e.g., layer 0 has a lower resolution picture compared to layer 1), the first AU has a POC value of n-1, the second AU has a POC value of n, and the third AU has a POC value of n+1), the new POC value can be used. Figure 3 Examples include two layers, but in each example, other numbers of layers can be used.

[0098] The examples described herein utilize syntax and operational structures for efficiently allowing data from other layers to be used as reference pictures by pictures in different layers, improving the operation of VVC devices and networks. As described herein, such improvements allow for efficient signaling for identifying reference pictures in another layer of a shared access unit when weighted prediction is enabled, and for efficient signaling for using reference pictures with an offset of zero signaled in other frames when weighted prediction is not enabled.

[0099] VVC operates using both short-term and long-term reference pictures. Some examples described herein operate using signaling of reference pictures for short-term reference pictures. In various systems, short-term reference pictures are signaled as a POC increment relative to the previous short-term reference picture as the POC of the previous reference picture. The first POC value can be initialized to be equal to the current picture POC. Allowing the incremental POC to be equal to 0 (e.g., there may be multiple reference pictures with the same POC value, such as different resolutions of the same picture), but the POC value cannot be equal to the current picture POC (e.g., the current picture cannot be inserted as a reference picture into the reference picture structure or list). In each example, the long-term reference picture POC is signaled as the most significant bit (MSB) and the least significant bit (LSB), and the LSB part can be equal to 0.

[0100] In some examples, as Figure 2 shown, the table 209 referring to reference pictures is stored in the DPB (e.g., using memory or asFigure 2The reference picture storage space shown (208). In some examples, the table 209 is a list of picture order count offset values from the current picture or a previous picture, which identifies pictures from the reference picture storage space 208 (e.g., as part of a picture memory such as picture memory 92). In some examples, such a table 209 can be generated in a loop, where the initial table value is equal to the initial picture POC value. A base POC value (e.g., a picture order count offset value) is assigned to identify the POC value of the reference picture for the initial picture. The POC value of the first reference picture for the initial picture associated with the initial POC value is assigned as the base POC value for the table. Then, all subsequent reference pictures are identified as incremental values from the base POC value (e.g., the POC value of the previous reference picture), with an incremental POC and a new base POC value for each entry in the table. The process loops until the table 209 is complete, which has all the reference pictures for each picture associated with the table. As described above, the initial incremental POC is added to the first picture POC, and subsequent incremental POCs are added to the previous reference picture POC. Then, the bitstream signal is constructed using the incremental POC and the base POC to limit the data usage in the bitstream while allowing the decoder to reconstruct the table 209. Then, the repeated reference pictures from the table 209 described above can be used for prediction. In some examples using weighted prediction, when the same picture is being used for prediction but with different weights, the incremental POC for the reference picture and the picture being processed can be zero. In some examples, certain types of prediction can use a modified table or modified incremental POC values. For example, two pictures with the same POC can have different weights. In an illustrative example, one picture in the pictures can have a weight value of 1.5 and a reference picture with the same weight having a value of 2.5. Thus, the same picture with different weighting parameters can be used as a reference with a zero value incremental POC (also referred to as "zero value picture order count offset" or "zero value POC offset"), which identifies that the reference picture is being used for weighted prediction. In another example, the modified incremental POC value is used when weighted prediction is disabled. For example, in some cases, weighted prediction can be turned off by a flag (such as weighted_pred_flag), which can be signaled in the picture parameter set (PPS). The weighted_pred_flag can be referred to as the PPS weighted prediction flag. In some examples, when weighted prediction is disabled, signaling of the same value POC for more than one reference picture is not allowed.To achieve such a restriction or constraint, weighted prediction flags (e.g., sps_weighted_pred_flag and sps_weighted_bipred_flag) may be signaled in a parameter set (e.g., in the SPS since the RPS information may be signaled in the SPS or other parameter sets). The sps_weighted_pred_flag and / or sps_weighted_bipred_flag may be referred to as sequence parameter set (SPS) weighted prediction flags.

[0101] Allowing a zero delta POC value when weighted prediction is disabled results in inefficiency in delta POC signaling because there is no reason to insert the same reference picture into the reference picture list multiple times. In such a case, the delta POC value will be at least equal to 1, and the delta POC value 0 is not used. However, in some examples, a zero delta POC value may be associated with a reserved codeword. The use of PPS and SPS flags allows for different handling of delta POC signaling in weighted prediction and non-weighted prediction cases as described herein. Thus, flags may be used to allow a zero picture order count offset (e.g., zero delta POC value) based on a determination of enabling weighted prediction for a picture slice (or other part of a picture such as a CTU, CU, or other block), and to implement decoding a zero delta POC signaled as a different value (e.g., the signaled value plus 1) when weighted prediction is disabled. Thus, for weighted prediction, the same picture with different weights is used for weighted prediction. For inter-layer prediction, different layer pictures from the same access unit may be used as reference pictures. The conditional use of zero delta POC signaling values provides efficiency in reference picture signaling and improved system and device operation, while limiting the additional overhead for signaling and improving system and device performance when weighted prediction is disabled.

[0102] For a zero picture order count offset where the reference picture is in the same access unit as the current picture being processed, a reference picture sampling tool may be used to generate the necessary reference data for processing the current picture. In some examples, the reference picture sampling tool may be part of a filter unit (e.g., filter unit 91), or in other examples may be part of any aspect of a device for encoding or decoding as described herein.

[0103] An illustrative example of a modification with respect to the VVC draft (e.g., Versatile Video Coding (Draft 5), Version 7) for such flags is shown below. Additional examples of syntax tables, syntax terms, syntax logic and values, and example implementations are shown below. In cases where an addition is shown, in “ <insert>”and" <insertend>" The symbols are described using underlines and text between them (e.g., " <insert>Additional text <insertend>”). An example for weighted prediction tagging is as follows:

[0104] <insert>When sps_weighted_pred_flag is equal to 0, it specifies that weighted prediction is not applied to P slices. When weighted_pred_flag is equal to 1, it specifies that weighted prediction can be applied to P slices.

[0105] When sps_weighted_bipred_flag is equal to 0, it specifies that default weighted prediction is applied to B slices. When weighted_bipred_flag is equal to 1, it specifies that weighted prediction can be applied to B slices. <insertend>

[0106] In such an example, the flag signaled in the PPS may be constrained by the SPS weighted prediction flag (e.g., the PPS flag may be equal to 1 only if the SPS flag is equal to 1). In another illustrative example:

[0107] weighted_pred_flag being equal to 0 specifies that weighted prediction is not applied to P slices. weighted_pred_flag being equal to 1 specifies that weighted prediction may be applied to P slices; <insert>The weighted_pred_flag can be equal to 1 only when the corresponding sps_weighted_pred_flag is equal to 1 <insertend>。

[0108] When weighted_bipred_flag is equal to 0, it specifies that the default weighted prediction is applied to B slices. When weighted_bipred_flag is equal to 1, it specifies that weighted prediction can be applied to B slices; <insert>The weighted_bipred_flag can be equal to 1 only when the corresponding sps_weighted_bipred_flag is equal to 1 <insertend>。

[0109] In some examples, the incremental POC equal to zero and the POC LSB equal to zero may be allowed only when weighted prediction is enabled (e.g., as indicated by a flag in the SPS). In another example, when inter-layer prediction is enabled, zero incremental POC and zero POC LSB are required. The enabling of the inter-layer prediction flag may be signaled in a parameter set such as the SPS. Taking the above two cases together, the incremental POC equal to zero and the POC LSB equal to zero may be allowed only when weighted prediction or inter-layer prediction is enabled.

[0110] In VVC, the POC values of the reference pictures are signaled in the ref_pic_list_struct (RPS) as follows:

[0111]

[0112] Table 1

[0113] Syntax examples include:

[0114] abs_delta_poc_st[listIdx][rplsIdx][i], which specifies the absolute difference between the picture order count value of the current picture and the picture referenced by the i-th entry when the i-th entry is the first STRP entry in the ref_pic_list_struct (listIdx, rplsIdx) syntax structure, or, when the i-th entry is a STRP entry but not the first STRP entry in the ref_pic_list_struct (listIdx, rplsIdx) syntax structure, specifies the absolute difference between the picture order count value of the picture referenced by the i-th entry and the picture order count value of the picture referenced by the previous STRP entry in the ref_pic_list_struct (listIdx, rplsIdx) syntax structure. The value of abs_delta_poc_st[listIdx][rplsIdx][i] shall be in the range of 0 to 2 15 –1 (inclusive).

[0115] rpls_poc_lsb_lt[listIdx][rplsIdx][i] specifies the value of the picture order count modulo MaxPicOrderCntLsb of the picture referenced by the i-th entry in the ref_pic_list_struct(listIdx,rplsIdx) syntax structure. The length of the rpls_poc_lsb_lt[listIdx][rplsIdx][i] syntax element is log2_max_pic_order_cnt_lsb_minus4 + 4 bits.

[0116] In an illustrative implementation example, when both weighted prediction and inter-layer prediction are disabled, the syntax elements that can indicate zero delta POC (e.g., using a zero picture order count offset to identify a reference picture) or the same value POC for a reference picture (such as abs_delta_poc_st and rpls_poc_lsb_lt) cannot be equal to 0 or cannot indicate the same POC value as that already used for other reference pictures. In some examples, zero delta POC is not allowed when weighted prediction is disabled. In some examples, zero delta POC is not allowed when inter-layer prediction is disabled.

[0117] In another illustrative implementation example, when both weighted prediction and inter-layer prediction are disabled, the values of these syntax elements (e.g., the values used to signal in the bitstream to identify a reference picture, such as the picture order count offset) can be considered as the syntax element minus 1. In some such examples, the value of the syntax element is reconstructed as the signaled value plus 1. In such examples, it is not possible to signal a 0 value. An implementation example is shown below:

[0118] abs_delta_poc_st[listIdx][rplsIdx][i], when the i-th entry is the first STRP entry in the ref_pic_list_struct(listIdx,rplsIdx) syntax structure, specifies the absolute difference between the picture order count value of the current picture and the picture referenced by the i-th entry, or, when the i-th entry is a STRP entry but not the first STRP entry in the ref_pic_list_struct(listIdx,rplsIdx) syntax structure, specifies the absolute difference between the picture order count value of the picture referenced by the i-th entry and the picture order count value of the picture referenced by the previous STRP entry in the ref_pic_list_struct(listIdx,rplsIdx) syntax structure <insert>Or when weighted prediction and inter-layer prediction are disabled, add 1 to the absolute difference <insertend>。

[0119] The value of abs_delta_poc_st[listIdx][rplsIdx][i] shall be in the range of 0 to 2 15 –1 (inclusive).

[0120] rpls_poc_lsb_lt[listIdx][rplsIdx][i] specifies the value of the picture order count modulo MaxPicOrderCntLsb of the picture referred to by the i-th entry in the ref_pic_list_struct(listIdx,rplsIdx) syntax structure <insert>Alternatively, when weighted prediction and inter-layer prediction are disabled, the value of Picture Order Count modulo MaxPicOrderCntLsb is incremented by 1 <insertend>The length of the rpls_poc_lsb_lt[listIdx][rplsIdx][i] syntax element is log2_max_pic_order_cnt_lsb_minus4 + 4 bits.

[0121] In some examples, it is recommended to add other layer pictures to the RPS. However, since the RPS only has POC values to identify pictures, the layer ID is signaled additionally for each reference picture.

[0122] In another example, to avoid signaling the layer ID for each reference picture, the layer ID is signaled only when the POC value of the reference picture is equal to the POC value of the current picture. If the incremental POC relative to the current POC value is equal to 0 (e.g., the reference picture POC is equal to the current picture POC), then the layer ID is signaled to identify the layer. In this example, pictures with layer IDs different from the current picture's layer ID can be used for prediction only when the picture has the same POC as the current picture (or belongs to the same access unit as the current picture). Additionally, inter-layer reference pictures can be regarded as only long-term reference pictures (LTRPs), and can be marked as LTRPs to avoid motion vector scaling. In this case, the layer ID can be signaled only for the LTRP type. In another example, the layer ID is signaled only for LTRPs and when the incremental POC is equal to 0 or the reference picture has the same POC value as the current picture. In another example, before signaling the POC value of the reference picture, the layer ID can be signaled for the reference picture or only for LTRPs. In this case, the POC value may be signaled conditionally based on the signaled layer ID and the current layer ID. For example, the POC is not signaled when the layer ID is not equal to the current layer ID (e.g., inter-layer reference picture), and then the POC value is inferred to be equal to the current picture POC. In some examples, the layer ID can be an indicator for identifying the layer used for prediction, and can be a layer number, layer index, or another indicator within the list of layers.

[0123]

[0124]

[0125] Table 2

[0126] According to the above syntax table for ref_pic_list_struct shown in Table 2, an example may include: num_ref_entries may include the total number of both customary and inter-layer reference pictures. In an alternative, the number of inter-layer reference pictures num_inter_layer_ref_entries[listIdx][rplsIdx] is signaled separately. For example, a high-level flag signaled in a parameter set (PS) (such as sps_interlayer_ref_pics_present_flag signaled in SPS) may indicate whether there are inter-layer reference pictures.

[0127] <insert>When sps_inter_layer_ref_pics_present_flag is equal to 1, it specifies that inter-layer prediction can be used when decoding pictures that reference the SPS. When sps_inter_layer_ref_pics_present_flag is equal to 0, it specifies that inter-layer prediction is not used when decoding pictures that reference the SPS. <insertend>

[0128] In some such examples, instead of signaling the nuh_layer_id directly, the layer index may be signaled. For example, layer 10 may use layer 0 and layer 9 as dependent layers. Instead of signaling the values 0 and 9 directly, the layer indices 0 (corresponding to layer ID 0) and 1 (corresponding to layer ID 9) may instead be signaled. The mapping between the layer ID and the layer index may be inferred or derived at the decoder.

[0129] In some examples, the separate flag il_ref_pic_flag may be signaled as shown in Table 3 below to indicate that a picture is used for inter-layer prediction. The flag may be signaled for each picture or only for the LTRP (e.g., when st_ref_pic_flag is equal to 0). In some such examples, the layer ID or layer index may be signaled only when il_ref_pic_flag is equal to 1. Such a flag may be needed when signaling the layer index because the current layer ID is not included as a dependent layer and may not have an associated layer index.

[0130]

[0131]

[0132] Table 3

[0133] Then, Table 4 below shows an example of the syntax table for ref_pic_list_struct, with the following details:

[0134] <insert>layer_dependency_idc[listIdx][rplsIdx][i] specifies the layer dependency index for the i-th entry in the ref_pic_list_struct(listIdx,rplsIdx) syntax structure. The value of layer_dependency_idc[listIdx][rplsIdx][i] shall be in the range of 0 to the current layer idc LayerIdc[nuh_layer_id] minus 1 minus 1 (inclusive).

[0135] il_ref_pic_flag[listIdx][rplsIdx][i] being equal to 1 specifies that the i-th entry in the ref_pic_list_struct(listIdx,rplsIdx) syntax structure is an inter-layer reference picture. il_ref_pic_flag[listIdx][rplsIdx][i] being equal to 0 specifies that the i-th entry in the ref_pic_list_struct(listIdx,rplsIdx) syntax structure is not an inter-layer reference picture entry. When not present, the value of il_ref_pic_flag[listIdx][rplsIdx][i] is inferred to be equal to 0.

[0136] num_inter_layer_ref_entries[listIdx][rplsIdx] specifies the number of direct inter-layer entries in the ref_pic_list_struct(listIdx,rplsIdx) syntax structure. The value of num_inter_layer_ref_entries[listIdx][rplsIdx] shall be in the range of 0 to sps_max_dec_pic_buffering_minus1 + 14 – num_ref_entries[listIdx][rplsIdx] (inclusive). <insertend>

[0137] In some such examples, the number of inter-layer reference pictures num_inter_layer_ref_entries can be limited to not more than the maximum number of reference pictures minus the number of customary inter-layer reference pictures. For example, the value of num_inter_layer_ref_entries[listIdx][rplsIdx] should be in the range from 0 to sps_max_dec_pic_buffering_minus 1 + 14 – num_ref_entries[listIdx][rplsIdx] (inclusive), which is added to the semantics of num_inter_layer_ref_entries. The num_inter_layer_ref_entries syntax element can be referred to as the inter-layer reference picture entry syntax element. The ref_pic_list_struct(listIdx,rplsIdx) syntax structure can be referred to as the reference picture list syntax structure. In some examples, when inter-layer reference pictures are allowed, il_ref_pic_flag[][] is signaled for all long-term reference pictures.

[0138]

[0139]

[0140] Table 4

[0141] Then, Tables 5 and 6 show example syntax tables for ref_pic_list_struct and the associated slice headers. In such additional examples, the layer information can be signaled in the slice header. This signaling can be outside the signaling allowed in the ref_pic_list_struct() syntax structure. Such a definition helps in the case where the layer information is signaled only when the LSB of the reference picture is equal to the LSB of the current picture.

[0142]

[0143]

[0144] Table 5

[0145]

[0146]

[0147] Table 6

[0148] In other embodiments described by Table 7, the il_ref_pic_flag[][] in the slice header is signaled only when the POC LSB of the long-term reference picture is equal to the POC LSB of the current picture. In some examples, rpls_poc_lsb_lt can also be used.

[0149]

[0150] Table 7

[0151] In some examples, Table 8 below can operate as an alternative to the examples of Table 6 and Table 7 above:

[0152]

[0153] Table 8

[0154] In some examples, when the indication of layer_depdendency_idc is not necessary for identifying the inter-layer reference picture, the layer_dependency_idc syntax element may not be sent. In some examples, if there is only one inter-layer reference picture that can be used for reference and the il_ref_pic_flag[][][] indicates that the reference picture can be used for reference to meet condition A (e.g., as defined below), the layer_dependency_idc is not signaled and is inferred to be the inter-layer reference picture that can be used for reference. In some cases, when there is more than one inter-layer reference picture that can be used for reference to meet condition A (e.g., n in number), no more than n - 1 layer_dependency_idc values are signaled to indicate that the specific inter-layer reference picture at the i position in the list may be sufficient. In such examples, condition A can be any condition that allows inter-layer prediction for the current picture. For example, in some cases, condition A checks whether the inter-layer reference picture shares the same POC LSB as the current picture. In some cases, condition A checks whether the inter-layer reference picture shares the same POC as the current picture. In some cases, condition A checks that the difference between the POC of the inter-layer reference picture and the POC of the current picture does not exceed a threshold value. In some cases, condition A can be restricted to be applied within the access unit. In some examples, when there is no picture that meets condition A for the current picture, the indication of the inter-layer reference picture (e.g., the relevant syntax element) may not be signaled.

[0155] In some instances, when a reference picture is indicated as an inter-layer reference picture (e.g., using il_ref_pic_flag[] or similar syntax), the picture notation can be modified to indicate those pictures as "for inter-layer reference pictures". In some cases, this notation can be combined with other notations (e.g., for short-term reference or for long-term reference), or not include other notations (e.g., the picture can be marked as only two of the following cases: for short-term reference or for long-term reference or for inter-layer reference pictures). In some cases, the notation of the picture can be modified as follows:

[0156] - After the picture is decoded, it is marked as "for inter-layer reference".

[0157] - After the access unit is completed, or when starting the next access unit, all pictures in the current access unit are marked for short-term reference.

[0158] In some cases, the discardable_flag is signaled to indicate that the picture may not be used for inter-layer reference. In this case, some modifications can be applied to some of the methods disclosed above. In one example, a picture with a discardable_flag equal to 1 can be marked as "not for reference" after the decoding of the picture. In another example, a picture with a discardable_flag equal to 1 may never be marked as anything other than "not for reference". In another example, a picture with a discardable_flag equal to 1 may not be considered for any prediction; thus, when determining certain conditions that could originally be applied to an inter-layer reference picture or a reference picture, it may be disregarded. For example, if there are two pictures A and B in different layers from the current picture in the current access unit, and only one of them (e.g., picture A) has a discardable_flag equal to 1, then when indicating the inter-layer reference picture, the system can infer that the inter-layer reference picture is picture B, in which case picture A is not added to any reference picture list in the reference picture list.

[0159] The mapping table mentioned above (e.g., vps_direct_dependency_flag[i][j]) can be signaled, for example, in a parameter set (such as the VPS shown in Example Table 10).

[0160]

[0161] Table 10

[0162] Example Table 10 can operate with the following syntax elements:

[0163]

[0164] vps_independent_layer_flag[i] being equal to 1 specifies that the layer with index i does not use inter-layer prediction. vps_independent_layer_flag[i] being equal to 1 specifies that the layer with index i can use inter-layer prediction and that vps_layer_dependency_flag exists in the VPS.

[0165] vps_direct_dependency_flag[i][j] being equal to 0 specifies that the layer with index j is not a direct reference layer for the layer with index i. vps_direct_dependency_flag[i][j] being equal to 1 specifies that the layer with index j is a direct reference layer for the layer with index i. When vps_direct_dependency_flag[i][j] does not exist for i and j in the range from 0 to MaxLayersMinus1, it is inferred to be equal to 0.

[0166] The variable DependencyIdc[i][j] is derived as follows:

[0167]

[0168] In the reference picture list according to some examples, when the POC value is not signaled (e.g., when the layer ID is not equal to the current layer ID, such as for inter-layer reference pictures), the POC value is inferred to be equal to the POC of the current picture. The layer ID is an indicator for identifying a layer and can be the layer index for prediction within a list of layers. In other examples, other indicators can be used. Num_ref_entries can include the total number of both conventional and inter-layer reference pictures. In some examples, a separate reference flag can be signaled to indicate that a picture is used for inter-layer prediction in each example, and this flag can be signaled for each reference picture.

[0169] In some examples, the inter-layer reference picture is not indicated by the POC value. In such examples, the inter-layer reference picture can be indicated by a flag (e.g., inter_layer_ref_pic_flag) and a layer index. The POC value of the inter-layer reference picture can be derived to be equal to the POC of the current picture according to the following condition: there is a reference picture in the DPB that has a nuh_layer_id equal to the reference picture layer ID and the same picture order count value as the current picture. Thus, as described above, in some examples, the reference picture list structure can be constructed for signaling inter-layer prediction using the layer ID and an inferred POC value of 0. In other examples, other reference picture list structures can be used.

[0170] In some such examples, the reference picture lists RefPicList[0] and RefPicList[1] are constructed as follows:

[0171]

[0172]

[0173] In some cases, if nuh_layer_id[i][RplsIdx[i]][j] is not signaled directly, it can be derived from the LayerIdc and DependencyIdc variables using layer_dependency_idc. For example, nuh_layer_id[listIdx][rplsIdx][j] = DependencyIdc[LayerIdc[nuh_layer_id]][layer_dependency_idc[listIdx][rplsIdx][j]], where nuh_layer_id is the layer ID of the current picture. This value can be used as the layer ID in the reference picture list derivation.

[0174] In some cases, the following constraint can be added, i.e., when signaling an inter-layer reference picture in the RPS, the POC value of the inter-layer reference picture derived from the RPS should be equal to the POC value of the current picture. When using inter-layer prediction, a zero motion vector can be assumed, in which case there is no displacement of the content in pictures of different representations (resolutions). The motion information can be signaled for the inter-frame prediction mode. To save overhead, the motion information can be signaled when (in some cases, only when) it is not used for inter-layer prediction. Otherwise (if for inter-layer prediction), the motion vector is inferred to be equal to 0 without signaling.

[0175] Table 11 is an example of an implementation based on the VVC draft (e.g., Versatile Video Coding (draft 5), version 7), which has variations in coding_unit(x0,y0,cbWidth,cbHeight,treeType) as follows:

[0176]

[0177]

[0178] Table 11

[0179] Among them, as shown in example Table 12, in the slice header, the following described syntax elements can be used to signal information about the number of active inter-layer pictures:

[0180]

[0181] Table 12

[0182] <insert>The num_inter_layer_active_override_flag being equal to 1 specifies the existence of the syntax element num_inter_layer_ref_pics_minus1[0] for P and B slices and the syntax element num_inter_layer_ref_pics_minus1[1] for B slices. The num_inter_layer_active_override_flag being equal to 0 specifies the non-existence of the syntax elements num_inter_layer_ref_pics_minus1[0] and num_inter_layer_ref_pics_minus1[1]. When they do not exist, the value of num_inter_layer_active_override_flag is inferred to be equal to 0.

[0183] num_inter_layer_ref_pics_minus1[i] is used for the derivation of the variable NumInterLayerActive[i]. The value of num_ref_idx_active_minus1[i] shall be in the range of 0 to num_inter_layer_ref_entries[i][RplsIdx[i]]. The variable NumInterLayerActive[i] is derived as follows:

[0184] for(i = 0; i < 2; i++)

[0185] if(slice_type == B || (slice_type == P && i == 0))

[0186] NumInterLayerActive[i] = num_inter_layer_active_override_flag?

[0187] num_inter_layer_ref_pics_minus1[i] + 1 : num_inter_layer_ref_entries[i][RplsIdx[i]] <insertend>

[0188] In some examples, a spatial ID similar to the temporal ID can be defined, which indicates the spatial representation of pictures that may not be captured by the nuh_layer_id. One of the purposes of signaling the nuh_layer_id in the NAL unit header is to easily identify the NAL units belonging to a specific layer, which can be removed in the case of sub-bitstream extraction purposes. Easily identifying the NAL units belonging to a specific layer provides benefits for the operation of the system, and thus the spatial_id is also considered to be signaled in the NAL unit header. For this purpose, the nuh_reserved_zero_bit can be renamed to layer_id_interpret_flag, and the semantics can be modified as follows, or in some cases, a new syntax element spatial_id can also be sent.

[0189] - When the layer_id_interpret_flag is equal to 0, the syntax element nuh_layer_id_plus1 is currently interpreted as specifying the layer ID of the current picture.

[0190] - When the layer_id_interpret_flag is equal to 1, the syntax element nuh_layer_id_plus1 is interpreted as containing some bits for indicating the layer ID and some bits for specifying the spatial ID. For example, the first four bits can indicate the layer ID, and the last three bits can indicate the spatial ID. Generally, in this case, the actual layer ID and spatial ID may be derived from the syntax element nuh_layer_id_plus1 using a predetermined method.

[0191] The difference between the spatial_id and the layer ID is that the consistency of the "single-layer" decoder is also involved when decoding a bitstream containing multiple spatial_ids. The value of the spatial_id can be used to identify the reference pictures in the "single-layer" decoder without reusing the layer concept, and in this way, the concept of the layer will be clearly separated from the temporal and spatial layers. For simplicity, the consistency can be further simplified so that a simple decoder does not have to consider bitstreams with non-zero spatial_ids: that is, single-layer bitstreams that conform to a type with only a single layer and one spatial ID (e.g., main-simple) and single-layer bitstreams that have a single layer but can contain multiple spatial IDs (e.g., main-representation).

[0192] Within a single layer, pictures with multiple spatial ID values can be considered as part of different spatial access units: a picture with a spatial_id equal to 0 and POC 0 belongs to one spatial access unit, while a picture with spatial ID equal to 1 and POC 0 (in the same access unit) belongs to another spatial access unit. The concept of spatial access units may not necessarily have a different relationship with the output time (similar to layer access units), but they are important in the different representations within the CPB and DPB. When one or more spatial access units within a layer access unit are to be output, the output order may also be affected by the spatial access units. Pictures belonging to a spatial_id may depend on other pictures with lower spatial_id; pictures belonging to spatial ID S may not reference pictures belonging to the same layer with a spatial ID greater than S.

[0193] The maximum number of spatial layers can be specified for a bitstream, and the POC value for each picture can be derived based on the picture order count LSB (syntax element) and the spatial ID. For example, the POC of pictures spanning layers can be constrained to be the same; while the POC of pictures within the same access unit in a layer but with different spatial IDs may not be the same. In some cases, this constraint may not be explicitly signaled, but is derived in the decoder based on the pic_order_cnt_lsb and spatial_id values.

[0194] Figure 4 FIG. is a flow chart illustrating an example process 400 according to some examples. In some examples, process 400 is performed by a decoding device (e.g., decoding device 112). In some examples, process 400 may be performed by an encoding device (e.g., encoding device 104), such as when performing a decoding process for storing one or more reference pictures in a decoded picture buffer (DPB). In other examples, process 400 may be implemented as instructions in a non-transitory storage medium, which when executed by a processor of a device, cause the device to perform process 400. In some cases, when process 400 is performed by a video decoder, the video data may include a decoded picture or a portion of a decoded picture (e.g., one or more blocks) included in an encoded video bitstream, or may include multiple decoded pictures included in the encoded video bitstream.

[0195] At block 402, process 400 obtains at least a portion of a picture from the bitstream. In some examples, this portion of the picture is a slice of the picture. As described herein, this can be a bitstream obtained by a mobile device (such as a smart phone, computer, any device with a screen, a television, or any device with decoding hardware).

[0196] At block 404, process 400 determines, based on the bitstream, that weighted prediction is enabled for the portion of the picture. Process 400 may determine that weighted prediction is enabled by parsing the bitstream for a flag that indicates weighted prediction. As described herein, the parsing may be performed using SPS flags and / or PPS flags. In some cases, the PPS flags are constrained by the SPS flags. In some examples, the flags (e.g., SPS flags and / or PPS flags) may specify a unidirectionally predicted frame (e.g., a P-frame) or a bidirectionally predicted frame (e.g., a B-frame).

[0197] At block 406, based on determining that weighted prediction is enabled for the portion of the picture, process 400 identifies a zero picture order count offset that indicates a reference picture from a reference picture list. As described herein, the zero picture order count offset may identify a reference picture from a table, and the reference picture may then be used for weighted prediction. The reference picture for weighted prediction, referenced by the zero picture order count offset at block 404, may be a version of the picture that includes the portion of the picture obtained at block 402, where the picture from block 402 has different weighting parameters (e.g., a copy with different weights) than the reference picture indicated by the zero picture order count offset.

[0198] At block 408, process 400 reconstructs at least that portion of the picture using at least a portion of the reference picture identified by the zero picture order count offset. The reconstruction of the picture may be performed as part of the reconstruction of the picture according to the bitstream (as part of a decoding operation compliant with VVC), or may be part of any similar decoding operation. The process 400 may be repeated for additional portions of the picture or for other pictures obtained from the bitstream. As part of the process, different respective picture order count offsets are identified for multiple portions of the picture. The process may also be repeated for any number of pictures, some of which use weighted prediction and others for which weighted prediction is disabled. When weighted prediction is enabled, some reference pictures of the video will have an associated picture order count offset value of zero, and other reference pictures will have non-zero values, where frames from other access units are used for weighted prediction of the current frame. Then, as the process is repeated for additional pictures, the associated pictures are reconstructed.

[0199] In some examples, an extra frame flag is used to indicate a flag that disables weighted prediction. In some such embodiments, an extra picture or picture part or picture slice may be processed by reconstructing the value of a syntax element in the bitstream that indicates a plurality of corresponding picture order count offsets to a value signaled plus one. In some such examples, when weighted prediction is disabled, the abs_delta_poc_st short-term reference picture syntax value specifies the absolute difference between the picture order count value of a second picture and the previous short-term reference picture entry in the reference picture list for a second reference picture as the abs_delta_poc_st short-term reference picture syntax value plus one.

[0200] In some examples, both weighted prediction and inter-layer prediction may be applied to a portion of a picture such that the reference picture has a different size than the picture, and the reference picture is associated with a set of weights that is different from a second different set of weights associated with the picture (e.g., from block 402). In some such examples, the reference picture is signaled as a short-term reference picture having a least significant bit value of the picture order count equal to zero.

[0201] In some examples, the picture is a non-instantaneous decoding refresh (non-IDR) picture (e.g., the picture may include another type of picture or a random access picture), as described above.

[0202] Figure 5 FIG. 500 is a flow diagram illustrating an example of a process 500 according to some examples. In some examples, process 500 is performed by an encoding device (e.g., encoding device 104). In other examples, process 500 may be implemented as instructions in a non-transitory storage medium that, when executed by a processor of a device, cause the device to perform process 500. In some cases, when process 500 is performed by a video encoder, the video data may include a picture or a portion of a picture (e.g., one or more blocks) to be encoded in an encoded video bitstream, or may include a plurality of pictures to be encoded in an encoded video bitstream.

[0203] At block 502, process 500 includes identifying at least a portion of a picture. At block 504, process 500 includes selecting weighted prediction to be enabled for the portion of the picture. At block 506, process 500 includes identifying a reference picture for the portion of the picture. Identifying the reference picture may include selecting weights for each picture and associating the picture with a copy having different weights as a reference picture for a first picture.

[0204] At block 508, process 500 includes generating a zero-valued picture order count offset for indicating a reference picture from a reference picture list. At block 510, process 500 includes generating a bitstream. The bitstream includes that portion of the picture and the zero-valued picture order count offset as associated with that portion of the picture.

[0205] As with process 400 described above, process 500 may be implemented in a variety of devices. In some implementations, process 500 is performed by a processing circuit of a mobile device that has a camera coupled to a processor and a memory for capturing and storing at least one picture. In other examples, process 500 is performed by any device having a display coupled to a processor, where the display is configured to display at least one picture before processing and transmitting the picture in a bitstream using process 400.

[0206] In some implementations, the processes (or methods) described herein (including processes 400 and 500) may be performed by a computing device or apparatus (such as system 100 shown in Figure 1 ). For example, the processes may be performed by the encoding device 104 shown in Figure 1 and Figure 6 , by another video source side device or video transmission device, by the decoding device 112 shown in Figure 1 and Figure 7 and / or by another client side device (such as a player device, a display, or any other client side device). In some cases, the computing device or apparatus may include one or more input devices, one or more processors, one or more microprocessors, one or more microcomputers, and / or other components of a device configured to perform the steps of the processes described herein.

[0207] In some examples, a computing device or apparatus can include or can be a mobile device, a desktop computer, a server computer and / or server system, or other types of computing devices. Components of the computing device (e.g., one or more input devices, one or more output devices, one or more processors, one or more microprocessors, one or more microcomputers, and / or other components) can be implemented in circuitry. For example, the components can include electronic circuits or other electronic hardware and / or can be implemented using electronic circuits or other electronic hardware, which can include one or more programmable circuits (e.g., a microprocessor, a graphics processing unit (GPU), a digital signal processor (DSP), a central processing unit (CPU), and / or other suitable electronic circuits), and / or can include computer software, firmware, or any combination thereof and / or can be implemented using computer software, firmware, or any combination thereof to perform the various operations described herein. In some examples, the computing device or apparatus can include a camera configured to capture video data (e.g., a video sequence) including video frames. In some examples, the camera or other capture device that captures the video data is separate from the computing device, in which case the computing device receives or obtains the captured video data. The computing device can include a network interface configured to transmit the video data. The network interface can be configured to transmit data based on the Internet Protocol (IP) or other types of data. In some examples, the computing device or apparatus can include a display for displaying output video content (such as samples of pictures of a video bitstream).

[0208] A process can be described with respect to a logic flowchart, the operations of which represent a series of operations that can be implemented with hardware, computer instructions, or a combination thereof. In the context of computer instructions, the operations represent computer-executable instructions stored on one or more computer-readable storage media that, when executed by one or more processors, implement the recited operations. Generally, computer-executable instructions include routines, programs, objects, components, data structures, etc. that perform a particular function or implement a particular data type. The order in which the operations are described is not intended to be construed as a limitation, and any number of the described operations can be combined in any order and / or in parallel to implement the process.

[0209] Additionally, the process may be under the control of one or more computer systems configured to have executable instructions, and may be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) executed jointly on one or more processors, implemented by hardware, or a combination thereof. As previously mentioned, the code may be stored, for example, in the form of a computer program including a plurality of instructions executable by one or more processors, on a computer-readable or machine-readable storage medium. The computer-readable storage medium or machine-readable storage medium may be non-transitory.

[0210] The decoding techniques discussed herein may be implemented in an example video encoding and decoding system (e.g., system 100). In some examples, the system includes a source device that provides encoded video data to be decoded later by a destination device. Specifically, the source device provides the video data to the destination device via a computer-readable medium. The source device and the destination device may include any of a variety of devices, including desktop computers, notebook (i.e., laptop) computers, tablet computers, set-top boxes, cellular phones (such as so-called "smart" phones), so-called "smart" pads, televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, etc. In some cases, the source device and the destination device may be equipped for wireless communication.

[0211] The destination device may receive the encoded video data to be decoded via a computer-readable medium. The computer-readable medium may include any type of medium or device capable of moving the encoded video data from the source device to the destination device. In one example, the computer-readable medium may include a communication medium that enables the source device to send the encoded video data directly and in real time to the destination device. The encoded video data may be modulated according to a communication standard such as a wireless communication protocol and sent to the destination device. The communication medium may include any wireless or wired communication medium, such as the radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium may form part of a packet-based network, such as a local area network, a wide area network, or a global network such as the Internet. The communication medium may include routers, switches, base stations, or any other device that may be useful for facilitating communication from the source device to the destination device.

[0212] In some examples, the encoded data can be output from the output interface to a storage device. Similarly, the encoded data can be accessed from a storage device via the input interface. The storage device can include any of a variety of distributed or local access data storage media, such as a hard disk drive, a Blu-ray disc, a DVD, a CD-ROM, a flash memory, a volatile or non-volatile memory, or any other suitable digital storage media for storing the encoded video data. In further examples, the storage device can correspond to a file server or another intermediate storage device that can store the encoded video generated by the source device. The destination device can access the stored video data from the storage device via streaming or downloading. The file server can be any type of server capable of storing the encoded video data and sending the encoded video data to the destination device. Example file servers include a web server (e.g., for a website), an FTP server, a network attached storage (NAS) device, or a local disk drive. The destination device can access the encoded video data via any standard data connection, including an Internet connection. This can include a wireless channel (e.g., a Wi-Fi connection), a wired connection (e.g., DSL, cable modem, etc.), or a combination of both suitable for accessing the encoded video data stored on the file server. The transmission of the encoded video data from the storage device can be a streaming transmission, a download transmission, or a combination thereof.

[0213] The techniques of the present disclosure are not necessarily limited to wireless applications or setups. The techniques can be applied to video decoding to support any of a variety of multimedia applications, such as over-the-air television broadcasting, cable television transmission, satellite television transmission, Internet streaming video transmission (such as Dynamic Adaptive Streaming over HTTP (DASH)), digital video encoded onto a data storage medium, decoding of digital video stored on a data storage medium, or other applications. In some examples, the system can be configured to support one-way or two-way video transmission to support applications such as video streaming, video playback, video broadcasting, and / or video telephony.

[0214] In one example, the source device includes a video source, a video encoder, and an output interface. The destination device can include an input interface, a video decoder, and a display device. The video encoder of the source device can be configured to apply the techniques disclosed herein. In other examples, the source device and the destination device can include other components or arrangements. For example, the source device can receive video data from an external video source, such as an external camera. Similarly, the destination device can interface with an external display device instead of including an integrated display device.

[0215] The example systems above are merely examples. Techniques for processing video data in parallel can be performed by any digital video encoding and / or decoding device. Although generally the techniques of the present disclosure are performed by a video encoding device, the techniques can also be performed by a video encoder / decoder commonly referred to as a "CODEC". Additionally, the techniques of the present disclosure can also be performed by a video pre-processor. The source device and the destination device are merely examples of such encoding devices in which the source device generates encoded video data for transmission to the destination device. In some examples, the source device and the destination device can operate in a substantially symmetric manner such that each of the devices includes video encoding and decoding components. Thus, the example systems can support one-way or two-way video transmission between video devices, e.g., for video streaming, video playback, video broadcast, or video telephony.

[0216] The video source can include a video capture device such as a camera, a video archive containing previously captured video, and / or a video feed interface for receiving video from a video content provider. As a further alternative, the video source can generate computer graphics-based data as the source video, or generate a combination of live video, archived video, and computer-generated video. In some cases, if the video source is a camera, the source device and the destination device can form a so-called videophone or video telephone. However, as described above, the techniques described in the present disclosure can generally be applied to video encoding and can be applied to wireless and / or wired applications. In each case, the captured, pre-captured, or computer-generated video can be encoded by a video encoder. Then, the encoded video information can be output onto a computer-readable medium via an output interface.

[0217] As mentioned, the computer-readable medium can include transient media such as wireless broadcast or wired network transmissions, or storage media (i.e., non-transitory storage media) such as hard disks, flash drives, compact discs, digital versatile discs, Blu-ray discs, or other computer-readable media. In some examples, a network server (not shown) can receive the encoded video data from the source device via a network transmission, for example, and provide the encoded video data to the destination device. Similarly, a computing device of a media production facility such as a compact disc stamping facility can receive the encoded video data from the source device and manufacture a compact disc containing the encoded video data. Thus, in various examples, the computer-readable medium can be understood to include one or more computer-readable media in various forms.

[0218] The input interface of the destination device receives information from a computer-readable medium. The information of the computer-readable medium may include syntax information defined by a video encoder (which is also used by a video decoder), and the syntax information includes syntax elements that describe the characteristics and / or processing of blocks and other decoding units (e.g., group of pictures (GOP)). The display device displays the decoded video data to the user, and may include any one of various display devices, such as a cathode ray tube (CRT), a liquid crystal display (LCD), a plasma display, an organic light-emitting diode (OLED) display, or another type of display device. Various examples of the present application have been described. In Figure 6 and 7 the specific details of the encoding device 104 and the decoding device 112 are shown respectively.

[0219] The specific details of the encoding device 104 and the decoding device 112 are shown in Figure 6 and Figure 7 respectively. Figure 6 is a block diagram showing an example encoding device 104 that can implement one or more of the techniques described in the present disclosure. The encoding device 104 may generate, for example, the syntax structures described herein (e.g., the syntax structures of VPS, SPS, PPS, or other syntax elements). The encoding device 104 may perform intra prediction and inter prediction decoding of video blocks within a video. As previously described, intra decoding depends at least in part on spatial prediction to reduce or remove spatial redundancy within a given video frame or picture. Inter decoding depends at least in part on temporal prediction to reduce or remove temporal redundancy within adjacent or surrounding frames of a video sequence. The intra mode (I mode) may refer to any one of several spatial-based compression modes. Inter modes such as unidirectional prediction (P mode) or bidirectional prediction (B mode) may refer to any one of several time-based compression modes.

[0220] The encoding device 104 includes a splitting unit 35, a prediction processing unit 41, a filter unit 63, a picture memory 64, an adder 50, a transform processing unit 52, a quantization unit 54, and an entropy encoding unit 56. The prediction processing unit 41 includes a motion estimation unit 42, a motion compensation unit 44, and an intra prediction processing unit 46. For video block reconstruction, the encoding device 104 further includes an inverse quantization unit 58, an inverse transform processing unit 60, and an adder 62. The filter unit 63 is intended to represent one or more loop filters, such as a deblocking filter, an adaptive loop filter (ALF), and a sample adaptive offset (SAO) filter. Although in Figure 6 The middle filter unit 63 is shown as a filter in the loop, but in other configurations, the filter unit 63 can be implemented as a post-loop filter. The post-processing device 57 can perform additional processing on the encoded video data generated by the encoding device 104. In some instances, the techniques of the present disclosure can be implemented by the encoding device 104. However, in other instances, one or more of the techniques of the present disclosure can be implemented by the post-processing device 57.

[0221] As Figure 6 shown, the encoding device 104 receives video data, and the segmentation unit 35 segments the data into video blocks. The segmentation can also include, for example, segmentation into slices, slice segments, tiles, or other larger units according to the quadtree structure of the LCU and CU, as well as video block segmentation. The encoding device 104 generally shows components for encoding video blocks within a video slice to be encoded. A slice can be divided into multiple video blocks (and possibly into sets of video blocks called tiles). The prediction processing unit 41 can select one decoding mode from multiple possible decoding modes for the current video block based on error results (e.g., coding rate and distortion level, etc.), such as one intra-prediction decoding mode from multiple intra-prediction decoding modes or one inter-prediction decoding mode from multiple inter-prediction decoding modes. The prediction processing unit 41 can provide the resulting intra- or inter-coded block to the adder 50 to generate residual block data, and to the adder 62 to reconstruct the encoded block for use as a reference picture.

[0222] The intra-prediction processing unit 46 within the prediction processing unit 41 can perform intra-prediction decoding of the current video block relative to one or more adjacent blocks in the same frame or slice as the current video block to be decoded, to provide spatial compression. The motion estimation unit 42 and the motion compensation unit 44 within the prediction processing unit 41 perform inter-prediction decoding of the current video block relative to one or more prediction blocks in one or more reference pictures, to provide temporal compression.

[0223] The motion estimation unit 42 can be configured to determine an inter-prediction mode for a video slice according to a predetermined pattern for the video sequence. The predetermined pattern can specify video slices in the sequence as P slices, B slices, or GPB slices. The motion estimation unit 42 and the motion compensation unit 44 can be highly integrated, but are shown separately for conceptual purposes. The motion estimation performed by the motion estimation unit 42 is a process of generating a motion vector that estimates the motion for a video block. The motion vector can, for example, indicate the displacement of a prediction unit (PU) of a video block within the current video frame or picture relative to a prediction block within a reference picture.

[0224] A prediction block is a block that is found to closely match the PU of the video block to be decoded in terms of pixel differences, which can be determined by the sum of absolute differences (SAD), the sum of squared differences (SSD), or other difference metrics. In some examples, the encoding device 104 can calculate values for sub-integer pixel positions of the reference pictures stored in the picture memory 64. For example, the encoding device 104 can interpolate values for quarter-pixel positions, eighth-pixel positions, or other fractional pixel positions of the reference pictures. Thus, the motion estimation unit 42 can perform motion search with respect to full pixel positions and fractional pixel positions and output motion vectors with fractional pixel accuracy.

[0225] The motion estimation unit 42 calculates a motion vector for the PU by comparing the position of the PU of the video block in the inter-frame decoded slice with the position of the prediction block of the reference picture. The reference picture can be selected from the first reference picture list (list 0) or the second reference picture list (list 1), and each of these two reference picture lists identifies one or more reference pictures stored in the picture memory 64. The motion estimation unit 42 sends the calculated motion vector to the entropy encoding unit 56 and the motion compensation unit 44.

[0226] The motion compensation performed by the motion compensation unit 44 can involve extracting or generating a prediction block based on the motion vector determined by motion estimation, possibly performing interpolation with sub-pixel accuracy. When receiving the motion vector for the PU of the current video block, the motion compensation unit 44 can locate the prediction block pointed to by the motion vector in the reference picture list. The encoding device 104 forms a residual video block by subtracting the pixel values of the prediction block from the pixel values of the currently decoded video block to form pixel differences. The pixel differences form the residual data for the block and can include both a luminance difference component and a chrominance difference component. The adder 50 represents the component or components that perform this subtraction operation. The motion compensation unit 44 can also generate syntax elements associated with the video block and the video slice for use by the decoding device 112 when decoding the video block of the video slice.

[0227] As described above, the intra prediction processing unit 46 may perform intra prediction on the current block as an alternative to the inter prediction performed by the motion estimation unit 42 and the motion compensation unit 44. Specifically, the intra prediction processing unit 46 may determine an intra prediction mode to be used for encoding the current block. In some examples, the intra prediction processing unit 46 may use various intra prediction modes to encode the current block (e.g., during a separate encoding pass), and the intra prediction processing unit 46 may select a suitable intra prediction mode to use from the tested modes. For example, the intra prediction processing unit 46 may calculate rate-distortion values using rate-distortion analysis for various tested intra prediction modes, and may select an intra prediction mode having the best rate-distortion characteristics among the tested modes. Rate-distortion analysis generally determines the amount of distortion (or error) between the encoded block and the original unencoded block that was encoded to produce the encoded block, and the bit rate (i.e., the number of bits) used to produce the encoded block. The intra prediction processing unit 46 may calculate a ratio based on the distortion and rate for various encoded blocks to determine which intra prediction mode exhibits the best rate-distortion value for the block.

[0228] In any case, after selecting an intra prediction mode for the block, the intra prediction processing unit 46 may provide information indicating the selected intra prediction mode for the block to the entropy coding unit 56. The entropy coding unit 56 may encode the information used to indicate the selected intra prediction mode. The encoding device 104 may include, in the transmitted bitstream configuration data, a definition of the coding context for various blocks, and an indication of the most likely intra prediction mode, the intra prediction mode index table, and the modified intra prediction mode index table to be used for each context in the context. The bitstream configuration data may include multiple intra prediction mode index tables and multiple modified intra prediction mode index tables (also referred to as codeword mapping tables).

[0229] After the prediction processing unit 41 generates a prediction block for the current video block via inter prediction or intra prediction, the encoding device 104 forms a residual video block by subtracting the prediction block from the current video block. The residual video data in the residual block may be included in one or more TUs and applied to the transform processing unit 52. The transform processing unit 52 uses a transform (such as a discrete cosine transform (DCT) or a conceptually similar transform) to transform the residual video data into residual transform coefficients. The transform processing unit 52 may convert the residual video data from the pixel domain to the transform domain (such as the frequency domain).

[0230] The transform processing unit 52 may send the obtained transform coefficients to the quantization unit 54. The quantization unit 54 quantizes the transform coefficients to further reduce the bit rate. The quantization process may reduce the bit depth associated with some or all of the coefficients in the coefficients. The degree of quantization may be modified by adjusting quantization parameters. In some examples, the quantization unit 54 may then perform a scan of the matrix including the quantized transform coefficients. Alternatively, the entropy coding unit 56 may perform the scan.

[0231] After quantization, the entropy coding unit 56 performs entropy coding on the quantized transform coefficients. For example, the entropy coding unit 56 may perform context - adaptive variable - length coding (CAVLC), context - adaptive binary arithmetic coding (CABAC), syntax - based context - adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or another entropy coding technique. After being entropy - coded by the entropy coding unit 56, the encoded bitstream may be sent to the decoding device 112, or archived for later transmission or retrieved by the decoding device 112. The entropy coding unit 56 may also perform entropy coding on the motion vectors and other syntax elements for the current video slice being decoded.

[0232] The inverse quantization unit 58 and the inverse transform processing unit 60 respectively apply inverse quantization and inverse transform to reconstruct the residual block in the pixel domain for later use as a reference block in a reference picture. The motion compensation unit 44 may calculate a reference block by adding the residual block to a predicted block of one of the reference pictures in the reference picture list. The motion compensation unit 44 may also apply one or more interpolation filters to the reconstructed residual block to calculate sub - integer pixel values for use in motion estimation. The adder 62 adds the reconstructed residual block to the motion - compensated predicted block produced by the motion compensation unit 44 to produce a reference block for storage in the picture memory 64. The reference block may be used by the motion estimation unit 42 and the motion compensation unit 44 as a reference block for inter - frame prediction of blocks in subsequent video frames or pictures.

[0233] In this way, Figure 6 the encoding device 104 represents an example of a video encoder configured to perform one or more of the techniques described herein (including the processes described above with respect to Figure 4 and Figure 5 ). In some cases, some of the techniques in the present disclosure may also be implemented by a post - processing device 57.

[0234] Figure 7 FIG. 0 is a block diagram showing an example decoding device 112. The decoding device 112 includes an entropy decoding unit 80, a prediction processing unit 81, an inverse quantization unit 86, an inverse transform processing unit 88, an adder 90, a filter unit 91, and a picture memory 92. The prediction processing unit 81 includes a motion compensation unit 82 and an intra prediction processing unit 84. In some examples, the decoding device 112 may perform a decoding path that is generally inverse to the encoding path described with respect to the encoding device 104 from Figure 6 the encoding device 104.

[0235] During the decoding process, the decoding device 112 receives an encoded video bitstream transmitted by the encoding device 104. The encoded video bitstream represents video blocks of an encoded video slice and associated syntax elements. In some examples, the decoding device 112 may receive the encoded video bitstream from the encoding device 104. In some examples, the decoding device 112 may receive the encoded video bitstream from a network entity 79, such as a server, a media-aware network element (MANE), a video editor / splicer, or other such device configured to implement one or more of the techniques described above. The network entity 79 may or may not include the encoding device 104. Some of the techniques described in this disclosure may be implemented by the network entity 79 before the network entity 79 transmits the encoded video bitstream to the decoding device 112. In some video decoding systems, the network entity 79 and the decoding device 112 may be parts of separate devices, while in other instances, the functions described with respect to the network entity 79 may be performed by the same device that includes the decoding device 112.

[0236] The entropy decoding unit 80 of the decoding device 112 performs entropy decoding on the bitstream to generate quantization coefficients, motion vectors, and other syntax elements. The entropy decoding unit 80 forwards the motion vectors and other syntax elements to the prediction processing unit 81. The decoding device 112 may receive syntax elements at the video slice level and / or the video block level. The entropy decoding unit 80 may process and parse both fixed-length syntax elements and variable-length syntax elements in more parameter sets such as VPS, SPS, and PPS.

[0237] When a video slice is encoded as a slice encoded intra (I), the intra prediction processing unit 84 of the prediction processing unit 81 may generate prediction data for a video block of the current video slice based on the intra prediction mode signaled and data from previously decoded blocks of the current frame or picture. When a video frame is decoded as a slice encoded inter (i.e., B, P, or GPB), the motion compensation unit 82 of the prediction processing unit 81 generates a prediction block for a video block of the current video slice based on the motion vector received from the entropy decoding unit 80 and other syntax elements. The prediction block may be generated from one of the reference pictures in the reference picture list within the reference pictures. The decoding device 112 may construct the reference frame lists, list 0 and list 1, using a construction technique based on the reference pictures stored in the picture memory 92.

[0238] The motion compensation unit 82 determines the prediction information for a video block of the current video slice by parsing the motion vector and other syntax elements, and uses the prediction information to generate a prediction block for the current video block being decoded. For example, the motion compensation unit 82 may use one or more syntax elements in the parameter set to determine the prediction mode (e.g., intra or inter prediction) for decoding the video block of the video slice, the inter prediction slice type (e.g., B slice, P slice, or GPB slice), the construction information for one or more reference picture lists for the slice, the motion vector for each inter-coded video block of the slice, the inter prediction state for each inter-decoded video block of the slice, and other information for decoding the video block in the current video slice.

[0239] The motion compensation unit 82 may also perform interpolation based on an interpolation filter. The motion compensation unit 82 may use the interpolation filter as used by the encoding device 104 during the encoding of the video block to calculate the interpolated values for pixels below the integer for the reference block. In this case, the motion compensation unit 82 may determine the interpolation filter used by the encoding device 104 based on the received syntax elements, and may use the interpolation filter to generate the prediction block.

[0240] The inverse quantization unit 86 inverse quantizes or dequantizes the quantized transform coefficients provided in the bitstream and decoded by the entropy decoding unit 80. The inverse quantization process may include using the quantization parameter calculated by the encoding device 104 for each video block in the video slice to determine the degree of quantization, and similarly determining the degree of inverse quantization to be applied. The inverse transform processing unit 88 applies an inverse transform (e.g., inverse DCT or other suitable inverse transform), inverse integer transform, or conceptually similar inverse transform process to the transform coefficients to generate a residual block in the pixel domain.

[0241] After the motion compensation unit 82 generates a prediction block for the current video block based on the motion vector and other syntax elements, the decoding device 112 forms a decoded video block by adding the residual block from the inverse transform processing unit 88 to the corresponding prediction block generated by the motion compensation unit 82. The adder 90 represents the component or components that perform this summation operation. If desired, a loop filter (either in the encoding loop or after the encoding loop) can also be used to smooth pixel transitions or otherwise improve video quality. The filter unit 91 is intended to represent one or more loop filters, such as a deblocking filter, an adaptive loop filter (ALF), and a sample adaptive offset (SAO) filter. Although the filter unit 91 is shown as a filter in the loop in Figure 7 in other configurations, the filter unit 91 can be implemented as a post-loop filter. The decoded video blocks in a given frame or picture are then stored in the picture memory 92, which stores reference pictures for subsequent motion compensation. The picture memory 92 also stores the decoded video for later presentation on a display device (such as the video destination device 122 shown in Figure 1 ).

[0242] In this way, Figure 7 the decoding device 112 of Figure 3 and Figure 4 is an example of a video decoder configured to perform one or more of the techniques described herein (including the processes described above with respect to

[0243] The filter unit 91 filters the reconstructed block (e.g., the output of the adder 90), and stores the filtered reconstructed block in the DPB 94 for use as a reference block and / or outputs the filtered reconstructed block (the decoded video). The reference block can be used by the motion compensation unit 82 as a reference block for performing inter prediction on blocks in subsequent video frames or pictures. The filter unit 91 can perform any type of filtering, such as deblocking filtering, SAO filtering, peak SAO filtering, ALF and / or GALF, and / or other types of loop filters. The deblocking filter can, for example, apply deblocking filtering to filter block boundaries to remove blocking artifacts from the reconstructed video. The peak SAO filter can apply an offset to the reconstructed pixel values in order to improve the overall decoding quality. Additional loop filters (either in the loop or after the loop) can also be used.

[0244] Additionally, filter unit 91 may be configured to perform any of the techniques related to adaptive loop filtering in the present disclosure. For example, as described above, filter unit 91 may be configured to determine parameters for filtering a current block based on parameters for filtering a previous block included in the same APS as the current block, a different APS, or a predefined filter.

[0245] As used herein, the term "computer-readable medium" includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other media capable of storing, containing, or carrying instructions and / or data. A computer-readable medium may include non-transitory media in which data can be stored and which does not include carrier waves and / or transient electronic signals propagating wirelessly or over a wired connection. Examples of non-transitory media may include, but are not limited to, magnetic disks or tapes, optical storage media such as compact discs (CDs) or digital versatile discs (DVDs), flash memory, memory, or memory devices. A computer-readable medium may have code and / or machine-executable instructions stored thereon, and the code and / or machine-executable instructions may represent a process, function, subroutine, program, routine, subroutine, module, software package, class, or any combination of instructions, data structures, or program statements. Code segments may be coupled to another code segment or hardware circuit by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc. may be passed, forwarded, or transmitted via any suitable means including memory sharing, message passing, token passing, network transmission, etc.

[0246] In some examples, a computer-readable storage device, medium, and memory may include a cable or wireless signal that includes a bitstream, etc. However, when mentioned, non-transitory computer-readable storage media explicitly exclude media such as energy, carrier signals, electromagnetic waves, and signals themselves.

[0247] Specific details are provided in the above description to provide a thorough understanding of the examples provided herein. However, one of ordinary skill in the art will understand that the examples may be practiced without these specific details. For clarity, in some instances, the techniques herein may be presented as separate functional blocks including functional blocks that include devices, device components, steps or routines in methods embodied in software, or combinations of hardware and software. In addition to the components shown in the figures and / or described herein, additional components may be used. For example, circuits, systems, networks, processes, and other components may be shown as components in block diagram form so as not to obscure the examples with unnecessary detail. In other instances, well-known circuits, processes, algorithms, structures, and techniques are shown without unnecessary detail so as not to obscure the examples.

[0248] The above may describe individual examples as processes or methods, which are depicted as flowcharts, process schematics, data flow diagrams, structure diagrams, or block diagrams. Although a flowchart may describe operations as a sequential process, many of the operations may be performed in parallel or simultaneously. In addition, the order of the operations may be rearranged. A process terminates when its operations are complete, but may have additional steps not included in the figure. A process may correspond to a method, function, procedure, subroutine, subprogram, etc. When a process corresponds to a function, its termination may correspond to the function returning to the calling function or the main function.

[0249] The processes and methods according to the examples described above may be implemented using computer-executable instructions that are stored in a computer-readable medium or otherwise obtainable from a computer-readable medium. Such instructions may include, for example, instructions or data that cause or otherwise configure a general-purpose computer, a special-purpose computer, or a processing device to perform a certain function or group of functions. Part of the computer resources used may be accessible via a network. The computer-executable instructions may be, for example, binary, intermediate format instructions such as assembly language, firmware, source code, etc. Examples of computer-readable media that may be used to store instructions, the information used, and / or the information created during the methods according to the described examples include magnetic or optical disks, flash memory, USB devices with non-volatile memory, storage devices on a network, etc.

[0250] Devices implementing the processes and methods according to these disclosures can include hardware, software, firmware, middleware, microcode, hardware description language, or any combination thereof, and can be in any of a variety of form factors. When implemented in software, firmware, middleware, or microcode, the program code or code segments (e.g., computer program product) for performing the necessary tasks can be stored in a computer-readable or machine-readable medium. The processor can perform the necessary tasks. Typical examples of form factors include laptop computers, smart phones, mobile phones, tablet devices, or other small form factor personal computers, personal digital assistants, rack-mounted devices, stand-alone devices, etc. The functions described herein can also be embodied in peripheral devices or plug-in cards. By way of further example, such functions can also be implemented on a circuit board among different chips or different processes executed in a single device.

[0251] Instructions, the media for conveying such instructions, the computing resources for executing them, and other structures for supporting such computing resources are example units for providing the functions described in this disclosure.

[0252] In the foregoing description, aspects of the present application are described with reference to specific examples of the present application, but those skilled in the art will recognize that the present application is not limited thereto. Thus, although illustrative examples of the present application have been described in detail herein, it should be understood that the inventive concept can be embodied and employed differently in other ways, and the appended claims are intended to be construed to include such variations, except those variations limited by the prior art. The various features and aspects of the applications described above can be used individually or jointly. Further, without departing from the broader spirit and scope of this specification, the examples can be utilized in any number of environments and applications other than those described herein. Therefore, the specification and drawings are considered illustrative rather than restrictive. For purposes of illustration, the methods are described in a particular order. It should be understood that in alternative examples, the methods can be performed in an order different from the order described.

[0253] Those of ordinary skill in the art will understand that, without departing from the scope of this specification, the less than ("<") and greater than (">") symbols or terms used herein can be replaced with the less than or equal to ("≤") and greater than or equal to ("≥") symbols, respectively.

[0254] In cases where a component is described as "configured to" perform certain operations, such configuration can be accomplished, for example, by designing an electronic circuit or other hardware to perform the operations, by programming a programmable electronic circuit (e.g., a microprocessor or other suitable circuit) to perform the operations, or any combination thereof.

[0255] The phrase "coupled to" refers to any component that is physically connected, directly or indirectly, to another component, and / or any component that communicates, directly or indirectly, with another component (e.g., connected to another component via a wired or wireless connection and / or other suitable communication interface).

[0256] Claim language or other language that recites "at least one of" in a recited set and / or "one or more of" in a set indicates that one member of the set or multiple members of the set (in any combination) satisfy the claim. For example, claim language that recites "at least one of A and B" means A, B, or A and B. In another example, claim language that recites "at least one of A, B, and C" means A, B, C, or A and B, or A and C, or B and C, or A and B and C. The language "at least one of" in a set and / or "one or more of" in a set does not limit the set to the items listed in the set. For example, claim language that recites "at least one of A and B" can mean A, B, or A and B, and can additionally include items not listed in the set of A and B.

[0257] The various illustrative logical blocks, modules, circuits, and algorithmic steps described in connection with the examples disclosed herein can be implemented as electronic hardware, computer software, firmware, or combinations thereof. To clearly illustrate this interchangeability of hardware and software, the various illustrative components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. Skilled artisans may implement the described functionality in a variable manner for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present application.

[0258] The techniques described herein may also be implemented using electronic hardware, computer software, firmware, or any combination thereof. Such techniques may be implemented in any of a variety of devices, such as a general-purpose computer, a wireless communication device handset, or an integrated circuit device having multiple uses, including applications in wireless communication device handsets and other devices. Any features described as modules or components may be implemented together in an integrated logic device or separately as discrete but interoperable logic devices. If implemented in software, the techniques may be at least partially implemented by a computer-readable data storage medium that includes program code, the program code including instructions that, when executed, implement one or more of the methods described above. The computer-readable data storage medium may form part of a computer program product, which may include packaging materials. The computer-readable medium may include a memory or data storage medium, such as random access memory (RAM) (e.g., synchronous dynamic random access memory (SDRAM)), read-only memory (ROM), non-volatile random access memory (NVRAM), electrically erasable programmable read-only memory (EEPROM), flash memory, magnetic or optical data storage media, and the like. Additionally or alternatively, the techniques may be at least partially implemented by a computer-readable communication medium (e.g., a propagated signal or wave) that carries or transports program code in the form of instructions or data structures and that can be accessed, read, and / or executed by a computer.

[0259] The program code may be executed by a processor, which may include one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other equivalent integrated or discrete logic circuitry. Such a processor may be configured to perform any of the techniques described in this disclosure. A general-purpose processor may be a microprocessor, but in the alternative, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration. Thus, as used herein, the term "processor" may refer to any of the foregoing structures, any combination of the foregoing structures, or any other structure or device suitable for implementing the techniques described herein. Additionally, in some aspects, the functions described herein may be provided within a dedicated software module or hardware module configured for encoding and decoding, or incorporated into a combined video encoder-decoder (CODEC).

[0260] Exemplary embodiments of this disclosure include:

[0261] Example 1. A method for processing video data, the method comprising: obtaining video data including one or more pictures; determining to disable weighted prediction for at least a portion of the one or more pictures; and in response to determining to disable weighted prediction for at least a portion of the one or more pictures, determining that it is not allowed for more than one reference picture to have the same picture order count (POC) value.

[0262] Example 2. The method according to Example 1, wherein at least a portion of the one or more pictures includes slices of a picture.

[0263] Example 3. The method according to any one of Examples 1 to 2, further comprising: obtaining a weighted prediction flag; and determining to disable weighted prediction for the one or more pictures based on the weighted prediction flag.

[0264] Example 4. The method according to Example 3, wherein the weighted prediction flag indicates whether weighted prediction is applied to a slice of unidirectional prediction (P slice).

[0265] Example 5. The method according to Example 3, wherein the weighted prediction flag indicates whether weighted prediction is applied to a slice of bidirectional prediction (B slice).

[0266] Example 6. An apparatus, comprising: a memory configured to store video data; and a processor configured to process the video data according to any one of Examples 1 to 5.

[0267] Example 7. The apparatus according to Example 6, wherein the apparatus includes an encoder.

[0268] Example 8. The apparatus according to Example 6, wherein the apparatus includes a decoder.

[0269] Example 9. The apparatus according to any one of Examples 6 to 8, wherein the apparatus is a mobile device.

[0270] Example 10. The apparatus according to any one of Examples 6 to 9, further comprising: a display configured to display the video data.

[0271] Example 11. The apparatus according to any one of Examples 6 to 10, further comprising: a camera configured to capture one or more pictures.

[0272] Example 12. A computer-readable medium having instructions stored thereon, the instructions when executed by a processor implement the method according to any one of Examples 1 to 5.

[0273] Example 13. A method for processing video data, the method comprising: obtaining video data including a plurality of reference pictures in a plurality of layers; and generating a reference picture structure (RPS), the RPS including a picture order count (POC) value and a layer identifier (ID) for at least one of the plurality of reference pictures, wherein the layer ID for a reference picture identifies the layer of the reference picture.

[0274] Example 14. The method according to Example 13, further comprising: generating the RPS to include the layer ID for each of the plurality of reference pictures.

[0275] Example 15. The method according to Example 13, further comprising: when the POC value of a reference picture is equal to the POC value of the current picture, generating the RPS to include the layer ID for the reference picture.

[0276] Example 16. An apparatus, comprising: a memory configured to store video data; and a processor configured to process the video data according to any one of Examples 13 to 15.

[0277] Example 17. The apparatus according to Example 16, wherein the apparatus includes an encoder.

[0278] Example 18. The apparatus according to any one of Examples 16 to 17, wherein the apparatus is a mobile device.

[0279] Example 19. The apparatus according to any one of Examples 16 to 18, further comprising: a display configured to display the video data.

[0280] Example 20. The apparatus according to any one of Examples 16 to 19, further comprising: a camera configured to capture one or more pictures.

[0281] Example 21. The apparatus according to any one of Examples 16 to 20, further comprising: a communication circuit configured to transmit the processed video data.

[0282] Example 22. A computer-readable medium having instructions stored thereon, the instructions when executed by a processor implement the method according to any one of Examples 13 to 15.

[0283] Example 23. A method for processing video data, the method comprising: obtaining video data including a plurality of reference pictures in a plurality of layers; and processing a reference picture structure (RPS) according to the video data, the RPS including a picture order count (POC) value and a layer identifier (ID) for at least one of the plurality of reference pictures, wherein the layer ID for a reference picture identifies the layer of the reference picture.

[0284] Example 24. The method according to Example 23, wherein the RPS includes layer IDs for each of a plurality of reference pictures.

[0285] Example 25. The method according to Example 23, wherein when the POC value of a reference picture is equal to the POC value of the current picture, the RPS includes the layer ID for the reference picture.

[0286] Example 26. An apparatus, comprising: a memory configured to store video data; and a processor configured to process the video data according to any one of Examples 23 to 25.

[0287] Example 27. The apparatus according to Example 26, wherein the apparatus includes a decoder.

[0288] Example 28. The apparatus according to any one of Examples 26 to 27, wherein the apparatus is a mobile device.

[0289] Example 29. The apparatus according to any one of Examples 26 to 28, further comprising: a display configured to display the video data.

[0290] Example 30. The apparatus according to any one of Examples 26 to 29, further comprising: a camera configured to capture one or more pictures.

[0291] Example 31. A computer-readable medium having instructions stored thereon that, when executed by a processor, implement the method according to any one of Examples 23 to 25.

[0292] Example 32. A method of processing video data, the method comprising: obtaining video data including a plurality of pictures in a plurality of layers, the plurality of layers corresponding to a plurality of representations of video content; determining that inter-layer prediction is enabled for the plurality of pictures; and inferring a zero motion vector based on determining that inter-layer prediction is enabled for the plurality of pictures, the zero motion vector indicating that there is no content displacement among the plurality of pictures of the plurality of representations.

[0293] Example 33. The method according to Example 33, wherein when motion information is not used for inter-layer prediction, the motion information is signaled in an encoded video bitstream including a plurality of pictures.

[0294] Example 34. An apparatus, comprising: a memory configured to store video data; and a processor configured to process the video data according to any one of Examples 32 to 33.

[0295] Example 35. The apparatus according to Example 34, wherein the apparatus includes an encoder.

[0296] Example 36. The apparatus according to Example 34, wherein the apparatus includes a decoder.

[0297] Example 37. The apparatus according to any one of Examples 34 to 36, wherein the apparatus is a mobile device.

[0298] Example 38. The apparatus according to any one of Examples 34 to 37, further comprising: a display configured to display video data.

[0299] Example 39. The apparatus according to any one of Examples 34 to 38, further comprising: a camera configured to capture one or more pictures.

[0300] Example 40. A computer-readable medium having instructions stored thereon, the instructions when executed by a processor implement the method according to any one of Examples 32 to 33.

[0301] Example 41. A method of processing video data, the method comprising: obtaining video data; and generating an encoded video bitstream according to the video data, the encoded video bitstream including one or more syntax elements and a plurality of pictures in a plurality of layers, the plurality of layers corresponding to a plurality of representations of video content, wherein the number of entries in the inter-layer reference picture entry syntax element is limited to not exceed the maximum number of reference pictures minus the number of inter-frame prediction reference pictures.

[0302] Example 42. The method according to Example 41, wherein the inter-layer reference picture entry syntax element specifies the number of direct inter-layer reference picture entries in the reference picture list syntax structure.

[0303] Example 43. An apparatus comprising: a memory configured to store video data; and a processor configured to process video data according to any one of Examples 41 to 42.

[0304] Example 44. The apparatus according to Example 43, wherein the apparatus includes an encoder.

[0305] Example 45. The apparatus according to Example 43, wherein the apparatus is a mobile device.

[0306] Example 46. The apparatus according to any one of Examples 43 to 45, further comprising: a display configured to display video data.

[0307] Example 47. The apparatus according to any one of Examples 43 to 46, further comprising: a camera configured to capture one or more pictures.

[0308] Example 48. A computer-readable medium having instructions stored thereon, the instructions when executed by a processor implement the method according to any one of Examples 41 to 42.

[0309] Example 49. A method for processing video data, the method comprising: obtaining an encoded video bitstream, the encoded video bitstream including one or more syntax elements and multiple pictures in multiple layers, the multiple layers corresponding to multiple representations of video content; and processing an inter-layer reference picture entry syntax element according to the encoded video bitstream, wherein the number of entries in the inter-layer reference picture entry syntax element is limited to not exceed the maximum number of reference pictures minus the number of inter-frame prediction reference pictures.

[0310] Example 50. The method according to Example 33, wherein the inter-layer reference picture entry syntax element specifies the number of direct inter-layer reference picture entries in a reference picture list syntax structure.

[0311] Example 51. An apparatus, comprising: a memory configured to store video data; and a processor configured to process video data according to any one of Examples 49 to 50.

[0312] Example 52. The apparatus according to Example 51, wherein the apparatus includes a decoder.

[0313] Example 53. The apparatus according to Example 51, wherein the apparatus is a mobile device.

[0314] Example 54. An example apparatus according to any one of Examples 51 to 53, further comprising: a display configured to display video data.

[0315] Example 55. An example apparatus according to any one of Examples 51 to 54, further comprising: a camera configured to capture one or more pictures.

[0316] Example 56. An example computer-readable medium having instructions stored thereon, the instructions when executed by a processor implementing the method according to any one of Examples 49 to 50.

[0317] Example 57. A method for decoding video data, the method comprising: obtaining at least a portion of a picture from a bitstream; determining, according to the bitstream, that weighted prediction is enabled for a portion of the picture; based on determining that weighted prediction is enabled for a portion of the picture, identifying a zero-valued picture order count offset for indicating a reference picture from a reference picture list; and reconstructing at least a portion of the picture using at least a portion of the reference picture identified by the zero-valued picture order count offset.

[0318] Example 58. The method according to Example 57 further includes: obtaining multiple parts of a picture from a bitstream, the multiple parts including at least a part of the picture reconstructed using a reference picture identified by a zero-valued picture order count offset; identifying multiple corresponding picture order count offsets for the multiple parts of the picture, wherein the zero-valued picture order count offset is the corresponding picture order count offset for the picture among the multiple corresponding picture order count offsets; and reconstructing the picture using multiple reference pictures identified by the multiple corresponding picture order count offsets.

[0319] Example 59. The method according to Example 57 further includes: determining to enable weighted prediction by parsing the bitstream to identify one or more weighted prediction flags for the picture.

[0320] Example 60. The method according to Example 59, wherein the one or more weighted prediction flags for the picture include a sequence parameter set weighted prediction flag and a picture parameter set weighted prediction flag.

[0321] Example 61. The method according to Example 60, wherein the sequence parameter set weighted prediction flag and the picture parameter set weighted prediction flag are flags for unidirectional prediction frames.

[0322] Example 62. The method according to Example 60, wherein the sequence parameter set weighted prediction flag and the picture parameter set weighted prediction flag are flags for bidirectional prediction frames.

[0323] Example 63. The method according to Example 60, wherein the picture parameter set weighted prediction flag is constrained by the sequence parameter set weighted prediction flag.

[0324] Example 64. The method according to Example 57 further includes: obtaining at least a part of a second picture included in the bitstream; determining to disable weighted prediction for the second picture according to the bitstream; based on disabling weighted prediction for the second picture, parsing a syntax element by reconstructing a value of the syntax element indicating a picture order count offset for a second reference picture in the bitstream as the signal-transmitted value plus 1; and reconstructing the second picture using the reconstructed value of the syntax element.

[0325] Example 65. The method according to Example 64, wherein when weighted prediction is disabled, the value of the reference picture syntax element specifies the absolute difference between the picture order count value of the second picture and the previous reference picture entry for the second reference picture in the reference picture list as the value plus one.

[0326] Example 66. The method according to Example 57, wherein a reference picture is associated with a first set of weights, and a picture is associated with a second set of weights different from the first set of weights; wherein the reference picture has a different size from the picture; and wherein the reference picture is signaled as a short-term reference picture having a picture order count least significant bit value equal to zero.

[0327] Example 67. The method according to Example 57, wherein the picture is a non-instantaneous decoding refresh (non-IDR) picture.

[0328] Example 68. The method according to Example 57, wherein at least a portion of the picture is a slice.

[0329] Example 69. The method according to Example 57, further comprising: determining, according to a bitstream, a layer identifier for a reference picture, the layer identifier indicating a layer of a layer index for inter-layer prediction; wherein when the layer identifier is different from a current layer identifier, a zero picture order count offset is identified by inferring a zero value based on that the picture order count for the reference picture is not signaled.

[0330] Example 70. An apparatus comprising: a memory configured to store video data; and a processor configured to process the video data according to any one of Examples 57 to 69.

[0331] Example 71. An apparatus for encoding video data, the apparatus comprising: a memory; and a processor implemented in circuitry and configured to: identify at least a portion of a picture; select weighted prediction to be enabled for a portion of the picture; identify a reference picture for a portion of the picture; generate a zero picture order count offset for indicating a reference picture from a reference picture list; and generate a bitstream including the portion of the picture and the zero picture order count offset associated with the portion of the picture.

[0332] Example 72. A method for encoding video data, the method comprising: identifying at least a portion of a picture; determining that weighted prediction is selected to be enabled for a portion of the picture; identifying a reference picture for a portion of the picture; generating a zero picture order count offset for indicating a reference picture from a reference picture list; and generating a bitstream including the portion of the picture and the zero picture order count offset associated with the portion of the picture.< / insertend> < / insert> < / insertend> < / insert> < / insertend> < / insert> < / insertend> < / insert> < / insertend> < / insert> < / insertend> < / insert> < / insertend> < / insert> < / insertend> < / insert> < / insertend> < / insert> < / insertend> < / insert>

Claims

1. An apparatus for decoding video data, the apparatus comprising: a memory; and a processor, implemented in circuitry and configured to: obtain at least a portion of a first picture included in a bitstream; determine, based on the bitstream, whether weighted prediction is enabled for at least a portion of the first picture; when weighted prediction is enabled for at least a portion of the first picture: determine, based on the bitstream, a picture order count offset associated with a portion of the first picture, wherein a zero value picture order count offset indicates that the same reference picture is being used for prediction multiple times; and reconstruct at least a portion of the first picture using at least a portion of the reference picture identified by the picture order count offset; and when weighted prediction is disabled for at least a portion of the first picture: determine, based on the bitstream, the picture order count offset associated with a portion of the first picture by generating a reconstructed value of a syntax element indicating the picture order count offset as the sum of the value signaled, parsed from the bitstream, of the syntax element and 1; and reconstruct at least a portion of the first picture using the reconstructed value of the syntax element.

2. The device according to claim 1, wherein, The processor is further configured to: obtain multiple portions of the first picture from the bitstream, the multiple portions including at least a portion of the first picture reconstructed using the reference picture identified by the picture order count offset; identify multiple respective picture order count offsets for the multiple portions of the first picture, wherein the picture order count offset identified for a first portion is the respective picture order count offset associated with the reference picture among the multiple respective picture order count offsets; and reconstruct the picture using multiple reference pictures identified by the multiple respective picture order count offsets.

3. The device according to claim 1, wherein The processor is further configured to determine that weighted prediction can be enabled by parsing the bitstream to identify one or more weighted prediction flags for the picture.

4. The apparatus according to claim 3, wherein, The one or more weighted prediction flags for the picture include a sequence parameter set weighted prediction flag and a picture parameter set weighted prediction flag.

5. The apparatus according to claim 4, wherein The sequence parameter set weighted prediction flag and the picture parameter set weighted prediction flag are flags for unidirectional prediction frames.

6. The apparatus according to claim 4, wherein, The sequence parameter set weighted prediction flag and the picture parameter set weighted prediction flag are flags for bidirectional prediction frames.

7. The apparatus according to claim 4, wherein, The picture parameter set weighted prediction flag is constrained by the sequence parameter set weighted prediction flag.

8. The apparatus according to claim 1, wherein When weighted prediction is disabled, the value of a reference picture syntax element specifies as the signaled value plus one the absolute difference between the picture order count value of the second picture and the previous reference picture entry for the second reference picture in a reference picture list.

9. The device according to claim 1, wherein The reference picture is associated with a first weight set, and the picture is associated with a second weight set different from the first weight set; wherein the reference picture has a different size from the picture; and Wherein, the reference picture is signaled as a short-term reference picture having a least significant bit value of the picture order count equal to zero.

10. The device according to claim 1, wherein, The picture is a non-instantaneous decoding refresh (non-IDR) picture.

11. The apparatus according to claim 1, wherein, The processor is configured to: Determine a layer identifier for the reference picture according to the bitstream, the layer identifier indicating a layer of a layer index for inter-layer prediction; And Wherein, when the layer identifier is different from the current layer identifier, the zero picture order count offset is identified by inferring a zero value according to the picture order count for the reference picture not being signaled.

12. The apparatus according to claim 1, wherein The apparatus includes a mobile device having a camera coupled to the processor and the memory for capturing and storing the picture.

13. The apparatus according to claim 1, further comprising a display.

14. A method for decoding video data, the method comprising: Obtain at least a portion of a first picture from a bitstream; Determine that weighted prediction is enabled for a portion of the first picture according to the bitstream; Determine a picture order count offset associated with a portion of the first picture according to the bitstream, wherein a zero picture order count offset indicates that the same reference picture is being used for prediction with different weights; Based on determining that weighted prediction is enabled for a portion of the first picture, reconstruct at least a portion of the first picture using at least a portion of the reference picture identified by the picture order count offset; Obtain at least a portion of a second picture included in the bitstream; Determine that weighted prediction is disabled for a portion of the second picture according to the bitstream; Parse the second syntax element by generating a reconstructed value of the second syntax element indicating a second picture order count offset for a second reference picture in the bitstream as a signaled value plus 1 based on disabling the weighted prediction for a portion of the second picture; and Reconstruct a portion of the second picture using the reconstructed value of the second syntax element.

15. The method according to claim 14, further comprising: Obtain a plurality of portions of the picture from the bitstream, the plurality of portions including at least a portion of the picture reconstructed using the reference picture identified by the zero picture order count offset; Identify a plurality of corresponding picture order count offsets for the plurality of portions of the picture, wherein the zero picture order count offset is the corresponding picture order count offset for the picture among the plurality of corresponding picture order count offsets; And Reconstruct the picture using a plurality of reference pictures identified by the plurality of corresponding picture order count offsets.

16. The method according to claim 14 further comprises: Determine that weighted prediction is enabled by parsing the bitstream to identify one or more weighted prediction flags for the picture.

17. The method according to claim 16, wherein, The one or more weighted prediction flags for the picture include a sequence parameter set weighted prediction flag and a picture parameter set weighted prediction flag.

18. The method according to claim 17, wherein, The sequence parameter set weighted prediction flag and the picture parameter set weighted prediction flag are flags for unidirectional prediction frames.

19. The method according to claim 17, wherein, The weighted prediction flags of the sequence parameter set and the weighted prediction flags of the picture parameter set are flags for bidirectional prediction frames.

20. The method according to claim 17, wherein The weighted prediction flags of the picture parameter set are constrained by the weighted prediction flags of the sequence parameter set.

21. The method according to claim 14, wherein, When weighted prediction is disabled, the value of the reference picture syntax element is specified as the signaled value plus one by the absolute difference between the picture order count value of the second picture and the previous reference picture entry for the second reference picture in the reference picture list.

22. The method according to claim 14, wherein The reference picture is associated with a first set of weights, and the picture is associated with a second set of weights different from the first set of weights; wherein the reference picture has a different size from the picture; and wherein the reference picture is signaled as a short-term reference picture having a picture order count least significant bit value equal to zero.

23. The method according to claim 14, wherein The picture is a non-instantaneous decoding refresh (non-IDR) picture.

24. The method according to claim 14, wherein At least a portion of the picture is a slice.

25. The method according to claim 14, further comprising: Determining, from the bitstream, a layer identifier for the reference picture, the layer identifier indicating the layer of the layer index for inter-layer prediction; and wherein when the layer identifier is different from the current layer identifier, the zero picture order count offset is identified by inferring a zero value based on the picture order count for the reference picture not being signaled.

26. An apparatus for encoding video data, the apparatus comprising: a memory; and a processor implemented in circuitry and configured to: Identify at least a portion of a first picture; Select weighted prediction to be enabled for a portion of the first picture; Identify a reference picture for a portion of the first picture; Based on determining that weighted prediction is enabled for a portion of the first picture, generate a picture order count offset for indicating the reference picture from the reference picture list, wherein a zero picture order count offset indicates that the same reference picture is being used for prediction with different weights; Identify at least a portion of a second picture; Select weighted prediction to be disabled for a portion of the second picture; Identify a second reference picture for a portion of the second picture; Generate a picture order count offset for indicating the second reference picture from the reference picture list; Generate a second syntax element indicating a decrement of the picture order count offset for the second reference picture based on disabling the weighted prediction for the second picture; and Generate a bitstream that includes a portion of the first picture and the picture order count offset associated with the portion of the first picture, a portion of the second picture, and the second syntax element.

27. A method for encoding video data, the method comprising: Identify at least a portion of a first picture; Select weighted prediction to be enabled for a portion of the first picture; Identify a reference picture for a portion of the first picture; Generate a picture order count offset for indicating the reference picture from a reference picture list based on determining that weighted prediction is enabled for a portion of the first picture, wherein a picture order count offset of zero indicates that the same reference picture is being used for prediction with different weights; Identify at least a portion of a second picture; Select weighted prediction to be disabled for a portion of the second picture; Identify a second reference picture for a portion of the second picture; Generate a picture order count offset for indicating the second reference picture from the reference picture list; Generate a second syntax element indicating a decrement by one of the picture order count offset for the second reference picture based on disabling the weighted prediction for the second picture; and Generate a bitstream that includes a portion of the first picture and the picture order count offset associated with the portion of the first picture, a portion of the second picture, and the second syntax element.