Encoder, decoder, and corresponding method

The method enhances video coding efficiency by effectively managing reference pictures in inter-layer prediction within the video coding system, addressing the challenges of resource utilization and image quality.

JP2025089323AActive Publication Date: 2025-06-12HUAWEI TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2025041525
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-05-30
Filing Date
2025-03-14
Publication Date
2025-06-12
Estimated Expiration
2040-05-19

AI Technical Summary

Technical Problem

Existing video coding systems face challenges in efficiently managing reference pictures, particularly in inter-layer prediction, which affects coding efficiency and resource utilization.

Method used

The proposed solution involves a method implemented in a decoder that receives a bitstream containing a current picture and a reference picture list structure with an inter-layer reference picture flag. This method determines if an entry in the reference picture list structure is an inter-layer reference picture entry and decodes the current picture based on the indicated inter-layer reference picture.

Benefits of technology

This approach improves coding efficiency by effectively managing reference pictures in inter-layer prediction, reducing processor, memory, and network resource usage while maintaining image quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025089323000001_ABST
    Figure 2025089323000001_ABST
Patent Text Reader

Abstract

To disclose a video coding mechanism.SOLUTION: The mechanism includes receiving a bit stream having a reference picture list structure having a current picture and an inter-layer reference picture flag. The mechanism determines, on the basis of the inter-layer reference picture flag, that an entry in the reference picture list structure related to the current picture is an inter-layer reference picture (ILRP) entry. The current picture is decoded on the basis of the inter-layer reference picture shown by an entry in the reference picture list structure when the entry is an ILRP entry. The current picture is transferred for display as a part of a decoding video sequence.SELECTED DRAWING: Figure 8
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - Reference to Related Applications This patent application claims the benefit of U.S. Provisional Patent Application No. 62 / 854,827, filed May 30, 2019, entitled "Reference Picture Management In Layered Video Coding", which is incorporated herein by reference.

[0002] The present disclosure generally relates to video coding, and more particularly to reference picture management when employed in inter - layer prediction in video coding.

Background Art

[0003] Even the amount of video data required to depict relatively short videos can be significant, which can pose difficulties when data is to be streamed or otherwise communicated across a communication network with limited bandwidth capacity. Thus, video data is generally compressed before being communicated across modern telecommunications networks. The size of the video can also be a problem when the video is stored on a storage device because memory resources may be limited. Video compression devices often code video data using software and / or hardware at the source prior to transmission or storage, thereby reducing the amount of data required to represent the digital video image. The compressed data is then received at the destination by a video decompression device that decodes the video data. As network resources are limited and the demand for higher video quality constantly increases, improved compression and restoration techniques that improve the compression ratio with little sacrifice in image quality are desirable.

Summary of the Invention

Means for Solving the Problems

[0004] In one embodiment, the present disclosure includes a method implemented in a decoder, the method comprising receiving, by a receiver of the decoder, a bitstream comprising a current picture and a reference picture list structure comprising an inter-layer reference picture flag; determining, by a processor of the decoder based on the inter-layer reference picture flag, that an entry in the reference picture list structure associated with the current picture is an inter-layer reference picture (ILRP) entry; and decoding, by the processor, the current picture based on the inter-layer reference picture indicated by the entry in the reference picture list structure when the entry is an ILRP entry.

[0005] A video coding system may encode pictures according to inter prediction. In inter prediction, a picture is coded with reference to another picture. The picture being coded is called the current picture, and the picture used as a reference is called the reference picture. Some video coding systems track reference pictures by adopting a reference picture list. Some video coding systems adopt scalable video coding. In scalable video coding, a video sequence is coded as a base layer and one or more enhancement layers. In this context, a picture may be divided into portions existing in different layers. For example, a lower layer version of a picture may be of lower quality than a higher layer version of the picture. Further, a lower layer version of a picture may be smaller than a higher layer version of the picture (e.g., may have a smaller width and / or height). When using layers, a picture in the current layer may be coded according to intra prediction (without using a reference picture), according to inter prediction by reference to a reference picture in the same layer, or according to inter-layer prediction by reference to a reference picture in a different layer. Some reference picture management systems may not be configured to manage reference pictures used in inter-layer prediction. This example includes a mechanism for managing a reference picture list when adopting inter-layer prediction. The reference picture list may be configured to include an entry for each reference picture used by the coded video. Then, a flag may be adopted to indicate whether each entry includes an inter prediction reference picture or an inter-layer reference picture. In one example, a group of reference picture flags for a syntax structure including a reference picture list may be indicated as inter_layer_ref_pic_flag[ listIdx ][ rplsIdx ][ i ]. Further, an ILRP layer indicator may be used to indicate which layer includes the indicated inter-layer reference picture.The decoder can then use the reference picture flag and the ILRP layer indicator to select an appropriate inter-layer reference picture to perform inter-layer prediction. Thus, the disclosed mechanism creates additional functionality in the encoder and / or decoder. Further, the disclosed mechanism may improve coding efficiency, which may reduce the usage of processor, memory, and / or network resources in the encoder and / or decoder.

[0006] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the reference picture list structure further comprises an ILRP layer indicator, and the method further comprises the step of determining, by a processor, the layer of the inter-layer reference picture based on the ILRP layer indicator when the entry is an ILRP entry.

[0007] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the reference picture list structure is represented as ref_pic_list_struct( listIdx, rplsIdx ), where listIdx identifies the reference picture list, rplsIdx identifies an entry in the reference picture list, and ref_pic_list_struct is a syntax structure.

[0008] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the inter-layer reference picture flag is indicated as inter_layer_ref_pic_flag[ listIdx ][ rplsIdx ][ i ], and when the i-th entry in ref_pic_list_struct( listIdx, rplsIdx ) is an ILRP entry, inter_layer_ref_pic_flag[ listIdx ][ rplsIdx ][ i ] is equal to 1, and when the i-th entry in ref_pic_list_struct( listIdx, rplsIdx ) is not an ILRP entry, inter_layer_ref_pic_flag[ listIdx ][ rplsIdx ][ i ] is equal to 0.

[0009] Optionally, in any of the foregoing aspects, another implementation of the aspect further provides that when the entry is not an ILRP entry, the current picture is decoded by the processor in accordance with intra prediction based on the intra-layer reference picture indicated by the entry in the reference picture list structure.

[0010] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that ref_pic_list_struct( listIdx, rplsIdx ) and inter_layer_ref_pic_flag[ listIdx ][ rplsIdx ][ i ] are included in the bitstream in the sequence parameter set (SPS: sequence parameter set).

[0011] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the inter-layer reference picture is in the same access unit (AU: access unit) as the current picture and that the inter-layer reference picture is associated with a lower layer identifier than the current picture.

[0012] In one embodiment, the present disclosure includes a method implemented in an encoder, the method comprising: encoding, by a processor of the encoder, a current picture into a bitstream, wherein the current picture is encoded according to inter-layer prediction based on an inter-layer reference picture; encoding, by the processor, a reference picture list structure into the bitstream, wherein the reference picture list structure comprises a plurality of entries for a plurality of reference pictures including an entry related to the current picture and indicates an inter-layer reference picture; encoding, by the processor, an inter-layer reference picture flag into the bitstream, wherein the inter-layer reference picture flag indicates that the entry related to the current picture is an ILRP entry; and storing, by a memory coupled to the processor, the bitstream for communication to a decoder.

[0013] A video coding system may encode pictures according to inter prediction. In inter prediction, a picture is coded with reference to another picture. The picture being coded is called the current picture, and the picture used as a reference is called the reference picture. Some video coding systems track reference pictures by adopting a reference picture list. Some video coding systems adopt scalable video coding. In scalable video coding, a video sequence is coded as a base layer and one or more enhancement layers. In this context, a picture may be divided into portions existing in different layers. For example, a lower layer version of a picture may be of lower quality than a higher layer version of the picture. Further, a lower layer version of a picture may be smaller than a higher layer version of the picture (e.g., may have a smaller width and / or height). When using layers, a picture in the current layer may be coded according to intra prediction (without using a reference picture), according to inter prediction by reference to a reference picture in the same layer, or according to inter-layer prediction by reference to a reference picture in a different layer. Some reference picture management systems may not be configured to manage reference pictures used in inter-layer prediction. This example includes a mechanism for managing a reference picture list when adopting inter-layer prediction. The reference picture list may be configured to include an entry for each reference picture used by the coded video. Then, a flag may be adopted to indicate whether each entry includes an inter prediction reference picture or an inter-layer reference picture. In one example, a group of reference picture flags for a syntax structure including a reference picture list may be indicated as inter_layer_ref_pic_flag[ listIdx ][ rplsIdx ][ i ]. Further, an ILRP layer indicator may be used to indicate which layer includes the indicated inter-layer reference picture.The decoder can then use the reference picture flag and the ILRP layer indicator to select an appropriate inter-layer reference picture to perform inter-layer prediction. Thus, the disclosed mechanism creates additional functionality in the encoder and / or decoder. Further, the disclosed mechanism may improve coding efficiency, which may reduce the usage of processor, memory, and / or network resources in the encoder and / or decoder.

[0014] Optionally, in any of the foregoing aspects, another implementation of the aspect further comprises encoding, by a processor, the ILRP layer indicator into the bitstream, where the ILRP layer indicator indicates the layer of the inter-layer reference picture.

[0015] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the reference picture list structure is represented as ref_pic_list_struct( listIdx, rplsIdx ), where listIdx identifies the reference picture list, rplsIdx identifies an entry in the reference picture list, and ref_pic_list_struct is a syntax structure that returns an entry based on listIdx and rplsIdx.

[0016] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the inter-layer reference picture flag is indicated as inter_layer_ref_pic_flag[ listIdx ][ rplsIdx ][ i ], and when the i-th entry in ref_pic_list_struct( listIdx, rplsIdx ) is an ILRP entry, inter_layer_ref_pic_flag[ listIdx ][ rplsIdx ][ i ] is equal to 1, and when the i-th entry in ref_pic_list_struct( listIdx, rplsIdx ) is not an ILRP entry, inter_layer_ref_pic_flag[ listIdx ][ rplsIdx ][ i ] is equal to 0.

[0017] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that ref_pic_list_struct( listIdx, rplsIdx ) and inter_layer_ref_pic_flag[ listIdx ][ rplsIdx ][ i ] are encoded into the bitstream in the SPS.

[0018] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that ref_pic_list_struct( listIdx, rplsIdx ) and inter_layer_ref_pic_flag[ listIdx ][ rplsIdx ][ i ] are encoded into the header associated with the current picture.

[0019] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the inter-layer reference picture is in the same AU as the current picture, and the inter-layer reference picture is associated with a lower layer identifier than the current picture.

[0020] In one embodiment, the present disclosure includes a video coding device comprising a processor, a receiver coupled to the processor, a memory coupled to the processor, and a transmitter coupled to the processor, wherein the processor, the receiver, the memory, and the transmitter are configured to execute any of the methods of the foregoing aspects.

[0021] In one embodiment, the present disclosure includes a non-transitory computer-readable medium comprising a computer program product for use by a video coding device, the computer program product comprising computer-executable instructions stored on the non-transitory computer-readable medium that, when executed by a processor, cause the video coding device to execute any of the methods of the foregoing aspects.

[0022] In one embodiment, the present disclosure includes a decoder comprising receiving means for receiving a bitstream comprising a current picture, a reference picture list structure, and an inter-layer reference picture flag; determining means for determining, based on the inter-layer reference picture flag, whether an entry in the reference picture list structure associated with the current picture is an ILRP entry; decoding means for decoding the current picture according to an inter-layer prediction based on an inter-layer reference picture indicated by the entry in the reference picture list structure when the entry is an ILRP entry; and transfer means for transferring the current picture for display as part of a decoded video sequence.

[0023] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the decoder is further configured to execute any of the methods of the foregoing aspects.

[0024] In one embodiment, the present disclosure is for encoding a current picture into a bitstream, wherein the current picture is encoded according to inter-layer prediction based on an inter-layer reference picture, for encoding, and for encoding a reference picture list structure into the bitstream, wherein the reference picture list structure comprises a plurality of entries for a plurality of reference pictures including an entry related to the current picture and indicates an inter-layer reference picture, for encoding, and for encoding an inter-layer reference picture flag into the bitstream, wherein the inter-layer reference picture flag indicates that the entry related to the current picture is an ILRP entry, for encoding, and includes an encoder comprising encoding means for performing the above and storage means for storing the bitstream for communication towards a decoder.

[0025] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the encoder is further configured to execute any of the methods of the foregoing aspects.

[0026] For clarity, any one of the above embodiments may be combined with any one or more of the other above embodiments to produce new embodiments within the scope of the present disclosure.

[0027] These and other features will be more clearly understood from the following forms for carrying out the invention, which are to be understood in conjunction with the accompanying drawings and the claims.

[0028] For a more complete understanding of the present disclosure, reference is now made to the following brief description, which is to be understood in connection with the accompanying drawings and forms for carrying out the invention, in which like reference numerals represent like parts.

Brief Description of the Drawings

[0029]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

DETAILED DESCRIPTION OF THE INVENTION

[0030] Exemplary implementations of one or more embodiments are provided below, but it should first be understood that the disclosed system and / or method may be implemented using any number of techniques, whether currently known or existing. The present disclosure should in no way be limited to the exemplary implementations, drawings, and techniques shown below, including the exemplary designs and implementations illustrated and described herein, but may be modified within the scope of the appended claims, together with the full scope of their equivalents.

[0031] The following terms are defined as follows unless used in the reverse context herein. In particular, the following definitions are intended to provide additional clarity to the present disclosure. However, terms may be interpreted differently in different contexts. Accordingly, the following definitions should be regarded as supplementary and should not be regarded as limiting any other definitions given to such terms herein.

[0032] A bitstream is a sequence of bits that includes video data compressed for transmission between an encoder and a decoder. An encoder is a device configured to employ an encoding process to compress video data into a bitstream. A decoder is a device configured to employ a decoding process to reconstruct video data from a bitstream for display. A picture is an array of luma samples and / or an array of chroma samples that creates a frame or a field of a frame. For clarity of explanation, an encoded or decoded picture can be called the current picture. A reference picture is a picture that includes reference samples that can be used when coding other pictures by reference via inter prediction and / or inter-layer prediction. A reference picture list is a list of reference pictures used for inter prediction and / or inter-layer prediction. Some video coding systems refer to two picture lists that can be denoted as reference picture list 1 and reference picture list 0. A reference picture list structure is an addressable syntax structure that includes multiple reference picture lists. A layer is a group of pictures all of which are related to characteristics of similar values, such as similar size, quality, resolution, signal-to-noise ratio, capabilities, etc. A layer identifier (ID) is an item of data that is associated with a picture and indicates that the picture is part of the layer being shown. Inter prediction is a mechanism for coding samples of a current picture by reference to samples shown in a reference picture different from the current picture, provided that the reference picture and the current picture are in the same layer. In an inter-layer context, since both the current picture and the reference picture are in the same layer, inter prediction can also be called intra-layer prediction. Inter-layer prediction is a mechanism for coding samples of a current picture by reference to samples shown in a reference picture, provided that the current picture and the reference picture are in different layers and thus have different layer IDs. An inter-layer reference picture is a reference picture used for inter-layer prediction.Some video coding systems may require that the current picture and the associated inter-layer reference pictures be included within the same access unit (AU). A reference picture list structure entry is an addressable location within the reference picture list structure that indicates a reference picture associated with the reference picture list. An inter-layer reference picture (ILRP) entry is an entry that includes a reference picture used for inter-layer prediction. An inter-layer reference picture flag is data that indicates that the reference picture within an entry of the reference picture list structure is an inter-layer reference picture. An ILRP layer indicator is data that indicates the layer associated with the inter-layer reference picture referenced by the current picture. A slice header is a part of the coded slice that includes data elements related to all video data within the tile represented within the slice. A sequence parameter set (SPS) is a parameter set that includes data related to a sequence of pictures. An AU is a set of one or more coded pictures associated with the same presentation time (e.g., the same picture order count) with respect to the output from a decoded picture buffer (DPB) (e.g., for presentation to a user). A decoded video sequence is a sequence of pictures that has been reconstructed by a decoder in preparation for presentation to a user.

[0033] In this specification, the following acronyms are used, namely, Coding Tree Block (CTB), Coding Tree Unit (CTU), Coding Unit (CU), Coded Video Sequence (CVS), Decoded Picture Buffer (DPB), Joint Video Experts Team (JVET), Motion-Constrained Tile Set (MCTS), Maximum Transfer Unit (MTU), Network Abstraction Layer (NAL), Picture Order Count (POC), Random Access Decodable Leading (RADL) picture, Random Access Skipped Leading (RASL) picture, Raw Byte Sequence Payload (RBSP), Reference Picture List (RPL), Sequence Parameter Set (SPS), Video Coding Layer (VCL), Versatile Video Coding (VVC), and Working Draft (WD).

[0034] To reduce the size of video files with only minimal loss of data, many video compression techniques can be employed. For example, video compression techniques can include performing spatial (e.g., intra-picture) prediction and / or temporal (e.g., inter-picture) prediction to reduce or eliminate data redundancy in a video sequence. In the case of block-based video coding, a video slice (e.g., a video picture or a portion of a video picture) may be partitioned into video blocks, also sometimes referred to as tree blocks, coding tree blocks (CTBs), coding tree units (CTUs), coding units (CUs), and / or coding nodes. Video blocks in an intra-coded (I) slice of a picture are coded using spatial prediction with respect to reference samples in adjacent blocks in the same picture. Video blocks in an inter-coded uni-directional prediction (P) or bi-directional prediction (B) slice of a picture can be coded by employing spatial prediction with respect to reference samples in adjacent blocks in the same picture, or temporal prediction with respect to reference samples in other reference pictures. A picture may sometimes be referred to as a frame and / or an image, and a reference picture may sometimes be referred to as a reference frame and / or a reference image. Spatial or temporal prediction results in a prediction block representing the image block. Residual data represents the pixel difference between the original image block and the prediction block. Thus, an inter-coded block is coded according to a motion vector indicating a block of reference samples forming the prediction block and residual data indicating the difference between the coded block and the prediction block. An intra-coded block is coded according to an intra-coding mode and residual data. For further compression, the residual data can be transformed from the pixel domain to a transform domain. These result in residual transform coefficients, and the residual transform coefficients can be quantized. The quantized transform coefficients may initially be arranged in a two-dimensional array. The quantized transform coefficients can be scanned to generate a one-dimensional vector of transform coefficients. To achieve further compression, entropy coding can be applied.Such video compression techniques are described in more detail below.

[0035] To ensure that the encoded video can be accurately decoded, the video is encoded and decoded according to the corresponding video coding standard. Video coding standards include International Telecommunication Union (ITU) Telecommunication Standardization Sector (ITU-T) H.261, International Organization for Standardization / International Electrotechnical Commission (ISO / IEC) Moving Picture Experts Group (MPEG)-1 Part 2, ITU-T H.262 or ISO / IEC MPEG-2 Part 2, ITU-T H.263, ISO / IEC MPEG-4 Part 2, ITU-T H.264 or ISO / IEC MPEG-4 Part 10, also known as Advanced Video Coding (AVC), and ITU-T H.265 or MPEG-H Part 2, also known as High Efficiency Video Coding (HEVC). AVC includes extensions such as Scalable Video Coding (SVC), Multi-View Video Coding (MVC), and Multi-View Video Coding Plus Depth (MVC+D), as well as three-dimensional (3D) AVC (3D-AVC). HEVC includes extensions such as Scalable HEVC (SHVC), Multi-View HEVC (MV-HEVC), and 3D HEVC (3D-HEVC). The Joint Video Expert Team (JVET) of ITU-T and ISO / IEC has started formulating a video coding standard called Versatile Video Coding (VVC). VVC is included in the working draft (WD) including JVET-N1001-v6.

[0036] A video coding system may encode pictures according to inter prediction. In inter prediction, a picture is coded by reference to another picture. The picture being coded is called the current picture, and the picture used as a reference is called the reference picture. Some video coding systems track reference pictures by employing a reference picture list. Further, some video coding systems employ scalable video coding. In scalable video coding, a video sequence is coded as a base layer and one or more enhancement layers. In this context, a picture may be divided into multiple pictures existing in different layers. For example, a lower layer version of a picture may be of lower quality than a higher layer version of the picture. Further, a lower layer version of a picture may be smaller than a higher layer version of the picture (e.g., may have a smaller width and / or height). When using layers, a picture in the current layer may be coded according to intra prediction (without using a reference picture), according to inter prediction by reference to a reference picture in the same layer, or according to inter-layer prediction by reference to a reference picture in a different layer. However, the reference picture list may not be designed to describe reference pictures across multiple layers.

[0037] Exemplary mechanisms for managing reference picture lists when adopting inter-layer prediction are disclosed herein. For example, the reference picture list may be configured to include an entry for each reference picture used by the coded video. Then, a flag may be adopted to indicate whether each entry includes an inter-prediction reference picture (for reference within the same level) or an inter-layer reference picture (for reference between layers). In one example, a group of reference picture flags for a syntax structure including a reference picture list may be denoted as inter_layer_ref_pic_flag[ listIdx ][ rplsIdx ][ i ]. The inter_layer_ref_pic_flag[ listIdx ][ rplsIdx ][ i ] flag may be set equal to 1 when the i-th entry in the reference picture structure is an ILRP entry, or may be set to 0 when the i-th entry in the reference picture structure is not an ILRP entry. Further, an ILRP layer indicator may be used to indicate which layer includes the inter-layer reference picture indicated by an entry in the reference picture structure. The decoder can then use the reference picture flag and the ILRP layer indicator to select an appropriate inter-layer reference picture for performing inter-layer prediction. Thus, the disclosed mechanisms create additional functionality in the encoder and / or decoder. Further, the disclosed mechanisms may improve coding efficiency, which may reduce the usage of processor, memory, and / or network resources in the encoder and / or decoder.

[0038] FIG. 1 is a flowchart of an exemplary operating method 100 for coding a video signal. Specifically, the video signal is encoded in an encoder. The encoding process compresses the video signal by employing various mechanisms to reduce the video file size. The smaller file size enables the compressed video file to be transmitted to the user while reducing the associated bandwidth overhead. The decoder then decodes the compressed video file to reconstruct the original video signal for display to the end user. The decoding process generally mimics the encoding process precisely to enable the decoder to reconstruct the video signal without contradictions.

[0039] In step 101, a video signal is input into the encoder. For example, the video signal may be an uncompressed video file stored in memory. As another example, the video file may be captured by a video capture device such as a video camera and encoded to support live video streaming. The video file may include both an audio component and a video component. The video component includes a series of image frames that, when viewed in sequence, give a visual impression of movement. The frames include pixels that are represented with respect to light, referred to herein as the luma component (or luma samples), and color, referred to as the chroma component (or color samples). In some examples, the frames may also include depth values to support 3D displays.

[0040] In step 103, the video is segmented into blocks. Segmenting includes re - dividing the pixels in each frame into square blocks and / or rectangular blocks for compression. For example, in High Efficiency Video Coding (HEVC) (also known as H.265 and MPEG - H Part 2), a frame can first be divided into Coding Tree Units (CTUs) which are blocks of a predefined size (e.g., 64 pixels × 64 pixels). A CTU contains both luma samples and chroma samples. A coding tree can be employed to divide the CTU into blocks and then recursively re - divide the blocks until a configuration that supports further encoding is obtained. For example, the luma component of a frame may be re - divided until the individual blocks contain relatively uniform illumination values. Further, the chroma component of a frame may be re - divided until the individual blocks contain relatively uniform color values. Thus, the segmentation mechanism varies according to the content of the video frame.

[0041] In step 105, various compression mechanisms are employed to compress the image blocks segmented in step 103. For example, inter prediction and / or intra prediction may be employed. Inter prediction is designed to utilize the fact that objects in a common scene tend to appear in consecutive frames. Therefore, blocks representing objects in a reference frame do not need to be repeatedly described in adjacent frames. In particular, an object such as a table may remain in a fixed position over multiple frames. Therefore, the table is described once, and adjacent frames can refer back to the reference frame. To align objects over multiple frames, a pattern matching mechanism may be employed. Further, for example, due to object movement or camera movement, a moving object may be represented across multiple frames. As a specific example, a video may show a car moving across the screen over multiple frames. Motion vectors may be employed to describe such movement. A motion vector is a two-dimensional vector that provides an offset from the coordinates of an object in one frame to the coordinates of the object in a reference frame. Therefore, inter prediction can encode an image block in the current frame as a set of motion vectors indicating the offsets from corresponding blocks in the reference frame.

[0042] Intra prediction encodes blocks within a common frame. Intra prediction utilizes the fact that luma and chroma components tend to cluster within a frame. For example, green fragments within a portion of a tree tend to be placed adjacent to similar green fragments. Intra prediction employs multiple directional prediction modes (e.g., 33 in HEVC), a planar mode, and a direct current (DC) mode. The directional mode indicates that the current block is similar / same as the samples of adjacent blocks in the corresponding direction. The planar mode indicates that a series of blocks along a row / column (e.g., a plane) can be interpolated based on adjacent blocks at the edge of the row. The planar mode effectively indicates a smooth transition of light / color across the row / column by adopting a relatively constant gradient when changing values. The DC mode is employed for boundary smoothing and indicates that the block is similar / same as the average value related to the samples of all adjacent blocks related to the angular direction of the directional prediction mode. Thus, an intra prediction block can represent an image block as various prediction mode values of relationships rather than actual values. Further, an inter prediction block can represent an image block as motion vector values rather than actual values. In either case, the prediction block may not exactly represent the image block in some cases. Any difference is stored in the residual block. A transformation may be applied to the residual block to further compress the file.

[0043] In step 107, various filter processing techniques may be applied. In HEVC, the filter is applied according to the in-loop filter processing method. The block-based prediction described above may result in the generation of block-shaped images in the decoder. Further, the block-based prediction method may encode a block and then reconstruct the encoded block for later use as a reference block. The in-loop filter processing method repeatedly applies a noise suppression filter, a deblocking filter, an adaptive loop filter, and a sample adaptive offset (SAO) filter to blocks / frames. These filters reduce such blocking artifacts so that the encoded file can be accurately reconstructed. Further, among the subsequent blocks encoded based on the reconstructed reference block, these filters reduce the artifacts in the reconstructed reference block so that the artifacts are less likely to generate additional artifacts.

[0044] When the video signal has been segmented, compressed, and filtered, in step 109, the obtained data is encoded in a bitstream. The bitstream includes the data described above, as well as any signaling data desired to support proper video signal reconstruction in the decoder. For example, such data may include segmentation data, prediction data, residual blocks, and various flags that provide coding instructions to the decoder. The bitstream may be stored in memory for transmission to the decoder upon request. The bitstream may also be broadcast and / or multicast to multiple decoders. The creation of the bitstream is an iterative process. Thus, steps 101, 103, 105, 107, and 109 may be performed continuously and / or simultaneously over many frames and blocks. The order shown in FIG. 1 is presented for clarity and simplicity of explanation and is not intended to limit the video coding process to a particular order.

[0045] The decoder receives the bitstream at step 111 and starts the decoding process. Specifically, the decoder adopts an entropy decoding method to convert the bitstream into corresponding syntax and video data. At step 111, the decoder adopts the syntax data from the bitstream to determine the segmentation for the frame. The segmentation should be consistent with the result of the block segmentation at step 103. The entropy coding / decoding as adopted at step 111 is described next. The encoder makes many selections during the compression process, such as selecting a block segmentation method from several possible options based on the spatial arrangement of the values in the input image. Signaling exact options may involve adopting a large number of bins. As used herein, a bin is a binary value treated as a variable (for example, a bit value that may change according to the context). Entropy coding enables the encoder to discard any option that is not clearly viable for a particular case and leave a set of acceptable options. Each acceptable option is then assigned a codeword. The length of the codeword is based on the number of acceptable options (for example, 1 bin for 2 options, 2 bins for 3 - 4 options, etc.). The encoder then encodes the codeword for the selected option. This method reduces the size of the codeword when the codeword is of a size similar to what is desired to uniquely indicate an option from a small subset of acceptable options, as opposed to uniquely indicating an option from a potentially large set of all possible options. The decoder then decodes the option by determining the set of acceptable options in a method similar to the encoder. By determining the set of acceptable options, the decoder can read the codeword and determine the selection made by the encoder.

[0046] In step 113, the decoder performs block decoding. Specifically, the decoder employs inverse transformation to generate residual blocks. Next, the decoder employs the residual blocks and corresponding prediction blocks to reconstruct image blocks according to the partition. The prediction blocks may include both intra prediction blocks and inter prediction blocks such as those generated in the encoder in step 105. The reconstructed image blocks are then arranged into the frame of the reconstructed video signal according to the partition data determined in step 111. The syntax for step 113 may also be signaled in the bitstream via entropy coding as described above.

[0047] In step 115, filtering is performed on the frame of the reconstructed video signal in a manner similar to step 107 in the encoder. For example, a noise suppression filter, a deblocking filter, an adaptive loop filter, and an SAO filter may be applied to the frame to remove blocking artifacts. Once the frame is filtered, the video signal may be output to a display in step 117 for viewing by an end user.

[0048] Figure 2 is a schematic diagram of an exemplary coding and decoding (codec) system 200 for video coding. Specifically, the codec system 200 provides functionality for implementing the operation method 100. The codec system 200 is generalized to show components employed in both the encoder and the decoder. The codec system 200 receives and partitions a video signal as described with respect to steps 101 and 103 in the operation method 100, which results in a partitioned video signal 201. The codec system 200 then compresses the partitioned video signal 201 into a coded bitstream when acting as an encoder as described with respect to steps 105, 107, and 109 in method 100. When acting as a decoder, the codec system 200 generates an output video signal from the bitstream as described with respect to steps 111, 113, 115, and 117 in the operation method 100. The codec system 200 includes a general-purpose codec control component 211, a transform scaling and quantization component 213, an intra-picture estimation component 215, an intra-picture prediction component 217, a motion compensation component 219, a motion estimation component 221, a scaling and inverse transform component 229, a filter control analysis component 227, an in-loop filter component 225, a decoded picture buffer component 223, as well as a header formatting and context-adaptive binary arithmetic coding (CABAC) component 231. Such components are coupled as shown. In FIG. 2, the solid lines indicate the movement of data to be encoded / decoded, and the dashed lines indicate the movement of control data that controls the operation of other components. All components of the codec system 200 may be present in the encoder. The decoder may include a subset of the components of the codec system 200. For example, the decoder may include the intra-picture prediction component 217, the motion compensation component 219, the scaling and inverse transform component 229, the in-loop filter component 225, as well as the decoded picture buffer component 223. These components are described next.

[0049] The segmented video signal 201 is a captured video sequence that is segmented into blocks of pixels by a coding tree. The coding tree adopts various splitting modes to split the blocks of pixels into smaller blocks of pixels. These blocks can then be further split into even smaller blocks. The blocks are sometimes referred to as nodes on the coding tree. A larger parent node is split into smaller child nodes. The number of times a node is split is called the depth of the node / coding tree. The split blocks can optionally be included within coding units (CUs). For example, a CU can be a sub - part of a CTU that includes a luma block, a red difference chroma (Cr) block, and a blue difference chroma (Cb) block, along with the corresponding syntax instructions for the CU. The splitting modes may include a binary tree (BT), a triple tree (TT), and a quad tree (QT) that are adopted to split each node into two, three, or four child nodes with shapes that vary according to the adopted splitting mode. The segmented video signal 201 is transferred to a general - purpose coder control component 211, a transform scaling and quantization component 213, an intra - picture prediction component 215, a filter control analysis component 227, and a motion estimation component 221 for compression.

[0050] The general coder control component 211 is configured to make decisions related to the coding of images of a video sequence into a bitstream according to application constraints. For example, the general coder control component 211 manages the optimization of bitrate / bitstream size versus reconstructed quality. Such decisions may be made based on storage space / bandwidth availability and image resolution requirements. The general coder control component 211 also manages buffer utilization in light of the transmission speed to mitigate buffer underrun and overrun problems. To manage these problems, the general coder control component 211 manages the partitioning, prediction, and filtering processes by other components. For example, the general coder control component 211 may dynamically increase the compression computation amount to increase the resolution and increase the bandwidth usage, or decrease the compression computation amount to decrease the resolution and bandwidth usage. Therefore, the general coder control component 211 controls other components of the codec system 200 to reconcile the video signal reconstructed quality with the bitrate importance. The general coder control component 211 creates control data for controlling the operations of other components. The control data is also transferred to the header formatting and CABAC component 231 for header formatting and for signaling parameters for decoding in the decoder, which are encoded in the bitstream.

[0051] The partitioned video signal 201 is also sent to the motion estimation component 221 and the motion compensation component 219 for inter prediction. A frame or slice of the partitioned video signal 201 may be divided into a plurality of video blocks. The motion estimation component 221 and the motion compensation component 219 perform inter prediction coding of the received video blocks with respect to one or more blocks in one or more reference frames for temporal prediction. The codec system 200 may execute a plurality of coding paths, for example, to select an appropriate coding mode for each block of video data.

[0052] The motion estimation component 221 and the motion compensation component 219 may be highly integrated but are shown separately for conceptual purposes. The motion estimation performed by the motion estimation component 221 is a process of generating a motion vector that estimates the motion for a video block. The motion vector may indicate, for example, the displacement of a coded object with respect to a prediction block. The prediction block is a block that is recognized as closely matching the block to be coded from the perspective of pixel differences. The prediction block may also be called a reference block. Such pixel differences may be determined by the sum of absolute differences (SAD), the sum of square differences (SSD), or other difference metrics. HEVC employs several coded objects including CTUs, coding tree blocks (CTBs), and CUs. For example, a CTU may be divided into CTBs, and then the CTBs may be divided into CUs for inclusion in the CUs. A CU may be coded as a prediction unit (PU) that includes prediction data and / or a transform unit (TU) that includes transformed residual data for the CU. The motion estimation component 221 generates the motion vector, PU, and TU by using rate-distortion analysis as part of the rate-distortion optimization process. For example, the motion estimation component 221 may determine a plurality of reference blocks, a plurality of motion vectors, etc. for the current block / frame, and may select the reference block, motion vector, etc. with the best rate-distortion characteristics. The best rate-distortion characteristics reconcile both the quality of video reconstruction (e.g., the amount of data loss due to compression) with the coding efficiency (e.g., the size of the final encoding).

[0053] In some examples, the codec system 200 may calculate values for sub-integer pixel positions of reference pictures stored in the decoded picture buffer component 223. For example, the video codec system 200 may interpolate values at 1 / 4 pixel positions, 1 / 8 pixel positions, or other fractional pixel positions of the reference picture. Accordingly, the motion estimation component 221 may perform motion search for full pixel positions and fractional pixel positions and may output motion vectors with fractional pixel accuracy. The motion estimation component 221 calculates a motion vector for a prediction unit (PU) of a video block in an inter-coded slice by comparing the position of the PU with the position of a prediction block of the reference picture. The motion estimation component 221 outputs the calculated motion vector as motion data for header formatting and to the CABAC component 231 for encoding, and also outputs the motion to the motion compensation component 219.

[0054] Motion compensation performed by the motion compensation component 219 may involve fetching or generating a prediction block based on the motion vector determined by the motion estimation component 221. Again, in some examples, the motion estimation component 221 and the motion compensation component 219 may be functionally integrated. Upon receiving a motion vector for the current video block's PU, the motion compensation component 219 may identify the position of the prediction block pointed to by the motion vector. Then, a residual video block is formed by subtracting the pixel values of the prediction block from the pixel values of the currently coded video block, forming pixel difference values. Generally, the motion estimation component 221 performs motion estimation with respect to the luma component, and the motion compensation component 219 uses the motion vector calculated based on the luma component for both the chroma component and the luma component. The prediction block and the residual block are transferred to the transform scaling and quantization component 213.

[0055] The segmented video signal 201 is also sent to the intra prediction component 215 and the intra prediction component 217 within the picture. Similar to the motion estimation component 221 and the motion compensation component 219, the intra prediction component 215 and the intra prediction component 217 within the picture may be highly integrated, but are shown separately for conceptual purposes. The intra prediction component 215 and the intra prediction component 217 perform intra prediction of the current block for the block within the current frame as an alternative to inter prediction performed by the motion estimation component 221 and the motion compensation component 219 between frames as described above. Specifically, the intra prediction component 215 determines the intra prediction mode to be used for encoding the current block. In some examples, the intra prediction component 215 selects an appropriate intra prediction mode for encoding the current block from among a plurality of tested intra prediction modes. The selected intra prediction mode is then transferred to the header formatting and CABAC component 231 for encoding.

[0056] For example, the intra prediction component 215 uses rate distortion analysis to calculate rate distortion values for various tested intra prediction modes and selects the intra prediction mode with the best rate distortion characteristics among the tested modes. Rate distortion analysis generally determines the amount of distortion (i.e., error) between the encoded block and the original unencoded block that was encoded to generate the encoded block, as well as the bit rate (e.g., the number of bits) used to generate the encoded block. The intra prediction component 215 calculates a ratio from the distortion and rate for various encoded blocks to determine which intra prediction mode for the block presents the best rate distortion value. In addition, the intra prediction component 215 may be configured to code depth blocks of the depth map using a depth modeling mode based on rate distortion optimization (RDO).

[0057] When the in-picture prediction component 217 is implemented on the encoder, it may generate a residual block from a prediction block based on the selected intra prediction mode determined by the in-picture estimation component 215, or when implemented on the decoder, it may read the residual block from the bitstream. The residual block includes the difference in values between the prediction block and the original block, represented as a matrix. The residual block is then transferred to the transform scaling and quantization component 213. The in-picture estimation component 215 and the in-picture prediction component 217 can operate on both the luma component and the chroma component.

[0058] The transform scaling and quantization component 213 is configured to further compress the residual block. The transform scaling and quantization component 213 applies a transform such as a discrete cosine transform (DCT), a discrete sine transform (DST), or a conceptually similar transform to the residual block to generate a video block with residual transform coefficient values. A wavelet transform, an integer transform, a subband transform, or other types of transforms may also be used. The transform can transform the residual information from the pixel value domain to a transform domain such as the frequency domain. The transform scaling and quantization component 213 is also configured to scale the transformed residual information, for example, based on frequency. Such scaling involves applying a scale factor to the residual information, as a result of which different frequency information is quantized at different granularities, which may affect the final visual quality of the reconstructed video. The transform scaling and quantization component 213 is also configured to quantize the transform coefficients to further reduce the bitrate. The quantization process may reduce the bit depth associated with some or all of the coefficients. The degree of quantization may be modified by adjusting the quantization parameter. In some examples, the transform scaling and quantization component 213 may then perform a scan of the matrix containing the quantized transform coefficients. The quantized transform coefficients are transferred to the header formatting and CABAC component 231 to be encoded in the bitstream.

[0059] The scaling and inverse transformation component 229 applies the inverse operations of the transform scaling and quantization component 213 to support motion estimation. The scaling and inverse transformation component 229 applies inverse scaling, transformation, and / or quantization to reconstruct the residual block in the pixel domain, for example, for later use as a reference block that can be a prediction block for another current block. The motion estimation component 221 and / or the motion compensation component 219 may calculate the reference block by adding the residual block back to the corresponding prediction block for use in motion estimation of subsequent blocks / frames. A filter is applied to the reconstructed reference block to reduce artifacts that occur during scaling, quantization, and transformation. Such artifacts may cause (and generate additional artifacts) inaccurate predictions in some cases when subsequent blocks are predicted.

[0060] The filter control analysis component 227 and the in-loop filter component 225 apply a filter to the residual block and / or the reconstructed image block. For example, to reconstruct the original image block, the transformed residual block from the scaling and inverse transform component 229 may be combined with the corresponding prediction block from the intra-picture prediction component 217 and / or the motion compensation component 219. Then, a filter may be applied to the reconstructed image block. In some examples, the filter may instead be applied to the residual block. Similar to the other components in FIG. 2, the filter control analysis component 227 and the in-loop filter component 225 may be highly integrated and implemented together, but are shown separately for conceptual purposes. The filter applied to the reconstructed reference block is applied to a specific spatial region and includes a plurality of parameters for adjusting how such a filter is applied. The filter control analysis component 227 analyzes the reconstructed reference block to determine where such a filter should be applied and sets the corresponding parameters. Such data is transferred as filter control data for encoding to the header formatting and CABAC component 231. The in-loop filter component 225 applies such a filter based on the filter control data. The filter may include a deblocking filter, a noise suppression filter, a SAO filter, and an adaptive loop filter. Such a filter may be applied, depending on the example, in the spatial domain / pixel domain (e.g., to the reconstructed pixel block) or in the frequency domain.

[0061] When operating as an encoder, the filtered and reconstructed image blocks, residual blocks, and / or prediction blocks are stored in the decoded picture buffer component 223 for later use in motion estimation, as described above. When operating as a decoder, the decoded picture buffer component 223 stores the reconstructed and filtered blocks and transfers them towards the display as part of the output video signal. The decoded picture buffer component 223 may be any memory device capable of storing prediction blocks, residual blocks, and / or reconstructed image blocks.

[0062] The header formatting and CABAC component 231 receives data from various components of the codec system 200 and encodes such data into the coded bitstream for transmission towards the decoder. Specifically, the header formatting and CABAC component 231 generates various headers for encoding control data such as general control data and filter control data. Further, all of the prediction data including intra prediction and motion data as well as the residual data in the form of quantized transform coefficient data are encoded in the bitstream. The final bitstream contains all the information desired by the decoder for reconstructing the segmented original video signal 201. Such information may also include an intra prediction mode index table (also called a codeword mapping table), the definition of coding contexts for various blocks, the indication of the most accurate intra prediction mode, the indication of segmentation information, etc. Such data may be encoded by employing entropy coding. For example, the information may be encoded by adopting context adaptive variable length coding (CAVLC), CABAC, syntax-based context adaptive binary arithmetic coding (SBAC), probability interval segmentation entropy (PIPE) coding, or another entropy coding technique. Following the entropy coding, the coded bitstream may be transmitted to another device (e.g., a video decoder) or may be archived so that it can be transmitted or retrieved later.

[0063] FIG. 3 is a block diagram showing an exemplary video encoder 300. The video encoder 300 may be employed to implement the encoding function of the codec system 200 and / or to perform steps 101, 103, 105, 107, and / or 109 of the operation method 100. The encoder 300 divides an input video signal, resulting in a divided video signal 301 that is substantially similar to the divided video signal 201. The divided video signal 301 is then compressed and encoded into a bitstream by components of the encoder 300.

[0064] Specifically, the divided video signal 301 is transferred to an intra-prediction in-picture prediction component 317 for intra prediction. The in-picture prediction component 317 may be substantially similar to the in-picture estimation component 215 and the in-picture prediction component 217. The divided video signal 301 is also transferred to a motion compensation component 321 for inter prediction based on reference blocks in a decoded picture buffer component 323. The motion compensation component 321 may be substantially similar to the motion estimation component 221 and the motion compensation component 219. The prediction blocks and residual blocks from the in-picture prediction component 317 and the motion compensation component 321 are transferred to a transform and quantization component 313 for transformation and quantization of the residual blocks. The transform and quantization component 313 may be substantially similar to the transform scaling and quantization component 213. The transformed and quantized residual blocks, as well as the corresponding prediction blocks, are transferred to an entropy coding component 331 (along with relevant control data) for coding into the bitstream. The entropy coding component 331 may be substantially similar to the header formatting and CABAC component 231.

[0065] The transformed and quantized residual blocks, and / or the corresponding prediction blocks, are also transferred from the transform and quantization component 313 to the inverse transform and dequantization component 329 for reconstruction into reference blocks for use by the motion compensation component 321. The inverse transform and dequantization component 329 may be substantially similar to the scaling and inverse transform component 229. The in-loop filter within the in-loop filter component 325 is also applied to the residual block and / or the reconstructed reference block, for example, according to an example. The in-loop filter component 325 may be substantially similar to the filter control analysis component 227 and the in-loop filter component 225. The in-loop filter component 325 may include multiple filters as described for the in-loop filter component 225. The filtered block is then stored in the decoded picture buffer component 323 for use as a reference block by the motion compensation component 321. The decoded picture buffer component 323 may be substantially similar to the decoded picture buffer component 223.

[0066] FIG. 4 is a block diagram showing an exemplary video decoder 400. The video decoder 400 may be employed to implement the decoding function of the codec system 200 and / or to perform steps 111, 113, 115, and / or 117 of the operation method 100. The decoder 400 receives, for example, a bitstream from the encoder 300 and generates a reconstructed output video signal based on the bitstream for display to an end user.

[0067] The bitstream is received by the entropy decoding component 433. The entropy decoding component 433 is configured to implement an entropy decoding method, such as CAVLC, CABAC, SBAC, PIPE coding, or other entropy coding techniques. For example, the entropy decoding component 433 may employ header information to provide a context for interpreting additional data encoded as codewords in the bitstream. The decoded information includes any desired information for decoding a video signal, such as general control data, filter control data, segmentation information, motion data, prediction data, and quantized transform coefficients from residual blocks. The quantized transform coefficients are transferred to the inverse transform and quantization component 429 for reconstruction of the residual blocks. The inverse transform and quantization component 429 may be similar to the inverse transform and quantization component 329.

[0068] The reconstructed residual block and / or prediction block is transferred to the in-picture prediction component 417 for the reconstruction of an image block based on an intra prediction operation. The in-picture prediction component 417 may be similar to the in-picture estimation component 215 and the in-picture prediction component 217. Specifically, the in-picture prediction component 417 adopts a prediction mode to identify the position of a reference block within a frame, and applies the residual block to the result to reconstruct an intra-predicted image block. The reconstructed, intra-predicted image block and / or residual block, as well as the corresponding inter prediction data, are transferred to the decoded picture buffer component 423 via the in-loop filter component 425, which may be substantially similar to the decoded picture buffer component 223 and the in-loop filter component 225, respectively. The in-loop filter component 425 filters the reconstructed image block, residual block, and / or prediction block, and such information is stored in the decoded picture buffer component 423. The reconstructed image block from the decoded picture buffer component 423 is transferred to the motion compensation component 421 for inter prediction. The motion compensation component 421 may be substantially similar to the motion estimation component 221 and / or the motion compensation component 219. Specifically, the motion compensation component 421 adopts a motion vector from a reference block to generate a prediction block, and applies the residual block to the result to reconstruct an image block. The resulting reconstructed block may also be transferred to the decoded picture buffer component 423 via the in-loop filter component 425. The decoded picture buffer component 423 continues to store the additional reconstructed image blocks, and such image blocks may be reconstructed into a frame via the segmentation information. Such a frame may also be placed within a sequence. The sequence is output towards a display as a reconstructed output video signal.

[0069] FIG. 5 is a schematic diagram showing an example of a unidirectional inter prediction 500 that is executed to determine a motion vector (MV) in, for example, block compression step 105, block decoding step 113, motion estimation component 221, motion compensation component 219, motion compensation component 321, and / or motion compensation component 421.

[0070] The unidirectional inter prediction 500 employs a reference picture 530 having a reference block 531 to predict a current block 511 in a current picture 510. The reference picture 530 may be temporally located after the current picture 510 (e.g., like a subsequent reference picture) as shown, but in some examples may be temporally located before the current picture 510 (e.g., like a previous reference picture). The current picture 510 is an exemplary frame / picture being encoded / decoded at a particular time. The current picture 510 includes an object in the current block 511 that matches an object in the reference block 531 of the reference picture 530. The reference picture 530 is a picture employed as a reference for encoding the current picture 510, and the reference block 531 is a block in the reference picture 530 that includes an object also included in the current block 511 of the current picture 510.

[0071] Currently, block 511 is any coding unit that is being encoded / decoded at a specified point during the coding process. The current block 511 may be the entire segmented block or a sub-block when the affine inter prediction mode is adopted. The current picture 510 is separated from the reference picture 530 by some temporal distance (TD: temporal distance) 533. TD 533 indicates the amount of time between the current picture 510 and the reference picture 530 in the video sequence and can be measured in picture units. The prediction information for the current block 511 may refer to the reference picture 530 and / or the reference block 531 by a reference index indicating the direction and temporal distance between pictures. Over the time period represented by TD 533, the object in the current block 511 moves from a certain position in the current picture 510 to another position in the reference picture 530 (e.g., the position of the reference block 531). For example, the object may move along a motion trajectory 513 which is the direction of the temporal movement of the object. The motion vector 535 represents the direction and magnitude of the movement of the object along the motion trajectory 513 over TD 533. Therefore, the encoded motion vector 535, the reference block 531, and the residual including the difference between the current block 511 and the reference block 531 provide sufficient information to reconstruct the current block 511 and identify the position of the current block 511 in the current picture 510.

[0072] FIG. 6 is a schematic diagram showing an example of a bidirectional inter prediction 600, such as may be executed to determine an MV, in, for example, block compression step 105, block decoding step 113, motion estimation component 221, motion compensation component 219, motion compensation component 321, and / or motion compensation component 421.

[0073] The bidirectional inter prediction 600 is similar to the unidirectional inter prediction 500, but employs a pair of reference pictures to predict the current block 611 in the current picture 610. Thus, the current picture 610 and the current block 611 are substantially similar to the current picture 510 and the current block 511, respectively. The current picture 610 is temporally placed between a preceding reference picture 620 that appears before the current picture 610 in the video sequence and a subsequent reference picture 630 that appears after the current picture 610 in the video sequence. The preceding reference picture 620 and the subsequent reference picture 630 are substantially similar to the reference picture 530, originally.

[0074] Current block 611 is aligned with a preceding reference block 621 in a preceding reference picture 620 and a succeeding reference block 631 in a succeeding reference picture 630. Such alignment indicates that, over the course of the video sequence, along motion trajectory 613 and through current block 611, an object moves from a position in preceding reference block 621 to a position in succeeding reference block 631. Current picture 610 is separated from preceding reference picture 620 by some preceding temporal distance (TD0) 623 and from succeeding reference picture 630 by some succeeding temporal distance (TD1) 633. TD0 623 indicates the amount of time, in picture units, between preceding reference picture 620 and current picture 610 in the video sequence. TD1 633 indicates the amount of time, in picture units, between current picture 610 and succeeding reference picture 630 in the video sequence. Thus, over the time period indicated by TD0 623, along motion trajectory 613, the object moves from preceding reference block 621 to current block 611. The object also moves from current block 611 to succeeding reference block 631 over the time period indicated by TD1 633 along motion trajectory 613. Prediction information for current block 611 can refer to preceding reference picture 620 and / or preceding reference block 621, as well as succeeding reference picture 630 and / or succeeding reference block 631, by a pair of reference indices indicating the direction and temporal distance between pictures.

[0075] The preceding motion vector (MV0) 625 represents the direction and magnitude of the movement of an object along the motion trajectory 613 over TD0 623 (e.g., between the preceding reference picture 620 and the current picture 610). The subsequent motion vector (MV1) 635 represents the direction and magnitude of the movement of the object along the motion trajectory 613 over TD1 633 (e.g., between the current picture 610 and the subsequent reference picture 630). Thus, in bidirectional inter prediction 600, the current block 611 can be coded and reconstructed by adopting the preceding reference block 621 and / or the subsequent reference block 631, MV0 625, and MV1 635.

[0076] FIG. 7 is a schematic diagram showing an example of layer-based prediction 700, such as may be performed to determine an MV in, for example, block compression step 105, block decoding step 113, motion estimation component 221, motion compensation component 219, motion compensation component 321, and / or motion compensation component 421. Layer-based prediction 700 is similar to unidirectional inter prediction 500 and / or bidirectional inter prediction 600, but is also performed between pictures in different layers.

[0077] Layer-based prediction 700 is applied between pictures 711, 712, 713, and 714 and pictures 715, 716, 717, and 718 in different layers. In the illustrated example, pictures 711, 712, 713, and 714 are part of layer N+1 732, and pictures 715, 716, 717, and 718 are part of layer N 731. Layers such as layer N 731 and / or layer N+1 732 are groups of pictures all related to characteristics of similar values, such as similar size, quality, resolution, signal-to-noise ratio, capabilities, etc. In the illustrated example, layer N+1 732 is related to a larger image size than layer N 731. Thus, in this example, pictures 711, 712, 713, and 714 in layer N+1 732 have a larger picture size (e.g., greater height and width, and thus more samples) than pictures 715, 716, 717, and 718 in layer N 731. However, such pictures can be separated between layer N+1 732 and layer N 731 by other characteristics. Only two layers, namely layer N+1 732 and layer N 731, are illustrated, but a set of pictures can be separated into any number of layers based on the relevant characteristics. Layer N+1 732 and layer N 731 may also be indicated by layer IDs. A layer ID is an item of data related to a picture and indicates that the picture is part of the layer being shown. Thus, each picture 711 - 718 may be related to a corresponding layer ID to indicate which of layer N+1 732 or layer N 731 contains the corresponding picture.

[0078] Pictures 711 - 718 in different layers 731 - 732 are configured to be displayed as alternatives. Thus, pictures 711 - 718 in different layers 731 - 732 can share the same picture order count and can be included in the same AU. As used herein, an AU is a set of one or more coded pictures related to the same display time for the output from the DPB. For example, if a smaller picture is desired, the decoder may decode and display picture 715 at the current display time, or if a larger picture is desired, the decoder may decode and display picture 711 at the current display time. Thus, pictures 711 - 714 in the higher layer N + 1 732 contain substantially the same image data as the corresponding pictures 715 - 718 in the lower layer N 731 (despite the difference in picture size). Specifically, picture 711 contains substantially the same image data as picture 715, picture 712 contains substantially the same image data as picture 716, and so on.

[0079] Pictures 711 to 718 can be coded by reference to other pictures 711 to 718 in the same layer N731 or N+1 732. Coding a picture by reference to another picture in the same layer results in inter prediction 723, which is substantially similar to unidirectional inter prediction 500 and / or bidirectional inter prediction 600. Inter prediction 723 is indicated by solid arrows. For example, picture 713 may be coded by adopting inter prediction 723 using one or two of pictures 711, 712, and / or 714 in layer N+1 732 as references, where one picture is referenced for unidirectional inter prediction and / or two pictures are references for bidirectional inter prediction. Further, picture 717 may be coded by adopting inter prediction 723 using one or two of pictures 715, 716, and / or 718 in layer N731 as references, where one picture is referenced for unidirectional inter prediction and / or two pictures are references for bidirectional inter prediction. When a picture is used as a reference for another picture in the same layer when performing inter prediction 723, that picture may be called a reference picture. For example, picture 712 may be a reference picture used to code picture 713 according to inter prediction 723. Inter prediction 723 can also be called intra-layer prediction in a multi-layer context. Thus, inter prediction 723 is a mechanism for coding samples of the current picture by reference to samples shown in a reference picture different from the current picture, provided that the reference picture and the current picture are in the same layer.

[0080] Pictures 711 to 718 can also be coded by reference to other pictures 711 to 718 in different layers. This process is also known as inter-layer prediction 721 and is indicated by the dashed arrows. Inter-layer prediction 721 is a mechanism for coding samples of the current picture by reference to the samples shown in the reference picture, provided that the current picture and the reference picture are in different layers and thus have different layer IDs. For example, a picture in the lower layer N731 can be used as a reference picture for coding the corresponding picture in the higher layer N+1 732. As a specific example, picture 711 can be coded by reference to picture 715 by inter-layer prediction 721. In such a case, picture 715 is used as an inter-layer reference picture. An inter-layer reference picture is a reference picture used for inter-layer prediction 721. In most cases, inter-layer prediction 721 is restricted such that the current picture, such as picture 711, can only use an inter-layer reference picture that is in the lower layer and is included in the same AU as, for example, picture 715. When multiple layers (e.g., three or more) are available, inter-layer prediction 721 can encode / decrypt the current picture based on multiple inter-layer reference pictures at a level lower than the current picture.

[0081] The video encoder can employ layer-based prediction 700 to encode pictures 711-718 through many different combinations and / or permutations of inter prediction 723 and inter-layer prediction 721. For example, picture 715 may be coded according to intra prediction. Pictures 716-718 may then be coded according to inter prediction 723 by using picture 715 as a reference picture. Further, picture 711 may be coded according to inter-layer prediction 721 by using picture 715 as an inter-layer reference picture. Pictures 712-714 may then be coded according to inter prediction 723 by using picture 711 as a reference picture. Thus, the reference picture can serve both as a single-layer reference picture and an inter-layer reference picture for different coding mechanisms. By coding the picture of the higher layer N+1 732 based on the picture of the lower layer N 731, the higher layer N+1 732 can avoid adopting intra prediction that has a much lower coding efficiency than inter prediction 723 and inter-layer prediction 721. Thus, the poor coding efficiency of intra prediction may be limited to the picture with the minimum / lowest quality and thus may be limited to coding the minimum amount of video data. The picture used as a reference picture and / or an inter-layer reference picture may be indicated in the entries of the reference picture list included in the reference picture list structure.

[0082] Figure 8 is a schematic diagram showing an exemplary reference picture list structure (RPL structure) 800. The RPL structure 800 may be adopted to store the indication of the reference picture and / or the inter-layer reference picture used in unidirectional inter prediction 500, bidirectional inter prediction 600, and / or layer-based prediction 700. Thus, the RPL structure 800 may be adopted by the codec system 200, the encoder 300, and / or the decoder 400 when executing method 100.

[0083] The RPL structure 800 is an addressable syntax structure that includes multiple reference picture lists such as RPL0 811 and RPL1 812. The RPL structure 800 may be stored, for example, in the SPS and / or slice headers of the bitstream according to an example. Reference picture lists such as RPL0 811 and RPL1 812 are lists of reference pictures used for inter prediction and / or inter-layer prediction. RPL0 811 and RPL1 812 may each include a plurality of entries 815. The reference picture list structure entry 815 is an addressable location in the RPL structure 800 that indicates a reference picture related to a reference picture list such as RPL0 811 and / or RPL1 812. Each entry 815 may include a picture order count (POC) value (or other pointer value) that references a picture used for inter prediction. Specifically, the reference to the picture used by the unidirectional inter prediction 500 is stored in RPL0 811, and the reference to the picture used by the bidirectional inter prediction 600 is stored in both RPL0 811 and RPL1 812. For example, the bidirectional inter prediction 600 may use one reference picture indicated by RPL0 811 and one reference picture indicated by RPL1 812.

[0084] In a specific example, the RPL structure 800 may be represented as ref_pic_list_struct( listIdx, rplsIdx ), where listIdx 821 identifies the reference picture list RPL0 811 and / or RPL1 812, and rplsIdx 825 identifies the entry 815 in the reference picture list. Thus, ref_pic_list_struct is a syntax structure that returns the entry 815 based on listIdx 821 and rplsIdx 825. The encoder can encode a part of the RPL structure 800 for each non-intra-coded slice in the video sequence. The decoder can then resolve the corresponding part of the RPL structure 800 before decoding each non-intra-coded slice in the coded video sequence.

[0085] As described above, the reference pictures can be referenced for inter prediction. Further, the reference pictures can be used as inter-layer reference pictures for inter-layer prediction. Thus, the RPL structure 800 is modified by including an ILRP flag 833 and an ILRP layer indicator 835 to support inter-layer prediction. The ILRP flag 833 is data indicating whether the picture referenced by the corresponding entry 815 of the RPL structure 800 is an inter-layer reference picture used for inter-layer prediction. Thus, the encoder can use the ILRP flag 833 to indicate whether each entry 815 should be treated as an ILRP entry. An ILRP entry is any entry 815 that references an inter-layer reference picture used for inter-layer prediction. Further, the decoder can use the ILRP flag 833 to determine whether each entry 815 in the RPL structure 800 is an ILRP entry. In a specific example, the ILRP flag 833 is indicated as inter_layer_ref_pic_flag[ listIdx ][ rplsIdx ][ i ], where i is a counter variable for which each value indicates the corresponding entry 815. When the i-th entry in ref_pic_list_struct( listIdx, rplsIdx ) is an ILRP entry, inter_layer_ref_pic_flag[ listIdx ][ rplsIdx ][ i ] can be set equal to 1. Further, when the i-th entry in ref_pic_list_struct( listIdx, rplsIdx ) is not an ILRP entry, inter_layer_ref_pic_flag[ listIdx ][ rplsIdx ][ i ] can be set equal to 0.

[0086] As described above, the inter-layer reference picture may be located within the same access unit as the currently encoded / decoded picture and may have the same POC value. Therefore, the RPL structure 800 may not need to be modified to add the POC value of the corresponding inter-layer reference picture. However, the decoder may not be able to infer which one or more layers of an access unit contain the appropriate inter-layer reference picture for decoding the current picture. For this purpose, the ILRP layer indicator 835 is included. The ILRP layer indicator 835 is data indicating the layer associated with the inter-layer reference picture referred to by the current picture. Specifically, the ILRP layer indicator 835 can indicate one or more layers for each entry 815 that is an ILRP entry as indicated by the ILRP flag 833. Therefore, the encoder can encode the ILRP layer indicator 835 into the bitstream to indicate the layer of the one or more inter-layer reference pictures associated with the entry 815. In addition, when the entry 815 is an ILRP entry as indicated by the ILRP flag 833, the decoder can determine the layer of the inter-layer reference picture based on the ILRP layer indicator 835. Therefore, the addition of the ILRP flag 833 and the ILRP layer indicator 835 to the RPL structure 800 provides sufficient information to enable the RPL structure 800 to manage the reference pictures for inter-layer prediction.

[0087] FIG. 9 is a schematic diagram showing an exemplary bitstream 900 including coding tool parameters to support inter-layer prediction. For example, the bitstream 900 may be generated by the codec system 200 and / or the encoder 300 for decoding by the codec system 200 and / or the decoder 400. As another example, the bitstream 900 may be generated by the encoder in step 109 of method 100 for use by the decoder in step 111. Further, the bitstream 900 may be an encoded video sequence that may be coded according to unidirectional inter prediction 500, bidirectional inter prediction 600, and / or layer-based prediction 700. Additionally, the bitstream 900 may be employed to communicate the RPL structure 800.

[0088] The bitstream 900 includes a sequence parameter set (SPS) 910, a plurality of picture parameter sets (PPS) 911, a plurality of slice headers 915, and picture data 920. The SPS 910 includes sequence data common to all pictures in the video sequence included in the bitstream 900. Such data can include picture size determination, bit depth, coding tool parameters, bitrate constraints, and the like. The PPS 911 includes parameters applied to the entire picture. Thus, each picture in the video sequence can refer to the PPS 911. Note that while each picture refers to the PPS 911, in some examples a single PPS 911 can include data for multiple pictures. For example, multiple similar pictures may be coded according to similar parameters. In such a case, a single PPS 911 may include data for such similar pictures. The PPS 911 can indicate coding tools, quantization parameters, offsets, etc. available for slices in the corresponding picture. The slice header 915 includes parameters specific to each slice in the picture. Thus, there can be one slice header 915 per slice in the video sequence. The slice header 915 can include slice type information, picture order count (POC), reference picture list, prediction weight, tile entry point, deblocking parameters, and the like. Note that in some contexts, the slice header 915 may also be referred to as a tile group header. In some examples, the bitstream 900 may also include a picture header, which is a syntax structure that includes parameters applied to all slices in a single picture. For this reason, in some contexts, the picture header and the slice header 915 can be used interchangeably. For example, some parameters can be moved between the slice header 915 and the picture header depending on whether such parameters are common to all slices in the picture.

[0089] The picture data 920 includes video data encoded according to inter prediction, intra prediction, and / or inter-layer prediction, as well as corresponding residual data that has been transformed and quantized. For example, a video sequence includes a plurality of pictures 923. A picture 923 is an array of luma samples and / or an array of chroma samples that creates a frame or a field of a frame. A frame is a complete image intended for a complete or partial display to the user at a corresponding instant in the video sequence. A picture 923 may be included within a single AU921. An AU921 is a coding unit configured to store all coded pictures 923 having the same picture order count, and optionally, one or more headers such as a slice header 915 that represents the coding mechanism employed to code the coded picture 923. Thus, an AU921 can include a single picture 923 for each picture level included in the bitstream 900. A picture 923 includes one or more slices 925. A slice 925 may be defined as an integral number of complete tiles of the picture 923 or an integral number of consecutive complete CTU rows (e.g., within a tile) that are exclusively included within a single NAL unit. A slice 925 is further divided into CTUs and / or coding tree blocks (CTBs). A CTU is a group of samples of a predefined size that can be partitioned by a coding tree. A CTB is a subset of a CTU and includes the luma or chroma components of the CTU. A CTU / CTB is further divided into coding blocks based on a coding tree. The coding blocks can then be encoded / decoded according to a prediction mechanism.

[0090] The bitstream 900 includes various coding tool parameters to support inter-layer prediction. Specifically, the bitstream includes a reference picture list structure 931, an ILRP flag 933, and an ILRP layer indicator 935, which may be substantially similar to the RPL structure 800, the ILRP flag 833, and the ILRP layer indicator 835, respectively. The reference picture list structure 931, the ILRP flag 933, and the ILRP layer indicator 935 may be coded in the SPS 910, the slice header 915, and / or the corresponding picture header. Thus, the encoder can enumerate the reference pictures among the entries in the reference picture list structure 931, can indicate in the ILRP flag 933 which of the reference pictures are inter-layer reference pictures, and can indicate in the ILRP layer indicator 935 which layer includes the associated inter-layer reference pictures. Further, the decoder can decode the SPS 910, the slice header 915, and / or the corresponding picture header to obtain the reference picture list structure 931, the ILRP flag 933, and the ILRP layer indicator 935. The decoder can then determine the reference pictures for the current picture from the entries in the reference picture list structure 931. The decoder can also determine which of the reference pictures are inter-layer reference pictures by adopting the ILRP flag 933. Further, the decoder can determine one or more levels associated with the inter-layer reference pictures by adopting the ILRP layer indicator 935. Thus, the bitstream 900 is configured to provide sufficient information for managing reference pictures when performing a combination of intra prediction, inter prediction, and / or inter-layer prediction to code a video sequence.

[0091] The foregoing information will be described in more detail below in this specification. To change the spatial resolution of a coded picture in the middle of a bitstream, a reference picture resampling (RPR) mechanism may be employed. This resolution change can be achieved in the current picture even when the current picture is not intra-coded. To enable this function, the current picture may refer to one or more reference pictures for inter-prediction purposes, provided that the reference pictures employ a spatial resolution different from that of the current picture. Thus, the encoding and decoding of the current picture may involve resampling such a reference picture or a portion of the reference picture. This function is sometimes referred to as adaptive resolution change (ARC). Resampling of the reference picture can be realized either at the picture level or at the coding block level.

[0092] Some implementations may benefit from RPR. For example, video telephony and video conferencing may adopt rate adaptation. The rate of the encoded video can be adapted to changing network conditions. When the network conditions deteriorate and the available bandwidth becomes smaller, the encoder may adapt by encoding the picture at a lower resolution picture. As another example, a change in the active speaker in a multiparty video conference may benefit from RPR. In a multiparty video conference, the active speaker may be shown at a larger video size than the videos for the other people among the conference participants. When the active speaker changes, the picture resolution for each participant can also be adjusted. When the active speaker changes more frequently, the RPR / ARC mechanism becomes increasingly more beneficial. In another example, a quick start in streaming may benefit from RPR. Streaming applications may buffer decoded pictures up to some length before starting to display. Starting the bitstream at a lower resolution may enable the application to buffer enough pictures to start displaying more quickly. Once the display has started, the resolution can then be increased. In another example, adaptive stream switching in streaming may benefit from RPR. Dynamic Adaptive Streaming over HTTP (DASH) adopts a function called @mediaStreamStructureId. This function enables switching between different representations at an open picture group (GOP) random access point having a non-decodable leading picture, sometimes called a clean random access (CRA) picture with an associated RASL picture in HEVC. For example, two different representations of the same video may have different bitrates and the same spatial resolution while they have the same value of @mediaStreamStructureId. In such a case, a switch between the two representations at the CRA picture with the associated RASL picture can be performed.The RASL pictures related to the CRA pictures can be decoded with acceptable quality, which enables seamless switching. When switching between DASH representations with different spatial resolutions, it becomes possible to adopt the @mediaStreamStructureId function by ARC / RPR.

[0093] ARC / RPR can be implemented by adopting hierarchical video coding, sometimes also called scalable video coding and / or video coding with scalability. Scalability in video coding can be supported by using multi-layer coding techniques. The multi-layer bitstream comprises a base layer (BL) and one or more enhancement layers (EL). Scalability may include spatial scalability, quality / signal-to-noise ratio (SNR) scalability, multi-view scalability, etc. When multi-layer coding techniques are adopted, a picture (or a part of a picture) can be coded (1) without using a reference picture (intra prediction), (2) by reference to one or more reference pictures within the same layer (inter prediction), or (3) by reference to one or more reference pictures within another layer (inter-layer prediction). The reference picture used for inter-layer prediction of the current picture is called an inter-layer reference picture (ILRP).

[0094] The H.26x video coding family may provide support for scalability in separate profiles from the profiles for single-layer coding. Scalable Video Coding (SVC) is a scalable extension of AVC / H.264 that provides support for spatial scalability, temporal scalability, and quality scalability. In the case of SVC, a flag is signaled within each macroblock (MB) in a picture to indicate whether an MB is predicted using a collocated block from a lower layer. Prediction from a collocated block may include a texture mode, a motion vector mode, and / or a coding mode.

[0095] Scalable HEVC (SHVC) is an extension of HEVC / H.265 that provides support for spatial scalability and quality scalability. Multi-View HEVC (MV-HEVC) is an extension of HEVC / H.265 that provides support for multi-view scalability. 3D HEVC (3D-HEVC) is an extension of HEVC / H.264 that provides support for 3D video coding. Temporal scalability may be adopted in a single-layer HEVC codec. The multi-layer extension of HEVC adopts a mechanism in which the decoded pictures used for inter-layer prediction are taken only from the same access unit and are treated as long-term reference pictures (LTRP). Such pictures are assigned reference indexes in the reference picture list together with other temporal reference pictures in the current layer. Inter-layer prediction (ILP) is achieved at the prediction unit (PU) level by setting the value of the reference index to refer to an inter-layer reference picture in the reference picture list.

[0096] When the ILRP has a spatial resolution different from that of the currently picture being symbolized or decoded, spatial scalability may employ resampling of the reference picture or a portion of the reference picture. Reference picture resampling may be implemented at either the picture level or the coding block level.

[0097] In a video codec specification, a picture may be identified for multiple purposes. For example, a picture may be identified for use as a reference picture in inter prediction, for output from the DPB, for scaling of motion vectors, for weighted prediction, etc. In some video coding systems, a picture may be identified by a picture order count (POC). Further, a picture in the DPB may be marked as being used for short-term reference, used for long-term reference, or not used for reference. A picture may no longer be used for prediction if it is marked as not being used for reference. When such a picture is no longer needed for output, the picture may be removed from the DPB.

[0098] AVC may adopt short-term reference pictures and long-term reference pictures. When a picture is no longer needed as a prediction reference, the reference picture may be marked as not being used for reference. The transition between these three statuses (short-term, long-term, and not used for reference) is controlled by the marking process of the decoded reference pictures. For reference picture marking in such a system, an implicit sliding window process and an explicit memory management control operation (MMCO) process may be adopted. The sliding window process marks the short-term reference pictures as not being used for reference when the number of reference frames is equal to the maximum number (max_num_ref_frames) in the SPS. The short-term reference pictures are stored in a first-in-first-out manner so that the most recently decoded short-term picture is held in the DPB. The explicit MMCO process may include a plurality of MMCO commands. The MMCO commands may mark one or more short-term reference pictures or long-term reference pictures as not being used for reference, may mark all pictures as not being used for reference, or may mark the current reference picture or an existing short-term reference picture as long-term and may assign a long-term picture index to that long-term reference picture. In AVC, the reference picture marking operation and the processes of output and removal of pictures from the DPB are performed after the pictures are decoded.

[0099] HEVC adopts a reference picture set (RPS) for reference picture marking. When adopting the RPS, a complete set of reference pictures used by the current picture or any subsequent picture is provided for each slice. Therefore, a complete set of all pictures that should be held in the DPB for use by the current or future pictures is signaled in the RPS. This is different from the AVC method in which only relative changes in the DPB are signaled. The RPS does not need to remember information from pictures earlier in the decoding order when maintaining the exact status of the reference pictures in the DPB.

[0100] In AVC, picture marking and buffer operations (both output and removal of decoded pictures from the DPB) may be applied after the current picture has been decoded. In HEVC, first the RPS is decoded from the slice header of the current picture. Then, picture marking and buffer operations may be applied before decoding the current picture.

[0101] VVC adopts reference picture list 0 and reference picture list 1 for reference picture management. Using that approach, the reference picture list for a picture is configured immediately without using a reference picture list initialization process and a reference picture list modification process. Furthermore, reference picture marking is performed immediately based on the two reference picture lists.

[0102] The aforementioned system has several problems. The VVC decoder should be able to derive inter-layer reference pictures in order to enable a multi-layer video codec with inter-layer prediction. To implement such a mechanism, the codec should adopt a mechanism for signaling the RPL, deriving the RPL, and performing reference picture marking in the multi-layer context.

[0103] The present disclosure includes several techniques for reference picture management in hierarchical video coding. This includes signaling of RPLs, derivation of RPLs, and reference picture marking. The description of the techniques is based on VVC, but applies to other hierarchical video coding specifications. For example, the present disclosure includes a mechanism for signaling of RPLs in a multi-layer video codec with inter-layer prediction based on VVC. The following constraints may apply to the examples described below. All VCL NAL units having the same layer ID and related to the same presentation time may comprise one picture. Further, all VCL NAL units within a picture may have the same POC value. Such POC values may also be referred to as the POC value of the picture.

[0104] The first example is summarized as follows. All pictures related to the same presentation time belong to one access unit. Pictures in different layers and within the same access unit have the same POC value. A first flag, such as inter_layer_ref_pics_flag, may be added to the SPS to specify whether ILRP is used for inter-prediction of any coded picture in the CVS. When the first flag specifies that ILRP can be used for inter-prediction of one or more coded pictures in the CVS, a second flag may be signaled for each entry in the RPL structure to specify whether the entry is an ILRP entry. When the second flag for an entry specifies that the entry is an ILRP entry, a first delta value may be signaled that specifies the difference -1 between the layer ID of the current picture and the picture referenced by the entry. In some examples, the first delta value specifies the difference -1 between the layer index of the layer containing the current picture and the layer index of the layer containing the picture referenced by the entry. At the start of decoding of the current slice, the RPL can be constructed by the decoder according to the RPL signaling in the bitstream. When the entry is an ILRP entry, the entry is derived to reference a picture that has the same PicOrderCntVal as the current picture and has a layer ID equal to the layer ID of the current picture - the first delta value - 1. At the start of decoding of the current picture, each ILRP (if present) is marked as being used for long-term reference. At the end of decoding of the current picture, each ILRP (if present) is marked as being used for short-term reference. When decoding pictures in the same layer, the decoded pictures may be marked only as not being used for reference.

[0105] The method of executing the above example is as follows. In one example, a method for decoding a video bitstream is disclosed. Each layer comprises a plurality of pictures, and the bitstream comprises a plurality of layers. One or more of the pictures belong to different layers and have the same presentation time, and thus form one access unit. The method comprises deriving a POC value for each picture. The POC values of all the pictures within one access unit are the same. Several RPLs are derived for the current slice. Each entry is associated with a flag that specifies whether the entry is an ILRP entry. The current slice is decoded based on the derived RPLs. In an exemplary aspect, the bitstream comprises an SPS flag that specifies whether an ILRP entry is used for inter prediction of any coded picture in the CVS. In an exemplary aspect, when an entry in the RPL is specified to be an ILRP entry, the bitstream comprises a first delta value that specifies a difference - 1 between the layer ID of the current picture containing the current slice and the picture referenced by the entry. In an exemplary aspect, the ILRP entry references a picture that has the same PicOrderCntVal as the current picture and has a layer ID equal to the layer ID of the current picture - the first delta value - 1. In an exemplary aspect, at the start of decoding of the current picture, each ILRP (if present) is marked as being used for long - term reference. At the end of decoding of the current picture, each ILRP (if present) is marked as being used for short - term reference.

[0106] A second way to execute the above example is as follows. In one example, all pictures related to the same presentation time belong to one access unit. Each picture may be related to an intra-layer POC value and an inter-layer POC value. Pictures within one CVS may have different inter-layer POC values. Pictures in different layers and within the same access unit may have the same intra-layer POC value, but may also have different inter-layer POC values. Any picture does not have to refer to a picture in a higher layer, and thus does not have to refer to a picture with a larger layer ID value. When decoding pictures in the same layer, the decoded pictures may be marked only as not being used for reference.

[0107] A third way to execute the above example is as follows. In one example, all pictures related to the same presentation time belong to one access unit. Pictures within a CVS may have different POC values. Pictures within one CVS may be identified only by the POC value. Any picture does not have to refer to a picture in a higher layer, and thus does not have to refer to a picture with a larger layer ID value. When decoding pictures in the same layer, the decoded pictures may be marked only as not being used for reference.

[0108] A fourth way to execute the above example is as follows. Each access unit may contain only one picture. Thus, any two pictures related to the same presentation time but belonging to different layers may belong to two different access units. Pictures within one CVS may have different POC values. Pictures within one CVS may be identified only by the POC value. Any picture does not have to refer to a picture in a higher layer, and thus does not have to refer to a picture with a larger layer ID value. Whether the current picture and the reference picture belong to the same layer or not, the reference picture may be marked as not being used for reference.

[0109] A first exemplary implementation of the above method is described below. The exemplary definitions are as follows. An ILRP is a picture in the same access unit as the current picture, having a nuh_layer_id smaller than the nuh_layer_id of the current picture, and marked as being used for long-term reference. A long-term reference picture (LTRP) is a picture having a nuh_layer_id equal to the nuh_layer_id of the current picture and marked as being used for long-term reference. A reference picture is a picture that is a short-term reference picture or a long-term reference picture or an inter-layer reference picture. A reference picture contains samples that can be used for inter prediction in the decoding process of subsequent pictures in decoding order. A short-term reference picture (STRP: short-term reference picture) is a picture having a nuh_layer_id equal to the nuh_layer_id of the current picture and marked as being used for short-term reference.

[0110] The exemplary sequence parameter set RBSP syntax is as follows.

[0111] [Table 1]

[0112] The exemplary reference picture list structure syntax is as follows.

[0113] [Table 2]

[0114] The semantic of an exemplary sequence parameter set RBSP is as follows. The long_term_ref_pics_flag may be set equal to 0 to specify that the LTRP is not used for inter prediction of any coded picture in the CVS. The long_term_ref_pics_flag may be set equal to 1 to specify that the LTRP may be used for inter prediction of one or more coded pictures in the CVS. The inter_layer_ref_pics_flag may be set equal to 0 to specify that the ILRP is not used for inter prediction of any coded picture in the CVS. The inter_layer_ref_pics_flag may be set equal to 1 to specify that the ILRP may be used for inter prediction of one or more coded pictures in the CVS.

[0115] The semantic of an exemplary general slice header is as follows. The slice_type specifies the coding type of the slice according to the following table.

[0116] [Table 3]

[0117] When the NalUnitType is a value of NalUnitType in the range of IDR_W_RADL to CRA_NUT including both end values, and the current picture is the first picture in the access unit, the slice_type should be equal to 2.

[0118] An exemplary reference picture list structure semantics is as follows. To specify that the i-th entry in the ref_pic_list_struct( listIdx, rplsIdx ) syntax structure is an ILRP entry, inter_layer_ref_pic_flag[ listIdx ][ rplsIdx ][ i ] may be set equal to 1. To specify that the i-th entry in the ref_pic_list_struct( listIdx, rplsIdx ) syntax structure is not an ILRP entry, inter_layer_ref_pic_flag[ listIdx ][ rplsIdx ][ i ] may be set equal to 0. When it does not exist, the value of inter_layer_ref_pic_flag[ listIdx ][ rplsIdx ][ i ] is assumed to be equal to 0. To specify that the i-th entry in the ref_pic_list_struct( listIdx, rplsIdx ) syntax structure is an STRP entry, st_ref_pic_flag[ listIdx ][ rplsIdx ][ i ] may be set equal to 1. To specify that the i-th entry in the ref_pic_list_struct( listIdx, rplsIdx ) syntax structure is an LTRP entry, st_ref_pic_flag[ listIdx ][ rplsIdx ][ i ] may be set equal to 0. When inter_layer_ref_pic_flag[ listIdx ][ rplsIdx ][ i ] is equal to 0 and st_ref_pic_flag[ listIdx ][ rplsIdx ][ i ] does not exist, the value of st_ref_pic_flag[ listIdx ][ rplsIdx ][ i ] may be assumed to be equal to 1.

[0119] The variable NumLtrpEntries[ listIdx ][ rplsIdx ] may be derived as follows. For (i = 0, NumLtrpEntries[listIdx][rplsIdx] = 0; i < num_ref_entries[listIdx][rplsIdx]; i++) if (!inter_layer_ref_pic_flag[listIdx][rplsIdx][i] &&!st_ref_pic_flag[listIdx][rplsIdx][i]) NumLtrpEntries[listIdx][rplsIdx]++

[0120] In order to specify that the i-th entry in the syntax structure ref_pic_list_struct(listIdx, rplsIdx) has a value greater than or equal to 0, strp_entry_sign_flag[listIdx][rplsIdx][i] may be set equal to 1. In order to specify that the i-th entry in the syntax structure ref_pic_list_struct(listIdx, rplsIdx) has a value less than 0, strp_entry_sign_flag[listIdx][rplsIdx][i] may be set equal to 0. When it does not exist, the value of strp_entry_sign_flag[listIdx][rplsIdx][i] may be presumed to be equal to 1.

[0121] List DeltaPocSt[listIdx][rplsIdx] may be derived as follows. for (i = 0; i < num_ref_entries[listIdx][rplsIdx]; i++) if (!inter_layer_ref_pic_flag[listIdx][rplsIdx][i] && st_ref_pic_flag[listIdx][rplsIdx][i]) { (7-93) DeltaPocSt[listIdx][rplsIdx][i] = (strp_entry_sign_flag[listIdx][rplsIdx][i])? abs_delta_poc_st[listIdx][rplsIdx][i] : 0 - abs_delta_poc_st[listIdx][rplsIdx][i]

[0122] rpls_entry_layer_id_delta_minus1[listIdx][rplsIdx][i] + 1 specifies the difference between the nuh_layer_id of the current picture and the picture referred to by the i-th entry. The value of rpls_entry_layer_id_delta_minus1[listIdx][rplsIdx][i] may be in the range of 0 to 125, inclusive of both end values.

[0123] An exemplary decoding process for a coded picture is as follows. The decoding process operates as follows for the current picture CurrPic. The NAL unit is decoded. The following decoding process may use the syntax from the following elements in the slice header layer and above. Variables and functions related to the picture order count are derived. This may be called only for the first slice of a picture. At the start of the decoding process for each slice of an instantaneous decoding refresh (IDR) picture, a decoding process for reference picture list construction may be called for the derivation of reference picture list 0 (RefPicList

[0000] ) and reference picture list 1 (RefPicList

[0001] ). A decoding process for reference picture marking is called, and the reference pictures may be marked as not being used for reference or being used for long-term reference. This may be called only for the first slice of a picture. When the current picture is a CRA picture with NoIncorrectPicOutputFlag equal to 1, or a Gradual Random Access (GRA) picture with NoIncorrectPicOutputFlag equal to 1, a decoding process for generating unavailable reference pictures is called, and such a process may be called only for the first slice of a picture. PictureOutputFlag may be set as follows. PictureOutputFlag may be set equal to 0 if one of the following conditions is true. When the current picture is a RASL picture and the NoIncorrectPicOutputFlag of the related IRAP picture is equal to 1, PictureOutputFlag may be set equal to 0. When gra_enabled_flag is equal to 1 and the current picture is a GRA picture with NoIncorrectPicOutputFlag equal to 1, PictureOutputFlag may be set equal to 0.When gra_enabled_flag is equal to 1, the current picture is related to a GRA picture having a NoIncorrectPicOutputFlag equal to 1, and the PicOrderCntVal of the current picture is smaller than the RpPicOrderCntVal of the related GRA picture, PictureOutputFlag may be set equal to 0. Otherwise, PictureOutputFlag is set equal to 1. After all slices of the current picture have been decoded, the current decoded picture is marked as being used for short-term reference, and each ILRP entry in RefPicList

[0000] or RefPicList

[0001] is marked as being used for short-term reference.

[0124] An exemplary decoding process for constructing a reference picture list is as follows. This process is called at the start of the decoding process for each slice of a non-IDR picture. The reference pictures are addressed through reference indices. A reference index is an index into the reference picture list. When decoding an intracoding (I) slice, the reference picture list is not used in decoding the slice data. When decoding a uni-directional inter prediction (P) slice, only the reference picture list 0 (RefPicList

[0000] ) is used in decoding the slice data. When decoding a bi-directional inter prediction (B) slice, both the reference picture list 0 and the reference picture list 1 (RefPicList

[0001] ) are used in decoding the slice data. At the start of the decoding process for each slice of a non-IDR picture, the reference picture lists RefPicList

[0000] and RefPicList

[0001] are derived. The reference picture lists are used in marking the reference pictures or in decoding the slice data. For an I slice of a non-IDR picture that is not the first slice of the picture, RefPicList

[0000] and RefPicList

[0001] may be derived for bitstream conformance checking purposes, but their derivation is not required for decoding the current picture or pictures subsequent to the current picture in decoding order. For a P slice that is not the first slice of the picture, RefPicList

[0001] may be derived for bitstream conformance checking purposes, but it is not required for decoding the current picture or pictures subsequent to the current picture in decoding order.

[0125] The reference picture lists RefPicList

[0000] and RefPicList

[0001] may be constructed as follows. for( i = 0; i < 2; i++ ) { for( j = 0, k = 0, pocBase = PicOrderCntVal; j < num_ref_entries[ i ][ RplsIdx[ i ] ]; j++) { if( !inter_layer_ref_pic_flag[ i ][ RplsIdx[ i ] ][ j ] ) { if( st_ref_pic_flag[ i ][ RplsIdx[ i ] ][ j ] ) { RefPicPocList[ i ][ j ] = pocBase - DeltaPocSt[ i ][ RplsIdx[ i ] ][ j ] if( there is a reference picture picA in the DPB that has the same nuh_layer_id as the current picture and a PicOrderCntVal equal to RefPicPocList[ i ][ j ] ) RefPicList[ i ][ j ] = picA else RefPicList[ i ][ j ] = "There is no reference picture" (8 - 5) pocBase = RefPicPocList[ i ][ j ] } else { if( !delta_poc_msb_cycle_lt[ i ][ k ] ) { if( there is a reference picA in the DPB that has the same nuh_layer_id as the current picture and PicOrderCntVal & ( MaxPicOrderCntLsb - 1 ) is equal to PocLsbLt[ i ][ k ] ) RefPicList[ i ][ j ] = picA else RefPicList[ i ][ j ] = "There is no reference picture" } else { if( there is a reference picA in the DPB that has the same nuh_layer_id as the current picture and a PicOrderCntVal equal to FullPocLt[ i ][ k ] ) RefPicList[ i ][ j ] = picA else RefPicList[ i ][ j ] = "There is no reference picture" } k++ } } else { refPicLayerId = nuh_layer_id - rpls_entry_layer_id_delta_minus1[ i ][ RplsIdx[ i ] ][ j ] - 1 if (there is a reference picture picA in the DPB that has the same nuh_layer_id equal to refPicLayerId and the same PicOrderCntVal as the current picture) RefPicList[ i ][ j ] = picA else RefPicList[ i ][ j ] = "There is no reference picture" } } } }

[0126] For each i equal to 0 or 1, the first NumRefIdxActive[ i ] entries in RefPicList[ i ] are referred to as the active entries in RefPicList[ i ], and the other entries in RefPicList[ i ] are referred to as the non - active entries in RefPicList[ i ]. It is possible for a particular picture to be referred to by entries in both RefPicList

[0000] and RefPicList

[0001] . It is also possible for a particular picture to be referred to by more than two entries in RefPicList

[0000] or by more than two entries in RefPicList

[0001] . The active entries in RefPicList

[0000] and the active entries in RefPicList

[0001] collectively refer to all the reference pictures that can be used for the inter - prediction of the current picture and one or more pictures following the current picture in decoding order. The non - active entries in RefPicList

[0000] and the non - active entries in RefPicList

[0001] collectively refer to all the reference pictures that are not used for the inter - prediction of the current picture but can be used in the inter - prediction for one or more pictures following the current picture in decoding order. Since the corresponding picture does not exist in the DPB, there may be one or more entries in RefPicList

[0000] or RefPicList

[0001] that are not equal to any reference picture. Each non - active entry in RefPicList

[0000] or RefPicList

[0001] that is not equal to any reference picture should be ignored. For each active entry in RefPicList

[0000] or RefPicList

[0001] that is not equal to any reference picture, an accidental picture loss should be inferred.

[0127] For bitstream conformity, the following constraints should apply. For each i equal to 0 or 1, num_ref_entries[ i ][ RplsIdx[ i ] ] should not be less than NumRefIdxActive[ i ]. Each picture referenced by an active entry in RefPicList

[0000] or RefPicList

[0001] should be present in the DPB and should have a TemporalId less than or equal to the TemporalId of the current picture. Each picture referenced by an entry in RefPicList

[0000] or RefPicList

[0001] should not be the current picture. The STRP entry in RefPicList

[0000] or RefPicList

[0001] of a slice of a picture, and the LTRP entry in RefPicList

[0000] or RefPicList

[0001] of the same slice or a different slice of the same picture, should not reference the same picture. There should not be an LTRP entry in RefPicList

[0000] or RefPicList

[0001] for which the difference between the PicOrderCntVal of the current picture and the PicOrderCntVal of the picture referenced by the entry is 224 or more. Let setOfRefPics be the set of unique pictures referenced by all entries in RefPicList

[0000] that have the same nuh_layer_id as the current picture, and all entries in RefPicList

[0001] that have the same nuh_layer_id as the current picture. The number of pictures in setOfRefPics should be less than or equal to sps_max_dec_pic_buffering_minus1, and setOfRefPics should be the same for all slices of the picture. Each picture referenced by an ILRP entry in RefPicList

[0000] or RefPicList

[0001] of a slice of the current picture should be in the same access unit as the current picture.The pictures referred to by each ILRP entry in RefPicList

[0000] or RefPicList

[0001] of the current picture slice should be present in the DPB and should have a nuh_layer_id smaller than that of the current picture. Each ILRP entry in RefPicList

[0000] or RefPicList

[0001] of the slice should be an active entry.

[0128] An exemplary decoding process for reference picture marking is as follows. This process is called once per picture after the decoding of the slice header and the decoding process for the reference picture list construction for the slice, but before the decoding of the slice data. This process can result in one or more of the reference pictures in the DPB being marked as not used for reference or marked as being used for long-term reference. The decoded pictures in the DPB can be marked as not used for reference, used for short-term reference, or used for long-term reference. However, a decoded picture can be marked as only one of these three at any given instant during the operation of the decoding process. Assigning one of these markings to a picture implicitly removes the others when applicable. When a picture is referred to as being marked as used for reference, this collectively refers to the picture being marked as used for short-term reference or used for long-term reference, but not both. The STRP and ILRP may be identified by their nuh_layer_id and PicOrderCntVal values. The LTRP may be identified by their nuh_layer_id values and the Log2(MaxLtPicOrderCntLsb) least significant bits (LSBs) of their PicOrderCntVal values. If the current picture is the coded layer video sequence start (CLVSS) picture, then all the reference pictures in the current DPB that have the same nuh_layer_id as the current picture, if any, are marked as not used for reference. Otherwise, the following applies. For each LTRP entry in RefPicList

[0000] or RefPicList

[0001] , when the picture being referred to is an STRP that has the same nuh_layer_id as the current picture, the picture is marked as being used for long-term reference.Each reference picture having the same nuh_layer_id as the current picture in the DPB that is not referenced by any entry in RefPicList

[0000] or RefPicList

[0001] is marked as not used for reference. For each ILRP entry in RefPicList

[0000] or RefPicList

[0001] , the referenced picture is marked as being used for long-term reference.

[0129] A second exemplary implementation of the above method is described below. Exemplary definitions are as follows. An ILRP is a picture in the same access unit as the current picture, having a nuh_layer_id smaller than that of the current picture, and marked as being used for long-term reference. An LTRP is a picture having a nuh_layer_id equal to the nuh_layer_id of the current picture and marked as being used for long-term reference. A reference picture is a picture that is a short-term reference picture or a long-term reference picture or an inter-layer reference picture. A reference picture contains samples that can be used for inter prediction in the decoding process of subsequent pictures in decoding order. An STRP is a picture having a nuh_layer_id equal to the nuh_layer_id of the current picture and marked as being used for short-term reference.

[0130] An exemplary video parameter set syntax is as follows.

[0131]

Table 4

[0132] An exemplary video parameter set RBSP semantics is as follows. vps_video_parameter_set_id provides an identifier for the VPS for reference by other syntax elements. vps_max_layer_id specifies the maximum allowable value of nuh_layer_id in each CVS that refers to a video parameter set (VPS). Once determined, it is expected that the value of vps_max_layer_id remains unchanged when generating the original bitstream. This includes cases where the original bitstream is rewritten or the rewritten bitstream is further rewritten. Otherwise, the picture order count value may be interrupted and unexpected behavior may occur. Alternatively, instead, the maximum allowable value of nuh_layer_id, which is signaled in the SPS and referred to as sps_max_layer_id.

[0133] An exemplary general slice header semantics is as follows. slice_type specifies the coding type of the slice according to the following table.

[0134] [Table 5]

[0135] When the NalUnitType is a value of NalUnitType in the range of IDR_W_RADL to CRA_NUT including both end values, and the current picture is the first picture in the access unit, slice_type should be equal to 2. slice_pic_order_cnt_lsb specifies the value of PicOrderCntVal modulo MaxPicOrderCntLsb for the current picture. The length of the slice_pic_order_cnt_lsb syntax element is log2_max_pic_order_cnt_lsb_minus4 + 4 bits. The value of slice_pic_order_cnt_lsb should be in the range of 0 to MaxPicOrderCntLsb - 1 including both end values. slice_poc_lsb_lt[ i ][ j ] specifies the value of PicOrderCntVal modulo MaxPicOrderCntLsb for the j-th LTRP entry in the i-th reference picture list. The length of the slice_poc_lsb_lt[ i ][ j ] syntax element is log2_max_pic_order_cnt_lsb_minus4 + 4 bits. st_ref_pic_flag[ listIdx ][ rplsIdx ][ i ] may be set equal to 1 to specify that the i-th entry in the ref_pic_list_struct( listIdx, rplsIdx ) syntax structure is a non-LTRP entry (STRP entry or ILRP entry). st_ref_pic_flag[ listIdx ][ rplsIdx ][ i ] may be set equal to 0 to specify that the i-th entry in the ref_pic_list_struct( listIdx, rplsIdx ) syntax structure is an LTRP entry. When it does not exist, the value of st_ref_pic_flag[ listIdx ][ rplsIdx ][ i ] can be presumed to be equal to 1.

[0136] The variable NumLtrpEntries[ listIdx ][ rplsIdx ] may be derived as follows. for( i = 0, NumLtrpEntries[ listIdx ][ rplsIdx ] = 0; i < num_ref_entries[ listIdx ][ rplsIdx ]; i++ ) if( !st_ref_pic_flag[ listIdx ][ rplsIdx ][ i ] ) (7-87) NumLtrpEntries[ listIdx ][ rplsIdx ]++

[0137] abs_delta_poc_st[ listIdx ][ rplsIdx ][ i ] specifies the absolute difference between the CrossLayerPoc value of the current picture and the CrossLayerPoc value of the picture referred to by the i-th entry when the i-th entry is the first non-LTRP entry in the ref_pic_list_struct(listIdx, rplsIdx) syntax structure. abs_delta_poc_st[ listIdx ][ rplsIdx ][ i ] specifies the absolute difference between the CrossLayerPoc value of the picture referred to by the i-th entry and the CrossLayerPoc value of the picture referred to by the previous non-LTRP entry in the ref_pic_list_struct(listIdx, rplsIdx) syntax structure when the i-th entry is a non-LTRP entry but not the first non-LTRP entry in the ref_pic_list_struct(listIdx, rplsIdx) syntax structure. The value of abs_delta_poc_st[ listIdx ][ rplsIdx ][ i ] is in the range of 0 to (2 15- It may be within the range of (vps_max_layer_id + 1). To specify that the i-th entry in the syntax structure ref_pic_list_struct(listIdx, rplsIdx) has a value of 0 or greater, entry_sign_flag[listIdx][rplsIdx][i] may be set equal to 1. To specify that the i-th entry in the syntax structure ref_pic_list_struct(listIdx, rplsIdx) has a value less than 0, entry_sign_flag[listIdx][rplsIdx] may be set equal to 0. When it does not exist, the value of entry_sign_flag[i][j] is presumed to be equal to 1.

[0138] List DeltaPoc[listIdx][rplsIdx] can be derived as follows. for (i = 0; i < num_ref_entries[listIdx][rplsIdx]; i++) { if (st_ref_pic_flag[listIdx][rplsIdx][i]) { (7-88) DeltaPoc[listIdx][rplsIdx][i] = (entry_sign_flag[listIdx][rplsIdx][i])? abs_delta_poc_st[listIdx][rplsIdx][i] : 0 - abs_delta_poc_st[listIdx][rplsIdx][i] } }

[0139] rpls_poc_lsb_lt[ listIdx ][ rplsIdx ][ i ] specifies the value of PicOrderCntVal modulo MaxPicOrderCntLsb for the picture referred to by the i-th entry in the ref_pic_list_struct( listIdx, rplsIdx ) syntax structure. The length of the rpls_poc_lsb_lt[ listIdx ][ rplsIdx ][ i ] syntax element is log2_max_pic_order_cnt_lsb_minus4 + 4 bits.

[0140] An exemplary decoding process for a coded picture is as follows. The decoding process operates on the current picture CurrPic as follows. Decoding of the NAL unit is as specified herein. The following decoding process uses syntax elements in layers above the slice header layer. Variables and functions related to the picture order count are derived. Such derivation may be called only for the first slice of a picture. At the start of the decoding process for each slice of a non-IDR picture, a decoding process for the reference picture list construction is called for the derivation of reference picture list 0 (RefPicList

[0000] ) and reference picture list 1 (RefPicList

[0001] ). A decoding process for reference picture marking is called. The reference pictures may be marked as not used for reference or used for long-term reference. This mechanism may be called only for the first slice of a picture. When the current picture is a CRA picture with NoIncorrectPicOutputFlag equal to 1, or a GRA picture with NoIncorrectPicOutputFlag equal to 1, a decoding process for generating unavailable reference pictures is called. This process may be called only for the first slice of a picture.

[0141] The PictureOutputFlag may be set as follows. If one of the following conditions is true, the PictureOutputFlag is set equal to 0. When the current picture is a RASL picture and the NoIncorrectPicOutputFlag of the related IRAP picture is equal to 1, the PictureOutputFlag is set equal to 0. When gra_enabled_flag is equal to 1 and the current picture is a GRA picture having a NoIncorrectPicOutputFlag equal to 1, the PictureOutputFlag is set equal to 0. When gra_enabled_flag is equal to 1, the current picture is related to a GRA picture having a NoIncorrectPicOutputFlag equal to 1, and the PicOrderCntVal of the current picture is less than the RpPicOrderCntVal of the related GRA picture, the PictureOutputFlag is set equal to 0. Otherwise, the PictureOutputFlag is set equal to 1. After all slices of the current picture have been decoded, the current decoded picture may be marked as being used for short-term reference, and each ILRP entry in RefPicList

[0000] or RefPicList

[0001] may be marked as being used for short-term reference.

[0142] An exemplary decoding process for picture order count is as follows. The output of this process is CrossLayerPoc, which is the picture order count of the current picture. Each coded picture is associated with a picture order count variable denoted as CrossLayerPoc. When the current picture is not a CLVSS picture, the variables prevPicOrderCntLsb and prevPicOrderCntMsb can be derived as follows. Let prevTid0Pic be the previous picture in decoding order that has an nuh_layer_id equal to the nuh_layer_id of the current picture and a TemporalId equal to 0, and is not a RASL picture or a RADL picture. The variable prevPicOrderCntLsb may be set equal to slice_pic_order_cnt_lsb of prevTid0Pic. The variable prevPicOrderCntMsb may be set equal to PicOrderCntMsb of prevTid0Pic. The variable PicOrderCntMsb of the current picture may be derived as follows. If the current picture is a CLVSS picture, PicOrderCntMsb may be set equal to 0. Otherwise, PicOrderCntMsb may be derived as follows. if( ( slice_pic_order_cnt_lsb < prevPicOrderCntLsb ) && ( ( prevPicOrderCntLsb - slice_pic_order_cnt_lsb ) >= ( MaxPicOrderCntLsb / 2 ) ) ) PicOrderCntMsb = prevPicOrderCntMsb + MaxPicOrderCntLsb else if( (slice_pic_order_cnt_lsb > prevPicOrderCntLsb ) && ( ( slice_pic_order_cnt_lsb - prevPicOrderCntLsb ) > ( MaxPicOrderCntLsb / 2 ) ) ) PicOrderCntMsb = prevPicOrderCntMsb - MaxPicOrderCntLsb else PicOrderCntMsb = prevPicOrderCntMsb

[0143] PicOrderCntVal can be derived as follows. PicOrderCntVal = PicOrderCntMsb + slice_pic_order_cnt_lsb For CLVSS pictures, all CLVSS pictures may have a PicOrderCntVal equal to slice_pic_order_cnt_lsb since PicOrderCntMsb is set equal to 0. The value of PicOrderCntVal may be in the range of -2 31 ~2 31 -1 including both end values. In one CVS, the PicOrderCntVal values for any two coded pictures having the same value of nuh_layer_id need not be the same. All pictures in any particular access unit may have the same value of PicOrderCntVal.

[0144] The variable CrossLayerPoc may be derived as follows. CrossLayerPoc = PicOrderCntVal * (vps_max_layer_id + 1) + nuh_layer_id The function PicOrderCnt(picX) may be specified as follows. PicOrderCnt(picX) = PicOrderCntVal of picture picX The function DiffPicOrderCnt(picA, picB) may be specified as follows. DiffPicOrderCnt(picA, picB) = PicOrderCnt(picA) - PicOrderCnt(picB) The bitstream may not include data that results in a value of DiffPicOrderCnt( picA, picB ) used in the decoding process that is not within the range of -2 15 ~2 15 -1. Let the current picture be X, and let two other pictures in the same CVS be Y and Z. When both DiffPicOrderCnt( X, Y ) and DiffPicOrderCnt( X, Z ) are positive or both are negative, Y and Z may be considered to be in the same output order direction from X.

[0145] An exemplary decoding process for constructing a reference picture list is as follows. This process may be invoked at the start of the decoding process for each slice of a non-IDR picture. Reference pictures are addressed through reference indices. A reference index is an index into the reference picture list. When decoding an I slice, the reference picture list is not used in decoding the slice data. When decoding a P slice, only reference picture list 0 (RefPicList

[0000] ) is used in decoding the slice data. When decoding a B slice, both reference picture list 0 and reference picture list 1 (RefPicList

[0001] ) are used in decoding the slice data. At the start of the decoding process for each slice of a non-IDR picture, reference picture lists RefPicList

[0000] and RefPicList

[0001] are derived. The reference picture lists may be used in marking reference pictures or in decoding slice data. For an I slice of a non-IDR picture that is not the first slice of the picture, RefPicList

[0000] and RefPicList

[0001] may be derived for bitstream conformance checking purposes. However, their derivation may not be essential for decoding the current picture or pictures subsequent to the current picture in decoding order. For a P slice that is not the first slice of the picture, RefPicList

[0001] may be derived for bitstream conformance checking purposes. However, the derivation of RefPicList

[0001] may not be essential for decoding the current picture or pictures subsequent to the current picture in decoding order. The reference picture lists RefPicList

[0000] and RefPicList

[0001] may be configured as follows. for( i = 0; i < 2; i++ ) { for( j = 0, k = 0, pocBase = CrossLayerPoc; j < num_ref_entries[ i ][ RplsIdx[ i ] ]; j++) { if( st_ref_pic_flag[ i ][ RplsIdx[ i ] ][ j ] ) { RefPicPocList[ i ][ j ] = pocBase - DeltaPoc[ i ][ RplsIdx[ i ] ][ j if(there is a reference picture picA in the DPB that has a CrossLayerPoc equal to RefPicPocList[ i ][ j ]) RefPicList[ i ][ j ] = picA else RefPicList[ i ][ j ] = "no reference picture" (8-5) pocBase = RefPicPocList[ i ][ j ] } else { if( !delta_poc_msb_cycle_lt[ i ][ k ] ) { if(there is a reference picA in the DPB that has the same nuh_layer_id as the current picture and PicOrderCntVal & (MaxPicOrderCntLsb - 1) is equal to PocLsbLt[ i ][ k ]) RefPicList[ i ][ j ] = picA else RefPicList[ i ][ j ] = "no reference picture" } else { if(there is a reference picA in the DPB that has the same nuh_layer_id as the current picture and a PicOrderCntVal equal to FullPocLt[ i ][ RplsIdx[ i ] ][ k ]) RefPicList[ i ][ j ] = picA else RefPicList[ i ][ j ] = "no reference picture" } k++ } } }

[0146] For each i equal to 0 or 1, the first NumRefIdxActive[ i ] entries in RefPicList[ i ] are referred to as active entries in RefPicList[ i ], and the other entries in RefPicList[ i ] are referred to as non - active entries in RefPicList[ i ]. When the picture referred to by an entry has a nuh_layer_id with a value different from that of the current picture, the entry in the reference picture list may be referred to as an ILRP entry. It is possible for a particular picture to be referred to by both an entry in RefPicList

[0000] and an entry in RefPicList

[0001] . It is also possible for a particular picture to be referred to by more than two entries in RefPicList

[0000] or by more than two entries in RefPicList

[0001] . The active entries in RefPicList

[0000] and the active entries in RefPicList

[0001] collectively refer to all the reference pictures that can be used for inter - prediction of the current picture and one or more pictures following the current picture in decoding order. The non - active entries in RefPicList

[0000] and the non - active entries in RefPicList

[0001] collectively refer to all the reference pictures that are not used for inter - prediction of the current picture but can be used in inter - prediction for one or more pictures following the current picture in decoding order. Since the corresponding picture does not exist in the DPB, there may be one or more entries in RefPicList

[0000] or RefPicList

[0001] that are not equal to any reference picture. Each non - active entry in RefPicList

[0000] or RefPicList

[0000] that is not equal to any reference picture may be ignored. For each active entry in RefPicList

[0000] or RefPicList

[0001] that is not equal to any reference picture, an accidental picture loss may be inferred.

[0147] Bitstream conformance may require that the following constraints apply. For each i equal to 0 or 1, num_ref_entries[ i ][ RplsIdx[ i ] ] need not be less than NumRefIdxActive[ i ]. The picture referenced by each active entry in RefPicList

[0000] or RefPicList

[0001] may be present in the DPB and may have a TemporalId less than or equal to the TemporalId of the current picture. The picture referenced by each entry in RefPicList

[0000] or RefPicList

[0001] need not be the current picture. The STRP entry in RefPicList

[0000] or RefPicList

[0001] of a slice of a picture, and the LTRP entry in RefPicList

[0000] or RefPicList

[0001] of the same slice or a different slice of the same picture need not reference the same picture. There need not be an LTRP entry in RefPicList

[0000] or RefPicList

[0001] for which the difference between the PicOrderCntVal of the current picture and the PicOrderCntVal of the picture referenced by the entry is 224 or more. Let setOfRefPics be the set of unique pictures referenced by all entries in RefPicList

[0000] that have the same nuh_layer_id as the current picture, and all entries in RefPicList

[0001] that have the same nuh_layer_id as the current picture. The number of pictures in setOfRefPics may be less than or equal to sps_max_dec_pic_buffering_minus1, and setOfRefPics may be the same for all slices of a picture. The picture referenced by each ILRP entry in RefPicList

[0000] or RefPicList

[0001] of a slice of the current picture may be present in the DPB and may be in the same access unit as the current picture.The pictures referenced by each ILRP entry in RefPicList

[0000] or RefPicList

[0001] of the current picture slice may have a nuh_layer_id smaller than the nuh_layer_id of the current picture. Each ILRP entry in RefPicList

[0000] or RefPicList

[0001] of the slice may be an active entry.

[0148] An exemplary decoding process for reference picture marking is as follows. This process may be called once per picture. This may be done after decoding of the slice header and the decoding process for reference picture list construction for the slice, but before decoding of the slice data. This process may result in one or more reference pictures in the DPB being marked as not used for reference or marked as being used for long-term reference. Decoded pictures in the DPB may be marked as not used for reference, used for short-term reference, or marked as being used for long-term reference. Decoded pictures in the DPB may be marked as only one of these three at any given instant during the operation of the decoding process. Assigning one of these markings to a picture implicitly removes another of these markings when applicable. When a picture is referred to as being marked as being used for reference, this collectively refers to the picture being marked as being used for short-term reference or long-term reference (but not both). STRPs are identified by their nuh_layer_id and PicOrderCntVal values. LTRPs are identified by their nuh_layer_id values and the Log2(MaxLtPicOrderCntLsb) least significant bits of their PicOrderCntVal values. If the current picture is a CLVSS picture, all reference pictures in the current DPB (if any) having the same nuh_layer_id as the current picture may be marked as not used for reference. Otherwise, the following applies. For each LTRP entry in RefPicList

[0000] or RefPicList

[0001] , when the picture being referred to is an STRP having the same nuh_layer_id as the current picture, the picture being referred to is marked as being used for long-term reference.Each reference picture in the DPB that is not referenced by any entry in RefPicList

[0000] or RefPicList

[0001] and has the same nuh_layer_id as the current picture may be marked as not used for reference. For each ILRP entry in RefPicList

[0000] or RefPicList

[0001] , the referenced picture may be marked as used for long-term reference.

[0149] A third exemplary implementation of the above method is described below. An example of the definition is as follows. ILRP is a picture in the same access unit as the current picture, has a nuh_layer_id smaller than the nuh_layer_id of the current picture, and is marked as used for long-term reference. LTRP is a picture that has a nuh_layer_id equal to the nuh_layer_id of the current picture and is marked as used for long-term reference. A reference picture is a picture that is a short-term reference picture or a long-term reference picture or an inter-layer reference picture. A reference picture contains samples that can be used for inter prediction in the decoding process of subsequent pictures in decoding order. STRP is a picture that has a nuh_layer_id equal to the nuh_layer_id of the current picture and is marked as used for short-term reference.

[0150] Exemplary general slice header semantics are as follows. slice_type may specify the coding type of the slice according to the following table.

[0151]

Table 6

[0152] When the NalUnitType is a value of NalUnitType within the range of IDR_W_RADL to CRA_NUT including both end values, and the current picture is the first picture in the access unit, the Slice_type may be set equal to 2.

[0153] An exemplary decoding process for a coded picture is as follows. The decoding process operates on the current picture CurrPic as follows. The decoding of NAL units is as specified herein. The following decoding process uses syntax elements in the slice header layer and above. Variables and functions related to the picture order count are derived. Such derivation may be called only for the first slice of a picture. At the start of the decoding process for each slice of a non-IDR picture, a decoding process for reference picture list construction is called for the derivation of reference picture list 0 (RefPicList

[0000] ) and reference picture list 1 (RefPicList

[0001] ). A decoding process for reference picture marking is called. The reference pictures may be marked as not used for reference or used for long-term reference. This mechanism may be called only for the first slice of a picture. When the current picture is a CRA picture with NoIncorrectPicOutputFlag equal to 1, or a GRA picture with NoIncorrectPicOutputFlag equal to 1, a decoding process for generating unavailable reference pictures is called. This process may be called only for the first slice of a picture.

[0154] The PictureOutputFlag may be set as follows. If one of the following conditions is true, the PictureOutputFlag is set equal to 0. When the current picture is a RASL picture and the NoIncorrectPicOutputFlag of the related IRAP picture is equal to 1, the PictureOutputFlag is set equal to 0. When gra_enabled_flag is equal to 1 and the current picture is a GRA picture having a NoIncorrectPicOutputFlag equal to 1, the PictureOutputFlag is set equal to 0. When gra_enabled_flag is equal to 1, the current picture is related to a GRA picture having a NoIncorrectPicOutputFlag equal to 1, and the PicOrderCntVal of the current picture is less than the RpPicOrderCntVal of the related GRA picture, the PictureOutputFlag is set equal to 0. Otherwise, the PictureOutputFlag is set equal to 1. After all slices of the current picture have been decoded, the current decoded picture may be marked as being used for short-term reference, and each ILRP entry in RefPicList

[0000] or RefPicList

[0001] may be marked as being used for short-term reference.

[0155] An exemplary decoding process for picture order count is as follows. The output of this process is PicOrderCntVal, which is the picture order count of the current picture. Each coded picture may be associated with a picture order count variable denoted as PicOrderCntVal. When the current picture is not a CLVSS picture that is the first picture in the access unit in decoding order, the variables prevPicOrderCntLsb and prevPicOrderCntMsb may be derived as follows. Let prevTid0Pic be the previous picture in decoding order that has a TemporalId equal to 0 and a nuh_layer_id less than or equal to the nuh_layer_id of the current picture and is not a RASL picture or a RADL picture. The variable prevPicOrderCntLsb may be set equal to slice_pic_order_cnt_lsb of prevTid0Pic. The variable prevPicOrderCntMsb may be set equal to PicOrderCntMsb of prevTid0Pic. The variable PicOrderCntMsb of the current picture may be derived as follows. If the current picture is a CLVSS picture that is the first picture in the access unit in decoding order, PicOrderCntMsb may be set equal to 0. Otherwise, PicOrderCntMsb is derived as follows. if( ( slice_pic_order_cnt_lsb < prevPicOrderCntLsb ) && ( ( prevPicOrderCntLsb - slice_pic_order_cnt_lsb ) >= ( MaxPicOrderCntLsb / 2 ) ) ) PicOrderCntMsb = prevPicOrderCntMsb + MaxPicOrderCntLsb else if( (slice_pic_order_cnt_lsb > prevPicOrderCntLsb ) && ( ( slice_pic_order_cnt_lsb - prevPicOrderCntLsb ) > ( MaxPicOrderCntLsb / 2 ) ) ) PicOrderCntMsb = prevPicOrderCntMsb - MaxPicOrderCntLsb else PicOrderCntMsb = prevPicOrderCntMsb

[0156] PicOrderCntVal may be derived as follows. PicOrderCntVal = PicOrderCntMsb + slice_pic_order_cnt_lsb After PicOrderCntMsb is set equal to 0, each CLVSS picture that is the first picture in the access unit in decoding order may have a PicOrderCntVal set equal to slice_pic_order_cnt_lsb. The value of PicOrderCntVal may be in the range of -2 31 ~2 31 -1, inclusive. Within one CVS, the PicOrderCntVal values for any two coded pictures need not be the same. The function PicOrderCnt( picX ) may be specified as follows. PicOrderCnt( picX ) = PicOrderCntVal of picture picX The function DiffPicOrderCnt( picA, picB ) may be specified as follows. DiffPicOrderCnt( picA, picB ) = PicOrderCnt( picA ) - PicOrderCnt( picB ) The bitstream is in the range of -2 15 ~2 15It may not include data that results in a value of DiffPicOrderCnt( picA, picB ) used in the decoding process that is not within the range of -1. Let the current picture be X, and let two other pictures in the same CVS be Y and Z. When both DiffPicOrderCnt( X, Y ) and DiffPicOrderCnt( X, Z ) are positive or both are negative, Y and Z may be regarded as being in the same output order direction from X.

[0157] An exemplary decoding process for constructing a reference picture list is as follows. This process may be called at the start of the decoding process for each slice of a non - IDR picture. The reference pictures are addressed through reference indices. The reference index is an index into the reference picture list. When decoding an I - slice, the reference picture list is not used in decoding the slice data. When decoding a P - slice, only the reference picture list 0 (RefPicList

[0000] ) is used in decoding the slice data. When decoding a B - slice, both the reference picture list 0 and the reference picture list 1 (RefPicList

[0001] ) are used in decoding the slice data. At the start of the decoding process for each slice of a non - IDR picture, the reference picture lists RefPicList

[0000] and RefPicList

[0001] are derived. The reference picture lists are used in marking the reference pictures or in decoding the slice data. For an I - slice of a non - IDR picture that is not the first slice of the picture, RefPicList

[0000] and RefPicList

[0001] may be derived for bit - stream conformity checking purposes. Such derivation is not essential for decoding the current picture or a picture following the current picture in decoding order. For a P - slice that is not the first slice of the picture, RefPicList

[0001] may be derived for bit - stream conformity checking purposes. However, such derivation is not essential for decoding the current picture or a picture following the current picture in decoding order. The reference picture lists RefPicList

[0000] and RefPicList

[0001] may be constructed as follows. for( i = 0; i < 2; i++ ) { for( j = 0, k = 0, pocBase = PicOrderCntVal; j < num_ref_entries[ i ][ RplsIdx[ i ] ]; j++) { if( st_ref_pic_flag[ i ][ RplsIdx[ i ] ][ j ] ) { RefPicPocList[ i ][ j ] = pocBase - DeltaPocSt[ i ][ RplsIdx[ i ] ][ j if (there is a reference picture picA with PicOrderCntVal equal to RefPicPocList[ i ][ j ] in the DPB) RefPicList[ i ][ j ] = picA else RefPicList[ i ][ j ] = "There is no reference picture" (8-5) pocBase = RefPicPocList[ i ][ j ] } else { if (!delta_poc_msb_cycle_lt[ i ][ k ] ) { if (there is a reference picA with PicOrderCntVal & (MaxPicOrderCntLsb - 1) equal to PocLsbLt[ i ][ k ] in the DPB) RefPicList[ i ][ j ] = picA else RefPicList[ i ][ j ] = "There is no reference picture" } else { if (there is a reference picA with PicOrderCntVal equal to FullPocLt[ i ][ k ] in the DPB) RefPicList[ i ][ j ] = picA else RefPicList[ i ][ j ] = "There is no reference picture" } k++ } } }

[0158] For each i equal to 0 or 1, the first NumRefIdxActive[ i ] entries in RefPicList[ i ] are referred to as active entries in RefPicList[ i ], and the other entries in RefPicList[ i ] are referred to as non - active entries in RefPicList[ i ]. When the picture referred to by an entry has a nuh_layer_id with a value different from that of the current picture, the entry in the reference picture list may be referred to as an ILRP entry. It is possible for a particular picture to be referred to by both an entry in RefPicList

[0000] and an entry in RefPicList

[0001] . It is also possible for a particular picture to be referred to by more than two entries in RefPicList

[0000] or by more than two entries in RefPicList

[0001] . The active entries in RefPicList

[0000] and the active entries in RefPicList

[0001] may collectively refer to all the reference pictures that can be used for the inter - prediction of the current picture and one or more pictures following the current picture in decoding order. The non - active entries in RefPicList

[0000] and the non - active entries in RefPicList

[0001] may collectively refer to all the reference pictures that are not used for the inter - prediction of the current picture but can be used in the inter - prediction for one or more pictures following the current picture in decoding order. Since the corresponding picture does not exist in the DPB, there may be one or more entries in RefPicList

[0000] or RefPicList

[0001] that are not equal to any reference picture. Each non - active entry in RefPicList

[0000] or RefPicList

[0000] that is not equal to any reference picture may be ignored. For each active entry in RefPicList

[0000] or RefPicList

[0001] that is not equal to any reference picture, an accidental picture loss may be inferred.

[0159] The following constraints may apply to bitstream compliance. For each i equal to 0 or 1, num_ref_entries[ i ][ RplsIdx[ i ] ] need not be less than NumRefIdxActive[ i ]. The pictures referenced by each active entry in RefPicList

[0000] or RefPicList

[0001] may be present in the DPB and may have a TemporalId less than or equal to the TemporalId of the current picture. The picture referenced by each entry in RefPicList

[0000] or RefPicList

[0001] need not be the current picture. The STRP entries in RefPicList

[0000] or RefPicList

[0001] of a slice of a picture, and the LTRP entries in RefPicList

[0000] or RefPicList

[0001] of the same slice or a different slice of the same picture need not reference the same picture. There need not be any LTRP entries in RefPicList

[0000] or RefPicList

[0001] for which the difference between the PicOrderCntVal of the current picture and the PicOrderCntVal of the picture referenced by the entry is 224 or more. Let setOfRefPics be the set of unique pictures referenced by all entries in RefPicList

[0000] that have the same nuh_layer_id as the current picture, and all entries in RefPicList

[0001] that have the same nuh_layer_id as the current picture. The number of pictures in setOfRefPics may be less than or equal to sps_max_dec_pic_buffering_minus1, and setOfRefPics may be the same for all slices of a picture. The picture referenced by each ILRP entry in RefPicList

[0000] or RefPicList

[0001] of a slice of the current picture may be in the same access unit as the current picture.The pictures referred to by each ILRP entry in RefPicList

[0000] or RefPicList

[0001] of the current picture slice may be present in the DPB and may have a nuh_layer_id smaller than the nuh_layer_id of the current picture. Each ILRP entry in RefPicList

[0000] or RefPicList

[0001] of the slice may be an active entry.

[0160] An exemplary decoding process for reference picture marking is as follows. This process may be called once per picture after the decoding of the slice header and the decoding process for the reference picture list construction for the slice. The process may also be called before the decoding of the slice data. This process can result in one or more reference pictures in the DPB being marked as not used for reference or marked as being used for long-term reference. The decoded pictures in the DPB can be marked as not used for reference, used for short-term reference, or marked as being used for long-term reference. The decoded pictures in the DPB can only be marked as one of these three at any given moment during the operation of the decoding process. Assigning one of these markings to a picture implicitly removes the other of these markings when applicable. When a picture is referred to as being marked as being used for reference, this collectively refers to the picture being marked as being used for short-term reference or long-term reference (but not both). STRP and ILRP may be identified by their nuh_layer_id and PicOrderCntVal values. LTRP may be identified by their nuh_layer_id values and the Log2(MaxLtPicOrderCntLsb) least significant bits of their PicOrderCntVal values. If the current picture is a CLVSS picture, all reference pictures in the current DPB (if any) having the same nuh_layer_id as the current picture may be marked as not used for reference. Otherwise, the following applies. For each LTRP entry in RefPicList

[0000] or RefPicList

[0001] , when the picture being referred to is an STRP having the same nuh_layer_id as the current picture, the picture is marked as being used for long-term reference.Each reference picture in the DPB that has the same nuh_layer_id as the current picture and is not referenced by any entry in RefPicList

[0000] or RefPicList

[0001] may be marked as not used for reference. For each ILRP entry in RefPicList

[0000] or RefPicList

[0001] , the referenced picture may be marked as used for long-term reference.

[0161] A fourth exemplary implementation of the method described above is described below. An exemplary decoding process for picture order count is as follows. The output of this process is PicOrderCntVal, which is the picture order count of the current picture. Each coded picture is associated with a picture order count variable shown as PicOrderCntVal. When the current picture is not a CLVSS picture, the variables prevPicOrderCntLsb and prevPicOrderCntMsb may be derived as follows. Let prevTid0Pic be the previous picture in decoding order that has a nuh_layer_id less than or equal to the nuh_layer_id of the current picture and a TemporalId equal to 0 and is not a RASL picture or a RADL picture. The variable prevPicOrderCntLsb may be set equal to slice_pic_order_cnt_lsb of prevTid0Pic. The variable prevPicOrderCntMsb may be set equal to PicOrderCntMsb of prevTid0Pic. The variable PicOrderCntMsb of the current picture may be derived as follows. When the current picture is a CLVSS picture, PicOrderCntMsb may be set equal to 0. Otherwise, PicOrderCntMsb may be derived as follows. if( ( slice_pic_order_cnt_lsb < prevPicOrderCntLsb ) && ( ( prevPicOrderCntLsb - slice_pic_order_cnt_lsb ) >= ( MaxPicOrderCntLsb / 2 ) ) ) PicOrderCntMsb = prevPicOrderCntMsb + MaxPicOrderCntLsb else if( (slice_pic_order_cnt_lsb > prevPicOrderCntLsb ) && ( ( slice_pic_order_cnt_lsb - prevPicOrderCntLsb ) > ( MaxPicOrderCntLsb / 2 ) ) ) PicOrderCntMsb = prevPicOrderCntMsb - MaxPicOrderCntLsb else PicOrderCntMsb = prevPicOrderCntMsb

[0162] PicOrderCntVal may be derived as follows. PicOrderCntVal = PicOrderCntMsb + slice_pic_order_cnt_lsb For CLVSS pictures, after PicOrderCntMsb is set equal to 0, all CLVSS pictures may have a PicOrderCntVal equal to slice_pic_order_cnt_lsb. The value of PicOrderCntVal may be in the range of -2 31 ~2 31 -1, inclusive. Within one CVS, the PicOrderCntVal values for any two coded pictures need not be the same. The function PicOrderCnt( picX ) may be specified as follows. PicOrderCnt( picX ) = PicOrderCntVal of picture picX The function DiffPicOrderCnt( picA, picB ) may be specified as follows. DiffPicOrderCnt( picA, picB ) = PicOrderCnt( picA ) - PicOrderCnt( picB ) The bitstream is -2, inclusive15 ~2 15 Data that does not result in a value of DiffPicOrderCnt( picA, picB ) used in the decoding process that is not within the range of -1 may not be included. Let the current picture be X, and two other pictures in the same CVS be Y and Z. When both DiffPicOrderCnt(X, Y) and DiffPicOrderCnt( X, Z ) are positive or both are negative, Y and Z are considered to be in the same output order direction from X.

[0163] An exemplary decoding process for constructing a reference picture list is as follows. This process may be called at the start of the decoding process for each slice of a non-IDR picture. A reference picture is addressed through a reference index. The reference index is an index into the reference picture list. When decoding an I slice, the reference picture list is not used in decoding the slice data. When decoding a P slice, only the reference picture list 0 (RefPicList

[0000] ) is used in decoding the slice data. When decoding a B slice, both the reference picture list 0 and the reference picture list 1 (RefPicList

[0001] ) are used in decoding the slice data. At the start of the decoding process for each slice of a non-IDR picture, the reference picture lists RefPicList

[0000] and RefPicList

[0001] may be derived. The reference picture lists may be used in marking reference pictures or in decoding slice data. For an I slice of a non-IDR picture that is not the first slice of the picture, RefPicList

[0000] and RefPicList

[0001] may be derived for bitstream conformance checking purposes. Such derivation may not be essential for decoding the current picture or a picture following the current picture in decoding order. For a P slice that is not the first slice of the picture, RefPicList

[0001] may be derived for bitstream conformance checking purposes. However, such derivation may not be essential for decoding the current picture or a picture following the current picture in decoding order.

[0164] The reference picture lists RefPicList

[0000] and RefPicList

[0001] may be constructed as follows. for( i = 0; i < 2; i++ ) { for( j = 0, k = 0, pocBase = PicOrderCntVal; j < num_ref_entries[ i ][ RplsIdx[ i ] ]; j++) { if( st_ref_pic_flag[ i ][ RplsIdx[ i ] ][ j ] ) { RefPicPocList[ i ][ j ] = pocBase - DeltaPocSt[ i ][ RplsIdx[ i ] ][ j if( there is a reference picture picA with PicOrderCntVal equal to RefPicPocList[ i ][ j ] in the DPB ) RefPicList[ i ][ j ] = picA else RefPicList[ i ][ j ] = "There is no reference picture" (8-5) pocBase = RefPicPocList[ i ][ j ] } else { if( !delta_poc_msb_cycle_lt[ i ][ k ] ) { if( there is a reference picA with PicOrderCntVal & ( MaxPicOrderCntLsb - 1 ) equal to PocLsbLt[ i ][ k ] in the DPB ) RefPicList[ i ][ j ] = picA else RefPicList[ i ][ j ] = "There is no reference picture" } else { if( there is a reference picA with PicOrderCntVal equal to FullPocLt[ i ][ k ] in the DPB ) RefPicList[ i ][ j ] = picA else RefPicList[ i ][ j ] = "There is no reference picture" } k++ } } }

[0165] For each i equal to 0 or 1, the first NumRefIdxActive[ i ] entries in RefPicList[ i ] are referred to as active entries in RefPicList[ i ], and the other entries in RefPicList[ i ] are referred to as non - active entries in RefPicList[ i ]. It is possible for a particular picture to be referred to by entries in both RefPicList

[0000] and RefPicList

[0001] . It is also possible for a particular picture to be referred to by more than two entries in RefPicList

[0000] or by more than two entries in RefPicList

[0001] . The active entries in RefPicList

[0000] and the active entries in RefPicList

[0001] may collectively refer to all the reference pictures that can be used for inter - prediction of the current picture and one or more pictures following the current picture in decoding order. The non - active entries in RefPicList

[0000] and the non - active entries in RefPicList

[0001] may collectively refer to all the reference pictures that are not used for inter - prediction of the current picture but may be used in inter - prediction for one or more pictures following the current picture in decoding order. Since the corresponding pictures do not exist in the DPB, there may be one or more entries in RefPicList

[0000] or RefPicList

[0001] that are not equal to any reference picture. Each non - active entry in RefPicList

[0000] or RefPicList

[0001] that is not equal to any reference picture may be ignored. For each active entry in RefPicList

[0000] or RefPicList

[0001] that is not equal to any reference picture, an accidental picture loss may be inferred.

[0166] Bitstream conformance may require that the following constraints apply. For each i equal to 0 or 1, num_ref_entries[ i ][ RplsIdx[ i ] ] need not be less than NumRefIdxActive[ i ]. The pictures referred to by each active entry in RefPicList

[0000] or RefPicList

[0001] may be present in the DPB and may have a TemporalId less than or equal to the TemporalId of the current picture. The picture referred to by each entry in RefPicList

[0000] or RefPicList

[0001] need not be the current picture. The STRP entries in RefPicList

[0000] or RefPicList

[0001] of a slice of a picture, and the LTRP entries in RefPicList

[0000] or RefPicList

[0001] of the same slice or a different slice of the same picture, need not refer to the same picture. There need not be an LTRP entry in RefPicList

[0000] or RefPicList

[0001] for which the difference between the PicOrderCntVal of the current picture and the PicOrderCntVal of the picture referred to by the entry is 224 or more. Let the set of unique pictures referred to by all entries in RefPicList

[0000] and all entries in RefPicList

[0001] be setOfRefPics. The number of pictures in setOfRefPics may be less than or equal to sps_max_dec_pic_buffering_minus1, and setOfRefPics may be the same for all slices of a picture. The pictures referred to by each active entry in RefPicList

[0000] or RefPicList

[0001] may be present in the DPB and may have a nuh_layer_id less than or equal to the nuh_layer_id of the current picture.

[0167] An exemplary decoding process for reference picture marking is as follows. This process may be called once per picture after the decoding of the slice header and the decoding process for reference picture list construction for the slice. The process may be called before the decoding of the slice data. This process may result in one or more reference pictures in the DPB being marked as not used for reference or being marked as used for long-term reference. The decoded pictures in the DPB may be marked as not used for reference, used for short-term reference, or used for long-term reference. The decoded pictures in the DPB may be marked as only one of these three at any given instant during the operation of the decoding process. Assigning one of these markings to a picture implicitly removes any other of these markings when applicable. When a picture is referred to as being marked as used for reference, this collectively refers to the picture being marked as used for short-term reference or used for long-term reference (but not both). STRPs may be identified by their PicOrderCntVal values. LTRPs may be identified by the Log2(MaxLtPicOrderCntLsb) least significant bits of their PicOrderCntVal values. If the current picture is a CLVSS picture, all reference pictures in the current DPB (if any) may be marked as not used for reference. Otherwise, the following applies. For each LTRP entry in RefPicList

[0000] or RefPicList

[0001] , when the referenced picture is an STRP, the referenced picture is marked as used for long-term reference. Each reference picture in the DPB that is not referenced by any entry in RefPicList

[0000] or RefPicList

[0001] is marked as not used for reference.

[0168] FIG. 10 is a schematic diagram of an exemplary video coding device 1000. The video coding device 1000 is suitable for implementing the disclosed examples / embodiments as described herein. The video coding device 1000 includes a transceiver unit (Tx / Rx) 1010 that includes a downstream port 1020, an upstream port 1050, and / or a transmitter and / or receiver for communicating data upstream and / or downstream via a network. The video coding device 1000 also includes a processor 1030 that includes a logic unit and / or a central processing unit (CPU) for processing data, and a memory 1032 for storing data. The video coding device 1000 may also include electrical components, optoelectronic (OE) components, electro-optical (EO) components, and / or wireless communication components coupled to the upstream port 1050 and / or the downstream port 1020 for communication of data via an electrical communication network, an optical communication network, or a wireless communication network. The video coding device 1000 may also include an input and / or output (I / O) device 1060 for communicating data with a user. The I / O device 1060 may include output devices such as a display for displaying video data, a speaker for outputting audio data, etc. The I / O device 1060 may also include input devices such as a keyboard, a mouse, a trackball, etc., and / or corresponding interfaces for interacting with such output devices.

[0169] Processor 1030 is implemented by hardware and software. Processor 1030 can be implemented as one or more CPU chips, cores (e.g., multi-core processor), field programmable gate array (FPGA), application specific integrated circuit (ASIC), and digital signal processor (DSP). Processor 1030 communicates with downstream port 1020, Tx / Rx 1010, upstream port 1050, and memory 1032. Processor 1030 includes a coding module 1014. Coding module 1014 implements the disclosed embodiments described herein, such as methods 100, 1100, and / or 1200, which may employ a bitstream 900 including pictures that may be coded according to an RPL structure 800, as well as unidirectional inter-prediction 500, bidirectional inter-prediction 600, and / or layer-based prediction 700. Coding module 1014 may also implement any other method / mechanism described herein. Further, coding module 1014 may implement a codec system 200, an encoder 300, and / or a decoder 400. For example, coding module 1014 may be employed to code an ILRP flag and / or an ILRP layer indicator in a reference picture structure to manage reference pictures to support inter-layer prediction as described above. Thus, coding module 1014 provides additional functionality and / or coding efficiency to video coding device 1000 when coding video data. Thus, coding module 1014 improves the functionality of video coding device 1000 and addresses problems specific to video coding techniques. Further, coding module 1014 affects the conversion of video coding device 1000 to different states. Alternatively, coding module 1014 may be implemented as instructions stored in memory 1032 and executed by processor 1030 (e.g., as a computer program product stored on a non-transitory medium).

[0170] The memory 1032 includes one or more memory types such as a disk, a tape drive, a solid state drive, a read-only memory (ROM), a random access memory (RAM), a flash memory, a ternary content-addressable memory (TCAM), a static random access memory (SRAM). The memory 1032 can be used as an overflow data storage device for storing such a program when the program is selected for execution and for storing instructions and data read during program execution.

[0171] FIG. 11 is a flowchart of an exemplary method 1100 for encoding a video sequence into a bitstream such as bitstream 900 according to inter-layer prediction. The method 1100 can be employed by an encoder such as codec system 200, encoder 300, and / or video coding device 1000 when executing method 100 for encoding pictures according to unidirectional inter prediction 500, bidirectional inter prediction 600, and / or layer-based prediction 700 by adopting the RPL structure 800.

[0172] When an encoder receives a video sequence including a plurality of pictures and determines, for example, based on user input, that the video sequence should be encoded into a bitstream, method 1100 may start. In step 1101, the encoder encodes the current picture into the bitstream. For example, the current picture may be encoded according to inter-layer prediction based on an inter-layer reference picture. For example, such inter-layer prediction may be performed according to inter-layer prediction 721. As described above, the inter-layer reference picture may be in the same AU as the current picture. Thus, the inter-layer reference picture may include the same POC as the current picture. Further, the inter-layer reference picture is arranged in a different layer from the current picture. For example, the inter-layer reference picture may be associated with a layer lower than the layer of the current picture. Thus, the inter-layer reference picture may be associated with a layer ID lower than the layer ID of the current picture.

[0173] In step 1103, the encoder can encode a reference picture list structure such as RPL structure 800 into the bitstream. The reference picture list structure includes a plurality of entries for a plurality of reference pictures. Such entries include entries related to the current picture. The entries indicate inter-layer reference pictures. For example, the reference picture list structure may be denoted as ref_pic_list_struct( listIdx, rplsIdx ), where listIdx identifies the reference picture list, rplsIdx identifies the entry in the reference picture list, and ref_pic_list_struct is a syntax structure that returns an entry based on listIdx and rplsIdx.

[0174] In step 1105, the encoder can encode the inter-layer reference picture flag into the bitstream. The inter-layer reference picture flag indicates that the entry related to the current picture is an ILRP entry. For example, the inter-layer reference picture flag may be indicated as inter_layer_ref_pic_flag[listIdx][rplsIdx][i]. Specifically, when the i-th entry in ref_pic_list_struct(listIdx, rplsIdx) is an ILRP entry, inter_layer_ref_pic_flag[listIdx][rplsIdx][i] may be set equal to 1. Further, when the i-th entry in ref_pic_list_struct(listIdx, rplsIdx) is not an ILRP entry, inter_layer_ref_pic_flag[listIdx][rplsIdx][i] may be set equal to 0. For example, ref_pic_list_struct(listIdx, rplsIdx) and inter_layer_ref_pic_flag[listIdx][rplsIdx][i] may be encoded into the bitstream in the SPS. In another example, ref_pic_list_struct(listIdx, rplsIdx) and inter_layer_ref_pic_flag[listIdx][rplsIdx][i] may be encoded into the bitstream in the header related to the current picture, such as the slice header and / or the picture header.

[0175] In step 1107, the encoder can encode the ILRP layer indicator into the bitstream. The ILRP layer indicator indicates the layer of the inter-layer reference picture. For example, the ILRP layer indicator can indicate one or more layers for each entry of the reference picture list structure related to the inter-layer reference picture as indicated by the inter-layer reference picture flag.

[0176] In step 1109, the encoder can store a bitstream for communication to the decoder, for example, for on-demand communication.

[0177] FIG. 12 is a flowchart of an exemplary method for decoding a video sequence from a bitstream such as bitstream 900 when adopting inter-layer prediction. Method 1200 may be employed by a decoder such as codec system 200, decoder 400, and / or video coding device 1000 when executing method 100 for decoding pictures according to unidirectional inter-prediction 500, bidirectional inter-prediction 600, and / or layer-based prediction 700 by adopting RPL structure 800.

[0178] For example, as a result of method 1100, when the decoder begins to receive a bitstream of coded data representing a video sequence, method 1200 may start. In step 1201, the decoder receives the bitstream. The bitstream may comprise a current picture, a reference picture list structure, an inter-layer reference picture flag, and / or an ILRP layer indicator. For example, the reference picture list structure may comprise an inter-layer reference picture flag and / or an ILRP layer indicator.

[0179] In step 1203, the decoder can determine, based on the inter-layer reference picture flag, that an entry in the reference picture list structure related to the current picture is an ILRP entry. For example, the reference picture list structure includes a plurality of entries for a plurality of reference pictures. Such an entry includes an entry related to the current picture. The entry indicates a reference picture. For example, the reference picture list structure may be denoted as ref_pic_list_struct( listIdx, rplsIdx ), where listIdx identifies the reference picture list, rplsIdx identifies an entry in the reference picture list, and ref_pic_list_struct is a syntax structure that returns an entry based on listIdx and rplsIdx. Further, the inter-layer reference picture flag indicates whether an entry related to the current picture is an ILRP entry. For example, the inter-layer reference picture flag may be denoted as inter_layer_ref_pic_flag[ listIdx ][ rplsIdx ][ i ]. Specifically, when the i-th entry in ref_pic_list_struct( listIdx, rplsIdx ) is an ILRP entry, inter_layer_ref_pic_flag[ listIdx ][ rplsIdx ][ i ] may be set equal to 1. Further, when the i-th entry in ref_pic_list_struct( listIdx, rplsIdx ) is not an ILRP entry, inter_layer_ref_pic_flag[ listIdx ][ rplsIdx ][ i ] may be set equal to 0. For example, ref_pic_list_struct( listIdx, rplsIdx ) and inter_layer_ref_pic_flag[ listIdx ][ rplsIdx ][ i ] may be encoded into the bitstream in the SPS.In another example, ref_pic_list_struct( listIdx, rplsIdx ) and inter_layer_ref_pic_flag[ listIdx ][ rplsIdx ][ i ] may be encoded in the bitstream in headers related to the current picture, such as the slice header and / or the picture header.

[0180] In step 1205, when the entry is an ILRP entry, the decoder may determine the layer of the inter-layer reference picture based on the ILRP layer indicator. For example, the ILRP layer indicator may indicate one or more layers for each entry of the reference picture list structure related to the inter-layer reference picture as indicated by the inter-layer reference picture flag. For example, the inter-layer reference picture may be included in the same AU as the current picture. Thus, the inter-layer reference picture may include the same POC as the current picture. Further, the inter-layer reference picture is arranged in a layer different from the current picture. For example, the inter-layer reference picture may be related to a layer lower than the layer of the current picture. Thus, the inter-layer reference picture may be related to a layer ID lower than the layer ID of the current picture.

[0181] In step 1207, when the entry is an ILRP entry, the decoder can decode the current picture according to the inter-layer prediction based on the inter-layer reference picture indicated by the entry in the reference picture list structure at the level indicated by the ILRP layer indicator.

[0182] In step 1209, when the entry is not an ILRP entry, the decoder can decode the current picture according to the intra-layer prediction based on the intra-layer reference picture indicated by the entry in the reference picture list structure. The intra-layer reference picture is simply referred to as a reference picture in a single-layer context. Further, the intra-layer prediction may include intra prediction and / or single-layer inter prediction according to the context. When a reference picture is adopted, the intra-layer prediction indicates a single-layer inter prediction.

[0183] In step 1211, the decoder transfers the decoded / reconstructed current picture for display as part of the decoded video sequence.

[0184] FIG. 13 is a schematic diagram of an exemplary system 1300 for coding a video sequence of an image in a bitstream such as bitstream 900 when adopting inter-layer prediction. System 1300 can be implemented by an encoder and a decoder, such as codec system 200, encoder 300, decoder 400, and / or video coding device 1000. Further, system 1300 can be adopted when implementing methods 100, 1100, and / or 1200 for decoding a picture according to unidirectional inter prediction 500, bidirectional inter prediction 600, and / or layer-based prediction 700 by adopting the RPL structure 800.

[0185] System 1300 includes a video encoder 1302. The video encoder 1302 includes an encoding module 1301 for encoding a current picture into a bitstream, where the current picture is encoded according to inter-layer prediction based on inter-layer reference pictures. The encoding module 1301 is further for encoding a reference picture list structure into the bitstream, where the reference picture list structure includes a plurality of entries for a plurality of reference pictures related to the current picture and indicates an inter-layer reference picture. The encoding module 1301 is further for encoding an inter-layer reference picture flag into the bitstream, where the inter-layer reference picture flag indicates that an entry related to the current picture is an ILRP entry. The video encoder 1302 further includes a storage module 1303 for storing the bitstream for communication to a decoder. The video encoder 1302 further includes a transmission module 1305 for transmitting the bitstream towards a video decoder 1310. The video encoder 1302 may be further configured to execute any of the steps of method 1100.

[0186] System 1300 also includes a video decoder 1310. The video decoder 1310 includes a receiving module 1311 for receiving a bitstream having a current picture, a reference picture list structure, and an inter-layer reference picture flag. The video decoder 1310 further includes a determining module 1313 for determining, based on the inter-layer reference picture flag, whether an entry in the reference picture list structure related to the current picture is an ILRP entry. The video decoder 1310 further includes a decoding module 1315 for decoding the current picture according to inter-layer prediction based on an inter-layer reference picture indicated by an entry in the reference picture list structure when the entry is an ILRP entry. The video decoder 1310 further includes a transfer module 1317 for transferring the current picture for display as part of a decoded video sequence. The video decoder 1310 may be further configured to perform any of the steps of method 1200.

[0187] Except for a line, trace, or another medium, when there is no intervening component between a first component and a second component, the first component is directly coupled to the second component. When there is an intervening component other than a line, trace, or another medium between the first component and the second component, the first component is indirectly coupled to the second component. The term "coupled" and its variations include both being directly coupled and being indirectly coupled. The use of the term "about" means a range including ± 10% of the number that follows, unless otherwise specified.

[0188] It should also be understood that the steps of the exemplary methods described herein are not necessarily required to be executed in the order described, and that such order of steps is for example only. Similarly, additional steps may be included in such methods, and some steps may be omitted or combined in ways consistent with various embodiments of the present disclosure.

[0189] While several embodiments are provided in the present disclosure, it can be understood that the disclosed systems and methods may be embodied in many other specific forms without departing from the spirit or scope of the present disclosure. This example is to be considered illustrative and not restrictive, and its intent is not limited to the details given herein. For example, various elements or components may be combined or integrated in another system, or some features may be omitted or not implemented.

[0190] In addition, the techniques, systems, subsystems, and methods described and illustrated individually or separately in various embodiments may be combined with or integrated into other systems, components, techniques, or methods without departing from the scope of the present disclosure. Other examples of changes, substitutions, and modifications can be identified by those skilled in the art and may be made without departing from the spirit and scope disclosed herein.

Description of Reference Numerals

[0191] 200 codec system 201 segmented video signal 211 general coder control component 213 conversion scaling and quantization component 215 intra-picture estimation component 217 intra-picture prediction component 219 motion compensation component 221 motion estimation component 223 decoded picture buffer component 225 in-loop filter component 227 filter control analysis component 229 scaling and inverse conversion component 231 header formatting and context-adaptive binary arithmetic coding component 300 video encoder 301 segmented video signal 313 conversion and quantization component 317 Intra-prediction component in picture 321 Motion compensation component 323 Decoded picture buffer component 325 In-loop filter component 329 Inverse transform and quantization component 331 Entropy coding component 400 Video decoder 417 Intra-prediction component in picture 421 Motion compensation component 423 Decoded picture buffer component 425 In-loop filter component 429 Inverse transform and quantization component 433 Entropy decoding component 500 Unidirectional inter-prediction 510 Current picture 511 Current block 513 Motion trajectory 530 Reference picture 531 Reference block 533 Temporal distance 535 Motion vector 600 Bidirectional inter-prediction 610 Current picture 611 Current block 613 Motion trajectory 620 Previous reference picture 621 Previous reference block 623 Previous temporal distance 625 Previous motion vector 630 Subsequent reference picture 631 Subsequent reference block 633 Subsequent temporal distance 635 Subsequent motion vector 700 Layer-based prediction 711 - 718 Pictures 721 Inter-layer prediction 723 Inter-prediction 731 Layer N 732 Layer N + 1 800 Reference Picture List Structure (RPL Structure) 811 RPL0 812 RPL1 815 Reference Picture List Structure Entry 821 listIdx 825 rplsIdx 833 ILRP Flag 835 ILRP Layer Indicator 900 Bitstream 910 Sequence Parameter Set (SPS) 911 Picture Parameter Set (PPS) 915 Slice Header 920 Image Data 921 Access Unit 923 Picture 925 Slice 931 Reference Picture List Structure 933 ILRP Flag 935 ILRP Layer Indicator 1000 Video Coding Device 1010 Transceiver Unit (Tx / Rx) 1014 Coding Module 1020 Downstream Port 1030 Processor 1032 Memory 1050 Upstream Port 1060 Input and / or Output (I / O) Device 1300 System 1301 Encoding Module 1302 Video Encoder 1303 Memory Module 1305 Transmission Module 1310 Video Decoder 1311 Reception Module 1313 Decision Module 1315 Decoding Module 1317 Transfer Module

Claims

1. A method implemented in a decoder, comprising the steps of: receiving, by a decoder receiver, a bitstream comprising a current picture and a reference picture list structure comprising an inter-layer reference picture flag; determining, by a processor of the decoder, based on the inter-layer reference picture flag, that an entry in the reference picture list structure associated with the current picture is an inter-layer reference picture (ILRP) entry; when the entry is the ILRP entry, decoding, by the processor, the current picture based on an inter-layer reference picture indicated by the entry in the reference picture list structure; A method for providing the above.

2. 2. The method of claim 1, wherein the reference picture list structure further comprises an ILRP layer indicator, and the method further comprises, when the entry is the ILRP entry, determining, by the processor, a layer of the inter-layer reference picture based on the ILRP layer indicator.

3. The method of claim 1 or 2, wherein the reference picture list structure is denoted as ref_pic_list_struct(listIdx, rplsIdx), where listIdx identifies a reference picture list, rplsIdx identifies an entry in the reference picture list, and ref_pic_list_struct is a syntax structure.

4. The inter-layer reference picture flag is indicated as inter_layer_ref_pic_flag[listIdx][rplsIdx][i], and when the i-th entry in the ref_pic_list_struct(listIdx,rplsIdx) is the ILRP entry, the inter_layer_ref_pic_flag[listIdx][rplsIdx 4. The method of claim 1, wherein inter_layer_ref_pic_flag[listIdx][rplsIdx][i] is equal to 1 and the inter_layer_ref_pic_flag[listIdx][rplsIdx][i] is equal to 0 when the i-th entry in the ref_pic_list_struct(listIdx,rplsIdx) is not the ILRP entry.

5. 5. The method of claim 1, further comprising: when the entry is not the ILRP entry, decoding the current picture by the processor according to intra-layer prediction based on an intra-layer reference picture indicated by the entry in the reference picture list structure.

6. 6. The method of claim 1, wherein the ref_pic_list_struct(listIdx, rplsIdx) and the inter_layer_ref_pic_flag[listIdx][rplsIdx][i] are included in the bitstream in a sequence parameter set (SPS).

7. 7. The method of claim 1, wherein the inter-layer reference picture is in the same access unit (AU) as the current picture, and the inter-layer reference picture is associated with a lower layer identifier than the current picture.

8. 1. A method implemented in an encoder, comprising: encoding, by a processor of the encoder, a current picture into a bitstream, the current picture being encoded according to inter-layer prediction based on an inter-layer reference picture; encoding, by the processor, a reference picture list structure into the bitstream, the reference picture list structure comprising a plurality of entries for a plurality of reference pictures including an entry associated with the current picture and indicating the inter-layer reference picture; encoding, by the processor, an inter-layer reference picture flag into the bitstream, the inter-layer reference picture flag indicating that the entry associated with the current picture is an inter-layer reference picture (ILRP) entry; storing the bitstream by a memory coupled to the processor for communication to a decoder; A method for providing the above.

9. 10. The method of claim 8, further comprising: encoding, by the processor, an ILRP layer indicator into the bitstream, the ILRP layer indicator indicating a layer of the inter-layer reference picture.

10. The method of claim 8 or 9, wherein the reference picture list structure is denoted as ref_pic_list_struct( listIdx, rplsIdx ), where listIdx identifies a reference picture list, rplsIdx identifies an entry in the reference picture list, and ref_pic_list_struct is a syntax structure that returns the entry based on listIdx and rplsIdx.

11. The inter-layer reference picture flag is indicated as inter_layer_ref_pic_flag[listIdx][rplsIdx][i], and when the i-th entry in the ref_pic_list_struct(listIdx,rplsIdx) is the ILRP entry, the inter_layer_ref_pic_flag[listIdx][rplsIdx 11. The method of claim 8, wherein inter_layer_ref_pic_flag[listIdx][rplsIdx][i] is equal to 1 and the inter_layer_ref_pic_flag[listIdx][rplsIdx][i] is equal to 0 when the i-th entry in the ref_pic_list_struct(listIdx,rplsIdx) is not the ILRP entry.

12. 12. The method of claim 8, wherein the ref_pic_list_struct(listIdx, rplsIdx) and the inter_layer_ref_pic_flag[listIdx][rplsIdx][i] are encoded in the bitstream within a sequence parameter set (SPS).

13. The ref_pic_list_struct( listIdx, rplsIdx ) and the inter_layer_ref_pic_flag[ listIdx ][ rplsIdx 12. The method according to claim 8, wherein [ i ] is encoded in a header associated with the current picture.

14. 14. The method of claim 8, wherein the inter-layer reference picture is in the same access unit (AU) as the current picture, and the inter-layer reference picture is associated with a lower layer identifier than the current picture.

15. 1. A video coding device, comprising: A method for implementing a method according to any one of claims 1 to 14, comprising: a processor; a receiver coupled to the processor; a memory coupled to the processor; and a transmitter coupled to the processor, wherein the processor, the receiver, the memory, and the transmitter are configured to perform the method according to any one of claims 1 to 14. Video coding device.

16. A non-transitory computer-readable medium comprising a computer program product for use by a video coding device, the computer program product comprising computer-executable instructions stored on the non-transitory computer-readable medium that, when executed by a processor, causes the video coding device to perform a method according to any one of claims 1 to 14.

17. receiving means for receiving a bitstream comprising a current picture and a reference picture list structure comprising an inter-layer reference picture flag; determining means for determining, based on the inter-layer reference picture flag, that an entry in the reference picture list structure associated with the current picture is an inter-layer reference picture (ILRP) entry; a decoding means for decoding the current picture based on an inter-layer reference picture indicated by the entry in the reference picture list structure when the entry is the ILRP entry; transfer means for transferring the current picture for display as part of a decoded video sequence; A decoder comprising:

18. A decoder according to claim 17, further configured to perform the method according to any one of claims 1 to 7.

19. An encoding means, encoding a current picture into a bitstream, the current picture being encoded according to inter-layer prediction based on an inter-layer reference picture; encoding a reference picture list structure into the bitstream, the reference picture list structure comprising a plurality of entries for a plurality of reference pictures including an entry associated with the current picture and indicating the inter-layer reference picture; encoding an inter-layer reference picture flag into the bitstream, the inter-layer reference picture flag indicating that the entry associated with the current picture is an inter-layer reference picture (ILRP) entry; and encoding means for performing storage means for storing said bitstream for communication to a decoder; An encoder comprising:

20. 20. An encoder according to claim 19, further configured to perform a method according to any one of claims 8 to 14.

Citation Information

Patent Citations

  • Systems and methods for inter-layer RPS derivation based on sub-layer reference prediction dependency

    US20150103904A1

  • Apparatus, a method and a computer program for video coding and decoding

    US20150195573A1

  • Method for encoding video, method for decoding video, and apparatus using same

    US20150334399A1

  • Disabling inter-view prediction for reference picture list in video coding

    WO2014113669A1

  • Image decoding device and image encoding device

    WO2015005331A1