Time motion vector prediction, interlayer reference, video coding aspect of time sublayer indication

By aligning tree root block sizes and using compatible inter-layer references, the inefficiencies in TMVP and sub-block TMVP are addressed, improving coding efficiency and reducing memory bandwidth in multi-layer video coding.

JP2025108754AActive Publication Date: 2025-07-23FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2025073245
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2020-06-09
Filing Date
2025-04-25
Publication Date
2025-07-23
Estimated Expiration
2041-06-08

AI Technical Summary

Technical Problem

Existing video coding technologies face inefficiencies in temporal motion vector prediction (TMVP) due to unaligned CTU sizes between layers, leading to inefficient buffer usage and memory bandwidth issues, particularly in multi-layer video streams.

Method used

Implementing constraints on the division of images into tree root blocks to ensure they have equal or integral multiples of the reference image's block size, allowing for efficient TMVP and sub-block TMVP by aligning CTU boundaries, and using inter-layer reference images with compatible block sizes for prediction.

Benefits of technology

This approach enhances coding efficiency by reducing buffer management issues and memory bandwidth requirements, enabling effective TMVP and sub-block TMVP even with varying CTU sizes across layers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025108754000001_ABST
    Figure 2025108754000001_ABST
Patent Text Reader

Abstract

To provide a method of determining an interlayer reference image for inter-predicting a multilayer video data stream image, for video coding, a method of using an interlayer prediction tool for a multilayer video data stream, and a method of determining a maximum time sublayer of an output layer set (OLS) or a decoded maximum time sublayer.SOLUTION: An encoder 10 encodes a video sequence 12 into a video bitstream 14. The video sequence includes a sequence of an image 21. The image is arranged in a presentation order or an image order 17. Each access unit, which forms a coded video sequence 20 in a shape of an access unit 22, encodes video data, which belong to a common time point, thereinto. The encoder encodes the video sequence coded according to a coding order 19.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to a video encoder, a video decoder, a method of encoding a video sequence into a video bitstream, and a method of decoding a video sequence from a video bitstream. Further embodiments relate to a video bitstream.

Background Art

[0002] In the encoding or decoding of images of a video sequence, prediction is used to reduce the amount of information signaled in the video bitstream in which the images are encoded / decoded. Prediction may be used on the image data itself, such as sample values or coefficients for which the sample values of the image are coded. Alternatively or additionally, prediction may be used on syntax elements used for coding of the image, such as motion vectors. To predict the motion vector of the image to be coded, a reference image may be selected, and from the reference image, a predictor of the motion vector of the image to be coded is determined.

Summary of the Invention

Means for Solving the Problems

[0003] A first aspect of the present disclosure provides a concept for selecting a reference image to be used for temporal motion vector prediction. Two lists of reference images are populated for a given image, e.g., the image to be coded. Each list may be empty or non-empty. The TMVP reference image is determined by selecting one of the two lists of reference images as the TMVP image list and selecting the TMVP reference image from that TMVP image list. According to the first aspect, when one of the two lists is empty and the other is non-empty, the reference image from the non-empty list is used for temporal motion vector prediction (TMVP). Thus, TMVP can be used regardless of which of the two lists is empty, providing high coding efficiency in both cases where only the first list or only the second list is empty.

[0004] The second aspect of the present disclosure is based on the idea that the tree root block into which an image of a coded video sequence is divided is smaller than or equal to the size of the tree root block into which the reference image of that image is divided. Imposing such a constraint on the division of an image into tree root blocks can ensure that the dependency of the image on the reference image does not cross the boundaries of the tree root block, or at least does not cross the row boundaries of the rows of the tree root block. Thus, the constraint can limit the dependencies between different tree root blocks and can be beneficial for buffer management. In particular, the dependencies between tree root blocks of different rows of a tree root block can lead to inefficient buffer usage because adjacent tree root blocks belonging to different rows may be separated by further tree root blocks in the coding order. Thus, by avoiding such dependencies, it is possible to eliminate the need to hold the entire row of tree root blocks between the currently coded tree root block and the referenced tree root block.

[0005] The third aspect of the present disclosure provides a concept for determining the maximum temporal sublayer up to which layer of an output layer set represented in a multi-layer video bitstream should be decoded. Thus, this concept enables a decoder to determine which part of a video bitstream to decode even without an indication of the maximum temporal sublayer to be decoded. Further, according to this concept, when each indication is absent and the maximum temporal sublayer to be decoded corresponds to that estimated by the decoder, the encoder can omit signaling an indication of the maximum temporal sublayer to be decoded and avoid an unnecessarily high signaling overhead.

[0006] Embodiments and advantageous aspects of the present disclosure will be described in more detail below with reference to the drawings.

Brief Description of the Drawings

[0007]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

[0008] In the following, embodiments will be described in detail. However, of course, the embodiments provide many applicable concepts that can be embodied in various video coding concepts. The specific embodiments described are merely examples of specific ways to implement and use this concept, and do not limit the scope of the embodiments. In the following description, a plurality of details are described to provide a more complete explanation of the embodiments of the present disclosure. However, it will be apparent to those skilled in the art that other embodiments can be practiced without these specific details. In other examples, well-known structures and devices are shown in block diagram form rather than in detail to avoid obscuring the examples described herein. Furthermore, the features of different embodiments described herein may be combined with each other unless otherwise specifically stated.

[0009] In the following description of embodiments, elements that are the same or similar, or have the same function, are given the same reference numerals or are identified by the same names, and the repetition of the description of elements given the same reference numerals or identified by the same names is usually omitted. Accordingly, the descriptions provided for elements having the same reference numerals or identified by the same names are interchangeable with each other or may be applied to each other in different embodiments.

[0010] The detailed description of embodiments of the disclosed concepts begins with descriptions of examples of encoders, decoders, and video bitstreams, which provide a framework into which embodiments of the present invention may be incorporated. The following provides a description of embodiments of the concepts of the present invention, along with a description of how such concepts can be incorporated into the encoders and decoders of FIG. 1. However, encoders and decoders that do not operate according to the framework described with respect to FIG. 1 may be formed using the embodiments described with respect to FIG. 2 and subsequent figures. It should also be noted that the encoders and decoders, although described together in FIG. 1 for illustrative purposes, may be implemented separately from each other. It should also be noted that the encoders and decoders may be combined within one device or one of the two may be implemented as part of the other. Also, some of the embodiments of the present invention are described with reference to FIG. 1.

[0011] FIG. 1 shows an example of an encoder 10 and a decoder 50. The encoder 10 (which may also be referred to as an encoding device) encodes a video sequence 12 into a video bitstream 14 (which may be referred to as a bitstream, data stream, video data stream, or stream). The video sequence 12 includes a sequence of pictures 21, and the pictures 21 are arranged in a presentation order or picture order 17. In other words, each of the pictures 21 may represent a frame of the video sequence 12 and may be associated with a point in the presentation order of the video sequence 12. Based on the video sequence 12, the encoder 10 may encode a coded video sequence 20 into the video bitstream 14. The encoder 10 may form the coded video sequence 20 in the form of access units 22, and each access unit 22 encodes video data belonging to a common point in time therein. In other words, each access unit 22 may encode one of the pictures 21 of the video sequence 12, i.e., one of the frames. The encoder 10 encodes the coded video sequence 20 according to a coding order 19, and the coding order 19 may be different from the picture order 17 of the video sequence 12.

[0012] Encoder 10 may encode the coded video sequence 20 into one or more layers. That is, the video bitstream 14 may be a single-layer or multi-layer video bitstream including one or more layers. Each access unit 22 includes one or more coded pictures 26 (e.g., pictures 260, 261 in FIG. 1, where the apostrophe and asterisk are used to refer to specific ones, and the subscript below indicates the layer to which the picture belongs). Each picture 26 belongs to one of the layers 24 of the coded video sequence, e.g., layers 240, 241 in FIG. 1. In FIG. 1, an exemplary number of two layers, i.e., the first layer 241 and the second layer 240, are shown. In embodiments according to the disclosed concepts, the coded video sequence 20 and the video bitstream 14 do not necessarily include a plurality of layers, but may include one, two, or more layers. In the example of FIG. 1, each access unit 22 includes the coded picture 261 of the first layer 241 and the coded picture 260 of the second layer 240. However, it should be noted that each access unit 22 may, but does not necessarily, include coded pictures for each layer of the coded video sequence 20. For example, layers 240, 241 may have different frame rates (or picture rates) and / or may include pictures for complementary subsets of the access units of the access unit 22.

[0013] As described above, one image 260, 261 of an access unit represents image content at the same point in time. For example, the images 260, 261 of the same access unit 22 may represent the same image content with different qualities, such as resolution or fidelity. In other words, layer 240 may represent a first version of the coded video sequence 20, and layer 241 may represent a second version of the coded sequence 20. Thus, a decoder such as decoder 50, or an extractor, may select between different versions of the coded video sequence 20 decoded or extracted from the video bitstream 14. For example, layer 240 may be decoded independently of further layers of the coded video sequence to provide a decoded video sequence of a first quality, and joint decoding of the first layer 241 and the second layer 240 may provide a decoded video sequence of a second quality higher than the first quality. For example, the first layer 241 may be encoded independently of the second layer 240. In other words, the second layer 240 may be a reference layer for the first layer 241. For example, in this scenario, the first layer 241 may be referred to as an enhancement layer, and the second layer 240 may be referred to as a base layer. Image 260 may have a smaller image size, an equal image size, or a larger image size than image 261. For example, the image size may refer to the number of samples in the two-dimensional array of the image. Note that images 260, 261 do not necessarily have to represent equal image content. For example, image 261 may represent an excerpt of the image content of image 260. For example, in some scenarios, different layers of the video bitstream 14 may include different sub-images of the images coded in the video bitstream.

[0014] Encoder 10 encodes access unit 22 into bitstream portion 16 of video bitstream 14. For example, each access unit 22 may be encoded into one or more bitstream portions 16. For example, image 26 may be subdivided into slices of tiles, and each slice may be encoded into one bitstream portion 16. The bitstream portion 16 into which image 26 is encoded may be referred to as a video coding layer (VCL) NAL unit. Video bitstream 14 may further include non-VCL NAL units, for example, bitstream portions 23, 29 in which descriptive data is coded. The descriptive data may provide information for decoding or information regarding coded video sequence 20. The bitstream portions in which the descriptive data is encoded may be associated with individual bitstream portions. For example, they may refer to individual slices, or may be associated with one of images 26, or one of access units 22, or may be associated with a sequence of access units, i.e., may be related to coded video sequence 20. Note that video 12 may be coded into a sequence of coded video sequences 20.

[0015] The decoder 50 (which may also be referred to as a decoding device) decodes the video bitstream 14 so as to obtain the decoded video sequence 20'. Note that the video bitstream 14 provided to the decoder 50 does not necessarily correspond to the video bitstream 14 provided by the encoder, and may be extracted from the video bitstream provided by the encoder. Therefore, the video bitstream decoded by the decoder 50 may be a sub-bitstream of the video bitstream encoded by an encoder such as the encoder 10. As described above, the decoder 50 may decode the entire coding video sequence 20 coded in the video data stream 14, or a part thereof, for example, a subset of the layers of the coding video sequence 20 and / or a temporal subset of the coding video sequence 20 (i.e., a video sequence having a frame rate lower than the maximum frame rate provided by the video sequence 20). Therefore, the decoded video sequence 20' does not necessarily correspond to the video sequence 12 encoded by the encoder 10. It should also be noted that the decoded video sequence 20' may be further different from the video sequence 12 due to coding losses such as quantization losses.

[0016] The picture 26 may be encoded using a prediction tool for predicting a signal or coefficient representing the picture in the video bitstream 14 from a previously encoded picture. That is, the encoder 10 may use a prediction tool to encode the picture to be currently encoded using, for example, a previously encoded picture. Accordingly, the decoder 50 may use a prediction tool to predict the picture 26 to be currently decoded from a previously decoded picture. * In the following description, a given picture or block, for example, the picture or block currently being coded, is denoted by a reference sign ([[]] * * For example, the encoder 10 may use a prediction tool to encode the picture to be currently encoded using a previously encoded picture. Accordingly, the decoder 50 may use a prediction tool to predict the picture 26 to be currently decoded from a previously decoded picture. *) may be represented using. For example, the image 261 in FIG. 1 * is considered to be the image currently being coded, where the image 26 currently being coded * may equally refer to the image currently being encoded by the encoder 10 and the image currently being decoded in the decoding process executed by the decoder 50.

[0017] The prediction of an image from other images in the coding video sequence 20 may also be called inter prediction. For example, the image 261 * is the image 261 * may be encoded using temporal inter prediction from an image 261' belonging to one of the access units different from that of the image 261. Thus, the image 261 * is the image 261 * belongs to the same layer as the image 261 * but may include a reference 32 to an image 261' belonging to a different access unit from that of the image 261. Additionally or alternatively, the image 261 * may be predicted using inter-layer (inter) prediction from an image in another layer, for example, a lower layer (lowered by the layer index that can be associated with each of the layers 24). For example, the image 261 * may include a reference 34 to an image 260' belonging to the same access unit but a different layer. In other words, in FIG. 1, the images 261', 260' may be examples of possible reference images for the image 261 currently being coded * . It should be noted that the prediction may be used to predict the coefficients of the image itself, such as the determination of the transform coefficients signaled in the video bitstream 14, or may be used for the prediction of the syntax elements used in the encoding of the image. For example, the image may be encoded using motion vectors that may represent the motion of the image content of the image 26 currently being coded with respect to a previously coded image or an image previous in the image order * . For example, the motion vectors may be signaled in the video bitstream 14. The image 261 *The motion vector may be predicted from a reference image, e.g., either the above and alternative images, using temporal motion vector prediction (TMVP).

[0018] Image 26 may be coded block by block. In other words, image 26 may be subdivided into blocks and / or sub-blocks, e.g., as described with respect to FIG. 2.

[0019] The embodiments described herein may be implemented in the context of versatile video coding (VVC) or other video codecs.

[0020] In the following, some concepts and embodiments will be described with reference to FIG. 1 and the features described with respect to FIG. 1. It is pointed out that the features described with respect to the encoder, video bitstream, or decoder are also to be understood as descriptions of other ones of these entities. For example, a feature described as being present in a video data stream is to be understood as a description of an encoder configured to encode this feature into a video bitstream and a decoder or extractor configured to read this feature from the video bitstream. It is further pointed out that the estimation of information based on an indication coded in the video bitstream may be performed equally on the encoder side and the decoder side. It is further noted that the features described with respect to individual aspects may optionally be combined with each other.

[0021] FIG. 2 shows an example of dividing one of the images 26 into a block 74 and sub-blocks 76. For example, the image 26 may be pre-divided into tree root blocks 72, and then, as exemplarily shown in a tree root block 72' which is one of the tree root blocks 72 in FIG. 2, the tree root block 72 may be recursively subdivided. That is, the tree root block 72 may be divided into blocks, and those blocks may then be divided into sub-blocks and so on. The recursive subdivision may also be called a multi-tree split. The tree root block 72 may be rectangular, and optionally quadratic. The tree root block 72 may also be called a coding tree unit (CTU).

[0022] For example, the above-described motion vector (MV) may be determined, and optionally, in the video bitstream 14, it may be signaled in units of blocks or sub-blocks. In other words, the motion vector may refer to the entire block 74 or the sub-blocks 76. For example, for each block 74 of the image 26, a motion vector may be determined. Alternatively, the motion vector may be determined for each of the sub-blocks 76 of the block 74. In the example, whether one motion vector is determined for the entire block 74 or one motion vector is determined for each of the sub-blocks 76 of the block 74 may vary from block to block. For example, all the images of the coding video sequence 20 belonging to the same layer of the layer 24 may be divided into tree root blocks of equal size.

[0023] Embodiments according to the first and second aspects may be related to temporal motion vector prediction.

[0024] Figure 3 shows a TMVP reference picture determination module 53, hereinafter named the TMVP module 53, according to an embodiment of the first aspect. This may also be optionally implemented in embodiments of the second and third aspects. The TMVP reference picture determination module 53 may be implemented in a video decoder that supports TMVP, for example, a video decoder configured to decode a sequence of coded pictures from a data stream, for example, the decoder 50 of FIG. 1. The TMVP module 53 may also be implemented in a video encoder supporter TMVP, for example, a video encoder configured to decode a sequence of pictures into a data stream, for example, the encoder 10 of FIG. 1. The module 53 is a module for determining the TMVP reference picture 59 of a predetermined picture, for example, the picture 26 being currently coded * of the TMVP reference picture 59 * is a module for determining. For example, the TMVP reference picture for a predetermined picture 26 * is the reference picture of the predetermined picture 26 * from which a predictor of the motion vector of the predetermined picture 26 * is selected.

[0025] TMVP reference picture 59 * To determine, the TMVP module 53 determines a first list 561 and a second list 562 of reference pictures from a plurality of previously decoded pictures. For example, the plurality of previously decoded pictures may include the pictures 26 of the previously decoded access units 22, for example, the picture 261' of FIG. 1 for a predetermined picture 261 * , for example, and may optionally include previously decoded pictures of the same access unit as a predetermined picture 26, such as the picture 260' of FIG. 1 * . This picture is the predetermined picture 26 * and may include previously decoded pictures of the same access unit. This picture is the predetermined picture 26 *It belongs to a layer lower than, i.e., layer 240. Thus, referring to the example of FIG. 1, for example, images 260’ and 261’ may be part of the first list 561 of reference images. In other examples, these two images may be part of the second list of reference images. The first list 561 may optionally include additional reference images. In the example shown in FIG. 3, the second list 562 of reference images is empty. In general, note that either or both of the first and second lists may be empty or may not be empty. The reference images from the first and second lists may be for the inter-prediction of a given image 26’. Since the encoder 10 can determine the first and second lists and signal them in the video bitstream 14, the decoder 50 can derive them from the video bitstream. Alternatively, the decoder 50 may determine the first and second lists independently of or at least partially independently of the explicit signaling. For example, the encoder 10 may signal the first and second lists in the video bitstream 14.

[0026] The TMVP reference image determination module 53 selects one reference image from the first list 561 and the second list 562 of reference images of a given image 26 * such that, for example, when at least one of the first and second lists of reference images is not empty, the TMVP reference image 59 * of the given image 26 * is designated. For this purpose, the module 53 may determine, for example, by the TMVP list selection module 57, one of the first list 561 and the second list 562 of reference images as the TMVP image list 56 * .

[0027] To determine the TMVP image list 56 * (57), when the second list 562 of reference images is empty for a given image 26 * , the encoder 10 sets the first list 561 as the TMVP image list 56 *may be selected. Thus, when the second list 562 of reference pictures is empty for the predetermined picture 26 * , the decoder 50 can assume that the TMVP picture list 56 * is the first list 561 of reference pictures. When the first list of reference pictures is empty and the second list of reference pictures is not empty, the encoder 10 may select the second list 562 as the TMVP picture list 56 * . Thus, in this case, the decoder 50 can assume that the TMVP picture list 56 * is the second list 562 of reference pictures. For the predetermined picture 26 * , when neither the first list nor the second list of reference pictures is empty, the TMVP list selection module 57 of the encoder 10 may select the TMVP picture list 56 * from the first list 561 and the second list 562. The encoder 10 may encode the list selector 58 into the video bitstream 14, and the list selector 58 indicates which of the first and second lists is the TMVP picture list 56 * for the predetermined picture 26 * . For example, the list selector 58 may correspond to the following ph_collocated_from_l0_flag syntax element. The decoder 50 may read the list selector 58 from the video bitstream 14 and select the TMVP picture list 56 * accordingly.

[0028] The TMVP module 53 further performs the selection 59 of the TMVP reference picture 59 * from the TMVP picture list 56 * . For example, the encoder 10 may signal the index of the TMVP reference picture 59 * within the TMVP picture list 56 * to signal the selected TMVP reference picture 59 *It may signal. In other words, the encoder 10 may signal the picture selector 61 in the video bitstream 14. For example, the picture selector 61 may correspond to the following ph_collocated_ref_idx syntax element. The decoder 50 reads the picture selector 61 from the video bitstream 14 and, accordingly, the picture list 56 * from the TMVP reference pictures 59 * and may select.

[0029] The encoder 10 and the decoder 50 may use the TMVP reference picture 59 * to predict the motion vector of a given picture 26 * and may use.

[0030] For example, the TMVP list selection module 57 of the decoder 50 may start TMVP list selection by detecting whether the video bitstream 14 indicates the list indicator 58. If it does, it may select the TMVP picture list 56 * as indicated by the list indicator 58. If the video bitstream 14 does not indicate the list selector 58, the TMVP list selection module 57 may select the first list 561 as the TMVP picture list 56 * if the second list 562 of reference pictures is empty for the given picture 26 * Otherwise, if the first list of reference pictures is empty for the given picture, the TMVP list selection module 57 may select the second list 562 of reference pictures as the TMVP picture list 56 * Or, if the second list 562 is empty for the given picture, the TMVP list selection module 57 may select the second list 562 as the TMVP picture list 56 * if the second list 562 is not empty.

[0031] In other words, in the embodiment of the first aspect, when the first list 561 is empty but the second list 562 is not empty, interference of the list selector 58, for example, PH_collocated_from_L0, can be considered. The first list 561 may be called L0, and the second list 562 may be called L1.

[0032] In other words, in the current specification (of VVC), two syntax elements, namely, ph_collocated_from_l0_flag and ph_collocated_ref_idx, are used to control the picture used for temporal motion vector prediction (TMVP) (or sub-block TMVP). The first syntax element specifies whether the picture used for TMVP is selected from L0 or from L1, and the second syntax element specifies which picture in the selected list is used. These syntax elements exist in either the picture header or the slice header. In the latter case, the prefix has "sh_" instead of "ph_". The picture header is shown as an example in Table 1.

[0033]

Table 1

[0034] That ph_collocated_from_l0_flag is equal to 1 specifies that the collocated picture used for temporal motion vector prediction is derived from reference picture list 0. That ph_collocated_from_l0_flag is equal to 0 specifies that the collocated picture used for temporal motion vector prediction is derived from reference picture list 1. When both ph_temporal_mvp_enabled_flag and pps_rpl_info_in_ph_flag are equal to 1 and num_ref_entries[1][RplsIdx[1]] is equal to 0, the value of ph_collocated_from_l0_flag is presumed to be equal to 1. ph_collocated_ref_idx specifies the reference index of the collocated image used for temporal motion vector prediction. When ph_collocated_from_l0_flag is equal to 1, ph_collocated_ref_idx refers to an entry in reference picture list 0, and the value of ph_collocated_ref_idx ranges from 0 to num_ref_entries[0][RplsIdx[0]] - 1. When ph_collocated_from_l0_flag is equal to 0, ph_collocated_ref_idx refers to an entry in reference picture list 1, and the value of ph_collocated_ref_idx ranges from 0 to num_ref_entries[1][RplsIdx[1]] - 1. If it does not exist, the value of ph_collocated_ref_idx is assumed to be equal to 0.

[0035] There are specific cases that the decoder needs to consider, which is when the reference picture list is empty, that is, when both L0 and L1 have zero entries. If either the L0 or L1 list is empty, currently the specification does not signal ph_collocated_from_l0_flag, assumes that the value of ph_collocated_from_l0_flag is equal to 1, and considers the collocated image to be in L0. However, this is associated with efficiency issues. In fact, regarding the states of L0 and L1, the following scenarios are possible. · L0 is empty and L1 is also empty: ph_collocated_from_l0_flag is assumed to be 1 => OK · L0 is not empty and L1 is empty: ph_collocated_from_l0_flag is assumed to be 1 => OK. · Neither L0 nor L1 is empty: ph_collocated_from_l0_flag is signaled as => OK. · L0 is empty and L1 is not empty: ph_collocated_from_l0_flag is assumed to be 1 => NOT OK.

[0036] When L0 is empty but L1 is not, the presumption that the value of ph_collocated_from_l0_flag is equal to 1 means that when the image of L1 is selected, TMVP (or sub-block TMVP) can still be used, but it means that TMVP is not used, thus resulting in a loss of efficiency.

[0037] Therefore, in one embodiment, the decoder (or encoder) determines the value of ph_collocated_from_l0_flag according to which list is empty, for example, as described above with respect to the TMVP list selection module 57 in FIG. 3. That ph_collocated_from_l0_flag is equal to 1 specifies that the collocated image used for temporal motion vector prediction is derived from reference picture list 0. That ph_collocated_from_l0_flag is equal to 0 specifies that the collocated image used for temporal motion vector prediction is derived from reference picture list 1. When both ph_temporal_mvp_enabled_flag and pps_rpl_info_in_ph_flag are equal to 1, the following occurs. · When num_ref_entries[1][RplsIdx[1]] is equal to 0, the value of ph_collocated_from_l0_flag is presumed to be equal to 1. · Otherwise, the value of ph_collocated_from_l0_flag is presumed to be equal to 0.

[0038] As an alternative to determining the TMVP picture list according to which list is empty, in another embodiment, when L0 is empty, there is a bitstream constraint that L1 needs to be empty. This is due to bitstream constraints or syntax prohibitions (see Table 2).

[0039] Thus, according to an alternative embodiment of the TMVP module 53 in FIG. 3, when the second list 562 is empty, the encoder 10 selects the TMVP reference image 59 from the first list 561 * and when the second list 562 is not empty, the TMVP reference image 59 is selected from the first list 561 or the second list 562. * In other words, the above-described TMVP list selection module 57 may select the TMVP image list 56 * as the first list 56 * 1 when the second list 562 is empty, and may select either the first list or the second list when the second list 562 is not empty.

[0040] Thus, when determining the list of reference images (55), the decoder 50 can presume that the second list 562 is empty when the first list 561 is empty. Thus, in the TMVP list selection 57, the decoder 50 reads the list selector 58 from the video bitstream 14, and when neither the first list 561 nor the second list 562 is empty, may select the TVMP reference image 59 according to the list selector 58. * When the first list 561 and the second list 562 are not empty, the decoder 50 may select the first list 561 as the TMVP image list 56 * and otherwise, the decoder 50 may select the first list 561 as the TMVP image list 56 * and may select.

[0041] According to an embodiment, the decoder 50 may perform the list determination (55) by reading from the video bitstream 14 information regarding how to populate the first list 561 from a plurality of previous decoder images for the first list of reference images. When the decoder 50 does not presume that the second list 562 is empty, the decoder 50 may read information regarding how to populate the second list 562 from the video bitstream 14.

[0042] According to the latter embodiment, when the decoder estimates that the second list 562 is empty, if the first list 561 is empty, the encoder 10 and the decoder 50 may encode a predetermined image 26 without using TMVP when the first list of reference images is empty. * without using TMVP when the first list of reference images is empty.

[0043] As described above, the bitstream constraint may instead be implemented in the form of syntax, for example, in the construction of the list of reference images. An implementation example is shown in Table 2.

[0044]

Table 2

[0045] Therefore, when it is estimated that ph_collocated_from_l0_flag is 1, if both lists are empty or only list L1 is empty and there is one non-empty list, that list is used for TMVP.

[0046] Therefore, according to a further alternative embodiment of the TMVP module 53 in FIG. 3, the encoder 10 may perform TMVP list selection 57 by selecting a TMVP reference image 56 from the first list 561 when the second list 572 is empty, and selecting a TMVP reference image 56 from the first list 561 or the second list 562 when the second list 562 is not empty. * from the first list 561 when the second list 572 is empty, and selecting a TMVP reference image 56 from the first list 561 or the second list 562 when the second list 562 is not empty. * from the first list 561 or the second list 562 when the second list 562 is not empty.

[0047] Hereinafter, embodiments according to the second aspect will be described with reference to FIGS. 1 to 3 and FIG. 4.

[0048] FIG. 4 shows a first image 261 *, for example, shows an example of dividing the currently coded image into the tree root block 721. The tree root block 721 is recursively divided into the block 74 and the sub-block 76, for example, as described with respect to FIG. 2. FIG. 4 further shows the second image 260', for example, the image 260' of FIG. 1 that can be an inter-layer reference image of the first image 261'. The second image 260' is divided into the tree root block 720. The first example of the division into the tree root block 720 is indicated by the dashed line. The dotted line further shows an alternative second example for dividing the second image 260' into the tree root block 720', and the tree root block 720' of the second example is smaller than the tree root block 720 of the first example. According to the first example of the division, the tree root block 720 has the same size as the tree root block 721 of the first image 261 * . In the second example of the division, the tree root block 720' is smaller than the tree root block 721 of the first image 261 * .

[0049] For the TMVP of the first image 261 * , the encoder 10 and the decoder 50 may determine one or more MV candidates. For this purpose, one or more MV candidates may be determined from each of one or more reference images of the image 261 * . For example, the encoder 10 and the decoder 50 may determine one or more MV candidates from one or more of one or more lists of reference images, for example, two lists, for example, the lists 561 and 562 described with respect to FIG. 3. In addition to the reference images of the first image 261 * , as shown in FIG. 1, there may be inter-layer reference images such as the second image 260' that are temporally collocated (in the same position), that is, belong to the same access unit 22, in the first image. As described with respect to FIG. 2, the image 26 may be coded in block units or sub-block units, and the encoder 10 is the currently coding tree root block 721 of the currently coded image 261 * ​* block 74 currently being coded * may determine one MV for block 74 currently being coded, or, alternatively, for each sub-block 76 of block 74 currently being coded * may determine one MV. The sub-blocks currently being coded may be referenced using reference numeral 76 * The encoder 10, and optionally also the decoder 50, may determine one or more MV candidates for block 74 currently being coded * or sub-block 76 currently being coded * from one reference picture, e.g., a second picture 260’, for block 74 currently being coded or sub-block 76 currently being coded. One or more MV candidates may be determined from different positions or locations within the reference picture. For example, for the bottom-right MV candidate of block 74 currently being coded * or sub-block 76 currently being coded * the MV of the reference picture located at a reference position 71’ within the reference picture may be selected as the MV candidate, and the reference position 71’ of the bottom-right MV candidate is collocated (at the same position) with the bottom-right position 71 within the picture 261’ currently being coded. For example, the bottom-right position 71 of block 74 currently being coded * or sub-block 76 currently being coded * may be a position adjacent downward and to the right of block 74 * or sub-block 76 *

[0050] In FIG. 4, for a first example of subdividing a second picture 260’ into a tree root block 720, the collocated block of block 74 currently being coded in a first picture 261’ * is indicated using reference numeral 714 * As shown in FIG. 4, in this first example of subdivision where the size of the tree root block 720 is equal to the size of a tree root block 721 of a first picture 261 * the reference position 71’ is within the bottom-right block 714’ of block 714 * and block 714’ is within the first picture 261 * ​The currently coding block 74 * The collocated block 714 * And the same tree root block 720 * Inside. For example, the first image 261 * Of the block 74 * One MV candidate for the coding of may be the MV of block 714'.

[0051] The first image 261 * In a second example of subdividing the reference image 260' into a tree root block 720' smaller than the tree root block 721 of, the reference position 71 may be located outside the collocated tree root block 720'. Specifically, the reference position 71 may be located outside the row of the tree root block where the collocated tree root block 720' of the currently coding block 74 * Is located. Therefore, from the tree root block where the reference position 71 is located, the currently coding block 74 * Or the currently coding sub-block 76 * To use the MV as an MV candidate for, the encoder 10 and the decoder 50 may need to hold one or more tree root blocks in the image buffer that exceed the currently coding tree root block of the second image 260' or exceed the current row of the currently coding tree root block. Therefore, in this second example of tree root block subdivision, using the reference position 71 for MV prediction may involve inefficient buffer usage.

[0052] In other words, as described with respect to FIG. 3, the current specification uses two syntax elements, namely ph_collocated_from_l0_flag and ph_collocated_ref_idx, to control the images used for TMVP (or sub-block TMVP), where the latter specifies which entry in each list is used as the reference image, as described for example with respect to FIG. 3. When TMVP (e.g., TMVP of one of the MVs in block 74 of FIG. 2) or sub-block TMVP (e.g., TMVP of one of the MVs in sub-block 76 of FIG. 2) is used in the bitstream, the motion vector (MV) of the current block or sub-block of image 26 * is derived from the MV of the reference image indicated by ph_collocated_ref_idx, and that MV is added to a candidate list from which it can be selected as a predictor of the actually used motion vector of the block / sub-block. This MV prediction selects from among MVs at different positions according to the boundaries of the largest block (CTU) of image 26 * , for example, the tree root block 72 of FIG. 2. In most scenarios, the reference image indicates the same CTU boundary.

[0053] More specifically, in the case of TMVP, for example, the lower right TMVP MV candidate 71 of (e.g., block 74 * or sub-block 76 that is currently being coded) crosses the CTU row boundary of the current block (i.e., for example, tree root block 74 of FIG. 4 **For the case where the block 74 or sub-block 76 is located beyond the row boundary of the tree root block 72 to which it belongs, or does not exist (for example, beyond the image boundary or intra-coded without using motion vectors), the TMVP candidate is not obtained from the bottom-right block. Instead, an alternative (collocated) MV candidate of the CTU row, that is, from the collocated block of (the block 74 or sub-block 76 for which the MV candidate is to be determined), is derived. Further, when the sub-block TMVP MV candidate is obtained from a position outside the collocated CTU, the MV used to determine the position of the sub-block TMVP MV is corrected (clipped) during derivation so that it points to the position where the sub-block TMVP MV candidate is obtained from a position belonging to the collocated CTU. For the sub-block TMVP, it should be noted that first an MV is obtained (for example, first from the spatial candidates), and that MV is used to identify the position of the block used to select the sub-block TMVP candidate within the collocated image. If the temporal MV used to identify the position points to a position outside the collocated CTU, the MV is clipped so that the sub-block TMVP is obtained from a position within the collocated CTU.

[0054] However, there are cases where the reference picture does not have the same CTU size, e.g., the tree root block 721’ in FIG. 4, and thus has no boundary. The described TMVP or sub-block TMVP process is not clear when the CTU size is changed. When applied in such a case, in the state-of-the-art, buffer-unfriendly TMVP and sub-block TMVP candidates outside the current CTU (row) or the CTU (row) of the reference picture, e.g., at position 71’, are selected. By this aspect of the present invention aiming to lead the MV candidate selection accordingly or to prevent such cases from occurring, the embodiments are freed from this burden. The described scenario occurs in a quality scalable multi-layer bitstream having two (or more) layers. Here, as shown on the left side of FIG. 5, one layer depends on the other layer, and the dependent layer has a larger CTU size than the reference layer. In such a case, the decoder needs to fetch the associated MV candidates from four small CTUs of the reference picture, and these CTUs do not necessarily occupy contiguous memory regions that have an adverse effect on the memory bandwidth during access. On the other hand, on the right side of FIG. 5, the CTU size of the reference picture is larger than that of the dependent layer. In such a case, when decoding the current block of the dependent layer, there is no scattered data among the relevant data of the reference layer.

[0055] As described, when the CTU sizes of the two layers (parameters of each SPS) are different, the CTU boundaries of the current picture and the reference picture (ILRP in this scenario) are not aligned, but TMVP and sub-block TMVP are activated and are not prohibited from being used by the encoder.

[0056] According to an embodiment of the second aspect, the encoder 10, for example, the encoder 10 in FIG. 1, is for layer video coding as described with respect to FIG. 1, and as described with respect to FIGS. 2 and 4, each image is pre-divided into one or more tree root blocks 72, and the image 26 is subdivided into blocks 74 by performing recursive block division on each tree root block 72 of each image 26, for encoding the image 26 of the video 20 into a multi-layer data stream 14 (or a multi-layer video bitstream 14). Each image 26 is associated with one of the layers 24 as described with respect to FIG. 1. The sizes of the tree root blocks 72 may be equal for all images belonging to the same layer 24. The encoder 10 according to an embodiment of the second aspect, for example, as described with respect to FIG. 3, for the first image 261 of the first layer 241 * (see FIG. 1), populates a list of reference images, for example, list 561 or list 562, from a plurality of previously coded images. As described with respect to FIG. 4, the list of reference images may include images of the same layer and different timestamps in the multi-layer data stream 14 upstream of the first image 261 * in which the image is encoded, for example, image 261' in FIG. 1. Further, the list of reference images may include one or more second images, the second images belonging to different layers, for example, the second layer 240, and being temporally aligned with the first image 261 * , for example, the second image 260'. The encoder may signal in the multi-layer video bitstream 14 the size of the tree root block of the first layer 241 to which the first image 261 * belongs and the size of a different layer, for example, the second layer 240 to which the second image belongs. The encoder 10 and the decoder 50 may use motion compensation prediction to inter-predict the inter-prediction blocks of the first image 261 * and may use the list of reference images to predict the motion vectors of the inter-prediction blocks.

[0057] According to the first embodiment of the second aspect, the size of the tree root block of the second image 260’ is equal to or an integral multiple of the size of the tree root block of the first image 260 * For example, the video encoder may populate the list 56 of reference images, or the two lists 561, 562 of reference images described with respect to FIG. 3, such that for each reference image in the list of reference images, the size of the tree root block is equal to or an integral multiple of the size of the tree root block of the first image 261 * For example, the video encoder may populate the list 56 of reference images, or the two lists 561, 562 of reference images described with respect to FIG. 3, such that for each reference image in the list of reference images, the size of the tree root block is equal to or an integral multiple of the size of the tree root block of the first image 261

[0058] In other words, in the example of the first embodiment, the use of TMVP and sub-block TMVP is disabled by imposing constraints on the ILRP of the reference image list of the current image as follows. - The following constraints apply to the image referred to by each ILRP entry if it exists in RefPicList[0] or RefPicList[1] of the slice of the current image. o The image is assumed to be in the same AU as the current image. o The image is assumed to be present in the DPB. o The image has a refPicLayerId with a nuh_layer_id smaller than that of the current image. o The image has the same value as the current image, or a value of sps_log2_ctu_size_minus5 greater than that of the current image. o One of the following constraints applies. o The image is assumed to be an IRAP image. o The image has a TemporalId less than or equal to Max(0, vps_max_tid_il_ref_pics_plus1[currLayerIdx][refLayerIdx]-1). Here, currLayerIdx and refLayerIdx are equal to GeneralLayerIdx[nuh_layer_id] and GeneralLayerIdx[refpicLayerId], respectively.

[0059] The worst case of the above problem occurs when the CTU size of the reference image is small. This is because the memory bandwidth requirements of TMVP and sub-block TMVP increase. When the CTU of the reference image is small, it is not so important to obtain candidates for TMVP or sub-block TMVP. Therefore, the foregoing embodiments are applicable only when the CTU size of the reference image is small. Thus, when the CTU size of the reference image is larger than the CTU size of the current image as in the foregoing embodiments, there is no problem.

[0060] Nevertheless, according to an example of the first embodiment of the second aspect, the constraint is that for each inter-layer reference image in the list of reference images, that is, for each of the second images (and, for example, images of the same layer may have tree root blocks of the same size, and as a result, for all reference images), the size of its tree root block needs to be equal to the size of the tree root block of the first image 261 * or an integer multiple thereof.

[0061] Thus, in the example of the first embodiment, the use of TMVP and sub-block TMVP is invalidated by imposing the following constraints on the ILRP of the reference image list of the current image. · The following constraints apply to the image referenced by the entry when each ILRP entry exists in RefPicList[0] or RefPicList[1] of the slice of the current image. o Assume that the image is in the same AU as the current image. o Assume that the image exists in the DPB. o Assume that the image has a refPicLayerId with a nuh_layer_id smaller than the nuh_layer_id of the current image. o Assume that the image has the same value of sps_log2_ctu_size_minus5 as the current image (i.e., the image has the same tree root block size as the image 261 currently being coded). * ). One of the following constraints applies. o Assume that the image is an IRAP image. o Assume that the image has a TemporalId less than or equal to Max(0, vps_max_tid_il_ref_pics_plus1[currLayerIdx][refLayerIdx] - 1). Here, currLayerIdx and refLayerIdx are equal to GeneralLayerIdx[nuh_layer_id] and GeneralLayerIdx[refpicLayerId], respectively.

[0062] However, this embodiment is a very restrictive constraint that does not recognize different CTU sizes in the dependent layer and the reference layer and hinders any form of prediction such as sample prediction.

[0063] Therefore, according to the second embodiment of the second aspect, the list of reference images may include images of a second type of image that do not necessarily have a size smaller than or equal to the tree root block 720, i.e., images of different layers such as the second layer 240. However, the list of reference images may include any of the second images (images of the second layer 240). According to this embodiment, the encoder 10 * specifies one image from the list of reference images as the TVMP reference image for the first image 261, for example, as described with respect to FIG. 3. According to this embodiment, the encoder 10 specifies the TVMP reference image so that the TVMP reference image is not the second image 260, and the size of its tree root block 720 is * smaller than the size of the tree root block 721 of the first image 261. In other words, when the encoder 10 selects one of the second images 260 as the TVMP reference image, i.e., when the encoder 10 selects an image from a different layer such as the layer 240 as the TVMP reference image for the first image 261 * the TVMP reference image is the first image 261 *It has a tree root block of the same size as or larger than that. The encoder 10 may signal, in the multi-layer video data stream 14, a pointer for identifying a TVMP reference image from the images in the list of reference images. The encoder 10 uses the TVMP reference image to predict the motion vector of the inter-prediction block of the first image 261 * For predicting the motion vector of the inter-prediction block of * .

[0064] Instead of restricting the population of the list, by restricting the designation of the TVMP reference image from the list of reference images, the use of the second image 260 having a tree root block size smaller than that of the first image becomes possible for other prediction tools.

[0065] For example, the list of reference images may be one of the first list 561 and the second list 562 as described with respect to FIG. 3. However, the method of selecting the TVMP reference list does not necessarily have to follow the method described with respect to FIG. 3 and may be implemented in another way, for example, as described for the prior art.

[0066] For example, when the TVMP reference image is an image of a layer different from the second image, i.e., the image being currently coded, the encoder 10 may select the TVMP reference image so that none of the criteria in the following set of criteria are satisfied. - The size of the tree root block 720 of the second image 260 is smaller than the size of the tree root block 721 of the first image 261 * - The size of the TVMP reference image is different from the size of the first image 261 * - The scaling window of the TVMP image, which is the scaling window used for motion vector scaling and offset, is different from the scaling window of the first image 261 * - The sub-image subdivision of the TVMP reference image 260’ is different from the image subdivision of the first image 261 *

[0067] ​​​​Encoder 10, when inter-predicting the inter-prediction blocks of the first image 261 * for each inter-prediction, activates a set of one or more inter-prediction refinement tools according to the reference images in the list of reference images from which each inter-prediction block is inter-predicted, where the reference images satisfy any of the above criteria sets. For example, the set of inter-prediction refinement tools may include one or more of TVMP, PROF, wrap-around, VDOF, and DVMR.

[0068] In other words, according to an example of the second embodiment, the problem is solved by imposing constraints on the syntax element sh_collocated_ref_idx indicating the reference image used for TVMP and sub-block TVMP. sh_collocated_ref_idx specifies the reference index of the collocated image used for temporal motion vector prediction. When sh_slice_type is equal to P, or when sh_slice_type is equal to B and sh_collocated_from_l0_flag is equal to 1, sh_collocated_ref_idx refers to an entry in reference picture list 0, and the value of sh_collocated_ref_idx ranges from 0 to NumRefIdxActive[0] - 1. When sh_slice_type is equal to B and sh_collocated_from_l0_flag is equal to 0, sh_collocated_ref_idx refers to an entry in reference picture list 1, and the value of sh_collocated_ref_idx ranges from 0 to NumRefIdxActive[1] - 1. When sh_collocated_ref_idx does not exist, the following applies. - When pps_rpl_info_in_ph_flag is equal to 1, the value of sh_collocated_ref_idx is presumed to be equal to ph_collocated_ref_idx. - Otherwise (when pps_rpl_info_in_ph_flag is equal to 0), the value of sh_collocated_ref_idx is assumed to be equal to 0. Set colPicList to be equal to sh_collocated_from_10_flag?0:1. The image referred to by sh_collocated_ref_idx is the same for all non-I slices of the coded image, the value of RprConstraintsActiveFlag[colPicList][sh_collocated_ref_idx] is equal to 0, and the value of sps_log2_ctu_size_minus5 of the image referred to by sh_collocated_ref_idx is greater than or equal to the value of sps_log2_ctu_size_minus5 of the current image, which is a requirement for bitstream compliance. Note - In the above constraints, the collocated image needs to have the same spatial resolution, the same scaling window offset, and the same or smaller CTU size as the current image.

[0069] Again, the above example can prevent the case where the selected inter-layer reference image has a smaller tree root block size than the first image. In another example, the encoder 10 may specify that the TVMP reference image of the first image 261 * has a tree root block size equal to the size of the tree root block 721 of the image 261 where the reference image is the selected TVMP reference image * .

[0070] Thus, another exemplary embodiment is as follows. sh_collocated_ref_idx specifies the reference index of the collocated image used for temporal motion vector prediction. When sh_slice_type is equal to P, or when sh_slice_type is equal to B and sh_collocated_from_l0_flag is equal to 1, sh_collocated_ref_idx refers to an entry in reference picture list 0, and the value of sh_collocated_ref_idx ranges from 0 to NumRefIdxActive[0] - 1. When sh_slice_type is equal to B and sh_collocated_from_l0_flag is equal to 0, sh_collocated_ref_idx refers to an entry in reference picture list 1, and the value of sh_collocated_ref_idx ranges from 0 to NumRefIdxActive[1] - 1. When sh_collocated_ref_idx does not exist, the following applies. - When pps_rpl_info_in_ph_flag is equal to 1, the value of sh_collocated_ref_idx is presumed to be equal to ph_collocated_ref_idx. - Otherwise (pps_rpl_info_in_ph_flag is equal to 0), the value of sh_collocated_ref_idx is presumed to be equal to 0. Set colPicList to be equal to sh_collocated_from_10_flag? 0:1. The picture referred to by sh_collocated_ref_idx is the same for all non-I slices of the coded picture, the value of RprConstraintsActiveFlag[colPicList][sh_collocated_ref_idx] is equal to 0, and the value of sps_log2_ctu_size_minus5 of the picture referred to by sh_collocated_ref_idx is equal to the value of sps_log2_ctu_size_minus5 of the current picture, which is a requirement for bitstream conformance. Note - In the above constraints, the collocated picture needs to have the same spatial resolution, the same scaling window offset, and the same CTU size as the current picture.

[0071] According to the third embodiment of the second aspect, the encoder 10 and the decoder 50 use a predetermined image 261 according to the reference image used to inter-predict each inter-prediction block that satisfies any of the above reference sets * when predicting the inter-prediction block of. Here, this constraint for using the inter-prediction refinement tool is not necessarily limited to the multi-layer case where the reference image is an inter-layer reference image of the coded image 26 * .

[0072] In an example, the encoder 10 and the decoder 50 can derive a list of reference images from which the reference image is selected as described with respect to FIG. 3, but different approaches for signaling or selecting the reference image may also be possible.

[0073] In other words, according to the embodiment, as described above, the encoder 10 and the decoder 50 that code (i.e., encode in the case of the encoder 10 and decode in the case of the decoder 50) an image in units of blocks obtained by recursive block division of the tree root block * may use a set of one or more inter-prediction refinement tools depending on whether any of the following reference sets are satisfied when inter-predicting the inter-prediction block of the currently coded image 26 - The size of the tree root block of the currently coded image 26 * is smaller than the size of the tree root block of the reference image. - The size of the reference image is equal to the size of the currently coded image 26 * . - For example, the scaling window of the reference image in which the scaling window is used for the scaling of the motion vector and the offset is equal to the scaling window of a predetermined image, and is different with respect to the offset of the scaling window. -The sub-image subdivision of the reference image is equal to the sub-image subdivision of a predetermined image. That is, for example, the sub-image subdivision differs with respect to the number of sub-images of the sub-image subdivision.

[0074] In an example, the reference set may include that the size of the tree root block of the reference image is equal to the size of the tree root block of a predetermined image 26 * In an example, the encoder 10 and the decoder 50 may selectively activate a set of inter-prediction refinement tools when all of a subset of the reference set are satisfied. In other examples, the encoder 10 and the decoder 50 may activate an inter-prediction refinement tool when all of the reference set are satisfied.

[0075] For example, according to the third embodiment, the constraint is represented by the derived variable RprConstraintsActiveFlag[refPicutre][currentPic], and this derived variable is derived by comparing the characteristics of the current image and the reference image (image size, scaling window offset, number of sub-images, etc.). This variable is used to impose constraints on the specified ph_collocated_ref_idx of the image header of sh_collocated_ref_idx of the slice header. In this embodiment, the sizes of the CTUs in the reference image and the current image (sps_log2_ctu_size_minus5) are incorporated into each derivation of RprConstraintsActiveFlag[refPicutre][currentPic], and as a result, when the CTU size of the reference image is larger, the CTU size RprConstraintsActiveFlag[refPicutre][currentPic] of the current image is derived as 1.

[0076] Similarly, but where the criterion is that the size of the tree root block of the reference image is a predetermined image 26

[0077] In an example, the reference set may include that the size of the tree root block of the reference image is equal to the size of the tree root block of a predetermined image 26 *When including being equal to the size of the tree root block of, the constraint may be represented by the derived variable RprConstraintsActiveFlag[refPicutre][currentPic], and this derived variable is derived by comparing the characteristics of the current image and the reference image (image size, scaling window offset, number of sub-images, etc.). This variable is used to impose constraints on the specified ph_collocated_ref_idx of the image header of sh_collocated_ref_idx in the slice header. In this embodiment, the size of the CTU (sps_log2_ctu_size_minus5) in the reference image and the current image is incorporated into each derivation of RprConstraintsActiveFlag[refPicutre][currentPic], and as a result, when the CTU sizes are different, RprConstraintsActiveFlag[refPicutre][currentPic] is derived as 1.

[0078] In such a case, in this embodiment, tools such as PROF, wraparound, BDOF, and DMVR are not permitted because it is not desirable to allow different CTU sizes when such tools are used.

[0079] The foregoing description has focused only on TMVP and sub-block TMVP, but further problems have been identified that apply when different CTUs are used in different layers. For example, TMVP and sub-block TMVP are permitted for images with different CTU sizes with associated "drawbacks", but problems occur when combined with sub-images. In fact, when sub-images are used and used together with layer coding, there is a series of constraints necessary to align the sub-image grid. This is done for the layer having sub-images in the dependency tree.

[0080] FIG. 6 shows an example of sub-image subdivision into two sub-images 281′ and 281″ of image 261. Image 261 may belong to a first layer such as layer 241 and may belong to access unit 22 as described with respect to FIG. 1. In other words, image 261 of FIG. 6 may be one of the images 261 of FIG. 1. Image 261 may depend on a reference image 260, e.g., the second layer 240 of FIG. 1. In other words, image 260 may be an inter-layer reference image of image 261, and the layer to which image 260 belongs may be referred to as a reference layer of the layer to which image 261 belongs. For example, all images belonging to the same coding video sequence 20 and the same layer may be subdivided with the same sub-image subdivision, i.e., subdivided into the same number of sub-images having equal sizes. It should also be noted that all images of coding video sequence 20 belonging to the same layer 24 may be divided into tree root blocks of equal size. Encoder 10 may be configured to subdivide image 261 of the first layer 241 and a reference layer of the first layer 261, i.e., image 260 belonging to the second layer 240, into two or more common numbers of sub-images 28. Such subdivision into two or more common numbers of sub-images is shown in FIG. 6, where image 260 of reference layer 240 is subdivided into two sub-images 280′ and 280″. Encoder 10 may encode sub-images 28′ and 28″, i.e., sub-images belonging to a common layer, e.g., layer 241 or layer 240, independently of each other. That is, encoder 10 may encode independently coded sub-images without spatial prediction of one part of sub-image 28′ by another sub-image 28″ of the sub-images. For example, encoder 10 may apply sample padding at the boundary between images 28′ and 28″ such that the coding of the two sub-images is independent of each other. Also, encoder 10 may clip motion vectors at the boundary between sub-images 28′ and 28″. Optionally, encoder 10 may signal in video data stream 14 that sub-images of a layer are coded independently, e.g., by signaling sps_subpic_treated_as_pic_flag = 1.Encoder 10 may code sub - image 28 into video bit - stream 14 in units of blocks 74, 76. Blocks 74, 76 result from the division of one or more tree - root blocks 72 into which sub - image 28 is divided, as described with respect to FIGS. 2 and 4 for image 26. Encoder 10 may divide the first layer 241 of image 26 and the sub - images 281, 280 of the reference layer 240 of the first layer 241, which are subdivided into a common number of sub - images, into tree - root blocks of equal size. In other words, when the image 261 of the first layer 241 and the reference image 260 of the reference layer 240 are subdivided into a common number of two or more sub - images, encoder 10 may subdivide sub - images 281, 280 such that the tree - root blocks of sub - image 281 of the first layer 241 and sub - image 280 of the reference layer 240 are of the same size.

[0081] In other words, in another alternative embodiment, by extending the sub - image - related constraints as follows, the problem is solved only for the case of independent sub - images. That sps_subpic_treated_as_pic_flag[i] is equal to 1 specifies that the i - th sub - image of each coded image in CLVS is treated as an image in the decoding process excluding the in - loop filtering operation. That sps_subpic_treated_as_pic_flag[i] is equal to 0 specifies that the i - th sub - image of each coded image in CLVS is not treated as an image in the decoding process excluding the in - loop filtering operation. If it does not exist, the value of sps_subpic_treated_as_pic_flag[i] is presumed to be equal to 1. When sps_num_subpics_minus1 is greater than 0 and sps_subpic_treated_as_pic_flag[i] is equal to 1, for each CLVS of the current layer that refers to the SPS, if targetAuSet is set to all AUs starting from the AU containing the first picture of the CLVS in decoding order and ending at the AU containing the last picture of the CLVS in decoding order, then for the targetLayerSet composed of the current layer and all layers that have the current layer as the reference layer, it is a requirement for bitstream compliance that all of the following conditions are true. - For each AU in targetAuSet, assume that all pictures of the layers in targetLayerSet have the same value of pps_pic_width_in_luma_samples and the same value of pps_pic_height_in_luma_samples. - Assume that all SPSs referred to by the layers in targetLayerSet have the same value of sps_num_subpics_minus1, and have the same values of sps_subpic_ctu_top_left_x[j], sps_subpic_ctu_top_left_y[j], sps_subpic_width_minus1[j], sps_subpic_height_minus1[j], and sps_subpic_treated_as_pic_flag[j] respectively, for each value of j in the range from 0 to sps_num_subpics_minus1 and for each value of j of sps_log2_ctu_size_minus5. - For each AU in targetAuSet, assume that all pictures of the layers in targetLayerSet have the same value of SubpicIdVal[j] for each value of j in the range from 0 to sps_num_subpics_minus1.

[0082] According to an embodiment, the encoder 10 and the decoder 50 are configured to process a block, for example, the block 74 that is currently being coded in the picture 26 that is currently being coded in the first layer 241 * that is currently being coded *or the sub-block 76 currently being coded * may be inter-predicted using coding parameters such as the motion vector of the corresponding block of the image in the reference layer 240. For example, block 714 * may be the corresponding block of block 74 in FIG. 4 * . In other words, the corresponding block may be collocated with (at the same position as) the currently-coded block at a position within each image, i.e., within each reference image, and may, for example, correspond to, or be referred to by, the motion vector of the currently-coded block. The encoder 10 and the decoder 50 may use the coding parameters of the corresponding block to predict coding parameters such as the motion vector of the currently-coded block.

[0083] By subdividing the sub-images of the reference layer 240 into tree root blocks having the same size as the tree root block of the first layer that depends on the reference layer, it is possible to ensure that the sub-images 28 are coded independently of each other. In this regard, the same considerations described with respect to FIG. 4 regarding the boundaries of the tree root blocks apply. In other words, the considerations regarding FIG. 4 explaining the advantages of the equally-sized tree root blocks for the subdivision of the image 26 may also apply to the subdivision of the sub-images 28.

[0084] FIG. 7 shows an example of a decoder 50 and a video bit stream 14 according to an embodiment of a third aspect of the present invention. The decoder 50 and the video bit stream 14 may optionally correspond to the decoder 50 and the video bit stream 14 of FIG. 1. According to an embodiment of the third aspect, the video bit stream 14 is a multi-layer video bit stream including access units 22, and each access unit 22 includes one or more pictures 26 of a coded video sequence 20 coded in the video bit stream 14. For example, the description of layer 24 and access unit 22 in FIG. 1 may be applicable. Thus, each picture 26 belongs to one of the layers 24, and the association between the picture and the layer is indicated by a subscript. That is, picture 261 belongs to the first layer 241, and picture 260 belongs to the second layer 240. In FIG. 7, the pictures 26 and the access units 22 to which the pictures belong are shown in the picture order of the coded video sequence 20, i.e., the presentation order thereof. Each access unit 22 belongs to a temporal sublayer of a set of temporal sublayers. For example, in FIG. 7, the access unit 221 to which a picture having a superscript 1 belongs belongs to the first temporal sublayer, and the access unit 222 to which a picture referred to using a reference sign having a superscript 2 belongs belongs to the second temporal sublayer. As shown in FIG. 7, access units belonging to different temporal sublayers do not necessarily have the same set of pictures of layer 24. For example, in FIG. 7, the access unit 221 may have pictures of the first layer 241 and the second layer 240, and the access unit 222 may have only pictures of the first layer 241. The temporal sublayers may be hierarchically ordered and may be indexed by an index representing the hierarchical order. In other words, the highest and lowest temporal sublayers may be defined within the set of temporal sublayers by the hierarchical order of the temporal sublayers. Note that the video bit stream 14 may optionally have additional temporal sublayers and / or additional layers.

[0085] According to the embodiment of the third aspect, the video bitstream 14 encodes therein an output layer set indication 81 indicating one or more output layer sets 83. The OLS indication 81, for the OLS 83, is a subset of the layers 24 of the multi-layer video bitstream 14 belonging to the OLS (not necessarily a proper subset, i.e., the OLS may indicate that all of the layers of the multi-layer audio bitstream 14 belong to the OLS). For example, the OLS may be an indication of a sub-bitstream extractable or decodable (not necessarily proper) from the video bitstream 14, and the sub-bitstream includes a subset of the layers 24. For example, by extracting or decoding a subset of the layers of the video bitstream 14, the decoded coded video sequence may be quality scalable and thus bitrate scalable.

[0086] According to an embodiment of the third aspect, the video bitstream 14 further includes a video parameter set (VPS) 91. The VPS 91 includes one or more bitstream compliance sets 86, such as a virtual reference decoder (HRD) parameter set. The video parameter set 91 further includes one or more buffer requirement sets 84, such as a decoded picture buffer (DPB) parameter set. The video parameter set 91 further includes one or more decoder requirement sets 82, such as a profile tier level parameter set (PTL set). Each of the bitstream compliance set 86, the buffer requirement set 84, and the decoder requirement set 82 is associated with each of the respective temporal subset indications 96, 94, 92 indicated by the video parameter set 91. The constraint for the maximum temporal sublayer for each of the parameter sets, i.e., the bitstream compliance set 86, the buffer requirement set 84, and the decoder requirement set 82, may represent the upper limit of the number of temporal sublayers that the parameters of each parameter set refer to. In other words, the parameters signaled by the parameter set may be valid for a (not necessarily appropriate) subsequence of the coded video sequence 20 defined by the set of layers and the set of temporal sublayers, or for a sub-bitstream of the video bitstream 14, and the constraint for the maximum temporal sublayer for each parameter set indicates the maximum temporal sublayer of the sub-bitstream or subsequence that each parameter set refers to.

[0087] According to the first embodiment of the third aspect, the decoder 50 may be configured to receive a maximum temporal sublayer indication 99. The maximum temporal sublayer indication 99 indicates the maximum temporal sublayer of the multi-layer video bitstream 14 decoded by the decoder 50. In other words, the maximum temporal sublayer indication 99 may signal to the decoder 50 which set or subset of the temporal layers of the video bitstream 14 the decoder 50 is to decode. Thus, the decoder 50 may receive the maximum temporal sublayer indication 99 from an external signal. For example, the maximum temporal sublayer indication 99 may be included in the video bitstream 14 or provided to the decoder 50 via an API. Upon receiving the maximum temporal sublayer indication 99, the decoder 50 may decode the video bitstream 14 or a portion thereof, as long as it belongs to the set of temporal sublayers indicated by the maximum temporal sublayer indication 99. For example, the decoder 50 may be further configured to receive an indication of the OLS to be decoded. Upon receiving the indication of the OLS to be decoded and the maximum temporal sublayer indication 99, the decoder 50 may decode the layer indicated by the OLS to be decoded up to the temporal sublayer indicated by the maximum temporal sublayer indication 99. However, there may be situations or scenarios where the decoder 50 does not receive one or both of the external indications of the temporal sublayer, namely, the maximum temporal sublayer indication 99 and the indication of the OLS to be decoded. In such situations, the decoder 50 may determine a missing indication, for example, based on the information available in the video bitstream 14.

[0088] According to the first embodiment of the third aspect, when the maximum time sublayer indication 99 cannot be obtained, for example, cannot be obtained from the data stream and / or cannot be obtained through other means, the decoder 50 determines that the maximum time sublayer to be decoded is equal to the maximum time sublayer indicated by the constraint 92 on the maximum time sublayer of the decoder requirement set 82 associated with the OLS 83. In other words, the decoder 50 may use the constraint 92 indicated for the decoder requirement set 82 associated with the OLS decoded by the decoder 50 for the maximum time sublayer.

[0089] When decoding the multi-layer video bitstream 14, if the time sublayer does not exceed the maximum time sublayer to be decoded, the decoder 50 considers the time sublayer of the multi-layer video bitstream for decoding. If it exceeds, the decoder 50 may omit the time sublayer during decoding and use the information regarding the maximum time sublayer to be decoded.

[0090] For example, the video bitstream 14 may indicate one or more OLSs 83, and each OLS 83 is associated with one of the bitstream compatibility set 86, buffer requirement set 84, and decoder requirement set 82 signaled in the video parameter set 91. The decoder 50 may select the OLS 83 for decoding based on an external indication (e.g., provided via an API or the video bitstream 14), or, for example, based on a selection rule in the absence of an external indication. When the maximum time sublayer indication 99 cannot be obtained, that is, when the decoder 50 does not receive the maximum time sublayer indication, the decoder 50 uses the constraint 92 on the maximum time sublayer of the decoder requirement set 82 associated with the OLS to be decoded.

[0091] For example, the constraints 82 of the decoder requirement set 82 associated with OLS are, for example, due to the constraints 94, 96 on the maximum temporal sublayer associated with the buffer requirement set 84 and the bitstream compatibility set 86 and the following bitstream constraints. Therefore, selecting the constraints 92 associated with the decoder capability set for decoding amounts to selecting the minimum value exceeding the maximum temporal sublayer shown for the decoder requirement set 82, buffer requirement set 84, and bitstream compatibility set 86 associated with OLS. By selecting the minimum value exceeding the constraints of the maximum temporal sublayer, it can be guaranteed that each parameter set 82, 84, 86 contains parameters valid for the selected bitstream for decoding. Therefore, by selecting the constraints 92 associated with the decoder requirement set 82, it is possible to surely select a bitstream for decoding for which all parameters of the parameter sets 82, 84, 86 are available.

[0092] For example, each of the parameter sets from the bitstream compatibility set 86, buffer requirement set 84, and decoder requirement set 82 associated with OLS may include one or more sets of parameters, and each of the parameter sets is associated with a temporal sublayer or a maximum temporal sublayer. The decoder 50 may select a parameter set associated with the maximum temporal sublayer to be decoded, as estimated or received, from each of the parameter sets. For example, if the parameter set is associated with the maximum temporal sublayer, or if the parameter set is associated with a temporal sublayer below the maximum temporal sublayer, the parameter set may be associated with the maximum temporal sublayer. The decoder 50 may use the selected parameter set to adjust one or more of the buffer size of the coded image, the buffer size of the decoded image, buffer scheduling, for example, HRD timing (AU / DU removal time, DPB output time).

[0093] As described above, the indication 99 of the maximum temporal sublayer to be decoded may be signaled in the video bitstream 14. According to the first embodiment, the encoder 10, e.g., the encoder 10 of FIG. 1, may be configured to optionally omit signaling of the indication 99 when the maximum temporal sublayer to be decoded corresponds to a constraint 92 associated with a decoder requirement set 82 of an OLS, e.g., an OLS instructed to decode. In other words, in this case, the encoder 10 may or may not signal the indication 99. In the example, the encoder 10 may determine whether to signal the indication 99 or omit signaling of the indication 99 in this case.

[0094] For example, the VPS 91 may be as follows. Currently, there are three syntax structures of the VPS that are generally defined and then mapped to a specific OLS. · Profile Tier Level (PTL), e.g., the decoder requirement set 82, · DPB parameters, e.g., the buffer requirement set 84, · HRD parameters, e.g., the bitstream compliance set 86, And another syntax element vps_max_sublayers_minus1 that is used when extracting the OLS sub-bitstream, i.e., in some cases, when deriving the variable NumSublayerInLayer[i][j].

[0095] FIG. 8 shows an example of the definitions of the PTL 82, DPB 84, and HRD 86 and their mapping to the OLS 83. The mapping from the PTL to the OLS is performed for the VPS of all OLSs (single-layer or multi-layer). However, the mapping of the DPB and HRD parameters to the OLS is performed only for the VPS of the OLS with multiple layers. As shown in FIG. 8, the parameters of the PTL, DPB, and HRD are first described in the VPS and then mapped to indicate which parameters the OLS uses.

[0096] In the example shown in FIG. 8, there are two OLSs, and there are two parameters for each of them. However, since the definitions and mappings are specified so that multiple OLSs can share the same parameters, for example, as shown in FIG. 9, it is not necessary to repeat the same information multiple times. FIG. 9 shows an example of sharing between different OLSs with different definitions of PTL, DPB, and HRD. Here, in the example of FIG. 9, OLS2 and OLS3 have the same PTL and DBP parameters but different HRD parameters.

[0097] In the examples of FIGS. 8 and 9, for a given OLS83, the values of vps_ptl_max_temporal_id[ptlIdx] (e.g., the constraint 92 for the maximum temporal sublayer of the decoder requirement set 82), vps_dpb_max_temporal_id[dpbIdx] (e.g., the constraint 94 for the maximum temporal sublayer of the buffer requirement set 84), and vps_hrd_max_tid[hrdIdx] (e.g., the constraint 96 for the maximum temporal sublayer of the bitstream conformity set 86) are aligned, but this is not currently necessary. These three values associated with the same OLS are not currently restricted to having the same value (however, they may be optionally restricted in some examples of the present disclosure). ptlIdx, dpbIdx, and hrdIdx are indices to each syntax structure signaled for each OLS.

[0098] According to an example of the first embodiment, the value of vps_ptl_max_temporal_id[ptlIdx] is used to set the variable HTid in the decoding process when it is not set by external means as follows. That is, when there is no external means for setting the value of HTid (e.g., via a decoder API), the value of vps_ptl_max_temporal_id[ptlIdx] is obtained by default to set HTid, that is, the minimum value of the three syntax elements described above. In other words, when the maximum temporal sublayer indication 99 is available, the decoder 50 may set the variable HTid accordingly, and when it is not available, HTid may be set to the value of vps_ptI_max_temporal_id[ptlldx].

[0099] According to a second embodiment of the third aspect, the decoder 50 estimates that the maximum temporal sublayer of the set of temporal sublayers to which each image 26 of layer 24 included in the OLS belongs is the minimum value of the maximum temporal sublayers indicated by the bitstream compliance set 86, buffer requirement set 84, and decoder requirement set 82 associated with the OLS. For example, the maximum temporal sublayer estimated for the set of temporal sublayers may correspond to or be indicated by the variable maximum TID WITHINOLS. In other words, the set of temporal sublayers to which each image of the layer of the OLS belongs may be the set of temporal sublayers that accommodate all the images belonging to the OLS. In this regard, all the images of the layers included in the OLS may belong to the OLS. For example, referring to FIG. 7, the OLS 83 including the first layer 241 and the second layer 240 may include the first temporal sublayer to which the access unit 221 belongs and the second temporal sublayer to which the access unit 222 belongs.

[0100] In an example, for one or more layers, or all layers of the OLS to be decoded, the decoder 50 may detect whether the video bitstream 14 indicates a constraint on the maximum temporal sublayer of the reference layer on which each layer depends. For example, the constraint may indicate that each layer depends only on the temporal sublayers of the reference layer up to the maximum temporal sublayer. If the video bitstream 14 does not indicate such a constraint, the decoder 50 estimates that the maximum temporal sublayer included in the OLS is equal to the maximum temporal sublayer estimated for the set of temporal sublayers to which each image of the layers of the OLS belongs.

[0101] In an example, the OLS indication 81 may further indicate one or more output layers for the OLS 83. In other words, one or more of the layers included in the OLS 83 may be indicated as output layers of the OLS 83. The decoder 50 can estimate that the maximum temporal sublayer included in the OLS is equal to the maximum temporal sublayer estimated for the set of temporal sublayers to which each image of the layers of the OLS belongs, for each layer indicated as an output layer of the OLS 83.

[0102] The decoder 50 may decode an image belonging to one of the layers included in the decoded OLS and an image belonging to a temporal sublayer below the decoded maximum temporal sublayer, from the image 26 of the video bitstream 14. The decoded maximum temporal sublayer may be the maximum temporal sublayer estimated for the set of temporal sublayers to which each image of the layers of the OLS belongs, or the maximum temporal sublayer decoded as described for the above embodiments.

[0103] In other words, according to the second embodiment, the above variable MaxTidWithiOls (which may be partially dropped and is not necessarily a bitstream) indicating the maximum number of temporal sublayers existing in the OLS is derived on the decoder side from the minimum value of the three values of the syntax elements, vps_ptl_max_temporal_id[ptlIdx], vps_dpb_max_temporal_id[dpbIdx], and vps_hrd_max_tid[hrdIdx]. This minimizes the sharing restrictions of any of the PTL, HRD, or DPB parameters and prohibits indicating an OLS where not all three parameters are defined. MaxTidWithinOls = min(vps_ptl_max_temporal_id[ptlIdx], min(vps_dpb_max_temporal_id[dpbIdx], vps_hrd_max_tid[hrdIdx]))

[0104] Furthermore, in that embodiment, NumSublayerlnLayer[i][j] representing the maximum sublayer included in the i-th OLS of layer j is set equal to the above-derived MaxTidWithinOls when vps_max_tid_il_ref_pics_plus1[m][k] does not exist or when layer j is the output layer of the i-th OLS.

[0105] For example, the decoder 50 can estimate that the maximum temporal sublayer of the multi-layer video bitstream to be decoded is equal to the maximum temporal sublayer estimated for the set of temporal sublayers to which each picture of the OLS layer belongs, that is, what is called MaxTidWithinOLS.

[0106] In other words, in another embodiment, if not set by external means as follows, the derived value of MaxTidWithinOls can be used to set the variable HTid of the decoding process. If there is no external means for setting the value of HTid (e.g., via the decoder API), the value of MaxTidWithinOls is obtained by default to be set to the value of HTid, i.e., the minimum value of the above three syntax elements.

[0107] In an example of an embodiment according to the third aspect, the decoder 50 selectively considers, for decoding, an image belonging to one of the temporal layers of the multi-layer video bitstream 14 if each image belongs to an access unit 22 associated with a temporal sublayer that does not exceed the maximum temporal sublayer to be decoded.

[0108] A further embodiment according to the third aspect includes an encoder 10, e.g., the encoder 10 of FIG. 1, for encoding a multi-layer video bitstream 14 according to FIG. 7. For this purpose, when the encoder 10 associates the OLS 83 with the bitstream compatibility set 86, the buffer requirement set 84, and the decoder requirement set 82, the smallest of the constraints 96, 94, 92 associated with each parameter set may form an OLS instruction 81 so as to accommodate the subset of layers indicated by the OLS 83. That is, for example, the minimum value among the maximum time sub-layers indicated by the bitstream compatibility set 86, the buffer requirement set 84, and the decoder requirement set 82 associated with the OLS 83 is greater than or equal to the maximum time sub-layer of the time sub-layers included in the subset of layers indicated by the OLS. The encoder 10 may further form the OLS instruction 81 so that the OLS 83 is valid as long as the parameters in the bitstream compatibility set 86, the buffer requirement set 84, and the decoder requirement set 82 refer to time sub-layers that are less than or equal to the minimum value among the maximum time sub-layers 92, 94, 96 indicated for the bitstream compatibility set 86, the buffer requirement set 84, and the decoder requirement set 82 associated with the OLS 83.

[0109] Although some aspects have been described as features in the context of an apparatus, it is clear that such descriptions may also be considered as descriptions of corresponding features of a method. Although some aspects have been described as features in the context of a method, it is clear that such descriptions may also be considered as descriptions of corresponding features related to the functions of an apparatus.

[0110] Some or all of the method steps may be performed by (or using) a hardware device such as a microprocessor, a programmable computer, or an electronic circuit. In some embodiments, one or more of the most important method steps may be performed by such a device.

[0111] The encoded image signal of the invention can be stored in a digital storage medium or transmitted over a transmission medium such as a wireless transmission medium like the Internet or a wired transmission medium. In other words, further embodiments provide a video bitstream product that includes a video bitstream according to any of the embodiments described herein, such as a digital storage medium storing the video bitstream.

[0112] Depending on specific implementation requirements, embodiments of the present invention can be implemented in hardware or software, or at least partially in hardware or at least partially in software. This implementation can be carried out using a digital storage medium such as a floppy disk, DVD, Blu-Ray, CD, ROM, PROM, EPROM, EEPROM or FLASH (registered trademark) memory storing electronically readable control signals, which cooperate (or can cooperate) with a programmable computer system such that each method is executed. Thus, the digital storage medium may be computer-readable.

[0113] Some embodiments according to the present invention include a data carrier having electronically readable control signals, and these control signals can cooperate with a programmable computer system such that one of the methods described herein is executed.

[0114] Generally, embodiments of the present invention can be implemented as a computer program product having program code, and this program code is operable to execute one of the methods when the computer program product is executed on a computer. The program code may be stored, for example, in a machine-readable carrier.

[0115] Other embodiments include a computer program that executes one of the methods described herein and is stored in a machine-readable carrier.

[0116] In other words, an embodiment of the method of the present invention is thus a computer program having program code for executing one of the methods described herein when the computer program is executed on a computer.

[0117] A further embodiment of the method of the present invention is thus a data carrier (or digital storage medium, or computer-readable medium), which data carrier contains a computer program for executing one of the methods described herein recorded thereon. The data carrier, digital storage medium, or recorded medium is usually tangible and / or non-transitory.

[0118] A further embodiment of the method of the present invention is thus a data stream or signal sequence representing a computer program for executing one of the methods described herein. The data stream or signal sequence may be configured to be transferred via a data communication connection, for example via the Internet.

[0119] A further embodiment includes, for example, a computer or processing means such as a programmable logic device configured or adapted to execute one of the methods described herein.

[0120] A further embodiment includes a computer on which a computer program for executing one of the methods described herein is installed.

[0121] A further embodiment according to the present invention includes an apparatus or system configured to transfer (e.g., electronically or optically) a computer program for executing one of the methods described herein to a receiver. The receiver may be, for example, a computer, a mobile device, a memory device, etc. The apparatus or system may include, for example, a file server for transferring the computer program to the receiver.

[0122] In some embodiments, a programmable logic device (e.g., a field programmable gate array) may be used to perform some or all of the functions of the methods described herein. In some embodiments, the field programmable gate array may cooperate with a microprocessor to perform one of the methods described herein. Generally, the methods are preferably performed by any hardware device.

[0123] The apparatuses described herein may be implemented using a hardware device, or using a computer, or using a combination of a hardware device and a computer.

[0124] The methods described herein may be performed using a hardware device, or using a computer, or using a combination of a hardware device and a computer.

[0125] In the embodiments for carrying out the above invention, it can be seen that various features are grouped together by way of example for the purpose of rationalizing the disclosure. This method of disclosure should not be construed as indicating that the claimed examples require more features than those explicitly recited in each claim. Rather, as the following claims show, the subject matter of the invention lies in less than all of the features of a single disclosed example. Thus, the following claims are incorporated into the embodiments for carrying out the invention by reference herein, and each claim may stand on its own as a separate example. Although each claim may stand on its own as a separate example, dependent claims may refer to a particular combination with one or more other claims within the scope of the claims, but it should be noted that other examples may also include combinations of the subject matter of dependent claims with each other, or combinations of each feature with other dependent or independent claims. Such combinations are proposed herein unless it is stated that a particular combination is not intended. Further, even if a claim is not directly dependent on an independent claim, it is also intended to include the features of this claim in any other independent claim.

[0126] The above embodiments are merely illustrative of the principles of the present disclosure. Naturally, modifications and variations of the configurations and details described herein will be apparent to those skilled in the art. Therefore, it is intended to be limited only by the scope of the pending claims, rather than by the specific details presented as descriptions and explanations of the embodiments herein.

Claims

1. A method for decoding a video bitstream, comprising: receiving, from the video bitstream, a video parameter set (VPS) including one or more profile layer level parameters, one or more decoded picture buffer (DPB) parameters, and one or more virtual reference decoder (HRD) parameters; receiving, from the video bitstream, an indication of the maximum temporal sublayer corresponding to the one or more profile layer level parameters, an indication of the maximum temporal sublayer corresponding to the one or more DPB parameters, and an indication of the maximum temporal sublayer corresponding to the one or more HRD parameters; setting, based on determining that an indication of the maximum temporal sublayer of the video bitstream is not available, the indication of the maximum temporal sublayer of the video bitstream equal to the indication of the maximum temporal sublayer corresponding to the one or more profile layer level parameters.

2. The method of claim 1, further comprising associating an output layer set (OLS) with the one or more profile layer level parameters, the one or more DPB parameters, and the one or more HRD parameters.

3. determining that a temporal sublayer of the video bitstream exceeds the maximum temporal sublayer of the video bitstream; omitting decoding of the temporal sublayer of the video bitstream based on the determination.

4. determining that a temporal sublayer of the video bitstream does not exceed the maximum temporal sublayer of the video bitstream; decoding the temporal sublayer of the video bitstream based on the determination.

5. A method for encoding a video bitstream, comprising: providing, to the video bitstream, a video parameter set (VPS) including one or more profile layer level parameters, one or more decoded picture buffer (DPB) parameters, and one or more virtual reference decoder (HRD) parameters; providing in the video bitstream an indication of a maximum temporal sublayer corresponding to the one or more profile tier level parameters, an indication of a maximum temporal sublayer corresponding to the one or more DPB parameters, and an indication of a maximum temporal sublayer corresponding to the one or more HRD parameters; based on determining whether the indication of the maximum temporal sublayer of the video bitstream is equal to the indication of the maximum temporal sublayer corresponding to the one or more profile tier level parameters, omitting or providing the indication of the maximum temporal sublayer of the video bitstream in the video bitstream, the method comprising.

6. The method according to claim 5, further comprising, when it is determined that the indication of the maximum temporal sublayer of the video bitstream is equal to the indication of the maximum temporal sublayer corresponding to the one or more profile tier level parameters, omitting the indication of the maximum temporal sublayer of the video bitstream in the video bitstream.

7. The method according to claim 5, further comprising, when it is determined that the indication of the maximum temporal sublayer of the video bitstream is not equal to the indication of the maximum temporal sublayer corresponding to the one or more profile tier level parameters, providing the indication of the maximum temporal sublayer of the video bitstream in the video bitstream.

8. The method according to claim 5, further comprising associating an output layer set (OLS) with the one or more profile tier level parameters, the one or more DPB parameters, and the one or more HRD parameters.

9. A video coding apparatus comprising one or more processors configured to execute the method according to any one of claims 1 to 8.

10. A non-transitory computer-readable medium storing a computer program for implementing the method according to any one of claims 1 to 8 when executed by a computer or a signal processor.

Citation Information

Patent Citations

  • Temporal sub-layer descriptor

    US20190158880A1

  • Avoidance of redundant signaling in multi-layer video bitstreams

    WO2021022267A2

  • Signaling of DPB parameters for multi-layer video bitstreams

    WO2021061489A1

  • Coding output layer set data and conformance window data of high level syntax for video coding

    WO2021174098A1

  • Systems and methods for decoding based on inferred video parameter sets

    WO2021236312A1