Video Coding Aspects of Temporal Motion Vector Prediction, Inter-Layer Referencing, and Temporal Sub-Layer Indication
By adopting the method of selecting suitable reference images for TMVP and limiting tree root block segmentation strategies in video encoding, the problem of inefficiency of TMVP and tree root block segmentation strategies in the prior art is solved, and more efficient encoding and decoding performance is achieved.
Patent Information
- Application Number
- JP2022571832
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-06-09
- Filing Date
- 2021-06-08
- Publication Date
- 2025-05-12
- Estimated Expiration
- 2041-06-08
AI Technical Summary
Existing video encoding technologies have inefficient problems in time motion vector prediction (TMVP) and tree root block segmentation strategies, especially in different levels of video bitstreams, resulting in improper cache management and reduced encoding efficiency.
A method of selecting a reference image for time motion vector prediction (TMVP) is proposed, by creating two reference image lists for a given image and selecting the appropriate reference image in the list for TMVP. At the same time, the segmentation strategy of tree root blocks is limited so that its size does not exceed the size of the tree root block of the reference image to control dependencies and optimize cache management.
Through these technical means, encoding efficiency is improved, unnecessary cache load is reduced, and efficient encoding and decoding performance in different levels of video bitstreams is ensured.
Smart Images

Figure 0007675107000003 
Figure 0007675107000004 
Figure 0007675107000005
Abstract
Description
[Technical field]
[0001] FIELD OF THE DISCLOSURE Embodiments of the present disclosure relate to a video encoder, a video decoder, a method for encoding a video sequence into a video bitstream, and a method for decoding a video sequence from a video bitstream. Further embodiments relate to video bitstreams. [Background technology]
[0002] In encoding or decoding an image of a video sequence, prediction is used to reduce the amount of information signaled in the video bitstream in which the image is encoded / decoded. Prediction may be used on the image data itself, such as the sample values or the coefficients on which the sample values of the image are coded. Alternatively or additionally, prediction may be used on syntax elements used to code the image, such as motion vectors. To predict the motion vector of the image to be coded, a reference image may be selected, from which a predictor of the motion vector of the image to be coded is determined. Summary of the Invention [Means for solving the problem]
[0003] The first aspect of the present disclosure provides a concept for selecting a reference image used for temporal motion vector prediction. Two lists of reference images are populated for a given image, e.g., an image to be coded. Each list may be empty or non-empty. A TMVP reference image is determined by selecting one from the two lists of reference images as a TMVP image list and selecting a TMVP reference image from the TMVP image list. According to the first aspect, if one of the two lists is empty and the other is non-empty, a reference image from the non-empty list is used for temporal motion vector prediction (TMVP). Thus, TMVP can be used regardless of which of the two lists is empty, providing high coding efficiency in both cases where only the first list or only the second list is empty.
[0004] A second aspect of the present disclosure is based on the idea that the tree root blocks into which an image of a coded video sequence is divided are smaller or equal in size to the tree root blocks into which the reference images of that image are divided. Imposing such a constraint on the division of an image into tree root blocks may ensure that the dependency of an image on a reference image does not cross the boundaries of the tree root blocks, or at least does not cross the row boundaries of the rows of tree root blocks. Thus, the constraint may limit the dependency between different tree root blocks, benefiting buffer management. In particular, the dependency between tree root blocks of different rows of tree root blocks may lead to inefficient buffer usage, since adjacent tree root blocks belonging to different rows may be separated by further tree root blocks in coding order. Avoiding such dependencies may therefore eliminate the need to keep an entire row of tree root blocks between the tree root block currently being coded and the referenced tree root block.
[0005] The third aspect of the present disclosure provides a concept for determining a maximum temporal sub-layer up to which layers of an output layer set indicated in a multi-layer video bitstream should be decoded. This concept thus allows a decoder to determine which part of a video bitstream to decode even without an indication of the maximum temporal sub-layer to be decoded. Furthermore, this concept allows an encoder to omit signaling an indication of the maximum temporal sub-layer to be decoded if the maximum temporal sub-layer to be decoded corresponds to the one estimated by the decoder in the absence of the respective indication, thus avoiding unnecessarily high signaling overhead.
[0006] Embodiments and advantageous implementations of the present disclosure are described in more detail below with reference to the drawings. [Brief description of the drawings]
[0007] [Figure 1] 1 illustrates an encoder, a decoder, and a video bitstream according to an embodiment; [Diagram 2] 1 shows an example of image segmentation. [Diagram 3] 1 illustrates the determination of a TMVP reference image according to an embodiment. [Figure 4] 13 shows an example of motion vector candidate determination depending on tree root splits. [Diagram 5] Two examples of different tree root block sizes for dependent and reference layers are given. [Figure 6] 1 shows an example of dependent and reference layer sub-image segmentation. [Figure 7] 4 shows a decoder and a video bitstream according to an embodiment of the third aspect; [Figure 8] 13 shows an example of a mapping between output layer sets and video parameter sets. [Figure 9] 13 shows an example of a mapping between output layer sets and video parameter sets that share parameters. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0008] In the following, the embodiments are described in detail, but it should be understood that the embodiments provide many applicable concepts that can be embodied in various video coding concepts. The specific embodiments described are merely illustrative of specific ways to implement and use the concepts, and do not limit the scope of the embodiments. In the following description, a number of details are described to provide a more complete description of the embodiments of the present disclosure. However, it will be apparent to one skilled in the art that other embodiments may be practiced without these specific details. In other instances, well-known structures and devices are shown in block diagram form, rather than in detail, to avoid obscuring the examples described herein. Furthermore, features of different embodiments described herein may be combined with each other, unless specifically stated otherwise.
[0009] In the following description of the embodiments, the same or similar elements or elements having the same functions are given the same reference numbers or identified by the same names, and the repeated description of the elements given the same reference numbers or identified by the same names is usually omitted. Therefore, the descriptions provided for the elements having the same reference numbers or identified by the same names can be mutually exchanged or applied to each other in different embodiments.
[0010] A detailed description of embodiments of the disclosed concepts begins with a description of examples of an encoder, a decoder, and a video bitstream, which provide a framework in which embodiments of the present invention can be incorporated. Below, a description of embodiments of the present concepts is provided along with a description of how such concepts can be incorporated into the encoder and decoder of FIG. 1. However, the embodiments described with respect to the subsequent FIG. 2 and subsequent figures may be used to form encoders and decoders that do not operate according to the framework described with respect to FIG. 1. It should also be noted that the encoder and decoder, although shown together for illustrative purposes in FIG. 1, may be implemented separately from each other. It should also be noted that the encoder and decoder may be combined in one device, or one of the two may be implemented as part of the other. Some embodiments of the present invention are also described with reference to FIG. 1.
[0011] FIG. 1 shows an example of an encoder 10 and a decoder 50. The encoder 10 (which may also be called an encoding device) encodes a video sequence 12 into a video bitstream 14 (which may also be called a bitstream, a data stream, a video data stream, or a stream). The video sequence 12 includes a sequence of images 21, which are arranged in a presentation order or image order 17. In other words, each of the images 21 may represent a frame of the video sequence 12 and may be associated with an instant of the presentation order of the video sequence 12. Based on the video sequence 12, the encoder 10 may encode a coded video sequence 20 into a video bitstream 14. The encoder 10 may form the coded video sequence 20 in the form of access units 22, each access unit 22 having encoded therein video data belonging to a common instant of time. In other words, each access unit 22 may encode therein one of the images 21 of the video sequence 12, i.e. one of the frames. Encoder 10 encodes a coded video sequence 20 according to a coding order 19 , which may differ from the picture order 17 of the video sequence 12 .
[0012] The encoder 10 may encode the coded video sequence 20 into one or more layers. That is, the video bitstream 14 may be a single-layer or multi-layer video bitstream containing one or more layers. Each access unit 22 contains one or more coded pictures 26 (e.g., pictures 260, 261 in FIG. 1, where the apostrophe and star are used to refer to a particular one and the subscript indicates the layer to which the picture belongs). Each picture 26 belongs to one of the layers 24 of the coded video sequence, e.g., layers 240, 241 in FIG. 1. In FIG. 1, an exemplary number of two layers is shown, namely, a first layer 241 and a second layer 240. In an embodiment according to the disclosed concept, the coded video sequence 20 and the video bitstream 14 may contain one, two or more layers, but not necessarily multiple layers. 1, each access unit 22 includes a coded image 261 of a first layer 241 and a coded image 260 of a second layer 240. However, it should be noted that each access unit 22 may, but does not necessarily, include a coded image for each layer of the coded video sequence 20. For example, layers 240, 241 may have different frame rates (or image rates) and / or may include images for complementary subsets of the access units of access unit 22.
[0013] As mentioned above, the images 260, 261 of one of the access units represent image content at the same time. For example, the images 260, 261 of the same access unit 22 may represent the same image content at different qualities, e.g., resolution or fidelity. In other words, the layer 240 may represent a first version of the coded video sequence 20, and the layer 241 may represent a second version of the coded sequence 20. Thus, a decoder, such as the decoder 50, or an extractor, may select between different versions of the coded video sequence 20 that are decoded or extracted from the video bitstream 14. For example, the layer 240 may be decoded independently of further layers of the coded video sequence to provide a decoded video sequence of a first quality, and a joint decoding of the first layer 241 and the second layer 240 may provide a decoded video sequence of a second quality that is higher than the first quality. For example, the first layer 241 may be encoded independently of the second layer 240. In other words, the second layer 240 may be a reference layer for the first layer 241. For example, in this scenario, the first layer 241 may be referred to as an enhancement layer and the second layer 240 may be referred to as a base layer. The image 260 may have a smaller image size than the image 261, an equal image size, or a larger image size. For example, the image size may refer to the number of samples in a two-dimensional array of the image. Note that the images 260, 261 do not necessarily represent equal image content, for example, the image 261 may represent an excerpt of the image content of the image 260. For example, in some scenarios, different layers of the video bitstream 14 may include different sub-images of an image coded into the video bitstream.
[0014] The encoder 10 encodes the access units 22 into bitstream portions 16 of the video bitstream 14. For example, each access unit 22 may be encoded into one or more bitstream portions 16. For example, a picture 26 may be subdivided into tiles of slices, and each slice may be encoded into one bitstream portion 16. The bitstream portions 16 into which the pictures 26 are encoded may be referred to as video coding layer (VCL) NAL units. The video bitstream 14 may further include non-VCL NAL units, e.g., bitstream portions 23, 29 into which description data is coded. The description data may provide information for decoding or about the coded video sequence 20. The bitstream portions into which the description data is encoded may be associated with individual bitstream portions. For example, they may refer to individual slices, or may be associated with one of the pictures 26, or one of the access units 22, or may be associated with a sequence of access units, i.e., related to the coded video sequence 20. It should be noted that the video 12 may be coded into a sequence of coded video sequences 20 .
[0015] The decoder 50 (which may also be referred to as a decoding device) decodes the video bitstream 14 to obtain a decoded video sequence 20'. It should be noted that the video bitstream 14 provided to the decoder 50 does not necessarily correspond to the video bitstream 14 provided by the encoder, but may be extracted from the video bitstream provided by the encoder, so that the video bitstream decoded by the decoder 50 may be a sub-bitstream of the video bitstream encoded by an encoder, such as the encoder 10. As mentioned above, the decoder 50 may decode the entire coding video sequence 20 coded in the video data stream 14, or may decode a portion thereof, for example a subset of the layers of the coding video sequence 20 and / or a temporal subset of the coding video sequence 20 (i.e. a video sequence having a frame rate lower than the maximum frame rate provided by the video sequence 20). Thus, the decoded video sequence 20' does not necessarily correspond to the video sequence 12 encoded by the encoder 10. It should also be noted that the decoded video sequence 20' may further differ from the video sequence 12 due to coding losses, such as quantization losses.
[0016] Image 26 may be encoded using a predictive tool to predict signals or coefficients representing the image in video bitstream 14 from previously coded images. * , for example, may use a predictive tool to encode the image to be currently encoded using previously encoded images. In response, the decoder 50 may derive the image to be currently decoded 26 from the previously decoded images. * In the following description, a given image or block, e.g., the image or block currently being coded, is referred to as ( *) may be used. For example, image 261 in FIG. * is considered as the image currently being coded, where the image currently being coded 26 * may equally refer to the currently encoding picture being encoded by the encoder 10 and the picture currently being decoded in the decoding process performed by the decoder 50.
[0017] The prediction of an image from another image in coded video sequence 20 is sometimes referred to as inter-prediction. * is image 261 * 2. Thus, image 261 may be encoded using temporal inter-prediction from image 261′ belonging to one of the access units different from image 261. * is image 261 * It belongs to the same layer as image 261 * Additionally or alternatively, image 261 may include a reference 32 to an image 261' that belongs to a different access unit than image 261. * may be predicted using inter-layer (inter) prediction from an image of another layer, e.g., a lower layer (lower by a layer index that can be associated with each of the layers 24). * may contain a reference 34 to a picture 260' that belongs to the same access unit but to another layer. In other words, in FIG. 1, pictures 261', 260' are the same as the picture 261 currently being coded. * It should be noted that prediction may be used to predict coefficients of the picture itself, such as determining transform coefficients signaled in the video bitstream 14, or may be used to predict syntax elements used in encoding the picture. For example, a picture may be a reference picture for the currently being coded picture 26 with respect to previously coded pictures or previous pictures in picture order. * The image content of image 261 may be encoded using motion vectors that may represent the motion of the image content of image 261. For example, the motion vectors may be signaled in video bitstream 14. *The motion vectors may be predicted from a reference image, for example any of the above and alternative images, using temporal motion vector prediction (TMVP).
[0018] The image 26 may be coded on a block-by-block basis, in other words, the image 26 may be subdivided into blocks and / or sub-blocks, as described, for example, with respect to FIG.
[0019] The embodiments described herein may be implemented in the context of Versatile Video Coding (VVC) or other video codecs.
[0020] In the following, some concepts and embodiments are described with reference to Fig. 1 and the features described with respect to Fig. 1. It is pointed out that features described with respect to an encoder, a video bitstream, or a decoder shall be understood as descriptions with respect to other of these entities. For example, a feature described as being present in a video data stream shall be understood as a description of an encoder configured to encode this feature into a video bitstream, and a decoder or extractor configured to read this feature from the video bitstream. It is further pointed out that the estimation of information based on instructions coded in the video bitstream may be performed equally on the encoder side and on the decoder side. It is further noted that features described with respect to individual aspects may be optionally combined with each other.
[0021] FIG. 2 shows an example of dividing one of the images 26 into blocks 74 and sub-blocks 76. For example, the image 26 may be previously divided into tree root blocks 72, which may then be subjected to recursive subdivision, as exemplarily shown in one of the tree root blocks 72 in FIG. 2, a tree root block 72′. That is, the tree root blocks 72 may be subdivided into blocks, which may then be subdivided into sub-blocks, and so on. Recursive subdivision may also be referred to as multi-tree division. The tree root blocks 72 may be rectangular, and optionally quadratic. The tree root blocks 72 may be referred to as coding tree units (CTUs).
[0022] For example, the above-mentioned motion vectors (MVs) may be determined and optionally signaled in the video bitstream 14 on a block-by-block or sub-block-by-subblock basis. In other words, the motion vectors may refer to the entire block 74 or to the sub-blocks 76. For example, a motion vector may be determined for each block 74 of the image 26. Alternatively, a motion vector may be determined for each of the sub-blocks 76 of the block 74. In an example, whether one motion vector is determined for the entire block 74 or one motion vector is determined for each sub-block 76 of the block 74 may vary from block to block. For example, all images of the coded video sequence 20 belonging to the same layer of the layer 24 may be divided into tree root blocks of equal size.
[0023] The embodiments according to the first and second aspects may relate to temporal motion vector prediction.
[0024] 3 shows a TMVP reference picture determination module 53, hereafter named TMVP module 53, according to an embodiment of the first aspect, which may also be optionally implemented in the embodiments of the second and third aspects. The TMVP reference picture determination module 53 may be implemented in a video decoder supporting TMVP, e.g. a video decoder configured to decode a sequence of coded pictures from a data stream, e.g. the decoder 50 of FIG. 1. The TMVP module 53 may also be implemented in a video encoder supporter TMVP, e.g. a video encoder configured to decode a sequence of pictures into a data stream, e.g. the encoder 10 of FIG. 1. The module 53 determines the reference picture of a given picture, e.g. the currently coded picture 26. * TMVP Reference Image 59 * For example, a module for determining a given image 26 * The TMVP reference image for is the given image 26 * From this reference image, a given image 26 * A predictor of the motion vector of is selected.
[0025] TMVP Reference Image 59 * To determine, the TMVP module 53 determines a first list 561 and a second list 562 of reference pictures from a plurality of previously decoded pictures. For example, the plurality of previously decoded pictures may include image 26 of previously decoded access unit 22. * , for example, a given image 261 * 1 for the image 261′ of FIG. 1 for the image 260′ of FIG. 1 for the image 261′ of FIG. * This image may include a previously decoded image of the same access unit as the given image 26. *2 belongs to a lower layer than , i.e. layer 240. Thus, referring to the example of FIG. 1, for example, images 260', 261' may be part of a first list of reference images 561. In other examples, these two images may be part of a second list of reference images. The first list 561 may optionally include further reference images. In the example shown in FIG. 3, the second list of reference images 562 is empty. Note that in general, either or both of the first and second lists may be empty or non-empty. The reference images from the first and second lists may be for inter-prediction of a given image 26'. The encoder 10 may determine the first and second lists and signal them in the video bitstream 14, so that the decoder 50 may derive them from the video bitstream. Alternatively, the decoder 50 may determine the first and second lists independently of explicit signaling, or at least partially independently of explicit signaling. For example, encoder 10 may signal the first and second lists in video bitstream 14.
[0026] The TMVP reference image determination module 53 determines the predetermined image 26 * 5. Select one reference image from the first list 561 and the second list 562 of reference images, for example, select a given image 26 if at least one of the first and second lists of reference images is not empty. * TMVP Reference Image 59 * For this purpose, module 53, for example by means of a TMVP list selection module 57, selects one of the first list 561 and the second list 562 of reference images as a TMVP image list 56. * It may be determined as follows.
[0027] TMVP Image List 56 * To determine (57), the encoder 10 determines whether the second list of reference images 562 corresponds to a given image 26 * If the first list 561 is empty, the second list 561 is added to the TMVP image list 56 *Thus, the decoder 50 may select the second list of reference pictures 562 as the first list of reference pictures for a given picture 26. * If it is empty for TMVP image list 56 * It can be assumed that the first list of reference pictures 561 is the first list of reference pictures. If the first list of reference pictures is empty and the second list of reference pictures is not empty, the encoder 10 may infer that the TMVP picture list 56 * Therefore, the decoder 50 may select the second list 562 as the TMVP image list 56 in this case. * It can be assumed that the second list of reference images 562 is the given image 26 * If neither the first list nor the second list of reference pictures is empty, the TMVP list selection module 57 of the encoder 10 selects a TMVP picture list 56 from the first list 561 and the second list 562. * Encoder 10 may encode a list selector 58 into video bitstream 14, where list selector 58 determines whether the first or second list corresponds to a given image 26. * TMVP Image List 56 * For example, list selector 58 may correspond to the following syntax element: ph_collocated_from_l0_flag. Decoder 50 reads list selector 58 from video bitstream 14 and selects TMVP picture list 56 accordingly. * You may select:
[0028] The TMVP module 53 further includes a TMVP image list 56 * TMVP reference image 59 from * For example, the encoder 10 may select 59 a TMVP image list 56. * TMVP reference image 59 * TMVP reference picture 59 in the video bitstream 14 by signaling the index of the selected TMVP reference picture 59 *In other words, the encoder 10 may signal a picture selector 61 in the video bitstream 14. For example, the picture selector 61 may correspond to the ph_collocated_ref_idx syntax element below. The decoder 50 reads the picture selector 61 from the video bitstream 14 and selects the picture list 56 accordingly. * TMVP Reference Image 59 * You may select:
[0029] The encoder 10 and the decoder 50 receive a given image 26 * TMVP reference image 59 is used to predict the motion vectors of * may be used.
[0030] For example, TMVP list selection module 57 of decoder 50 may initiate TMVP list selection by detecting whether video bitstream 14 indicates list indicator 58, and if so, selecting a TMVP picture list 56 as indicated by list indicator 58. * If the video bitstream 14 does not indicate a list selector 58, the TMVP list selection module 57 may select a second list of reference pictures 562 for the given picture 26. * If it is empty, the first list 561 is added to the TMVP image list 56 * Otherwise, if the first list of reference images is empty for a given image, the TMVP list selection module 57 may select the second list of reference images 562 as the TMVP image list 56. * Alternatively, if second list 562 is empty for a given image, TMVP list selection module 57 may select second list 562 as TMVP image list 56 if second list 562 is not empty. * may be selected as.
[0031] In other words, an embodiment of the first aspect may consider the interference of the list selector 58, e.g., PH_collocated_from_L0, when the first list 561 is empty but the second list 562 is not empty. The first list 561 may be referred to as L0 and the second list 562 may be referred to as L1.
[0032] In other words, the current specification (for VVC) uses two syntax elements to control the pictures used for temporal motion vector prediction (TMVP) (or sub-block TMVP): ph_collocated_from_l0_flag and ph_collocated_ref_idx. The first syntax element specifies whether the picture used for TMVP is selected from L0 or from L1, and the second syntax element specifies which of the pictures from the selected list is used. These syntax elements are present in either the picture header or the slice header. In the latter case, the prefix has "sh_" instead of "ph_". The picture header is shown as an example in Table 1.
[0033] [Table 1]
[0034] ph_collocated_from_l0_flag equal to 1 specifies that the collocated picture used for temporal motion vector prediction is derived from reference picture list 0. ph_collocated_from_l0_flag equal to 0 specifies that the collocated picture used for temporal motion vector prediction is derived from reference picture list 1. If ph_temporal_mvp_enabled_flag and pps_rpl_info_in_ph_flag are both equal to 1 and num_ref_entries[1][RplsIdx[1]] is equal to 0, the value of ph_collocated_from_l0_flag is inferred to be equal to 1. ph_collocated_ref_idx specifies the reference index of the collocated picture used for temporal motion vector prediction. If ph_collocated_from_l0_flag is equal to 1, then ph_collocated_ref_idx refers to an entry in reference image list 0, and the value of ph_collocated_ref_idx ranges from 0 to num_ref_entries[0][RplsIdx[0]]-1. If ph_collocated_from_l0_flag is equal to 0, ph_collocated_ref_idx refers to an entry in reference image list 1, and the value of ph_collocated_ref_idx ranges from 0 to num_ref_entries[1][RplsIdx[1]]-1. If not present, the value of ph_collocated_ref_idx is inferred to be equal to 0.
[0035] There is a specific case that the decoder needs to take into account, which is when the reference picture list is empty, i.e. it has zero entries in both L0 and L1. If either the L0 or L1 list is empty, the specification currently does not signal ph_collocated_from_l0_flag, infers that the value of ph_collocated_from_l0_flag is equal to 1, and considers the collocated picture to be in L0. However, this comes with efficiency issues. In fact, the following scenarios are possible with respect to the L0 and L1 states: L0 is empty and L1 is empty: ph_collocated_from_l0_flag is assumed to be 1 => OK · L0 is not empty and L1 is empty: ph_collocated_from_l0_flag is inferred as 1 => OK. · Neither L0 nor L1 is empty: ph_collocated_from_l0_flag is signaled => OK. · L0 is empty and L1 is not empty: ph_collocated_from_l0_flag is inferred as 1 => NOT OK.
[0036] If L0 is empty but L1 is not, the value of ph_collocated_from_l0_flag is inferred to be equal to 1, which means that TMVP (or sub-block TMVP) can still be used if an image in L1 is selected, but TMVP is not used, leading to a loss of efficiency.
[0037] Thus, in one embodiment, the decoder (or encoder) determines the value of ph_collocated_from_l0_flag depending on which list is empty, for example, as described above with respect to TMVP list selection module 57 of FIG. ph_collocated_from_l0_flag equal to 1 specifies that the collocated picture used for temporal motion vector prediction is derived from reference picture list 0. ph_collocated_from_l0_flag equal to 0 specifies that the collocated picture used for temporal motion vector prediction is derived from reference picture list 1. When ph_temporal_mvp_enabled_flag and pps_rpl_info_in_ph_flag are both equal to 1: If num_ref_entries[1][RplsIdx[1]] is equal to 0, the value of ph_collocated_from_l0_flag is inferred to be equal to 1. Otherwise, the value of ph_collocated_from_l0_flag is inferred to be equal to 0.
[0038] As an alternative to determining the TMVP picture list depending on which of the lists is empty, in another embodiment, there is a bitstream constraint that if L0 is empty, then L1 must be empty, either due to a bitstream constraint or syntax prohibition (see Table 2).
[0039] Thus, according to an alternative embodiment of the TMVP module 53 of FIG. 3, the encoder 10 selects a TMVP reference image 59 from the first list 561 if the second list 562 is empty. * If the second list 562 is not empty, select the TMVP reference image 59 from the first list 561 or the second list 562. * In other words, the TMVP list selection module 57 selects the TMVP image list 56 if the second list 562 is empty. * The first list 56 * 1, and if the second list 562 is not empty, then either the first list or the second list may be selected.
[0040] Thus, in determining (55) the list of reference pictures, decoder 50 can infer that if first list 561 is empty, then second list 562 is empty. Thus, in TMVP list selection 57, decoder 50 reads list selector 58 from video bitstream 14, and if neither first list 561 nor second list 562 is empty, selects TVMP reference picture 59 according to list selector 58. * If the first list 561 and the second list 562 are not non-empty, the decoder 50 may select the first list 561 as the TMVP image list 56. * If not, the decoder 50 may select the first list 561 as the TMVP image list 56. * may be selected as.
[0041] According to an embodiment, the decoder 50 may perform the list determination (55) by reading from the video bitstream 14, for a first list of reference pictures, information on how to populate the first list 561 from a number of previous decoder pictures. If the decoder 50 does not estimate that the second list 562 is empty, the decoder 50 may read from the video bitstream 14 information on how to populate the second list 562.
[0042] According to the latter embodiment, if the first list 561 is empty, following the decoder's estimation that the second list 562 is empty, the encoder 10 and the decoder 50 may select a given image 26 without using TMVP if the first list of reference images is empty. * may be encoded.
[0043] As mentioned above, the bitstream constraints may alternatively be implemented in syntax, for example in the construction of a list of reference pictures. An example implementation is shown in Table 2.
[0044] [Table 2]
[0045] Thus, if ph_collocated_from_l0_flag is estimated to be 1, then either both lists are empty or only list L1 is empty, and if there is one non-empty list, that list is used for TMVP.
[0046] Thus, according to a further alternative embodiment of the TMVP module 53 of FIG. 3, the encoder 10 selects a TMVP reference image 56 from the first list 561 if the second list 572 is empty. * If the second list 562 is not empty, select the TMVP reference image 56 from the first list 561 or the second list 562. * A TMVP list selection 57 may be performed by selecting
[0047] In the following, an embodiment according to the second aspect will be described with reference to FIGS.
[0048] FIG. 4 shows the first image 261 *4 shows an example of subdivision of the image currently being coded, for example, into tree root blocks 721. The tree root blocks 721 are recursively subdivided, for example, into blocks 74 and sub-blocks 76, as described with respect to FIG. 2. FIG. 4 further shows the image 260' of FIG. 1, which may be an interlayer reference image for a second image 260', for example, for the first image 261'. The second image 260' is subdivided into tree root blocks 720. A first example of subdivision into tree root blocks 720 is indicated by dashed lines. The dotted lines further show an alternative second example for subdivision of the second image 260' into tree root blocks 720', the tree root blocks 720' of the second example being smaller than the tree root blocks 720 of the first example. According to the first example of subdivision, the tree root blocks 720 are sub-divided into the tree root blocks 720 of the first image 261. * In the second example of subdivision, the tree root block 720′ has the same size as the tree root block 721 in the first image 261. * is smaller than the tree root block 721.
[0049] First image 261 * For the TMVP of image 261, the encoder 10 and the decoder 50 may determine one or more MV candidates. * For example, the encoder 10 and the decoder 50 may determine one or more MV candidates from one or more lists of reference pictures, e.g., one or more of the reference pictures from two lists, e.g., list 561 and list 562 described with respect to FIG. 3. * 1, there may be inter-layer reference pictures, such as a second picture 260′, that are temporally collocated with the first picture, i.e., belong to the same access unit 22. As explained with respect to FIG. 2, the pictures 26 may be coded block-wise or sub-block-wise, and the encoder 10 may use the currently coded picture 261 as a reference picture for the currently coded picture 262. * Currently coding tree root block 721* Block 74 currently being coded * Alternatively, the MV of the currently coded block 74 may be determined. * One MV may be determined for each sub-block 76 of the current coding. The sub-block currently being coded is denoted by reference numeral 76. * The encoder 10, and optionally the decoder 50, may refer to the currently coded block 74 using the * or subblock 76 currently being coded * For the currently coded block 74, one or more MV candidates may be determined from one reference image, e.g., the second image 260′. The one or more MV candidates may be determined from different positions or locations within the reference image. * or subblock 76 currently being coded * For the bottom right MV candidate of the currently coded block 74, the MV of the reference image located at the reference position 71′ in the reference image may be selected as the MV candidate, and the reference position 71′ of the bottom right MV candidate is collocated with the bottom right position 71 in the currently coded image 261′. * or subblock 76 currently being coded * The bottom right position 71 is block 74 * or subblock 76 * It may be adjacent to below and to the right of.
[0050] In FIG. 4, for a first example of subdividing the second image 260′ into tree root blocks 720, the currently coded block 74 of the first image 261′ is * The collocated block is designated by reference numeral 714. * As shown in FIG. 4, the size of the tree root block 720 is * In this first example of a subdivision equal to the size of the tree root block 721, the reference position 71' is the block 714 * , which is in the bottom right block 714' of the first image 261. *Block 74 currently being coded * Colocated block of 714 * Same as tree root block 720 * For example, the first image 261 * Block 74 of * One MV candidate for coding of may be the MV of block 714'.
[0051] First image 261 * In a second example of subdividing the reference image 260′ into tree root blocks 720′ smaller than the tree root block 721 of the current coding block 74, the reference position 71′ may be located outside the collocated tree root block 720′. * Therefore, the reference position 71′ may be located outside the row of the tree root block in which the colocated tree root block 720′ is located. * or subblock 76 currently being coded * Using MV as an MV candidate for may require the encoder 10 and decoder 50 to keep in the image buffer one or more tree root blocks beyond the currently coded tree root block of the second image 260' or beyond the current row of the currently coded tree root block. Thus, in this second example of tree root block subdivision, using reference position 71 for MV prediction may involve inefficient buffer usage.
[0052] In other words, as described with respect to Figure 3, the current specification uses two syntax elements to control the picture used for TMVP (or sub-block TMVP), namely ph_collocated_from_l0_flag and ph_collocated_ref_idx, where the latter specifies which entry in each list is used as a reference picture, e.g., as described with respect to Figure 3. If TMVP (e.g., the TMVP of one of the MVs of block 74 in Figure 2) or sub-block TMVP (e.g., the TMVP of one of the MVs of sub-block 76 in Figure 2) is used in the bitstream, the current picture 26 * The motion vector (MV) of the block or sub-block of the current image is derived from the MV of the reference image indicated by ph_collocated_ref_idx, and the MV is added to a list of candidates that can be selected as predictors of the block / sub-block's actually used motion vector. This MV prediction is done using the 26 * The MVs are selected from among the MVs at different positions according to the boundaries of the largest block (CTU) of the frame (e.g., the tree root block 72 in FIG. 2). In most scenarios, the reference images show the same CTU boundaries.
[0053] More specifically, in the case of TMVP (e.g., block 74 currently being coded), * 4. If the bottom-right TMVP MV candidate 71 of the current block (or sub-block 76) crosses the CTU row boundary of the current block (i.e., for example, the tree root block 74 of FIG. 4), **For , if the block 74 or sub-block 76 is located beyond the row boundary of the tree root block 72 to which it belongs) or does not exist (e.g., because it crosses an image boundary or is intra-coded without using a motion vector), then the TMVP candidate is not obtained from the bottom right block, but instead an alternative (co-located) MV candidate of the CTU row is derived, i.e., from the co-located block (of the block 74 or sub-block 76 for which the MV candidate is to be determined). Furthermore, if the sub-block TMVP MV candidate is obtained from a position outside the co-located CTU, the MV used to determine the position of the sub-block TMVP MV is modified (clipped) during derivation to point to a position where the sub-block TMVP MV candidate is obtained from a position belonging to the co-located CTU. Note that in the case of sub-block TMVP, the MV is obtained first (e.g., from the spatial candidate first), and that MV is used to identify the position of the block in the co-located image that is used to select the sub-block TMVP candidate. If the temporal MV used to identify the position points to a position outside the collocated CTU, the MV is clipped so that the sub-block TMVP is obtained from a position within the collocated CTU.
[0054] However, there are cases where the reference picture does not have the same CTU size, e.g., tree root block 721′ in FIG. 4, and thus does not have a boundary. The described TMVP or sub-block TMVP process is not clear for such cases where the CTU size is changed. When so applied, the state of the art would select non-buffer-friendly TMVP and sub-block TMVP candidates that are outside the current CTU (row) or CTU (row) of the reference picture, e.g., at position 71′. This aspect of the invention, which aims to guide the MV candidate selection accordingly or prevent such cases from occurring, relieves the implementation from this burden. The described scenario occurs in quality scalable multi-layer bitstreams with two (or more) layers. Here, one layer depends on the other layer, and the dependent layer has a larger CTU size than the reference layer, as shown on the left side of FIG. 5. In such a case, the decoder would have to fetch the associated MV candidates from four small CTUs of the reference picture, which do not necessarily occupy a contiguous memory region that would negatively impact memory bandwidth when accessed. On the other hand, on the right side of Figure 5, the CTU size of the reference image is larger than that of the dependent layer. In such a case, when decoding the current block of the dependent layer, data is not interspersed among the related data of the reference layer.
[0055] As explained, if the CTU sizes of the two layers (parameters for each SPS) are different, the CTU boundaries of the current picture and the reference picture (ILRP in this scenario) are not aligned, but this does not prohibit TMVP and Sub-block TMVP from being activated and used by the encoder.
[0056] According to an embodiment of the second aspect, the encoder 10, e.g. the encoder 10 of FIG. 1, is for layered video coding, e.g. as described with respect to FIG. 1, for encoding images 26 of a video 20 into a multi-layer data stream 14 (or a multi-layer video bitstream 14) in units of blocks 74 into which the images 26 are subdivided, e.g. by pre-dividing each image into one or more tree root blocks 72 and performing a recursive block division on each tree root block 72 of each image 26, e.g. as described with respect to FIGS. 2 and 4. Each image 26 is associated with one of the layers 24, e.g. as described with respect to FIG. 1. The size of the tree root block 72 may be equal for all images belonging to the same layer 24. The encoder 10 according to the embodiment of the second aspect encodes a first image 261 of a first layer 241, e.g. as described with respect to FIG. 3. * (see FIG. 1), populate a list of reference images, e.g., list 561 or list 562, from multiple previously coded images. As described with respect to FIG. 4, the list of reference images includes the first image 261 * 1. The list of reference images may include images of the same layer and different time stamps, for example image 261′ of FIG. 1, where the image is encoded in the multi-layer data stream 14 upstream with respect to the first image 261. Furthermore, the list of reference images may include one or more second images, which belong to a different layer, for example the second layer 240, and which are encoded in the multi-layer data stream 14 upstream with respect to the first image 261. * The encoder then outputs the first image 261 to the second image 260', which is aligned in time with the first image 261. * The encoder 10 and the decoder 50 may signal in the multi-layer video bitstream 14 the size of the tree root block of the first layer 241 to which the first image 261 belongs, and the size of a different layer, e.g., the second layer 240 to which the second image belongs. * The inter-prediction block may be inter-predicted and the list of reference images may be used to predict the motion vector of the inter-prediction block.
[0057] According to a first embodiment of the second aspect, the size of the tree root block of the second image 260' is * For example, the video encoder may populate the list of reference pictures 56, or the two lists of reference pictures 561, 562 described with respect to FIG. 3, such that for each reference picture in the list of reference pictures, the size of the tree root block is equal to or an integer multiple of the size of the tree root block in the first picture 261. * It is equal to the size of the tree root block or an integer multiple of it.
[0058] In other words, in the first embodiment example, the use of TMVP and sub-block TMVP is disabled by imposing a constraint on the ILRP of the reference picture list of the current picture as follows: The following constraints apply to each ILRP entry referenced by RefPicList[0] or RefPicList[1] of a slice of the current image, if present: o The image shall be in the same AU as the current image. o Images shall be present in the DPB. o The image shall have a nuh_layer_id refPicLayerId that is less than the nuh_layer_id of the current image. o The image shall have a value of sps_log2_ctu_size_minus5 that is the same as or greater than the current image. o One of the following constraints applies: o The image shall be an IRAP image. o The Image shall have a TemporalId less than or equal to Max(0,vps_max_tid_il_ref_pics_plus1[currLayerIdx][refLayerIdx]-1), where currLayerIdx and refLayerIdx are equal to GeneralLayerIdx[nuh_layer_id] and GeneralLayerIdx[refpicLayerId], respectively.
[0059] The worst case of the above problem occurs when the CTU size of the reference image is small. This is because the memory bandwidth requirements of TMVP and sub-block TMVP are high. When the CTU of the reference image is small, it is not important to obtain candidates for TMVP or sub-block TMVP, so the above embodiment is only applicable when the CTU size of the reference image is small. Therefore, it is not a problem when the CTU size of the reference image is larger than the CTU size of the current image, as in the above embodiment.
[0060] Nevertheless, according to an example of a first embodiment of the second aspect, the constraint is that for each inter-layer reference image in the list of reference images, i.e. for each second image (and, as a consequence, for all reference images, since e.g. images of the same layer may have tree root blocks of the same size), the size of its tree root block must be less than or equal to the size of the first image 261. * It is required that the size of the tree root block be equal to, or an integer multiple of, the size of the tree root block.
[0061] Therefore, in the first example embodiment, the use of TMVP and sub-block TMVP is disabled by imposing the following constraints on the ILRP of the reference picture list of the current picture: The following constraints apply to each ILRP entry referenced by RefPicList[0] or RefPicList[1] of a slice of the current image, if present: o The image shall be in the same AU as the current image. o Images shall be present in the DPB. o The image shall have a nuh_layer_id refPicLayerId that is less than the nuh_layer_id of the current image. o The image shall have the same value of sps_log2_ctu_size_minus5 as the current image (i.e. the image shall be the same as the image currently being coded 261 * (It shall have the same tree root block size as o One of the following constraints applies: o The image shall be an IRAP image. o The Image shall have a TemporalId less than or equal to Max(0,vps_max_tid_il_ref_pics_plus1[currLayerIdx][refLayerIdx]-1), where currLayerIdx and refLayerIdx are equal to GeneralLayerIdx[nuh_layer_id] and GeneralLayerIdx[refpicLayerId], respectively.
[0062] However, this embodiment is a very restrictive constraint that prevents any form of prediction, such as sample prediction, that does not recognize different CTU sizes in the dependent and reference layers.
[0063] Thus, according to a second embodiment of the second aspect, the list of reference images may include images of the type of second image that do not necessarily have a smaller or equal size than the tree root block 720, i.e. images of different layers such as the second layer 240, but the list of reference images may include any of the second images (images of the second layer 240). According to this embodiment, the encoder 10 performs the first image 261 * 3, the encoder 10 designates one image from the list of reference images as the TVMP reference image, so that the TVMP reference image is not the second image 260, and the size of the tree root block 720 of the first image 261 is smaller than that of the first image 261. * In other words, when the encoder 10 selects one of the second images 260 as a TVMP reference image, i.e., when the encoder 10 selects an image from a different layer, such as layer 240, as the first image 261, the size of the tree root block 721 is smaller than the size of the tree root block 721 of the first image 261. * When the first image 261 is selected as the TVMP reference image, the TVMP reference image is *The encoder 10 may signal in the multi-layer video data stream 14 a pointer that identifies a TVMP reference picture from a picture in the list of reference pictures. The encoder 10 may signal in the multi-layer video data stream 14 a pointer that identifies a TVMP reference picture from a picture in the list of reference pictures. * A TVMP reference image is used to predict the motion vector of the inter prediction block.
[0064] By restricting the specification of TMVP reference pictures from the list of reference pictures instead of restricting the population of the list, it is possible to allow other predictors to use a second picture 260 that has a smaller tree root block size than the first picture.
[0065] For example, the list of reference images may be one of the first list 561 and the second list 562 as described with respect to Figure 3. However, the method of selecting the TMVP reference list does not necessarily have to follow the method described with respect to Figure 3, but may be performed in another way, e.g. as described in the state of the art.
[0066] For example, encoder 10 may select a TVMP reference picture such that if the TVMP reference picture is a second picture, i.e., a picture of a different layer than the picture currently being coded, it does not meet any of the criteria in the following set of criteria: The size of the tree root block 720 of the second image 260 is * The size of the tree root block 721 is smaller than that of the tree root block 721. - TVMP reference image size is 261 for the first image * Different size. The scaling window of the TVMP image, which is the scaling window used for scaling and offsetting the motion vectors, is the first image 261 * The scaling window is different from that of The sub-image subdivision of the TVMP reference image 260' is * This is different from image segmentation.
[0067] The encoder 10 receives the first image 261 * When inter predicting the inter predicted blocks, activate a set of one or more inter prediction refinement tools according to a reference image of the list of reference images from which each inter predicted block is inter predicted that satisfies any of the above criteria sets for each inter prediction. For example, the set of inter prediction refinement tools may include one or more of TVMP, PROF, wraparound, VDOF, and DVMR.
[0068] In other words, according to the example of the second embodiment, the problem is solved by imposing constraints on the syntax element sh_collocated_ref_idx, which indicates the reference image used for the TVMP and the sub-block TVMP. sh_collocated_ref_idx specifies the reference index of the collocated picture used for temporal motion vector prediction. If sh_slice_type is equal to P, or if sh_slice_type is equal to B and sh_collocated_from_l0_flag is equal to 1, sh_collocated_ref_idx refers to an entry in reference image list 0, and the value of sh_collocated_ref_idx ranges from 0 to NumRefIdxActive[0]-1. If sh_slice_type is equal to B and sh_collocated_from_l0_flag is equal to 0, sh_collocated_ref_idx refers to an entry in reference image list 1, and the value of sh_collocated_ref_idx ranges from 0 to NumRefIdxActive[1]-1. If sh_collocated_ref_idx is not present, the following applies: - If pps_rpl_info_in_ph_flag is equal to 1, the value of sh_collocated_ref_idx is inferred to be equal to ph_collocated_ref_idx. - Otherwise (pps_rpl_info_in_ph_flag is equal to 0), the value of sh_collocated_ref_idx is inferred to be equal to 0. Set colPicList equal to sh_collocated_from_10_flag?0:1. The picture referenced by sh_collocated_ref_idx is the same for all non-I slices of the coded picture, the value of RprConstraintsActiveFlag[colPicList][sh_collocated_ref_idx] is equal to 0, and the value of sps_log2_ctu_size_minus5 of the picture referenced by sh_collocated_ref_idx is greater than or equal to the value of sps_log2_ctu_size_minus5 of the current picture, however this is a bitstream conformance requirement. NOTE - The above constraints require that the collocated image has the same spatial resolution, the same scaling window offset, and the same or smaller CTU size as the current image.
[0069] Again, the above example can prevent the case where the selected inter-layer reference picture has a smaller tree root block size than the first picture. * TVMP reference image, the reference image is TVMP reference image is selected image 261 * The tree root block size may be specified to be equal to the size of the tree root block 721 of the first embodiment.
[0070] Thus, another exemplary implementation is as follows. sh_collocated_ref_idx specifies the reference index of the collocated picture used for temporal motion vector prediction. If sh_slice_type is equal to P, or if sh_slice_type is equal to B and sh_collocated_from_l0_flag is equal to 1, sh_collocated_ref_idx refers to an entry in reference image list 0, and the value of sh_collocated_ref_idx ranges from 0 to NumRefIdxActive[0]-1. If sh_slice_type is equal to B and sh_collocated_from_l0_flag is equal to 0, sh_collocated_ref_idx refers to an entry in reference image list 1, and the value of sh_collocated_ref_idx ranges from 0 to NumRefIdxActive[1]-1. If sh_collocated_ref_idx is not present, the following applies: - If pps_rpl_info_in_ph_flag is equal to 1, the value of sh_collocated_ref_idx is inferred to be equal to ph_collocated_ref_idx. - Otherwise (pps_rpl_info_in_ph_flag is equal to 0), the value of sh_collocated_ref_idx is inferred to be equal to 0. Set colPicList equal to sh_collocated_from_10_flag?0:1. The picture referenced by sh_collocated_ref_idx is the same for all non-I slices of the coded picture, the value of RprConstraintsActiveFlag[colPicList][sh_collocated_ref_idx] is equal to 0, and the value of sps_log2_ctu_size_minus5 of the picture referenced by sh_collocated_ref_idx is equal to the value of sps_log2_ctu_size_minus5 of the current picture, which are bitstream conformance requirements. NOTE - The above constraints require that the collocated image has the same spatial resolution, the same scaling window offset, and the same CTU size as the current image.
[0071] According to a third embodiment of the second aspect, the encoder 10 and the decoder 50 determine whether a given image 261 satisfies any of the above criteria sets, depending on the reference image used to inter predict each inter prediction block. * In predicting an inter-predicted block of the image 26 to be coded, the above set of one or more inter-prediction refinement tools may be employed. Note that this constraint for using the inter-prediction refinement tools is based on the fact that the reference image is the image 26 to be coded. * The present invention is not necessarily limited to the multi-layer case where the inter-layer reference picture is a multi-layer reference picture.
[0072] In an example, the encoder 10 and decoder 50 may derive a list of reference pictures from which a reference picture is selected, as described with respect to FIG. 3, although different approaches for signaling or selecting a reference picture may also be possible.
[0073] In other words, according to the embodiment, the encoder 10 and the decoder 50 code (i.e., encode in the case of the encoder 10 and decode in the case of the decoder 50) an image currently being coded in units of blocks obtained by recursive block division of the tree root block as described above. * When inter predicting an inter predicted block, a set of one or more inter prediction refinement tools may be used depending on whether any of the following sets of criteria are met: -Image 26 currently being coded * The size of the tree root block of the image is smaller than the size of the tree root block of the reference image. -The size of the reference image is 26 * is equal to the size of For example, the scaling window of the reference image, whose scaling window is used for scaling and offsetting the motion vectors, is equal to the scaling window of the given image, but differs with respect to the offset of the scaling window. the sub-image subdivision of the reference image is equal to the sub-image subdivision of the given image, i.e. for example the sub-image subdivisions differ with respect to the number of sub-images of the sub-image subdivision;
[0074] In the example, the reference set is such that the size of the tree root block of the reference image is 26 * may include being equal to the size of the tree root block.
[0075] In an example, in inter prediction of an inter prediction block, the encoder 10 and the decoder 50 may selectively activate a set of inter prediction refinement tools if all of a subset of a set of criteria are satisfied. In another example, the encoder 10 and the decoder 50 may activate an inter prediction refinement tool if all of a set of criteria are satisfied.
[0076] For example, according to the third embodiment, the constraint is represented by a derivation variable RprConstraintsActiveFlag[refPicutre][currentPic], which is derived by comparing the characteristics of the current and reference images (image size, scaling window offset, number of sub-images, etc.). This variable is used to constrain the specified ph_collocated_ref_idx of the image header on the sh_collocated_ref_idx of the slice header. In this embodiment, the size of the CTU in the reference and current images (sps_log2_ctu_size_minus5) is incorporated into each derivation of RprConstraintsActiveFlag[refPicutre][currentPic], so that if the CTU size of the reference image is larger, the CTU size of the current image RprConstraintsActiveFlag[refPicutre][currentPic] is derived as 1.
[0077] Similarly, except that the criterion is that the size of the tree root block of the reference image is *In the case where the constraint includes being equal to the size of the tree root block of the current picture, the constraint may be represented by a derivation variable RprConstraintsActiveFlag[refPicutre][currentPic], which is derived by comparing the characteristics of the current picture and the reference picture (picture size, scaling window offset, number of sub-pictures, etc.). This variable is used to constrain the specified ph_collocated_ref_idx of the picture header on the sh_collocated_ref_idx of the slice header. In this embodiment, the size of the CTU in the reference picture and the current picture (sps_log2_ctu_size_minus5) is incorporated into each derivation of RprConstraintsActiveFlag[refPicutre][currentPic], so that if the CTU sizes are different, RprConstraintsActiveFlag[refPicutre][currentPic] is derived as 1.
[0078] In such cases, this embodiment does not allow tools such as PROF, wraparound, BDOF, and DMVR, as it is undesirable to allow different CTU sizes when such tools are used.
[0079] While the above discussion has focused only on TMVP and sub-block TMVP, further issues have been identified that apply when different CTUs are used in different layers. For example, TMVP and sub-block TMVP are permitted for pictures with different CTU sizes with associated "drawbacks", but problems arise when combined with sub-pictures. Indeed, when sub-pictures are used and used together with layer coding, there is a set of constraints that are required to ensure that the sub-picture grids are aligned. This is done for layers that have sub-pictures in their dependency tree.
[0080] FIG. 6 shows an example of a sub-image subdivision of an image 261 into two sub-images 281′, 281″. The image 261 may belong to a first layer, such as the layer 241, and may belong to an access unit 22 as described with respect to FIG. 1. In other words, the image 261 of FIG. 6 may be one of the images 261 of FIG. 1. The image 261 may depend on a reference image 260, for example the second layer 240 of FIG. 1. In other words, the image 260 may be an inter-layer reference image of the image 261, and the layer to which the image 260 belongs may be called the reference layer of the layer to which the image 261 belongs. For example, all images belonging to the same coding video sequence 20 and belonging to the same layer may be subjected to the same sub-image subdivision, i.e., may be subdivided into the same number of sub-images having equal sizes. It should also be noted that all images of the coding video sequence 20 belonging to the same layer 24 may be divided into tree root blocks of equal size. The encoder 10 may be configured to subdivide the image 261 of the first layer 241 and the image 260 belonging to a reference layer of the first layer 261, i.e. the second layer 240, into two or more common numbers of sub-images 28. Such a subdivision into two or more common sub-images is illustrated in FIG. 6, where the image 260 of the reference layer 240 is subdivided into two sub-images 280′, 280″. The encoder 10 may encode the sub-images 28′, 28″, i.e. the sub-images belonging to a common layer, e.g. the layer 241 or the layer 240, independently of each other. That is, the encoder 10 may encode the independently coded sub-images without spatial prediction of one part of the sub-image 28′ by another sub-image 28″ of the sub-image. For example, the encoder 10 may apply sample padding to the boundary between the images 28′ and 28″, such that the coding of the two sub-images is independent of each other. The encoder 10 may also clip the motion vectors at the boundary between sub-pictures 28' and 28". Optionally, the encoder 10 may signal in the video data stream 14 that the sub-pictures of a layer are coded independently, for example by signaling sps_subpic_treated_as_pic_flag=1.The encoder 10 may code the sub-pictures 28 into the video bitstream 14 in units of blocks 74, 76, which may result from, for example, the division of one or more tree root blocks 72 into which the sub-pictures 28 are divided, as described with respect to Figures 2 and 4 with respect to the picture 26. The encoder 10 may divide the first layer 241 of the picture 26 and the sub-pictures 281, 280 of the reference layer 240 of the first layer 241, which are subdivided into a common number of sub-pictures, into tree root blocks of equal size. In other words, if the picture 261 of the first layer 241 and the reference picture 260 of the reference layer 240 are subdivided into a common number of two or more sub-pictures, the encoder 10 may subdivide the sub-pictures 281, 280 such that the tree root blocks of the sub-picture 281 of the first layer 241 and the sub-picture 280 of the reference layer 240 are of the same size.
[0081] In other words, in another alternative embodiment, the problem is solved only for the case of independent sub-images by extending the sub-image related constraints as follows: sps_subpic_treated_as_pic_flag[i] equal to 1 specifies that the i-th sub-picture of each coded picture in the CLVS is treated as a picture of the decoding process except for in-loop filtering operations. sps_subpic_treated_as_pic_flag[i] equal to 0 specifies that the i-th sub-picture of each coded picture in the CLVS is not treated as a picture of the decoding process except for in-loop filtering operations. If not present, the value of sps_subpic_treated_as_pic_flag[i] is inferred to be equal to 1. If sps_num_subpics_minus1 is greater than 0 and sps_subpic_treated_as_pic_flag[i] is equal to 1, for each CLVS of the current layer that references an SPS, let targetAuSet be all AUs starting with the AU containing the first picture of the CLVS in decoding order through the AU containing the last picture of the CLVS in decoding order. Bitstream conformance requires that all of the following conditions are true for the targetLayerSet consisting of the current layer and all layers that have the current layer as a reference layer: - For each AU in targetAuSet, all images of layers in targetLayerSet shall have the same value of pps_pic_width_in_luma_samples and the same value of pps_pic_height_in_luma_samples. -All SPS referenced by layers in targetLayerSet shall have the same value of sps_num_subpics_minus1 and shall have the same values of sps_subpic_ctu_top_left_x[j], sps_subpic_ctu_top_left_y[j], sps_subpic_width_minus1[j], sps_subpic_height_minus1[j], and sps_subpic_treated_as_pic_flag[j], respectively, in the range from 0 to sps_num_subpics_minus1 and for each value of j of sps_log2_ctu_size_minus5. - For each AU in targetAuSet, all images in the layers in targetLayerSet shall have the same value of SubpicIdVal[j] for each value of j in the range 0 to sps_num_subpics_minus1.
[0082] According to an embodiment, the encoder 10 and the decoder 50 select a block, e.g. a currently coded image 26 of the first layer 241. * Block 74 currently being coded *or subblock 76 currently being coded * may be inter-predicted using coding parameters such as motion vectors of corresponding blocks of images in the reference layer 240. * is block 74 in FIG. * In other words, the corresponding block may be collocated (located at the same position) with the block currently being coded, e.g., in a position within each image, i.e., in each reference image, and may correspond to or be referenced by a motion vector of the block currently being coded. Encoder 10 and decoder 50 may use the coding parameters of the corresponding block to predict coding parameters, such as a motion vector, of the block currently being coded.
[0083] By subdividing the sub-pictures of the reference layer 240 into tree root blocks having the same size as the tree root blocks of the first layer that depend on the reference layer, it is possible to ensure that the sub-pictures 28 are coded independently of each other. In this respect, similar considerations apply as described with respect to Figure 4 for the boundaries of the tree root blocks. In other words, the considerations with respect to Figure 4 that explain the advantages of equal sized tree root blocks for the subdivision of the picture 26 may also be applied to the subdivision of the sub-pictures 28.
[0084] Fig. 7 shows an example of a decoder 50 and a video bitstream 14 according to an embodiment of the third aspect of the invention. The decoder 50 and the video bitstream 14 may optionally correspond to the decoder 50 and the video bitstream 14 of Fig. 1. According to an embodiment of the third aspect, the video bitstream 14 is a multi-layer video bitstream comprising access units 22, each of which comprises one or more images 26 of the coded video sequence 20 coded in the video bitstream 14. For example, the description of the layers 24 and the access units 22 of Fig. 1 may be applied. Each image 26 thus belongs to one of the layers 24, the association between the image and the layer being indicated by a subscript. That is, the image 261 belongs to the first layer 241 and the image 260 belongs to the second layer 240. In Fig. 7, the images 26 and the access units 22 to which they belong are shown according to the image order of the coded video sequence 20, i.e. their presentation order. Each access unit 22 belongs to a temporal sublayer of a set of temporal sublayers. For example, in FIG. 7, the access unit 221 to which the image with the superscript 1 belongs belongs to the first temporal sublayer, and the access unit 222 to which the image referenced using the reference number with the superscript 2 belongs belongs to the second temporal sublayer. As shown in FIG. 7, the access units that belong to different temporal sublayers do not necessarily have the same set of images of the layer 24. For example, in FIG. 7, the access unit 221 may have images of the first layer 241 and the second layer 240, and the access unit 222 may only have images of the first layer 241. The temporal sublayers may be hierarchically ordered and indexed with an index that represents the hierarchical order. In other words, the hierarchical order of the temporal sublayers may define the highest and lowest temporal sublayers within the set of temporal sublayers. It should be noted that the video bitstream 14 may optionally have further temporal sublayers and / or further layers.
[0085] According to an embodiment of the third aspect, the video bitstream 14 has encoded therein an output layer set indication 81 indicating one or more output layer sets 83. The OLS indication 81 indicates for the OLS 83 a subset (not necessarily a proper subset) of layers 24 of the multi-layer video bitstream 14 that belong to the OLS, i.e., the OLS may indicate that all of the layers of the multi-layer audio bitstream 14 belong to the OLS. For example, the OLS may be an indication of a sub-bitstream (not necessarily a proper) extractable or decodable from the video bitstream 14, the sub-bitstream comprising a subset of layers 24. For example, by extracting or decoding a subset of the layers of the video bitstream 14, the decoded coded video sequence may be quality scalable and therefore bitrate scalable.
[0086] According to an embodiment of the third aspect, the video bitstream 14 further includes a video parameter set (VPS) 91. The VPS 91 includes one or more bitstream conformance sets 86, e.g., hypothetical reference decoder (HRD) parameter sets. The video parameter set 91 further includes one or more buffer requirement sets 84, e.g., decoded picture buffer (DPB) parameter sets. The video parameter set 91 further includes one or more decoder requirement sets 82, e.g., profile layer level parameter sets (PTL sets). Each of the bitstream conformance set 86, the buffer requirement set 84, and the decoder requirement set 82 is associated with each temporal subset indication 96, 94, 92, respectively, indicated in the video parameter set 91. A constraint on the maximum temporal sublayers for each of the parameter sets, i.e., the bitstream conformance set 86, the buffer requirement set 84, and the decoder requirement set 82, may represent an upper limit on the number of temporal sublayers that the parameters of each parameter set refer to. In other words, the parameters signaled by a parameter set may be valid for a (not necessarily proper) sub-sequence of the coded video sequence 20, or a sub-bitstream of the video bitstream 14, defined by a set of layers and a set of temporal sublayers, and the constraint on the maximum temporal sublayer of each parameter set indicates the maximum temporal sublayer of the sub-bitstream or sub-sequence to which each parameter set refers.
[0087] According to a first embodiment of the third aspect, the decoder 50 may be configured to receive a maximum temporal sublayer indication 99. The maximum temporal sublayer indication 99 indicates a maximum temporal sublayer of the multi-layer video bitstream 14 to be decoded by the decoder 50. In other words, the maximum temporal sublayer indication 99 may signal to the decoder 50 which set or which subset of temporal layers of the video bitstream 14 the decoder 50 will decode. Thus, the decoder 50 may receive the maximum temporal sublayer indication 99 from an external signal. For example, the maximum temporal sublayer indication 99 may be included in the video bitstream 14 or may be provided to the decoder 50 via an API. Upon receiving the maximum temporal sublayer indication 99, the decoder 50 may decode the video bitstream 14 or a portion thereof as long as it belongs to the set of temporal sublayers indicated by the maximum temporal sublayer indication 99. For example, the decoder 50 may further be configured to receive an indication of the OLS to be decoded. Upon receiving an indication of the OLS to be decoded and receiving the maximum temporal sublayer indication 99, the decoder 50 may decode the layers indicated by the OLS to be decoded up to the temporal sublayer indicated by the maximum temporal sublayer indication 99. However, there may be situations or scenarios in which the decoder 50 does not receive an external indication of the temporal sublayers, i.e., one or both of the maximum temporal sublayer indication 99 and the indication of the OLS to be decoded. In such situations, the decoder 50 may determine a dropping indication based on information available in the video bitstream 14, for example.
[0088] According to a first embodiment of the third aspect, if the maximum temporal sublayer indication 99 is not available, e.g., not available from the data stream and / or not available via other means, the decoder 50 determines that the maximum temporal sublayer to be decoded is equal to the maximum temporal sublayer indicated by the constraint 92 on the maximum temporal sublayer in the decoder requirement set 82 associated with the OLS 83. In other words, the decoder 50 may use the constraint 92 for the maximum temporal sublayer indicated for the decoder requirement set 82 associated with the OLS to be decoded by the decoder 50.
[0089] To decode the multi-layered video bitstream 14, the decoder 50 may consider a temporal sub-layer of the multi-layered video bitstream for decoding if the temporal sub-layer does not exceed the maximum temporal sub-layer to be decoded, and if so, may omit the temporal sub-layer when decoding and use information regarding the maximum temporal sub-layer to be decoded.
[0090] For example, video bitstream 14 may indicate one or more OLSs 83, each OLS 83 associated with one of a bitstream conformance set 86, a buffer requirement set 84, and a decoder requirement set 82 signaled in a video parameter set 91. Decoder 50 may select an OLS 83 for decoding based on an external instruction (e.g., provided via an API or video bitstream 14) or, for example, in the absence of an external instruction, based on a selection rule. If a maximum temporal sublayer indication 99 is not available, i.e., if decoder 50 does not receive a maximum temporal sublayer indication, decoder 50 uses constraint 92 on maximum temporal sublayer of decoder requirement set 82 associated with the OLS to decode.
[0091] For example, the constraint 82 of the decoder requirement set 82 associated with the OLS is due to a bitstream constraint below, for example, the constraints 94, 96 on the maximum temporal sublayer associated with the buffer requirement set 84 and the bitstream conformance set 86. Thus, selecting the constraint 92 associated with the decoder capability set for decoding would select a minimum value that exceeds the maximum temporal sublayer indicated for the decoder requirement set 82, the buffer requirement set 84, and the bitstream conformance set 86 associated with the OLS. Selecting a minimum value that exceeds the maximum temporal sublayer constraint may ensure that each parameter set 82, 84, 86 contains parameters that are valid for the bitstream selected for decoding. Thus, selecting the constraint 92 associated with the decoder requirement set 82 may ensure that a bitstream for decoding is selected for which all parameters of the parameter sets 82, 84, 86 are available.
[0092] For example, each of the parameter sets from the bitstream conformance set 86, the buffer requirement set 84, and the decoder requirement set 82 associated with the OLS may include one or more sets of parameters, and each of the parameter sets is associated with a temporal sublayer or a maximum temporal sublayer. From each of the parameter sets, the decoder 50 may select a parameter set associated with the maximum temporal sublayer to be decoded, as estimated or received. For example, if the parameter set is associated with the maximum temporal sublayer, or if the parameter set is associated with a temporal sublayer less than or equal to the maximum temporal sublayer, the parameter set may be associated with the maximum temporal sublayer. The decoder 50 may use the selected parameter set to adjust one or more of the coded picture buffer size, the decoded picture buffer size, buffer scheduling, e.g., HRD timing (AU / DU removal time, DPB output time).
[0093] As mentioned above, the indication 99 of the maximum temporal sublayer to be decoded may be signaled in the video bitstream 14. According to a first embodiment, the encoder 10, e.g., the encoder 10 of FIG. 1, may be configured to optionally omit signaling of the indication 99 if the maximum temporal sublayer to be decoded corresponds to an OLS, e.g., a constraint 92 associated with the decoder requirement set 82 of the OLS instructed to decode, since in this case the decoder 50 can accurately estimate the maximum temporal sublayer to be decoded. In other words, in this case the encoder 10 may signal the indication 99 or may not signal the indication 99. In an example, the encoder 10 may decide whether to signal the indication 99 or whether to omit signaling of the indication 99 in this case.
[0094] For example, VPS91 may follow the example below: Currently, there are three syntax structures for VPS that are generically defined and then mapped to specific OLS. Profile Hierarchy Level (PTL), e.g., Decoder Requirement Set 82; DPB parameters, e.g. buffer requirement set 84, HRD parameters, e.g., Bitstream Conformance Set 86, and another syntax element vps_max_sublayers_minus1, which is used when extracting the OLS sub-bitstreams, i.e., possibly deriving the variable NumSublayerInLayer[i][j].
[0095] Figure 8 shows an example of the definition of PTL 82, DPB 84, and HRD 86 and their mapping to OLS 83. The mapping from PTL to OLS is done in the VPS of all OLS (single layer or multi-layer). However, the mapping of DPB and HRD parameters to OLS is done only in the VPS of OLS with multiple layers. As shown in Figure 8, the parameters of PTL, DPB, and HRD are first described in the VPS and then mapped to indicate which parameters OLS uses.
[0096] In the example shown in Figure 8, there are two OLSs and two of each of these parameters. However, the definitions and mappings are specified to allow multiple OLSs to share the same parameters, so there is no need to repeat the same information multiple times, for example as shown in Figure 9. Figure 9 shows an example of the definitions of PTL, DPB, and HRD and their sharing between different OLSs. Here, in the example of Figure 9, OLS2 and OLS3 have the same PTL and DBP parameters, but different HRD parameters.
[0097] In the examples of Figures 8 and 9, the values of vps_ptl_max_temporal_id[ptlIdx] (e.g., constraint on maximum temporal sublayers 92 in decoder requirement set 82), vps_dpb_max_temporal_id[dpbIdx] (e.g., constraint on maximum temporal sublayers 94 in buffer requirement set 84), and vps_hrd_max_tid[hrdIdx] (e.g., constraint on maximum temporal sublayers 96 in bitstream conformance set 86) for a given OLS 83 are aligned, although this is not currently required. These three values associated with the same OLS are not currently restricted to have the same value (but may optionally be restricted in some examples of the present disclosure). ptlIdx, dpbIdx, and hrdIdx are indices into each syntax structure signaled for each OLS.
[0098] According to the first embodiment example, the value of vps_ptl_max_temporal_id[ptlIdx] is used to set the variable HTid of the decoding process if not set by external means as follows: if no external means are provided to set the value of HTid (e.g. via the decoder API), the value of vps_ptl_max_temporal_id[ptlIdx] is taken by default to set HTid, i.e. the minimum value of the three syntax elements described above. In other words, the decoder 50 may set the variable HTid according to the maximum temporal sublayer indication 99 if it is available, otherwise HTid may be set to the value of vps_ptI_max_temporal_id[ptlldx].
[0099] According to a second embodiment of the third aspect, the decoder 50 estimates that the maximum temporal sublayer of the set of temporal sublayers to which each image 26 of the layer 24 included in the OLS belongs is the minimum of the maximum temporal sublayers indicated by the bitstream conformance set 86, the buffer requirement set 84, and the decoder requirement set 82 associated with the OLS. For example, the estimated maximum temporal sublayer for the set of temporal sublayers may correspond to or be indicated by the variable MAX-TID WITHINOLS. In other words, the set of temporal sublayers to which each image of a layer of the OLS belongs may be a set of temporal sublayers that contains all images that belong to the OLS. In this respect, all images of the layers included in the OLS may belong to the OLS. For example, referring to FIG. 7, the OLS 83 including the first layer 241 and the second layer 240 may include a first temporal sublayer to which the access unit 221 belongs and a second temporal sublayer to which the access unit 222 belongs.
[0100] In an example, decoder 50 may detect, for one or more layers, or all layers, of an OLS to be decoded, whether video bitstream 14 indicates a constraint on the maximum temporal sublayers of a reference layer on which each layer depends. For example, the constraint may indicate that each layer depends only on temporal sublayers of a reference layer up to a maximum temporal sublayer. If video bitstream 14 does not indicate such a constraint, decoder 50 estimates that the maximum temporal sublayers included in the OLS are equal to the maximum temporal sublayers estimated for the set of temporal sublayers to which each image of a layer of the OLS belongs.
[0101] In an example, the OLS instruction 81 may further indicate one or more output layers for the OLS 83. In other words, one or more of the layers included in the OLS 83 may be indicated as output layers of the OLS 83. The decoder 50 may estimate, for each layer indicated to be an output layer of the OLS 83, that the maximum temporal sublayer included in the OLS is equal to the maximum temporal sublayer estimated for the set of temporal sublayers to which each image of the layer of the OLS belongs.
[0102] The decoder 50 may decode, from the images 26 of the video bitstream 14, images belonging to one of the layers included in the OLS to be decoded and images belonging to a temporal sublayer less than or equal to the maximum temporal sublayer to be decoded, which may be the maximum temporal sublayer estimated for the set of temporal sublayers to which each image of a layer of the OLS belongs, or the maximum temporal sublayer to be decoded as described with respect to the above embodiment.
[0103] In other words, according to the second embodiment, the above variable MaxTidWithiOls indicating the maximum number of temporal sublayers present in the OLS (not necessarily the bitstream, since some may have been dropped) is derived at the decoder side from the minimum of the three values of the syntax elements vps_ptl_max_temporal_id[ptlIdx], vps_dpb_max_temporal_id[dpbIdx], vps_hrd_max_tid[hrdIdx]. This minimizes the restriction of sharing any of the parameters PTL, HRD or DPB, and prohibits indicating an OLS where all three parameters are not defined. MaxTidWithinOls=min(vps_ptl_max_temporal_id[ptlIdx],min(vps_dpb_max_temporal_id[dpbIdx],vps_hrd_max_tid[hrdIdx]))
[0104] Furthermore, in this embodiment, NumSublayerlnLayer[i][j], which represents the maximum sublayer included in the i-th OLS of layer j, is set equal to the derived MaxTidWithinOls above when vps_max_tid_il_ref_pics_plus1[m][k] is not present or when layer j is an output layer of the i-th OLS.
[0105] According to an example, the decoder 50 may estimate that the maximum temporal sublayer of the decoded multi-layer video bitstream is equal to the maximum temporal sublayer estimated for the set of temporal sublayers to which each image of the layer of the OLS belongs, i.e., referred to as MaxTidWithinOLS.
[0106] In other words, in another embodiment, the derived value of MaxTidWithinOls may be used to set the variable HTid of the decoding process if not set by external means, as follows: If no external means are provided for setting the value of HTid (e.g., via the decoder API), then the value of MaxTidWithinOls is taken by default to be set to HTid, i.e., the minimum value of the three syntax elements above.
[0107] In an example embodiment according to the third aspect, the decoder 50 selectively considers for decoding images belonging to one of the temporal layers of the multi-layer video bitstream 14 if each image belongs to an access unit 22 associated with a temporal sublayer that does not exceed the maximum temporal sublayer to be decoded.
[0108] A further embodiment according to the third aspect includes an encoder 10, e.g., the encoder 10 of Fig. 1, for encoding a multi-layer video bitstream 14 according to Fig. 7. To this end, the encoder 10 may form an OLS indication 81 such that, in associating an OLS 83 with a bitstream conformance set 86, a buffer requirement set 84, and a decoder requirement set 82, a minimum of the constraints 96, 94, 92 associated with each parameter set accommodates a subset of layers indicated by the OLS 83. That is, for example, the minimum among the maximum temporal sublayers indicated by the bitstream conformance set 86, the buffer requirement set 84, and the decoder requirement set 82 associated with the OLS 83 is equal to or greater than the maximum temporal sublayer of the temporal sublayers included in the subset of layers indicated by the OLS. The encoder 10 may further form the OLS instruction 81 such that parameters in the bitstream compatibility set 86, buffer requirement set 84, and decoder requirement set 82 are valid for the OLS 83 as long as those parameters reference a temporal sublayer that is less than or equal to the minimum among the maximum temporal sublayers 92, 94, 96 indicated for the bitstream compatibility set 86, buffer requirement set 84, and decoder requirement set 82 associated with the OLS 83.
[0109] Although some aspects have been described as features in the context of an apparatus, it will be apparent that such descriptions may also be considered as descriptions of corresponding features of a method. Although some aspects have been described as features in the context of a method, it will be apparent that such descriptions may also be considered as descriptions of corresponding features with respect to the functionality of the apparatus.
[0110] Some or all of the method steps may be performed by (or using) a hardware apparatus, such as, for example, a microprocessor, a programmable computer or an electronic circuit. In some embodiments, one or more of the most important method steps may be performed by such an apparatus.
[0111] The inventive encoded image signal can be stored on a digital storage medium or can be transmitted over a transmission medium, such as a wireless transmission medium or a wired transmission medium, such as the Internet. In other words, a further embodiment provides a video bitstream product, including a video bitstream according to any of the embodiments described herein, such as a digital storage medium having stored thereon the video bitstream.
[0112] Depending on the specific implementation requirements, the embodiments of the present invention can be implemented in hardware or software, or at least partly in hardware or at least partly in software. This implementation can be performed using a digital storage medium, such as a floppy disk, DVD, Blu-Ray, CD, ROM, PROM, EPROM, EEPROM or FLASH memory, storing electronically readable control signals, which cooperate (or can cooperate) with a programmable computer system to execute the respective methods. Thus, the digital storage medium can be computer readable.
[0113] Some embodiments according to the invention include a data carrier having electronically readable control signals that can cooperate with a programmable computer system to perform one of the methods described herein.
[0114] Generally, embodiments of the present invention can be implemented as a computer program product having program code operable to perform one of the methods when the computer program product runs on a computer. The program code may for example be stored on a machine readable carrier.
[0115] Other embodiments comprise the computer program for performing one of the methods described herein, stored on a machine readable carrier.
[0116] In other words, an embodiment of the inventive method is, therefore, a computer program having a program code for performing one of the methods described herein, when the computer program runs on a computer.
[0117] A further embodiment of the inventive method is therefore a data carrier (or digital storage medium, or computer readable medium) comprising, recorded thereon, the computer program for performing one of the methods described herein. The data carrier, digital storage medium, or recorded medium is typically tangible and / or non-transitory.
[0118] A further embodiment of the inventive method is therefore a data stream or a sequence of signals representing the computer program for performing one of the methods described herein, the data stream or the sequence of signals being adapted to be transferred via a data communication connection, for example via the Internet.
[0119] A further embodiment comprises a processing means, such as for example a computer, or a programmable logic device, configured to or adapted to perform one of the methods described herein.
[0120] A further embodiment comprises a computer having installed thereon the computer program for performing one of the methods described herein.
[0121] Further embodiments according to the invention include an apparatus or system configured to transfer (e.g. electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver may be, for example, a computer, a mobile device, a memory device, etc. The apparatus or system may, for example, include a file server for transferring the computer program to the receiver.
[0122] In some embodiments, a programmable logic device (e.g., a field programmable gate array) may be used to perform some or all of the functions of the methods described herein. In some embodiments, a field programmable gate array may cooperate with a microprocessor to perform one of the methods described herein. In general, the methods are preferably performed by any hardware apparatus.
[0123] The apparatus described herein may be implemented using a hardware apparatus, or using a computer, or using a combination of a hardware apparatus and a computer.
[0124] The methods described herein may be performed using a hardware apparatus, or using a computer, or using a combination of a hardware apparatus and a computer.
[0125] In the above detailed description, it can be seen that various features are grouped together in examples for the purpose of streamlining the disclosure. This method of disclosure should not be interpreted as expressly expressly indicating that the claimed examples require more features than are expressly recited in each claim. Instead, as the following claims indicate, inventive subject matter resides in less than all features of a single disclosed example. As such, the following claims are hereby incorporated into the detailed description, with each claim standing on its own as a separate example. Although each claim may stand on its own as a separate example, it should be noted that although a dependent claim may refer to a specific combination with one or more other claims within the scope of the claim, other examples may also include combinations of the dependent claim with the subject matter of each other dependent claim, or combinations of each feature with other dependent or independent claims. Such combinations are suggested herein unless it is stated that a particular combination is not intended. Furthermore, it is intended to include features of a claim in any other independent claim, even if that claim is not directly dependent on the independent claim.
[0126] The above-described embodiments are merely illustrative of the principles of the present disclosure. Of course, modifications and variations of the configurations and details described herein will be apparent to those skilled in the art. It is therefore intended to be limited only by the scope of the appended claims and not by the specific details presented as descriptions and explanations of the embodiments herein.
Claims
1. 1. A decoder (50) for decoding a multi-layered video bitstream (14) representing a coded video sequence (20), the multi-layered video bitstream (14) comprising access units (22), each access unit (22) comprising one or more images of the coded video sequence (20), each of the images belonging to one of the layers of the multi-layered video bitstream (14), each of the access units (22) comprising a temporal sub-layer (22) of a set of temporal sub-layers of the coded video sequence (20). 1 , 22 2 ), and the decoder (50) From the multi-layer video bitstream (14), a video parameter set (91) including one or more bitstream conformance sets (86), one or more buffer requirement sets (84), and one or more decoder requirement sets (82); an OLS (83) for the multi-layer video bitstream (14), the OLS having OLS indications (81) for the OLS indicating a subset of layers of the multi-layer video bitstream (14), the OLS indications (81) associating the OLS with corresponding ones of the bitstream conformance set (86), the buffer requirement set (84), and the decoder requirement set (82); Derive In each of the bitstream conformance set (86), the buffer requirement set (84), and the decoder requirement set (82), a temporal subset indication indicates a constraint (92, 94, 96) on a maximum temporal sub-layer. The derivation, if a maximum temporal sub-layer indication (99) is not available, estimating that a maximum temporal sub-layer of the multi-layer video bitstream (14) to be decoded is equal to the maximum temporal sub-layer indicated by the constraint (92) on the maximum temporal sub-layer in the decoder requirement set (82) associated with the OLS; The decoder (50) is configured as follows.
2. selecting a set of parameters associated with the greatest temporal sub-layer of the multi-layer video bitstream (14) to be decoded from each of the bitstream conformance set (86), the buffer requirement set (84), and the decoder requirement set (82) associated with the OLS; and using the set of parameters to adjust one or more of a coded picture buffer size, a decoded picture buffer size, and buffer scheduling; A decoder (50) according to claim 1, configured so as to
3. selectively considering one of the temporal sub-layers of the multi-layer video bitstream for decoding if each temporal sub-layer does not exceed the maximum temporal sub-layer of the multi-layer video bitstream (14) to be decoded. A decoder (50) according to claim 1 or 2, configured so as to
4. An encoder (10) for providing a multi-layered video bitstream (14) representing a coded video sequence (20), the multi-layered video bitstream (14) comprising access units (22), each access unit (22) comprising one or more images of the coded video sequence (20), each of the images belonging to one of the layers of the multi-layered video bitstream (14), each of the access units (22) comprising a temporal sublayer (22) of a set of temporal sublayers of the coded video sequence (20). 1 , 22 2 ), and the encoder (10) The multi-layer video bitstream (14), a video parameter set (91) including one or more bitstream conformance sets (86), one or more buffer requirement sets (84), and one or more decoder requirement sets (82); an OLS (83) for the multi-layer video bitstream (14), the OLS having OLS indications (81) for the OLS indicating a subset of layers of the multi-layer video bitstream (14), the OLS indications (81) associating the OLS with corresponding ones of the bitstream conformance set (86), the buffer requirement set (84), and the decoder requirement set (82); providing In each of the bitstream conformance set (86), the buffer requirement set (84), and the decoder requirement set (82), a temporal subset indication indicates a constraint (92, 94, 96) on a maximum temporal sub-layer. The provision of selecting a maximum temporal sublayer of the multi-layer video bitstream (14) to be decoded by a decoder (50), and omitting signaling an indication (99) of the maximum temporal sublayer to be decoded in the multi-layer video bitstream (14) if the maximum temporal sublayer to be decoded is equal to the maximum temporal sublayer indicated by the constraint (92) on the maximum temporal sublayer of the decoder requirement set (82) associated with the OLS. The encoder (10) is configured as follows.
5. 5. The encoder of claim 4, configured to provide the indication of the maximum temporal sublayer of the multi-layer video bitstream if the maximum temporal sublayer to be decoded is not equal to the maximum temporal sublayer indicated by the constraint on the maximum temporal sublayer of the decoder requirement set associated with the OLS.
6. determining whether to omit signaling in the multi-layer video bitstream (14) an indication of the maximum temporal sublayer to be decoded (99) if the maximum temporal sublayer to be decoded is equal to the maximum temporal sublayer indicated by the constraint (92) on the maximum temporal sublayer of the decoder requirement set (82) associated with the OLS; An encoder (10) according to claim 4 or 5, configured as follows:
7. A method for decoding (50) a multi-layered video bitstream (14) representing a coded video sequence (20), the multi-layered video bitstream (14) comprising access units (22), each access unit (22) comprising one or more images of the coded video sequence (20), each of the images belonging to one of the layers of the multi-layered video bitstream (14), each of the access units (22) comprising a temporal sub-layer (22) of a set of temporal sub-layers of the coded video sequence (20). 1 , 22 2 ) , wherein the method comprises: From the multi-layer video bitstream (14), a video parameter set (91) including one or more bitstream conformance sets (86), one or more buffer requirement sets (84), and one or more decoder requirement sets (82); an OLS (83) for the multi-layer video bitstream (14), the OLS having OLS indications (81) for the OLS indicating a subset of layers of the multi-layer video bitstream (14), the OLS indications (81) associating the OLS with corresponding ones of the bitstream conformance set (86), the buffer requirement set (84), and the decoder requirement set (82); Derive In each of the bitstream conformance set (86), the buffer requirement set (84), and the decoder requirement set (82), a temporal subset indication indicates a constraint (92, 94, 96) on a maximum temporal sub-layer. The deriving; if a maximum temporal sub-layer indication (99) is unavailable, estimating that a maximum temporal sub-layer of the multi-layer video bitstream (14) to be decoded is equal to the maximum temporal sub-layer indicated by a constraint (92) on the maximum temporal sub-layer of the decoder requirement set (82) associated with the OLS; The method comprising:
8. A method for providing a multi-layered video bitstream (14) representing a coded video sequence (20), the multi-layered video bitstream (14) comprising access units (22), each access unit (22) comprising one or more images of the coded video sequence (20), each of the images belonging to one of the layers of the multi-layered video bitstream (14), each of the access units (22) comprising a temporal sub-layer (22) of a set of temporal sub-layers of the coded video sequence (20). 1 , 22 2 ) , wherein the method comprises: The multi-layer video bitstream (14), a video parameter set (91) including one or more bitstream conformance sets (86), one or more buffer requirement sets (84), and one or more decoder requirement sets (82); an OLS (83) for the multi-layer video bitstream (14), the OLS having OLS indications (81) for the OLS indicating a subset of layers of the multi-layer video bitstream (14), the OLS indications (81) associating the OLS with corresponding ones of the bitstream conformance set (86), the buffer requirement set (84), and the decoder requirement set (82); providing In each of the bitstream conformance set (86), the buffer requirement set (84), and the decoder requirement set (82), a temporal subset indication indicates a constraint (92, 94, 96) on a maximum temporal sub-layer. The providing; selecting a maximum temporal sublayer of the multi-layer video bitstream (14) to be decoded by a decoder (50), and omitting signaling an indication (99) of the maximum temporal sublayer to be decoded in the multi-layer video bitstream (14) if the maximum temporal sublayer to be decoded is equal to the maximum temporal sublayer indicated by the constraint (92) on the maximum temporal sublayer of the decoder requirement set (82) associated with the OLS; The method comprising:
9. A computer program for implementing any of the methods according to claims 7 or 8 when the computer program is run on a computer or signal processor.
Citation Information
Patent Citations
Avoidance of redundant signaling in multi-layer video bitstreams
WO2021022267A2
Systems and methods for decoding based on inferred video parameter sets
WO2021236312A1