Video encoding and decoding method and apparatus for determining the maximum time sublayer
By optimizing TMVP reference image selection within tree root block boundaries and determining the maximum time sublayer without explicit signaling, the solution addresses inefficiencies in video encoding and decoding, enhancing coding efficiency and reducing transmission overhead.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2026-04-02
Smart Images

Figure 0007839924000003 
Figure 0007839924000004 
Figure 0007839924000005
Abstract
Description
[Technical Field]
[0001] Embodiments of this disclosure relate to a video encoder, a video decoder, a method for encoding a video sequence into a video bitstream, and a method for decoding a video sequence from a video bitstream. Further embodiments relate to a video bitstream. [Background technology]
[0002] In encoding or decoding images in a video sequence, predictions are used to reduce the amount of information transmitted in the video bitstream from which the image is encoded / decoded. Predictions may be used on the image data itself, such as sample values or coefficients on which the sample values of an image are coded. Alternatively or additionally, predictions may be used on syntactic elements used to code an image, such as motion vectors. A reference image may be selected to predict the motion vector of the image to be coded, from which predictors for the motion vector of the image to be coded are determined. [Overview of the project] [Means for solving the problem]
[0003] A first aspect of this disclosure provides a concept for selecting reference images to be used for time motion vector prediction. Two lists of reference images are populated for a given image, e.g., an image to be coded. Each list may be empty or not. The TMVP reference image is determined by selecting one of the two lists of reference images as the TMVP image list, and then selecting the TMVP reference image from that TMVP image list. According to the first aspect, if one of the two lists is empty and the other is not, the reference image from the non-empty list is used for time motion vector prediction (TMVP). Thus, TMVP can be used regardless of which of the two lists is empty, and it provides high coding efficiency in both cases where only the first list or only the second list is empty.
[0004] A second aspect of this disclosure is based on the idea that the tree root block into which an image in a coded video sequence is divided is such that the reference image of that image is smaller in size than or equal to the tree root block into which it is divided. Imposing such constraints on the division of images into tree root blocks can ensure that the dependency of an image to a reference image does not extend beyond the boundaries of the tree root block, or at least not beyond the row boundaries of a row in the tree root block. Thus, the constraints can limit dependencies between different tree root blocks and benefit buffer management. In particular, dependencies between tree root blocks of different rows in a tree root block can lead to inefficient buffer usage because adjacent tree root blocks belonging to different rows may be separated by further tree root blocks in the coding order. Thus, by avoiding such dependencies, it is possible to eliminate the need to maintain an entire row of a tree root block between the tree root block currently being coded and the tree root block being referenced.
[0005] A third aspect of this disclosure provides a concept for determining the maximum time sublayer, which is the maximum layer of an output layer set represented by a multilayer video bitstream that should be decoded. This concept enables the decoder to determine which portion of the video bitstream to decode without any instruction about the maximum time sublayer to be decoded. Furthermore, this concept allows the encoder to omit signaling instructions for the maximum time sublayer to be decoded when the decoder estimates that the maximum time sublayer to be decoded in the absence of such instructions, thereby avoiding unnecessarily high signal transmission overhead.
[0006] Embodiments and advantageous embodiments of this disclosure are described in more detail below with reference to the drawings. [Brief explanation of the drawing]
[0007] [Figure 1] An encoder, decoder, and video bitstream according to an embodiment are shown. [Figure 2] An example of image subdivision is shown. [Figure 3] This demonstrates the determination of the TMVP reference image according to one embodiment. [Figure 4] This example demonstrates how to determine motion vector candidates that depend on tree root partitioning. [Figure 5] Two examples are shown with different tree root block sizes for the dependent layer and the reference layer. [Figure 6] This shows an example of sub-image subdivision of dependent and reference layers. [Figure 7] A decoder and video bitstream according to an embodiment of the third aspect are shown. [Figure 8] This shows an example of mapping between the output layer set and the video parameter set. [Figure 9] This shows an example of mapping between an output layer set and a video parameter set that share parameters. [Modes for carrying out the invention]
[0008] The embodiments described below are detailed, but naturally, the embodiments provide many applicable concepts that can be embodied in a variety of video coding concepts. The specific embodiments described are merely illustrative of specific ways of implementing and using the concepts and do not limit the scope of the embodiments. Several details are provided below to provide a more complete description of the embodiments of the disclosure. However, it will be apparent to those skilled in the art that other embodiments can be practiced without these specific details. In other examples, well-known structures and devices are shown in block diagram form rather than in detail, in order to avoid obscuring the examples described herein. Furthermore, features of different embodiments described herein may be combined with each other unless otherwise specified.
[0009] In the following descriptions of embodiments, identical or similar elements, or elements having the same function, are given the same reference numeral or identified by the same name, and repetition of the description of elements given the same reference numeral or identified by the same name is usually omitted. Therefore, the descriptions provided for elements having the same reference numeral or identified by the same name are interchangeable or applicable to each other in different embodiments.
[0010] A detailed description of the embodiments of the disclosed concept begins with a description of examples of encoders, decoders, and video bitstreams, which provide a framework in which embodiments of the present invention can be incorporated. Hereafter, a description of embodiments of the concept of the present invention is provided, along with a description of how such concepts can be incorporated into the encoder and decoder of Figure 1. However, embodiments described with respect to the subsequent Figure 2 and subsequent figures may be used to form encoders and decoders that do not operate according to the framework described with respect to Figure 1. It should also be noted that although the encoder and decoder are shown together in Figure 1 for illustrative purposes, they may be implemented separately. It should also be noted that the encoder and decoder may be combined within a single device, or one of the two may be implemented as part of the other. Furthermore, some embodiments of the present invention are described with reference to Figure 1.
[0011] Figure 1 shows an example of an encoder 10 and a decoder 50. The encoder 10 (which may also be called an encoding device) encodes a video sequence 12 into a video bitstream 14 (which may also be called a bitstream, data stream, video data stream, or stream). The video sequence 12 includes a sequence of images 21, which are arranged in presentation order or image order 17. In other words, each of the images 21 may represent a frame of the video sequence 12 and may be associated with a point in time in the presentation order of the video sequence 12. Based on the video sequence 12, the encoder 10 may encode a coded video sequence 20 into a video bitstream 14. The encoder 10 may form the coded video sequence 20 in the form of access units 22, each access unit 22 encoding video data belonging to a common point in time. In other words, each access unit 22 may encode one of the images 21 of the video sequence 12, i.e., one of the frames, within it. The encoder 10 encodes the video sequence 20, which has been coded according to the coding order 19, and the coding order 19 may differ from the image order 17 of the video sequence 12.
[0012] The encoder 10 may encode the coded video sequence 20 into one or more layers. That is, the video bitstream 14 may be a single-layer or multi-layer video bitstream containing one or more layers. Each access unit 22 contains one or more coded images 26 (for example, images 260, 261 in Figure 1, where apostrophes and asterisks are used to refer to specific ones, and subscripts indicate the layer to which the image belongs). Each image 26 belongs to one of the layers 24 of the coded video sequence, for example, layers 240, 241 in Figure 1. Figure 1 shows an exemplary number of two layers, namely the first layer 241 and the second layer 240. In embodiments of the disclosed concept, the coded video sequence 20 and video bitstream 14 do not necessarily contain multiple layers, but may contain one, two, or more layers. In the example in Figure 1, each access unit 22 includes the coded image 261 of the first layer 241 and the coded image 260 of the second layer 240. However, it should be noted that each access unit 22 may, but not necessarily, include coded images for each layer of the coded video sequence 20. For example, layers 240 and 241 may have different frame rates (or image rates), and / or may include images for a complementary subset of the access units of access unit 22.
[0013] As mentioned above, one image 260, 261 in an access unit represents the same image content at the same point in time. For example, images 260, 261 in the same access unit 22 may represent the same image content at different qualities, such as resolution or fidelity. In other words, layer 240 may represent a first version of the coded video sequence 20, and layer 241 may represent a second version of the coded sequence 20. Thus, a decoder, such as decoder 50, or extractor may select between different versions of the coded video sequence 20 to be decoded or extracted from the video bitstream 14. For example, layer 240 may be decoded independently of further layers of the coded video sequence to provide a decoded video sequence of a first quality, and joint decoding of the first layer 241 and the second layer 240 may provide a decoded video sequence of a second quality that is higher than the first quality. For example, the first layer 241 may be encoded independently of the second layer 240. In other words, the second layer 240 may be a reference layer for the first layer 241. For example, in this scenario, the first layer 241 may be called the enhancement layer, and the second layer 240 may be called the base layer. Image 260 may have a smaller image size than, equal to, or larger than, image 261. For example, image size may refer to the number of samples in a two-dimensional array of images. Note that images 260 and 261 do not necessarily have to represent the same image content; for example, image 261 may represent an excerpt of the image content of image 260. For example, in some scenarios, different layers of the video bitstream 14 may contain different sub-images of the image coded into the video bitstream.
[0014] The encoder 10 encodes the access units 22 into bitstream portions 16 of the video bitstream 14. For example, each access unit 22 may be encoded into one or more bitstream portions 16. For example, an image 26 may be subdivided into tiles of slices, and each slice may be encoded into one bitstream portion 16. The bitstream portion 16 into which the image 26 is encoded may be called a video coding layer (VCL) NAL unit. The video bitstream 14 may further include non-VCL NAL units, for example, bitstream portions 23, 29 into which descriptive data is encoded. The descriptive data may provide information for decoding or information about the coded video sequence 20. The bitstream portions into which the descriptive data is encoded may be associated with individual bitstream portions. For example, they may refer to individual slices, or to one of the images 26, or one of the access units 22, or to a sequence of access units, i.e., to the coded video sequence 20. Note that video 12 may be coded into the sequence of the coded video sequence 20.
[0015] The decoder 50 (which may also be called a decoding device) decodes the video bitstream 14 to obtain the decoded video sequence 20'. Note that the video bitstream 14 provided to the decoder 50 does not necessarily correspond to the video bitstream 14 provided by the encoder, and may be an extract from the video bitstream provided by the encoder. Therefore, the video bitstream decoded by the decoder 50 may be a sub-bitstream of the video bitstream encoded by an encoder such as the encoder 10. As described above, the decoder 50 may decode the entire coded video sequence 20 coded in the video data stream 14, or a part thereof, for example, a subset of the layers of the coded video sequence 20, and / or a time subset of the coded video sequence 20 (i.e., a video sequence with a frame rate lower than the maximum frame rate provided by the video sequence 20). Therefore, the decoded video sequence 20' does not necessarily correspond to the video sequence 12 encoded by the encoder 10. It should also be noted that the decoded video sequence 20' may differ further from video sequence 12 due to coding losses such as quantization losses.
[0016] Image 26 may be encoded using a prediction tool to predict signals or coefficients representing images in the video bitstream 14 from previously coded images. That is, the encoder 10 may encode a given image 26 * For example, a prediction tool may be used to encode the image to be encoded now using a previously encoded image. Accordingly, the decoder 50 will use the previously decoded image to encode the image to be decoded now 26 * A prediction tool may be used to predict. In the following description, a given image or block, for example, the image or block currently being coded, is referred to by the reference code ( *) may be represented using. For example, the image 261 in FIG. 1 * is considered to be the image currently being coded, where the image 26 currently being coded * may equally refer to the image currently being encoded by the encoder 10 and the image currently being decoded in the decoding process executed by the decoder 50.
[0017] The prediction of an image from other images in the coding video sequence 20 may also be called inter-prediction. For example, the image 261 * is the image 261 * may be encoded using temporal inter-prediction from an image 261’ belonging to a different access unit from the image 261. Thus, the image 261 * is the image 261 * belongs to the same layer as the image 261 * but may include a reference 32 to an image 2' belonging to a different access unit from the image 261. Additionally or alternatively, the image 261 * may be predicted using inter-layer (inter) prediction from an image in another layer, for example, a lower layer (lowered by a layer index that can be associated with each of the layers 24). For example, the image 261 * may include a reference 34 to an image 260' belonging to the same access unit but a different layer. In other words, in FIG. 1, the images 261', 260' may be examples of possible reference images for the image 261 currently being coded * of the current coding. Note that the prediction may be used to predict the coefficients of the image itself, such as the determination of the transform coefficients signaled in the video bitstream 14, or may be used for the prediction of the syntax elements used in the encoding of the image. For example, the image may be encoded using a motion vector that may represent the motion of the image content of the image 26 currently being coded with respect to a previously coded image or an image previous in image order. For example, the motion vector may be signaled in the video bitstream 14. The image 261 * * The motion vector may be predicted using time-motion vector prediction (TMVP) from a reference image, for example, either the above or alternative image.
[0018] Image 26 may be coded block by block. In other words, Image 26 may be subdivided into blocks and / or subblocks, for example, as described with respect to Figure 2.
[0019] The embodiments described herein may be implemented in the context of versatile video coding (VVC) or other video codecs.
[0020] The following describes several concepts and embodiments with reference to Figure 1 and the features described in relation to Figure 1. It is noted that features described in relation to an encoder, video bitstream, or decoder should be understood as also being descriptions of other entities within these entities. For example, a feature described as being present in a video data stream should be understood as a description of an encoder configured to encode this feature into a video bitstream, and a decoder or extractor configured to read this feature from the video bitstream. It is further noted that the inference of information based on instructions coded in the video bitstream may be performed equally on the encoder and decoder sides. It should also be noted that features described in relation to individual embodiments may be combined with each other at will.
[0021] Figure 2 shows an example of dividing one of the images 26 into block 74 and subblock 76. For example, image 26 may be pre-divided into tree root block 72, and then recursive subdivision may be performed on tree root block 72, as illustrated in tree root block 72', which is one of the tree root blocks 72 in Figure 2. That is, tree root block 72 may be subdivided into blocks, and those blocks may then be subdivided into subblocks, and so on. Recursive subdivision is sometimes called multi-tree partitioning. Tree root block 72 may be rectangular, or optionally quadratic. Tree root block 72 may also be called a coding tree unit (CTU).
[0022] For example, the motion vector (MV) described above may be determined and optionally transmitted in the video bitstream 14 in block units or sub-block units. In other words, the motion vector may refer to the entire block 74 or to a sub-block 76. For example, a motion vector may be determined for each block 74 of image 26. Alternatively, a motion vector may be determined for each of the sub-blocks 76 of block 74. In the example, whether one motion vector is determined for the entire block 74 or for each sub-block 76 of block 74 may differ from block to block. For example, all images of the coding video sequence 20 belonging to the same layer of layer 24 may be divided into tree root blocks of equal size.
[0023] Embodiments according to the first and second aspects may relate to time motion vector prediction.
[0024] Figure 3 shows a TMVP reference image determination module 53, hereafter referred to as TMVP module 53, according to an embodiment of the first aspect. This may also be optionally implemented in embodiments of the second and third aspects. The TMVP reference image determination module 53 may be implemented in a video decoder that supports TMVP, for example, a video decoder configured to decode a sequence of images coded from a data stream, for example, decoder 50 in Figure 1. The TMVP module 53 may also be implemented in a video encoder supporter TMVP, for example, a video encoder configured to decode a sequence of images into a data stream, for example, encoder 10 in Figure 1. Module 53 takes a predetermined image, for example, the image 26 currently being coded. * TMVP reference image 59 * This is a module for determining a given image 26 * The TMVP reference image for this purpose is the given image 26. * This is a reference image, and from this reference image, a predetermined image 26 * A predictor for the motion vector is selected.
[0025] TMVP Reference Image 59 * To determine this, the TMVP module 53 determines the first list 561 and the second list 562 of reference images from multiple previously decoded images. For example, multiple previously decoded images include image 26 of the previously decoded access unit 22. * For example, a given image 261 * It may include image 261' of Figure 1, and optionally includes a predetermined image 26 such as image 260' of Figure 1. * This may include previously decoded images of the same access unit as the given image 26. *It belongs to a lower layer, i.e., layer 240. Therefore, referring to the example in Figure 1, for example, images 260' and 261' may be part of the first list 561 of reference images. In other examples, these two images may be part of the second list of reference images. The first list 561 may optionally contain further reference images. In the example shown in Figure 3, the second list 562 of reference images is empty. Note that in general, either the first and second lists, or one or both, may be empty or not empty. The reference images from the first and second lists may be for interpretation of a given image 26'. The encoder 10 can determine the first and second lists and signal them in the video bitstream 14 so that the decoder 50 can derive them from the video bitstream. Alternatively, the decoder 50 may determine the first and second lists independently of explicit signal transmission, or at least partially independently of explicit signal transmission. For example, the encoder 10 may transmit the first and second lists in the video bitstream 14.
[0026] The TMVP reference image determination module 53 determines a predetermined image 26 * From the first list 561 and the second list 562 of reference images, one reference image is selected, for example, if at least one of the first and second lists of reference images is not empty, a predetermined image 26 * TMVP reference image 59 * It is specified as follows. For this purpose, module 53, for example, by TMVP list selection module 57, selects one of the first list 561 and the second list 562 of reference images from TMVP image list 56. * It is acceptable to decide this.
[0027] TMVP Image List 56 * To determine (57), the encoder 10 determines that the second list 562 of reference images is a predetermined image 26 * If empty, the first list 561 is TMVP image list 56 *It may be selected as such. Therefore, the decoder 50 determines that the second list 562 of reference images is a predetermined image 26 * If empty, TMVP image list 56 * It can be inferred that this is the first list of reference images 561. If the first list of reference images is empty and the second list of reference images is not empty, the encoder 10 will call TMVP image list 56 * The second list 562 may be selected. Therefore, the decoder 50, in this case, selects TMVP image list 56 * It can be inferred that this is the second list 562 of the reference images. The given image 26 * If neither the first nor the second list of reference images is empty, the TMVP list selection module 57 of the encoder 10 selects the TMVP image list 56 from the first list 561 and the second list 562. * The encoder 10 may select either the first or second list, and encode the list selector 58 into a video bitstream 14, and the list selector 58 may select either the first or second list, which is a predetermined image 26 * TMVP Image List 56 * This indicates whether or not. For example, list selector 58 may correspond to the following ph_collocated_from_l0_flag syntax element. Decoder 50 reads list selector 58 from video bitstream 14 and, accordingly, TMVP image list 56 * You may choose this option.
[0028] TMVP module 53 further includes TMVP image list 56 * TMVP reference image 59 * Perform selection 59. For example, encoder 10 selects TMVP image list 56. * TMVP reference image 59 * By transmitting the index of the selected TMVP reference image 59 in the video bitstream 14, *The encoder 10 may transmit the image selector 61 in the video bitstream 14. For example, the image selector 61 may correspond to the following ph_collocated_ref_idx syntax element. The decoder 50 reads the image selector 61 from the video bitstream 14 and, accordingly, the image list 56 * See TMVP reference image 59 * You may choose this option.
[0029] The encoder 10 and decoder 50 are connected to a predetermined image 26 * TMVP reference image 59 to predict the motion vector * You may use it.
[0030] For example, the TMVP list selection module 57 of the decoder 50 may initiate TMVP list selection by detecting whether the video bitstream 14 indicates a list indicator 58, and if so, the TMVP image list 56 as indicated by the list indicator 58. * You may select the following. If the video bitstream 14 does not indicate the list selector 58, the TMVP list selection module 57 will select the second list 562 of reference images, which is the given image 26 * If empty, the first list 561 is TMVP image list 56 * It may be selected as such. Otherwise, if the first list of reference images is empty for a given image, the TMVP list selection module 57 selects the second list of reference images 562 from the TMVP image list 56 * It may be selected as such. Alternatively, if the second list 562 is empty for a given image, the TMVP list selection module 57 will select the second list 562 as the TMVP image list 56 if the second list 562 is not empty. * You may choose this option.
[0031] In other words, an embodiment of the first aspect may consider interference of the list selector 58, for example PH_collocated_from_L0, when the first list 561 is empty but the second list 562 is not. The first list 561 may be called L0 and the second list 562 may be called L1.
[0032] In other words, the current VVC specification uses two syntax elements, ph_collocated_from_l0_flag and ph_collocated_ref_idx, to control the image used for time-motion vector prediction (TMVP) (or subblock TMVP). The first syntax element specifies whether the image used for TMVP is selected from L0 or L1, and the second syntax element specifies which image from the selected list is used. These syntax elements reside in either the image header or the slice header. In the latter case, the prefix is "sh_" instead of "ph_". An example image header is shown in Table 1.
[0033] [Table 1]
[0034] A value of ph_collocated_from_l0_flag equal to 1 specifies that the collocated image used for time motion vector prediction is derived from reference image list 0. A value of ph_collocated_from_l0_flag equal to 0 specifies that the collocated image used for time motion vector prediction is derived from reference image list 1. If both ph_temporal_mvp_enabled_flag and pps_rpl_info_in_ph_flag are equal to 1 and num_ref_entries[1][RplsIdx[1]] is equal to 0, then the value of ph_collocated_from_l0_flag is assumed to be equal to 1. ph_collocated_ref_idx specifies the reference index of the collocated image used for time motion vector prediction. If ph_collocated_from_l0_flag is equal to 1, then ph_collocated_ref_idx refers to an entry in reference image list 0, and the value of ph_collocated_ref_idx is in the range of 0 to num_ref_entries[0][RplsIdx[0]]-1. If ph_collocated_from_l0_flag is equal to 0, ph_collocated_ref_idx refers to an entry in reference image list 1, and the value of ph_collocated_ref_idx is in the range of 0 to num_ref_entries[1][RplsIdx[1]]-1. If it does not exist, the value of ph_collocated_ref_idx is assumed to be equal to 0.
[0035] There is a specific case that the decoder must consider, which is when the reference image list is empty, i.e., when both L0 and L1 have zero entries. If either the L0 or L1 list is empty, the specification currently does not signal ph_collocated_from_l0_flag and assumes that the value of ph_collocated_from_l0_flag is equal to 1, and the collocated image is considered to be in L0. However, this comes with efficiency issues. In fact, with respect to the states of L0 and L1, the following scenarios are possible: L0 is empty and L1 is also empty: ph_collocated_from_l0_flag is estimated to be 1 => OK L0 is not empty and L1 is empty: ph_collocated_from_l0_flag is presumed to be 1 => OK. If neither L0 nor L1 is empty, the signal ph_collocated_from_l0_flag is transmitted as => OK. • L0 is empty and L1 is not empty: ph_collocated_from_l0_flag is presumed to be 1 => NOT OK.
[0036] If L0 is empty but L1 is not, the estimated value of ph_collocated_from_l0_flag is equal to 1, which means that TMVP (or subblock TMVP) is still available if the image in L1 is selected, but TMVP will not be used, leading to a loss of efficiency.
[0037] Therefore, in one embodiment, the decoder (or encoder) determines the value of ph_collocated_from_l0_flag depending on which list is empty, for example, as described above with respect to the TMVP list selection module 57 in Figure 3. If ph_collocated_from_l0_flag is equal to 1, it specifies that the collocated image used for time motion vector prediction is derived from reference image list 0. If ph_collocated_from_l0_flag is equal to 0, it specifies that the collocated image used for time motion vector prediction is derived from reference image list 1. If both ph_temporal_mvp_enabled_flag and pps_rpl_info_in_ph_flag are equal to 1, the following occurs: If num_ref_entries[1][RplsIdx[1]] is equal to 0, then the value of ph_collocated_from_l0_flag is presumed to be equal to 1. Otherwise, the value of ph_collocated_from_l0_flag is assumed to be equal to 0.
[0038] As an alternative to determining the TMVP image list based on which of the lists is empty, in another embodiment there is a bitstream constraint that if L0 is empty, then L1 must also be empty. This is due to a bitstream constraint or syntax prohibition (see Table 2).
[0039] Therefore, according to the alternative embodiment of the TMVP module 53 in Figure 3, the encoder 10, if the second list 562 is empty, uses the TMVP reference image 59 from the first list 561. * Select and if the second list 562 is not empty, select TMVP reference image 59 from either the first list 561 or the second list 562. * Select the TMVP list selection module 57 mentioned above. If the second list 562 is empty, select the TMVP image list 56 * List 56 * You may select 1, and if the second list 562 is not empty, you may select either the first list or the second list.
[0040] Therefore, when the decoder 50 determines the list of reference images (55), if the first list 561 is empty, it can infer that the second list 562 is also empty. Thus, in the TMVP list selection 57, the decoder 50 reads the list selector 58 from the video bitstream 14, and if neither the first list 561 nor the second list 562 is empty, it selects the TVMP reference image 59 according to the list selector 58. * You may select the following. If the first list 561 and the second list 562 are not empty, the decoder 50 selects the first list 561 as the TMVP image list 56 * It may be selected as such, otherwise the decoder 50 selects the first list 561 as the TMVP image list 56 * You may choose this option.
[0041] According to one embodiment, the decoder 50 may perform list determination (55) by reading from the video bitstream 14 information on how to populate a first list 561 from a plurality of previous decoder images for a first list of reference images. If the decoder 50 does not assume that a second list 562 is empty, the decoder 50 may read from the video bitstream 14 information on how to populate a second list 562.
[0042] According to the latter embodiment, if the first list 561 is empty, the encoder 10 and decoder 50, in accordance with the decoder's estimation that the second list 562 is empty, will use a predetermined image 26 if the first list of reference images is empty, without using TMVP. * You may encode it.
[0043] As mentioned above, bitstream constraints may instead be implemented in the form of syntax, for example, by constructing a list of reference images. An example implementation is shown in Table 2.
[0044] [Table 2]
[0045] Therefore, if ph_collocated_from_l0_flag is presumed to be 1, both lists are empty, or only list L1 is empty and there is one non-empty list, that list will be used for TMVP.
[0046] Therefore, according to a further alternative embodiment of the TMVP module 53 in Figure 3, the encoder 10 is obtained from the first list 561 of the TMVP reference image 56 if the second list 572 is empty. * Select and if the second list 562 is not empty, select TMVP reference image 56 from either the first list 561 or the second list 562. * By selecting this option, you may perform TMVP list selection 57.
[0047] The following describes an embodiment according to the second aspect, with reference to Figures 1 to 3 and Figure 4.
[0048] Figure 4 is the first image 261 *For example, an example is shown of subdividing the image currently being coded into a tree root block 721. The tree root block 721 is recursively subdivided into block 74 and subblock 76, for example, as described with respect to Figure 2. Figure 4 further shows image 260' of Figure 1, which may be an interlayer reference image of a second image 260', for example, the first image 261'. The second image 260' is subdivided into a tree root block 720. A first example of subdivision into tree root block 720 is shown by a dashed line. A dotted line further shows an alternative second example for subdividing the second image 260' into a tree root block 720', where the tree root block 720' in the second example is smaller than the tree root block 720 in the first example. According to the first example of subdivision, the tree root block 720 is the same as the first image 261 * It has the same size as tree root block 721. In the second example of subdivision, tree root block 720' is the same size as in the first image 261. * It is smaller than the tree root block size of 721.
[0049] Image 261 (first image) * For the TMVP, the encoder 10 and decoder 50 may determine one or more MV candidates. For this purpose, see Figure 261 * One or more MV candidates may be determined from each of one or more reference images. For example, the encoder 10 and decoder 50 may determine one or more MV candidates from one or more reference images from one or more lists of reference images, for example, two lists, for example, list 561 and list 562 described with respect to Figure 3. First image 261 * In addition to the reference image, there may be interlayer reference images such as a second image 260' that are temporally collated (at the same location) to the first image, i.e., belong to the same access unit 22, as shown in Figure 1. As explained with respect to Figure 2, the image 26 may be coded in block units or subblock units, and the encoder 10 is currently coding the image 261 * Tree root block 721 currently being coded* Block 74 currently being coded * You may decide on one MV, or instead, block 74 which is currently being coded. * One MV may be determined for each subblock 76. The subblock currently being coded is reference numeral 76. * The encoder 10, and optionally the decoder 50, may be referenced using the block 74 currently being coded. * Or subblock 76 currently being coded * For this, one or more MV candidates may be determined from a single reference image, for example, the second image 260'. One or more MV candidates may be determined from different positions or locations within the reference image. For example, block 74 currently being coded * Or subblock 76 currently being coded * Regarding the MV candidate in the lower right, the MV of the reference image located at reference position 71' in the reference image may be selected as the MV candidate, and the reference position 71' of the lower right MV candidate is collated to position 71 in the lower right of the image 261' currently being coded (they are in the same position). For example, block 74 currently being coded * Or subblock 76 currently being coded * The lower right position 71 is block 74 * or subblock 76 * It may be in a position adjacent to it in the downward and rightward directions.
[0050] Figure 4 shows the first example of subdividing the second image 260' into the tree root block 720, and the block 74 currently being coded in the first image 261'. * The collated block is reference numeral 714. * As shown in Figure 4, the size of the tree root block 720 is shown in the first image 261. * In this first example of subdivision equal to the size of the tree root block 721, reference position 71' is block 714 * It is located in block 714' in the lower right, and block 714' is in the first image 261 *The currently-coded block 74 * The collocated block 714 * Within the same tree root block 720 * For example, for the block 74 of the first image 261 * One MV candidate for coding of the block 74 * May be the MV of the block 714'.
[0051] The first image 261 * In a second example where the reference image 260' is subdivided into a tree root block 720' smaller than the tree root block 721 of the first image 261, the reference position 71 may be located outside the collocated tree root block 720'. Specifically, the reference position 71 may be located outside the row of the tree root block in which the collocated tree root block 720' of the currently-coded block 74 * Is located. Thus, from the tree root block in which the reference position 71 is located, for the currently-coded block 74 * Or the currently-coded sub-block 76 * To use the MV as an MV candidate, the encoder 10 and the decoder 50 may need to hold one or more tree root blocks in the image buffer that exceed the currently-coded tree root block of the second image 260' or exceed the current row of the currently-coded tree root block. Thus, in this second example of tree root block subdivision, using the reference position 71 for MV prediction may involve inefficient buffer usage.
[0052] In other words, as described with respect to Figure 3, the current specification uses two syntax elements, namely ph_collocated_from_l0_flag and ph_collocated_ref_idx, which control the image used for the TMVP (or subblock TMVP), where the latter specifies which entry in each list is used as the reference image, as described, for example, with respect to Figure 3. When the TMVP (e.g., the TMVP of one MV in block 74 in Figure 2) or subblock TMVP (e.g., the TMVP of one MV in subblock 76 in Figure 2) is used in the bitstream, the current image 26 * The motion vector (MV) of a block or subblock is derived from the MV of the reference image indicated by ph_collocated_ref_idx, and that MV is added to a list of candidates that can be selected as predictors for the motion vector actually used for the block / subblock. This MV prediction is based on 26 images of the current image. * Select from MVs at different locations depending on the boundary of the largest block (CTU) (for example, tree root block 72 in Figure 2). In most scenarios, the reference images will show the same CTU boundary.
[0053] More specifically, in the case of TMVP, (for example, block 74 currently being coded) * Or the lower right TMVP MV candidate 71 (of subblock 76) crosses the CTU row boundary of the current block (i.e., tree root block 74 in Figure 4). **If, regarding block 74 or subblock 76, it is located beyond the row boundary of the tree root block 72 to which block 74 or subblock 76 belongs, or does not exist (for example, because it crosses the image boundary or is intracoded without using motion vectors), the TMVP candidate is not taken from the lower right block, and instead, an alternative (collocated) MV candidate for the CTU row is derived from the collocated block (of block 74 or subblock 76 to which the MV candidate is to be determined). Furthermore, if the subblock TMVP MV candidate is taken from a location outside the collocated CTU, the MV used to determine the location of the subblock TMVP MV is modified (clipped) during derivation to point to a location from which the subblock TMVP MV candidate is taken from a location to which the collocated CTU belongs. Note that in the case of subblock TMVP, the MV is taken first (for example, from the spatial candidate first), and that MV is used to identify the location of the block used to select the subblock TMVP candidate within the collocated image. If the time MV used to determine the location points to a location outside the collocated CTU, the MV is clipped so that the subblock TMVP is taken from a location inside the collocated CTU.
[0054] However, there are cases where the reference image does not have the same CTU size, for example, the tree root block 721' in Figure 4, and therefore does not have a boundary. The TMVP or subblock TMVP process described is not clear when the CTU size changes. When applied in this way, state-of-the-art techniques select non-buffer-friendly TMVP and subblock TMVP candidates located outside the current CTU (row) or the CTU (row) of the reference image, for example, at position 71'. This aspect of the invention, which aims to guide MV candidate selection accordingly or prevent such cases from occurring, relieves the embodiment of this burden. The scenario described occurs in a quality-scalable multilayer bitstream with two (or more) layers. Here, as shown on the left side of Figure 5, one layer depends on other layers, and the dependent layer has a larger CTU size than the reference layer. In such a case, the decoder needs to fetch associated MV candidates from four smaller CTUs of the reference image, and these CTUs do not necessarily occupy contiguous memory regions that negatively impact memory bandwidth at access. On the other hand, on the right side of Figure 5, the CTU size of the reference image is larger than that of the dependent layer. In such cases, when decoding the current block of the dependent layer, the data is not scattered among the related data of the reference layer.
[0055] As explained, if the CTU sizes of the two layers (parameters for each SPS) are different, the CTU boundaries of the current image and the reference image (ILRP in this scenario) will not be aligned, but the TMVP and subblock TMVP will still be activated and not prohibited from being used by the encoder.
[0056] According to an embodiment of the second aspect, the encoder 10, for example, the encoder 10 in Figure 1, is for layer video coding, as described with respect to Figure 1, for example, and encodes the images 26 of the video 20 into a multilayer data stream 14 (or multilayer video bitstream 14) by pre-dividing each image into one or more tree root blocks 72, as described with respect to Figures 2 and 4, and performing recursive block division on each tree root block 72 of each image 26, thereby subdividing the image 26 into units of blocks 74. Each image 26 is associated with one of the layers 24, as described with respect to Figure 1, for example. The size of the tree root block 72 may be equal for all images belonging to the same layer 24. The encoder 10 according to an embodiment of the second aspect, for example, as described with respect to Figure 3, the first image 261 of the first layer 241 * Regarding this (see Figure 1), a list of reference images, e.g., list 561 or list 562, is populated from multiple previously coded images. As described with respect to Figure 4, the list of reference images is the first image 261 * The upstream multilayer data stream 14 may include images of the same layer and different timestamps on which the image is encoded, for example, image 261' in Figure 1. Furthermore, the list of reference images may include one or more second images, the second image belonging to a different layer, for example, a second layer 240 and the first image 261'. * For example, the second image 260' is temporally aligned with the first image 261. * The size of the tree root block of the first layer 241 to which the first image belongs, and the size of a different layer, for example, the second layer 240 to which the second image belongs, may be transmitted in the multilayer video bitstream 14. The encoder 10 and decoder 50 use motion compensation prediction to transmit the first image 261 * A list of reference images may be used to predict the interpretation blocks and predict the motion vectors of the interpretation blocks.
[0057] According to the first embodiment of the second aspect, the size of the tree root block in the second image 260' is the same as in the first image 260 * It is equal to or an integer multiple of the size of the tree root block. For example, a video encoder may populate a list of reference images 56, or two lists of reference images 561 and 562 as described in Figure 3, and as a result, for each reference image in the list of reference images, the size of the tree root block is equal to the size of the first image 261 * It is equal to the size of the tree root block, or an integer multiple thereof.
[0058] In other words, in the example of the first embodiment, the use of TMVP and subblock TMVP is disabled by imposing constraints on the ILRP of the current image's reference image list, as follows: -The following constraints apply to the images referenced by each ILRP entry if each ILRP entry exists in RefPicList[0] or RefPicList[1] of the current image slice. The image is assumed to be in the same AU as the current image. The image is assumed to exist in the DPB. The image shall have a nuh_layer_id refPicLayerId that is smaller than the nuh_layer_id of the current image. The image must have a value of sps_log2_ctu_size_minus5 that is the same as or greater than the current image. o One of the following constraints applies: The image is assumed to be an IRAP image. The image shall have a TemporalId less than or equal to Max(0,vps_max_tid_il_ref_pics_plus1[currLayerIdx][refLayerIdx]-1). Here, currLayerIdx and refLayerIdx are equal to GeneralLayerIdx[nuh_layer_id] and GeneralLayerIdx[refpicLayerId], respectively.
[0059] The worst-case scenario for the above problem occurs when the CTU size of the reference image is small. This is because the memory bandwidth requirements for the TMVP and subblock TMVP become high. When the CTU size of the reference image is small, obtaining a candidate for the TMVP or subblock TMVP is not so critical, so the above embodiment applies only when the CTU size of the reference image is small. Therefore, as in the above embodiment, there is no problem when the CTU size of the reference image is larger than the CTU size of the current image.
[0060] Nevertheless, according to an example of the first embodiment of the second aspect, the constraint is that for each interlayer reference image in the list of reference images, i.e., each of the second images (and, for example, for all reference images, since images in the same layer may have the same size tree root block), the size of the tree root block is the size of the first image 261 * It is required that the size of the tree root block be equal to, or an integer multiple of, that size.
[0061] Therefore, in the example of the first embodiment, the use of TMVP and subblock TMVP is disabled by imposing the following constraints on the ILRP of the current image's reference image list. The following constraints apply to the images referenced by each ILRP entry, provided that each ILRP entry exists in RefPicList[0] or RefPicList[1] of the current image slice. The image is assumed to be in the same AU as the current image. The image is assumed to exist in the DPB. The image shall have a nuh_layer_id refPicLayerId that is smaller than the nuh_layer_id of the current image. The image shall have the same sps_log2_ctu_size_minus5 value as the current image (i.e., the image is the same as image 261 currently being coded). * (Assuming it has the same tree root block size as [another tree].) o One of the following constraints applies: The image is assumed to be an IRAP image. The image shall have a TemporalId less than or equal to Max(0,vps_max_tid_il_ref_pics_plus1[currLayerIdx][refLayerIdx]-1). Here, currLayerIdx and refLayerIdx are equal to GeneralLayerIdx[nuh_layer_id] and GeneralLayerIdx[refpicLayerId], respectively.
[0062] However, this embodiment has a very restrictive constraint that prevents any form of prediction, such as sample prediction, because it does not recognize different CTU sizes in the dependent and reference layers.
[0063] Therefore, according to the second embodiment of the second aspect, the list of reference images may include images of a second type of image that does not necessarily have a size smaller than or equal to the tree root block 720, i.e., images of different layers such as the second layer 240, but the list of reference images may include any of the second images (images of the second layer 240). According to this embodiment, the encoder 10 includes the first image 261 * For example, as described with respect to Figure 3, one image is specified as the TVMP reference image from the list of reference images. According to this embodiment, the encoder 10 specifies the TVMP reference image so that it is not the second image 260, and the size of its tree root block 720 is the same as the first image 261. * It is smaller than the size of the tree root block 721. In other words, if encoder 10 selects one of the second images 260 as the TVMP reference image, i.e., encoder 10 selects images from different layers such as layer 240 as the first image 261. * When selected as the TVMP reference image, the TVMP reference image is the first image 261. *It has a tree root block of the same or larger size. The encoder 10 may signal a pointer in the multilayer video data stream 14 that identifies a TVMP reference image from an image in a list of reference images. The encoder 10 has a first image 261 * A TVMP reference image is used to predict the motion vector of the interpretation block.
[0064] Instead of restricting the popularity of the list, restricting the specification of TMVP reference images from the list of reference images allows other prediction tools to use a second image 260 with a smaller tree root block size than the first image.
[0065] For example, the list of reference images may be one of the first list 561 and the second list 562, as described with respect to Figure 3. However, the method for selecting the TMVP reference list does not necessarily have to follow the method described with respect to Figure 3, and may be carried out in a different way, for example, as described for the latest technology.
[0066] For example, if the encoder 10 is a second image, i.e., an image on a different layer from the image currently being coded, it may select a TVMP reference image such that it does not satisfy any of the criteria in the following set of criteria. - The size of the tree root block 720 in the second image 260 is the same as in the first image 261 * It is smaller than the size of the tree root block 721. -The size of the TVMP reference image is 261 for the first image. * It is a different size. - The scaling window for the TVMP image, which is used for scaling and offsetting motion vectors, is shown in the first image 261. * This is different from the scaling window. -The sub-image subdivision of TVMP reference image 260' is the first image 261 * This is different from image subdivision.
[0067] Encoder 10 is shown in the first image 261 * When interpreting an interpretation block, for each interpretation, activate one or more sets of interpretation refinement tools depending on the reference image in the list of reference images from which each interpretation block is interpreted, provided that it satisfies one of the above set of criteria. For example, the set of interpretation refinement tools may include one or more of TVMP, PROF, wraparound, VDOF, and DVMR.
[0068] In other words, according to the example of the second embodiment, the problem is solved by imposing constraints on the syntax element sh_collocated_ref_idx, which indicates the reference image used for the TVMP and subblock TVMP. sh_collocated_ref_idx specifies the reference index of the collocated image used for time motion vector prediction. If sh_slice_type is equal to P, or if sh_slice_type is equal to B and sh_collocated_from_l0_flag is equal to 1, sh_collocated_ref_idx refers to an entry in reference image list 0, and the value of sh_collocated_ref_idx is in the range of 0 to NumRefIdxActive[0]-1. If sh_slice_type is equal to B and sh_collocated_from_l0_flag is equal to 0, then sh_collocated_ref_idx refers to an entry in reference image list 1, and the value of sh_collocated_ref_idx is in the range of 0 to NumRefIdxActive[1]-1. If sh_collocated_ref_idx does not exist, the following applies: -If pps_rpl_info_in_ph_flag is equal to 1, the value of sh_collocated_ref_idx is presumed to be equal to ph_collocated_ref_idx. -If not (pps_rpl_info_in_ph_flag is equal to 0), the value of sh_collocated_ref_idx is presumed to be equal to 0. Set colPicList to equal sh_collocated_from_10_flag?0:1. The image referenced by sh_collocated_ref_idx is the same for all non-I slices of the coded image, the value of RprConstraintsActiveFlag[colPicList][sh_collocated_ref_idx] is equal to 0, and the value of sps_log2_ctu_size_minus5 of the image referenced by sh_collocated_ref_idx is greater than or equal to the value of sps_log2_ctu_size_minus5 of the current image, which is a bitstream conformance requirement. Note - Under the above constraints, the collated image must have the same spatial resolution, the same scaling window offset, and the same or smaller CTU size as the current image.
[0069] Again, the aforementioned example can prevent the case where the selected interlayer reference image has a smaller tree root block size than the first image. In another example, encoder 10 uses the first image 261 * The TVMP reference image is selected as image 261. * You may specify that the tree root block size is equal to the size of the tree root block 721.
[0070] Therefore, another exemplary embodiment is as follows: sh_collocated_ref_idx specifies the reference index of the collocated image used for time motion vector prediction. If sh_slice_type is equal to P, or if sh_slice_type is equal to B and sh_collocated_from_l0_flag is equal to 1, sh_collocated_ref_idx refers to an entry in reference image list 0, and the value of sh_collocated_ref_idx is in the range of 0 to NumRefIdxActive[0]-1. If sh_slice_type is equal to B and sh_collocated_from_l0_flag is equal to 0, then sh_collocated_ref_idx refers to an entry in reference image list 1, and the value of sh_collocated_ref_idx is in the range of 0 to NumRefIdxActive[1]-1. If sh_collocated_ref_idx does not exist, the following applies: - If pps_rpl_info_in_ph_flag is equal to 1, the value of sh_collocated_ref_idx is presumed to be equal to ph_collocated_ref_idx. - If not (pps_rpl_info_in_ph_flag is equal to 0), the value of sh_collocated_ref_idx is presumed to be equal to 0. Set colPicList to equal sh_collocated_from_10_flag?0:1. The image referenced by sh_collocated_ref_idx is the same for all non-I slices of the coded image, the value of RprConstraintsActiveFlag[colPicList][sh_collocated_ref_idx] is equal to 0, and the value of sps_log2_ctu_size_minus5 of the image referenced by sh_collocated_ref_idx is equal to the value of sps_log2_ctu_size_minus5 of the current image, which is a bitstream conformance requirement. Note - Under the above constraints, the collated image must have the same spatial resolution, the same scaling window offset, and the same CTU size as the current image.
[0071] According to a third embodiment of the second aspect, the encoder 10 and decoder 50, depending on the reference image used to interpret each interpret block that satisfies any of the above reference sets, a predetermined image 261 * When predicting interpretation blocks, one or more of the above set of interpretation refinement tools may be employed. Herein, this constraint for using interpretation refinement tools is that the reference image is coded in image 26. * This is not necessarily limited to the multi-layer case where the interlayer reference image is multilayer.
[0072] In this example, the encoder 10 and decoder 50 may derive a list of reference images from which a reference image is selected, as described with respect to Figure 3, but other approaches to signaling or selecting reference images may also be possible.
[0073] In other words, according to the embodiment, the encoder 10 and decoder 50, which code the image in units of blocks obtained by recursive block division of the tree root block as described above (i.e., encode in the case of encoder 10, and decode in the case of decoder 50), currently code the image 26 * When predicting interpretation blocks, one or more sets of interpretation refinement tools may be used depending on whether any of the following set of criteria are met. -Image 26 currently being coded * The size of the tree root block is smaller than the size of the tree root block in the reference image. - The size of the reference image is image 26, which is currently being coded. * It is equal to the size of [this]. -For example, the scaling window of a reference image used for scaling and offsetting a motion vector is equal to the scaling window of a given image, but differs with respect to the offset of the scaling window. - The sub-image subdivision of a reference image is equal to the sub-image subdivision of a given image. That is, for example, the sub-image subdivision differs in terms of the number of sub-images in the subdivision.
[0074] In the example, the reference set is defined as the size of the tree root block of the reference image being a given image 26 * This may include being equal to the size of the tree root block.
[0075] In one example, during interpretation of an interpretation block, the encoder 10 and decoder 50 may selectively activate a set of interpretation refinement tools if all of a subset of the reference set are met. In another example, the encoder 10 and decoder 50 may activate an interpretation refinement tool if all of the reference set are met.
[0076] For example, according to the third embodiment, the constraint is represented by the derived variable RprConstraintsActiveFlag[refPicutre][currentPic], which is derived by comparing the characteristics of the current image and the reference image (image size, scaling window offset, number of sub-images, etc.). This variable is used to impose a constraint on the specified ph_collocated_ref_idx of the image header in the slice header sh_collocated_ref_idx. In this embodiment, the CTU size (sps_log2_ctu_size_minus5) in the reference image and the current image is incorporated into each derivation of RprConstraintsActiveFlag[refPicutre][currentPic], and as a result, if the CTU size of the reference image is larger, the CTU size of the current image RprConstraintsActiveFlag[refPicutre][currentPic] is derived as 1.
[0077] Similarly, however the criterion is that the size of the tree root block of the reference image is given image 26 *If the constraint includes being equal to the size of the tree root block, the constraint may be represented by the derived variable RprConstraintsActiveFlag[refPicutre][currentPic], which is derived by comparing the characteristics of the current image and the reference image (image size, scaling window offset, number of subimages, etc.). This variable is used to impose a constraint on the specified ph_collocated_ref_idx of the image header in the slice header sh_collocated_ref_idx. In this embodiment, the size of the CTU in the reference image and the current image (sps_log2_ctu_size_minus5) is incorporated into each derivation of RprConstraintsActiveFlag[refPicutre][currentPic], and as a result, if the CTU sizes are different, RprConstraintsActiveFlag[refPicutre][currentPic] is derived as 1.
[0078] In such cases, this embodiment does not permit tools such as PROF, Wraparound, BDOF, and DMVR, as it is undesirable to allow different CTU sizes when such tools are used.
[0079] The above explanation focused only on TMVP and subblock TMVP, but further issues have been identified that apply when different CTUs are used in different layers. For example, TMVP and subblock TMVP are permitted for images with different CTU sizes, with associated "disadvantages," but problems arise when combined with subimages. In fact, when subimages are used and used together with layer coding, there are a set of constraints required to ensure that the subimage grid is aligned. This is done for layers that have subimages in their dependency tree.
[0080] Figure 6 shows an example of sub-image subdivision of image 261 into two sub-images 281' and 281''. Image 261 may belong to a first layer, such as layer 241, and may belong to an access unit 22 as described with respect to Figure 1. In other words, image 261 in Figure 6 may be one of the images 261 in Figure 1. Image 261 may depend on a reference image 260, for example, the second layer 240 in Figure 1. In other words, image 260 may be an interlayer reference image of image 261, and the layer to which image 260 belongs may be called the reference layer of the layer to which image 261 belongs. For example, all images belonging to the same coding video sequence 20 and belonging to the same layer may undergo the same sub-image subdivision, i.e., they may be subdivided into the same number of sub-images of equal size. Note also that all images of a coding video sequence 20 belonging to the same layer 24 may be divided into tree root blocks of equal size. The encoder 10 may be configured to subdivide the image 261 of the first layer 241 and the image 260 belonging to the reference layer of the first layer 261, i.e., the second layer 240, into two or more common sub-images 28. Such subdivision into two or more common sub-images is shown in Figure 6, where the image 260 of the reference layer 240 is subdivided into two sub-images 280' and 280''. The encoder 10 may encode the sub-images 28' and 28'', i.e., the sub-images belonging to a common layer, for example, layer 241 or layer 240, independently of each other. That is, the encoder 10 may encode independently coded sub-images without spatial prediction of one part of sub-image 28' by another sub-image 28''. For example, the encoder 10 may apply sample padding to the boundary between images 28' and 28'' so that the coding of the two sub-images is independent of each other. Furthermore, the encoder 10 may clip the motion vector at the boundary between sub-images 28' and 28''. Optionally, the encoder 10 may signal in the video data stream 14 that the sub-images of the layer are coded independently, for example, by signaling sps_subpic_treated_as_pic_flag=1.The encoder 10 may encode the subimages 28 into the video bitstream 14 in units of blocks 74, 76, where blocks 74, 76 result from the division of one or more tree root blocks 72 into which the subimages 28 are divided, for example, as described with respect to Figures 2 and 4 with respect to the image 26. The encoder 10 may divide the subimages 281, 280 of the first layer 241 of the image 26 and the reference layer 240 of the first layer 241 into tree root blocks of equal size. In other words, if the image 261 of the first layer 241 and the reference image 260 of the reference layer 240 are subdivided into two or more subimages with respect to a common number, the encoder 10 may subdivide the subimages 281, 280 such that the tree root blocks of the subimage 281 of the first layer 241 and the subimage 280 of the reference layer 240 are the same size.
[0081] In other words, in another alternative embodiment, the problem is solved only for the case of independent subimages by extending the subimage-related constraints as follows: If sps_subpic_treated_as_pic_flag[i] is equal to 1, it specifies that the i-th subimage of each coded image in the CLVS is treated as an image in the decoding process, excluding the in-loop filtering operation. If sps_subpic_treated_as_pic_flag[i] is equal to 0, it specifies that the i-th subimage of each coded image in the CLVS is not treated as an image in the decoding process, excluding the in-loop filtering operation. If it does not exist, the value of sps_subpic_treated_as_pic_flag[i] is assumed to be equal to 1. If sps_num_subpics_minus1 is greater than 0 and sps_subpic_treated_as_pic_flag[i] is equal to 1, then for each CLVS in the current layer that references an SPS, if targetAuSet is all AUs, starting from the AU containing the first image of the CLVS in decode order and ending with the AU containing the last image of the CLVS in decode order, then for targetLayerSet, which consists of the current layer and all layers that reference the current layer, all of the following conditions must be true: -For each AU in targetAuSet, all images in the layers within targetLayerSet shall have the same values for pps_pic_width_in_luma_samples and pps_pic_height_in_luma_samples. -All SPS referenced by layers in targetLayerSet shall have the same value of sps_num_subpics_minus1, and the same values of sps_subpic_ctu_top_left_x[j], sps_subpic_ctu_top_left_y[j], sps_subpic_width_minus1[j], sps_subpic_height_minus1[j], and sps_subpic_treated_as_pic_flag[j], respectively, in the range of 0 to sps_num_subpics_minus1 and for each value of j in sps_log2_ctu_size_minus5. -For each AU in targetAuSet, all images in the layers of targetLayerSet shall have the same value of SubpicIdVal[j] for each value of j in the range of 0 to sps_num_subpics_minus1.
[0082] According to the embodiment, the encoder 10 and decoder 50 are configured to process a block, for example, the image 26 currently being coded on the first layer 241. * Block 74 currently being coded *Or subblock 76 currently being coded * This may be interpreted using coding parameters such as the motion vector of the corresponding block in the image of reference layer 240. For example, block 714 * This is block 74 in Figure 4. * It may be the corresponding block. In other words, the corresponding block may collate (be in the same position as) the block currently being coded, for example, within each image, i.e., within each reference image, and may be, for example, the corresponding block or referenced by the motion vector of the block currently being coded. The encoder 10 and decoder 50 may use the coding parameters of the corresponding block to predict coding parameters such as the motion vector of the block currently being coded.
[0083] By subdividing the subimages of the reference layer 240 into tree root blocks having the same size as the tree root block of the first layer that depends on the reference layer, it is possible to ensure that the subimages 28 are coded independently of each other. In this regard, the same considerations as those described with respect to Figure 4 apply to the boundaries of the tree root blocks. In other words, the considerations with respect to Figure 4 that illustrate the advantages of equal-sized tree root blocks for the subdivision of image 26 may also apply to the subdivision of subimages 28.
[0084] Figure 7 shows an example of a decoder 50 and a video bitstream 14 according to an embodiment of a third aspect of the present invention. The decoder 50 and video bitstream 14 may optionally correspond to the decoder 50 and video bitstream 14 of Figure 1. According to the embodiment of the third aspect, the video bitstream 14 is a multilayer video bitstream including access units 22, each access unit 22 containing one or more images 26 of a coded video sequence 20 coded into the video bitstream 14. For example, the description of the layers 24 and access units 22 in Figure 1 may apply. Thus, each image 26 belongs to one of the layers 24, and the association between the image and the layer is indicated by a subscript. That is, image 261 belongs to the first layer 241, and image 260 belongs to the second layer 240. In Figure 7, the images 26 and the access units 22 to which the images belong are shown according to the image order of the coded video sequence 20, i.e., their presentation order. Each access unit 22 belongs to a time sublayer of a set of time sublayers. For example, in Figure 7, access unit 221, to which the image with superscript 1 belongs, belongs to the first time sublayer, and access unit 222, to which the image referenced using a reference code with superscript 2 belongs, belongs to the second time sublayer. As shown in Figure 7, access units belonging to different time sublayers do not necessarily have the same set of images in layer 24. For example, in Figure 7, access unit 221 may have images from the first layer 241 and the second layer 240, while access unit 222 may have images only from the first layer 241. Time sublayers may be hierarchically ordered and may be indexed by an index representing the hierarchical order. In other words, the hierarchical order of time sublayers may define the highest and lowest time sublayers within a set of time sublayers. Note that the video bitstream 14 may optionally have further time sublayers and / or further layers.
[0085] According to an embodiment of the third aspect, the video bitstream 14 encodes within an output layer set instruction 81 indicating one or more output layer sets 83. The OLS instruction 81 indicates, for OLS 83, a subset of layers 24 of the multilayer video bitstream 14 belonging to the OLS (not necessarily an appropriate subset; i.e., OLS may indicate that all layers of the multilayer audio bitstream 14 belong to the OLS). For example, OLS may be an instruction for a sub-bitstream that is extractable or decodeable (not necessarily appropriate) from the video bitstream 14, and the sub-bitstream includes a subset of layers 24. For example, by extracting or decoding a subset of layers of the video bitstream 14, the coded video sequence to be decoded may be scalable in quality and therefore in bitrate.
[0086] According to an embodiment of the third aspect, the video bitstream 14 further includes a video parameter set (VPS) 91. The VPS 91 includes one or more bitstream compatibility sets 86, for example, a virtual reference decoder (HRD) parameter set. The video parameter set 91 further includes one or more buffer requirement sets 84, for example, a decoded image buffer (DPB) parameter set. The video parameter set 91 further includes one or more decoder requirement sets 82, for example, a profile layer level parameter set (PTL set). Each of the bitstream compatibility sets 86, buffer requirement sets 84, and decoder requirement sets 82 is associated with each time subset instruction 96, 94, and 92 indicated in the video parameter set 91, respectively. A constraint on the maximum time sublayer for each of the parameter sets, i.e., the bitstream compatibility sets 86, buffer requirement sets 84, and decoder requirement sets 82, may represent an upper limit on the number of time sublayers referenced by the parameters of each parameter set. In other words, the parameters transmitted by the parameter sets may be valid for (not necessarily appropriate) subsequences of the coded video sequence 20, defined by a set of layers and a set of time sublayers, or for subbitstreams of the video bitstream 14, and the constraint on the maximum time sublayer of each parameter set indicates the maximum time sublayer of the subbitstream or subsequence referenced by each parameter set.
[0087] According to the first embodiment of the third aspect, the decoder 50 may be configured to receive a maximum time sublayer instruction 99. The maximum time sublayer instruction 99 indicates the maximum time sublayer of the multilayer video bitstream 14 to be decoded by the decoder 50. In other words, the maximum time sublayer instruction 99 may signal to the decoder 50 which set or which subset of the time layers of the video bitstream 14 the decoder will decode. Thus, the decoder 50 may receive the maximum time sublayer instruction 99 from an external signal. For example, the maximum time sublayer instruction 99 may be included in the video bitstream 14 or provided to the decoder 50 via an API. Upon receiving the maximum time sublayer instruction 99, the decoder 50 may decode the video bitstream 14 or any portion thereof, insofar as it belongs to the set of time sublayers indicated by the maximum time sublayer instruction 99. For example, the decoder 50 may be further configured to receive instructions for the OLS to be decoded. Upon receiving the instruction for the OLS to be decoded and the maximum time sublayer instruction 99, the decoder 50 may decode the layer indicated by the OLS to be decoded up to the time sublayer indicated by the maximum time sublayer instruction 99. However, there may be situations or scenarios in which the decoder 50 does not receive one or both of the external instructions for the time sublayer, i.e., the maximum time sublayer instruction 99 and the instruction for the OLS to be decoded. In such situations, the decoder 50 may determine the missing instruction, for example, based on information available in the video bitstream 14.
[0088] According to the first embodiment of the third aspect, if the decoder 50 is unable to obtain the maximum time sublayer instruction 99, for example, from the data stream and / or via other means, the decoder 50 determines that the maximum time sublayer to be decoded is equal to the maximum time sublayer indicated by the constraint 92 for the maximum time sublayer in the decoder requirements set 82 associated with the OLS 83. In other words, the decoder 50 may use the constraint 92 indicated for the decoder requirements set 82 associated with the OLS to be decoded by the decoder 50 for the maximum time sublayer.
[0089] The decoder 50, in order to decode the multilayer video bitstream 14, may consider the time sublayers of the multilayer video bitstream for decoding if the time sublayers do not exceed the maximum time sublayer to be decoded, and if they do exceed the maximum time sublayer, it may omit the time sublayers during decoding and use information about the maximum time sublayer to be decoded.
[0090] For example, the video bitstream 14 may represent one or more OLS 83s, each OLS 83 associated with one of the bitstream compatibility set 86, buffer requirements set 84, and decoder requirements set 82, which are signaled in the video parameter set 91. The decoder 50 may select an OLS 83 for decoding based on an external instruction (e.g., provided via an API or the video bitstream 14) or, for example, if there is no external instruction, based on a selection rule. If a maximum time sublayer instruction 99 is unavailable, i.e., the decoder 50 does not receive a maximum time sublayer instruction, the decoder 50 uses the constraint 92 on the maximum time sublayer of the decoder requirements set 82 associated with the OLS to decode.
[0091] For example, the constraint 82 of the decoder requirements set 82 associated with the OLS is due to bitstream constraints less than or equal to the constraints 94 and 96 for the maximum time sublayer associated with the buffer requirements set 84 and the bitstream compatibility set 86. Therefore, selecting the constraint 92 associated with the decoder capability set for decoding means selecting the minimum value that exceeds the maximum time sublayer shown for the decoder requirements set 82, buffer requirements set 84, and bitstream compatibility set 86 associated with the OLS. By selecting the minimum value that exceeds the maximum time sublayer constraint, it can be ensured that each parameter set 82, 84, and 86 contains parameters that are valid for the bitstream selected for decoding. Thus, by selecting the constraint 92 associated with the decoder requirements set 82, it is possible to ensure that a bitstream for decoding is selected for which all parameters of parameter sets 82, 84, and 86 are available.
[0092] For example, each of the parameter sets from the bitstream compatibility set 86, buffer requirements set 84, and decoder requirements set 82 associated with the OLS may contain one or more sets of parameters, and each parameter set may be associated with a time sublayer or a maximum time sublayer. The decoder 50 may select from each parameter set a parameter set associated with the maximum time sublayer to be decoded, so as to be estimated or received. For example, if a parameter set is associated with the maximum time sublayer, or if a parameter set is associated with a time sublayer less than or equal to the maximum time sublayer, the parameter set may be associated with the maximum time sublayer. The decoder 50 may use the selected parameter set to adjust one or more of the following: the buffer size of the coded image, the buffer size of the decoded image, buffer scheduling, for example, HRD timing (AU / DU removal time, DPB output time).
[0093] As described above, the instruction 99 for the maximum time sublayer to be decoded may be signaled in the video bitstream 14. According to the first embodiment, the encoder 10, for example, the encoder 10 in Figure 1, may be configured to optionally omit the signaling of instruction 99 if the maximum time sublayer to be decoded corresponds to a constraint 92 associated with a decoder requirement set 82 of an OLS, for example, an OLS instructed to be decoded, in which case the decoder 50 can accurately estimate the maximum time sublayer to be decoded. In other words, in this case the encoder 10 may or may not signal instruction 99. In the example, the encoder 10 may decide whether to signal instruction 99 or to omit the signaling of instruction 99 in this case.
[0094] For example, VPS91 may follow the following example. Currently, there are three syntactic structures for VPS that are generally defined and then mapped to specific OLS. • Profile hierarchy level (PTL), for example, decoder requirement set 82, • DPB parameters, e.g., buffer requirement set 84, • HRD parameters, e.g., bitstream compatibility set 86, And another syntax element, vps_max_sublayers_minus1, is used when extracting the OLS sub-bitstream, that is, when deriving the variable NumSublayerInLayer[i][j] in some cases.
[0095] Figure 8 shows an example of the definitions of PTL82, DPB84, and HRD86 and their mapping to OLS83. The mapping from PTL to OLS is performed in the VPS for all OLS (single-layer or multi-layer). However, the mapping of DPB and HRD parameters to OLS is performed only in the VPS for OLS with multiple layers. As shown in Figure 8, the PTL, DPB, and HRD parameters are first written in the VPS and then mapped to indicate which parameters the OLS will use.
[0096] In the example shown in Figure 8, there are two OLSs and two of each of their parameters. However, the definitions and mappings are specified so that multiple OLSs can share the same parameters, so it is not necessary to repeat the same information multiple times, as shown in Figure 9, for example. Figure 9 shows an example of the definitions of PTL, DPB, and HRD and their sharing among different OLSs. Here, in the example in Figure 9, OLS2 and OLS3 have the same PTL and DBP parameters, but different HRD parameters.
[0097] In the examples in Figures 8 and 9, the values of vps_ptl_max_temporal_id[ptlIdx] (e.g., constraint 92 for the maximum time sublayer of decoder requirement set 82), vps_dpb_max_temporal_id[dpbIdx] (e.g., constraint 94 for the maximum time sublayer of buffer requirement set 84), and vps_hrd_max_tid[hrdIdx] (e.g., constraint 96 for the maximum time sublayer of bitstream compatibility set 86) for a given OLS 83 are aligned, but this is not currently required. These three values associated with the same OLS are not currently restricted to have the same value (however, in some examples of this disclosure, they may be restricted at will). ptlIdx, dpbIdx, and hrdIdx are indices to each syntax structure that is signaled for each OLS.
[0098] According to an example of the first embodiment, the value of vps_ptl_max_temporal_id[ptlIdx] is used to set the variable HTid in the decoding process unless it is set by an external means as follows: That is, if there is no external means to set the value of HTid (e.g., via the decoder API), the value of vps_ptl_max_temporal_id[ptlIdx] is obtained by default to set HTid, i.e., the minimum value of the three syntax elements described above. In other words, the decoder 50 may set the variable HTid according to the maximum time sublayer instruction 99 if it is available, and if it is not available, HTid may be set to the value of vps_ptI_max_temporal_id[ptlldx].
[0099] According to a second embodiment of the third aspect, the decoder 50 estimates that the maximum time sublayer of the set of time sublayers to which each image 26 of layer 24 included in the OLS belongs is the minimum of the maximum time sublayers indicated by the bitstream compatibility set 86, buffer requirements set 84, and decoder requirements set 82 associated with the OLS. For example, the maximum time sublayer estimated for a set of time sublayers may correspond to or be indicated by the variable Maximum TID WITHINOLS. In other words, the set of time sublayers to which each image of a layer in the OLS belongs may be the set of time sublayers that contain all images belonging to the OLS. In this regard, all images of layers included in the OLS may belong to the OLS. For example, referring to Figure 7, an OLS 83 including a first layer 241 and a second layer 240 may include a first time sublayer to which access unit 221 belongs and a second time sublayer to which access unit 222 belongs.
[0100] In the example, the decoder 50 may detect whether the video bitstream 14 indicates a constraint on the maximum time sublayer of the reference layer to which each layer depends for one or more, or all, layers of the OLS being decoded. For example, the constraint may indicate that each layer depends only on the time sublayer of the reference layer up to the maximum time sublayer. If the video bitstream 14 does not indicate such a constraint, the decoder 50 infers that the maximum time sublayer included in the OLS is equal to the maximum time sublayer estimated for the set of time sublayers to which each image of the layers in the OLS belongs.
[0101] In the example, the OLS instruction 81 may further indicate one or more output layers for OLS83. In other words, one or more of the layers included in OLS83 may be indicated as the output layers of OLS83. For each layer indicated as an output layer of OLS83, the decoder 50 can estimate that the maximum time sublayer included in the OLS is equal to the estimated maximum time sublayer for the set of time sublayers to which each image of the layer in the OLS belongs.
[0102] The decoder 50 may decode from the video bitstream 14 image 26 images belonging to one of the layers included in the OLS to be decoded, and images belonging to time sublayers less than or equal to the maximum time sublayer to be decoded. The maximum time sublayer to be decoded may be the maximum time sublayer estimated for the set of time sublayers to which each image of the OLS layer belongs, or the maximum time sublayer to be decoded as described in relation to the above embodiment.
[0103] In other words, according to the second embodiment, the above variable MaxTidWithiOls (not necessarily a bitstream, as some may have been dropped), which indicates the maximum number of time sublayers present in the OLS, is derived on the decoder side from the minimum value of the three values of the syntax element: vps_ptl_max_temporal_id[ptlIdx], vps_dpb_max_temporal_id[dpbIdx], and vps_hrd_max_tid[hrdIdx]. This minimizes the constraint of sharing any one of the PTL, HRD, or DPB parameters and prevents the representation of an OLS where all three parameters are undefined. MaxTidWithinOls=min(vps_ptl_max_temporal_id[ptlIdx],min(vps_dpb_max_temporal_id[dpbIdx],vps_hrd_max_tid[hrdIdx]))
[0104] Furthermore, in this embodiment, NumSublayerlnLayer[i][j], which represents the maximum sublayer included in the i-th OLS of layer j, is set to equal the derived MaxTidWithinOls when vps_max_tid_il_ref_pics_plus1[m][k] does not exist, or when layer j is the output layer of the i-th OLS.
[0105] For example, the decoder 50 can estimate that the maximum time sublayer of the multilayer video bitstream being decoded is equal to the maximum time sublayer estimated for the set of time sublayers to which each image in the OLS layer belongs, i.e., called MaxTidWithinOLS.
[0106] In other words, in another embodiment, the variable HTid in the decoding process can be set using the derived value of MaxTidWithinOls, unless set by an external means as follows: If there is no external means to set the value of HTid (e.g., via the decoder API), the value of MaxTidWithinOls is obtained by default to be set to HTid, i.e., the minimum value of the three syntax elements described above.
[0107] In an example of an embodiment according to the third aspect, the decoder 50 selectively considers for decoding images belonging to one of the time layers of the multilayer video bitstream 14, provided that each image belongs to an access unit 22 associated with a time sublayer that does not exceed the maximum time sublayer to be decoded.
[0108] A further embodiment according to a third aspect includes an encoder 10, for example, the encoder 10 of Figure 1, which encodes a multilayer video bitstream 14 according to Figure 7. For this purpose, when associating the OLS 83 with the bitstream compatibility set 86, the buffer requirements set 84, and the decoder requirements set 82, the encoder 10 may form the OLS instruction 81 such that the smallest of the constraints 96, 94, and 92 associated with each parameter set accommodates a subset of layers indicated by the OLS 83. That is, for example, the smallest of the maximum time sublayers indicated by the bitstream compatibility set 86, the buffer requirements set 84, and the decoder requirements set 82 associated with the OLS 83 is greater than or equal to the maximum time sublayer of the time sublayers included in the subset of layers indicated by the OLS. The encoder 10 may further form the OLS instruction 81 such that the parameters in the bitstream compatibility set 86, buffer requirements set 84, and decoder requirements set 82 are valid for the OLS 83 insofar as those parameters refer to time sublayers less than or equal to the minimum value among the maximum time sublayers 92, 94, 96 indicated for the bitstream compatibility set 86, buffer requirements set 84, and decoder requirements set 82 associated with the OLS 83.
[0109] Some embodiments are described as features in the context of the apparatus, but it is clear that such descriptions may also be considered as descriptions of corresponding features of the method. Some embodiments are described as features in the context of the method, but it is clear that such descriptions may also be considered as descriptions of corresponding features relating to the function of the apparatus.
[0110] Some or all method steps may be performed by (or using) hardware devices such as a microprocessor, a programmable computer, or an electronic circuit. In some embodiments, one or more of the most important method steps may be performed by such devices.
[0111] The encoded image signal of the invention can be stored in a digital storage medium or transmitted over a transmission medium such as a wireless transmission medium such as the Internet or a wired transmission medium. In other words, further embodiments provide a video bitstream product including a video bitstream according to any of the embodiments described herein, for example, a digital storage medium storing a video bitstream.
[0112] Depending on specific implementation requirements, embodiments of the present invention may be implemented in hardware or software, or at least partially in hardware or at least partially in software. These implementations may be performed using digital storage media such as floppy disks, DVDs, Blu-rays, CDs, ROMs, PROMs, EPROMs, EEPROMs, or FLASH® memory, which store electronically readable control signals, and which cooperate (or can cooperate) with a programmable computer system to perform each method. Therefore, the digital storage media may be computer-readable.
[0113] In some embodiments of the present invention, one of the methods described herein is performed by including a data carrier having electronically readable control signals, wherein these control signals can cooperate with a programmable computer system.
[0114] Generally, embodiments of the present invention can be implemented as a computer program product having program code, which is operable to perform one of the methods when the computer program product is executed on a computer. The program code may be stored, for example, in a machine-readable carrier.
[0115] Other embodiments include a computer program that performs one of the methods described herein and is stored in a machine-readable carrier.
[0116] In other words, embodiments of the method of the present invention are computer programs having program code for performing one of the methods described herein when the computer program is executed on a computer.
[0117] Further embodiments of the methods of the present invention are, therefore, a data carrier (or digital storage medium, or computer-readable medium), which includes a computer program recorded thereon for performing one of the methods described herein. The data carrier, digital storage medium, or recorded medium is typically tangible and / or non-temporary.
[0118] A further embodiment of the method of the present invention is a data stream or signal sequence representing a computer program for performing one of the methods described herein. The data stream or signal sequence may be configured to be transmitted over a data communication connection, such as over the Internet.
[0119] Further embodiments include, for example, processing means such as a computer or programmable logic device configured or adapted to perform one of the methods described herein.
[0120] Further embodiments include a computer on which a computer program for performing one of the methods described herein is installed.
[0121] Further embodiments of the present invention include an apparatus or system configured to transfer (e.g., electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver may be, for example, a computer, a mobile device, a memory device, etc. The apparatus or system may include, for example, a file server for transferring the computer program to the receiver.
[0122] In some embodiments, a programmable logic device (e.g., a field-programmable gate array) may be used to perform some or all of the functions of the method herein. In some embodiments, a field-programmable gate array may cooperate with a microprocessor to perform one of the methods herein. Generally, the method is preferably performed by any hardware device.
[0123] The apparatus described herein may be implemented using hardware devices, a computer, or a combination of hardware devices and a computer.
[0124] The methods described herein may be performed using hardware devices, or using a computer, or using a combination of hardware devices and a computer.
[0125] In the modes for carrying out the invention described above, various features are grouped together in examples for the purpose of streamlining the disclosure. This method of disclosure should not be interpreted as indicating an intention that the claimed examples require more features than are explicitly described in each claim. Rather, as the following claims demonstrate, the subject matter of the invention consists of fewer features than all the features of a single disclosed example combined. Thus, the following claims are incorporated into the modes for carrying out the invention by this specification, and each claim can stand alone as a separate example. While each claim can stand alone as a separate example, it should be noted that dependent claims may refer to a specific combination with one or more other claims within the claims, but other examples may also include combinations of a dependent claim with the subject matter of each other dependent claim, or combinations of each feature with other dependent or independent claims. Such combinations are proposed herein unless otherwise stated that a particular combination is not intended. Furthermore, even if a claim is not directly dependent on an independent claim, it is also intended that the features of that claim be included in any other independent claim.
[0126] The embodiments described above are merely illustrative of the principles of this disclosure. Naturally, modifications and variations of the configurations and details described herein will be obvious to those skilled in the art. Therefore, it is intended that the invention be limited only by the pending claims and not by the specific details presented in the description and explanation of the embodiments herein.
Claims
1. A method for decoding a video bitstream, The process involves receiving a video parameter set (VPS) from the video bitstream, which includes one or more profile hierarchy level parameters, one or more decoded image buffer (DPB) parameters, and one or more virtual reference decoder (HRD) parameters. Receiving from the video bitstream an instruction for a first maximum time sublayer corresponding to one or more profile hierarchy level parameters, an instruction for a second maximum time sublayer corresponding to one or more DPB parameters, and an instruction for a third maximum time sublayer corresponding to one or more HRD parameters, A method comprising: setting the instruction for the stream maximum time sublayer based on the instruction for the first maximum time sublayer when it is determined that there is no instruction for the stream maximum time sublayer included in the video bitstream.
2. The method according to claim 1, further comprising associating an output layer set (OLS) with one or more profile hierarchy level parameters, one or more DPB parameters, and one or more HRD parameters.
3. Determining whether the time sublayer of the video bitstream exceeds the maximum stream time sublayer of the video bitstream, The method according to claim 1, further comprising omitting decoding the time sublayer of the video bitstream based on the determination.
4. Determining that the time sublayer of the video bitstream does not exceed the maximum stream time sublayer of the video bitstream, The method according to claim 1, further comprising decoding the time sublayer of the video bitstream based on the determination.
5. A method for encoding a video bitstream, The video bitstream is provided with a video parameter set (VPS) which includes one or more profile hierarchy level parameters, one or more decoded image buffer (DPB) parameters, and one or more virtual reference decoder (HRD) parameters. The video bitstream is provided with instructions for a first maximum time sublayer corresponding to one or more profile hierarchy level parameters, instructions for a second maximum time sublayer corresponding to one or more DPB parameters, and instructions for a third maximum time sublayer corresponding to one or more HRD parameters. If it is determined that there is no instruction for a maximum stream time sublayer included in the video bitstream, the instruction for the maximum stream time sublayer is set as the instruction for the maximum stream time sublayer based on the instruction for the first maximum time sublayer. A method comprising providing the video bitstream with a stream maximum time sublayer instruction if such instruction exists within the video bitstream.
6. The method according to claim 5, further comprising omitting the instruction for the stream maximum time sublayer of the video bitstream in the video bitstream if it is determined that the instruction for the stream maximum time sublayer of the video bitstream is equal to the instruction for the first maximum time sublayer.
7. The method of claim 5, further comprising providing the instruction for the stream maximum time sublayer of the video bitstream in the video bitstream if it is determined that the instruction for the stream maximum time sublayer of the video bitstream is not equal to the instruction for the first maximum time sublayer.
8. The method according to claim 5, further comprising associating an output layer set (OLS) with one or more profile hierarchy level parameters, one or more DPB parameters, and one or more HRD parameters.
9. A video coding apparatus comprising one or more processors configured to perform the method described in any one of claims 1 to 8.
10. A non-temporary computer-readable medium storing a computer program for carrying out the method described in any one of claims 1 to 8 when executed by a computer or signal processor.
Citation Information
Patent Citations
Temporal sub-layer descriptor
US20190158880A1
Avoidance of redundant signaling in multi-layer video bitstreams
WO2021022267A2
Signaling of DPB parameters for multi-layer video bitstreams
WO2021061489A1
Coding output layer set data and conformance window data of high level syntax for video coding
WO2021174098A1
Systems and methods for decoding based on inferred video parameter sets
WO2021236312A1