The concept of video coding related to subpictures

JP2026139874APending Publication Date: 2026-09-01FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2026112397
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2020-05-22
Filing Date
2026-06-23
Publication Date
2026-09-01

Smart Images

  • Figure 2026139874000001_ABST
    Figure 2026139874000001_ABST
Patent Text Reader

Abstract

The system appropriately controls whether or not the inter-layer prediction tool can be used, even when the configuration settings of the current picture and the reference picture differ. [Solution] The current picture of the data stream's enhancement layer (L1) is compared with the first and second configuration settings associated with the reference picture of the base layer (L0), and variables are set based on this comparison. Based on the variables, it is determined that at least one cross-layer prediction tool is unavailable, and the blocks of the current picture from the reference picture are predicted without using that tool.
Need to check novelty before this filing date? Find Prior Art

Description

SUMMARY OF THE INVENTION

[0001] Embodiments of the present disclosure relate to inter-layer prediction in multi-layer video data streams, and in particular, to encoding and decoding video data for controlling the availability of an inter-layer prediction tool based on a comparison of configuration settings between a current picture encoded in an enhancement layer and a reference picture encoded in a base layer.

[0002] In multi-layer video coding, inter-layer prediction using a reference picture obtained from a base layer can be performed to encode or decode a current picture of an enhancement layer. The reference picture may be, for example, a picture that is temporally co-located with the current picture.

[0003] The current picture and the reference picture may have mutually different configuration settings regarding independent coding in units of subpictures. For example, one picture may include a plurality of subpictures, while the other picture may include one subpicture or may not be subdivided into subpictures.

[0004] If the availability of an inter-layer prediction tool is determined without considering the relationship between these configuration settings, the availability or operation of the tool will not match between encoding and decoding, which may cause inconsistency between the encoding side and the decoding side. Such a problem may particularly occur when the derivation result of configuration settings changes along with subpicture extraction or the like.

[0005] Accordingly, an object of the present disclosure is to provide a concept for appropriately controlling the availability of an inter-layer prediction tool based on a comparison between the configuration settings associated with the current picture and the configuration settings associated with the reference picture.

[0006] A decoding method according to one embodiment includes acquiring a data stream including an enhancement layer and a base layer, and identifying a reference picture decoded from the base layer for use in inter-layer prediction of the current picture, for the current picture encoded in the enhancement layer.

[0007] The decoding method further includes setting variables based on a comparison between a first configuration setting associated with the current picture and a second configuration setting associated with a reference picture, and determining, based on these variables, that at least one cross-layer prediction tool is unavailable for predicting blocks of the current picture using the reference picture.

[0008] The decoding method further includes predicting blocks in the current picture using a reference picture without using the at least one interlayer prediction tool that has been determined to be unavailable.

[0009] In one embodiment, a first configuration setting indicates the number of first sub-pictures included in the current picture, and a second configuration setting indicates the number of second sub-pictures included in the reference picture. If the number of first sub-pictures and the number of second sub-pictures are different, the variable may be set to a first value, and based on the fact that the variable has a first value, it may be determined that the at least one inter-layer prediction tool is unavailable.

[0010] The at least one inter-layer prediction tool may include at least one of inter-layer motion vector prediction, optical flow refinement of vector-based inter-layer prediction, and motion vector wrap-around in vector-based inter-layer prediction.

[0011] In one embodiment, the current picture and the reference picture may have the same spatial resolution and the same scaling window offset. The current picture and the reference picture may be at the same position in time.

[0012] Further embodiments relate to a decoder comprising a processor configured to perform the decoding method, and a non-temporary computer-readable medium including a program that, when executed by the processor, causes the processor to perform the decoding method.

[0013] Further embodiments relate to an encoding method and encoder that sets variables based on a comparison of first and second configuration settings of a current picture encoded in an enhancement layer and a reference picture encoded in a base layer, and encodes blocks of the current picture using the reference picture without using at least one cross-layer prediction tool that is determined to be unavailable based on those variables. Further embodiments relate to a non-temporary computer-readable medium including a program that causes a processor to perform the encoding method.

[0014] This ensures that even when the configuration settings of the current picture and the reference picture differ, the availability of the inter-layer prediction tool on the encoding and decoding sides is consistent, allowing for proper inter-layer prediction.

[0015] Embodiments and advantageous embodiments of this disclosure are described in further detail below with reference to the figures. [Brief explanation of the drawing]

[0016] [Figure 1] Examples of encoders, extractors, and decoders according to embodiments are shown. [Figure 2] Examples of encoders, video data streams, and extractors according to embodiments are shown. [Figure 3] This shows an example of a video data stream and an example of incorrect extraction of video data streams by subpicture. [Figure 4] An example of video data streams separated by sub-picture is shown. [Figure 5] An example of an encoder, extractor, and multilayer video data stream according to a second aspect of the present invention is shown. [Figure 6]This demonstrates the adaptation of the scaling window according to the embodiment. [Figure 7] Here is another example of extracting video data streams by subpicture. [Figure 8] This shows an example of the output picture area percentage for each sub-picture in the video data stream. [Figure 9] Examples of the ratio of subdivided and non-subdivided layers in the standard decoder functional requirements are shown. [Figure 10] Examples of encoders, video data streams, devices for mixing and / or extracting video data streams, and decoders are shown. [Figure 11] This example shows a bitstream that combines boundary processing on both the decoder and encoder sides. [Figure 12] This example shows a single-layer bitstream with mixed boundary processing between the decoder and encoder sides. [Figure 13] Examples of bitstream portions depending on the type of independence are shown. [Modes for carrying out the invention]

[0017] The embodiments described below will be described in detail, but it should be understood that the embodiments provide many applicable concepts that can be embodied in a wide variety of video coding concepts. The specific embodiments described are merely illustrative of specific ways of carrying out and using the concepts and do not limit the scope of the embodiments. Several details are given in the following description to provide a more complete description of the embodiments of the present invention. However, it will be apparent to those skilled in the art that other embodiments can be carried out without these specific details. In other examples, well-known structures and devices are shown in block diagram form rather than in detail to avoid obscuring the examples described herein. Furthermore, features of the different embodiments described herein can be combined with each other unless otherwise noted.

[0018] In the following description of embodiments, identical or similar elements, or elements having the same function, are denoted by the same reference signs or identified by the same names. Repeated descriptions of elements denoted by the same reference signs or identified by the same names are generally omitted. Therefore, the descriptions provided for elements having the same or similar reference signs or identified by the same names are mutually interchangeable, or may be applied to each other in different embodiments.

[0019] The following description of the figures starts with the presentation of an encoder, an extractor and a decoder with reference to Figure 1. The encoder and extractor of Figure 1 provide an example of a framework that may incorporate embodiments of the present invention. Hereinafter, a description of embodiments of the concepts of the present invention is presented together with a description of how such concepts may be incorporated into the encoder and extractor of Figure 1. However, the embodiments described with reference to Figure 2 and subsequent figures may be used to form encoders and extractors that do not operate according to the framework described with reference to Figure 1. Furthermore, it should be noted that although the encoder, the extractor, and the decoder are described together for illustrative purposes in Figure 1, they may be implemented separately from each other.

[0020] 0. Encoder 40, extractor 10, decoder 50 and video data stream 14 according to Figure 1 Figure 1 shows an example of an encoder 40, an extractor 20, and a decoder 50. The encoder 40 encodes the video 24 into a video data stream 14. The video data stream 14 (the video data stream may also be referred to herein as a bitstream) can be transmitted, for example, or stored in a data carrier. The video 24 may contain a sequence of pictures, each of which can be associated with a presentation time in presentation time order. The video 24 may contain a plurality of pictures 26, represented by pictures 260 and 261 in Figure 1. Each of the pictures 26 can be associated with a layer, such as a first layer L1 or a second layer L0 in Figure 1. For example, in Figure 1, picture 260 is associated with the second layer L0, and picture 261 is associated with the first layer L1. Picture 260 on layer L0 can form a video sequence 240. Picture 261 on layer L1 can form a video sequence 241. In the example, video sequences 240 and 241 of video 24 may represent the same content and be associated with the same presentation time, but may contain pictures of different resolutions. However, it should be noted that video sequences 240 and 241 associated with layers L0 and L1 of video 24 may have different frame rates. Therefore, for example, video sequence 240 may not necessarily contain picture 260 for each presentation time in which video sequence 241 contains picture 261. Encoder 40 can encode picture 261 of video sequence 241 into video data stream 14, depending on, i.e., associated with, picture 260 of video sequence 240 that is at the same temporal position as picture 261, i.e., associated with the same presentation time (e.g., called a reference picture for inter-layer prediction). In other words, a picture in layer L0 may be a reference picture for a picture in layer L1. Therefore, decoding picture 261 from video data stream 14 may require picture 260 encoded in video data stream 14.In these examples, layer L0 may be referred to as a base layer, and pictures associated with layer L0 may have a first resolution. Picture 261 associated with a first layer L1, which may be referred to as an enhancement layer, can have a second resolution higher than the first resolution. For example, encoder 40 can encode video 24 into video data stream 14 such that layer L0 representing video 24 of the first resolution can be decoded from video data stream 14 independently of layer L1 to obtain video of the first resolution. When picture 261 of layer L1 depends on picture 260 of layer L0, video sequences 240 and 241 can optionally be decoded together, resulting in video having a resolution higher than the first resolution, for example, the second resolution. Accordingly, there may be multiple options for decoding video data stream 14, with each individual option involving different data rates of the data stream to be decoded and resulting in video having different resolutions.

[0021] It is pointed out that the number of layers shown in FIG. 1 is exemplary, and video data stream 14 may have three or more layers. It should also be noted that video data stream 14 does not necessarily include a plurality of layers, but in some examples, as in the exemplary embodiment described in section 6, it may include only a single layer.

[0022] Extractor 10 can receive video data stream 14 to extract video data stream 12, and can extract from it a not-so-appropriate subset of the bitstream portion of video data stream 14. In other words, the extracted video data stream 12 may correspond to video data stream 14 or contain a distribution of video data stream 14. It is noted that when transferring the bitstream portion in the extracted video data stream 12, extractor 10 may optionally modify the content of the bitstream portion of video data stream 14. For example, extractor 10 may modify descriptive data containing information about how to decode the encoded video data of the extracted video data stream 12. Extractor 10 may select the bitstream portion of video data stream 14 to transfer to the extracted video data stream 12 based on the output layer set (OLS) extracted from video data stream 14. For example, extractor 10 may receive an OLS display indicating the OLS to be extracted or presented. Video data stream 14 may contain an OLS display indicating the set of OLS that can be extracted from video data stream 14. The OLS may indicate one or more or all of the layers of the video data stream 14 that should be transferred in whole or in part to the extracted video data stream 12.

[0023] In other words, an OLS may represent, for example, a (not necessarily appropriate) subset of the layers of a multilayer video data stream 14. An OLS may be represented in the multilayer data stream itself, such as an OLS representation 18, which may be included in the video parameter set (VPS) of the bitstream 14. In fact, there may be two or more such OLSs represented within the data stream 14, which are determined by external means, for example via an API 20, to determine which ones are used for extraction. Note that while an OLS representation 18 primarily represents one or more layers of the output or presented video, it may also represent non-output / non-presented output reference layers that belong to one or more outputs in that one or more output layers depend on the layer directly or indirectly (through another reference layer). An OLS representation 18 may represent one or more OLSs.

[0024] Optionally, the OLS display 18 may show video parameters of the OLS in, for example, a video parameter set (VPS). For example, the video parameters may show reference decoder functional requirements (also called decoder functional requirements, (reference) level information, or (reference) level display) that raise the requirements of a decoder to enable decoding of the bitstream described by the OLS. Note that the video parameters may show one or more reference decoder functional requirements for a single OLS, since the bitstream described by the OLS may still be scalable by selecting / extracting one or more time sublayers and / or one or more layers of the OLS after extraction by the extractor 10. For example, a mixer or merger (e.g., device 100 in Figure 10) may form a bitstream using the extracted video data stream 12, or a decoder may select sub-bitstreams of the extracted video data stream 12 for decoding.

[0025] The extracted video data stream 12 is transferred to the decoder 50, which decodes the extracted video data stream 12 to obtain the decoded video 24'. The decoded video 24' may differ from video 24 in that it does not necessarily contain the entire content of video 24, and / or may have a different resolution, and / or may have distortions with respect to video 24, for example, based on quantization loss.

[0026] A picture 26 in one of the layers of the video data stream 14 may contain one or more subpictures 28, or may be subdivided into multiple subpictures 28. Note that, below, reference numeral 26 is used to refer to the pictures of a layer. Picture 26 may refer, for example, to picture 261 of layer L1 or picture 260 of layer L0. The encoder 40 can encode the subpictures 28 of picture 26 independently of each other. That is, one of the subpictures 28 of picture 26 can be decoded without requiring another subpicture 28 of picture 26. For example, the extractor 10 does not need to transfer all of the subpictures 28 of a layer, but can transfer only a subset or one of the subpictures 28 of each picture 26 in the layer. Therefore, in this case, the data rate of the extracted video data stream 12, which can be called the sub-picture video data stream 12, may be lower than the data rate of the video data stream 14 so as to reduce the decoder resources required to decode the extracted video data stream 12. In the example in Figure 1, the extractor 10 transfers the picture 240 of layer L0 and the sub-picture 28 of the picture 261 of layer L1. In this scenario, the decoded video 24' includes the decoded video sequence 24'0, which represents the decoded picture of video sequence 240. Furthermore, the decoded video sequence 24' in this example includes the decoded video sequence 24'1, which represents the decoded picture 26'1, which includes the sub-picture 28 of the picture 261 of video sequence 241.

[0027] The encoder not only indicates one or more valid / extractable layer sets in the OLS display, but, according to the embodiment, also provides the data stream with information that the decoder can use to determine whether a particular OLS shown is decodeable by the decoder, such as, for example, buffer memory available for DPB and / or CPB, kernel processing, and desired decoding delay. This information may be included in decoder functional requirements information.

[0028] After a general explanation of the concepts of multilayer video data streams, subpictures, bitstream scalability, and reference pictures, several embodiments for carrying out the extraction process of the extracted video data stream 12 from the video data stream 14, the associated representation within the video data stream 14, and / or the encoding method of the video data stream 14 will be described below. It should be noted that the features described with respect to the extraction process also correspond to a description of the corresponding video data stream from which the extracted video data stream is extracted, and a description of the corresponding encoding process of the video data stream. For example, a feature that defines the extractor 10 to derive information from the video data stream 14 should also be understood as a feature of the video data stream 14 that information can be derived from the video data stream, and should also be understood as a feature of the encoder 40 in terms of properly encoding the video data stream 14.

[0029] 1. Encoder 40, extractor 10 and video data stream 14 according to the first embodiment This section describes an embodiment according to a first aspect with reference to Figure 1. The details described in Section 0 may be optionally applied to the embodiment according to the first aspect.

[0030] According to an embodiment of the first aspect, the extractor 10 in Figure 1, also called a device for extracting sub-picture video data streams, is configured to extract sub-picture video data streams 12 from a multi-layer video data stream 14. That is, according to the first aspect, the video data stream 14 includes multiple layers and also includes bitstream portions 16, each bitstream portion 16 belonging to one of the layers of the multi-layer video data stream 14, for example, layers L0 and L1. For example, each bitstream portion 16 may include an indication such as a layer ID that shows the layer to which each bitstream portion 16 belongs. According to an embodiment of the first aspect, the extractor 10 is configured to check, for each layer from the layer set, whether the video 24 is encoded in each layer such that each picture 261, 262 of the video 24 is subdivided into two or more subpictures 28 that are independently encoded in each layer, so that each subpicture 28 of each picture is encoded in a different bitstream portion 16 of each layer, or whether the video 24 is encoded in each layer without being subdivided into subpictures.

[0031] For example, the layer set is a layer set represented by OLS that the OLS extractor 10 is instructed to extract by an external means such as the API 20, or by OLS that the OLS extractor infers will be extracted if, for example, there are no such instructions, as explained with respect to Figure 1.

[0032] For example, the fact that subpictures 28 are encoded independently of each other may mean that each subpicture 28 is encoded independently of any other subpictures (of the picture to which it belongs), and that each bitstream portion 16 has only one (no more than one, or only a part of one) subpicture encoded therein. In other words, an independently encoded subpicture 28 may not require any bitstream portions from the bitstream portion 16 into which the picture to which the subpicture 28 belongs is encoded, other than the bitstream portion into which the independently encoded subpicture 28 is encoded. Encoding video 24 so as not to be subdivided into subpictures may mean, for example, that each layer has only one sub-part encoded, or that each layer has only one sub-part encoded. In other words, encoding so as not to be subdivided into subpictures may mean that each picture is encoded so as to be encoded as one subpicture.

[0033] For example, in the case of layer L0 in Figure 1, where the video is encoded on each layer so as not to be subdivided into subpictures, the extractor 10 can transfer the bitstream portion 16 belonging to each layer (the dot-shaped bitstream portion 16 in the video data streams 14 and 12 in Figure 1) from the multi-layer video data stream 14 to the subpicture-specific video data stream 12, so that the video 24 of each layer, for example the video sequence 240 in Figure 1, is fully encoded in the subpicture-specific video data stream 12. Fully encoding the video of each layer may mean, for example, that all of the bitstream portion 16 of each layer, for example layer L0 of the multi-layer video data stream 14, is included in the subpicture-specific video data stream 12. Alternatively or additionally, the fact that the video 24 of each layer is fully encoded into a separate video data stream for each subpicture may mean that the bitstream portion 16 taken over by the separate video data stream 12 is independent of the selection in the separate video data stream 12 of which of the two or more subpictures of another layer is transferred, i.e., which belongs to a predetermined set of one or more subpictures described later.

[0034] If the extractor finds that a video 24, for example, a video sequence 241 of layer L1, is encoded in each layer, i.e., the currently checked layer, such that a video picture 26 of the layer, for example, picture 261 of layer L1, is subdivided into two or more subpictures 28, the extractor 10 can read information from each bitstream portion 16 belonging to each layer, such as layer L1 in Figure 1, that reveals which of the two or more subpictures is encoded in each bitstream portion 16. If each bitstream portion 16 contains subpictures belonging to a predetermined set of one or more subpictures, the extractor can transfer each bitstream portion 16 of each layer from the multilayer video data stream 14 to the subpicture-specific video data stream 12. For example, in Figure 1, the predetermined set of subpictures includes only the subpicture 28 shown by cross-hatching. For example, extractor 10 may drop, leave, or remove—that is, not transfer—each bitstream portion that encodes a subpicture that belongs to a particular layer but does not belong to a predetermined set of one or more subpictures.

[0035] Figure 2 illustrates in more detail the information signaled in the video data stream 14 and the sub-picture video data stream 12 according to an embodiment of the first aspect of the present invention, with respect to the scenario shown in Figure 1. According to the exemplary diagram in Figure 2, a given set of sub-pictures includes only one sub-picture 28, i.e., a cross-hatched sub-picture 28 of picture 261 of layer L1. As shown in Figure 2, based on the discovery that the video, i.e., the video sequence 240, in layer L0 is encoded in such a way that it is not subdivided into sub-pictures, the extractor 10 transfers or takes over all bitstream portions 16 (dot portions 16) associated with layer L0. The bitstream portions associated with layer L1 are encoded in the video data stream 14 using a sub-picture subdivision scheme. In the example in Figure 2, the bitstream portion into which the sub-picture 28 of picture 261 of layer L1 is encoded is shown with cross-hatching, and the bitstream portions into which further sub-pictures of picture 261 are encoded are shown with simple hatching. As shown in Figure 2, bitstream portions belonging to pictures of equal presentation time may be part of a common access unit (AU). Furthermore, note that each picture 26 or subpicture 28 may be encoded into one or more bitstream portions 16. For example, picture 26 or subpicture 28 may be further subdivided into slices, each of which may be encoded into one or more bitstream portions. In the exemplary example of Figure 2, a given set of subpictures 28, i.e., a set of one or more subpictures transferred in the subpicture-specific video data stream 12 for decoding, has only one subpicture, i.e., a cross-hatched subpicture 28. Therefore, of the bitstream portions associated with picture 261 of layer L1, only the cross-hatched bitstream portion of subpicture 28 is transferred in the subpicture-specific video data stream 12.

[0036] In other words, when the device 10 takes over the bitstream portion 16 of the layer to the substream 12, it may be sensitive to whether or not the layer is encoded in units of two or more independently encoded subpictures 28.

[0037] For example, a predetermined set of one or more subpictures may be provided to the extractor 10 by an external means such as the API 20. That is, the extractor 10 may receive information on which subpictures of the picture 26 of the layer of the multilayer video data stream 14 should be transferred to the decoder 50.

[0038] For example, information that reveals which bitstream portion a subpicture 28 of two or more subpictures of a layer encoded to be subdivided into subpictures belongs to can be provided by a subpicture identifier, such as SH_subpic_ID. For example, this information, i.e., the subpicture identifier, can be provided in the header of each bitstream portion, such as a sliced ​​header. In other words, a bitstream portion 16 in which the video data of either picture 26 or subpicture 28 is encoded may have a subpicture identifier associated in its bitstream portion description data, for example by an index, indicating the subpicture to which that bitstream portion belongs.

[0039] In the example, the video data stream 14 may include an association table 30, such as subpicIDVal, which is used to associate subpicture identifiers in the bitstream portion with the spatial location of the sub-parts in the picture of the video 24.

[0040] According to one embodiment, the extractor 10 can perform a check to determine whether video has been encoded in a layer, either without subdividing it into subpictures or with subdividing it into subpictures, by evaluating the syntax elements signaled in the video data stream 14 for each layer. For example, each syntax element may be signaled in the sequence parameter set to which each layer is associated. For example, the syntax element may reveal several subpictures or subparts encoded in each layer. The syntax element may be included in picture or subpicture configuration data 22, and in this example, may be signaled in the video data stream 14, for example, the sequence parameter set (SPS). For example, each syntax element may be the sps_num_subpics_minus1 syntax element.

[0041] As shown in Figure 2, the video data stream 14 may optionally include an OLS display 18, which indicates a set of layers, e.g., an OLS, e.g., a set of layers that are indicated to be transferred at least partially in the sub-picture video data stream 12. As previously mentioned, the OLS display 18 may include, for example, a set of OLS that includes a set of layers indicated to be transferred in the sub-picture video data stream 12. For example, the API 20 indicates to the extractor 10 which of the OLS in the OLS display 18 should be transferred in the sub-picture video data stream 12 by, for example, indicating an index that presents one of the OLS in the OLS display 18.

[0042] The video data stream 14 may optionally further include decoder function requirement information 60, which may include, for example, buffer sizes such as picture size and CPB and / or DPB buffer sizes, and similar information such as HRD, DPD, and TPL information. The information may be provided in the video data stream 14 in the form of a list of its various versions, and the version applicable to a particular OLS is referenced by indexing. That is, for a particular extractable OLS, for example, an OLS indicated by an OLS display 18, an index presenting the corresponding HRD, DPD, and / or TPL information may be signaled.

[0043] The decoder function requirement information 60 may be information relating to sub-picture-specific video data streams 12 that can be extracted from the multi-layer video data stream 14. For example, the decoder function requirement information 60 or a part thereof may be determined by the encoder 40 by actually performing an extraction process, such that the extractor 10 performs an extraction process to extract each extractable sub-picture-specific video data stream 12. The encoder 40 can perform the extraction process at least to some extent to determine specific parameters of the decoder function requirement information 60.

[0044] It should be noted that the extractor 10 can adapt some data when forming data stream 14 from data stream 12. This possibility is shown in Figure 1 by using an apostrophe in the reference numeral 16 of the bitstream portion of the video data stream 12. The same applies to the components of stream 12 shown in Figure 2, which may also be, optionally, part of the video data stream 12 in Figure 1. These components may also be present in embodiments of video data streams 14 and 12 according to further embodiments described herein. For example, picture configuration data 22 in stream 14 indicates the picture size of a full picture 26 encoded in data stream 14, and picture configuration data 22' in stream 12 indicates the picture size of picture 26' encoded in stream 12 that is not then subdivided into subpictures, i.e., picture 261' from one subpicture 28, which is encoded in data stream 14. The bitstream packet 16, in which the actual picture content is encoded, such as the VCL NAL unit, can be carried over as is without any modification, at least with respect to its arithmetically coded portion. As can be seen from the figure, even the OLS representation 18' can be carried over from stream 14 to stream 12. It can also be left as is in stream 14. The related table 30 may also be modified accordingly. A modified version of any of the data items 18, 20, and 22 may be hidden / nested in stream 12, i.e., provided therein by encoder 40 and therefore simply used by the extractor to replace the corresponding rejected version that occurs in stream 12 when forming stream 14, or it may be interpreted on the fly by extractor 10 based on the overall information in stream 14 and placed in stream 12. When decoder 50 receives or is supplied with stream 12, it may no longer be able to see how the picture 261 of subpicture subdivision layer L1 once appeared in stream 14.

[0045] Depending on the check, the extractor 10 further determines whether specific layer-specific parameter sets, such as PPS and SPS, are adapted from scratch or newly generated, and it should be noted that as a result, certain parameters among them are adapted to the sub-picture data stream 12. Such parameters include the aforementioned picture size and sub-picture configuration, cropping window offset, and level indicator in the picture or sub-picture configuration data 22. Therefore, adaptations, such as the adaptation of the picture or sub-picture configuration data 22 in the data stream 12 in terms of picture size and sub-picture configuration, are not performed for layers that are not subdivided into subpictures, i.e., "when the video is encoded in each layer so that it is not subdivided into subpictures," but rather "when the video (24) is encoded in each layer so that the picture (26) of the video is subdivided into two or more subpictures (28)."

[0046] For example, when picture 260 is transferred in a separate video data stream 12 for subpictures, these pictures can be used as a reference layer for, for example, subpicture 28 of layer L1. For example, as described with respect to Figure 1, a multi-layer video data stream 14 can enable, for example, scalability of the data rate of the data stream. For this purpose, a base layer, for example layer L0, may contain video encoded at a first resolution, and an enhancement layer, for example layer L1, may contain information about video at a second resolution higher than the first resolution. For example, decoding picture 261 encoded in the enhancement layer may require information about the temporally arranged pictures in the base layer. In other words, the base layer can be a reference layer for the enhancement layer. Thus, in this scenario, when both the base layer and the enhancement layer are part of a layer set, it can be ensured that the reference pictures of the subpictures of the enhancement layer, which are part of the base layer, are available in the separate video data stream for subpictures. Furthermore, in examples where low-resolution video is encoded into a base layer so as not to be subdivided into subpictures, the bitstream portion of the base layer is consequently taken over by the subpicture-specific video data stream, according to the presented concept. Thus, it can be ensured that at least the low-resolution video of the entire video content is available in the subpicture-specific video data stream, and consequently, if the subpicture to be presented, and therefore to be decoded, is changed, at least the low-resolution picture in the base layer is immediately available to the decoder, allowing for a quick change of the subpicture.

[0047] In other words, when scalable coding is used in combination with subpictures, for example in a viewport-dependent 360-degree video streaming scenario, one common setup is to have a low-resolution base layer depicting the entire 360-degree video scene and an enhancement layer containing the scene at higher fidelity or spatial resolution, but the enhancement layer image is subdivided into several independently coded subpictures. In such a setup, it is possible to extract a single subpicture, for example, a portion of the picture corresponding to the client's line of sight, from the enhancement layer and the non-subpicture base layer and decode it together with the full 360-degree video base layer.

[0048] In the case of the bitstream configuration described above (subpictures including non-subpicture reference layers), state-of-the-art extraction processes such as the VVC specification will not produce a compliant bitstream. For example, when creating a temporary outBitstream from a direct copy of the inBitstream (for example, to derive OLS display 18' from OLS display 18) and then extracting the subpictures by index subpicIdx, the current VVC draft specification's subpicture subbitstream extraction process performs the following steps for each layer i of the extracted OLS:

[0049] - Removes all VCL NAL units from the outBitstream whose nuh_layer_id is not equal to the nuh_layer_id of the i-th layer and whose sh_subpic_id is not equal to SubpicIdVal[subpicIdx]. In other words, only VCL NAL units whose nuh_layer_id is equal to the nuh_layer_id of the i-th layer and whose sh_subpic_id is equal to SubpicIdVal[subpicIdx] are included in the extracted bitstream.

[0050] Please note that all VCL NAL units with a subpicture ID different from the subpicture to be extracted will be deleted.

[0051] Here, SubpicIdVal is an AU-specific mapping from the index Idx order of the subpictures in which they appear in the bitstream and associated signaling structure, and is an identifier ID value that is brought in as the syntax element sh_subpic_id in the slice header of the slice belonging to the subpicture and is used, for example, by the extractor to identify the subpicture during extraction.

[0052] Figure 3 shows an example of a multilayer video data stream 14, which may be an example of the video data stream 14 described in relation to Figures 1 and 2. According to Figure 3, a sequence of four access units of video is shown. Each of the pictures 261 in layer L1 is subdivided into two subpictures, a first subpicture indexed at index 0 and a second subpicture indexed at index 1, while the picture 260 in layer L0 is encoded without subdivision.

[0053] Conventionally, if the original bitstream shown at the top of Figure 3 is to be used for extracting a subpicture, e.g., subpicIdx 1, a non-compatible bitstream 12* is created because L1 lacks an interlayer reference picture (ILRP) from L0, as shown at the bottom of Figure 3. Note that the extraction is index-based because the subpicture ID may change during the bitstream process. Performing such a step will obviously result in the removal of all NAL units, e.g., the bitstream portion, from the non-subpicture layer because the subpicture index is 0.

[0054] Therefore, the embodiment according to the first aspect can enable the extraction of a subpicture-specific video data stream in a scenario in which the subpicture-specific video data stream includes a non-subpicture layer. In other words, the embodiment can enable the extraction of a subpicture-specific video data stream in which the subpicture may have a reference to a non-subpicture layer.

[0055] In other words, in contrast, embodiments according to the first aspect of the present invention can enable the extraction (i.e., removal) of NAL units from each layer in the OLS, depending on its subpicture configuration. For example, in one embodiment, the above removal step is conditioned, through the presence of subpictures, for example, the number of subpictures in layer SPS, as follows:

[0056] -If sps_num_subpics_minus1 is greater than 0, remove all VCL NAL units from outBitstream where nuh_layer_id is equal to the nuh_layer_id of the i-th layer and sh_subpic_id is not equal to SubpicIdVal[subpicIdx].

[0057] In particular, pay attention to the part that restricts the deletion of VCL NAL units to layers or pictures that have at least two subpictures, specifically "if sps_num_subpics_minus1 is greater than 0".

[0058] In another embodiment, additional steps in the extraction process described above for non-VCL NAL units of each layer, such as SPS / PPS adjustments for picture size, cropping window offset, level idc, and sub-picture configuration, are omitted based on the presence of sub-pictures, as follows:

[0059] The output subbitstream outBitstream is derived as follows: - The sub-bitstream extraction process specified in Appendix C.6 is invoked with inBitstream, targetOlsIdx, and tIdTarget as inputs, and the process's output is assigned to outBitstream. -If external means not specified herein are available for providing a replacement parameter set for the subbitstream outBitstream, replace all parameter sets with the replacement parameter set. - Otherwise, if a subpicture-level information SEI message is present in the inBitstream, the following applies: [...] For the i-th layer whose NumLayersInOls[targetOlsIdx] ranges from -1 to 0, the following applies: - If sps_num_subpics_minus1 is greater than 0, the following applies: [...] Layer-specific tasks such as adjusting SPS / PPS for levels, picture size, / / fit window, and sub-picture parameters, and deleting VCL NAL units. In this case as well, please note the part that restricts layer-specific tasks to layers or pictures that have at least two subpictures, specifically "when sps_num_subpics_minus1 is greater than 0".

[0060] Figure 4 shows an exemplary result of extracting the sub-picture video data stream 12 when the above embodiment is used in the extraction process that yields a suitable bitstream.

[0061] 2. Encoder 40, extractor 10 and video data stream 14 according to a second embodiment This section describes embodiments of the encoder 40, extractor 10, and video data stream 14 according to a second aspect of the present invention. The encoder 40, extractor 10, and video data stream 14 according to the second aspect may conform to the encoder 40, extractor 10, and video data stream 14 described with respect to Figure 1, or optionally conform to the first aspect described with respect to Figure 2 in Section 1.

[0062] Figure 5 shows an embodiment of the encoder 40, extractor 10, decoder 50, and video data stream 14, the video data stream 14 including at least a first layer L1 and a second layer L0, as described with respect to Figures 1 and 2, for example. The video sequence 24, picture 26, and subpicture 28 of the scenario illustrated in Figure 5 can follow the description of the scenario shown in Figures 1 and 2.

[0063] According to an embodiment of the second aspect, the multilayer video data stream 14 is composed of or includes bitstream portions 16, for example, NAL units, each of which belongs to one of the layers of the multilayer video data stream 14, as indicated, for example, by a layer ID associated with each bitstream portion. The multilayer video data stream 14 includes a first layer, for example layer L1 in Figure 5, and the bitstream portion 16 on its left side contains the first video 241 encoded such that the picture 261 of the first video 241 is subdivided into two or more subpictures 28, for example, half of the picture 261 of layer L1, which is divided into simple hatching and half cross-hatching. The subpictures 28 are encoded independently of each other in the bitstream portion 16 of the first layer L1, so that for each picture 261 of the first video 241, the subpicture 28 of each picture of the first video is encoded in different bitstream portions 16 of the first layer L1. For example, as already explained with respect to Figure 2, a simply hatched subpicture 28 is encoded in the simply hatched bitstream portion of the video data stream 14, and a cross-hatched subpicture 28 is encoded in the cross-hatched bitstream portion of the video data stream 14. The multi-layer video data stream 14 further includes a second layer, for example layer L0 in Figure 5, in which the second video 240 is encoded in its left bitstream portion. For example, the second video 240 covers at least the video content of the first video 241 and relates to the temporally aligned video content, so that inter-layer prediction of the first video based on the second video is feasible. For example, the video content of picture 260 in layer L0 can cover the video content of the temporally aligned picture 261 in layer L1. In this example, the second video 240 is encoded into the second layer L0 of the video data stream 14 so as not to be subdivided into subpictures with respect to the spatial picture region of the second video 240 that spatially corresponds to the first video 241, but this is not necessarily the case.In other words, the number of sub-parts encoded in the second layer L0 may be one. To put it another way, each picture in the second layer L0 can be encoded as one sub-picture, or, for example, if subdivided into sub-parts, one sub-part can completely cover the footprint of the first video. The bitstream portion of the first layer, i.e., the cross-hatched portion and the simple hatched portion, encodes the first video 241 using vector-based predictions from a reference picture. Furthermore, the bitstream portion of the first layer is encoded such that the first video 241 includes (for example, at least) the picture of the second video 240 as a reference picture, and the vector composed of or signaled by the bitstream portion of the first layer is encoded using the vector encoded by the bitstream portion of the first layer, such that the picture 261 and the reference picture 260 of the first video are scaled and offset according to the size and position of the scaling window 72 in the multilayer video data stream 14, respectively, for use in vector-based user prediction.

[0064] For example, the picture of the second video 240 is included in at least the reference picture, using a reference picture encoded using vector-based prediction in the bitstream portion of the first layer. For example, picture 260 in Figure 5 could be the reference picture of picture 261. In other words, picture 260 of the second video 240 can function as an inter-layer reference picture for encoding picture 261 of the first video 241. For example, vector-based prediction may be included in a motion vector-based prediction framework, where the term “motion” generally means the fact that picture content in the first video 241 predicted from a certain reference picture is spatially moved by some vector from its corresponding picture content in the reference picture from which it is predicted. That is, the picture content is moved, and the reference picture could be, for example, a previously encoded picture in the first video 241, or a concurrently or temporally aligned picture in the second video 240. For example, the picture 26 drawn overlapping in the first video 241 and the second video 240 in Figure 5 could be such a picture.

[0065] The offset of vector 70, depending on the size and position of the scaling window 72, may be performed according to the positional offset between scaling windows, for example, between the scaling window 720 of layer L0 and the scaling window 721 of layer L1. Furthermore, for example, the scaling of vector 70, depending on the size and position of the scaling window 72, may be performed according to the size ratio between scaling windows, for example, between scaling windows 720 and 721.

[0066] For example, the size and position of the scaling window 72 may be signaled within the multilayer video data stream 14 by the scaling window display 72, for example, for each picture. For example, the size and position of the scaling window 72 may be signaled by the positions of two opposing corners of the scaling window, or the position of one corner, as well as the display of the scaling window's dimensions in x and y. Note that it is also possible to signal a particular position of the scaling window in a way other than using one or more coordinates and their dimensions, for example, by indicating that the scaling window coincides with the picture boundary using a corresponding flag. In the latter case, the example shows that the position of the scaling window is indicated using coordinates, etc., only if the flag indicates that the position of the scaling window does not coincide with the picture boundary.

[0067] The extractor 10 according to the second embodiment is configured to take over the bitstream portion belonging to the second layer L0 from the multilayer video data stream 14 to the sub-picture video data stream 12 so that the video 240 of each layer is fully encoded in the sub-picture video data stream 12. For example, the extractor 10 can transfer the bitstream portion associated with layer L0 regardless of whether it transfers the entire bitstream portion 16 of layer L0 or which of the two or more subpictures of the first layer L1 is transferred to the sub-picture video data stream 12, that is, regardless of which subpicture of the first layer L1 belongs to a predetermined set of one or more subpictures described below. In other words, the device 10 according to the second embodiment can take over the bitstream portion of the second layer L0, for example, as described with respect to Figures 1 and 2 for the case of layer L0 that is encoded so as not to be subdivided into subpictures.

[0068] The extractor 10 in the second embodiment is further configured to transfer each bitstream portion in which a subpicture 28 belonging to a predetermined set of one or more subpictures that belongs to the first layer L1 is encoded from the multilayer video data stream 14 to the subpicture-specific video data stream 12. For example, the extractor 10 may drop, retain, or remove, i.e., not transfer, each bitstream portion in which a subpicture 28 belonging to each layer, for example, layer L1, but not belonging to a predetermined set of one or more subpictures is encoded. In other words, the extractor 10 in the second embodiment can selectively transfer bitstream portions of the first layer L1 that belong to a predetermined set of subpictures, as described with respect to the first embodiment (see Figures 1 and 2). As described with respect to Figures 1 and 2, a predetermined set of one or more subpictures may be indicated by information provided to the extractor 10 by the API.

[0069] The extractor 10 is further configured to adapt the scaling window signal 74 to the first layer L1 and / or the second layer L0 in the subpicture-specific video data stream 12 such that the spatial region of the scaling window 720' for the second video picture spatially corresponds to the spatial region of the scaling window 721' for a predetermined set of one or more subpictures.

[0070] As shown in Figure 5, after the extraction of the subpicture-specific video data stream 12, a subpicture defined by a predetermined set of one or more subpictures is adapted as a result of the extraction to become or constitute the picture 261' of the first layer L1 of the data stream 12, for example, the subpicture 28 in Figure 5. The extractor 10 can adapt the scaling window signaling such that the scaling of the vector 70 of the second layer L0 of the subpicture-specific video data stream 12 after scaling using the scaling windows 720', 721' is associated with or mapped to the vector 70 having size and position in the array picture 261' of the first layer L1 of the subpicture-specific video data stream 12, the size and position of which correspond to the size and position of the vector 70 of the picture in the second layer L0 in the multi-layer video data stream 14, to which the vector 70 is associated or mapped using the scaling windows 720, 721. In other words, the extractor 10 can adapt the scaling window signal 74 so that the scaling windows 720', 721' determine the vector 70 of a given position within picture 261' of the first layer L1 of the data stream 12. As a result of this determination, the vector 70 corresponds to the size and position vector of the corresponding position within picture 261 of layer L1 of the multilayer video data stream 14 (corresponding to the position within picture 261' in terms of content, not relative to the picture boundary). However, since the size of picture 261' may change relative to picture 261, the scaling window may need to be adapted or shifted with respect to those positions relative to the picture boundary 261' if it is defined relative to the picture boundary.

[0071] According to one embodiment, the extractor 10 can adapt the scaling window signal 74 of the first layer L1 such that the scaling window 721' of the first layer L1 corresponds in position and size to the scaling window 720' of the second layer. This example is shown in Figure 5 by the scaling windows 720' and 721' drawn with solid lines.

[0072] Alternatively, the extractor 10 can adapt the scaling window signaling

[0073] In a further embodiment, the extractor 10 can adapt scaling window signals to both the first and second layers, that is, it can adapt the scaling windows 720' and 721' so that their perimeters coincide.

[0074] For example, the extractor 10 can adapt the scaling window signal 74 by using the sub-picture configuration data 22 in the data stream 14. In other words, the extractor 10 can adapt the sub-picture configuration data 22 in the data stream 14 to provide the sub-picture configuration data 22' in the data stream 12. For example, the sub-picture configuration data 22 may include a representation of the position and size of the scaling window. The extractor 10 can adapt the position and size of the scaling window in the configuration data 22 to obtain the sub-picture configuration data 22'.

[0075] Note that the same notes provided for Figure 2 regarding the concealed adaptation, incorporation, or new generation of the bitstream portion by the extractor 10 also apply to Figure 5 (see the apostrophe used for the bitstream portion of the data stream). Here, information 74 is another example of information that is concealed by the encoder or newly adapted or generated by the extractor on the fly.

[0076] Figure 6 shows the above example where the scaling window 720' of the second layer L0 is applied by the extractor 10. In other words, Figure 6 shows an example of adjusting the scaling window when extracting the right-hand subpictures of layers L1 and L0.

[0077] The following is an exemplary specification of an embodiment for adapting the scaling window.

[0078] The output subbitstream outBitstream is derived as follows:

[0079] - The sub-bitstream extraction process specified in Appendix C.6 is invoked with inBitstream, targetOlsIdx, and tIdTarget as inputs, and the process's output is assigned to outBitstream. - If external means not specified herein are available for providing a replacement parameter set for the subbitstream outBitstream, replace all parameter sets with the replacement parameter set. - Otherwise, if a subpicture-level information SEI message is present in the inBitstream, the following applies: [...] For the i-th layer whose NumLayersInOls[targetOlsIdx] ranges from -1 to 0, the following applies: - If sps_num_subpics_minus1 is greater than 0, the following applies: [...] Adjusting SPS / PPS for level, picture / / size, fitting window, and subpicture parameters, / / and other layer-specific tasks such as deleting VCL NAL units, the variables subPicLeftPos, subPicRightPos, subPicTopPos, and subPicBottomPos are derived as follows: subPicLeftPos=sps_subpic_ctu_top_left_x[subpicIdx]*CtbSizeY (C.XX) subPicRightPos=subPicLeftPos+(sps_subpic_width_minus1[subpicIdx]+1)*CtbSizeY (C.XX) subPicTopPos=sps_subpic_ctu_top_left_y[subpicIdx]*CtbSizeY (C.XX) subPicBottomPos=subPicTopPos+(sps_subpic_height_minus1[subpicIdx]+1)*CtbSizeY (C.XX) [...] / / Adjustments to SPS / PPS regarding levels, picture size, fitting window, and sub-picture parameters, / / Remaining layer-specific tasks such as deleting VCL NAL units - Otherwise (if sps_num_subpics_minus1 is 0, the following applies): - Rewrite the values ​​of pps_scaling_win_left_offset, pps_scaling_win_right_offset, pps_scaling_win_top_offset, and pps_scaling_win_bottom_offset for all referenced PPSNAL units to be equal to subPicLeftPos, (ppsPicWidth-subPicRightPos), subPicTopPos, and (ppsPicHeight-subPicBottom), respectively.

[0080] The second embodiment described with respect to Figures 5 and 6 allows for the combination of cross-layer prediction and extraction of sub-picture video data streams such as data stream 12, since the vector 17 used in vector-based prediction can be correctly mapped in the decoding of the extracted sub-picture video data stream 12.

[0081] In other words, the original bitstream contains control information that guides the inter-layer prediction to apply the correct upscale filter and MV offset through the so-called scaling window of each PPS, but in the sub-bitstream resulting from the extraction of one EL subpicture + full BL, it is important to adjust so that these higher-level parameters can be correctly decoded. Therefore, as part of the aspects of this chapter, the scaling window offset parameter is adjusted to compensate for the dropped subpicture of the enhancement layer with respect to the derivation of the MV offset and scaling factor. Figure 6 shows an example in which the enhancement layer contains two subpictures. How the scaling windows of the base layer L0 and the enhancement layer L1 are adjusted so that the correct inter-layer prediction is performed on the decoder side as was initially performed on the encoder side, i.e., the scaling window of the layer containing the subpicture (L1) is rewritten to include only the extracted full coded image region corresponding to the extracted subpicture, as shown on the right side of Figure 6. On the other hand, the layer without a subpicture (L0) is rewritten to include only the picture region corresponding to the extracted subpicture of the higher layer. In other words, it includes the samples necessary for predicting the remaining L1 coded picture region or subpictures. The same applies to the example above, and accordingly, a scaling window 721' of the first layer L1 is applied, in which one or more subpictures defined by a given set of one or more subpictures are extracted by the extractor 10.

[0082] 3. Encoder 40 and decoder 50 according to a third embodiment This section describes an embodiment according to a third aspect with reference to Figure 1. The details described in Section 0 may be optionally applied to the embodiment according to the third aspect. Furthermore, the details of data streams 12 and 14 shown in Figures 2 and 5 may also be included in the embodiment according to the third aspect, in which case the corresponding descriptions in Sections 1 and 2 shall apply as appropriate.

[0083] As mentioned above, encoder 40 can apply cross-layer prediction to encoding picture 26 of video 24. For example, encoder 14 can use cross-layer prediction for encoding picture 261 of layer L1. Encoder 40 can use picture 261 and picture 260, which are time-positioned, for cross-layer prediction in encoding picture 261. For example, encoder 40 can use time-motion vector prediction (TMVP) to predict the motion vector for encoding picture 261. Other examples of methods that can be used for cross-layer prediction are optical flow prediction refinement (PROF) or motion vector wraparound, which will be explained in more detail below. As mentioned above, the picture used to predict the picture to be encoded is sometimes called the reference picture of the picture to be encoded.

[0084] If TMVP requires resampling of reference picture motion vector (MV) storage, for example, due to single-layer resolution changes or multi-layer spatial scalability (which may be provided by multi-layer video sequences 24), the gating variable RprConstraintsActive[currentPic][refPic] is used to disable the use of TMVP as follows:

[0085] The picture referenced by sh_collocated_ref_idx is assumed to be the same across all slices of the coded picture, and RprConstraintsActive[sh_collocated_from_l0_flag?0:1][sh_collocated_ref_idx] being equal to 0 is a requirement for bitstream conformance.

[0086] Note - The above constraints require that the array picture has the same spatial resolution and scaling window offset as the current picture.

[0087] Figure 7 shows an example of extracting video data streams by subpicture from a multilayer bitstream, where picture 261 of the first layer L1 is subdivided into subpictures 28, and array picture 260 of the second layer L0 is not subdivided, i.e., encoded into a single subpicture. As previously stated, embodiments within this section are described with respect to Figure 1, so that the descriptions of the elements and corresponding reference numerals of Figure 1 may apply. Furthermore, the descriptions of the corresponding elements with respect to Figure 2 may also apply to embodiments according to a third aspect.

[0088] As explained above (sections 1 and 2, etc.), when layers containing subpictures are combined in OLS with layers that do not contain subpictures, and spatial scalability is not used, as shown via the original bitstream on the left side of Figure 7, the encoder can also use TMVP.

[0089] However, a problem arises when subpictures are extracted, as shown on the right side of Figure 7, and the scaling window is adjusted during the extraction process so that cross-layer prediction of sample values ​​can work properly. After extraction, the derivation conditions for RprConstraintsActive[currentPic][refPic] have been changed and TMVP is prohibited, as can be seen from the following section of the current VVC draft specification, in particular where "RprConstraintsActive[i][j]" is defined.

[0090] Otherwise, the reference picture lists RefPicList[0] and RefPicList[1], the scaling ratios of the reference picture RefPicScale[i][j][0] and RefPicScale[i][j][1], and the scaling flag RprConstraintsActive for the reference picture are used.

[0091] [j] and RprConstraintsActive[1][j] are derived as follows: for(i=0;i<2;i++){ for(j=0,k=0,pocBase=PicOrderCntVal;j <num_ref_entries[i][RplsIdx[i]];j++){ [...] fRefWidth is set to be equal to PicOutputWidthL of the reference picture RefPicList[i][j]. fRefHeight is set to be equal to PicOutputHeightL of the reference picture RefPicList[i][j]. refPicWidth, refPicHeight, refScalingWinLeftOffset, refScalingWinRightOffset, refScalingWinTopOffset, And refScalingWinBottomOffset is the base picture pps_pic_width_in_luma_samples in RefPicList[i][j], pps_pic_height_in_luma_samples, pps_scaling_win_left_offset, pps_scaling_win_right_offset, These values ​​are set to be equal to the values ​​of pps_scaling_win_top_offset and pps_scaling_win_bottom_offset, respectively. RefPicScale[i][j][0]=((fRefWidth<<14)+(PicOutputWidthL>>1)) / PicOutputWidthL RefPicScale[i][j][1]=((fRefHeight<<14)+(PicOutputHeightL>>1)) / PicOutputHeightL RprConstraintsActive[i][j]=( pps_pic_width_in_luma_samples!=refPicWidth||| pps_pic_height_in_luma_samples!=refPicHeight||| pps_scaling_win_left_offset!=refScalingWinLeftOffset||| pps_scaling_win_right_offset!=refScalingWinRightOffset||| pps_scaling_win_top_offset!=refScalingWinTopOffset||| pps_scaling_win_bottom_offset!=refScalingWinBottomOffset) } } The part where RprConstraintsActive[i][j] is defined defines the conditions corresponding to the note mentioned earlier in this section, "The above constraints (reference picture constraints) require that the sequence picture has the same spatial resolution and the same scaling window offset as the current picture." Any syntax-based prediction, such as TMVP, is prohibited if the sizes of the current picture and the reference picture are different, for example, in single-layer resolution changes of CVS (encoded video sequences) or inter-layer prediction of spatially scalable bitstreams, because it requires special precautions for the prediction mechanism involved. For example, syntax such as MV candidates is stored in a special normal storage grid, such as 16x16 samples, for the purpose of syntax prediction, and the (CTU-) block boundaries of the current picture are aligned with the boundaries of this normal storage without changing the resolution. Thus, it is relatively easy to find the MV storage location in the same place, but if a sampling ratio that is not equal to 1 occurs (i.e., a change in resolution between the reference and current picture), or if an arbitrary position offset is used (via different scaling window positions between the reference and current picture), the above boundaries may no longer be aligned. Deriving the correct MV storage location for prediction would be an additional implementation burden. In the case of VVC, since this is assumed to rarely occur in the primary use case of single-layer resolution changes, a design trade-off was made to avoid incorporating special precautions for this case. Multilayer applications are not the primary target of the codec.

[0092] Also, note that motion vector prediction currently does not use scaling windows. Scaling windows are only used for sample prediction based on motion vectors used for blocks. In other words, scaling windows define how motion vectors are offset. However, in motion vector prediction, scaling windows are ignored, and the top-left block of one picture is mapped to the top-left block of another picture (a block placed in the same location). Therefore, in the illustrated example, once extraction is performed, completely different blocks will be placed in the same location compared to before extraction.

[0093] Another tool affected similarly to TMVP is Optical Flow Prediction Refinement (PROF). This is part of the VVC's affine motion prediction model and is introduced to more efficiently represent certain types of motion (e.g., zoom in or zoom out, rotation). When this mode is selected, CU-based affine motion compensation prediction is applied. The CU's affine motion model is described by an MV (4-parameter model) with two control points at the top-left and top-right corners, or an MV (6-parameter model) with three control points at the top-left, top-right, and bottom corners. To achieve finer granularity of motion compensation, Optical Flow Prediction Refinement (PROF) is used to refine each luminance prediction subblock, aiming for the effect of sample-level motion compensation. Prediction samples within a luminance subblock are refined by adding a difference derived based on the gradient and a sample-based motion vector difference. PROF is not applied to chroma samples.

[0094] However, in affine motion prediction processing, the PROF-related flag cbProfFlagLX is derived based on RprConstraintsActive[i][j], and differences in the derivation before and after subpicture extraction can lead to encoder / decoder mismatches.

[0095] Another tool affected is motion vector wraparound, used for certain types of content (360-degree video in equirectangular projection format). For example, when an object leaves the picture plane coded at the vertical boundary of the picture, it will typically re-enter the picture plane at the opposite vertical boundary due to the nature of the underlying projection. This fact facilitates motion vector wraparound by using the opposite boundary for sample padding when the motion vector points outside the picture.

[0096] However, in the sample prediction process, a flag (refWraparoundEnabledFlag) that controls motion vector wraparound is derived based on RprConstraintsActive[i][j]. If there is a difference in the derivation before and after subpicture extraction, it will cause an encoder / decoder mismatch, similar to the case of TMVP or PROF.

[0097] A third aspect provides a concept for controlling the use of inter-layer prediction tools such as TMVP, PROF, and MV wraparound, thereby avoiding the above-mentioned problems that may occur when separate video data streams are extracted for each subpicture. Embodiments according to the third aspect can avoid these problems by controlling the use of these inter-layer prediction tools in encoding the original bitstream, i.e., the multi-layer video data stream 14, so as to avoid different behavior before and after subpicture extraction.

[0098] According to an embodiment of the third aspect, the encoder 40 for encoding video 24 into a multi-layer video data stream 14 supports independent coding on a sub-picture basis. Furthermore, the encoder 40 is configured to encode a first version 241 of video 24 into the first layer L1 of the multi-layer video data stream 14 using a set of inter-layer prediction tools for the second layer L0 of the multi-layer video data stream 14 and using a first configuration setting for independent coding on a sub-picture basis. The encoder 40 according to the third aspect is further configured to encode a second version 240 of video 24 into the second layer L0 of the multi-layer video data stream 14 using a second configuration setting for independent coding on a sub-picture basis. The encoder 40 is configured to check whether the first and second configuration settings have a predetermined relationship. If the first and second configuration settings do not have a predetermined relationship, the encoder 40 disables or refrains from using a predetermined subset of one or more inter-layer prediction tools.

[0099] For example, the first and second configuration settings of the first layer L1 and the second layer L0 may take into account whether (or not) each layer is encoded in such a way that it is subdivided into subpictures. In other words, the first and second configuration settings may include information about several subpictures (e.g., identifying one and multiple, where the number of subpictures means that the picture is coded in such a way that it is not subdivided). For example, the first and second configuration settings of the first layer L1 and the second layer L0 may include information about one or more of the subpicture boundaries, subpicture identifiers, and subpicture boundary processing.

[0100] For example, a predetermined relationship may be one in which a first configuration setting and a second configuration setting have a predetermined relationship if one or more or all of the following conditions are met. - The first video and the second video are each subdivided into multiple subpictures and encoded into the first and second layers, respectively. - The first video and the second video are encoded into the first and second layers, respectively, such that they are subdivided into an equal number of subpictures between the first and second videos. - The first video and the second video are encoded into the first and second layers, respectively, so that they are subdivided into an equal number of subpictures between them, and the boundaries of the subpictures are modified to coincide spatially. - The first and second videos are encoded into the first and second layers, respectively, so that they are subdivided into several subpictures, and the subpicture IDs signaled to the subpictures in the multilayer video data stream match. - The first and second videos are encoded into the first and second layers, respectively, so that each is subdivided into several subpictures, and the boundary processing of the subpictures matches.

[0101] For example, a given subset of one or more interlayer predictions includes one or more interlayer motion vector predictions, such as TMVP, optical flow refinement of vector-based interlayer predictions, and motion vector wraparound in vector-based interlayer predictions.

[0102] For example, the encoder 40 can signal a first configuration setting and a second configuration setting to a multi-layer video data stream.

[0103] In the example, the encoder 40 can enable or utilize a predetermined subset of one or more inter-layer prediction tools, or a subset thereof, if the first configuration setting and the second configuration setting have a predetermined relationship.

[0104] According to an embodiment of the third aspect, the decoder 50 for decoding a multilayer video data stream is configured to decode a first version 241 of video 24 from the first layer L1 of the multilayer video data stream, using a set of interlayer prediction tools for prediction from the second layer L0 of the multilayer video data stream, and using a first configuration setting for independent coding on a sub-picture basis. Note that the multilayer video data stream may be a sub-picture video data stream 12 or a multilayer video data stream 14. In other words, although Figure 1 shows the decoder 50 decoding a sub-picture video data stream 12, the decoder 50 may also decode a multilayer video data stream 14, for example, as it may be transferred by the extractor 10 when another OLS is selected, as in the scenario shown in Figure 1. When the multilayer video data stream decoded by decoder 50 may correspond to a sub-picture-specific video data stream 12, the first version 241 of video 24 may differ from the version described with respect to encoder 40, since the multilayer video data stream may have undergone an extraction process as described with respect to Figures 1 to 7. Thus, the first version 241 of video referred to with respect to decoder 50 may differ from the one referred to with respect to encoder 40, in the example, due to the omission of one or more sub-pictures. The same applies to the first configuration setting with respect to decoder 50, relating to the picture region of the picture encoded in the extracted / remaining / normally removed portion of the data stream decoded by decoder 50. For example, if the multilayer video data stream entering the decoder relates to only one sub-picture 28 of the two previous sub-pictures containing the picture, the first configuration data will only indicate a non-sub-picture subdivision.

[0105] A decoder 50 according to a third embodiment may be configured to decode a second version 240' of video 24' from a multilayer video data stream, i.e., a second layer L0 of the video data stream provided to the decoder 50, using a second configuration setting for independent coding on a sub-picture basis. The decoder 50 is configured to check whether the first and second configuration settings have a predetermined relationship. If the first and second configuration settings do not have a predetermined relationship, the decoder 50 may disable a predetermined subset of one or more inter-layer prediction tools as described with respect to the encoder 40.

[0106] According to one embodiment, the decoder 50 may be configured for independent decoding on a subpicture basis, including vector clipping and / or boundary padding at subpicture boundaries.

[0107] For example, the variable RprConstrainsActive introduced above may represent a comparison between the first configuration setting and the second configuration setting. For instance, the encoder 40 and decoder 50 can derive this variable based on whether the first configuration setting and the second configuration setting have a predetermined relationship, and can use this variable to determine whether to use a predetermined subset of the inter-layer prediction tool by setting the variable as appropriate.

[0108] In one embodiment, the derivation of RprConstraintsActive[i][j] is adjusted as follows: Otherwise, the reference picture lists RefPicList[0] and RefPicList[1], the scaling ratios of the reference picture RefPicScale[i][j][0] and RefPicScale[i][j][1], and the scaling flag RprConstraintsActive for the reference picture are used.

[0109] [j] and RprConstraintsActive[1][j] are derived as follows: for(i=0;i<2;i++){ for(j=0,k=0,pocBase=PicOrderCntVal;j <num_ref_entries[i][RplsIdx[i]];j++){ [...] fRefWidth is set to equal to PicOutputWidthL of the reference picture RefPicList[i][j]. fRefHeight is set to be equal to PicOutputHeightL of the reference picture RefPicList[i][j]. refPicWidth, refPicHeight, refScalingWinLeftOffset, refScalingWinRightOffset, refScalingWinTopOffset, And refScalingWinBottomOffset is the base picture pps_pic_width_in_luma_samples in RefPicList[i][j], pps_pic_height_in_luma_samples, pps_scaling_win_left_offset, pps_scaling_win_right_offset, Set to be equal to the values ​​of pps_scaling_win_top_offset and pps_scaling_win_bottom_offset, respectively. fRefSubpicsEnabled is set to equal to sps_subpic_info_present_flag of the base picture RefPicList[i][j] (i.e., it checks whether the ILMVP base picture is subdivided into subpictures). RefPicScale[i][j]

[0110] =((fRefWidth<<14)+(PicOutputWidthL>>1)) / PicOutputWidthL RefPicScale[i][j]

[0111] =((fRefHeight<<14)+(PicOutputHeightL>>1)) / PicOutputHeightL RprConstraintsActive[i][j]=( pps_pic_width_in_luma_samples!=refPicWidth|| pps_pic_height_in_luma_samples!=refPicHeight|| pps_scaling_win_left_offset!=refScalingWinLeftOffset|| pps_scaling_win_right_offset!=refScalingWinRightOffset|| pps_scaling_win_top_offset!=refScalingWinTopOffset|| pps_scaling_win_bottom_offset!=refScalingWinBottomOffset|| sps_subpic_info_present_flag!=fRefSubpicsEnabled(Therefore, if subpicture information exists for only one of the current picture or the reference picture, i.e., if the presence or absence of subpicture information differs between the two, then certain inter-layer prediction tools such as ILMVP will be disabled.) } } Alternatively, in one embodiment, a requirement for bitstream conformance is that the bitstream does not enable TMVP, PROF, or MvWrapAround in predictions between a current picture and a reference picture having different subpicture configurations (enabled vs. disabled), and the bitstream carries indications of the above constraints.

[0112] In the case of TMVP, the above derivation does not affect the decoder in the example, but in the case of PROF and MvWrapAround, the decoder may behave differently when the modified derivation is executed and checked by the respective tools.

[0113] Alternatively, such checks may be performed depending on the number or identifiers of subpictures within the current picture and the reference picture, the characteristics of the boundary processing of such subpictures, or other further features.

[0114] 4. Encoder 40 and extractor 10 according to a fourth embodiment This section describes an embodiment according to the fourth aspect with reference to Figure 1. The details described in Section 0 may be optionally applied to the embodiment according to the fourth aspect. Furthermore, the details of data streams 12 and 14 shown in Figures 2 and 5 may be optionally included in the embodiment according to the third aspect, in which case the corresponding descriptions in Sections 1 and 2 shall apply as appropriate. Optimally, features described with respect to the first, second, and third aspects may also be combined with those of the fourth aspect.

[0115] When subpictures are extracted from a multilayer video data stream, such as video data stream 14, a percentage (e.g., indicated by ref_level_fraction_minus1) used to determine the limits on a particular level of the extracted subbitstream (e.g., bitrate, CPB size (encoded picture buffer), number of tiles, etc.) may be signaled or derived. For example, these level constraints may raise or relate to decoder functional requirements that the decoder must satisfy in order to decode the video data stream or the video sequence described by the OLS on which the functional requirements are based.

[0116] Traditionally, if this ratio does not exist, it is estimated that the size of the sub-picture is equal to the size of the complete picture within that layer, as follows:

[0117] If it does not exist, the value of ref_level_fraction_minus1[i][j] is presumed to be equal to Ceil(256*SubpicSizeY[j]÷PicSizeInSamplesY*MaxLumaPs(general_level_idc)÷MaxLumaPs(ref_level_idc[i])-1).

[0118] However, this mechanism has several problems. Firstly, the above estimation does not take into account the fact that a bitstream can host multiple layers. For example, one level limit has a bitstream range (e.g., bitrate and CPB size), while another level limit has a layer range (e.g., number of tiles), so the ratio SubpicSizeY[j]÷PicSizeInSamplesY is insufficient.

[0119] Therefore, this is part of the present embodiment, in one embodiment the estimation is performed in a manner that takes into account all layers in the OLS, i.e., by derived bitstream.

[0120] Therefore, as part of this embodiment, in one example, the estimation of the percentage per subpicture is modified to also incorporate the effect of the level of the base layer without subpictures, as follows:

[0121] If it does not exist, the value of ref_level_fraction_minus1[i][j] is assumed to be equal to max(255,Ceil(256*SumOfSubpicSizeY[j]÷SumOfPicSizeInSamplesY*MaxLumaPs(general_level_idc)÷MaxLumaPs(ref_level_idc[i])-1)). If the maximum function is omitted, it can be Ceil(256*SumOfSubpicSizeY[j]÷SumOfPicSizeInSamplesY*MaxLumaPs(general_level_idc)÷MaxLumaPs(ref_level_idc[i])-1).

[0122] Here, SumOfSubpicSizeY and SumOfPicSizeInSamplesY are the sum of all samples associated with each subpicture in all layers within the OLS.

[0123] Figure 8 shows an example of extracting a video data stream 14 by subpicture, following the scenarios shown in Figures 1, 2, and 5, for example. Figure 8 shows the relative change in the output picture area between the original data stream, e.g., the video 24 of the data stream 14 before extraction, and the video sequence 24′ of the data stream 12 by subpicture after extraction.

[0124] The second issue concerns the fact that subpictures can be used in combination with a base layer that does not contain a corresponding subpicture, as shown in Figure 8 where L0 does not contain a subpicture.

[0125] To cover this case, or to provide a more accurate level limit, in another embodiment, for example, the estimation of ref_level_fraction_minus1 is performed as follows:

[0126] Here, SumOfSubpicSizeY and SumOfPicSizeInSamplesY are the sum of all samples associated with each subpicture for all layers in the OLS and all layers in the OLS that do not have subpictures.

[0127] In another embodiment, the estimation of ref_level_fraction_minus1[i][j] is invariant with respect to the state of the technology (it remains the layer-specific fraction of the layer containing the subpictures). Instead, instead of the above modification, the OLS-specific fraction variable is derived from the specific fraction per layer as follows:

[0128] Ols_fraction_nominator[i][j][k]=0 Ols_fraction_denominator[i][j][k]=0 k-Ols{ Ols_fraction_nominator[i][j][k]=+PicSizeInSamplesY[layer]*(sps_subpic_info_present_flag[layer]?ref_level_fraction_minus1[i][j]:255) Ols_fraction_denominator[i][j][k]=+PicSizeInSamplesY[layer] } OlsRefLevelFraction[i][j][k]=Ols_fraction_nominator[i][j][k] / Ols_fraction_denominator[i][j][k] Here, k is the OLS index, i is the base level index, and j is the sub-picture index.

[0129] In the above embodiments, an equal rate distribution can be expected among all layers. To allow the encoder to freely determine the rate distribution between subpicture layers and non-subpicture layers, in another embodiment of the present invention, the proportion of layers that do not contain subpictures within the bitstream's OLS is explicitly signaled. That is, at a given reference level, the proportion of the bitstream's OLS that all non-subpicture layers together is signaled. [Table 1] non_subpic_layers_fraction[i] specifies the percentage of the bitstream / OLS-specific level limit associated with the bitstream / OLS layers for which sps_num_subpics_minus1 is equal to 0. non_subpic_layers_fraction[i] is equal to 0 if vps_max_layers_minus1 is equal to 0, or if there are no layers in the bitstream / OLS for which sps_num_subpics_minus1 is equal to 0.

[0130] The variable OlsRefLevelFraction[i][j] for the j-th subpicture of the i-th ref_level_idc is set to equal to non_subpic_layers_fraction[i] + (255 - non_subpic_layers_fraction[i]) ÷ 255 * ref_level_fraction_minus1[i][j] + 1.

[0131] OlsRefLevelFraction[i][j] is used per OLS, as described in Section 5.

[0132] According to an embodiment of the fourth aspect, the multilayer video data stream 14 includes a set of layers, for example, an OLS such as an OLS selected for decoding. For example, the OLS may include layers L0 and L1 as shown in Figure 1. It should be noted that the multilayer video data stream encoded by the encoder 40 may consist not only of this set of layers, i.e., this one set of layers, but also of further OLSs, and may consist of further layers other than layers L0 and L1 shown in Figure 1. The encoder 40 according to the fourth aspect is configured to encode unreduced versions of the set of layers, for example, layers L0 and L1, into the multilayer video data stream 14. Video versions, for example, video sequences 240 and 241, are encoded into unreduced versions of the set of layers in units of one or more independently coded subpictures per layer. The expression "in units of one independently coded subpicture" means, for example, a layer coded without being subdivided into subpictures. That is, a picture such as picture 260 of video version 240 forms one subpicture. For example, a layer set may contain one or more such unsubdivided layers and one or more subpicture-divided layers.

[0133] The encoder 40 according to the fourth embodiment is further configured to encode reference decoder functional requirements related to decoding the un-reduced version of the set of layers into the multi-layer video data stream 14. Furthermore, the encoder 14 according to the fourth embodiment is configured to encode into the multi-layer video data stream (14) for each layer of the set of layers information about the picture size of the picture encoded in each layer of the un-reduced version of the set of layers (such as the information 22 shown in Figure 2), and for each of the one or more independently encoded subpictures of each layer, the subpicture size. The encoder 40 according to the fourth embodiment is configured to determine decoder functional requirements, e.g., OlsRefLevelFraction[i][j], relating to decoding each reduced version of the set of layers, for each of the set of layers from which the subpicture-related portion of at least one layer has been removed, e.g., subpicture-specific video data stream 12 or further versions of the video data stream 12 that do not necessarily constitute the entire multi-layer video data stream 14 (note that OlsRefLevelFraction[i][j] is specific to the reduced version relating to the j-th subpicture).

[0134] The encoder 40 determines the decoder functionality requirements for each reduced version of the set of layers by scaling the reference decoder functionality requirements with a coefficient determined by using the quotient obtained by dividing the sum over the subpicture size of independently coded subpictures contained in each reduced version of the set of layers by the sum over the picture size. In other words, a reduced version of a set of layers means a version of the set of layers from which bitstream portions related to other subpictures, such as subpictures that are not extracted, i.e., bitstream portions that are not part of a given set of sublayers to be transmitted in the subpicture-specific video data stream 14, have been removed.

[0135] For example, the encoder 40 provides the decoder function requirements for the multilayer video data stream 14, which can be compared with the decoder function requirements 60, for example, as shown in Figure 2.

[0136] Decoder functionality requirements information can be used as an indicator of whether the extracted bitstream can be decoded by the decoder receiving the bitstream. For example, decoder functionality requirements information can be used by a decoder receiving the bitstream to set up or initialize the decoder.

[0137] By determining specific decoder functional requirements for each reused version, it is possible to provide more accurate decoder functional requirements, thus avoiding information that indicates unnecessarily high decoder functional requirements. For example, the decoder can select the highest quality OLS appropriate to its capabilities.

[0138] Accordingly, an embodiment of the extractor 10 according to the fourth aspect is an extractor 10 for extracting a subpicture-specific video data stream 12 from a multilayer video data stream 14 constituting a set of layers, wherein the multilayer video data stream 14 has an unreduced version of the set of layers encoded therein, and the video versions 240, 241 are encoded in the unreduced version of the set of layers in units of one or more independently coded subpictures per layer, which is configured to derive a reference decoder functional requirement related to decoding the unreduced version of the set of layers from the multilayer video data stream 14. According to this embodiment, the extractor 10 is configured to derive from the multilayer video data stream, for each layer of the set of layers, information such as information 22 illustrated in Figure 2 regarding the picture size of the picture encoded in each layer of the unreduced version of the set of layers, and for each of the one or more independently coded subpictures of each layer, the subpicture size. Furthermore, according to this embodiment, the extractor 10 is configured to determine the decoder functionality requirements related to decoding a predetermined reduced version of a set of layers, for a predetermined (e.g., predetermined by an OLS display 20 provided by the API) reduced version of a set of layers from which the subpicture-related portions of at least one layer have been removed, by scaling a reference decoder functionality requirement by a coefficient determined using the quotient obtained by dividing the sum over the subpicture size of independently coded subpictures contained in the predetermined reduced version of the set of layers by the sum over the picture size.

[0139] In the following, an alternative embodiment of the fourth aspect will be described with reference to Figures 1 and 2.

[0140] According to these alternative embodiments, the multilayer video data stream 14 includes multiple layers, such as layer L0 and layer L1. Furthermore, the encoder 40 is configured to encode video pictures 26, such as picture 260 and picture 261, into the layers of the multilayer video data stream such that they are subdivided into subpictures independently encoded with respect to one or more first layers, such as layer L1 in Figure 1, but not subdivided with respect to one or more second layers, such as layer L0. In other words, the encoder 14 may encode the pictures of video 24 into first type layers, such as the first layer L1, so that they are subdivided into independently encoded subpictures, or it may encode the pictures so that they are not subdivided into one or more second type layers, such as the second layer L0. Note that the multilayer video data stream 14 may include one or more layers of the first type and one or more layers of the second type. The encoder 40 in these alternative embodiments is configured to encode a display of the multilayer video data stream 14, such as the display 18 shown in Figure 2, of a layer set in which at least one first layer is included and at least one second layer is included. For example, the display of the set of layers may be an OLS display 18. The multilayer video data stream 14 may include displays of multiple sets of layers, as described with respect to Figure 1.

[0141] The encoder 40 in these alternative embodiments is configured to encode information into the multilayer video data stream 14 of a first percentage (e.g., the first percentage 134 in Figure 9) of each of several reference decoder functional requirements sufficient to decode a set of layers, and a section percentage 132 (e.g., the second percentage 134 in Figure 9) of each reference decoder functional requirement, each of which is at least one first layer. For example, the first portion 134 may be represented by the syntax element ref_level_fraction_minus1[i][j] introduced above, and the second portion 132 may be represented by the syntax element non_subpic_layer_fraction[i][j] introduced above.

[0142] For example, the reference decoder capability requirements sufficient to decode a layer set may mean that each of the reference decoder capability requirements is defined by, at least, a minimum CPB size. Some reference decoder capability requirements may differ from one another with respect to their indicated minimum CPB size. Note that for some reference decoder capability requirements, a layer set may consume only a small fraction of the overall CPB size, for example, only a small fraction of the total available CPB size, so a mixer like mixer 100 described with respect to Figure 10 may use such information, i.e., the reference decoder capability requirements, to determine how many OLSs of different streams will be mixed together, even though it satisfies the level / reference decoder capability requirements.

[0143] For example, decoder function requirements, also known as decoder function information (DCI), can indicate the requirements of a decoder for decoding the layer set to which the decoder function requirements are associated. For example, a layer set may have several reference decoder function requirements associated with it. For example, each reference decoder function requirement associated with a layer set may refer to a different version of video encoded in the layer set, such as a different sub-picture-specific video data stream, where different versions of the video containing different sub-pictures are encoded.

[0144] Embodiments of the fourth aspect include an apparatus for processing a multilayer video data stream 14, which can process a video data stream provided by the alternative embodiment of the encoder 40 described above. For example, the apparatus for processing the multilayer video data stream 14 includes a mixer such as the mixer 100 described below, an extractor such as the extractor 10 in Figures 1, 2 and 5, a decoder such as the decoder 50, or a combination thereof. Thus, the apparatus for processing the data stream is referred to as apparatus 10, 100. For example, the decoder 50 and the extractor 10 may be combined, or one may include the other. The apparatus for processing a multilayer video data stream according to these embodiments is configured to decode from the multilayer video data stream 14 a representation of a layer set including at least one first layer and one second layer, as described above. The device is configured to decode from the multilayer video data stream 14 information regarding, for each of the multiple reference decoder functional requirements, a first proportion of each reference decoder functional requirement attributable to at least one first layer, and a second proportion of each reference decoder functional requirement attributable to at least one second layer.

[0145] Figure 9 shows examples of a first layer picture 261 and a second layer picture 260, and the first and second proportions 134 and 132 to which they belong, respectively, which can be signaled in the video data stream 14 according to the alternative embodiment just described in the fourth aspect. In the example of Figure 9, the first layer picture 261 is subdivided into two subpictures 28. Figure 9a shows the determination of the first proportions 1341 and 1321 to the first reference decoder function requirement, which can be associated with a first reference level, as shown in Figure 9. In the exemplary example of Figure 9, the first reference level is associated with a first CPB 1301 having an exemplary value of 1000 (which can represent an exemplary decoder function requirement in Figure 9). In Figure 9, for each picture or subpicture, its relative share to the total picture size of the video stream associated with its respective reference decoder function requirement is shown. For example, according to Figure 9a, the picture size of picture 260 has a 20% share of the total picture size of the video stream. In other words, Figure 9a shows the determination of the reference decoder function requirement for a video data stream containing pictures 260 and 261, which contain both subpictures. The second portion 1321 is determined as the share of picture 260 of the second layer to the total CPB shown for each decoder function requirement. The first percentage 1341 related to the decoder function requirement of the video data stream shown in Figure 9a is the remaining percentage of the CPB, i.e., the share not used by picture 260 of the second layer, which is distributed over the subpictures on picture 261 of the first layer.

[0146] Figure 9b shows an example of a video data stream similar to that shown in Figure 9a, but containing only one of the subpictures 28 of picture 261 in the first layer. Therefore, the corresponding CPB level is smaller by the size of one of the subpictures 28 in Figure 9a. In the example of Figure 9b, since the video data stream contains only one of the subpictures of the first layer, the picture relative to the entire picture area signaled by the video data stream 260The share is higher, namely 33% instead of 20%. The second and first percentages 1322 and 1342 related to the reference decoder functional requirements in the example of Figure 9b are determined accordingly, taking into account that in the case of Figure 9b, picture 261 contains only one of the two subpictures 28, as explained with respect to Figure 9a.

[0147] According to one embodiment, the encoder 40 can encode information about the first and second percentages by writing the first percentage to the multilayer video data stream on the one hand, and the second percentage divided by 1 minus the first percentage on the other hand. Therefore, devices 10 and 100 can decode information about the first and second percentages by reading the first percentage and the second percentage from the data stream. Alternatively, devices 10 and 100 can decode information about the first and second percentages by reading the first percentage and the further percentage from the data stream and deriving the second percentage by multiplying the further percentage by the difference obtained by subtracting the first percentage from 1.

[0148] According to some examples, the encoder 40 can encode information about one or more of several reference decoder functional requirements into a multilayer video data stream, such as the decoder functional requirement 60 shown in Figure 1. In other words, the encoder 40 can encode information about one or more of several reference decoder functional requirements into a multilayer video data stream, though not necessarily all of them. For example, information about other of several reference decoder functional requirements may be derivable by default rules. For example, in the example of Figure 9, information about one of the decoder functional requirements in Figures 9a and 9b can be derived from the other based on information about the relative sizes of the subpictures into which the picture 261 of the first layer is subdivided.

[0149] For example, the encoder 40 may be configured to encode information such that, for at least one of several reference decoder functional requirements, the first percentage 134 and the second percentage 132 relate to a reduced version of the multilayer video data stream, e.g., the first percentage 1342 and the second percentage 1322 in Figure 9b. The encoder may encode information such that, for each of one or more independently encoded subpictures present in the reduced version, the first percentage is derivable from the multilayer video data stream. For example, referring to Figure 9a, recognizing parameters in the video encoded by the encoder 40 into the data stream 14, e.g., the width and / or height of the subpictures, e.g., the ratio between two subpictures of picture 261, e.g., 50% in Figure 9a, which may be derivable from measurements taken in a sample or tile, the device 10, 100 may derive the first percentage of the right subpicture of picture 261 from the first percentage of the left subpicture.

[0150] Furthermore, or alternatively, the encoder 14 may be configured to decode information such that, for at least one further of several reference decoder functional requirements, the first and second percentages relate to the unreduced version of the multilayer video data stream, e.g., the first percentage 1341 and the second percentage 1321 in Figure 9a, and the first percentage 134 is derived from the multilayer video data stream 14 for each of the independently coded subpictures of at least one first layer.

[0151] As shown in Figure 9, devices 10 and 100 can sum the first ratio 134 and the second ratio 132 to obtain the actual decoder function requirements, such as CPB, for each layer set. As illustrated, the actual decoder function requirements may be specific to each sub-picture video data stream.

[0152] As described above, devices 10 and 100 may be mixers. That is, devices 10 and 100 may mix the multilayer video data stream 14 with another video data stream. For this purpose, devices 10 and 100 may use information on the first and second proportions to form a further data stream based on the data stream 14 and further input video data streams, as described with respect to Figure 10, for example.

[0153] 5. Encoder 40 and extractor 10 according to a fifth embodiment This section describes an embodiment according to the fifth aspect with reference to Figure 1, and the details described in Section 0 may be optionally applied to the embodiment according to the fifth aspect. Furthermore, the details of data streams 12 and 14 shown in Figures 2 and 5 may be optionally included in an embodiment according to the third aspect, in which case the corresponding descriptions in Sections 1 and 2 shall apply as appropriate. Optionally, features described with respect to the first, second, third, and especially fourth aspects may also be combined with the fifth aspect. In particular, examples of embodiments described in Section 5 may be relevant to embodiments described in Section 4, especially with respect to decoder functional requirements and the first and second proportions.

[0154] Embodiments of the fifth aspect may relate to the selective application of constraints to layers of the OLS.

[0155] As mentioned above, the syntax element ref_level_fraction_minus1[i][j] is used in modern techniques to impose specific layer-specific constraints, such as tile count, whether explicitly signaled or inferred. However, for (reference) layers without subpictures, these constraints can be unnecessarily strict or difficult to satisfy from the encoder's perspective. This is because the extraction process does not affect each layer individually, but rather reinforces the constraints that each layer must satisfy.

[0156] Therefore, in another embodiment, part of this embodiment is to selectively impose the following constraints only on layers in the OLS that contain subpictures, i.e., layers that undergo picture size reduction from the extraction process. Thus, the encoder can take advantage of the fewer restrictions on non-subpicture subdivided layers.

[0157] According to an embodiment of the fifth aspect, the encoder 10 may encode picture 261 in one or more first layers, for example, a first layer L1, so that it is subdivided into independently encoded subpictures 28, and encode picture 260 in one or more second layers, for example, layer L0, so that it is not subdivided, as shown in Figure 1 or Figure 2. The encoder may selectively check layer-specific constraints for each layer set of layer sets, for example, an OLS, or one or more OLS layers shown in the OLS display 18, for the first layer of the layer set, i.e., layers that encode a picture subdivided into subpictures, i.e., layers that encode a picture subdivided into at least two subpictures. In other words, checks may be performed for layers in which the picture is subdivided and encoded, and checks may be omitted for layers in which the picture is not subdivided and encoded. In other words, checks may be performed only for layers in which the picture is encoded in a subpicture subdivision manner.

[0158] The constraints selectively checked for a first layer relate to a predetermined set of parameters for each subpicture of the first layer. In other words, the encoder may check constraints for each of the one or more reference decoder functional requirements for each of the subpictures of each first layer, the constraints being unique to each of these requirements. For example, the one or more reference decoder functional requirements for which constraints are checked for each subpicture may refer to the subpicture sequence resulting from extracting each subpicture from each first layer, or to the contribution of this subpicture sequence to the video sequence or video data stream resulting from extracting each subpicture from each first layer. As described in Section 0, the bitstream described by the OLS may remain scalable by the selection / extraction of one or more time sublayers and / or one or more layers of the OLS even after extraction by the extractor 10. The same applies when the extractor 10 extracts video data streams for each subpicture, and one or more reference decoder functional requirements signaled by the encoder may be applied to each subpicture in the video data stream to the extent defined by the proportion of each reference decoder functional requirement specific to the subpicture, for example, the first proportion described in Section 4. In other words, one or more reference decoder functional requirements may be associated with each subpicture, each of the reference decoder functional requirements may be applied to the subbit streams that can be extracted from each video data stream for each subpicture, and each of the reference decoder functional requirements may impose constraints specific to the reference decoder functional requirement.

[0159] For example, the parameters that are subject to constraints are: - Number of samples, e.g., LUMA samples, e.g., MaxLumaPs, - Picture size, e.g., picture width and / or picture height, e.g., number of samples, - The number of tiles per row and / or per column, e.g., MaxTileRows and / or MaxTileCols, - May include one or more, or all, of the total number of tiles.

[0160] For example, the encoder 10 may determine the parameters or properties of individual subpictures using a proportion of the respective reference decoder capability requirements that belong to each subpicture, for example, a first proportion 134 as described in Section 4, for example, RefLevelFraction[i][j] in the following example (where i represents the respective reference decoder capability requirement and j represents the respective subpicture).

[0161] The following is an example of the current VVC draft specification, including the highlighted changes.

[0162] Each layer in the bitstream, extracted with the j-th subpicture set to j in the range from 0 to sps_num_subpics_minus1 and sps_num_subpics_minus1 greater than 0, is a bitstream conformance requirement. Furthermore, if general_tier_flag is 0, level is ref_level_idc[i], and i conforms to a profile from 0 to num_ref_level_minus1, the following constraints must be observed in each bitstream conformance test specified in Appendix C of the VVC specification.

[0163] - Ceil(256*SubpicSizeY[j]÷RefLevelFraction[i][j]) must be less than or equal to MaxLumaPs. Here, MaxLumaPs is specified in Table A.1 (of VVC) for the level ref_level_idc[i]. - The value of Ceil(256*(sps_subpic_width_minus1[j]+1)*CtbSizeY÷RefLevelFraction[i][j]) must be less than or equal to Sqrt(MaxLumaPs*8). - The value of Ceil(256*(sps_subpic_height_minus1[j]+1)*CtbSizeY÷RefLevelFraction[i][j]) must be less than or equal to Sqrt(MaxLumaPs*8). - The value of SubpicWidthInTiles[j] must be less than or equal to MaxTileCols, and the value of SubpicHeightInTiles[j] must be less than or equal to MaxTileRows. Here, MaxTileCols and MaxTileRows are defined in Table A.1 (of VVC) of level ref_level_idc[i]. - The value of SubpicWidthInTiles[j]*SubpicHeightInTiles[j] shall be less than or equal to MaxTileCols*MaxTileRows*RefLevelFraction[i][j], where MaxTileCols and MaxTileRows are as defined in Table A.1 (of VVC) of level ref_level_idc[i].

[0164] In another embodiment, for a layer of OLS where sps_num_subpics_minus1 is equal to 0, RefLevelFraction[i][j] is derived to be equal to 255 (for example, having a percentage of 100%).

[0165] In another embodiment, the limits on MaxLumaPs, maximum and minimum aspect ratios, or the number of tile columns or rows, slices or subpictures are layer-specific and use a layer-specific ref_level_fraction (i.e., 255 for layers without subpictures), whereas the layer-specific limits on CPB size, bitrate and MinCR are applied based on an OLS-specific OLSRefLevelFraction corresponding to OLS_fraction[k] (for example, as derived above in Section 4).

[0166] The variables SubpicCpbSizeVcl[i][j] and SubpicCpbSizeNal[i][j] are derived as follows: SubpicCpbSizeVcl[i][j]=Floor(CpbVclFactor*MaxCPB*OLSRefLevelFraction[i][j]÷256) (D.5) SubpicCpbSizeNal[i][j]=Floor(CpbNalFactor*MaxCPB*OLSRefLevelFraction[i][j]÷256) (D.6) Use MaxCPB derived from ref_level_idc[i] as defined in Section A.4.2. The variables SubpicBitRateVcl[i][j] and SubpicBitRateNal[i][j] are derived as follows:

[0167] SubpicBitRateVcl[i][j]=Floor(CpbVclFactor*ValBR*OLSRefLevelFraction[i][j]÷256) (D.7) SubpicBitRateNal[i][j]=Floor(CpbNalFactor*ValBR*OLSRefLevelFraction[i][j]÷256) (D.8) Note 1 - When a subpicture is extracted, the resulting bitstream consists of a CpbSize (shown or estimated in SPS) greater than or equal to SubpicCpbSizeVcl[i][j] and SubpicCpbSizeNal[i][j], and a BitRate (shown or estimated in SPS) greater than or equal to SubpicBitRateVcl[i][j] and SubpicBitRateNal[i][j].

[0168] The sum of the NumBytesInNalUnit variables of AU0 corresponding to the j-th subpicture must be less than or equal to FormatCapabilityFactor*(Max(SubpicSizeY[j],fR*MaxLumaSr*OLSRefLevelFraction[i][j]÷256)+MaxLumaSr*(AuCpbRemovalTime[0]-AuNominalRemovalTime[0])*OLSRefLevelFraction[i][j])÷(256*MinCr) with respect to the value of SubpicSizeInSamplesY of AU0. Here, MaxLumaSr and FormatCapabilityFactor are the values ​​shown in Tables A.2 and A.3, respectively, applied to AU0 and level ref_level_idc[i], and MinCr is derived as shown in A.4.2 (VVC draft).

[0169] The sum of the NumBytesInNalUnit variables of AUn (where n is greater than 0) corresponding to the j-th subpicture must be less than or equal to FormatCapabilityFactor*MaxLumaSr*(AuCpbRemovalTime[n]-AuCpbRemovalTime[n-1])*OLSRefLevelFraction[i][j]÷(256*MinCr). Here, MaxLumaSr and FormatCapabilityFactor are the values ​​shown in Tables A.2 and A.3, respectively, applied to the level ref_level_idc[i] of AUn, and MinCr is derived as shown in A.4.2 (VVC draft).

[0170] Embodiments of the encoder 40, extractor 10, and multilayer video data stream 14 according to a fifth aspect are described below.

[0171] According to the fifth aspect, the multilayer video data stream includes multiple layers. The multilayer video data stream 14 encodes video pictures 26 into layers of the multilayer video data stream, for example, by the encoder 40, such that the video pictures 26 are subdivided into subpictures independently encoded with respect to one or more first layers, e.g., layer L1, and into one or more second layers, e.g., not subdivided with respect to layer L1. Furthermore, the multilayer video data stream 14 includes a representation of a layer set, such as an OLS as described in the previous section, and according to the fifth aspect, the layer set includes at least one first layer, i.e., a layer encoded such that the picture is subdivided into subpictures independently encoded. According to the fifth aspect, the multilayer video data stream includes, for example, reference decoder function requirement information encoded therein, as described in Section 4, the reference decoder function requirement information relating to decoding an uncompressed version of the layer set, and includes a first decoder function requirement parameter for each layer of the layer set and a second decoder function requirement parameter for the layer set. As described in Section 4, the uncompressed version may refer to a version of the video data stream encoded in the layer set, which includes all the video data encoded in the layer set, i.e., all the subpictures of at least one first layer of the layer set. For example, the first decoder function requirement parameter could be the number of tiles into which the layer's picture is subdivided. For example, a tile is a mutually independent coded portion of a picture. For example, the second decoder function requirement parameter may refer to required or necessary picture buffer sizes, such as the ACPB and / or decode picture buffer (DPB) size and / or minimum compression ratio (minCR).

[0172] According to a fifth aspect, the encoder 40 is configured to further determine reference decoder function requirement information related to decoding each un-reduced or reduced version of the layer set for each of one or more reduced versions of the layer set from which the sub-picture related portion of at least one first layer has been removed. The reduced version of the layer set may refer to a video data stream that does not contain all of the video data of the layer set, for example, a sub-picture related video data stream 12, i.e., a video data stream from which the sub-picture related portions, i.e., portions not related to a given set of sub-pictures as described with respect to Figures 1, 2, and 5, have been omitted when extracting each reduced version of the layer set. The encoder 40 is configured to determine further reference decoder function requirement information by scaling the first decoder function requirement parameter using a first percentage in order to obtain a third decoder function requirement parameter for each layer of the layer set. For example, the first percentage may be determined based on the ratio of the picture size between the picture of each layer in the layer set and the picture in the reduced version of that layer set. The encoder 40 can optionally encode the first percentage into a multi-layer video data stream 14. Therefore, the extractor 10 can optionally derive a first percentage from the multilayer video data stream 14, or alternatively determine the first percentage as performed by the encoder 40.

[0173] The encoder 40 is further configured to determine further reference decoder function requirement information by scaling the second decoder function requirement parameter using a second proportion in order to obtain a fourth decoder function requirement parameter for each reduced version of the layer set. The second proportion may be different from, for example, the first proportion.

[0174] The second percentage can be determined based on the ratio coding to Section 4, such as percentage 132 and percentage 134. The encoder 40 can optionally encode the second percentage into the multilayer video data stream 14. Thus, the extractor 10 can optionally derive the second percentage from the multilayer video data stream 14 or determine a second percentage similar to that performed by the encoder 40.

[0175] The encoder 10 forms the subpicture-specific video data stream 12 based on the multilayer video data stream 14 by removing the subpicture-related portion of at least one first layer of the multilayer video data stream 14 and providing the subpicture-specific video data stream 12 with a third decoder function requirement parameter and a fourth decoder function requirement parameter. Alternatively, the encoder 10 may form the subpicture-specific video data stream 12 based on the multilayer video data stream 14 by removing the subpicture-related portion of at least one first layer and providing the subpicture-specific video data stream 12 with a first percentage and a second percentage for determining further reference decoder function requirement information related to decoding a predetermined un-reduced version of the layer set of layers by the decoder.

[0176] 6. Encoder 40 and decoder 110 as shown in Figure 10 Figure 10 shows a first encoder 401 and a second encoder 402, each of which can optionally correspond to the encoder 40 described in relation to Figure 1 and sections 1-5. Encoder 401 is configured to encode video data stream 121, and encoder 402 is configured to encode video data stream 122. Video data streams 121 and 122 do not necessarily have to contain multiple layers, although this is optional. In other words, video data streams 121 and 122 in Figure 10 may optionally be multi-layer video data streams such as video data stream 14, but in this example they may contain only one single layer. Separately, video data streams 121 and 122 in Figure 10 may optionally correspond to video data stream 14 described in the previous section.

[0177] Encoder 401 is configured to encode picture 261 into video data stream 121. Encoder 402 is configured to encode picture 262 into video data stream 122. Picture 261 and picture 262 may optionally be encoded into their respective video streams 121 and 122 by one or more independently coded subpictures 28. Encoders 401 and 402 are configured to encode an indicator 80 into their respective video data streams 121 and 122 indicating whether picture 261 and 262 are encoded into their respective video data streams by one or more independently coded subpictures. The indicator 80 identifies different types of coding independence between one or more independently coded subpictures 28 and the area surrounding one or more independently coded subpictures. For example, the area surrounding one or more independently coded subpictures may refer to the picture area surrounding each subpicture.

[0178] For example, display 80 may be an m array syntax element where m > 1, or it may contain two or more syntax elements to indicate whether such independently coded subpictures exist, and if so, further syntax elements to clarify the type of decoding independence.

[0179] Figure 10 further illustrates an apparatus 100 for providing a video data stream 20 based on one or more input video data streams 12, such as video data stream 121 and video data stream 122. For example, the apparatus 100 determines the video data stream 20 based on video data stream 12 1、 122 can be mixed. Therefore, the device 100 may be called a mixer 100. Optionally, the device 100 may be further configured to extract video data stream 20 from one or more input video data streams, such as video data stream 121 and / or video data stream 122. For example, the mixer 100 may extract some or all of the video data from each of its input video data streams and provide the video data extracted from different input video data streams to video data stream 20. For example, as shown in Figure 10, the device 100 may extract one or more subpictures from each of video data streams 121 and 122 and then combine them with one or more subpictures associated with the same presentation time to obtain video data stream 20.

[0180] Figure 10 further shows a decoder 110 which may optionally correspond to the decoder 50 described in the previous section. The decoder 110 is configured to decode from the video data stream 20 in display 80.

[0181] In the example, the encoders 401, 402, and decoder 110 may, regardless of any of the different types of encoding independence indicated by the indication 80, perform one or more of the following actions in response to indication 80 indicating that the picture encoded in the video data stream is encoded in a manner of one or more independently encoded subpictures (for example, only if otherwise, or not otherwise):

[0182] - From each bitstream packet of a video data stream, read a subpicture identifier that indicates which of one or more independently encoded subpictures is encoded in each bitstream packet, and / or - To prevent in-loop filtering from crossing sub-picture boundaries, and / or - Deriving a picture location from the block address of a block of a picture, where this block address is encoded in a given bitstream portion of a video data stream, depending on which of one or more independently encoded subpictures is encoded into a given bitstream packet (and deriving such a picture location from the block address by referring to a given picture location, such as its upper left corner, for example), and / or - To allow different NAL unit types occurring in one access unit of a video data stream. For example, the mixer 100 in Figure 10 allows bitstream 12 in the mixed data stream 90 such that one sub-picture 28 has one NAL unit type and another sub-picture 28 within the same picture is coded using a different NAL unit type. 1,2Mixing is permitted. The NAL unit type is encoded in bitstream packet 16 and remains there unmodified when it becomes packet 16' of stream 90. Decoder 110 finds no conflict because whatever type of encoding independence is used for the subpicture 28 of the decoded (mixed) picture 26', the picture is fully encoded using it. and / or - The encoder must derive decoder function requirement information for each of the one or more independently coded subpictures.

[0183] For example, different types of coding independence between subpictures of a picture include: - An encoder-constrained encoding type, wherein the decoder is configured such that, with respect to independently encoded subpictures, when an encoder-constrained encoding type is indicated by the display, it does not perform isolation processing at the boundaries of independently encoded subpictures. - A decoder-aware encoding type, wherein the decoder is configured to perform isolation processing at the boundaries of independently encoded subpictures by vector clipping for vector-based prediction and / or boundary padding at the boundaries of independently encoded subpictures, when the decoder-aware encoding type is indicated by indication, with respect to independently encoded subpictures.

[0184] In other words, embodiments according to the sixth aspect can enable the mixing of subpictures or MCTS.

[0185] The current VVC draft specification includes numerous measures to enable splitting into subpictures and to enable extraction and merge functionality. For example, · Carriage of subpicture identifier in slice header, or • Block addressing schemes that depend on subpicture identifiers, or • Compatibility information for individual sub-pictures, or • Mixing of NAL unit types within a picture (e.g., IRAP and TRAIL).

[0186] However, all of these measurements depend on the boundary processing characteristics of the subpicture. That is, outside the subpicture boundary, the sample values ​​and syntax predictions are extrapolated in the same way as at the picture boundary, i.e., they depend on special boundary padding processing on the decoder side. This enables independent coding of subpictures, and subsequent extraction and merging. This characteristic is also used to enable all of the above means of extraction and merging functions. Figure 11 shows a bitstream that uses such decoder-processed subpictures only for the first two access units of the bitstream. That is, Figure 6 shows an example of a bitstream that mixes decoder-side boundary processing and encoder-side boundary processing in independently coded subpictures.

[0187] However, there are further methods for the independent coding of such rectangular regions, i.e., constraints on the movement followed during coding, in other words, boundary processing on the encoder side. The encoder can easily avoid references that would lead to reconstruction errors after extraction or merging by ensuring that it does not point outside the circumference of each subpicture. Based on state-of-the-art signaling, the encoder should indicate that the right-hand subpicture of L1 is not independently coded, since no decoder-side boundary processing occurs from the second IRAP onward in Figure 11, and therefore some of the measures regarding the extraction and merging functions described above can be used efficiently. Most importantly, the conformance information is associated with the decoder-side boundary processing. However, merely indicating other means for independent coding should enable the encoder and allow the network device or decoder to facilitate the above means.

[0188] As shown in Figure 12, it is important to note that a bitstream that mixes different types of independent coding between adjacent subpictures does not necessarily need to contain multiple layers. Figure 12 shows an example of a single-layer mixing of subpicture independence types.

[0189] Therefore, part of this embodiment is to indicate in the bitstream that a region is independently coded by the constrained encoding of a particular subpicture, so that the decoder can recognize that all the previously described measures for independent region coding (address schemes, conformance information) can be used despite the absence of boundary padding. In one embodiment of this embodiment, a flag used to indicate the decoder-side padding procedure (sps_subpic_treated_as_pic_flag[i] equal to 1) is changed to sps_subpic_treated_as_pic_mode[i], which indicates that the new state (equal to 2) is sps_subpic_treated_as_pic_mode[i], indicating that the i-th subpicture is independently coded by the encoder-side boundary processing constraint.

[0190] An example is shown below. [Table 2] If sps_subpic_treated_as_pic_mode[i] is equal to 12, it specifies that the i-th subpicture of each coded picture in the CLVS should be treated as a picture during decoding, excluding the in-loop filtering operation. If sps_subpic_treated_as_pic_mode[i] is equal to 0 or 1, it specifies that the i-th subpicture of each coded picture in the CLVS should not be treated as a picture during decoding, excluding the in-loop filtering operation. If it does not exist, the value of sps_subpic_treated_as_pic_mode[i] is assumed to be equal to 12. The value of sps_subpic_treated_as_pic_mode[i] of 3 is reserved for future use by ITU-T|ISO / IEC.

[0191] Furthermore, a constraint is imposed that different types of NAL units cannot be mixed unless the subpicture is a constraint on the encoder side or a boundary treatment on the decoder side.

[0192] If any two subpictures within a picture have different NAL unit types, the value of sps_subpic_treated_as_pic_mode[] shall not be equal to 0 for all subpictures within the picture that contain at least one P-slice or B-slice.

[0193] Figure 13 shows examples where the mixing of NAL units is permitted and prohibited by subpictures that do not depend on encoder-side constraints or decoder-side boundary processing.

[0194] 7. Further Embodiments In the previous sections 0-6, some embodiments were described as features in the context of the apparatus, but it is clear that such descriptions can also be considered as descriptions of corresponding features of the method.

[0195] Some or all of the method steps may be performed by (or using) a hardware device, such as a microprocessor, a programmable computer, or an electronic circuit. In some embodiments, one or more of the most important method steps may be performed by such a device.

[0196] The encoded image signal of the invention may be stored in a digital storage medium, or it may be transmitted over a transmission medium such as a wireless transmission medium such as the Internet or a wired transmission medium.

[0197] Depending on certain implementation requirements, embodiments of the present invention may be implemented in hardware or software or at least partially in hardware or at least partially in software. These implementations may be performed using digital storage media such as floppy disks, DVDs, Blu-rays, CDs, ROMs, PROMs, EPROMs, EEPROMs, or FLASH® memory, which have electronically readable control signals, and which are stored thereon and which work (or can work) with a programmable computer system to perform each method. Thus, the digital storage media may be computer-readable.

[0198] In some embodiments of the present invention, one of the methods described herein is performed by including a data carrier having electronically readable control signals, wherein these control signals can cooperate with a programmable computer system.

[0199] In general, embodiments of the present invention may be implemented as a computer program product having program code, which is operable to perform one of the methods when the computer program product is executed on a computer. The program code may be stored, for example, in a machine-readable carrier.

[0200] Other embodiments include a computer program that performs one of the methods described herein and is stored in a machine-readable carrier.

[0201] In other words, one embodiment of the method of the present invention is a computer program having program code for performing one of the methods described herein when the computer program is executed on a computer.

[0202] Further embodiments of the methods of the present invention are, therefore, a data carrier (or digital storage medium, or computer-readable medium), which includes a computer program recorded thereon for performing one of the methods described herein. The data carrier, digital storage medium, or recorded medium is typically tangible and / or non-temporary.

[0203] A further embodiment of the method of the present invention is a data stream or signal sequence representing a computer program for performing one of the methods described herein. The data stream or signal sequence may be configured to be transmitted over a data communication connection, such as over the Internet.

[0204] Further embodiments include, for example, processing means such as a computer or programmable logic device configured or adapted to perform one of the methods described herein.

[0205] Further embodiments include a computer on which a computer program for performing one of the methods described herein is installed.

[0206] Further embodiments of the present invention include an apparatus or system configured to transfer (e.g., electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver may be, for example, a computer, a mobile device, a memory device, etc. The apparatus or system may include, for example, a file server for transferring the computer program to the receiver.

[0207] In some embodiments, a programmable logic device (e.g., a field-programmable gate array) may be used to perform some or all of the functions of the method herein. In some embodiments, a field-programmable gate array may cooperate with a microprocessor to perform one of the methods herein. Generally, the method is preferably performed by any hardware device.

[0208] The apparatus described herein may be implemented using hardware devices, or using a computer, or using a combination of hardware devices and a computer.

[0209] The methods described herein may be performed using hardware devices, or using a computer, or using a combination of hardware devices and a computer.

[0210] As can be seen in the detailed description above, various features are grouped into examples for the purpose of streamlining the disclosure. This method of disclosure should not be interpreted as reflecting the intention that the examples described in the claims require more features than are explicitly stated in each claim. Rather, as reflected in the following claims, the subject matter may consist of fewer features than all of the single disclosed examples combined. Therefore, the following claims are incorporated within the detailed description herein, and each claim can stand alone as a separate example. While each claim can stand alone as a separate example, dependent claims may refer in a particular combination with one or more other claims, but it should be noted that other examples may also include combinations of a dependent claim with the subject matter of each other dependent claim, or combinations of each feature with other dependent or independent claims. Such combinations are proposed herein unless otherwise stated that a particular combination is not intended. Furthermore, even if a claim is not directly dependent on an independent claim, it is intended to include features of the claim relative to other independent claims.

[0211] The embodiments described above are merely illustrative of the principles of this disclosure. It will be understood that modifications and variations of the configurations and details described herein will be obvious to those skilled in the art. Accordingly, it is intended that the invention be limited only by the pending claims and not by the specific details presented as part of the description and explanation of the embodiments herein.

Claims

1. A method for decoding video data from data streams (12, 14), The acquisition of the data stream (12, 14), wherein the data stream (12, 14) includes an enhancement layer (L1) and a base layer (L0), With respect to the current picture encoded in the enhancement layer (L1), a reference picture decoded from the base layer (L0) is identified for use in inter-layer prediction of the current picture. The variables are set based on a comparison between the first configuration setting associated with the current picture and the second configuration setting associated with the reference picture. Based on the aforementioned variables, it is determined that at least one inter-layer prediction tool is unavailable for predicting the blocks of the current picture using the reference picture, Predicting the blocks of the current picture using the reference picture without using the at least one interlayer prediction tool that has been determined to be unavailable, A method that includes this.

2. The first configuration setting indicates the number of first sub-pictures included in the current picture, The second configuration setting indicates the number of second sub-pictures included in the reference picture. The method according to claim 1.

3. Comparing the number of the first sub-pictures with the number of the second sub-pictures, When the number of the first sub-pictures and the number of the second sub-pictures are different, the variable is set to the first value, Based on the fact that the variable has the first value, it is determined that the at least one inter-layer prediction tool is unavailable, The method according to claim 2, further comprising:

4. The aforementioned at least one interlayer prediction tool is: Inter-layer motion vector prediction and, Refinement of optical flow in vector-based inter-layer prediction, or, Motion vector wrap-around in vector-based inter-layer prediction, The method according to claim 1, comprising at least one of the following.

5. The method according to claim 1, wherein the current picture and the reference picture have the same spatial resolution and the same scaling window offset.

6. The method according to claim 1, wherein the current picture and the reference picture are at the same position in time.

7. A non-temporary computer-readable medium comprising a program, when executed by a processor, that causes the processor to perform the method according to claim 1.

8. A decoder for decoding video data from data streams (12, 14), The decoder mentioned above is The process involves acquiring the data stream (12, 14), wherein the data stream (12, 14) includes an enhancement layer (L1) and a base layer (L0). For the current picture encoded in the enhancement layer (L1), a reference picture decoded from the base layer (L0) is identified for use in inter-layer prediction of the current picture. A variable is set based on a comparison between the first configuration setting associated with the current picture and the second configuration setting associated with the reference picture. Based on the aforementioned variables, it is determined that at least one inter-layer prediction tool is unavailable for predicting the blocks of the current picture using the reference picture. Without using the at least one interlayer prediction tool that has been determined to be unavailable, predict the blocks of the current picture using the reference picture. A decoder equipped with a processor configured in such a way.

9. The first configuration setting indicates the number of first sub-pictures included in the current picture, The second configuration setting indicates the number of second sub-pictures included in the reference picture. The decoder according to claim 8.

10. The aforementioned processor further, The number of the first sub-pictures and the number of the second sub-pictures are compared, If the number of the first sub-pictures and the number of the second sub-pictures are different, the variable is set to the first value. Based on the fact that the variable has the first value, it is determined that the at least one interlayer prediction tool is unavailable. The decoder according to claim 9, configured as described above.

11. The aforementioned at least one interlayer prediction tool is: Inter-layer motion vector prediction and, Refinement of optical flow in vector-based inter-layer prediction, or, Motion vector wrap-around in vector-based inter-layer prediction, Including at least one of the following: The decoder according to claim 8.

12. The decoder according to claim 8, wherein the current picture and the reference picture include the same spatial resolution and the same scaling window offset.

13. The decoder according to claim 8, wherein the current picture and the reference picture are at the same time position.

14. A method for encoding video data into a data stream (12, 14), For the current picture encoded in the enhancement layer (L1), a reference picture encoded in the base layer (L0) is identified for use in inter-layer prediction of the current picture, The variables are set based on a comparison between the first configuration setting associated with the current picture and the second configuration setting associated with the reference picture. Based on the aforementioned variables, it is determined that at least one inter-layer prediction tool is unavailable for predicting the blocks of the current picture using the reference picture, Encoding the blocks of the current picture using the reference picture without using the at least one interlayer prediction tool that has been determined to be unavailable, The process involves generating the data streams (12, 14), wherein the data streams (12, 14) include the enhancement layer (L1) and the base layer (L0). A method that includes this.

15. The first configuration setting indicates the number of first sub-pictures included in the current picture, The second configuration setting indicates the number of second sub-pictures included in the reference picture. The method according to claim 14.

16. Comparing the number of the first sub-pictures with the number of the second sub-pictures, When the number of the first sub-pictures and the number of the second sub-pictures are different, the variable is set to the first value, Based on the fact that the variable has the first value, it is determined that the at least one inter-layer prediction tool is unavailable, The method according to claim 15, further comprising:

17. The aforementioned at least one interlayer prediction tool is: Inter-layer motion vector prediction and, Refinement of optical flow in vector-based inter-layer prediction, or, Motion vector wrap-around in vector-based inter-layer prediction, The method according to claim 14, comprising at least one of the following.

18. The method according to claim 14, wherein the current picture and the reference picture have the same spatial resolution and the same scaling window offset.

19. The method according to claim 14, wherein the current picture and the reference picture are at the same position in time.

20. A non-temporary computer-readable medium comprising a program, when executed by a processor, that causes the processor to perform the method according to claim 14.

21. An encoder for encoding video data into a data stream (12, 14), The encoder described above is For the current picture encoded in the enhancement layer (L1), a reference picture encoded in the base layer (L0) is identified for use in inter-layer prediction of the current picture. A variable is set based on a comparison between the first configuration setting associated with the current picture and the second configuration setting associated with the reference picture. Based on the aforementioned variables, it is determined that at least one inter-layer prediction tool is unavailable for predicting the blocks of the current picture using the reference picture. Without using the at least one interlayer prediction tool that has been determined to be unavailable, encode the block of the current picture using the reference picture, The process involves generating the data streams (12, 14), wherein the data streams (12, 14) include the enhancement layer (L1) and the base layer (L0). An encoder equipped with a processor configured in such a way.

22. The first configuration setting indicates the number of first sub-pictures included in the current picture, The second configuration setting indicates the number of second sub-pictures included in the reference picture. The encoder according to claim 21.

23. The aforementioned processor further, The number of the first sub-pictures and the number of the second sub-pictures are compared, If the number of the first sub-pictures and the number of the second sub-pictures are different, the variable is set to the first value. Based on the fact that the variable has the first value, it is determined that the at least one interlayer prediction tool is unavailable. The encoder according to claim 22, configured as described above.

24. The aforementioned at least one interlayer prediction tool is: Inter-layer motion vector prediction and, Refinement of optical flow in vector-based inter-layer prediction, or, Motion vector wrap-around in vector-based inter-layer prediction, The encoder according to claim 23, comprising at least one of the following.

25. The encoder according to claim 24, wherein the current picture and the reference picture have the same spatial resolution and the same scaling window offset, and the current picture and the reference picture are at the same position in time.