Video coding in relation to subpicture

By dividing images into sub-pictures with boundary expansion and using aligned boundaries and IDs, the solution addresses the inefficiencies in existing video coding standards, enabling efficient decoding and extraction of sub-pictures on devices with limited resources.

JP2025157506APending Publication Date: 2025-10-15FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2025123809
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-12-20
Filing Date
2025-07-24
Publication Date
2025-10-15

AI Technical Summary

Technical Problem

Existing video coding standards like HEVC and VVC lack efficient mechanisms for handling the extraction and decoding of sub-pictures, particularly in applications involving 360-degree video streaming and multi-party conferencing, where sub-pictures need to be processed independently and efficiently on devices with limited computational resources.

Method used

The proposed solution involves dividing images into sub-pictures with boundary expansion for motion compensation, constraining the number of bins to be coded, using CABAC with zero words, and providing indications for treating sub-pictures as images, ensuring aligned boundaries and IDs across layers, and inserting dummy data for constant bit rate processing.

Benefits of technology

This approach enables efficient decoding and extraction of sub-pictures, optimizing computational resources and maintaining bit rate consistency, thereby enhancing the performance of video coding on devices with limited capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025157506000001_ABST
    Figure 2025157506000001_ABST
Patent Text Reader

Abstract

To provide additional signaling, which improves the current available mechanisms.SOLUTION: A data encoding method includes subdividing a specific image into a plurality of subpictures, determining constraints for bins and bits so that the number of coded bins for the subpictures shall be less than or equal to (32÷3)×NumByteslnVclNalUnits+(RawMinCuBits×PicSizelnMinCbsY)÷32, and encoding the specific image into a bitstream 2605 by CABAC (context-adaptive binary arithmetic coding). The encoding includes inserting one or more zero words into the bitstream so that the constraints for bins and bits are satisfied, corresponding to subpictures in the plurality of subpictures, and provides a display that treats the subpictures as an image with the one or more zero words inserted thereto by CABAC encoding.SELECTED DRAWING: Figure 26
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This application relates to the concept of video coding, and in particular to sub-pictures.

[0002] There are certain video-based applications in which multiple coded video bitstreams or data streams are jointly decoded, for example merged into a joint bitstream and fed to a single decoder, such as in multi-party conferencing where coded video streams from multiple participants are processed at a single endpoint, or in tile-based streaming for example for 360-degree tiled video playback in VR (Virtual Reality) applications.

[0003] For High Efficiency Video Coding (HEVC), a motion-constrained tile set is defined, constraining motion vectors to not reference a tile set (or tile) different from the current tile set (or tile). Thus, the tile set (or tile) in question can be extracted from the bitstream or merged into another bitstream without affecting the decoding results; for example, decoded samples will match exactly regardless of whether the tile set (or tile) in question is decoded alone or as part of a bitstream with more tile sets (or tiles) for each image.

[0004] The latter provides an example of 360-degree video where such techniques are useful. The video is spatially segmented, and each spatial segment is provided to the streaming client in multiple representations at various spatial resolutions, as shown in Figure 1. The figure shows a cubemap projected 360-degree video divided into 6x4 spatial segments at two resolutions. For simplicity, these independently decodable spatial segments are referred to as tiles in this description.

[0005] Users typically view only a subset of the tiles that make up the entire 360-degree video when using modern head-mounted displays, via a solid blue field of view boundary representing a 90x90-degree field of view, as shown on the left side of Figure 2. The corresponding tiles, shaded green in Figure 2, are downloaded at the highest resolution.

[0006] However, the client application also needs to download and decode representations of other tiles outside the current viewport, shaded red in Figure 2, to handle sudden user orientation changes. Thus, a client of such an application downloads tiles covering the current viewport at the highest resolution and tiles outside the current viewport at a lower resolution, always adjusting the tile resolution selection to the user's orientation. After client-side downloading, merging the downloaded tiles into a single bitstream for processing by a single decoder addresses the constraints of typical mobile devices with limited computational and power resources. Figure 3 illustrates a possible tile arrangement for the joint bitstream in the above example. The merging operation to generate the joint bitstream avoids pixel-domain processing, such as transcoding, which would otherwise have to be performed via compressed-domain processing.

[0007] The HEVC bitstream is coded according to some constraints that are primarily related to the inter-coding tool, so that the merging process can be performed, for example constraining the motion vectors as described above.

[0008] The emerging codec VVC offers another, more effective means to achieve the same goal: subpictures. Using subpictures, regions smaller than the entire image can be treated similarly to images, in the sense that their boundaries are treated as if they were images, e.g., by applying boundary extension for motion compensation, e.g., if a motion vector points outside the region, the last sample of the region (crossing the boundary) is repeated to generate a sample in the reference block used for prediction, exactly as is done at the image boundary. This means that motion vectors are not constrained at the encoder as in HEVC MCTS, with a corresponding loss of efficiency.

[0009] The new VVC coding standard also envisages providing scalable coding tools in the main profile for multi-layer support. Therefore, an even more efficient implementation of the above application scenarios can be achieved by encoding the entire low-resolution content with less frequent RAPs. However, this always requires the use of a layered coding structure, where the base layer contains the low-resolution content and the enhancement layers contain somewhat higher-resolution content. The layered coding structure is shown in Figure 1.

[0010] However, there may still be interest in some use cases that allow for the extraction of single regions of the bitstream: for example, a user with a high end-to-end delay downloads the entire 360-degree video in low resolution (all tiles), while a user with a low end-to-end delay downloads fewer tiles of low-resolution content, e.g., the same number of high-resolution tiles that they download.

[0011] Therefore, the extraction of layered sub-pictures needs to be properly handled by the video coding standard, which requires additional signaling to ensure proper knowledge on the decoder or extractor side.

[0012] It is therefore an object of the present invention to provide this additional signaling that improves upon currently available mechanisms.

[0013] This object is achieved by the subject matter of the independent claims of the present application.

[0014] According to a first aspect of the present application, a particular image is divided into multiple sub-pictures, each of which includes a boundary for boundary expansion for motion compensation, allowing each sub-picture to be processed independently as an image.

[0015] According to a second aspect of the present application, the number of bins to be coded for each sub-picture is constrained to be equal to or less than (32 ÷ 3) × NumBytesInVclNalUnits + (RawMinCuBits × PicSizeInMinCbsY) ÷ 32. Based on this constraint, the bin-to-bit ratio is adjusted to maintain an appropriate value when coding the sub-picture.

[0016] According to a third aspect of the present application, the image is encoded into a bitstream using CABAC (Context-Adaptive Binary Arithmetic Coding), and one or more zero words corresponding to the sub-pictures are inserted into the encoded bitstream to satisfy the constraint, thereby maintaining the bin-to-bit ratio of each sub-picture within a predetermined range.

[0017] According to a fourth aspect of the present application, an indication is provided for indicating that each subpicture is treated as a picture with a zero word, and this indication is realized, for example, by sps_subpic_treated_as_pic_flag. If the flag is set to 1, the subpicture is treated the same as a picture in the decoding process. If the flag is not specified, the value can be considered to be 1.

[0018] According to a fifth aspect of the present application, with the above indication that a sub-picture is treated as an image, boundary extension for motion compensation is applied to said sub-picture, i.e., if a motion vector points outside the sub-picture, the last sample of the sub-picture is repeated to form a reference block, just like an image.

[0019] According to a sixth aspect of the present application, the plurality of images includes at least two layers, and the images of at least one layer are divided into a plurality of sub-pictures, and one or more images or sub-pictures of one layer may correspond to one image or sub-picture of one or more other layers.

[0020] According to a seventh aspect of the present application, the bitstream includes an indication that at least one of the boundaries of corresponding sub-pictures or corresponding images in the images of the plurality of layers are aligned with each other, which aligns the positional relationship of sub-pictures between different layers and allows for efficient encoding and extraction processes.

[0021] According to an eighth aspect of the present application, a bitstream containing multiple images is received, and in a configuration in which a specific image among the images is divided into multiple sub-pictures, an indication that the sub-picture should be treated as an image into which one or more zero words corresponding to the sub-picture have been inserted is decoded or inferred from the bitstream.

[0022] According to the ninth aspect of the present application, the indication is interpreted based on sps_subpic_treated_as_pic_flag, and if the flag is set to 1 or is presumed to be 1 even if not explicitly stated, the subpicture is subject to decoding processing as an image.

[0023] According to a tenth aspect of the present application, in decoding a sub-picture, the bin-to-bit ratio constraints imposed in encoding are applied, and the bitstream is analyzed for the presence of zero words to verify satisfaction of the ratio constraints. Zero words are typically inserted at the end of one or more slices of a sub-picture and are identified as part of the decoding process.

[0024] According to an eleventh aspect of the present application, in a bitstream including multiple layers, an image of at least one layer is divided into multiple sub-pictures, and the sub-pictures have correspondence relationships with images or sub-pictures of other layers, thereby enabling a decoder to efficiently limit the scope of interest and perform processing using an indication that sub-picture boundaries between different layers coincide with each other.

[0025] According to a twelfth aspect of the present application, the various processes are realized by a non-transitory computer-readable recording medium in which instructions are recorded as a computer-executable program, the recording medium storing instructions including the various encoding processes, decoding processes, sub-picture identification processes, and zero word analysis processes described above.

[0026] According to a thirteenth aspect of the present application, there is provided an encoding device configured to perform the above method, comprising at least one processor and control means for causing the processor to perform the above-mentioned sequence of operations, the processor being configured to perform bitstream generation, zero word insertion, indication addition, etc.

[0027] Each aspect of the present invention can be implemented in any combination with respect to the encoding, decoding, recording medium, and device configuration described above, and multiple functions may be integrated and realized in a single video codec or processing system.

[0028] Preferred embodiments of the present application are described below with reference to the figures. [Brief explanation of the drawings]

[0029] [Figure 1] It shows 360-degree video in a cubic map projection, arranged in 6x4 tiles at two resolutions. [Figure 2] Showing user viewport and tile selection for 360° video streaming. [Figure 3] 10 shows the resulting tile arrangement (packing) in the joint bitstream after the merge operation. [Figure 4] 1 shows a scalable sub-picture-based bitstream. [Figure 5] 1 shows exemplary syntax elements. [Figure 6] 1 shows exemplary syntax elements. [Figure 7] 1 shows exemplary syntax elements. [Figure 8] 1 shows exemplary syntax elements. [Figure 9] 1 shows exemplary syntax elements. [Figure 10] 1 shows exemplary syntax elements. [Figure 11] 1 shows exemplary syntax elements. [Figure 12] 1 shows an image of different layers divided into sub-pictures. [Figure 13] It shows the boundaries that are aligned boundaries between the lower and upper layers, and the boundaries of the upper layer that have no correspondence in the lower layer. [Figure 14] 1 shows exemplary syntax elements. [Figure 15] 1 shows an exemplary region of interest (RoI) provided in the low resolution layer but not in the high resolution layer. [Figure 16] 1 illustrates an exemplary layer and sub-picture configuration. [Figure 17] 1 shows exemplary syntax elements. [Figure 18] 1 shows exemplary syntax elements. [Figure 19] 1 shows a data stream with constant bit rate where the subpictures are filled with dummy data for subpicture extraction. [Figure 20] 1 shows a data stream with constant bit rate where the subpictures are filled with dummy data for subpicture extraction. [Figure 21] 1 shows exemplary syntax elements. [Figure 22] 1 shows exemplary syntax elements. [Figure 23] 1 shows exemplary syntax elements. [Figure 24] 1 shows exemplary syntax elements. [Figure 25] 1 shows exemplary syntax elements. [Figure 26] Illustrates encoding using cabac zero words. DETAILED DESCRIPTION OF THE INVENTION

[0030] Below are described additional embodiments and aspects of the present invention, which can be used individually or in combination with any of the features and functions and details described herein.

[0031] The first embodiment concerns layers and sub-pictures, in particular sub-picture boundary alignment (2.1).

[0032] The signaling indicates that subpicture boundaries are aligned across layers, e.g., there are the same number of subpictures per layer, and all subpicture boundaries are sample-accurately collocated. This is shown, for example, in Figure 4. Such signaling can be implemented, for example, by sps_subpic_treatment_as_pic_flag. If this flag sps_subpic_treatment_as_pic_flag[i] is set, e.g., has a value equal to 1, it specifies that the i-th subpicture of each coded picture in the coded layer-wise video sequence CLVS is to be treated as a picture in the decoding process, except for in-loop filtering operations. If sps_subpic_treatment_as_pic_flag[i] is not set, e.g., has a value of 0, it specifies that the i-th subpicture of each coded picture in the CLVS is not to be treated as a picture in the decoding process, except for in-loop filtering operations. If the flag is not present, it can be considered to be set, e.g., the value of sps_subpic_treatment_as_pic_flag[i] is inferred to be equal to 1.

[0033] The signaling is useful for extraction because the extractor / receiver can look up the subpicture IDs present in the slice header and drop all NAL units that do not belong to the subpicture of interest. When subpictures are aligned, subpicture IDs have a one-to-one mapping between layers. This means that for every sample in enhancement layer subpicture A identified by subpicture ID A, the collocated sample in base layer subpicture B belongs to a single subpicture identified by subpicture ID Bn.

[0034] It should be noted that without the alignment as described, for each sub-picture in the reference layer there may be multiple sub-pictures in the reference layer, which is much worse from the point of view of bitstream extraction or parallelization, as the sub-pictures of the layers overlap partially.

[0035] In a particular implementation, the indication of subpicture alignment can be done by a constraint flag in the general_constraint_info() structure of the profile_tier_level() indicated in the VPS, as this is layer-specific signaling. This is shown, for example, in Figure 5.

[0036] The syntax can indicate that a layer in an Output Layer Set (OLS) is subpicture-aligned. Another option is to indicate alignment of all layers in the bitstream, for example by signaling OLS independently for all OLSs.

[0037] In other words, in this embodiment, the data stream is decoded into multiple images of video, and the data stream includes the multiple images in at least two layers. In this example, one layer is a base layer ("low resolution" in FIG. 4) and the other layer is an enhancement layer ("high resolution" in FIG. 4). The images in both layers are divided into sub-pictures, sometimes called tiles, and at least in the example of FIG. 4, each sub-picture in the base layer has a corresponding sub-picture in the enhancement layer. The sub-picture boundaries shown in FIG. 4 are used for boundary extending for motion compensation.

[0038] In decoding, there is an indication that at least one of the boundaries of corresponding subpictures of different layers is interpreted as being aligned with one another. As a brief note, it should be noted that throughout this description, subpictures of different layers are understood to correspond to one another due to their co-location, i.e., they are located at the same position in the image in which they are contained. A subpicture is coded independently of any other subpictures in the same image in which it is contained. However, subpictures of a particular layer, the term describing corresponding subpictures of an image of that layer, and these corresponding subpictures are coded using motion-compensated prediction and temporal prediction so that no coding dependency is required from areas outside these corresponding subpictures, and so that, for example, boundary extensions are used for motion-compensated reference portions pointed to by motion vectors of blocks in these corresponding subpictures, i.e., portions of these reference portions that extend beyond the boundaries of the subpicture. The subdivision of the image of such a layer into subpictures is done in such a way that overlapping subpictures correspond to one another and form a kind of independently coded subvideo. It should be noted that subpictures may also function in conjunction with inter-layer prediction. When inter-layer prediction is available for coding / decoding, one image of a layer is predicted from another image of the layer below, so that both images are at the same time. A region of one image of a layer, called a block, extends beyond the boundaries of a subpicture of a reference image of another layer, and the regions where a block is collocated, i.e., the regions where the blocks are overlaid, are filled in by extension, i.e., by not using content from outside the collocated subpicture of the reference image.

[0039] If an image is divided into only one sub-picture, i.e., if the image is treated as a sub-picture, the images correspond to each other across layers, indicating that the image boundaries are aligned and used for boundary expansion for motion compensation.

[0040] Each of the multiple images in the at least two layers may also be divided into one or more sub-pictures, and the display may be interpreted to the extent that all boundaries of corresponding sub-pictures in different layers are aligned with each other.

[0041] The following is about the correspondence of subpicture IDs (2.2).

[0042] In addition to aligning subpicture boundaries, it is also necessary to facilitate the detection of corresponding subpicture IDs between layers when the subpictures are aligned.

[0043] In the first option, the subpicture IDs in the bitstream are different, i.e., each subpicture ID can only be used in one layer of the OLS. This allows each layer subpicture to be uniquely identified by its own ID. This is shown, for example, in Figure 6.

[0044] Note that this may also be true for IDs that are unique across the bitstream, e.g., for any OLS in the bitstream, regardless of whether two layers belong to one OLS. Thus, in one embodiment, the value of any subpicture ID is unique within the bitstream (e.g., unique_ids_in_bitstream_flag).

[0045] In the second option, subpictures in a layer have the same ID value as the corresponding aligned subpicture in another layer (in OLS or the entire bitstream). This makes it very easy to identify corresponding subpictures by simply matching the subpicture ID values. The extraction process is also simplified, as only one ID per subpicture of interest is needed by the extractor. This is shown, for example, in Figure 7. Instead of aligned_subpictures_ids_in_ols_unique_flag, the SubpicIdVal constraint can also be used.

[0046] The following concerns the signaling order of sub-pictures (2.3):

[0047] The handling of subpicture ID correspondence can be greatly simplified if the signaling order, position, and dimensions of subpicture IDs are also constrained. This means that if corresponding aligned subpictures and unique subpicture IDs are indicated (e.g., the flags subpictures_in_ols_aligned_flag and aligned_subpictures_ids_in_ols_unique_flag are set to 1), the subpicture definitions (including their positions, widths, heights, border handling, and other properties) must be aligned across all SPSs on the OLS.

[0048] For example, the syntax element sps_num_subpics_minus1 must have the same value in a layer and all its reference layers, as well as subpic_ctu_top_left_x[i], subpic_ctu_top_left_y[i], subpic_width_minus1[i], and subpic_height_minus1[i], taking into account possible resampling factors.

[0049] In this embodiment, the order of subpictures in the SPS of all layers (see the green marking of syntax elements in the SPS) is constrained to be the same. This allows for easy matching and verification of the position between subpictures and layers. This is shown for example in Figure 8.

[0050] Subpicture ID signaling may occur in the respective override functions of SPS, PPS, or PH, as shown below, and if aligned_subpictures_ids_in_ols_flag is 1, the subpicture ID values ​​MUST be the same in all layers. If unique_subpictures_ids_in_ols_flag is 1, the subpicture ID values ​​MUST NOT occur in multiple layers. This is shown, for example, in Figures 9 to 11.

[0051] In other words, in this embodiment, the indication is interpreted as meaning that the relative subpicture positions and subpicture dimensions are the same for corresponding subpictures in all layers. For example, the indices of syntax elements indicating the positions and dimensions of corresponding subpictures, representing the signaling order within the data stream, are the same. The indication may include one flag per layer of the data stream, a flag indicating subpicture alignment for the layer whose flag is present in the data stream, and one or more upper layers, such as one that uses any of the images of that flag's layer as a reference image.

[0052] This is illustrated in Figure 12, which shows the layers of a layered data stream. Layer 1 image 1211, for example, is not divided into subpictures, i.e., it is divided into only one subpicture, so there is a (sub)picture L1(1). Layer 2 image 1212 is divided into four subpictures L2(1) to L2(4). The images of the three to four higher layers are divided in the same way. Note that the images are coded using inter-layer prediction, i.e., inter-layer prediction can be used for coding / decoding. One image of one layer can be predicted from another image of the layer below, and both images are at the same time. Again, subpicture boundaries cause cross-prediction, i.e., an area called a block of one image of one layer extends beyond the boundary of a subpicture of a reference image of another layer, and a single block is co-located and filled by not using content from outside the co-located subpicture of the reference image. Furthermore, for each layer divided into subpictures, the image may need to be a constant size; that is, the size of the image for each layer does not change; that is, RPR is not used. For a single subpicture layer, such as layer L1, various picture sizes and RPRs may be allowed. Next, the display flag may indicate that the subpictures for layers 2 and above are aligned, as described above. That is, the images 1213 and 1214 in the higher layers, layers 3 and above, all have the same subpicture division as layer 2. The subpictures for layers 3 and above also have the same subpicture ID as layer 3. As a result, layers above 2 also have subpictures LX(1) through LX(4), where X is the layer number. Figure 12 exemplarily illustrates this for several higher layers 3 and 4, but there could also be one or more higher layers, in which case the display also relates. Therefore, the subpictures for each layer above the layer for which the display is made (here, layer 2), i.e., the subpictures for which the flag is set, have the same index. 1, 2, 3, or 4 as the corresponding subpicture in layer 2, and has the same size and juxtaposed borders as the corresponding subpicture in layer 2.Corresponding sub-pictures are those that are juxtaposed to each other within a picture. Again, such indication flags may also exist in other layers, where the same sub-picture subdivision may also be set in higher layers such as Layer 3 and Layer 4.

[0053] When each of the multiple images of at least two layers is divided into one or more sub-pictures, the display indicates that, for each of the one or more layers, the images of one or more higher layers higher than the respective layer are divided into sub-pictures, such that the number of predetermined layer-specific sub-pictures is equal for each layer and the one or more higher layers, and the boundaries of the predetermined layer-specific number of sub-pictures spatially coincide between each layer and the one or more higher layers.

[0054] In particular, applying this to the image of FIG. 12, it becomes clear that if the sub-pictures in layers 2 and above are shown to be aligned, then layers 3 and above will also have correspondingly aligned sub-pictures.

[0055] One or more higher layers can use the images of the respective layer for prediction.

[0056] Also, the subpicture IDs are the same between each layer and one or more upper layers.

[0057] Also shown in Figure 12 is a subpicture extractor 1210 for coding, from the bitstream in which the just-discussed layers are coded, an extracted bitstream specific to a particular set of one or more subpictures in layers 2-4. From the complete bitstream, it extracts portions related to regions of interest in the layered image, here extracting the upper right corner of the image as an example of a subpicture. Thus, extractor 1210 extracts portions of the entire bitstream related to the tie Route of Interest (ROI), i.e., portions related to the entire image of any layer for which the above indication does not indicate subpicture alignment, i.e., the above flag is not set, as well as portions related to subpictures of images in mutually aligned layers, i.e., portions for which the flag is set, and any upper layers whose subpictures overlay the RoI, here L2(2), L3(2), and L4(2) as examples. In other words, any single subpicture layer, such as L1 in the case of FIG. 12, that is in the output layer set and serves as a reference for subpicture layers, here 2-4, is included in the bitstream extracted by extractor 1210, and in the extracted bitstream, the HLS inter-layer prediction parameters are adjusted accordingly; that is, in the extracted bitstream, the scaling window conditions are adjusted to reflect that the higher layer images split into previous subpictures in layers 2-4 are small enough to be cropped only to subpictures within the RoI. Extraction is straightforward because the indication shows that the subpictures are aligned with each other. Another way to utilize the indication of subpicture alignment could be a decoder that organizes decoding more easily.

[0058] The following is about the border subset of the sub-picture in the lower layer (2.4).

[0059] There may be other use cases, such as fully parallel coding of regions within a layer to speed up coding, where the higher the resolution and the higher the layer, the greater the number of sub-pictures (spatial scalability). In such cases, alignment is also desirable. A further embodiment is that when the number of sub-pictures in an upper layer is greater than the number of sub-pictures in a lower (reference) layer, the boundaries of all sub-pictures in the lower layer have a counterpart in the upper layer (co-located boundaries). Figure 13 shows boundaries that are aligned between the lower and upper layers, and boundaries in the upper layer that have no counterpart in the lower layer.

[0060] An exemplary syntax for reporting the property is shown in FIG.

[0061] The following concerns the impact of layered motion compensated prediction on sub-picture boundaries (2.5).

[0062] Whether a sub-picture is only within a layer or is an independent sub-picture between layers is indicated in the bitstream. More specifically, this indication indicates whether the motion compensated prediction performed between layers also takes sub-picture boundaries into account. In one example, as shown in Figure 15, an RoI may be provided in a low-resolution version (lower layer) of content (e.g., 720p RoI in 1080p content) but not in a high-resolution content (e.g., a higher layer with 4k resolution).

[0063] In another embodiment, whether motion compensated prediction performed between layers also takes subpicture boundaries into account depends on whether the subpictures are aligned between layers. If so, the boundaries are taken into account when considering motion compensation, such as inter-layer prediction. For example, motion compensation is not allowed to use sample positions outside the boundaries, or sample values ​​at such sample positions are not allowed to be extrapolated from sample values ​​within the boundaries. Otherwise, boundaries are ignored in inter-layer motion compensated prediction.

[0064] The following is for subpicture downsizing reference OLS (2.6).

[0065] Another use case that can be considered for the use of layered sub-pictures is RoI scalability. A diagram of a potential layer and sub-picture configuration for such a use case is shown in Figure 16. In such a case, only the RoI portion or sub-picture of the lower layer is needed for the higher layer, meaning that fewer samples need to be decoded when only the RoI is of interest, i.e., when decoding the 720p version at the base layer or enhancement layer in the given example. Nevertheless, it is necessary to indicate to the decoder that only a subset of the bitstream needs to be decoded, and therefore the level of the sub-bitstream associated with the RoI is lower (fewer samples decoded) than the entire base layer and enhancement layer.

[0066] In this example, instead of 1080+4K, only 720+4K needs to be decoded.

[0067] In one embodiment, the OLS signaling indicates that only a subpicture is needed, rather than a complete layer. For each output layer, when it is indicated (reduced_subpic_reference_flag equals 1), a list of associated subpicture IDs to be used for reference is given (num_sub_pic_ids, subPicIdToDecodeForReference). This is shown, for example, in Figure 17.

[0068] In another embodiment, additional PTL signaling is presented to indicate the PTLs that would be needed if unnecessary subpictures of layers in the OLS are removed from the OLS bitstream. The options are shown in Figure 18, where the respective syntax is added to the VPS.

[0069] The following concerns constant bitrate (CBR) and subpictures (3):

[0070] According to an aspect, a video processing device may be configured to process multiple images of a video decoded from a data stream including multiple images, where each of the multiple images, e.g., all layers, is divided into a predetermined number of sub-pictures, each sub-picture including a boundary for boundary extension for motion compensation, and the video processing device is configured to generate at least one sub-picture from the data stream at a constant bit rate by maintaining dummy data, examples of which are FD_NUT and filler payload Supplemental Enhancement Information (SEI) messages included in the data stream for a sub-picture immediately following a sub-picture, or included in the data stream for a sub-picture that is not adjacent to a sub-picture but includes an indication of the sub-picture, e.g., sub-picture identification information, until another sub-picture occurs in the data stream. Note that if the data stream is a layered data stream, this coded image may be associated with one layer, and per-sub-picture bit rate control may be applied to each layer.

[0071] According to another aspect, a video encoder may be configured to encode a plurality of images of a video into a data stream including the plurality of images, each of the plurality of images of every layer being divided into a predetermined number of sub-pictures, each sub-picture including a boundary for boundary extension for motion compensation, and the video encoder is configured to generate at least one sub-picture in the data stream at a constant bitrate by including dummy data in the data stream, e.g., FD_NUT and a filler payload Supplemental Enhancement Information, SEI, message, and for each sub-picture, a sub-picture immediately following the respective sub-picture or not adjacent to the respective sub-picture but including an indication of the sub-picture, e.g., sub-picture identification information.

[0072] According to another aspect, a method for processing video may include processing a plurality of images of the video decoded from a data stream including a plurality of images, wherein each of the plurality of images, e.g., all layers, is divided into a predetermined number of sub-pictures, each sub-picture including a boundary for boundary extension for motion compensation, and the method may include generating at least one sub-picture from the data stream at a constant bit rate by maintaining dummy data, examples of which are FD_NUT and filler payload supplemental enhancement information, SEI, messages included in the data stream for a sub-picture immediately following the sub-picture, or included in the data stream for a sub-picture that is not adjacent to the sub-picture but includes an indication of the sub-picture, e.g., sub-picture identification information, until another sub-picture occurs in the data stream.

[0073] According to another aspect, a method for encoding a video may include encoding a plurality of images of the video into a data stream including the plurality of images, wherein each of the plurality of images, e.g., the plurality of images of all layers, is divided into a predetermined number of sub-pictures, each sub-picture including a boundary for boundary extension for motion compensation, and the method also includes encoding at least one sub-picture into the data stream at a constant bit rate by including dummy data, e.g., FD_NUT and filler payload Supplemental Enhancement Information, SEI, messages in the data stream, wherein for each sub-picture, there is an indication of the sub-picture, e.g., sub-picture identification information, either immediately following the sub-picture or not adjacent to the sub-picture.

[0074] A bitstream containing subpictures may be coded and the associated HRD syntax element defines a constant bitrate (CBR) bitstream, e.g., cbr_flag is set to 1 for at least one scheduling value to indicate that the bitstream corresponds to a constant bitrate bitstream, frequently by using so-called filler data, VCL NAL units or filler payload SEI (non-VCL NAL units), e.g., FD_NUT and filler payload SEI messages.

[0075] However, it is unclear whether this property still applies when a subpicture bitstream is extracted.

[0076] In one embodiment, the sub-picture extraction process is defined in a way that always results in a VBR bitstream. There is no possibility to indicate / guarantee the CBR case. In that case, the FD_NUT and filler payload SEI messages are simply discarded during the extraction process.

[0077] In another embodiment, it is indicated that the CBR "operation point" of each subpicture is guaranteed by placing the respective FD_NUT and filler payload SEI messages immediately after the VCL NAL units that constitute the subpicture, and thus during the extraction of the subpicture's CBR "operation point," and the respective FD_NUT and filler payload SEI messages associated with the subpicture VCL NAL units are retained when the sub-bitstream of the subpicture is extracted during the extraction process, and are retained during the extraction process aimed at another subpicture or at a non-CBR "operation point." Such an indication can be performed, for example, using the sli_cbr_constraint_flag. Thus, for example, if sli_cbr_constraint_flag is equal to 0, all NAL units with nal_unit_type equal to FD_NUT and SEI messages containing filler payload SEI messages are removed.

[0078] In the described process, to know whether an FD_NUT or Filler Payload SEI message has been dropped, it is necessary to maintain the state of the VCL NAL unit and its associated non-VCL NAL units, and it is necessary to know what the subpicture ID of the immediately preceding VCL NAL unit was. To facilitate this process, signaling is added to the FD_NUT and Filler Payload SEI to indicate that they belong to a particular subpicture ID. Alternatively, in another embodiment, an SEI message is added to the bitstream indicating that the next SEI message or NAL unit belongs to a given subpicture with a subpicture ID, until the presence of another SEI message indicating a different subpicture ID.

[0079] In other words, a constant bit rate data stream having multiple images encoded therein is processed so that each of the multiple images is divided into a predetermined number of sub-pictures, and each sub-picture includes a boundary for boundary extension for motion compensation, and through the processing, a sub-picture data stream related to at least one sub-picture of the constant bit rate is generated from the data stream by retaining dummy data included in the data stream of at least one sub-picture and removing dummy data included in the data stream of another sub-picture to which the extracted data stream does not relate.

[0080] This can be seen in Figures 19 and 20, which show two exemplary images 1910 and 1920 of video coded into a processed data stream. These images are split into sub-pictures 1911, 1912, 1913, and 1921, 1922, and 1923, respectively, which are then encoded by some encoder 1930 into corresponding access units 1940 and 1950 of a bitstream 1945.

[0081] The subdivision into sub-pictures is shown for illustrative purposes only. For purposes of explanation, the access units AU are each shown to also include several header information portions, HI1 and HI2, but this is for illustrative purposes only, and the picture is shown to be coded into a bitstream 1945 that is fragmented into several portions, such as VCL NAL units 1941, 1942, 1943, 1951, 1952, and 1953, respectively, and although there is one VCL unit per sub-picture of pictures 1940 and 1950, this is done for illustrative purposes only, and one sub-picture may result from fragmenting / coding into multiple such portions or VCL NAL units.

[0082] The encoder creates a constant bit rate data stream by including dummy data d at the end of each data portion representing a subpicture, i.e., dummy data 1944, 1945, 1946, 1954, 1955, and 1956 for each subpicture. The dummy data is indicated using crosses in Figure 19, where dummy data 1944 corresponds to subpicture 1911, 1945 corresponds to subpicture 1912, 1946 corresponds to subpicture 1913, 1954 corresponds to subpicture 1921, 1955 corresponds to subpicture 1922, and 1956 corresponds to subpicture 1923. For some subpictures, dummy data may not be necessary.

[0083] Optionally, as mentioned above, a subpicture can be distributed across multiple NAL units, in which case dummy data can be inserted at the end of each NAL unit or at the end of the last NAL unit of that subpicture.

[0084] The constant bitrate data stream 1945 may then be processed by a subpicture extractor 1960, which extracts, by way of example, information relating to subpicture 1911 and corresponding subpicture 1921, i.e., the co-located and therefore mutually corresponding video subpictures coded in bitstream 1945. There may be multiple sets of extracted subpictures. The extraction results in extracted bitstream 1955. For extraction, NAL units 1941 and 1951 of extracted subpictures 1911 / 1921 are extracted and carried over to extracted bitstream 1955, while other VCL NAL units of other subpictures that are not extracted are ignored or dropped. Only dummy data d corresponding to extracted subpicture (2) is carried over to extracted bitstream 1955, and dummy data for all other subpictures is also dropped. That is, for other VCL NAL units, dummy data d is deleted. If necessary, the header portions HI3 and HI4 may be modified as detailed below: they may relate to or include parameter sets such as PPS and / or SPS and / or VPS and / or APS.

[0085] The extracted bitstream 1955 therefore contains access units 1970 and 1980, which correspond to access units 1940 and 1950 of bitstream 1945, and is at a constant bitrate. The bitstream thus produced by extractor 1960 can then be decoded by decoder 1990 to produce a sub-video, i.e., a video containing sub-pictures consisting only of extracted sub-pictures 1911 and 1921. The extraction may, of course, affect multiple sub-pictures of the images of the video of bitstream 1945.

[0086] As another alternative, the dummy data may be ordered within each access unit at the end of the data stream, regardless of whether one or more subpictures are coded into one or more NAL units, but with each dummy data associated with the corresponding subpicture coded into the respective access unit. This is shown in Figure 20, labeled access units 2040 and 2050, instead of 1940 and 1950.

[0087] Note that the dummy data may be NAL units of a particular NAL unit type so that they can be distinguished from VCL NAL units.

[0088] Another way of obtaining bitstream 1945 as being at a constant bit rate can also be illustrated in Figure 20, but here the constant bit rate is not created for each sub-picture. Rather, it is created globally. As shown in Figure 20, dummy data can now be included entirely in all sub-pictures at the end of an AU, and only the data stream portions corresponding to Figure 19 are illustratively shown therein, in access units 2040 and 2050. The data portions representing sub-pictures, 2044, 2045, 2046, 2054, 2055, and 2056, respectively, directly follow each other, with dummy data added at the end of the sub-pictures, at the end of their respective access units.

[0089] It is indicated in the bitstream processed by the extractor whether subpicture selection dummy data removal should be performed, i.e., whether all dummy data should be removed when extracting the subpicture bitstream, or whether the dummy data(s) of the subpicture(s) of interest, i.e., the one(s) to be extracted, should be retained as shown in Figure 19.

[0090] Also, if the data stream is indicated to be not constant bit rate per subpicture, the extractor 1960 can remove all dummy data d in all NAL units, and the resulting extracted subpicture bitstream will not be constant bit rate.

[0091] Note that the placement of the dummy data within an access unit may vary compared to the example just described, and may be placed between VCL units or at the beginning of an access unit.

[0092] It should be noted that the above and below processes of the processor or extractor 1960 can also be performed by, for example, a decoder.

[0093] In other words, the data stream can be processed by interpretation, e.g., decoding and parsing, to indicate that at least one extractable set of sub-pictures of the data stream can be extracted such that the bitstream of the extracted sub-pictures is at a constant bitrate. This indication can be processed from or included in, for example, a video parameter set, VPS, or SEI message, and can be, for example, cbr_flag or sli_cbr_constraint_flag.

[0094] The data stream may also interpret and process an indication, for example in a video parameter set, VPS, or SEI message, that an extractable set of one or more subpictures of the data stream is not coded at a constant bit rate, and if so indicated, all dummy data is removed.

[0095] As mentioned above, the dummy data may include FD_NUT, a special NAL unit. Additionally or alternatively, the dummy data may include filler payload supplemental enhancement information, SEI, messages.

[0096] The data stream can also be processed using interpretations, for example decoding an indication from the data stream for generation of an extracted data stream whether the generation will result in a constant bit rate extracted data stream, in which case the video processing device will retain and delete dummy data d, or whether the generation will not result in a constant bit rate extracted data stream, in which case the video processing device will delete all dummy data.

[0097] Finally, the aforementioned dummy data may include, for example, FD_NUT and filler payload supplemental extension information, SEI, messages.

[0098] The following is the completion of the subpicture extraction and the rewriting of the parameter set (4).

[0099] According to another aspect, a video processing device may be configured to perform sub-picture extraction on a data stream including multiple images of a video, where each of the multiple images of all layers is divided into a predetermined number of sub-pictures, each sub-picture including a border for border extension for motion compensation, and interpret an indication in the data stream that it is permitted to rewrite the parameter set and / or picture header of the data stream when performing sub-picture extraction on the data stream.

[0100] According to another aspect when referring to the previous aspect, the video processing device may be further configured to interpret a representation of a data stream providing additional information for rewriting information regarding HRD-related information, image size, and / or subdivision of multiple images into sub-pictures.

[0101] According to yet another aspect, a video encoder may be configured to encode a plurality of images of a video into a data stream, the data stream including the plurality of images, where each of the plurality of images of all layers is divided into a predetermined number of sub-pictures, each sub-picture including a boundary for boundary extension for motion compensation, and indicating that rewriting a parameter set and / or a picture header of the data stream is permitted when performing sub-picture extraction in the data stream.

[0102] According to another aspect when referring to the previous aspect, the video encoder may be further configured to interpret a representation of a data stream that provides additional information for rewriting information regarding HRD-related information, image size, and / or subdivision of multiple images into sub-pictures.

[0103] According to another aspect, a method for processing video may be configured to perform sub-picture extraction on a data stream including multiple images of the video, where each of the multiple images of all layers is divided into a predetermined number of sub-pictures, each sub-picture including a border for border extension for motion compensation, and interpret an indication in the data stream that a parameter set and / or a picture header of the data stream is allowed to be rewritten when performing sub-picture extraction on the data stream.

[0104] According to another aspect when referring to the previous aspect, the method for processing video may be further configured to display additional information in the data stream for rewriting information regarding HRD-related information, image size, and / or subdivision of multiple images into sub-pictures.

[0105] According to another aspect, a method for processing video may include performing sub-picture extraction on a data stream including a plurality of images of the video, where each of the plurality of images of all layers is divided into a predetermined number of sub-pictures, each sub-picture including a boundary for boundary extension for motion compensation, and interpreting an indication in the data stream that rewriting a parameter set and / or picture header of the data stream is permitted when performing sub-picture extraction on the data stream.

[0106] According to another aspect when referring to the previous aspect, the method for processing video may further comprise a step of interpreting a representation of a data stream providing additional information for rewriting information regarding HRD-related information, image size, and / or subdivision of multiple images into sub-pictures.

[0107] Another aspect may refer to a data stream for which video is encoded, the video including multiple images, where each of the multiple images of all layers is divided into a predetermined number of sub-pictures, each sub-picture including a boundary for boundary extension for motion compensation, and where it is permitted to rewrite the parameter set and / or picture header of the data stream when performing sub-picture extraction in the data stream.

[0108] Typically, the bitstream extraction process is specified for a (sub)layer as follows: - Drop NAL units that do not correspond to the target (sub)layer - Dropping SEI messages (Picture Timing SEI, Buffering Period SEI, etc.) - Optionally, retrieve the appropriate SEI message from the nested SEI messages associated with the target (sub)layer

[0109] However, as explained above, dropping NAL units at a finer granularity than per (sub)layer, e.g., sub-picture extraction, can cause problems that have not been solved so far. See the above description of FD_NUT for CBR for an example. Another problem arises from parameter sets. In (sub)layer extraction, a parameter set may contain additional information that is not related to the extracted (sub)layer. For example, it may contain information about a dropped (sub)layer, but still be correct in the sense that it describes the (sub)layer of the extracted sub-bitstream. This additional information for the dropped (sub)layer may simply be ignored. For sub-picture extraction, this design principle is too complex to maintain; therefore, for example, in HEVC, replacement parameter sets are included in the SEI message that nests the MCTS. However, VVC defines a picture header that further complicates the situation and makes an HEVC-style solution of nesting SEI messages infeasible.

[0110] Therefore, the parameter set is modified when performing subpicture extraction. Note that the parameter set contains information such as image size, profile_level information, tiling / slice grid (which may also be present in the image header) that needs to be modified when performing subpicture extraction.

[0111] In one embodiment, the extraction process is defined as follows: - Drop NAL units that do not correspond to the target subpicture - Dropping SEI messages (Picture Timing SEI, Buffering Period SEI, etc.) - Retrieve the appropriate SEI message from the nested SEI messages associated with the target (sub)layer - Drop the parameter set - Add the appropriate parameter set - Drop Image Header - Add the appropriate image header

[0112] "Proper" parameter sets and image headers need to be generated, most likely by rewriting the parameter sets and image headers before extraction. The information that needs to be changed is as follows: - Level and HRD related - Image size - Tiling / Slicing / Subpicture Grid

[0113] The resulting level and HRD related information can be extracted from the respective SEI messages (sub-picture level information SEI messages) and used to rewrite the parameter sets.

[0114] The size of each subpicture can be easily found in the SPS, so if there is a single subpicture extracted in the process, the resulting image size can be easily derived. However, if multiple subpictures are extracted from the bitstream, e.g., the resulting bitstream contains multiple subpictures, the resulting image size will vary depending on the placement defined by external means.

[0115] In one embodiment, the bitstream is constrained such that only extraction of sub-pictures corresponding to rectangular regions in the original bitstream is permitted, and their relative placement in the original bitstream is kept unchanged, so no reordering or gaps are allowed within the extracted regions.

[0116] The coordinates of a subpicture are currently defined by the top_left coordinate and parameters representing the width and height, as shown in Figure 21. This allows the top left, top right, bottom left, and bottom right corners or areas to be easily derived. Therefore, the extraction process is performed for all extracted subpictures, and if a smaller coordinate is found (e.g., searching for a minimum value), the minimum "x" and "y" coordinates of the top_left coordinate are updated, and if a smaller coordinate is found (e.g., searching for a maximum value), the maximum "x" and "y" coordinates of the bottom right coordinate are updated. For example:

[0117] MinTopLeftX=PicWidth MinTopLeftY= PicHeight MaxBottomRightX=0 MaxBottomRightY=0 For i=0..NumExtSubPicMinus1 If(subpic_ctu_top_left_x[ i ] <MinTopLeftX) MinTopLeftX= subpic_ctu_top_left_x[ i ] If(subpic_ctu_top_left_y[ i ] <MinTopLeftY) MinTopLeftY= subpic_ctu_top_left_y[ i ] If(subpic_ctu_top_left_x[ i ]+ subpic_width_minus1[i]>MaxBottomRightX) MaxBottomRightX= subpic_ctu_top_left_x[ i ]+ subpic_width_minus1[i] If(subpic_ctu_top_left_y[ i ]+ subpic_height_minus1[i]> MaxBottomRightY) MaxBottomRightY= subpic_ctu_top_left_y[ i ]+ subpic_height_minus1[i]

[0118] This value can then be used to derive the maximum image size and new values ​​for subpic_ctu_top_left_x[i], subpic_ctu_top_left_y[i] for each extracted subpicture by subtracting the respective value of MinTopLeftX or MinTopLeftY.

[0119] In an alternative embodiment, signaling is provided that allows rewriting of parameter sets and picture headers without requiring or deriving the discussed values. For each potential combination of extracted sub-pictures, the picture sizes that need to be rewritten in the parameter sets can be provided. This can be done, for example, in the form of an SEI message. The SEI message can contain a parameter_type syntax element that indicates that there is information inside that is useful for rewriting. For example, type 0 could be the picture size, type 1 level of the extracted bitstream, type 2 information in the picture header, or a combination thereof.

[0120] In other words, sub-picture extraction is performed on a data stream having multiple images of video coded into multiple layers, where each of the multiple images of all layers is divided into a predetermined number of sub-pictures, and each sub-picture includes a boundary for boundary extension for motion compensation. Sub-picture extraction is performed on the data stream by dropping NAL units of the data stream that do not correspond to one or more sub-pictures and by rewriting parameter sets and / or picture headers into an extracted data stream related to one or more sub-pictures of interest.

[0121] This will be explained again with respect to Figure 19. Sub-pictures 1912, 1913, 1922, 1923 do not belong to the sub-picture(s) of interest and are therefore not required in sub-picture extraction 1960, and the corresponding NAL units 1942, 1943, 1952, 1953 can be dropped entirely by the extraction. This time, the presence of dummy data in Figure 19 should be understood to be optional, or if present, the same can be non-specific by sub-picture, as explained in Figure 20.

[0122] 19, the extractor 1960 can rewrite the header information HI1 and HI2 to HI3 and HI4, respectively. In this process, which can occur simultaneously when extracting sub-pictures, the header information, which can be, for example, a video parameter set, VPS, and / or a sequence parameter set, SPS, and / or a picture parameter set, PPS, etc., can be changed by rewriting certain variable values ​​thereof.

[0123] The information for rewriting can be derived from or present in the data stream, and in addition to the actual header information HI1 and HI2, additional information may relate to HRD-related information, picture size and / or subdivision into sub-pictures. An indication that rewriting is allowed or that additional and / or additional information is present can be indicated in the data stream, for example with supplemental extension information, SEI, messages.

[0124] The sub-picture extraction may be performed such that the extracted data stream consists of a set of extracted sub-pictures, each of which is further configured to be rectangular in that it corresponds to a rectangular region in the data stream.

[0125] The values ​​that can be rewritten include the sub-picture width of the luminance sample, the sub-picture height of the luminance sample, level information, HRD-related parameters, and picture size.

[0126] The following is about the removal of emulation prevention for sub_pic_id(5).

[0127] Subpicture IDs are currently signaled in slice headers in VVC, as shown in Figure 22 for example.

[0128] However, since they are listed as one of the first syntax elements in the slice header, their values ​​are easily accessible. Note that the extraction process (or any merging process that requires changing / checking the IDs) needs to read and write the subpicture ID values, so easy access is desirable. One aspect that is still missing for easy access to their values ​​is emulation prevention. The syntax element slice_pic_order_cnt_lsb is a maximum of 16 bits long and can take the value 0, e.g., 0x0000. The slice_sub_pic_id is 16 bits long. Therefore, depending on the combination of the slice_pic_order_cnt_lsb and slice_sub_pic_id values, emulation prevention may be performed. This makes parsing the subpicture IDs for higher layer applications more complex and makes changing or modifying the higher layer applications even more difficult, because the length of the emulated prevention in the slice_header may change if the value of slice_sub_pic_id is changed to another value of slice_sub_pic_id*.

[0129] In one embodiment, slice_pic_order_cnt_lsb is changed to slice_pic_order_cnt_lsb_plus1 and slice_sub_pic_id is changed to slice_sub_pic_id_plus1, so emulation prevention does not occur between slice_pic_order_cnt_lsb and slice_sub_pic_id. This is useful for parsing slice_sub_pic_id, but does not solve use cases where the value of slice_sub_pic_id needs to be changed. Note that slice_sub_pic_id_plus1 does not guarantee that the least significant bit of slice_sub_pic_id_plus1 in the last byte of the slice header containing slice_sub_pic_id_plus1 is non-zero. Therefore, 0x00 can occur in that byte, which may trigger emulation prevention depending on the values ​​of the next two syntax bytes.

[0130] In another embodiment, one 1-bit syntax element follows slic_sub_pic_id_plus1, as in the example of Figure 23, which would solve the described problem.

[0131] Obviously, the offered solution comes with overhead, e.g., one additional bit and the x_plus1 syntax element. Moreover, the described problem only occurs if:

[0132] · For emulation prevention that occurs between slice_pic_order_cnt_lsb and slice_sub_pic_id: When the combined length of both syntax elements is greater than 24 bits. - To prevent emulation that occurs with slice_sub_pic_id_plus1 and the following syntax elements: slice_sub_pic_id_plus1 spans multiple bytes.

[0133] Since the described problem is likely to occur infrequently, the described changes could be conditioned on some gating flags in one of the parameter sets of the picture header. This is shown for example in Figure 24, where Val (offset to subpic_id) is determined from the combination of slice_pic_order_cnt_lsb_plus1 and slice_subpic_id_plusVal coded length as follows:

[0134] Val = codedLength(slice_pic_order_cnt_lsb_plus1 + slice_subpic_id_plusVal) < 31 ? 1 : 4.

[0135] This can be read as: if the coded length of slice_pic_order_cnt_lsb_plus1 and slice_subpic_id_plusVal together is less than 31 bits, then Val is set to the value of "1", otherwise it is set to the value of "4".

[0136] The following is about the CABAC zero word (6).

[0137] cabac_zero_words can be inserted at the end of each slice. They can be placed in any slice of an image. The current video specification describes a picture-level constraint that ensures that cabac_zero_words are inserted if the bin-to-bit ratio is too high.

[0138] When subpictures are extracted and shown to conform to a particular profile and level, it MUST be a requirement that the following conditions that apply to AUs also apply to each subpicture individually:

[0139] The variable RawMinCuBits is derived as follows: RawMinCuBits = MinCbSizeY*MinCbSizeY* (BitDepth + 2 * BitDepth / ( SubWidthC * SubHeightC ))

[0140] The value of BinCountsInNalUnits must be less than or equal to (32÷3)*NumBytesInVclNalUnits +(RawMinCuBits * PicSizeInMinCbsY)÷32.

[0141] Therefore, the bitstream constraint or indication indicates that each sub-picture individually satisfies the above conditions.

[0142] As a first embodiment, a flag can be signaled as a constraint flag indicating that the sub-picture satisfies the mentioned constraint, as shown for example in Figure 25.

[0143] In another embodiment, enforcement of such constraints is based on one or both of the following conditions: 1) The subpic_treatment_as_pic_flag of the subpicture is equal to 1 2) The bitstream includes a subpicture level conformance indication, for example, via an SEI message (existing subpic_level_info SEI message).

[0144] In other words, in encoding multiple images of a video into a data stream, each of the multiple images is divided into a predetermined number of sub-pictures, each sub-picture including a boundary for boundary extension for motion compensation, and each sub-picture may be coded using context-adaptive binary arithmetic coding.

[0145] In this case, the data stream provides one or more slices for each subpicture of zero words or at least one subpicture at the end of a VCL NAL unit to avoid any of the subpictures exceeding a predetermined bin-to-bit ratio.

[0146] As can be seen in Figure 26, the encoder can generate the bitstream 2605 in the following manner. Portions of two access units 2620 or 2630 are exemplarily shown, each of which has an image coded therein that is subdivided into sub-pictures, as exemplarily shown for the images in Figure 19. However, Figure 26 concentrates on those portions of the bitstream that relate to exactly one sub-picture of each image. In the case of AU 2621, the coding exemplarily results in only one slice 2621, in which the sub-picture of the image of that AU 2610 is coded; in the case of AU 260, three slices, namely 2631, 2632, and 2633, in which the sub-picture of that AU 2630 is coded.

[0147] Encoders use binary arithmetic coding, such as CABAC. That is, they use syntax elements to encode generated syntax elements that describe the content of a video or image. Those that don't already have a binary value are binarized into a string of bins. Thus, an image is encoded into a syntax element, which is then sequentially encoded into a series of bins, which are arithmetically coded into a sequential data stream, resulting in portions of the bitstream, each of which has a specific portion of the video encoded therein, such as a subpicture. This portion consumes a specific number of bits, which are generated by arithmetically coding a specific number for each sequential bin, resulting in a bin-to-bit ratio. The decoder does the reverse: the portion associated with one subpicture is arithmetically decoded to generate a sequence of bins, which are then binarized to generate syntax elements from which the sequential decoder can reconstruct the subpicture. Thus, the "bin" in "bin-to-bitratio" refers to the number of bits or digits of the binarized symbol being coded, and the "bit" refers to the number of bits written / read in the CABAC bitstream.

[0148] To avoid exceeding a predetermined bin-to-bit ratio, the encoder checks, for each subpicture, the ratio of bins to bits in the CABAC-coded portion of the slice or slices into which the respective subpicture is coded. If the ratio is too high, the encoder supplies as many CABAC-coded zero words as possible to the CABAC encoding engine at one or more ends of the slice or slices into which the respective subpicture is coded, so that the ratio is no longer exceeded. These CABAC zero words are shown as zeros, or "0," in Figure 26. They can be added to the end of the subpicture 2621, i.e., the end of the last slice, or they can be added in a distributed manner to one or more ends of the subpicture's slices 2631, 2632, and 2633. That is, while Figure 26 shows a distributed variant of Figure 26, zero words could alternatively be added entirely to the end of the last slice 2613.

[0149] When the data stream 2605 reaches the decoder 2640, the decoder decodes the data stream, or at least the portion thereof related to the subpicture, and analyzes the corresponding data portions in AUs 2620 and 2630. In doing so, the decoder 2640 can determine the quantity or number of zero words coded at the end of one or more slices of a particular subpicture by determining the number of zero words necessary to avoid exceeding the bin-to-bit ratio, and can discard the CABAC zero words. However, it would also be possible for the decoder 2640 to analyze the slices in such a way that it can distinguish the zero words from other syntax elements by other means and then discard the zero words. That is, the decoder may or may not check the bin-to-bit ratio of one or more slices of a particular subpicture to continue CABAC decoding the number of zero words necessary to achieve a ratio within a certain predetermined range. Otherwise, the decoder can syntactically distinguish the CABAC-decoded zero words from other syntax elements.

[0150] If the ratio condition is not met, the decoder 2640 may enter a particular error mode and certain error handling may be triggered by the decoder.

[0151] Finally, it should be noted that the number of sub-pictures can be any number, in other words, at least one can in particular mean two or more sub-pictures.

[0152] Each subpicture can be evaluated for independence properties, e.g., subpic_treatment_as_pic_flag, and for corresponding level of conformance indications, e.g., Supplemental Enhancement Information, SEI, messages. Then, depending on the evaluation, the subpicture can be provided with a zero word, and only according to the evaluation.

[0153] The data stream may be provided with zero words at the end of one or more slices of each subpicture of at least one subpicture, and the number of coded bins for each subpicture in the data stream using context-adaptive arithmetic coding is less than or equal to a number determined using a product between a predetermined coefficient and the byte length of one or more VCL NAL units of the data stream associated with the respective subpicture.

[0154] The data stream may also be provided with zero words at the end of one or more slices of each subpicture of at least one subpicture, and the number of coded bins for each subpicture in the data stream using context-adaptive arithmetic coding is less than or equal to a number determined using the sum of a first product between a first predetermined coefficient and the byte length of one or more VCL NAL units of the data stream associated with the respective subpicture, and a second product between a second predetermined coefficient and the minimum number of bits per coding block and the number of coding blocks that make up the respective subpicture.

[0155] While some aspects are described in terms of apparatus, it will be apparent that these aspects also represent descriptions of corresponding methods, where blocks or devices correspond to method steps or features of method steps. Similarly, aspects described in terms of method steps also represent descriptions of corresponding blocks or items or features of corresponding apparatus. Some or all method steps may be performed by (or using) a hardware apparatus, such as, for example, a microprocessor, a programmable computer, or electronic circuitry. In some embodiments, one or more of the most significant method steps may be performed by such an apparatus.

[0156] The inventive data stream may be stored on a digital storage medium or transmitted over a transmission medium, such as a wireless or wired transmission medium, such as the Internet.

[0157] Depending on specific implementation requirements, embodiments of the invention may be implemented in hardware or software. Implementations may be performed using a digital storage medium, such as a floppy disk, DVD, Blu-ray, CD, ROM, PROM, EPROM, EEPROM, or flash memory, having electronically readable control signals stored thereon that cooperate (or are capable of cooperating) with a programmable computer system so that the respective methods are performed. Thus, the digital storage medium may be computer-readable.

[0158] Some embodiments according to the invention include a data carrier having electronically readable control signals that can cooperate with a programmable computer system to perform one of the methods described herein.

[0159] Generally, embodiments of the present invention may be implemented as a computer program product having program code operable to perform one of the methods when the computer program product runs on a computer, which may for example be stored on a machine-readable carrier.

[0160] Other embodiments comprise the computer program for performing one of the methods described herein, stored on a machine readable carrier.

[0161] In other words, an embodiment of the inventive methods is, therefore, a computer program having a program code for performing one of the methods described herein, when the computer program runs on a computer.

[0162] A further embodiment of the inventive method is therefore a data carrier (or digital storage medium or computer readable medium) comprising, recorded thereon, the computer program for performing one of the methods described herein. The data carrier, digital storage medium or recorded medium is typically tangible and / or non-transitory.

[0163] A further embodiment of the inventive method is, therefore, a data stream or sequence of signals representing the computer program for performing one of the methods described herein. The data stream or sequence of signals can be adapted to be transferred via a data communication connection, for example via the Internet.

[0164] A further embodiment comprises a processing means, such as for example a computer, or a programmable logic device, configured to or adapted to perform one of the methods described herein.

[0165] A further embodiment comprises a computer having installed thereon the computer program for performing one of the methods described herein.

[0166] Further embodiments according to the invention include an apparatus or system configured to transfer (e.g., electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver may be, for example, a computer, a mobile device, a memory device, etc. The apparatus or system may, for example, include a file server for transferring the computer program to the receiver.

[0167] In some embodiments, a programmable logic device (e.g., a field programmable gate array) may be used to perform some or all of the functions of the methods described herein. In some embodiments, a field programmable gate array may cooperate with a microprocessor to perform one of the methods described herein. In general, the methods are preferably performed by any hardware apparatus.

[0168] The devices described herein may be implemented using a hardware device, or using a computer, or using a combination of a hardware device and a computer.

[0169] The apparatus described herein, or any components of the apparatus described herein, may be implemented at least partially in hardware and / or software.

[0170] The methods described herein may be performed using a hardware apparatus, or using a computer, or using a combination of a hardware apparatus and a computer.

[0171] The methods described herein, or any components of the apparatus described herein, may be implemented at least in part in hardware and / or software.

[0172] The above-described embodiments are merely illustrative of the principles of the present invention. It is understood that modifications and variations in arrangement and detail will be apparent to those skilled in the art. It is therefore the intention to be limited only by the scope of the appended claims and not by the specific details presented by way of description and illustration of the embodiments herein.

Claims

1. Divide a particular image into multiple sub-pictures, determining bin and bit constraints such that the number of coded bins for the subpicture is less than or equal to (32 ÷ 3) × NumBytesInVclNalUnits + (RawMinCuBits × PicSizeInMinCbsY) ÷ 32; encoding the particular image into a bitstream using context-adaptive binary arithmetic coding (CABAC), the encoding including inserting one or more zero words into the bitstream corresponding to a subpicture among the plurality of subpictures such that the bin and bit constraints are satisfied; providing a display that treats the sub-picture as an image into which one or more zero words have been inserted by the CABAC encoding; Data encoding method.

2. the indication includes sps_subpic_treated_as_pic_flag set to 1, or an inference that sps_subpic_treated_as_pic_flag is set to 1 if the sps_subpic_treated_as_pic_flag is omitted from the bitstream. The method of claim 1.

3. the display treating the sub-picture as an image includes a display applying motion-compensated boundary expansion to the sub-picture. The method of claim 1.

4. the zero words are inserted at the end of one or more slices of the sub-picture; The method of claim 1.

5. the plurality of images includes at least two layers; The method further includes dividing the image of at least one layer into a predetermined, layer-specific number of sub-pictures; one or more of the images or sub-pictures of one layer correspond to an image or sub-picture of one or more other layers; At least one of the sub-pictures includes a boundary for boundary extension for motion compensation; the method further comprising encoding an indication that at least one of the boundaries of corresponding sub-pictures or corresponding images of different layers are aligned with one another; The method of claim 1.

6. receiving a bitstream including a plurality of images, wherein a particular image among the plurality of images is divided into a plurality of sub-pictures; Decomposing or estimating, based on the bitstream, a representation of a subpicture corresponding to one of a plurality of subpictures as a picture in which one or more zero words are inserted into the bitstream by CABAC (context-adaptive binary arithmetic coding); determining bin and bit constraints such that the number of coded bins for the subpicture is less than or equal to (32 ÷ 3) × NumBytesInVclNalUnits + (RawMinCuBits × PicSizeInMinCbsY) ÷ 32; decoding from the bitstream to the particular picture by CABAC (context-adaptive binary arithmetic coding), the decoding including parsing one or more zero words from the bitstream corresponding to the subpicture such that the bin and bit constraints are satisfied; Data Combining Methods.

7. the indication includes sps_subpic_treated_as_pic_flag set to 1; The method of claim 6.

8. the display treating the sub-picture as an image includes a display applying motion-compensated boundary expansion to the sub-picture. The method of claim 6.

9. the zero words are parsed from the end of one or more slices of the sub-picture; The method of claim 6.

10. the plurality of images includes at least two layers; The method further includes dividing the image of at least one layer into a predetermined, layer-specific number of sub-pictures; one or more of the images or sub-pictures of one layer correspond to an image or sub-picture of one or more other layers; At least one of the sub-pictures includes a boundary for boundary extension for motion compensation; the method further comprising encoding an indication that at least one of the boundaries of corresponding sub-pictures or corresponding images of different layers are aligned with one another; The method of claim 6.

11. When executed, causes at least one processor to perform the method of any one of claims 1 to 10. A non-transitory computer-readable storage medium.

12. 11. A method for implementing a method according to claim 1, comprising: Encoding device.

Citation Information

Patent Citations

  • Parallel processing for video coding

    JP2016514427A