Video coding related to sub-images
By segmenting video images into sub-images and using boundary expansion and pseudo-data processing, the problem of low video coding efficiency in existing technologies is solved, achieving efficient video downloading and decoding, adapting to changes in user field of view, and meeting the latency requirements of different user terminals.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
- Filing Date
- 2020-12-18
- Publication Date
- 2026-05-05
AI Technical Summary
Existing video coding technologies suffer from inefficiency and resource waste when jointly decoding multiple encoded video bitstreams, especially in 360-degree video playback in multi-party conferencing and VR applications. In particular, when the user's field of view changes, it is necessary to download and decode low-resolution tiles in addition to high-resolution tiles.
By employing sub-image coding technology, video images are segmented into a predetermined number of sub-images. Through boundary extension and pseudo-data processing, the alignment of sub-image boundaries and the proper use of pseudo-data are ensured, thereby achieving effective image layering and decoding.
It improves video encoding efficiency, reduces resource consumption, adapts to changes in user field of view, achieves more efficient video downloading and decoding, and meets the latency requirements of different user terminals.
Smart Images

Figure CN121985137A_ABST
Abstract
Description
Technical Field
[0001] This application is a divisional application of application number 202080088665.6, entitled "Video Coding Related to Sub-images". This application relates to the concept of video coding, and particularly to sub-images. Background Technology
[0002] There are certain video-based applications where multiple encoded video bitstreams or data streams need to be jointly decoded (e.g., merged into a joint bitstream and fed into a single decoder), such as multi-party conferencing where encoded video streams from multiple participants are processed on a single endpoint, or tile-based streaming for 360-degree tiled video playback in VR (virtual reality) applications.
[0003] For High Efficiency Video Coding (HEVC), a set of motion-constrained tiles is defined, where motion vectors are constrained to a non-reference set of tiles (or tiles) that is different from the current set of tiles (or tiles). Therefore, the set of tiles (or tiles) in question can be extracted from or incorporated into another bitstream without affecting the decoding result. For example, regardless of whether the set of tiles (or tiles) in question is decoded individually or as part of a bitstream with more tile sets (or tiles) for each image, the decoded samples will be exact matches.
[0004] In the latter, an example is shown in 360-degree video, where such techniques are useful. The video is spatially segmented, and each spatial segment is defined as follows: Figure 1 Multiple representations of the varying spatial resolutions illustrated in the figure are provided to the streaming client. The figure shows a cube of a 360-degree video projection divided into 6×4 spatial segments at two different resolutions. For simplicity, these individual, decodeable spatial segments are referred to as tiles in this description.
[0005] When using, such as Figure 2 The state-of-the-art head-mounted display shown on the left side of the image typically displays only a subset of the tiles that make up the entire 360-degree video when viewed through the solid blue viewport borders representing a 90×90-degree field of view. Figure 2 The corresponding tiles highlighted in green with shading were downloaded at the highest resolution.
[0006] However, the client application will also have to download and decode representations of other tiles outside the current viewport. Figure 2(Used in red with shading) to handle sudden changes in user orientation. In such applications, the client therefore downloads tiles covering its current viewport at the highest resolution and tiles outside its current viewport at a relatively lower resolution, while the tile resolution selection is consistently adapted to the user's orientation. After downloading on the client side, merging the downloaded tiles into a single bitstream for processing with a single decoder is a means to address the constraints of typical mobile devices with limited computing and power resources. Figure 3 The illustration shows a possible tile arrangement in the joint bitstream for the example above. The merging operation used to generate the joint bitstream must be performed through compression field processing, for example, avoiding pixel field processing through transcoding.
[0007] Because the HEVC bitstream is encoded in accordance with some constraints that mainly involve inter-encoding tools (such as the constrained motion vectors described above), a merging process can be performed.
[0008] The emerging VVC codec offers a more efficient alternative for achieving the same goal: sub-images. With sub-images, regions smaller than the full image can be treated as similar to images in such a sense that their boundaries are treated as if they were images. For example, when applying boundary extensions for motion compensation, such as when motion vectors point outside a region, the last sample of that region (the boundary crossed) is repeated to generate samples in the reference block for prediction, just as it would be done at the image boundary. Thus, motion vectors at the encoder are not constrained by the corresponding efficiency loss as in HEVC MCTS.
[0009] The emerging VVC coding standard also envisions providing scalable coding tools for multi-layer support in its main profile. Therefore, by encoding the entire low-resolution content with less frequent RAP, further efficient configurations for the above application scenarios can be achieved. However, this requires the use of a layered coding structure that includes low-resolution content always in the base layer and some high-resolution content in enhancement layers. Layered coding structures in... Figure 1 As shown in the image.
[0010] However, there may still be some use cases where a single region of the bitstream can be extracted. For example, a user with higher end-to-end latency might download the entire 360-degree video (all tiles) at low resolution, while a user with lower end-to-end latency might download fewer tiles with lower resolution content, such as the same number as the downloaded high-resolution tiles.
[0011] Therefore, the extraction of layered sub-images should be handled appropriately by the video coding standard. This requires additional signaling on the decoder or extractor side to ensure appropriate knowledge. Summary of the Invention
[0012] Therefore, the object of the present invention is to provide such additional signaling, which improves upon currently available mechanisms.
[0013] This objective is achieved through the subject matter of the independent claims of this application.
[0014] According to a first aspect of this application, a data stream is decoded into a plurality of images of a video, wherein the data stream comprises a plurality of images in at least two layers. The images in at least one layer are segmented into a predetermined number of sub-images specific to that layer, one or more sub-images of one layer corresponding to sub-images or one image in one or more other layers, and at least one of the sub-images includes a boundary extension for motion compensation. In this process, an indication that at least one boundary of corresponding images or corresponding sub-images in different layers is aligned with each other is interpreted, in other words, parsed or decoded from the data stream.
[0015] According to a second aspect of this application, a constant bit rate data stream having multiple images encoded therein is processed, wherein each of the multiple images is segmented into a predetermined number of sub-images, each sub-image including a boundary extension for motion compensation. A sub-image data stream is generated from the constant bit rate data stream associated with at least one sub-image by retaining pseudo-data included in the data stream of at least one sub-image and removing pseudo-data included in the data stream of another sub-image not associated with it from the extracted data stream.
[0016] According to a third aspect of the invention, sub-image extraction is performed on a data stream having multiple images of a video encoded therein in multiple layers, wherein each of the multiple images in all layers is segmented into a predetermined number of sub-images, each sub-image including a boundary extension for motion compensation. Sub-image extraction is performed on the data stream to obtain an extracted data stream associated with one or more sub-images of interest by discarding NAL units of the data stream that do not correspond to one or more sub-images and rewriting the parameter set and / or image header.
[0017] According to a fourth aspect of the invention, a plurality of images of a video are encoded into a data stream, wherein each of the plurality of images is divided into a predetermined number of sub-images, each sub-image including a boundary for motion compensation. The plurality of images are encoded in units of slices, and a first syntax element and a sub-image identifier parameter are written into the slice header of the slice, wherein in the slice header of at least one slice, the first syntax element and the sub-image identifier parameter are written into the data stream in an incrementing manner and separated by bits set to one.
[0018] According to a fifth aspect of the invention, a plurality of images of a video are encoded into a data stream, wherein each of the plurality of images is divided into a predetermined number of sub-images, each sub-image including a boundary extension for motion compensation. The plurality of images are encoded in units of slices, and a first syntax element and a sub-image identifier parameter are written into the slice header of the slice. The first syntax element and the sub-image identifier parameter are written into the data stream such that the first syntax element precedes the sub-image identifier parameter, the first syntax element is written into the data stream by incrementing by one, and the sub-image identifier parameter is written into the data stream by incrementing by a certain value and followed by a bit set to one, such that the first syntax element is written with a first bit length and the sub-image identifier parameter is written with a second bit length.
[0019] The value is determined by checking whether the sum of the first length and the second length is less than 31 bits. If the sum of the first length and the second length is less than 31 bits, the value is set to 1. If the sum of the first length and the second length is not less than 31 bits, the value is set to 4.
[0020] According to a sixth aspect of the invention, a plurality of images of a video are encoded into a data stream, wherein each of the plurality of images is segmented into a predetermined number of sub-images, each sub-image including a boundary for boundary extension for motion compensation. Context-adaptive binary arithmetic coding is used to encode each sub-image, wherein the video encoder is configured to provide zero-words for at least one sub-image at the end of one or more slices of each sub-image to the data stream in order to prevent any sub-image from exceeding a predetermined bin-to-bit ratio.
[0021] All the aspects mentioned above are not limited to encoding or decoding; the corresponding other aspect in encoding and decoding is based on the same principle. Regarding the foregoing aspects of this application, it should be noted that the foregoing aspects of this application can be combined to enable the simultaneous implementation of more than one of the foregoing aspects, such as all of them, in a video codec. Attached Figure Description
[0022] The preferred embodiments of this application are described below with reference to the accompanying drawings, in which: Figure 1 The video is shown in a cubic projection at two resolutions and divided into 6×4 tiles, displaying a 360-degree video. Figure 2 This demonstrates the user viewport and tile selection for 360-degree video streaming; Figure 3The resulting tile arrangement (packing) in the joint bitstream after the merge operation is shown. Figure 4 A bitstream based on scalable sub-images is shown; Figures 5 to 11 An example syntax element is shown; Figure 12 The illustration shows an image segmented into sub-images at different layers; Figure 13 The diagram shows the boundaries aligned between lower and higher layers, as well as the boundaries of higher layers that do not have a corresponding relationship in the lower layers. Figure 14 An example syntax element is shown; Figure 15 An exemplary region of interest (RoI) is shown in a low-resolution layer and not provided in a high-resolution layer. Figure 16 An exemplary layer and sub-image configuration is shown; Figure 17 and 18 An example syntax element is shown; Figure 19 and 20 The diagram illustrates a data stream with a constant bit rate, where sub-image extraction is performed and the sub-images are filled with pseudo-data. Figures 21 to 25 Exemplary syntax elements are shown; and Figure 26 This demonstrates encoding using the cabac zero word. Detailed Implementation
[0024] In the following, additional embodiments and aspects of the invention will be described, which may be used alone or in combination with any of the details, functionalities and features described herein.
[0025] The first embodiment relates to layers and sub-images, and in particular to sub-image boundary alignment (2.1).
[0026] The signal indicates that sub-image boundaries are aligned across layers; for example, each layer contains the same number of sub-images, and the boundaries of all sub-images are sample-to-sample juxtapositions. This is, for example, in... Figure 4As illustrated, this type of signaling can be implemented, for example, via `sps_subpic_treated_as_pic_flag`. If this flag `sps_subpic_treated_as_pic_flag[i]` is set, for example having a value equal to 1, it specifies that the i-th sub-image in each encoded image of the CLVS encoded layer-by-layer video sequence is considered an image during the decoding process of the exclusion loop filtering operation. If `sps_subpic_treated_as_pic_flag[i]` is not set, for example by having a value equal to 0, it specifies that the i-th sub-image of each encoded image in the CLVS is not considered an image during the decoding process of the exclusion loop filtering operation. If the flag does not exist, it can be assumed that it is to be set; for example, the value of `sps_subpic_treated_as_pic_flag[i]` is inferred to be equal to 1.
[0027] The signaling is useful for extraction because the extractor / receiver can look for sub-image IDs present in the slice header to discard all NAL units that do not belong to the sub-image of interest. When sub-images are aligned, the sub-image IDs are mapped one-to-one between layers. This means that for all samples in enhancement layer sub-image A, identified by sub-image ID An, juxtaposed samples in base layer sub-image B belong to a single sub-image identified by sub-image ID Bn.
[0028] Note that if the alignment as described does not exist, for each sub-image at the reference layer, there may be more than one sub-image in the reference layer, or even worse, from the perspective of extraction or parallelization from the bitstream: partial overlap of sub-images in the layer.
[0029] In certain implementations, sub-image alignment can be indicated via constraint flags in the `general_constraint_info()` structure within `profile_tier_level()` in the VPS, as this is layer-specific signaling. This is, for example, in... Figure 5 As shown in the diagram.
[0030] The syntax allows indicating that layers within an Output Layer Set (OLS) have aligned sub-images. Another option is to indicate alignment for all layers in the bitstream, such as alignment for all OLS, to make signaling OLS independent.
[0031] In other words, in this embodiment, the data stream is decoded into multiple images of the video, and the data stream includes multiple images across at least two layers. In this example, one layer is the base layer ( Figure 4 The “low resolution” in the text, and another layer is an enhancement layer ( Figure 4(High resolution in the context). The images from the two layers are segmented into sub-images, sometimes called tiles, and at least in... Figure 4 In the example, each sub-image in the base layer has a corresponding sub-image in the enhancement layer. Figure 4 The boundaries of the sub-images shown are used for boundary extension of motion compensation. During decoding, there is an indication that at least one boundary of the corresponding sub-images in different layers is aligned with each other. For brevity, note that throughout the description, sub-images of different layers are understood to correspond to each other due to their common location, i.e., they are located at the same location within the image (which they are part of). Sub-images are encoded independently of any other sub-images within the same image (which the sub-image is part of). However, a sub-image of a certain layer is also a term describing the mutually corresponding sub-images of the image of that layer, and these mutually corresponding sub-images are encoded using motion compensation prediction and temporal prediction in a manner that makes the encoding dependency of regions outside these mutually corresponding sub-images unnecessary, and, for example, boundary extension is used for the motion compensation reference portion pointed to by the motion vectors of blocks within these mutually corresponding sub-images, i.e., portions of these reference portions extending beyond the sub-image boundaries. Subdividing such layered images into sub-images is done in a manner that makes overlapping sub-images correspond to others and form an independently coded sub-video. Note that sub-images can also function in relation to inter-layer prediction. If interlayer prediction is available for encoding / decoding, an image of a layer can be predicted from another image of a lower layer, with both images having the same time instant. A region of an image, referred to as a layer, is filled by boundary expansion (i.e., by not using content outside the juxtaposed sub-image from the reference image), the region extending beyond the boundary of the block and the sub-image in the reference image of the other juxtaposed layer; that is, the region is covered by it.
[0032] If an image is segmented into only one sub-image, that is, the image is treated as a sub-image, then the images correspond to each other across layers, and the indications indicate the boundary alignment of the images, and they are used for boundary extension for motion compensation.
[0033] Each of the multiple images in at least two layers can also be segmented into one or more sub-images, and the indication can be interpreted to the extent that all boundaries of the corresponding sub-images in different layers are aligned with each other.
[0034] The following relates to the sub-image ID correspondence (2.2).
[0035] In addition to sub-image boundary alignment, it is also necessary to facilitate the detection of corresponding sub-image IDs across layers, provided that the sub-images are aligned.
[0036] In the first option, the sub-image IDs in the bitstream are distinct; for example, each sub-image ID can be used only in one layer of the OLS. Thus, each layer-sub-image is uniquely identifiable by its own ID. This is, for example, in... Figure 6 As shown in the diagram.
[0037] Note that this can also be applied to IDs that are unique across the entire bitstream, for example, to any OLS within the bitstream, regardless of whether two layers belong to the same OLS. Therefore, in one embodiment, any subpicture ID value will be unique within the bitstream (e.g., unique_subpicture_ids_in_bitstream_flag).
[0038] In the second option, the sub-images in the layer carry the same ID value as their corresponding aligned sub-images in another layer (either in the OLS or the entire bitstream). This makes it very simple to identify the corresponding sub-images simply by matching their ID values. Furthermore, the extraction process is simplified because only one ID is required per sub-image of interest at the extractor. This is, for example, in… Figure 7 As shown in the diagram. Instead of aligned_subpictures_ids_in_ols_unique_flag, you can also use the SubpicIdVal constraint.
[0039] The following pertains to the sub-image signaling sequence (2.3).
[0040] When the signaling order of subpicture IDs, positions, and sizes is also constrained, the processing of subpicture ID correspondences can be greatly simplified. This means that when indicating the corresponding aligned subpictures and unique subpicture IDs (e.g., the flags subpictures_in_ols_aligned_flag and aligned_subpictures_ids_in_ols_unique_flag are set to equal to 1), the definition of subpictures (including position, width, height, boundary treatment, and other attributes) should span all SPS alignments on the OLS.
[0041] For example, considering possible resampling factors, the syntax element sps_num_subpics_minus1 should have the same value in the layer and all its reference layers, as should subpic_ctu_top_left_x[i], subpic_ctu_top_left_y[i], subpic_width_minus1[i], and subpic_height_minus1[i].
[0042] In this embodiment, the order of sub-images in the SPS of all layers (see the green markers for syntax elements in the SPS) is constrained to be the same. This allows for simple matching and verification of sub-images and their positions across layers. This is, for example, in... Figure 8 As shown in the diagram.
[0043] Subpicture ID signaling can be implemented using the corresponding rewrite functionality in SPS, PPS, or PH (as shown below), and with `aligned_subpictures_ids_in_ols_flag` equal to 1, the subpicture ID value should be identical across all layers. With `unique_subpictures_ids_in_ols_flag` equal to 1, the subpicture ID value should not appear in more than a single layer. This is, for example, in... Figures 9 to 11 As shown in the diagram.
[0044] In other words, in this embodiment, the interpretation of relative sub-image position and sub-image size is the same for corresponding sub-images in all layers. For example, the indices representing the signaling order in the data stream for the syntax elements indicating the position and size of corresponding sub-images are equal. The indication may include a flag in each layer of the data stream that indicates sub-image alignment for that layer, including the layer in the data stream where the flag is present, and one or more higher layers, such as those layers that use any image in the image of the layer with that flag as a reference image.
[0045] Figure 12This is illustrated. In it, images of layers of a hierarchical data stream are shown. Image 1211 of layer 1 is exemplary not segmented into sub-images, or in other words, only segmented into one sub-image, thus having a (sub)image L1(1). Image 1212 of layer 2 is segmented into four sub-images L2(1) to L2(4). Images of higher layers (here, 3 to 4) are segmented in the same manner. Note that inter-layer prediction is used to encode the images; that is, inter-layer prediction can be used for encoding / decoding. An image of a layer can be predicted based on another image of a lower layer, with both images at the same time. Even here, sub-image boundaries trigger boundary extension in inter-layer prediction, i.e., by filling a region of an image called a layer with content outside the juxtaposed sub-image from the reference image, the region extending beyond the boundary of the sub-image in the reference image of the other layer juxtaposed with the block. Furthermore, in each layer subdivided into sub-images, it may be required that the images have a constant size, meaning that the images of each layer do not change in size, i.e., RPR is not used. For a single sub-image layer such as layer L1, varying image sizes and the application of RPR are permissible. The indicated flags can then indicate that sub-images of layers 2 and above are aligned, as discussed above. That is, images 1213, 1214 of higher layers (layers 3 and above) have all the same sub-image divisions as layer 2. Sub-images of layers 3 and above also have the same sub-image IDs as layer 3. Therefore, layers above layer 2 also have sub-images LX(1) to LX(4), where X is the layer number. Figure 12 This is illustrated, exemplarily for some higher layers 3 and 4, but there could also be only one or more higher layers, and then the indication would also be associated with them. Thus, the sub-image of each layer above the layer to which the indication is made (here, layer 2) has the same index as the corresponding sub-image in layer 2, such as 1, 2, 3, or 4, and has the same size and juxtaposition boundary as the corresponding sub-image in layer 2. The corresponding sub-images are those juxtaposed with each other within the image. Similarly, such indications could also exist for each of the other layers, and they could also be set for higher layers (here, such as layers 3 and 4 with the same sub-image subdivisions).
[0046] When each of the multiple images in at least two layers is segmented into one or more sub-images, the indication indicates for each of the one or more layers that the image above the corresponding layer is segmented into sub-images such that a predetermined layer-specific number of sub-images is equal for the corresponding layer and the one or more higher layers, and the boundaries of the predetermined layer-specific number of sub-images spatially coincide between the corresponding layer and the one or more higher layers.
[0047] In particular, if this is applied Figure 12 The image then becomes clear that if the sub-images of layers 2 and above are aligned, then layers 3 and above also have corresponding aligned sub-images.
[0048] One or more higher layers can use images from the corresponding layers for prediction.
[0049] Furthermore, the ID of the sub-image is the same across the corresponding layer and one or more higher layers.
[0050] Figure 12 Also shown is a sub-image extractor 1210, which is used to extract the extracted bitstream from the bitstream encoded by the layers just discussed, the extracted bitstream being specific to a certain set of one or more sub-images of layers 2-4. It extracts from the complete bitstream those portions that are related to the region of interest of the layered image, exemplarily the sub-image of the upper right corner of the image. Thus, extractor 1210 extracts those portions of the entire bitstream that are related to the associated RoI, that is, those portions that are related to the entire image of any layer for which the indication described above does not indicate sub-image adjustment (i.e., for which the flags mentioned above are not set), and those portions that are related to the sub-images of images within mutually aligned layers (i.e., for which the flags are set and for any higher layers whose sub-images cover the RoI, exemplarily L2(2), L3(2), and L4(2)). In other words, extractor 1210 will extract any single sub-image layer (such as...) in the output layer set and used as a reference to the sub-image layers (here 2-4) Figure 12 In the case of L1, this is included in the extracted bitstream, where the HLS inter-layer prediction parameters are adjusted accordingly (i.e., based on the scaling window) so that the higher-layer images of layers 2-4, previously subdivided by sub-images, are smaller in the extracted bitstream because they are cropped into sub-images only within the RoI. Extraction is straightforward because the indications specify the alignment of the sub-images with each other. Another way to utilize the indications of sub-image alignment is to create a decoder that is easier to organize for decoding.
[0051] The following pertains to subsets of sub-image boundaries in lower layers (2.4).
[0052] Other use cases may exist, such as fully parallel coding of intra-layer regions to accelerate coding, where the number of sub-images is greater for higher layers with higher resolution (spatial scalability). In such cases, alignment would also be desirable. A further embodiment is that if the number of sub-images in a higher layer is greater than the number of sub-images in a lower (reference) layer, then all sub-image boundaries in the lower layer have corresponding counterparts (juxtaposed boundaries) at the higher layer. Figure 13The diagram illustrates the boundaries, which are the alignment boundaries between lower and higher layers, as well as the boundaries of higher layers that do not have a corresponding relationship in the lower layers.
[0053] Figure 14 An example syntax for signaling this property is shown in the figure.
[0054] The following section discusses the impact of motion compensation prediction on sub-image boundary layering (2.5).
[0055] The presence of a sub-image in the bitstream indicates whether it is an independent sub-image only within a layer or also across layers. More specifically, it indicates whether motion compensation predictions performed across layers also take into account sub-image boundaries. An example could be when an RoI is provided in a low-resolution version (lower layer) of content (e.g., a 720p RoI within 1080p content) and not in high-resolution content (e.g., a higher layer with 4k resolution), such as... Figure 15 As shown in the diagram.
[0056] In another embodiment, does the motion compensation prediction performed across layers also consider the fact that sub-image boundaries depend on whether the sub-images are aligned across layers? If so, the boundaries are considered for motion compensation in inter-layer predictions; for example, motion compensation is not allowed to use sample locations outside the boundaries, or to infer sample values for such sample locations from sample values inside the boundaries. Otherwise, the boundaries are ignored for inter-layer motion compensation predictions.
[0057] The following relates to the sub-image reduction reference OLS (2.6).
[0058] For the use of sub-images with layers, a further use case to consider is RoI scalability. Figure 16 The diagram illustrates a potential layer and sub-image configuration for this type of use case. In such cases, only the sub-image or RoI portion from the lower layers is needed for higher layers. This means that when only the RoI is of interest, i.e., when decoding the 720p version in the base layer or enhancement layer in a given example, fewer samples need to be decoded. However, it is necessary to instruct the decoder that only a subset of the bitstream needs to be decoded, and therefore the level of the sub-bitstream associated with the RoI is lower than (fewer decoded samples) the entire base layer and enhancement layer.
[0059] In the example, instead of 1080+4K, only 720+4K decoding is required.
[0060] In one embodiment, the OLS signaling indication does not require an entire layer, but only its subpicks. For each output layer, when the indication is given (reduced_subpic_reference_flag equals 1), a list of relevant subpick IDs for reference is provided (num_sub_pic_ids, subPicIdToDecodeForReference). This is, for example, in... Figure 17 The diagram is shown in the image.
[0061] In another embodiment, additional PTL signaling is provided, which indicates that a PTL will be requested if unnecessary sub-images of layers within the OLS will be removed from the OLS bitstream. Figure 18 The options are shown, with the corresponding syntax added to the VPS.
[0062] The following deals with constant bit rate (CBR) and sub-images (3).
[0063] According to the present invention, a video processing apparatus can be configured to process multiple images of a video decoded from a data stream comprising multiple images, wherein, for example, each of the multiple images in all layers is segmented into a predetermined number of sub-images, each sub-image including a boundary extension for motion compensation, and wherein the video processing apparatus is configured to generate at least one sub-image with a constant bit rate from the data stream by maintaining pseudo-data (e.g., FD_NUT and filler payload supplemental enhancement information SEI messages) until another sub-image appears in the data stream, wherein the pseudo-data is included in the data stream of the sub-image immediately following the sub-image, or is included in the data stream of the sub-image that is not adjacent to the sub-image but has an indication of the sub-image (e.g., a sub-image identifier). Note that in the case where the data stream is a hierarchical data stream, the images thus encoded can be associated with a layer, and per-sub-image bit rate control can be applied to each layer.
[0064] According to another aspect, a video encoder can be configured to encode multiple images of a video into a data stream comprising multiple images, wherein each of the multiple images in all layers is segmented into a predetermined number of sub-images, each sub-image including a boundary extension for motion compensation, and wherein the video encoder is configured to encode at least one sub-image with a constant bit rate into a data stream by including pseudo-data (e.g., FD_NUT and padding payload supplemental enhancement information SEI message) for each sub-image immediately following each sub-image or not adjacent to the sub-image but having an indication of the sub-image (e.g., sub-image identifier) into the data stream.
[0065] According to another aspect, a method for processing video may have the following steps: processing multiple images of a video decoded from a data stream comprising multiple images, wherein, for example, each of the multiple images in all layers is segmented into a predetermined number of sub-images, each sub-image including a boundary for boundary extension for motion compensation, and wherein the method includes generating at least one sub-image with a constant bit rate from the data stream by maintaining pseudo-data (e.g., FD_NUT and padding payload supplemental enhancement information SEI messages) until another sub-image appears in the data stream, wherein the pseudo-data is included in the data stream of the sub-image immediately following the sub-image, or is included in the data stream of the sub-image that is not adjacent to the sub-image but has an indication of the sub-image (e.g., a sub-image identifier).
[0066] According to another aspect, a method for encoding video may include the steps of: encoding multiple images of a video into a data stream, wherein the data stream includes multiple images, wherein each image in multiple images across all layers is segmented into a predetermined number of sub-images, each sub-image including a boundary for motion compensation, and wherein the method includes encoding at least one sub-image with a constant bit rate into a data stream by including pseudo-data (e.g., FD_NUT and padding payload supplemental enhancement information SEI message) for each sub-image immediately following or not adjacent to the sub-image but having an indication of the sub-image (e.g., sub-image identifier) into the data stream.
[0067] There are cases where a bitstream containing a sub-image is encoded and the associated HRD syntax element defines a constant bit rate (CBR) bitstream. For example, for at least one scheduling value, cbr_flag is set to 1 to indicate that the bitstream often corresponds to a constant bit rate bitstream by using so-called padding data VCL NAL units or padding payload SEI (non-VCL NAL units) (e.g., FD_NUT and padding payload SEI messages).
[0068] However, it is unclear whether this property still applies when the sub-image bitstream is extracted.
[0069] In one embodiment, the sub-image extraction process is defined in a way that always induces a VBR bitstream. There is no possibility of indicating / ensuring a CBR condition. In this case, the FD_NUT and padding payload SEI messages are simply discarded during the extraction process.
[0070] In another embodiment, the bitstream indicates that the CBR “operation point” of each sub-image is ensured by placing the corresponding FD_NUT and padding payload SEI message immediately after the VCL NAL unit constituting the sub-image. Therefore, during the extraction of the CBR “operation point” of the sub-image, the padding payload SEI message and the corresponding FD_NUT associated with the sub-image VCL NAL unit are maintained while the sub-image bitstream is being extracted during the extraction process, and are also maintained during the extraction process when the extraction process targets another sub-image or a non-CBR “operation point.” Such an indication can be performed, for example, using sli_cbr_constraint_flag. Thus, for example, when sli_cbr_constraint_flag equals 0, all NAL units with nal_unit_type equal to FD_NUT and the SEI NAL unit containing the padding payload SEI message are removed.
[0071] The described process requires maintaining the state of the VCL NAL cell and its associated non-VCL NAL cells because knowing the sub-image ID directly preceding the VCL NAL cell is necessary to determine whether to discard the FD_NUT or the padding payload SEI message. To simplify this process, signaling indicating that they belong to a given sub-image ID is added to the FD_NUT and the padding payload SEI. Alternatively, in another embodiment, an SEI message is added to the bitstream indicating that subsequent SEI messages or NAL cells belong to a given sub-image with a sub-image ID, until another SEI message indicating another sub-image ID is present.
[0072] In other words, a constant bit rate data stream having multiple images encoded therein is processed in a manner in which each of the multiple images is segmented into a predetermined number of sub-images, each sub-image including a boundary for motion compensation, and through the processing, a sub-image data stream associated with at least one sub-image having a constant bit rate is generated from the data stream by retaining the pseudo-data included in the data stream for at least one sub-image and removing the pseudo-data included in the data stream for another sub-image that is not associated with the extracted data stream.
[0073] This can be done Figure 19 and 20 As seen in the example, two images 1910 and 1920 of a video encoded into a processed data stream are shown. The images are respectively segmented into sub-images 1911, 1912, 1913 and 1921, 1922 and 1923, and encoded by some encoder 1930 into corresponding access units 1940 and 1950 of bit stream 1945.
[0074] The subdivision into sub-images is shown only illustratively. For illustrative purposes, the AU and access unit are shown to also include some header information portions HI1 and HI2, respectively, but this is only for illustration, and the image is shown to be encoded into a bitstream 1945 that is segmented into several parts (such as VCL NAL units 1941, 1942, 1943, 1951, 1952, and 1953), such that for each sub-image in images 1940 and 1950, there is one VCL unit, but this is only for illustrative purposes, and a sub-image may be segmented / encoded into more than one such part or VCL NAL unit.
[0075] The encoder creates a constant bit rate data stream by including pseudo-data d (i.e., pseudo-data 1944, 1945, 1946, 1954, 1955, and 1956 for each sub-image) at the end of each data portion representing the sub-image. Figure 19 The dummy data is represented by a cross. In this diagram, dummy data 1944 corresponds to sub-image 1911, 1945 to sub-image 1912, 1946 to sub-image 1913, 1954 to sub-image 1921, 1955 to sub-image 1922, and 1956 to sub-image 1923. It is possible that dummy data is not needed for some sub-images.
[0076] Alternatively, as indicated, the sub-image can also be distributed across more than one NAL unit. In this case, pseudo-data can be inserted at the end of each NAL unit, or at the end of the last NAL unit of the sub-image.
[0077] Then, the constant bit rate data stream 1945 can be processed by the sub-image extractor 1960, which, as an example, extracts information related to sub-image 1911 and the corresponding sub-image 1921, that is, the mutually corresponding sub-images of the video encoded in the bit stream 1945 since juxtaposition. The set of extracted sub-images can be more than one. The extraction causes the extracted bit stream 1955. For the extraction, the NAL units 1941 and 1951 of the extracted sub-images 1911 / 1921 are extracted and taken over into the extracted bit stream 1955, while other VCL NAL units of other sub-images that are not to be extracted are ignored or discarded, and only the corresponding pseudo-data d of the extracted sub-image (2) is taken over into the extracted bit stream 1955, while the pseudo-data of all other sub-images is also discarded. That is, for other VCL NAL units, pseudo-data d is removed. Optionally, the header portions HI3 and HI4 can be modified as described in detail below. They may be associated with, or include, a set of parameters such as PPS and / or SPS and / or VPS and / or APS.
[0078] Therefore, the extracted bitstream 1955 includes access units 1970 and 1980 corresponding to access units 1940 and 1950 of bitstream 1945, and has a constant bit rate. Thus, the bitstream generated by extractor 1960 can then be decoded by decoder 1990 to produce a sub-video, that is, a video comprising only the extracted sub-images 1911 and 1921. Naturally, the extraction can affect more than one sub-image of the video of bitstream 1945.
[0079] As another alternative, regardless of whether one or more sub-images are encoded in one or more NAL units, the pseudo-data can be sorted within each access unit at its end, but in a manner that associates each piece of pseudo-data with the corresponding sub-image encoded into the respective access unit. This is in Figure 20 The diagram uses Named Access Units 2040 and 2050 instead of 1940 and 1950.
[0080] Note that pseudo data can be a NAL cell of a certain NAL cell type, so as to be distinguishable from VCL NAL cells.
[0081] exist Figure 20 This can also be interpreted as an alternative to bitstream 1945 with a constant bit rate, but here the constant bit rate is not performed per sub-image. Instead, it is performed globally. For example... Figure 20 As shown in the diagram, for all sub-images, the pseudo-data can be included together at the end of the AU, which is only shown illustratively. Figure 19 The corresponding data stream portion. In access units 2040 and 2050, the data portions representing sub-images 2044, 2045, 2046, 2054, 2055 and 2056 respectively follow each other directly, and pseudo data is appended to the end of the sub-image at the end of the corresponding access unit.
[0082] The instruction in the bitstream to be processed by the extractor indicates whether to perform sub-image selective pseudo-data removal; that is, when extracting the sub-image bitstream, should all pseudo-data be stripped, or should it be removed as before? Figure 19 The diagram shows pseudo-data that preserves the sub-images of interest (i.e., the ones to be extracted).
[0083] Furthermore, if the data stream is indicated to not have a constant bit rate per sub-image, the extractor 1960 can remove all pseudo data d for all NAL cells, and the resulting extracted sub-image bit stream will not have a constant bit rate.
[0084] Note that the arrangement of pseudo-data in the access cell can be modified relative to the example just discussed, and they can be placed between VCL cells or at the beginning of the access cell.
[0085] It should be noted that the processing described above and below by the processing device or extractor 1960 can also be performed by, for example, a decoder.
[0086] In other words, the data stream can be processed by an extractable set of at least one sub-image from the sub-images of the data stream, using an indication that makes the extracted sub-image bitstream have a constant bit rate. This indication can be processed, for example, from or contained in the Video Parameter Set (VPS) or SEI message, and can be, for example, cbr_flag or sli_cbr_constraint_flag.
[0087] The data stream can also be processed by interpreting an indication that the extractable set of one or more sub-images of the data stream is not encoded at a constant bit rate, such as in the video parameter set VPS or in the SEI message, and if so indicated, all pseudo data is removed.
[0088] As indicated, pseudo-data may include FD_NUT and special NAL units. Additionally or alternatively, pseudo-data may include padding payload supplemental enhancement information SEI messages.
[0089] The data stream can also be processed by interpreting (e.g., decoding) an instruction from which the generation of the extracted data stream is performed, whether the generation should end in an extracted data stream with a constant bit rate (in which case the video processing device performs retention and removal of spurious data) or the generation should not end in an extracted data stream with a constant bit rate (in which case the video processing device performs removal of all spurious data).
[0090] Finally, the pseudo data mentioned above can include, for example, FD_NUT and padding payload supplemental enhancement information SEI messages.
[0091] The following involves the completion of sub-image extraction: parameter set rewriting (4).
[0092] According to another aspect, the video processing device can be configured to perform sub-image extraction on a data stream comprising multiple images of video, wherein each of the multiple images in all layers is segmented into a predetermined number of sub-images, each sub-image including a boundary for motion compensation and an indication in the data stream that allows rewriting of the data stream's image header and / or parameter set when performing sub-image extraction on the data stream.
[0093] According to another aspect, when referring to the preceding aspect, the video processing device can be further configured to interpret additional information in the data stream that provides information for rewriting information about HRD, image size, and / or the subdivision of multiple images into sub-images.
[0094] According to another aspect, a video encoder can be configured to encode multiple images of a video into a data stream, wherein the data stream includes multiple images, wherein each of the multiple images in all layers is segmented into a predetermined number of sub-images, each sub-image including a boundary for boundary extension for motion compensation, and an image header and / or parameter set in the data stream indicating that the data stream is allowed to be rewritten when sub-image extraction is performed on the data stream.
[0095] According to another aspect, when referring back to the previous aspect, the video encoder can be further configured to indicate additional information in the data stream for rewriting information about HRD-related information, image size, and / or information about the subdivision of multiple images into sub-images.
[0096] According to another aspect, a method for processing video may have the following steps: performing sub-image extraction on a data stream comprising multiple images of video, wherein each of the multiple images in all layers is segmented into a predetermined number of sub-images, each sub-image including a boundary for boundary extension for motion compensation, and an indication of an image header and / or parameter set that allows the data stream to be rewritten when sub-image extraction is performed on the data stream.
[0097] According to another aspect, when referring back to the previous aspect, the method for processing video may further include the following steps: interpreting the instructions in the data stream that provide additional information for rewriting information about HRD-related information, image size, and / or multiple images subdivided into sub-images.
[0098] According to another aspect, a method for encoding video may have the following steps: encoding multiple images of a video into a data stream, wherein the data stream includes multiple images, wherein each of the multiple images in all layers is segmented into a predetermined number of sub-images, each sub-image including a boundary for boundary extension for motion compensation, and indicating in the data stream an image header and / or parameter set that allows rewriting of the data stream when performing sub-image extraction on the data stream.
[0099] According to another aspect, when referring to the previous aspect, the method for encoding video may further include the following steps: indicating additional information in the data stream for rewriting information related to HRD, image size, and / or information on the subdivision of multiple images into sub-images.
[0100] On the other hand, it can involve a data stream into which video is encoded, wherein the video comprises multiple images, wherein each of the multiple images in all layers is segmented into a predetermined number of sub-images, each sub-image including a boundary for motion compensation and an indication that the image header and / or parameter set of the data stream can be rewritten when sub-image extraction is performed on the data stream.
[0101] Typically, for (sub)layers, the bitstream extraction process is specified as follows: - Discard NAL units that do not correspond to the (sub-layer) of interest. - Discard SEI messages (such as image timing SEI, buffer periodic SEI, etc.) - Optionally, obtain the appropriate SEI message from the nested SEI message associated with the (sub)layer of interest. However, as discussed above, discarding NAL units at a finer granularity than per (sub)layer in sub-image extraction, for example, can introduce problems that have remained unresolved to date. See, for example, the description of FD_NUT for CBR above. Another problem arises from the parameter set. In (sub)layer extraction, the parameter set can potentially contain additional information not related to the extracted (sub)layer; for example, it may contain some information about the (sub)layer that has been discarded, but it will be correct in the sense that it still describes the (sub)layer in the extracted sub-bitstream. The additional information about the discarded (sub)layer can then be ignored. This design principle is too complex to maintain when it comes to sub-image extraction, which is why, for example, in HEVC, MCTS nested SEI messages carry replacement parameter sets. However, in VVC, the image header is defined, which further complicates the situation and makes HEVC-type solutions with nested SEI messages infeasible.
[0102] Therefore, the parameter set is modified when performing sub-image extraction. Note that the parameter set contains information such as image size, profile_level information, and tile / slice grid (which may also exist in the image header), and this information needs to be changed when performing sub-image extraction.
[0103] In one embodiment, the extraction process is defined as: - Discard NAL units that do not correspond to the sub-image of interest. - Discard SEI messages (such as image timing SEI, buffer periodic SEI, etc.) - Obtain the appropriate SEI message from the nested SEI messages associated with the (sub)level of interest. - Discard parameter set - Add appropriate parameter sets - Discard image header - Add appropriate image headers The "appropriate" parameter set and image header need to be generated most likely by rewriting the pre-extracted parameter set and image header.
[0104] The information that needs to be changed is: -Related to level and HRD -Image Size - Block / Slice / Sub-image Grid Information related to level and HRD can be extracted from the corresponding SEI message (sub-image level information SEI message) to be used for rewriting the parameter set.
[0105] If a single sub-image is extracted during the process, the size of the resulting image can be easily derived because the size of each sub-image is readily available in SPS. However, when more than one sub-image is extracted from the bitstream (e.g., the resulting bitstream contains more than one sub-image), the size of the resulting image depends on the arrangement defined by the external components.
[0106] In one embodiment, bitstream constraints only allow the extraction of sub-images corresponding to rectangular regions in the original bitstream, and their relative arrangement in the original bitstream remains unchanged. Therefore, reordering or gaps within the extracted regions are not permitted.
[0107] The current sub-image coordinates are as follows: Figure 21 The diagram illustrates the top_left coordinates and the parameters for width and height. This allows for the easy extraction of the top-left, top-right, bottom-left, and bottom-right corners or regions. Therefore, the extraction process runs on all extracted sub-images, updating the minimum "x" and "y" coordinates of the top_left coordinate if a smaller value is found (e.g., searching for the minimum), and updating the maximum "x" and "y" coordinates of the bottom-right coordinate if a smaller value is found (e.g., searching for the maximum). For example, MinTopLeftX=PicWidth MinTopLeftY=PicHeight MaxBottomRightX=0 MaxBottomRightY=0 For i=0..NumExtSubPicMinus1 If(subpic_ctu_top_left_x[i] <MinTopLeftX) MinTopLeftX= subpic_ctu_top_left_x[ i ] If(subpic_ctu_top_left_y[i] <MinTopLeftY) MinTopLeftY= subpic_ctu_top_left_y[ i ] If(subpic_ctu_top_left_x[i]+subpic_width_minus1[i]>MaxBottomRightX) MaxBottomRightX=subpic_ctu_top_left_x[i]+subpic_width_minus1[i] If(subpic_ctu_top_left_y[i]+subpic_height_minus1[i]>MaxBottomRightY) MaxBottomRightY=subpic_ctu_top_left_y[i]+subpic_height_minus1[i] Then, using these values, the maximum image size can be derived, and by subtracting the corresponding values of MinTopLeftX or MinTopLeftY, the new values of subpic_ctu_top_left_x[i] and subpic_ctu_top_left_y[i] of each subimage in the extracted subimages can be derived.
[0108] In an alternative embodiment, signaling is provided that allows the parameter set and image header to be rewritten without requiring or deriving the values in question. For each potential combination of sub-images to be extracted, the image size that needs to be rewritten in the parameter set can be provided. This can be done, for example, in the form of an SEI message. The SEI message may include a parameter_type syntax element that indicates what information is internally helpful for rewriting. For example, type 0 could be the image size, type 1 could be the level of the extracted bitstream, type 2 could be information in the image header, combinations thereof, etc.
[0109] In other words, sub-image extraction is performed on a data stream of multiple images encoded into a video across multiple layers, wherein each image in all layers is segmented into a predetermined number of sub-images, each sub-image including a boundary extension for motion compensation. Sub-image extraction is performed on the data stream by discarding NAL units of the data stream that do not correspond to one or more sub-images and rewriting the parameter set and / or image header to obtain an extracted data stream associated with one or more sub-images of interest.
[0110] Let's talk about Figure 19 To reiterate, subimage extractor 1960 does not need subimages 1912, 1913, 1922, and 1923 because they do not belong to the subimage(s) of interest, and the corresponding NAL units 1942, 1943, 1952, and 1953 can be discarded together with the extracted images. At this point, Figure 19 The presence of pseudo-data in the image should be understood as optional, or if it exists, it can also be derived from sub-images that are not specific, such as... Figure 20 As explained in the text.
[0111] However, in the present regarding Figure 19 In the described example, extractor 1960 can rewrite header information HI1 and HI2 into HI3 and HI4, respectively. During this process (which can occur simultaneously when extracting sub-images), header information, such as video parameter sets (VPS) and / or sequence parameter sets (SPS) and / or image parameter sets (PPS), can be modified by rewriting their specific variable values.
[0112] Information used for rewriting can be derived from the data stream and can exist in the data stream in addition to the actual header information HI1 and HI2. Additional information can be related to HRD-related information, image size, and / or subdivision into sub-images. For example, supplementary enhancement information (SEI) messages can be used to indicate in the data stream that rewriting is permitted or that additional indications and / or additional information are present.
[0113] Sub-image extraction can be performed such that the extracted data stream consists of a set of extracted sub-images, wherein each of the sub-images corresponds to a rectangular region in the data stream.
[0114] The rewritable values include the sub-image width in the luminance sample, the sub-image height in the luminance sample, level information, HRD-related parameters, and image size.
[0115] The following section addresses the simulation prevention of obsolescence for sub_pic_id (5).
[0116] The sub-image ID is currently signaled in the slice header within the VVC. This is, for example, in... Figure 22 As shown in the diagram.
[0117] However, these are written as one of the first syntax elements in the slice header to make accessing the value easy. Note that the extraction process (or the merging process in which the id needs to be changed / checked) will require reading and / or writing the value of the sub-image id, and therefore easy access is expected. One aspect still missing for easy access to this value is emulation blocking. The syntax element slice_pic_order_cnt_lsb is at most 16 bits long and can take the value 0, such as 0x0000. The length of slice_sub_pic_id is 16 bits. Therefore, some combinations of slice_pic_order_cnt_lsb and slice_sub_pic_id values can cause emulation blocking to occur, which makes parsing the sub-image id by upper-layer applications more complex and even more difficult for upper-layer applications to change or modify, because if the value of slice_sub_pic_id changes to a different value of slice_sub_pic_id*, the length of the slice_header that blocks emulation may change.
[0118] In one embodiment, `slice_pic_order_cnt_lsb` is changed to `slice_pic_order_cnt_lsb_plus1`, and `slice_sub_pic_id` is changed to `slice_sub_pic_id_plus1`, so that emulation blocking is impossible between `slice_pic_order_cnt_lsb` and `slice_sub_pic_id` under any circumstances. This facilitates `slice_sub_pic_id` resolution but does not provide a solution for use cases where the value of `slice_sub_pic_id` needs to be changed. Note that `slice_sub_pic_id_plus1` does not guarantee that the least significant bit of `slice_sub_pic_id_plus1` is not zero in the last byte of the slice header containing `slice_sub_pic_id_plus1`. Therefore, depending on the value of the syntax element in the following two bytes, 0x00 may occur at the byte that could potentially trigger emulation blocking.
[0119] In another embodiment, such as in Figure 23 In the example, a single-bit syntax element followed by slice_sub_pic_id_plus1 resolves the described problem.
[0120] Clearly, the provided solution comes with overhead, such as the additional bit and the x_plus1 syntax element. Furthermore, the described problem only applies to the following: • To prevent simulation from occurring between slice_pic_order_cnt_lsb and slice_sub_pic_id: when the combined length of the two syntax elements is greater than 24 bits.
[0121] • For simulation prevention, the following occurs when slice_sub_pic_id_plus1 and the following syntax element: when slice_sub_pic_id_plus1 spans more than one byte.
[0122] Since the described problem may not occur very frequently, the described changes can be conditional on some gating flags in any set of parameters of the image header. This is, for example, in Figure 24 The diagram shows that Val (the offset from the sub-image id) is determined from the combined encoded length of slice_pic_order_cnt_lsb_plus1 and slice_subpic_id_plusVal, as follows: Val = codedLength(slice_pic_order_cnt_lsb_plus1 + slice_subpic_id_plusVal)<31 1:4. If the combined encoded length of slice_pic_order_cnt_lsb_plus1 and slice_subpic_id_plusVal is less than 31 bits, then this reads that Val is set to the value "1", and otherwise it is set to the value "4".
[0123] The following involves the CABAC zero word (6).
[0124] Cabac_zero_words can be inserted at the end of each slice. They can be located in any slice of the image. The current video specification describes image-level constraints that ensure that a cabac_zero_word is inserted if the bit ratio is too high.
[0125] When sub-images are extracted and indicated to conform to a specific profile and level, the following conditions applied to AU should also be applied individually to each of the sub-images: Export the variable RawMinCuBits as follows: RawMinCuBits = MinCbSizeY * MinCbSizeY * (BitDepth + 2 * BitDepth / (SubWidthC * SubHeightC ) ) The value of BinCountslnNalUnits should be less than or equal to (32 ÷ 3) * NumByteslnVclNalUnits + (RawMinCuBits * PicSizelnMinCbsY) ÷ 32. Therefore, the bitstream constraint or indication shows that each sub-image individually satisfies the conditions described above.
[0126] As a first embodiment, a flag can be signaled as a constraint flag, indicating that the sub-image satisfies the mentioned constraint. This is, for example, in... Figure 25 As shown in the diagram.
[0127] In another embodiment, such constraints are satisfied based on one or both of the following conditions: 1) The subpic has a subpic_treated_as_pic_flag equal to 1. 2) The bitstream contains level consistency indicators for subpics, for example, by means of SEI messages (existing subpic_level_info SEI messages).
[0128] In other words, in encoding multiple images of a video into a data stream, where each of the multiple images is segmented into a predetermined number of sub-images, and each sub-image includes a boundary for motion compensation, each of the sub-images can be encoded using context-adaptive binary arithmetic coding.
[0129] To this end, a zero word is provided for the data stream at the end of one or more slices or VCL NAL units of each sub-image to prevent any sub-image from exceeding a predetermined binary bit ratio.
[0130] like Figure 26 As can be seen, the encoder can generate bitstream 2605 in the following manner. Exemplary examples are shown portions of two access units 2620 or 2630. Each of them has an image encoded therein, subdivided into sub-images, such as... Figure 19 The image is shown as an example. However, Figure 26 The focus is on those portions of the bitstream associated with an exact sub-image of each image. For AU 2621, the encoding exemplarily results in only one slice 2621 into which the sub-image of the image of AU 2610 has been encoded, and in the case of AU 260, three slices are obtained, namely, the sub-image of AU 2630 has been encoded into 2631, 2632, and 2633.
[0131] The encoder uses binary arithmetic encoding such as CABAC. That is, it encodes syntax elements generated to describe the content of a video or image using syntax elements. Those not binary values are binaryized into binary strings. Thus, an image is encoded into syntax elements, which are then encoded into binary sequences, which are arithmetically encoded into a data stream, resulting in a bitstream, each of which has a portion of the video encoded therein, such as a sub-image. This portion consumes a certain number of bits, which are generated by arithmetically encoding a certain number of binary bits, producing a binary bit ratio. The decoder performs the opposite operation: the portion associated with a sub-image is arithmetically decoded to produce a binary sequence, which is then debinded to produce syntax elements, from which the decoder can reconstruct the sub-image. Therefore, in "binary bit ratio," "binary" refers to the number of bits or digits of the binary symbols to be encoded, and "bit" refers to the number of bits written / read from the CABAC bitstream.
[0132] To avoid exceeding a predetermined bit ratio, the encoder checks the bit ratio of the CABAC-coded portions of one or more slices into which the corresponding sub-image has been encoded for each sub-image. If the ratio is too high, it feeds the CABAC encoding engine with so many zeros to be CABAC-coded at the ends of one or more of the slices into which the corresponding sub-image is encoded, so that the ratio is no longer exceeded. These cabac zeros are... Figure 26 They are indicated by zero (i.e., "0"). They can be added at the end of sub-image 2621 (i.e., at the end of the last slice), or added in a manner distributed to the ends of one or more slices in slices 2631, 2632, and 2633 of the sub-image. That is, although Figure 26 It shows Figure 26 The distributed variant in the middle, but alternatively, the zero word can be appended completely to the end of the last slice 2613.
[0133] When the data stream 2605 reaches decoder 2640, it causes the decoder to decode the data stream or at least a portion thereof related to the aforementioned sub-image, and to parse the corresponding data portions within AUs 2620 and 2630. In doing so, decoder 2640 can determine the amount or number of zero codewords encoded at the end of one or more slices of a sub-image by determining how many zero codewords are needed to not exceed the binary bit ratio, and can discard cabac zeros. However, it is also possible that decoder 2640 can parse the slices in a manner that allows the zeros to be distinguished from other syntax elements by other means, and then discard the zeros. That is, the decoder may or may not check the binary bit ratio of one or more slices of a sub-image in order to continue CABAC decoding with the number of zeros required to bring the ratio into a predetermined range. If not, the decoder can syntactically distinguish the zeros that have already been CABAC decoded from other syntax elements.
[0134] If the ratio condition is not met, the decoder 2640 may fall into an error mode, and the predetermined error handling may be triggered by the decoder.
[0135] Finally, it's important to note that the number of sub-images can be any number. In other words, specifically, at least one can mean two or more sub-images.
[0136] Each subpic can be evaluated for independence attributes (e.g., subpic_treated_as_pic_flag) and / or for corresponding level consistency indicators (e.g., in the Supplemental Enhancement Information (SEI) message). Then, a zero word can be provided for the subpic based solely on and in response to the evaluation.
[0137] A zero word may be provided at the end of one or more slices of each subimage of at least one subimage for the data stream, such that the number of bits encoded into the data stream for the corresponding subimage using context-adaptive arithmetic coding is less than or equal to the number determined by the product of the byte length of one or more VCL NAL units of the data stream associated with the corresponding subimage and a predetermined factor.
[0138] Alternatively, a zero word may be provided at the end of one or more slices of each sub-image of at least one sub-image, such that the number of bits encoded into the data stream for the corresponding sub-image using context-adaptive arithmetic coding is less than or equal to the sum of a first product between the byte length of one or more VCL NAL units of the data stream associated with the corresponding sub-image and a first predetermined factor, and a second product between a second predetermined factor, the minimum number of bits per coded block, and the number of coded blocks constituting the corresponding sub-image.
[0139] Although some aspects have been described in the context of the device, it is clear that these aspects also represent a description of the corresponding method, where blocks or devices correspond to method steps or features of method steps. Similarly, aspects described in the context of method steps also represent a description of corresponding blocks or items or features of the corresponding device. Some or all of the method steps can be executed by (or using) hardware devices, such as microprocessors, programmable computers, or electronic circuits. In some embodiments, one or more of the most important method steps can be executed by such devices.
[0140] The data stream of the present invention can be stored on a digital storage medium or transmitted on a transmission medium such as the Internet, such as a wireless transmission medium or a wired transmission medium.
[0141] Depending on certain implementation requirements, embodiments of the present invention can be implemented in hardware or software. The implementation can be performed using a digital storage medium (e.g., floppy disk, DVD, Blu-ray, CD, ROM, PROM, EPROM, EEPROM, or FLASH memory) having electrically readable control signals stored thereon, which cooperates with (or is capable of cooperating with) a programmable computer system to perform the corresponding methods. Therefore, the digital storage medium can be computer-readable.
[0142] According to some embodiments of the invention, a data carrier having electronically readable control signals is included, the data carrier being capable of cooperating with a programmable computer system to perform one of the methods described herein.
[0143] Generally, embodiments of the present invention can be implemented as a computer program product having program code that, when run on a computer, operates to perform one of the methods. The program code may, for example, be stored on a machine-readable medium.
[0144] Other embodiments include a computer program stored on a machine-readable medium for performing one of the methods described herein.
[0145] Therefore, in other words, an embodiment of the method of the present invention is a computer program having program code that, when run on a computer, performs one of the methods described herein.
[0146] Therefore, a further embodiment of the method of the present invention is a data carrier (or digital storage medium, or computer-readable medium) comprising a computer program recorded thereon for performing one of the methods described herein. The data carrier, digital storage medium, or recording medium is typically tangible and / or non-transitory.
[0147] Therefore, a further embodiment of the method of the present invention represents a sequence of signals or a stream of data for a computer program to perform one of the methods described herein. The sequence of signals or the stream of data may, for example, be configured to be transmitted via a data communication connection (e.g., via the Internet).
[0148] Further embodiments include processing components (e.g., a computer or programmable logic device) configured or adapted to perform one of the methods described herein.
[0149] A further embodiment includes a computer having a computer program mounted thereon for performing one of the methods described herein.
[0150] Further embodiments of the invention include a device or system configured to transmit (e.g., electronically or optically) a computer program for performing one of the methods described herein to a receiver. For example, the receiver may be a computer, a mobile device, a memory device, etc. The device or system may, for example, include a file server for transmitting the computer program to the receiver.
[0151] In some embodiments, a programmable logic device (e.g., a field-programmable gate array) may be used to perform some or all of the functionalities of the methods described herein. In some embodiments, the field-programmable gate array may cooperate with a microprocessor to perform one of the methods described herein. Generally, the methods are preferably performed by any hardware device.
[0152] The devices described in this article can be implemented using hardware devices, computers, or a combination of hardware devices and computers.
[0153] The device described herein or any component of the device described herein may be implemented, at least in part, in hardware and / or software.
[0154] The methods described in this article can be performed using hardware devices, computers, or a combination of hardware devices and computers.
[0155] The methods described herein or any component of the devices described herein may be performed, at least in part, by hardware and / or by software.
[0156] The embodiments described above are merely illustrative of the principles of the invention. It is understood that modifications and variations of the arrangements and details described herein will be apparent to those skilled in the art. Therefore, the intent is limited only by the scope of the impending patent claims and not by the specific details presented through the description and explanation of the embodiments herein.
Claims
1. A data encoding method, comprising: Divide a specific image into multiple sub-images; Determine the binary bit constraint such that the number of binary bits for encoding the sub-image is less than or equal to (32÷3)*NumByteslnVclNalUnits + (RawMinCuBits * PicSizelnMinCbsY)÷32; The specific image is encoded into a bitstream via Context Adaptive Binary Arithmetic Encoding (CABAC), the encoding including inserting one or more zero words into the bitstream such that a binary bit constraint corresponding to one of the plurality of sub-images is satisfied; and Instructions are provided to treat the sub-image as an image that has been inserted with one or more zeros via CABAC encoding.
2. The method of claim 1, wherein the instruction comprises: Set sps_subpic_treated_as_pic_flag to 1; or It is inferred that when sps_subpic_treated_as_pic_flag is omitted from the bitstream, sps_subpic_treated_as_pic_flag is set to 1.
3. The method of claim 1, wherein the indication for treating a sub-image as an image includes an indication for applying motion-compensated boundary extension to the sub-image.
4. The method according to claim 1, wherein, The zero is inserted at the end of one or more slices of the sub-image.
5. The method according to claim 1, The plurality of images comprise at least two layers; The method further includes segmenting an image of at least one layer into a predetermined number of sub-images specific to each layer; One or more of the images or sub-images in one layer correspond to one or more images or sub-images in other layers; At least one of the sub-images includes a boundary for boundary extension for motion compensation, and The method further includes encoding instructions such that at least one of the boundaries of corresponding sub-images or corresponding images in different layers is aligned with each other.
6. A non-transitory computer-readable medium having instructions that, when executed, cause at least one processor to perform the method of claim 1.
7. An encoder, comprising: At least one processor is configured to perform operations including the following: Divide a specific image into multiple sub-images; Determine the binary bit constraint such that the number of binary bits for encoding the sub-image is less than or equal to (32÷3)*NumByteslnVclNalUnits + (RawMinCuBits * PicSizelnMinCbsY)÷32; The specific image is encoded into a bitstream via Context Adaptive Binary Arithmetic Encoding (CABAC), the encoding including inserting one or more zero words into the bitstream such that a binary bit constraint corresponding to one of the plurality of sub-images is satisfied; and Instructions are provided to treat the sub-image as an image that has been inserted with one or more zeros via CABAC encoding.
8. The encoder of claim 7, wherein the indication includes: Set sps_subpic_treated_as_pic_flag to 1; or It is inferred that when sps_subpic_treated_as_pic_flag is omitted from the bitstream, sps_subpic_treated_as_pic_flag is set to 1.
9. The encoder of claim 7, wherein the indication for treating a sub-image as an image includes an indication for applying motion-compensated boundary extension to the sub-image.
10. The encoder according to claim 7, wherein, The zero is inserted at the end of one or more slices of the sub-image.
11. The encoder according to claim 7, The plurality of images comprise at least two layers; An image of at least one layer is segmented into a predetermined number of sub-images specific to each layer; One or more of the images or sub-images in one layer correspond to one or more images or sub-images in other layers; At least one of the sub-images includes a boundary for boundary extension for motion compensation, and At least one processor is also configured to encode instructions such that at least one of the boundaries of corresponding sub-images or corresponding images in different layers are aligned with each other.
12. A data decoding method, comprising: Receive a bitstream comprising multiple images, wherein a particular image among the multiple images is divided into multiple sub-images; Based on the bitstream, a decoding or inference instruction is provided, which is used to treat a sub-image as an image in which one or more zero words have been inserted into the bitstream by context-adaptive binary arithmetic encoding CABAC, wherein the sub-image corresponds to one of a plurality of sub-images; Determine binary bit constraints such that the number of binary bits encoding the sub-image is less than or equal to (32 ÷ 3) * NumByteslnVclNalUnits + (RawMinCuBits * PicSizelnMinCbsY) ÷ 32; and The specific image is decoded from the bit stream via context-adaptive binary arithmetic coding (CABAC), the decoding including parsing one or more zero words from the bit stream such that the binary bit constraints corresponding to the sub-image are satisfied.
13. The method of claim 12, wherein the instruction comprises: sps_subpic_treated_as_pic_flag is set to 1.
14. The method of claim 12, wherein the indication for treating a sub-image as an image includes an indication for applying motion-compensated boundary extension to the sub-image.
15. The method of claim 12, wherein the zero word is parsed from the end of one or more slices of the sub-image.
16. The method according to claim 12, The plurality of images comprise at least two layers; An image of at least one layer is segmented into a predetermined number of sub-images specific to each layer; One or more of the images or sub-images in one layer correspond to one or more images or sub-images in other layers; At least one of the sub-images includes a boundary for boundary extension for motion compensation, and The method further includes interpreting instructions such that at least one of the boundaries of corresponding sub-images or corresponding images in different layers is aligned with each other.
17. A non-transitory computer-readable medium having instructions that, when executed, cause at least one processor to perform the method of claim 12.
18. A decoder, comprising: At least one processor is configured to perform operations including the following: Receive a bitstream comprising multiple images, wherein a particular image among the multiple images is divided into multiple sub-images; Based on the bitstream, a decoding or inference instruction is provided, which is used to treat a sub-image as an image in which one or more zero words have been inserted into the bitstream by context-adaptive binary arithmetic encoding CABAC, wherein the sub-image corresponds to one of a plurality of sub-images; Determine binary bit constraints such that the number of binary bits encoding the sub-image is less than or equal to (32 ÷ 3) * NumByteslnVclNalUnits + (RawMinCuBits * PicSizelnMinCbsY) ÷ 32; and The specific image is decoded from the bit stream via context-adaptive binary arithmetic coding (CABAC), the decoding including parsing one or more zero words from the bit stream such that the binary bit constraints corresponding to the sub-image are satisfied.
19. The decoder of claim 18, wherein the indication includes: sps_subpic_treated_as_pic_flag is set to 1.
20. The decoder according to claim 18, The indications for treating a sub-image as an image include indications for applying motion compensation to the boundary extension of the sub-image.
21. The decoder according to claim 18, The zero word is parsed from the end of one or more slices of the sub-image.
22. The decoder according to claim 18, The plurality of images comprise at least two layers; The image of at least one layer is segmented into a predetermined number of sub-images for each layer; One or more of the images or sub-images in one layer correspond to one or more images or sub-images in other layers; At least one of the sub-images includes a boundary for boundary extension for motion compensation, and The at least one processor is configured to interpret instructions such that at least one of the boundaries of corresponding sub-images or corresponding images in different layers are aligned with each other.