Method and device for encoding and decoding a video stream using sub-pictures
By encoding the CTB splitting information of sub-pictures in the video bitstream, the problem of inferring sub-pictures CTB splitting in the prior art is solved, and more efficient video bitstream decoding is achieved.
Patent Information
- Application Number
- CN202080065219.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-12-17
- Filing Date
- 2020-09-07
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2040-09-07
AI Technical Summary
When decoding in a video bitstream, the split of the right or bottom CTB of the sub-picture cannot be effectively inferred, resulting in the decoding failure.
By encoding information indicating that the sub-picture has a CTB size multiple size and can be decoded independently in the bitstream, and encoding the encoded tree blocks that make up the sub-picture, to allow the decoder to infer the split of the CTB.
It realizes that when the sub-picture is not located on the right or bottom of the image, the decoder allows the decoder to infer CTB splitting, thereby solving the problem of decoding failure and improving the decoding efficiency of the video bit stream.
Smart Images

Figure CN114556936B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to a method and apparatus for encoding and decoding a video bitstream, which facilitates the displacement of sub-pictures. The present invention more particularly relates to the encoding and decoding of a video bitstream obtained by merging sub-pictures from different video bitstreams. Background Art
[0002] The size of an image in a video bitstream may not be a multiple of the size of a coding tree block (CTB) used in the encoding process. The CTB can be recursively split during encoding, particularly to find a coding block size that optimizes the encoding process. Thus, the CTB can be split until the minimum coding block size. When the size of the image is not a multiple of the size of the CTB, the right or bottom boundary of the image straddles the rightmost or bottommost CTB. In this case, an inferred split of the CTB is provided. Coding blocks that fall outside the image are generally not encoded.
[0003] When considering splitting an image into sub-pictures, the rightmost or bottommost sub-picture may contain an incomplete CTB that has undergone an inferred split, which includes some unencoded coding blocks.
[0004] The size of the image in pixels and the size of the CTB are provided to the decoder in the bitstream. Thus, the decoder can determine the exact positions of the right and bottom boundaries of the image and perform an inferred split of the rightmost and bottommost CTBs in the image.
[0005] When considering the displacement of a sub-picture located at the right or bottom boundary of an image elsewhere in the image, since the size of the sub-picture is provided as an integer number of CTBs, the decoder cannot infer the split of the rightmost or bottommost CTB and the identification of missing coding blocks. Decoding fails.
[0006] The present invention aims to solve one or more of the above problems. An encoding method is proposed that includes encoding information that allows a decoder to infer the split of a CTB located on the right or bottom of a sub-picture when the width or height of the sub-picture is not a multiple of the CTB size and the sub-picture is not located on the right or bottom of the image. A corresponding decoding method for the generated bitstream is also proposed. Summary of the Invention
[0007] According to one aspect of the present invention, there is provided a method for encoding video data including pictures into a bitstream, where the pictures are segmented into sub-pictures, and the method includes: for at least one sub-picture, encoding in the bitstream information indicating that the picture has a size that is a multiple of the size of the coding tree block and that the sub-picture is independently decodable; encoding the coding tree blocks constituting the sub-picture into the bitstream.
[0008] In an embodiment, the information is related to all sub-pictures.
[0009] In an embodiment, the information is defined in a sequence parameter set, which is a syntax structure containing syntax elements applied to the pictures of the bitstream.
[0010] In an embodiment, the sub-pictures are further identified by sub-picture identifiers.
[0011] In an embodiment, it is prohibited to define the sub-picture identifier in the picture header.
[0012] In an embodiment, the sub-picture identifiers defined in the picture parameter set must be the same in all picture parameter sets.
[0013] In an embodiment, the information is associated with a specific profile.
[0014] According to another aspect of the present invention, there is provided a method for encoding video data including pictures into a bitstream, where the pictures are segmented into sub-pictures, and the method includes: for at least one sub-picture, encoding in the bitstream information indicating the consistency window of the sub-picture; encoding the coding tree blocks constituting the sub-picture into the bitstream.
[0015] In an embodiment, the information indicating the consistency window is defined in an SEI message.
[0016] In an embodiment, the information defines a left offset, a right offset, an upper offset, and a lower offset for the sub-picture.
[0017] According to another aspect of the present invention, there is provided a method for encoding video data including pictures into a bitstream, where the pictures are segmented into sub-pictures, and the method includes: for at least one sub-picture, encoding in the bitstream first information indicating that the sub-picture has a size that is a multiple of the size of the coding tree block and that the sub-picture is independently decodable; encoding in the bitstream second information indicating the consistency window of the sub-picture; encoding the coding tree blocks constituting the sub-picture into the bitstream.
[0018] According to another aspect of the present invention, there is provided a computer program product for a programmable device, the computer program product comprising a sequence of instructions which, when loaded into and executed by the programmable device, are for implementing the method according to the present invention.
[0019] According to another aspect of the present invention, there is provided a computer-readable storage medium storing instructions of a computer program for implementing the method according to the present invention.
[0020] According to another aspect of the present invention, there is provided a computer program which, when executed, causes the method according to the present invention to be performed.
[0021] According to another aspect of the present invention, there is provided an apparatus for encoding video data including pictures into a bitstream, the pictures being segmented into sub-pictures, the apparatus comprising a processor configured to perform, for at least one sub-picture: encoding, in the bitstream, information indicating that the sub-picture has a size that is a multiple of the size of an encoding tree block and that the sub-picture is independently decodable; and encoding the encoding tree blocks constituting the sub-picture into the bitstream.
[0022] According to another aspect of the present invention, there is provided an apparatus for encoding video data including pictures into a bitstream, the pictures being segmented into sub-pictures, the apparatus comprising a processor configured to perform, for at least one sub-picture: encoding, in the bitstream, information indicating a coherence window of the sub-picture; and encoding the encoding tree blocks constituting the sub-picture into the bitstream.
[0023] According to another aspect of the present invention, there is provided an apparatus for encoding video data including pictures into a bitstream, the pictures being segmented into sub-pictures, the apparatus comprising a processor configured to perform, for at least one sub-picture: encoding, in the bitstream, first information indicating that the sub-picture has a size that is a multiple of the size of an encoding tree block and that the sub-picture is independently decodable; encoding, in the bitstream, second information indicating a coherence window of the sub-picture; and encoding the encoding tree blocks constituting the sub-picture into the bitstream.
[0024] At least part of the method according to the present invention may be computer-implemented. Accordingly, the present invention may take the form of a full hardware embodiment, a full software embodiment (including firmware, resident software, microcode, etc.) or an embodiment combining software and hardware aspects, which software and hardware are generally referred to herein as "circuitry", "module" or "system". Furthermore, the present invention may take the form of a computer program product embodied in any tangible expression medium having computer-usable program code embodied in the medium.
[0025] Since the present invention can be implemented in software, the present invention can be embodied as computer-readable code for providing to a programmable device on any suitable carrier medium. The tangible non-transitory carrier medium may include storage media such as floppy disks, CD-ROMs, hard disk drives, tape devices, or solid-state memory devices, etc. The transient carrier medium may include signals such as electrical signals, electronic signals, optical signals, acoustic signals, magnetic signals, or electromagnetic signals (e.g., microwave or RF signals). BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Embodiments of the present invention will now be described, by way of example only and with reference to the following drawings, in which:
[0027] Figure 1a and 1b show two different application examples for combining regions of interest;
[0028] Figure 2a show some partitions in an encoding system;
[0029] Figure 2b show an example of partitioning a picture into sub-pictures;
[0030] Figure 2c show an example of partitioning a picture into titles and bricks; Figure 3 show the organization of a bitstream in an exemplary encoding system VVC;
[0031] Figure 4 schematically show a quadtree inference mechanism for encoding a tree block spanning an image boundary in VVC;
[0032] Figure 5a and 5b show the creation of a bitstream in which a boundary sub-picture is moved to a non-boundary position;
[0033] Figure 6 show a method of encoding a video image into a bitstream according to a first aspect of the present invention;
[0034] Figure 7 show the general decoding process of an embodiment of the present invention;
[0035] Figure 8 show the decoding process of a CTB encoded in a slice;
[0036] Figure 9 show an example of a picture being segmented into 16 sub-pictures;
[0037] Figure 10An example is shown where an image is split into 16 sub-images with splitting inference boundaries of the image width;
[0038] Figure 11 The concept of a picture-wide boundary is shown;
[0039] Figure 12 It is a schematic block diagram of a computing device for implementing one or more embodiments of the present invention. Detailed implementation manners
[0040] Figure 1a and 1b Two different application examples for combining regions of interest are shown.
[0041] For example, Figure 1a An example is shown where picture (or frame) 100 from the first video bitstream and picture 101 from the second video bitstream are merged into picture 102 of the resulting bitstream. Each picture is composed of four regions of interest numbered from 1 to 4. Picture 100 has been encoded using encoding parameters that result in high-quality encoding. Picture 101 has been encoded using encoding parameters that result in low-quality encoding. As is well known, pictures encoded with low quality are associated with a lower bitrate compared to pictures encoded with high quality. The resulting picture 102 combines regions of interest 1, 2, and 4 from picture 101 (and thus encoded with low quality) with region of interest 3 from picture 100 (encoded with high quality). The goal of this combination is typically to obtain a high-quality region of interest, here region 3, while keeping the resulting bitrate reasonable by encoding regions 1, 2, and 4 with low quality. This scenario may occur particularly in the context of omnidirectional content, allowing the actually visible content to have higher quality while the rest has lower quality.
[0042] Figure 1b A second example is shown where four different videos A, B, C, and D are merged to form the resulting video. Picture 103 of video A is composed of regions of interest A1, A2, A3, and A4. Picture 104 of video B is composed of regions of interest B1, B2, B3, and B4. Picture 105 of video C is composed of regions of interest C1, C2, C3, and C4. Picture 106 of video D is composed of regions of interest D1, D2, D3, and D4. Picture 107 of the resulting video is composed of regions B4, A3, C3, and D1. In this example, the resulting video is a mosaic video of different regions of interest of the original video streams. The regions of interest of the original video streams are rearranged and combined in new positions in the resulting video stream.
[0043] The compression of video relies on block - based video coding in most coding systems such as the HEVC (representing High Efficiency Video Coding) or the emerging VVC (representing Versatile Video Coding) standard. In these coding systems, a video is composed of a sequence of frames or pictures or images or samples that can be displayed at several different times. In the case of multi - layer video (e.g., scalable, stereoscopic, 3D video), several pictures can be decoded to form the resulting image to be displayed at a certain moment. A picture can also be composed of different image components. For example, for encoding luminance, chrominance, or depth information.
[0044] The compression of a video sequence relies on several partitioning techniques for each picture. Figure 2a Some partitions in the coding system are shown. Pictures 201 and 202 are divided by coding tree units (CTUs) shown by dashed lines. A CTU is the basic unit of encoding and decoding. For example, a CTU can encode an area of 128×128 pixels.
[0045] A coding tree unit (CTU) can also be called a block, a macro - block, or a coding block. Different image components can be encoded simultaneously, or it can be limited to only one image component. When an image contains several components, a CTU corresponds to a CTB for each component. In the following, the present invention applies to both the CTU and CTB levels.
[0046] As Figure 2a shown, a picture can be partitioned according to a grid of tiles shown by thin solid lines. A tile is a picture part and thus a rectangular area of pixels that can be defined independently of the CTU partitioning. The boundaries of a tile and a CTU can be different. As in the example shown, a tile can also correspond to a sequence of CTUs, meaning that the boundaries of the tile coincide with those of the CTUs.
[0047] The tile definition stipulates that the tile boundaries break the spatial coding dependencies. This means that the encoding of CTUs within a tile is not based on pixel data from another tile in the picture.
[0048] Some coding systems (e.g., VVC) provide the concept of a slice. This mechanism allows a picture to be partitioned into one or more slice groups. Each slice is composed of one or more slices. As shown in Pictures 201 and 202, two different types of slices are provided. The first type of slice is limited to slices that form a rectangular region in the picture. Picture 201 shows the partitioning of a picture into five different rectangular slices. The second type of slice is limited to consecutive slices in raster scan order. Picture 202 shows the partitioning of a picture into three different slices composed of consecutive slices in raster scan order. Rectangular slices are structures used for the selection of regions of interest in video processing. Slices can be encoded as one or more NAL units in the bitstream. A NAL unit, which stands for Network Abstraction Layer unit, is a logical unit for encapsulating data in an encoded bitstream. In an example of a VVC coding system, a slice is encoded as a single NAL unit. When a slice is encoded as multiple NAL units in the bitstream, each NAL unit of the slice is a slice segment. A slice segment includes a slice segment header containing the encoding parameters of the slice segment. The header of the first segment NAL unit of a slice contains all the encoding parameters of the slice. The slice segment headers of subsequent NAL units of the slice can contain fewer parameters than the first NAL unit. In this case, the first slice segment is an independent slice segment, and the subsequent segments are dependent slice segments.
[0049] In OMAF v2 ISO / IEC 23090-2, a subpicture is a part of a picture that represents a spatial subset of the original video content, which has been split into spatial subsets before video encoding on the content generation side. A subpicture is, for example, one or more slices that form a rectangular region.
[0050] Figure 2b An example of partitioning a picture into subpictures is shown. A subpicture represents a part of the picture that covers a rectangular region of the picture. Each subpicture can have different sizes and encoding parameters. For example, different slice grids and slice partitions can be defined for each subpicture. In Figure 2b Picture 204 is subdivided into 24 subpictures including Subpictures 205 and 206. These two subpictures are similar to the slice grids and partitions in the slices further described in Figure 2a Pictures 201 and 202. In a second example, the slice and block partitions are not defined for each subpicture but at the picture level. Then, a subpicture is defined as one or more slices that form a rectangular region.
[0051] Figure 2c An example of partitioning using block partitioning is shown. Each slice can include a set of blocks. A block is a consecutive set of CTUs that form a row in a slice. For example, Figure 2cFrame 207 is segmented into 25 slices. Each slice contains exactly one block, except for the rightmost column of slices which contains two-block slices. For example, slice 208 contains two blocks 209 and 210. When block partitioning is employed, a stripe contains blocks from one slice or several blocks from other slices. In other words, a VCL NAL unit is a collection of blocks rather than a collection of slices.
[0052] Figure 3 Shows the organization of the bitstream in an exemplary coding system VVC.
[0053] The bitstream 300 according to the VVC coding system consists of an ordered sequence of syntax elements and coded data. The syntax elements and coded data are placed into NAL units 301 - 305. There are different NAL unit types. The network abstraction layer provides the ability to encapsulate the bitstream into different protocols (such as RTP / IP (representing Real-Time Protocol / Internet Protocol), ISO base media file format, etc.). The network abstraction layer also provides a framework for packet loss resilience.
[0054] NAL units are divided into VCL NAL units and non-VCL NAL units, where VCL stands for Video Coding Layer. VCL NAL units contain the actual coded video data. Non-VCL NAL units contain additional information. This additional information can be parameters required for decoding the coded video data, or supplementary data that can enhance the usability of the decoded video data. NAL unit 305 corresponds to a stripe and constitutes the VCL NAL unit of the bitstream. Different NAL units 301 - 304 correspond to different parameter sets, and these NAL units are non-VCL NAL units. The VPS NAL unit 301 (VPS stands for Video Parameter Set) contains parameters defined for the entire video and thus contains parameters defined for the entire bitstream. The naming of the VPS can be changed and for example become the DPS in VVC. In an alternative, the VPS and DPS are different parameter set NAL units. The DPS (which stands for Decoder Parameter Set) NAL unit can define more static parameters than those in the VPS. In other words, the parameters of the DPS change less frequently than the parameters of the VPS. The SPS NAL unit 302 (SPS stands for Sequence Parameter Set) contains parameters defined for a video sequence. Specifically, the SPS NAL unit can define sub-pictures of the video sequence. The syntax of the SPS contains, for example, the following syntax elements:
[0055]
[0056] The descriptor list gives the coding of syntax elements. u(1) means that one bit is used to code the syntax element, and ue(v) means that the syntax element is coded using unsigned integer order-0 Exp-Golomb coding, where the first left bit is a variable length coding.
[0057] The presence of sub-pictures in an image depends on the value of subpics_present_flag. When this flag is equal to 0, it indicates that the image does not contain sub-pictures. When equal to 1, a set of syntax elements specifies the sub-pictures in the frame. The syntax element max_subpics_minus1 specifies the maximum number of sub-pictures in a picture of the video sequence. Then, the SPS defines the sub-picture partition with a grid of sub-picture grid elements whose size is defined by subpic_grid_col_width_minus1 and subpic_grid_row_height_minus1. Each grid element specifies subpic_grid_idx[i][j] (where i and j are the coordinates of the element in the grid), which is the index of the sub-picture. There are as many subpic_grid_idx[i][j] values as there are sub-pictures in a picture of the video sequence. The subpic_grid_idx[i][j] syntax element is the identifier of the sub-picture. All grid elements sharing the same index value form a rectangular region corresponding to the sub-picture with the index (or identifier) equal to the index value. The subpic_treated_as_pic_flag[i] syntax element indicates whether the sub-picture boundary should be treated as a picture boundary except for loop filter processing. The loop_filter_across_subpic_enabled_flag[i] syntax element states whether the loop filter is applied across the sub-picture boundary.
[0058] The PPS NAL unit 303 (PPS stands for Picture Parameter Set) contains parameters defined for a picture or a group of pictures. The APS NAL unit 304 (APS stands for Adaptation Parameter Set) contains parameters of loop filters (typically Adaptive Loop Filter (ALF) or shaper model (or luminance mapping with chroma scaling model)) defined at the slice level. The bitstream may also contain SEI (stands for Supplemental Enhancement Information) NAL units. The occurrence periods of these parameter sets in the bitstream are variable. The VPS defined for the whole bitstream needs to occur only once in the bitstream. In contrast, the APS defined for a slice may occur once for each slice in each picture. In fact, different slices may depend on the same APS, and thus there are usually fewer APSs than the slices in each picture. When a picture is partitioned into sub-pictures, new parameter sets or PPSs may be defined for each sub-picture or group of sub-pictures.
[0059] Each of the VCL NAL units 305 contains a slice. A slice may correspond to the whole picture or a sub-picture, a single tile or multiple tiles, or a single block or multiple blocks. A slice consists of a slice header 310 and an RBSP (raw byte sequence payload) 311 containing blocks.
[0060] The syntax of the PPS proposed in the current version of VVC includes syntax elements that specify the size of the picture and the luma samples, as well as the partitioning of each picture into tiles, blocks, and slices.
[0061] The syntax of the PPS proposed in the current version of VVC is organized as follows:
[0062]
[0063]
[0064] The descriptor column gives the encoding of the syntax elements, u(1) means encoding the syntax element using one bit, ue(v) means encoding the syntax element using unsigned integer order-0 Exp-Golomb coding, where the first left bit is variable length coded. The syntax elements pic_width_in_luma_samples and pic_height_in_luma_samples specify the width and height of the picture in luma samples.
[0065] When the number of tiles in the picture is greater than one (single_tile_in_pic_flag equals 0), the PPS definition specifies several syntax elements (not shown in the table above) that partition the tiles in the frame into a tile grid.
[0066] When there are bricks (brick_spliting_present_flag equals 1), the PPS includes loops over the tiles in the tile grid to indicate whether a tile is split into bricks. When a tile contains bricks, the brick configuration of the tile is encoded in the PPS.
[0067] The slice partitioning is represented by the following syntax elements:
[0068] The syntax element single_tile_in_pic_flag indicates whether the picture contains a single tile. In other words, when this flag is true, there is only one tile and one slice in the picture.
[0069] single_brick_per_slice_flag indicates whether each slice contains a single brick. In other words, when this flag is true, all the bricks in the picture belong to different slices.
[0070] The syntax element rect_slice_flag indicates that the slices of the picture form a rectangular shape as shown in Figure 201.
[0071] When present, the syntax element num_slices_in_pic_minus1 equals the number of rectangular slices in the picture minus one.
[0072] Then, syntax elements (not shown in the table above) encode the position of each slice relative to the brick partitioning in a "for loop" over all the slices in the picture. The index of the slice position parameter in this for loop is the index of the slice.
[0073] When signalled_slice_id_flag equals 1, a slice identifier is specified. In this case, the signalled_slice_id_length_minus1 syntax element indicates the number of bits used to encode the value of each slice identifier. The slice_id[] association table is indexed by the slice index and contains the identifiers of the slices. When signalled_slice_id_flag equals 0, slice_id is indexed by the slice index and contains the slice index of the slice.
[0074] In summary, the PPS contains syntax elements that allow the determination of the slice positions in the frame. Since the sub-pictures form rectangular regions in the frame, it is possible to determine the sets of slices, tiles, and bricks that belong to a sub-picture.
[0075] The slice header includes a slice address according to the following syntax in the current VVC version:
[0076]
[0077] When the slice is not rectangular, the slice header indicates the number of slices in the slice NAL unit by means of the num_tiles_in_slice_minus1 syntax element.
[0078] Each slice 320 may include a slice segment header 330 and slice segment data 331. The slice segment data 331 includes encoded coding blocks 340. In the current version of the VVC standard, there is no slice segment header, and the slice segment data contains the coding block data 340.
[0079] In a variant, the video sequence includes sub-pictures; the syntax of the slice header may be as follows:
[0080]
[0081] The slice header includes a slice_subpic_id syntax element that specifies the identifier of the sub-picture to which it belongs (e.g., one of the values corresponding to subpic_grid_idx[i][j] defined in the SPS). As a result, all slices in the video sequence that share the same slice_sub_pic_id belong to the same sub-picture.
[0082] For illustrative purposes only, Figure 4 A quadtree inference mechanism for coding tree blocks that cross image boundaries used in VVC is schematically shown. In VVC, an image is not limited to having a width and height that are multiples of the coding tree block size. Then, the rightmost coding tree block of a frame may cross the right boundary 401 of the image, and the bottommost coding tree block of a frame may cross the bottom boundary 402 of the image. In these cases, VVC defines a quadtree inference mechanism for coding tree blocks that cross the boundary. The mechanism includes recursively splitting any coding block of a coding tree block that crosses the image boundary until there are no more coding blocks that cross the boundary, or until these coding tree blocks reach the maximum quadtree depth. For example, the coding tree block 403 is not automatically split, while the coding tree blocks 404, 405, and 406 are automatically split. The absence of an inferred quadtree is signaled: the decoder must infer the same quadtree at the image boundary. However, by signaling the split information of these coding tree blocks (if the maximum quadtree depth is not reached), the automatically obtained quadtree can be further refined for coding tree blocks within the frame, e.g., as shown in 407.
[0083] When splitting the CTB into coding blocks, there is a smallest coding block that cannot be split. The size of this smallest coding block (which is square) is given by MinCBSizeY. In some examples, MinCBSizeY is equal to 4.
[0084] During encoding, the coding blocks of coding tree blocks located outside the image are generally not encoded in the bitstream. During decoding, the decoder uses the same quadtree inference mechanism and knows that these coding blocks are not encoded. The decoder can decode other coding blocks that fall within the image and have been encoded in the bitstream. The obtained coding tree is encoded in the bitstream by the encoder. The decoder relies on this encoded coding tree information to correctly identify the blocks for decoding and reconstructing the image. Specifically, the encoded coding tree contains parameters that specify whether a coding unit is split, such as a parameter called split_cu_flag.
[0085] Figure 5a and 5b illustrates the creation of a bitstream in which a border sub-picture is moved to a non-border position.
[0086] In this example, the first bitstream 500 is composed of 4 sub-pictures 1 HQ to 4 HQ . This bitstream represents a high-quality version of the video. The second bitstream 501 is composed of 4 sub-pictures 1 LQ to 4 LQ . This bitstream represents a low-quality version of the same video. The bitstream 502 is created by merging and rearranging some sub-pictures emitted from the bitstreams 500 and 501.
[0087] Specifically, the bitstream 502 is composed of sub-pictures 2 HQ , 1 LQ , 4 HQ and 3 LQ , where sub-picture 2 HQ is moved from the upper-right position in the bitstream 500 to the upper-left position in the bitstream 502, sub-picture 4 HQ is moved from the lower-right position in the bitstream 500 to the lower-left position in the bitstream 502, sub-picture 1 LQ is moved from the upper-left position in the bitstream 501 to the upper-right position in the bitstream 502, and sub-picture 3 LQ is moved from the lower-left position in the bitstream 501 to the lower-right position in the bitstream 502.
[0088] As Figure 5b shown, the bitstream 500 is composed of an image whose width is not a multiple of the coding tree block (CTB) size. Therefore, the rightmost coding tree block undergoes inferred splitting. The rightmost coding block in these CTBs is not encoded in the bitstream. These unencoded coding blocks are represented by the shaded part 503. In the bitstream 502, the unencoded coding block 504 is located in sub-picture 2HQ and 4 HQ The right boundary of, in the middle of the image.
[0089] Encode the size of the image in terms of pixels (or luminance samples) in the bitstream. This information allows the decoder to know the boundaries of the image and infer the inferred splitting of the rightmost CTB in the image. Then, the decoder knows the coded blocks that are not coded and is able to decode the coded blocks that make up the image.
[0090] When like subpicture 2 HQ and 4 HQ The subpictures of move from the rightmost part of the image to elsewhere in the bitstream 502, problems occur. The size of the subpicture is encoded as an integer number of CTBs in the bitstream. Due to this particularity in the standard, it is impossible for the decoder to infer the right boundary of the subpicture and the corresponding inferred splitting of the rightmost CTB of the subpicture. The decoder expects all the coded blocks of the rightmost CTB of the subpicture to be coded in the bitstream. When some of them are missing, the decoding fails.
[0091] A method for rearranging subpictures in an image without re-encoding them needs to be found.
[0092] This problem can be solved by proposing an encoding method that includes encoding information that allows the decoder to infer the splitting of CTBs located to the right or bottom of the subpicture, when the width or height of the subpicture is not a multiple of the CTB size when the subpicture is not located on the right or bottom of the image. A corresponding decoding method for the generated bitstream is also proposed.
[0093] A subpicture is a rectangular region in a picture represented by one or more slices. Subpictures are defined, for example, in SPS or PPS NAL units. The encoder can indicate for each subpicture that their boundaries are considered as image boundaries, which means that subpictures can be decoded independently of each other. Intra and inter prediction processing is constrained to use prediction information only from the same subpicture in the current frame and reference frames. The filtering of subpicture boundaries is controlled by a flag defined for each subpicture. This flag makes it possible to apply or not apply the loop filter at the boundaries of the subpicture. When decoding a CTB from a slice, the decoder determines the index or identifier of the subpicture to which the CTB belongs.
[0094] In the following, the use of sub - pictures is to allow decoding regions of a video sequence at different positions and may also move sub - pictures to new decoding positions. For this purpose, in a preferred embodiment, the sub - picture boundary is considered as a picture boundary (since it ensures that motion prediction is constrained to allow independent decoding of sub - pictures), and the loop filter is disabled at the sub - picture boundary. For VVC, when targeting sub - pictures, the flags subpic_treated_as_pic_flag[i] and loop_filter_across_subpic_enabled_flag[i] are equal to 1 and 0 respectively, which satisfy these two conditions. For HEVC, a sub - picture is, for example, a set of strips belonging to a Motion Constrained Tile Set according to the HEVC specification.
[0095] Figure 6 A method for encoding a video picture into a bitstream according to a first aspect of the present invention is shown.
[0096] In a first step 601, the image is segmented into sub - pictures.
[0097] In step 602, the size of the sub - pictures is determined. The width and height of each sub - picture are a function of the region of interest present in the input video sequence. Generally, the size of each sub - picture includes a region of interest. The size of the sub - picture is determined in pixels (or luminance samples). When the size of the sub - picture is not a multiple of the CTB size, it is determined that specific split - inference processing is required for the sub - picture. In particular, this may occur when, as Figure 5a shown when the encoder merges two bitstreams.
[0098] In step 603, the size of the sub - pictures is encoded in the bitstream. This size is information indicating the position of the split - inference boundary of the sub - picture when split - inference processing is required. If necessary, information indicating that split - inference processing is required for this sub - picture is encoded in the bitstream. The information provided in this step in the bitstream allows the decoder to perform split - inference processing.
[0099] In step 604, at least one strip constituting the sub - picture is encoded in the bitstream. When split - inference processing is required, the encoder divides the sub - picture into two parts. The first part is the region between the top and left boundaries of the sub - picture and the horizontal and vertical split - inference boundaries. The second part is the region between the split - inference boundaries and the bottom and right boundaries of the sub - picture. For example, Figure 5b the 2 HQ white area in the sub - picture corresponds to the first part, and the shaded area corresponds to the second part. Herein, we refer to the first part of the sub - picture as the useful part of the sub - picture. Coding units for which this split - inference processing falls outside the useful part of the sub - picture are not encoded.
[0100] In an alternative embodiment, coding blocks that fall outside the useful part of the sub-picture are coded with padding data.
[0101] In another alternative embodiment, a flag is inserted in the bitstream to indicate whether coding blocks that fall outside the useful part of the sub-picture are coded with padding data or not. Typically, the encoder specifies in the picture parameter set NAL unit that the sub-picture contains coded coding blocks using padding data for coding units that fall outside the useful part of the sub-picture. For example, the flag is associated with each sub-picture identifier to specify whether padding coded data is provided for coding units outside the useful part of the sub-picture.
[0102] In some embodiments, it may be known from the context of the bitstream that split inference processing is used for each slice. In these embodiments, it is not necessary to code a flag to signal the use of split inference processing for the sub-picture.
[0103] It is proposed to introduce a new syntax element in one parameter set NAL unit (e.g., in the SPS) that enables the same split inference for CTBs at the right and / or bottom boundaries of a sub-picture to be obtained when moving at different positions in the merged bitstream.
[0104] When the sub-picture boundary is considered as the picture boundary, these syntax elements will enable the decoder to determine the size of the skipped coding blocks in the last CTB row / or last CTB column of the sub-picture.
[0105] Thus, when moving the sub-picture from the rightmost position in the picture to another position, the merge or coding operation includes determining the size of the skipped coding blocks in the last CTB row and column and then specifying, for example, in the SPS associated with the sub-picture. Decoding of the merged bitstream will use these values to determine the available size of the sub-picture.
[0106] For example, the syntax of the SPS includes the following syntax elements:
[0107]
[0108]
[0109] According to the proposed embodiment, some new syntax elements are introduced that are shown in bold in the table. The semantics of these new syntax elements can be as follows.
[0110] subpic_split_inference_flag being equal to 1 indicates that the last CTB row or the last CTB column of the sub - picture is incomplete. The inference process for splitting the CTBs in the last CTB row and column of the sub - picture takes into account the available size of the sub - picture to determine the value of split_cu_flag.
[0111] subpic_split_inference_flag being equal to 0 indicates that all CTBs of the sub - picture are complete. No split inference process is required for the CTBs on the last CTB row and column of the sub - picture.
[0112] subpic_split_inference_ctb_width[i] specifies the actual width of the CTB in the right - most column CTB of the i - th sub - picture (when it exists), i.e., the width of the pixels of the CTB in the useful part of the i - th sub - picture. The value of subpic_split_inference_ctb_width[i] is specified in units of coded blocks of MinCbSizeY width and can be in the range of [0, CtbSizeY / MinCbSizeY - 1] (inclusive). When it does not exist, it is inferred that the value of subpic_split_inference_ctb_width[i] is equal to CtbSizeY / MinCbSizeY corresponding to the width of the CTB. This coded syntax element is encoded, for example, using a 7 - bit or a fixed - length code equal to log2(CtbSizeY / MinCbSizeY - 1). An Exp - Golomb code can also be used.
[0113] subpic_split_inference_ctb_height[i] specifies the actual height of the CTB in the last row CTB of the i - th sub - picture (when it exists). The value of subpic_split_inference_ctb_height[i] is specified in units of coded blocks of MinCbSizeY height and can be in the range of [0, CtbSizeY / MinCbSizeY - 1] (inclusive). When it does not exist, it is inferred that the value of subpic_split_inference_ctb_height[i] is equal to CtbSizeY / MinCbSizeY corresponding to the height of the CTB. This coded syntax element is encoded, for example, using a 7 - bit or a fixed - length code equal to log2(CtbSizeY / MinCbSizeY - 1). An Exp - Golomb code can also be used.
[0114] The subpic_split_inference_flag enables the encoder to indicate to the decoder that specific processing is required to handle the CTBs in the last CTB rows and columns of some sub-pictures. This means that the rightmost and / or bottom CTB rows contain coded blocks that have not been encoded by the encoder. These CTBs are split into blocks to be handled by the coding tree split process, which can infer the split of each block based on the values of the syntax elements subpic_split_inference_ctb_width[i] and subpic_split_inference_ctb_height[i] for the i-th sub-picture.
[0115] The actual sizes of the rightmost and bottom CTBs in the i-th sub-picture (specified by subpic_split_inference_ctb_width[i] and subpic_split_inference_ctb_height[i]) are in units of the minimum coded block size (MinCbSizeY, MinCbSizeY) specified in the SPS. The actual size of a CTB cannot exceed the maximum size of a coding tree block (CtbSizeY, CtbSizeY) as specified in the SPS. Therefore, the range of subpic_split_inference_ctb_height and subpic_split_inference_ctb_width is from 0 to the ratio of the maximum size of a CTB to the minimum size of a coded block minus one.
[0116] In an alternative, subpic_split_inference_ctb_width[i] and subpic_split_inference_ctb_height[i] are expressed in units of a number of luma samples (typically four luma samples) to simplify the parsing of the SPS. In fact, the sizes of the CTB and the minimum coded block are specified in the SPS and can be defined after the syntax elements subpic_split_inference_ctb_width[i] and subpic_split_inference_ctb_height[i]. In this case, the position of the split inference boundary is represented as an integer multiple of the minimum coded block size of a coding tree block. Additionally, when subpic_split_inference_ctb_width[i] and subpic_split_inference_ctb_height[i] are defined in different parameter set NAL units, the actual size of each sub-picture can be calculated independently of the CTB and coded block sizes.
[0117] The following pseudocode determines the boundaries in the i-th sub-picture that trigger the splitting of the coding tree. The position of the vertical boundary is represented by the variable SubPicRightSplitInferenceBoundary[i], and the position of the horizontal boundary is represented by the variable SubPicBotSplitInferenceBoundary[i].
[0118] The position of the vertical boundary in the i-th sub-picture in units of luminance samples is equal to the position of the right boundary aligned on the CTB boundary of the i-th sub-picture minus the syntax element subpic_split_inference_ctb_width[i] defined in the SPS. The position of the horizontal boundary in the i-th sub-picture in units of luminance samples is equal to the position of the bottom boundary aligned on the CTB boundary of the i-th sub-picture minus the syntax element subpic_split_inference_ctb_height[i] defined in the SPS.
[0119] The variables SubPicRightSplitInferenceBoundary[i] and SubPicBotSplitInferenceBoundary[i] are derived as follows:
[0120]
[0121] Therefore, the proposed new syntax elements subpic_split_inference_flag, subpic_split_inference_ctb_width[i], and subpic_split_inference_ctb_height[i] associated with the sub-picture allow the encoder to indicate to the decoder the information needed to be able to infer the splitting of an incomplete CTB and determine the coding blocks that fall outside the useful part of the sub-picture.
[0122] The coding tree syntax includes a flag split_cu_flag that indicates whether a coding block is split. The split is inferred when the coding tree block is not entirely within the boundaries of the useful part of the sub-picture and thus the flag is not coded.
[0123] A split_cu_flag equal to 0 specifies that the coding block is not split. A split_cu_flag equal to 1 specifies that the coding block is split into four coding blocks using a quadtree split as indicated by the syntax element split_qt_flag, or split into two coding blocks using a binary split as indicated by the syntax element mtt_split_cu_binary_flag, or split into three coding blocks using a ternary split as indicated by the syntax element mtt_split_cu_binary_flag. As indicated by the syntax element mtt_split_cu_vertical_flag, the binary or ternary split can be vertical or horizontal.
[0124] When split_cu_flag does not exist in the coding of the coding block, the value of split_cu_flag is inferred as follows:
[0125] - If one or more of the following conditions are true, the value of split_cu_flag is inferred to be equal to 1:
[0126] - x0 + cbWidth is greater than SubPicRightSplitInferenceBoundary[SubPicIdx].
[0127] - y0 + cbHeight is greater than SubPicBotSplitInferenceBoundary[SubPicIdx].
[0128] - Otherwise, the value of split_cu_flag is inferred to be equal to 0.
[0129] The above pseudocode determines whether a coding block located at the position (x0, y0) of the luma sample coordinates with a width (or height) equal to cbWidth (or cbHeight) is completely within the useful part of the current sub-picture with an index equal to SubPicIdx. When outside this area, the encoder (and decoder) infers that the coding block is split (split_cu_flag equal to 1).
[0130] When it is inferred that the coding block is split because the boundary of the coding block is outside the useful part of the sub-picture actual pixel boundary, there are three possible splits. The first is a quadtree split, i.e., the coding block is further divided into 4 equal-sized square coding blocks. The second is a split into two coding blocks separated by a vertical or horizontal boundary across the entire coding block. This split is a binary tree split. Finally, the third split is a ternary tree split that divides the coding block into three blocks separated by a vertical or horizontal boundary.
[0131] In addition to the split inference processing, the encoder (and similarly, the decoder) applies specific encoding processing to coding blocks outside the useful part of the sub-picture. In the first embodiment, these coding blocks are not encoded. For each split coding block, the encoder checks whether the y-axis (correspondingly, x-axis) coordinate of the top boundary (correspondingly, left boundary) of the block is greater than or equal to the SubPicBotSplitInferenceBoundary (correspondingly, SubPicRightSplitInferenceBoundary) value of the current sub-picture. In this case, the encoder skips the encoding of the block. Therefore, no encoded syntax elements are provided for these coding blocks. Thus, when encoding adjacent coding units, these skipped coding blocks are considered unavailable for inter-frame and intra-frame prediction. Similarly, temporal prediction is restricted to avoid using pixel information from these skipped coding blocks.
[0132] Figure 7 Illustrates the general decoding process of an embodiment of the present invention. In step 701, the decoder determines the size of the sub-pictures of the frame by parsing the SPS and PPS of the stream, typically the width and height of the sub-pictures. Specifically, the slice NAL units belonging to each sub-picture are determined. In step 702, the decoder decodes the slices of the sub-pictures that form the picture. Refer to Figure 8 Gives more details of the decoding process of the encoded CTBs in each slice. In step 703, generating an output picture from the decoded picture is included. In particular, the decoder may optionally apply some cropping operations to the decoded sub-pictures.
[0133] Figure 8 Illustrates the decoding process 702 of the CTBs encoded in a slice. In step 801, it is checked whether there is a slice to be decoded. When all slices have been decoded, the process ends. For a given slice, the decoder determines the identifier of the sub-picture to which the slice belongs in step 802.
[0134] Then, in step 803, it is checked whether there are remaining CTBs to be decoded in the current slice. For a given CTB, in step 804, the decoder parses the syntax elements of the CTB. When the CTB is on the rightmost or bottom boundary of the sub-picture with the identifier determined in 802, if the current slice requires this processing, the decoder applies the split inference processing to the CTB. Then, when the CTB is split, in step 805, the decoder decodes each coding block obtained by the inference processing and located within the sub-picture. The coding blocks located outside the useful part of the sub-picture are not decoded because they do not exist in the bitstream. In some embodiments, if the coding blocks located outside the useful part of the sub-picture have been encoded with padding data, these coding blocks are decoded.
[0135] For sub - pictures whose size is not a multiple of the CTB size, the decoding process of the sub - picture results in the right column or bottom row of an incomplete CTB. When arranging sub - pictures to form the final frame to be rendered, the decoder must determine the values of the pixels located in these pixel bands. In Figure 5b an example of the band of undetermined pixels is given, namely band 504. This is due to the fact that the sub - pictures within the frame are restricted to being composed of an integer number of CTBs. As a result, the sub - picture contains a consistent set of decoded pixels corresponding to the region between the origin of the sub - picture and the inferred split boundary. The remaining region corresponds to pixel values that are undefined or not intended to be displayed. In some embodiments, the decoder may set the values of the pixels in the remaining bands to zero. In some other embodiments, the decoder may set the values of the pixels in the remaining bands by copying the value of the nearest consistent pixel.
[0136] Since these bands of undetermined pixels are not intended to be displayed, these bands can be suppressed in the resulting frame by shifting adjacent sub - pictures. For example, Figure 5b sub - picture 1 LQ and 3 LQ in the bitstream 502 of
[0137] can be shifted left by the width of band 504. In this embodiment, the decoder shifts the decoded pixels of the sub - pictures to the right and bottom of the current sub - picture so that they are aligned with the right and bottom inferred split boundaries. The decoded positions of the encoded blocks in terms of the luminance samples take into account that some sub - pictures are incomplete.
[0138] The decoder has all the information required to identify the bands of undetermined pixels. In one embodiment, the decoder performs the shift of the sub - pictures to eliminate these bands in post - decoding processing after decoding all sub - pictures.
[0139] In another embodiment, the decoder may integrate the shift operation into the decoding process. In this embodiment, new variables are defined by the decoder and used during the decoding of each CTB to decode the encoded blocks at their exact final positions.
[0140] The SubPicOffSetX[i] and SubPicOffSetY[i] variables respectively indicate the offsets to be subtracted from the X - axis coordinate and Y - axis coordinate of the pixels of the encoded block of the i - th sub - picture, and the pixels of the encoded block of the i - th sub - picture will be placed at their correct positions.
[0141] For example, the xCtbShifted and yCtbShifted variables are the new coordinates of the top - left pixel of the CTB at the CtbAddrInRs address in the raster - scan order of the picture after the shift operation.
[0142] The following pseudocode calculates these coordinates, where CtbAddressInRs is the raster scan address of the CTB in the picture; PicWidthInCtbsY is the width of the picture in CTBs, and CtbLog2SizeY is the log2 of the size of the CTB in luminance samples; SubPicIdx is the index of the subpicture containing the current CTB.
[0143]
[0144] The first two lines calculate the x-axis and y-axis coordinates of the origin of the coding tree block in luminance samples. These values do not take into account the skipped pixels in adjacent subpictures (if any). Then, if subpic_split_inference_flag is equal to 1, which indicates that the subpicture may contain skipped blocks, the sum of the widths of the skipped pixels of the subpicture located to the left (correspondingly, above) of the coding tree block is subtracted from the xCtbShifted (correspondingly, yCtbShifted) coordinates.
[0145] The encoder determines the shift offsets of the x-axis (SubPicOffsetX[i]) and y-axis (SubPicOffsetY[i]) coordinates of each subpicture as follows. First, for a given subpicture, the set of subpictures to the left of the current subpicture is determined. For each subpicture in this set, the shift offset is equal to the difference between the right boundary of the last CTB row and the right subpicture inference boundary. Typically, this value is equal to CtbSizeY - Subpic_split_inference_ctb_width[j]*MinCbSizeY, where j is the index of the subpicture in the set, and CtbSizeY is the size of the CTB in luminance samples or pixels. The value of the shift offset of the x-axis is equal to the sum of CtbSizeY - Subpic_split_inference_ctb_width[j]*MinCbSizeY, where each j is the index of the subpicture in the set of left subpictures. Similarly, the value of the shift offset of the y-axis is equal to the sum of CtbSizeY - Subpic_split_inference_ctb_height[j]*MinCbSizeY, where each j is the index of the subpicture in the set of top subpictures above the current subpicture.
[0146] Figure 9Shows the case where the picture is divided into 16 sub - pictures, and the bands of undetermined pixels are represented by shaded bands. The arrows indicate the shift of the sub - pictures to eliminate the bands of undetermined pixels. If all vertical bands (correspondingly, horizontal bands) have the same size, it can be understood that the resulting image is rectangular. This is also the case because each column of the sub - pictures includes the same number of horizontal bands, and each row of the sub - pictures includes the same number of vertical bands.
[0147] It can be understood that if these conditions are not met, the resulting image may not be rectangular.
[0148] In one embodiment, the possibility of specifying different inference boundaries for sub - pictures is restricted to avoid defining a picture with a non - rectangular picture shape after the shift operation. In particular, after the shift or cropping of the decoded frame, the height (correspondingly, width) of the luminance samples in each column (correspondingly, row) of the sub - pictures must be equal.
[0149] The possibility of specifying different inference boundaries in a given row (or column) of a sub - picture may lead to a complex operation of shifting the decoded samples with different shift offsets at each sub - picture. For this reason, in an embodiment, a constraint on the alignment of the sub - picture inference boundaries across the image can be defined. For example, Figure 10 Shows the case where the picture is divided into 16 sub - pictures that comply with this constraint.
[0150] Figure 11 Shows the concept of the wide picture boundary.
[0151] In an alternative, the encoder determines two types of boundaries for sub - pictures. First, the sub - picture boundaries are aligned with the boundaries of other sub - pictures that jointly span the entire width or height of the picture. We call this type of sub - picture boundary the wide picture boundary. For example, Figure 11 The bottom boundary of sub - picture 0 in is the wide picture boundary. In contrast, the bottom boundary of sub - picture 1 is not picture - wide.
[0152] In an embodiment, the encoder constrains the sub-picture layout such that the last CTB row of a sub-picture having a bottom boundary that is not the picture width does not use inferred splitting. In other words, the splitting inference boundary is aligned with the CTB boundary, and for example, the size of subpic_split_inference_ctb_height[i] is equal to the size of the CTB. The same principle applies to the last CTB column of a sub-picture having a right boundary that is the picture width. In an alternative, the encoder determines the type of the boundary of each sub-picture, and when the right (correspondingly, bottom) boundary of the sub-picture is a picture-width boundary, encodes the value of subpic_split_inference_ctb_width[i] (correspondingly, subpic_split_inference_ctb_height[i]). When the right (correspondingly, bottom) boundary is not the picture width, subpic_split_inference_ctb_width[i] (correspondingly, subpic_split_inference_ctb_height[i]) is not encoded and is inferred to be equal to the size of the CTB.
[0153] The encoder may apply a second constraint to sub-pictures whose bottom (correspondingly, right) boundaries are the same picture-width boundary. For example, the constraint is that when the boundary is at the bottom of the sub-picture, for all the i-th sub-pictures of the sub-picture set, subpic_split_inference_height[i] is equal. Similarly, for all the i-th sub-pictures of the sub-picture set having a common right picture-width boundary, subpic_split_inference_width[i] may be equal.
[0154] In addition, when a sub-picture cannot be independently decoded, the encoder may apply a third constraint that the splitting inference boundary is aligned with the CTB boundary. Typically, for the i-th sub-picture, when subpic_treated_as_pic_flag[i] is equal to 0 or loop_filter_across_subpic_enabled_flag[i] is equal to 1. This constraint ensures that the splitting inference mechanism is only used in the context of sub-picture merging. Embodiments according to these constraints will now be described.
[0155] In another embodiment, the syntax of the SPS is changed to specify the positions of the vertical splitting inference boundary and the horizontal splitting inference boundary across the entire picture. In this embodiment, since the sub-picture boundaries subject to splitting inference only involve picture-width boundaries, signaling is used at the picture level.
[0156] There are several alternative ways to specify the positions of these boundaries. The positions of the split inference boundaries can be signaled in a parameter set NAL unit such as SPS or PPS. The parameter set shall describe information related to sub-picture and picture information.
[0157] In one alternative, the positions of the inference boundaries are defined, for example, relative to the origin of the picture in terms of luma samples.
[0158] For example, the PPS syntax includes the following elements:
[0159]
[0160] where:
[0161] The pps_subpic_split_inference_flag is a flag indicating the use of sub-picture split inference. Typically, split inference applies to the picture wide boundaries. The pps_subpic_split_inference_flag being equal to 1 indicates the existence of pps_split_inference_boundary_pos_x and pps_split_inference_boundary_pos_y; the pps_subpic_split_inference_flag being equal to 0 indicates the non-existence of pps_split_inference_boundary_pos_x and pps_split_inference_boundary_pos_y;
[0162] The pps_split_inference_boundary_pos_x is used to calculate the value of PpsSplitInferenceBoundaryPosX, which specifies the position of the vertical split inference boundary in terms of luma samples. The pps_split_inference_boundary_pos_x can be in the range from 1 to Ceil(pic_width_in_luma_samples÷4)-1 (inclusive). For example, this coded syntax element is coded using a 13-bit or a fixed-length code equal to log2(pic_width_in_luma_samples / 4). An Exp-Golomb code can also be used.
[0163] The position of the vertical split inference boundary PpsSplitInferenceBoundaryPosX is derived as follows:
[0164] PpsSplitInferenceBoundaryPosX = pps_split_inference_boundary_pos_x * 4
[0165] pps_split_inference_boundary_pos_y is used to calculate the value of PpsSplitInferenceBoundaryPosY, which specifies the position of the horizontal split inference boundary in terms of luma samples. pps_split_inference_boundary_pos_y can range from 1 to Ceil(pic_height_in_luma_samples÷4) - 1 (inclusive). For example, this coded syntax element is coded using a 13-bit or fixed-length code equal to log2(pic_width_in_luma_samples / 4). Exp-Golomb codes can also be used.
[0166] The position of the horizontal split inference boundary PpsSplitInferenceBoundaryPosY is derived as follows:
[0167] PpsSplitInferenceBoundaryPosY = pps_split_inference_boundary_pos_y * 4
[0168] The coding tree split inference process compares the coding unit boundary with the position of the split inference boundary. For example, when the right or bottom boundary of a coding block is greater than one of the split inference boundaries spanning the current sub-picture, the splitting of the coding unit is inferred. Thus, the split inference of slit_cu_flag is as follows, for example:
[0169] When split_cu_flag does not exist, the value of split_cu_flag is inferred as follows:
[0170] - If one or more of the following conditions are true, then the value of split_cu_flag is inferred to be equal to 1:
[0171] -x0 + cbWidth is greater than PpsSplitInferenceBoundaryPosX and (SubPicLeft[SubPicIdx]) * (subpic_grid_col_width_minus1 + 1) * 4) < PpsSplitInferenceBoundaryPosX and (SubPicLeft[SubPicIdx] + SubPicWidth[i]) * (subpic_grid_col_width_minus1 + 1) * 4) > PpsSplitInferenceBoundaryPosX
[0172] -y0 + cbHeight is greater than PpsSplitInferenceBoundaryPosY and (SubPicTop[SubPicIdx]) * (subpic_grid_row_height_minus1 + 1) * 4) < PpsSplitInferenceBoundaryPosY and (SubPicTop[SubPicIdx] + SubPicHeight[i]) * (subpic_grid_row_height_minus1 + 1) * 4) > PpsSplitInferenceBoundaryPosY
[0173] - Otherwise, infer that the value of split_cu_flag is equal to 0.
[0174] Another equivalent algorithm for determining the value of split_cu_flag is as follows:
[0175] When split_cu_flag does not exist, the value of split_cu_flag is inferred as follows:
[0176] - If one or more of the following conditions are true, then infer that the value of split_cu_flag is equal to 1:
[0177] - SubPicInferenceSplitFlag is equal to 0 and x0 + cbWidth is greater than pic_width_in_luma_samples.
[0178] - SubPicInferenceSplitFlag is equal to 0 and y0 + cbHeight is greater than pic_height_in_luma_samples.
[0179] - SubPicInferenceSplitFlag is equal to 1 and x0 + cbWidth is greater than SubPicInferenceBoundaryPosX.
[0180] - SubPicInferenceSplitFlag is equal to 1 and y0 + cbHeight is greater than SubPicInferenceBoundaryPosY.
[0181] - Otherwise, infer that the value of split_cu_flag is equal to 0.
[0182] When the sub - picture containing the coding block is crossed by the split inference boundary, SubPicInferenceSplitFlag is equal to 1. SubPicInferenceBoundaryPosX is the horizontal coordinate of the vertical split inference boundary crossing the current sub - picture. When the sub - picture is not crossed by the vertical split inference boundary, SubPicInferenceBoundaryPosX is set to a value greater than or equal to the horizontal coordinate of the right boundary of the sub - picture to avoid unnecessary split inference. Similarly, SubPicInferenceBoundaryPosY is the vertical coordinate of the horizontal split inference boundary crossing the current sub - picture (if any). Otherwise, when the sub - picture is not crossed by the horizontal split inference boundary, SubPicInferenceBoundaryPosY is set to a value greater than or equal to the horizontal coordinate of the right boundary of the sub - picture. The factor "4" is introduced because the split inference boundary is constrained to correspond to the minimum coding block boundary. Assume the size of the minimum coding block is 4×4. If the size of the minimum coding block is different, another factor can be used in these equations.
[0183] In another example, the syntax elements described above in the PPS can be defined at the SPS level, which can be advantageous when the split inference boundary does not change at each new PPS NAL unit. Thus, the syntax of the SPS includes, for example, the following elements:
[0184]
[0185] Among them, subpic_split_inference_flag, split_inference_boundary_pos_x, and split_inference_boundary_pos_y have similar semantics to pps_subpic_split_inference_flag, pps_split_inference_boundary_pos_x, and pps_split_inference_boundary_pos_y:
[0186] subpic_split_inference_flag is a flag indicating the use of sub-picture split inference. When subpic_split_inference_flag is equal to 1, it indicates the existence of split_inference_boundary_pos_x and split_inference_boundary_pos_y; when subpic_split_inference_flag is equal to 0, it indicates the non-existence of split_inference_boundary_pos_x and split_inference_boundary_pos_y;
[0187] split_inference_boundary_pos_x is used to calculate the value of SplitInferenceBoundaryPosX, which specifies the position of the vertical split inference boundary in terms of luma samples. split_inference_boundary_pos_x can be in the range from 1 to Ceil(pic_width_in_luma_samples÷4)-1 (inclusive).
[0188] The position of the vertical split inference boundary SplitInferenceBoundaryPosX is derived as follows:
[0189] SplitInferenceBoundaryPosX = split_inference_boundary_pos_x * 4
[0190] split_inference_boundary_pos_y is used to calculate the value of SplitInferenceBoundaryPosY, which specifies the position of the horizontal split inference boundary in units of luma samples. split_inference_boundary_pos_y can be in the range from 1 to Ceil(pic_height_in_luma_samples÷4)-1 (inclusive).
[0191] The position of the horizontal split inference boundary SplitInferenceBoundaryPosY is derived as follows:
[0192] SplitInferenceBoundaryPosY = split_inference_boundary_pos_y * 4;
[0193] Syntax elements split_inference_boundary_pos_x and split_inference_boundary_pos_y are encoded using, for example, a 13-bit or a fixed-length code with log2(pic_width_in_luma_samples / 4) for split_inference_boundary_pos_x and log2(pic_height_in_luma_samples / 4) for split_inference_boundary_pos_y. Exp-Golomb codes can also be used.
[0194] In another embodiment, split_inference_boundary_pos_y can be in the range from 1 to Ceil(pic_height_in_luma_samples÷4) (inclusive), and split_inference_boundary_pos_x can be in the range from 1 to Ceil(pic_width_in_luma_samples÷4) (inclusive). In this case, the maximum value of the range indicates that the split inference boundary is aligned with or outside the picture boundary. As a result, when set to the maximum value, it indicates that no vertical or horizontal split boundary is used. In a variant, two different flags are used (one flag for the vertical and one for the horizontal split inference boundary respectively). In this variant, the existence of split_inference_boundary_pos_x and split_inference_boundary_pos_y is conditional on these flag values.
[0195] In another embodiment, the PPS defines one or more horizontal and vertical split inference boundaries. The encoder specifies the number of horizontal split inference boundaries and the number of vertical split inference boundaries. These boundaries are limited to the picture width boundaries. For example, the syntax of the PPS contains the following elements:
[0196]
[0197] pps_num_ver_split_inference_boundaries specifies the number of pps_split_inference_boundaries_pos_x[i] syntax elements present in the PPS. When pps_num_ver_split_inference_boundaries is not present, it is inferred to be equal to 0.
[0198] pps_split_inference_boundaries_pos_x[i] is used to calculate the value of PpsSplitInferenceBoundaryPosX[i], which specifies the position of the i-th vertical split inference boundary in terms of luma samples. pps_split_inference_boundary_pos_x[i] can range from 1 to Ceil(pic_width_in_luma_samples÷4)-1 (inclusive). For example, this coded syntax element is coded using a 13-bit or a fixed-length code equal to log2(pic_width_in_luma_samples / 4). An Exp-Golomb code can also be used.
[0199] The position of the i-th vertical split inference boundary PpsSplitInferenceBoundaryPosX[i] is derived as follows:
[0200] PpsSplitInferenceBoundaryPosX[i] = pps_split_inference_boundary_pos_x[i] * 4
[0201] The distance between any two vertical split inference boundaries can be greater than or equal to CtbSizeY luma samples, which is the size of the CTB in terms of luma samples.
[0202] pps_num_hor_split_inference_boundaries specifies the number of the pps_split_inference_boundaries_pos_y[i] syntax elements present in the PPS. When pps_num_hor_split_inference_boundaries is not present, it is inferred to be equal to 0.
[0203] pps_split_inference_boundaries_pos_y[i] is used to calculate the value of PpsSplitInferenceBoundaryPosY[i], which specifies the position of the i-th horizontal split inference boundary in terms of luma samples. pps_split_inference_boundary_pos_y[i] can range from 1 to Ceil(pic_height_in_luma_samples÷4)-1 (inclusive). This coded syntax element is encoded, for example, using 13 bits or a fixed-length code equal to log2(pic_height_in_luma_samples / 4). Exp-Golomb codes can also be used.
[0204] The position of the i-th horizontal split inference boundary PpsSplitInferenceBoundaryPosY[i] is derived as follows:
[0205] PpsSplitInferenceBoundaryPosX[i] = pps_split_inference_boundary_pos_y[i] * 4
[0206] The distance between any two horizontal split inference boundaries can be greater than or equal to CtbSizeY luma samples, which is the size of a CTB in terms of luma samples.
[0207] In another embodiment, the encoder specifies the position of picture-wide split inference boundaries relative to the grid formed by the CTBs of the picture. Typically, the SPS or PPS includes syntax elements to indicate the number and positions of the vertical and horizontal boundaries of the split inference process as indices of CTB rows or CTB columns. For each CTB row (correspondingly, column) described in the PPS, the encoder indicates the width (correspondingly, height) of the CTB used for the inference process.
[0208] In a variant of this embodiment, the encoder and decoder may determine picture-wide sub-picture boundaries according to the sub-picture definition. Specifically, the encoder determines NumVerSplitInferenceBoundaries, which is the number of vertical picture-wide sub-picture boundaries, and NumHorSplitInferenceBoundaries, which is the number of horizontal picture-wide sub-picture boundaries. It also determines the position of the i-th (in raster scan order) vertical picture-wide sub-picture boundary on the x-axis and stores it in the PictureWideSubPictureBoundaryPosX[i] variable. It further determines the position of the i-th (in raster scan order) horizontal picture-wide sub-picture boundary on the y-axis and stores it in the PictureWideSubPictureBoundaryPosY[i] variable.
[0209] Then, the encoder encodes the width (and correspondingly, height) used for split inference for the CTB rows (and correspondingly, columns) at the left (and correspondingly, top) of the vertical (and correspondingly, horizontal) picture-wide sub-picture boundaries.
[0210] For example, the PPS includes the following syntax elements:
[0211]
[0212] Where:
[0213] pps_split_inference_ctb_width[i] is used to calculate the value of PpsSplitInferenceBoundaryPosX[i], which specifies the position of the i-th vertical split inference boundary in terms of luma samples. pps_split_inference_ctb_width[i] is specified in units of coded blocks of width MinCbSizeY and can range from 0 to CtbSizeY / MinCbSizeY - 1 (inclusive). For example, this coded syntax element is encoded using a fixed-length code equal to log2(CtbSizeY / MinCbSizeY - 1). An Exp-Golomb code can also be used.
[0214] The position (in terms of luma samples) of the i-th vertical split inference boundary PpsSplitInferenceBoundaryPosX[i] is derived as follows:
[0215] PpsSplitInferenceBoundaryPosX[i] = PictureWideSubPictureBoundaryPosX[i] - CtbSizeY + pps_split_inference_ctb_width[i] * MinCbSizeY
[0216] pps_split_inference_ctb_height[i] is used to calculate the value of PpsSplitInferenceBoundaryPosY[i], and PpsSplitInferenceBoundaryPosY[i] specifies the position of the i-th horizontal split inference boundary in terms of luma samples. pps_split_inference_ctb_height[i] is specified in units of coded blocks of MinCbSizeY height and can range from 0 to CtbSizeY / MinCbSizeY - 1 (inclusive). For example, a fixed-length code equal to log2(CtbSizeY / MinCbSizeY - 1) is used to code this coded syntax element. An Exp-Golomb code can also be used.
[0217] The position (in terms of luma samples) of the i-th horizontal split inference boundary PpsSplitInferenceBoundaryPosY[i] is derived as follows:
[0218] PpsSplitInferenceBoundaryPosY[i] = PictureWideSubPictureBoundaryPosY[i] - CtbSizeY + pps_split_inference_ctb_height[i] * MinCbSizeY
[0219] In another embodiment, the split inference boundary is derived from sub-picture partitioning (e.g., from a sub-picture grid). Specifically, the encoder may describe sub-picture boundaries that do not align with CTB boundaries. For example, the sub-picture partitioning defines the size of sub-picture grid elements that are smaller than the CTB size. When two consecutive grid elements have two different sub-picture indices and belong to the same CTB, this indicates that the first sub-picture has a right (correspondingly, bottom) boundary that does not align with the CTB boundary. The second sub-picture has a left (correspondingly, top) boundary that does not align with the CTB boundary. In this case, the inferred split boundary aligns with the right (correspondingly, bottom) boundary of the first sub-picture and thus with the left (correspondingly, top) boundary of the second sub-picture. In a variant, different signaling is used to indicate the right and left boundaries of the sub-picture, for example, by explicitly indicating the width and height of each sub-picture in units smaller than the CTB size. In this case, the sub-picture width and height are compared with their values in CTB units to determine the sub-pictures that do not align with the CTB boundary. In another variant, a specific value of the sub-picture element is reserved to indicate that the sub-picture grid element does not have encoded data. Then, the split inference boundary is derived from the sub-picture partitioning, which avoids explicit signaling of the split inference boundary in the SPS, PPS, or any parameter set or SEI.
[0220] The VVC specification defines a conformance window in the PPS as shown in the following table:
[0221]
[0222] This conformance window is a rectangular area in each picture represented by left, right, top, and bottom offsets from the picture boundary. At the end of the decoding process, the decoder applies a clipping process to remove the pixels outside the conformance window.
[0223] In the previous embodiment, we described that the encoder skips some coding blocks in the CTB that are straddled by the split inference boundary. The skipped coding blocks are located to the right of the vertical split inference boundary and below the horizontal split inference boundary. In one embodiment, the decoder will decode a CTB that has undefined pixels for these coding blocks. To this end, it is proposed to add a syntax element in the PPS that will define the conformance window for each sub-picture.
[0224] In one embodiment, the parameter set includes a new syntax element for indicating a rectangular area of luma samples in each consistent sub-picture. It is all the pixels decoded by the decoder whose values are equal to the values encoded by the encoder, i.e., corresponding to the useful part of the sub-picture.
[0225] Typically, SPS or PPS defines four consistency window offset parameters for each sub - picture of a stream with the flag subpic_treated_as_pic_flag equal to true. For example, the syntax can be as follows:
[0226]
[0227] The syntax elements subpic_conf_win_left_offset[i], subpic_conf_win_right_offset[i], subpic_conf_win_top_offset[i], subpic_conf_win_bottom_offset[i], subpic_conf_win_left_offset[i] specify the four consistency window offset parameters.
[0228] In another embodiment, for example, when the encoder constrains the split inference boundary to span the picture width and height, undefined pixels form one or more pixel bands in the picture. Thus, instead of specifying the consistency regions in each sub - picture, the PPS describes the pixel bands to be excluded from the consistency window and that should be cropped.
[0229] For example, the PPS defines the number of pixel bands to be excluded from the consistency window defined for the picture in the PPS. For each pixel band, the encoder specifies the width of the band. For example, the syntax of the PPS can include the following elements:
[0230]
[0231] With the following semantics:
[0232] conformance_exclusion_band_flag equal to 1 indicates the presence of a consistency exclusion band in the PPS. conformance_exclusion_band_flag equal to 0 indicates the absence of a consistency exclusion band in the PPS.
[0233] num_ver_conformance_exclusion_bands specifies the number of the conformance_exclusion_ver_band_pos_x[i] and conformance_exclusion_ver_band_width_minus1[i] syntax elements present in the PPS. When num_ver_conformance_exclusion_bands is absent, it is inferred to be equal to 0.
[0234] conformance_exclusion_ver_band_width_minus1[i] plus 1 specifies the width of the i-th vertical conformance exclusion band for luma samples and is used to calculate PpsConformanceVerBandPosX[i], which specifies the position of the boundary of the i-th vertical conformance exclusion band in terms of luma samples. Conformance. conformance_exclusion_ver_band_width_minus1[i] can range from 0 to CtbSizeY - 2. This coded syntax element is coded, for example, using 8 bits or a fixed-length code equal to log2(pic_width_in_luma_samples / CtbSizeY) or equal to log2(CtbSizeY - 2). Exp-Golomb codes can also be used.
[0235] conformance_exclusion_ver_band_pos_x[i] is used to calculate the value of PpsConformanceVerBandPosX[i], which specifies the position of the boundary of the i-th vertical conformance exclusion band in terms of luma samples. conformance_exclusion_ver_band_pos_x[i] can range from 0 to PicWidthInCtbsY (inclusive). In a variant, the range is 1 to PicWidthInCtbsY - 1 (inclusive) to avoid specifying vertical bands starting at the first or last pixel column of the picture. For example, this coded syntax element is coded using 8 bits or a fixed-length code equal to log2(pic_width_in_luma_samples / CtbSizeY). Exp-Golomb codes can also be used.
[0236] The position of the i-th vertical exclusion band PpsConformanceVerBandPosX[i] is derived as follows:
[0237] PpsConformanceVerBandPosX[i] = conformance_exclusion_ver_band_pos_x[i] * CtbSizeY – (conformance_exclusion_ver_band_width_minus1[i] + 1)
[0238] The distance between any two vertical conformance exclusion band boundaries can be greater than or equal to CtbSizeY luma samples, which is the size of a CTB in luma samples. In a variant, there is no restriction on the distance between two vertical conformance exclusion band boundaries, and each band can overlap with another band. In this case, the decoder must determine the overlapping bands to determine the actual conformance region of the picture.
[0239] num_hor_conformance_exclusion_bands specifies the number of conformance_exclusion_hor_band_pos_y[i] and conformance_exclusion_hor_band_height_minus1[i] syntax elements present in the PPS. When num_hor_conformance_exclusion_bands is not present, it is inferred to be equal to 0.
[0240] conformance_exclusion_hor_band_height_minus1[i] plus 1 specifies the height of the i-th horizontal conformance exclusion band in luma samples and is used to calculate PpsConformanceHorBandPosY[i], which specifies the position of the boundary of the i-th horizontal conformance exclusion band in luma samples. conformance_exclusion_hor_band_height_minus1[i] can range from 0 to CtbSizeY - 2. This coded syntax element is encoded, for example, using a 7-bit or a fixed-length code equal to log2(pic_height_in_luma_samples / CtbSizeY). An Exp-Golomb code can also be used.
[0241] conformance_exclusion_hor_band_pos_y[i] is used to calculate the value of PpsConformanceHorBandPosY[i], which specifies the position of the i-th horizontal conformance exclusion band boundary in terms of luma samples. conformance_exclusion_hor_band_pos_y[i] can range from 0 to PicHeightInCtbsY (inclusive). PicHeightInCtbsY is the height of the picture in CTBs. In a variant, the range is from 1 to PicHeightInCtbsY - 1 (inclusive) to avoid specifying horizontal bands starting at the first or last pixel row of the picture. This coded syntax element is encoded, for example, using 8 bits or a fixed-length code equal to log2(pic_height_in_luma_samples / CtbSizeY). Exp-Golomb codes can also be used.
[0242] The position of the i-th horizontal exclusion band PpsConformanceHorBandPosY[i] is derived as follows:
[0243] PpsConformanceHorBandPosY[i] = conformance_exclusion_hor_band_pos_y[i] * CtbSizeY - (conformance_exclusion_hor_band_height_minus1[i] + 1)
[0244] The distance between any two horizontal conformance exclusion band boundaries can be greater than or equal to CtbSizeY luma samples, which is the size of a CTB in luma samples. In a variant, there is no restriction on the distance between two horizontal conformance exclusion band boundaries, and the bands can overlap each other. In this case, the decoder must determine the overlapping bands to determine the actual conformance region of the picture.
[0245] In a variant, the number of horizontal and vertical exclusion bands is optional. In this case, when the num_ver_conformance_exclusion_bands and num_hor_conformance_exclusion_bands syntax elements are absent, their values are inferred to be equal to 1. The syntax of the PPS is, for example, as follows:
[0246]
[0247] In a variant, the description of horizontal or / and vertical split inference boundary information is optional. For example, the syntax of PPS (or SPS) is as follows:
[0248]
[0249] The semantics of the new syntax (bold) elements are as follows:
[0250] conformance_exclusion_ver_band_flag being equal to 1 indicates the presence of conformance_exclusion_ver_band_width_minus1 and conformance_exclusion_ver_band_pos_x in the PPS, which specify the vertical conformance exclusion band. conformance_exclusion_band_flag being equal to 0 indicates the absence of conformance_exclusion_ver_band_width_minus1 and conformance_exclusion_ver_band_pos_x in the PPS.
[0251] conformance_exclusion_ver_band_width_minus1 plus 1 specifies the width of the vertical conformance exclusion band in terms of luminance samples and is used to calculate PpsConformanceVerBandPosX. PpsConformanceVerBandPosX specifies the position of the left boundary of the vertical conformance exclusion band. conformance_exclusion_ver_band_width_minus1 can be in the range of 0 to CtbSizeY - 2.
[0252] conformance_exclusion_ver_band_pos_x is used to calculate the value of PpsConformanceVerBandPosX. conformance_exclusion_ver_band_pos_x can be in the range of 1 to PicWidthInCtbsY - 1 (inclusive).
[0253] The position PpsConformanceVerBandPosX of the left boundary of the vertical exclusion band is derived as follows:
[0254] PpsConformanceVerBandPosX = conformance_exclusion_ver_band_pos_x * CtbSizeY – (conformance_exclusion_ver_band_width_minus1 + 1)
[0255] A value of 1 for conformance_exclusion_hor_band_flag indicates the presence of conformance_exclusion_hor_band_height_minus1 and conformance_exclusion_hor_band_pos_y in the PPS, which specify a horizontal conformance exclusion band. A value of 0 for conformance_exclusion_band_flag indicates the absence of conformance_exclusion_hor_band_height_minus1 and conformance_exclusion_hor_band_pos_y in the PPS.
[0256] conformance_exclusion_hor_band_height_minus1 plus 1 specifies the height of the horizontal conformance exclusion band in terms of luma samples and is used to calculate PpsConformanceHorBandPosY. PpsConformanceHorBandPosY specifies the position of the top boundary of the horizontal conformance exclusion band. conformance_exclusion_hor_band_height_minus1 can range from 0 to CtbSizeY - 2.
[0257] conformance_exclusion_hor_band_pos_y is used to calculate the value of PpsConformanceHorBandPosY. conformance_exclusion_hor_band_pos_y can range from 1 to PicHeightInCtbsY - 1 (inclusive).
[0258] The position of the top boundary of the horizontal exclusion band, PpsConformanceHorBandPosY, is derived as follows:
[0259] PpsConformanceHorBandPosY = conformance_exclusion_hor_band_pos_y * CtbSizeY – (conformance_exclusion_hor_band_height_minus1 + 1)
[0260] In a variant, the width, height, and coordinates of the exclusion band are expressed in terms of chroma samples. The values for the luma samples are obtained by multiplying the values expressed in the chroma samples by SubWidthC and SubHeightC. SubWidthC and SubHeightC represent the horizontal and vertical sampling ratios between the luma component and the chroma components, respectively. For example, when the chroma format is 4:2:0, SubWidthC and SubHeightC are equal to 2. In another variant, the width, height, and coordinates of the exclusion band are expressed in terms of the minimum CTB size units to further reduce the length of the syntax elements.
[0261] In another embodiment, the width and height of the exclusion band are greater than the CTB. This allows two consecutive pixel exclusion bands to be merged into a single exclusion band. For example, when two adjacent sub - pictures define two sets of adjacent non - conforming pixels. This allows for a more compact description of the exclusion band.
[0262] In another embodiment, the number of vertical conformance exclusion bands and the number of horizontal conformance exclusion bands are respectively equal to the number of vertical split inference boundaries and horizontal split inference boundaries.
[0263] In another embodiment, the positions of the vertical and horizontal conformance exclusion bands are respectively equal to the positions of the vertical split inference boundaries and horizontal split inference boundaries.
[0264] In another embodiment, the width of a given vertical conformance exclusion band is equal to the number of pixels between the vertical split inference boundary at the same position and the right - most CTB boundary. The same applies to the horizontal exclusion band with a horizontal split inference boundary.
[0265] In another embodiment, the sub - picture split inference is derived from the sub - picture partitioning. The pixels excluded from the conformance window are inferred to correspond to the region between the right (correspondingly, bottom) boundary of the sub - picture that does not align with the CTB boundary and the right (correspondingly, bottom) boundary of these CTBs.
[0266] In a variant, these regions are grouped in different sub - pictures. In this case, the width or height of these sub - pictures is less than the CTB size. Thus, these non - conforming sub - pictures can be associated with a specific index. The SPS or PPS can provide a list of sub - picture indices that are non - conforming and should be excluded from the conformance window. In another embodiment, the number and size of the bands excluded from the conformance window are determined from the inferred split - boundary positions.
[0267] In another alternative, the number and size of the bands excluded from the conformance window are determined according to the inferred split - boundary positions.
[0268] The conformance cropping window contains luma samples having horizontal picture coordinates from SubWidthC * conf_win_left_offset to pic_width_in_luma_samples-(SubWidthC * conf_win_right_offset + 1) and vertical picture coordinates from SubHeightC * conf_win_top_offset to pic_height_in_luma_samples-(SubHeightC * conf_win_bottom_offset + 1) (inclusive).
[0269] Additionally, luma samples having horizontal picture coordinates from PpsConformanceVerBandPosX[i] to PpsConformanceVerBandPosX[i]+conformance_exclusion_ver_band_width_minus1[i]+1 are excluded from the conformance window, where i ranges from 0 to num_ver_conformance_exclusion_bands.
[0270] Additionally, luma samples having vertical picture coordinates from PpsConformanceHorBandPosY[i] to PpsConformanceHorBandPosY[i]+conformance_exclusion_hor_band_height_minus1[i]+1 are excluded from the conformance window, where i ranges from 0 to num_ver_conformance_exclusion_bands.
[0271] The width and height of the picture after cropping correspond to the following exported variables PicOutputWidthL and PicOutputHeightL: The width (correspondingly, height) of the picture is equal to the width (correspondingly, height) of the cropping window minus the width (correspondingly, height) of the pixel exclusion band, which corresponds to the following pseudocode:
[0272]
[0273] In an embodiment, the pictures and their sub-pictures that can be displaced in the merge operation are constrained to be of a size that is a multiple of the CTB size in two directions. Generally, when this constraint has not been complied with, it can be done by extending the input picture with padding pixels. Therefore, there is no inference boundary mechanism to be used. This constraint is signaled in the bitstream. For example, a flag can be used to indicate that the sub-pictures are subject to this constraint and can thus be freely merged, and the sub-pictures can be displaced in the resulting image at any position without any boundary problems. To ensure that the sub-pictures can be freely merged, an additional constraint is associated with the flag, i.e., the sub-pictures are independently decodable. This flag can be defined to apply to all sub-pictures in the image, or can be defined at the sub-picture level to indicate that the associated sub-pictures can be freely merged. In one embodiment, the flag indicating that the size of the image and its sub-pictures is a multiple of the CTB size is defined in the sequence parameter set and is valid for all sub-pictures of the sequence.
[0274] In some embodiments, a consistency window can be defined at the sub-picture level (e.g., in an SEI message).
[0275] In one embodiment, the syntax of the SPS is changed to include an additional flag indicating whether the encoding of the sub-pictures of the bitstream constrains the encoding tools and video sequence characteristics to allow the merge operation. When set, this flag constrains the sub-pictures to be independently decodable or not independently decodable. The syntax can be as follows:
[0276]
[0277]
[0278] For example, the semantics of the elements are as follows:
[0279] subpic_mergeable_flag equal to 1 indicates that all sub-pictures of the picture are constrained for the merge operation. subpic_mergeable_flag equal to 0 indicates that the sub-pictures may or may not be constrained.
[0280] pic_width_max_in_luma_samples specifies the maximum width (in luma samples) of each decoded picture that refers to the SPS. pic_width_max_in_luma_samples shall not be equal to 0 and shall be an integer multiple of Max(8, MinCbSizeY). When subpic_mergeable_flag is equal to 1, pic_width_max_in_luma_samples shall be an integer multiple of CtbSizeY. Therefore, when sub-pictures are constrained for merge operations (subpic_mergeable_flag equal to 1), when the original width of the picture is not a multiple of the CTB size, the encoded picture is constrained to use padding.
[0281] pic_height_max_in_luma_samples specifies the maximum height (in luma samples) of each decoded picture that refers to the SPS. pic_height_max_in_luma_samples shall not be equal to 0 and shall be an integer multiple of Max(8, MinCbSizeY). When subpic_mergeable_flag_mergeable_flag is equal to 1, pic_height_max_in_luma_samples shall be an integer multiple of CtbSizeY. Therefore, in the case where sub-pictures are constrained for merge operations (subpic_mergeable_flag equal to 1), when the original height of the picture is not a multiple of the CTB size, the encoded picture is constrained to use padding.
[0282] In this embodiment, the semantics of some elements of the PPS can be constrained to ensure that when sub-pictures are constrained for merge operations, each picture that refers to the PPS has a picture size that is a multiple of the CTB size. For example, the semantics of pic_width_in_luma_samples and pic_height_in_luma_samples are as follows:
[0283] pic_width_in_luma_samples specifies the width (in luma samples) of each decoded picture that refers to the PPS. pic_width_in_luma_samples shall not be equal to 0, shall be an integer multiple of Max(8, MinCbSizeY), and shall be less than or equal to pic_width_max_in_luma_samples. When subpic_mergeable_flag is equal to 1 (e.g., in an SPS with an identifier signaled in the PPS), pic_width_in_luma_samples shall be an integer multiple of CtbSizeY.
[0284] pic_height_in_luma_samples specifies the height of each decoded picture of the reference PPS in luma samples. pic_height_in_luma_samples shall not be equal to 0 and shall be an integer multiple of Max(8, MinCbSizeY), and shall be less than or equal to pic_height_max_in_luma_samples. When subpic_mergeable_flag is equal to 1, pic_height_in_luma_samples shall be an integer multiple of CtbSizeY.
[0285] This first set of constraints on the picture ensures that when subpic_mergeable_flag is equal to 1, for sub-pictures in the pictures of the merge stream, any merge position is acceptable. The encoder may have to use padding data to make the width and height of the picture multiples of the CTB size. The encoder may signal the region with padding data by defining a consistency window in the bitstream. Typically, a consistency window for each sub-picture can be defined in an SEI message.
[0286] In another embodiment, when sub-pictures are constrained for merge operations (subpic_mergeable_flag is equal to 1), the merge flag also constrains the temporal prediction mechanism within each sub-picture of all sub-pictures described in the parameter set NAL unit.
[0287] For example, subpic_treated_as_pic_flag[i] being equal to 1 specifies that the i-th sub-picture of each encoded picture in the CLVS is treated as a picture in the decoding process other than the in-loop filtering operation. subpic_treated_as_pic_flag[i] being equal to 0 specifies that the i-th sub-picture of each encoded picture in the CLVS is not treated as a picture in the decoding process other than the in-loop filtering operation. When not present, it is inferred that the value of subpic_treated_as_pic_flag[i] is equal to subpic_mergeable_flag. Therefore, when sub-pictures are constrained for merge operations (subpic_mergeable_flag is equal to 1), the subpic_treated_as_pic_flag[i] syntax element does not exist in the parameter set NAL unit and is inferred to be equal to 1, which indicates that the temporal prediction of the sub-picture is constrained to each sub-picture boundary. Otherwise, the temporal prediction may or may not be constrained.
[0288] In another embodiment, when sub-pictures are constrained for merge operation (subpic_mergeable_flag equals 1), the loop filter mechanism is disabled across sub-picture boundaries for all sub-pictures described in the parameter set NAL unit. For example, loop_filter_across_subpic_enabled_flag[i] equals 1 specifies that in-loop filtering operations can be performed across the boundaries of the i-th sub-picture in each encoded picture in the CLVS. loop_filter_across_subpic_enabled_flag[i] equals 0 specifies that in-loop filtering operations are not performed across the boundaries of the i-th sub-picture in each encoded picture in the CLVS. When not present, it is inferred that the value of loop_filter_across_subpic_enabled_pic_flag[i] equals!subpic_mergeable_flag. Thus, when sub-pictures are constrained for merge operation (subpic_mergeable_flag equals 1), the loop_filter_across_subpic_enabled_pic_flag[i] syntax element does not exist in the parameter set NAL unit and is inferred to be equal to 0, which indicates that the loop filter is constrained not to be enabled across sub-picture boundaries. Otherwise, the loop filter may or may not be constrained.
[0289] In some embodiments, spatial access may be provided only at some time intervals. The proposed flag may be defined accordingly in the PPS or even in the picture header. The picture header is a non-VCL NAL unit that defines certain syntax elements defined at the picture level and applicable to all slices of the picture.
[0290] In another embodiment, the signaling of sub-pictures associates the identifier of a sub-picture with a sub-picture index. This identifier is unique for a given sub-picture and simplifies the merge operation because when a sub-picture is moved to a new position after a merge operation, it allows avoiding rewriting the sub-picture index in the slice header. For this reason, when sub-pictures are constrained for merge operation, one of the non-VCL NAL units may signal the sub-picture identifier in the bitstream.
[0291] Generally, the sub-picture identifier may be present in the SPS, PPS, or picture header. Specifically, sps_subpic_id_present_flag is a flag in the SPS that, when equal to 1, indicates the presence of a sub-picture identifier in the bitstream. The semantics of this syntax element are as follows, for example:
[0292] The sps_subpic_id_present_flag being equal to 1 indicates the presence of sub-picture ID mapping in the SPS. The sps_subpic_id_present_flag being equal to 0 indicates the absence of sub-picture ID mapping in the SPS. When subpic_mergeable_flag is equal to 1, the sps_subpic_id_present_flag must be equal to 1. As a result, when sub-pictures are constrained for merge operations, signaling of sub-picture identifiers is provided in the bitstream. In a variant, when subpic_mergeable_flag is equal to 1, the sps_subpic_id_present_flag does not exist in the SPS and is inferred to be equal to 1. The corresponding syntax of the SPS may include the following syntax elements:
[0293]
[0294]
[0295] In another embodiment, the presence of sub-picture identifiers in the picture header may complicate the merge operation. In fact, when signaling identifiers in the picture header, it overrides the mapping of sub-picture identifiers performed in the parameter set NAL unit. Since different picture headers are sent for each frame, there is a possibility of mapping changes at each picture. As a result, a typical merge operation needs to check whether the picture header modifies the sub-picture identifier mapping. To avoid this check operation, when sub-pictures are constrained for merge operations, the mapping of identifiers in the picture header is disabled. The ph_subpic_id_signalling_present_flag syntax element is used to control the presence of sub-picture identifier mapping in the picture header. In this embodiment, the semantics of ph_subpic_id_signalling_present_flag are as follows:
[0296] The ph_subpic_id_signalling_present_flag being equal to 1 indicates signaling of sub-picture ID mapping in the PH. The ph_subpic_id_signalling_present_flag being equal to 0 indicates no signaling of sub-picture ID mapping in the PH. When subpic_mergeable_flag is equal to 1, the ph_subpic_id_signalling_present_flag must be equal to 0. In a variant, when subpic_mergeable_flag is equal to 1, the ph_subpic_id_signalling_present_flag does not exist in the picture header and is inferred to be equal to 0.
[0297] PPS can also signal the mapping of sub - picture identifiers. For the picture header, it is necessary to check whether the mapping between two PPS NAL units has changed. This additional check increases the complexity of the merge operation. Thus, in one embodiment, when a sub - picture is constrained for the merge operation, the sub - picture is constrained to be the same in each and every PPS where subpic_mergeable_flag equals 1. The pps_subpic_id[i] syntax element of the PPS specifies the sub - picture ID (or identifier) of the i - th sub - picture corresponding to the mapping of the sub - picture identifier to the sub - picture index. When subpic_mergeable_flag equals 1, all PPSs referring to the same SPS should have the same value of pps_subpic_id[i], where i ranges from 0 to pps_num_subpics_minus1 (inclusive).
[0298] In another embodiment, the information indicating that a sub - picture is constrained for the merge operation corresponds to a specific profile of the VVC specification. Typically, a specific value of the general_profile_idc syntax element (e.g., 3) indicates that the output layer conforms to the profile for sub - picture merge operation. In a variant, the information indicating that a sub - picture is constrained for the merge operation corresponds to a sub - profile. In this case, a specific value of the general_sub_profile_idc[i] (e.g., 3) indicates that the bitstream is constrained for sub - picture merge operation.
[0299] The general constrained information structure of VVC is set by flags described in the profile, tier, and level information, which enables disabling one or more encoding tools. In one embodiment, the general constrained information structure includes any of the following syntax elements:
[0300] - The no_dependent_subpicture_flag syntax element, when equal to 1, specifies that for any value of i, subpic_treated_as_pic_flag[i] should be equal to 1. no_dependent_subpicture_flag being equal to 0 does not impose such a constraint. In a variant, when equal to 1, it specifies that for any value of i, subpic_treated_as_pic_flag[i] and loop_filter_across_subpic_enabled_flag[i] should be equal to 1.
[0301] - The no_picture_header_subpicture_id_mapping syntax element, when equal to 1, specifies that for all picture header NAL units, the ph_subpic_id_signalling_present_flag shall be equal to 0. When no_picture_header_subpicture_id_mapping is equal to 0, no such constraint is imposed.
[0302] - The no_pps_subpicture_id_mapping_change syntax element, when equal to 1, specifies that for all PPS NAL units, for any i value, the pps_subpic_id[i] values shall be equal. When no_pps_subpicture_id_mapping_change is equal to 0, no such constraint is imposed.
[0303] In an embodiment, an SEI message is proposed to handle in-picture regions that require post-decoding cropping operations. This may result in, for example, bitstream extraction and merging operations, where some sub-pictures initially located at the right or bottom boundary of the picture contain padding data. The SEI indicates to the decoder a consistency window for at least some of the sub-pictures, where the sub-pictures define a picture region that contains padding or even unreliable or useless data that the content creator deems should be removed from the image to be displayed.
[0304] The SEI may provide the following syntax, for example:
[0305]
[0306]
[0307] Having the following semantics:
[0308] subpic_conf_win_cancel_flag equal to 1 indicates that the SEI message cancels the persistence of any previous sub-picture consistency window SEI message applied to the current layer in output order. subpic_conf_win_cancel_flag equal to 0 indicates that sub-picture consistency window information follows.
[0309] subpic_conf_win_num_subpics_minus1 plus 1 specifies the number of sub-picture consistency windows present in the SEI message. This value is a function of the number of sub-pictures present in the image. Typically, it is required that the value of subpic_conf_win_num_subpics_minus1 shall be equal to sps_num_subpics_minus1 to allow the definition of a consistency window for each sub-picture.
[0310] subpic_conf_win_left_offset[i], subpic_conf_win_right_offset[i], subpic_conf_win_top_offset[i], and subpic_conf_win_bottom_offset[i] specify the samples of the i-th sub-picture in the picture in the CLV output from the decoding process (with respect to the rectangular region specified in the output picture coordinates relative to the origin of the i-th sub-picture as described in the SPS NAL unit).
[0311] The sub-picture consistency cropping window for the i-th sub-picture contains luma samples with horizontal picture coordinates from SubPictureLuma_X[i] + SubWidthC * conf_win_left_offset to SubPictureLuma_X[i] + SubPictureLuma_Width[i] - (SubWidthC * subpic_conf_win_right_offset[i] + 1) and vertical picture coordinates from SubPictureLuma_Y[i] + SubHeightC * subpic_conf_win_top_offset[i] to SubPictureLuma_Y[i] + SubPictureLuma_Height[i] - (SubHeightC * subpic_conf_win_bottom_offset[i] + 1) (inclusive). Where SubPictureLuma_X[i] and SubPictureLuma_Y[i] specify the horizontal and vertical picture coordinates of the first pixel in the i-th sub-picture, and SubPictureLuma_Width[i] and SubPictureLuma_Height[i] specify the width and height of the i-th sub-picture described in the SPS in terms of luma samples. For example, these variables are calculated as follows:
[0312] SubPictureLuma_X[i] = subpic_ctu_top_left_x[i] * CtbSizeY
[0313] SubPictureLuma_Y[i] = subpic_ctu_top_left_y[i] * CtbSizeY
[0314] SubPictureLuma_Width[i] = (subpic_width_minus1[i] + 1) * CtbSizeY
[0315] SubPictureLuma_Height[i] = (subpic_height_minus1[i] + 1) * CtbSizeY
[0316] In some cases, only a subset of sub - pictures (typically, sub - pictures with padding data) need a consistency window. In such cases, a new syntax element indicates for each sub - picture whether the consistency window is signalled. For example, a "for" loop over each sub - picture specifies the subpic_conf_win_signalled_flag[i] syntax element for the i - th sub - picture described in the sub - picture. When equal to 1, there is an offset parameter and a consistency window is specified for the sub - picture. Otherwise, subpic_conf_win_signalled_flag[i] is equal to 0, no consistency is signalled for the i - th sub - picture, and the offset parameter does not exist and is inferred to be equal to 0. In a variant, the number of sub - picture consistency windows described in the SEI message is different from the number of sub - pictures in the picture, and for each signalled sub - picture consistency window signalled in the SEI, a list of one or more sub - picture indices in the picture is associated with the index of the sub - picture consistency window. This index list indicates the sub - pictures that use the sub - picture consistency window. For example, the "for" loop in the SEI message indicates the subpic_conf_win_num_subpics_minus1[i] syntax element, which is the number of sub - picture indices minus 1 associated with the i - th sub - picture consistency window in the SEI message. Then, a processing loop for j in the range from 0 to subpic_conf_win_num_subpics_minus1[i] (inclusive) defines the subpic_conf_win_subpic_index[i][j] syntax element. subpic_conf_win_subpic_index[i][j] specifies the j - th index of the sub - picture that uses the i - th sub - picture consistency window. In a variant, the SEI message can define sub - picture identifiers instead of sub - picture indices. In this case, the bit - length of the sub - picture identifier is optionally described in the SEI message.
[0317] Figure 12 is a schematic block diagram of a computing device 120 for implementing one or more embodiments of the present invention. The computing device 120 can be a device such as a microcomputer, a workstation, or a lightweight portable device.
[0318] The computing device 120 includes a communication bus that is connected to:
[0319] - a central processing unit 121 labeled CPU, such as a microprocessor;
[0320] - A random access memory 122 labeled as RAM, which is used to store the executable code of the method according to the embodiments of the present invention and registers suitable for recording variables and parameters necessary for implementing the method according to the embodiments of the present invention. The memory capacity of this random access memory can be expanded, for example, by an optional RAM connected to an expansion port;
[0321] - A read-only memory 123 labeled as ROM, which is used to store computer programs for implementing the embodiments of the present invention;
[0322] - A network interface 124, which is usually connected to a communication network, through which digital data to be processed is sent or received. The network interface 124 can be a single network interface or composed of a collection of different network interfaces (for example, wired and wireless interfaces, or different types of wired or wireless interfaces). Under the control of a software application running in the CPU 121, data packets are written to the network interface for sending, or read from the network interface for receiving;
[0323] - A user interface 125, which can be used to receive input from a user or display information to the user;
[0324] - A hard disk 126 labeled as HD, which can be set as a large storage device;
[0325] - An I / O module 127, which can be used to receive / send data from / to external devices (such as a video source or a display, etc.).
[0326] The executable code can be stored either in the read-only memory 123, on the hard disk 126, or on a removable digital medium (for example, a disk). According to a variant, the executable code of the program can be received via the network interface 124 by means of a communication network so that the executable code of the program is stored in one of the storage components (such as the hard disk 126, etc.) of the communication device 120 before being executed.
[0327] The central processing unit 121 is adapted to control and direct the execution of instructions or parts of software code of one or more programs according to the embodiments of the present invention, and these instructions are stored in one of the aforementioned storage components. After power-on, the CPU 121 can execute these instructions after, for example, loading instructions related to the software application from the program ROM 123 or the hard disk (HD) 126 into the main RAM memory 122. When such a software application is executed by the CPU 121, the steps of the flowchart of the present invention are carried out.
[0328] Any step of the algorithm of the present invention can be implemented in software by executing instructions or a set of programs by a programmable computing machine such as a PC ("personal computer"), DSP ("digital signal processor") or microcontroller, etc.; or implemented in hardware by a machine or a dedicated component such as an FPGA ("field programmable gate array") or an ASIC ("application specific integrated circuit"), etc.
[0329] Although the present invention has been described above with reference to specific embodiments, the present invention is not limited to the specific embodiments, and modifications within the scope of the present invention will be apparent to those skilled in the art.
[0330] When referring to the foregoing illustrative embodiments, those skilled in the art will think of many further modifications and variations. These embodiments are given only as examples and are not intended to limit the scope of the present invention, which is determined only by the appended claims. In particular, where appropriate, different features from different embodiments may be interchanged.
[0331] The various embodiments of the present invention described above can be implemented separately or implemented as a combination of multiple embodiments. In addition, features from different embodiments can be combined when necessary, or combinations of elements or features from separate embodiments can be combined when beneficial in a single embodiment.
[0332] Each feature disclosed in this specification (including any appended claims, abstract and drawings) can be replaced by alternative features for the same, equivalent or similar purposes, unless otherwise expressly specified. Thus, unless otherwise expressly specified, each disclosed feature is only one example of a general series of equivalent or similar features.
[0333] In the claims, the word "comprising" does not exclude other elements or steps, and the indefinite article "a" or "an" does not exclude a plurality. The mere fact that different features are defined in mutually different dependent claims does not indicate that a combination of these features cannot be used advantageously.
Claims
1. A method for encoding video data including pictures into a bitstream, where the pictures are segmented into sub-pictures, the method comprising: Encode a first flag in a sequence parameter set of the bitstream, wherein the existence of information of a sub - picture depends on the first flag; In the case where the first flag is set to 1, encode a second flag in the sequence parameter set of the bitstream, wherein, in the case where the second flag is set to 1, the second flag indicates that all sub - picture boundaries related to the sequence parameter set are regarded as picture boundaries and loop filtering across sub - picture boundaries is not performed; In the case where the second flag is set to 0, encode a third flag in the bitstream that indicates whether a loop filter across a sub - picture boundary is enabled, wherein in the case where the second flag is set to 1, the third flag is not encoded in the bitstream; And Encode coding tree blocks constituting the picture into the bitstream.
2. The method according to claim 1, wherein, The method includes: In the case where the second flag is set to 0, encode the third flag and a fourth flag in the bitstream for each sub - picture, the fourth flag indicating whether the sub - picture has a boundary that is constrained to be a picture boundary.
3. The method according to claim 1, wherein, The second flag is related to all sub - pictures.
4. The method according to claim 1, wherein, The second flag is defined in the sequence parameter set, which is a syntax structure containing syntax elements of pictures applied to the bitstream.
5. The method according to claim 1, wherein, Sub - pictures are further identified using sub - picture identifiers.
6. The method according to claim 5, wherein, It is prohibited to define the sub - picture identifier in a picture header.
7. The method according to claim 5, wherein, The sub - picture identifier defined in a picture parameter set must be the same in all picture parameter sets.
8. The method according to claim 1, wherein, The second flag is associated with a specific profile.
9. The method according to claim 1, wherein, The method further includes: for at least one sub - picture, Encode information indicating a consistency window of the sub - picture in the bitstream.
10. A method for decoding a bitstream of video data including pictures, where the pictures are segmented into sub-pictures, the method comprising: Decode the first flag from the sequence parameter set of the bitstream, wherein the existence of information of a sub - picture depends on the first flag; In the case where the first flag is set to 1, decode the second flag from the sequence parameter set of the bitstream, wherein, in the case where the second flag is set to 1, the second flag indicates that all sub - picture boundaries related to the sequence parameter set are regarded as picture boundaries and loop filtering across sub - picture boundaries is not performed; In the case where the second flag is set to 0, decode the third flag from the bitstream that indicates whether a loop filter across a sub - picture boundary is enabled, wherein in the case where the second flag is set to 1, the third flag is not decoded from the bitstream; And Decode the bitstream based at least on the first flag.
11. The method according to claim 10, wherein, The method includes: In the case where the second flag is set to 0, decode the third flag and the fourth flag from the bitstream for each sub - picture, the fourth flag indicating whether the sub - picture has a boundary that is constrained to be a picture boundary.
12. The method according to claim 10, wherein, The second flag is related to all sub - pictures.
13. The method according to claim 10, wherein, The second flag is defined in the sequence parameter set, which is a syntax structure containing syntax elements of pictures applied to the bitstream.
14. A computer program product for a programmable device, the computer program product comprising a sequence of instructions which, when loaded into and executed by the programmable device, are used to implement the method according to any one of claims 1 to 13.
15. A non-transitory computer-readable storage medium storing computer program instructions which, when executed by a processor, are used to implement the method according to any one of claims 1 to 13.
16. An apparatus for encoding video data including pictures into a bitstream, the pictures being segmented into sub-pictures, the apparatus comprising a processor configured to: encode a first flag in a sequence parameter set of the bitstream, wherein the presence of information of the sub-pictures depends on the first flag; in a case where the first flag is set to 1, encode a second flag in the sequence parameter set of the bitstream, wherein, When the second flag is set to 1, the second flag indicates that all sub - picture boundaries related to the sequence parameter set are regarded as picture boundaries and loop filtering across sub - picture boundaries is not performed; When the second flag is set to 0, a third flag indicating whether the loop filter across sub - picture boundaries is enabled is coded in the bitstream, where when the second flag is set to 1, the third flag is not coded in the bitstream; And Encode the coding tree blocks constituting the picture into the bitstream.
17. An apparatus for decoding a bitstream of video data including pictures, the pictures being segmented into sub-pictures, the apparatus comprising a processor configured to: decode a first flag from a sequence parameter set of the bitstream, wherein the presence of information of the sub-pictures depends on the first flag; When the first flag is set to 1, decode the second flag from the sequence parameter set of the bitstream, where When the second flag is set to 1, the second flag indicates that all sub - picture boundaries related to the sequence parameter set are regarded as picture boundaries and loop filtering across sub - picture boundaries is not performed; When the second flag is set to 0, decode the third flag indicating whether the loop filter across sub - picture boundaries is enabled from the bitstream, where when the second flag is set to 1, the third flag is not decoded from the bitstream; And decode the bitstream at least based on the first flag.
Citation Information
Patent Citations
Tile alignment signaling and conformance constraints
CN105359525A
Concept for picture / video data streams allowing efficient reducibility or efficient random access
CN109076247A