Method and apparatus for encoding video data in a bitstream, method and apparatus for decoding a bitstream of video data, computer program product and storage medium

By encoding the information of the sub-picture in the video bitstream, the decoding failure problem caused by the sub-picture size not corresponding to the multiple of the encoded tree block is solved, and the independent decoding of the sub-picture and the effective transmission of video data are realized.

CN120499401APending Publication Date: 2025-08-15CANON KK
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510832531.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2019-12-17
Filing Date
2020-09-07
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

In a video bitstream, when the size of the sub-picture does not correspond to multiples of the encoded tree block, the decoder cannot correctly infer the split of the rightmost or bottommost CTB, resulting in the decoding failure.

Method used

By encoding information indicating sub-pictures in the bitstream, including their size and consistency windows, the decoder allows the decoder to infer the splitting of the sub-pictures, ensuring that the sub-pictures can be decoded independently.

Benefits of technology

The decoding failure problem caused by the sub-image size does not correspond to the multiple of the encoded tree block is solved, and the correct decoding of the sub-image and the effective transmission of video data is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120499401A_ABST
    Figure CN120499401A_ABST
Patent Text Reader

Abstract

The present invention relates to a method and apparatus for encoding video data in a bitstream, a method and apparatus for decoding a bitstream of video data, a computer program product and a storage medium. The invention also relates to an encoding method comprising encoding information that allows a decoder to infer a splitting of a CTB located on the right side or the lower portion of a sub-picture whose width or height is not a multiple of the size of the CTB when the sub-picture is not located on the right side or the lower portion of the image. A corresponding decoding method for the generated bitstream is also presented.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] (This application is a divisional application of an application filed on September 7, 2020, with application number 2020800652193, and entitled “Method and Apparatus for Encoding and Decoding Video Streams Using Sub-Pictures.”) Technical Field

[0002] The present disclosure relates to a method and apparatus for encoding and decoding a video bitstream that facilitates displacement of sub-pictures. More particularly, the present invention relates to encoding and decoding a video bitstream resulting from merging sub-pictures from different video bitstreams. Background Art

[0003] The size of a picture in a video bitstream may not correspond to a multiple of the size of the coding tree block (CTB) used in the encoding process. The CTB can be recursively split during encoding, particularly to find a coding block size that optimizes the encoding process. Thus, the CTB can be split up to a minimum coding block size. When the size of a picture is not a multiple of the size of the CTB, the right or bottom boundary of the picture spans the rightmost or bottommost CTB. In this case, an inferred split of the CTB is provided. Coding blocks that fall outside the picture are generally not encoded.

[0004] When considering partitioning an image into sub-pictures, the rightmost or bottommost sub-picture may contain an incomplete CTB that is subject to inferred splitting, the incomplete CTB including some uncoded coding blocks.

[0005] The decoder is provided with the size of the image in pixels and the size of the CTB in the bitstream. Thus, the decoder can determine the exact location of the right and bottom boundaries of the image and perform inferred splitting of the rightmost and bottommost CTBs in the image.

[0006] When considering the displacement of a sub-picture at the right or bottom boundary of a picture that is located elsewhere in the picture, since the size of the sub-picture is provided as an integer number of CTBs, the decoder cannot infer the splitting of the right-most or bottom-most CTBs and identify the missing coded blocks. Decoding fails.

[0007] The present invention aims to solve one or more of the above-mentioned problems. A coding method is proposed, which includes encoding information that allows a decoder to infer the splitting of CTBs located at the right or bottom of a sub-picture, when the sub-picture is not located at the right or bottom of the image and its width or height is not a multiple of the CTB size. A corresponding decoding method for the generated bitstream is also proposed. Summary of the Invention

[0008] According to one aspect of the present invention, a method for encoding video data including a picture into a bitstream is provided, wherein the picture is divided into sub-pictures, the method comprising: for at least one sub-picture, encoding the following information in the bitstream, the information indicating that the picture has a size that is a multiple of the size of a coding tree block and that the sub-picture is independently decodable; and encoding the coding tree blocks constituting the sub-picture into the bitstream.

[0009] In an embodiment, the information is related to all sub-pictures.

[0010] In an embodiment, said information is defined in a sequence parameter set, which is a syntax structure containing syntax elements that apply to pictures of said bitstream.

[0011] In an embodiment, the sub-picture is further identified using a sub-picture identifier.

[0012] In an embodiment, it is forbidden to define the sub-picture identifier in the picture header.

[0013] In an embodiment, the sub-picture identifier defined in a picture parameter set must be the same in all picture parameter sets.

[0014] In an embodiment, the information is associated with a particular profile.

[0015] According to another aspect of the present invention, a method for encoding video data including a picture into a bitstream is provided, wherein the picture is divided into sub-pictures, the method comprising: encoding, for at least one sub-picture, information indicating a consistency window of the sub-picture in the bitstream; and encoding a coding tree block constituting the sub-picture into the bitstream.

[0016] In an embodiment, said information indicating the consistency window is defined in a SEI message.

[0017] In an embodiment, the information defines a left offset, a right offset, an upper offset and a lower offset for the sub-picture.

[0018] According to another aspect of the present invention, a method for encoding video data including a picture into a bitstream is provided, wherein the picture is divided into sub-pictures, the method comprising: encoding first information in the bitstream for at least one sub-picture, the first information indicating that the sub-picture has a size that is a multiple of the size of a coding tree block and that the sub-picture is independently decodable; encoding second information in the bitstream indicating a consistency window of the sub-picture; and encoding the coding tree blocks constituting the sub-picture into the bitstream.

[0019] According to another aspect of the present invention, a computer program product for a programmable device is provided, the computer program product comprising a sequence of instructions for implementing the method according to the present invention when the sequence of instructions is loaded into and executed by the programmable device.

[0020] According to another aspect of the present invention, a computer-readable storage medium is provided, which stores instructions of a computer program for implementing the method according to the present invention.

[0021] According to another aspect of the invention, a computer program is provided which, when executed, causes the method according to the invention to be performed.

[0022] According to another aspect of the present invention, a device for encoding video data including a picture into a bitstream is provided, wherein the picture is divided into sub-pictures, the device including a processor configured to perform, for at least one sub-picture: encoding the following information in the bitstream, the information indicating that the sub-picture has a size that is a multiple of the size of a coding tree block and that the sub-picture is independently decodable; and encoding the coding tree blocks constituting the sub-picture into the bitstream.

[0023] According to another aspect of the present invention, a device for encoding video data including a picture into a bitstream is provided, wherein the picture is divided into sub-pictures, and the device includes a processor configured to perform, for at least one sub-picture: encoding information indicating a consistency window of the sub-picture in the bitstream; and encoding a coding tree block constituting the sub-picture into the bitstream.

[0024] According to another aspect of the present invention, a device for encoding video data including a picture into a bitstream is provided, wherein the picture is divided into sub-pictures, and the device includes a processor configured to perform, for at least one sub-picture: encoding first information in the bitstream, the first information indicating that the sub-picture has a size that is a multiple of the size of a coding tree block and that the sub-picture is independently decodable; encoding second information in the bitstream indicating a consistency window of the sub-picture; and encoding the coding tree blocks constituting the sub-picture into the bitstream.

[0025] At least part of the method according to the present invention may be computer-implemented. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment (including firmware, resident software, microcode, etc.), or an embodiment combining software and hardware aspects, which may all be generally referred to herein as "circuits," "modules," or "systems." Furthermore, the present invention may take the form of a computer program product embodied in any tangible medium of expression having computer-usable program code embodied in the medium.

[0026] Since the present invention can be implemented in software, the present invention can be embodied as computer-readable code for providing to a programmable device on any suitable carrier medium. Tangible, non-transitory carrier media can include storage media such as floppy disks, CD-ROMs, hard drives, magnetic tape devices, or solid-state memory devices. Transient carrier media can include signals such as electric, electronic, optical, acoustic, magnetic, or electromagnetic signals (e.g., microwave or RF signals). BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Embodiments of the present invention will now be described, by way of example only, with reference to the following drawings, in which:

[0028] Figure 1a and 1b Two different application examples for combining regions of interest are shown;

[0029] Figure 2a Some of the partitions in the coding system are shown;

[0030] Figure 2b An example of partitioning a picture into sub-pictures is shown;

[0031] Figure 2c An example of partitioning a picture into titles and bricks is shown; Figure 3 shows the organization of the bitstream in an exemplary coding system VVC;

[0032] Figure 4 Schematically illustrates the quadtree inference mechanism used in VVC to encode treeblocks that cross image boundaries;

[0033] Figure 5a and 5b shows the creation of a bitstream where a boundary sub-picture is moved to a non-border position;

[0034] Figure 6 A method of encoding a video image into a bitstream according to a first aspect of the present invention is shown;

[0035] Figure 7 shows the general decoding process of an embodiment of the present invention;

[0036] Figure 8 shows the decoding process of CTBs coded in a slice;

[0037] Figure 9 An example of a picture being divided into 16 sub-pictures is shown;

[0038] Figure 10shows an example of a picture being partitioned into 16 sub-pictures with picture-wide splitting inference boundaries;

[0039] Figure 11 The concept of picture-wide boundary is shown;

[0040] Figure 12 is a schematic block diagram of a computing device for implementing one or more embodiments of the present invention. DETAILED DESCRIPTION

[0041] Figure 1a and 1b Two different application examples for combining regions of interest are shown.

[0042] For example, Figure 1a An example is shown of a picture (or frame) 100 from a first video bitstream and a picture 101 from a second video bitstream being merged into a resulting picture 102 of the bitstream. Each picture consists of four regions of interest, numbered 1 to 4. Picture 100 has been encoded using coding parameters that result in high-quality encoding. Picture 101 has been encoded using coding parameters that result in low-quality encoding. As is well known, pictures encoded at low quality are associated with a lower bitrate than pictures encoded at high quality. The resulting picture 102 combines regions of interest 1, 2, and 4 from picture 101 (thus encoded at low quality) with region of interest 3 from picture 100 (encoded at high quality). The goal of this combination is generally to obtain a high-quality region of interest, here region 3, while keeping the resulting bitrate reasonable by encoding regions 1, 2, and 4 at low quality. This scenario may particularly occur in the context of omnidirectional content, allowing the actually visible content to have higher quality while the rest has lower quality.

[0043] Figure 1b A second example of four different videos A, B, C, and D being merged to form a resulting video is shown. Picture 103 of video A consists of regions of interest A1, A2, A3, and A4. Picture 104 of video B consists of regions of interest B1, B2, B3, and B4. Picture 105 of video C consists of regions of interest C1, C2, C3, and C4. Picture 106 of video D consists of regions of interest D1, D2, D3, and D4. Picture 107 of the resulting video consists of regions B4, A3, C3, and D1. In this example, the resulting video is a mosaic video of the different regions of interest of each original video stream. The regions of interest of the original video streams are rearranged and combined in the new positions of the resulting video stream.

[0044] The compression of video relies on block-based video coding in most coding systems such as the HEVC (standing for High Efficiency Video Coding) or the emerging VVC (standing for Versatile Video Coding) standards. In these coding systems, a video consists of a sequence of frames or pictures or images or samples that can be displayed at several different times. In the case of multi-layer video (e.g., scalable, stereo, 3D video), several pictures can be decoded to compose the resulting image to be displayed at a certain moment. A picture can also be composed of different image components. For example, to encode brightness, chrominance or depth information.

[0045] Compression of video sequences relies on several partitioning techniques for each picture. Figure 2a Some partitions in the coding system are shown. Pictures 201 and 202 are divided into coding tree units (CTUs) shown by dashed lines. A CTU is the basic unit for encoding and decoding. For example, a CTU can encode an area of 128×128 pixels.

[0046] A coding tree unit (CTU) can also be called a block, macroblock, or coding block. Different image components can be coded simultaneously, or they can be limited to just one image component. When an image contains several components, a CTU corresponds to a CTB for each component. Hereinafter, the present invention applies to both the CTU and CTB levels.

[0047] like Figure 2a As shown, a picture can be partitioned according to a grid of tiles, shown by thin solid lines. A tile is a portion of a picture and is therefore a rectangular region of pixels that can be defined independently of CTU partitioning. The boundaries of a tile and a CTU can be different. As in the example shown, a tile can also correspond to a sequence of CTUs, meaning that the boundaries of a tile and a CTU coincide.

[0048] The slice definition specifies that slice boundaries break spatial coding dependencies, which means that the coding of a CTU in a slice is not based on pixel data from another slice in the picture.

[0049] Some coding systems (such as VVC) provide the concept of slices. This mechanism allows a picture to be partitioned into one or several groups of slices. Each slice consists of one or several slices. As shown in pictures 201 and 202, two different slices are provided. The first type of slice is limited to slices that form rectangular areas in the picture. Picture 201 shows the picture being partitioned into five different rectangular slices. The second type of slice is limited to consecutive slices in raster scan order. Picture 202 shows the picture being partitioned into three different slices consisting of consecutive slices in raster scan order. Rectangular slices are a structure for processing the selection of regions of interest in a video. A slice can be encoded in the bitstream as one or several NAL units. The NAL unit, which stands for Network Abstraction Layer unit, is a logical unit of data used to encapsulate data in the encoded bitstream. In the example of the VVC coding system, a slice is encoded as a single NAL unit. When a slice is encoded in the bitstream as several NAL units, each NAL unit of the slice is a slice segment. A slice segment includes a slice segment header that contains the coding parameters of the slice segment. The header of the first segment NAL unit of a slice contains all the coding parameters of the slice. The slice segment headers of subsequent NAL units of a slice can contain fewer parameters than the first NAL unit. In this case, the first slice segment is an independent slice segment, and the subsequent segments are dependent slice segments.

[0050] In OMAF v2 ISO / IEC 23090-2, a sub-picture is a portion of a picture that represents a spatial subset of the original video content, which has been split into spatial subsets before video encoding at the content production side. A sub-picture is, for example, one or more slices forming a rectangular area.

[0051] Figure 2b An example of partitioning a picture into sub-pictures is shown. A sub-picture represents a portion of a picture that covers a rectangular area of the picture. Each sub-picture can have different sizes and encoding parameters. For example, a different slice grid and stripe partition can be defined for each sub-picture. Figure 2b In FIG, picture 204 is subdivided into 24 sub-pictures including sub-pictures 205 and 206. These two sub-pictures are Figure 2a The slice grid and partitioning of pictures 201 and 202 are similar to those described further below. In the second example, the slice and block partitioning is not defined per sub-picture, but rather at the picture level. A sub-picture is then defined as one or more slices forming a rectangular area.

[0052] Figure 2c An example of partitioning using block partitioning is shown. Each slice may include a set of blocks. A block is a contiguous set of CTUs that form a row in a slice. For example, Figure 2cFrame 207 is partitioned into 25 slices. Each slice contains exactly one block, except for the slice in the rightmost column, where each slice contains two blocks. For example, slice 208 contains two blocks 209 and 210. When block partitioning is used, a slice contains blocks from one slice or several blocks from other slices. In other words, a VCL NAL unit is a set of blocks, not a set of slices.

[0053] Figure 3 The organization of the bitstream in an exemplary coding system VVC is shown.

[0054] The bitstream 300 according to the VVC coding system consists of an ordered sequence of syntax elements and coded data. The syntax elements and coded data are placed into NAL units 301-305. There are different NAL unit types. The network abstraction layer provides the ability to encapsulate the bitstream into different protocols (such as RTP / IP (standing for Real Time Protocol / Internet Protocol), ISO base media file format, etc.). The network abstraction layer also provides a framework for packet loss resilience.

[0055] NAL units are divided into VCL NAL units and non-VCL NAL units, where VCL stands for Video Coding Layer. VCL NAL units contain the actual coded video data. Non-VCL NAL units contain additional information. This additional information can be parameters required for decoding the coded video data or supplementary data that can enhance the usability of the decoded video data. NAL unit 305 corresponds to a slice and constitutes the VCL NAL unit of the bitstream. Different NAL units 301-304 correspond to different parameter sets, and these NAL units are non-VCL NAL units. The VPS NAL unit 301 (VPS stands for Video Parameter Set) contains parameters defined for the entire video and, therefore, for the entire bitstream. The naming of the VPS can be changed and, for example, become the DPS in VVC. In an alternative, the VPS and DPS are different parameter set NAL units. The DPS (which stands for Decoder Parameter Set) NAL unit can define parameters that are more static than those in the VPS. In other words, the parameters of the DPS change less frequently than the parameters of the VPS. The SPS NAL unit 302 (SPS stands for Sequence Parameter Set) contains parameters defined for a video sequence. Specifically, the SPS NAL unit can define a sub-picture of a video sequence. The syntax of the SPS contains, for example, the following syntax elements:

[0056]

[0057]

[0058] The descriptor column gives the encoding of the syntax element, u(1) means that the syntax element is encoded using one bit, ue(v) means that the syntax element is encoded using an unsigned integer order 0 Exp-Golomb encoding, where the first left bit is a variable length code.

[0059] The presence of subpictures in an image depends on the value of subpics_present_flag. When this flag is equal to 0, it indicates that the image does not contain subpictures. When equal to 1, the syntax element set specifies the subpictures in a frame. The syntax element max_subpics_minus1 specifies the maximum number of subpictures in a picture of the video sequence. The SPS then defines the subpicture partitions using a grid of subpicture grid elements of a size defined by subpic_grid_col_width_minus1 and subpic_grid_row_height_minus1. Each grid element specifies subpic_grid_idx[i][j] (where i and j are the coordinates of the element in the grid), which is the index of the subpicture. There are as many subpic_grid_idx[i][j] values as there are subpictures in the picture of the video sequence. The subpic_grid_idx[i][j] syntax element is an identifier for the subpicture. All grid elements sharing the same index value form a rectangular area corresponding to the subpicture with an index (or identifier) equal to the index value. The subpic_treated_as_pic_flag[i] syntax element indicates whether sub-picture boundaries should be treated as picture boundaries except for loop filtering processing. The loop_filter_across_subpic_enabled_flag[i] syntax element states whether the loop filter is applied across sub-picture boundaries.

[0060] The PPS NAL unit 303 (PPS stands for Picture Parameter Set) contains the parameters defined for a picture or group of pictures. The APS NAL unit 304 (APS stands for Adaptation Parameter Set) contains the parameters of the loop filter (typically an adaptive loop filter (ALF) or a shaper model (or a luma map with a chroma scaling model)) defined at the slice level. The bitstream may also contain SEI (standing for Supplemental Enhancement Information) NAL units. The periodicity of occurrence of these parameter sets in the bitstream is variable. A VPS defined for the entire bitstream needs to appear only once in the bitstream. In contrast, an APS defined for a slice may appear once for each slice in each picture. In practice, different slices may rely on the same APS, and therefore there are typically fewer APSs than slices in each picture. When a picture is partitioned into sub-pictures, a new parameter set or PPS may be defined for each sub-picture or group of sub-pictures.

[0061] Each VCL NAL unit 305 contains a slice. A slice can correspond to an entire picture or a sub-picture, a single slice or multiple slices, or a single block or multiple blocks. A slice consists of a slice header 310 and an RBSP (raw byte sequence payload) 311 containing the block.

[0062] The syntax of the PPS as proposed in the current version of VVC includes syntax elements that specify the size of pictures and luma samples, and the partitioning of pictures into slices, blocks, and slices.

[0063] The syntax of PPS as proposed in the current version of VVC is organized as follows:

[0064]

[0065]

[0066] The descriptor column gives the encoding of the syntax element, u(1) means that the syntax element is encoded using one bit, ue(v) means that the syntax element is encoded using unsigned integer order 0 Exp-Golomb coding syntax elements, where the first left bit is variable length coded. The syntax elements pic_width_in_luma_samples and pic_height_in_luma_samples specify the width and height of the picture in luma samples.

[0067] When the number of slices in a picture is greater than one (single_tile_in_pic_flag is equal to 0), the PPS defines several syntax elements (not represented in the table above) that specify the partitioning of slices in a frame into a slice grid.

[0068] When tiles are present (brick_spliting_present_flag is equal to 1), the PPS contains a loop over each tile of the tile grid to indicate whether the tile is split into tiles. When a tile contains tiles, the tile configuration of the tile is encoded in the PPS.

[0069] Stripe partitions are represented using the following syntax elements:

[0070] The syntax element single_tile_in_pic_flag specifies whether the picture contains a single slice. In other words, when this flag is true, there is only one slice and one slice in the picture.

[0071] single_brick_per_slice_flag specifies whether each slice contains a single block. In other words, when this flag is true, all blocks of the picture belong to different slices.

[0072] The syntax element rect_slice_flag indicates that the slice of the picture forms a rectangular shape as represented in picture 201 .

[0073] When present, the syntax element num_slices_in_pic_minus1 is equal to the number of rectangular slices in the picture minus one.

[0074] Then, a syntax element (not shown in the table above) encodes the position of each slice relative to the block partition in a "for loop" over all slices of the picture. The index of the slice position parameter in the for loop is the index of the slice.

[0075] When signaled_slice_id_flag is equal to 1, a slice identifier is specified. In this case, the signaled_slice_id_length_minus1 syntax element indicates the number of bits used to encode each slice identifier value. The slice_id[] association table is indexed by the slice index and contains the identifier of the slice. When signaled_slice_id_flag is equal to 0, slice_id is indexed by the slice index and contains the slice index of the slice.

[0076] In summary, the PPS contains syntax elements that allow the location of slices in a frame to be determined. Since a sub-picture forms a rectangular area in a frame, the set of slices, slices, and blocks belonging to a sub-picture can be determined.

[0077] The slice header includes the slice address according to the following syntax in the current VVC version:

[0078]

[0079] When a slice is not rectangular, the slice header indicates the number of slices in the slice NAL unit by means of the num_tiles_in_slice_minus1 syntax element.

[0080] Each slice 320 may include a slice segment header 330 and slice segment data 331. The slice segment data 331 includes an encoded coding block 340. In the current version of the VVC standard, the slice segment header is not present, and the slice segment data contains the coding block data 340.

[0081] In a variant, the video sequence includes sub-pictures; the syntax of the slice header may be as follows:

[0082]

[0083] The slice header includes a slice_subpic_id syntax element that specifies the identifier of the sub-picture to which it belongs (e.g., corresponding to one of the values of subpic_grid_idx[i][j] defined in the SPS). As a result, all slices in a video sequence that share the same slice_sub_pic_id belong to the same sub-picture.

[0084] For illustration purposes only, Figure 4 The quadtree inference mechanism used in VVC for coding tree blocks that cross picture boundaries is schematically shown. In VVC, pictures are not restricted to having width and height multiples of the coding tree block size. Then, the rightmost coding tree block of a frame can cross the right side boundary 401 of the picture, and the bottommost coding tree block of a frame can cross the bottom boundary 402 of the picture. In these cases, VVC defines a quadtree inference mechanism for coding tree blocks that cross boundaries. The mechanism consists of recursively splitting any coding blocks of the coding tree blocks that cross picture boundaries until there are no more coding blocks that cross the boundary, or until the maximum quadtree depth is reached for those coding tree blocks. For example, coding tree block 403 is not automatically split, while coding tree blocks 404, 405 and 406 are automatically split. There is no signaling of the absence of an inferred quadtree: the decoder must infer the same quadtree across picture boundaries. However, the automatically obtained quadtree may be further refined for the coding treeblocks within a frame by signaling the splitting information of these coding treeblocks (if the maximum quadtree depth is not reached), for example, as shown in 407 .

[0085] When splitting a CTB into coding blocks, there is a minimum coding block that cannot be split. The size of this minimum coding block (which is a square) is given by MinCBSizeY. In some examples, MinCBSizeY is equal to 4.

[0086] When encoding, the coding blocks of the coding tree blocks that are outside the image are usually not encoded in the bitstream. When decoding, the decoder uses the same quadtree inference mechanism and knows that these coding blocks are not encoded. The decoder is able to decode other coding blocks that fall into the image that have been encoded in the bitstream. The obtained coding tree is encoded in the bitstream by the encoder. The decoder relies on this encoded coding tree information to correctly identify the blocks to decode and reconstruct the image. Specifically, the encoded coding tree contains a parameter that specifies whether the coding unit is split, for example called split_cu_flag.

[0087] Figure 5a and 5b The creation of a bitstream in which a boundary sub-picture is moved to a non-border position is shown.

[0088] In this example, the first bitstream 500 consists of 4 sub-pictures 1 HQ to 4 HQ The second bitstream 501 consists of 4 sub-pictures 1 LQ to 4 LQ This bitstream represents a low-quality version of the same video. Bitstream 502 is created by merging and rearranging some sub-pictures from bitstreams 500 and 501.

[0089] Specifically, the bitstream 502 consists of sub-picture 2 HQ 、1 LQ , 4 HQ and 3 LQ Composition, where sub-image 2 HQ Moving from the upper right position in bitstream 500 to the upper left position in bitstream 502, sub-picture 4 HQ Moving from the lower right position in bitstream 500 to the lower left position in bitstream 502, sub-picture 1 LQ Moving from the upper left position in bitstream 501 to the upper right position in bitstream 502, sub-picture 3 LQ Moves from the lower left position in bitstream 501 to the lower right position in bitstream 502 .

[0090] like Figure 5b As shown, the bitstream 500 consists of an image whose width is not a multiple of the coding tree block (CTB) size. Therefore, the rightmost coding tree blocks are subjected to inferred splitting. The rightmost coding blocks in these CTBs are not encoded in the bitstream. These uncoded coding blocks are represented by the shaded area 503. In the bitstream 502, the uncoded coding block 504 is located in sub-picture 2.HQ and 4 HQ The right border of , in the middle of the image.

[0091] The size of the image in pixels (or luma samples) is encoded in the bitstream. This information allows the decoder to know the image boundaries and infer the split of the rightmost CTB in the image. The decoder then knows the uncoded coded blocks and can decode the coded blocks that make up the image.

[0092] Dang Ruzi picture 2 HQ and 4 HQ A problem arises when a sub-picture, such as that in bitstream 502, is moved from the rightmost part of an image to another part. The size of the sub-picture is encoded in the bitstream as an integer number of CTBs. Due to this peculiarity in the standard, it is impossible for a decoder to infer the right boundary of the sub-picture and the corresponding inferred split of the rightmost CTB of the sub-picture. The decoder expects all coded blocks of the rightmost CTB of the sub-picture to be encoded in the bitstream. When some of these are missing, decoding fails.

[0093] What is needed is a way to rearrange sub-pictures in an image without re-encoding them.

[0094] This problem can be solved by proposing a coding method that includes encoding information that allows the decoder to infer the splitting of the CTB located at the right or bottom of the sub-picture, when the sub-picture is not located at the right or bottom of the image and its width or height is not a multiple of the CTB size. A corresponding decoding method for the generated bitstream is also proposed.

[0095] A sub-picture is a rectangular area in a picture represented by one or more slices. Sub-pictures are defined, for example, in SPS or PPS NAL units. The encoder can indicate for each sub-picture that their boundaries are to be treated as picture boundaries, which means that sub-pictures can be decoded independently of each other. Intra- and inter-frame prediction processes are constrained to use prediction information only from the same sub-picture in the current frame and the reference frame. Filtering of sub-picture boundaries is controlled by a flag defined for each sub-picture. This flag makes it possible to apply a loop filter or not to apply a loop filter at the boundaries of the sub-picture. When decoding a CTB from a slice, the decoder determines the index or identifier of the sub-picture to which the CTB belongs.

[0096] In the following, the use of sub-pictures is to allow areas of a video sequence to be decoded at different positions, and it is also possible to move a sub-picture at a new decoding position. For this reason, in a preferred embodiment, sub-picture boundaries are treated as picture boundaries (because it ensures that motion prediction is constrained to allow independent decoding of sub-pictures), and the loop filter is disabled at sub-picture boundaries. For VVC, these two conditions are met when the flags subpic_treated_as_pic_flag[i] and loop_filter_across_subpic_enabled_flag[i] are equal to 1 and 0, respectively, for a sub-picture. For HEVC, a sub-picture is, for example, a set of slices that belong to a Motion Constrained Tile Set according to the HEVC specification.

[0097] Figure 6 A method of encoding a video picture into a bitstream according to a first aspect of the invention is shown.

[0098] In a first step 601, an image is segmented into sub-pictures.

[0099] In step 602, the size of the sub-pictures is determined. The width and height of each sub-picture is a function of the region of interest present in the input video sequence. Typically, each sub-picture is sized to encompass a region of interest. The size of the sub-picture is determined in pixels (or luminance samples). When the size of the sub-picture is not a multiple of the CTB size, it is determined that a specific splitting inference process is required for the sub-picture. In particular, when Figure 5a This can happen when the encoder shown merges two bitstreams.

[0100] In step 603, the size of the sub-picture is encoded in the bitstream. This size indicates the location of the sub-picture splitting inference boundary when splitting inference processing is required. If required, information indicating that splitting inference processing is required for the sub-picture is encoded in the bitstream. This step provides information in the bitstream that allows the decoder to perform splitting inference processing.

[0101] In step 604, at least one slice constituting the sub-picture is encoded in the bitstream. When split inference processing is required, the encoder divides the sub-picture into two parts. The first part is the area between the top and left boundaries of the sub-picture and the horizontal and vertical split inference boundaries. The second part is the area between the split inference boundaries and the bottom and right boundaries of the sub-picture. For example, Figure 5b 2 HQ The white area in the sub-image corresponds to the first portion, and the shaded area corresponds to the second portion. In this paper, we refer to the first portion of the sub-image as the useful portion of the sub-image. Based on this split inference process, coding units that fall outside the useful portion of the sub-image are not encoded.

[0102] In an alternative embodiment, coding blocks that fall outside the useful part of the sub-picture are encoded with padding data.

[0103] In another alternative embodiment, a flag is inserted into the bitstream to indicate whether coding blocks that fall outside the useful portion of a sub-picture are coded with or without padding data. Typically, the encoder specifies in a picture parameter set (NAL) unit that a sub-picture contains coded blocks that are coded with padding data for coding units that fall outside the useful portion of the sub-picture. For example, a flag is associated with each sub-picture identifier to specify whether padding coded data is provided for coding units that fall outside the useful portion of the sub-picture.

[0104] In some embodiments, it may be known from the context of the bitstream that split inference processing is used for each slice. In these embodiments, no flag needs to be encoded to signal the use of split inference processing for sub-pictures.

[0105] It is proposed to introduce new syntax elements in a parameter set NAL unit (e.g. in an SPS) that enable the same splitting inference for CTBs at the right and / or bottom boundaries of a sub-picture to be obtained when moving at different positions in the merged bitstream.

[0106] When sub-picture boundaries are considered as picture boundaries, these syntax elements will allow the decoder to determine the size of the skipped coding blocks in the last CTB row and / or last CTB column of the sub-picture.

[0107] Therefore, when moving a sub-picture from the rightmost position in a picture to another position, the merging or encoding operation includes determining the size of the skipped coding block in the last CTB row and column, which is then specified, for example, in the SPS associated with the sub-picture. Decoding of the merged bitstream will use these values to determine the available size of the sub-picture.

[0108] For example, the syntax of SPS includes the following syntax elements:

[0109]

[0110]

[0111] According to the proposed embodiment, some new syntax elements have been introduced which are indicated in bold in the table. The semantics of these new syntax elements may be as follows.

[0112] subpic_split_inference_flag equal to 1 indicates that the last CTB row or last CTB column of the sub-picture is incomplete. The inference process used to split the CTBs of the last CTB row and column of the sub-picture takes into account the available size of the sub-picture to determine the value of split_cu_flag.

[0113] subpic_split_inference_flag is equal to 0 to indicate that all CTBs of the sub-picture are complete. No split inference processing is required for the CTBs on the last CTB row and column of the sub-picture.

[0114] subpic_split_inference_ctb_width[i] specifies the actual width of the CTB in the rightmost column of the CTB of the i-th sub-picture (when present) (i.e., the width in pixels of the CTB in the useful part of the i-th sub-picture). The value of subpic_split_inference_ctb_width[i] is specified in units of coding blocks of width MinCbSizeY and can be in the range of [0, CtbSizeY / MinCbSizeY-1] (inclusive). When not present, the value of subpic_split_inference_ctb_width[i] is inferred to be equal to CtbSizeY / MinCbSizeY, which corresponds to the width of the CTB. This coding syntax element is encoded, for example, using a fixed-length code of 7 bits or equal to log2(CtbSizeY / MinCbSizeY-1). Exp-Golomb codes may also be used.

[0115] subpic_split_inference_ctb_height[i] specifies the actual height of the CTBs in the last row of CTBs of the i-th sub-picture (when present). The value of subpic_split_inference_ctb_height[i] is specified in units of coding blocks of MinCbSizeY height and can be in the range of [0, CtbSizeY / MinCbSizeY-1] (inclusive). When not present, the value of subpic_split_inference_ctb_height[i] is inferred to be equal to CtbSizeY / MinCbSizeY, which corresponds to the height of the CTB. This coding syntax element is encoded, for example, using a fixed-length code of 7 bits or equal to log2(CtbSizeY / MinCbSizeY-1). Exp-Golomb codes may also be used.

[0116] The subpic_split_inference_flag allows the encoder to indicate to the decoder that special processing is required for the last CTB rows and columns of some sub-pictures. This means that the rightmost and / or bottom CTB rows contain coding blocks that have not yet been encoded by the encoder. The splitting of these CTBs into blocks is handled by the coding tree splitting process, which can infer the splitting of each block based on the values of the subpic_split_inference_ctb_width[i] and subpic_split_inference_ctb_height[i] syntax elements of the i-th sub-picture.

[0117] The actual size of the rightmost and bottom CTBs in the i-th sub-picture (specified by subpic_split_inference_ctb_width[i] and subpic_split_inference_ctb_height[i]) is in units of the minimum coding block size (MinCbSizeY, MinCbSizeY) specified in the SPS. The actual size of the CTB cannot exceed the maximum size of the coding tree block (CtbSizeY, CtbSizeY) as specified in the SPS. Therefore, subpic_split_inference_ctb_height and subpic_split_inference_ctb_width range from 0 to the ratio of the maximum size of a CTB to the minimum size of a coding block minus one.

[0118] In an alternative, subpic_split_inference_ctb_width[i] and subpic_split_inference_ctb_height[i] are expressed in units of a number of luma samples (typically four luma samples) to simplify the parsing of the SPS. In practice, the size of the CTB and the minimum size of the coding block are specified in the SPS and can be defined after the subpic_split_inference_ctb_width[i] and subpic_split_inference_ctb_height[i] syntax elements. In this case, the position of the split inference boundary is expressed as an integer multiple of the size of the minimum coding block of the coding tree block. In addition, when subpic_split_inference_ctb_width[i] and subpic_split_inference_ctb_height[i] are defined in different parameter set NAL units, the actual size of each sub-picture can be calculated independently of the CTB and coding block size.

[0119] The following pseudo code determines the boundary in the i-th sub-picture that triggers the split of the coding tree. The position of the vertical boundary is represented by the variable SubPicRightSplitInferenceBoundary[i], and the position of the horizontal boundary is represented by the variable SubPicBotSplitInferenceBoundary[i].

[0120] The position of the vertical boundary in the i-th sub-picture in units of luma samples is equal to the position of the right boundary aligned on the CTB boundary of the i-th sub-picture minus the subpic_split_inference_ctb_width[i] syntax element defined in the SPS. The position of the horizontal boundary in the i-th sub-picture in units of luma samples is equal to the position of the bottom boundary aligned on the CTB boundary of the i-th sub-picture minus the subpic_split_inference_ctb_height[i] syntax element defined in the SPS.

[0121] The variables SubPicRightSplitInferenceBoundary[i] and SubPicBotSplitInferenceBoundary[i] are derived as follows:

[0122]

[0123] Therefore, the proposed new syntax elements subpic_split_inference_flag, subpic_split_inference_ctb_width[i] and subpic_split_inference_ctb_height[i] associated with a sub-picture allow the encoder to indicate to the decoder the information needed to be able to infer the splitting of the incomplete CTB and determine the coded blocks that fall outside the useful part of the sub-picture.

[0124] The coding tree syntax includes a flag split_cu_flag that indicates whether the coding block is split. A split is inferred when the coding tree block is not completely within the boundaries of the useful part of the sub-picture and therefore the flag is not coded.

[0125] The value of split_cu_flag is 0, which specifies that the coding block is not to be split. The value of split_cu_flag is 1, which specifies that the coding block is to be split into four coding blocks using a quadtree split as indicated by the syntax element split_qt_flag, or into two coding blocks using a binary split as indicated by the syntax element mtt_split_cu_binary_flag, or into three coding blocks using a ternary split as indicated by the syntax element mtt_split_cu_binary_flag. The binary or ternary split can be vertical or horizontal, as indicated by the syntax element mtt_split_cu_vertical_flag.

[0126] When split_cu_flag is not present in the encoding of a coding block, the value of split_cu_flag is inferred as follows:

[0127] - The value of split_cu_flag is inferred to be equal to 1 if one or more of the following conditions are true:

[0128] -x0+cbWidth is greater than SubPicRightSplitInferenceBoundary[SubPicIdx].

[0129] -y0+cbHeight is greater than SubPicBotSplitInferenceBoundary[SubPicIdx].

[0130] - Additionally, the value of split_cu_flag is inferred to be equal to 0.

[0131] The above pseudo code determines whether the coded block at position (x0, y0) in luma sample coordinates with width (or height) equal to cbWidth (or cbHeight) is completely within the useful part of the current sub-picture with index equal to SubPicIdx. When outside this area, the encoder (and decoder) infers that the coded block is split (split_cu_flag is equal to 1).

[0132] When a coding block is inferred to be split because its boundaries are outside the useful portion of the sub-picture's actual pixel boundaries, there are three possible splits. The first is a quadtree split, where the coding block is further split into four equally sized square coding blocks. The second is a split into two coding blocks separated by a vertical or horizontal boundary that spans the entire coding block. This split is a binary tree split. Finally, the third split is a ternary tree split, which splits the coding block into three blocks separated by a vertical or horizontal boundary.

[0133] In addition to the split inference process, the encoder (and similarly, the decoder) applies a specific coding process to the coding blocks outside the useful part of the sub-picture. In a first embodiment, these coding blocks are not encoded. For each split coding block, the encoder checks whether the y-axis (respectively, x-axis) coordinate of the top boundary (respectively, the left boundary) of the block is greater than or equal to the SubPicBotSplitInferenceBoundary (respectively, SubPicRightSplitInferenceBoundary) value of the current sub-picture. In this case, the encoder skips the encoding of the block. Therefore, no encoded syntax elements are provided for these coding blocks. Therefore, when encoding adjacent coding units, these skipped coding blocks are considered unavailable for inter-frame and intra-frame prediction. Similarly, temporal prediction is restricted to avoid using pixel information from these skipped coding blocks.

[0134] Figure 7 The general decoding process of an embodiment of the present invention is shown. In step 701, the decoder determines the size of the sub-pictures of the frame by parsing the SPS and PPS of the stream, typically the width and height of the sub-picture. Specifically, the slice NAL units belonging to each sub-picture are determined. In step 702, the decoder decodes the slices of the sub-pictures that form the picture. Figure 8 More details of the decoding process of the coded CTBs in each slice are given. In step 703, an output picture is generated from the decoded picture. In particular, the decoder can optionally apply some cropping operations to the decoded sub-picture.

[0135] Figure 8 The decoding process 702 of the CTBs encoded in a slice is shown. In step 801, it is checked whether there is a slice to be decoded. When all slices have been decoded, the process ends. For a given slice, the decoder determines the identifier of the sub-picture to which the slice belongs in step 802.

[0136] Then, in step 803, a check is performed to determine whether there are any remaining CTBs to be decoded in the current slice. For a given CTB, in step 804, the decoder parses the syntax elements of the CTB. When the CTB is at the rightmost or bottom boundary of a sub-picture with the identifier determined in step 802, the decoder applies a split inference process to the CTB if this process is required for the current slice. Then, in step 805, the decoder decodes each coded block derived from the inference process and located within the sub-picture when the CTB is split. Coded blocks located outside the useful portion of the sub-picture are not decoded because they are not present in the bitstream. In some embodiments, if coded blocks located outside the useful portion of the sub-picture have been encoded with padding data, these coded blocks are decoded.

[0137] For sub-pictures whose size is not a multiple of the CTB size, the decoding process of the sub-picture results in incomplete right columns or bottom rows of CTBs. When the sub-pictures are arranged to form the final frame to be rendered, the decoder must determine the values of the pixels located in these pixel strips. Figure 5b An example of a band of undetermined pixels, band 504, is given in

[15] . This is due to the fact that sub-pictures within a frame are constrained to consist of an integer number of CTBs. As a result, a sub-picture contains a consistent set of decoded pixels corresponding to the area between the sub-picture's origin and the inferred split boundary. The remaining area corresponds to pixel values that are undefined or not intended to be displayed. In some embodiments, the decoder may set the values of the pixels in the remaining band to zero. In some other embodiments, the decoder may set the values of the pixels in the remaining band by copying the value of the nearest consistent pixel.

[0138] Since these bands of undetermined pixels are not intended to be displayed, they can be suppressed in the resulting frame by shifting adjacent sub-pictures. Figure 5b Sub-picture 1 in the bitstream 502 of LQ and 3 LQ The width of the tape 504 may be shifted left.

[0139] In this embodiment, the decoder shifts the decoded pixels of the sub-pictures to the right and bottom of the current sub-picture so that they align with the right and bottom inferred split boundaries. The decoded positions of luma samples of the coding block take into account that some sub-pictures are incomplete.

[0140] The decoder has all the information needed to identify the bands of undetermined pixels.In one embodiment, the decoder performs a shift of the sub-pictures to eliminate these bands in a post-decoding process after decoding all sub-pictures.

[0141] In another embodiment, the decoder can integrate the shift operation in the decoding process. In this embodiment, a new variable is defined by the decoder and used during the decoding of each CTB to decode the coded block just in its final position.

[0142] The SubPicOffSetX[i] and SubPicOffSetY[i] variables indicate the offsets to be subtracted from the X-axis coordinates and Y-axis coordinates of the pixels of the coding block of the i-th sub-picture, respectively, so that the pixels of the coding block of the i-th sub-picture will be placed at their correct positions.

[0143] For example, the xCtbShifted and yCtbShifted variables are the new coordinates of the top-left pixel of the CTB at the CtbAddrInRs address in the raster scan order of the picture after the shift operation.

[0144] The following pseudocode calculates these coordinates, where CtbAddressInRs is the raster scan address of the CTB in the picture; PicWidthInCtbsY is the width of the picture in CTBs, and CtbLog2SizeY is the log2 of the size of the CTB in luma samples; SubPicIdx is the index of the sub-picture containing the current CTB.

[0145]

[0146] The first two lines calculate the x- and y-coordinates of the origin of the coding tree block in units of luma samples. These values do not take into account skipped pixels in neighboring sub-pictures (if any). Then, if subpic_split_inference_flag is equal to 1, which indicates that the sub-picture may contain skipped blocks, the sum of the widths of the skipped pixels of the sub-picture located to the left (respectively, top) of the coding tree block is subtracted from the xCtbShifted (respectively, yCtbShifted) coordinate.

[0147] The encoder determines the shift offsets for the x-axis (SubPicOffsetX[i]) and y-axis (SubPicOffsetY[i]) coordinates of each sub-picture as follows. First, for a given sub-picture, determine the set of sub-pictures to the left of the current sub-picture. For each sub-picture in this set, the shift offset is equal to the difference between the right boundary of the last CTB row and the inferred boundary of the right sub-picture. Typically, this value is equal to CtbSizeY - Subpic_split_inference_ctb_width[j] * MinCbSizeY, where j is the index of the sub-picture in the set and CtbSizeY is the size of the CTB in units of luma samples or pixels. The value of the x-axis shift offset is equal to the sum of CtbSizeY - subpic_split_inference_ctb_width[j] * MinCbSizeY, where each j is equal to the index of the sub-picture in the left sub-picture set. Similarly, the value of the shift offset in the y-axis is equal to the sum of CtbSizeY-subpic_split_inference_ctb_height[j]*MinCbSizeY, where each j is equal to the index of the sub-picture in the top sub-picture set above the current sub-picture.

[0148] Figure 9The image is shown as being divided into 16 sub-images, where bands of undetermined pixels are represented by shaded bands. The arrows indicate the shifting of the sub-images to eliminate the bands of undetermined pixels. If all vertical bands (and, accordingly, horizontal bands) are the same size, the resulting image will be rectangular. This is also the case because each column of the sub-image contains the same number of horizontal bands, and each row of the sub-image contains the same number of vertical bands.

[0149] It will be appreciated that if these conditions are not met, the resulting image may not be rectangular.

[0150] In one embodiment, the possibility of specifying different inferred boundaries for sub-pictures is constrained to avoid defining pictures with non-rectangular picture shapes after the shift operation. In particular, after the shift or cropping of the decoded frame, the height (respectively, width) of each column (respectively, row) of the sub-picture must be equal in terms of luma samples.

[0151] The possibility of specifying different inferred boundaries in a given row (or column) of a sub-picture may result in a complex operation of shifting decoded samples with different shift offsets at each sub-picture. To this end, in an embodiment, a constraint that the sub-picture inferred boundaries are aligned across the image may be defined. For example, Figure 10 The case where a picture is partitioned into 16 sub-pictures that comply with this constraint is shown.

[0152] Figure 11 The concept of picture-wide borders is shown.

[0153] In an alternative approach, the encoder determines two types of sub-picture boundaries. First, sub-picture boundaries are aligned with other sub-picture boundaries that collectively span the entire picture width or height. We refer to such sub-picture boundaries as picture-wide boundaries. For example, Figure 11 The bottom border of sub-image 0 in the image is the image-wide border. Conversely, the bottom border of sub-image 1 is not the image-wide border.

[0154] In an embodiment, the encoder constrains the sub-picture layout so that the last CTB row of a sub-picture with a bottom boundary that is not picture-wide does not use inferred splitting. In other words, the split inference boundary is aligned with the CTB boundary, and, for example, the size of subpic_split_inference_ctb_height[i] is equal to the size of the CTB. The same principle applies to the last CTB column of a sub-picture with a picture-wide right boundary. In an alternative approach, the encoder determines the boundary type of each sub-picture and encodes the value of subpic_split_inference_ctb_width[i] (respectively, subpic_split_inference_ctb_height[i]) when the right (respectively, bottom) boundary of the sub-picture is a picture-wide boundary. When the right (respectively, bottom) boundary is not picture-wide, subpic_split_inference_ctb_width[i] (respectively, subpic_split_inference_ctb_height[i]) is not encoded and is inferred to be equal to the size of the CTB.

[0155] The encoder may apply a second constraint to sub-pictures whose bottom (respectively, right) boundaries are the same picture-wide boundary. For example, the constraint is that when the boundary is at the bottom of the sub-picture, subpic_split_inference_height[i] is equal for all i-th sub-pictures in the sub-picture set. Similarly, subpic_split_inference_width[i] may be equal for all i-th sub-pictures in the sub-picture set that have a common right picture-wide boundary.

[0156] Additionally, when a sub-picture is not independently decodable, the encoder can apply a third constraint that the split inference boundary is aligned with the CTB boundary. Typically, for the i-th sub-picture, this occurs when subpic_treated_as_pic_flag[i] is equal to 0 or loop_filter_across_subpic_enabled_flag[i] is equal to 1. This constraint ensures that the split inference mechanism is only used in the context of sub-picture merging. An embodiment based on these constraints will now be described.

[0157] In another embodiment, the syntax of the SPS is changed to specify the location of the vertical split inference boundary and the location of the horizontal split inference boundary across the entire picture. In this embodiment, since the sub-picture boundaries subject to split inference only involve picture-wide boundaries, picture-level signaling is used.

[0158] There are several alternatives to specify the location of these boundaries. The location of the split inferred boundaries can be signaled in parameter set NAL units such as SPS or PPS. The parameter set should describe information related to sub-picture and picture information.

[0159] In one alternative, the position of the inferred boundary is defined relative to the origin of the picture, for example in units of luma samples.

[0160] For example, the PPS syntax includes the following elements:

[0161]

[0162] in:

[0163] pps_subpic_split_inference_flag is a flag that indicates the use of sub-picture split inference. Typically, split inference is applied to picture-wide boundaries. pps_subpic_split_inference_flag is equal to 1 to indicate the presence of pps_split_inference_boundary_pos_x and pps_split_inference_boundary_pos_y; pps_subpic_split_inference_flag is equal to 0 to indicate the absence of pps_split_inference_boundary_pos_x and pps_split_inference_boundary_pos_y.

[0164] pps_split_inference_boundary_pos_x is used to calculate the value of PpsSplitInferenceBoundaryPosX, which specifies the position of the vertical split inference boundary in units of luma samples. pps_split_inference_boundary_pos_x can be in the range of 1 to Ceil(pic_width_in_luma_samples ÷ 4) - 1 (inclusive). For example, this coding syntax element is encoded using a fixed-length code of 13 bits or equal to log2(pic_width_in_luma_samples / 4). Exp-Golomb codes can also be used.

[0165] The position of the vertical split inference boundary PpsSplitInferenceBoundaryPosX is derived as follows:

[0166] PpsSplitInferenceBoundaryPosX=pps_split_inference_boundary_pos_x*4

[0167] pps_split_inference_boundary_pos_y is used to calculate the value of PpsSplitInferenceBoundaryPosY, which specifies the position of the horizontal split inference boundary in units of luma samples. pps_split_inference_boundary_pos_y can be in the range of 1 to Ceil(pic_height_in_luma_samples ÷ 4) - 1 (inclusive). For example, this coding syntax element is encoded using a fixed-length code of 13 bits or equal to log2(pic_width_in_luma_samples / 4). Exp-Golomb codes can also be used.

[0168] The position of the horizontal split inference boundary PpsSplitInferenceBoundaryPosY is derived as follows:

[0169] PpsSplitInferenceBoundaryPosY=pps_split_inference_boundary_pos_y*4

[0170] The coding tree split inference process compares the coding unit boundaries with the locations of the split inference boundaries. For example, when the right or bottom boundary of the coding block is greater than one of the split inference boundaries that span the current sub-picture, the coding unit split is inferred. Therefore, the split inference of slit_cu_flag is as follows:

[0171] When split_cu_flag is not present, the value of split_cu_flag is inferred as follows:

[0172] - The value of split_cu_flag is inferred to be equal to 1 if one or more of the following conditions are true:

[0173] -x0 + cbWidth is greater than PpsSplitInferenceBoundaryPosX and (SubPicLeft[SubPicIdx]) * (subpic_grid_col_width_minus1 + 1) * 4) < PpsSplitInferenceBoundaryPosX and (SubPicLeft[SubPicIdx] + SubPicWidth[i]) * (subpic_grid_col_width_minus1 + 1) * 4) > PpsSplitInferenceBoundaryPosX

[0174] -y0 + cbHeight is greater than PpsSplitInferenceBoundaryPosY and (SubPicTop[SubPicIdx]) * (subpic_grid_row_height_minus1 + 1) * 4) < PpsSplitInferenceBoundaryPosY and (SubPicTop[SubPicIdx] + SubPicHeight[i]) * (subpic_grid_row_height_minus1 + 1) * 4) > PpsSplitInferenceBoundaryPosY

[0175] - Otherwise, infer that the value of split_cu_flag is equal to 0.

[0176] Another equivalent algorithm for determining the value of split_cu_flag is as follows:

[0177] When split_cu_flag does not exist, the value of split_cu_flag is inferred as follows:

[0178] - If one or more of the following conditions are true, then infer that the value of split_cu_flag is equal to 1:

[0179] - SubPicInferenceSplitFlag is equal to 0 and x0 + cbWidth is greater than pic_width_in_luma_samples.

[0180] - SubPicInferenceSplitFlag is equal to 0 and y0 + cbHeight is greater than pic_height_in_luma_samples.

[0181] -SubPicInferenceSplitFlag is equal to 1 and x0+cbWidth is greater than SubPicInferenceBoundaryPosX.

[0182] -SubPicInferenceSplitFlag is equal to 1 and y0+cbHeight is greater than SubPicInferenceBoundaryPosY.

[0183] Otherwise, the value of split_cu_flag is inferred to be equal to 0.

[0184] SubPicInferenceSplitFlag is equal to 1 when the sub-picture containing the coding block is spanned by a split inference boundary. SubPicInferenceBoundaryPosX is the horizontal coordinate of the vertical split inference boundary spanning the current sub-picture. When the sub-picture is not spanned by a vertical split inference boundary, SubPicInferenceBoundaryPosX is set to a value greater than or equal to the horizontal coordinate of the right side boundary of the sub-picture to avoid unnecessary split inference. Similarly, SubPicInferenceBoundaryPosY is the vertical coordinate of the horizontal split inference boundary spanning the current sub-picture (if any). Otherwise, when the sub-picture is not spanned by a horizontal split inference boundary, SubPicInferenceBoundaryPosY is set to a value greater than or equal to the horizontal coordinate of the right side boundary of the sub-picture. The factor "4" is introduced because the split inference boundary is constrained to correspond to the minimum coding block boundary. Assume that the size of the minimum coding block is 4×4. If the size of the minimum coding block is different, another factor can be used in these equations.

[0185] In another example, the syntax elements described above in the PPS may be defined at the SPS level, which may be advantageous when the split inference boundary does not change at each new PPS NAL unit. Thus, the syntax of the SPS includes, for example, the following elements:

[0186]

[0187] Among them, subpic_split_inference_flag, split_inference_boundary_pos_x, and split_inference_boundary_pos_y have similar semantics to pps_subpic_split_inference_flag, pps_split_inference_boundary_pos_x, and pps_split_inference_boundary_pos_y:

[0188] subpic_split_inference_flag is a flag indicating the use of sub-picture split inference. subpic_split_inference_flag is equal to 1 to indicate the presence of split_inference_boundary_pos_x and split_inference_boundary_pos_y; subpic_split_inference_flag is equal to 0 to indicate the absence of split_inference_boundary_pos_x and split_inference_boundary_pos_y;

[0189] split_inference_boundary_pos_x is used to calculate the value of SplitInferenceBoundaryPosX, which specifies the position of the vertical split inference boundary in units of luma samples. split_inference_boundary_pos_x can be in the range of 1 to Ceil(pic_width_in_luma_samples ÷ 4) - 1 (inclusive).

[0190] The position of the vertical split inference boundary SplitInferenceBoundaryPosX is derived as follows:

[0191] SplitInferenceBoundaryPosX=split_inference_boundary_pos_x*4

[0192] split_inference_boundary_pos_y is used to calculate the value of SplitInferenceBoundaryPosY, which specifies the position of the horizontal split inference boundary in units of luma samples. split_inference_boundary_pos_y can be in the range of 1 to Ceil(pic_height_in_luma_samples ÷ 4) - 1 (inclusive).

[0193] The position of the horizontal split inference boundary SplitInferenceBoundaryPosY is derived as follows:

[0194] SplitInferenceBoundaryPosY=split_inference_boundary_pos_y*4;

[0195] The split_inference_boundary_pos_x and split_inference_boundary_pos_y syntax elements are coded, for example, using 13 bits or a fixed length code equal to log2(pic_width_in_luma_samples / 4) for split_inference_boundary_pos_x and log2(pic_height_in_luma_samples / 4) for split_inference_boundary_pos_y. Exp-Golomb codes may also be used.

[0196] In another embodiment, split_inference_boundary_pos_y can be in the range of 1 to Ceil(pic_height_in_luma_samples÷4) (inclusive), and split_inference_boundary_pos_x can be in the range of 1 to Ceil(pic_width_in_luma_samples÷4) (inclusive). In this case, the maximum value of the range indicates that the split inference boundary is aligned with the picture boundary or outside the picture boundary. As a result, when set to the maximum value, it indicates that no vertical or horizontal split boundary is used. In a variant, two different flags are used (one flag each for vertical and horizontal split inference boundaries). In this variant, the presence of split_inference_boundary_pos_x and split_inference_boundary_pos_y is conditional on these flag values.

[0197] In another embodiment, the PPS defines one or more horizontal and vertical splitting inference boundaries. The encoder specifies the number of horizontal splitting inference boundaries and the number of vertical splitting inference boundaries. These boundaries are limited to picture-wide boundaries. For example, the syntax of the PPS contains the following elements:

[0198]

[0199] pps_num_ver_split_inference_boundaries specifies the number of pps_split_inference_boundaries_pos_x[i] syntax elements present in the PPS. When pps_num_ver_split_inference_boundaries is not present, it is inferred to be equal to 0.

[0200] pps_split_inference_boundaries_pos_x[i] is used to calculate the value of PpsSplitInferenceBoundaryPosX[i], which specifies the position of the i-th vertical split inference boundary in units of luma samples. pps_split_inference_boundary_pos_x[i] can be in the range of 1 to Ceil(pic_width_in_luma_samples ÷ 4) - 1 (inclusive). For example, this coding syntax element is encoded using a fixed-length code of 13 bits or equal to log2(pic_width_in_luma_samples / 4). Exp-Golomb codes can also be used.

[0201] The position of the i-th vertical split inference boundary PpsSplitInferenceBoundaryPosX[i] is derived as follows:

[0202] PpsSplitInferenceBoundaryPosX[i]=pps_split_inference_boundary_pos_x[i]*4

[0203] The distance between any two vertical split inferred boundaries may be greater than or equal to CtbSizeY luma samples, which is the size of the CTB in luma samples.

[0204] pps_num_hor_split_inference_boundaries specifies the number of pps_split_inference_boundaries_pos_y[i] syntax elements present in the PPS. When pps_num_hor_split_inference_boundaries is not present, it is inferred to be equal to 0.

[0205] pps_split_inference_boundaries_pos_y[i] is used to calculate the value of PpsSplitInferenceBoundaryPosY[i], which specifies the position of the inferred boundary for the i-th horizontal split of luma samples. pps_split_inference_boundary_pos_y[i] can be in the range of 1 to Ceil(pic_height_in_luma_samples ÷ 4) - 1 (inclusive). This coding syntax element is encoded, for example, using a fixed-length code of 13 bits or equal to log2(pic_height_in_luma_samples / 4). Exp-Golomb codes can also be used.

[0206] The position of the i-th horizontal split inference boundary PpsSplitInferenceBoundaryPosY[i] is derived as follows:

[0207] PpsSplitInferenceBoundaryPosX[i]=pps_split_inference_boundary_pos_y[i]*4

[0208] The distance between any two horizontal split inferred boundaries may be greater than or equal to CtbSizeY luma samples, which is the size of the CTB in luma samples.

[0209] In another embodiment, the encoder specifies the location of the picture-wide split inference boundaries relative to the grid formed by the CTBs of the picture. Typically, the SPS or PPS includes syntax elements to indicate the number and location of the vertical and horizontal boundaries of the split inference process as indices of CTB rows or CTB columns. For each CTB row (respectively, column) described in the PPS, the encoder indicates the width (respectively, height) of the CTB to be used for the inference process.

[0210]

[00136] In a variation of this embodiment, the encoder and decoder can determine picture-wide sub-picture boundaries based on the sub-picture definition. Specifically, the encoder determines NumVerSplitInferenceBoundaries, which is the number of vertical picture-wide sub-picture boundaries, and NumHorSplitInferenceBoundaries, which is the number of horizontal picture-wide sub-picture boundaries. The position of the i-th (in raster scan order) vertical picture-wide sub-picture boundary on the x-axis is also determined and stored in the PictureWideSubPictureBoundaryPosX[i] variable. The position of the i-th (in raster scan order) horizontal picture-wide sub-picture boundary on the y-axis is also determined and stored in the PictureWideSubPictureBoundaryPosY[i] variable.

[0211] The encoder then encodes the width (respectively, height) used for split inference for the CTB row (respectively, column) at the left (respectively, top) of the vertical (respectively, horizontal) picture-wide sub-picture boundary.

[0212] For example, PPS includes the following syntax elements:

[0213]

[0214] in:

[0215] pps_split_inference_ctb_width[i] is used to calculate the value of PpsSplitInferenceBoundaryPosX[i], which specifies the position of the i-th vertical split inference boundary in units of luma samples. pps_split_inference_ctb_width[i] is specified in units of coding blocks of width MinCbSizeY and can be in the range of 0 to CtbSizeY / MinCbSizeY-1 (inclusive). For example, this coding syntax element is encoded using a fixed-length code equal to log2(CtbSizeY / MinCbSizeY-1). Exp-Golomb codes can also be used.

[0216] The position (in terms of luma samples) of the i-th vertical split inference boundary PpsSplitInferenceBoundaryPosX[i] is derived as follows:

[0217] PpsSplitInferenceBoundaryPosX[i]=PictureWideSubPictureBoundaryPosX[i]-CtbSizeY+pps_split_inference_ctb_width[i]*MinCbSizeY

[0218] pps_split_inference_ctb_height[i] is used to calculate the value of PpsSplitInferenceBoundaryPosY[i], which specifies the position of the i-th horizontal split inference boundary in units of luma samples. pps_split_inference_ctb_height[i] is specified in units of coding blocks of MinCbSizeY height and can be in the range of 0 to CtbSizeY / MinCbSizeY-1 (inclusive). For example, this coding syntax element is encoded using a fixed-length code equal to log2(CtbSizeY / MinCbSizeY-1). Exp-Golomb codes can also be used.

[0219] The position (in terms of luminance samples) of the i-th horizontal split inference boundary PpsSplitInferenceBoundaryPosY[i] is derived as follows:

[0220] PpsSplitInferenceBoundaryPosY[i]=PictureWideSubPictureBoundaryPosY[i]-CtbSizeY+pps_split_inference_ctb_height[i]*MinCbSizeY

[0221] In another embodiment, the inferred split boundary is derived from the sub-picture partition (e.g., from the sub-picture grid). Specifically, the encoder can describe sub-picture boundaries that are not aligned with CTB boundaries. For example, the sub-picture partition defines a sub-picture grid element size that is lower than the CTB size. When two consecutive grid elements have two different sub-picture indices and belong to the same CTB, this indicates that the first sub-picture has a right (respectively, bottom) boundary that is not aligned with the CTB boundary. The second sub-picture has a left (respectively, top) boundary that is not aligned with the CTB boundary. In this case, the inferred split boundary is aligned with the right (respectively, bottom) boundary of the first sub-picture and, therefore, with the left (respectively, top) boundary of the second sub-picture. In a variant, different signaling is used to indicate the right and left boundaries of the sub-pictures, for example, by explicitly indicating the width and height of each sub-picture in units lower than the CTB size. In this case, the sub-picture width and height are compared to their values in CTB units to determine the sub-pictures that are not aligned with the CTB boundary. In another variant, a specific value of the sub-picture element is reserved to indicate that the sub-picture grid element does not have coded data. The split inferred boundaries are then derived from the sub-picture partitions, which avoids explicit signaling of the split inferred boundaries in the SPS, PPS, or any parameter set or SEI.

[0222] The VVC specification defines the conformance window in PPS, as shown in the following table:

[0223]

[0224] The consistency window is a rectangular area in each picture represented by the left, right, top and bottom offsets to the picture boundaries. At the end of the decoding process, the decoder applies a cropping process to remove pixels outside the consistency window.

[0225] In the previous embodiment, we described how the encoder skips some coding blocks in a CTB that are crossed by an inferred split boundary. The skipped coding blocks are located to the right of the inferred vertical split boundary and below the inferred horizontal split boundary. In one embodiment, the decoder will decode CTBs with undefined pixels for these coding blocks. To this end, it is proposed to add a syntax element to the PPS that will define the consistency window for each sub-picture.

[0226] In one embodiment, the parameter set includes a new syntax element for indicating a consistent rectangular region of luma samples in each sub-picture. This is all pixels decoded by the decoder whose value is equal to the value encoded by the encoder, i.e., corresponding to the useful part of the sub-picture.

[0227] Typically, an SPS or PPS defines four consistency window offset parameters for each sub-picture of a stream that has the flag subpic_treated_as_pic_flag equal to true. For example, the syntax may be as follows:

[0228]

[0229] The syntax elements subpic_conf_win_left_offset[i], subpic_conf_win_right_offset[i], subpic_conf_win_top_offset[i], subpic_conf_win_bottom_offset[i], subpic_conf_win_left_offset[i] specify the four consistency window offset parameters.

[0230] In another embodiment, for example, when the encoder constrains the split inferred boundary to span the picture width and height, the undefined pixels form one or more pixel strips in the picture. Therefore, instead of specifying a conformance region in each sub-picture, the PPS describes the pixel strips that are excluded from the conformance window and should be cropped.

[0231] For example, the PPS defines the number of pixel strips that are excluded from the consistency window defined for the pictures in the PPS. For each pixel strip, the encoder specifies the width of the strip. For example, the syntax of the PPS may include the following elements:

[0232]

[0233] has the following semantics:

[0234] conformance_exclusion_band_flag equal to 1 indicates that a conformance exclusion band is present in the PPS. conformance_exclusion_band_flag equal to 0 indicates that no conformance exclusion band is present in the PPS.

[0235] num_ver_conformance_exclusion_bands specifies the number of conformance_exclusion_ver_band_pos_x[i] and conformance_exclusion_ver_band_width_minus1[i] syntax elements present in the PPS. When num_ver_conformance_exclusion_bands is not present, it is inferred to be equal to 0.

[0236] conformance_exclusion_ver_band_width_minus1[i] plus 1 specifies the width of the i-th vertical conformance exclusion band in luma samples and is used to calculate PpsConformanceVerBandPosX[i], which specifies the position of the i-th vertical conformance exclusion band boundary in luma samples. Conformance. conformance_exclusion_ver_band_width_minus1[i] can be in the range of 0 to CtbSizeY-2. This coding syntax element is encoded, for example, using an 8-bit fixed-length code or equal to log2(pic_width_in_luma_samples / CtbSizeY) or equal to log2(CtbSizeY-2). Exp-Golomb codes may also be used.

[0237] conformance_exclusion_ver_band_pos_x[i] is used to calculate the value of PpsConformanceVerBandPosX[i], which specifies the position of the i-th vertical conformance exclusion band boundary in units of luma samples. conformance_exclusion_ver_band_pos_x[i] can be in the range of 0 to PicWidthInCtbsY (inclusive). In a variant, the range is 1 to PicWidthInCtbsY-1 (inclusive) to avoid specifying a vertical band that starts at the first or last pixel column of the picture. For example, this coding syntax element is encoded using a fixed length code of 8 bits or equal to log2(pic_width_in_luma_samples / CtbSizeY). Exp-Golomb codes may also be used.

[0238] The position of the i-th vertical exclusion band PpsConformanceVerBandPosX[i] is derived as follows:

[0239] PpsConformanceVerBandPosX[i]=conformance_exclusion_ver_band_pos_x[i]*CtbSizeY-(conformance_exclusion_ver_band_width_minus1[i]+1)

[0240] The distance between any two vertical consistency exclusion band boundaries can be greater than or equal to CtbSizeY luma samples, which is the size of the CTB in luma samples. In a variant, there is no restriction on the distance between two vertical consistency exclusion band boundaries, and each band can overlap another band. In this case, the decoder must determine the overlapping bands to determine the actual consistency area of the picture.

[0241] num_hor_conformance_exclusion_bands specifies the number of conformance_exclusion_hor_band_pos_y[i] and conformance_exclusion_hor_band_height_minus1[i] syntax elements present in the PPS. When num_hor_conformance_exclusion_bands is not present, it is inferred to be equal to 0.

[0242] conformance_exclusion_hor_band_height_minus1[i] plus 1 specifies the height of the i-th horizontal conformance exclusion band in luma samples and is used to calculate PpsConformanceHorBandPosY[i], which specifies the position of the i-th horizontal conformance exclusion band boundary in luma samples. Conformance. conformance_exclusion_hor_band_height_minus1[i] can be in the range of 0 to CtbSizeY-2. This coding syntax element is encoded, for example, using a fixed-length code of 7 bits or equal to log2(pic_height_in_luma_samples / CtbSizeY). Exp-Golomb codes may also be used.

[0243] conformance_exclusion_hor_band_pos_y[i] is used to calculate the value of PpsConformanceHorBandPosY[i], which specifies the position of the i-th horizontal consistency exclusion band boundary in units of luma samples. conformance_exclusion_hor_band_pos_y[i] can be in the range of 0 to PicHeightInCtbsY (inclusive). PicHeightInCtbsY is the height of the picture in units of CTB. In a variant, the range is 1 to PicHeightInCtbsY-1 (inclusive) to avoid specifying a horizontal band that starts at the first or last pixel row of the picture. This coding syntax element is encoded, for example, using a fixed length code of 8 bits or equal to log2(pic_height_in_luma_samples / CtbSizeY). Exp-Golomb codes can also be used.

[0244] The position of the i-th horizontal exclusion band PpsConformanceHorBandPosY[i] is derived as follows:

[0245] PpsConformanceHorBandPosY[i]=conformance_exclusion_hor_band_pos_y[i]*CtbSizeY-(conformance_exclusion_hor_band_height_minus1[i]+1)

[0246] The distance between any two horizontal conformance exclusion band boundaries can be greater than or equal to CtbSizeY luma samples, which is the size of the CTB in luma samples. In a variant, there is no restriction on the distance between two horizontal conformance exclusion band boundaries, and each band can overlap another band. In this case, the decoder must determine the overlapping bands to determine the actual conformance region of the picture.

[0247] In a variant, the number of horizontal and vertical exclusion bands is optional. In this case, when the num_ver_conformance_exclusion_bands and num_hor_conformance_exclusion_bands syntax elements are not present, their values are inferred to be equal to 1. The syntax of the PPS is, for example, as follows:

[0248]

[0249] In a variant, the description of the horizontal or / and vertical split inference boundary information is optional. For example, the syntax of PPS (or SPS) is as follows:

[0250]

[0251] The semantics of the new syntactic (bold) elements are as follows:

[0252] conformance_exclusion_ver_band_flag equal to 1 indicates that conformance_exclusion_ver_band_width_minus1 and conformance_exclusion_ver_band_pos_x are present in the PPS, which specify the vertical conformance exclusion band. conformance_exclusion_band_flag equal to 0 indicates that conformance_exclusion_ver_band_width_minus1 and conformance_exclusion_ver_band_pos_x are not present in the PPS.

[0253] conformance_exclusion_ver_band_width_minus1 plus 1 specifies the width of the vertical conformance exclusion band in luma samples and is used to calculate PpsConformanceVerBandPosX. PpsConformanceVerBandPosX specifies the position of the left edge of the vertical conformance exclusion band. conformance_exclusion_ver_band_width_minus1 can be in the range of 0 to CtbSizeY-2.

[0254] conformance_exclusion_ver_band_pos_x is used to calculate the value of PpsConformanceVerBandPosX. conformance_exclusion_ver_band_pos_x can be in the range of 1 to PicWidthInCtbsY-1 (inclusive).

[0255] The position of the left border of the vertical exclusion band, PpsConformanceVerBandPosX, is derived as follows:

[0256] PpsConformanceVerBandPosX=conformance_exclusion_ver_band_pos_x*CtbSizeY-(conformance_exclusion_ver_band_width_minus1+1)

[0257] conformance_exclusion_hor_band_flag equal to 1 indicates that conformance_exclusion_hor_band_height_minus1 and conformance_exclusion_hor_band_pos_y are present in the PPS, which specify the horizontal conformance exclusion band. conformance_exclusion_band_flag equal to 0 indicates that conformance_exclusion_hor_band_height_minus1 and conformance_exclusion_hor_band_pos_y are not present in the PPS.

[0258] conformance_exclusion_hor_band_height_minus1 plus 1 specifies the height of the horizontal conformance exclusion band in luma samples and is used to calculate PpsConformanceHorBandPosY. PpsConformanceHorBandPosY specifies the position of the top edge of the horizontal conformance exclusion band. conformance_exclusion_hor_band_height_minus1 can be in the range of 0 to CtbSizeY-2.

[0259] conformance_exclusion_hor_band_pos_y is used to calculate the value of PpsConformanceHorBandPosY. conformance_exclusion_hor_band_pos_y can be in the range of 1 to PicHeightInCtbsY-1 (inclusive).

[0260] The position of the top edge of the horizontal exclusion band, PpsConformanceHorBandPosY, is derived as follows:

[0261] PpsConformanceHorBandPosY=conformance_exclusion_hor_band_pos_y*CtbSizeY-(conformance_exclusion_hor_band_height_minus1+1)

[0262] In a variant, the width, height, and coordinates of the exclusion band are expressed in chroma samples. These values in luma samples are obtained by multiplying the values expressed in chroma samples by SubWidthC and SubHeightC. SubWidthC and SubHeightC represent the horizontal and vertical sampling ratios between the luma and chroma components, respectively. For example, when the chroma format is 4:2:0, SubWidthC and SubHeightC are equal to 2. In another variant, the width, height, and coordinates of the exclusion band are expressed in units of the minimum CTB size to further reduce the length of the syntax elements.

[0263] In another embodiment, the width and height of the exclusion band are larger than the CTB. This allows two consecutive pixel exclusion bands to be merged into a single exclusion band. For example, when two adjacent sub-pictures define two adjacent groups of non-uniform pixels. This allows for a more compact description of the exclusion band.

[0264] In another embodiment, the number of inferred vertical consistency exclusion bands and the number of horizontal consistency exclusion bands are equal to the number of vertical split inferred boundaries and the number of horizontal split inferred boundaries, respectively.

[0265] In another embodiment, the positions of the vertical and horizontal consistency exclusion bands are inferred to be equal to the positions of the vertical split inferred boundary and the horizontal split inferred boundary, respectively.

[0266] In another embodiment, the width of a given vertical consistency exclusion band is inferred to be equal to the number of pixels between the vertical split inferred boundary and the nearest right CTB boundary at the same location. The same applies to horizontal exclusion bands with horizontal split inferred boundaries.

[0267] In another embodiment, the sub-image split inference is derived from the sub-image partitions. The pixels excluded from the consistency window are inferred to correspond to the area between the right (respectively, bottom) boundary of the sub-image that is not aligned with the CTB boundary and the right (respectively, bottom) boundary of these CTBs.

[0268] In a variation, these regions are grouped into different sub-pictures. In this case, the width or height of these sub-pictures is smaller than the CTB size. Therefore, these non-consistent sub-pictures can be associated with specific indices. The SPS or PPS can provide a list of sub-picture indices that are inconsistent and should be excluded from the consistency window. In another embodiment, the number of bands excluded from the consistency window and their sizes are determined from the inferred split boundary locations.

[0269] In another alternative, the number of bands excluded from the consensus window and their sizes are determined based on the inferred split boundary positions.

[0270] The conforming cropping window includes luma samples with horizontal picture coordinates from SubWidthC*conf_win_left_offset to pic_width_in_luma_samples-(SubWidthC*conf_win_right_offset+1) and vertical picture coordinates from SubHeightC*conf_win_top_offset to pic_height_in_luma_samples-(SubHeightC*conf_win_bottom_offset+1), inclusive.

[0271] Additionally, luma samples with horizontal picture coordinates from PpsConformanceVerBandPosX[i] to PpsConformanceVerBandPosX[i]+conformance_exclusion_ver_band_width_min us1[i]+1 are excluded from the conformance window, where i is in the range of 0 to num_ver_conformance_exclusion_bands.

[0272] Additionally, luma samples with vertical picture coordinates from PpsConformanceHorBandPosY[i] to PpsConformanceHorBandPosY[i]+conformance_exclusion_hor_band_height_minus1[i]+1 are excluded from the conformance window, where i is in the range of 0 to num_ver_conformance_exclusion_bands.

[0273] The width and height of the image after the cropping process correspond to the variables PicOutputWidthL and PicOutputHeightL derived as follows: the width (respectively, height) of the image is equal to the cropping window width (respectively, height) minus the width (respectively, height) of the pixel exclusion band, which corresponds to the following pseudo code:

[0274]

[0275] In an embodiment, the pictures and their sub-pictures that can be displaced in a merge operation are constrained to have sizes that are multiples of the CTB size in both directions. Typically, when this constraint is not adhered to, it can be accomplished by extending the input picture with padding pixels. Therefore, there is no inferred boundary mechanism to be used. This constraint is signaled in the bitstream. For example, a flag can be used to indicate that sub-pictures are subject to this constraint, and therefore sub-pictures can be freely merged and displaced at any position in the resulting image without any boundary issues. To ensure that sub-pictures can be freely merged, an additional constraint is associated with the flag, namely that the sub-pictures are independently decodable. This flag can be defined to apply to all sub-pictures in the picture, or it can be defined at the sub-picture level to indicate that the associated sub-pictures can be freely merged. In one embodiment, the flag indicating that the size of the picture and its sub-pictures is a multiple of the CTB size is defined in the sequence parameter set and is valid for all sub-pictures of the sequence.

[0276] In some embodiments, the consistency window may be defined at the sub-picture level (eg, in a SEI message).

[0277] In one embodiment, the syntax of the SPS is changed to include an additional flag indicating whether the encoding of the sub-pictures of the bitstream constrains the coding tools and video sequence characteristics to allow merge operations. When set, the flag constrains the sub-pictures to be independently decodable or not. The syntax may be as follows:

[0278]

[0279]

[0280] For example, the semantics of the element are as follows:

[0281] subpic_mergeable_flag equal to 1 indicates that all sub-pictures of a picture are constrained for merge operations. subpic_mergeable_flag equal to 0 indicates that a sub-picture may or may not be constrained.

[0282] pic_width_max_in_luma_samples specifies the maximum width (in luma samples) of each decoded picture that references the SPS. pic_width_max_in_luma_samples shall not be equal to 0 and shall be an integer multiple of Max(8, MinCbSizeY). When subpic_mergeable_flag is equal to 1, pic_width_max_in_luma_samples shall be an integer multiple of CtbSizeY. Therefore, when a sub-picture is constrained for merge operations (subpic_mergeable_flag is equal to 1), the coded picture is constrained to use padding when the original width of the picture is not a multiple of the CTB size.

[0283] pic_height_max_in_luma_samples specifies the maximum height (in luma samples) of each decoded picture that references the SPS. pic_height_max_in_luma_samples shall not be equal to 0 and shall be an integer multiple of Max(8, MinCbSizeY). When subpic_mergeable_flag_mergeable_flag is equal to 1, pic_height_max_in_luma_samples shall be an integer multiple of CtbSizeY. Therefore, when sub-pictures are constrained for merge operations (subpic_mergeable_flag is equal to 1), the coded picture is constrained to use padding when the original height of the picture is not a multiple of the CTB size.

[0284] In this embodiment, the semantics of some elements of the PPS can be constrained to ensure that when sub-pictures are constrained for merging operations, each picture that references the PPS has a picture size that is a multiple of the CTB size. For example, the semantics of pic_width_in_luma_samples and pic_height_in_luma_samples are as follows:

[0285] pic_width_in_luma_samples specifies the width (in luma samples) of each decoded picture that references the PPS. pic_width_in_luma_samples shall not be equal to 0, shall be an integer multiple of Max(8, MinCbSizeY), and shall be less than or equal to pic_width_max_in_luma_samples. When subpic_mergeable_flag is equal to 1 (e.g., in an SPS with an identifier signaled in the PPS), pic_width_in_luma_samples shall be an integer multiple of CtbSizeY.

[0286] pic_height_in_luma_samples specifies the height (in luma samples) of each decoded picture that references the PPS. pic_height_in_luma_samples shall not be equal to 0 and shall be an integer multiple of Max(8, MinCbSizeY) and shall be less than or equal to pic_height_max_in_luma_samples. When subpic_mergeable_flag is equal to 1, pic_height_in_luma_samples shall be an integer multiple of CtbSizeY.

[0287] This first set of constraints on pictures ensures that when subpic_mergeable_flag is equal to 1, any merge position is possible for subpictures in a picture of the merged stream. The encoder may have to use padding data to make the width and height of the picture a multiple of the CTB size. The encoder can signal areas with padding data by defining a consistency window in the bitstream. Typically, one consistency window for each subpicture can be defined in the SEI message.

[0288] In another embodiment, when a sub-picture is constrained for merging operations (subpic_mergeable_flag is equal to 1), the mergeable flag also constrains the temporal prediction mechanism within each sub-picture of all sub-pictures described in the parameter set NAL unit.

[0289] For example, subpic_treated_as_pic_flag[i] equal to 1 specifies that the i-th sub-picture of each coded picture in the CLVS is to be treated as a picture in the decoding process except for in-loop filtering operations. subpic_treated_as_pic_flag[i] equal to 0 specifies that the i-th sub-picture of each coded picture in the CLVS is not to be treated as a picture in the decoding process except for in-loop filtering operations. When not present, the value of subpic_treated_as_pic_flag[i] is inferred to be equal to subpic_mergeable_flag. Therefore, when a sub-picture is constrained for merge operations (subpic_mergeable_flag equal to 1), the subpic_treated_as_pic_flag[i] syntax element is not present in the parameter set NAL unit and is inferred to be equal to 1, which indicates that the temporal prediction of the sub-picture is constrained to each sub-picture boundary. Otherwise, the temporal prediction may or may not be constrained.

[0290] In another embodiment, when a sub-picture is constrained for merging (subpic_mergeable_flag is equal to 1), the loop filter mechanism is disabled across sub-picture boundaries for all sub-pictures described in the parameter set NAL unit. For example, loop_filter_across_subpic_enabled_flag[i] equal to 1 specifies that in-loop filtering operations may be performed across the boundaries of the i-th sub-picture in each coded picture in the CLVS. loop_filter_across_subpic_enabled_flag[i] equal to 0 specifies that in-loop filtering operations are not performed across the boundaries of the i-th sub-picture in each coded picture in the CLVS. When not present, the value of loop_filter_across_subpic_enabled_pic_flag[i] is inferred to be equal to !subpic_mergeable_flag. Therefore, when a sub-picture is constrained for merging operations (subpic_mergeable_flag is equal to 1), the loop_filter_across_subpic_enabled_pic_flag[i] syntax element is not present in the parameter set NAL unit and is inferred to be equal to 0, which indicates that the loop filter is constrained to not be enabled across sub-picture boundaries. Otherwise, the loop filter may or may not be constrained.

[0291] In some embodiments, spatial access may be provided only over some time intervals. The proposed flag may be defined in the PPS or even in the picture header accordingly. The picture header is a non-VCL NAL unit that defines certain syntax elements defined at the picture level and applicable to all slices of the picture.

[0292] In another embodiment, the signaling of a sub-picture associates an identifier of the sub-picture with a sub-picture index. This identifier is unique for a given sub-picture and simplifies the merge operation because it allows avoiding the need to rewrite the sub-picture index in the slice header when the sub-picture is moved to a new location after the merge operation. For this reason, when a sub-picture is constrained for a merge operation, one of the non-VCL NAL units can signal the sub-picture identifier in the bitstream.

[0293] Typically, the sub-picture identifier can be present in the SPS, PPS or picture header. Specifically, sps_subpic_id_present_flag is a flag of the SPS, which, when equal to 1, indicates that the sub-picture identifier is present in the bitstream. The semantics of this syntax element are, for example, as follows:

[0294] sps_subpic_id_present_flag equal to 1 specifies that a sub-picture ID map is present in the SPS. sps_subpic_id_present_flag equal to 0 specifies that a sub-picture ID map is not present in the SPS. When subpic_mergeable_flag is equal to 1, sps_subpic_id_present_flag must be equal to 1. As a result, when a sub-picture is constrained for merge operations, signaling of the sub-picture identifier is provided in the bitstream. In a variant, when subpic_mergeable_flag is equal to 1, sps_subpic_id_present_flag is not present in the SPS and is inferred to be equal to 1. The corresponding syntax of the SPS may contain the following syntax elements:

[0295]

[0296]

[0297] In another embodiment, the presence of sub-picture identifiers in the picture header may complicate the merge operation. In practice, when the identifier is signaled in the picture header, the mapping of sub-picture identifiers performed in the parameter set NAL unit is overridden. Since a different picture header is sent for each frame, there is a possibility that the mapping changes at each picture. As a result, a typical merge operation requires checking whether the picture header modifies the sub-picture identifier mapping. To avoid this check, the mapping of identifiers in the picture header is disabled when a sub-picture is constrained for a merge operation. The presence of the sub-picture identifier mapping in the picture header is controlled by the ph_subpic_id_signalling_present_flag syntax element. In this embodiment, the semantics of ph_subpic_id_signalling_present_flag are as follows:

[0298] ph_subpic_id_signalling_present_flag equal to 1 specifies that sub-picture ID mapping is signaled in the PH. ph_subpic_id_signalling_present_flag equal to 0 specifies that sub-picture ID mapping is not signaled in the PH. When subpic_mergeable_flag is equal to 1, ph_subpic_id_signalling_present_flag must be equal to 0. In a variant, when subpic_mergeable_flag is equal to 1, ph_subpic_id_signalling_present_flag is not present in the picture header and is inferred to be equal to 0.

[0299] The PPS can also signal the mapping of sub-picture identifiers. For the picture header, it is necessary to check whether the mapping between two PPS NAL units has changed. This additional check increases the complexity of the merge operation. Therefore, in one embodiment, when sub-pictures are constrained for merge operations, the sub-pictures are constrained to be identical in each and every PPS for which subpic_mergeable_flag is equal to 1. The pps_subpic_id[i] syntax element of the PPS specifies the sub-picture ID (or identifier) of the i-th sub-picture corresponding to the mapping of sub-picture identifiers to sub-picture indices. When subpic_mergeable_flag is equal to 1, all PPSs referencing the same SPS shall have the same value of pps_subpic_id[i], where i is in the range of 0 to pps_num_subpics_minus1 (inclusive).

[0300] In another embodiment, the information indicating that a sub-picture is constrained for a merge operation corresponds to a specific profile of the VVC specification. Typically, a specific value (e.g., 3) of the general_profile_idc syntax element indicates that the output layer conforms to the profile for sub-picture merge operations. In a variant, the information indicating that a sub-picture is constrained for a merge operation corresponds to a sub-profile. In this case, a specific value (e.g., 3) of general_sub_profile_idc[i] indicates that the bitstream is constrained for sub-picture merge operations.

[0301] The general constrained information structure of VVC is set by the flags described in the profile, tier and level information, which makes it possible to disable one or more coding tools. In one embodiment, the general constrained information structure includes any one of the following syntax elements:

[0302] - no_dependent_subpicture_flag syntax element, when equal to 1, specifies that for any value of i, subpic_treated_as_pic_flag[i] shall be equal to 1. no_dependent_subpicture_flag equal to 0 does not impose such a constraint. In a variant, when equal to 1, specifies that for any value of i, subpic_treated_as_pic_flag[i] and loop_filter_across_subpic_enabled_flag[i] shall be equal to 1.

[0303] - The no_picture_header_subpicture_id_mapping syntax element, when equal to 1, specifies that all picture header NAL units shall have ph_subpic_id_signalling_present_flag equal to 0. no_picture_header_subpicture_id_mapping equal to 0 imposes no such constraint.

[0304] - no_pps_subpicture_id_mapping_change syntax element, when equal to 1, specifies that all PPS NAL units shall have equal pps_subpic_id[i] values for any value of i. no_pps_subpicture_id_mapping_change equal to 0 imposes no such constraint.

[0305] In an embodiment, SEI messages are proposed to handle areas within a picture that require post-decoding cropping operations. This may result in, for example, bitstream extraction and merging operations where some sub-pictures originally located at the right or bottom border of the picture contain padding data. The SEI indicates to the decoder a consistency window for at least some sub-pictures that define picture areas containing padding or even unreliable or useless data that the content creator believes should be removed before the image is displayed.

[0306] SEI can provide the following syntax, for example:

[0307]

[0308]

[0309] has the following semantics:

[0310] subpic_conf_win_cancel_flag equal to 1 indicates that the SEI message cancels the persistence of any previous sub-picture consistency window SEI message in output order that applies to the current layer. subpic_conf_win_cancel_flag equal to 0 indicates that sub-picture consistency window information follows.

[0311] subpic_conf_win_num_subpics_minus1 plus 1 specifies the number of sub-picture consistency windows present in the SEI message. This value is a function of the number of sub-pictures present in the picture. Typically, the value of subpic_conf_win_num_subpics_minus1 is required to be equal to sps_num_subpics_minus1 to allow the definition of a consistency window for each sub-picture.

[0312] subpic_conf_win_left_offset[i], subpic_conf_win_right_offset[i], subpic_conf_win_top_offset[i], and subpic_conf_win_bottom_offset[i] specify the samples of the i-th sub-picture of the picture in the CLV that is output from the decoding process (in terms of a rectangular area specified in picture coordinates for the output relative to the origin of the i-th sub-picture as described in the SPS NAL unit).

[0313] The sub-picture consistent cropping window for the i-th sub-picture contains luma samples with horizontal picture coordinates from SubPictureLuma_X[i]+SubWidthC*conf_win_left_offset to SubPictureLuma_X[i]+SubPictureLuma_Width[i]-(SubWidthC*subpic_conf_win_right_offset[i]+1) and vertical picture coordinates from SubPictureLuma_Y[i]+SubHeightC*subpic_conf_win_top_offset[i] to SubPictureLuma_Y[i]+SubPictureLuma_Height[i]-(SubHeightC*sub pic_conf_win_bottom_offset[i]+1), inclusive. Where SubPictureLuma_X[i] and SubPictureLuma_Y[i] specify the horizontal and vertical picture coordinates of the first pixel in the i-th sub-picture, and SubPictureLuma_Width[i] and SubPictureLuma_Height[i] specify the width and height in luma samples of the i-th sub-picture described in the SPS. For example, these variables are calculated as follows:

[0314] SubPictureLuma_X[i]=subpic_ctu_top_left_x[i]*CtbSizeY

[0315] SubPictureLuma_Y[i]=subpic_ctu_top_left_y[i]*CtbSizeY

[0316] SubPictureLuma_Width[i]=(subpic_width_minus1[i]+1)*CtbSizeY

[0317] SubPictureLuma_Height[i]=(subpic_height_minus1[i]+1)*CtbSizeY

[0318] In some cases, only a subset of sub-pictures (typically, sub-pictures with padding data) require a consistency window. In this case, a new syntax element indicates for each sub-picture whether a consistency window is signaled. For example, a "for" loop over each sub-picture specifies the subpic_conf_win_signalled_flag[i] syntax element for the i-th sub-picture described in the sub-picture. When equal to 1, the offset parameter is present and the consistency window is specified for the sub-picture. Otherwise, subpic_conf_win_signalled_flag[i] is equal to 0, consistency is not signaled for the i-th sub-picture, and the offset parameter is not present and is inferred to be equal to 0. In a variant, the number of sub-picture consistency windows described in the SEI message is different from the number of sub-pictures in the picture, and for each signaled sub-picture consistency window signaled in the SEI, a list of one or more sub-picture indices in the picture is associated with the index of the sub-picture consistency window. The list of indices indicates the sub-pictures that use the sub-picture consistency window. For example, the for loop of the SEI message indicates the subpic_conf_win_num_subpics_minus1[i] syntax element, which is the number of sub-picture indices associated with the i-th sub-picture consistency window of the SEI message minus 1. Next, the processing loop for j in the range of 0 to subpic_conf_win_num_subpics_minus1[i] (inclusive) defines the subpic_conf_win_subpic_index[i][j] syntax element. subpic_conf_win_subpic_index[i][j] specifies the j-th index of the sub-picture using the i-th sub-picture consistency window. In a variant, the SEI message may define a sub-picture identifier instead of a sub-picture index. In this case, the bit length of the sub-picture identifier is optionally described in the SEI message.

[0319] Figure 12 FIG1 is a schematic block diagram of a computing device 120 for implementing one or more embodiments of the present invention. The computing device 120 may be a device such as a microcomputer, a workstation, or a lightweight portable device.

[0320] The computing device 120 includes a communication bus connected to:

[0321] - a central processing unit 121 denoted as CPU, such as a microprocessor;

[0322] A random access memory 122, designated as RAM, is used to store executable code of the method according to the embodiment of the present invention and registers suitable for recording variables and parameters necessary for implementing the method according to the embodiment of the present invention. The memory capacity of the random access memory can be expanded, for example, by connecting an optional RAM to an expansion port;

[0323] - a read-only memory 123 marked as ROM, for storing a computer program for implementing an embodiment of the present invention;

[0324] - A network interface 124, which is typically connected to a communications network through which digital data to be processed is sent or received. The network interface 124 can be a single network interface or a collection of different network interfaces (e.g., wired and wireless interfaces, or different types of wired or wireless interfaces). Under the control of a software application running in the CPU 121, data packets are written to the network interface for transmission or read from the network interface for reception;

[0325] - a user interface 125 , which may be used to receive input from a user or display information to a user;

[0326] A hard disk 126 , denoted as HD, which may be provided as a mass storage device;

[0327] An I / O module 127 , which may be used to receive / send data from / to an external device (such as a video source or a display, etc.).

[0328] The executable code can be stored in the read-only memory 123, on the hard disk 126, or on a removable digital medium (e.g., a disk). According to a variant, the executable code of the program can be received via the network interface 124 by means of a communication network so that the executable code of the program is stored in one of the storage components of the communication device 120 (e.g., the hard disk 126) before being executed.

[0329] The central processing unit 121 is adapted to control and direct the execution of instructions or portions of software code of one or more programs according to embodiments of the present invention, which instructions are stored in one of the aforementioned storage components. After being powered on, the CPU 121 is capable of executing instructions related to software applications from the main RAM memory 122 after loading these instructions, for example, from the program ROM 123 or the hard disk (HD) 126. Such software applications, when executed by the CPU 121, cause the steps of the flowchart of the present invention to be performed.

[0330] Any step of the algorithm of the present invention may be implemented in software by executing instructions or a program set by a programmable computing machine (such as a PC ("personal computer"), a DSP ("digital signal processor") or a microcontroller, etc.); or may be implemented in hardware by a machine or a dedicated component (such as an FPGA ("field programmable gate array") or an ASIC ("application-specific integrated circuit"), etc.).

[0331] Although the present invention has been described above with reference to specific embodiments, the present invention is not limited to the specific embodiments, and modifications within the scope of the invention will be apparent to those skilled in the art.

[0332] Many further modifications and variations will occur to those skilled in the art upon reference to the foregoing illustrative embodiments, which are given by way of example only and are not intended to limit the scope of the invention, which is determined solely by the appended claims. In particular, different features from different embodiments may be interchanged where appropriate.

[0333] The various embodiments of the present invention described above may be implemented individually or as a combination of multiple embodiments. In addition, features from different embodiments may be combined where necessary or a combination of elements or features from separate embodiments may be combined in a single embodiment where it is beneficial.

[0334] Each feature disclosed in this specification (including any accompanying claims, abstract and drawings) may be replaced by an alternative feature serving the same, equivalent or similar purpose, unless expressly stated otherwise. Therefore, unless expressly stated otherwise, each feature disclosed is only one example of a general series of equivalent or similar features.

[0335] In the claims, the word "comprising" does not exclude other elements or steps, and the indefinite article "a" or "an" does not exclude a plurality. The mere fact that different features are recited in mutually different dependent claims does not indicate that a combination of these features cannot be used to advantage.

Claims

1. A method for encoding video data comprising a picture into a bitstream, the picture being partitioned into sub-pictures, the method comprising: encoding a first flag in a sequence parameter set of the bitstream, wherein the presence of sub-picture information depends on the first flag; If the first flag is set to 1, encoding a second flag in the sequence parameter set of the bitstream, wherein if the second flag is set to 1, the second flag indicates that all sub-picture boundaries associated with the sequence parameter set are regarded as picture boundaries and loop filtering across sub-picture boundaries is not performed; if the second flag is set to 0, encoding a third flag and a fourth flag in the bitstream, the third flag indicating whether a loop filter across sub-picture boundaries for a given sub-picture is enabled, and the fourth flag indicating whether the sub-picture is treated as a picture in decoding processing other than in-loop filtering operations, wherein if the second flag is set to 1, the third flag and the fourth flag are not encoded in the bitstream; as well as Coding tree blocks constituting the picture are encoded into the bitstream, wherein information specifying a maximum width in units of luma samples of each coded picture referring to the sequence parameter set is encoded in the sequence parameter set.

2. The method according to claim 1, wherein The second flag is associated with all sub-pictures.

3. The method according to claim 1, wherein The second flag is defined in the sequence parameter set, which is a syntax structure containing syntax elements applied to pictures of the bitstream.

4. The method according to claim 1, wherein Sub-pictures are further identified using sub-picture identifiers.

5. The method according to claim 4, wherein It is forbidden to define the sub-picture identifier in the picture header.

6. The method according to claim 4, wherein: The sub-picture identifier defined in a picture parameter set must be the same in all picture parameter sets.

7. The method according to claim 1, wherein The second mark is associated with a specific grade.

8. The method according to claim 1, wherein The method further comprises: for at least one sub-picture, Information indicating the consistency window of the sub-picture is encoded in the bitstream.

9. A method for decoding a bitstream of video data comprising a picture, the picture being partitioned into sub-pictures, the method comprising: decoding a first flag from a sequence parameter set of the bitstream, wherein the presence of sub-picture information depends on the first flag; decoding, if the first flag is set to 1, a second flag from the sequence parameter set of the bitstream, wherein, if the second flag is set to 1, the second flag indicates that all sub-picture boundaries associated with the sequence parameter set are considered as picture boundaries and loop filtering across sub-picture boundaries is not performed; if the second flag is set to 0, decoding a third flag and a fourth flag from the bitstream, the third flag indicating whether a loop filter across a sub-picture boundary is enabled for a given sub-picture, and the fourth flag indicating whether the sub-picture is treated as a picture in decoding processing other than an in-loop filtering operation, wherein if the second flag is set to 1, the third flag and the fourth flag are not decoded from the bitstream; as well as The bitstream is decoded based on at least the first flag, wherein information specifying a maximum width in units of luma samples of each decoded picture that references the sequence parameter set is decoded from the sequence parameter set.

10. The method according to claim 9, wherein: The second flag is associated with all sub-pictures.

11. The method according to claim 9, wherein The second flag is defined in the sequence parameter set, which is a syntax structure containing syntax elements applied to pictures of the bitstream.

12. A computer program product for a programmable device, the computer program product comprising a sequence of instructions for implementing the method according to any one of claims 1 to 11 when the sequence of instructions is loaded into and executed by the programmable device. 13 . A non-transitory computer-readable storage medium storing computer program instructions, wherein the computer program instructions are used to implement the method according to claim 1 when executed by a processor.

14. An apparatus for encoding video data comprising a picture into a bitstream, the picture being partitioned into sub-pictures, the apparatus comprising a processor configured to: encoding a first flag in a sequence parameter set of the bitstream, wherein the presence of sub-picture information depends on the first flag; In case the first flag is set to 1, a second flag is encoded in the sequence parameter set of the bitstream, wherein In a case where the second flag is set to 1, the second flag indicates that all sub-picture boundaries associated with the sequence parameter set are regarded as picture boundaries, and loop filtering across sub-picture boundaries is not performed; if the second flag is set to 0, encoding a third flag and a fourth flag in the bitstream, the third flag indicating whether a loop filter across sub-picture boundaries for a given sub-picture is enabled, and the fourth flag indicating whether the sub-picture is treated as a picture in decoding processing other than in-loop filtering operations, wherein if the second flag is set to 1, the third flag and the fourth flag are not encoded in the bitstream; as well as Coding tree blocks constituting the picture are encoded into the bitstream, wherein information specifying a maximum width in units of luma samples of each coded picture referring to the sequence parameter set is encoded in the sequence parameter set.

15. An apparatus for decoding a bitstream of video data comprising a picture, the picture being partitioned into sub-pictures, the apparatus comprising a processor configured to: decoding a first flag from a sequence parameter set of the bitstream, wherein the presence of sub-picture information depends on the first flag; In case the first flag is set to 1, a second flag is decoded from the sequence parameter set of the bitstream, wherein In a case where the second flag is set to 1, the second flag indicates that all sub-picture boundaries associated with the sequence parameter set are regarded as picture boundaries, and loop filtering across sub-picture boundaries is not performed; if the second flag is set to 0, decoding a third flag and a fourth flag from the bitstream, the third flag indicating whether a loop filter across a sub-picture boundary is enabled for a given sub-picture, and the fourth flag indicating whether the sub-picture is treated as a picture in decoding processing other than an in-loop filtering operation, wherein if the second flag is set to 1, the third flag and the fourth flag are not decoded from the bitstream; as well as The bitstream is decoded based on at least the first flag, wherein information specifying a maximum width in units of luma samples of each decoded picture that references the sequence parameter set is decoded from the sequence parameter set.