Method and apparatus for encoding and decoding a video stream with sub-pictures

By encoding specific information in the bitstream to infer CTB splits at image boundaries, the method addresses decoding failures in sub-picture merging, ensuring accurate decoding of rearranged sub-pictures without re-encoding.

JP7804806B2Active Publication Date: 2026-01-22CANON KK
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2025034026
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-12-17
Filing Date
2025-03-04
Publication Date
2026-01-22
Estimated Expiration
2040-09-07

AI Technical Summary

Technical Problem

Decoders face challenges in inferring the split of coding tree blocks (CTBs) at the rightmost or bottommost positions of sub-pictures when their sizes are not multiples of the CTB size, leading to incomplete encoding and decoding failures during the merging of sub-pictures from different video bitstreams.

Method used

Encoding methods that include information in the bitstream to allow decoders to infer the splits of CTBs at the right and bottom of sub-pictures, using syntax elements like subpic_split_inference_flag, subpic_split_inference_ctb_width, and subpic_split_inference_ctb_height to determine the partition boundaries, ensuring complete decoding.

Benefits of technology

Enables successful decoding of sub-pictures rearranged within images without re-encoding, by providing necessary information for the decoder to handle incomplete CTBs at image boundaries, thus maintaining decoding integrity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007804806000018
    Figure 0007804806000018
  • Figure 0007804806000019
    Figure 0007804806000019
  • Figure 0007804806000020
    Figure 0007804806000020
Patent Text Reader

Abstract

To provide a method and a device for encoding and decoding a video bitstream that facilitates sub-picture replacement.SOLUTION: An encoding method includes dividing a picture into sub-pictures, determining a sub-picture size where the size of each sub-picture contains one region of interest present in the input video sequence, and when a division inference process is required, encoding the sub-picture size, which is information indicating the position of the division inference boundary of the sub-picture, into a bitstream, and encoding at least one slice constituting the sub-picture into the bitstream.SELECTED DRAWING: Figure 6
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This disclosure relates to methods and devices for encoding and decoding video bitstreams that facilitate sub-picture replacement, and more particularly to encoding and decoding video bitstreams that result from the merging of sub-pictures coming from different video bitstreams. [Background technology]

[0002] The size of an image in a video bitstream may not correspond to a multiple of the size of the coding tree block (CTB) used in the encoding process. The CTB may be recursively divided during encoding to find a coding block size that specifically optimizes the encoding process. Thus, the CTB may be divided down to the smallest coding block size. If the size of an image is not a multiple of the size of the CTB, the right or bottom boundary of the image intersects with the rightmost or bottommost CTB. In this case, an estimated split of the CTB is provided. Coding blocks outside the image are typically not encoded.

[0003] When considering the division of an image into sub-pictures, the right-most or bottom-most sub-picture may contain an incomplete CTB that is subject to an inferred split, including some uncoded coding blocks.

[0004] The decoder is provided with the size of the image in pixels and the size of the CTB in the bitstream, so it can determine the exact location of the right and bottom boundaries of the image and perform a speculative split of the rightmost and bottommost CTB in the image.

[0005] Because the size of a subpicture is provided as an integer number of CTBs, given the displacement of a subpicture located at the right or bottom boundary of the image elsewhere in the image, the decoder cannot infer the split of the rightmost or bottommost CTB and the identity of the missing coding block, and decoding fails. Summary of the Invention

[0006] The present invention is devised to address one or more of the aforementioned problems. An encoding method is proposed that includes encoding information that allows a decoder to infer splits of CTBs located to the right and bottom of subpictures whose width and height are not multiples of the size of the CTBs when the subpictures are not located to the right and bottom of the image. A corresponding decoding method for the generated bitstream is also proposed.

[0007] According to one aspect of the present invention, there is provided a method for encoding video data comprising pictures into a bitstream, the method comprising: dividing the pictures into sub-pictures; and for at least one sub-picture: encoding, in the bitstream, information indicating that a picture has a size that is a multiple of the size of a coding tree block and that a sub-picture is independently decodable; and encoding the coding tree blocks that make up the sub-picture into a bitstream.

[0008] In one embodiment, the information is associated with all sub-pictures.

[0009] In one embodiment, the information is defined in a sequence parameter set, which is a syntax structure that contains syntax elements that apply to pictures in a bitstream.

[0010] In one embodiment, the sub-picture is further identified using a sub-picture identifier.

[0011] In one embodiment, defining sub-picture identifiers in picture headers is prohibited.

[0012] In one embodiment, the sub-picture identifiers defined in a picture parameter set must be the same in all picture parameter sets.

[0013] In one embodiment, the information is associated with a particular profile.

[0014] According to another aspect of the present invention, there is provided a method of encoding video data comprising pictures into a bitstream, the method comprising: dividing the pictures into sub-pictures; and for at least one sub-picture: encoding, in the bitstream, information indicative of a conformance window of the sub-picture; and encoding the coding tree blocks that make up the sub-picture into a bitstream.

[0015] In one embodiment, information indicating the compatibility window is defined in the SEI message.

[0016] In one embodiment, the information defines a left offset, a right offset, a top offset, and a bottom offset for the subpicture.

[0017] According to another aspect of the present invention, there is provided a method of encoding video data comprising pictures into a bitstream, the method comprising: dividing the pictures into sub-pictures; and for at least one sub-picture: encoding, in the bitstream, first information indicating that the sub-picture has a size that is a multiple of a size of a coding tree block and that the sub-picture is independently decodable; encoding second information indicative of a conformance window of the sub-picture in the bitstream; and encoding the coding tree blocks that make up the sub-picture into a bitstream.

[0018] According to another aspect of the invention, a computer program product for a programmable device is proposed, the computer program product comprising a series of instructions for carrying out the method according to the invention when loaded into and executed by the programmable device.

[0019] According to another aspect of the invention, there is provided a computer readable storage medium storing computer program instructions for carrying out a method according to the invention.

[0020] According to another aspect of the invention there is provided a computer program product which when executed causes the method of the invention to be carried out.

[0021] According to another aspect of the present invention, there is provided a device for encoding video data including pictures into a bitstream, the picture being divided into sub-pictures, and for at least one sub-picture: encoding, in the bitstream, information indicating that the sub-picture has a size that is a multiple of the size of the coding tree block and that the sub-picture is independently decodable; A device is provided having a processor configured to encode coding tree blocks that constitute a sub-picture into a bitstream.

[0022] According to another aspect of the present invention, there is provided a device for encoding video data including pictures into a bitstream, the picture being divided into sub-pictures, and for at least one sub-picture: encoding, in the bitstream, information indicating a conformance window of the sub-picture; A device is provided having a processor configured to encode coding tree blocks that constitute a sub-picture into a bitstream.

[0023] According to another aspect of the present invention, there is provided a device for encoding video data including pictures into a bitstream, the picture being divided into sub-pictures, and for at least one sub-picture: encoding, in the bitstream, first information indicating that the sub-picture has a size that is a multiple of the size of the coding tree block and that the sub-picture is independently decodable; encoding second information indicative of a conformance window of the sub-picture in the bitstream; A device is provided having a processor configured to encode coding tree blocks that constitute a sub-picture into a bitstream.

[0024] At least a portion of the methods according to the present invention may be computer-implemented. Accordingly, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment (including firmware, resident software, microcode, etc.), or an embodiment combining software and hardware aspects, all of which may be referred to generally herein as a "circuit," "module," or "system." Furthermore, the present invention may take the form of a computer program product embodied in any tangible medium of expression having computer-usable program code embodied in the medium.

[0025] Since the present invention can be implemented in software, the present invention can be embodied as computer-readable code for provision to a programmable device on any suitable carrier medium. Tangible, non-transitory carrier media can include storage media such as floppy disks, CD-ROMs, hard disk drives, magnetic tape devices, or solid-state memory devices. Transit carrier media can include signals such as electrical, electronic, optical, acoustic, magnetic, or electromagnetic signals, e.g., microwave or RF signals. [Brief explanation of the drawings]

[0026] Embodiments of the present invention will now be described, by way of example only, with reference to the following drawings, in which: [Figure 1a] FIG. 1a shows two different application examples for combining regions of interest. [Figure 1b] FIG. 1b shows two different application examples for combining regions of interest. [Figure 2a] Figure 2a shows some divisions in the coding system. [Figure 2b] FIG. 2b shows an example of division of a picture into sub-pictures. [Figure 2c] FIG. 2c shows an example of a division of a picture into tiles and bricks. [Figure 3] FIG. 3 shows the structure of a bitstream in an exemplary coding system VVC. [Figure 4] FIG. 4 shows a schematic representation of the quadtree inference mechanism used in VVC for coding tree blocks that cross image boundaries. [Figure 5a] FIG. 5a shows the generation of a bitstream in which a border subpicture is moved to a non-border position. [Figure 5b] FIG. 5b shows the generation of a bitstream in which a border subpicture is moved to a non-border position. [Figure 6] FIG. 6 illustrates a method for encoding pictures of a video into a bitstream according to a first aspect of the invention. [Figure 7] FIG. 7 illustrates the general decoding process of one embodiment of the present invention. [Figure 8] FIG. 8 shows the decoding process of a CTB coded into slices. [Figure 9] FIG. 9 shows an example where a picture is divided into 16 sub-pictures. [Figure 10] FIG. 10 shows an example where a picture is split into 16 sub-pictures with a picture-wide split inference boundary. [Figure 11]FIG. 11 illustrates the concept of a picture-wide boundary. [Figure 12] FIG. 12 is a schematic block diagram of a computing device for implementing one or more embodiments of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0027] 1a and 1b show two different application examples for combining regions of interest.

[0028] For example, FIG. 1a illustrates an example in which picture (or frame) 100 from a first video bitstream and picture 101 from a second video bitstream are merged into picture 102 in a resulting bitstream. Each picture consists of four regions of interest, numbered 1 through 4. Picture 100 has been coded using coding parameters that result in high-quality coding. Picture 101 has been coded using coding parameters that result in low-quality coding. As is well known, pictures coded at low quality are associated with lower bitrates than pictures coded at high quality. The resulting picture 102 combines regions of interest 1, 2, and 4 from picture 101, coded at low quality, with region of interest 3 from picture 100, coded at high quality. The goal of such a combination is generally to obtain a high-quality region of interest, here region 3, while keeping the resulting bitrate reasonable by coding regions 1, 2, and 4 at low quality. Such kind of scenarios can occur especially in the context of omnidirectional content, which allows for higher quality for the content that is actually viewed while the rest has lower quality.

[0029] FIG. 1b shows a second example in which four different videos A, B, C, and D are merged to form a resulting video. Picture 103 of video A is composed of attention regions A1, A2, A3, and A4. Picture 104 of video B is composed of attention regions B1, B2, B3, and B4. Picture 105 of video C is composed of attention regions C1, C2, C3, and C4. Picture 106 of video D is composed of attention regions D1, D2, D3, and D4. Picture 107 of the resulting video is composed of regions B4, A3, C3, and D1. In this example, the resulting video is a mosaic video of different attention regions from each original video stream. The attention regions from the original video streams are repositioned and combined into new positions in the resulting video stream.

[0030] Video compression relies on block-based video coding in most coding systems, such as HEVC, short for High Efficiency Video Coding, or the emerging VVC, short for Versatile Video Coding. In these coding systems, video is composed of a sequence of frames, pictures, images, or samples that can be displayed at several different times. In the case of multi-layered video (e.g., scalable, stereo, or 3D video), several pictures can be decoded to compose the resulting image and display it at the same time. A picture can also be composed of different image components, for example, encoding luminance, chrominance, or depth information.

[0031] Compression of video sequences relies on several partitioning techniques for each picture. Figure 2a illustrates several partitioning techniques in a coding system. Pictures 201 and 202 are partitioned into coding tree units (CTUs), indicated by dotted lines. CTUs are the basic units of encoding and decoding. For example, a CTU can encode a region of 128x128 pixels.

[0032] A coding tree unit (CTU) can also be named block, macroblock, or coding block. It can simultaneously code different image components or can be limited to only one image component. If an image contains several components, a CTU corresponds to one CTB for each component. In the following, the invention applies to both the CTU and CTB level.

[0033] As shown in Figure 2a, a picture can be divided according to a grid of tiles, indicated by the thin solid lines. A tile is a portion of a picture and is therefore a rectangular region of pixels that can be defined independently of the CTU division. The boundaries of the tiles and the CTUs can be different. A tile can also correspond to a sequence of CTUs, as in the depicted example, meaning that the boundaries of the tiles and CTUs coincide.

[0034] The tile definition provides that tile boundaries break spatial coding dependencies, meaning that the coding of a CTU within a tile is not based on pixel data from another tile in the picture.

[0035] Some coding systems, such as VVC, provide the concept of slices. This mechanism allows the division of a picture into one or several groups of tiles. Each slice consists of one or several tiles. Two different types of slices are provided, as shown by pictures 201 and 202. The first type of slice is limited to slices that form a rectangular area within the picture. Picture 201 illustrates the division of a picture into five different rectangular slices. The second type of slice is limited to consecutive tiles in raster scan order. Picture 202 illustrates the division of a picture into three different slices consisting of consecutive tiles in raster scan order. Rectangular slices are a structure chosen to handle regions of interest within a video. Slices can be coded into a bitstream as one or several NAL units. A NAL unit, short for Network Abstraction Layer unit, is a logical unit of data for encapsulation of data in a coded bitstream. In the example of the VVC coding system, a slice is coded as a single NAL unit. When a slice is bi-stream coded as several NAL units, each NAL unit of the slice is a slice segment. A slice segment includes a slice segment header that contains the coding parameters of the slice segment. The header of the first segment NAL unit of a slice includes all coding parameters of the slice. The slice segment headers of subsequent NAL units of the slice may contain fewer parameters than the first NAL unit. In such a case, the first slice segment is an independent slice segment and the subsequent segments are dependent slice segments.

[0036] In OMAF v2 ISO / IEC 23090-2, a subpicture is a part of a picture that represents a spatial subset of the original video content, divided into spatial subsets by content creators before video encoding. A subpicture can be, for example, one or more slices that form a rectangular area.

[0037] Figure 2b shows an example of picture division into sub-pictures. A sub-picture represents a picture portion covering a rectangular area of ​​the picture. Each sub-picture can have different size and coding parameters. For example, a different tile grid and slice division can be defined for each sub-picture. In Figure 2b, picture 204 is subdivided into 24 sub-pictures, including sub-pictures 205 and 206. These two sub-pictures further describe divisions in tile grids and slices similar to pictures 201 and 202 in Figure 2a. In the second example, tile and brick divisions are defined at the picture level, rather than per sub-picture. A sub-picture is then defined as one or more slices that form a rectangular area.

[0038] Figure 2c shows an example of partitioning using brick partitioning. Each tile can contain a set of bricks. A brick is a contiguous set of CTUs that form a line within the tile. For example, frame 207 in Figure 2c is divided into 25 tiles. Each tile contains exactly one brick, except for the rightmost column of tiles, which contains two bricks per tile. For example, tile 208 contains two bricks, 209 and 210. When brick partitioning is used, a slice contains either a brick from one tile or several bricks from other tiles. In other words, a VCL NAL unit is a set of bricks instead of a set of tiles.

[0039] FIG. 3 shows the structure of a bitstream in an exemplary coding system VVC.

[0040] A bitstream 300 from a VVC coding system consists of an ordered sequence of syntax elements and coded data. The syntax elements and coded data are arranged in NAL units 301-305. There are different NAL unit types. The network abstraction layer provides the ability to encapsulate the bitstream in different protocols, such as RTP / IP, which stands for Real Time Protocol / Internet Protocol, ISO Base Media File Format, etc. The network abstraction layer also provides a framework for packet loss resilience.

[0041] NAL units are divided into VCL NAL units and non-VCL NAL units, where VCL stands for Video Coding Layer. VCL NAL units contain the actual coded video data. Non-VCL NAL units contain additional information, such as parameters required for decoding the coded video data or supplementary data that can improve the usability of the decoded video data. NAL unit 305 corresponds to a slice and constitutes the VCL NAL units of the bitstream. Different NAL units 301-304 correspond to different parameter sets; these NAL units are non-VCL NAL units. VPS NAL unit 301, for which VPS stands for Video Parameter Set, contains parameters defined for the entire video, i.e., the entire bitstream. The naming of VPS can vary; for example, in VVC it ​​becomes DPS. Alternatively, VPS and DPS are different parameter set NAL units. DPS (Decoder Parameter Set) NAL units can define more static parameters than those in the VPS. In other words, the parameters of a DPS change less frequently than the parameters of a VPS. The SPS NAL unit 302, where SPS stands for Sequence Parameter Set, contains parameters defined for a video sequence. In particular, the SPS NAL unit can define subpictures of a video sequence. The syntax of the SPS includes, for example, the following syntax elements: [Table 1] The descriptor sequence gives the encoding of the syntax element, where u(1) means that the syntax element is encoded using 1 bit, and ue(v) means that the syntax element is encoded using a variable length encoding, left-bit first, unsigned integer zeroth order Exp-Golomb encoded syntax element.

[0042] The presence of subpictures in a picture depends on the value of subpics_present_flag. If this flag is equal to 0, it indicates that the picture does not contain subpictures. If it is equal to 1, the set of syntax elements specifies the subpictures in the frame. The syntax element max_subpics_minus1 specifies the maximum number of subpictures in a picture of the video sequence. The SPS then defines the subpictures divided into a grid of subpicture grid elements of size defined by subpic_grid_col_width_minus1 and subpic_grid_row_height_minus1. Each grid element specifies a subpicture index, subpic_grid_idx[i][j] (where i and j are the coordinates of the element within the rectangle). The value of subpic_grid_idx[i][j] is the same number as the number of subpictures in the picture of the video sequence. The subpic_grid_idx[i][j] syntax element is the identifier of the subpicture. All grid elements that share the same index value form a rectangular area corresponding to the subpicture whose index (or identifier) ​​is equal to the index value. The subpic_treated_as_pic_flag[i] syntax element indicates whether the subpicture boundary should be treated as a picture boundary other than for loop filter processing. The loop_filter_across_subpic_enabled_flag[i] syntax element indicates whether the loop filter is applied across subpicture boundaries.

[0043] The PPS NAL unit 303, where PPS stands for Picture Parameter Set, contains parameters defined for a picture or a group of pictures. The APS NAL unit 304, where APS stands for Adaptation Parameter Set, typically contains parameters for the loop filter of an adaptive loop filter (ALF) or reshaper model (or luma mapping with chroma scaling model) defined at the slice level. The bitstream may also contain an SEI NAL unit, which stands for Supplemental Enhancement Information. The periodicity with which these parameter sets occur within the bitstream is variable. A VPS defined for the entire bitstream need only occur once within the bitstream. Conversely, an APS defined for a slice can occur once for each slice within each picture. In practice, different slices can rely on the same APS; therefore, there are typically fewer APSs than slices within each picture. If a picture is divided into subpictures, a new parameter set or PPS can be defined for each subpicture or group of subpictures.

[0044] The VCL NAL unit 305 contains each slice, which can correspond to an entire picture or subpicture, a single tile or multiple tiles, or a single brick or multiple bricks. A slice consists of a slice header 310 and a raw byte sequence payload RBSP 311 that contains the bricks.

[0045] The syntax of PPS as proposed in the current version of VVC also includes syntax elements that specify the size of the picture and luma samples, and the division of each picture in tiles, bricks, and slices.

[0046] The syntax of the PPS proposed in the current version of VVC is organized as follows: [Table 2] The descriptor string gives the encoding of the syntax element, where u(1) means that the syntax element is encoded using 1 bit, and ue(v) means that the syntax element is encoded using a left-bit-first unsigned integer zero-order Exp-Golomb encoding syntax element, which is a variable length encoding. The syntax elements pic_width_in_luma_samples and pic_height_in_luma_samples specify the width and height of the picture in luma samples.

[0047] If the number of tiles in a picture is greater than 1 (single_tile_in_pic_flag is equal to 0), the PPS defines several syntax elements (not listed in the table above) that specify the tile division within the frame as a grid of tiles.

[0048] If bricks are present (brick_splitting_present_flag is equal to 1), the PPS contains a loop for each tile in the tile grid, indicating whether the tile is split into bricks. If the tile contains bricks, the brick configuration of the tile is coded in the PPS.

[0049] Slice division is expressed by the following syntax elements:

[0050] The syntax element single_tile_in_pic_flag indicates whether the picture contains a single tile. In other words, if this flag is true, there is only one tile and one slice in the picture.

[0051] single_brick_per_slice_flag indicates whether each slice contains a single brick. In other words, if this flag is true, all the bricks of the picture belong to different slices.

[0052] The syntax element rect_slice_flag indicates that the slices of the picture form a rectangular shape as represented in picture 201.

[0053] If present, the syntax element num_slices_in_pic_minus1 is equal to the number of rectangular slices in the picture minus one.

[0054] Next, a syntax element (not represented in the table above) encodes the position of each slice relative to the brick division in a "for loop" over all slices of the picture, where the index of the slice position parameter is the index of the slice.

[0055] Slice identifiers are specified if signalled_slice_id_flag is equal to 1. In this case, the signalled_slice_id_length_minus1 syntax element indicates the number of bits used to encode each slice identifier value. The slice_id[] association table is indexed by slice index and contains the identifiers of the slices. If signalled_slice_id_flag is equal to 0, slice_id is indexed by slice index and contains the slice index of the slice.

[0056] In summary, the PPS contains syntax elements that allow determining the location of slices within a frame. Since a sub-picture forms a rectangular area within a frame, it is possible to determine the set of slices, tiles and bricks that belong to the sub-picture.

[0057] The slice header contains the slice address in the current VVC version with the following syntax: [Table 3] If the slice is not rectangular, the slice header indicates the number of tiles in the slice NAL unit with the help of the num_tiles_in_slice_minus1 syntax element.

[0058] Each tile 320 may include a tile segment header 330 and tile segment data 331. The tile segment data 331 includes encoded coding blocks 340. In the current version of the VVC standard, there are no tile segment headers, and the tile segment data includes coding block data 340.

[0059] In a variant, the video sequence includes sub-pictures and the syntax of the slice header is as follows: [Table 4] A slice header contains a slice_subpic_id syntax element that specifies the identifier of the subpicture to which it belongs (e.g., corresponding to one of the values ​​of subpic_grid_idx[i][j] defined in the SPS). Consequently, all slices that share the same slice_sub_pic_id in a video sequence belong to the same subpicture.

[0060] For illustrative purposes only, Figure 4 schematically illustrates the quadtree inference mechanism used in VVC for coding tree blocks that cross image boundaries. In VVC, images are not restricted to having widths and heights that are multiples of the coding tree block size. Then, a coding tree block at the right edge of a frame can cross the right image boundary 401, and a coding tree block at the bottom edge of a frame can cross the bottom image boundary 402. In these cases, VVC defines a quadtree inference mechanism for coding tree blocks that cross boundaries. This mechanism recursively splits any coding blocks of coding tree blocks that cross image boundaries until there are no more coding blocks that cross the boundary or until the maximum quadtree depth for these coding tree blocks is reached. For example, coding tree block 403 is not automatically split, while coding tree blocks 404, 405, and 406 are split. There is no signaling of the inferred quadtree; the decoder must infer the same quadtree on the image boundary. However, the automatically obtained quadtree may be further refined for coding tree blocks that are within the frame, for example, by signaling split information for these coding tree blocks (if the maximum quadtree depth is not reached), as in 407.

[0061] When dividing a CTB into coding blocks, there is a smallest coding block that cannot be divided. The size of this smallest square coding block is given by MinCBSizeY. In some examples, MinCBSizeY is equal to 4.

[0062] During encoding, coding blocks of a coding tree located outside the image are typically not coded in the bitstream. During decoding, the decoder uses the same quadtree inference mechanism and knows that these coding blocks are not coded. The decoder can then decode other coding blocks within the image that were coded in the bitstream. The resulting coding tree is coded into the bitstream by the encoder. The decoder relies on this coded coding tree information to correctly identify the blocks to decode and reconstruct the image. In particular, the coded coding tree includes a parameter, such as a parameter called split_cu_flag, that specifies whether the coding unit is split.

[0063] 5a and 5b show the generation of a bitstream in which a border sub-picture is moved to a non-border position.

[0064] In this example, the first bitstream 500 contains four subpictures 1 HQ ~4 HQ This bitstream represents a high quality version of the video. The second bitstream 501 is made up of four subpictures 1 LQ ~4 LQ This bitstream represents a lower quality version of the same video. Bitstream 502 is created by merging and rearranging some subpictures emitted from bitstreams 500 and 501.

[0065] In particular, bitstream 502 contains subpicture 2 HQ , 1 LQ , 4 HQ and 3 LQ where subpicture 2 consists of HQ is moved from its top right position in bitstream 500 to its top left position in bitstream 502, and subpicture 4 HQis moved from the bottom right position in bitstream 500 to the bottom left position in bitstream 502, and subpicture 1 LQ is moved from the top left position in bitstream 501 to the top right position in bitstream 502, and subpicture 3 LQ is moved from the bottom left position in bitstream 501 to the bottom right position in bitstream 502.

[0066] Bitstream 500 is composed of pictures with widths that are not multiples of the coding tree block (CTB) size, as shown in Figure 5b. Therefore, the rightmost coding tree blocks are subject to inferred splitting. The coding blocks at the right end of these CTBs are not coded in the bitstream. These uncoded coding blocks are represented by hatched area 503. In bitstream 502, subpicture 2 HQ and 4 HQ An uncoded coding block 504 is placed in the center of the image on the right boundary of the image.

[0067] The size of the pixels (or luma samples) of the image is coded into the bitstream. This information allows the decoder to know the image boundaries and estimate the division of the rightmost CTB within the image. The decoder then knows which coding blocks are not coded and can decode the coding blocks that make up the image.

[0068] Subpicture 2 HQ and 4 HQA problem occurs when a subpicture such as (1) is moved from the rightmost part of the image to another location, as in bitstream 502. The size of the subpicture is coded in the bitstream as an integer number of CTBs. Because of this peculiarity in the standard, it is impossible for a decoder to infer the right boundary of the subpicture and the corresponding inferred division of the CTBs at the rightmost part of the subpicture. The decoder expects all coding blocks of the CTBs at the rightmost part of the subpicture to be coded in the bitstream. Since some of them are missing, decoding fails.

[0069] There is a need to find a way to rearrange the subpictures within an image without re-encoding them.

[0070] This problem can be solved by proposing an encoding method that includes encoding information that allows a decoder to infer the division of CTBs located at the right and bottom of a subpicture whose width and height are not multiples of the CTB size when the subpictures are not located at the right and bottom of the image. A corresponding decoding method for the generated bitstream is also proposed.

[0071] A subpicture is a rectangular region within a picture represented by one or more slices. Subpictures are defined, for example, in SPS or PPS NAL units. For each subpicture, the encoder may indicate that its boundary is treated as a picture boundary, meaning that the subpictures can be decoded independently of each other. Intra- and inter-prediction processes are constrained to use only prediction information from the same subpicture in the current and reference frames. Subpicture boundary filtering is controlled by a flag defined for each subpicture. This flag allows applying or not applying a loop filter at the subpicture boundary. When decoding a CTB from a slice, the decoder determines the index or identifier of the subpicture to which the CTB belongs.

[0072] In the following, the use of sub-pictures allows for decoding regions of a video sequence at different positions, and potentially also for moving sub-pictures at new decoding positions. For this reason, in a preferred embodiment, sub-picture boundaries are treated as picture boundaries (to ensure that motion prediction is constrained to allow independent decoding of sub-pictures), and the loop filter is disabled at sub-picture boundaries. For VVC, these two conditions are met when the sub-picture's flags subpic_treated_as_pic_flag[i] and loop_filter_across_subpic_enabled_flag[i] are equal to 1 and 0, respectively. For HEVC, a sub-picture is, for example, a set of slices that belong to a motion-constrained tile set according to the HEVC specification.

[0073] FIG. 6 illustrates a method for encoding pictures of a video into a bitstream according to a first aspect of the invention.

[0074] In a first step 601, the picture is divided into sub-pictures.

[0075] In step 602, the size of the subpictures is determined. The width and height of each subpicture is a function of the regions of interest present in the input video sequence. Typically, each subpicture is sized to contain one region of interest. The size of the subpicture is determined in pixels (or luma samples). If the size of the subpicture is not a multiple of the size of the CTB, it is determined that the subpicture requires a specific split inference process. In particular, this can occur when the encoder merges two bitstreams as shown in Figure 5a.

[0076] In step 603, the size of the subpicture is coded into the bitstream. This size is information indicating the location of the subpicture's partition inference boundary if a partition inference process is required. If necessary, information indicating the need for partition inference process for this subpicture is coded into the bitstream. The information provided in the bitstream in this step allows the decoder to perform the partition inference process.

[0077] In step 604, at least one slice constituting the subpicture is encoded into a bitstream. If a partition inference process is required, the encoder divides the subpicture into two parts. The first part is the area between the left and top boundaries of the subpicture and the horizontal and vertical partition inference boundaries. The second part is the area between the partition inference boundaries and the bottom and right boundaries of the subpicture. For example, the white area in the 2HQ subpicture of Figure 5b corresponds to the first part, and the hatched area corresponds to the second part. In this document, the first part of the subpicture is referred to as the useful part of the subpicture. Coding units outside the useful part of the subpicture due to this partition inference process are not coded.

[0078] In an alternative embodiment, coding blocks that lie outside the useful portion of the sub-picture are encoded with padding data.

[0079] In another alternative embodiment, a flag is inserted into the bitstream to indicate whether coding blocks outside the useful portion of the subpicture are coded with padding data or not. Typically, an encoder specifies in a picture parameter set NAL unit that a subpicture includes coded coding blocks of padding data for coding units that are outside the useful portion of the subpicture. For example, a flag is associated with each subpicture identifier to specify whether to provide padded coding data for coding units outside the useful portion of the subpicture.

[0080] In some embodiments, the use of the split inference process for each tile can be known from the bitstream context, and in these embodiments, no flag needs to be coded to signal the use of the sub-picture split inference process.

[0081] It is proposed to introduce a new syntax element in one parameter set NAL unit, e.g., SPS, which when moved to different positions in the merged bitstream allows to obtain the same split inference for CTBs at the right and / or bottom boundaries of one subpicture.

[0082] These syntax elements allow the decoder to determine the size of the skipped coding block in the last CTB row and / or last CTB column of the sub-picture when the sub-picture boundary is treated as a picture boundary.

[0083] As a result, when moving a subpicture from its rightmost position in a picture to another position, the merging or encoding operation consists of determining the size of the skipped coding block, for example, in the last CTB row and column, and then specifying it in the SPS associated with the subpicture. Decoding of the merged bitstream uses these values ​​to determine the usable size of the subpicture.

[0084] For example, the syntax of the SPS includes the following syntax elements: [Table 5] According to the proposed embodiment, some new syntax elements are introduced, which are marked in bold in the table. The semantics of these new syntax elements are as follows:

[0085] subpic_split_inference_flag equal to 1 indicates that the last CTB row or last CTB column of the subpicture is not complete. The inference process for splitting the CTBs of the last CTB row and column of the subpicture takes into account the available size of the subpicture to determine the value of split_cu_flag.

[0086] subpic_split_inference_flag equal to 0 indicates that all CTBs of the subpicture are complete. Split inference processing is not required for the CTBs on the last CTB row and column of the subpicture.

[0087] If present, subpic_split_inference_ctb_width[i] specifies the actual width of the CTB in the rightmost column of the CTB of the ith subpicture (i.e., the width in pixels of the CTB in the useful part of the ith subpicture). The value of subpic_split_inference_ctb_width[i] is specified in units of MinCbSizeY-wide coding blocks and can be in the range [0, CtbSizeY / MinCbSizeY-1]. If not present, the value of subpic_split_inference_ctb_width[i] is inferred to be equal to CtbSizeY / MinCbSizeY, which corresponds to the width of the CTB. This coding syntax element is coded using a fixed-length code, for example, 7 bits or log2(CtbSizeY / MinCbSizeY-1,). Exp-Golomb coding can also be used.

[0088] If present, subpic_split_inference_ctb_height[i] specifies the actual height of the CTB in the last row of the CTB for the ith subpicture. The value of subpic_split_inference_ctb_height[i] is specified in coding block units of MinCbSizeY height and can be in the range [0, CtbSizeY / MinCbSizeY-1]. If not present, the value of subpic_split_inference_ctb_height[i] is inferred to be equal to CtbSizeY / MinCbSizeY, which corresponds to the CTB height. This coding syntax element is coded using a fixed-length code, e.g., 7 bits or log2(CtbSizeY / MinCbSizeY-1,). Exp-Golomb coding is also possible.

[0089] The subpic_split_inference_flag allows the encoder to indicate to the decoder that a specific process is required to process the CTBs in the last CTB rows and columns of some subpictures. The rightmost and / or bottommost CTB rows are meant to contain coding blocks that the encoder has not coded. The division of these CTBs into blocks is handled by a coding tree partitioning process that can infer the division of each block as a function of the values ​​of the subpic_split_inference_ctb_width[i] and subpic_split_inference_ctb_height[i] syntax elements for the i-th subpicture.

[0090] The actual size of the right- and bottom-most CTBs of the ith subpicture (specified by subpic_split_inference_ctb_width[i] and subpic_split_inference_ctb_height[i]) is in units of the minimum coding block size (MinCbSizeY, MinCbSizeY) specified in the SPS. The actual size of the CTB cannot exceed the maximum coding tree block size specified in the SPS (CtbSizeY, CtbSizeY). As a result, the range of subpic_split_inference_ctb_height and subpic_split_inference_ctb_width is 0 to 1 minus the ratio of the maximum CTB size to the minimum coding block size.

[0091] In one alternative, to simplify the analysis of the SPS, subpic_split_inference_ctb_width[i] and subpic_split_inference_ctb_height[i] are expressed in units of several luma samples (usually four luma samples). In fact, the size of the CTB and the minimum size of the coding block can be specified in the SPS and defined after the subpic_split_inference_ctb_width[i] and subpic_split_inference_ctb_height[i] syntax elements. In this case, the position of the split inference boundary is expressed as an integer multiple of the size of the smallest coding block in the coding tree block. Furthermore, if subpic_split_inference_ctb_width[i] and subpic_split_inference_ctb_height[i] are defined in different parameter set NAL units, the actual size of each subpicture can be calculated regardless of the CTB and coding block sizes.

[0092] The following pseudocode determines the boundary of the ith subpicture that triggers a split in the coding tree: The position of the vertical boundary is represented by the variable SubPicRightSplitInferenceBoundary[i], and the horizontal boundary is represented by the variable SubPicBotSplitInferenceBoundary[i].

[0093] The position, in luma samples, of the vertical border of the ith subpicture is equal to the position of the CTB-aligned right border of the ith subpicture minus the subpic_split_inference_ctb_width[i] syntax element defined in the SPS. The position, in luma samples, of the horizontal border of the ith subpicture is equal to the position of the CTB-aligned bottom border of the ith subpicture minus the subpic_split_inference_ctb_height[i] syntax element defined in the SPS.

[0094] The variables SubPicRightSplitInferenceBoundary[i] and SubPicBotSplitInferenceBoundary[i] are derived as follows: for (i=0;i <NumSubPics;i++){ SubPicRightSplitInferenceBoundary[i]=(SubPicLeft[i]+SubPicWidth[i])*(subpic_grid_col_width_minus1+1)*4-CtbSizeY+subpic_split_inference_ctb_width[i]*MinCbSizeY; SubPicBotSplitInferenceBoundary[i]=(SubPicTop[i]+SubPicHeight[i])*(subpic_grid_row_height_minus1+1)*4-CtbSizeY+subpic_split_inference_ctb_height[i]*MinCbSizeY;} Therefore, the proposed new syntax elements related to subpictures, subpic_split_inference_flag, subpic_split_inference_ctb_width[i], and subpic_split_inference_ctb_height[i], allow the encoder to indicate to the decoder the information it needs to be able to infer the splitting of incomplete CTBs and to determine coding blocks that lie outside the useful part of the subpicture.

[0095] The coding tree syntax includes a flag, split_cu_flag, that indicates whether a coding block is split. A split is inferred when a coding tree block is not entirely within the bounds of the useful portion of a subpicture, and as a result, the flag is not coded.

[0096] split_cu_flag equal to 0 specifies that the coding block is not split. split_cu_flag equal to 1 specifies that the coding block is split into four coding blocks using quad partitioning, as indicated by the syntax element split_qt_flag, or into three coding blocks using ternary partitioning, or into two coding blocks using binary partitioning, as indicated by the syntax element mtt_split_cu_binary_flag. The binary or ternary partitioning can be either vertical or horizontal, as indicated by the syntax element mtt_split_cu_vertical_flag.

[0097] If split_cu_flag is not present in the encoding of a coding block, the value of split_cu_flag is inferred as follows:

[0098] The value of split_cu_flag is inferred to be equal to 1 if one or more of the following conditions are true: x0+cbWidth is greater than SubPicRightSplitInferenceBoundary[SubPicIdx].

[0099] y0+cbHeight is greater than SubPicBotSplitInferenceBoundary[SubPicIdx].

[0100] Otherwise, the value of split_cu_flag is inferred to be equal to 0.

[0101] The above pseudocode determines whether a coding block located at position (x0, y0) in luma sample coordinates with width (or height) equal to cbWidth (or cbHeight) is completely inside the useful portion of the current subpicture with index equal to SubPicIdx. If it is outside this region, the encoder (and decoder) infers that the coding block is split (split_cu_flag equals 1).

[0102] When a coding block is inferred to be split because its boundary lies outside the useful portion of the actual pixel boundary of the subpicture, there are three possible partitions: the first is a quad-tree partition, i.e., the coding block is further divided into four square coding blocks of equal size; the second is a partition into two coding blocks separated by either a vertical or horizontal boundary across the entire coding block; this partition is a binary tree partition; and finally, the third partition is a ternary tree partition, which divides the coding block into three blocks separated by a vertical or horizontal boundary.

[0103] In addition to the split inference process, the encoder (and similarly the decoder) applies a specific encoding process to coding blocks outside the useful portion of the subpicture. In the first embodiment, these coding blocks are not encoded. For each split coding block, the encoder checks whether the y-axis (or x-axis) coordinate of the block's top boundary (or left boundary) is equal to the SubPicBotSplitInferenceBoundary (or SubPicRightSplitInferenceBoundary) value of the current subpicture. In such a case, the encoder skips encoding the block. As a result, these coding blocks are not provided with coded syntax elements. Consequently, when encoding neighboring coding units, these skipped coding blocks are considered unavailable for inter-prediction and intra-prediction. Similarly, temporal prediction is restricted to avoid using pixel information from these skipped coding blocks.

[0104] Figure 7 shows the general decoding process of one embodiment of the present invention. In step 701, the decoder determines the size of the subpictures of a frame, typically their width and height, by parsing the SPS and PPS of the stream. In particular, it determines the slice NAL units belonging to each subpicture. In step 702, the decoder decodes the slices of the subpictures that form the picture. The decoding process of the coded CTBs in each slice is explained in more detail with reference to Figure 8. Step 703 consists of generating an output picture from the decoded picture. In particular, the decoder can optionally apply some cropping operations to the decoded subpictures.

[0105] Figure 8 shows the process 702 for decoding CTBs coded into slices. In step 801, it is checked whether any slices remain to be decoded. If all slices have been decoded, the process ends. For a given slice, the decoder determines in step 802 the identifier of the subpicture to which the slice belongs.

[0106] Next, in step 803, it is checked whether any CTBs remain to be decoded within the current slice. For a given CTB, in step 804, the decoder parses the syntax elements of the CTB. If the CTB is on the rightmost or bottommost boundary of the subpicture with the identifier determined in 802, the decoder applies a split inference process to the CTB if the process is required for the current slice. Next, once the CTB is split, in step 805, the decoder decodes each coding block resulting from the inference process that is within the subpicture. Coding blocks that are outside the useful portion of the subpicture are not decoded because they are not present in the bitstream. In some embodiments, coding blocks that are outside the useful portion of the subpicture are decoded if they are coded with padding data.

[0107] The subpicture decoding process results in subpictures whose size is not a multiple of the CTB size in the right column or row below the incomplete CTB. When composing the subpicture to compose the final frame for rendering, the decoder must determine the values ​​of the pixels in these bands of pixels. An example of a band of undetermined pixels is shown in Figure 5b, band 504. This is because subpictures in a frame are restricted to being composed of an integer number of CTBs. As a result, the subpicture contains a set of matching decoded pixels corresponding to the region between the subpicture origin and the split inference boundary. The remaining region corresponds to undefined or pixel values ​​not intended to be displayed. In some embodiments, the decoder may set the values ​​of pixels in the remaining bands to zero. In some other embodiments, the decoder may set the values ​​of pixels in the remaining bands by replicating the value of the closest matching pixel.

[0108] Since these bands of undetermined pixels are not intended to be displayed, they can be suppressed in the resulting frame by shifting adjacent subpictures. For example, subpicture 1 in bitstream 502 of Figure 5b LQ and 3 LQ may be left shifted by the width of band 504.

[0109] In this embodiment, the decoder shifts the decoded pixels of the subpictures to the right and below the current subpicture so that they align with the right and bottom inferred split boundaries. The decoding position in the luma samples of the coding block takes into account that some subpictures are not complete.

[0110] The decoder has all the information necessary to identify the bands of undetermined pixels, and in one embodiment, in post-decoding processing after the decoder has decoded all the sub-pictures, it proceeds to shift the sub-pictures to remove these bands.

[0111] In another embodiment, the decoder can integrate the shifting operation in the decoding process, where a new variable is defined by the decoder and used during the decoding of each CTB to correctly decode the coding block in its final position.

[0112] The SubPicOffsetX[i] and SubPicOffsetY[i] variables indicate the offsets to be subtracted from the X-axis and Y-axis coordinates, respectively, of the pixel of the coding block of the ith subpicture to be placed in the correct position.

[0113] For example, the xCtbShifted and yCtbShifted variables are the new coordinates of the top left pixel of the CTB at the CtbAddrInRs address in the raster scan order of the picture after the shift operation.

[0114] The following pseudocode calculates these coordinates, where CtbAddressInRs is the raster scan address of the CTB in the picture, PicWidthInCtbsY is the width of the picture in CTB units, CtbLog2SizeY is the log2 of the size of the CTB in luma samples, and SubPicIdx is the index of the subpicture containing the current CTB.

[0115] xCtbShifted=(CtbAddrInRs % PicWidthInCtbsY)< <CtbLog2SizeY yCtbShifted=(CtbAddrInRs / PicWidthInCtbsY)< <CtbLog2SizeY if(subpic_split_inference_flag){ xCtbShifted-=SubPicOffsetX[SubPicIdx] yCtbShifted-=SubPicOffsetY[SubPicIdx] } The first two lines calculate the x-axis and y-axis coordinates of the origin of the coding tree block in luma sample units. These values ​​do not take into account skipped pixels, if any, in adjacent subpictures. Next, if subpic_split_inference_flag is equal to 1, indicating that the subpicture may contain skipped blocks, the sum of the widths of the skipped pixels of the subpictures located to the left (or above) of the coding tree block is subtracted from the xCtbShifted (or yCtbShifted) coordinate.

[0116] The encoder determines the shift offsets for the x-axis (SubPicOffsetX[i]) and y-axis (SubPicOffsetY[i]) coordinates of each SubPicture as follows: First, for a given subpicture, it determines the set of subpictures to the left of the current subpicture. For each subpicture in this set, the shift offset is equal to the difference between the right boundary of the last CTB row and the right subpicture inference boundary. Typically, this value is equal to CtbSizeY-subpic_split_inferernce_ctb_width[j]*MinCbSizeY, where j is the index of the subpicture in the set and CtbSizeY is the size of the CTB in luma samples or pixels. The value of the shift offset for the x-axis is equal to the sum of CtbSizeY-subpic_split_inferernce_ctb_width[j]*MinCbSizeY for each j equal to the index of the subpicture in the set of left subpictures. Similarly, the value of the y-axis shift offset is equal to the sum of CtbSizeY-subpic_split_inferernce_ctb_height[j]*MinCbSizeY for each j equal to the index of the subpicture in the set of upper subpictures above the current subpicture.

[0117] Figure 9 shows a picture divided into 16 subpictures, where the bands of undetermined pixels are represented by hatched bands. The arrows represent the shifting of the subpictures to remove the bands of undetermined pixels. If all vertical bands, and each horizontal band, have the same size, the resulting image can be considered rectangular. It is also true that each column of a subpicture contains the same number of horizontal bands, and each row of a subpicture contains the same number of vertical bands.

[0118] It should be understood that if these conditions are not met, the resulting image will not be rectangular.

[0119] In one embodiment, the possibility of specifying different inferred boundaries for subpictures is constrained to avoid defining pictures with non-rectangular picture shapes after shifting operations: in particular, the height (or width) in luma samples of each column (or row) of a subpicture must be equal after shifting or cropping of the decoded frame.

[0120] The possibility of specifying different inference boundaries in a given row (or column) of a subpicture can lead to complex operations of shifting the decoded samples in each subpicture using different shift offsets. For this reason, in one embodiment, a constraint can be defined that the subpicture inference boundaries be aligned across the picture. For example, Figure 10 shows the case where a picture is divided into 16 subpictures that respect this constraint.

[0121] FIG. 11 illustrates the concept of a picture-wide boundary.

[0122] Alternatively, the encoder determines two types of boundaries for subpictures. First, a subpicture boundary is aligned with other subpicture boundaries that collectively span the entire picture width or height. This type of subpicture boundary is called a picture-wide boundary. For example, the bottom boundary of Subpicture 0 in Figure 11 is a picture-wide boundary. Conversely, the bottom boundary of Subpicture 1 is not picture-wide.

[0123] In one embodiment, the encoder constrains the subpicture arrangement so that the last CTB row of a subpicture with a bottom boundary that is not picture-wide does not use an inferred split. In other words, the split inference boundary is aligned to the CTB boundary, e.g., the size of subpic_split_inference_ctb_height[i] is equal to the size of the CTB. The same principle applies to the last CTB column of a subpicture with a right boundary that is picture-wide. As an alternative, the encoder determines the type of boundary for each subpicture and encodes the value of subpic_split_inference_ctb_width[i] (or subpic_split_inference_ctb_height[i]) if the right (or bottom) boundary of the subpicture is a picture-wide boundary. If the right (or bottom) boundary is not picture-wide, subpic_split_inference_ctb_width[i] (or subpic_split_inference_ctb_height[i]) is not coded and is inferred to be equal to the size of the CTB.

[0124] The encoder may apply a second constraint to subpictures whose bottom (or right) boundary is the same picture-wide boundary. For example, this constraint may be that subpic_split_inference_height[i] is equal for all i-th subpictures in a set of subpictures when the boundary is below the subpicture. Similarly, subpic_split_inference_width[i] may be equal for all i-th subpictures in a set of subpictures that have a common right picture-wide boundary.

[0125] Additionally, the encoder may apply a third constraint: that the split inference boundary be aligned with the CTB boundary when the subpictures are not independently decodable. Typically, this is the case when subpic_treated_as_pic_flag[i] is equal to 0 or loop_filter_across_subpic_enabled_flag[i] is equal to 1 for the i-th subpicture. This constraint ensures that the split inference mechanism is used only in the context of subpicture merging. Next, we describe an embodiment that complies with these constraints.

[0126] In another embodiment, the syntax of the SPS is modified to specify the location of the vertical split inference boundary and the location of the horizontal split inference boundary across the entire picture. In this embodiment, picture-level signaling is used because the sub-picture boundaries that are subject to split inference relate only to picture-wide boundaries.

[0127] There are several alternatives for specifying the location of these boundaries: The location of the split inference boundary can be signaled in a parameter set NAL unit such as an SPS or PPS. The parameter set must describe information related to subpicture and picture information.

[0128] In one alternative, the position of the inference boundary is defined relative to the origin of the picture, for example in luma sample units.

[0129] For example, the PPS syntax includes the following elements: [Table 6] where: pps_subpic_split_inference_flag is a flag indicating the use of subpicture split inference. Typically, split inference is applied to the picture-wide boundary. pps_subpic_split_inference_flag equal to 1 indicates the presence of pps_split_inference_boundary_pos_x and pps_split_inference_boundary_pos_y, while pps_subpic_split_inference_flag equal to 0 indicates the absence of pps_split_inference_boundary_pos_x and pps_split_inference_boundary_pos_y.

[0130] pps_split_inference_boundary_pos_x is used to calculate the value of PpsSplitInferenceBoundaryPosX, which specifies the position of the vertical split inference boundary in luma samples. pps_split_inference_boundary_pos_x is in the range of 1 to Ceil(pic_width_in_luma_samples÷4)-1. This coding syntax element is coded, for example, using a fixed-length code equal to 13 bits or log2(pic_width_in_luma_samples / 4). Exp-Golomb coding is also possible.

[0131] The position of the vertical split inference boundary PpsSplitInferenceBoundaryPosX is derived as follows: PpsSplitInferenceBoundaryPosX=pps_split_inference_boundary_pos_x * 4 pps_split_inference_boundary_pos_y is used to calculate the value of PpsSplitInferenceBoundaryPosY, which specifies the position of the horizontal split inference boundary in luma samples. pps_split_inference_boundary_pos_y is in the range from 1 to Ceil(pic_height_in_luma_samples÷4)-1. This coding syntax element is coded, for example, using a fixed-length code equal to 13 bits or log2(pic_width_in_luma_samples / 4). It is also possible to use Exp-Golomb coding.

[0132] The position of the horizontal split inference boundary PpsSplitInferenceBoundaryPosY is derived as follows: PpsSplitInferenceBoundaryPosY=pps_split_inference_boundary_pos_y * 4; The coding tree split inference process compares the coding unit boundary with the location of the split inference boundary. For example, a coding unit split is inferred if the right or bottom boundary of the coding block is greater than one of the split inference boundaries that cross the current subpicture. As a result, the split inference for split_cu_flag is, for example, as follows: If split_cu_flag is not present, the value of split_cu_flag is inferred as follows: - The value of split_cu_flag is inferred to be equal to 1 if one or more of the following conditions are true: - x0+cbWidth is PpsSplitInferenceBoundaryPosX and (SubPicLeft[SubPicIdx])*(subpic_grid_col_width_minus1 + 1)*4) <PpsSplitInferenceBoundaryPosXおよび(SubPicLeft[SubPicIdx]+ SubPicWidth[i])*(subpic_grid_col_width_minus1 + 1)* 4)> Greater than PpsSplitInferenceBoundaryPosX - y0+cbHeight is PpsSplitInferenceBoundaryPosY and (SubPicTop[SubPicIdx])*(subpic_grid_row_height_minus1 + 1)*4) <PpsSplitInferenceBoundaryPosYおよび(SubPicTop[SubPicIdx]+SubPicHeight[i])*(subpic_grid_row_height_minus1+1)*4)> Greater than PpsSplitInferenceBoundaryPosY - Otherwise, the value of split_cu_flag is inferred to be equal to 0.

[0133] Another equivalent algorithm for determining the value of split_cu_flag is as follows:

[0134] If split_cu_flag is not present, the value of split_cu_flag is inferred as follows: - The value of split_cu_flag is inferred to be equal to 1 if one or more of the following conditions are true: - SubPicInferenceSplitFlag is equal to 0 and x0+cbWidth is greater than pic_width_in_luma_samples.

[0135] - SubPicInferenceSplitFlag is equal to 0 and y0+cbHeight is greater than pic_height_in_luma_samples.

[0136] - SubPicInferenceSplitFlag is equal to 1 and x0+cbWidth is greater than SubPicInferenceBoundaryPosX.

[0137] - SubPicInferenceSplitFlag is equal to 1 and y0+cbHeight is greater than SubPicInferenceBoundaryPosY.

[0138] - Otherwise, the value of split_cu_flag is inferred to be equal to 0.

[0139] SubPicInferenceSplitFlag is equal to 1 if the subpicture containing the coding block intersects the split inference boundary. SubPicInferenceBoundaryPosX is the horizontal coordinate of the vertical split inference boundary that intersects the current subpicture. If the subpicture does not intersect a vertical split inference boundary, SubPicInferenceBoundaryPosX is set to a value greater than or equal to the horizontal coordinate of the subpicture's right boundary to avoid unnecessary split inference. Similarly, SubPicInferenceBoundaryPosY is the vertical coordinate of the horizontal split inference boundary that intersects the current subpicture, if any. Otherwise, if the subpicture does not intersect a horizontal split inference boundary, SubPicInferenceBoundaryPosY is set to a value greater than or equal to the horizontal coordinate of the subpicture's right boundary. The factor "4" is introduced because the split inference boundary is constrained to correspond to the smallest coding block boundary. Assume that the smallest coding block size is 4x4. If the size of the smallest coding block is different, different factors can be used in these equations.

[0140] In another example, the syntax elements mentioned above in the PPS can be defined at the SPS level, which may be advantageous if the split inference boundary does not change with each new PPS NAL unit. The syntax of the SPS includes, for example, the following elements: [Table 7] Here, subpic_split_inference_flag, split_inference_boundary_pos_x, and split_inference_boundary_pos_y have similar semantics as pps_subpic_split_inference_flag, pps_split_inference_boundary_pos_x, and pps_split_inference_boundary_pos_y.

[0141] subpic_split_inference_flag is a flag indicating the use of subpicture split inference. When subpic_split_inference_flag is equal to 1, it indicates the presence of split_inference_boundary_pos_x and split_inference_boundary_pos_y, and when subpic_split_inference_flag is equal to 0, it indicates the absence of split_inference_boundary_pos_x and split_inference_boundary_pos_y.

[0142] split_inference_boundary_pos_x is used to calculate the value of SplitInferenceBoundaryPosX, which specifies the position of the vertical split estimation boundary in luma samples. split_inference_boundary_pos_x can be in the range of 1 to Ceil(pic_width_in_luma_samples÷4)-1.

[0143] The position of the vertical split inference boundary SplitInferenceBoundaryPosX is derived as follows: SplitInferenceBoundaryPosX=split_inference_boundary_pos_x * 4 split_inference_boundary_pos_y is used to calculate the value of SplitInferenceBoundaryPosY, which specifies the position of the horizontal split inference boundary in luma samples. split_inference_boundary_pos_y can be in the range of 1 to Ceil(pic_height_in_luma_samples÷4)-1.

[0144] The position of the horizontal split inference boundary SplitInferenceBoundaryPosY is derived as follows: SplitInferenceBoundaryPosY=split_inference_boundary_pos_y * 4 The Split_inference_boundary_pos_x and split_inference_boundary_pos_y syntax elements are coded using a fixed length code, e.g., equal to 13 bits or log2(pic_width_in_luma_samples / 4) for Split_inference_boundary_pos_x and equal to log2(pic_height_in_luma_samples / 4) for Split_inference_boundary_pos_x. It is also possible to use Exp-Golomb codes.

[0145] In another embodiment, split_inference_boundary_pos_y may range from 1 to Ceil(pic_height_in_luma_samples÷4), and split_inference_boundary_pos_x may range from 1 to Ceil(pic_width_in_luma_samples÷4). In such a case, the maximum value of the range indicates that the split inference boundary is aligned with or off the picture boundary. Consequently, when set to the maximum value, it indicates that no vertical or horizontal split boundary is used. In a variant, two separate flags are used, one for each vertical and horizontal split inference boundary. In this variant, the presence of split_inference_boundary_pos_x and split_inference_boundary_pos_y is conditional on these flag values.

[0146] In another embodiment, the PPS defines one or more horizontal and vertical split inference boundaries. The encoder specifies the number of horizontal split inference boundaries and the number of vertical split inference boundaries. These boundaries are limited to picture-wide boundaries. For example, the syntax of the PPS includes the following elements: [Table 8] pps_num_ver_split_inference_boundaries specifies the number of pps_split_inference_boundaries_pos_x[i] syntax elements present in the PPS. If pps_num_ver_split_inference_boundaries is not present, it is inferred to be equal to 0.

[0147] pps_split_inference_boundaries_pos_x[i] is used to calculate the value of PpsSplitInferenceBoundaryPosX[i], which specifies the position of the ith vertical split inference boundary in luma samples. pps_split_inference_boundary_pos_x[i] can be in the range of 1 to Ceil(pic_width_in_luma_samples÷4)-1. This coding syntax element is coded, for example, using a fixed-length code equal to 13 bits or log2(pic_width_in_luma_samples / 4). Exp-Golomb coding is also possible.

[0148] The position of the i-th vertical split inference boundary PpsSplitInferenceBoundaryPosX[i] is derived as follows:

[0149] PpsSplitInferenceBoundaryPosX[i]=pps_split_inference_boundary_pos_x[i]* 4 The distance between any two vertical split inference boundaries can be greater than or equal to CtbSizeY luma samples, which is the size of the CTB in luma samples.

[0150] pps_num_hor_split_inference_boundaries specifies the number of pps_split_inference_boundaries_pos_y[i] syntax elements present in the PPS. If pps_num_hor_split_inference_boundaries is not present, it is inferred to be equal to 0.

[0151] pps_split_inference_boundaries_pos_y[i] is used to calculate the value of PpsSplitInferenceBoundaryPosY[i], which specifies the position of the ith horizontal split inference boundary in luma samples. pps_split_inference_boundary_pos_y[i] can be in the range of 1 to Ceil(pic_height_in_luma_samples÷4)-1. This coding syntax element is coded, for example, using a fixed-length code equal to 13 bits or log2(pic_height_in_luma_samples / 4). Exp-Golomb coding is also possible.

[0152] The position of the i-th horizontal split inference boundary PpsSplitInferenceBoundaryPosY[i] is derived as follows:

[0153] PpsSplitInferenceBoundaryPosX[i]=pps_split_inference_boundary_pos_y[i]* 4 The distance between any two horizontal split inference boundaries may be greater than or equal to CtbSizeY luma samples, which is the size of the CTB in luma samples.

[0154] In another embodiment, the encoder specifies the location of picture-wide split inference boundaries / boundaries relative to the grid formed by the CTBs of the picture. Typically, the SPS or PPS includes syntax elements indicating their location as CTB row or CTB column indices and the number of vertical and horizontal boundaries for the split inference process. For each CTB row (or column) described in the PPS, the encoder indicates the width (or height) of the CTB for the inference process.

[0155] In a variation of this embodiment, the encoder and decoder can determine picture-wide subpicture boundaries from the subpicture definition. In particular, the encoder determines NumVerSplitInferenceBoundaries, the number of vertical picture-wide subpicture boundaries, and NumHorSplitInferenceBoundaries, the number of horizontal picture-wide subpicture boundaries. It also determines the x-axis position of the i-th (raster scan order) vertical picture-wide subpicture boundary and stores it in the PictureWideSubPictureBoundaryPosX[i] variable. It also determines the y-axis position of the i-th (raster scan order) horizontal picture-wide subpicture boundary and stores it in the PictureWideSubPictureBoundaryPosY[i] variable.

[0156] Next, the encoder encodes the width (or height) used by the split inference of the CTB row (or column) to the left (or above) of the vertical (or horizontal) picture-picture-wide sub-picture boundary.

[0157] For example, a PPS contains the following syntax elements: [Table 9]

[0158] where: pps_split_inference_ctb_width[i] is used to calculate the value of PpsSplitInferenceBoundaryPosX[i], which specifies the position of the i-th vertical split inference boundary in luma samples. pps_split_inference_ctb_width[i] is specified in units of MinCbSizeY-width coding blocks and may be in the range 0..CtbSizeY / MinCbSizeY-1. This coding syntax element is coded, for example, using a fixed-length code equal to log2(CtbSizeY / MinCbSizeY-1). Exp-Golomb coding is also possible.

[0159] The position (in luma samples) of the i-th vertical split inference boundary PpsSplitInferenceBoundaryPosX[i] is derived as follows:

[0160] PpsSplitInferenceBoundaryPosX[i]=PictureWideSubPictureBoundaryPosX[i]-CtbSizeY + pps_split_inference_ctb_width[i]*MinCbSizeY pps_split_inference_ctb_height[i] is used to calculate the value of PpsSplitInferenceBoundaryPosY[i], which specifies the position of the i-th horizontal split inference boundary in luma samples. pps_split_inference_ctb_height[i] is specified in coding block units of MinCbSizeY height and may be in the range 0..CtbSizeY / MinCbSizeY-1. This coding syntax element is coded, for example, using a fixed-length code equal to log2(CtbSizeY / MinCbSizeY-1). Exp-Golomb coding may also be used.

[0161] The position (in luma samples) of the i-th horizontal split inference boundary PpsSplitInferenceBoundaryPosY[i] is derived as follows:

[0162] PpsSplitInferenceBoundaryPosY[i]=PictureWideSubPictureBoundaryPosY[i]- CtbSizeY + pps_split_inference_ctb_height[i]*MinCbSizeY In another embodiment, the inferred split boundary is inferred from the subpicture division (e.g., from the subpicture grid). In particular, the encoder can describe subpicture boundaries that are not aligned with the CTB boundary. For example, the subpicture division defines a subpicture grid element size that is smaller than the CTB size. If two consecutive grid elements have two different subpicture indices and belong to the same CTB, this indicates that the first subpicture has a right (or bottom) boundary that is not aligned with the CTB boundary. The second subpicture has a left (or top) boundary that is not aligned with the CTB boundary. In such a case, the inferred split boundary is inferred to align with the right (or bottom) boundary of the first subpicture, and therefore with the left (or top) boundary of the second subpicture. In a variant, different signaling is used to indicate the right and left boundaries of the subpictures, for example, by explicitly indicating the width and height of each subpicture in units smaller than the CTB size unit. In such cases, the width and height of the subpicture are compared with their values ​​in CTB units to determine subpictures that are not aligned with CTB boundaries. In another variation, a specific value of a subpicture element is reserved to indicate that this subpicture grid element has no coded data. Inferring split inference boundaries from subpicture splits avoids explicit signaling of split inference boundaries in the SPS, PPS, or any parameter set or SEI.

[0163] The VVC specification defines conformance windows within a PPS, as shown in the table below. [Table 10] This conformance window is a rectangular region within each picture represented by left, right, top, and bottom offsets relative to the picture boundary. At the end of the decoding process, the decoder applies a cropping process to remove pixels outside the conformance window.

[0164] In the previous embodiment, we described that the encoder skips some coding blocks in the CTB that are crossed by the split inference boundary. The skipped coding blocks are located to the right of the vertical split inference boundary and below the horizontal split inference boundary. In one embodiment, the decoder decodes the CTB using undefined pixels for these coding blocks. For this reason, we propose adding syntax elements to the PPS that define the adaptation window for each subpicture.

[0165] In one embodiment, the parameter set includes a new syntax element to indicate a rectangular region within the luma samples in each conforming sub-picture: it is all pixels decoded by the decoder whose values ​​are equal to the values ​​coded by the encoder, i.e., the values ​​corresponding to the useful part of the sub-picture.

[0166] Typically, the SPS or PPS defines four adaptive window offset parameters for each subpicture of the stream that has the subpic_treated_as_pic_flag flag equal to true. For example, the syntax is: [Table 11] The syntax elements subpic_conf_win_left_offset[i], subpic_conf_win_right_offset[i], subpic_conf_win_top_offset[i], subpic_conf_win_bottom_offset[i], subpic_conf_win_left_offset[i] specify the four adaptation window offset parameters.

[0167] In another embodiment, for example, if the encoder constrains the split inference boundary to span the picture width and height, the undefined pixels form one or more bands of pixels within the picture. As a result, instead of specifying a conformance region within each subpicture, the PPS describes bands of pixels that should be excluded from the conformance window and cropped.

[0168] For example, a PPS defines the number of bands of pixels to exclude from the conformance window defined for a picture in the PPS. For each band of pixels, the encoder specifies the width of the band. For example, the syntax of a PPS may include elements: [Table 12] With the following semantics: conformance_exclusion_band_flag equal to 1 indicates that there is a conformance exclusion band in the PPS. conformance_exclusion_band_flag equal to 0 indicates that there is no conformance exclusion band in the PPS.

[0169] num_ver_conformance_exclusion_bands specifies the number of conformance_exclusion_ver_band_pos_x[i] and conformance_exclusion_ver_band_width_minus1[i] syntax elements present in the PPS. If num_ver_conformance_exclusion_bands is not present, it is inferred to be equal to 0.

[0170] conformance_exclusion_ver_band_width_minus1[i] plus 1 specifies the width of the i-th vertical conformance exclusion band in luma samples and is used to calculate PpsConformanceVerBandPosX[i], which specifies the position of the i-th vertical conformance exclusion band boundary in luma samples. conformance. conformance_exclusion_ver_band_width_minus1[i] is in the range 0..CtbSizeY-2. This coding syntax element is coded, for example, using 8-bit or fixed-length codes equal to log2(pic_width_in_luma_samples / CtbSizeY) or equal to log2(CtbSizeY-2). Exp-Golomb codes can also be used.

[0171] conformance_exclusion_ver_band_pos_x[i] is used to calculate the value of PpsConformanceVerBandPosX[i], which specifies the position of the i-th vertical conformance exclusion band boundary in luma samples. conformance_exclusion_ver_band_pos_x[i] is in the range 0..PicWidthInCtbsY. In a variant, the range is 1..PicWidthInCtbsY-1 to avoid specifying vertical bands starting at the first or last pixel column of the picture. This coding syntax element is coded, for example, using a fixed-length code equal to 8 bits or log2(pic_width_in_luma_samples / CtbSizeY). Exp-Golomb coding is also possible.

[0172] The position of the i-th vertical exclusion band PpsConformanceVerBandPosX[i] is derived as follows:

[0173] PpsConformanceVerBandPosX[i]=conformance_exclusion_ver_band_pos_x[i]* CtbSizeY - (conformance_exclusion_ver_band_width_minus1[i]+ 1) The distance between any two vertical fit exclusion band boundaries may be greater than or equal to CtbSizeY luma samples, which is the size of the CTB in luma samples. In a variant, there is no limit on the distance between two vertical fit exclusion band boundaries, and each band may overlap another band. In such cases, the decoder must determine the overlapping bands to determine the actual fit area of ​​the picture.

[0174] num_hor_conformance_exclusion_bands specifies the number of conformance_exclusion_hor_band_pos_y[i] and conformance_exclusion_hor_band_height_minus1[i] syntax elements present in the PPS. If num_hor_conformance_exclusion_bands is not present, it is inferred to be equal to 0.

[0175] conformance_exclusion_hor_band_height_minus1[i] plus 1 specifies the height of the ith horizontal conformance exclusion band in luma samples and is used to calculate PpsConformanceHorBandPosY[i], which specifies the position of the ith horizontal conformance exclusion band boundary in luma samples. Conformance. conformance_exclusion_hor_band_height_minus1[i] is in the range 0..CtbSizeY-2. This coding syntax element is coded, for example, using a fixed-length code equal to 7 bits or log2(pic_height_in_luma_samples / CtbSizeY). It is also possible to use Exp-Golomb coding.

[0176] conformance_exclusion_hor_band_pos_y[i] is used to calculate the value of PpsConformanceHorBandPosY[i], which specifies the position of the ith horizontal conformance exclusion band boundary in luma samples. conformance_exclusion_hor_band_pos_y[i] is in the range 0..PicHeightInCtbsY, where PicHeightInCtbsY is the picture height in CTBs. In a variant, the range is 1..PicHeightInCtbsY-1 to avoid specifying horizontal bands starting on the first or last pixel row of the picture. This coding syntax element is coded, for example, using a fixed-length code equal to 8 bits or log2(pic_height_in_luma_samples / CtbSizeY). It is also possible to use Exp-Golomb coding.

[0177] The position of the i-th horizontal exclusion band PpsConformanceHorBandPosY[i] is derived as follows:

[0178] PpsConformanceHorBandPosY [i]=conformance_exclusion_hor_band_pos_y[i]* CtbSizeY -(conformance_exclusion_hor_band_height_minus1[i]+1) The distance between any two horizontal fit exclusion band boundaries may be greater than or equal to CtbSizeY luma samples, which is the size of the CTB in luma samples. In a variant, there is no limit on the distance between two horizontal fit exclusion band boundaries, and each band may overlap another band. In such cases, the decoder must determine the overlapping bands to determine the actual fit area of ​​the picture.

[0179] In a variant, the number of horizontal and vertical exclusion bands is arbitrary. In this case, if the num_ver_conformance_exclusion_bands and num_hor_conformance_exclusion_bands syntax elements are not present, their value is inferred to be equal to 1. The syntax of a PPS could be, for example: [Table 13] In a variant, the description of horizontal and / or vertical split inference boundary information is optional. For example, the syntax of a PPS (or SPS) is as follows: [Table 14] The semantics of the new syntax (in bold) elements are as follows:

[0180] conformance_exclusion_ver_band_flag equal to 1 indicates the presence of conformance_exclusion_ver_band_width_minus1 and conformance_exclusion_ver_band_pos_x in the PPS that specify the vertical conformance exclusion band. conformance_exclusion_band_flag equal to 0 indicates the absence of conformance_exclusion_ver_band_width_minus1 and conformance_exclusion_ver_band_pos_x in the PPS.

[0181] conformance_exclusion_ver_band_width_minus1 plus 1 specifies the width of the vertical conformance exclusion band in luma samples and is used to calculate PpsConformanceVerBandPosX. PpsConformanceVerBandPosX specifies the position of the left boundary of the vertical conformance exclusion band. conformance_exclusion_ver_band_width_minus1 may be in the range 0..CtbSizeY-2.

[0182] conformance_exclusion_ver_band_pos_x is used to calculate the value of PpsConformanceVerBandPosX. conformance_exclusion_ver_band_pos_x may be in the range 1..PicWidthInCtbsY-1.

[0183] The position of the left boundary of the vertical exclusion band PpsConformanceVerBandPosX is derived as follows:

[0184] PpsConformanceVerBandPosX=conformance_exclusion_ver_band_pos_x* CtbSizeY-(conformance_exclusion_ver_band_width_minus1 + 1)

[0185] conformance_exclusion_hor_band_flag equal to 1 indicates the presence of conformance_exclusion_hor_band_height_minus1 and conformance_exclusion_hor_band_pos_y in the PPS that specify the horizontal conformance exclusion band. conformance_exclusion_band_flag equal to 0 indicates the absence of conformance_exclusion_hor_band_height_minus1 and conformance_exclusion_hor_band_pos_y in the PPS.

[0186] conformance_exclusion_hor_band_height_minus1 plus 1 specifies the height of the horizontal conformance exclusion band in luma samples and is used to calculate PpsConformanceHorBandPosY. PpsConformanceHorBandPosY specifies the position of the upper boundary of the horizontal conformance exclusion band. conformance_exclusion_hor_band_height_minus1 may be in the range 0..CtbSizeY-2.

[0187] conformance_exclusion_hor_band_pos_y is used to calculate the value of PpsConformanceHorBandPosY. conformance_exclusion_hor_band_pos_y may be in the range 1..PicHeightInCtbsY-1.

[0188] The position of the upper boundary of the horizontal exclusion band PpsConformanceHorBandPosY is derived as follows.

[0189] PpsConformanceHorBandPosY=conformance_exclusion_hor_band_pos_y * CtbSizeY-(conformance_exclusion_hor_band_height_minus1 + 1)

[0190] In a variant, the width, height, and coordinates of the exclusion band are expressed in chroma samples. These values ​​in luma samples are obtained by multiplying the values ​​expressed in chroma samples by SubWidthC and SubHeightC. SubWidthC and SubHeightC represent the horizontal and vertical sampling ratios between the luma and chroma elements, respectively. For example, if the chroma format is 4:2:0, SubWidthC and SubHeightC are equal to 2. In another variant, the width, height, and coordinates of the exclusion band are expressed in minimum CTB size units to further reduce the length of the syntax element.

[0191] In another embodiment, the width and height of the exclusion band are greater than the CTB. Two consecutive exclusion bands of pixels can be merged into one exclusion band, for example, when two adjacent subpictures define two adjacent sets of non-matching pixels. This allows for a more compact description of the exclusion band.

[0192] In another embodiment, the number of vertical and horizontal fit exclusion bands is inferred to be equal to the number of vertical and horizontal split inference boundaries, respectively.

[0193] In another embodiment, the positions of the vertical and horizontal fitted exclusion bands are inferred to be equal to the positions of the vertical and horizontal split inference boundaries, respectively.

[0194] In another embodiment, the width of a given vertical fitted exclusion band is inferred to be equal to the number of pixels between the vertical split inference boundary and the nearest right CTB boundary at the same location. The same applies to horizontal exclusion bands with horizontal split inference boundaries.

[0195] In another embodiment, subpicture split inference is inferred from the subpicture division: pixels to be excluded from the matching window are inferred to correspond to the region between the right (or bottom) borders of subpictures that are not aligned with CTB boundaries and the right (or bottom) borders of these CTBs.

[0196] In a variant, these regions are grouped into different subpictures. In such a case, the width or height of these subpictures is less than the CTB size. Therefore, it is possible to associate these non-conforming subpictures with specific indices. The SPS or PPS can provide a list of subpicture indices that are not conforming and should be excluded from the conformance window. In another embodiment, the number of bands to be excluded from the conformance window and their size are determined from the inferred split boundary positions.

[0197] In another alternative, the number of bands excluded from the fitting window and their size are determined from the inferred split boundary positions.

[0198] The adaptive cropping window includes luma samples with horizontal picture coordinates from SubWidthC * conf_win_left_offset to pic_width_in_luma_samples-(SubWidthC * conf_win_right_offset + 1) and vertical picture coordinates from SubHeightC * conf_win_top_offset to pic_height_in_luma_samples-(SubHeightC * conf_win_bottom_offset+1).

[0199] Additionally, luma samples with horizontal picture coordinates from PpsConformanceVerBandPosX[i] to PpsConformanceVerBandPosX[i]+conformance_exclusion_ver_band_width_minus1[i]+1 are excluded from the conformance window for i in the range 0..num_ver_conformance_exclusion_bands.

[0200] Additionally, luma samples with vertical picture coordinates from PpsConformanceHorBandPosY[i] to PpsConformanceHorBandPosY[i]+conformance_exclusion_hor_band_height_minus1[i]+1 are excluded from the conformance window for i in the range 0..num_hor_conformance_exclusion_bands.

[0201] The width and height of the picture after cropping corresponds to the variables PicOutputWidthL and PicOutputHeightL, which are derived as follows: The width (or height) of this picture is equal to the width (or height) of the cropping window minus the width (or height) of the pixel exclusion band, which corresponds to the following pseudocode: PicOutputWidthL = pic_width_in_luma_samples -SubWidthC*(conf_win_right_offset + conf_win_left_offset) for(i=0;i <num_ver_conformance_exclusion_bands; i++) PicOutputWidthL -= conformance_exclusion_ver_band_width_minus1[i]+1 PicOutputHeightL = pic_height_in_pic_size_units - SubHeightC*(conf_win_bottom_offset + conf_win_top_offset) for(i=0;i <num_hor_conformance_exclusion_bands; i++) PicOutputHeightL -= conformance_exclusion_hor_band_height_minus1[i]+1 In one embodiment, a picture and its subpictures that can be displaced in a merge operation are constrained to have a size that is a multiple of the CTB size in both directions. Typically, this can be done by extending the input picture with padding pixels if it is not already taking this constraint into account. Therefore, there is no inferred boundary mechanism to be used. This constraint is signaled in the bitstream. For example, a flag can be used to indicate that a subpicture is subject to this constraint and can therefore be freely merged and displaced to any position within the resulting image without boundary issues. To ensure that subpictures can be freely merged, an additional constraint is associated with the flag, which is that the subpictures are independently decodable. The flag may be defined to apply to all subpictures in a picture, or it may be defined at the subpicture level to indicate that the associated subpictures can be freely merged. In one embodiment, a flag indicating that a picture and its subpictures have a size that is a multiple of the CTB size is defined in the sequence parameter set and is valid for all subpictures in the sequence.

[0202] In some embodiments, the adaptation window may be defined at the sub-picture level, for example in an SEI message.

[0203] In one embodiment, the syntax of the SPS is modified to include one additional flag that indicates whether the encoding of sub-pictures in the bitstream constrains the coding tool and video sequence characteristics to allow merging operations. When set, this flag restricts whether the sub-pictures are independently decodable. The syntax becomes: [Table 15] For example, the semantics of the element are:

[0204] subpic_mergeable_flag equal to 1 indicates that all subpictures of the picture are constrained for merge operations. subpic_mergeable_flag equal to 0 indicates whether the subpicture is constrained.

[0205] pic_width_max_in_luma_samples specifies the maximum width, in luma samples, of each decoded picture that references the SPS. pic_width_max_in_luma_samples is not equal to 0 and is an integer multiple of Max(8, MinCbSizeY). If subpic_mergeable_flag is equal to 1, pic_width_max_in_luma_samples is an integer multiple of CtbSizeY. As a result, if a subpicture is constrained for a merge operation (subpic_mergeable_flag is equal to 1), the coded picture is constrained to use padding when the picture's original width is not a multiple of the CTB size.

[0206] pic_height_max_in_luma_samples specifies the maximum height, in luma samples, of each decoded picture that references the SPS. pic_height_max_in_luma_samples shall not be equal to 0 and shall be an integer multiple of Max(8, MinCbSizeY). subpic_mergeable_flag If mergeable_flag is equal to 1, pic_height_max_in_luma_samples shall be an integer multiple of CtbSizeY. As a result, if a subpicture is constrained for a merge operation (subpic_mergeable_flag is equal to 1), the coded picture is constrained to use padding if the picture's original height is not a multiple of the CTB size.

[0207] In this embodiment, the semantics of some elements of the PPS may be constrained to ensure that each picture that references the PPS has a picture size that is a multiple of the CTB size when subpictures are constrained for merge operations. For example, the semantics of pic_width_in_luma_samples and pic_height_in_luma_samples are as follows:

[0208] pic_width_in_luma_samples specifies the width, in luma samples, of each decoded picture that references the PPS. pic_width_in_luma_samples is not equal to 0, is an integer multiple of Max(8, MinCbSizeY), and is less than or equal to pic_width_max_in_luma_samples. If subpic_mergeable_flag is equal to 1 (e.g., in an SPS with an identifier signaled in the PPS), pic_width_in_luma_samples MUST be an integer multiple of CtbSizeY.

[0209] pic_height_in_luma_samples specifies the height, in luma samples, of each decoded picture referencing the PPS. pic_height_in_luma_samples is not equal to 0 and is an integer multiple of Max(8, MinCbSizeY) and is less than or equal to pic_height_max_in_luma_samples. If subpic_mergeable_flag is equal to 1, pic_height_in_luma_samples is an integer multiple of CtbSizeY.

[0210] This first set of constraints on a picture ensures that any merge position is possible for a subpicture of a picture in the merged stream if subpic_mergeable_flag is equal to 1. The encoder may have to use padding data to make both the width and height of the picture a multiple of the CTB size. The encoder can signal the area with the padding data by defining adaptation windows in the bitstream. Typically, one adaptation window can be defined for each subpicture in the SEI message.

[0211] In another embodiment, the mergeable flag also constrains the temporal prediction mechanism within each subpicture for all subpictures described in the parameter set NAL unit if the subpicture is constrained for merging operations (subpic_mergeable_flag is equal to 1).

[0212] For example, subpic_treated_as_pic_flag[i] equal to 1 specifies that the i-th subpicture of each coded picture in the CLVS is treated as a picture in the decoding process excluding in-loop filtering operations. subpic_treated_as_pic_flag[i] equal to 0 specifies that the i-th subpicture of each coded picture in the CLVS is not treated as a picture in the decoding process excluding in-loop filtering operations. If not present, the value of subpic_treated_as_pic_flag[i] is inferred to be equal to subpic_mergeable_flag. Consequently, if a subpicture is constrained for merging operations (subpic_mergeable_flag equal to 1), the subpic_treated_as_pic_flag[i] syntax element is not present in the parameter set NAL unit and is inferred to be equal to 1, indicating that the temporal prediction of the subpicture is constrained to each subpicture boundary. Otherwise, the temporal prediction may or may not be constrained.

[0213] In another embodiment, if a subpicture is constrained for merging operations (subpic_mergeable_flag equals 1), the loop filter mechanism is disabled across subpicture boundaries for all subpictures described in the parameter set NAL unit. For example, loop_filter_across_subpic_enabled_flag[i] equal to 1 specifies that in-loop filtering operations are performed across the boundaries of the i-th subpicture in each coded picture in the CLVS. loop_filter_across_subpic_enabled_flag[i] equal to 0 specifies that in-loop filtering operations are not performed across the boundaries of the i-th subpicture in each coded picture in the CLVS. If not present, the value of loop_filter_across_subpic_enabled_pic_flag[i] is inferred to be equal to !subpic_mergeable_flag. As a result, if a subpicture is constrained for merging operations (subpic_mergeable_flag is equal to 1), the loop_filter_across_subpic_enabled_pic_flag[i] syntax element is inferred to be absent in the parameter set NAL unit and equal to 0, indicating that the loop filter is constrained to not be enabled across subpicture boundaries. Otherwise, the loop filter may or may not be constrained.

[0214] In some embodiments, spatial access may be provided only at certain time intervals. The proposed flags may be defined correspondingly in the PPS or even in the picture header, which is a non-VCL NAL unit that is defined at the picture level and defines some syntax elements that apply to all slices of the picture.

[0215] In another embodiment, subpicture signaling associates a subpicture identifier with a subpicture index. This identifier is unique for a given subpicture, simplifying the merge operation because it allows avoiding rewriting the subpicture index in the slice header when the subpicture is moved to a new position after the merge operation. Thus, if a subpicture is constrained for a merge operation, one of the non-VCL NAL units can signal the subpicture identifier in the bitstream.

[0216] Typically, a subpicture identifier can be present in the SPS, PPS, or picture header. In particular, sps_subpic_id_present_flag is a flag in the SPS that, when equal to 1, indicates the presence of a subpicture identifier in the bitstream. The semantics of this syntax element are, for example, as follows:

[0217] sps_subpic_id_present_flag equal to 1 specifies that a subpicture ID mapping is present in the SPS. sps_subpic_id_present_flag equal to 0 specifies that a subpicture ID mapping is not present in the SPS. If subpic_mergeable_flag is equal to 1, sps_subpic_id_present_flag MUST be equal to 1. Consequently, if a subpicture is constrained for merging operations, signaling of the subpicture identifier is provided in the bitstream. In a variant, if subpic_mergeable_flag is equal to 1, sps_subpic_id_present_flag is not present in the SPS and is inferred to be 1. The corresponding syntax of the SPS may include the following syntax elements: [Table 16]

[0218] In another embodiment, the presence of sub-picture identifiers in picture headers may make merging operations more complicated. In fact, if the identifier signaling is done in the picture header, it overrides the sub-picture identifier mapping done in the parameter set NAL unit. Because a different picture header is transmitted for each frame, the mapping may change for each picture. As a result, a typical merging operation needs to check whether the picture header changes the sub-picture identifier mapping. To avoid this checking operation, if a sub-picture is constrained for merging operations, the identifier mapping in the picture header is disabled. The presence of sub-picture identifier mapping in the picture header is controlled by the ph_subpic_id_signalling_present_flag syntax element. In this embodiment, the semantics of ph_subpic_id_signalling_present_flag are as follows:

[0219] ph_subpic_id_signalling_present_flag equal to 1 specifies that subpicture ID mapping is signaled to the PH. ph_subpic_id_signalling_present_flag equal to 0 specifies that subpicture ID mapping is not signaled to the PH. If subpic_mergeable_flag is equal to 1, ph_subpic_id_signalling_present_flag MUST be equal to 0. In a variant, if subpic_mergeable_flag is equal to 1, ph_subpic_id_signalling_present_flag is not present in the picture header and is inferred to be equal to 0.

[0220] A PPS can also signal the mapping of subpicture identifiers. For picture headers, it is necessary to check whether the mapping has changed between two PPS NAL units. This additional check increases the complexity of the merging operation. As a result, in one embodiment, if a subpicture is constrained for a merging operation, the subpicture is constrained to be identical in each PPS and in each PPS for which subpic_mergeable_flag is equal to 1. The pps_subpic_id[i] syntax element of a PPS specifies the subpicture ID (or identifier) ​​of the ith subpicture corresponding to the mapping of subpicture identifiers with the subpicture index. If subpic_mergeable_flag is equal to 1, all PPSs that refer to the same SPS have the same value of pps_subpic_id[i] for i ranging from 0 to pps_num_subpics_minus1.

[0221] In another embodiment, the information indicating that a sub-picture is constrained for a merging operation corresponds to a specific profile of the VVC specification. Typically, a specific value (e.g., 3) of the general_profile_idc syntax element indicates the profile to which the output layer conforms for a sub-picture merging operation. In a variant, the information indicating that a sub-picture is constrained for a merging operation corresponds to a sub-profile. In that case, a specific value (e.g., 3) of general_sub_profile_idc[i] indicates that the bitstream is constrained for a sub-picture merging operation.

[0222] The VVC general constrained information structure is configured by flags described in the profile, layer and level information that allow one or more coding tools to be disabled. In one embodiment, the general constrained information structure includes any of the following syntax elements:

[0223] - a no_dependent_subpicture_flag syntax element which, when equal to 1, specifies that subpic_treated_as_pic_flag[i] is equal to 1 for any value of i. A no_dependent_subpicture_flag equal to 0 imposes no such constraint. In a variant, when equal to 1, it specifies that subpic_treated_as_pic_flag[i] and loop_filter_across_subpic_enabled_flag[i] are equal to 1 for any value of i.

[0224] - a no_picture_header_subpicture_id_mapping syntax element which, when equal to 1, specifies that all picture header NAL units have ph_subpic_id_signalling_present_flag equal to 0. no_picture_header_subpicture_id_mapping equal to 0 imposes no such constraint.

[0225] - a no_pps_subpicture_id_mapping_change syntax element that, when equal to 1, specifies that all PPS NAL units have equal values ​​of pps_subpic_id[i] for any value of i. no_pps_subpicture_id_mapping_change equal to 0 imposes no such constraint.

[0226] In one embodiment, an SEI message is proposed to handle regions within a picture that require a post-decoding crop operation. This may be the result of a bitstream extraction and merging operation, for example, where some sub-pictures at the right or bottom border of the picture originally contained padding data. This SEI indicates to the decoder a conformance window for at least some sub-pictures that defines picture regions containing padding, or unreliable or even meaningless data, that the content creator believes should be removed on the image to be displayed.

[0227] For example, the SEI might provide the following syntax: [Table 17]

[0228] With the following semantics: subpic_conf_win_cancel_flag equal to 1 indicates that the SEI message cancels the persistence of any previous subpicture adaptation window SEI message in output order that applies to the current layer. subpic_conf_win_cancel_flag equal to 0 indicates that subpicture adaptation window information follows.

[0229] subpic_conf_win_num_subpics_minus1 plus 1 specifies the number of subpicture adaptation windows present in the SEI message. This value is a function of the number of subpictures present in the picture. Typically, it is required that the value of subpic_conf_win_num_subpics_minus1 is equal to sps_num_subpics_minus1, allowing a adaptation window to be defined for each subpicture.

[0230] subpic_conf_win_left_offset[i], subpic_conf_win_right_offset[i], subpic_conf_win_top_offset[i], and subpic_conf_win_bottom_offset[i] specify the samples of the ith subpicture of the picture in the CLVS that are output from the decoding process, with respect to the rectangular area specified in picture coordinates for output relative to the origin of the ith subpicture, as described in the SPS NAL unit.

[0231] The subpicture adaptive cropping window for the i-th subpicture includes luma samples with horizontal picture coordinates from SubPictureLuma_X[i] + SubWidthC * conf_win_left_offset to SubPictureLuma_X[i] + SubPictureLuma_Width[i] - (SubWidthC * subpic_conf_win_right_offset[i] + 1) and vertical picture coordinates from SubPictureLuma_Y[i] + SubHeightC * subpic_conf_win_top_offset[i] to SubPictureLuma_Y[i] + SubPictureLuma_Height[i] - (SubHeightC * subpic_conf_win_bottom_offset[i] + 1). SubPictureLuma_X[i] and SubPictureLuma_Y[i] specify the horizontal and vertical picture coordinates of the first pixel of the ith subpicture, and SubPictureLuma_Width[i] and SubPictureLuma_Height[i] are the width and height of the luma sample of the ith subpicture as described in the SPS. For example, these variables are calculated as follows: SubPictureLuma_X[i]=subpic_ctu_top_left_x[i]* CtbSizeY SubPictureLuma_Y[i]=subpic_ctu_top_left_y[i]* CtbSizeY SubPictureLuma_Width[i]=(subpic_width_minus1[i]+1)* CtbSizeY SubPictureLuma_Height[i]=(subpic_height_minus1[i]+1)* CtbSizeY In some cases, only a subset of subpictures (typically subpictures containing padding data) require a conformance window. In such cases, if a conformance window is signaled, a new syntax element is specified for each subpicture. For example, a "for" loop for each subpicture specifies the subpic_conf_win_signalled_flag[i] syntax element for the i-th subpicture described in the subpicture. If equal to 1, the offset parameter is present and a conformance window is specified for the subpicture. Otherwise, subpic_conf_win_signalled_flag[i] is equal to 0, conformance is not signaled for the i-th subpicture, and the offset parameter is inferred to be absent and equal to 0. In a variant, the number of subpicture conformance windows described in the SEI message differs from the number of subpictures in the picture, and for each subpicture conformance window signaled in the SEI, a list of one or more subpicture indices within the picture is associated with the subpicture conformance window index. This list of indexes indicates the subpictures that use the subpicture adaptation window. For example, a for loop in an SEI message indicates a subpic_conf_win_num_subpics_minus1[i] syntax element, which is the number of subpicture indices associated with the i-th subpicture adaptation window of the SEI message minus 1. Then, a processing loop for j ranging from 0 to subpic_conf_win_num_subpics_minus1[i] defines a subpic_conf_win_subpic_index[i][j] syntax element, where subpic_conf_win_subpic_index[i][j] specifies the j-th index of the subpicture that uses the i-th subpicture adaptation window. In a variant, an SEI message can define a subpicture identifier instead of a subpicture index. In that case, the bit length of the subpicture identifier is optionally described in the SEI message.

[0232] 12 is a schematic block diagram of a computing device 120 for implementing one or more embodiments of the present invention. The computing device 120 may be a device such as a microcomputer, a workstation, or a light portable device. The computing device 120 may include: a central processing unit 121 such as a microprocessor, denoted CPU; a random access memory 122, denoted RAM, for storing the executable code of the method according to an embodiment of the present invention, as well as registers adapted to record variables and parameters necessary for carrying out the method according to an embodiment of the present invention, the memory capacity of which can be expanded, for example, by an optional RAM connected to an expansion port; a read-only memory 123, denoted ROM, for storing a computer program for implementing an embodiment of the present invention; The network interface 124 is typically connected to a communications network over which the digital data to be processed is transmitted and received. The network interface 124 may be a single network interface or may consist of a set of different network interfaces (e.g., wired and wireless interfaces, or different types of wired or wireless interfaces). Data packets are read from the network interface for reception or written to the network interface for transmission under the control of software applications running within the CPU 121; The user interface 125 may be used to receive input from a user or to display information to a user; A hard disk 126, denoted HD, may be provided as mass storage; The I / O module 127 may be used to receive / send data from / to external devices such as video sources or displays. A communication bus is connected to the

[0233] The executable code may be stored either in the read-only memory 123, on the hard disk 126 or on a removable digital medium such as a disk. According to a variant, the executable code of the program may be received by the communications network via the network interface 124 in order to be stored in one of the storage means of the communications device 120, such as the hard disk 126, before being executed.

[0234] The central processing unit 121 is adapted to control and direct the execution of instructions or portions of software code of a program or programs according to embodiments of the present invention stored in one of the aforementioned storage means. After power-on, the CPU 121 is able to execute instructions from the main RAM memory 122 relating to a software application after those instructions have been loaded, for example, from a program ROM 123 or a hard disk (HD) 126. Such a software application, when executed by the CPU 121, causes the steps of the flowcharts of the present invention to be performed.

[0235] Any step of an algorithm of the present invention may be implemented in software by execution of a program or set of instructions by a programmable computing machine such as a PC ("personal computer"), a DSP ("digital signal processor") or a microcontroller, or may be implemented in hardware by a machine or dedicated component such as an FPGA ("field programmable gate array") or an ASIC ("application specific integrated circuit").

[0236] Although the present invention has been described herein with reference to specific embodiments, the present invention is not limited to those embodiments, and modifications within the scope of the present invention will be apparent to those skilled in the art.

[0237] Many further modifications and variations will suggest themselves to those skilled in the art by reference to the foregoing exemplary embodiments, which are given by way of example only and are not intended to limit the scope of the invention, which is determined solely by the appended claims. In particular, different features from different embodiments may be interchanged where appropriate.

[0238] Each of the above-described embodiments of the present invention may be implemented alone or in combination in multiple embodiments, and features from various embodiments may be combined where necessary or where a combination of elements or features from the individual embodiments in a single embodiment is beneficial.

[0239] Each feature disclosed in this specification (including any accompanying claims, abstract, and drawings), unless stated otherwise, may be replaced by alternative features serving the same, equivalent, or similar purpose. Thus, unless stated otherwise, each feature disclosed is only an example of a generic series of equivalent or similar features.

[0240] In the claims, the word "comprise" does not exclude other elements or steps, and the indefinite articles "a" or "an" do not exclude a plurality. The mere fact that different features are recited in mutually different dependent claims does not indicate that a combination of these features cannot be used to advantage.

Claims

1. 1. A method of encoding video data including pictures into a bitstream, wherein the pictures are divided into sub-pictures including one or more slices, the method comprising: encoding a first flag in a sequence parameter set of the bitstream to indicate whether sub-picture information is present; encoding a second flag relating to a sub-picture constraint into the sequence parameter set of the bitstream when the value of the first flag is 1; encoding, when the value of the second flag is 0, a third flag indicating whether in-loop filtering is enabled at a sub-picture boundary for a sub-picture within a picture, and a fourth flag indicating whether the sub-picture is to be treated as a picture in a decoding process excluding in-loop filtering, into the sequence parameter set of the bitstream; encoding coding tree blocks constituting a picture into the bitstream; and When the value of the second flag is 1, the third flag and the fourth flag are not coded in the sequence parameter set of the bitstream; Information for specifying the width of a sub-picture can be coded into the sequence parameter set. A method characterized by:

2. 2. The method of claim 1, wherein when the value of the second flag is 1, the third flag is inferred as 0.

3. 3. The method of claim 1, wherein the second flags are associated with a plurality of sub-pictures included in the picture.

4. 4. The method of claim 3, wherein the second flag is associated with all sub-pictures contained in the picture.

5. 5. The method of claim 1, wherein when the second flag has a value of 1, the sub-picture boundary is treated as an image boundary and no in-loop filtering is performed at the sub-picture boundary.

6. 6. The method of claim 1, wherein the first flag, the second flag, and the third flag are defined in the sequence parameter set.

7. 7. A method according to any one of claims 1 to 6, wherein the sub-pictures are identified by sub-picture identifiers.

8. 8. A method according to any one of claims 1 to 7, wherein the sub-picture is a rectangular region made up of one or more slices.

9. 2. The method of claim 1, wherein when the value of the second flag is one, the fourth flag is inferred as one.

10. 10. A method according to any one of claims 1 to 9, comprising the step of encoding information indicative of a conformance window for a sub-picture.

11. 1. A method of decoding a bitstream of video data including pictures, the pictures being divided into sub-pictures including one or more slices, the method comprising: decoding a first flag indicating whether sub-picture information is present from a sequence parameter set of the bitstream; decoding a second flag relating to a sub-picture constraint from the sequence parameter set of the bitstream when the value of the first flag is 1; decoding, when the value of the second flag is 0, a third flag indicating whether in-loop filtering is enabled at a sub-picture boundary for a sub-picture within a picture, and a fourth flag indicating whether the sub-picture is to be treated as a picture in a decoding process excluding in-loop filtering, from the sequence parameter set of the bitstream; and decoding the bitstream based on at least the first flag; when the value of the second flag is 1, the third flag and the fourth flag are not decoded from the sequence parameter set of the bitstream; Information for specifying the width of a sub-picture can be decoded from the sequence parameter set. A method characterized by:

12. 12. The method of claim 11, wherein when the value of the second flag is 1, the third flag is inferred as 0.

13. 13. The method of claim 11 or 12, wherein the second flags are associated with a plurality of sub-pictures included in the picture.

14. 14. The method of claim 13, wherein the second flag is associated with all sub-pictures contained in the picture.

15. 15. The method of any one of claims 11 to 14, wherein when the value of the second flag is 1, the sub-picture boundary is treated as an image boundary and no in-loop filtering is performed at the sub-picture boundary.

16. 16. The method of any one of claims 11 to 15, wherein the first flag, the second flag, and the third flag are defined in the sequence parameter set.

17. 17. A method according to any one of claims 11 to 16, wherein the sub-pictures are identified by sub-picture identifiers.

18. 18. A method according to any one of claims 11 to 17, wherein the sub-picture is a rectangular region made up of one or more slices.

19. 12. The method of claim 11, wherein when the value of the second flag is one, the fourth flag is inferred as one.

20. 20. A method according to any one of claims 11 to 19, comprising the step of decoding information indicative of a conformance window for a sub-picture.

21. 1. An apparatus for encoding video data including pictures into a bitstream, the picture being divided into sub-pictures including one or more slices, the apparatus comprising: means for encoding a first flag indicating whether sub-picture information is present in a sequence parameter set of the bitstream; means for encoding a second flag relating to a sub-picture constraint into the sequence parameter set of the bitstream when the value of the first flag is 1; means for encoding, when the value of the second flag is 0, a third flag indicating whether in-loop filtering is enabled at a sub-picture boundary for a sub-picture within a picture, and a fourth flag indicating whether the sub-picture is to be treated as a picture in a decoding process excluding in-loop filtering, into the sequence parameter set of the bitstream; means for encoding coding tree blocks constituting a picture into the bitstream; When the value of the second flag is 1, the third flag and the fourth flag are not coded in the sequence parameter set of the bitstream; Information for specifying the width of a sub-picture can be coded into the sequence parameter set. An apparatus characterized in that

22. 1. An apparatus for decoding a bitstream of video data including pictures, the pictures being divided into sub-pictures including one or more slices, the apparatus comprising: means for decoding a first flag indicating whether sub-picture information is present from a sequence parameter set of a bitstream; means for decoding a second flag relating to a sub-picture constraint from the sequence parameter set of the bitstream when the value of the first flag is 1; means for decoding, from the sequence parameter set of the bitstream, a third flag indicating whether in-loop filtering is enabled at a sub-picture boundary for a sub-picture within a picture when the value of the second flag is 0, and a fourth flag indicating whether the sub-picture is to be treated as a picture in a decoding process excluding in-loop filtering; means for decoding the bitstream based on at least the first flag; when the value of the second flag is 1, the third flag and the fourth flag are not decoded from the sequence parameter set of the bitstream; Information for specifying the width of a sub-picture can be decoded from the sequence parameter set. An apparatus characterized in that

23. A computer program product for causing a computer to carry out the method according to any one of claims 1 to 20.

Citation Information

Patent Citations

  • Method and device for decoding of compensation offsets for set of reconstructed samples of image

    JP2019110587A

  • Signaling in a video bitstream dependent on subpictures - Patent Application 20070122997

    JP2022543705A