Syntax for signaling video sub-pictures

By modifying the VVC standard to allow more subpictures and optimizing syntax element handling, the challenges of subpicture signaling are addressed, enhancing video coding efficiency and flexibility.

JP2026012196APending Publication Date: 2026-01-23BYTEDANCE INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025172473
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-12-27
Filing Date
2025-10-14
Publication Date
2026-01-23

AI Technical Summary

Technical Problem

Existing video coding standards, particularly Versatile Video Coding (VVC), face issues with signaling subpictures, tiles, and slices, including limitations on the number of subpictures per picture, redundancy in syntax elements, and inefficient handling of subpicture IDs, leading to undesirable modifications and signaling inconsistencies.

Method used

Modifications to the VVC standard include changing the coding of sps_num_subpics_minus1 to ue(v) to allow more than 256 subpictures, restricting syntax elements based on conditions, removing subpicture ID signaling in the picture header, and reordering syntax elements to reduce redundancy and improve flexibility in subpicture handling.

Benefits of technology

The proposed modifications enhance the VVC standard by allowing for more subpictures per picture, reducing unnecessary signaling, and eliminating redundant syntax elements, thereby improving video coding efficiency and flexibility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026012196000001_ABST
    Figure 2026012196000001_ABST
Patent Text Reader

Abstract

To provide a method for signaling usage of sub-pictures in coded video pictures.SOLUTION: One example method of video processing includes performing a conversion between a video and a bitstream of the video, wherein the bitstream conforms to a format rule, wherein the format rule specifies that a first syntax element signaled in a sequence parameter set (SPS) of the bitstream indicates a length of an identifier of a subpicture in the SPS, and wherein a signaling of the first syntax element is independent of a value of a second syntax element that indicates an explicit signaling of the identifier of the subpicture in the SPS or a picture parameter set (PPS).SELECTED DRAWING: Figure 10
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS Under applicable patent laws and / or regulations pursuant to the Paris Convention, this application is filed to timely claim priority to and the benefit of U.S. Provisional Patent Application No. 62 / 954,364, filed December 27, 2019. For all purposes under law, the entire disclosure of the aforementioned application is incorporated by reference as part of the disclosure of this application.

[0002] This specification relates to image and video coding and decoding. [Background technology]

[0003] Digital video accounts for the largest bandwidth usage on the Internet and other digital communication networks, and as the number of connected user devices capable of receiving and displaying video increases, the bandwidth demands for digital video usage are expected to continue to increase. Summary of the Invention

[0004] This specification discloses a method, apparatus and system for subpicture signaling that can be used by a video encoder and decoder for video encoding and decoding, respectively.

[0005] In one exemplary aspect, a video processing method is disclosed that includes performing a conversion between a video including a picture and a video bitstream, where a number of sub-pictures in the picture is signaled in a Sequence Parameter Set (SPS) of the bitstream as a field whose bitwidth is based on a value of the number of sub-pictures, the field being an unsigned integer zeroth-order Exponential-Golomb (Exp-Golomb) coded syntax element with left bit first.

[0006] In another example aspect, a video processing method is disclosed that includes performing a conversion between a video and a video bitstream, the bitstream conforming to a format rule that specifies that a first syntax element indicating whether a picture of the video can be partitioned is conditionally included in a Picture Parameter Set (PPS) of the bitstream based on values ​​of a second syntax element indicating whether a sub-picture identifier is signaled in the PPS and a third syntax element indicating the number of sub-pictures in the PPS.

[0007] In yet another example aspect, a video processing method is disclosed. The method includes performing a conversion between a video and a bitstream of the video, the bitstream conforming to a format rule, the format rule specifying that a first syntax element indicating whether a picture of the video can be partitioned is included in a Picture Parameter Set (PPS) of the bitstream before a set of syntax elements in the PPS that indicate identifiers of sub-pictures of the picture.

[0008] In yet another exemplary aspect, a video processing method is disclosed, the method including performing a conversion between a video domain of a video and a bitstream of the video, the bitstream conforming to a format rule, the format rule specifying that a first syntax element indicating whether sub-picture identifier information is included in a parameter set of the bitstream is conditionally included in a sequence parameter set (SPS) based on a value of a second syntax element indicating whether the sub-picture information is included in the SPS.

[0009] In yet another exemplary aspect, a video processing method is disclosed, the method including performing a conversion between a picture of a video and a bitstream of the video, the bitstream conforming to a format rule, the format rule specifying a mapping between identifiers of one or more sub-pictures of the picture, the one or more sub-pictures not being included in a picture header of the picture, the format rule further specifying that the identifiers of the one or more sub-pictures are derived based on syntax elements of a sequence parameter set (SPS) and a picture parameter set (PPS) referenced by the picture.

[0010] In yet another exemplary aspect, a video processing method is disclosed, the method including performing a conversion between a video and a video bitstream, the bitstream conforming to a format rule, the format rule specifying that if a value of a first syntax element indicates that a mapping between a subpicture identifier and one or more subpictures of a picture is explicitly signaled for the one or more subpictures, the mapping is signaled in either a sequence parameter set (SPS) or a picture parameter set (PPS).

[0011] In yet another exemplary aspect, a video processing method is disclosed, the method including performing a conversion between a video and a video bitstream, the bitstream conforming to a format rule, the format rule specifying that signaling a length of a sub-picture identifier in a sequence parameter set (SPS) is not based on a value of a syntax element indicating whether the identifier is signaled in the SPS.

[0012] In yet another exemplary aspect, a video processing method is disclosed, the method including performing a conversion between a video and a video bitstream, the bitstream conforming to a format rule, the format rule specifying signaling a length of a sub-picture identifier in a Picture Parameter Set (PPS) by a syntax element indicating that the identifier is explicitly signaled in the PPS.

[0013] In yet another example aspect, a video processing method is disclosed, the method including performing a conversion between video and a video bitstream, the bitstream conforming to a format rule, the format rule specifying that a first syntax element signaled in a Sequence Parameter Set (SPS) of the bitstream indicates a length of a sub-picture identifier in the SPS, the signaling of the first syntax element being independent of a value of a second syntax element indicating explicit signaling of the sub-picture identifier in the SPS or a Picture Parameter Set (PPS).

[0014] In yet another exemplary aspect, a video encoder apparatus is disclosed, the video encoder having a processor configured to implement the above-described method.

[0015] In yet another exemplary aspect, a video decoder apparatus is disclosed, the video decoder having a processor configured to implement the above-described method.

[0016] In yet another exemplary aspect, a computer-readable medium having stored thereon code, the code embodying one of the methods described herein in the form of processor-executable code, is disclosed.

[0017] These and other features are described throughout this specification. [Brief explanation of the drawings]

[0018] [Figure 1]An example of partitioning a picture by luma coding tree units (CTUs) is shown.

[0019] [Figure 2] Another example of partitioning a picture by luma CTU is shown.

[0020] [Figure 3] 1 illustrates an example partitioning of a picture.

[0021] [Figure 4] 1 illustrates another example partitioning of a picture.

[0022] [Figure 5] FIG. 1 is a block diagram of an exemplary video processing system in which the disclosed techniques may be implemented.

[0023] [Figure 6] FIG. 1 is a block diagram of an exemplary hardware platform used for video processing.

[0024] [Figure 7] 1 is a block diagram illustrating a video coding system according to some embodiments of the present disclosure.

[0025] [Figure 8] FIG. 2 is a block diagram illustrating an encoder according to some embodiments of the present disclosure.

[0026] [Figure 9] FIG. 2 is a block diagram illustrating a decoder according to some embodiments of the present disclosure.

[0027] [Figure 10] 1 shows a flowchart of an exemplary method for video processing.

[0028] [Figure 11]1 shows a flowchart of an exemplary method for video processing.

[0029] [Figure 12] 1 shows a flowchart of an exemplary method for video processing. DETAILED DESCRIPTION OF THE INVENTION

[0030] Section headings are used herein for ease of understanding and are not intended to limit the applicability of the techniques and embodiments disclosed in each section to that section alone. Additionally, H.266 terminology is used for ease of understanding only and is not intended to limit the scope of the disclosed techniques. As such, the techniques described herein are applicable to other video codec protocols and designs.

[0031] 1. Overview This specification relates to video coding techniques, particularly to signaling of subpictures, tiles, and slices. These ideas can be applied, individually or in various combinations, to any video coding standard or non-standard video codec that supports multi-layer video coding, such as Versatile Video Coding (VVC), which is under development. 2. Abbreviations APS Adaptation Parameter Set AU Access Unit AUD Access Unit Delimiter AVC Advanced Video Coding CLVS Coded Layer Video Sequence CPB Coded Picture Buffer CRA Clean Random Access CTU Coding Tree Unit CVS Coded Video Sequence DPB Decoded Picture Buffer DPS Decoding Parameter Set EOB End Of Bitstream EOS End of Sequence GDR Gradual Decoding Refresh HEVC High Efficiency Video Coding HRD Hypothetical Reference Decoder IDR Instantaneous Decoding Refresh JEM Joint Exploration Model MCTS Motion-Constrained Tile Sets NAL Network Abstraction Layer OLS Output Layer Set PH Picture Header PPS Picture Parameter Set PTL Profile, Tier and Level PU Picture Unit RBSP Raw Byte Sequence Payload SEI Supplemental Enhancement Information SPS Sequence Parameter Set SVC Scalable Video Coding VCL Video Coding Layer VPS Video Parameter Set VTM VVC Test Model VUI Video Usability Information VVC Versatile Video Coding

[0032] 3. Initial discussion Video coding standards have evolved primarily through the development of well-known ITU-T and ISO / IEC standards. ITU-T created H.261 and H.263, while ISO / IEC created MPEG-1 and MPEG-4 Visual. The two organizations jointly developed the H.262 / MPEG-2 Video, H.264 / MPEG-4 Advanced Video Coding (AVC), and H.265 / HEVC standards. Since H.262, video coding standards have been based on a hybrid video coding architecture, which utilizes temporal prediction and transform coding. To explore future video coding technologies beyond HEVC, the Joint Video Exploration Team (JVET) was jointly established by VCEG and MPEG in 2015. Since then, many new methods have been adopted by the JVET and incorporated into reference software named the Joint Exploration Model (JEM). JVET meetings are held quarterly, and the new coding standard aims to achieve a 50% bitrate reduction compared to HEVC. The new video coding standard was officially named Versatile Video Coding (VVC) at the April 2018 JVET meeting, and the first version of the VVC Test Model (VTM) was released at that time. As ongoing efforts contribute to VVC standardization, new coding techniques are adopted into the VVC standard at each JVET meeting. The VVC Working Draft and Test Model VTM are updated at each subsequent meeting. The VVC project is currently aiming for technical completion (FDIS) at the July 2020 meeting.

[0033] 3.1. Picture Partitioning Scheme in HEVC HEVC includes four different picture partitioning schemes, namely, regular slice, dependent slice, tile, and Wavefront Parallel Processing (WPP), which can be applied for maximum transmission unit (MTU) size matching, parallel processing, and reduced end-to-end delay.

[0034] Regular slices are similar to H.264 / AVC: each regular slice is encapsulated in its own NAL unit, and in-picture prediction (intra-sample prediction, motion information prediction, coding mode prediction) and entropy coding dependencies across slice boundaries are disabled. Therefore, regular slices can be reconstructed independently from other regular slices in the same picture (although there may still be interdependencies due to loop filtering operations).

[0035] Regular slices are the only tool available for parallelization, and they are also available in H.264 / AVC in nearly identical formats. Parallelization based on regular slices does not require much inter-processor or inter-core communication (except for inter-processor or inter-core data sharing for motion compensation when decoding predictively coded pictures, which is typically much heavier than inter-processor or inter-core data sharing for in-picture prediction). However, for the same reasons, the use of regular slices can incur significant coding overhead due to the bit cost of slice headers and the lack of prediction across slice boundaries. Furthermore, regular slices (in contrast to the other tools described below) also serve as a key mechanism for bitstream partitioning to accommodate MTU size requirements due to their in-picture independence and the fact that each regular slice is encapsulated in its own NAL unit. Often, the goals of parallelization and MTU size matching impose conflicting requirements on the slice layout in a picture. Recognizing this situation, the parallelization tool described below was developed.

[0036] Dependent slices have short slice headers and allow for partitioning of the bitstream at treeblock boundaries without interrupting any in-picture prediction. Essentially, dependent slices provide fragmentation of regular slices into multiple NAL units to reduce end-to-end delay by allowing transmission of parts of a regular slice before the encoding of the entire regular slice is finished.

[0037] In WPP, pictures are partitioned into single-row coding treeblocks (CTBs). Entropy decoding and prediction are enabled using data from CTBs in other partitions. Parallel processing is enabled by parallel decoding of CTB rows; the start of decoding of a CTB row is delayed by two CTBs, thereby ensuring that data related to the CTB immediately above and to the right of the subject CTB is available before the subject CTB is decoded. This staggered start (graphically resembling a wavefront) allows parallelization across processors / cores up to the number of CTB rows a picture contains. Because in-picture prediction between adjacent treeblock rows within a picture is possible, the inter-processor / inter-core communication required to enable in-picture prediction may be sufficient. WPP partitioning does not result in the generation of additional NAL units compared to when it is not applied, and therefore WPP is not a tool for MTU size matching. However, if MTU size matching is required, regular slices can be used in WPP, with certain coding overhead.

[0038] Tiles define horizontal and vertical boundaries that partition a picture into tile columns and tile rows. Tile columns extend from the top to the bottom of the picture. Similarly, tile columns extend from the left to the right of the picture. The number of tiles in a picture can be derived by simply multiplying the number of tile columns by the number of tile rows.

[0039] The scan order of the CTBs is changed to be local within a tile (in the order of the tile's CTB raster scan) before decoding the top-left CTB of the next tile in the order of the tile's raster scan of the picture. Like regular slices, tiles impair intra-picture prediction and entropy decoding dependencies. However, they do not need to be included in individual NAL units (similar to WPP in this respect), and therefore tiles cannot be used for MTU size matching. Each tile can be processed by one processor / core, and the inter-processor / inter-core communication required for in-picture prediction between processing units decoding adjacent tiles is limited to conveying a shared slice header if the slice spans two or more tiles, and sharing related to loop filtering of reconstructed samples and metadata. If a slice contains more than one tile or WPP segment, the entry point byte offset of each tile or WPP segment in the slice, except for the first tile or WPP segment, is signaled in the slice header.

[0040] For simplicity, HEVC specifies restrictions on the application of four different picture partitioning schemes. A given coded video sequence cannot contain both tiles and wavefronts for most profiles specified in HEVC. For each slice and tile, one or both of the following conditions must be met: 1) All coded treeblocks in a slice belong to the same tile. 2) All coded treeblocks in a tile belong to the same slice. Finally, a wavefront segment contains exactly one CTB row, and when WPP is used, if a slice starts in a CTB row, it must end in the same CTB row.

[0041] A recent amendment to HEVC is specified in the JCT-VC output document JCTVC-AC1005, J. Boyce, A. Ramasubramonian, R. Skupin, G.J. Sullivan, A. Tourapis, and Y.-K. Wang (editors), "HEVC Additional Supplemental Enhancement Information (Draft 4)", October 24, 2017, which is available here: http: / / phenix.int-evry.fr / jct / doc_end_user / documents / 29_Macau / wg11 / JCTVC-AC1005-v2.zip. HEVC, including this amendment, specifies three MCTS-related SEI messages: the Temporal MCTS SEI message, the MCTS Extraction Information Set SEI message, and the MCTS Extraction Information Nest SEI message.

[0042] The temporal MCTS SEI message indicates the presence of an MCTS in the bitstream and signals the MCTS. For each MCTS, motion vectors are restricted to point to full sample locations within the MCTS and fractional-sample locations that require only full sample locations within the MCTS for interpolation, and the use of motion vector candidates derived from blocks outside the MCTS for temporal motion vector prediction is not permitted. In this way, each MCTS can be decoded independently, without the presence of tiles not included in the MCTS.

[0043] The MCTS Extraction Information Set SEI message provides supplemental information that can be used in MCTS sub-bitstream extraction (specified as part of the semantics of the SEI message) to generate a conforming bitstream for an MCTS set. This information consists of several extraction information sets, each of which defines several MCTS sets and contains the RBSP bytes of alternative VPS, SPS, and PPS used during the MCTS sub-bitstream extraction process. When extracting a sub-bitstream according to the MCTS sub-bitstream extraction process, the parameter sets (VPS, SPS, and PPS) need to be rewritten or replaced, and the slice header needs to be slightly updated because one or all of the slice address-related syntax elements (including first_slice_segment_in_pic_flag and slice_segment_address) usually need to have different values.

[0044] 3.2. Picture Partitioning in VVC In VVC, a picture is divided into one or more tile rows and one or more tile columns. A tile is a sequence of CTUs that cover a rectangular area of ​​the picture. The CTUs within a tile are scanned in raster scan order within that tile.

[0045] A slice consists of an integer number of complete tiles or an integer number of contiguous complete CTU rows within a tile of a picture.

[0046] Two modes of slicing are supported: raster scan slice mode and rectangular slice mode. In raster scan slice mode, a slice contains a sequence of complete tiles in the tile raster scan of the picture. In rectangular slice mode, a slice contains either several complete tiles that collectively form a rectangular area of ​​the picture, or several contiguous complete CTU rows of one tile that collectively form a rectangular area of ​​the picture. The tiles within a rectangular slice are scanned in tile raster scan order within the rectangular area corresponding to the slice.

[0047] A subpicture contains one or more slices that collectively cover a rectangular area of ​​the picture.

[0048] Figure 1 shows an example of raster scan slice partitioning of a picture, where the picture is divided into 12 tiles and 3 raster scan slices.

[0049] FIG. 2 shows an example of rectangular slice partitioning of a picture, where the picture is divided into 24 tiles (6 tile rows and 4 tile columns) and 9 rectangular slices.

[0050] FIG. 3 shows an example of a picture partitioned into tiles and rectangular slices, where the picture is divided into four tiles (two tile rows and two tile columns) and four rectangular slices.

[0051] Figure 4 shows an example of subpicture partitioning of a picture, where the picture is partitioned into 18 tiles, 12 on the left covering one slice of 4x4 CTU and 6 on the right covering two vertically stacked slices of 2x2 CTU, resulting in a total of 24 slices and 24 subpictures of various dimensions (each slice is a subpicture).

[0052] 3.3. Signaling Subpictures, Tiles, and Slices in VVC In the latest VVC draft text, subpicture information, including the subpicture layout (i.e., the number of subpictures for each picture and the position and size of each picture) and other sequence-level subpicture information, is signaled in the SPS. The order of subpictures signaled in the SPS defines the subpicture index. A list of subpicture IDs, one for each subpicture, can be explicitly signaled, for example, in the SPS or PPS.

[0053] Tiles in VVC are conceptually the same as in HEVC, i.e., each picture is partitioned into tile columns and tile rows, but the PPS syntax for signaling tiles is different.

[0054] In VVC, the slice mode is also signaled in the PPS. When the slice mode is rectangular slice mode, the slice layout of each picture (i.e., the number of slices for each picture and the position and size of each slice) is signaled in the PPS. The order of rectangular slices within a picture signaled in the PPS defines the picture-level slice index. The sub-picture level slice index is defined as the order of slices within a sub-picture in ascending order of their picture-level slice index. The position and size of rectangular slices are signaled / derived based on the sub-picture position and size signaled in the SPS (if each sub-picture contains only one slice) or based on the tile position and size signaled in the PPS (if a sub-picture can contain more than one slice). When the slice mode is raster scan slice mode, the layout of slices within a picture, along with various details, is signaled in the slice itself, similar to HEVC.

[0055] The syntax and semantics of the SPS, PPS and slice header in the latest VVC draft text that is most relevant to the present invention are as follows: [Table 1] TIFF2026012196000003.tif240170 TIFF2026012196000004.tif87170 [Table 2] TIFF2026012196000006.tif221170 TIFF2026012196000007.tif226170 TIFF2026012196000008.tif208170 [Table 3] TIFF2026012196000010.tif231170 TIFF2026012196000011.tif92170

[0056] 4. Examples of technical problems solved by the solutions herein Existing designs for signaling subpictures, tiles, and slices in VVC have the following problems: 1) The coding of sps_num_subpics_minus1 is u(8), which does not allow more than 256 subpictures per picture. However, certain applications may require the maximum number of subpictures per picture to exceed 256. 2) It is allowed to have subpics_present_flag equal to 0 and sps_subpic_id_present_flag equal to 1. However, this does not make sense because subpics_present_flag equal to 0 means that CLVS has no information about the subpicture at all. 3) A list of sub-picture IDs, one for each sub-picture, can be signaled in the picture header (PH). However, when a list of sub-picture IDs is signaled in the PH and a subset of sub-pictures is extracted from the bitstream, all PHs need to be completely modified. This is undesirable. 4) Currently, when subpicture IDs are indicated to be explicitly signaled by sps_subpic_id_present_flag equal to 1 (or the syntax element is renamed to subpic_ids_explicitly_signalled_flag), the subpicture IDs may not be signaled anywhere. This is problematic because when subpicture IDs are indicated to be explicitly signaled, they need to be explicitly signaled in either the SPS or PPS. 5) When subpicture IDs are not explicitly signaled, the slice header syntax element slice_subpic_id still needs to be signaled as long as subpics_present_flag is equal to 1, including when sps_num_subpics_minus1 is equal to 0. However, the length of slice_subpics_id is currently specified as Ceil(Log2(sps_num_subpics_minus1+1)) bits, or 0 bits if sps_num_subpics_minus1 is equal to 0. This is problematic because the length of current syntax elements cannot be 0 bits. 6) The sub-picture layout, including the number, size, and position of sub-pictures, remains unchanged across CLVS. Even if the sub-picture ID is not explicitly signaled in the SPS or PPS, the length of the sub-picture ID must be signaled for the sub-picture ID syntax element in the slice header. 7) Whenever rect_slice_flag is equal to 1, the syntax element slice_address is signaled in the slice header and specifies the slice index within the subpicture that contains the slice, including when the number of slices in the subpicture (i.e., NumSlicesInSubpic[SubPicIdx]) is equal to 1. However, currently, when rect_slice_flag is equal to 1, slice_address specifies that its length is Ceil(Log2(NumSlicesInSubpic[SubPicIdx])) bits, which is 0 bits when NumSlicesInSubpic[SubPicIdx] is equal to 1. This is problematic because the length of a current syntax element cannot be 0 bits. 8) There is redundancy between the syntax element no_pic_partition_flag and the syntax element pps_num_subpics_minus1, subject to the following constraints of the latest VVC text: if sps_num_subpics_minus1 is greater than 0, the value of no_pic_partition_flag must be equal to 1.

[0057] 5. Exemplary Embodiments and Solutions In order to solve the above problems and others, the following summarized methods are disclosed. The present invention should be considered as an example to explain the general concept and should not be interpreted narrowly. Furthermore, these inventions can be applied alone or in any combination.

[0058] To solve the first problem, change the coding of sps_num_subpics_minus1 from u(8) to ue(v), enabling more than 256 subpictures per picture. a. Additionally, the value of sps_num_subpics_minus1 is restricted to the range 0 to Ceil(pic_width_max_in_luma_samples÷CtbSizeY)*Ceil(pic_height_max_in_luma_samples÷CtbSizeY)-1. b. Furthermore, the number of sub-pictures per picture is further restricted in the definition of a level. 2) To solve the second problem, "if(subpics_present_flag)" is used as a condition for signaling the syntax element sps_subpic_id_present_flag, i.e., the syntax element sps_subpic_id_present_flag is not signaled when subpics_present_flag is equal to 0, and the value of sps_subpic_id_present_flag is inferred to be equal to 0 when not present. a. Alternatively, the syntax element sps_subpic_id_present_flag is still signaled when subpics_present_flag is equal to 0, but is required to have a value equal to 0 when subpics_present_flag is equal to 0. b. Additionally, the names of the syntax elements subpics_present_flag and sps_subpic_id_present_flag are changed to subpic_info_present_flag and subpic_ids_explicity_signalled_flag, respectively. 3) To solve the third problem, the signaling of subpicture IDs in the PH syntax is removed. Thus, the list SubpicIdList[i] for i ranging from 0 to sps_num_subpics_minus1 is derived as follows: for(i=0; i<=sps_num_subpics_minus1; i++) if(subpic_ids_explicitly_signalled_flag) SubpicIdList[i]=subpic_ids_in_pps_flag?pps_subpic_id[i]:sps_subpic_id[i] else SubpicIdList[i]=i 4) To solve the fourth problem, when a sub-picture is indicated to be explicitly signaled, the sub-picture ID is signaled in either the SPS or PPS. a. This is achieved by adding the following constraint: if subpic_ids_explicity_signalled_flag is 0 or subpic_ids_in_sps_flag is equal to 1, then subpic_ids_in_pps_flag MUST be equal to 0. Otherwise (subpic_ids_explicity_signalled_flag is 1 and subpic_ids_in_sps_flag is equal to 0), then subpic_ids_in_pps_flag MUST be equal to 1. 5) To solve the fifth and sixth problems, the length of the subpicture ID is signaled in the SPS regardless of the value of the SPS flag sps_subpic_id_present_flag (or renamed to subpic_ids_explicitly_signaled_flag), but if the subpicture ID is explicitly signaled in the PPS to avoid analyzing the dependency of the PPS on the SPS, the length may also be signaled in the PPS. In this case, the length also specifies the length of the subpicture ID in the slice header, and even the subpicture ID is not explicitly signaled in the SPS or PPS. Therefore, the length of slice_subpic_id is also specified by the length of the subpicture ID signaled in the SPS, if present. 6) Alternatively, to solve the fifth and sixth problems, a flag is added to the SPS syntax, with a value of 1 to specify the presence of subpicture ID length in the SPS syntax. This flag exists independently of the value of the flag indicating whether subpicture IDs are explicitly signaled in the SPS or PPS. The value of this flag can be 1 or 0 if subpic_ids_explicitly_signalled_flag is equal to 0, but the value of the flag must be equal to 1 if subpic_ids_expliclity_signalled_flag is equal to 1. If this flag is equal to 0, i.e., if no subpicture length exists, the length of slice_subpic_id is specified as Max(Log2(sps_num_subpics_minus1 + 1)), 1) bits (instead of Ceil(Log2(sps_num_subpics_minus1 + 1)) bits in the latest VVC draft text). Alternatively, this flag is present only if subpic_ids_explicity_signalled_flag is equal to 0, and the value of this flag is inferred to be equal to 1 if subpic_ids_explicitly_signalled_flag is equal to 1. 7) To solve the seventh problem, if rect_slice_flag is equal to 1, the length of slice_address is Max( Ceil( Log2( NumSlicesInSubpic[SubPicIdx])),1) It is specified in bits. a. Alternatively, and further, if rect_slice_flag is equal to 0, the length of slice_address is specified to be Max(Ceil(Log2(NumTilesInPic))), 1) bits instead of Ceil(Log2(NumTilesInPic)) bits. 8) To solve the eighth problem, the signaling of no_pic_partition_flag is conditioned on "if(subpic_ids_in_pps_flag && pps_num_subpics_minus1 > 0)" and the following presumption is added: if not present, the value of no_pic_partition_flag is presumed to be equal to 1. a. Alternatively, move the subpicture ID syntax (all four syntax elements) after the tile and slice syntax in the PPS, for example, just before the syntax element entropy_coding_sync_enabled_flag, and then condition the signaling of pps_num_subpics_minus1 to "if(no_pic_partition_flag)". 6. Implementation Below are some illustrative examples of all aspects of the invention, except for item 8, summarized in section 5 above, that can be applied to the VVC specification. The modified text is based on the latest VVC text in JVET-P2001-v14. The most relevant parts that have been added or modified are: Underlined, bold, and italicized text and the most relevant deletions are highlighted in bold double brackets, e.g., [[a]] indicates that an "a" was deleted. Some changes are not highlighted because they are editorial in nature. 6.1. First embodiment [Table 4] TIFF2026012196000013.tif233170 TIFF2026012196000014.tif123170 [Table 5] TIFF2026012196000016.tif206170 [Table 6] TIFF2026012196000018.tif36170

[0059] 5 is a block diagram illustrating an example video processing system 500 in which various techniques disclosed herein may be implemented. Various implementations may include some or all of the components of system 500. System 500 may include an input 502 for receiving video content. The video content may be received in a raw or uncompressed format, e.g., 8- or 10-bit multi-component pixel values, or may be received in a compressed or encoded format. Input 502 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, passive optical networks (PONs), and wireless interfaces such as Wi-Fi or cellular interfaces.

[0060] System 500 may include a coding component 504 that may implement various coding or encoding methods described herein. Coding component 504 may reduce the average bitrate of video from input 502 to the output of coding component 504 to generate a coded representation of the video. Accordingly, coding techniques are sometimes referred to as video compression or video transcoding techniques. The output of coding component 504 may be stored or transmitted via a connected communication, as represented by component 506. The stored or communicated bitstream (or coded) representation of the video received at input 502 may be used by component 508 to generate pixel values ​​or displayable video sent to display interface 510. The process of generating user-viewable video from the bitstream representation is sometimes referred to as video decompression. Furthermore, while certain video processing operations are referred to as “coding” operations or tools, it will be understood that the coding tools or operations are used in an encoder, and that corresponding decoding tools or operations that reverse the results of the coding are performed by a decoder.

[0061] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB) or High-Definition Multimedia Interface (HDMI®) or DisplayPort, etc. Examples of storage interfaces include SATA (Serial Advanced Technology Attachment), PCI, IDE interfaces, etc. The techniques described herein may be embodied in various electronic devices such as mobile phones, laptops, smartphones, or other devices capable of performing digital data processing and / or video display.

[0062] FIG. 6 is a block diagram of a video processing device 600. The device 600 may be used to implement one or more of the methods described herein. The device 600 may be embodied in a smartphone, a tablet, a computer, an Internet of Things (IoT) receiver, etc. The device 600 may include one or more processors 602, one or more memories 604, and video processing hardware 606. The processor(s) 602 may be configured to implement one or more of the methods described herein. The memory(s) 604 may be used to store data and code used to implement the methods and techniques described herein. The video processing hardware 606 may be used to implement some of the techniques described herein in a hardware circuit. In some embodiments, the hardware 606 may be partially or entirely within the processor 602, e.g., a graphics processor.

[0063] FIG. 7 is a block diagram illustrating an example video coding system 100 that may utilize the techniques of this disclosure.

[0064] 7, video coding system 100 may include source device 110 and destination device 120. Source device 110 generates encoded video data, which may be referred to as a video encoding device. Destination device 120 may decode the encoded video data generated by source device 110, which may be referred to as a video decoding device.

[0065] The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.

[0066] The video source 112 may include a source such as a video capture device, an interface for receiving video data from a video content provider, and / or a computer graphics system for generating video data, or a combination of such sources. The video data may include one or more pictures. The video encoder 114 encodes the video data from the video source 112 to generate a bitstream. The bitstream may include a sequence of bits forming a coded representation of the video data. The bitstream may include coded pictures and associated data. A coded picture is a coded representation of a picture. The associated data may include sequence parameter sets, picture parameter sets, and other syntax structures. The I / O interface 116 may include a modulator / demodulator (modem) and / or transmitter. The encoded video data may be transmitted directly to the destination device 120 via the I / O interface 116 over the network 130a. The encoded video data may also be stored on a storage medium / server 130b for access by the destination device 120.

[0067] The destination device 120 may include an I / O interface 126 , a video decoder 124 , and a display device 122 .

[0068] I / O interface 126 may include a receiver and / or a modem. I / O interface 126 may obtain encoded video data from source device 110 or storage medium / server 130b. Video decoder 124 may decode the encoded video data. Display device 122 may display the decoded video data to a user. Display device 122 may be integrated with destination device 120 or may be external to destination device 120 and configured to interface with an external display device.

[0069] Video encoder 114 and video decoder 124 may operate according to video compression standards such as the High Efficiency Video Coding (HEVC) standard, the Versatile Video Coding (VVM) standard, and other current and / or future standards.

[0070] FIG. 8 is a block diagram illustrating an example of a video encoder 200, which may be video encoder 114 in system 100 shown in FIG.

[0071] Video encoder 200 may be configured to perform any or all of the techniques of this disclosure. In the example of FIG. 8, video encoder 200 includes multiple functional components. The techniques described in this disclosure may be shared among various components of video encoder 200. In some examples, a processor may be configured to perform any or all of the techniques described in this disclosure.

[0072] The functional components of the video encoder 200 may include a partition unit 201, a prediction unit 202, which may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra prediction unit 206, a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213, and an entropy encoding unit 214.

[0073] In other examples, video encoder 200 may include more, fewer, or different functional components. In one example, prediction unit 202 may include an intra block copy (IBC) unit. The IBC unit may perform prediction in an IBC mode in which at least one reference picture is the picture in which the current video block is located.

[0074] Furthermore, some components, such as the motion estimation unit 204 and the motion compensation unit 205, may be highly integrated, but are shown separately in the example of FIG. 8 for illustrative purposes.

[0075] Partition unit 201 may partition a picture into one or more video blocks. Video encoder 200 and video decoder 300 may support a variety of video block sizes.

[0076] Mode select unit 203 may select one of intra- or inter-coding modes, for example, based on an error result, and may provide the resulting intra- or inter-coded block to residual generation unit 207 to generate residual block data and to reconstruction unit 212 to reconstruct an encoded block for use as a reference picture. In some examples, mode select unit 203 may select a combination of intra and inter predication (CIIP) mode, in which prediction is based on an inter-prediction signal and an intra-prediction signal. Mode select unit 203 may also select the resolution of motion vectors for blocks (e.g., sub-pixel or integer-pixel precision) in the case of inter prediction.

[0077] To perform inter prediction on the current video block, motion estimation unit 204 may generate motion information for the current video block by comparing one or more reference frames from buffer 213 with the current video block. Motion compensation unit 205 may determine a predicted video block for the current video block based on the motion information and decoded samples of pictures from buffer 213 other than the picture associated with the current video block.

[0078] Motion estimation unit 204 and motion compensation unit 205 may perform different operations on the current video block depending on whether the current video block is in an I slice, a P slice, or a B slice, for example.

[0079] In some examples, motion estimation unit 204 may perform uni-directional prediction on the current video block, and motion estimation unit 204 may search reference pictures in list 0 or list 1 for a reference video block for the current video block. Motion estimation unit 204 may then generate a reference index indicating a reference picture in list 0 or list 1 that includes the reference video block and a motion vector that indicates a spatial displacement between the current video block and the reference video block. Motion estimation unit 204 may output the reference index, the prediction direction indicator, and the motion vector as motion information for the current video block. Motion compensation unit 205 may generate a predictive video block for the current block based on the reference video block indicated by the motion information of the current video block.

[0080] In another example, motion estimation unit 204 may perform bi-directional prediction on the current video block, and motion estimation unit 204 may search reference pictures in list 0 for reference video blocks for the current video block and may also search reference pictures in list 1 for reference video blocks for the current video block. Motion estimation unit 204 may then generate reference indexes indicating reference pictures in lists 0 and 1 that include the reference video blocks and motion vectors that indicate spatial displacements between the reference video blocks and the current video block. Motion estimation unit 204 may output the reference index and motion vector of the current video block as motion information for the current video block. Motion compensation unit 205 may generate a predictive video block for the current video block based on the reference video blocks indicated by the motion information of the current video block.

[0081] In some examples, the motion estimation unit 204 may output a full set of motion information for the decoding process of the decoder.

[0082] In some examples, motion estimation unit 204 may not output a complete set of motion information for the current video. Rather, motion estimation unit 204 may signal the motion information of the current video block with reference to motion information of other video blocks. For example, motion estimation unit 204 may determine that the motion information of the current video block is sufficiently similar to the motion information of neighboring video blocks.

[0083] In one example, motion estimation unit 204 may indicate, in a syntax structure associated with the current video block, a value that indicates to video decoder 300 that the current video block has the same motion information as another video block.

[0084] In another example, motion estimation unit 204 may identify another video block and a motion vector difference (MVD) in a syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. Video decoder 300 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.

[0085] As mentioned above, video encoder 200 may predictively signal motion vectors. Two examples of predictive signaling techniques that may be implemented by video encoder 200 include advanced motion vector prediction (AMVP) and merge mode signaling.

[0086] Intra prediction unit 206 may perform intra prediction on the current video block. When intra prediction unit 206 performs intra prediction on the current video block, intra prediction unit 206 may generate predictive data for the current video block based on decoded samples of other video blocks within the same picture. The predictive data for the current video block may include the video block to be predicted and various syntax elements.

[0087] Residual generation unit 207 may generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) the predicted video block(s) of the current video block from the current video block. The residual data for the current video block may include residual video blocks that correspond to different sample components of the samples in the current video block.

[0088] In other examples, for example, in skip mode, residual data for the current video block may not exist and residual generation unit 207 may not perform a subtraction operation.

[0089] Transform processing unit 208 may generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video block related to the current video block.

[0090] After transform processing unit 208 generates a transform coefficient video block associated with the current video block, quantization unit 209 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values ​​associated with the current video block.

[0091] Inverse quantization unit 210 and inverse transform unit 211 may apply inverse quantization and inverse transform, respectively, to the transform coefficient video block to reconstruct a residual video block from the transform coefficient video block. Reconstruction unit 212 may add the reconstructed residual video block to corresponding samples from one or more predicted video blocks generated by prediction unit 202 to generate a reconstructed video block related to the current block for storage in buffer 213.

[0092] After reconstruction unit 212 reconstructs the video blocks, it may perform a loop filtering operation to reduce video blocking artifacts in the video blocks.

[0093] Entropy encoding unit 214 may receive data from other functional components of video encoder 200. If entropy encoding unit 214 receives data, it may perform one or more entropy encoding operations to generate entropy encoded data and output a bitstream including the entropy encoded data.

[0094] FIG. 9 is a block diagram illustrating an example of a video decoder 300, which may be video decoder 114 in system 100 shown in FIG.

[0095] Video decoder 300 may be configured to perform any or all of the techniques of this disclosure. In the example of FIG. 8, video decoder 300 includes multiple functional components. The techniques described in this disclosure may be shared among various components of video decoder 300. In some examples, a processor may be configured to perform any or all of the techniques described in this disclosure.

[0096] 9, video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. Video decoder 300 may, in some examples, perform a decoding path that is generally the reverse of the encoding path described with respect to video encoder 200 (FIG. 8).

[0097] The entropy decoding unit 301 may retrieve an encoded bitstream. The encoded bitstream may include entropy-coded video data (e.g., encoded blocks of video data). The entropy decoding unit 301 may decode the entropy-coded video data, and from the entropy-decoded video data, the motion compensation unit 302 may determine motion information including motion vectors, motion vector precision, reference picture list indexes, and other motion information. The motion compensation unit 302 may determine such information by, for example, implementing AMVP and merge mode.

[0098] The motion compensation unit 302 may generate motion-compensated blocks and may perform interpolation based on interpolation filters. Identifiers for the interpolation filters to be used with sub-pixel precision may be included in the syntax elements.

[0099] Motion compensation unit 302 may calculate interpolated values ​​for sub-integer pixels of the reference block using interpolation filters used by video encoder 200 during encoding of the video block. Motion compensation unit 302 may determine the interpolation filters used by video encoder 200 according to received syntax information and use the interpolation filters to generate the predictive block.

[0100] The motion compensation unit 302 may use some of the syntax information to determine the size of the blocks used to encode the frame(s) and / or slice(s) of the encoded video sequence, partition information describing how each macroblock of a picture of the encoded video sequence is partitioned, a mode indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-encoded block, and other information for decoding the encoded video sequence.

[0101] The intra prediction unit 303 may form a prediction block from spatially neighboring blocks, for example, using an intra prediction mode received in the bitstream. The inverse quantization unit 303 inverse quantizes, i.e., dequantizes, the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 301. The inverse transform unit 303 applies an inverse transform.

[0102] Reconstruction unit 306 may sum the residual blocks with corresponding prediction blocks generated by motion compensation unit 202 or intra prediction unit 303 to form decoded blocks. Optionally, a deblocking filter may be applied to filter the decoded blocks to remove blockiness artifacts. The decoded video blocks are then stored in buffer 307, which provides reference blocks for subsequent motion compensation / intra prediction and also generates decoded video for presentation on a display device.

[0103] 10-12 illustrate exemplary ways in which the technical solutions described above in the embodiments illustrated in, for example, FIGS. 5-9 can be implemented.

[0104] 10 shows a flowchart for an example method of video processing 1000. The method 1000 includes, at operation 1010, performing a conversion between video and a video bitstream, where the bitstream complies with a format rule that specifies that signaling a length of a sub-picture identifier in a sequence parameter set (SPS) is not based on a value of a syntax element indicating whether the identifier is signaled in the SPS.

[0105] 11 shows a flowchart for an example method of video processing 1100. The method 1100 includes, at operation 1110, performing a conversion between video and a video bitstream, the bitstream complying with format rules that specify signaling a length of a sub-picture identifier in a Picture Parameter Set (PPS) by a syntax element indicating that the identifier is explicitly signaled in the PPS.

[0106] 12 shows a flowchart for an example method 1200 of video processing. The method 1200 includes, at operation 1210, performing a conversion between video and a video bitstream, where the bitstream complies with a format rule that specifies that a first syntax element signaled in a Sequence Parameter Set (SPS) of the bitstream indicates a length of a sub-picture identifier in the SPS, and the signaling of the first syntax element is independent of a value of a second syntax element that indicates explicit signaling of a sub-picture identifier in the SPS or Picture Parameter Set (PPS).

[0107] Next, a list of preferred solutions is provided according to some embodiments.

[0108] A1. A video processing method comprising: performing a conversion between a video including a picture and a bitstream of the video, wherein the number of sub-pictures in the picture is signaled in a sequence parameter set (SPS) of the bitstream as a field whose bit width is based on the value of the number of sub-pictures, and the field is an unsigned integer zeroth-order Exponential-Golomb (Exp-Golomb) coded syntax element with the left bit first.

[0109] A2. The method of Solution A1, wherein the value of the field is constrained to be in the range of zero to a maximum value based on the picture's maximum width in luma samples and the picture's maximum height in luma samples.

[0110] A3. The method of solution A2, wherein the maximum value is equal to an integer number of coding tree blocks that fit within a picture.

[0111] A4. The method of Solution A1, wherein the number of sub-pictures is limited based on a coding level associated with the bitstream.

[0112] A5. A video processing method comprising: performing a conversion between video and a video bitstream, wherein the bitstream conforms to a format rule that specifies that a first syntax element indicating whether a picture of the video can be partitioned is conditionally included in a picture parameter set (PPS) of the bitstream based on the values ​​of a second syntax element indicating whether a sub-picture identifier is signaled in the PPS, and a third syntax element indicating the number of sub-pictures in the PPS.

[0113] A6. The method of Solution A5, wherein the first syntax element is no_pic_partition_flag, the second syntax element is subpic_ids_in_pps_flag, and the third syntax element is pps_num_subpics_minus1.

[0114] A7. The method of solution A5 or A6, wherein the first syntax element is excluded from the PPS and is inferred to indicate that picture partitioning is not applied to each picture that references the PPS.

[0115] A8. The method of solution A5 or A6, wherein the first syntax element is not signaled in the PPS and is presumed to be equal to 1.

[0116] A9. The method of Solution A5 or A6, wherein the second syntax element is signaled after one or more tile and / or slice syntax elements of the PPS.

[0117] A10. A video processing method comprising: performing a conversion between video and a video bitstream, wherein the bitstream conforms to a format rule, and wherein the format rule specifies that a first syntax element indicating whether a picture of the video can be partitioned is included in a picture parameter set (PPS) of the bitstream, before a set of syntax elements in the PPS that indicate identifiers of sub-pictures of the picture.

[0118] A11. The method of Solution A10, wherein a second syntax element indicating a number of sub-pictures is conditionally included in the set of syntax elements based on the value of the first syntax element.

[0119] A12. The method of any of Solutions A1-A11, wherein the converting includes decoding the video from the bitstream.

[0120] A13. The method of any of Solutions A1-A11, wherein the converting includes encoding the video into the bitstream.

[0121] A14. A method for storing a bitstream representing a video on a computer-readable recording medium, the method comprising generating a bitstream from the video according to a method described in any one or more of Solutions A1 to A11, and writing the bitstream to a computer-readable recording medium.

[0122] A15. A video processing device including a processor configured to implement the methods described in any one or more of Solutions A1-A14.

[0123] A16. A computer-readable medium having stored thereon instructions that, when executed, cause a processor to implement a method according to any one or more of solutions A1-A14.

[0124] A17. A computer-readable medium storing a bitstream generated according to any one or more of solutions A1-A14.

[0125] A18. A video processing device for storing a bitstream, said video processing device being configured to implement the methods described in any one or more of Solutions A1 to A14.

[0126] Next, we provide another list of preferred solutions according to some embodiments.

[0127] B1. A video processing method comprising: performing a conversion between a video domain of a video and a bitstream of the video, wherein the bitstream conforms to a format rule, and wherein the format rule specifies that a first syntax element indicating whether sub-picture identifier information is included in a parameter set of the bitstream is conditionally included in a sequence parameter set (SPS) based on the value of a second syntax element indicating whether sub-picture information is included in the SPS.

[0128] B2. The method of Solution B1, wherein the format rule further specifies that the value of the second syntax element is zero, indicating that the information for the sub-pictures is omitted from the SPS, such that each of the video regions associated with the SPS is not divided into multiple sub-pictures, and based on the value being zero, the first syntax element is omitted from the SPS.

[0129] B3. The method of Solution B1 or B2, wherein the first syntax element is subpic_ids_explicity_signalled_flag and the second syntax element is subpic_info_present_flag.

[0130] B4. The method of Solution B1 or B3, wherein the formatting rules further specify that if the value of the first syntax element is zero, then the first syntax element is included in the SPS with a zero value.

[0131] B5. The method of any of Solutions B1 to B4, wherein the video region is a video picture.

[0132] B6. A video processing method comprising: performing a conversion between pictures of a video and a bitstream of the video, wherein the bitstream conforms to a format rule, wherein the format rule specifies that a mapping between identifiers of one or more subpictures of the picture and the one or more subpictures is not included in a picture header of the picture, and wherein the format rule further specifies that identifiers of the one or more subpictures are derived based on syntax elements of a sequence parameter set (SPS) and a picture parameter set (PPS) referenced by the picture.

[0133] B7. The method of Solution B6, wherein a flag in the SPS takes a first value indicating that the identifiers of the one or more subpictures are derived based on syntax elements of the PPS, or a second value indicating that the identifiers of the one or more subpictures are derived based on syntax elements of the SPS.

[0134] B8. The method of Solution B7, wherein the flag corresponds to a subpic_ids_in_pps_flag field, the first value is 1, and the second value is zero.

[0135] B9. The identifier (denoted SubpicIdList[i]) is derived as follows: for(i = 0; i <= sps_num_subpics_minus1; i++) if(subpic_ids_explicitly_signalled_flag) SubpicIdList[i] = subpic_ids_in_pps_flag ?pps_subpic_id[i] sps_subpic_id[i] else SubpicIdList[i] = i Solution B6 method.

[0136] B10. A method of video processing, comprising performing a conversion between video and a bitstream of the video, wherein the bitstream conforms to a format rule, and wherein the format rule specifies that when a value of a first syntax element indicates that a mapping between an identifier of the subpicture and one or more subpictures of a picture is explicitly signaled for the one or more subpictures, the mapping is signaled in either a sequence parameter set (SPS) or a picture parameter set (PPS).

[0137] B11. The method of Solution B10, wherein the first syntax element is subpic_ids_explicity_signalled_flag.

[0138] B12. The method of Solution B11, wherein the identifier is signaled in the SPS based on the value of a second syntax element (denoted subpic_ids_in_sps_flag) and the identifier is signaled in the PPS based on the value of a third syntax element (denoted subpic_ids_in_pps_flag).

[0139] B13.subpic_ids_in_pps_flag is equal to zero because subpic_ids_explicity_signalled_flag is zero or subpic_ids_in_sps_flag is 1, Solution B12 method.

[0140] B14. The method of solution B13, wherein subpic_ids_in_pps_flag equal to zero indicates that no identifiers are signaled in the PPS.

[0141] B15.subpic_ids_in_pps_flag is equal to 1 because subpic_ids_explicity_signalled_flag is 1 and subpic_ids_in_sps_flag is zero, Solution B12 method.

[0142] B16. The method of solution B15, wherein subpic_ids_in_pps_flag equal to 1 indicates that identifiers for each of the one or more subpictures are explicitly signaled in the PPS.

[0143] B17. The method of any of Solutions B1-B15, wherein the converting includes decoding the video from the bitstream.

[0144] B18. The method of any of Solutions B1-B15, wherein the converting includes encoding the video into the bitstream.

[0145] B19. A method for storing a bitstream representing a video on a computer-readable recording medium, the method comprising generating a bitstream from a video according to a method described in one or more of Solutions B1 to B15, and writing the bitstream to the computer-readable recording medium.

[0146] B20. A video processing device including a processor configured to implement the methods described in any one or more of Solutions B1-B19.

[0147] B21. A computer-readable medium having stored thereon instructions that, when executed, cause a processor to implement a method according to one or more of solutions B1-B19.

[0148] B22. A computer-readable medium storing a bitstream generated according to any one or more of solutions B1-B19.

[0149] B23. A video processing device for storing a bitstream, said video processing device being configured to implement the methods described in any one or more of Solutions B1 to B19.

[0150] A further list of solutions preferred by some embodiments is provided below.

[0151] C1. A method of video processing comprising performing a conversion between video and a video bitstream, wherein the bitstream conforms to a format rule, and wherein the format rule specifies that signaling the length of a sub-picture identifier in a sequence parameter set (SPS) is not based on the value of a syntax element indicating whether the identifier is signaled in the SPS.

[0152] C2. A video processing method comprising: performing a conversion between video and a video bitstream, wherein the bitstream complies with a format rule, and wherein the format rule specifies that a length of a sub-picture identifier is signaled in a picture parameter set (PPS) by a syntax element indicating that the identifier is explicitly signaled in the PPS.

[0153] C3. The method of solution C1 or C2, wherein the length further corresponds to the length of a sub-picture identifier in a slice header.

[0154] C4. Any of Solutions C1-C3, wherein the syntax element is subpic_ids_explicity_signalled_flag.

[0155] C5. The method of solution C3 or C4, wherein the identifier is not signaled in the SPS and the identifier is not signaled in the PPS.

[0156] C6. The method of Solution C5, wherein the length of the sub-picture identifier corresponds to a sub-picture identifier length signaled in the SPS.

[0157] C7. A method of video processing, comprising performing a conversion between video and a video bitstream, wherein the bitstream conforms to a format rule, and wherein the format rule specifies that a first syntax element signaled in a sequence parameter set (SPS) of the bitstream indicates a length of an identifier of a sub-picture in the SPS, and wherein the signaling of the first syntax element is independent of a value of a second syntax element indicating explicit signaling of the identifier of the sub-picture in the SPS or a picture parameter set (PPS).

[0158] C8. The method of solution C7, wherein the second syntax element is subpic_ids_explicity_signalled_flag.

[0159] C9. The method of Solution C7 or C8, wherein the second syntax element equal to 1 indicates that a set of identifiers, one for each of the sub-pictures, is explicitly signaled for each of the sub-pictures in the SPS or the PPS, and the second syntax element equal to zero indicates that no identifiers are explicitly signaled in the SPS or the PPS.

[0160] C10. The method of any of solutions C7 or C8, wherein the value of the first syntax element is either zero or one, such that the value of the second syntax element is zero.

[0161] C11. The method of solution C7 or C8, wherein the value of the first syntax element is 1 because the value of the second syntax element is 1.

[0162] C12. The method of solution C7 or C8, wherein the value of the second syntax element is zero.

[0163] C13. The method of any of Solutions C1-C12, wherein the converting includes decoding the video from the bitstream.

[0164] C14. The method of any of Solutions C1-C12, wherein the converting includes encoding the video into the bitstream.

[0165] C15. A method for storing a bitstream representing a video on a computer-readable recording medium, the method comprising generating the bitstream from the video according to a method described in one or more of Solutions C1 to C12, and writing the bitstream to the computer-readable recording medium.

[0166] C16. A video processing device having a processor configured to implement the methods described in any one or more of solutions C1 to C15.

[0167] C17. A computer-readable medium having instructions stored thereon that, when executed, cause a processor to implement a method according to any one or more of solutions C1 to C15.

[0168] C18. A computer-readable medium storing the bitstream generated according to any one or more of solutions C1 to C15.

[0169] C19. A video processing device for storing a bitstream, said video processing device being configured to implement the method described in any one or more of solutions C1 to C15.

[0170] A further list of solutions preferred by some embodiments is provided below.

[0171] P1. A video processing method comprising performing a conversion between a picture of a video and a coded representation of said video, wherein the number of sub-pictures in said picture is included in said coded representation as a field whose bit width depends on the value of the number of sub-pictures.

[0172] P2. The method of solution P1, wherein the field represents the number of the sub-pictures using a codeword.

[0173] P3. The method of solution P2, wherein said codewords comprise Golomb codewords.

[0174] P4. The method of any of Solutions P1-P3, wherein the value of the number of sub-pictures is constrained to be less than or equal to an integer number of coding tree blocks that fit within the picture.

[0175] P5. The method of any of solutions P1 to P4, wherein the field depends on a coding level associated with the coded representation.

[0176] P6. A video processing method comprising performing a conversion between a video region of a video and a coded representation of said video, said coded representation conforming to a format rule, said format rule specifying that syntax elements indicating sub-picture identifiers are to be omitted by said video region not containing sub-pictures.

[0177] P7. The method of solution P6, wherein the coded representation includes a field having a value of 0 indicating that the video region does not contain a sub-picture.

[0178] P8. A method of video processing comprising performing a conversion between a video region of a video and a coded representation of said video, said coded representation conforming to a format rule, said format rule specifying that identifiers of sub-pictures within said video region are to be omitted at a video region header level in said coded representation.

[0179] P9. The method of solution P8, wherein the coded representation numerically identifies sub-pictures according to the order in which the sub-pictures are listed in the video region header.

[0180] P10. A video processing method comprising performing a conversion between a video domain of a video and a coded representation of said video, said coded representation conforming to a format rule, said format rule specifying the inclusion of a sub-picture identifier and / or a length of a sub-picture identifier in said video domain at a sequence parameter set level or a picture parameter set level.

[0181] P11. The method of solution P10, wherein the length is included at the picture parameter set level.

[0182] P12. A method of video processing comprising performing a conversion between a video domain of a video and a coded representation of said video, said coded representation conforming to a format rule, said format rule specifying that a field in said coded representation be included at a video sequence level to indicate whether a sub-picture identifier length field is included in said coded representation at said video sequence level.

[0183] P13. The method of solution P12, wherein the formatting rules specify that the field is set to '1' if another field in the coded representation indicates that a length identifier for the video region is included in the coded representation.

[0184] P14. The method of any of the preceding solutions, wherein the video region includes a sub-picture of the video.

[0185] P15. The method of any of the preceding solutions, wherein the converting includes parsing and decoding the coded representation to generate the video.

[0186] P16. The method of any one of the preceding solutions, wherein the transforming includes encoding the video to generate the coded representation.

[0187] P17. A video decoding device having a processor configured to implement the methods described in one or more of solutions P1 to P16.

[0188] P18. A video encoding device having a processor configured to implement the methods described in one or more of solutions P1 to P16.

[0189] P19. A computer program product storing computer code which, when executed by a processor, causes the processor to implement a method according to any of solutions P1 to P16.

[0190] The disclosed solutions, examples, embodiments, modules, and functional operations described herein can be implemented in digital electronic circuitry, or computer software, firmware, or hardware, including the structures disclosed herein and equivalents of those structures, or one or more combinations thereof. The disclosed and other embodiments can be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer-readable medium for execution by or to control the operation of a data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition of matter affecting a machine-readable propagated signal, or one or more combinations thereof. The term "data processing apparatus" encompasses all apparatus, devices, and machines for processing data, including, for example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, an apparatus can include code that creates an execution environment for the computer program in question, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, or one or more combinations thereof. The propagated signal may be an artificially generated signal, for example, a machine-generated electrical, optical, or electromagnetic signal generated to encode information for transmission to an appropriate receiver device.

[0191] A computer program (also known as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored as part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files storing one or more modules, subprograms, or portions of code). A computer program can be deployed to be executed on one computer or one site, or distributed across multiple sites and interconnected by a communications network.

[0192] The processes and logic flows described herein may be performed by one or more programmable processors executing one or more computer programs to perform functions by manipulating input data and generating output. The processes and logic flows may also be performed by, and an apparatus may be implemented as, special purpose logic circuitry, such as an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit).

[0193] Processors suitable for the execution of a computer program include, by way of example, both general-purpose and special-purpose microprocessors, and any one or more processors of any kind of digital computer. Typically, a processor receives instructions and data from a read-only memory or a random-access memory, or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also be operatively coupled to receive data from, or transfer data to, one or more mass storage devices, e.g., magnetic disks, magneto-optical disks, or optical disks, for storing data. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices, such as EPROMs, EEPROMs, and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and memory may be supplemented by, or incorporated in, special-purpose logic circuitry.

[0194] While this patent document contains many details, these should not be construed as limiting the scope of any subject matter or what may be claimed, but rather as descriptions of features specific to particular embodiments of a particular technology. Certain features described in this patent document in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented in multiple embodiments separately or in any suitable subcombination. Furthermore, while features may be described above as acting in particular combinations and initially claimed as such, one or more features from a claimed combination may, in some cases, be carved out of the combination, and the claimed combination may be directed to subcombinations or variations of the subcombinations.

[0195] Similarly, although the figures depict acts in a particular order, this should not be understood as requiring such acts to be performed in a particular order or sequential order, or to perform all of the illustrated acts, to achieve desired results. Furthermore, the separation of various system components in the embodiments described in this patent document should not be understood as requiring such separation in all embodiments.

[0196] Although only a few embodiments and examples have been described, other embodiments, extensions and variations can be made based on what is described and explained in this patent document.

Claims

1. 1. A method for storing a video bitstream, comprising: generating the bitstream of the video including one or more pictures including one or more sub-pictures; storing the bitstream on a non-transitory computer-readable recording medium; the bitstream conforms to format rules; the format rule specifies that signaling a length of a sub-picture identifier in a sequence parameter set (SPS) of the bitstream is based on a value of a first syntax element, subpic_info_present_flag, in the SPS and is independent of a value of a second syntax element, subpic_info_present_flag, in the SPS; the first syntax element subpic_info_present_flag indicates whether subpicture information is present, and the second syntax element indicates whether the identifier is explicitly signaled; the length is signaled by a sixth syntax element sps_subpic_id_len_minus1, and the value of sps_subpic_id_len_minus1 plus 1 is equal to the length; method.

2. the length corresponds to the number of bits used to represent a third syntax element indicating a sub-picture identifier in the SPS, a fourth syntax element indicating a sub-picture identifier in a picture parameter set (PPS), if present, and a fifth syntax element indicating a sub-picture identifier in a slice header, if present; The method of claim 1.

3. The value of sps_subpic_id_len_minus1 is in the range of 0 to 15. The method of claim 1.

4. The second syntax element equal to one specifies that a set of identifiers, one for each sub-picture, is explicitly signaled for each sub-picture in either the SPS or a Picture Parameter Set (PPS), and the second syntax element equal to zero specifies that no identifiers are explicitly signaled in the SPS or the PPS. The method of claim 1.