Video processing method and device

By modifying the design of sub-pictures, slices and stripes signaling in VVC, the number of sub-pictures, complex signaling notification logic, ID length problems and redundancy problems are solved, and a more flexible and consistent signaling mechanism is achieved.

CN114902677BActive Publication Date: 2025-05-23DOUYIN CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080090816.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-12-27
Filing Date
2020-12-27
Publication Date
2025-05-23
Estimated Expiration
2040-12-27

AI Technical Summary

Technical Problem

There are several problems with the signaling design of sub-pictures, slices and stripes in VVC, including the low number of sub-pictures, the complex logic of sub-picture ID signaling notification, the problem of sub-picture ID length, the problem of strip address length, and the redundancy between no_pic_partition_flag and pps_num_subpics_minus1.

Method used

By modifying the encoding method of sps_num_subpics_minus1, it supports more than 256 sub-pictures; adjusting the conditions and location of sub-picture ID signaling notifications; signaling notification sub-picture ID length in SPS; correcting the length problem of slice_address; and redefining the signaling conditions of no_pic_partition_flag.

Benefits of technology

It solves the problem of sub-picture number limitation, complex signaling notification logic, ID length problems and redundancy problems, and improves the flexibility and consistency of VVC neutron pictures, slices and stripe signaling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114902677B_ABST
    Figure CN114902677B_ABST
Patent Text Reader

Abstract

A method, apparatus, and system for signaling the use of sub-pictures in a coded video picture is described. An example of a video processing method includes: performing conversion between a video and a bitstream of the video, wherein the bitstream conforms to a format rule, wherein the format rule specifies that a first syntax element signaled in a sequence parameter set (SPS) of the bitstream indicates a length of an identifier of a sub-picture in the SPS, and wherein the signaling of the first syntax element is independent of a value of a second syntax element, the value of the second syntax element indicating that the identifier of the sub-picture is explicitly signaled in the SPS or a picture parameter set (PPS).
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application timely claims priority to and the benefit of U.S. Provisional Patent Application No. 62 / 954,364 filed on December 27, 2019, pursuant to applicable provisions of the Patent Act and / or the Paris Convention. The entire disclosure of the foregoing application is incorporated herein by reference as a part of the disclosure of this application for all legal purposes. Technical Field

[0003] This application document relates to image and video encoding and decoding. Background Art

[0004] Digital video accounts for the largest use of bandwidth on the Internet and other digital communications networks. As the number of connected user devices capable of receiving and displaying video increases, bandwidth requirements for digital video usage are expected to continue to grow. Summary of the invention

[0005] This document discloses methods, devices and systems for sub-picture signaling for video encoding and decoding by video encoders and decoders, respectively.

[0006] In one example aspect, a video processing method is disclosed. The method includes performing conversion between a video including a picture and a bitstream of the video, wherein the number of sub-pictures in the picture is signaled as a field in a sequence parameter set (SPS) of the bitstream, the bit width of the field is based on the value of the number of sub-pictures, and wherein the field is a left bit first unsigned integer 0-order exponential Golomb (Exp-Golomb) codec syntax element.

[0007] In another example aspect, a video processing method is disclosed. The method includes: performing conversion between a video and a bitstream of the video, wherein the bitstream conforms to a format rule, and wherein the format rule specifies that a first syntax element indicating whether a picture of the video can be split is conditionally included in a picture parameter set (PPS) of the bitstream based on values ​​of a second syntax element and a third syntax element, wherein the second syntax element indicates whether an identifier of a sub-picture is signaled in the PPS, and the third syntax element indicates the number of sub-pictures in the PPS.

[0008] In yet another example aspect, a video processing method is disclosed. The method includes: performing conversion between a video and a bitstream of the video, wherein the bitstream conforms to a format rule, and wherein the format rule specifies that a first syntax element indicating whether a picture of the video can be split is included in a picture parameter set (PPS) of the bitstream, located before a set of syntax elements in the PPS indicating identifiers of sub-pictures of the picture.

[0009] In yet another example aspect, a video processing method is disclosed. The method includes: performing conversion between a video region of a video and a bitstream of the video, wherein the bitstream conforms to a format rule, wherein the format rule specifies that a first syntax element is conditionally included in a sequence parameter set (SPS) based on a value of a second syntax element indicating whether information of a sub-picture is included in the SPS, wherein the first syntax element indicates whether information of a sub-picture identifier is included in a parameter set of the bitstream.

[0010] In yet another example aspect, a video processing method is disclosed. The method includes: performing conversion between a picture of a video and a bitstream of the video, wherein the bitstream conforms to a format rule, wherein the format rule specifies that a mapping between an identifier of one or more sub-pictures of the picture and the one or more sub-pictures is not included in a picture header of the picture, wherein the format rule further specifies that the identifier of the one or more sub-pictures is derived based on syntax elements in a picture parameter set (PPS) and a sequence parameter set (SPS) referenced by the picture.

[0011] In yet another example aspect, a video processing method is disclosed. The method includes: performing conversion between a video and a bitstream of the video, wherein the bitstream conforms to a format rule, and wherein the format rule specifies that when a value of a first syntax element indicates that a mapping between an identifier of a sub-picture and one or more sub-pictures of a picture is explicitly signaled for the one or more sub-pictures, the mapping is signaled in a sequence parameter set (SPS) or a picture parameter set (PPS).

[0012] In yet another example aspect, a video processing method is disclosed. The method includes: performing conversion between a video and a bitstream of the video, wherein the bitstream conforms to a format rule, and wherein the format rule specifies that a length of an identifier of a sub-picture signaled in a sequence parameter set (SPS) is not based on a value of a syntax element indicating whether the identifier is signaled in the SPS.

[0013] In yet another example aspect, a video processing method is disclosed. The method includes: performing conversion between a video and a bitstream of the video, wherein the bitstream conforms to a format rule, and wherein the format rule specifies that due to a syntax element indicating that an identifier of a sub-picture is explicitly signaled in a picture parameter set (PPS), a length of the identifier is signaled in the PPS.

[0014] In yet another example aspect, a video processing method is disclosed. The method includes: performing conversion between a video and a bitstream of the video, wherein the bitstream conforms to a format rule, wherein the format rule specifies that a first syntax element signaled in a sequence parameter set (SPS) of the bitstream indicates a length of an identifier of a sub-picture in the SPS, and wherein the signaling of the first syntax element is independent of a value of a second syntax element, the value of the second syntax element indicating that the identifier of the sub-picture is explicitly signaled in the SPS or a picture parameter set (PPS).

[0015] In yet another exemplary aspect, a video encoder apparatus is disclosed. The video encoder includes a processor configured to implement the above method.

[0016] In yet another exemplary aspect, a video decoder apparatus is disclosed. The video decoder includes a device configured to implement the above method.

[0017] In yet another exemplary aspect, a computer readable medium is disclosed, on which code is stored. The code is in the form of processor executable code to implement one of the methods described herein.

[0018] These and other features are described in this document. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 An example of partitioning a picture by luma codec tree units (CTUs) is shown.

[0020] Figure 2 Another example of partitioning a picture with luma CTUs is shown.

[0021] Figure 3 An example of picture segmentation is shown.

[0022] Figure 4 Another example of picture segmentation is shown.

[0023] Figure 5 is a block diagram of an example video processing system in which the disclosed techniques may be implemented.

[0024] Figure 6 is a block diagram of an example hardware platform for video processing.

[0025] Figure 7 is a block diagram illustrating a video encoding and decoding system according to some embodiments of the present disclosure.

[0026] Figure 8 is a block diagram illustrating an encoder according to some embodiments of the present disclosure.

[0027] Fig. 9 is a block diagram illustrating a decoder according to some embodiments of the present disclosure.

[0028] Figure 10-12 A flow chart illustrating an example method of video processing is shown. DETAILED DESCRIPTION

[0029] In this article, the section titles are used for ease of understanding and do not limit the applicability of the techniques and embodiments disclosed in each section to only that section. In addition, the use of H.266 terminology in some descriptions is only for ease of understanding and is not intended to limit the scope of the disclosed technology. Therefore, the technology described here is also applicable to other video codec protocols and designs.

[0030] 1. Overview

[0031] This article relates to video coding techniques. Specifically, the signaling of sub-pictures, slices and strips. These concepts can be applied alone or in various combinations to any video coding standard or non-standard video codec that supports multi-layer video coding, such as the Versatile Video Codec (VVC) under development.

[0032] 2. Abbreviations

[0033] APS Adaptive Parameter Set

[0034] AU Access Unit

[0035] AUD Access Unit Delimiter

[0036] AVC Advanced Video Codec

[0037] CLVS Codec Layer Video Sequence

[0038] CPB codec picture buffer

[0039] CRA Clear Random Access

[0040] CTU Codec Tree Unit

[0041] CVS codec video sequence

[0042] DPB decoded picture buffer

[0043] DPS Decoding Parameter Set

[0044] EOB End of bitstream

[0045] EOS sequence ends

[0046] GDR Progressive Decode Refresh

[0047] HEVC High Efficiency Video Codec

[0048] HRD Virtual Reference Decoder

[0049] IDR instant decoding refresh

[0050] JEM Joint Exploration Model

[0051] MCTS Motion Constraints Episode Set

[0052] NAL Network Abstraction Layer

[0053] OLS output layer set

[0054] PH Image Header

[0055] PPS Picture Parameter Set

[0056] PTL Profiles, Tiers and Levels

[0057] PU picture unit

[0058] RBSP Raw Byte Sequence Payload

[0059] SEI Supplemental Enhancement Information

[0060] SPS Sequence Parameter Set

[0061] SVC Scalable Video Codec

[0062] VCL video codec layer

[0063] VPS Video Parameter Set

[0064] VTM VVC test model

[0065] VUI Video Availability Information

[0066] VVC Multi-functional Video Codec

[0067] 3. Preliminary Discussion

[0068] Video codec standards have evolved primarily through the development of the known ITU-T and ISO / IEC standards. ITU-T developed H.261 and H.263, ISO / IEC developed MPEG-1 and MPEG-4 video, and the two organizations jointly developed H.262 / MPEG-2 video and H.264 / MPEG-4 Advanced Video Codec (AVC) and H.265 / HEVC standards. Since H.262, video codec standards have been based on a hybrid video codec structure, which uses temporal prediction plus transform codec. In order to explore future video codec technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Team (JVET) in 2015. Since then, JVET has adopted many new methods and applied them to reference software called the Joint Exploration Model (JEM). JVET meetings are held simultaneously every quarter, and the goal of the new codec standard is to reduce the bit rate by 50% compared to HEVC. The new video codec standard was officially named Versatile Video Coding (VVC) at the JVET meeting in April 2018, when the first version of the VVC Test Model (VTM) was released. With the continuous efforts of VVC standardization, new codec technologies are adopted into the VVC standard at each JVET meeting. The VVC working draft and test model VTM are updated after each meeting. The VVC project now aims to achieve technical completion (FDIS) at the July 2020 meeting.

[0069] 3.1. Picture segmentation scheme in HEVC

[0070] HEVC includes four different picture segmentation schemes, namely normal slices, dependent slices, tiles, and wavefront parallel processing (WPP), which can be applied to maximum transmission unit (MTU) size matching, parallel processing, and reduction of end-to-end delay.

[0071] Regular slices are similar to those in H.264 / AVC. Each regular slice is encapsulated in its own NAL unit, and intra-picture prediction (intra-sample prediction, motion information prediction, codec mode prediction) and entropy coding dependencies across slice boundaries are disabled. Therefore, regular slices can be reconstructed independently of other regular slices in the same picture (although there may still be interdependencies due to loop filtering operations).

[0072] Regular slices are the only tool available for parallelization, which is also available in H.264 / AVC in almost the same form. Regular slice-based parallelization does not require much inter-processor or inter-core communication (except for inter-processor or inter-core data sharing for motion compensation when decoding predictively coded pictures, which is usually much heavier than inter-processor or inter-core data sharing due to intra-picture prediction). However, for the same reason, the use of regular slices may incur a large codec overhead due to the bit cost of the slice header and the lack of prediction across slice boundaries. In addition, regular slices (in contrast to the other tools mentioned below) can also serve as a key mechanism for bitstream segmentation to match MTU size requirements due to their intra-picture independence and the fact that each regular slice is encapsulated in its own NAL. In many cases, the goals of parallelization and MTU size matching have conflicting requirements on the layout of slices in a picture. The realization of such situations led to the development of the parallelization tools mentioned below.

[0073] Dependent slices have short slice headers and allow the bitstream to be partitioned at treeblock boundaries without breaking any intra-picture prediction. Basically, dependent slices split a regular slice into multiple NAL units, reducing end-to-end latency by allowing part of a regular slice to be sent before encoding of the entire regular slice is complete.

[0074] In WPP, a picture is partitioned into a single row of codec blocks (CTBs). Entropy decoding and prediction are allowed to use data from CTBs in other partitions. Parallel processing is possible through parallel decoding of CTB rows, where the start of decoding of a CTB row is delayed by two CTBs to ensure that data associated with CTBs above and to the right of the main CTB is available before the main CTB being decoded. Using this staggered start (which looks like a wavefront when represented graphically), as many processors / cores as pictures contain CTB rows can be parallelized. Because intra-picture prediction between adjacent treeblock rows within a picture is allowed, the inter-processor / inter-core communication required to implement intra-picture prediction can be substantial. WPP partitioning does not result in the generation of additional NAL units compared to when WPP partitioning is not applied, so WPP is not a tool for MTU size matching. However, if MTU size matching is required, regular slices can be used with WPP, but with some codec overhead.

[0075] Slices define the horizontal and vertical boundaries that divide the image into slice columns and slice rows. Slice columns extend from the top of the image to the bottom of the image. Similarly, slice rows extend from the left side of the image to the right side of the image. The number of slices in a picture can be simply calculated by multiplying the number of slice columns by the number of slice rows.

[0076] Before decoding the top left CTB of the next slice in the order of the slice raster scan of one picture, the scan order of the CTBs is changed to the local scan order within the slice (in the order of the CTB raster scan of the slice). Similar to regular slices, slices break intra-picture prediction dependencies and entropy decoding dependencies. However, they do not need to be contained in independent NAL units (same as WPP in this respect); therefore slices cannot be used for MTU size matching. Each slice can be processed by one processor / core, and in the case where a slice spans multiple slices, the inter-processor / inter-core communication required for intra-picture prediction between processing units decoding adjacent slices is limited to the transmission of shared slice headers and loop filtering related to reconstruction samples and metadata sharing. When more than one slice or WPP segment is contained in a slice, the entry point byte offset of each slice or WPP segment in the slice except the first is signaled in the slice header.

[0077] For simplicity, restrictions on the application of four different picture partitioning schemes are specified in HEVC. For most profiles specified in HEVC, a given codec video sequence cannot contain both slices and wavefronts. For each slice and slice, one or both of the following conditions must be met: 1) all codec blocks in a slice belong to the same slice; 2) all codec blocks in a slice belong to the same slice. Finally, a wavefront segment contains exactly one CTB row, and when using WPP, if a slice starts within a CTB row, the slice must end in the same CTB row.

[0078] The most recent revision of HEVC is specified in the JCT-VC output document JCTVC-AC1005, J. Boyce, A. Ramasubramonian, R. Skupin, G. J. Sullivan, A. Tourapis, Y.-K. Wang (eds.), "HEVC Additional Supplemental Enhancement Information (Draft 4)", published on October 24, 2017 at http: / / phenix.intevry.fr / jct / doc_end_user / documents- / 29_Macau / wg11 / JCTVC-AC1005-v2.zip. Contained within this revision, HEVC specifies three MCT-related SEI messages, namely the time-domain MCTS SEI message, the MCTS Extraction Information Set SEI message, and the MCTS Extraction Information Nesting SEI message.

[0079] The temporal MCTS SEI message indicates the presence of MCTS in the bitstream and signals the MCTS. For each MCTS, the motion vector is restricted to point to the full sample position within the MCTS and the fractional sample position that only needs the full sample position within the MCTS for interpolation, and the motion vector candidate predicted by the temporal motion vector of the block outside the MCTS is not allowed to be used. In this way, each MCTS can be decoded independently in the absence of slices not included in the MCTS.

[0080] The MCTS extraction information set SEI message provides supplementary information that can be used for MCTS sub-bitstream extraction (specified as part of the semantics of the SEI message) to generate a bitstream that conforms to the MCTS set. The information consists of multiple extraction information sets, each of which defines multiple MCTS sets and contains RBSP bytes that replace the VPS, SPS, and PPS to be used in the MCTS sub-bitstream extraction process. When extracting a sub-bitstream according to the MCTS sub-bitstream extraction process, the parameter sets (VPS, SPS, and PPS) need to be rewritten or replaced, and the slice header needs to be slightly updated because one or all of the slice address related syntax elements (including first_slice_segment_in_pic_flag and slice_segment_address) usually need to have different values.

[0081] 3.2. Image Segmentation in VVC

[0082] In VVC, a picture is divided into one or more slice rows and one or more slice columns. A slice is a sequence of CTUs covering a rectangular area of ​​a picture. The CTUs in a slice are scanned in raster scan order within the slice.

[0083] A slice consists of an integer number of complete slices or an integer number of consecutive complete CTU rows within a picture.

[0084] Two stripe modes are supported, namely raster scan stripe mode and rectangular stripe mode. In raster scan stripe mode, a stripe contains a complete sequence of stripes in a stripe raster scan of a picture. In rectangular stripe mode, a stripe contains multiple complete slices that together form a rectangular area of ​​a picture or multiple consecutive complete CTU rows that together form a slice of a rectangular area of ​​a picture. Stripes within a rectangular stripe are scanned in a stripe raster scan order within the rectangular area corresponding to the stripe.

[0085] A sub-picture consists of one or more strips that together cover a rectangular area of ​​the picture.

[0086] Figure 1 An example of raster scan stripe partitioning of a picture is shown, where the picture is partitioned into 12 slices and 3 raster scan stripes.

[0087] Figure 2 An example of rectangular strip partitioning of a picture is shown, where the picture is partitioned into 24 slices (6 slice columns and 4 slice rows) and 9 rectangular slices.

[0088] Figure 3 An example of a picture partitioned into slices and rectangular strips is shown, where the picture is partitioned into 4 slices (2 slice columns and 2 slice rows) and 4 rectangular strips.

[0089] Figure 4 An example of sub-picture partitioning of a picture is shown, where the picture is partitioned into 18 slices, 12 slices on the left, each covering a stripe with 4x4 CTUs, and 6 slices on the right, each covering 2 vertically stacked stripes with 2x2 CTUs, resulting in a total of 24 stripes and 24 sub-pictures of different dimensions (each stripe is a sub-picture).

[0090] 3.3. Signaling of sub-pictures, slices, and stripes in VVC

[0091] In the latest VVC draft text, sub-picture information is signaled in SPS, including sub-picture layout (i.e., the number of sub-pictures per picture and the position and size of each picture) and other sequence-level sub-picture information. The order of sub-pictures signaled in SPS defines the sub-picture index. The list of sub-picture IDs that each sub-picture has can be explicitly signaled in SPS or PPS, for example.

[0092] Slices in VVC are conceptually the same as in HEVC, ie, each picture is partitioned into slice columns and slice rows, but have a different syntax for signaling slices in the PPS.

[0093] In VVC, the slice mode is also signaled in the PPS. When the slice mode is the rectangular slice mode, the strip layout of each picture (i.e., the number of strips per picture and the position and size of each strip) is signaled in the PPS. The order of the rectangular strips within the picture signaled in the PPS defines the picture level strip index. The sub-picture level strip index is defined as the order of the strips within the sub-picture in ascending order of their picture level strip indexes. The position and size of the rectangular strips are sent / derived based on the sub-picture position and size signaled in the SPS (when each sub-picture contains only one stripe), or based on the slice position and size signaled in the PPS (when the sub-picture may contain multiple stripes). When the strip mode is the raster scan strip mode, similar to HEVC, the strip layout within the picture is signaled in different details in the strips themselves.

[0094] The SPS, PPS and slice headers and semantics in the latest VVC draft text most relevant to the present invention are as follows.

[0095] 7.3.2.3 Sequence parameter set RBSP syntax

[0096]

[0097]

[0098] 7.4.3.3 Sequence parameter set RBSP semantics ...

[0099] subpics_present_flag equal to 1 specifies that sub-picture parameters are present in the SPS RBSP syntax. subpics_present_flag equal to 0 specifies that sub-picture parameters are not present in the SPS RBSP syntax.

[0100] NOTE 2 – When the bitstream is the result of a sub-bitstream extraction process and contains only a subset of the sub-pictures of the input bitstream of the sub-bitstream extraction process, it may be necessary to set the value of subpics_present_flag in the RBSP of the SPS equal to 1.

[0101] sps_num_subpics_minus1 plus 1 specifies the number of sub-pictures. sps_num_subpics_minus1 shall be in the range of 0 to 254. When not present, the value of sps_num_subpics_minus1 is inferred to be equal to 0.

[0102] subpic_ctu_top_left_x[i] specifies the horizontal position of the top left CTU of the i-th sub-picture in units of CtbSizeY. The length of the syntax element is Ceil(Log2(pic_width_max_in_luma_samples / CtbSizeY)) bits. When it is not present, the value of subpic_ctu_top_left_x[i] is inferred to be equal to 0.

[0103] subpic_ctu_top_left_y[i] specifies the vertical position of the top left CTU of the i-th sub-picture in units of CtbSizeY. The length of the syntax element is Ceil(Log2(pic_height_max_in_luma_samples / CtbSizeY)) bits. When it is not present, the value of subpic_ctu_top_left_y[i] is inferred to be equal to 0.

[0104] subpic_width_minus1[i] plus 1 specifies the width of the i-th sub-picture in units of CtbSizeY. The length of the syntax element is Ceil(Log2(pic_width_max_in_luma_samples / CtbSizeY)) bits. When it is not present, the value of subpic_width_minus1[i] is inferred to be equal to Ceil(pic_width_max_in_luma_samples / CtbSizeY)-1.

[0105] subpic_height_minus1[i] plus 1 specifies the height of the i-th sub-picture in units of CtbSizeY. The length of the syntax element is Ceil(Log2(pic_height_max_in_luma_samples / CtbSizeY)) bits. When it is not present, the value of subpic_height_minus1[i] is inferred to be equal to Ceil(pic_height_max_in_luma_samples / CtbSizeY)-1.

[0106] subpic_treated_as_pic_flag[i] equal to 1 indicates that the i-th subpicture of each codec picture in the CLVS is treated as a picture in the decoding process except for loop filtering operations. subpic_treatment_as_pic_flag[i] equal to 0 indicates that the i-th subpicture of each codec picture in the CLVS is not treated as a picture in the decoding process except for loop filtering operations. When it is not present, the value of subpic_treatment_as_pic_flag[i] is inferred to be equal to 0.

[0107] loop_filter_across_subpic_enabled_flag[i] equal to 1 specifies that loop filtering operations may be performed across the boundary of the i-th sub-picture in each coded picture in the CLVS. loop_filter_across_subpic_enabled_flag[i] equal to 0 specifies that loop filtering operations are not performed across the boundary of the i-th sub-picture in each coded picture in the CLVS. When it is not present, the value of loop_filter_across_subpic_enabled_pic_flag[i] is inferred to be equal to 1.

[0108] The following constraints apply to the bitstream conformance requirements:

[0109] For any two sub-pictures subpicA and subpicB, when the sub-picture index of subpicA is less than the sub-picture index of subpicB, any codec slice NAL unit of subPicA shall precede any codec slice NAL unit of subPicB in decoding order.

[0110] The shape of the sub-pictures should be such that when each sub-picture is decoded, its entire left border and its entire top border consist of picture boundaries or of boundaries of previously decoded sub-pictures.

[0111] sps_subpic_id_present_flag equal to 1 indicates that sub-picture ID mapping is present in the SPS. sps_subpic_id_present_flag equal to 0 specifies that sub-picture ID mapping is not present in the SPS.

[0112] sps_subpic_id_signalling_present_flag equal to 1 specifies that the sub-picture ID mapping is signaled in the SPS. sps_subpic_id_signalling_present_flag equal to 0 specifies that the sub-picture ID mapping is not signaled in the SPS. When not present, the value of sps_subpic_id_signalling_present_flag is inferred to be equal to 0.

[0113] sps_subpic_id_len_minus1 plus 1 specifies the number of bits used to represent the syntax element sps_subpic_id[i]. The value of sps_subpic_id_len_minus1 shall be in the range of 0 to 15, inclusive.

[0114] sps_subpic_id[i] specifies the sub-picture ID of the i-th sub-picture. The length of the sps_subpic_id[i] syntax element is sps_subpic_id_len_minus1+1 bits. When not present, and when sps_subpic_id_present_flag is equal to 0, the value of sps_subpic_id[i] is inferred to be equal to i, for each i in the range from 0 to sps_num_subpics_minus1, inclusive. ...

[0115] 7.3.2.4 Picture parameter set RBSP syntax

[0116]

[0117]

[0118] 7.4.3.4 Picture parameter set RBSP semantics ...

[0119] pps_subpic_id_signalling_present_flag equal to 1 specifies that the sub-picture ID mapping is signaled in the PPS. pps_subpic_id_signalling_present_flag equal to 0 specifies that the sub-picture ID mapping is not signaled in the PPS. pps_subpic_id_signalling_present_flag shall be equal to 0 when sps_subpic_id_present_flag is 0 or sps_subpic_id_signalling_present_flag is equal to 1.

[0120] pps_num_subpics_minus1 plus 1 specifies the number of subpictures in the coded picture that references the PPS.

[0121] The value of pps_num_subpic_minus1 shall be equal to sps_num_subpics_minus1 as a bitstream conformance requirement.

[0122] pps_subpic_id_len_minus1 plus 1 specifies the number of bits used to represent the syntax element pps_subpic_id[i]. The value of pps_subpic_id_len_minus1 shall be in the range of 0 to 15, inclusive.

[0123] The value of pps_subpic_id_len_minus1 for codec pictures referenced in CLVS shall be the same for all PPSs, as a requirement for bitstream conformance.

[0124] pps_subpic_id[i] specifies the sub-picture ID of the i-th sub-picture. The length of the pps_subpic_id[i] syntax element is pps_subpic_id_len_minus1+1 bits.

[0125] no_pic_partition_flag equal to 1 specifies that picture partitioning is not applied to each picture referencing a PPS. no_pic_partition_flag equal to 0 specifies that each picture referencing a PPS may be partitioned into multiple slices or slices.

[0126] The value of no_pic_partition_flag shall be the same for all PPSs referenced by a codec picture within a CLVS, as this is a requirement for bitstream conformance.

[0127] When the value of sps_num_subpics_minus1+1 is greater than 1, the value of no_pic_partition_flag cannot be equal to 1, which is a requirement for bitstream consistency.

[0128] pps_log2_ctu_size_minus5 plus 5 specifies the luma codec treeblock size for each CTU. pps_log2_ctu_size_minus5 shall be equal to sps_log2_ctu_size_minus5.

[0129] num_exp_tile_columns_minus1 plus 1 specifies the number of explicitly provided tile column widths. The value of num_exp_tile_columns_minus1 shall be in the range of 0 to PicWidthInCtbsY-1, inclusive. When no_pic_partition_flag is equal to 1, the value of num_exp_tile_columns_minus1 is inferred to be equal to 0.

[0130] num_exp_tile_rows_minus1 plus 1 specifies the number of explicitly provided tile row heights. The value of num_exp_tile_rows_minus1 shall be in the range of 0 to PicHeightInCtbsY-1, inclusive. When no_pic_partition_flag is equal to 1, the value of num_exp_tile_columns_minus1 is inferred to be equal to 0.

[0131] tile_column_width_minus1[i] plus 1 specifies the width of the i-th tile column in CTBs, i in the range 0 to num_exp_tile_columns_minus1-1, inclusive. tile_column_width_minus1[num_exp_tile_columns_minus1] is used to derive the width of tile columns with indices greater than or equal to num_exp_tile_columns_minus1, as described in clause 6.5.1. When it is not present, tile_column_width_minus1[0] is inferred to be equal to PicWidthInCtbsY-1.

[0132] tile_row_height_minus1[i] plus 1 specifies the height of the i-th tile row in CTBs, i in the range of 0 to num_exp_tile_rows_minus1-1, inclusive. tile_row_height_minus1[num_exp_tile_rows_minus1] is used to derive the height of tile rows with indices greater than or equal to num_exp_tile_rows_minus1, as described in clause 6.5.1. When it is not present, the value of tile_row_height_minus1[0] is inferred to be equal to PicHeightInCtbsY-1.

[0133] rect_slice_flag equal to 0 specifies that the slices within each slice are arranged in raster scan order and the slice information is not signaled in the PPS. rect_slice_flag equal to 1 specifies that the slices within each slice cover a rectangular area of ​​the picture and the slice information is signaled in the PPS. When it is not present, rect_slice_flag is inferred to be equal to 1. When subpics_present_flag is equal to 1, the value of rect_slice_flag shall be equal to 1.

[0134] single_slice_per_subpic_flag equal to 1 specifies that each sub-picture consists of one and only one rectangular slice. single_slice_per_subpic_flag equal to 0 specifies that each sub-picture may contain one or more rectangular slices. When subpics_present_flag is equal to 0, single_slice_per_subpic_flag shall be equal to 0. When single_slice_per_subpic_flag is equal to 1, num_slices_in_pic_minus1 is inferred to be equal to sps_num_subpics_minus1.

[0135] num_slices_in_pic_minus1 plus 1 specifies the number of rectangular slices in each picture that reference the PPS. The value of num_slices_in_pic_minus1 shall be in the range of 0 to MaxSlicesPerPicture-1, inclusive, where MaxSlicesPerPicture is specified in Annex A. When no_pic_partition_flag is equal to 1, the value of num_slices_in_pic_minus1 is inferred to be equal to 0.

[0136] tile_idx_delta_present_flag equal to 0 specifies that tile_idx_delta values ​​are not present in the PPS and that all rectangular slices in pictures referencing the PPS are specified in raster order according to the procedure defined in clause 6.5.1. tile_idx_delta_present_flag equal to 1 specifies that tile_idx_delta values ​​may be present in the PPS and that all rectangular slices in pictures referencing the PPS are specified in the order indicated by the tile_idx_delta values.

[0137] slice_width_in_tiles_minus1[i] plus 1 specifies the width of the i-th rectangular strip in units of tile columns. The value of slice_width_in_tiles_minus1[i] shall be in the range of 0 to NumTileColumns-1, inclusive. When it is not present, the value of slice_width_in_tiles_minus1[i] shall be inferred as specified in Section 6.5.1.

[0138] slice_height_in_tiles_minus1[i] plus 1 specifies the height of the i-th rectangular strip in tile rows. The value of slice_height_in_tiles_minus1[i] shall be in the range of 0 to NumTileRows-1, inclusive. When it is not present, the value of slice_height_in_tiles_minus1[i] shall be inferred as specified in Section 6.5.1.

[0139] num_slices_in_tile_minus1[i] plus 1 specifies the number of slices in the current slice, applicable when the i-th slice contains a subset of CTU rows from a single slice. The value of num_slices_in_tile_minus1[i] shall be in the range of 0 to RowHeight[tileY]-1, inclusive, where tileY is the index of the slice row containing the i-th slice. When it is not present, the value of num_slices_in_tile_minus1[i] is inferred to be equal to 0.

[0140] slice_height_in_ctu_minus1[i] plus 1 specifies the height of the i-th rectangular slice in units of CTU rows, applicable to the case where the i-th slice contains a subset of CTU rows from a single slice. The value of slice_height_in_ctu_minus1[i] shall be in the range of 0 to RowHeight[tileY]-1, inclusive, where tileY is the index of the slice row containing the i-th slice.

[0141] tile_idx_delta[i] specifies the tile index difference between the i-th rectangular strip and the (i+1)-th rectangular strip. The value of tile_idx_delta[i] shall be in the range of –NumTilesInPic+1 to NumTilesInPic-1, inclusive. When it is not present, the value of tile_idx_delta[i] is inferred to be equal to 0. In all other cases, the value of tile_idx_delta[i] shall not be equal to 0.

[0142] loop_filter_across_tiles_enabled_flag equal to 1 specifies that loop filtering operations may be performed across slice boundaries in pictures that reference the PPS. loop_filter_across_tiles_enabled_flag equal to 0 specifies that loop filtering operations are not performed across slice boundaries in pictures that reference the PPS. Loop filtering operations include deblocking filter, sample adaptive offset filter, and adaptive loop filter operations. When it is not present, the value of loop_filter_across_tiles_enabled_flag is inferred to be equal to 1.

[0143] loop_filter_across_slices_enabled_flag equal to 1 specifies that loop filtering operations may be performed across slice boundaries in pictures that reference the PPS. loop_filter_across_slice_enabled_flag equal to 0 specifies that loop filtering operations are not performed across slice boundaries in pictures that reference the PPS. Loop filtering operations include deblocking filter, sample adaptive offset filter, and adaptive loop filter operations. When it is not present, the value of loop_filter_across_slices_enabled_flag is inferred to be equal to 0.

[0144] 7.3.7.1 Generic Strip Header Syntax

[0145]

[0146]

[0147] 7.4.8.1 Generic Strip Header Semantics ...

[0148] slice_subpic_id specifies the sub-picture identifier of the sub-picture containing the slice. If slice_subpic_id exists, the value of the variable SubPicIdx is derived so that SubpicIdList[SubPicIdx] is equal to slice_subpic_id. Otherwise (slice_subpic_id does not exist), the variable SubPicIdx is derived to be equal to 0. The length of slice_subpic_id, in bits, is derived as follows:

[0149] - If sps_subpic_id_signalling_present_flag is equal to 1, the length of slice_subpic_id is equal to sps_subpic_id_len_minus1+1.

[0150] Otherwise, if ph_subpic_id_signalling_present_flag is equal to 1, the length of slice_subpic_id is equal to ph_subpic_id_len_minus1+1.

[0151] Otherwise, if pps_subpic_id_signalling_present_flag is equal to 1, the length of slice_subpic_id is equal to pps_subpic_id_len_minus1+1.

[0152] Otherwise, the length of slice_subpic_id is equal to Ceil(Log2(sps_num_subpics_minus1+1)).

[0153] slice_address specifies the slice address of the slice. When not present, the value of slice_address is inferred to be equal to 0.

[0154] If rect_slice_flag is equal to 0, the following applies:

[0155] —The strip address is the raster scan slice index.

[0156] —The length of slice_address is Ceil(Log2(NumTilesInPic)) bits.

[0157] —slice_address values ​​should be in the range of 0 to NumTilesInPic-1, inclusive.

[0158] Otherwise (rect_slice_flag is equal to 1), the following applies:

[0159] —The slice address is the slice index of the slice in the SubPicIdx-th sub-picture.

[0160] —The length of slice_address is Ceil(Log2(NumSlicesInSubpic[SubPicIdx])) bits.

[0161] —The value of slice_address shall be in the range of 0 to NumSlicesInSubpic[SubPicIdx]-1, inclusive.

[0162] The following constraints apply to the bitstream conformance requirements:

[0163] - If rect_slice_flag is equal to 0 or subpics_present_flag is equal to 0, the value of slice_address shall not be equal to the value of slice_address of any other codec slice NAL unit of the same codec picture.

[0164] — Otherwise, the pair of slice_subpic_id and slice_address values ​​shall not be equal to the pair of slice_subpic_id and slice_address values ​​of any other codec slice NAL unit of the same codec picture.

[0165] — When rect_slice_flag is equal to 0, the slices of the picture shall be arranged in ascending order of their slice_address values.

[0166] - The shape of a picture slice shall be such that the entire left and upper boundaries of each CTU, when decoded, shall consist of either a picture boundary or a boundary of a previously decoded CTU.

[0167] num_tiles_in_slice_minus1 plus 1, if present, specifies the number of tiles in a slice. The value of num_tiles_in_slice_minus1 should be in the range of 0 to NumTilesInPic-1, inclusive.

[0168] The variable NumCtuInCurrSlice specifies the number of CTUs in the current slice, and the list CtbAddrInCurrSlice[i] specifies the picture raster scan address of the i-th CTB in the slice, where i ranges from 0 to NumCtuInCurrSlice-1, including the end value, as follows:

[0169]

[0170]

[0171] The variables SubPicLeftBoundaryPos, SubPicTopBoundaryPos, SubPicRightBoundaryPos, and SubPicBotBoundaryPos are derived as follows:

[0172]

[0173] 4. Examples of technical problems solved by this solution

[0174] The existing design of signaling for sub-pictures, slices, and strips in VVC has the following problems:

[0175] 1) The codec for sps_num_subpics_minus1 is u(8), which means that each picture is not allowed to have more than 256 sub-pictures. However, in some applications, the maximum number of sub-pictures per picture may need to be greater than 256.

[0176] 2) Allow subpics_present_flag to be equal to 0 and sps_subpic_id_present_flag to be equal to 1. However, this does not make sense, because subpics_present_flag equal to 0 means that CLVS has no information about sub-pictures at all.

[0177] 3) For each sub-picture, a list of sub-picture IDs can be signaled in the picture header (PH). However, when the list of sub-picture IDs is signaled in the PH, and when a subset of sub-pictures is extracted from the bitstream, all PHs will need to be changed. This is undesirable.

[0178] 4) Currently, when a sub-picture ID is indicated as explicitly signaled, the sub-picture ID may not be signaled anywhere, via sps_subpic_id_present_flag (or the name of the syntax element is changed to subpic_ids_explicitly_signalled_flag) equal to 1. This is problematic because when a sub-picture ID is indicated as explicitly signaled, the sub-picture ID needs to be explicitly signaled in the SPS or PPS.

[0179] 5) When the sub-picture ID is not explicitly signaled, the slice header syntax element slice_subpic_id still needs to be signaled whenever subpics_present_flag is equal to 1, including when sps_num_subpics_minus1 is equal to 0. However, the length of slice_subpic_id is currently specified as Ceil(Log2(sps_num_subpics_minus1+1)) bits, which is 0 bits when sps_num_subpics_minus1 is equal to 0. This is problematic because any existing syntax element cannot be 0 bits long.

[0180] 6) The sub-picture layout, including the number, size and position of sub-pictures, remains unchanged throughout the CLVS. Even if the sub-picture ID is not explicitly signaled in the SPS or PPS, the sub-picture ID length still needs to be signaled for the sub-picture ID syntax element in the slice header.

[0181] 7) Whenever rect_slice_flag is equal to 1, the syntax element slice_address is signaled in the slice header and specifies the slice index within the sub-picture that contains the slice, including when the number of slices within the sub-picture (i.e., NumSlicesInSubpic[SubPicIdx]) is equal to 1. However, currently, when rect_slice_flag is equal to 1, the length of slice_address is specified as Ceil(Log2(NumSlicesInSubpic[SubPicIdx])) bits, and when NumSlicesInSubpic[SubPicIdx] is equal to 1, the length of slice_address is 0 bits. This is problematic because the length of any existing syntax element cannot be 0 bits.

[0182] 8) There is redundancy between the syntax elements no_pic_partition_flag and pps_num_subpics_minus1, although the latest VVC text has the following constraint: when sps_num_subpics_minus1 is greater than 0, the value of no_pic_partition_flag shall be equal to 1.

[0183] 5. Example embodiments and solutions

[0184] To solve the above problems and other problems, the methods summarized below are disclosed. The present invention should be regarded as an example to explain the general concept and should not be interpreted narrowly. In addition, these inventions can be applied alone or in combination in any way.

[0185] 1) To solve the first problem, change the codec of sps_num_subpics_minus1 from u(8) to ue(v) so that each picture can have more than 256 sub-pictures.

[0186] a. In addition, the value of sps_num_subpics_minus1 is limited to the range of 0 to Ceil(pic_width_max_in_luma_samples÷CtbSizeY)*Ceil(pic_height_max_in_-luma_samples÷CtbSizeY)-1.

[0187] b. In addition, the number of sub-pictures per picture is further restricted in the definition of the level.

[0188] 2) To solve the second problem, the condition for signaling the syntax element sps_subpic_id_present_flag is set to "if (subpics_present_flag)", that is, when subpics_present_flag is equal to 0, the sps_subpic_id_present_flag syntax element is not signaled, and when it does not exist, it is inferred that the value of sps_subpic_id_present_flag is equal to 0.

[0189] a. Alternatively, when subpics_present_flag is equal to 0, the syntax element sps_subpic_id_present_flag is still signaled, but when subpics_present_flag is equal to 0, this value needs to be equal to 0.

[0190] b. Additionally, the names of the syntax elements subpics_present_flag and sps_subpic_id_present_flag are changed to subpic_info_present_flag and subpic_ids_explicitly_signalled_flag respectively.

[0191] 3) To solve the third problem, the signalling of subpicture IDs in the PH syntax is removed. Therefore, for i in the range from 0 to sps_num_subpics_minus1 (including the end values), the list SubpicIdList[i] is derived as follows:

[0192]

[0193] 4) To solve the fourth problem, when subpictures are indicated by explicit signalling, the subpicture IDs are signalled in the SPS or PPS.

[0194] a. This is achieved by adding the following constraint: if subpic_ids_explicitly_signalled_flag is 0 or subpic_ids_in_sps_flag equals 1, then subpic_ids_in_pps_flag should equal 0. Otherwise (subpic_ids_explicitly_signalled_flag is 1 or subpic_ids_in_sps_flag equals 0), subpic_ids_in_pps_flag should equal 1.

[0195] 5) To solve the fifth and sixth problems, regardless of the value of the SPS flag sps_subpic_id_present_flag (or renamed to subpic_ids_explicitly_signalled_flag), the length of the subpicture IDs is signalled in the SPS. Although when the subpicture IDs are also explicitly signalled in the PPS, the length can also be signalled in the PPS to avoid parsing the PPS's dependence on the SPS. In this case, the length also specifies the length of the subpicture IDs in the slice header, even if the subpicture IDs are not explicitly signalled in the SPS or PPS. Therefore, when it exists, the length of slice_subpic_id is also specified by the length of the subpicture IDs signalled in the SPS.

[0196] 6) Alternatively, to address the fifth and sixth issues, a flag is added to the SPS syntax with a value of 1 to specify the presence of the sub-picture ID length in the SPS syntax. The presence of this flag is independent of the value of the flag indicating whether the sub-picture ID is explicitly signaled in the SPS or PPS. When subpic_ids_explicitly_signalled_flag is equal to 0, the value of this flag can be equal to 1 or 0, but when subpic_ids_explicitly_signalled_flag is equal to 1, the value of this flag must be equal to 1. When this flag is equal to 0, that is, the sub-picture length does not exist, the length of slice_subpic_id is specified as Max(Ceil(Log2(sps_num_subpics_minus1+1)),1) bits (instead of Ceil(Log2(sps_num_subpics_minus1+1)) bits in the latest VVC draft text).

[0197] a. Alternatively, this flag is present only when subpic_ids_explicitly_signalled_flag is equal to 0, and when subpic_ids_explicitly_signalled_flag is equal to 1, the value of this flag is inferred to be equal to 1.

[0198] 7) To solve the seventh problem, when rect_slice_flag is equal to 1, the length of slice_address is specified to be Max(Ceil(Log2(NumSlicesInSubpic[SubPicIdx])),1) bits.

[0199] a. Alternatively, further, when rect_slice_flag is equal to 0, the length of slice_address is specified as Max(Ceil(Log2(NumTilesInPic)),1) bits instead of Ceil(Log2(NumTilesInPic)) bits.

[0200] 8) To solve the eighth problem, the condition for signaling no_pic_partition_flag is set to "if (subpic_ids_in_pps_flag&&pps_num_subpics-_minus1>0)", and the following inference is added: when it does not exist, the value of no_pic_partition_flag is inferred to be equal to 1.

[0201] a. Alternatively, move the sub-picture ID syntax (all four syntax elements) after the slice and slice syntax in the PPS, for example, immediately before the syntax element entropy_coding_sync_enabled_flag, and then set the condition for signaling pps_num_subpics_minus1 to "if (no_pic_partition_flag)".

[0202] 6. Examples

[0203] The following are some example embodiments of all invention parts except the 8 items summarized in Section 5 above, which can be applied to the VVC specification. The changed text is based on the latest VVC text in JVET-P2001-v14. The most relevant parts that have been added or modified are The most relevant deleted parts are shown with bold double brackets to highlight, e.g. Indicates that "a" has been deleted. There are a few other changes that are editorial in nature and therefore not highlighted.

[0204] 6.1. First embodiment

[0205] 7.3.2.3 Sequence parameter set RBSP syntax

[0206]

[0207]

[0208] 7.4.3.3 Sequence parameter set RBSP semantics ...

[0209]

[0210] NOTE 2 – When the bitstream is the result of a sub-bitstream extraction process and contains only a subset of the sub-pictures of the input bitstream of the sub-bitstream extraction process, it may be necessary to include The value of is set to 1.

[0211] Add 1 to specify the number of sub-images. When it is not present, the value of sps_num_subpics_minus1 is inferred to be equal to 0.

[0212] subpic_ctu_top_left_x[i] specifies the horizontal position of the top left CTU of the i-th sub-picture in units of CtbSizeY. The length of the syntax element is Ceil(Log2(pic_width_max_in_luma_samples / CtbSizeY)) bits. When it is not present, the value of subpic_ctu_top_left_x[i] is inferred to be equal to 0.

[0213] subpic_ctu_top_left_y[i] specifies the vertical position of the top left CTU of the i-th subpicture in units of CtbSizeY. The length of the syntax element is Ceil(Log2(pic_height_max_in_luma_samples / CtbSizeY)) bits. When it is not present, the value of subpic_ctu_top_left_y[i] is inferred to be equal to 0.

[0214] subpic_width_minus1[i] plus 1 specifies the width of the i-th sub-picture in units of CtbSizeY. The length of the syntax element is Ceil(Log2(pic_width_max_in_luma_samples / CtbSizeY)) bits. When it is not present, the value of subpic_width_minus1[i] is inferred to be equal to Ceil(pic_width_max_in_luma_samples / CtbSizeY)-1.

[0215] subpic_height_minus1[i] plus 1 specifies the height of the i-th sub-picture in units of CtbSizeY. The length of the syntax element is Ceil(Log2(pic_height_max_in_luma_samples / CtbSizeY)) bits. When it is not present, the value of subpic_height_minus1[i] is inferred to be equal to Ceil(pic_height_max_in_luma_samples / CtbSizeY)-1.

[0216] subpic_treated_as_pic_flag[i] equal to 1 specifies that the i-th subpicture of each codec picture in the CLVS is treated as a picture in the decoding process that does not include loop filtering operations. subpic_treatment_as_pic_flag[i] equal to 0 specifies that the i-th subpicture of each codec picture in the CLVS is not treated as a picture in the decoding process that does not include loop filtering operations. When it is not present, the value of subpic_treatment_as_pic_flag[i] is inferred to be equal to 0.

[0217] loop_filter_across_subpic_enabled_flag[i] equal to 1 specifies that loop filtering operations may be performed across the boundary of the i-th sub-picture in each coded picture in the CLVS. loop_filter_across_subpic_enabled_flag[i] equal to 0 indicates that loop filtering operations are not performed across the boundary of the i-th sub-picture in each coded picture in the CLVS. When it is not present, the value of loop_filter_across_subpic_enabled_pic_flag[i] is inferred to be equal to 1.

[0218] The following constraints apply to the bitstream conformance requirements:

[0219] —For any two sub-pictures subpicA and subpicB, when the sub-picture index of subpicA is less than the sub-picture index of subpicB, any codec slice NAL unit of subPicA shall precede any codec slice NAL unit of subPicB in decoding order.

[0220] - The shape of sub-pictures shall be such that, when each sub-picture is decoded, its entire left and entire top borders consist of picture boundaries or of boundaries of previously decoded sub-pictures.

[0221]

[0222]

[0223] sps_subpic_id[i] specifies the sub-picture ID of the i-th sub-picture. The length of the sps_subpic_id[i] syntax element is sps_subpic_id_len_minus1+1 bits. ...

[0224] 7.3.2.4 Picture parameter set RBSP syntax

[0225]

[0226]

[0227] 7.4.3.4 Picture parameter set RBSP semantics...

[0228]

[0229] pps_num_subpics_minus1 shall be equal to sps_num_subpics_minus1.

[0230] pps_subpic_id_len_minus1 shall be equal to sps_subpic_id_len_minus1.

[0231] pps_subpic_id[i] specifies the sub-picture ID of the i-th sub-picture. The length of the pps_subpic_id[i] syntax element is pps_subpic_id_len_minus1+1 bits.

[0232]

[0233] A bitstream conformance requirement is that for any i and j in the range 0 to sps_num_subpics_minus1 (inclusive), when i is less than j, SubpicIdList[i] shall be less than SubpicIdList[j]. ...

[0234] rect_slice_flag equal to 0 specifies that the slices within each slice are arranged in raster scan order and the slice information is not signaled in the PPS. rect_slice_flag equal to 1 specifies that the slices within each slice cover a rectangular area of ​​the picture and the slice information is signaled in the PPS. When it is not present, rect_slice_flag is inferred to be equal to 1. When When equal to 1, the value of rect_slice_flag shall be equal to 1.

[0235] single_slice_per_subpic_flag equal to 1 specifies that each sub-picture consists of one and only one rectangular slice. single_slice_per_subpic_flag equal to 0 specifies that each sub-picture may contain one or more rectangular slices. When equal to 0, single_slice_per_subpic_flag shall be equal to 0. When single_slice_per_subpic_flag is equal to 1, num_slices_in_pic_minus1 is inferred to be equal to sps_num_subpics_minus1. ...

[0236] 7.3.7.1 General Strip Header Syntax

[0237]

[0238] 7.4.8.1 Generic Strip Header Semantics ...

[0239] slice_subpic_id specifies the sub-picture ID of the sub-picture containing the slice.

[0240] When it is not present, the value of slice_subpic_id is inferred to be equal to 0.

[0241] The variable SubPicIdx is derived such that SubpicIdList[SubPicIdx] is equal to the value of slice_subpic_id.

[0242] slice_address specifies the slice address of the slice. If it is not present, the value of slice_address is inferred to be equal to 0.

[0243] If rect_slice_flag is equal to 0, the following applies:

[0244] —The strip address is the raster scan slice index.

[0245] —The length of slice_address is Ceil(Log2(NumTilesInPic)) bits.

[0246] —slice_address values ​​should be in the range of 0 to NumTilesInPic-1, inclusive.

[0247] Otherwise (rect_slice_flag is equal to 1), the following applies:

[0248] —The slice address is the sub-picture level slice index of the slice.

[0249] —The length of slice_address is Bit.

[0250] —The value of slice_address shall be in the range of 0 to NumSlicesInSubpic[SubPicIdx]-1, inclusive.

[0251] The following constraints apply to the bitstream conformance requirements:

[0252] —If rect_slice_flag is equal to 0 or If equal to 0, the value of slice_address shall not be equal to the value of slice_address of any other codec slice NAL unit of the same codec picture.

[0253] — Otherwise, the pair of slice_subpic_id and slice_address values ​​shall not be equal to the pair of slice_subpic_id and slice_address values ​​of any other codec slice NAL unit of the same codec picture.

[0254] — When rect_slice_flag is equal to 0, the slices of the picture shall be arranged in ascending order of their slice_address values.

[0255] - The shape of the picture slice shall be such that the entire left and the entire top boundary of each CTU when decoded shall consist of a picture boundary or of the boundary of the previously decoded CTU(s). ...

[0256] Figure 5 A block diagram of an example video processing system 500 in which various techniques of the present disclosure may be implemented is shown. Various implementations may include some or all of the components of system 500. System 500 may include an input 502 for receiving video content. The video content may be received in a raw or uncompressed format (e.g., 8 or 10 bit multi-component pixel values), or may be received in a compressed or encoded format. Input 502 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, a passive optical network (PON), and wireless interfaces such as Wi-Fi or a cellular interface.

[0257] System 500 may include a codec component 504, which may implement various codecs or encoding methods described in the present disclosure. Codec component 504 may reduce the average bit rate of the video from input 502 to the output of codec component 504 to generate a codec representation of the video. Therefore, codec technology is sometimes referred to as video compression or video transcoding technology. The output of codec component 504 may be stored or transmitted via a communication connection as represented by component 506. The bitstream (or codec) representation of the storage or communication of the video received at input 502 may be used by component 508, which is used to generate pixel values ​​or displayable video sent to display interface 510. The process of generating a user-viewable video from a bitstream representation is sometimes referred to as video decompression. In addition, although some video processing operations are referred to as "codec" operations or tools, it should be understood that the encoding tool or operation is used at the encoder, and the corresponding decoding tool or operation of the inverse encoding result will be performed by the decoder.

[0258] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB) or High Definition Multimedia Interface (HDMI) or Display Port, etc. Examples of storage interfaces include SATA (Serial Advanced Technology Attachment), PCI, IDE interfaces, etc. The techniques described in the present disclosure may be implemented in various electronic devices, such as mobile phones, laptop computers, smart phones, or other devices capable of performing digital data processing and / or video display.

[0259] Figure 6 6 is a block diagram of a video processing device 600. Device 600 may be used to implement one or more methods described in the present disclosure. Device 600 may be located in a smart phone, a tablet computer, a computer, an Internet of Things (IoT) receiver, etc. Device 600 may include one or more processors 602, one or more memories 604, and video processing hardware 606. Processor 602 may be configured to implement one or more methods described in the present disclosure. Memory 604 may be used to store data and code for implementing the methods and techniques described in the present disclosure. Video processing hardware 606 may be used in hardware circuits to implement some of the techniques described in the present disclosure. In some embodiments, hardware 606 may be partially or entirely in processor 602, such as a graphics processor.

[0260] Figure 7 is a block diagram illustrating an example video coding system 100 that may utilize the techniques of this disclosure.

[0261] like Figure 7As shown, the video coding system 100 may include a source device 110 and a target device 120. The source device 110 may generate encoded video data, and the source device 110 may be referred to as a video encoding device. The target device 120 may decode the encoded video data generated by the source device 110, and the target device 120 may be referred to as a video decoding device.

[0262] Source device 110 may include a video source 112 , a video encoder 114 , and an input / output (I / O) interface 116 .

[0263] The video source 112 may include, for example, a source of a video capture device, an interface for receiving video data from a video content provider, and / or a computer graphics system for generating video data, or a combination of these sources. The video data may include one or more pictures. The video encoder 114 encodes the video data from the video source 112 to generate a bit stream. The bit stream may include a bit sequence that forms a codec representation of the video data. The bit stream may include a codec picture and associated data. The codec picture is a codec representation of the picture. The associated data may include a sequence parameter set, a picture parameter set, and other grammatical structures. The I / O interface 116 may include a modulator / demodulator (modem) and / or a transmitter. The encoded video data may be directly sent to the target device 120 via the network 130a via the I / O interface 116. The encoded video data may also be stored on a storage medium / server 130b for access by the target device 120.

[0264] Target device 120 may include an I / O interface 126 , a video decoder 124 , and a display device 122 .

[0265] The I / O interface 126 may include a receiver and / or a modem. The I / O interface 126 may obtain encoded video data from the source device 110 or the storage medium / server 130b. The video decoder 124 may decode the encoded video data. The display device 122 may display the decoded video data to the user. The display device 122 may be integrated with the target device 120, or may be external to the target device 120, and the target device 120 may be configured to interface with an external display device.

[0266] The video encoder 114 and the video decoder 124 may operate according to a video compression standard, such as the High Efficiency Video Coding (HEVC) standard, the Versatile Video Coding (VVM) standard, and other current and / or other standards.

[0267] Figure 8 is a block diagram illustrating an example of a video encoder 200, which may be Figure 7 The video encoder 114 in the system 100 illustrated in FIG.

[0268] Video encoder 200 may be configured to perform any or all of the techniques of this disclosure. Figure 8 In the example shown, the video encoder 200 includes multiple functional components. The techniques described in this disclosure may be shared among the various components of the video encoder 200. In some examples, a processor may be configured to perform any or all of the techniques described in this disclosure.

[0269] The functional components of the video encoder 200 may include a segmentation unit 201, a prediction unit 202 which may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205 and an intra-frame prediction unit 206, a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213 and an entropy coding unit 214.

[0270] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an intra-block copy (IBC) unit. The IBC unit may perform prediction in an IBC mode, where at least one reference picture is a picture where the current video block is located.

[0271] Furthermore, some components such as the motion estimation unit 204 and the motion compensation unit 205 may be highly integrated but are not shown in the figure for the purpose of description. Figure 8 are represented separately in the example.

[0272] The partitioning unit 201 may partition a picture into one or more video blocks. The video encoder 200 and the video decoder 300 may support various video block sizes.

[0273] The mode selection unit 203 may select one of the codec modes (intra or inter), for example, based on the error result, and provide the resulting intra or inter codec block to the residual generation unit 207 to generate residual block data and the reconstruction unit 212 to reconstruct the codec block for use as a reference picture. In some examples, the mode selection unit 203 may select a combination of intra and inter prediction (CIIP) modes, where the prediction is based on an inter prediction signal and an intra prediction signal. In the case of inter prediction, the mode selection unit 203 may also select a resolution of motion vectors for the block (e.g., sub-pixel or integer pixel precision).

[0274] In order to perform inter-frame prediction on the current video block, the motion estimation unit 204 may generate motion information of the current video block by comparing the current video block with one or more reference frames from the buffer 213. The motion compensation unit 205 may determine a predicted video block of the current video block based on the motion information of pictures other than the picture associated with the current video block from the buffer 213 and decoded samples.

[0275] Motion estimation unit 204 and motion compensation unit 205 may perform different operations on the current video block, eg, depending on whether the current video block is in an I slice, a P slice, or a B slice.

[0276] In some examples, the motion estimation unit 204 may perform unidirectional prediction on the current video block, and the motion estimation unit 204 may search the reference picture of list 0 or 1 for the reference video block of the current video block. The motion estimation unit 204 may then generate a reference index indicating the reference picture containing the reference video block in list 0 or list 1 and a motion vector indicating the spatial displacement between the current video block and the reference video block. The motion estimation unit 204 may output the reference index, the prediction direction indicator, and the motion vector as the motion information of the current video block. The motion compensation unit 205 may generate a predicted video block of the current block based on the reference video block indicated by the motion information of the current video block.

[0277] In other examples, the motion estimation unit 204 may perform bidirectional prediction on the current video block, the motion estimation unit 204 may search the reference picture in list 0 for the reference video block of the current video block, and may also search the reference picture in list 1 for the reference video block of the current video block. The motion estimation unit 204 may then generate reference indexes indicating the reference pictures in list 0 and list 1, which include reference video blocks and motion vectors indicating spatial displacements between the reference video blocks and the current video block. The motion estimation unit 204 may output the reference index and motion vector of the current video block as motion information of the current video block. The motion compensation unit 205 may generate a predicted video block of the current video block based on the reference video block indicated by the motion information of the current video block.

[0278] In some examples, motion estimation unit 204 may output a complete set of motion information for use in a decoding process by a decoder.

[0279] In some examples, motion estimation unit 204 may not output a complete set of motion information for the current video. Instead, motion estimation unit 204 may reference motion information of another video block to signal motion information for the current video block. For example, motion estimation unit 204 may determine that the motion information of the current video block is sufficiently similar to the motion information of a neighboring video block.

[0280] In one example, motion estimation unit 204 may indicate in a syntax structure associated with the current video block a value that indicates to video decoder 300 that the current video block has the same motion information as another video block.

[0281] In another example, motion estimation unit 204 may identify another video block and a motion vector difference (MVD) in a syntax structure associated with the current video block. The motion vector difference represents the difference between the motion vector of the current video block and the motion vector of the indicated video block. Video decoder 300 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.

[0282] As described above, the video encoder 200 may predictively signal motion vectors. Two examples of predictive signaling techniques that may be implemented by the video encoder 200 include advanced motion vector prediction (AMVP) and merge mode signaling.

[0283] The intra prediction unit 206 may perform intra prediction on the current video block. When the intra prediction unit 206 performs intra prediction on the current video block, the intra prediction unit 206 may generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block may include a prediction video block and various syntax elements.

[0284] The residual generation unit 207 may generate residual data for the current video block by subtracting (eg, indicated by a minus sign) the predicted video block of the current video block from the current video block. The residual data for the current video block may include residual video blocks corresponding to different sample components of samples in the current video block.

[0285] In other examples, for the current video block, such as in skip mode, there may be no residual data for the current video block, and the residual generation unit 207 may not perform a subtraction operation.

[0286] Transform processing unit 208 may generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video block associated with the current video block.

[0287] After transform processing unit 208 generates a transform coefficient video block associated with the current video block, quantization unit 209 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values ​​associated with the current video block.

[0288] The inverse quantization unit 210 and the inverse transform unit 211 may apply inverse quantization and inverse transform to the transform coefficient video block, respectively, to reconstruct a residual video block from the transform coefficient video block. The reconstruction unit 212 may add the reconstructed residual video block to corresponding samples of one or more predicted video blocks generated by the prediction unit 202 to generate a reconstructed video block associated with the current block stored in the buffer 213.

[0289] After the reconstruction unit 212 reconstructs the video block, a loop filtering operation may be performed to reduce video block artifacts in the video block.

[0290] The entropy coding unit 214 may receive data from other functional components of the video encoder 200. When the entropy coding unit 214 receives the data, the entropy coding unit 214 may perform one or more entropy coding operations to generate entropy-coded data and output a bitstream including the entropy-coded data.

[0291] Fig. 9 is a block diagram illustrating an example of a video decoder 300, which may be Figure 7 The video decoder 114 in the system 100 is shown.

[0292] Video decoder 300 may be configured to perform any or all of the techniques of this disclosure. Figure 8 In the example shown, video decoder 300 includes multiple functional components. The techniques described in this disclosure may be shared among various components of video decoder 300. In some examples, a processor may be configured to perform any or all of the techniques described in this disclosure.

[0293] In such Fig. 9 In the example shown, the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra-frame prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, and a reconstruction unit 306 and a buffer 307. In some examples, the video decoder 300 can perform operations generally related to the video encoder 200 ( Figure 8 ) is the decoding channel that is the opposite of the encoding channel described by .

[0294] The entropy decoding unit 301 may obtain a coded bitstream. The coded bitstream may include entropy coded video data (e.g., coded video data blocks). The entropy decoding unit 301 may decode the entropy coded video data, and based on the entropy decoded video data, the motion compensation unit 302 may determine motion information including motion vectors, motion vector precision, reference picture list index, and other motion information. For example, the motion compensation unit 302 may determine such information by performing AMVP and merge modes.

[0295] The motion compensation unit 302 may generate a motion compensated block, possibly performing interpolation based on an interpolation filter. An identifier for the interpolation filter at sub-pixel precision may be included in the syntax element.

[0296] Motion compensation unit 302 may calculate interpolated values ​​for sub-integer pixels of a reference block using interpolation filters used by video encoder 20 during encoding of the video block. Motion compensation unit 302 may determine the interpolation filters used by video encoder 200 according to received syntax information and use the interpolation filters to produce a prediction block.

[0297] The motion compensation unit 302 may use some syntax information to determine the size of blocks used to encode frames and / or slices of the encoded video sequence, partition information describing how each macroblock of a picture of the encoded video sequence is partitioned, a mode indicating how to encode each partition, one or more reference frames (and a list of reference frames) for each inter-frame coded block, and other information for decoding the encoded video sequence.

[0298] The intra prediction unit 303 may form a prediction block from spatially neighboring blocks using, for example, an intra prediction mode received in the bitstream. The inverse quantization unit 303 inversely quantizes, i.e., dequantizes, the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 301. The inverse transform unit 303 applies an inverse transform.

[0299] The reconstruction unit 306 may add the residual block to the corresponding prediction block generated by the motion compensation unit 202 or the intra prediction unit 303 to form a decoded block. If necessary, a deblocking filter may also be applied to filter the decoded block to remove blocky artifacts. The decoded video block is then stored in a buffer 307, which provides reference blocks for subsequent motion compensation / intra prediction and also produces decoded video for presentation on a display device.

[0300] Figure 10-12 An example method for implementing the above technical solution is shown, for example, Figure 5-9 The embodiment shown.

[0301] Fig.10 A flow chart of an example method 1000 of video processing is shown. The method 1000 includes, at operation 1010, performing conversion between a video and a bitstream of the video, wherein the bitstream conforms to a format rule, and wherein the format rule specifies that a length of an identifier of a sub-picture signaled in a sequence parameter set (SPS) is not based on a value of a syntax element indicating whether the identifier is signaled in the SPS.

[0302] Fig.11A flow chart of an example method 1100 of video processing is shown. The method 1100 includes, at operation 1110, performing conversion between a video and a bitstream of the video, wherein the bitstream conforms to a format rule, and wherein the format rule specifies that the length of the identifier is signaled in a picture parameter set (PPS) due to a syntax element indicating that an identifier of a sub-picture is explicitly signaled in the PPS.

[0303] Fig.12 A flow chart of an example method 1200 of video processing is shown. The method 1200 includes, at operation 1210, performing conversion between a video and a bitstream of the video, wherein the bitstream conforms to a format rule, wherein the format rule specifies that a first syntax element signaled in a sequence parameter set (SPS) of the bitstream indicates a length of an identifier of a sub-picture in the SPS, and wherein the signaling of the first syntax element is independent of a value of a second syntax element, the value of the second syntax element indicating that the identifier of the sub-picture is explicitly signaled in the SPS or a picture parameter set (PPS).

[0304] Next, a list of preferred solutions for some embodiments is provided.

[0305] A1. A video processing method, comprising: performing conversion between a video comprising a picture and a bitstream of the video, wherein the number of sub-pictures in the picture is signaled as a field in a sequence parameter set (SPS) of the bitstream, and the bit width of the field is based on the value of the number of sub-pictures, wherein the field is an unsigned integer 0-order exponential Golomb (Exp-Golomb) codec syntax element with the left bit first.

[0306] A2. A method according to solution A1, wherein the value of the field is restricted to a range from zero to a maximum value based on the maximum width of the picture in units of luma samples and the maximum height of the picture in units of luma samples.

[0307] A3. The method according to solution A2, wherein the maximum value is equal to the integer number of suitable codec tree blocks in the picture.

[0308] A4. The method according to solution A1, wherein the number of sub-pictures is limited based on a codec level associated with the bitstream.

[0309] A5. A video processing method, comprising performing conversion between a video and a bitstream of the video, wherein the bitstream conforms to a format rule, wherein the format rule specifies that a first syntax element indicating whether a picture of the video can be split is conditionally included in a picture parameter set (PPS) of the bitstream based on values ​​of a second syntax element and a third syntax element, wherein the second syntax element indicates whether an identifier of a sub-picture is signaled in the PPS, and the third syntax element indicates the number of sub-pictures in the PPS.

[0310] A6. The method according to solution A5, wherein the first syntax element is no_pic_partition_flag, the second syntax element is subpic_ids_in_pps_flag, and the third syntax element is pps_num_subpics_minus1.

[0311] A7. The method according to solution A5 or A6, wherein the first syntax element is excluded from the PPS and is inferred to indicate that no picture partitioning is applied to each picture of the reference PPS.

[0312] A8. The method according to solution A5 or A6, wherein the first syntax element is not signaled in the PPS and is inferred to be equal to one.

[0313] A9. A method according to solution A5 or A6, wherein the second syntax element is signaled after one or more slice and / or slice syntax elements in the PPS.

[0314] A10. A video processing method, comprising performing conversion between a video and a bitstream of the video, wherein the bitstream conforms to a format rule, wherein the format rule specifies that a first syntax element indicating whether a picture of the video can be split is included in a picture parameter set (PPS) of the bitstream, located before a set of syntax elements in the PPS indicating identifiers of sub-pictures of the picture.

[0315] A11. The method according to solution A10, wherein a second syntax element indicating the number of sub-pictures is conditionally included in the set of syntax elements based on a value of the first syntax element.

[0316] A12. A method according to any of solutions A1 to A11, wherein converting comprises decoding the video from a bitstream.

[0317] A13. A method according to any of solutions A1 to A11, wherein converting comprises encoding the video into a bitstream.

[0318] A14. A method for storing a bitstream representing a video to a computer-readable recording medium, comprising generating a bitstream from a video according to any one or more of the methods of solutions A1 to A11; and writing the bitstream to a computer-readable recording medium.

[0319] A15. A video processing device, comprising a processor configured to execute any one or more methods according to solutions A1 to A14.

[0320] A16. A computer-readable medium having stored thereon instructions which, when executed, cause a processor to perform the method of any one or more of Solutions A1 to A14.

[0321] A17. A computer-readable medium storing a bitstream generated according to any one or more of schemes A1 to A14.

[0322] A18. A video processing device for storing a bit stream, wherein the video processing device is configured to execute any one or more of the methods of schemes A1 to A14.

[0323] Another list of preferred solutions for some embodiments is provided next.

[0324] B1. A video processing method, comprising performing conversion between a video region of a video and a bitstream of the video, wherein the bitstream conforms to a format rule, the format rule specifying that a first syntax element is conditionally included in a sequence parameter set (SPS) based on a value of a second syntax element indicating whether information of a sub-picture is included in the SPS, wherein the first syntax element indicates whether information of a sub-picture identifier is included in a parameter set of the bitstream.

[0325] B2. A method according to solution B1, wherein the format rule further specifies that the value of the second syntax element is 0, the value of the second syntax element being 0 indicates that information of the sub-picture is omitted in the SPS, and thus each video region associated with the SPS is not divided into multiple sub-pictures, and based on the value of the second syntax element being 0, the first syntax element is omitted in the SPS.

[0326] B3. The method according to solution B1 or B2, wherein the first syntax element is subpic_ids_explicitly_signalled_flag and the second syntax element is subpic_info_present_flag.

[0327] B4. A method according to solution B1 or B3, wherein the format rule further specifies that, in the case where the value of the first syntax element is 0, the first syntax element with a value of 0 is included in the SPS.

[0328] B5. A method according to any one of solutions B1 to B4, wherein the video area is a video picture.

[0329] B6. A video processing method, comprising performing conversion between a picture of a video and a bitstream of the video, wherein the bitstream complies with a format rule, the format rule specifies that a mapping between identifiers of one or more sub-pictures of a picture and the one or more sub-pictures is not included in a picture header of the picture, wherein the format rule further specifies that the identifiers of the one or more sub-pictures are derived based on syntax elements in a picture parameter set (PPS) and a sequence parameter set (SPS) referenced by the picture.

[0330] B7. A method according to solution B6, wherein a flag in the SPS takes a first value to indicate that an identifier of one or more sub-pictures is derived based on a syntax element in the PPS, or takes a second value to indicate that an identifier of one or more sub-pictures is derived based on a syntax element in the SPS.

[0331] B8. A method according to solution B7, wherein the flag corresponds to the subpic_ids_in_pps_flag field, the first value is 1 and the second value is 0.

[0332] B9. A method according to solution B6, wherein the identifier (denoted as SubpicIdList[i]) is obtained as follows:

[0333]

[0334]

[0335] B10. A video processing method, comprising performing conversion between a video and a bitstream of the video, wherein the bitstream conforms to a format rule, wherein the format rule specifies that when a value of a first syntax element indicates that a mapping between an identifier of a sub-picture and one or more sub-pictures of a picture is explicitly signaled for the one or more sub-pictures, the mapping is signaled in a sequence parameter set (SPS) or a picture parameter set (PPS).

[0336] B11. The method according to solution B10, wherein the first syntax element is subpic_ids_explicitly_signalled_flag.

[0337] B12. A method according to solution B11, wherein the identifier is signaled in the SPS based on the value of a second syntax element (denoted subpic_ids_in_sps_flag) and the identifier is signaled in the PPS based on the value of a third syntax element (denoted subpic_ids_in_pps_flag).

[0338] B13. The method according to solution B12, wherein subpic_ids_in_pps_flag is equal to 0 because subpic_ids_explicitly_signalled_flag is 0 or subpic_ids_in_sps_flag is 1.

[0339] B14. Method according to solution B13, wherein subpic_ids_in_pps_flag equal to 0 indicates that the identifier is not signaled in the PPS.

[0340] B15. The method according to solution B12, wherein subpic_ids_in_pps_flag is equal to 1 because subpic_ids_explicitly_signalled_flag is 1 and subpic_ids_in_sps_flag is 0.

[0341] B16. The method according to solution B15, wherein subpic_ids_in_pps_flag equal to 1 indicates that an identifier of each sub-picture of the one or more sub-pictures is explicitly signaled in the PPS.

[0342] B17. A method according to any of solutions B1 to B15, wherein the conversion comprises decoding the video from a bitstream.

[0343] B18. A method according to any of solutions B1 to B15, wherein converting comprises encoding the video into a bitstream.

[0344] B19. A method for storing a bit stream representing a video to a computer-readable recording medium, comprising generating a bit stream from a video according to any one or more of the methods of solutions B1 to B15; and writing the bit stream to a computer-readable recording medium.

[0345] B20. A video processing device comprising a processor configured to execute any one or more methods of solutions B1 to B19.

[0346] B21. A computer-readable medium having instructions stored thereon, wherein the instructions, when executed, cause a processor to perform any one or more of the methods of schemes B1 to B19.

[0347] B22. A computer-readable medium storing a bitstream generated according to any one or more of schemes B1 to B19.

[0348] B23. A video processing device for storing a bitstream, wherein the video processing device is configured to perform any one or more methods of solutions B1 to B19.

[0349] Next, another list of preferred solutions for some embodiments is provided.

[0350] Cl. A video processing method, comprising performing conversion between a video and a bitstream of the video, wherein the bitstream conforms to a format rule, wherein the format rule specifies that the length of an identifier of a sub-picture signaled in a sequence parameter set (SPS) is not based on the value of a syntax element indicating whether the identifier is signaled in the SPS.

[0351] C2. A video processing method, comprising performing conversion between a video and a bitstream of the video, wherein the bitstream conforms to a format rule, wherein the format rule specifies that due to a syntax element indicating an identifier of a sub-picture is explicitly signaled in a picture parameter set (PPS), the length of the identifier is signaled in the PPS.

[0352] C3. The method according to solution C1 or C2, wherein the length also corresponds to the length of the sub-picture identifier in the slice header.

[0353] C4. A method according to any of solutions C1 to C3, wherein the syntax element is subpic_ids_explicitly_signalled_flag.

[0354] C5. A method according to solution C3 or C4, wherein the identifier is not signaled in the SPS and the identifier is not signaled in the PPS.

[0355] C6. A method according to solution C5, wherein the length of the sub-picture identifier corresponds to the sub-picture identifier length signaled in the SPS.

[0356] C7. A video processing method, comprising performing conversion between a video and a bitstream of the video, wherein the bitstream conforms to a format rule, wherein the format rule specifies that a first syntax element signaled in a sequence parameter set (SPS) of the bitstream indicates a length of an identifier of a sub-picture in the SPS, wherein the signaling of the first syntax element is independent of a value of a second syntax element, and the value of the second syntax element indicates an identifier of the sub-picture explicitly signaled in the SPS or a picture parameter set (PPS).

[0357] C8. The method according to solution C7, wherein the second syntax element is subpic_ids_explicitly_signalled_flag.

[0358] C9. A method according to solution C7 or C8, wherein the second syntax element equal to 1 indicates that a set of identifiers is signaled for each sub-picture in the SPS or PPS, wherein each sub-picture corresponds to an identifier, and wherein the second syntax element equal to 0 indicates that the identifier is not explicitly signaled in the SPS or PPS.

[0359] C10. A method according to solution C7 or C8, wherein the value of the first syntax element is 0 or 1 due to the value of the second syntax element being 0.

[0360] C11. A method according to solution C7 or C8, wherein the value of the first syntax element is 1 due to the value of the second syntax element being 1.

[0361] C12. A method according to solution C7 or C8, wherein the value of the second syntax element is 0.

[0362] C13. A method according to any of solutions C1 to C12, wherein the conversion comprises decoding the video from the bitstream.

[0363] C14. A method according to any of solutions C1 to C12, wherein the conversion comprises encoding the video into a bitstream.

[0364] C15. A method for storing a bit stream representing a video in a computer-readable recording medium, comprising generating a bit stream from a video according to any one or more of the methods of schemes C1 to C12; and writing the bit stream into the computer-readable recording medium.

[0365] C16. A video processing device, comprising a processor, wherein the processor is configured to execute any one or more methods of schemes C1 to C15.

[0366] C17. A computer-readable medium having instructions stored thereon, which when executed cause a processor to perform the method of any one or more of Solutions C1 to C15.

[0367] C18. A computer-readable medium storing a bitstream generated according to any one or more of schemes C1 to C15.

[0368] C19. A video processing device for storing a bit stream, wherein the video processing device is configured to execute any one or more methods of schemes C1 to C15.

[0369] Another list of preferred solutions for some embodiments is provided next.

[0370] P1. A video processing method, comprising performing conversion between a picture of a video and a codec representation of the video, wherein the number of sub-pictures in the picture is included as a field in the codec representation, and the bit width of the field depends on the value of the number of sub-pictures.

[0371] P2. The method according to solution P1, wherein the field indicates the number of sub-pictures using the codeword.

[0372] P3. A method according to solution P2, wherein the codewords comprise Golomb codewords.

[0373] P4. A method according to any of solutions P1 to P3, wherein the value of the number of sub-pictures is restricted to be less than or equal to the integer number of codec treeblocks that fit within the picture.

[0374] P5. The method according to any of solution solutions P1 to P4, wherein the field depends on a codec level associated with the codec representation.

[0375] P6. A video processing method, comprising performing a conversion between a video region of a video and a codec representation of the video, wherein the codec representation conforms to a format rule, wherein the format rule specifies that a syntax element indicating a sub-picture identifier is omitted because the video region does not contain any sub-pictures.

[0376] P7. A method according to solution P6, wherein the codec representation comprises a field with a value of 0, the field with a value of 0 indicating that the video region does not include any sub-picture.

[0377] P8. A video processing method, comprising performing conversion between a video region of a video and a codec representation of the video, wherein the codec representation conforms to a format rule, wherein the format rule specifies that an identifier of a sub-picture in the video region is omitted at a video region header level of the codec representation.

[0378] P9. A method according to solution P8, wherein the codec indicates that sub-pictures are identified numerically according to the order in which they are listed in the video region header.

[0379] P10. A video processing method, comprising performing conversion between a video region of a video and a codec representation of the video, wherein the codec representation complies with a format rule, wherein the format rule specifies an identifier of a sub-picture in the video region and / or the length of the identifier of the sub-picture at a sequence parameter set level or a picture parameter set level.

[0380] P11. A method according to solution P10, wherein the length is contained at the picture parameter set level.

[0381] P12. A video processing method, comprising performing conversion between a video region of a video and a codec representation of the video, wherein the codec representation conforms to a format rule, wherein the format rule specifies including a field in the codec representation at a video sequence level to indicate whether a sub-picture identifier length field is included in the codec representation at the video sequence level.

[0382] P13. A method according to solution P12, wherein the format rule specifies setting the field to 1 in case another field in the codec representation indicates that a length identifier of the video region is included in the codec representation.

[0383] P14. A method according to any of the above schemes, wherein the video area includes a sub-picture of the video.

[0384] P15. A method according to any of the above schemes, wherein the conversion comprises parsing and decoding the codec representation to generate a video.

[0385] P16. A method according to any of the above schemes, wherein the conversion comprises encoding the video to generate a codec representation.

[0386] P17. A video decoding device, comprising a processor, the processor being configured to execute any one or more methods of schemes P1 to P16.

[0387] P18. A video encoding device, comprising a processor, the processor being configured to execute any one or more methods of schemes P1 to P16.

[0388] P19. A computer program product having computer code stored thereon, which, when executed, causes a processor to perform the method of any one of schemes P1 to P16.

[0389] The disclosure and other solutions, examples, embodiments, modules and functional operations described in this application document can be implemented in digital electronic circuits, or computer software, firmware or hardware, including the structures disclosed in this specification and their structural equivalents, or a combination of one or more thereof. The contents and other embodiments disclosed in this specification can be implemented as one or more computer program products, that is, one or more modules of computer program instructions encoded on a tangible and non-volatile computer-readable medium for data processing devices to execute or control the operation of the data processing device. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a storage device, a material composition that affects a machine-readable propagation signal, or a combination of one or more thereof. The term "data processing unit" or "data processing device" includes all devices, equipment and machines for processing data, including, for example, a programmable processor, a computer or a multiprocessor or a computer group. In addition to hardware, the device may also include code that creates an execution environment for a computer program, for example, code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more thereof. The propagated signal is an artificially generated signal, for example, a machine-generated electrical, optical or electromagnetic signal, which is generated to encode information for transmission to a suitable receiver device.

[0390] A computer program (also referred to as a program, software, software application, script, or code) may be written in any form of programming language (including compiled or interpreted languages) and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program may be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program, or in multiple coordinated files (e.g., files that store one or more modules, subroutines, or portions of code). A computer program may be deployed for execution on one or more computers that are located at one site or distributed across multiple sites and interconnected by a communications network.

[0391] The processes and logic flows described in this specification may be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows may also be performed by special purpose logic circuits, and the apparatus may also be implemented as special purpose logic circuits, such as FPGAs (field programmable gate arrays) or ASICs (application specific integrated circuits).

[0392] For example, processors suitable for executing computer programs include general and special purpose microprocessors, and any one or more of any type of digital computer. Typically, the processor will receive instructions and data from a read-only memory or a random access memory or both. The basic components of a computer are a processor that executes instructions and one or more storage devices that store instructions and data. Typically, a computer will also include one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, or be operatively coupled to one or more mass storage devices to receive data from them or transfer data to one or more mass storage devices, or both. However, a computer does not necessarily have such a device. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices, such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable hard disks; magneto-optical disks; and CDROM and DVDROM disks. The processor and memory may be supplemented by, or incorporated into, dedicated logic circuits.

[0393] Although this patent document contains many details, they should not be construed as limitations on the scope of any invention or the claims, but rather as descriptions of features of particular embodiments of particular inventions. Certain features described in this patent document in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various functions described in the context of a single embodiment may also be implemented separately in multiple embodiments, or in any suitable subcombination. In addition, although the above-mentioned features may be described as working in certain combinations, or even initially claimed to be so, in some cases, one or more features in the claim combination may be removed from the combination, and the claim combination may be directed to subcombinations or variations of subcombinations.

[0394] Likewise, although operations are described in a particular order in the drawings, this should not be understood as requiring that such operations be performed in the particular order or order shown, or that all illustrated operations be performed, in order to achieve the desired results. In addition, the separation of various system components in the embodiments described in this patent document should not be understood as requiring such separation in all embodiments.

[0395] Only a few implementations and examples are described, and other implementations, enhancements, and variations can be made based on what is described and illustrated in this patent document.

Claims

1. A video processing method, include: performing conversion between a video and a bitstream of said video, wherein the bit stream complies with the format rules, wherein the format rule specifies that the length of an identifier of a sub-picture signaled in a sequence parameter set (SPS) of the bitstream is based on the value of a first syntax element indicating the presence of sub-picture information in the SPS and does not depend on the value of a second syntax element indicating whether the identifier is explicitly signaled in the SPS, Wherein, the length corresponds to the number of bits used to represent a third syntax element indicating a sub-picture identifier in the SPS; if present, a fourth syntax element indicating a sub-picture identifier in a picture parameter set (PPS); and if present, a fifth syntax element indicating a sub-picture identifier in a slice header.

2. The method according to claim 1, in, The length is signaled by the sixth syntax element sps_subpic_id_len_minus1, the value of which plus 1 equals the length.

3. The method according to claim 2, in, The value of sps_subpic_id_len_minus1 ranges from 0 to 15.

4. The method according to claim 1, in, The second syntax element being equal to 1 specifies that a set of identifiers is explicitly signaled for each sub-picture in the SPS or the picture parameter set (PPS), one identifier for each sub-picture, and the second syntax element being equal to 0 specifies that identifiers are not explicitly signaled in the SPS or the PPS.

5. The method according to claim 1, in, The converting includes decoding the video from the bitstream.

6. The method according to claim 1, in, The converting includes encoding the video into the bitstream.

7. An apparatus for processing video data, the apparatus comprising a processor and a non-transitory memory having instructions thereon, in, When the instructions are executed by the processor, the processor: performing conversion between a video and a bitstream of said video, wherein the bit stream complies with the format rules, wherein the format rule specifies that the length of an identifier of a sub-picture signaled in a sequence parameter set (SPS) of the bitstream is based on the value of a first syntax element indicating the presence of sub-picture information in the SPS and does not depend on the value of a second syntax element indicating whether the identifier is explicitly signaled in the SPS, Wherein, the length corresponds to the number of bits used to represent a third syntax element indicating a sub-picture identifier in the SPS; if present, a fourth syntax element indicating a sub-picture identifier in a picture parameter set (PPS); and if present, a fifth syntax element indicating a sub-picture identifier in a slice header.

8. The device according to claim 7, in, The length is signaled by the sixth syntax element sps_subpic_id_len_minus1, the value of which plus 1 equals the length.

9. The device according to claim 8, in, The value of sps_subpic_id_len_minus1 ranges from 0 to 15.

10. The device according to claim 7, in, The second syntax element being equal to 1 specifies that a set of identifiers is explicitly signaled for each sub-picture in the SPS or the picture parameter set (PPS), one identifier for each sub-picture, and the second syntax element being equal to 0 specifies that identifiers are not explicitly signaled in the SPS or the PPS.

11. A non-transitory computer-readable storage medium having stored therein instructions that cause a processor to: performing conversion between a video and a bitstream of said video, in, The bitstream complies with the format rules, wherein the format rule specifies that the length of an identifier of a sub-picture signaled in a sequence parameter set (SPS) of the bitstream is based on the value of a first syntax element indicating the presence of sub-picture information in the SPS and does not depend on the value of a second syntax element indicating whether the identifier is explicitly signaled in the SPS, Wherein, the length corresponds to the number of bits used to represent a third syntax element indicating a sub-picture identifier in the SPS; if present, a fourth syntax element indicating a sub-picture identifier in a picture parameter set (PPS); and if present, a fifth syntax element indicating a sub-picture identifier in a slice header.

12. The non-transitory computer-readable storage medium of claim 11, in, The length is signaled by a sixth syntax element, sps_subpic_id_len_minus1, whose value plus 1 is equal to the length, wherein the value of sps_subpic_id_len_minus1 ranges from 0 to 15.

13. The non-transitory computer-readable storage medium of claim 11, in, The second syntax element being equal to 1 specifies that a set of identifiers is explicitly signaled for each sub-picture in the SPS or the picture parameter set (PPS), one identifier for each sub-picture, and the second syntax element being equal to 0 specifies that identifiers are not explicitly signaled in the SPS or the PPS.

14. A non-transitory computer-readable recording medium storing a bit stream of a video generated by a method executed by a video processing device, in, The method comprises: generating a bitstream of the video, wherein the bit stream complies with the format rules, wherein the format rule specifies that the length of an identifier of a sub-picture signaled in a sequence parameter set (SPS) of the bitstream is based on the value of a first syntax element indicating the presence of sub-picture information in the SPS and does not depend on the value of a second syntax element indicating whether the identifier is explicitly signaled in the SPS, Wherein, the length corresponds to the number of bits used to represent a third syntax element indicating a sub-picture identifier in the SPS; if present, a fourth syntax element indicating a sub-picture identifier in a picture parameter set (PPS); and if present, a fifth syntax element indicating a sub-picture identifier in a slice header.

15. The non-transitory computer-readable recording medium according to claim 14, in, The length is signaled by a sixth syntax element, sps_subpic_id_len_minus1, whose value plus 1 is equal to the length, wherein the value of sps_subpic_id_len_minus1 ranges from 0 to 15.

16. The non-transitory computer-readable recording medium according to claim 14, in, The second syntax element being equal to 1 specifies that a set of identifiers is explicitly signaled for each sub-picture in the SPS or the picture parameter set (PPS), one identifier for each sub-picture, and the second syntax element being equal to 0 specifies that identifiers are not explicitly signaled in the SPS or the PPS.

17. A method for storing a bit stream of a video, include: generating a bitstream of the video; as well as storing the bitstream in a non-transitory computer-readable recording medium, wherein the bit stream complies with the format rules, wherein the format rule specifies that the length of an identifier of a sub-picture signaled in a sequence parameter set (SPS) of the bitstream is based on the value of a first syntax element indicating the presence of sub-picture information in the SPS and does not depend on the value of a second syntax element indicating whether the identifier is explicitly signaled in the SPS, Wherein, the length corresponds to the number of bits used to represent a third syntax element indicating a sub-picture identifier in the SPS; if present, a fourth syntax element indicating a sub-picture identifier in a picture parameter set (PPS); and if present, a fifth syntax element indicating a sub-picture identifier in a slice header.

Citation Information

Patent Citations

  • Bitstream conformance test in video coding

    CN104662914A

  • Bitstream properties in video coding

    CN104813671A