Restrictions on inter prediction of sub-pictures
By extending the sub-picture number limit in the VVC standard and optimizing the signaling notification conditions, the problems in the signaling notification of sub-pictures, slices and strips are solved, more flexible video encoding and decoding is achieved, multi-layer video encoding and decoding and scalable video encoding and decoding are supported, and the encoding and decoding efficiency and flexibility are improved.
Patent Information
- Application Number
- CN202180008179.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-01-04
- Filing Date
- 2021-01-04
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2041-01-04
AI Technical Summary
Existing video coding and decoding technologies have problems in the signaling notification design of sub-pictures, slices and strips in the VVC standard, such as sub-picture number restrictions, redundant signaling notification conditions, unclear sub-picture ID signaling, and signaling length of strip index ID. In addition, there is a lack of inter-layer prediction constraints, which limits the coding and decoding efficiency and flexibility.
By extending the codec of sps_num_subpics_minus1 to ue(v), more than 256 sub-pictures are allowed per picture, the signaling notification conditions of the syntax elements are adjusted to ensure the reasonable signaling of the sub-picture ID, the slice index length is optimized, and constraints on inter-layer prediction are introduced to support multi-layer video coding and decoding.
It achieves more flexible sub-picture segmentation and encoding and decoding, improves encoding and decoding efficiency, supports more application scenarios, especially multi-layer video encoding and decoding and scalable video encoding and decoding, and improves the flexibility and efficiency of encoding and decoding.
Smart Images

Figure CN114930837B_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to and the benefit of U.S. Provisional Patent Application No. 62 / 957,123, filed on January 4, 2020, in a timely manner under applicable patent laws and / or the Paris Convention. The entire disclosure of the above application is incorporated herein by reference and made a part of the disclosure of this application for all legal purposes. Technical Field
[0003] This patent document relates to image and video encoding and decoding. Background Art
[0004] Digital video consumes the largest amount of bandwidth on the Internet and other digital communications networks. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth demand for digital video usage is expected to continue to grow. Summary of the Invention
[0005] This document discloses techniques that can be used by video encoders and decoders for video encoding or decoding, and includes restrictions on inter prediction of sub-pictures.
[0006] In one example aspect, a video processing method is disclosed. The method includes performing conversion between a video including a plurality of pictures including one or more sub-pictures and a bitstream of the video, wherein the bitstream conforms to a format rule that specifies whether to disallow a current sub-picture from referencing a previous sub-picture for inter-frame prediction if a current sub-picture in a current picture has a current identifier that is different from an identifier of a previous sub-picture in a previous picture at the same position as the current sub-picture.
[0007] In another example aspect, a video processing method is disclosed. The method includes performing conversion between a video comprising a video region and a bitstream of the video comprising a plurality of codec layers, wherein the bitstream complies with a format rule, and wherein the format rule specifies whether inter-layer prediction (ILP) between the video region in different codec layers of the plurality of codec layers is allowed based on a condition.
[0008] In yet another example aspect, a video processing method is disclosed. The method includes performing conversion between a video including a current video region and a bitstream of the video including a plurality of codec layers, wherein the bitstream conforms to a format rule, and wherein the format rule provides that the bitstream includes an indication of whether inter-layer prediction (ILP) is allowed between the current video region and a video region in a reference layer.
[0009] In yet another exemplary aspect, a video encoder apparatus is disclosed. The video encoder includes a processor configured to implement the above method.
[0010] In yet another exemplary aspect, a video decoder apparatus is disclosed. The video decoder includes a processor configured to implement the above method.
[0011] In yet another exemplary aspect, a computer-readable medium having stored thereon code is disclosed. The code is in the form of processor-executable code embodying one of the methods described herein.
[0012] These and other features will be described throughout this document. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 An example of partitioning a picture using luma codec tree units (CTUs) is shown.
[0014] Figure 2 Another example of partitioning a picture using luma CTUs is shown.
[0015] Figure 3 An example segmentation of a picture is shown.
[0016] Figure 4 Another example segmentation of a picture is shown.
[0017] Figure 5 is a block diagram of an example video processing system in which the disclosed technology may be implemented.
[0018] Figure 6 is a block diagram of an example hardware platform for video processing.
[0019] Figure 7 is a block diagram illustrating a video encoding and decoding system according to some embodiments of the present disclosure.
[0020] Figure 8 is a block diagram illustrating an encoder according to some embodiments of the present disclosure.
[0021] Figure 9 is a block diagram illustrating a decoder according to some embodiments of the present disclosure.
[0022] Figure 10-12 A flow chart illustrating an example method for video processing is shown. DETAILED DESCRIPTION
[0023] The section headings used in this document are intended to facilitate understanding and are not intended to limit the applicability of the techniques and embodiments disclosed in each section to that section. Furthermore, the use of H.266 terminology in some descriptions is intended to facilitate understanding and is not intended to limit the scope of the disclosed techniques. Therefore, the techniques described herein are also applicable to other video codec protocols and designs.
[0024] 1. Initial Discussion
[0025] This document is about video codec technology. Specifically, it is about the signaling of sub-pictures, slices, and slices. These ideas can be applied, alone or in various combinations, to any video codec standard or non-standard video codec that supports multi-layer video coding, such as the Versatile Video Codec (VVC) under development.
[0026] 2. Abbreviation
[0027] APS Adaptive Parameter Set
[0028] AU Access Unit
[0029] AUD Access Unit Delimiter
[0030] AVC Advanced Video Codec
[0031] CLVS codec layer video sequence
[0032] CPB codec picture buffer
[0033] CRA Clean Random Access
[0034] CTU Codec Tree Unit
[0035] CVS codec video sequence
[0036] DPB decoded picture buffer
[0037] DPS decoding parameter set
[0038] EOB End of bitstream
[0039] EOS End of sequence
[0040] GDR Progressive Decode Refresh
[0041] HEVC High-Efficiency Video Codec
[0042] HRD Hypothesized Reference Decoder
[0043] IDR Instant Decode Refresh
[0044] JEM Joint Exploration Model
[0045] MCTS motion constraint set
[0046] NAL Network Abstraction Layer
[0047] OLS output layer set
[0048] PH Image Header
[0049] PPS Picture Parameter Set
[0050] PTL Profiles, Hierarchies, and Levels
[0051] PU picture unit
[0052] RBSP Raw Byte Sequence Payload
[0053] SEI Supplemental Enhancement Information
[0054] SPS sequence parameter set
[0055] SVC Scalable Video Codec
[0056] VCL video codec layer
[0057] VPS Video Parameter Set
[0058] VTM VVC test model
[0059] VUI Video Availability Information
[0060] VVC multifunctional video codec
[0061] 3. Introduction to Video Codec
[0062] Video codec standards have evolved primarily through the development of the well-known ITU-T and ISO / IEC standards. ITU-T produced H.261 and H.263, ISO / IEC produced MPEG-1 and MPEG-4 Visual, and the two organizations worked together to produce H.262 / MPEG-2 Video, H.264 / MPEG-4 Advanced Video Codec (AVC), and H.265 / HEVC. Since H.262, video codec standards have been based on a hybrid video codec architecture that utilizes temporal prediction plus transform coding. To explore future video codec technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Team (JVET) in 2015. Since then, JVET has adopted many new approaches and incorporated them into reference software called the Joint Exploration Model (JEM). JVET meetings are held quarterly, and the goal for new codec standards is to reduce bitrates by 50% compared to HEVC. The new video codec standard was officially named the Versatile Video Codec (VVC) at the JVET meeting in April 2018, and the first version of the VVC Test Model (VTM) was also released at that time. Due to the continuous efforts to contribute to VVC standardization, new codec technologies are adopted into the VVC standard at each JVET meeting. Subsequently, the VVC working draft and test model VTM are updated after each meeting. The VVC project is now aiming for technical completion (FDIS) at the July 2020 meeting.
[0063] 3.1. Image Segmentation Scheme in HEVC
[0064] HEVC includes four different picture partitioning schemes, namely normal slice, dependent slice, tile, and wavefront parallel processing (WPP), which can be applied to maximum transmission unit (MTU) size matching, parallel processing, and reduced end-to-end delay.
[0065] Regular slices are similar to those in H.264 / AVC. Each regular slice is encapsulated in its own NAL unit, and intra-picture prediction (intra sample prediction, motion information prediction, codec mode prediction) and entropy codec dependencies are disabled across slice boundaries. Therefore, regular slices can be reconstructed independently of other regular slices in the same picture (although there may still be interdependencies due to loop filtering operations).
[0066] Regular slices are the only tool that can be used for parallelization, and are available in H.264 / AVC in a nearly identical form. Regular slice-based parallelization does not require much inter-processor or inter-core communication (except for inter-processor or inter-core data sharing for motion compensation when decoding predicted codec pictures, which is typically much more onerous than inter-processor or inter-core data sharing due to intra-picture prediction). However, for the same reasons, the use of regular slices incurs significant codec overhead due to the bit cost of the slice header and the lack of prediction across slice boundaries. Furthermore, due to the intra-picture independence of regular slices and the fact that each regular slice is encapsulated in its own NAL unit, regular slices (compared to the other tools mentioned below) also serve as a key mechanism for bitstream segmentation to match MTU size requirements. In many cases, the goals of parallelization and MTU size matching place conflicting demands on the slice layout within a picture. Recognition of this situation led to the development of the parallelization tools mentioned below.
[0067] Dependent slices have short slice headers and allow the bitstream to be split at treeblock boundaries without breaking any intra-picture prediction. Essentially, dependent slices provide fragmentation of regular slices into multiple NAL units to provide reduced end-to-end latency by allowing part of a regular slice to be sent before coding of the entire regular slice is complete.
[0068] In WPP, a picture is partitioned into a single row of codec treeblocks (CTBs). Entropy decoding and prediction are allowed to use data from CTBs in other partitions. Parallel processing is possible through parallel decoding of CTB rows, where the start of decoding is delayed by two CTBs to ensure that data related to CTBs above and to the right of the subject CTB is available before the subject CTB is decoded. This staggered start (which, when represented graphically, looks like a wavefront) allows parallelization using up to as many processors / cores as the picture contains CTB rows. Because intra-picture prediction is allowed between adjacent treeblock rows within a picture, the inter-processor / inter-core communication required to enable intra-picture prediction can be extensive. WPP partitioning does not result in the generation of additional NAL units compared to when it is not used, so WPP is not a tool for MTU size matching. However, if MTU size matching is required, conventional slices can be used with WPP, with some codec overhead.
[0069] Slices define the horizontal and vertical boundaries that divide an image into slice columns and rows. Slice columns extend from the top to the bottom of the image. Similarly, slice rows extend from the left to the right of the image. The number of slices in an image can be simply derived by multiplying the number of slice columns by the number of slice rows.
[0070] Before decoding the top left CTB of the next slice in the order of the slice raster scan of the picture, the scan order of the CTBs is changed to be local within the slice (in the order of the slice's CTB raster scan). Similar to regular slices, slices break intra-picture prediction dependencies and entropy decoding dependencies. However, they do not need to be included in separate NAL units (the same as WPP in this respect); therefore, slices cannot be used for MTU size matching. Each slice can be processed by one processor / core, and the inter-processor / inter-core communication required for intra-picture prediction between processing units decoding adjacent slices is limited to transmitting a shared slice header when the slice spans more than one slice, and loop filtering related to sharing of reconstruction samples and metadata. When more than one slice or WPP segment is included in a slice, the entry point byte offset of each slice or WPP segment in the slice except the first is signaled in the slice header.
[0071] For simplicity, HEVC has specified restrictions on the application of four different picture partitioning schemes. A given codec video sequence cannot include both slices and wavefronts from most profiles specified in HEVC. For each slice and slice, one or both of the following conditions must be met: 1) all codec treeblocks in a slice belong to the same slice; 2) all coded treeblocks in a slice belong to the same slice. Finally, a wavefront segment contains exactly one CTB row, and when using WPP, if a slice starts within a CTB row, it must end on the same CTB row.
[0072] The latest revision to HEVC is specified in the JCT-VC output document JCTVC-AC1005, J. Boyce, A. Ramasubramonian, R. Skupin, G.J. Sullivan, A. Tourapis, Y.-K. Wang (eds.), "HEVC Additional Supplementary Enhancement Information (Draft 4)", publicly released on October 24, 2017: http: / / phenix.int-evry.fr / jct / doc_end_user / documents / 29_Macau / wg11 / JCTVC-AC1005-v2.zip. With the inclusion of this revision, HEVC specifies three SEI messages related to MCTS, namely the time-domain MCTS SEI message, the MCTS Extraction Information Set SEI message, and the MCTS Extraction Information Nesting SEI message.
[0073] The temporal MCTS SEI message indicates the presence of MCTS in the bitstream and signals the MCTS. For each MCTS, motion vectors are restricted to pointing to full sample positions within the MCTS and fractional sample positions that only require full sample positions within the MCTS for interpolation, and do not allow the use of motion vector candidates derived from blocks outside the MCTS for temporal motion vector prediction. In this way, each MCTS can be decoded independently without the presence of slices not included in the MCTS.
[0074] The MCTS extraction information set SEI message provides supplementary information (defined as part of the semantics of the SEI message) that can be used in MCTS sub-bitstream extraction to generate conforming bitstreams for MCTS sets. The information consists of multiple extraction information sets, each defining multiple MCTS sets and containing RBSP bytes that replace the VPS, SPS, and PPS to be used during the MCTS sub-bitstream extraction process. When extracting a sub-bitstream according to the MCTS sub-bitstream extraction process, the parameter sets (VPS, SPS, and PPS) need to be rewritten or replaced, and the slice header needs to be slightly updated because one or all slice address-related syntax elements (including first_slice_segment_in_pic_flag and slice_segment_address) usually need to have different values.
[0075] 3.2.VVC Image Segmentation
[0076] In VVC, a picture is divided into one or more slice rows and one or more slice columns. A slice is a sequence of CTUs covering a rectangular area of a picture. The CTUs in a slice are scanned in raster scan order within the slice.
[0077] A slice consists of an integer number of complete slices or an integer number of consecutive complete CTU rows within a slice of a picture.
[0078] Two striping modes are supported: raster scan striping mode and rectangular striping mode. In raster scan striping mode, a stripe contains a complete sequence of slices in a slice raster scan of a picture. In rectangular striping mode, a stripe contains multiple complete slices that together form a rectangular area of a picture, or multiple consecutive complete CTU rows that together form a slice of a rectangular area of a picture. Slices within a rectangular stripe are scanned in slice raster scan order within the rectangular area corresponding to the stripe.
[0079] A sub-picture consists of one or more strips that together cover a rectangular area of the picture.
[0080] Figure 1 An example of raster scan striping partitioning of a picture is shown, where the picture is divided into 12 slices and 3 raster scan strips.
[0081] Figure 2 An example of rectangular strip partitioning of a picture is shown, where the picture is divided into 24 slices (6 slice columns and 4 slice rows) and 9 rectangular strips.
[0082] Figure 3 An example of a picture partitioned into slices and rectangular strips is shown, where the picture is divided into 4 slices (2 slice columns and 2 slice rows) and 4 rectangular strips.
[0083] Figure 4 An example of sub-picture partitioning of a picture is shown, where the picture is partitioned into 18 slices, with each of the 12 left-hand slices covering a strip of 4×4 CTUs, and each of the 6 right-hand slices covering two vertically stacked strips of 2×2 CTUs, resulting in a total of 24 slices and 24 sub-pictures of different dimensions (each slice is a sub-picture).
[0084] 3.3. Signaling Notification of Sub-Pictures, Slices, and Strips in VVC
[0085] In the latest VVC draft text, sub-picture information includes sub-picture layout (i.e., the number of sub-pictures per picture and the position and size of each picture) and other sequence-level sub-picture information, which is signaled in the SPS. The order of sub-pictures signaled in the SPS defines the sub-picture index. For example, in the SPS or PPS, a list of sub-picture IDs (one ID for each sub-picture) can be explicitly signaled.
[0086] Slices in VVC are conceptually the same as in HEVC, ie, each picture is partitioned into slice columns and slice rows, but there is a different syntax for signaling slices in the PPS.
[0087] In VVC, the slice mode is also signaled in the PPS. When the slice mode is rectangular strip mode, the strip layout of each picture (i.e. the number of strips per picture and the position and size of each strip) is signaled in the PPS. The order of the rectangular strips within a picture signaled in the PPS defines the picture level strip index. The sub-picture level strip index is defined as the order of the strips within a sub-picture in ascending order of their picture level strip index. The position and size of the rectangular strips are signaled / derived based on the sub-picture position and size signaled in the SPS (when each sub-picture contains only one stripe) or based on the slice position and size signaled in the PPS (when a sub-picture may contain more than one stripe). When the strip mode is raster scan strip mode, the layout of the strips within the picture is signaled in the stripes themselves, similar to in HEVC, but with different details.
[0088] The SPS, PPS and slice header syntax and semantics in the latest VVC draft text that is most relevant to the invention herein are as follows.
[0089] 7.3.2.3 Sequence Parameter Set RBSP Syntax
[0090]
[0091]
[0092] 7.4.3.3 Sequence Parameter Set RBSP Semantics ...
[0094] subpics_present_flag equal to 1 specifies the presence of sub-picture parameters in the SPS RBSP syntax.
[0095] subpics_present_flag equal to 0 specifies that sub-picture parameters are not present in the SPS RBSP syntax.
[0096] NOTE 2 – When the bitstream is the result of a sub-bitstream extraction process and contains only a subset of the sub-pictures of the input bitstream to the sub-bitstream extraction process, it may be necessary to set the value of subpics_present_flag to 1 in the RBSP of the SPS.
[0097] sps_num_subpics_minus1 plus 1 specifies the number of sub-pictures. sps_num_subpics_minus1 shall be in the range 0 to 254. When not present, sps_num_subpics_minus1 is inferred to be equal to 0. subpic_ctu_top_left_x[i] specifies the horizontal position of the top-left CTU of the i-th sub-picture in units of CtbSizeY. The length of this syntax element is Ceil Ceil(Log2(pic_width_max_in_luma_samples / CtbSizeY)) bits. When not present, subpic_ctu_top_left_x[i] is inferred to be equal to 0.
[0098] subpic_ctu_top_left_y[i] specifies the vertical position of the top left CTU of the i-th sub-picture in units of CtbSizeY. The length of this syntax element is Ceil(Log2(pic_height_max_in_luma_samples / CtbSizeY)) bits. When not present, subpic_ctu_top_left_y[i] is inferred to be equal to 0.
[0099] subpic_width_minus1[i] plus 1 specifies the width of the i-th sub-picture in units of CtbSizeY. The length of this syntax element is Ceil(Log2(pic_width_max_in_luma_samples / CtbSizeY)) bits. When not present, the value of subpic_width_minus1[i] is inferred to be equal to Ceil(pic_width_max_in_luma_samples / CtbSizeY)-1.
[0100] subpic_height_minus1[i] plus 1 specifies the height of the i-th sub-picture in units of CtbSizeY. The length of this syntax element is Ceil(Log2(pic_height_max_in_luma_samples / CtbSizeY)) bits. When not present, the value of subpic_height_minus1[i] is inferred to be equal to Ceil(pic_height_max_in_luma_samples / CtbSizeY)-1.
[0101] subpic_treated_as_pic_flag[i] equal to 1 specifies that the i-th subpicture of each codec picture in the CLVS is treated as a picture during decoding without in-loop filtering operations. subpic_treated_as_pic_flag[i] equal to 0 specifies that the i-th subpicture of each codec picture in the CLVS is not treated as a picture during decoding without in-loop filtering operations. When not present, the value of subpic_treated_as_pic_flag[i] is inferred to be 0.
[0102] loop_filter_across_subpic_enabled_flag[i] equal to 1 specifies that in-loop filtering operations may be performed across the boundaries of the i-th sub-picture in each codec picture in the CLVS. loop_filter_across_subpic_enabled_flag[i] equal to 0 specifies that in-loop filtering operations are not performed across the boundaries of the i-th sub-picture in each codec picture in the CLVS. When not present, the value of loop_filter_across_subpic_enabled_pic_flag[i] is inferred to be equal to 1.
[0103] The requirements for bitstream conformance are that the following constraints apply:
[0104] - For any two sub-pictures subpicA and subpicB, when the sub-picture index of subpicA is less than the sub-picture index of subpicB, any codec slice NAL unit of subPicA should precede any codec slice NAL unit of subPicB in decoding order.
[0105] - The shape of sub-pictures shall be such that each sub-picture, when decoded, shall have its entire left and entire top borders consisting of picture boundaries, or of the boundaries of previously decoded sub-pictures.
[0106] sps_subpic_id_present_flag equal to 1 specifies that sub-picture ID mapping is present in the SPS. sps_subpic_id_present_flag equal to 0 specifies that sub-picture ID mapping is not present in the SPS.
[0107] sps_subpic_id_signalling_present_flag equal to 1 specifies that sub-picture ID mapping is signaled in the SPS. sps_subpic_id_signalling_present_flag equal to 0 specifies that sub-picture ID mapping is not signaled in the SPS. When not present, the value of sps_subpic_id_signalling_present_flag is inferred to be 0.
[0108] sps_subpic_id_len_minus1 plus 1 specifies the number of bits used to represent the syntax element sps_subpic_id[i]. The value of sps_subpic_id_len_minus1 shall be in the range of 0 to 15 (inclusive). sps_subpic_id[i] specifies the sub-picture ID of the i-th sub-picture. The length of the sps_subpic_id[i] syntax element is sps_subpic_id_len_minus1 + 1 bits. When not present, and when sps_subpic_id_present_flag is equal to 0, the value of sps_subpic_id[i] is inferred to be equal to i for each i in the range of 0 to sps_num_subpics_minus1 (inclusive). ...
[0110] 7.3.2.4 Picture Parameter Set RBSP Syntax
[0111]
[0112]
[0113]
[0114] 7.4.3.4 Picture Parameter Set RBSP Semantics ...
[0116] pps_subpic_id_signalling_present_flag equal to 1 specifies that sub-picture ID mapping is signaled in the PPS. pps_subpic_id_signalling_present_flag equal to 0 specifies that sub-picture ID mapping is not signaled in the PPS. When sps_subpic_id_present_flag is 0 or sps_subpic_id_signalling_present_flag is equal to 1, pps_subpic_id_signalling_present_flag shall be equal to 0.
[0117] pps_num_subpics_minus1 plus 1 specifies the number of subpictures in the codec picture that references the PPS. A bitstream conformance requirement is that the value of pps_num_subpic_minus1 shall be equal to sps_num_subpics_minus1.
[0118] pps_subpic_id_len_minus1 plus 1 specifies the number of bits used to represent the syntax element pps_subpic_id[i]. The value of pps_subpic_id_len_minus1 shall be in the range of 0 to 15 (inclusive). A bitstream conformance requirement is that the value of pps_subpic_id_len_minus1 shall be the same for all PPSs referenced by a codec picture in the CLVS.
[0119] pps_subpic_id[i] specifies the sub-picture ID of the i-th sub-picture. The length of the pps_subpic_id[i] syntax element is pps_subpic_id_len_minus1+1 bits.
[0120] no_pic_partition_flag equal to 1 specifies that no picture partitioning is applied to each picture that references the PPS.
[0121] no_pic_partition_flag equal to 0 specifies that each picture of the referenced PPS may be partitioned into more than one slice or slice.
[0122] The bitstream conformance requirement is that the value of no_pic_partition_flag shall be the same for all PPSs referenced by a coded picture within a CLVS.
[0123] A bitstream conformance requirement is that when the value of sps_num_subpics_minus1+1 is greater than 1, the value of no_pic_partition_flag shall not be equal to 1.
[0124] pps_log2_ctu_size_minus5 plus 5 specifies the luma codec treeblock size for each CTU. pps_log2_ctu_size_minus5 shall be equal to sps_log2_ctu_size_minus5.
[0125] num_exp_tile_columns_minus1 plus 1 specifies the number of explicitly provided tile column widths. The value of num_exp_tile_columns_minus1 should be in the range of 0 to PicWidthInCtbsY-1 (inclusive). When no_pic_partition_flag is equal to 1, the value of num_exp_tile_columns_minus1 is inferred to be equal to 0.
[0126] num_exp_tile_rows_minus1 plus 1 specifies the number of explicitly provided tile row heights. The value of num_exp_tile_rows_minus1 should be in the range of 0 to PicHeightInCtbsY-1 (inclusive). When no_pic_partition_flag is equal to 1, the value of num_tile_rows_minus1 is inferred to be equal to 0.
[0127] tile_column_width_minus1[i] plus 1 specifies the width of the i-th tile column in CTBs for i in the range of 0 to num_exp_tile_columns_minus1-1 (inclusive).
[0128] tile_column_width_minus1[num_exp_tile_columns_minus1] is used to derive the width of tile columns with indices greater than or equal to num_exp_tile_columns_minus1 (as specified in section 6.5.1). When not present, the value of tile_column_width_minus1[0] is inferred to be equal to PicWidthInCtbsY-1.
[0129] tile_row_height_minus1[i] plus 1 specifies the height of the i-th tile row in CTBs for i in the range 0 to num_exp_tile_rows_minus1-1 (inclusive). tile_row_height_minus1[num_exp_tile_rows_minus1] is used to derive the height of tile rows with indices greater than or equal to num_exp_tile_rows_minus1 (as specified in section 6.5.1). When not present, the value of tile_row_height_minus1[0] is inferred to be equal to PicHeightInCtbsY-1.
[0130] rect_slice_flag equal to 0 specifies that the slices within each slice are in raster scan order and that slice information is not signaled in the PPS. rect_slice_flag equal to 1 specifies that the slices within each slice cover a rectangular area of the picture and that slice information is signaled in the PPS. When not present, rect_slice_flag is inferred to be equal to 1. When subpics_present_flag is equal to 1, the value of rect_slice_flag shall be equal to 1.
[0131] single_slice_per_subpic_flag equal to 1 specifies that each sub-picture consists of one and only one rectangular slice. single_slice_per_subpic_flag equal to 0 specifies that each sub-picture may consist of one or more rectangular slices. When subpics_present_flag is equal to 0, single_slice_per_subpic_flag shall be equal to 0. When single_slice_per_subpic_flag is equal to 1, num_slices_in_pic_minus1 is inferred to be equal to sps_num_subpics_minus1.
[0132] num_slices_in_pic_minus1 plus 1 specifies the number of rectangular slices in each picture referencing the PPS. The value of num_slices_in_pic_minus1 shall be in the range of 0 to MaxSlicesPerPicture-1, inclusive, where MaxSlicesPerPicture is specified in Annex A. When no_pic_partition_flag is equal to 1, the value of num_slices_in_pic_minus1 is inferred to be equal to 0. tile_idx_delta_present_flag equal to 0 specifies that tile_idx_delta values are not present in the PPS and that all rectangular slices in pictures referencing the PPS are specified in raster order according to the procedure defined in section 6.5.1. tile_idx_delta_present_flag equal to 1 specifies that tile_idx_delta values may be present in the PPS and that all rectangular slices in pictures referencing the PPS are specified in the order indicated by the tile_idx_delta values.
[0133] slice_width_in_tiles_minus1[i] plus 1 specifies the width of the i-th rectangular slice in tile columns. The value of slice_width_in_tiles_minus1[i] shall be in the range of 0 to NumTileColumns-1 (inclusive). When not present, the value of slice_width_in_tiles_minus1[i] is inferred to be the value specified in section 6.5.1.
[0134] slice_height_in_tiles_minus1[i] plus 1 specifies the height of the i-th rectangular slice in tile rows. The value of slice_height_in_tiles_minus1[i] shall be in the range of 0 to NumTileRows-1 (inclusive). When not present, the value of slice_height_in_tiles_minus1[i] is inferred to be the value specified in section 6.5.1.
[0135] num_slices_in_tile_minus1[i] plus 1 specifies the number of slices in the current slice, for the case where the i-th slice contains a subset of CTU rows from a single slice. The value of num_slices_in_tile_minus1[i] shall be in the range of 0 to RowHeight[tileY]-1, inclusive, where tileY is the index of the slice row containing the i-th slice. When not present, the value of num_slices_in_tile_minus1[i] is inferred to be equal to 0.
[0136] slice_height_in_ctu_minus1[i] plus 1 specifies the height of the i-th rectangular slice in CTU rows, for the case where the i-th slice contains a subset of CTU rows from a single slice. The value of slice_height_in_ctu_minus1[i] shall be in the range of 0 to RowHeight[tileY]-1 (inclusive), where tileY is the index of the slice row containing the i-th slice.
[0137] tile_idx_delta[i] specifies the tile index difference between the i-th rectangular strip and the (i+1)-th rectangular strip. The value of tile_idx_delta[i] shall be in the range of –NumTilesInPic+1 to NumTilesInPic-1, inclusive. When not present, the value of tile_idx_delta[i] is inferred to be equal to 0. In all other cases, the value of tile_idx_delta[i] shall not be equal to 0.
[0138] loop_filter_across_tiles_enabled_flag equal to 1 specifies that in-loop filtering operations may be performed across slice boundaries in pictures that reference the PPS. loop_filter_across_tiles_enabled_flag equal to 0 specifies that in-loop filtering operations are not performed across slice boundaries in pictures that reference the PPS. In-loop filtering operations include deblocking filter, sample adaptive offset filter, and adaptive loop filter operations. When not present, the value of loop_filter_across_tiles_enabled_flag is inferred to be 1.
[0139] loop_filter_across_slices_enabled_flag equal to 1 specifies that in-loop filtering operations may be performed across slice boundaries in pictures that reference a PPS. loop_filter_across_slice_enabled_flag equal to 0 specifies that in-loop filtering operations are not performed across slice boundaries in pictures that reference a PPS. In-loop filtering operations include deblocking filter, sample adaptive offset filter, and adaptive loop filter operations. When not present, the value of loop_filter_across_slices_enabled_flag is inferred to be 0. ...
[0141] 7.3.7.1 General Strip Header Syntax
[0142]
[0143]
[0144] 7.4.8.1 Common Strip Header Semantics ...
[0146] slice_subpic_id specifies the sub-picture identifier of the sub-picture containing the slice. If slice_subpic_id exists, the value of the variable SubPicIdx is derived such that SubpicIdList[SubPicIdx] is equal to slice_subpic_id. Otherwise (slice_subpic_id does not exist), the variable SubPicIdx is derived to be equal to 0. The length of slice_subpic_id in bits is derived from:
[0147] If sps_subpic_id_signalling_present_flag is equal to 1, the length of slice_subpic_id is equal to sps_subpic_id_len_minus1+1.
[0148] Otherwise, if ph_subpic_id_signalling_present_flag is equal to 1, the length of slice_subpic_id is equal to ph_subpic_id_len_minus1+1.
[0149] Otherwise, if pps_subpic_id_signalling_present_flag is equal to 1, the length of slice_subpic_id is equal to pps_subpic_id_len_minus1+1.
[0150] Otherwise, the length of slice_subpic_id is equal to Ceil(Log2(sps_num_subpics_minus1+1)).
[0151] slice_address specifies the slice address of the slice. When not present, slice_address is inferred to be equal to 0.
[0152] If rect_slice_flag is equal to 0, the following applies:
[0153] - Strip address is the raster scan strip index.
[0154] - The length of slice_address is Ceil(Log2(NumTilesInPic)) bits.
[0155] -slice_address values should be in the range of 0 to NumTilesInPic-1 (inclusive).
[0156] Otherwise (rect_slice_flag is equal to 1), the following applies:
[0157] - The slice address is the slice index of the slice within the SubPicIdx-th sub-picture.
[0158] - The length of slice_address is Ceil(Log2(NumSlicesInSubpic[SubPicIdx])) bits.
[0159] - The value of slice_address should be in the range of 0 to NumSlicesInSubpic[SubPicIdx]-1 (inclusive).
[0160] The requirements for bitstream conformance are that the following constraints apply:
[0161] - If rect_slice_flag is equal to 0 or subpics_present_flag is equal to 0, the value of slice_address shall not be equal to the value of slice_address of any other codec slice NAL unit of the same codec picture.
[0162] Otherwise, the pair of slice_subpic_id and slice_address values shall not be equal to the pair of slice_subpic_id and slice_address values of any other codec slice NAL unit of the same codec picture.
[0163] - When rect_slice_flag is equal to 0, the slices of the picture will be arranged in ascending order of their slice_address values.
[0164] - The shape of a slice of a picture shall be such that each CTU, when decoded, shall have its entire left and entire top borders consisting of picture boundaries or of the boundaries of previously decoded CTU(s). num_tiles_in_slice_minus1, plus 1, when present, specifies the number of slices in the slice. The value of num_tiles_in_slice_minus1 shall be in the range of 0 to NumTilesInPic - 1, inclusive.
[0165] The variable NumCtuInCurrSlice specifies the number of CTUs in the current slice, and for i in the range 0 to NumCtuInCurrSlice-1 (inclusive), the list CtbAddrInCurrSlice[i] specifies the picture raster scan address of the i-th CTB in the slice, derived as follows:
[0166]
[0167]
[0168] The variables SubPicLeftBoundaryPos, SubPicTopBoundaryPos, SubPicRightBoundaryPos, and SubPicBotBoundaryPos are derived as follows:
[0169] ...
[0171] 4. Examples of technical problems solved by the technical solutions in this article
[0172] In VVC, existing designs for signaling sub-pictures, slices, and stripes have the following problems:
[0173] 1) The codec for sps_num_subpics_minus1 is u(8), which does not allow more than 256 subpictures per picture. However, in some applications, the maximum number of subpictures per picture may need to be greater than 256.
[0174] 2) It is allowed that subpics_present_flag is equal to 0 and sps_subpic_id_present_flag is equal to 1. However, it does not make sense because subpics_present_flag equal to 0 means that CLVS has no information about sub-pictures at all.
[0175] 3) A list of sub-picture IDs can be signaled in the picture header (PH), one for each sub-picture. However, when the list of sub-picture IDs is signaled in the PH, and when a subset of sub-pictures is extracted from the bitstream, all PHs need to be changed. This is undesirable.
[0176] 4) Currently, when the sub-picture ID is indicated to be explicitly signaled by sps_subpic_id_present_flag being equal to 1 (or the name of the syntax element is changed to subpic_ids_explicitly_signalled_flag), the sub-picture ID may not be signaled anywhere. This is problematic because when the sub-picture ID is indicated to be explicitly signaled, the sub-picture ID needs to be explicitly signaled in the SPS or PPS.
[0177] 5) When the sub-picture ID is not explicitly signaled, the slice header syntax element slice_subpic_id still needs to be signaled whenever subpics_present_flag is equal to 1 (including when sps_num_subpics_minus1 is equal to 0). However, the length of slice_subpic_id is currently specified to be Ceil(Log2(sps_num_subpics_minus1+1)) bits, which can be 0 bits when sps_num_subpics_minus1 is equal to 0. This is problematic because any existing syntax element cannot be 0 bits long.
[0178] 6) The sub-picture layout, including the number of sub-pictures and their size and position, remains unchanged for the entire CLVS. Even when the sub-picture ID is not explicitly signaled in the SPS or PPS, the sub-picture ID length still needs to be signaled for the sub-picture ID syntax element in the slice header.
[0179] 7) Whenever rect_slice_flag is equal to 1, the syntax element slice_address is signaled in the slice header and specifies the slice index within the sub-picture that contains the slice (including when the number of slices within the sub-picture (i.e., NumSlicesInSubpic[SubPicIdx]) is equal to 1). However, currently, when rect_slice_flag is equal to 1, slice_address is specified to be of length Ceil(Log2(NumSlicesInSubpic[SubPicIdx])) bits, which is 0 bits when NumSlicesInSubpic[SubPicIdx] is equal to 1. This is problematic because any existing syntax element cannot be 0 bits in length.
[0180] 8) There is redundancy between the syntax elements no_pic_partition_flag and pps_num_subpics_minus1, although the latest VVC text has the following constraint: when sps_num_subpics_minus1 is greater than 0, the value of no_pic_partition_flag shall be equal to 1.
[0181] 9) Within CLVS, the sub-picture ID value for a particular sub-picture position or index may vary from picture to picture. When this occurs, in principle, the sub-picture cannot use inter prediction by referencing reference pictures in the same layer. However, currently, there is a lack of constraints in the current VVC specification that prohibit this situation.
[0182] 10) In the current VVC design, reference pictures can be pictures in different layers to support various applications, such as scalable video codec and multi-view video codec. If sub-pictures exist in different layers, it is necessary to study whether to allow or not allow inter-layer prediction.
[0183] 5. Example Techniques and Embodiments
[0184] To solve the above problems and other problems, the following methods are disclosed. The present invention should be considered as an example to explain the general concept and should not be interpreted in a narrow way. In addition, these inventions can be applied alone or in combination in any way.
[0185] 1) To solve the first problem, change the codec of sps_num_subpics_minus1 from u(8) to ue(v) to enable more than 256 subpictures per picture.
[0186] a. In addition, the value of sps_num_subpics_minus1 is limited to 0 to
[0187] Ceil(pic_width_max_in_luma_samples÷CtbSizeY)*
[0188] The range is within Ceil(pic_height_max_in_luma_samples ÷ CtbSizeY)-1 (including the endpoints).
[0189] b. In addition, the number of sub-pictures per picture is further restricted in the definition of the level.
[0190] 2) To solve the second problem, the signaling condition of the syntax element sps_subpic_id_present_flag is set to "if(subpics_present_flag)", that is, when subpics_present_flag is equal to 0, the syntax element sps_subpic_id_present_flag is not signaled, and when it does not exist, the value of sps_subpic_id_present_flag is inferred to be equal to 0.
[0191] a. Alternatively, when subpics_present_flag is equal to 0, the syntax element sps_subpic_id_present_flag is still signaled, but when subpics_present_flag is equal to 0, then this value needs to be equal to 0.
[0192] b. In addition, the names of the syntax elements subpics_present_flag and sps_subpic_id_present_flag are changed to subpic_info_present_flag and subpic_ids_explicitly_signalled_flag, respectively.
[0193] 3) To address the third issue, the signaling of sub-picture IDs in the PH syntax is removed. Therefore, for i in the range of 0 to sps_num_subpics_minus1 (inclusive), the list SubpicIdList[i] is derived as follows:
[0194]
[0195]
[0196] 4) To address the fourth problem, when a sub-picture is indicated to be explicitly signaled, the sub-picture ID is signaled in the SPS or PPS.
[0197] a. This is achieved by adding the following constraint: If subpic_ids_explicitly_signalled_flag is 0 or subpic_ids_in_sps_flag is equal to 1, then subpic_ids_in_pps_flag shall be equal to 0. Otherwise (subpic_ids_explicitly_signalled_flag is 1 and subpic_ids_in_sps_flag is equal to 0), subpic_ids_in_pps_flag shall be equal to 1.
[0198] 5) To address the fifth and sixth issues, the length of the sub-picture ID is signaled in the SPS regardless of the value of the SPS flag sps_subpic_id_present_flag (or renamed subpic_ids_explicitly_signalled_flag). Although when the sub-picture ID is explicitly signaled in the PPS, the length can also be signaled in the PPS to avoid parsing the PPS's dependency on the SPS. In this case, the length also specifies the length of the sub-picture ID in the slice header, even if the sub-picture ID is not explicitly signaled in the SPS or PPS. Therefore, when present, the length of slice_subpic_id is also specified by the sub-picture ID length signaled in the SPS.
[0199] 6) Alternatively, to address the fifth and sixth issues, a flag is added to the SPS syntax with a value of 1 to specify the presence of the sub-picture ID length in the SPS syntax. The presence of this flag does not depend on the value of the flag indicating whether the sub-picture ID is explicitly signaled in the SPS or PPS. When subpic_ids_explicitly_signalled_flag is equal to 0, the value of this flag can be equal to 1 or 0, but when subpic_ids_explicitly_signalled_flag is equal to 1, the value of this flag must be equal to 1. When this flag is equal to 0 (i.e., when the sub-picture length does not exist), the length of slice_subpic_id is specified to be Max(Ceil(Log2(sps_num_subpics_minus1+1)), 1) bits (as opposed to Ceil(Log2(sps_num_subpics_minus1+1)) bits in the latest VVC draft text).
[0200] a. Alternatively, this flag is only present when subpic_ids_explicitly_signalled_flag is equal to 0, and when subpic_ids_explicitly_signalled_flag is equal to 1, the value of this flag is inferred to be equal to 1.
[0201] 7) To solve the seventh problem, when rect_slice_flag is equal to 1, the length of slice_address is specified to be Max(Ceil(Log2(NumSlicesInSubpic[SubPicIdx])), 1) bits.
[0202] a. Alternatively, furthermore, when rect_slice_flag is equal to 0, the length of slice_address is specified to be Max(Ceil(Log2(NumTilesInPic)), 1) bits, as opposed to Ceil(Log2(NumTilesInPic)) bits.
[0203] 8) To address the eighth issue, the signaling condition for no_pic_partition_flag is set to “if (subpic_ids_in_pps_flag&&pps_num_subpics_minus1>0)”, and the following inference is added: when not present, the value of no_pic_partition_flag is inferred to be equal to 1.
[0204] a. Alternatively, move the sub-picture ID syntax (all four syntax elements) after the slice and slice syntax in the PPS, e.g., immediately before the syntax element entropy_coding_sync_enabled_flag, and then set the condition for the signaling of pps_num_subpics_minus1 to "if (no_pic_partition_flag)".
[0205] 9) To address the ninth issue, the following constraint is specified: for each specific sub-picture index (or equivalently, sub-picture position), when the sub-picture ID value at picture picA changes compared to the sub-picture ID value of the previous picture in decoding order in the same layer of picA, unless picA is the first picture of CLVS, the sub-picture at picA shall only contain codec slice NAL units with nal_unit_type equal to IDR_W_RADL, IDR_N_LP or CRA_NUT.
[0206] a. Alternatively, the above constraint applies only to sub-picture indices whose value of subpic_treated_as_pic_flag[i] is equal to 1.
[0207] b. Alternatively, for items 9 and 9a, change “IDR_W_RADL, IDR_N_LP, or CRA_NUT” to “IDR_W_RADL, IDR_N_LP, CRA_NUT, RSV_IRAP_11, or RSV_IRAP_12”.
[0208] c. Alternatively, the sub-picture at picA may contain other types of codec slice NAL units, however, these codec slice NAL units only use one or more of intra prediction, intra block copy (IBC) prediction, and palette mode prediction.
[0209] d. Alternatively, a first video unit (such as a slice, a tile, or a block) in a sub-picture of picA can reference a second video unit in a previous picture. This constrains that, although the sub-picture IDs of the second video unit and the first video unit may be different, both can be in a sub-picture with the same sub-picture index. The sub-picture index is a unique number assigned to a sub-picture that cannot be changed in CLVS.
[0210] 10) For a specific sub-picture index (or equivalently, sub-picture position), an indication of which sub-pictures, identified by a layer ID value together with the sub-picture index or sub-picture ID value, are allowed to be used as reference pictures may be signaled in the bitstream.
[0211] 11) For the multi-layer case, inter-layer prediction (ILR) based on sub-pictures from different layers is allowed when certain conditions are met (e.g., may depend on the number of sub-pictures, the location of sub-pictures), and ILR is disabled when certain conditions are not met.
[0212] a. In one example, even when two sub-pictures in two layers have the same sub-picture index value but different sub-picture ID values, inter-layer prediction may still be allowed when certain conditions are met.
[0213] i. In one example, some of the conditions are "if two layers are associated with different view order index / view order ID values".
[0214] b. If the first sub-picture in the first layer and the second sub-picture in the second layer have the same sub-picture index, it can be constrained that the two sub-pictures must be in co-located positions and / or have reasonable widths / heights.
[0215] c. If the first sub-picture can refer to the second reference sub-picture, the first sub-picture in the first layer and the second sub-picture in the second layer can be constrained to be at the same position and / or have a reasonable width / height.
[0216] 12) In the bitstream (such as in the VPS / DPS / SPS / PPS / APS / sequence header / picture header), signaling whether the current sub-picture can use inter-layer prediction (ILP) from sample values and / or other values (e.g., motion information and / or coding mode information) associated with a region or sub-picture of the reference layer.
[0217] a. In one example, the reference regions or sub-pictures of the reference layer are those reference regions or sub-pictures that contain at least one co-located sample of a sample in the current sub-picture.
[0218] b. In one example, the reference region or sub-picture of the reference layer is outside the co-located region of the current sub-picture.
[0219] c. In one example, this indication is signaled in one or more SEI messages.
[0220] d. In one example, regardless of whether the reference layer has multiple sub-pictures, and when multiple sub-pictures are present in one or more reference layers, regardless of whether the partitioning of the picture into sub-pictures is aligned with the current picture such that each sub-picture in the current picture has a corresponding sub-picture in the reference picture that covers the co-located area, and further regardless of whether the corresponding / co-located sub-picture has the same sub-picture ID value as the current sub-picture, such an indication is signaled.
[0221] 6. Examples
[0222] The following are some example embodiments of all the inventive aspects except for item 8 summarized in Section 5 above, which are applicable to the VVC specification. The modified text is based on the latest VVC text in JVET-P2001-v14. The most relevant parts that have been added or modified are Underlined, bold, and italic text The most relevant deleted parts are shown, and the most relevant deleted parts are highlighted in bold double brackets, for example, [[a]] indicates that "a" has been deleted. There are also some other changes that are editorial in nature and are therefore not highlighted.
[0223] 6.1. First embodiment
[0224] 7.3.2.3 Sequence Parameter Set RBSP Syntax
[0225]
[0226] 7.4.3.3 Sequence Parameter Set RBSP Semantics ...
[0228] subpic_info_present_flag is equal to 1 to specify that sub-picture information is present. flag is equal to 0, which specifies that sub-picture information does not exist. sps_ref_pic_resampling_enabled_flag and subpic_ The value of info_present_flag shall not all be equal to 1.
[0229] NOTE 2 – When the bitstream is the result of a sub-bitstream extraction process and contains only a subset of the sub-pictures of the input bitstream of the sub-bitstream extraction process, it may be necessary to set
[0230] subpic_info_present_flag The value of is equal to 1.
[0231] sps_num_subpics_minus1 plus 1 specifies the number of sub-pictures. The value of sps_num_subpics_minus1 Should be between 0 and Ceil(pic_width_max_in_luma_samples÷CtbSizeY)*Ceil(pic_height_max_ in_luma_samples÷CtbSizeY)-1 (inclusive). Note that the maximum number of sub-pictures can be at level Further restricted in other definitions When not present, the value of sps_num_subpics_minus1 is inferred to be equal to 0.
[0232] subpic_ctu_top_left_x[i] specifies the horizontal position of the top left CTU of the i-th sub-picture in units of CtbSizeY. The length of this syntax element is Ceil(Log2(pic_width_max_in_luma_samples / CtbSizeY)) bits. When not present, subpic_ctu_top_left_x[i] is inferred to be equal to 0.
[0233] subpic_ctu_top_left_y[i] specifies the vertical position of the top left CTU of the i-th sub-picture in units of CtbSizeY. The length of this syntax element is Ceil(Log2(pic_height_max_in_luma_samples / CtbSizeY)) bits. When not present, subpic_ctu_top_left_y[i] is inferred to be equal to 0.
[0234] subpic_width_minus1[i] plus 1 specifies the width of the i-th sub-picture in units of CtbSizeY. The length of this syntax element is Ceil(Log2(pic_width_max_in_luma_samples / CtbSizeY)) bits. When not present, the value of subpic_width_minus1[i] is inferred to be equal to Ceil(pic_width_max_in_luma_samples / CtbSizeY)-1.
[0235] subpic_height_minus1[i] plus 1 specifies the height of the i-th sub-picture in units of CtbSizeY. The length of this syntax element is Ceil(Log2(pic_height_max_in_luma_samples / CtbSizeY)) bits. When not present, the value of subpic_height_minus1[i] is inferred to be equal to Ceil(pic_height_max_in_luma_samples / CtbSizeY)-1.
[0236] subpic_treated_as_pic_flag[i] equal to 1 specifies that the i-th subpicture of each codec picture in the CLVS is treated as a picture during decoding without in-loop filtering operations. subpic_treated_as_pic_flag[i] equal to 0 specifies that the i-th subpicture of each codec picture in the CLVS is not treated as a picture during decoding without in-loop filtering operations. When not present, the value of subpic_treated_as_pic_flag[i] is inferred to be 0.
[0237] loop_filter_across_subpic_enabled_flag[i] equal to 1 specifies that in-loop filtering operations may be performed across the boundaries of the i-th sub-picture in each codec picture in the CLVS. loop_filter_across_subpic_enabled_flag[i] equal to 0 specifies that in-loop filtering operations are not performed across the boundaries of the i-th sub-picture in each codec picture in the CLVS. When not present, the value of loop_filter_across_subpic_enabled_pic_flag[i] is inferred to be equal to 1.
[0238] The requirements for bitstream conformance are that the following constraints apply:
[0239] - For any two sub-pictures subpicA and subpicB, when the sub-picture index of subpicA is less than the sub-picture index of subpicB, any codec slice NAL unit of subPicA should precede any codec slice NAL unit of subPicB in decoding order.
[0240] - The shape of sub-pictures should be such that each sub-picture, when decoded, should have its entire left and entire top borders consisting of picture boundaries, or of the boundaries of previously decoded sub-pictures.
[0241] sps_subpic_id_len_minus1 plus 1 specifies the sub-picture ID and slice for explicit signaling notification The value of sps_subpic_id_len_minus1 shall be between 0 and 15 (inclusive of the endpoints).
[0242] subpic_ids_explicitly_signalled_flag is equal to 1 to specify that for each sub-picture in the SPS or PPS Explicitly signal the set of sub-picture IDs, one per sub-picture. subpic_ids_explicitly_signalled_flag Equal to 0 specifies that no sub-picture IDs are explicitly signaled in the SPS or PPS. When not present, subpic_ids_ The value of explicitly_signalled_flag is equal to 0.
[0243] subpic_ids_in_sps_flag equal to 1 specifies that the subpicture ids of each subpicture are explicitly signaled in the SPS. subpic_ids_in_sps_flag is equal to 0 and specifies that no sub-picture IDs are signaled in the SPS. The value of subpic_ids_in_sps_flag is equal to 0.
[0244] sps_subpic_id[i] specifies the sub-picture ID of the i-th sub-picture. The length of the sps_subpic_id[i] syntax element is sps_subpic_id_len_minus1+1 bits. ...
[0246] 7.3.2.4 Picture Parameter Set RBSP Syntax
[0247]
[0248]
[0249]
[0250] 7.4.3.4 Picture Parameter Set RBSP Semantics ...
[0252] subpic_ids_in_pps_flag equal to 1 specifies that the subpicture ids of each subpicture are explicitly signaled in the PPS. subpic_ids_in_pps_flag is equal to 0 and specifies that no sub-picture IDs are signaled in the PPS. ids_explicitly_signalled_flag is 0 or subpic_ids_in_sps_flag is equal to 1, then subpic_ids_ in_pps_flag shall be equal to 0. Otherwise (subpic_ids_explicitly_signalled_flag is 1 and subpic_ ids_in_sps_flag is equal to 0), subpic_ids_in_pps_flag shall be equal to 1.
[0253] pps_num_subpics_minus1 shall be equal to sps_num_subpics_minus1.
[0254] pps_subpic_id_len_minus1 shall be equal to sps_subpic_id_len_minus1.
[0255] pps_subpic_id[i] specifies the sub-picture ID of the i-th sub-picture. The length of the pps_subpic_id[i] syntax element is pps_subpic_id_len_minus1+1 bits.
[0256] For i in the range 0 to sps_num_subpics_minus1 (inclusive), the list SubpicIdList [i] The derivation is as follows:
[0257]
[0258]
[0259] The bitstream conformance requirement is that for any i and j in the range 0 to sps_num_subpics_minus1 (inclusive), when i is less than j, SubpicIdList[i] shall be less than SubpicIdList[j]. ...
[0261] rect_slice_flag equal to 0 specifies that the slices within each slice are in raster scan order, and slice information is not signaled in the PPS. rect_slice_flag equal to 1 specifies that the slices within each slice cover a rectangular area of the picture, and slice information is signaled in the PPS. When not present, rect_slice_flag is inferred to be equal to 1. When subpic_ info_present_flag When equal to 1, the value of rect_slice_flag shall be equal to 1.
[0262] single_slice_per_subpic_flag is equal to 1 and specifies that each sub-picture consists of one and only one rectangular slice. single_slice_per_subpic_flag is equal to 0 and specifies that each sub-picture may include one or more rectangular slices. subpic_info_present_flag When equal to 0, single_slice_per_subpic_flag shall be equal to 0. When single_slice_per_subpic_flag is equal to 1, num_slices_in_pic_minus1 is inferred to be equal to sps_num_subpics_minus1. ...
[0264] 7.3.7.1 General Strip Header Syntax
[0265]
[0266] 7.4.8.1 Common Strip Header Semantics ...
[0268] slice_subpic_id specifies the sub-picture identifier of the sub-picture containing the slice. Length of slice_subpic_id The length is sps_subpic_id_len_minus1+1 bits.
[0269] When not present, the value of slice_subpic_id is inferred to be equal to 0.
[0270] The variable SubPicIdxbe is derived so that SubpicIdList[SubPicIdx] is equal to the value of slice_subpic_id
[0271] slice_address specifies the slice address of the slice. When not present, slice_address is inferred to be equal to 0.
[0272] If rect_slice_flag is equal to 0, the following applies:
[0273] - The strip address is the raster scan slice index.
[0274] - The length of slice_address is Ceil(Log2(NumTilesInPic)) bits.
[0275] -slice_address values should be in the range of 0 to NumTilesInPic-1 (inclusive).
[0276] Otherwise (rect_slice_flag is equal to 1), the following applies:
[0277] - The slice address is the sub-picture level slice index of the slice.
[0278] -The length of slice_address is Max(Ceil(Log2(NumSlicesInSubpic[SubPicIdx])),
[0279] 1) bit.
[0280] -slice_address value should be in the range of 0 to NumSlicesInSubpic[SubPicIdx]-1,
[0281] (Inclusive endpoints).
[0282] The requirements for bitstream conformance are that the following constraints apply:
[0283] - If rect_slice_flag is equal to 0 or subpic_info_present_flag If slice_address is equal to 0, the value of slice_address shall not be equal to the value of slice_address of any other codec slice NAL unit of the same codec picture.
[0284] Otherwise, the pair of slice_subpic_id and slice_address values shall not be equal to the pair of slice_subpic_id and slice_address values of any other codec slice NAL unit of the same codec picture.
[0285] - When rect_slice_flag is equal to 0, the slices of the picture shall be arranged in ascending order of their slice_address values.
[0286] - The shape of a slice of a picture shall be such that each CTU, when decoded, shall have its entire left and entire top borders consisting of the picture boundary or of the boundaries of the previously decoded CTU(s). ...
[0288] Figure 5 is a block diagram illustrating an example video processing system 500 in which the various techniques disclosed herein may be implemented. Various implementations may include some or all of the components of system 500. System 500 may include an input 502 for receiving video content. The video content may be received in a raw or uncompressed format (e.g., 8 or 10 bit multi-component pixel values) or may be in a compressed or encoded format. Input 502 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, a passive optical network (PON), and wireless interfaces such as Wi-Fi or a cellular interface.
[0289] System 500 may include a codec component 504 that can implement various codecs or encoding methods described in this document. The codec component 504 can reduce the average bit rate of the video from the input 502 to the output of the codec component 504 to generate a codec representation of the video. Therefore, codec technology is sometimes referred to as video compression or video code conversion technology. As represented by component 506, the output of the codec component 504 can be stored or sent via a connected communication. Component 508 can use the stored or communicated bitstream (or codec) representation of the video received at the input 502 to generate pixel values or displayable video sent to the display interface 510. The process of generating user-visible video from the bitstream representation is sometimes referred to as video decompression. In addition, although some video processing operations are referred to as "codec" operations or tools, it should be understood that the codec tools or operations are used at the encoder, and the corresponding decoding tools or operations opposite to the results of the codec will be performed by the decoder.
[0290] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB), High-Definition Multimedia Interface (HDMI), or DisplayPort, etc. Examples of storage interfaces include SATA (Serial Advanced Technology Attachment), PCI, IDE interfaces, etc. The technology described in this document may be implemented in various electronic devices, such as mobile phones, laptop computers, smartphones, or other devices capable of performing digital data processing and / or video display.
[0291] Figure 6 6 is a block diagram of a video processing device 600. Device 600 can be used to implement one or more methods described herein. Device 600 can be embodied in a smartphone, tablet computer, computer, Internet of Things (IoT) receiver, etc. Device 600 can include one or more processors 602, one or more memories 604, and video processing hardware 606. Processor(s) 602 can be configured to implement one or more methods described herein. Memory(s) 604 can be used to store data and code used to implement the methods and techniques described herein. Video processing hardware 606 can be used to implement some of the techniques described herein in hardware circuitry.
[0292] Figure 7 is a block diagram illustrating an example video coding system 100 that may utilize the techniques of this disclosure.
[0293] like Figure 7 As shown, the video encoding and decoding system 100 may include a source device 110 and a target device 120. The source device 110 generates encoded video data and may be referred to as a video encoding device. The target device 120 may decode the encoded video data generated by the source device 110 and may be referred to as a video decoding device.
[0294] Source device 110 may include a video source 112 , a video encoder 114 , and an input / output (I / O) interface 116 .
[0295] The video source 112 may include a source such as a video capture device, an interface for receiving video data from a video content provider, and / or a computer graphics system for generating video data, or a combination of such sources. The video data may include one or more pictures. The video encoder 114 encodes the video data from the video source 112 to generate a bitstream. The bitstream may include a sequence of bits that form a codec representation of the video data. The bitstream may include a codec picture and associated data. The codec picture is a codec representation of the picture. The associated data may include a sequence parameter set, a picture parameter set, and other syntax structures. The I / O interface 116 may include a modulator / demodulator (modem) and / or a transmitter. The encoded video data may be sent directly to the target device 120 via the network 130a via the I / O interface 116. The encoded video data may also be stored on a storage medium / server 130b for access by the target device 120.
[0296] Target device 120 may include an I / O interface 126 , a video decoder 124 , and a display device 122 .
[0297] I / O interface 126 may include a receiver and / or a modem. I / O interface 126 may obtain encoded video data from source device 110 or storage medium / server 130b. Video decoder 124 may decode the encoded video data. Display device 122 may display the decoded video data to a user. Display device 122 may be integrated with target device 120 or may be external to target device 120 and configured to interface with an external display device.
[0298] The video encoder 114 and the video decoder 124 may operate according to a video compression standard, such as the High Efficiency Video Codec (HEVC) standard, the Versatile Video Codec (VVM) standard, and other current and / or future standards.
[0299] Figure 8 is a block diagram illustrating an example of a video encoder 200, which may be Figure 7 The video encoder 114 in the system 100 is shown.
[0300] Video encoder 200 may be configured to perform any or all of the techniques of this disclosure. Figure 8 In the example of , video encoder 200 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of video encoder 200. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.
[0301] Functional components of the video encoder 200 may include a segmentation unit 201, a prediction unit 202, a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213, and an entropy coding unit 214. The prediction unit 202 may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra-frame prediction unit 206.
[0302] In other examples, the video encoder 200 may include more, fewer, or different functional components. In an example, the prediction unit 202 may include an intra block copy (IBC) unit. The IBC unit may perform prediction in an IBC mode in which at least one reference picture is a picture in which the current video block is located.
[0303] Furthermore, some components, such as the motion estimation unit 204 and the motion compensation unit 205, may be highly integrated, but for explanation purposes, are not shown in FIG. Figure 8 In the example, they are shown separately.
[0304] The segmentation unit 201 may segment a picture into one or more video blocks. The video encoder 200 and the video decoder 300 may support multiple video block sizes.
[0305] The mode selection unit 203 may select one of the coding modes (intra or inter) (e.g., based on the error result) and provide the resulting intra or inter coded block to the residual generation unit 207 to generate residual block data and to the reconstruction unit 212 to reconstruct the coded block for use as a reference picture. In some examples, the mode selection unit 203 may select a combined intra and inter prediction (CIIP) mode in which prediction is based on an inter prediction signal and an intra prediction signal. In the case of inter prediction, the mode selection unit 203 may also select a resolution of motion vectors for the block (e.g., sub-pixel or integer pixel precision).
[0306] To perform inter-frame prediction on the current video block, the motion estimation unit 204 may generate motion information for the current video block by comparing the current video block with one or more reference frames from the buffer 213. The motion compensation unit 205 may determine a predicted video block for the current video block based on the motion information of pictures other than the picture associated with the current video block from the buffer 213 and decoded samples.
[0307] Motion estimation unit 204 and motion compensation unit 205 may perform different operations on the current video block, eg, depending on whether the current video block is in an I slice, a P slice, or a B slice.
[0308] In some examples, motion estimation unit 204 may perform unidirectional prediction on the current video block, and motion estimation unit 204 may search for a reference video block for the current video block in the reference pictures in list 0 or list 1. Motion estimation unit 204 may then generate a reference index indicating a reference picture in list 0 or list 1 that contains the reference video block and a motion vector indicating a spatial displacement between the current video block and the reference video block. Motion estimation unit 204 may output the reference index, the prediction direction indicator, and the motion vector as motion information for the current video block. Motion compensation unit 205 may generate a predicted video block for the current block based on the reference video block indicated by the motion information for the current video block.
[0309] In other examples, the motion estimation unit 204 may perform bidirectional prediction on the current video block. The motion estimation unit 204 may search for a reference video block for the current video block in the reference pictures in list 0 and may also search for another reference video block for the current video block in the reference pictures in list 1. The motion estimation unit 204 may then generate a reference index and a motion vector, where the reference index indicates the reference pictures in list 0 and list 1 that contain the reference video block, and the motion vector indicates the spatial displacement between the reference video block and the current video block. The motion estimation unit 204 may output the reference index and motion vector for the current video block as motion information for the current video block. The motion compensation unit 205 may generate a predicted video block for the current video block based on the reference video block indicated by the motion information for the current video block.
[0310] In some examples, motion estimation unit 204 may output a complete motion information set for use in the decoding process of a decoder.
[0311] In some examples, motion estimation unit 204 may not output a complete set of motion information for the current video. Instead, motion estimation unit 204 may reference motion information of another video block to signal the motion information for the current video block. For example, motion estimation unit 204 may determine that the motion information for the current video block is sufficiently similar to the motion information for the neighboring video block.
[0312] In one example, motion estimation unit 204 may indicate in a syntax structure associated with the current video block a value that indicates to video decoder 300 that the current video block has the same motion information as another video block.
[0313] In another example, the motion estimation unit 204 may identify another video block and a motion vector difference (MVD) in a syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 300 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[0314] As described above, the video encoder 200 may predictively signal motion vectors.Two examples of predictive signaling techniques that may be implemented by the video encoder 200 include Advanced Motion Vector Prediction (AMVP) and Merge mode signaling.
[0315] The intra-frame prediction unit 206 may perform intra-frame prediction on the current video block. When the intra-frame prediction unit 206 performs intra-frame prediction on the current video block, the intra-frame prediction unit 206 may generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block may include the predicted video block and various syntax elements.
[0316] The residual generation unit 207 can generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) the predicted video block(s) of the current video block from the current video block. The residual data for the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.
[0317] In other examples, there may be no residual data for the current video block, such as in skip mode, and the residual generation unit 207 may not perform a subtraction operation.
[0318] Transform processing unit 208 may generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video block associated with the current video block.
[0319] After transform processing unit 208 generates a transform coefficient video block associated with the current video block, quantization unit 209 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0320] The inverse quantization unit 210 and the inverse transform unit 211 may apply inverse quantization and inverse transform, respectively, to the transform coefficient video block to reconstruct a residual video block from the transform coefficient video block. The reconstruction unit 212 may add the reconstructed residual video block to corresponding samples from one or more prediction video blocks generated by the prediction unit 202 to generate a reconstructed video block associated with the current block for storage in the buffer 213.
[0321] After the reconstruction unit 212 reconstructs the video block, a loop filtering operation may be performed to reduce video block artifacts in the video block.
[0322] The entropy coding unit 214 may receive data from other functional components of the video encoder 200. When the entropy coding unit 214 receives data, the entropy coding unit 214 may perform one or more entropy coding operations to generate entropy-coded data and output a bitstream including the entropy-coded data.
[0323] Figure 9 is a block diagram illustrating an example of a video decoder 300, which may be Figure 7 The video decoder 114 in the system 100 is shown.
[0324] Video decoder 300 may be configured to perform any or all of the techniques of this invention. Figure 8 In the example of FIG, video decoder 300 includes various functional components. The techniques described in this disclosure can be shared among the various components of video decoder 300. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.
[0325] exist Figure 9 In the example of FIG, the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra-frame prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, the video decoder 300 can perform operations generally related to the video encoder 200 ( Figure 8 ) is the decoding pass that is the opposite of the encoding pass described.
[0326] The entropy decoding unit 301 can retrieve an encoded bitstream. The encoded bitstream can include entropy-encoded video data (e.g., encoded video data blocks). The entropy decoding unit 301 can decode the entropy-encoded video data, and from the entropy-decoded video data, the motion compensation unit 302 can determine motion information including motion vectors, motion vector precision, reference picture list index, and other motion information. For example, the motion compensation unit 302 can determine such information by performing AMVP and Merge modes.
[0327] The motion compensation unit 302 may generate a motion compensated block, possibly performing interpolation based on an interpolation filter. An identifier of an interpolation filter to be used with sub-pixel precision may be included in the syntax element.
[0328] Motion compensation unit 302 may calculate interpolated values for sub-integer pixels of a reference block using interpolation filters, such as those used during encoding of the video block by video encoder 200. Motion compensation unit 302 may determine the interpolation filters used by video encoder 200 based on received syntax information and use the interpolation filters to generate a prediction block.
[0329] The motion compensation unit 302 may use some syntax information to determine the size of blocks used to encode frame(s) and / or slice(s) of the coded video sequence, partitioning information describing how each macroblock of a picture of the coded video sequence is partitioned, a mode indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-coded block, and other information used to decode the coded video sequence.
[0330] The intra prediction unit 303 can form a prediction block from spatially neighboring blocks using, for example, an intra prediction mode received in the bitstream. The inverse quantization unit 303 inversely quantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 301. The inverse transform unit 303 applies an inverse transform.
[0331] The reconstruction unit 306 can add the residual block to the corresponding prediction block generated by the motion compensation unit 202 or the intra prediction unit 303 to form a decoded block. If necessary, a deblocking filter can also be applied to the decoded block to remove blocking artifacts. The decoded video block is then stored in a buffer 307, which provides reference blocks for subsequent motion compensation / intra prediction and also produces decoded video for presentation on a display device.
[0332] Figure 10-11 It is shown that it is possible to Figure 5-9 The illustrated embodiment is an example method for implementing the above technical solution.
[0333] Figure 10 A flowchart of an example method 1000 for video processing is shown. The method 1000 includes, at operation 1010, performing conversion between a video including a plurality of pictures including one or more sub-pictures and a bitstream of the video, the bitstream conforming to a format rule that specifies whether, if a current sub-picture in a current picture has a current identifier that is different from an identifier of a previous sub-picture in a previous picture at the same position as the current picture, the current sub-picture is not allowed to refer to a previous sub-picture for inter-frame prediction.
[0334] Figure 11A flow chart of an example method 1100 for video processing is shown. The method 1100 includes, at operation 1110, performing conversion between a video including a video region and a bitstream of video including a plurality of codec layers, the bitstream conforming to a format rule that specifies whether inter-layer prediction (ILP) between video regions in different codec layers of the plurality of codec layers is allowed based on a condition.
[0335] Figure 12 A flow chart of an example method 1200 for video processing is shown. The method 1200 includes, at operation 1210, performing conversion between a video including a current video region and a bitstream of the video including a plurality of codec layers, the bitstream conforming to a format rule that specifies that the bitstream includes an indication of whether inter-layer prediction (ILP) is allowed between the current video region and a video region in a reference layer.
[0336] The following provides a list of preferred solutions for some embodiments.
[0337] 1. A method of video processing, comprising performing conversion between a video comprising a plurality of pictures including one or more sub-pictures and a bitstream of the video, wherein the bitstream conforms to a format rule, the format rule specifying whether a current sub-picture in a current picture is not allowed to refer to a previous sub-picture for inter-frame prediction if the current sub-picture has a current identifier that is different from an identifier of a previous sub-picture in a previous picture at the same position as the current picture.
[0338] 2. The method according to solution 1, wherein the current identifier is a current sub-picture identifier or a current sub-picture position.
[0339] 3. The method according to solution 2, wherein, since the current identifier is different from the identifier of the previous sub-picture, the current sub-picture is not allowed to refer to the previous sub-picture for inter-frame prediction.
[0340] 4. The method according to solution 3, wherein, since the current identifier is different from the identifier of the previous sub-picture, the current sub-picture is encoded and decoded using intra-frame encoding and decoding.
[0341] 5. The method according to solution 3 or 4, wherein the current sub-picture only includes one or more codec slices of an Instantaneous Decoding Refresh (IDR) sub-picture or a Pure Random Access (CRA) sub-picture.
[0342] 6. The method according to solution 3 or 4, wherein the current sub-picture only includes one or more codec slices of an inner random access point (IRAP) sub-picture.
[0343] 7. The method according to solution 3 or 4, wherein the current sub-picture only includes codec slice Network Abstraction Layer (NAL) units with one or more of a predetermined set of NAL unit types.
[0344] 8. The method of solution 7, wherein the predetermined set of NAL unit types includes IDR_W_RADL, IDR_N_LP, and CRA_NUT.
[0345] 9. The method of solution 7, wherein the predetermined set of NAL unit types includes IDR_W_RADL, IDR_N_LP, CRA_NUT, RSV_IRAP_11, and RSV_IRAP_12.
[0346] 10. The method of solution 3 or 4, wherein the bitstream includes a syntax element indicating that a sub-picture is to be treated as a picture.
[0347] 11. The method according to solution 3 or 4, wherein the current sub-picture includes a codec slice network abstraction layer (NAL) unit using one or more of intra prediction, intra block copy (IBC) prediction, and palette mode prediction.
[0348] 12. A method according to solution 2, wherein a first video unit in a current sub-picture references a second video unit in a previous sub-picture, wherein the sub-picture index of the current sub-picture is the same as the sub-picture index of the previous sub-picture, and wherein the sub-picture index is a number assigned to a sub-picture that cannot be changed in a codec layer video sequence (CLVS).
[0349] 13. The method of solution 12, wherein the sub-picture identifier of the current sub-picture is the same as the sub-picture identifier of the previous sub-picture.
[0350] 14. The method of solution 12, wherein the sub-picture identifier of the current sub-picture is different from the sub-picture identifier of the previous sub-picture.
[0351] 15. The method of solution 2, wherein the current sub-picture refers to the previous sub-picture, and since the current sub-picture is identified by a layer identifier value and a sub-picture index or a sub-picture identifier, an indication of the current sub-picture is signaled in the bitstream.
[0352] 16. A method of video processing, comprising performing conversion between a video comprising a video region and a bitstream of the video comprising multiple codec layers, wherein the bitstream complies with a format rule, and wherein the format rule specifies whether inter-layer prediction (ILP) between video regions in different codec layers of the multiple codec layers is allowed based on a condition.
[0353] 17. The method of solution 16, wherein the video region is a sub-picture and wherein inter-layer prediction is allowed.
[0354] 18. The method of solution 17, wherein two sub-pictures in different codec layers include the same sub-picture index value and different sub-picture identifier values.
[0355] 19. The method of solution 18, wherein the condition dictates that the two layers are associated with different view order indices or different view order identifier values.
[0356] 20. The method according to solution 17, wherein the first sub-picture and the second sub-picture in different codec layers are in the same position or have a reasonable height or width because the first sub-picture and the second sub-picture have the same sub-picture index.
[0357] 21. The method according to solution 17, wherein the first sub-picture and the second sub-picture in different codec layers are in the same position or have a reasonable height or width because the first sub-picture refers to the second sub-picture.
[0358] 22. A method of video processing, comprising performing a conversion between a video comprising a current video region and a bitstream of the video comprising multiple codec layers, wherein the bitstream conforms to a format rule, and wherein the format rule specifies that the bitstream includes an indication of whether inter-layer prediction (ILP) is allowed between the current video region and a video region in a reference layer.
[0359] 23. The method of solution 22, wherein the video region is a sub-picture.
[0360] 24. The method of solution 23, wherein the indication is signaled in a video parameter set (VPS), a decoding parameter set (DPS), a sequence parameter set (SPS), a picture parameter set (PPS), an adaptation parameter set (APS), a sequence header, or a picture header.
[0361] 25. The method of solution 23, wherein the video region in the reference layer includes at least one sample that is co-located with a sample of the current video region.
[0362] 26. The method of solution 23, wherein the video area in the reference layer is outside the co-located area of the current video area.
[0363] 27. The method of solution 23, wherein the indication is signaled in one or more Supplemental Enhancement Information (SEI) messages.
[0364] 28. The method of solution 23, wherein the indication is signaled regardless of whether the reference layer includes multiple sub-pictures.
[0365] 29. A method according to solution 23, wherein the reference layer includes multiple sub-pictures, and wherein a notification indication is signaled regardless of whether the partitioning of the picture into multiple sub-pictures is aligned with the current picture so that each sub-picture in the reference layer is co-located with the corresponding sub-picture in the current picture.
[0366] 30. The method of any one of solutions 1 to 29, wherein converting comprises decoding the video from a bitstream.
[0367] 31. A method according to any of solutions 1 to 29, wherein converting includes encoding the video into a bitstream.
[0368] 32. A method for storing a bitstream representing a video to a computer-readable recording medium, comprising generating a bitstream from a video according to any one or more of the methods described in Solutions 1 to 29; and writing the bitstream to the computer-readable recording medium.
[0369] 33. A video processing device comprising a processor configured to implement the method described in any one or more of solutions 1 to 32.
[0370] 34. A computer-readable medium having instructions stored thereon, which, when executed, cause a processor to implement the method of one or more of solutions 1 to 32.
[0371] 35. A computer-readable medium storing a bitstream generated according to any one or more of solutions 1 to 32.
[0372] 36. A video processing device for storing a bitstream, wherein the video processing device is configured to implement any one or more of the methods described in solutions 1 to 32.
[0373] Another list of preferred solutions for some embodiments is provided next.
[0374] P1. A method of video processing, comprising performing a conversion between a picture of a video and a codec representation of the video, wherein a plurality of sub-pictures in the picture are included in the codec representation as fields whose bit width depends on a value of the number of sub-pictures.
[0375] P2. The method according to solution P1, wherein the field uses a codeword to represent the number of sub-pictures.
[0376] P3. The method of solution P2, wherein the codeword comprises a Golomb codeword.
[0377] P4. A method according to any of solutions P1 to P3, wherein the value of the number of sub-pictures is restricted to be less than or equal to an integer number of codec treeblocks that fit into the picture.
[0378] P5. The method according to any of the solutions P1 to P4, wherein the field depends on a codec level associated with the codec representation.
[0379] P6. A method of video processing, comprising performing a conversion between a video region of a video and a codec representation of the video, wherein the codec representation conforms to a format rule, wherein the format rule provides for omitting a syntax element indicating a sub-picture identifier because the video region does not include any sub-pictures.
[0380] P7. A method according to solution P6, wherein the codec representation includes a field with a value of 0 indicating that the video region does not include any sub-pictures.
[0381] P8. A method of video processing, comprising performing a conversion between a video region of a video and a codec representation of the video, wherein the codec representation conforms to a format rule, wherein the format rule provides for omitting identifiers of sub-pictures in the video region at a video region header level in the codec representation.
[0382] P9. The method of solution P8, wherein the codec indicates that the sub-pictures are numerically identified according to the order in which they are listed in the video region header.
[0383] P10. A method for video processing, comprising performing conversion between a video region of a video and a codec representation of the video, wherein the codec representation conforms to a format rule, wherein the format rule specifies an identifier of a sub-picture in the video region and / or the length of the identifier of the sub-picture at a sequence parameter set level or a picture parameter set level.
[0384] P11. The method of solution P10, wherein the length is included at the picture parameter set level.
[0385] P12. A method of video processing, comprising performing a conversion between a video region of a video and a codec representation of the video, wherein the codec representation conforms to a format rule, wherein the format rule specifies a field included in the codec representation at a video sequence level to indicate whether a sub-picture identifier length field is included in the codec representation at the video sequence level.
[0386] P13. The method of solution P12, wherein the format rule dictates that the field be set to "1" if another field in the codec representation indicates a length identifier for a video region included in the codec representation.
[0387] P14. A method of video processing, comprising performing a conversion between a video region of a video and a codec representation of the video, wherein the codec representation complies with format rules, and wherein the format rules provide for including in the codec representation an indication of whether the video region can be used as a reference picture.
[0388] P15. The method of solution P14, wherein the indication comprises a layer ID and an index or ID value associated with the video region.
[0389] P16. A method of video processing, comprising performing a conversion between a video region of a video and a codec representation of the video, wherein the codec representation conforms to a format rule, and wherein the format rule provides for including in the codec representation an indication of whether the video region can use inter-layer prediction (ILP) from a plurality of sample values associated with the video region of a reference layer.
[0390] P17. The method of solution P16, wherein the indication is included at a sequence level, a picture level, or a video level.
[0391] P18. The method of solution P16, wherein the video region of the reference layer includes at least one sample that is co-located with a sample within the video region.
[0392] P19. The method of solution P16, wherein the indication is included in one or more Supplemental Enhancement Information (SEI) messages.
[0393] P20. The method according to any of the preceding claims, wherein the video region comprises a sub-picture of the video.
[0394] P21. A method according to any of the preceding claims, wherein converting comprises parsing and decoding a codec representation to generate a video.
[0395] P22. A method according to any of the preceding claims, wherein converting comprises encoding the video to generate a codec representation.
[0396] P23. A video decoding device comprising a processor configured to implement the method described in one or more of solutions P1 to P22.
[0397] P24. A video encoding device comprising a processor configured to implement the method described in one or more of solutions P1 to P22.
[0398] P25. A computer program product having computer code stored thereon, which, when executed by a processor, causes the processor to implement the method of any one of solutions P1 to P22.
[0399] In some embodiments, the bitstream generated according to the above method may be stored on a computer-readable medium.
[0400] The disclosed and other solutions, examples, embodiments, modules, and functional operations described in this document may be implemented in digital electronic circuitry, or in computer software, firmware, or hardware, including the structures disclosed in this document and their structural equivalents, or in a combination of one or more thereof. The disclosed and other embodiments may be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer-readable medium, for execution by or to control the operation of a data processing apparatus. The computer-readable medium may be a machine-readable storage device, a machine-readable storage substrate, a memory device, a combination of matter that effects a machine-readable propagated signal, or a combination of one or more thereof. The term "data processing apparatus" encompasses all apparatus, devices, and machines for processing data, including, for example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, an apparatus may include code that creates an operating environment for the computer program in question, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more thereof. A propagated signal is an artificially generated signal, such as a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to a suitable receiver device.
[0401] A computer program (also referred to as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language file), in a single file dedicated to the program in question, or in multiple collaborating files (e.g., files that store one or more modules, subroutines, or code portions). A computer program can be deployed to run on one computer or on multiple computers located in one location or distributed across multiple locations and interconnected by a communications network.
[0402] The processes and logic flows described herein can be performed by one or more programmable processors running one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by, and apparatus can be implemented as, special purpose logic circuitry, such as an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit).
[0403] By way of example, processors suitable for running computer programs include general-purpose and special-purpose microprocessors, as well as any one or more processors of any type of digital computer. Typically, a processor will receive instructions and data from a read-only memory or a random access memory, or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include, or be operatively coupled to, one or more mass storage devices for storing data, such as magnetic, magneto-optical, or optical disks, to receive data from or transfer data to the mass storage devices, or both. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and storage devices, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD ROM and DVD-ROM disks. The processor and memory may be supplemented by, or incorporated into, special-purpose logic circuitry.
[0404] Although this patent document contains many details, they should not be interpreted as limitations on the scope of any subject matter or of what is claimed, but rather as descriptions of features unique to particular embodiments of particular technologies. Certain features described in the context of separate embodiments in this patent document may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented separately or in any suitable subcombination in multiple embodiments. Furthermore, although features may be described above as working in certain combinations, and even initially claimed as such, one or more features from a claimed combination may in some cases be deleted from the combination, and a claimed combination may be directed to a subcombination or variant of a subcombination.
[0405] Similarly, while operations are described in a particular order in the drawings, this should not be understood as requiring that these operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, in order to achieve the desired results. Furthermore, the separation of various system components in the embodiments described in this patent document should not be understood as requiring such separation in all embodiments.
[0406] Only a few implementations and examples are described, and other implementations, enhancements, and variations can be made based on what is described and illustrated in this patent document.
Claims
1. A video processing method, comprising: performing conversion between a video comprising a plurality of pictures including one or more sub-pictures and a bitstream of said video, wherein the bitstream complies with a format rule, wherein the format rule stipulates that, if a current sub-picture in a current picture has a current identifier that is different from an identifier of a previous sub-picture in a previous picture at the same position as the current sub-picture, the current sub-picture is not allowed to refer to the previous sub-picture for inter-frame prediction, and the current sub-picture is encoded and decoded using intra-frame coding, The current identifier is a current sub-picture identifier or a current sub-picture position.
2. The method according to claim 1, wherein The current sub-picture only includes one or more codec slices of an Instantaneous Decoding Refresh (IDR) sub-picture or a Pure Random Access (CRA) sub-picture.
3. The method according to claim 1, wherein The current sub-picture only includes one or more codec slices of an intra random access point (IRAP) sub-picture.
4. The method according to claim 1, wherein The current sub-picture includes only codec slice Network Abstraction Layer (NAL) units having one or more of a predetermined set of NAL unit types.
5. The method according to claim 4, wherein The predetermined NAL unit type set includes IDR_W_RADL, IDR_N_LP and CRA_NUT.
6. The method according to claim 4, wherein: The predetermined NAL unit type set includes IDR_W_RADL, IDR_N_LP, CRA_NUT, RSV_IRAP_11 and RSV_IRAP_12.
7. The method according to claim 1, wherein The bitstream includes a syntax element indicating that the sub-picture is to be treated as a picture.
8. The method according to claim 1, wherein The current sub-picture includes a codec slice Network Abstraction Layer (NAL) unit using one or more of intra prediction, intra block copy (IBC) prediction, and palette mode prediction.
9. The method according to claim 1, wherein: The first video unit in the current sub-picture refers to the second video unit in the previous sub-picture, wherein the sub-picture index of the current sub-picture is the same as the sub-picture index of the previous sub-picture, and wherein the sub-picture index is a number assigned to a sub-picture that cannot be changed in a codec layer video sequence (CLVS).
10. The method according to claim 9, wherein: The sub-picture identifier of the current sub-picture is the same as the sub-picture identifier of the previous sub-picture.
11. The method according to claim 9, wherein The sub-picture identifier of the current sub-picture is different from the sub-picture identifier of the previous sub-picture.
12. The method according to claim 1, wherein The current sub-picture refers to the previous sub-picture, and since the current sub-picture is identified by a layer identifier value and a sub-picture index or a sub-picture identifier, an indication of the current sub-picture is signaled in the bitstream.
13. The method according to claim 1, wherein The format rule specifies whether inter-layer prediction (ILP) between video regions in different codec layers of a plurality of codec layers is allowed based on a condition.
14. The method according to claim 13, wherein The video region is a sub-picture, and the inter-layer prediction is allowed therein.
15. The method according to claim 14, wherein The two sub-pictures in the different coding and decoding layers include the same sub-picture index value and different sub-picture identifier values.
16. The method according to claim 15, wherein The condition stipulates that two different codec layers are associated with different view order indices or different view order identifier values.
17. The method according to claim 14, wherein: The first sub-picture and the second sub-picture in the different coding and decoding layers are in the same position or have a reasonable height or width because the first sub-picture and the second sub-picture have the same sub-picture index.
18. The method according to claim 14, wherein The first sub-picture and the second sub-picture in the different coding and decoding layers are in the same position or have a reasonable height or width because the first sub-picture refers to the second sub-picture.
19. The method according to claim 1, wherein The format rules specify that the bitstream includes an indication of whether inter-layer prediction (ILP) between a current video region and a video region in a reference layer is allowed.
20. The method according to claim 19, wherein The video region is a sub-picture.
21. The method according to claim 20, wherein The indication is signaled in a video parameter set (VPS), a decoding parameter set (DPS), a sequence parameter set (SPS), a picture parameter set (PPS), an adaptation parameter set (APS), a sequence header, or a picture header.
22. The method according to claim 20, wherein The video area in the reference layer includes at least one sample that is co-located with a sample in the current video area.
23. The method according to claim 20, wherein The video area in the reference layer is outside a co-located area of the current video area.
24. The method according to claim 20, wherein The indication is signaled in one or more Supplemental Enhancement Information (SEI) messages.
25. The method according to claim 20, wherein The indication is signaled regardless of whether the reference layer includes multiple sub-pictures.
26. The method according to claim 20, wherein The reference layer comprises a plurality of sub-pictures, and wherein the indication is signaled regardless of whether partitioning of a picture into the plurality of sub-pictures is aligned with a current picture such that each sub-picture in the reference layer is co-located with a corresponding sub-picture in the current picture.
27. The method according to any one of claims 1 to 26, wherein The converting includes decoding the video from the bitstream.
28. The method according to any one of claims 1 to 26, wherein The converting includes encoding the video into the bitstream.
29. A method of storing a bitstream representing a video to a computer-readable recording medium, comprising: Generating a bitstream from a video according to the method of any one of claims 1 to 28; as well as The bit stream is written to the computer-readable recording medium.
30. A video processing device comprising a processor configured to implement the method according to any one of claims 1 to 14.
31. A computer readable medium having stored thereon instructions which, when executed, cause a processor to implement the method of one of claims 1 to 28.
32. A computer-readable medium storing a bit stream generated by the video processing apparatus according to claim 30.
Citation Information
Patent Citations
Error mitigation in sub-picture bitstream based viewport dependent video coding
WO2019195035A1