Using picture-level strip indexing in video encoding and decoding
By introducing regular video processing methods, optimizing sub-picture segmentation and tool activation and disabling, the efficiency and bandwidth utilization problems of multi-layer video encoding and decoding in the prior art are solved, and a more efficient video encoding and decoding effect is achieved.
Patent Information
- Application Number
- CN202180016696.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-02-24
- Filing Date
- 2021-02-23
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2041-02-23
AI Technical Summary
When handling multi-layer video encoding and decoding, it is difficult to effectively support the enablement and disabling of sub-picture, brightness mapping and chromaticity scaling (LMCS) tools, and the application of adaptive motion vector difference accuracy (AMVR) in affine inter-frame mode is not flexible enough, resulting in poor encoding and decoding efficiency and bandwidth utilization.
By introducing regular video processing methods, including deriving the use of slice indexes, strip indexes and control information, optimizing the segmentation of sub-pictures and the activation and disabling of tools during video encoding and decoding, especially the applications of LMCS and AMVR, to achieve more efficient video encoding and decoding.
It improves the efficiency and bandwidth utilization of video encoding and decoding, supports the flexibility of multi-layer video encoding and decoding, and improves the quality and performance of video processing.
Smart Images

Figure CN115244922B_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application is an application entering the Chinese national phase of International Patent Application No. PCT / US2021 / 019217 filed on February 23, 2021, which claims priority to U.S. Provisional Patent Application No. US 62 / 980,963 filed on February 24, 2020. The entire disclosure of the above application is incorporated by reference as part of the disclosure of this application. Technical Field
[0003] This patent document relates to image and video encoding and decoding. Background Art
[0004] Digital video accounts for the largest use of bandwidth on the Internet and other digital communications networks. As the number of connected user devices capable of receiving and displaying video increases, bandwidth demand for digital video usage is expected to continue to grow. Summary of the Invention
[0005] This document discloses techniques that can be used by video encoders and decoders for processing a codec representation of a video using control information useful for decoding the codec representation.
[0006] In one example aspect, a video processing method is disclosed. The method includes performing conversion between a video comprising one or more video pictures and a bitstream of the video, wherein each video picture comprises one or more slices, the one or more slices comprising one or more slice columns, wherein the bitstream conforms to a format rule, and wherein the format rule provides for derivation of a slice column index for each codec tree unit (CTU) column of a slice of the video picture.
[0007] In another example aspect, a video processing method is disclosed. The method includes performing conversion between a video comprising one or more video pictures and a bitstream of the video, wherein each video picture comprises one or more slices, the one or more slices comprising one or more slice columns, wherein the bitstream conforms to a format rule, and wherein the format rule provides for derivation of a slice row index for each codec tree unit (CTU) row of a slice of the video picture.
[0008] In another example aspect, a video processing method is disclosed. The method includes performing conversion between a video comprising at least one video picture and a bitstream of the video according to a rule, wherein the at least one video picture comprises one or more slices and one or more sub-pictures, and wherein the rule specifies an order of slice indexes indicating one or more slices in the at least one video picture in response to a syntax element associated with the at least one video picture indicating whether each sub-picture of the at least one video picture comprises a single slice.
[0009] In another example aspect, a video processing method is disclosed. The method includes performing conversion between video units of a video region of a video and a bitstream of the video, wherein the bitstream conforms to a format rule, wherein the format rule specifies that first control information of a first level of the video region in the bitstream controls whether a second level of the video unit in the bitstream includes second control information, wherein the second level is less than the first level, wherein the first control information and the second control information include information regarding whether or how a luma mapping and chroma scaling (LMCS) tool is applied to the video unit, and wherein the LMCS tool includes using chroma residual scaling (CRS) or a luma reformation process (RP) for the conversion.
[0010] In another example aspect, a video processing method is disclosed. The method includes performing conversion between a video and a bitstream of the video according to a rule, the rule specifying that a luma mapping and chroma scaling (LMCS) tool is enabled when a first syntax element in a reference sequence parameter set indicates that the LMCS tool is enabled, the rule specifying that the LMCS tool is not used when the first syntax element indicates that the LMCS tool is disabled, the rule specifying that when a second syntax element in the bitstream indicates that the LMCS tool is enabled at a picture header level of the video, the LMCS tool is enabled for all slices associated with a picture header of a video picture, the rule specifying that when the second syntax element indicates that the LMCS tool is disabled at the picture header level of the video, the LMCS tool is not used for all slices associated with the picture header, the rule specifying that when a third syntax element selectively included in the bitstream indicates that the LMCS tool is enabled at a slice header level of the video, the LMCS tool is used for a current slice associated with a slice header of the video picture, and the rule specifying that when the third syntax element indicates that the LMCS tool is disabled at the slice header level of the video, the LMCS tool is not used for the current slice.
[0011] In another example aspect, a video processing method is disclosed. The method includes performing conversion between a video including one or more video pictures and a bitstream of the video according to a rule, wherein the rule specifies whether adaptive motion vector difference precision (AMVR) is used in motion vector coding for affine inter mode based on a syntax element selectively included in a reference sequence parameter set (SPS) indicating whether AMVR is enabled, wherein the rule specifies that when the syntax element indicates that AMVR is disabled, AMVR is not used for motion vector coding for affine inter mode, and wherein the rule specifies that when the syntax element is not included in the SPS, AMVR is inferred not to be used for motion vector coding for affine inter mode.
[0012] In another example aspect, a video processing method is disclosed. The method includes performing conversion between a video including a video picture and a bitstream of the video according to a rule, wherein the video picture includes a sub-picture, a slice, and a slice, and wherein the rule specifies that because the sub-picture includes a slice divided from a slice, the conversion is performed by avoiding using a number of slices of the video picture to calculate the height of the sub-picture.
[0013] In another example aspect, a video processing method is disclosed that includes performing conversion between a video comprising a video picture and a bitstream of the video, wherein the bitstream indicates a height of a sub-picture of the video picture calculated based on a number of codec tree units (CTUs) of the video picture.
[0014] In another example aspect, a video processing method is disclosed that includes: determining, based on a rule, whether a height of a sub-picture of a video picture of the video is less than a height of a slice line of the video picture; and using the determination to convert between the video and a bitstream of the video.
[0015] In another example aspect, a video processing method is disclosed. The method includes performing conversion between a video comprising one or more video pictures and a codec representation of the video, wherein each video picture comprises one or more slices, wherein the codec representation conforms to a format rule; wherein the format rule specifies first information signaled in the codec representation and second information derived from the codec representation, wherein at least the first information or the second information is related to a row index or a column index of the one or more slices.
[0016] In another example aspect, a video processing method is disclosed. The method includes performing conversion between video units of a video region of a video and a codec representation of the video, wherein the codec representation conforms to a format rule; wherein the format rule specifies that first control information at the video region controls whether second control information is included at the video unit level; wherein the first control information and / or the second control information includes information about luma mapping and chroma scaling (LMCS) or chroma residual scaling (CRS) or a reforming process (RP) for the conversion.
[0017] In yet another exemplary aspect, a video encoder apparatus is disclosed. The video encoder includes a processor configured to implement the above method.
[0018] In yet another exemplary aspect, a video decoder apparatus is disclosed. The video decoder includes a processor configured to implement the above method.
[0019] In yet another exemplary aspect, a computer-readable medium having code stored thereon is disclosed. The code is in the form of processor-executable code embodying one of the methods described herein.
[0020] These and other features are described in this document. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 An example of raster scan stripe partitioning of a picture is shown, where the picture is divided into 12 slices and 3 raster scan stripes.
[0022] Figure 2 An example of rectangular strip partitioning of a picture is shown, where the picture is divided into 24 slices (6 slice columns and 4 slice rows) and 9 rectangular strips.
[0023] Figure 3 An example of a picture partitioned into slices and rectangular strips is shown, where the picture is divided into 4 slices (2 slice columns and 2 slice rows) and 4 rectangular strips.
[0024] Figure 4 A picture partitioned into 15 slices, 24 slices, and 24 sub-pictures is shown.
[0025] Figure 5 is a block diagram of an example video processing system.
[0026] Figure 6 A block diagram of a video processing device.
[0027] Figure 7 A flowchart of an example method for video processing is provided.
[0028] Figure 8 FIG. 1 is a block diagram illustrating a video encoding and decoding system according to some embodiments of the present disclosure.
[0029] Figure 9 A block diagram illustrating an encoder according to some embodiments of the present disclosure.
[0030] Figure 10 FIG. 4 is a block diagram illustrating a decoder according to some embodiments of the present disclosure.
[0031] Figures 11 to 19 A flowchart of an example method for video processing is provided. DETAILED DESCRIPTION
[0032] Section headings are used in this document for ease of understanding and do not limit the applicability of the techniques and embodiments disclosed in each section to only that section. In addition, the use of H.266 technical terms in some descriptions is solely for ease of understanding and is not intended to limit the scope of the disclosed techniques. Therefore, the techniques described herein are also applicable to other video codec protocols and designs. In this document, editorial changes to text relative to the current draft of the VVC specification are indicated by left and right double brackets (e.g., [[ ]]) to indicate deleted text between the brackets, and by bold italic text to indicate added text.
[0033] 1. Summary
[0034] This document relates to video codec technology. Specifically, it is about support for sub-pictures, LMCS, and AMVR. Aspects regarding sub-pictures include deriving the number of slice rows and slice columns included in a sub-picture, and deriving a list of raster scan CTU addresses of the CTUs included in a slice when each sub-picture contains only one slice. The LMCS aspect is about signaling notification of enabling LMCS at different levels. The AMVR aspect is about the semantics of the sps_affine_amvr_enabled_flag. These ideas can be applied alone or in various combinations to any video codec standard or non-standard video codec that supports single and / or multi-layer video codecs, such as the Versatile Video Codec (VVC) under development.
[0035] 2. Abbreviation
[0036] ALF Adaptive Loop Filter
[0037] AMVR Adaptive Motion Vector Interpolation Precision
[0038] APS Adaptive Parameter Set
[0039] AU Access Unit
[0040] AUD Access Unit Delimiter
[0041] AVC Advanced Video Codec
[0042] CLVS codec layer video sequence
[0043] CPB codec picture buffer
[0044] CRA Clean Random Access
[0045] CTU Codec Tree Unit
[0046] CVS codec video sequence
[0047] DPB decoded picture buffer
[0048] DPS decoding parameter set
[0049] EOB End of bitstream
[0050] EOS sequence ends
[0051] GDR Gradual Decode Refresh
[0052] HEVC High-Efficiency Video Codec
[0053] HRD Hypothesized Reference Decoder
[0054] IDR Instantaneous Decode Refresh
[0055] JEM Joint Exploration Model
[0056] LMCS Luma Mapping with Chroma Scaling
[0057] MCTS motion constraint set
[0058] NAL Network Abstraction Layer
[0059] OLS output layer set
[0060] PH Image Header
[0061] PPS picture parameter set
[0062] PTL grades, tiers, and levels
[0063] PU picture unit
[0064] RBSP Raw Byte Sequence Payload
[0065] SEI Supplemental Enhancement Information
[0066] SPS sequence parameter set
[0067] SVC Scalable Video Codec
[0068] VCL video codec layer
[0069] VPS Video Parameter Set
[0070] VTM VVC test model
[0071] VUI Video Availability Information
[0072] VVC multifunctional video codec
[0073] 3. Preliminary Discussion
[0074] Video codec standards are primarily developed through the well-known ITU-T and ISO / IEC standards. ITU-T developed H.261 and H.263, while ISO / IEC developed MPEG-1 and MPEG-4 Vision. The two organizations jointly developed the H.262 / MPEG-2 Video standard, the H.264 / MPEG-4 Advanced Video Codec (AVC) standard, and the H.265 / HEVC standard. Starting with H.262, video codec standards have been based on a hybrid video codec architecture that utilizes temporal prediction plus transform coding. To explore future video codec technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Team (JVET) in 2015. Since then, JVET has adopted many new approaches and incorporated them into reference software called the Joint Exploration Model (JEM). JVET meetings are held simultaneously every quarter, with the goal of achieving a 50% bitrate reduction compared to HEVC for the new codec standard. The new video codec standard was officially named Versatile Video Codec (VVC) at the JVET meeting in April 2018, and the first version of the VVC Test Model (VTM) was released at that time. Due to the continuous efforts to standardize VVC, new codec technologies are adopted into the VVC standard at each JVET meeting. The VVC working draft and test model VTM are updated after each meeting. The VVC project is now aiming for technical completion (FDIS) at the July 2020 meeting.
[0075] 3.1. Image Segmentation Scheme in HEVC
[0076] HEVC includes four different picture partitioning schemes, namely normal slices, dependent slices, tiles, and wavefront parallel processing (WPP), which can be applied to maximum transmission unit (MTU) size matching, parallel processing, and reduced end-to-end delay.
[0077] Regular slices are similar to those in H.264 / AVC. Each regular slice is encapsulated in its own NAL unit, and intra-picture prediction (intra-sample prediction, motion information prediction, codec mode prediction) and entropy codec dependencies are disabled across slice boundaries. Therefore, regular slices can be reconstructed independently of other regular slices in the same picture (although interdependencies may still exist due to loop filtering operations).
[0078] Regular slices are the only tool available for parallelization, and are also available in H.264 / AVC in a nearly identical form. Regular slice-based parallelization does not require much inter-processor or inter-core communication (except for inter-processor or inter-core data sharing for motion compensation when decoding predictive codec pictures, which is typically much more than inter-processor or inter-core data sharing due to intra-picture prediction). However, for the same reasons, using regular slices can incur significant codec overhead due to the bit cost of the slice header and the loss of prediction across slice boundaries. Furthermore, due to their intra-picture independence and the fact that each regular slice is encapsulated in its own NAL unit, regular slices (compared to the other tools mentioned below) can also serve as a key mechanism for bitstream segmentation to match MTU size requirements. In many cases, the goals of parallelization and MTU size matching conflict with respect to slice layout within a picture. The realization of such situations led to the development of the parallelization tools mentioned below.
[0079] Dependent slices have short slice headers and allow the bitstream to be split at treeblock boundaries without breaking any intra-picture prediction. Essentially, dependent slices split a regular slice into multiple NAL units, reducing end-to-end latency by allowing part of a regular slice to be sent before coding of the entire regular slice is complete.
[0080] In WPP, a picture is partitioned into a single row of codec treeblocks (CTBs). This allows entropy decoding and prediction to use data from CTBs in other partitions. Parallel processing is enabled by parallel decoding of CTB rows, where the start of decoding a CTB row is delayed by two CTBs to ensure that data associated with CTBs above and to the right of the main CTB is available before the main CTB being decoded. Using this staggered start (which appears as a wavefront when graphically represented), as many processors / cores as there are CTB rows in a picture can be parallelized. Because intra-picture prediction is allowed between adjacent treeblock rows within a picture, the inter-processor / inter-core communication required to enable intra-picture prediction can be extensive. WPP partitioning does not result in the generation of additional NAL units compared to when WPP partitioning was not applied, so WPP is not a tool for MTU size matching. However, if MTU size matching is required, regular slices can be used with WPP, but with some codec overhead.
[0081] Slices define the horizontal and vertical boundaries that divide an image into slice columns and slice rows. Slice columns extend from the top to the bottom of the image. Similarly, slice rows extend from the left to the right. The number of slices in an image can be simply derived by multiplying the number of slice columns by the number of slice rows.
[0082] Before decoding the top left CTB of the next slice in the order of the slice raster scan of the picture, the scan order of the CTBs is changed to the local scan order within the slice (in the order of the CTB raster scan of the slice). Similar to regular slices, slices break intra-picture prediction dependencies and entropy decoding dependencies. However, they do not need to be included in separate NAL units (same as WPP in this respect); therefore, slices cannot be used for MTU size matching. Each slice can be processed by one processor / core, and in the case where a slice spans multiple slices, the inter-processor / inter-core communication required for intra-picture prediction between processing units decoding adjacent slices is limited to transmitting a shared slice header and loop filtering related to sharing of reconstruction samples and metadata. When more than one slice or WPP segment is included in a slice, the entry point byte offset of each slice or WPP segment in the slice except the first is signaled in the slice header.
[0083] For simplicity, HEVC specifies restrictions on the application of four different picture partitioning schemes. For most profiles specified in HEVC, a given codec video sequence cannot contain both slices and wavefronts. For each slice and slice, one or both of the following conditions must be met: 1) all codec tree blocks in a slice belong to the same slice; 2) all codec tree blocks in a slice belong to the same slice. Finally, a wavefront segment contains exactly one CTB row, and when using WPP, if a slice starts within a CTB row, it must end within the same CTB row.
[0084] The most recent revision of HEVC is specified in the JCT-VC output document JCTVC-AC1005, J. Boyce, A. Ramasubramonian, R. Skupin, G.J. Sullivan, A. Tourapis, and Y.-Kwang (eds.), "HEVC Additional Supplemental Enhancement Information (Draft 4)", published on October 24, 2017, at: http: / / phenix.intevry.fr / jct / doc_end_user / documents- / 29_Macau / wg11 / JCTVC-AC1005-v2.zip. Within this revision, HEVC specifies three MCT-related SEI messages: the Time-Domain MCTS SEI message, the MCTS Extraction Information Set SEI message, and the MCTS Extraction Information Nesting SEI message.
[0085] The temporal MCTS SEI message indicates the presence of an MCTS in the bitstream and signals the MCTS. For each MCTS, motion vectors are restricted to pointing to full sample positions within the MCTS and fractional sample positions that only require full sample positions within the MCTS for interpolation, and motion vector candidates predicted from temporal motion vectors derived from blocks outside the MCTS are not allowed. In this way, each MCTS can be independently decoded in the absence of slices not included in the MCTS.
[0086] The MCTS extraction information set SEI message provides supplementary information that can be used for MCTS sub-bitstream extraction (specified as part of the semantics of the SEI message) to generate a conforming bitstream for the MCTS set. The information consists of multiple extraction information sets, each of which defines multiple MCTS sets and contains RBSP bytes that replace the VPS, SPS, and PPS to be used during the MCTS sub-bitstream extraction process. When extracting a sub-bitstream according to the MCTS sub-bitstream extraction process, the parameter sets (VPS, SPS, and PPS) need to be rewritten or replaced, and the slice headers need to be slightly updated because one or all of the slice addresses associated with the syntax elements (including first_slice_segment_in_pic_flag and slice_segment_address) usually need to have different values.
[0087] 3.2. Image Segmentation in VVC
[0088] In VVC, a picture is divided into one or more slice rows and one or more slice columns. A slice is a sequence of CTUs covering a rectangular area of the picture. The CTUs in a slice are scanned in raster scan order within the slice.
[0089] A slice consists of an integer number of complete slices or an integer number of consecutive complete CTU rows within a slice of a picture.
[0090] Two striping modes are supported: raster scan striping mode and rectangular striping mode. In raster scan striping mode, a stripe contains a complete sequence of slices in a slice raster scan of a picture. In rectangular striping mode, a stripe contains multiple complete slices that together form a rectangular region of a picture, or multiple consecutive complete CTU rows that together form a slice of a rectangular region of a picture. Slices within a rectangular stripe are scanned in slice raster scan order within the rectangular region corresponding to the stripe.
[0091] A sub-picture consists of one or more strips that together cover a rectangular area of the picture.
[0092] Figure 1 An example of raster scan striping partitioning of a picture is shown, where the picture is partitioned into 12 slices and 3 raster scan strips.
[0093] Figure 2 An example of rectangular strip partitioning of a picture is shown, where the picture is partitioned into 24 slices (6 slice columns and 4 slice rows) and 9 rectangular slices.
[0094] Figure 3 An example of a picture partitioned into slices and rectangular strips is shown, where the picture is divided into 4 slices (2 slice columns and 2 slice rows) and 4 rectangular strips.
[0095] Figure 4 An example of sub-picture partitioning of a picture is shown, where the picture is partitioned into 18 slices, 12 slices on the left (each covering a strip of 4x4 CTUs) and 6 slices on the right (each covering 2 vertically stacked strips of 2x2 CTUs), resulting in a total of 24 slices and 24 sub-pictures of different dimensions (each slice is a sub-picture).
[0096] 3.3. Signaling Notification of SPS / PPS / Picture Header / Slice Header in VVC (e.g. JVET-Q2001-vC)
[0097] 7.3.2.3 Sequence Parameter Set RBSP Syntax
[0098]
[0099]
[0100]
[0101]
[0102]
[0103]
[0104]
[0105]
[0106]
[0107] 7.3.2.4 Picture Parameter Set RBSP Syntax
[0108]
[0109]
[0110]
[0111]
[0112]
[0113]
[0114] 7.3.2.7 Picture header structure syntax
[0115]
[0116]
[0117]
[0118]
[0119]
[0120]
[0121] 7.3.7.1 General Strip Header Syntax
[0122]
[0123]
[0124]
[0125]
[0126]
[0127] 3.4. Specifications for slices, stripes, and sub-pictures in JVET-Q2001-vC
[0128] 3 definitions
[0129] The index of the slice into the slice list in the picture, in the order they are signaled in the PPS when rect_slice_flag is equal to 1.
[0130] The index of the slice into the slice list in the sub-picture, in the order they are signaled in the PPS when rect_slice_flag is equal to 1.
[0131] 6.5.1 CTB Raster Scanning, Slice Scanning, and Sub-Picture Scanning Process
[0132] The variable NumTileColumns specifies the number of tile columns, and the list colWidth[i], where i ranges from 0 to NumTileColumn-1 (inclusive), specifies the width of the i-th tile column in CTB units, as derived below:
[0133]
[0134] The variable NumTileRows specifies the number of tile rows, and the list RowHeight[j] of j in the range 0 to NumTileRows-1 (inclusive) specifies the height of the j-th tile row in units of CTB, as derived below:
[0135]
[0136]
[0137] The variable NumTilesInPic is set equal to NumTileColumns*NumTileRows.
[0138] The list tileColBd[i] for i in the range 0 to NumTileColumns (inclusive) specifies the location of the i-th tile column boundary in units of CTBs, as derived below:
[0139] for(tileColBd[0]=0,i=0;i <NumTileColumns;i++)
[0140] tileColBd[i+1]=tileColBd[i]+colWidth[i] (25)
[0141] NOTE 1 – The size of the array tileColBd[] is one greater than the actual number of tile columns in the derivation of CtbToTileColBd[].
[0142] The list tileRowBd[j] of j in the range 0 to NumTileRows (inclusive) specifies the location of the j-th tile row boundary in units of CTB, derived as follows:
[0143] for(tileRowBd[0]=0,j=0;j <NumTileRows;j++)
[0144] tileRowBd[j+1]=tileRowBd[j]+RowHeight[j] (26)
[0145] NOTE 2 – The size of the array tileRowBd[] in the above derivation is one larger than the actual number of tile rows in the derivation of CtbToTileRowBd[].
[0146] The list CtbToTileColBd[ctbAddrX] of ctbAddrX ranging from 0 to PicWidthInCtbsY (inclusive) specifies the translation from the horizontal CTB address to the left tile column boundary in units of CTB, as derived below:
[0147]
[0148] NOTE 3 – The size of the array CtbToTileColBd[] in the above derivation is one larger than the actual number of picture widths in the CTB in the derivation of slice_data() signalling.
[0149] The list CtbToTileRowBd[ctbAddrY] of ctbAddrY in the range from 0 to PicHeightInCtbsY (inclusive) specifies the translation from the vertical CTB address to the top tile column boundary in CTB units, derived as follows:
[0150]
[0151] NOTE 4 – The size of the array CtbToTileRowBd[] in the above derivation is one greater than the actual number of picture heights in the CTB signalled in slice_data().
[0152] For rectangular slices, the list NumCtusInSlice[i] for i ranging from 0 to num_slices_in_pic_minus1 (inclusive) specifies the number of CTUs in the i-th slice, the list SliceTopLeftTileIdx[i] for i ranging from 0 to num_slices_in_pic_minus1 (inclusive) specifies the index of the top left slice of the slice, and the matrix of i ranging from 0 to num_slices_in_pic_minus1 (inclusive) and j ranging from 0 to NumCtusInSlice[i]–1 (inclusive) specifies the picture raster scan address of the j-th CTB within the i-th slice, derived as follows:
[0153]
[0154]
[0155] The function AddCtbsToSlice(sliceIdx, startX, stopX, startY, stopY) is specified as follows:
[0156]
[0157] It is a bitstream conformance requirement that the value of NumCtusInSlice[i] for i in the range 0 to num_slices_in_pic_minus1 (inclusive) shall be greater than 0. Furthermore, it is a bitstream conformance requirement that the matrix CtbAddrInSlice[i][j] for i in the range 0 to num_slices_in_pic_minus1 (inclusive) and j in the range 0 to NumCtusInSlice[i]-1 (inclusive) shall include all CTB addresses in the range 0 to PicSizeInCtbsY-1 once and only once.
[0158] The list CtbToSubpicIdx[ctbAddrRs] of ctbAddrRs in the range 0 to PicSizeInCtbsY-1 (inclusive) specifies the conversion from CTB addresses in the picture raster scan to subpicture indices, derived as follows:
[0159]
[0160] The list NumSlicesInSubpic[i] specifies the number of rectangular strips in the i-th sub-picture, which is derived as follows:
[0161]
[0162] …
[0163] 7.3.4.3 Picture Parameter Set RBSP Semantics
[0164] …
[0165] Equal to 1 specifies that sub-picture ID mapping is signaled in PPS. subpic_id_mapping_in_pps_flag equal to 0 specifies that sub-picture ID mapping is not signaled in PPS. If subpic_id_mapping_explicitly_signaled_flag is 0 or subpic_id_mapping_in_sps_flag is equal to 1, the value of subpic_id_mapping_in_pps_flag shall be equal to 0. Otherwise (subpic_id_mapping_explicitly_signaled_flag is equal to 1 and subpic_id_mapping_in_sps_flag is equal to 0), the value of subpic_id_mapping_in_pps_flag shall be equal to 1.
[0166] Should be equal to sps_num_subpics_minus1.
[0167] Should be equal to sps_subpic_id_len_minus1.
[0168] pps_subpic_id[i] specifies the sub-picture ID of the i-th sub-picture. The length of the pps_subpic_id[i] syntax element is pps_subpic_id_len_minus1+1 bits.
[0169] For each value of i in the range from 0 to sps_num_subpics_minus1, the variable SubpicIdVal[i] is derived as follows:
[0170]
[0171] The requirement for bitstream conformance is that the following two constraints apply:
[0172] --For any two different values of i and j in the range of 0 to sps_num_subpics_minus1 (inclusive), SubpicIdVal[i] shall not be equal to SubpicIdVal[j].
[0173] --When the current picture is not the first picture of the CLVS, for each value of i in the range of 0 to sps_num_subpics_minus1 (inclusive), if the value of SubpicIdVal[i] is not equal to the value of SubpicIdVal[i] of the previous picture in the same layer in decoding order, the nal_unit_type of all codec slice NAL units of the subpictures in the current picture with subpicture index i shall be equal to the specified value in the range of IDR_W_RADL to CRA_NUT (inclusive).
[0174] Equal to 1 specifies that picture partitioning is not applied to each picture of the referenced PPS. no_pic_partition_flag equal to 0 specifies that each picture of the referenced PPS may be partitioned into more than one slice or slice.
[0175] A bitstream conformance requirement is that the value of no_pic_partition_flag shall be the same for all PPSs referenced by a codec picture within a CLVS.
[0176] A bitstream conformance requirement is that when the value of sps_num_subpics_minus1+1 is greater than 1, the value of no_pic_partition_flag shall not be equal to 1.
[0177] Specify the luma codec treeblock size for each CTU by adding 5. pps_log2_ctu_size_minus5 should be equal to sps_log2_ctu_size_minus5.
[0178] Add 1 to specify the number of explicitly provided tile column widths. The value of num_exp_tile_columns_minus1 should be in the range of 0 to PicWidthInCtbsY–1, inclusive. When no_pic_partition_flag is equal to 1, the value of num_exp_tile_columns_minus1 is inferred to be equal to 0.
[0179] Add 1 to specify the number of explicitly provided tile row heights. The value of num_exp_tile_rows_minus1 should be in the range of 0 to PicHeightInCtbsY–1, inclusive. When no_pic_partition_flag is equal to 1, the value of num_tile_rows_minus1 is inferred to be equal to 0.
[0180] [i] plus 1 specifies the width of the i-th tile column in units of CTB, and the range of i is 0 to num_exp_tile_columns_minus1-1 (including the end value).
[0181] tile_column_width_minus1[num_exp_tile_columns_minus1] is used to derive the width of tile columns with indices greater than or equal to num_exp_tile_columns_minus1, as specified in clause 6.5.1. The value of tile_column_width_minus1[i] shall be in the range of 0 to PicWidthInCtbsY–1, inclusive. When not present, the value of tile_column_width_minus1[0] is inferred to be equal to PicWidthInCtbsY-1.
[0182] [i] plus 1 specifies the height of the i-th tile row in CTB units, and the range of i is 0 to num_exp_tile_rows_minus1–1 (including the end value).
[0183] tile_row_height_minus1[num_exp_tile_rows_minus1] is used to derive the height of tile rows with indices greater than or equal to num_exp_tile_rows_minus1, as specified in clause 6.5.1. The values of tile_row_height_minus1[i] shall be in the range of 0 to PicHeightInCtbsY–1, inclusive. When not present, the value of tile_row_height_minus1[0] is inferred to be equal to PicHeightInCtbsY-1.
[0184] When equal to 0, the slices within each slice are in raster scan order and the slice information is not signaled in the PPS. When rect_slice_flag is equal to 1, the slices within each slice cover a rectangular area of the picture and the slice information is signaled in the PPS. When not present, rect_slice_flag is inferred to be equal to 1. When subpic_info_present_flag is equal to 1, the value of rect_slice_flag shall be equal to 1.
[0185] When equal to 1, specifies that each sub-picture consists of one and only one rectangular slice. When single_slice_per_subpic_flag is equal to 0, specifies that each sub-picture can consist of one or more rectangular slices. When single_slice_per_subpic_flag is equal to 1, num_slices_in_pic_minus1 is inferred to be equal to sps_num_subpics_minus1. When not present, the value of single_slice_per_subpic_flag is inferred to be equal to 0.
[0186] Specifies the number of rectangular slices in each picture of the referenced PPS, plus 1. The value of num_slices_in_pic_minus1 shall be in the range of 0 to MaxSlicesPerPicture–1, inclusive, where MaxSlicesPerPicture is specified in Annex A. When no_pic_partition_flag is equal to 1, the value of num_slices_in_pic_minus1 is inferred to be equal to 0.
[0187] Equal to 0 specifies that tile_idx_delta values are not present in the PPS and that all rectangular slices in pictures referencing the PPS are specified in raster order according to the procedure defined in clause 6.5.1.
[0188] tile_idx_delta_present_flag equal to 1 specifies that tile_idx_delta values may be present in the PPS, and all rectangular slices in pictures referencing the PPS are specified in the order indicated by the values of tile_idx_delta. When not present, the value of tile_idx_delta_present_flag is inferred to be equal to 0.
[0189] [i] plus 1 specifies the width of the i-th rectangular strip in units of tile columns. The value of slice_width_in_tiles_minus1[i] should be in the range of 0 to NumTileColumns–1 (inclusive).
[0190] When slice_width_in_tiles_minus1[i] is not present, the following applies:
[0191] --If NumTileColumns is equal to 1, the value of slice_width_in_tiles_minus1[i] is inferred to be equal to 0.
[0192] -- Otherwise, the value of slice_width_in_tiles_minus1[i] is inferred as specified in clause 6.5.1.
[0193] [i] plus 1 specifies the height of the i-th rectangular strip in units of tile rows. The value of slice_height_in_tiles_minus1[i] should be in the range of 0 to NumTileRows–1 (inclusive).
[0194] When slice_height_in_tiles_minus1[i] is not present, the following applies:
[0195] --If NumTileRows is equal to 1, or tile_idx_delta_present_flag is equal to 0, and tileIdx%NumTileColumns is greater than 0, then the value of slice_height_in_tiles_minus1[i] is inferred to be equal to 0.
[0196] Otherwise (NumTileRows is not equal to 1, and tile_idx_delta_present_flag is equal to 1 or tileIdx%NumTileColumns is equal to 0), when tile_idx_delta_present_flag is equal to 1 or tileIdx%NumTileColumns is equal to 0, the value of slice_height_in_tiles_minus1[i] is inferred to be equal to slice_height_in_tiles_minus1[i-1].
[0197] [i] specifies the number of explicitly provided strip heights in the current slice containing multiple rectangular strips. The value of num_exp_slices_in_tile[i] should be in the range of 0 to RowHeight[tileY]–1 (inclusive), where tileY is the index of the strip row containing the i-th strip. When not present, the value of num_exp_slices_in_tile[i] is inferred to be equal to 0. When num_exp_slices_in_tile[i] is equal to 0, the value of the variable NumSlicesInTile[i] is derived to be equal to 1.
[0198] [j] plus 1 specifies the height of the j-th rectangular slice in the current slice in units of CTU rows. The value of exp_slice_height_in_ctus_minus1[j] shall be in the range of 0 to RowHeight[tileY]–1 (inclusive), where tileY is the slice row index of the current slice.
[0199] When num_exp_slices_in_tile[i] is greater than 0, the variables NumSlicesInTile[i] and SliceHeightInCtusMinus1[i+k] for k in the range 0 to NumSlicesInTile[i]–1 are derived as follows:
[0200]
[0201] [i] specifies the difference between the tile index of the first tile in the i-th rectangular strip and the tile index of the first tile in the (i+1)-th rectangular strip. The value of tile_idx_delta[i] should be in the range of -NumTilesInPic+1 to NumTilesInPic–1 (inclusive). When not present, the value of tile_idx_delta[i] is inferred to be equal to 0. When present, the value of tile_idx_delta[i] shall not be equal to 0.
[0202] …
[0203] 7.4.2.4.5 Order of VCL NAL units and their association with coded and decoded pictures
[0204] The order of VCL NAL units within a codec picture is constrained as follows:
[0205] --For any two codec slice NAL units A and B of a codec picture, let subpicIdxA and subpicIdxB be their sub-picture level index values, and sliceAddrA and sliceddrB be their slice_address values.
[0206] -- Codec slice NAL unit A shall precede codec slice NAL unit B when any of the following conditions is true:
[0207] –subpicIdxA is less than subpicIdxB.
[0208] --subpicidxa is equal to subpicIdxB, and sliceAddrA is less than sliceAddrB.
[0209] 7.4.8.1 General Strip Header Semantics
[0210] The variable CuQpDeltaVal, which specifies the difference between the luma quantization parameter and its prediction for the codec unit containing cu_qp_delta_abs, is set equal to 0. The variables CuQpOffsetCb, CuQpOffsetCr, and CuQpOffsetCbCr, which specify the values to be used in determining the respective values of the Qp′Cb, Qp′Cr, and Qp′CbCr quantization parameters for the codec unit containing cu_chroma_qp_offset_flag, are all set equal to 0.
[0211] Equal to 1 specifies that the PH syntax structure is present in the slice header. picture_header_in_slice_header_flag equal to 0 specifies that the PH syntax structure is not present in the slice header.
[0212] A bitstream conformance requirement is that the value of picture_header_in_slice_header_flag shall be the same in all codec slices in a CLVS.
[0213] When picture_header_in_slice_header_flag is equal to 1 for a codec slice, the bitstream conformance requirement is that VCL NAL units with nal_unit_type equal to PH_NUT shall not appear in the CLVS.
[0214] When picture_header_in_slice_header_flag is equal to 0, all codec slices in the current picture shall have picture_header_in_slice_header_flag equal to 0, and the current PU shall have a PH NAL unit.
[0215] Specifies the sub-picture ID of the sub-picture containing the slice. If slice_subpic_id exists, the value of the variable CurrSubpicIdx is derived so that SubpicIdVal[CurrSubpicIdx] is equal to slice_subpic_id. Otherwise (slice_subpic_id does not exist), CurrSubpicIdx is derived to be equal to 0. The length of slice_subpic_id is sps_subpic_id_len_minus1+1 bits.
[0216] Specifies the slice address of the slice. When not present, the value of slice_address is inferred to be equal to 0. When rect_slice_flag is equal to 1 and NumSlicesInSubpic[CurrSubpicIdx] is equal to 1, the value of slice_address is inferred to be equal to 0.
[0217] If rect_slice_flag is equal to 0, the following applies:
[0218] - The strip address is the raster scan slice index.
[0219] - The length of slice_address is Ceil(Log2(NumTilesInPic)) bits.
[0220] -slice_address values should be in the range of 0 to NumTilesInPic–1 (inclusive).
[0221] Otherwise (rect_slice_flag is equal to 1), the following applies:
[0222] - The slice address is the sub-picture level slice index of the slice.
[0223] - The length of slice_address is Ceil(Log2(NumSlicesInSubpic[CurrSubpicIdx])) bits.
[0224] -slice_address values should be in the range of 0 to NumSlicesInSubpic[CurrSubpicIdx]–1 (inclusive).
[0225] The requirements for bitstream conformance are that the following constraints apply:
[0226] - If rect_slice_flag is equal to 0 or subpic_info_present_flag is equal to 0, the value of slice_address shall not be equal to the value of slice_address of any other codec slice NAL unit of the same codec picture.
[0227] Otherwise, the pair of slice_subpic_id and slice_address values shall not be equal to the pair of slice_subpic_id and slice_address values of any other codec slice NAL unit of the same codec picture.
[0228] - The shape of the slices of a picture shall be such that each CTU, when decoded, has its entire left and top borders consisting of picture boundaries or of the boundaries of previously decoded CTU(s).
[0229] [i] can be equal to 1 or 0. Decoders conforming to this version of this specification shall ignore the value of sh_extra_bit[i]. Its value does not affect the conformance of the decoder to the profile specified in this version of the specification.
[0230] Add 1 (if present) to specify the number of tiles in the stripe. The value of num_tiles_in_slice_minus1 should be in the range of 0 to NumTilesInPic–1, inclusive.
[0231] The variable NumCtusInCurrSlice specifies the number of CTUs in the current slice, and the list CtbAddrInCurrSlice[i], with i ranging from 0 to NumCtusInCurrSlice–1 (inclusive), specifies the picture raster scan address of the i-th CTB within the slice, derived as follows:
[0232]
[0233]
[0234] The variables SubpicLeftBoundaryPos, SubpicTopBoundaryPos, SubpicRightBoundaryPos, and SubpicBotBoundaryPos are derived as follows:
[0235]
[0236] …
[0237] 3.5. Luminance Mapping with Chroma Scaling (LMCS)
[0238] LMCS consists of two aspects: luma mapping (the transformation process, denoted by RP) and luma-dependent chroma residual scaling (CRS). For luma signals, the LMCS mode operates on two domains: a first domain that is the original domain and a second domain that is a shaped domain that maps luma samples to specific values according to a shaping model. In addition, for chroma signals, residual scaling can be applied, where the scaling factor is derived from the luma samples.
[0239] The relevant syntax elements and semantics in SPS, picture header (PH) and slice header (SH) are described as follows:
[0240] Syntax Table
[0241] 7.3.2.3 Sequence Parameter Set RBSP Syntax
[0242]
[0243] 7.3.2.7 Picture header structure syntax
[0244]
[0245]
[0246] 7.3.7 Strip Header Syntax
[0247] 7.3.7.1 General Strip Header Syntax
[0248]
[0249] Semantics
[0250] sps_lmcs_enabled_flag is equal to 0 and specifies that luma mapping with chroma scaling is not used in CLVS.
[0251] ph_lmcs_enabled_flag is inferred to be equal to 0 when it is not present.
[0252] Specifies the adaptation_parameter_set_id of the LMCS APS referenced by the slice associated with the PH. The TemporalId of the APS NAL unit with aps_params_type equal to LMCS_APS and adaptation_parameter_set_id equal to ph_lmcs_aps_id shall be less than or equal to the TemporalId of the picture associated with the PH.
[0253] 1 specifies that chroma residual scaling is enabled for all slices associated with the PH. ph_chroma_residual_scale_flag equal to 0 specifies that chroma residual scaling may be disabled for one, multiple, or all slices associated with the PH. When ph_chroma_residual_scale_flag is not present, it is inferred to be equal to 0.
[0254] slice_lmcs_enabled_flag is inferred to be equal to 0 when slice_lmcs_enabled_flag is not present.
[0255] 3.6. Adaptive Motion Vector Differentiation (AMVR) Precision for Affine Codec Blocks
[0256] Affine AMVR is a codec tool that allows the affine inter-frame codec block to send MV differences with different precisions, such as 1 / 4 luma sample (default, amvr_flag is set to 0), 1 / 16 luma sample, and 1 luma sample.
[0257] The relevant syntax elements and semantics in SPS are described as follows:
[0258] Syntax Table
[0259] 7.3.2.3 Sequence Parameter Set RBSP Syntax
[0260]
[0261]
[0262] Semantics
[0263] When equal to 1, specifies the use of adaptive motion vector difference resolution in motion vector coding for affine inter mode. When equal to 0, specifies the use of adaptive motion vector difference resolution in motion vector coding for affine inter mode. When not present, the value of sps_affine_amvr_enabled_flag is inferred to be equal to 0.
[0264] D.7 Sub-picture level information SEI message
[0265] D.7.1 Sub-picture level information SEI message syntax
[0266]
[0267] D.7.2 Sub-picture level information SEI message semantics
[0268] The sub-picture level information SEI message contains information about the level to which a sub-picture sequence in a bitstream conforms when the extracted bitstream containing the sub-picture sequence is tested for conformance according to Annex A.
[0269] When a sub-picture level information SEI message is present in any picture of a CLVS, the sub-picture level information SEI message shall be present in the first picture of the CLVS. Sub-picture level information SEI messages continue in decoding order from the current picture to the current layer until the end of the CLVS. All sub-picture level information SEI messages applicable to the same CLVS shall have the same content. A sub-picture sequence consists of all sub-pictures with the same sub-picture index value within the CLVS.
[0270] A bitstream conformance requirement is that, when the sub-picture level information SEI message of CLVS is present, the value of subpic_treated_as_pic_flag[i] shall be equal to 1 for every value of i in the range of 0 to sps_num_subpics_minus1, inclusive.
[0271] Plus 1 specifies the number of reference levels signaled for each sps_num_subpics_minus1+1 sub-picture.
[0272] sli_cbr_constraint_flag equal to 1 specifies that the HSS operates in constant bitrate (CBR) mode for decoding the sub-bitstream generated by any sub-picture of the extracted bitstream according to clause C.7 using the HRD of any CPB specification in the extracted sub-bitstream.
[0273] 1 specifies that the syntax element ref_level_fraction_minus1[i] is present. explicit_fraction_present_flag equal to 0 specifies that the syntax element ref_level_fraction_minus1[i] is not present.
[0274] Specifies the number of subpictures in the picture of the CLVS, plus 1. When present, the value of sli_num_subpics_minus1 shall be equal to the value of sps_num_subpics_minus1 in the SPS referenced by the picture in the CLVS.
[0275] Should be equal to 0.
[0276] [i] indicates the level to which each sub-picture conforms as specified in Annex A. The bitstream shall not contain values of ref_level_idc other than those specified in Annex A. Other values of ref_level_idc[i] are reserved for future use by ITU-T | ISO / IEC. A requirement for bitstream conformance is that, for any value of k greater than i, the value of ref_level_idc[i] shall be less than or equal to ref_level_idc[k].
[0277] [i][j] plus 1 specifies the score of the level constraint associated with ref_level_idc[i] to which the j-th sub-picture complies, as specified in clause A.4.1.
[0278] The variable SubpicSizeY[j] is set equal to (subpic_width_minus1[j]+1)*CtbSizeY*(subpic_height_minus1[j]+1)*CtbSizeY.
[0279] When not present, the value of ref_level_fraction_minus1[i][j] is inferred to be equal to Ceil(256*SubpicSizeY[j]÷PicSizeInSamplesY*MaxLumaPs(general_level_idc)÷MaxLumaPs(ref_level_idc[i])-1.
[0280] The variable RefLevelFraction[i][j] is set equal to ref_level_fraction_minus1[i][j]+1.
[0281] The variables SubpicNumTileCols[j] and SubpicNumTileRows[j] are derived as follows:
[0282]
[0283]
[0284] The variables SubpicCpbSizeVcl[i][j] and SubpicCpbSizeNal[i][j] are derived as follows:
[0285] SubpicCpbSizeVcl[i][j]=Floor(CpbVclFactor*MaxCPB*RefLevelFraction[i][j]÷256) (D.6)
[0286] SubpicCpbSizeNal[i][j]=Floor(CpbNalFactor*MaxCPB*RefLevelFraction[i][j]÷256) (D.7)
[0287] where MaxCPB is derived from ref_level_idc[i] as specified in clause A.4.2.
[0288] The variables SubpicBitRateVcl[i][j] and SubpicBitRateNal[i][j] are derived as follows:
[0289] SubpicBitRateVcl[i][j]=Floor(CpbVclFactor*MaxBR*RefLevelFraction[i][j]÷256) (D.8)
[0290] SubpicBitRateNal[i][j]=Floor(CpbNalFactor*MaxBR*RefLevelFraction[i][j]÷256) (D.9) Where MaxBR is derived from ref_level_idc[i] as specified in clause A.4.2.
[0291] NOTE 1 – When a sub-picture is extracted, the resulting bitstream has a CpbSize greater than or equal to SubpicCpbSizeVcl[i][j] and SubpicCpbSizeNal[i][j] (indicated in the SPS or inferred), and a bitrate greater than or equal to SubpicBitRateVcl[i][j] and SubpicBitRateNal[i][j] (indicated in the SPS or inferred).
[0292] The bitstream conformance requirement is that a bitstream resulting from extracting the j-th sub-picture in the range of 0 to sps_num_subpics_minus1 (inclusive), and conforming to the profile with general_tier_flag equal to 0 and level equal to ref_level_idc[i] for i in the range of 0 to num_ref_level_minus1 (inclusive), shall obey the following constraints for each bitstream conformance test specified in Annex C:
[0293] - Ceil(256*SubpicSizeY[j]÷RefLevelFraction[i][j]) shall be less than or equal to MaxLumaPs, where MaxLumaPs is specified in Table A.1 for level ref_level_idc[i].
[0294] -The value of Ceil(256*(subpic_width_minus1[j]+1)*CtbSizeY÷RefLevelFraction[i][j]) should be less than or equal to Sqrt(MaxLumaPs*8).
[0295] -The value of Ceil(256*(subpic_height_minus1[j]+1)*CtbSizeY÷RefLevelFraction[i][j]) should be less than or equal to Sqrt(MaxLumaPs*8).
[0296] - The value of SubpicNumTileCols[j] shall be less than or equal to MaxTileCols, and the value of SubpicNumTileRows[j] shall be less than or equal to MaxTileRows, where MaxTileCols and MaxTileRows are specified in Table A.1 for level ref_level_idc[i].
[0297] - The value of SubpicNumTileCols[j]*SubpicNumTileRows[j] shall be less than or equal to MaxTileCols*MaxTileRows*RefLevelFraction[i][j], where MaxTileCols and MaxTileRows are specified in Table A.1 for level ref_level_idc[i].
[0298] - For the value of SubpicSizeInSamplesY of AU 0, the sum of the NumBytesInNalUnit variables of AU 0 corresponding to the j-th sub-picture shall be less than or equal to FormatCapabilityFactor*(Max(SubpicSizeY[j],fR*MaxLumaSr*RefLevelFraction[i][j]÷256)+MaxLumaSr*(AuCpbRemovalTime[0]-AuNominalRemovalTime[0])*RefLevelFraction[i][j])÷(256*MinCr), where MaxLumaSr and FormatCapabilityFactor are the values specified in Tables A.2 and A.3, respectively, for AU 0 at level ref_level_idc[i], and MinCr is derived as described in A.4.2.
[0299] - The sum of the NumBytesInNalUnit variables for AU n (where n is greater than 0) corresponding to the j-th sub-picture shall be less than or equal to FormatCapabilityFactor*MaxLumaSr*(AuCpbRemovalTime[n]-AuCpbRemovalTime[n-1])*RefLevelFraction[i][j]÷(256*MinCr), where MaxLumaSr and FormatCapabilityFactor are the values specified in Table A.2 and Table A.3, respectively, for AU n at level ref_level_idc[i], and MinCr is derived as described in A.4.2.
[0300] For any subpicture set containing one or more subpictures and consisting of a number of subpictures in the subpicture index SubpicSetIndices list and the subpicture set NumSubpicsInSet, derive the level information of the subpicture set.
[0301] The variable SubpicSetAccLevelFraction[i] for the total level fraction relative to the reference level ref_level_idc[i], and the variables SubpicSetCpbSizeVcl[i], SubpicSetCpbSizeNal[i], SubpicSetBitRateVcl[i] and SubpicSetBitRateNal[i] for the subpicture set are derived as follows:
[0302]
[0303] The value of the subpicture set sequence level indicator SubpicSetLevelIdc is derived as follows:
[0304]
[0305] Among them, MaxTileCols and MaxTileRows are specified for ref_level_idc[i] in Table A.1.
[0306] A subpicture set bitstream conforming to a profile with general_tier_flag equal to 0 and level equal to SubpicSetLevelIdc shall obey the following constraints for each bitstream conformance test specified in Appendix C:
[0307] - For VCL HRD parameters, SubpicSetCpbSizeVcl[i] shall be less than or equal to CpbVclFactor*MaxCPB, where CpbVclFactor is specified in Table A.3 and MaxCPB is specified in Table A.1 in units of CpbVclFactor bits.
[0308] - For NAL HRD parameters, SubpicSetCpbSizeNal[i] shall be less than or equal to CpbNalFactor*MaxCPB, where CpbNalFactor is specified in Table A.3 and MaxCPB is specified in Table A.1 in units of CpbNalFactor bits.
[0309] - For VCL HRD parameters, SubpicSetBitRateVcl[i] shall be less than or equal to CpbVclFactor*MaxBR, where CpbVclFactor is specified in Table A.3 and MaxBR is specified in Table A.1 in units of CpbVclFactor bits.
[0310] - For NAL HRD parameters, SubpicSetBitRateNal[i] shall be less than or equal to CpbNalFactor*MaxCR, where CpbNalFactor is specified in Table A.3 and MaxBR is specified in Table A.1 in units of CpbNalFactor bits.
[0311] NOTE 2 – When a sub-picture set is extracted, the resulting bitstream has a CpbSize greater than or equal to SubpicSetCpbSizeVcl[i][j] and SubpicSetCpbSizeNal[i][j] (indicated in the SPS or inferred), and a bitrate greater than or equal to SubpicSetBitRateVcl[i][j] and SubpicSetBitRateNal[i][j] (indicated in the SPS or inferred).
[0312] 4. Technical problems solved by the disclosed technical solutions
[0313] The existing design of sub-images and LMCS in VVC has the following problems:
[0314] 1) The derivation of the list SubpicNumTileRows[] (specifying the number of tile rows included in a sub-picture) in equation D.5 is incorrect because the index value idx in CtbToTileRowBd[idx] in the equation may be greater than the maximum allowed value. In addition, the deviation between SubpicNumTileRows[] and SubpicNumTileCols[] (specifying the number of tile columns included in a sub-picture) uses CTU-based operations, which is unnecessarily complex.
[0315] 2) When single_slice_per_subpic_flag is equal to 1, the derivation of the array CtbAddrInSlice in Equation 29 is incorrect because the values of the raster scan CTB addresses in the array for each slice need to be in the decoding order of the CTU, rather than the raster scan order of the CTU.
[0316] 3) LMCS signaling is inefficient. When ph_lmcs_enabled_flag is equal to 1, LMCS will be enabled for all slices of the picture in most cases. However, in the current VVC design, when LMCS is enabled for all slices of the picture, not only does ph_lmcs_enabled_flag equal to 1, but slice_lmcs_enabled_flag with a value of 1 also needs to be signaled for each slice.
[0317] a.when When true, The semantics of conflict with the motivation of signaling slice-level LMCS flag. When true, it means that all slices should enable LMCS. Therefore, there is no need to signal the LMCS enable flag in the slice header.
[0318] b. Furthermore, when the picture header signals that LMCS is enabled, LMCS is typically enabled for all slices. Control of LMCS in the slice header is primarily intended to handle corner cases. Therefore, if the PH LMCS flag is true and the SH LMCS flag is always signaled, this may result in unnecessary bits being signaled for normal user scenarios.
[0319] 4) The semantics of the SPS affine AMVR flag is incorrect because for each affine inter-coded CU, affine AMVR can be enabled or disabled.
[0320] 5. Examples of Techniques and Embodiments
[0321] To solve the above problems and some other problems not mentioned, the following methods are disclosed. These items should be considered as examples to explain the general concept and should not be interpreted in a narrow way. In addition, these items can be applied alone or in any combination.
[0322] Related to the sub-images that solve the first and second problems
[0323] 1. One or more of the following methods are disclosed:
[0324] a. Derives the slice column index for each CTU column of the picture.
[0325] b. The number of slice columns included in a sub-picture is derived based on the slice column indexes of the leftmost and / or rightmost CTUs included in the sub-picture.
[0326] c. Derive the slice row index for each CTU row of the picture.
[0327] d. The number of slice rows included in a sub-picture is deduced based on the slice row indices of the top and / or bottom CTUs included in the sub-picture.
[0328] e. The term picture-level stripe index is defined as follows:
[0329] Indexes into the slice list of the picture for the slices defined when rect_slice_flag is equal to 1, in the order in which the slices are signaled in the PPS when single_slice_per_subpic_flag is equal to 0, or in the order of increasing sub-picture index of the sub-pictures corresponding to the slice when single_slice_per_subpic_flag is equal to 1.
[0330] f. In one example, when a sub-picture contains slices split from a tile, the height of the sub-picture cannot be calculated from the tile.
[0331] g. In one example, the height of a sub-picture can be calculated based on CTUs instead of slices.
[0332] h. Determine whether the height of the sub-image is less than one slice row.
[0333] i. In one example, when a sub-picture includes only CTUs from one slice row and when the top CTU in the sub-picture is not the top CTU of the slice row or the bottom CTU in the sub-picture is not the bottom CTU of the slice row, whether the height of the sub-picture is less than one slice row is derived as true.
[0334] ii. When it is indicated that each sub-picture contains only one slice and the height of the sub-picture is less than one slice row, for each slice of the picture with picture-level slice index i, the value of CtbAddrInSlice[i][j] for j in the range from 0 to the number of CTUs in the slice minus 1, inclusive, is derived as the picture raster scan CTU address of the j-th CTU in the CTU raster scan of the sub-picture.
[0335] iii. In one example, whether the height of the sub-picture is less than one slice row is derived as true when the distance between the top CTU in the sub-picture and the bottom CTU in the sub-picture is less than the height of a slice according to CTUs.
[0336] iv. When it is indicated that each sub-picture contains only one slice and the height of the sub-picture is greater than or equal to one slice row, for each slice of the picture with picture-level slice index i, the value of CtbAddrInSlice[i][j] for j in the range from 0 to the number of CTUs in the slice minus 1, inclusive, is derived as the picture raster scan CTU address of the j-th CTU, with the order of the CTUs being as follows:
[0337] 1) CTUs in different slices in a sub-picture are ordered such that the first CTU in the first slice with a smaller slice index value precedes the second CTU in the second slice with a larger slice index value.
[0338] 2) The CTUs within a slice in a sub-picture are sorted in the CTU raster scan of the slice.
[0339] Related to LMCS for solving the third problem (including sub-problems)
[0340] 2. Two-level control of LMCS is introduced (which includes two aspects: luma mapping (transformation process, denoted by RP) and luma-dependent chroma residual scaling (CRS)), where higher-level (e.g., picture-level) and lower-level (e.g., slice-level) control are used, and the presence of lower-level control information depends on the high-level control information. In addition, the following also applies:
[0341] a. In the first example, one or more of the following sub-items apply:
[0342] i. The first indicator (e.g., ) can be signaled at a higher level (e.g., in the picture header (PH)) to specify how to enable LMCS at a lower level that is not a binary value.
[0343] 1) In one example, when the first indicator is equal to X (e.g., X=2), it specifies that LMCS is enabled for all stripes associated with the PH; when the first indicator is equal to Y (Y!=X) (e.g., Y=1), it specifies that LMCS is enabled for one or more but not all stripes associated with the PH; when the first indicator is equal to Z (Z!=X and Z!=Y) (e.g., Z=0), it specifies that LMCS is disabled for all stripes associated with the PH.
[0344] a) Alternatively, furthermore, when the first indicator is not present, the value of the indicator is inferred to be equal to a default value, such as Z.
[0345] 2) In one example, when the first indicator is equal to X (e.g., X=2), it specifies that LMCS is disabled for all stripes associated with the PH; when the first indicator is equal to Y (Y!=X) (e.g., Y=1), it specifies that LMCS is disabled for one or more but not all stripes associated with the PH; when the first indicator is equal to Z (Z!=X and Z!=Y) (e.g., Z=0), it specifies that LMCS is enabled for all stripes associated with the PH.
[0346] a) Alternatively, furthermore, when the first indicator is not present, the value of the indicator is inferred to be equal to a default value, such as X.
[0347] 3) Alternatively, in addition, the first indicator may be based on the LMCS enable flag in the sequence level (e.g. ) value is conditionally signaled.
[0348] 4) Alternatively, in addition, the first indicator can be encoded and decoded using u(v), or u(2) or ue(v).
[0349] 5) Alternatively, furthermore, the first indicator may be encoded or decoded using a truncated unary code.
[0350] 6) Alternatively, additionally, by the stripe and / or CS enable flags (e.g., ) using LMCS APS information (e.g., ) can be signaled under a conditional check of the value of the first indicator.
[0351] ii. A second indicator (eg, ), and the second indicator can be conditionally signaled by checking the value of the first indicator.
[0352] 1) In one example, the second indicator may be signaled under a conditional check of "first indicator equals Y".
[0353] a) Alternatively, the second indicator may be signaled under a conditional check of "value of first indicator >> 1" or "value of first indicator / 2" or "value of first indicator & 0x01".
[0354] b) Alternatively, furthermore, when the first indicator is equal to X, it can be inferred that it is enabled; or when the first indicator is equal to z, it can be inferred that it is disabled.
[0355] b. In the second example, one or more of the following subitems apply:
[0356] i. More than one indicator may be signaled at a higher level (eg, in the picture header (PH)) to specify how to enable LMCS at a lower level other than a binary value.
[0357] 1) In one example, two indicators may be signaled in the PH.
[0358] a) In one example, the first indicator specifies whether there is at least one slice associated with the PH that has LMCS enabled, and the second indicator specifies whether all slices associated with the PH have LMCS enabled.
[0359] i. Alternatively, furthermore, the second indicator may be signaled conditionally depending on the value of the first indicator, eg, when the first indicator specifies that there is at least one slice with LMCS enabled.
[0360] i. Alternatively, further, when the second indicator is not present, it is inferred that all stripes have LMCS enabled.
[0361] ii. Alternatively, the third indicator may furthermore be conditionally signaled in the SH according to the value of the second indicator, for example when the second indicator specifies that not all slices have LMCS enabled.
[0362] i. Alternatively, furthermore, when the third indicator is not present, it may be inferred from the value of the first and / or second indicator (eg, inferred to be equal to the value of the first indicator).
[0363] b) Alternatively, the first indicator specifies whether there is at least one slice associated with the PH with LMCS disabled, and the second indicator specifies whether all slices associated with the PH have LMCS disabled.
[0364] i. Alternatively, furthermore, the second indicator may be signaled conditionally depending on the value of the first indicator, eg when the first indicator specifies that there is at least one slice with disabled LMCS.
[0365] i. Alternatively, furthermore, when the second indicator is not present, it is inferred that LMCS is disabled for all stripes associated with the PH.
[0366] ii. Alternatively, the third indicator may be conditionally signaled in the SH, in addition, depending on the value of the second indicator, eg, when the second indicator specifies that not all slices have LMCS disabled.
[0367] i. Alternatively, furthermore, when the third indicator is not present, it may be inferred from the value of the first and / or second indicator (eg, inferred to be equal to the value of the first indicator).
[0368] 2) Alternatively, in addition, the first indicator may be based on the LMCS enable flag in the sequence level (e.g. ) is conditionally signaled.
[0369] ii. A third indicator (eg, ), and the third indicator can be conditionally signaled by checking the value of the first indicator and / or the second indicator.
[0370] 1) In one example, under the conditional check of "not all slices have LMCS enabled" or "not all slices have LMCS disabled", the third indicator may be signaled.
[0371] c. In yet another example, the first and / or second and / or third indicators mentioned in the first / second example may be used to control the use of RP or CRS instead of LMCS.
[0372] 3. The semantics of the three LMCS flags in SPS / PH / SH are updated as follows:
[0373] Equal to 1 specified in CLVS [[Used]] Luma mapping with chroma scaling. sps_lmcs_enabled_flag equal to 0 specifies that luma mapping with chroma scaling is not used in CLVS.
[0374] A value of 1 specifies luma mapping with chroma scaling [[enabled]] For all slices associated with PH, ph_lmcs_enabled_flag equal to 0 specifies luma mapping with chroma scaling [[one or more may be disabled, or]] All slices associated with PH. When not present, the value of ph_lmcs_enabled_flag is inferred to be equal to 0.
[0375] A value of 1 specifies luma mapping with chroma scaling [[enabled]] slice_lmcs_enabled_flag equal to 0 specifies that luma mapping with chroma scaling is not enabled for the current slice. The current slice. When slice_lmcs_enabled_flag is not present, it is inferred to be equal to 0.
[0376] a. Change the PH and / or SH LMCS signaling so that when LMCS is used for all slices of a picture, there is no LMCS signaling in the SH.
[0377] i. Alternatively, furthermore, how to infer LMCS depends on the PH LMCS signaling.
[0378] 1) In one example, when LMCS is used for all slices of a picture, it is inferred to be enabled; and when LMCS is not used for all slices of a picture, it is inferred to be disabled.
[0379] Related to affine AMVR
[0380] 4. The semantics of the affine AMVR flag in SPS are updated as follows:
[0381] Equal to 1 specifies adaptive motion vector interpolation precision [[yes]] Used in motion vector coding for affine inter mode. sps_affine_amvr_enabled_flag equal to 0 specifies not to use adaptive motion vector difference precision in motion vector coding for affine inter mode. When not present, the value of sps_affine_amvr_enabled_flag is inferred to be equal to 0.
[0382] 6. Examples
[0383] 6.1. Example 1: Sub-image support
[0384] This embodiment is directed to Project 1 and its subprojects.
[0385] 3 definitions
[0386]
[0387] [[ Slice index into the slice list of the picture, in the order they are signaled in the PPS when rect_slice_flag is equal to 1. ]]
[0388] 6.5.1 CTB Raster Scanning, Slice Scanning, and Sub-Picture Scanning Process
[0389] …
[0390] For ctbAddrX in the range from 0 to PicWidthInCtbsY (inclusive) CtbToTileColBd[ctbAddrX] Specifies the range from the horizontal CTB address to the left column boundary in CTB units. The conversion is derived as follows:
[0391]
[0392] Note 3 – In the above derivation CtbToTileColBd[] The size is 1 larger than the actual picture width in the CTB.
[0393] For ctbAddrY in the range 0 to PicHeightInCtbsY (inclusive) CtbToTileRowBd[ctbAddrY] Specifies the vertical CTB address to the top slice column boundary in CTB units. The conversion is derived as follows:
[0394]
[0395] Note 4 – In the above derivation CtbToTileRowBd[] The size is 1 larger than the actual picture height in the CTB.
[0396]
[0397] When rect_slice_flag is equal to 1, the list NumCtusInSlice[i] for i in the range of 0 to num_slices_in_pic_minus1 (inclusive) specifies the number of CTUs in the i-th slice, the list SliceTopLeftTileIdx[i] for i in the range of 0 to num_slices_in_pic_minus1 (inclusive) specifies the slice index of the slice containing the first CTU in the slice, and the matrix CtbAddrInSlice[i][j] for i in the range of 0 to num_slices_in_pic_minus1 (inclusive) and j in the range of 0 to NumCtusInSlice[i]–1 (inclusive) specifies the picture raster scan address of the j-th CTB within the i-th slice, and the variable NumSlicesInTile[i] specifies the number of slices in the slice containing the i-th slice, as derived as follows:
[0398]
[0399]
[0400] ....
[0402] D.7.2 Sub-picture level information SEI message semantics
[0403] …
[0404] [i][j] plus 1 specifies the fraction of the level constraints associated with ref_level_idc[i] that the j-th sub-picture complies with, as specified in clause A.4.1.
[0405] The variable SubpicSizeY[j] is set equal to (subpic_width_minus1[j]+1)*CtbSizeY*(subpic_height_minus1[j]+1)*CtbSizeY.
[0406] When not present, the value of ref_level_fraction_minus1[i][j] is inferred to be equal to Ceil(256*SubpicSizeY[j]÷PicSizeInSamplesY*MaxLumaPs(general_level_idc)÷MaxLumaPs(ref_level_idc[i])-1.
[0407] The variable RefLevelFraction[i][j] is set equal to ref_level_fraction_minus1[i][j]+1.
[0408] The variables SubpicNumTileCols[j] and SubpicNumTileRows[j] are derived as follows:
[0409] ...
[0411] - The value of [j] should be less than or equal to MaxTileCols, and The value of [j] shall be less than or equal to MaxTileRows, where MaxTileCols and MaxTileRows are specified in Table A.1 for level ref_level_idc[i].
[0412] - [j]* The value of [j] shall be less than or equal to MaxTileCols*MaxTileRows*RefLevelFraction[i][j], where MaxTileCols and MaxTileRows are specified in Table A.1 for level ref_level_idc[i]. ...
[0414] The variable SubpicSetAccLevelFraction[i] for the total level fraction relative to the reference level ref_level_idc[i], and the variables SubpicSetCpbSizeVcl[i], SubpicSetCpbSizeNal[i], SubpicSetBitRateVcl[i] and SubpicSetBitRateNal[i] for the sub-picture set are derived as follows:
[0415] ...
[0417] 6.2. Example 2: LMCS support
[0418] In this embodiment, the syntax and semantics of the LMCS-related syntax elements in the picture header are modified so that when LMCS is used for all slices of a picture, there is no LMCS signaling in the SH.
[0419] 7.3.2.7 Picture header structure syntax
[0420]
[0421]
[0422]
[0423] 7.3.7.1 General Strip Header Syntax
[0424]
[0425] equal [[1]] specifies that luma mapping with chroma scaling is enabled for all slices associated with PH. ph_lmcs_enabled_flag equal to 0 specifies that [[luma mapping with chroma scaling may be disabled for one or more, or]] all slices associated with the PH. When not present, the value of ph_lmcs_enabled_flag is inferred to be equal to 0.
[0426] slice_lmcs_enabled_flag is equal to 0 and specifies that luma mapping with chroma scaling is not enabled for the current slice. When slice_lmcs_enabled_flag is not present, it is inferred to be equal to [[0]]
[0427] In the above example, the values of M and N can be set to 1 and 2, respectively. Alternatively, the values of M and N can be set to 2 and 1, respectively.
[0428] Figure 5 1 is a block diagram of an example video processing system 1900 that may implement the various techniques disclosed herein. Various implementations may include some or all of the components of system 1900. System 1900 may include an input 1902 for receiving video content. The video content may be received in a raw or uncompressed format (e.g., 8- or 10-bit multi-component pixel values), or may be received in a compressed or encoded format. Input 1902 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces (such as Ethernet, passive optical networks (PONs), etc.) and wireless interfaces (such as Wi-Fi or cellular interfaces).
[0429] System 1900 may include a codec component 1904 that can implement the various codecs or encoding methods described in this document. Codec component 1904 can reduce the average bit rate of the video from input 1902 to the output of codec component 1904 to produce a codec representation of the video. Therefore, codec technology is sometimes referred to as video compression or video transcoding technology. The output of codec component 1904 can be stored or sent via a connected communication, as represented by component 1906. The stored or communicated bitstream (or codec) representation of the video received at input 1902 can be used by component 1908 to generate pixel values or displayable video that is sent to display interface 1910. The process of generating user-visible video from the bitstream representation is sometimes referred to as video decompression. In addition, although some video processing operations are referred to as "codec" operations or tools, it should be understood that the codec tools or operations are used at the encoder, and the corresponding decoding tools or operations that invert the codec results will be performed by the decoder.
[0430] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB), High-Definition Multimedia Interface (HDMI), DisplayPort, etc. Examples of storage interfaces include SATA (Serial Advanced Technology Attachment), PCI, IDE interfaces, etc. The technology described in this document can be implemented in various electronic devices, such as mobile phones, laptop computers, smart phones, or other devices capable of digital data processing and / or video display.
[0431] Figure 6 36 is a block diagram of a video processing device 3600. Device 3600 can be used to implement one or more of the methods described herein. Device 3600 can be implemented in a smartphone, tablet, computer, Internet of Things (IoT) receiver, etc. Device 3600 may include one or more processors 3602, one or more memories 3604, and video processing hardware 3606. Processor(s) 3602 can be configured to implement one or more of the methods described herein. Memory(s) 3604 can be used to store data and code used to implement the methods and techniques described herein. Video processing hardware 3606 can be used to implement some of the techniques described in this document in hardware circuitry.
[0432] Figure 8 is a block diagram illustrating an example video coding system 100 that may utilize the techniques of this disclosure.
[0433] like Figure 8As shown, the video encoding and decoding system 100 may include a source device 110 and a target device 120. The source device 110 generates encoded video data and may be referred to as a video encoding device. The target device 120 may decode the encoded video data generated by the source device 110 and may be referred to as a video decoding device.
[0434] Source device 110 may include a video source 112 , a video encoder 114 , and an input / output (I / O) interface 116 .
[0435] The video source 112 may include a source such as a video capture device, an interface for receiving video data from a video content provider, and / or a computer graphics system that generates video data, or a combination of these sources. The video data may include one or more pictures. The video encoder 114 encodes the video data from the video source 112 to generate a bitstream. The bitstream may include a sequence of bits that form a codec representation of the video data. The bitstream may include a codec picture and associated data. The codec picture is a codec representation of the picture. The associated data may include a sequence parameter set, a picture parameter set, and other syntax elements. The I / O interface 116 includes a modulator / demodulator (modem) and / or a transmitter. The encoded video data may be sent directly to the target device 120 via the network 130a via the I / O interface 116. The encoded video data may also be stored on a storage medium / server 130b for access by the target device 120.
[0436] Target device 120 may include an I / O interface 126 , a video decoder 124 , and a display device 122 .
[0437] The I / O interface 126 may include a receiver and / or a modem. The I / O interface 126 may obtain encoded video data from the source device 110 or the storage medium / server 130b. The video decoder 124 may decode the encoded video data. The display device 122 may display the decoded video data to a user. The display device 122 may be integrated with the target device 120 or may be external to the target device 120 and configured to interface with an external display device.
[0438] The video encoder 114 and the video decoder 124 may operate according to a video compression standard, such as the High Efficiency Video Codec (HEVC) standard, the Versatile Video Codec (VVC) standard, and other current and / or other standards.
[0439] Figure 9 is a block diagram illustrating an example of a video encoder 200, which may be Figure 8 The video encoder 114 in the system 100 is shown in FIG.
[0440] Video encoder 200 may be configured to perform any or all of the techniques of this disclosure. Figure 9 In the example of FIG, video encoder 200 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of video encoder 200. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.
[0441] The functional components of the video encoder 200 may include a segmentation unit 201, a prediction unit 202 (which may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205 and an intra-frame prediction unit 206), a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213 and an entropy coding unit 214.
[0442] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an intra block copy (IBC) unit. The IBC unit may perform prediction in an IBC mode where at least one reference picture is a picture in which the current video block is located.
[0443] Furthermore, some components such as the motion estimation unit 204 and the motion compensation unit 205 may be highly integrated but are not shown for the purpose of explanation. Figure 9 In the example, they are shown separately.
[0444] The partitioning unit 201 may partition a picture into one or more video blocks. The video encoder 200 and the video decoder 300 may support various video block sizes.
[0445] The mode selection unit 203 can, for example, select one of the intra or inter coding modes based on the error result, and provide the resulting intra or inter coding block to the residual generation unit 207 to generate residual block data and to the reconstruction unit 212 to reconstruct the coding block for use as a reference picture. In some examples, the mode selection unit 203 can select a combined intra and inter prediction (CIIP) mode, where the prediction is based on an inter prediction signal and an intra prediction signal. The mode selection unit 203 can also select the resolution of the motion vector for the block in the case of inter prediction (e.g., sub-pixel or whole pixel accuracy).
[0446] To perform inter-frame prediction on the current video block, the motion estimation unit 204 may generate motion information for the current video block by comparing the current video block with one or more reference frames from the buffer 213. The motion compensation unit 205 may determine a predicted video block for the current video block based on the motion information and decoded samples of a picture from the buffer 213 (other than the picture associated with the current video block).
[0447] The motion estimation unit 204 and the motion compensation unit 205 may perform different operations for the current video block, for example, performing different operations depending on whether the current video block is in an I slice, a P slice, or a B slice.
[0448] In some examples, motion estimation unit 204 may perform unidirectional prediction of the current video block, and motion estimation unit 204 may search for a reference video block for the current video block in the reference pictures in list 0 or list 1. Motion estimation unit 204 may then generate a reference index indicating the reference video block in the reference pictures in list 0 or list 1, and a motion vector indicating the spatial displacement between the current video block and the reference video block. Motion estimation unit 204 may output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 205 may generate a predicted video block for the current block based on the reference video block indicated by the motion information for the current video block.
[0449] In other examples, the motion estimation unit 204 may perform bidirectional prediction for the current video block. The motion estimation unit 204 may search for a reference video block for the current video block in the reference pictures in list 0 and may also search for another reference video block for the current video block in the reference pictures in list 1. The motion estimation unit 204 may then generate a reference index indicating the reference video block in the reference pictures in list 0 or list 1 and a motion vector indicating the spatial displacement between the reference video block and the current video block. The motion estimation unit 204 may output the reference index and the motion vector of the current video block as motion information for the current video block. The motion compensation unit 205 may generate a predicted video block for the current video block based on the reference video block indicated by the motion information for the current video block.
[0450] In some examples, motion estimation unit 204 may output the entire set of motion information for use in a decoding process by a decoder.
[0451] In some examples, motion estimation unit 204 may not output the entire set of motion information for the current video. Instead, motion estimation unit 204 may reference the motion information of another video block to signal the motion information of the current video block. For example, motion estimation unit 204 may determine that the motion information of the current video block is sufficiently similar to the motion information of a neighboring video block.
[0452] In one example, motion estimation unit 204 may indicate in a syntax structure associated with the current video block a value that indicates to video decoder 300 that the current video block has the same motion information as another video block.
[0453] In another example, the motion estimation unit 204 may identify another video block and a motion vector difference (MVD) in a syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 300 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[0454] As discussed above, the video encoder 200 may predictively signal motion vectors. Two examples of predictive signaling techniques that may be implemented by the video encoder 200 include advanced motion vector prediction (AMVP) and merge mode signaling.
[0455] The intra-frame prediction unit 206 can perform intra-frame prediction on the current video block. When the intra-frame prediction unit 206 performs intra-frame prediction on the current video block, the intra-frame prediction unit 206 can generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block can include the predicted video block and various syntax elements.
[0456] The residual generation unit 207 can generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) the predicted video block(s) of the current video block from the current video block. The residual data for the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.
[0457] In other examples, such as in skip mode, there may be no residual data for the current video block, and the residual generation unit 207 may not perform a subtraction operation.
[0458] Transform processing unit 208 may generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video block associated with the current video block.
[0459] After transform processing unit 208 generates a transform coefficient video block associated with the current video block, quantization unit 209 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0460] The inverse quantization unit 210 and the inverse transform unit 211 may apply inverse quantization and inverse transform, respectively, to the transform coefficient video block to reconstruct a residual video block from the transform coefficient video block. The reconstruction unit 212 may add the reconstructed residual video block to corresponding samples from one or more predicted video blocks generated by the prediction unit 202 to generate a reconstructed video block associated with the current block for storage in the buffer 213.
[0461] After the reconstruction unit 212 reconstructs the video block, a loop filtering operation may be performed to reduce video blocking artifacts in the video block.
[0462] The entropy coding unit 214 may receive data from other functional components of the video encoder 200. When the entropy coding unit 214 receives data, the entropy coding unit 214 may perform one or more entropy coding operations to generate entropy-coded data and output a bitstream including the entropy-coded data.
[0463] Some embodiments of the technology of the present disclosure include deciding or determining to enable a video processing tool or mode. In an example, when a video processing tool or mode is enabled, the encoder will use or implement the tool or mode in the processing of the video block, but it is not necessary to modify the resulting bitstream based on the use of the tool or mode. In other words, when a video processing tool or mode is enabled based on a decision or determination, the conversion of the bitstream (or bitstream representation) from the video block to the video will use the video processing tool or mode. In another example, when a video processing tool or mode is enabled, the decoder will process the bitstream using the knowledge that the bitstream has been modified based on the video processing tool or mode. In other words, the conversion from the bitstream of the video to the video block will be performed using the video processing tool or mode enabled based on the decision or determination.
[0464] Figure 10 is a block diagram illustrating an example of a video decoder 300, which may be Figure 8 The video decoder 114 in the system 100 is shown in FIG.
[0465] Video decoder 300 may be configured to perform any or all of the techniques of this disclosure. Figure 10 In the example of FIG, video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of video decoder 300. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.
[0466] exist Figure 10 In the example of FIG, the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra-frame prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, the video decoder 300 can perform the same operation as that of the video encoder 200 ( Figure 9 ) is a decoding process that is the overall inverse of the encoding process described.
[0467] The entropy decoding unit 301 can retrieve a coded bitstream. The coded bitstream can include entropy-encoded video data (e.g., coded blocks of video data). The entropy decoding unit 301 can decode the entropy-encoded video data, and based on the entropy-encoded video data, the motion compensation unit 302 can determine motion information including motion vectors, motion vector precision, reference picture list index, and other motion information. The motion compensation unit 302 can determine such information, for example, by performing AMVP and merge modes.
[0468] The motion compensation unit 302 may generate a motion compensated block, possibly interpolated based on an interpolation filter. An identifier of the interpolation filter to be used with sub-pixel precision may be included in a syntax element.
[0469] The motion compensation unit 302 may calculate interpolated values of a sub-integer number of pixels of the reference block using the interpolation filter used by the video encoder 200 during encoding of the video block. The motion compensation unit 302 may determine the interpolation filter used by the video encoder 200 based on received syntax information and use the interpolation filter to generate a prediction block.
[0470] The motion compensation unit 302 may use some syntax information to determine: the size of the blocks used to encode the frame(s) and / or slice(s) of the coded video sequence, partitioning information describing how each macroblock of a picture of the coded video sequence is partitioned, a mode indicating how to encode each partition, one or more reference frames (and reference frame lists) for each inter-coded block, and other information for decoding the coded video sequence.
[0471] The intra prediction unit 303 can form a prediction block from spatially neighboring blocks using, for example, an intra prediction mode received in the bitstream. The inverse quantization unit 303 inversely quantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 301. The inverse transform unit 303 applies an inverse transform.
[0472] The reconstruction unit 306 can sum the residual block with the corresponding prediction block generated by the motion compensation unit 202 or the intra prediction unit 303 to form a decoded block. If desired, a deblocking filter can also be applied to filter the decoded block to remove blocking artifacts. The decoded video block is then stored in a buffer 307, which provides reference blocks for subsequent motion compensation / intra prediction and also produces decoded video for presentation on a display device.
[0473] A list of preferred solutions for some embodiments is provided below.
[0474] The following solution illustrates an example embodiment of the technique discussed in the previous section (eg, item 1).
[0475] 1. A video processing method (e.g., Figure 7 ), comprising: performing (902) a conversion between a video comprising one or more video pictures and a codec representation of the video, wherein each video picture comprises one or more slices, wherein the codec representation conforms to a format rule; wherein the format rule specifies first information signaled in the codec representation and second information derived from the codec representation, wherein at least the first information or the second information is associated with a row index or a column index of the one or more slices.
[0476] 2. The method of solution 1, wherein the format rules specify deriving a slice column index for each codec tree unit column of each video picture.
[0477] 3. The method of solution 1, wherein the format rules specify deriving a slice row index for each codec tree unit row of each video picture.
[0478] The following solutions illustrate example embodiments of the techniques discussed in the previous section (eg, item 2). In these solutions, the video region may be a video picture, and the video unit may be a video block or a codec tree unit or a video slice.
[0479] 4. A video processing method, comprising: performing conversion between video units of a video area of a video and a codec representation of the video, wherein the codec representation complies with format rules; wherein the format rules specify that first control information at the video area controls whether second control information is included at the video unit level; wherein the first control information and / or the second control information includes information about luminance mapping and chroma scaling (LMCS) or chroma residual scaling (CRS) or a transformation process (RP) for the conversion.
[0480] 5. The method of solution 4, wherein the first control information includes an indicator indicating whether the second control information is included in the codec representation.
[0481] 6. The method according to solution 4-5, wherein a specific value of the first control information indicates that LMCS is disabled for all video units in the video area.
[0482] 7. The method according to any one of solutions 4-6, wherein the second control information controls the enabling of LMCS at the video unit.
[0483] 8. The method according to solution 4, wherein the first control information includes a plurality of indicators.
[0484] 9. A method according to any of solutions 1 to 8, wherein the conversion includes encoding the video into a codec representation.
[0485] 10. A method according to any of solutions 1 to 8, wherein converting comprises decoding the codec representation to generate pixel values of the video.
[0486] 11. A video decoding device comprising a processor configured to implement the method described in one or more of solutions 1 to 10.
[0487] 12. A video encoding apparatus comprising a processor configured to implement the method described in one or more of solutions 1 to 10.
[0488] 13. A computer program product having computer code stored thereon, which, when executed by a processor, causes the processor to implement the method of any one of solutions 1 to 10.
[0489] 14. The methods, apparatus, or systems described in this document.
[0490] In the solution described herein, an encoder can conform to the format rules by generating a codec representation according to the format rules. In the solution described herein, a decoder can use the format rules to parse syntax elements in the codec representation, where the presence and absence of syntax elements are understood according to the format rules to produce decoded video.
[0491] Figure 11 A flow chart of an example method 1100 for video processing is provided. Operation 1102 includes performing conversion between a video comprising one or more video pictures and a bitstream of the video, wherein each video picture comprises one or more slices, the one or more slices comprising one or more slice columns, wherein the bitstream conforms to a format rule, and wherein the format rule provides for derivation of a slice column index for each codec tree unit (CTU) column of a slice of the video picture.
[0492] In some embodiments of the method 1100, the tile column index of the ctbAddrX-th tile column, denoted as ctbToTileColIdx[ctbAddrX], is derived as follows: tileX=0
[0493]
[0494] Where PicWidthInCtbsY represents the width of the video picture in units of codec tree blocks (CTBs), and where tileColBd[i] represents the position of the i-th slice column boundary in units of CTBs.
[0495] In some embodiments of method 1100, each video picture further includes one or more sub-pictures, each sub-picture including one or more slices that together form a rectangular subset of the video picture, and the format rules further provide that the width of the sub-picture in units of included slice columns is derived based on the slice column index of the leftmost CTU and / or the rightmost CTU included in the sub-picture.
[0496] In some embodiments of the method 1100 , the width of the i-th sub-picture in tiles, denoted as SubpicWidthInTiles[i], is derived as follows:
[0497]
[0498]
[0499] Where sps_num_subpics_minus1 represents the number of sub-pictures in the video picture, where sps_subpic_ctu_top_left_x[i] represents the horizontal position of the top-left CTU of the i-th sub-picture, where sps_subpic_width_minus1[i] plus 1 specifies the width of the i-th sub-picture, and where ctbToTileColIdx[rightX] and ctbToTileColIdx[leftX] represent the slice column indices of the leftmost CTU and the rightmost CTU contained in the sub-picture, respectively.
[0500] In some embodiments of method 1100 , in response to the tile being partitioned into a plurality of rectangular strips and only a subset of the tile's rectangular strips being included in the sub-picture, the tile is counted as one tile in the value of the width of the sub-picture.
[0501] Figure 12 12 is a flow chart of an example method 1200 for video processing. Operation 1202 comprises performing conversion between a video comprising one or more video pictures and a bitstream of the video, wherein each video picture comprises one or more slices, the one or more slices comprising one or more slice rows, wherein the bitstream conforms to a format rule, and wherein the format rule provides for derivation of a slice row index for each codec tree unit (CTU) row of a slice of the video picture.
[0502] In some embodiments of the method 1200 , the tile row index of the ctbAddrY th tile row, denoted as ctbToTileRowIdx[ctbAddrY], is derived as follows:
[0503]
[0504] Where PicHeightInCtbsY represents the height of the video picture in units of codec tree blocks (CTBs), and where tileRowBd[i] represents the position of the i-th slice row boundary in units of CTBs.
[0505] In some embodiments of method 1200, each video picture further includes one or more sub-pictures, each sub-picture including one or more slices that together form a rectangular subset of the video picture, and the format rules further provide that the height of the sub-picture in slice rows is derived based on the slice row index of the top CTU and / or bottom CTU included in the sub-picture.
[0506] In some embodiments of method 1200, the height of the i-th sub-picture in tiles, denoted as SubpicHeightInTiles[i], is derived as follows:
[0507]
[0508]
[0509] Where sps_num_subpics_minus1 represents the number of sub-pictures in the video picture, where sps_subpic_ctu_top_left_y[i] represents the vertical position of the top-left CTU of the i-th sub-picture, where sps_subpic_height_minus1[i] plus 1 specifies the height of the i-th sub-picture, and where ctbToTileRowIdx[botY] and ctbToTileRowIdx[topY] represent the slice row indices of the bottom CTU and top CTU, respectively, contained in the sub-picture.
[0510] In some embodiments of method 1200 , in response to the tile being partitioned into a plurality of rectangular strips and only a subset of the rectangular strips of the tile being included in the sub-picture, the tile is counted as one tile in the value of the height of the sub-picture.
[0511] Figure 13 A flow chart of an example method 1300 for video processing. Operation 1302 comprises performing conversion between a video comprising at least one video picture and a bitstream of the video according to a rule, wherein the at least one video picture comprises one or more slices and one or more sub-pictures, and wherein the rule specifies an order of slice indexes indicating one or more slices in the at least one video picture in response to a syntax element associated with the at least one video picture indicating whether each sub-picture of the at least one video picture comprises a single slice.
[0512] In some embodiments of the method 1300, the rules further provide that, in response to each slice in the at least one video picture being a rectangular slice, a slice index is indicated. In some embodiments of the method 1300, the rules provide that, where the syntax element indicates that each of the one or more sub-pictures comprises a single rectangular slice, the order corresponds to increasing values of sub-picture indices of the one or more sub-pictures in the video picture, and the sub-picture indices of the one or more sub-pictures are indicated in a sequence parameter set (SPS) referenced by the at least one video picture. In some embodiments of the method 1300, the rules provide that, where the syntax element indicates that each sub-picture comprises one or more rectangular slices, the order corresponds to an order in which the one or more slices are included in a picture parameter set (PPS) referenced by the at least one video picture. In some embodiments of the method 1300, the syntax element is included in a picture parameter set (PPS) referenced by the at least one video picture.
[0513] Figure 14 A flowchart of an example method 1400 for video processing is provided. Operation 1402 includes performing conversion between video units of a video region of a video and a bitstream of the video, wherein the bitstream conforms to a format rule, wherein the format rule specifies that first control information at a first level of the video region in the bitstream controls whether second control information is included at a second level of the video unit in the bitstream, wherein the second level is less than the first level, wherein the first control information and the second control information include information regarding whether or how a luma mapping and chroma scaling (LMCS) tool is applied to the video unit, and wherein the LMCS tool includes using a chroma residual scaling (CRS) or a luma reformation process (RP) for the conversion.
[0514] In some embodiments of method 1400, the first control information optionally includes a first indicator indicating whether the LMCS tool is enabled for one or more slices of a first level of the video region to specify whether the LMCS tool is enabled at a second level of the video unit, and the first indicator is a non-binary value. In some embodiments of method 1400, the first level of the video region includes a picture header. In some embodiments of method 1400, the first level of the video region includes a picture header, the first control information includes a first indicator, when the first indicator is equal to a first value, the LMCS tool is enabled for all slices of the picture header, when the first indicator is equal to a second value, the LMCS tool is enabled for less than all slices of the picture header, and when the first indicator is equal to a third value, the LMCS tool is disabled for all slices of the picture header, and the first value, the second value, and the third value are different from each other. In some embodiments of method 1400, when the first control information does not include the first indicator, the value of the first indicator is inferred to be a default value.
[0515] In some embodiments of the method 1400, the first level of the video region includes a picture header, the first control information includes a first indicator, when the first indicator is equal to a first value, the LMCS tool is disabled for all slices of the picture header, when the first indicator is equal to a second value, the LMCS tool is disabled for less than all slices of the picture header, when the first indicator is equal to a third value, the LMCS tool is enabled for all slices of the picture header, and the first value, the second value, and the third value are different from each other. In some embodiments of the method 1400, whether the first indicator is selectively included in the first control information is based on a value of a syntax element in the bitstream indicating whether the LMCS tool is enabled at a sequence level. In some embodiments of the method 1400, the first indicator is encoded or decoded using u(v) or u(2) or ue(v). In some embodiments of the method 1400, the first indicator is encoded or decoded using a truncated unary code.
[0516] In some embodiments of the method 1400, based on the value of a first indicator indicating whether the LMCS tool is enabled for one or more slices of a first level of the video region, adaptation parameter set (APS) information and / or chroma scaling syntax elements for the LMCS tool used by the one or more slices are included in the bitstream. In some embodiments of the method 1400, the second control information selectively includes a second indicator indicating whether the LMCS tool is enabled or disabled for one or more slices of a second level of the video unit, and the second indicator is included in the bitstream based on the value of the first indicator included in the first control information, and the first indicator indicates whether the LMCS tool is enabled or disabled for the one or more slices of the second level of the video unit. In some embodiments of the method 1400, the second control information includes a slice header. In some embodiments of the method 1400, in response to the first indicator being equal to the first value, the second indicator is included in the second control information. In some embodiments of method 1400, the second indicator is included in the second control information in response to performing the following conditional check: the first indicator >>1, or the first indicator / 2, or the first indicator &0x01, where >> describes a right shift operation, and where & describes a bitwise logical and operation.
[0517] In some embodiments of method 1400, in response to the first indicator being equal to a first value, the second indicator is inferred to indicate that the LMCS tool is enabled for one or more slices of the second level of the video unit, or in response to the first indicator being equal to a third value, the second indicator is inferred to indicate that the LMCS tool is disabled for one or more slices of the second level of the video unit, and the first value, the second value, and the third value of the first indicator are different from one another. In some embodiments of method 1400, the first control information includes a plurality of indicators indicating whether the LMCS tool is enabled for one or more slices of the first level of the video region to specify whether the LMCS tool is enabled at the second level of the video unit, and the plurality of indicators have non-binary values. In some embodiments of method 1400, the plurality of indicators include at least two indicators included in a picture header. In some embodiments of method 1400, the at least two indicators include a first indicator specifying whether the LMCS tool is enabled for at least one slice associated with the picture header, and the at least two indicators optionally include a second indicator specifying whether the LMCS tool is enabled for all slices associated with the picture header. In some embodiments of method 1400 , the second indicator is selectively present in the plurality of indicators based on a value of the first indicator.
[0518] In some embodiments of method 1400, the value of the first indicator indicates that the LMCS tool is enabled for at least one slice. In some embodiments of method 1400, in response to the absence of the second indicator in the bitstream, it is inferred that the LMCS tool is enabled for all slices associated with the picture header. In some embodiments of method 1400, the at least two indicators include a third indicator that is selectively included in the slice header based on a second value of the second indicator. In some embodiments of method 1400, the second value of the second indicator indicates that the LMCS tool is disabled for all slices. In some embodiments of method 1400, in response to the absence of the third indicator in the bitstream, the value of the third indicator is inferred based on the first value of the first indicator and / or the second value of the second indicator. In some embodiments of method 1400, the at least two indicators include a first indicator that specifies whether the LMCS tool is disabled for at least one slice associated with the picture header, and the at least two indicators selectively include a second indicator that specifies whether the LMCS tool is disabled for all slices associated with the picture header. In some embodiments of the method 1400, the second indicator is present in the plurality of indicators based on a value of the first indicator. In some embodiments of the method 1400, the value of the first indicator specifies that the LMCS tool is disabled for at least one slice. In some embodiments of the method 1400, in response to the absence of the second indicator in the bitstream, it is inferred that the LMCS tool is disabled for all slices associated with the picture header. In some embodiments of the method 1400, at least two indicators selectively include a third indicator in the slice header based on a second value of the second indicator.
[0519] In some embodiments of method 1400, the second value of the second indicator specifies that the LMCS tool is enabled for all slices. In some embodiments of method 1400, in response to the absence of the third indicator in the bitstream, the value of the third indicator is inferred based on the first value of the first indicator and / or the second value of the second indicator. In some embodiments of method 1400, the plurality of indicators selectively include a first indicator based on the value of a syntax element indicating whether the LMCS tool is enabled at the sequence level. In some embodiments of method 1400, the plurality of indicators selectively include a third indicator that indicates whether the LMCS tool is enabled or disabled at the second level of the video unit, and the third indicator is selectively present based on the first value of the first indicator and / or the second value of the second indicator. In some embodiments of method 1400, the third indicator is selectively present based on the second indicator that indicates that the LMCS tool is not enabled for all slices or is not disabled for all slices. In some embodiments of method 1400, the first indicator, the second indicator, and / or the third indicator control the use of CRS or luma RP.
[0520] Figure 15 A flow chart of an example method 1500 for video processing. Operation 1502 comprises performing conversion between a video and a bitstream of the video according to a rule, wherein the rule provides for enabling a luma mapping and chroma scaling (LMCS) tool when a first syntax element in a reference sequence parameter set indicates enabling the LMCS tool, wherein the rule provides for not using the LMCS tool when the first syntax element indicates disabling the LMCS tool, wherein the rule provides for enabling the LMCS tool for all slices associated with a picture header of a video picture when a second syntax element in the bitstream indicates enabling the LMCS tool at a picture header level of the video, wherein the rule provides for not using the LMCS tool for all slices associated with the picture header when the second syntax element indicates disabling the LMCS tool at the picture header level of the video, wherein the rule provides for using the LMCS tool for a current slice associated with a slice header of the video picture when a third syntax element selectively included in the bitstream indicates enabling the LMCS tool at a slice header level of the video, and wherein the rule provides for not using the LMCS tool for the current slice when the third syntax element indicates disabling the LMCS tool at the slice header level of the video.
[0521] In some embodiments of the method 1500, a rule specifies that when the LMCS tool is applied to all slices of a video picture, the third syntax element is not included in the slice header of the bitstream. In some embodiments of the method 1500, whether the LMCS tool is enabled or disabled is based on the second syntax element. In some embodiments of the method 1500, the LMCS tool is enabled when the LMCS tool is applied to all slices of the video picture, and is disabled when the LMCS tool is not applied to all slices of the video picture.
[0522] Figure 16 A flowchart of an example method 1600 for video processing is provided. Operation 1602 includes performing conversion between a video including one or more video pictures and a bitstream of the video according to a rule, wherein the rule specifies whether adaptive motion vector difference precision (AMVR) is used in motion vector coding for affine inter mode based on a syntax element selectively included in a reference sequence parameter set (SPS) indicating whether AMVR is enabled, wherein the rule specifies that when the syntax element indicates that AMVR is disabled, AMVR is not used for motion vector coding for affine inter mode, and wherein the rule specifies that when the syntax element is not included in the SPS, AMVR is inferred not to be used for motion vector coding for affine inter mode.
[0523] Figure 17A flow chart of an example method 1700 for video processing is provided. Operation 1702 comprises performing, according to a rule, conversion between a video comprising a video picture and a bitstream of the video, wherein the video picture comprises a sub-picture, a slice, and a slice, and wherein the rule specifies that because the sub-picture comprises a slice divided from a slice, the conversion is performed by avoiding using a number of slices of the video picture to calculate a height of the sub-picture.
[0524] In some embodiments of the method 1700, the height of the sub-picture is calculated based on the number of codec tree units (CTUs). In some embodiments of the method 1700, the height of the sub-picture is less than one slice line.
[0525] Figure 18 Flowchart of an example method 1800 for video processing. Operation 1802 includes performing conversion between a video including a video picture and a bitstream of the video, wherein the bitstream indicates a height of a sub-picture of the video picture calculated based on a number of codec tree units (CTUs) of the video picture.
[0526] In some embodiments of method 1800, the height of the sub-picture is not based on the number of slices of the video picture.In some embodiments of method 1800, the height of the sub-picture is less than one slice row.
[0527] Figure 19 Flowchart of an example method 1900 for video processing. Operation 1902 includes making a determination based on a rule as to whether a sub-picture of a video picture of the video is less than a slice-line height of the video picture. Operation 1904 includes using the determination to perform conversion between the video and a bitstream of the video.
[0528] In some embodiments of method 1900, a rule provides that the height of a sub-picture is less than one slice row when: the sub-picture includes only one slice row of codec tree units (CTUs), and a first group of CTUs at the top of the sub-picture is different from a second group of CTUs at the top of the one slice row, or a third group of CTUs at the bottom of the sub-picture is different from a fourth group of CTUs at the bottom of the one slice row. In some embodiments of method 1900, when each sub-picture of a video picture includes only one slice and the height of the sub-picture is less than one slice row, for each slice of the video picture with picture-level slice index i, the value of CtbAddrInSlice[i][j] is derived from the picture raster scan CTU address of the jth CTU in the CTU raster scan of the sub-picture, and j is in the range from 0 to the number of CTUs in the slice minus 1, inclusive.
[0529] In some embodiments of method 1900, the rule provides that when a distance between a first group of CTUs at a top of the sub-picture and a second group of CTUs at a bottom of the sub-picture is less than a second height of a slice of the video picture, the height of the sub-picture is less than one slice row, where the second height of the slice is based on the number of CTUs of the sub-picture. In some embodiments of method 1900, when each sub-picture of the video picture includes only one slice and the height of the sub-picture is greater than or equal to one slice row, for each slice of the video picture with picture-level slice index i, the value of CtbAddrInSlice[i][j] is derived from the picture raster scan CTU address of the jth CTU in the order of the CTUs in the sub-picture, and j is in the range from 0 to the number of CTUs in the slice minus 1, inclusive.
[0530] In some embodiments of method 1900, the order of the CTUs in the sub-picture is such that a first CTU in a first slice having a first slice index is placed before a second CTU in a second slice having a second slice index, and the value of the first slice index is less than the value of the second slice index. In some embodiments of method 1900, the order of the CTUs in the sub-picture is such that the CTUs within a slice in the sub-picture are ordered in a raster scan of the CTUs in the slice.
[0531] In some embodiments of method(s) 1100-1900, performing the conversion includes encoding the video into a bitstream. In some embodiments of method(s) 1100-1900, performing the conversion includes encoding the video into a bitstream, and the method further includes storing the bitstream in a non-transitory computer-readable recording medium. In some embodiments of method(s) 1100-1900, performing the conversion includes decoding the video from the bitstream.
[0532] In some embodiments, a video decoding device includes a processor configured to implement the operations described for any one or more of methods 1100 to 1900. In some embodiments, a video encoding device includes a processor configured to implement the operations described for any one or more of methods 1100 to 1900. In some embodiments, a computer program product having computer instructions stored thereon, which, when executed by a processor, causes the processor to implement the operations described for any one or more of methods 1100 to 1900. In some embodiments, a non-transitory computer-readable storage medium stores a bitstream generated according to the operations described for any one or more of methods 1100 to 1900. In some embodiments, a non-transitory computer-readable storage medium stores instructions that cause the processor to implement the operations described for any one or more of methods 1100 to 1900. In some embodiments, a bitstream generation method includes generating a video bitstream according to the operations described for any one or more of methods 1100 to 1900, and storing the bitstream on a computer-readable program medium. In some embodiments, a method, an apparatus, and a bitstream generated according to the disclosed method or system described in this document.
[0533] In this document, the term "video processing" may refer to video encoding, video decoding, video compression, or video decompression. For example, during the conversion from a pixel representation of a video to a corresponding bitstream representation, or vice versa, a video compression algorithm may be applied. The bitstream representation of the current video block may, for example, correspond to bits that are co-located or spread across different locations within the bitstream, as defined by the syntax. For example, a macroblock may be encoded based on error residual values from a transform and a codec, and also using bits in a header and other fields in the bitstream. Furthermore, during conversion, a decoder may, based on this determination, parse the bitstream knowing that some fields may or may not be present, as described in the solution above. Similarly, an encoder may determine whether to include or not include certain syntax fields, and generate the codec representation accordingly by including or excluding the syntax fields from the codec representation.
[0534] The disclosed and other aspects, examples, embodiments, modules, and functional operations described in this document may be implemented in digital electronic circuitry or in computer software, firmware, or hardware, including the structures disclosed in this document and their structural equivalents, or a combination of one or more thereof. The disclosed and other embodiments may be implemented as one or more computer program products, such as one or more modules of computer program instructions, encoded on a computer-readable medium for execution by a data processing apparatus or to control the operation of the data processing apparatus. The computer-readable medium may be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composite material that effects a machine-readable, propagable signal, or a combination of one or more thereof. The term "data processing apparatus" encompasses all apparatus, devices, and machines for processing data, including, for example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, the apparatus may also include code that creates an execution environment for the computer program in question, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more thereof. A propagated signal is an artificially generated signal, such as a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to a suitable receiver device.
[0535] A computer program (also referred to as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, and can be deployed in any form, including stand-alone programs or modules, components, subroutines, or other units suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored in a portion of a file preserving other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple collaborative files (e.g., files storing one or more modules, subroutines, or code portions). A computer program can be deployed to execute on one computer or on multiple computers, which are located at a site or distributed across multiple sites and interconnected by a communication network.
[0536] The processes and logic flows described in this document may be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows may also be performed by, and apparatus may be implemented as, special purpose logic circuitry (e.g., a field programmable gate array (FPGA) or an application specific integrated circuit (ASIC)).
[0537] Processors suitable for executing computer programs include, for example, both general-purpose and special-purpose microprocessors, as well as any one or more processors of any type of digital computer. Typically, a processor will receive instructions and data from read-only memory or random-access memory, or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include one or more mass storage devices (e.g., magnetic, magneto-optical, or optical disks) for storing data, or be operatively coupled to receive data from or transfer data to a mass storage device (e.g., magnetic, magneto-optical, or optical disks), or both. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of nonvolatile memory, media, and storage devices, including, for example, semiconductor memory devices (e.g., EPROM, EEPROM, and flash memory devices); magnetic disks (e.g., internal hard disks or removable disks); magneto-optical disks; and CD ROM and DVD-ROM disks. The processor and memory may be supplemented by, or incorporated into, special-purpose logic circuitry.
[0538] Although this patent document contains many details, these details should not be interpreted as limitations on any subject matter or the scope of the claims, but rather as descriptions of features specified for specific embodiments of a particular technology. In this patent document, certain features described in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented separately in multiple embodiments or in various suitable subcombinations. In addition, although features may be described above as working in certain combinations and even initially claimed in the same manner, in some cases one or more features from the claimed combination may be removed from the combination, and the claimed combination may be directed to a subcombination or a variation of the subcombination.
[0539] Similarly, while operations are depicted in a particular order in the drawings, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, in order to achieve desired results. Furthermore, the separation of various system components in the embodiments described in this patent document should not be understood as requiring such separation in all embodiments.
[0540] Only a few implementations and examples are described, and other implementations, enhancements, and variations can be made based on what is described and illustrated in this patent document.
Claims
1. A video processing method, comprising: performing, according to a rule, conversion between a video comprising at least one video picture and a bitstream of said video, The at least one video picture includes one or more slices and one or more sub-pictures, and wherein the rule specifies indicating an order of picture-level slice indexes of the one or more slices in the at least one video picture in response to a syntax element single_slice_per_subpic_flag associated with the at least one video picture, wherein the syntax element single_slice_per_subpic_flag indicates whether each sub-picture of the at least one video picture includes a single slice, The rule further provides that in response to each of the slices in the at least one video picture being a rectangular slice, the picture-level slice index is indicated, where the picture-level slice index indicates the index of the slice in the slice list. wherein the rules further specify that, in the case where the syntax element single_slice_per_subpic_flag indicates that each sub-picture of the one or more sub-pictures comprises a single rectangular slice, the order corresponds to increasing values of the sub-picture index of the one or more sub-pictures, The rule further stipulates that, when the syntax element single_slice_per_subpic_flag indicates that each sub-picture is allowed to include one or more rectangular slices, the order corresponds to the order in which the one or more slices are included in the picture parameter set PPS referenced by the at least one video picture. Wherein, the sub-picture indexes of the one or more sub-pictures are derived based on information in a sequence parameter set SPS referenced by the at least one video picture.
2. The method according to claim 1, wherein The syntax element is included in a picture parameter set (PPS) referenced by the at least one video picture.
3. The method according to any one of claims 1 to 2, wherein Performing the conversion includes encoding the video into the bitstream.
4. The method according to any one of claims 1 to 2, wherein Performing the conversion includes decoding the video from the bitstream.
5. A video data processing apparatus comprising a processor and a non-transitory memory having instructions thereon, wherein: The instructions, when executed by the processor, cause the processor to: performing, according to a rule, conversion between a video comprising at least one video picture and a bitstream of said video, The at least one video picture includes one or more slices and one or more sub-pictures, and wherein the rule specifies an order of picture-level slice indexes of the one or more slices in the at least one video picture in response to a syntax element single_slice_per_subpic_flag associated with the at least one video picture, wherein the syntax element single_slice_per_subpic_flag indicates whether each sub-picture of the at least one video picture includes a single slice The rule further provides that in response to each of the slices in the at least one video picture being a rectangular slice, the picture-level slice index is indicated, where the picture-level slice index indicates the index of the slice in the slice list. wherein the rules further specify that, in the case where the syntax element single_slice_per_subpic_flag indicates that each sub-picture of the one or more sub-pictures comprises a single rectangular slice, the order corresponds to increasing values of the sub-picture index of the one or more sub-pictures, The rule further stipulates that, when the syntax element single_slice_per_subpic_flag indicates that each sub-picture is allowed to include one or more rectangular slices, the order corresponds to the order in which the one or more slices are included in the picture parameter set PPS referenced by the at least one video picture. Wherein, the sub-picture indexes of the one or more sub-pictures are derived based on information in a sequence parameter set SPS referenced by the at least one video picture.
6. The device according to claim 5, in, The syntax element is included in a picture parameter set (PPS) referenced by the at least one video picture.
7. A non-transitory computer-readable storage medium having stored therein instructions that cause a processor to: performing, according to a rule, conversion between a video comprising at least one video picture and a bitstream of said video, in, The at least one video picture includes one or more slices and one or more sub-pictures, and wherein the rule specifies an order of picture-level slice indexes of the one or more slices in the at least one video picture in response to a syntax element single_slice_per_subpic_flag associated with the at least one video picture, wherein the syntax element single_slice_per_subpic_flag indicates whether each sub-picture of the at least one video picture includes a single slice The rule further provides that in response to each of the slices in the at least one video picture being a rectangular slice, the picture-level slice index is indicated, where the picture-level slice index indicates the index of the slice in the slice list. wherein the rules further specify that, in the case where the syntax element single_slice_per_subpic_flag indicates that each sub-picture of the one or more sub-pictures comprises a single rectangular slice, the order corresponds to increasing values of the sub-picture index of the one or more sub-pictures, The rule further stipulates that, when the syntax element single_slice_per_subpic_flag indicates that each sub-picture is allowed to include one or more rectangular slices, the order corresponds to the order in which the one or more slices are included in the picture parameter set PPS referenced by the at least one video picture. Wherein, the sub-picture indexes of the one or more sub-pictures are derived based on information in a sequence parameter set SPS referenced by the at least one video picture.
8. The non-transitory computer-readable storage medium of claim 7, wherein: The syntax element is included in a picture parameter set (PPS) referenced by the at least one video picture.
9. A non-transitory computer-readable recording medium storing a bit stream of a video generated by a method executed by a video processing device, wherein: The method comprises: generating a bitstream of the video comprising at least one video picture according to a rule, The at least one video picture includes one or more slices and one or more sub-pictures, and wherein the rule specifies an order of picture-level slice indexes of the one or more slices in the at least one video picture in response to a syntax element single_slice_per_subpic_flag associated with the at least one video picture, wherein the syntax element single_slice_per_subpic_flag indicates whether each sub-picture of the at least one video picture includes a single slice The rule further provides that in response to each of the slices in the at least one video picture being a rectangular slice, the picture-level slice index is indicated, where the picture-level slice index indicates the index of the slice in the slice list. wherein the rules further specify that, in the case where the syntax element single_slice_per_subpic_flag indicates that each sub-picture of the one or more sub-pictures comprises a single rectangular slice, the order corresponds to increasing values of the sub-picture index of the one or more sub-pictures, The rule further stipulates that, when the syntax element single_slice_per_subpic_flag indicates that each sub-picture is allowed to include one or more rectangular slices, the order corresponds to the order in which the one or more slices are included in the picture parameter set PPS referenced by the at least one video picture. Wherein, the sub-picture indexes of the one or more sub-pictures are derived based on information in a sequence parameter set SPS referenced by the at least one video picture.
10. The non-transitory computer-readable recording medium according to claim 9, wherein The syntax element is included in a picture parameter set (PPS) referenced by the at least one video picture.
11. A method for storing a bitstream of a video, comprising: generating a bitstream of the video comprising at least one video picture according to a rule, and storing the bitstream in a non-transitory computer-readable recording medium, The at least one video picture includes one or more slices and one or more sub-pictures, and wherein the rule specifies an order of picture-level slice indexes of the one or more slices in the at least one video picture in response to a syntax element single_slice_per_subpic_flag associated with the at least one video picture, wherein the syntax element single_slice_per_subpic_flag indicates whether each sub-picture of the at least one video picture includes a single slice The rule further provides that in response to each of the slices in the at least one video picture being a rectangular slice, the picture-level slice index is indicated, where the picture-level slice index indicates the index of the slice in the slice list. wherein the rules further specify that, in the case where the syntax element single_slice_per_subpic_flag indicates that each sub-picture of the one or more sub-pictures comprises a single rectangular slice, the order corresponds to increasing values of the sub-picture index of the one or more sub-pictures, The rule further stipulates that, when the syntax element single_slice_per_subpic_flag indicates that each sub-picture is allowed to include one or more rectangular slices, the order corresponds to the order in which the one or more slices are included in the picture parameter set PPS referenced by the at least one video picture. The sub-picture indexes of the one or more sub-pictures are derived based on information in a sequence parameter set (SPS) referenced by the at least one video picture.