Sub-picture level segmentation calculation
By simplifying sub-picture height calculation, optimizing LMCS signaling notification and correcting the semantics of SPS affine AMVR flags, the inefficiency problems in VVC standard neutron pictures and LMCS design are solved, and more efficient video encoding and decoding performance is achieved.
Patent Information
- Application Number
- CN202180016670.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-02-24
- Filing Date
- 2021-02-23
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2041-02-23
AI Technical Summary
In the prior art, sub-picture and LMCS design in the VVC standard have problems with inefficient efficiency, including complex sub-picture calculations, inefficient LMCS signaling notifications, and incorrect SPS affine AMVR flag semantics.
By deducing the number of slice rows and slice columns in the sub-picture, the calculation of sub-picture height is simplified; the two-level control of LMCS is introduced, and the control information at the picture level and strip level is used to improve efficiency; the semantics of SPS affine AMVR flags are corrected to ensure their correct application.
It improves the simplicity and efficiency of sub-picture height calculation, optimizes the LMCS signaling notification process, ensures the correct semantic application of SPS affine AMVR flag, thereby improving the overall performance of video encoding and decoding.
Smart Images

Figure CN115152209B_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This application is a national stage entry of International Patent Application No. PCT / US2021 / 019216, filed on February 23, 2021, which claims the priority of U.S. Provisional Patent Application No. US 62 / 980,963, filed on February 24, 2020. The entire disclosure of the above applications is incorporated by reference as part of the disclosure of this application. Technical Field
[0003] This patent document relates to image and video encoding, decoding, and transcoding. Background Art
[0004] Digital video accounts for the largest bandwidth usage on the Internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth demand for digital video use is expected to continue to grow. Summary of the Invention
[0005] This document discloses techniques that can be used by video encoders and decoders for processing encoded representations of video using control information useful for decoding the encoded representations.
[0006] In one example aspect, a video processing method is disclosed. The method includes: performing a conversion between a video including one or more video pictures and a bitstream of the video, where each video picture includes one or more slices, the one or more slices include one or more slice columns, the bitstream conforms to format rules, and the format rules specify deriving a slice column index for each coding tree unit (CTU) column of the slices of the video picture.
[0007] In another example aspect, a video processing method is disclosed. The method includes: performing a conversion between a video including one or more video pictures and a bitstream of the video, where each video picture includes one or more slices, the one or more slices include one or more slice columns, the bitstream conforms to format rules, and the format rules specify deriving a slice row index for each coding tree unit (CTU) row of the slices of the video picture.
[0008] In another example aspect, a video processing method is disclosed. The method includes: performing a conversion between a video including at least one video picture and a bitstream of the video according to rules, where the at least one video picture includes one or more strips and one or more sub - pictures, and the rules specify an order of strip indices of one or more strips in the at least one video picture in response to a syntax element associated with the at least one video picture indicating whether each sub - picture of the at least one video picture includes a single strip.
[0009] In another example aspect, a video processing method is disclosed. The method includes: performing a conversion between video units of a video region of a video and a bitstream of the video, where the bitstream conforms to format rules, where the format rules specify that first-level first control information of the video region in the bitstream controls whether a second level of video units in the bitstream includes second control information, where the second level is less than the first level, where the first control information and the second control information include information on whether or how a luminance mapping and chrominance scaling (LMCS) tool is applied to the video units, and where the LMCS tool includes using chrominance residual scaling (CRS) or a luminance remodelling process (RP) for the conversion.
[0010] In another example aspect, a video processing method is disclosed. The method includes: performing a conversion between a video and a bitstream of the video according to rules, where the rules specify that a luminance mapping and chrominance scaling (LMCS) tool is enabled when a first syntax element in a reference sequence parameter set indicates that the LMCS tool is enabled, where the rules specify that the LMCS tool is not used when the first syntax element indicates that the LMCS tool is disabled, where the rules specify that when a second syntax element in the bitstream indicates that the LMCS tool is enabled at the picture header level of the video, the LMCS tool is enabled for all slices associated with the picture header of the video picture, where the rules specify that when the second syntax element indicates that the LMCS tool is disabled at the picture header level of the video, the LMCS tool is not used for all slices associated with the picture header, where the rules specify that when a third syntax element selectively included in the bitstream indicates that the LMCS tool is enabled at the slice header level of the video, the LMCS tool is used for the current slice associated with the slice header of the video picture, and where the rules specify that when the third syntax element indicates that the LMCS tool is disabled at the slice header level of the video, the LMCS tool is not used for the current slice.
[0011] In another example aspect, a video processing method is disclosed. The method includes: performing a conversion between a video including one or more video pictures and a bitstream of the video according to rules, where the rules specify that whether adaptive motion vector difference precision (AMVR) is used in motion vector coding / decoding for an affine inter prediction mode is based on a syntax element selectively included in a reference sequence parameter set (SPS) indicating whether AMVR is enabled, where the rules specify that AMVR is not used in motion vector coding / decoding for the affine inter prediction mode when the syntax element indicates that AMVR is disabled, and where the rules specify that AMVR is inferred not to be used in motion vector coding / decoding for the affine inter prediction mode when the syntax element is not included in the SPS.
[0012] In another example aspect, a video processing method is disclosed. The method includes: performing a conversion between a video including video pictures and a bitstream of the video according to a rule, where the video pictures include sub-pictures, slices, and strips, and where the rule specifies that since the sub-pictures include strips segmented from slices, the conversion is performed by avoiding using the number of slices of the video pictures to calculate the height of the sub-pictures.
[0013] In another example aspect, a video processing method is disclosed. The method includes: performing a conversion between a video including video pictures and a bitstream of the video, where the bitstream indicates the height of sub-pictures of the video pictures calculated based on the number of coding tree units (CTUs) of the video pictures.
[0014] In another example aspect, a video processing method is disclosed. The method includes: making a determination according to a rule, the determination being about whether the height of sub-pictures of video pictures of the video is less than the height of slice rows of the video pictures; and using the determination to perform a conversion between the video and the bitstream of the video.
[0015] In another example aspect, a video processing method is disclosed. The method includes: performing a conversion between a video including one or more video pictures and a coded representation of the video, where each video picture includes one or more slices, and where the coded representation complies with format rules; where the format rules specify first information signaled in the coded representation and second information derived from the coded representation, where at least the first information or the second information is related to the row index or column index of one or more slices.
[0016] In another example aspect, a video processing method is disclosed. The method includes: performing a conversion between video units of a video region of a video and a coded representation of the video, where the coded representation complies with format rules; where the format rules specify that first control information at the video region controls whether second control information is included at the video unit level; where the first control information and / or the second control information includes information about luminance mapping and chrominance scaling (LMCS) or chrominance residual scaling (CRS) or remodelling process (RP) for the conversion.
[0017] In yet another example aspect, a video encoder device is disclosed. The video encoder includes a processor configured to implement the above method.
[0018] In yet another example aspect, a video decoder device is disclosed. The video decoder includes a processor configured to implement the above method.
[0019] In yet another example aspect, a computer-readable medium storing code is disclosed. The code embodies one of the methods described herein in the form of processor-executable code.
[0020] These and other features are described in this document. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 An example of raster scan stripe segmentation of an image is shown, where the image is divided into 12 slices and 3 raster scan stripes.
[0022] Figure 2 An example of rectangular stripe segmentation of an image is shown, where the image is divided into 24 slices (6 slice columns and 4 slice rows) and 9 rectangular stripes.
[0023] Figure 3 An example of an image segmented into slices and rectangular stripes is shown, where the image is divided into 4 slices (2 slice columns and 2 slice rows) and 4 rectangular stripes.
[0024] Figure 4 An image divided into 15 slices, 24 stripes, and 24 sub - images is shown.
[0025] Figure 5 Is a block diagram of an example video processing system.
[0026] Figure 6 Is a block diagram of a video processing device.
[0027] Figure 7 Is a flowchart of an example method for video processing.
[0028] Figure 8 Is a block diagram showing a video codec system according to some embodiments of the present disclosure.
[0029] Figure 9 Is a block diagram showing an encoder according to some embodiments of the present disclosure.
[0030] Figure 10 Is a block diagram showing a decoder according to some embodiments of the present disclosure.
[0031] Figures 11 to 19 Is a flowchart of an example method for video processing. DETAILED DESCRIPTION
[0032] In this document, chapter headings are used for ease of understanding, and the applicability of the technologies and embodiments disclosed in each chapter is not limited to that chapter only. Additionally, in some descriptions, H.266 technical terms are used merely for ease of understanding and not to limit the scope of the disclosed technologies. Therefore, the technologies described herein are also applicable to other video codec protocols and designs. In this document, with respect to the current draft of the VVC specification, deleted text between double brackets (e.g., [[]]) and added text in bold italic are used to indicate editorial changes to the text.
[0033] 1. Overview
[0034] This document relates to video coding and decoding technology. Specifically, it is about the support for sub - pictures, LMCS, and AMVR. Regarding the sub - picture aspect, it includes deriving the number of slice rows and slice columns included in a sub - picture, and deriving a list of raster - scanned CTU addresses of CTUs included in a strip when each sub - picture contains only one strip. The LMCS aspect is about signaling to enable LMCS at different levels. The AMVR aspect is about the semantics of the sps_affine_amvr_enabled_flag. These ideas can be applied individually or in various combinations to any video coding standard or non - standard video codec that supports single - layer and / or multi - layer video coding and decoding, such as the Versatile Video Coding (VVC) being developed.
[0035] 2. Abbreviations
[0036] ALF Adaptive Loop Filter
[0037] AMVR Adaptive Motion Vector Resolution
[0038] APS Adaptive Parameter Set
[0039] AU Access Unit
[0040] AUD Access Unit Delimiter
[0041] AVC Advanced Video Coding
[0042] CLVS Coding - Layer Video Sequence
[0043] CPB Coding - Picture Buffer
[0044] CRA Clean Random Access
[0045] CTU Coding - Tree Unit
[0046] CVS Coding - Video Sequence
[0047] DPB Decoded - Picture Buffer
[0048] DPS Decoding Parameter Set
[0049] EOB End - of - Bit - stream
[0050] EOS End - of - Sequence
[0051] GDR Gradual Decoding Refresh
[0052] HEVC High - Efficiency Video Coding
[0053] HRD Hypothetical Reference Decoder
[0054] IDR Instantaneous Decoding Refresh
[0055] JEM Joint Exploration Model
[0056] LMCS Luma Mapping with Chroma Scaling
[0057] MCTS Motion Constrained Tile Set
[0058] NAL Network Abstraction Layer
[0059] OLS Output Layer Set
[0060] PH Picture Header
[0061] PPS Picture Parameter Set
[0062] PTL Profile, Tier and Level
[0063] PU Picture Unit
[0064] RBSP Raw Byte Sequence Payload
[0065] SEI Supplemental Enhancement Information
[0066] SPS Sequence Parameter Set
[0067] SVC Scalable Video Coding
[0068] VCL Video Coding Layer
[0069] VPS Video Parameter Set
[0070] VTM VVC Test Model
[0071] VUI Video Usability Information
[0072] VVC Versatile Video Coding
[0073] 3. Preliminary Discussion
[0074] Video coding standards have mainly evolved through the well-known ITU-T and ISO / IEC standards. ITU-T developed H.261 and H.263, ISO / IEC developed MPEG-1 and MPEG-4 Visual, and the two organizations jointly developed the H.262 / MPEG-2 video standard, the H.264 / MPEG-4 Advanced Video Coding (AVC) standard, and the H.265 / HEVC standard. Since H.262, video coding standards have been based on a hybrid video coding structure, where temporal prediction plus transform coding is used. To explore future video coding technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Team (JVET) in 2015. Since then, JVET has adopted many new methods and incorporated them into a reference software called the Joint Exploration Model (JEM). JVET meetings are held quarterly simultaneously. The goal of the new coding standard is to reduce the bitrate by 50% compared to HEVC. The new video coding standard was officially named Versatile Video Coding (VVC) at the JVET meeting in April 2018, and the first version of the VVC Test Model (VTM) was released at that time. Due to the continuous efforts on VVC standardization, new coding technologies have been incorporated into the VVC standard at each JVET meeting. The working draft of VVC and the test model VTM are updated after each meeting. The VVC project now aims to be technically completed (FDIS) at the meeting in July 2020.
[0075] 3.1. Picture segmentation schemes in HEVC
[0076] HEVC includes four different picture segmentation schemes, namely regular slices, non-independent slices, tiles, and wavefront parallel processing (WPP), which can be applied to maximum transmission unit (MTU) size matching, parallel processing, and reducing end-to-end latency.
[0077] Regular slices are similar to those in H.264 / AVC. Each regular slice is encapsulated in its own NAL unit, and intra-picture prediction (intra-sample prediction, motion information prediction, coding mode prediction) and entropy coding dependency across slice boundaries are disabled. Therefore, a regular slice can be reconstructed independently of other regular slices within the same picture (although there may still be dependencies due to loop filter operations).
[0078] Regular stripes are the only tool available for parallelization and are also available in almost the same form in H.264 / AVC. Parallelization based on regular stripes does not require much inter-processor or inter-core communication (except for the inter-processor or inter-core data sharing for motion compensation when decoding predicted coded pictures, which is usually much more than the inter-processor or inter-core data sharing due to intra-picture prediction). However, for the same reason, using regular stripes may incur a large amount of coding and decoding overhead due to the bit cost of the stripe header and the loss of prediction across stripe boundaries. In addition, due to the intra-picture independence of regular stripes and each regular stripe being encapsulated in its own NAL unit, regular stripes (compared to other tools mentioned below) can also serve as a key mechanism for bitstream segmentation to match the MTU size requirements. In many cases, the goals of parallelization and MTU size matching have conflicting requirements for the stripe layout in the picture. The implementation of this situation led to the development of the parallelization tools mentioned below.
[0079] Non-independent stripes have short stripe headers and allow the bitstream to be segmented at tree block boundaries without breaking any intra-picture prediction. Basically, non-independent stripes divide a regular stripe into multiple NAL units, reducing the end-to-end latency by allowing a part of the regular stripe to be sent before the encoding of the entire regular stripe is completed.
[0080] In WPP, a picture is segmented into single rows of coding tree blocks (CTBs). Entropy decoding and prediction are allowed to use data from CTBs in other segments. Parallel processing can be performed through the parallel decoding of CTB rows, where the decoding of a CTB row starts with a delay of two CTBs to ensure that data related to the CTBs above and to the right of the main CTB can be obtained before the main CTB being decoded. Using this staggered start (which looks like a wavefront when represented graphically), as many processors / cores can be parallelized as there are CTB rows in the picture. Since intra-picture prediction between adjacent tree block rows within the picture is allowed, the inter-processor / inter-core communication required to enable intra-picture prediction can be substantial. Compared to not applying WPP segmentation, WPP segmentation does not result in the generation of additional NAL units, so WPP is not a tool for MTU size matching. However, if MTU size matching is required, regular stripes can be used together with WPP, but with some coding and decoding overhead.
[0081] Slice definitions define the horizontal and vertical boundaries that divide a picture into slice columns and slice rows. Slice columns extend from the top to the bottom of the picture. Similarly, slice rows extend from the left to the right of the picture. The number of slices in a picture can be simply derived by multiplying the number of slice columns by the number of slice rows.
[0082] Before decoding the top-left CTB of the next slice in the order of slice raster scan of the picture, the scan order of CTBs is changed to the local scan order within the slice (in the order of CTB raster scan of the slice). Similar to conventional stripes, slices break the prediction dependency and entropy decoding dependency within the picture. However, they do not need to be included in separate NAL units (the same as WPP in this regard); thus, slices cannot be used for MTU size matching. Each slice can be processed by one processor / core, and in the case where a stripe spans multiple slices, the inter-processor / inter-core communication required for intra-picture prediction between processing units decoding adjacent slices is limited to transmitting the shared stripe header and loop filtering related to the sharing of reconstructed samples and metadata. When a stripe contains more than one slice or WPP segment, the entry point byte offset for each slice or WPP segment except the first one in the stripe is signaled in the stripe header.
[0083] For simplicity, restrictions on the application of four different picture partitioning schemes are specified in HEVC. For most profiles specified in HEVC, a given coded video sequence cannot contain both slices and wavefronts simultaneously. For each stripe and slice, one or both of the following conditions must be satisfied: 1) all coded tree blocks in the stripe belong to the same slice; 2) all coded tree blocks in the slice belong to the same stripe. Finally, a wavefront segment contains exactly one CTB row, and when using WPP, if a stripe starts within a CTB row, the stripe must end in the same CTB row.
[0084] The most recent revision of HEVC is specified in the JCT-VC output document JCTVC-AC1005, J. Boyce, A. Ramasubramonian, R. Skupin, G. J. Sullivan, A. Tourapis, Y.-Kwang (editors). "HEVC Additional Supplemental Enhancement Information (Draft 4)", publicly available here on October 24, 2017: http: / / phenix.intevry.fr / jct / doc_end_user / documents- / 29_Macau / wg11 / JCTVC-AC1005-v2.zip. Included within this revision, HEVC specifies three SEI messages related to MCT, namely the Temporal MCTS SEI message, the MCTS Extraction Information Set SEI message, and the MCTS Extraction Information Nested SEI message.
[0085] The time-domain MCTS SEI message indicates the presence of MCTS in the bitstream and signals the MCTS. For each MCTS, the motion vectors are restricted to point to full-sample positions within the MCTS and fractional-sample positions that only require full-sample positions within the MCTS for interpolation, and motion vector candidates predicted from time-domain motion vectors derived from blocks outside the MCTS are not allowed. In this way, each MCTS can be independently decoded without the presence of slices not included in the MCTS.
[0086] The MCTS extraction information set SEI message provides supplementary information that can be used for MCTS sub-bitstream extraction (the part specified as the semantics of the SEI message) to generate a compliant bitstream for the MCTS set. This information consists of multiple extraction information sets, each extraction information set defining multiple MCTS sets and containing the RBSP bytes to replace the VPS, SPS, and PPS to be used during the MCTS sub-bitstream extraction process. When extracting the sub-bitstream according to the MCTS sub-bitstream extraction process, the parameter sets (VPS, SPS, and PPS) need to be rewritten or replaced, and the slice header needs to be slightly updated because the strip addresses associated with one or all of the syntax elements (including first_slice_segment_in_pic_flag and slice_segment_address) usually need to have different values.
[0087] 3.2. Segmentation of Pictures in VVC
[0088] In VVC, a picture is divided into one or more slice rows and one or more slice columns. A slice is a sequence of CTUs that cover a rectangular region of the picture. The CTUs within a slice are scanned in raster scan order within that slice.
[0089] A strip consists of an integer number of complete slices or an integer number of consecutive complete CTU rows within the slices of a picture.
[0090] Two strip modes are supported, namely the raster scan strip mode and the rectangular strip mode. In the raster scan strip mode, a strip contains the sequence of complete slices in the picture slice raster scan. In the rectangular strip mode, a strip contains multiple complete slices that together form a rectangular region of the picture or multiple consecutive complete CTU rows of a single slice that together form a rectangular region of the picture. The slices within a rectangular strip are scanned in slice raster scan order within the rectangular region corresponding to that strip.
[0091] A sub-picture contains one or more strips that together cover a rectangular region of the picture.
[0092] Figure 1 An example of the raster scan strip segmentation of a picture is shown, where the picture is divided into 12 slices and 3 raster scan strips.
[0093] Figure 2 An example of rectangular strip segmentation of a picture is shown, where the picture is segmented into 24 slices (6 slice columns and 4 slice rows) and 9 rectangular slices.
[0094] Figure 3 An example of a picture segmented into slices and rectangular strips is shown, where the picture is divided into 4 slices (2 slice columns and 2 slice rows) and 4 rectangular strips.
[0095] Figure 4 An example of sub - picture segmentation of a picture is shown, where the picture is segmented into 18 slices, 12 slices on the left (each covering a strip with 4x4 CTUs) and 6 slices on the right (each covering 2 vertically stacked strips with 2x2 CTUs), resulting in a total of 24 strips and 24 sub - pictures of different dimensions (each strip is a sub - picture).
[0096] 3.3. Signaling of SPS / PPS / Picture Header / Strip Header in VVC (such as JVET - Q2001 - vC)
[0097] 7.3.2.3 Sequence Parameter Set RBSP Syntax
[0098]
[0099]
[0100]
[0101]
[0102]
[0103]
[0104]
[0105]
[0106]
[0107] 7.3.2.4 Picture Parameter Set RBSP Syntax
[0108]
[0109]
[0110]
[0111]
[0112]
[0113]
[0114] 7.3.2.7 Picture Header Structure Syntax
[0115]
[0116]
[0117]
[0118]
[0119]
[0120]
[0121] 7.3.7.1 Regular Strip Header Syntax
[0122]
[0123]
[0124]
[0125]
[0126]
[0127] 3.4. Specifications for Slices, Strips, and Sub - pictures in JVET - Q2001 - vC
[0128] 3 Definitions
[0129] Picture - level strip index: Index of the strip into the list of strips in the picture, in the order they are signaled in the PPS when rect_slice_flag equals 1.
[0130] Sub - picture - level strip index: Index of the strip into the list of strips in the sub - picture, in the order they are signaled in the PPS when rect_slice_flag equals 1.
[0131] 6.5.1 CTB Raster Scan, Slice Scan, and Sub - picture Scan Processes
[0132] The variable NumTileColumns specifies the number of tile columns, and the list colWidth[i] for i ranging from 0 to NumTileColumn - 1 (including the end values) specifies the width of the i - th tile column in CTBs, derived as follows:
[0133]
[0134] The variable NumTileRows specifies the number of tile rows, and the list RowHeight[j] for j ranging from 0 to NumTileRows-1 (including the end values) specifies the height of the j-th tile row in units of CTB, and is derived as follows:
[0135]
[0136]
[0137] The variable NumTilesInPic is set to be equal to NumTileColumns * NumTileRows.
[0138] The list tileColBd[i] for i ranging from 0 to NumTileColumns (including the end values) specifies the position of the i-th tile column boundary in units of CTB, and is derived as follows:
[0139] for (tileColBd[0] = 0, i = 0; i < NumTileColumns; i++)
[0140] tileColBd[i + 1] = tileColBd[i] + colWidth[i] (25)
[0141] Note 1 – The size of the array tileColBd[] is 1 larger than the actual number of tile columns in the derivation of CtbToTileColBd[].
[0142] The list tileRowBd[j] for j ranging from 0 to NumTileRows (including the end values) specifies the position of the j-th tile row boundary in units of CTB, and is derived as follows:
[0143] for (tileRowBd[0] = 0, j = 0; j < NumTileRows; j++)
[0144] tileRowBd[j + 1] = tileRowBd[j] + RowHeight[j] (26)
[0145] Note 2 – The size of the array tileRowBd[] in the above derivation is 1 larger than the actual number of tile rows in the derivation of CtbToTileRowBd[].
[0146] The list CtbToTileColBd[ctbAddrX] of ctbAddrX ranging from 0 to PicWidthInCtbsY (including the end values) specifies the conversion from the horizontal CTB address to the left tile column boundary in terms of CTBs, and is derived as follows:
[0147]
[0148] Note 3 – The size of the array CtbToTileColBd[] in the above derivation is 1 larger than the actual number of picture widths in CTBs in the derivation of slice_data() signaling.
[0149] The list CtbToTileRowBd[ctbAddrY] of ctbAddrY ranging from 0 to PicHeightInCtbsY (including the end values) specifies the conversion from the vertical CTB address to the top tile column boundary in terms of CTBs, and is derived as follows:
[0150]
[0151] Note 4 – The size of the array CtbToTileRowBd[] in the above derivation is 1 larger than the actual number of picture heights in CTBs in the slice_data() signaling notification.
[0152] For rectangular stripes, the list NumCtusInSlice[i] of i ranging from 0 to num_slices_in_pic_minus1 (including the end values) specifies the number of CTUs in the i-th stripe, the list SliceTopLeftTileIdx[i] of i ranging from 0 to num_slices_in_pic_minus1 (including the end values) specifies the index of the left-top tile of the stripe, and the matrix of i ranging from 0 to num_slices_in_pic_minus1 (including the end values) and j ranging from 0 to NumCtusInSlice[i] – 1 (including the end values) specifies the picture raster scan address of the j-th CTB within the i-th stripe, and is derived as follows:
[0153]
[0154]
[0155] where the function AddCtbsToSlice(sliceIdx,startX,stopX,startY,stopY) is specified as follows:
[0156]
[0157] The requirement for bitstream consistency is that the value of NumCtusInSlice[i] for i ranging from 0 to num_slices_in_pic_minus1 (including the end values) should be greater than 0. Additionally, the requirement for bitstream consistency is that the matrix CtbAddrInSlice[i][j] for i ranging from 0 to num_slices_in_pic_minus1 (including the end values) and j ranging from 0 to NumCtusInSlice[i]-1 (including the end values) should include all CTB addresses in the range from 0 to PicSizeInCtbsY-1 exactly once and only once.
[0158] The list CtbToSubpicIdx[ctbAddrRs] for ctbAddrRs ranging from 0 to PicSizeInCtbsY-1 (including the end values) specifies the conversion from CTB addresses in the picture raster scan to sub-picture indices, which is derived as follows:
[0159]
[0160] The list NumSlicesInSubpic[i] specifies the number of rectangular stripes in the i-th sub-picture, which is derived as follows:
[0161]
[0162] …
[0163] 7.3.4.3 Picture Parameter Set RBSP Semantics
[0164] …
[0165] subpic_id_mapping_in_pps_flag being equal to 1 specifies signaling the sub-picture ID mapping in the pps. subpic_id_mapping_in_pps_flag being equal to 0 specifies not signaling the sub-picture ID mapping in the PPS. If subpic_id_mapping_explicitly_signaled_flag is 0 or subpic_id_mapping_in_sps_flag is equal to 1, the value of subpic_id_mapping_in_pps_flag should be equal to 0. Otherwise (sub pic_id_mapping_explicitly_signaled_flag is equal to 1 and subpic_id_mapping_in_sps_flag is equal to 0), the value of subpic_id_mapping_in_pps_flag should be equal to 1.
[0166] pps_num_subpics_minus1 shall be equal to sps_num_subpics_minus1.
[0167] pps_subpic_id_len_minus1 shall be equal to sps_subpic_id_len_minus1.
[0168] pps_subpic_id[i] specifies the sub-picture ID of the i-th sub-picture. The length of the pps_subpic_id[i] syntax element is pps_subpic_id_len_minus1 + 1 bits.
[0169] For each value of i in the range from 0 to sps_num_subpics_minus1, the variable SubpicIdVal[i] is derived as follows:
[0170]
[0171] The requirements for bitstream consistency are to apply the following two constraints:
[0172] -- For any two different values of i and j in the range from 0 to sps_num_subpics_minus1 (including the end values), SubpicIdVal[i] shall not be equal to SubpicIdVal[j].
[0173] -- When the current picture is not the first picture of CLVS, for each value of i in the range from 0 to sps_num_subpics_minus1 (including the end values), if the value of SubpicIdVal[i] is not equal to the value of SubpicIdVal[i] of the previous picture in decoding order in the same layer, then the nal_unit_type of all coded and decoded slice NAL units of the sub-picture in the current picture with sub-picture index i shall be equal to a specific value in the range from IDR_W_RADL to CRA_NUT (including the end values).
[0174] no_pic_partition_flag being equal to 1 specifies that picture partitioning is not applied to each picture of the reference PPS. no_pic_partition_flag being equal to 0 specifies that each picture of the reference PPS can be partitioned into more than one slice or strip.
[0175] The requirement for bitstream consistency is that for all PPSs referred to by coded pictures within CLVS, the value of no_pic_partition_flag shall be the same.
[0176] The requirement for bitstream consistency is that when the value of sps_num_subpics_minus1 + 1 is greater than 1, the value of no_pic_partition_flag shall not be equal to 1.
[0177] pps_log2_ctu_size_minus5 plus 5 specifies the size of the luma coding tree block for each CTU. pps_log2_ctu_size_minus5 shall be equal to sps_log2_ctu_size_minus5.
[0178] num_exp_tile_columns_minus1 plus 1 specifies the number of explicitly provided tile column widths. The value of num_exp_tile_columns_minus1 shall be in the range of 0 to PicWidthInCtbsY – 1 (including the end values). When no_pic_partition_flag is equal to 1, the value of num_exp_tile_columns_minus1 is inferred to be equal to 0.
[0179] num_exp_tile_rows_minus1 plus 1 specifies the number of explicitly provided tile row heights. The value of num_exp_tile_rows_minus1 shall be in the range of 0 to PicHeightInCtbsY – 1 (including the end values). When no_pic_partition_flag is equal to 1, the value of num_tile_rows_minus1 is inferred to be equal to 0.
[0180] tile_column_width_minus1[i] plus 1 specifies the width of the i-th tile column in units of CTB, where i ranges from 0 to num_exp_tile_columns_minus1 - 1 (including the end values).
[0181] tile_column_width_minus1[num_exp_tile_columns_minus1] is used to derive the width of tile columns with indices greater than or equal to num_exp_tile_columns_minus1, as specified in Clause 6.5.1. The value of tile_column_width_minus1[i] shall be in the range of 0 to PicWidthInCtbsY – 1 (including the end values). When not present, the value of tile_column_width_minus1[0] is inferred to be equal to PicWidthInCtbsY - 1.
[0182] tile_row_height_minus1[i] plus 1 specifies the height of the i-th tile row in units of CTB, where i ranges from 0 to num_exp_tile_rows_minus1–1 (inclusive).
[0183] tile_row_height_minus1[num_exp_tile_rows_minus1] is used to derive the height of tile rows with indices greater than or equal to num_exp_tile_rows_minus1, as specified in Clause 6.5.1. The value of tile_row_height_minus1[i] shall be in the range of 0 to PicHeightInCtbsY–1 (inclusive). When absent, the value of tile_row_height_minus1[0] is inferred to be equal to PicHeightInCtbsY - 1.
[0184] rect_slice_flag equal to 0 specifies that the tiles within each slice are in raster scan order and the slice information is not signaled in the PPS. rect_slice_flag equal to 1 specifies that the tiles within each slice cover a rectangular region of the picture and the slice information is signaled in the PPS. When absent, rect_slice_flag is inferred to be equal to 1. When subpic_info_present_flag is equal to 1, the value of rect_slice_flag shall be equal to 1.
[0185] single_slice_per_subpic_flag equal to 1 specifies that each sub-picture consists of one and only one rectangular slice. single_slice_per_subpic_flag equal to 0 specifies that each sub-picture can consist of one or more rectangular slices. When single_slice_per_subpic_flag is equal to 1, num_slices_in_pic_minus1 is inferred to be equal to sps_num_subpics_minus1. When absent, the value of single_slice_per_subpic_flag is inferred to be equal to 0.
[0186] num_slices_in_pic_minus1 plus 1 specifies the number of rectangular stripes in each picture of the reference PPS. The value of num_slices_in_pic_minus1 shall be in the range of 0 to MaxSlicesPerPicture – 1 (inclusive), where MaxSlicesPerPicture is specified in Annex A. When no_pic_partition_flag is equal to 1, the value of num_slices_in_pic_minus1 is inferred to be equal to 0.
[0187] tile_idx_delta_present_flag being equal to 0 specifies that the tile_idx_delta value does not exist in the PPS, and all rectangular stripes in the picture of the reference PPS are specified in raster scan order according to the process defined in Clause 6.5.1.
[0188] tile_idx_delta_present_flag being equal to 1 specifies that the tile_idx_delta value may exist in the PPS, and all rectangular stripes in the picture of the reference PPS are specified in the order indicated by the value of tile_idx_delta. When it does not exist, the value of tile_idx_delta_present_flag is inferred to be equal to 0.
[0189] slice_width_in_tiles_minus1[i] plus 1 specifies the width of the i-th rectangular stripe in terms of tile columns. The value of slice_width_in_tiles_minus1[i] shall be in the range of 0 to NumTileColumns – 1 (inclusive).
[0190] When slice_width_in_tiles_minus1[i] does not exist, the following applies:
[0191] -- If NumTileColumns is equal to 1, the value of slice_width_in_tiles_minus1[i] is inferred to be equal to 0.
[0192] -- Otherwise, the value of slice_width_in_tiles_minus1[i] is inferred as specified in Clause 6.5.1.
[0193] slice_height_in_tiles_minus1[i] plus 1 specifies the height of the i-th rectangular strip in terms of tile rows. The value of slice_height_in_tiles_minus1[i] shall be in the range of 0 to NumTileRows–1 (including the end values).
[0194] When slice_height_in_tiles_minus1[i] does not exist, the following applies:
[0195] -- If NumTileRows is equal to 1, or tile_idx_delta_present_flag is equal to 0, and tileIdx % NumTileColumns is greater than 0, the value of slice_height_in_tiles_minus1[i] is inferred to be equal to 0.
[0196] -- Otherwise (NumTileRows is not equal to 1, and tile_idx_delta_present_flag is equal to 1 or tileIdx % NumTileColumns is equal to 0), when tile_idx_delta_present_flag is equal to 1 or tileIdx % NumTileColumns is equal to 0, the value of slice_height_in_tiles_minus1[i] is inferred to be equal to slice_height_in_tiles_minus1[i-1].
[0197] num_exp_slices_in_tile[i] specifies the number of strip heights explicitly provided in the current tile that contains multiple rectangular strips. The value of num_exp_slices_in_tile[i] shall be in the range of 0 to RowHeight[tileY]–1 (including the end values), where tileY is the strip row index that contains the i-th strip. When it does not exist, the value of num_exp_slices_in_tile[i] is inferred to be equal to 0. When num_exp_slices_in_tile[i] is equal to 0, the value of the variable NumSlicesInTile[i] is derived to be equal to 1.
[0198] exp_slice_height_in_ctus_minus1[j] plus 1 specifies the height of the j-th rectangular stripe in the current slice in CTU rows. The value of exp_slice_height_in_ctus_minus1[j] shall be in the range of 0 to RowHeight[tileY] – 1 (inclusive), where tileY is the slice row index of the current slice.
[0199] When num_exp_slices_in_tile[i] is greater than 0, the variables NumSlicesInTile[i] and SliceHeightInCtusMinus1[i + k] for k in the range of 0 to NumSlicesInTile[i] – 1 are derived as follows:
[0200]
[0201] tile_idx_delta[i] specifies the difference between the slice index of the first slice in the i-th rectangular stripe and the slice index of the first slice in the (i + 1)-th rectangular stripe. The value of tile_idx_delta[i] shall be in the range of -NumTilesInPic + 1 to NumTilesInPic – 1 (inclusive). When not present, the value of tile_idx_delta[i] is inferred to be equal to 0. When present, the value of tile_idx_delta[i] shall not be equal to 0.
[0202] …
[0203] 7.4.2.4.5 Order of VCL NAL Units and Their Association with Decoded / Encoded Pictures
[0204] The order of VCL NAL units within a decoded / encoded picture is constrained as follows:
[0205] -- For any two coded slice NAL units A and B of a decoded / encoded picture, let subpicIdxA and subpicIdxB be their sub-picture level index values, and sliceAddrA and sliceAddrB be their slice_address values.
[0206] -- The coded slice NAL unit A shall be before the coded slice NAL unit B when any of the following conditions is true:
[0207] – subpicIdxA is less than subpicIdxB.
[0208] -- subpicidxa is equal to subpicIdxB, and sliceAddrA is less than sliceAddrB.
[0209] 7.4.8.1 General Strip Header Semantics
[0210] The variable CuQpDeltaVal, which specifies the difference between the luma quantization parameter of a coding unit containing cu_qp_delta_abs and its prediction, is set to be equal to 0. The variables CuQpOffset Cb , CuQpOffset Cr , and CuQpOffset CbCr , which specify the values to be used when determining the respective values of the quantization parameters Qp′ Cb , Qp′ Cr , and Qp′ CbCr are all set to be equal to 0.
[0211] picture_header_in_slice_header_flag being equal to 1 specifies that the PH syntax structure is present in the strip header. picture_header_in_slice_header_flag being equal to 0 specifies that the PH syntax structure is not present in the strip header.
[0212] The requirement for bitstream conformance is that the value of picture_header_in_slice_header_flag should be the same for all coded strips in the CLVS.
[0213] When picture_header_in_slice_header_flag of a coded strip is equal to 1, the requirement for bitstream conformance is that no VCL NAL unit with nal_unit_type equal to PH_NUT should appear in the CLVS.
[0214] When picture_header_in_slice_header_flag is equal to 0, all coded strips in the current picture should have picture_header_in_slice_header_flag equal to 0, and the current PU should have a PH NAL unit.
[0215] The slice_subpic_id specifies the sub-picture ID of the sub-picture containing the slice. If slice_subpic_id exists, the value of the variable CurrSubpicIdx is derived such that SubpicIdVal[CurrSubpicIdx] is equal to slice_subpic_id. Otherwise (slice_subpic_id does not exist), CurrSubpicIdx is derived to be equal to 0. The length of slice_subpic_id is sps_subpic_id_len_minus1 + 1 bits.
[0216] The slice_address specifies the slice address of the slice. When it does not exist, the value of slice_address is inferred to be equal to 0. When rect_slice_flag is equal to 1 and NumSlicesInSubpic[CurrSubpicIdx] is equal to 1, the value of slice_address is inferred to be equal to 0.
[0217] If rect_slice_flag is equal to 0, the following applies:
[0218] - The slice address is the raster scan slice index.
[0219] - The length of slice_address is Ceil(Log2(NumTilesInPic)) bits.
[0220] - The value of slice_address shall be in the range of 0 to NumTilesInPic – 1 (including the end values).
[0221] Otherwise (rect_slice_flag is equal to 1), the following applies:
[0222] - The slice address is the slice sub-picture level slice index.
[0223] - The length of slice_address is Ceil(Log2(NumSlicesInSubpic[CurrSubpicIdx])) bits.
[0224] - The value of slice_address shall be in the range of 0 to NumSlicesInSubpic[CurrSubpicIdx] – 1 (including the end values).
[0225] The requirements for bitstream consistency are to apply the following constraints:
[0226] - If rect_slice_flag is equal to 0 or subpic_info_present_flag is equal to 0, the value of slice_address shall not be equal to the value of slice_address of any other coded slice NAL unit of the same coded picture.
[0227] - Otherwise, a pair of slice_subpic_id and slice_address values shall not be equal to a pair of slice_subpic_id and slice_address values of any other coded slice NAL unit of the same coded picture.
[0228] - The shape of the slices of a picture shall be such that when each CTU is decoded, its entire left and top boundaries are formed by the picture boundary or by the boundaries of previously decoded CTU(s).
[0229] sh_extra_bit[i] may be equal to 1 or 0. Decoders compliant with this version of this specification shall ignore the value of sh_extra_bit[i]. Its value does not affect the decoder's compliance with the profiles specified in this version of the specification.
[0230] num_tiles_in_slice_minus1 plus 1 (if present) specifies the number of tiles in the slice. The value of num_tiles_in_slice_minus1 shall be in the range of 0 to NumTilesInPic–1 (inclusive).
[0231] The variable NumCtusInCurrSlice specifies the number of CTUs in the current slice, and the list CtbAddrInCurrSlice[i] for i in the range from 0 to NumCtusInCurrSlice–1 (inclusive) specifies the picture raster scan address of the i-th CTB within the slice, derived as follows:
[0232]
[0233]
[0234] …
[0235] 3.5. Luminance Mapping with Chroma Scaling (LMCS)
[0236] LMCS consists of two aspects: luminance mapping (the transformation process, denoted as RP) and luminance-dependent chrominance residual scaling (CRS). For the luminance signal, the LMCS mode operates based on two domains, including the first domain as the original domain and the second domain as the shaping domain that maps luminance samples to specific values according to the shaping model. Additionally, for the chrominance signal, residual scaling can be applied, where the scaling factor is derived from the luminance samples.
[0237] The relevant syntax elements and semantic descriptions in the SPS, Picture Header (PH), and Slice Header (SH) are as follows:
[0238] Syntax table
[0239] 7.3.2.3 Sequence Parameter Set RBSP Syntax
[0240]
[0241] 7.3.2.7 Picture Header Structure Syntax
[0242]
[0243]
[0244] 7.3.7 Slice Header Syntax
[0245] 7.3.7.1 General Slice Header Syntax
[0246]
[0247] Semantics
[0248] The sps_lmcs_enabled_flag being equal to 1 specifies the use of luminance mapping with chrominance scaling in CLVS. The sps_lmcs_enabled_flag being equal to 0 specifies that luminance mapping with chrominance scaling is not used in CLVS.
[0249] The ph_lmcs_enabled_flag being equal to 1 specifies that luminance mapping with chrominance scaling is enabled for all slices related to the PH. The ph_lmcs_enabled_flag being equal to 0 specifies that luminance mapping with chrominance scaling can be disabled for one, multiple, or all slices related to the PH. When absent, the value of the ph_lmcs_enabled_flag is inferred to be equal to 0.
[0250] The ph_lmcs_aps_id specifies the adaptation_parameter_set_id of the LMCS APS that the strips associated with the PH refer to. The TemporalId of the APS NAL unit with aps_params_type equal to LMCS_APS and adaptation_parameter_set_id equal to ph_lmcs_aps_id should be less than or equal to the TemporalId of the picture associated with the PH.
[0251] The ph_chroma_residual_scale_flag being equal to 1 specifies that chroma residual scaling is enabled for all strips associated with the PH. The ph_chroma_residual_scale_flag being equal to 0 specifies that chroma residual scaling can be disabled for one, multiple, or all strips associated with the PH. When the ph_chroma_residual_scale_flag is absent, it is inferred to be equal to 0.
[0252] The slice_lmcs_enabled_flag being equal to 1 specifies that luminance mapping with chroma scaling is enabled for the current strip. The slice_lmcs_enabled_flag being equal to 0 specifies that luminance mapping with chroma scaling is not enabled for the current strip. When the slice_lmcs_enabled_flag is absent, it is inferred to be equal to 0.
[0253] 3.6. Adaptive Motion Vector Difference Precision (AMVR) of Affine Coding / Decoding Blocks
[0254] Affine AMVR is a coding / decoding tool that allows affine inter-coding blocks to send MV differences with different precisions, such as with a precision of 1 / 4 luminance samples (default, amvr_flag set to 0), 1 / 16 luminance samples, and 1 luminance sample.
[0255] The relevant syntax elements and semantics in the SPS are described as follows:
[0256] Syntax table
[0257] 7.3.2.3 Sequence Parameter Set RBSP Syntax
[0258]
[0259]
[0260] Semantics
[0261] The sps_affine_amvr_enabled_flag being equal to 1 specifies the use of adaptive motion vector difference resolution in the motion vector coding / decoding of the affine inter prediction mode. The sps_affine_amvr_enabled_flag being equal to 0 specifies not to use the adaptive motion vector difference resolution in the motion vector coding / decoding of the affine inter prediction mode. When not present, the value of the sps_affine_amvr_enabled_flag is inferred to be equal to 0.
[0262] D.7 Sub-picture level information SEI message
[0263] D.7.1 Sub-picture level information SEI message syntax
[0264]
[0265] D.7.2 Sub-picture level information SEI message semantics
[0266] The sub-picture level information SEI message contains information about the level to which the sub-picture sequence in the bitstream conforms when checking the consistency of the bitstream containing the sub-picture sequence extracted according to Annex A testing.
[0267] When the sub-picture level information SEI message is present in any picture of the CLVS, the sub-picture level information SEI message will be present in the first picture of the CLVS. The sub-picture level information SEI message continues from the current picture to the current layer in decoding order until the end of the CLVS. All sub-picture level information SEI messages applicable to the same CLVS shall have the same content. A sub-picture sequence consists of all sub-pictures within the CLVS having the same sub-picture index value.
[0268] The requirement for bitstream consistency is that when the sub-picture level information SEI message of the CLVS is present, for each i value in the range from 0 to sps_num_subpics_minus1 (inclusive of the end values), the value of subpic_treated_as_pic_flag[i] shall be equal to 1.
[0269] num_ref_levels_minus1 plus 1 specifies the number of reference levels signaled for each of the sps_num_subpics_minus1 + 1 sub-pictures.
[0270] When sli_cbr_constraint_flag equals 0, it is specified that for decoding the sub-bitstream generated for any sub-picture of the extracted bitstream according to Clause C.7 by using the HRD of any CPB specification in the extracted sub-bitstream, the hypothetical stream scheduler (HSS) operates in the intermittent bitrate mode. When sli_cbr_constraint_flag equals 1, it is specified that the HSS operates in the constant bitrate (CBR) mode.
[0271] When explicit_fraction_present_flag equals 1, it is specified that the syntax element ref_level_fraction_minus1[i] exists. When explicit_fraction_present_flag equals 0, it is specified that the syntax element ref_level_fraction_minus1[i] does not exist.
[0272] sli_num_subpics_minus1 plus 1 specifies the number of sub-pictures in the picture of CLVS. When present, the value of sli_num_subpics_minus1 shall be equal to the value of sps_num_subpics_minus1 in the SPS referred to by the picture in CLVS.
[0273] sli_alignment_zero_bit shall be equal to 0.
[0274] ref_level_idc[i] indicates the level that each sub-picture conforms to as specified in Annex A. Except for the values specified in Annex A, the bitstream shall not contain the value of ref_level_idc. Other values of ref_level_idc[i] are reserved for future use by ITU-T|ISO / IEC. The requirement for bitstream conformance is that for any k value greater than i, the value of ref_level_idc[i] shall be less than or equal to ref_level_idc[k].
[0275] ref_level_fraction_minus1[i][j] plus 1 specifies the fraction of the level limit associated with ref_level_idc[i] that the j-th sub-picture conforms to, as specified in Clause A.4.1.
[0276] The variable SubpicSizeY[j] is set to be equal to (subpic_width_minus1[j] + 1) * CtbSizeY * (subpic_height_minus1[j] + 1) * CtbSizeY.
[0277] When it does not exist, the value of ref_level_fraction_minus1[i][j] is inferred to be equal to Ceil(256 * SubpicSizeY[j] ÷ PicSizeInSamplesY * MaxLumaPs(general_level_idc) ÷ MaxLumaPs(ref_level_idc[i]) - 1).
[0278] The variable RefLevelFraction[i][j] is set to be equal to ref_level_fraction_minus1[i][j] + 1.
[0279] The derivation of the variables SubpicNumTileCols[j] and SubpicNumTileRows[j] is as follows:
[0280]
[0281]
[0282] The derivation of the variables SubpicCpbSizeVcl[i][j] and SubpicCpbSizeNal[i][j] is as follows:
[0283] SubpicCpbSizeVcl[i][j] =
[0284] Floor(CpbVclFactor * MaxCPB * RefLevelFraction[i][j] ÷ 256) (D.6)
[0285] SubpicCpbSizeNal[i][j] =
[0286] Floor(CpbNalFactor * MaxCPB * RefLevelFraction[i][j] ÷ 256) (D.7) Wherein, MaxCPB is derived from ref_level_idc[i] as specified in Clause A.4.2.
[0287] The derivation of the variables SubpicBitRateVcl[i][j] and SubpicBitRateNal[i][j] is as follows:
[0288] SubpicBitRateVcl[i][j] =
[0289] Floor(CpbVclFactor*MaxBR*RefLevelFraction[i][j]÷256) (D.8)
[0290] SubpicBitRateNal[i][j] =
[0291] Floor(CpbNalFactor*MaxBR*RefLevelFraction[i][j]÷256) (D.9) where, as specified in Clause A.4.2, MaxBR is derived from ref_level_idc[i].
[0292] Note 1 – When extracting a sub-picture, the resulting bitstream has a CpbSize (indicated or inferred in the SPS) greater than or equal to SubpicCpbSizeVcl[i][j] and SubpicCpbSizeNal[i][j], and a bitrate (indicated or inferred in the SPS) greater than or equal to SubpicBitRateVcl[i][j] and SubpicBitRateNal[i][j].
[0293] The requirement for bitstream conformance is that the bitstream generated from the extraction of the j-th sub-picture with j ranging from 0 to sps_num_subpics_minus1 (including the end values), and conforming to the profile with general_tier_flag equal to 0 and level equal to ref_level_idc[i] with i ranging from 0 to num_ref_level_minus1 (including the end values), shall comply with the following constraints for each bitstream conformance test specified in Appendix C:
[0294] - Ceil(256*SubpicSizeY[j]÷RefLevelFraction[i][j]) shall be less than or equal to MaxLumaPs, where MaxLumaPs is specified in Table A.1 for level ref_level_idc[i].
[0295] - The value of Ceil(256*(subpic_width_minus1[j]+1)*CtbSizeY÷RefLevelFraction[i][j]) shall be less than or equal to Sqrt(MaxLumaPs*8).
[0296] - The value of Ceil(256*(subpic_height_minus1[j]+1)*CtbSizeY÷RefLevelFraction[i][j]) should be less than or equal to Sqrt(MaxLumaPs*8).
[0297] - The value of SubpicNumTileCols[j] should be less than or equal to MaxTileCols, and the value of SubpicNumTileRows[j] should be less than or equal to MaxTileRows, where MaxTileCols and MaxTileRows are specified in Table A.1 for level ref_level_idc[i].
[0298] - The value of SubpicNumTileCols[j]*SubpicNumTileRows[j] should be less than or equal to MaxTileCols*MaxTileRows*RefLevelFraction[i][j], where MaxTileCols and MaxTileRows are specified in Table A.1 for level ref_level_idc[i].
[0299] - For the value of SubpicSizeInSamplesY of AU 0, the sum of the NumBytesInNalUnit variables of AU 0 corresponding to the j-th subpicture should be less than or equal to FormatCapabilityFactor*(Max(SubpicSizeY[j],fR*MaxLumaSr*RefLevelFraction[i][j]÷256)+MaxLumaSr*(AuCpbRemovalTime[0]-AuNominalRemovalTime[0])*RefLevelFraction[i][j])÷(256*MinCr), where MaxLumaSr and FormatCapabilityFactor are the values specified in Tables A.2 and A.3 respectively, applicable to AU 0 of ref_level_idc[i] level, and MinCr is derived as shown in A.4.2.
[0300] The sum of the NumBytesInNalUnit variables for the AUs n (where n > 0) corresponding to the j-th sub-picture shall be less than or equal to FormatCapabilityFactor * MaxLumaSr * (AuCpbRemovalTime[n] - AuCpbRemovalTime[n - 1]) * RefLevelFraction[i][j] ÷ (256 * MinCr), where MaxLumaSr and FormatCapabilityFactor are the values specified in Tables A.2 and A.3, respectively, for the AU n at the ref_level_idc[i] level, and MinCr is derived as shown in A.4.2.
[0301] For any sub-picture set that contains one or more sub-pictures and consists of multiple sub-pictures in the sub-picture index SubpicSetIndices list and the sub-picture set NumSubpicsInSet, derive the level information of the sub-picture set.
[0302] The variable SubpicSetAccLevelFraction[i] for the total level fraction relative to the reference level ref_level_idc[i], and the variables SubpicSetCpbSizeVcl[i], SubpicSetCpbSizeNal[i], SubpicSetBitRateVcl[i], and SubpicSetBitRateNal[i] of the sub-picture set are derived as follows:
[0303]
[0304] The value of the sub-picture set sequence level indicator SubpicSetLevelIdc is derived as follows:
[0305]
[0306] where MaxTileCols and MaxTileRows are specified for ref_level_idc[i] in Table A.1.
[0307] The sub-picture set bitstream conforming to the configuration document with general_tier_flag equal to 0 and level equal to SubpicSetLevelIdc shall comply with the following constraints for each bitstream conformance test specified in Appendix C:
[0308] - For VCL HRD parameters, SubpicSetCpbSizeVcl[i] shall be less than or equal to CpbVclFactor * MaxCPB, where CpbVclFactor is specified in Table A.3 and MaxCPB is specified in Table A.1 in units of CpbVclFactor bits.
[0309] - For NAL HRD parameters, SubpicSetCpbSizeNal[i] shall be less than or equal to CpbNalFactor * MaxCPB, where CpbNalFactor is specified in Table A.3 and MaxCPB is specified in Table A.1 in units of CpbNalFactor bits.
[0310] - For VCL HRD parameters, SubpicSetBitRateVcl[i] shall be less than or equal to CpbVclFactor * MaxBR, where CpbVclFactor is specified in Table A.3 and MaxBR is specified in Table A.1 in units of CpbVclFactor bits.
[0311] - For NAL HRD parameters, SubpicSetBitRateNal[i] shall be less than or equal to CpbNalFactor * MaxCR, where CpbNalFactor is specified in Table A.3 and MaxBR is specified in Table A.1 in units of CpbNalFactor bits.
[0312] Note 2 – When extracting a sub-picture set, the resulting bitstream has a CpbSize (indicated or inferred in the SPS) greater than or equal to SubpicSetCpbSizeVcl[i][j] and SubpicSetCpbSizeNal[i][j], and a bitrate (indicated or inferred in the SPS) greater than or equal to SubpicSetBitRateVcl[i][j] and SubpicSetBitRateNal[i][j].
[0313] 4. Technical problems solved by the disclosed technical solution
[0314] The existing designs of sub-pictures and LMCS in VVC have the following problems:
[0315] 1) The derivation of the list SubpicNumTileRows[] (which specifies the number of tile rows included in a sub - picture) is incorrect in Equation D.5 because the index value idx in CtbToTileRowBd[idx] in the equation may be greater than the maximum allowed value. Additionally, the deviation for both SubpicNumTileRows[] and SubpicNumTileCols[] (which specifies the number of tile columns included in a sub - picture) uses CTU - based operations, which is unnecessarily complex.
[0316] 2) When single_slice_per_subpic_flag is equal to 1, the derivation of the array CtbAddrInSlice in Equation 29 is incorrect because the values of the raster - scanned CTB addresses in the array for each slice need to be in the CTU decoding order, rather than in the CTU raster - scan order.
[0317] 3) The LMCS signaling notification is inefficient. When ph_lmcs_enabled_flag is equal to 1, in most cases, LMCS will be enabled for all slices of a picture. However, in the current VVC design, for the case where LMCS is enabled for all slices of a picture, not only is ph_lmcs_enabled_flag equal to 1, but also the slice_lmcs_enabled_flag with a value of 1 needs to be signaled for each slice.
[0318] a. When ph_lmcs_enabled_flag is true, the semantics of ph_lmcs_enabled_flag conflict with the motivation of signaling the slice - level LMCS flag. In the current VVC, when ph_lmcs_enabled_flag is true, it means that all slices should have LMCS enabled. Therefore, there is no need to also signal the LMCS enable flag in the slice header.
[0319] b. Additionally, when the picture header indicates that LMCS is enabled, generally, LMCS is enabled for all slices. The control of LCMS in the slice header is mainly for handling extreme cases. Therefore, if the PH LMCS flag is true and the SH LMCS flag is always signaled, this may result in signaling unnecessary bits for common user cases.
[0320] 4) The semantics of the SPS affine AMVR flag is incorrect because for each CU with affine inter - coding, affine AMVR can be enabled or disabled.
[0321] 5. Examples of Techniques and Embodiments
[0322] To solve the above problems and some other problems not mentioned, the following summarized methods are disclosed. These items should be considered as examples to explain general concepts and should not be interpreted in a narrow way. In addition, these items can be applied individually or in any combination.
[0323] Related to sub - pictures for solving the first and second problems
[0324] 1. One or more of the following methods are disclosed:
[0325] a. Derive the slice column index for each CTU column of the picture.
[0326] b. The derivation of the number of slice columns included in the sub-picture is based on the slice column indices of the leftmost and / or rightmost CTUs included in the sub-picture.
[0327] c. Derive the slice row index for each CTU row of the picture.
[0328] d. The derivation of the number of slice rows included in the sub-picture is based on the slice row indices of the top and / or bottom CTUs included in the sub-picture.
[0329] e. The term picture-level strip index is defined as follows:
[0330] The index from the strip defined when rect_slice_flag equals 1 to the list of strips in the picture, in the order the strips are signaled in the PPS when single_slice_per_subpic_flag equals 0, or in the order of increasing sub-picture indices of the sub-pictures corresponding to the strips when single_slice_per_subpic_flag equals 1.
[0331] f. In one example, when the sub-picture contains strips split from slices, the height of the sub-picture cannot be calculated based on the slices.
[0332] g. In one example, the height of the sub-picture can be calculated based on CTUs instead of slices.
[0333] h. Derive whether the height of the sub-picture is less than one slice row.
[0334] i. In one example, when the sub-picture includes only CTUs from one slice row and when the top CTU in the sub-picture is not the top CTU of the slice row or the bottom CTU in the sub-picture is not the bottom CTU of the slice row, it is derived that the height of the sub-picture is less than one slice row is true.
[0335] ii. When it is indicated that each sub - picture contains only one stripe and the height of the sub - picture is less than one slice row, for each stripe of the picture with picture - level stripe index i, the value of CtbAddrInSlice[i][j] for j in the range from 0 to the number of CTUs in the stripe minus 1 (inclusive) is derived as the picture raster - scan CTU address of the j - th CTU in the CTU raster - scan of the sub - picture.
[0336] iii. In one example, when the distance between the top CTU and the bottom CTU in the sub - picture is less than the height of the slice according to the CTU, whether the height of the sub - picture is less than one slice row is derived as true.
[0337] iv. When it is indicated that each sub - picture contains only one stripe and the height of the sub - picture is greater than or equal to one slice row, for each stripe of the picture with picture - level stripe index i, the value of CtbAddrInSlice[i][j] for j in the range from 0 to the number of CTUs in the stripe minus 1 (inclusive) is derived as the picture raster - scan CTU address of the j - th CTU, and the order of the CTUs is as follows:
[0338] 1) The CTUs in different slices in the sub - picture are sorted such that the first CTU in the first slice with a smaller slice index value comes before the second CTU in the second slice with a larger slice index value.
[0339] 2) The CTUs within a slice in the sub - picture are sorted in the CTU raster - scan of the slice.
[0340] Related to LMCS for solving the third problem (including sub - problems)
[0341] 2. Two - level control of LMCS is introduced (which includes two aspects: luminance mapping (the transformation process, represented by RP) and luminance - dependent chrominance residual scaling (CRS)), where higher - level (e.g., picture - level) and lower - level (e.g., stripe - level) control are used, and the existence of lower - level control information depends on higher - level control information. In addition, the following also applies:
[0342] a. In the first example, one or more of the following sub - items are applied:
[0343] i. The first indicator (e.g., ph_lmcs_enabled_type) can be signaled at a higher level (e.g., in the picture header (PH)) to specify how to enable LMCS at a lower level for non - binary values.
[0344] 1) In one example, when the first indicator is equal to X (e.g., X = 2), it specifies that LMCS is enabled for all the stripes associated with PH; when the first indicator is equal to Y (Y != X) (e.g., Y = 1), it specifies that LMCS is enabled for one or more but not all of the stripes associated with PH; when the first indicator is equal to Z (Z != X and Z != Y) (e.g., Z = 0), it specifies that LMCS is disabled for all the stripes associated with PH.
[0345] a) Alternatively, in addition, when the first indicator is absent, the value of the indicator is inferred to be equal to a default value, e.g., Z.
[0346] 2) In one example, when the first indicator is equal to X (e.g., X = 2), it specifies that LMCS is disabled for all the stripes associated with PH; when the first indicator is equal to Y (Y != X) (e.g., Y = 1), it specifies that LMCS is disabled for one or more but not all of the stripes associated with PH; when the first indicator is equal to Z (Z != X and Z != Y) (e.g., Z = 0), it specifies that LMCS is enabled for all the stripes associated with PH.
[0347] a) Alternatively, in addition, when the first indicator is absent, the value of the indicator is inferred to be equal to a default value, e.g., X.
[0348] 3) Alternatively, in addition, the first indicator can be signaled conditionally according to the value of the LMCS enable flag (e.g., sps_lmcs_enabled_flag) in the sequence level.
[0349] 4) Alternatively, in addition, the first indicator can be coded / decoded using u(v), or u(2) or ue(v).
[0350] 5) Alternatively, in addition, the first indicator can be coded / decoded using a truncated unary code.
[0351] 6) Alternatively, in addition, the LMCS APS information (e.g., ph_lmcs_aps_id) used by the stripe and / or CS enable flag (e.g., ph_chroma_residual_scale_flag) can be signaled under the condition check of the value of the first indicator.
[0352] ii. A second indicator (e.g., slice_lmcs_enabled_flag) for enabling / disabling LMCS at a lower level (e.g., in the stripe header) can be signaled, and this second indicator can be signaled conditionally by checking the value of the first indicator.
[0353] 1) In one example, the second indicator may be signaled under the condition check of "the first indicator equals Y".
[0354] a) Alternatively, under the condition check of "the value of the first indicator >> 1" or "the value of the first indicator / 2" or "the value of the first indicator & 0x01", the second indicator may be signaled.
[0355] b) Alternatively, in addition, when the first indicator equals X, it can be inferred that it is enabled; or when the first indicator equals z, it can be inferred that it is disabled.
[0356] b. In a second example, one or more of the following sub-items are applied:
[0357] i. More than one indicator may be signaled at a higher level (e.g., in the picture header (PH)) to specify how to enable LMCS at a lower level for non-binary values.
[0358] 1) In one example, two indicators may be signaled in the PH.
[0359] a) In one example, the first indicator specifies whether there is at least one slice associated with the PH that enables LMCS. And the second indicator specifies whether all slices associated with the PH enable LMCS.
[0360] i. Alternatively, in addition, the second indicator may be signaled conditionally according to the value of the first indicator, e.g., when the first indicator specifies that there is at least one slice that enables LMCS.
[0361] i. Alternatively, in addition, when the second indicator does not exist, it is inferred that all slices enable LMCS.
[0362] ii. Alternatively, in addition, the third indicator may be signaled conditionally in the SH according to the value of the second indicator, e.g., when the second indicator specifies that not all slices enable LMCS.
[0363] i. Alternatively, in addition, when the third indicator does not exist, it can be inferred according to the value of the first and / or second indicator (e.g., inferred to be equal to the value of the first indicator).
[0364] b) Alternatively, the first indicator specifies whether there is at least one slice associated with the PH that disables LMCS. And the second indicator specifies whether all slices associated with the PH disable LMCS.
[0365] i. Alternatively, in addition, the second indicator may be signaled conditionally according to the value of the first indicator, e.g., when the first indicator specifies that there is at least one strip that disables LMCS.
[0366] i. Alternatively, in addition, when the second indicator does not exist, it is inferred that all strips related to PH disable LMCS.
[0367] ii. Alternatively, in addition, according to the value of the second indicator, the third indicator may be signaled conditionally in SH, e.g., when the second indicator specifies that not all strips disable LMCS.
[0368] i. Alternatively, in addition, when the third indicator does not exist, it may be inferred according to the value of the first and / or second indicator (e.g., inferred to be equal to the value of the first indicator).
[0369] 2) Alternatively, in addition, the first indicator may be signaled conditionally according to the value of the LMCS enable flag (e.g., sps_lmcs_enabled_flag) in the sequence level.
[0370] ii. The third indicator (e.g., slice_lmcs_enabled_flag) for enabling / disabling LMCS at a lower level may be signaled at a lower level (e.g., in the slice header), and the third indicator may be signaled conditionally by checking the value of the first and / or second indicator.
[0371] 1) In one example, under the condition check of "not all strips enable LMCS" or "not all strips disable LMCS", the third indicator may be signaled.
[0372] c. In yet another example, the first and / or second and / or third indicators mentioned in the first / second examples may be used to control the use of RP or CRS instead of LMCS.
[0373] 3. The semantic updates of the three LMCS flags in SPS / PH / SH are as follows:
[0374] sps_lmcs_enabled_flag being equal to 1 specifies that luminance mapping with chroma scaling can be used in CLVS. sps_lmcs_enabled_flag being equal to 0 specifies that luminance mapping with chroma scaling is not used in CLVS.
[0375] The ph_lmcs_enabled_flag being equal to 1 specifies that the luminance mapping with chroma scaling [[enabled]] is available for all slices related to PH. The ph_lmcs_enabled_flag being equal to 0 specifies that the luminance mapping with chroma scaling [[can be disabled for one or more, or]] is not used for all slices related to PH. When it is absent, the value of the ph_lmcs_enabled_flag is inferred to be equal to 0.
[0376] The slice_lmcs_enabled_flag being equal to 1 specifies that the luminance mapping with chroma scaling [[enabled]] is used for the current slice. The slice_lmcs_enabled_flag being equal to 0 specifies that the luminance mapping with chroma scaling is not [[enabled for]] used for the current slice. When the slice_lmcs_enabled_flag is absent, it is inferred to be equal to 0.
[0377] a. Change the PH and / or SH LMCS signaling such that when LMCS is used for all slices of a picture, there is no LMCS signaling in SH.
[0378] i. Alternatively, in addition, how LMCS is inferred depends on the PH LMCS signaling.
[0379] 1) In one example, when LMCS is used for all slices of a picture, it is inferred to be enabled; and when LMCS is not used for all slices of a picture, it is inferred to be disabled.
[0380] Related to affine AMVR
[0381] 4. The semantics of the affine AMVR flag in the SPS is updated as follows:
[0382] The sps_affine_amvr_enabled_flag being equal to 1 specifies that the adaptive motion vector difference precision [[is]] can be used in the motion vector coding / decoding for the affine inter prediction mode. The sps_affine_amvr_enabled_flag being equal to 0 specifies that the adaptive motion vector difference precision is not used in the motion vector coding / decoding for the affine inter prediction mode. When it is absent, the value of the sps_affine_amvr_enabled_flag is inferred to be equal to 0.
[0383] 6. Examples
[0384] 6.1. Example 1: Support for sub - pictures
[0385] This example is for Item 1 and its sub - items.
[0386] 3 definitions
[0387] Picture-level strip index: The index of the strip in the list of strips in the picture as defined when rect_slice_flag equals 1, in the order in which the strips are signaled in the PPS when single_slice_per_subpic_flag equals 0, or in the order of increasing subpicture index of the subpicture corresponding to the strip when single_slice_per_subpic_flag equals 1.
[0388] [[Picture-level strip index: The index of the strip in the list of strips in the picture, in the order in which they are signaled in the PPS when rect_slice_flag equals 1.]]
[0389] 6.5.1 CTB raster scan, slice scan, and subpicture scan processes
[0390] …
[0391] For the list CtbToTileColBd[ctbAddrX] and ctbToTileColIdx[ctbAddrX] of ctbAddrX ranging from 0 to PicWidthInCtbsY (including the end values), they respectively specify the conversion from the horizontal CTB address to the left tile column boundary in CTB units and to the tile column index, and are derived as follows:
[0392]
[0393] Note 3 – The sizes of the arrays CtbToTileColBd[] and ctbToTileColIdx[] in the above derivation are 1 larger than the actual picture width in CTBs.
[0394] For the list CtbToTileRowBd[ctbAddrY] and ctbToTileRowIdx[ctbAddrY] of ctbAddrY ranging from 0 to PicHeightInCtbsY (including the end values), they respectively specify the conversion from the vertical CTB address to the top tile column boundary in CTB units and to the tile row index, and are derived as follows:
[0395]
[0396] Note 4 – The sizes of the arrays CtbToTileRowBd[] and ctbToTileRowIdx[] in the above derivation are 1 larger than the actual picture height in CTBs.
[0397] For the list SubpicWidthInTiles[i] and SubpicHeightInTiles[i] of i ranging from 0 to sps_num_subpics_minus1 (including the end values), the width and height of the i-th sub-picture in the slice row and slice column are specified respectively, and for the list subpicHeightLessThanOneTileFlag[i] of i ranging from 0 to sps_num_subpics_minus1 (including the end values), it is deduced as follows whether the height of the i-th sub-picture is less than one slice row:
[0398]
[0399] Alternatively, in the above equation, the following condition:
[0400] if(SubpicHeightInTiles[i] == 1 && (bottomY + 1 - topY < RowHeight[ctbToTileRowIdx[topY]])
[0401] can be changed to the following:
[0402] if(SubpicHeightInTiles[i] == 1 && (topY != CtbToTileRowBd[topY] ||
[0403] bottomY + 1 != CtbToTileRowBd[bottomY + 1])
[0404] Note 5 – When a slice is divided into multiple rectangular strips and only a subset of the rectangular strips of the slice is included in the i-th sub-picture, the slice is counted as one slice in the value of SubpicHeightInTiles[i].
[0405] When rect_slice_flag is equal to 1, the list NumCtusInSlice[i] for i ranging from 0 to num_slices_in_pic_minus1 (including the end values) specifies the number of CTUs in the i-th slice, the list SliceTopLeftTileIdx[i] for i ranging from 0 to num_slices_in_pic_minus1 (including the end values) specifies the slice index of the slice containing the first CTU in the slice, and the matrix CtbAddrInSlice[i][j] for i ranging from 0 to num_slices_in_pic_minus1 (including the end values) and j ranging from 0 to NumCtusInSlice[i] – 1 (including the end values) specifies the picture raster scan address of the j-th CTB within the i-th slice, and the variable NumSlicesInTile[i] specifies the number of slices in the slice containing the i-th slice, which is derived as follows:
[0406]
[0407]
[0408] ....
[0409] D.7.2 Sub-picture level information SEI message semantics
[0410] …
[0411] ref_level_fraction_minus1[i][j] plus 1 specifies the fractional value of the level limit associated with ref_level_idc[i] that corresponds to the j-th sub-picture, as specified in Clause A.4.1.
[0412] The variable SubpicSizeY[j] is set to be equal to (subpic_width_minus1[j] + 1) * CtbSizeY * (subpic_height_minus1[j] + 1) * CtbSizeY.
[0413] When not present, the value of ref_level_fraction_minus1[i][j] is inferred to be equal to Ceil(256 * SubpicSizeY[j] ÷ PicSizeInSamplesY * MaxLumaPs(general_level_idc) ÷ MaxLumaPs(ref_level_idc[i]) - 1.
[0414] The variable RefLevelFraction[i][j] is set to be equal to ref_level_fraction_minus1[i][j]+1.
[0415] [[The derivation of the variables SubpicNumTileCols[j] and SubpicNumTileRows[j] is as follows:
[0416] ...
[0417] - The value of SubpicWidthInTiles[j] should be less than or equal to MaxTileCols, and the value of SubpicHeightInTiles[j] should be less than or equal to MaxTileRows, where MaxTileCols and MaxTileRows are specified in Table A.1 for level ref_level_idc[i].
[0418] - The value of SubpicWidthInTiles[j]*SubpicHeightInTiles[j] should be less than or equal to MaxTileCols*MaxTileRows*RefLevelFraction[i][j], where MaxTileCols and MaxTileRows are specified in Table A.1 for level ref_level_idc[i]. ...
[0419] For the variable SubpicSetAccLevelFraction[i] of the total level fraction relative to the reference level ref_level_idc[i], and the variables SubpicSetCpbSizeVcl[i], SubpicSetCpbSizeNal[i], SubpicSetBitRateVcl[i], and SubpicSetBitRateNal[i] of the sub-picture set, the derivation is as follows:
[0420] ...
[0421] 6.2. Example 2: Support for LMCS
[0422] In this example, the syntax and semantics of the LMCS-related syntax elements in the picture header are modified such that when LMCS is used for all stripes of a picture, there is no LMCS signaling notification in the SH.
[0423] 7.3.2.7 Picture Header Structure Syntax
[0424]
[0425]
[0426]
[0427] 7.3.7.1 General strip header syntax
[0428]
[0429] ph_lmcs_enabled_type [[flag]] equal to M (e.g., M = 1) [[1]] specifies that luminance mapping with chroma scaling is enabled for all strips associated with PH. ph_lmcs_enabled_flag equal to N (e.g., N = 2) specifies that luminance mapping with chroma scaling is enabled for at least one strip and disabled for at least one strip associated with PH. ph_lmcs_enabled_flag equal to 0 specifies [[that it is possible to disable luminance mapping with chroma scaling for one or more, or]] for all strips associated with PH. When absent, the value of ph_lmcs_enabled_flag is inferred to be 0.
[0430] slice_lmcs_enabled_flag equal to 1 specifies that luminance mapping with chroma scaling is enabled for the current strip. slice_lmcs_enabled_flag equal to 0 specifies that luminance mapping with chroma scaling is not enabled for the current strip. When slice_lmcs_enabled_flag is absent, it is inferred to be equal to [[0]] (ph_lmcs_enabled_type? 1:0).
[0431] In the above example, the values of M and N can be set to 1 and 2 respectively. Alternatively, the values of M and N can be set to 2 and 1 respectively.
[0432] Figure 5is a block diagram of an example video processing system 1900 that can implement the various technologies disclosed herein. Various implementations may include some or all of the components in system 1900. System 1900 may include an input 1902 for receiving video content. The video content may be received in a raw or uncompressed format (such as 8- or 10-bit multi-component pixel values), or it may be received in a compressed or encoded format. Input 1902 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces (such as Ethernet, Passive Optical Network (PON), etc.) and wireless interfaces (such as Wi-Fi or cellular interfaces).
[0433] System 1900 may include a codec component 1904 that can implement the various codec or encoding methods described in this document. The codec component 1904 may reduce the average bit rate of the video from the output of input 1902 to the output of the codec component 1904 to produce a coded representation of the video. Thus, codec technologies are sometimes referred to as video compression or video transcoding technologies. The output of the codec component 1904 may be stored or transmitted via the connected communication, as represented by component 1906. The stored or communicated bitstream (or coded) representation of the video received at input 1902 may be used by component 1908 to generate pixel values or a displayable video that is sent to the display interface 1910. The process of generating a user-visible video from the bitstream representation is sometimes referred to as video decompression. Additionally, although certain video processing operations are referred to as "codec" operations or tools, it should be understood that codec tools or operations are used at the encoder, and the corresponding decoding tools or operations that invert the results of the codec will be performed by the decoder.
[0434] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB), High-Definition Multimedia Interface (HDMI), or Displayport, etc. Examples of storage interfaces include SATA (Serial Advanced Technology Attachment), PCI, IDE interfaces, etc. The technologies described in this document may be implemented in various electronic devices, such as mobile phones, laptop computers, smartphones, or other devices capable of digital data processing and / or video display.
[0435] Figure 6is a block diagram of a video processing apparatus 3600. The apparatus 3600 can be used to implement one or more of the methods described herein. The apparatus 3600 can be implemented in a smart phone, a tablet computer, a computer, an Internet of Things (IoT) receiver, etc. The apparatus 3600 can include one or more processors 3602, one or more memories 3604, and video processing hardware 3606. The (multiple) processors 3602 can be configured to implement one or more of the methods described in this document. The (multiple) memories 3604 can be used to store data and code for implementing the methods and techniques described herein. The video processing hardware 3606 can be used to implement some of the techniques described in this document in hardware circuitry.
[0436] Figure 8 is a block diagram showing an example video codec system 100 that can utilize the techniques of the present disclosure.
[0437] As Figure 8 shown, the video codec system 100 can include a source device 110 and a destination device 120. The source device 110 generates encoded video data, which can be referred to as a video encoding device. The destination device 120 can decode the encoded video data generated by the source device 110, and the destination device 120 can be referred to as a video decoding device.
[0438] The source device 110 can include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.
[0439] The video source 112 can include sources such as a video capture device, an interface for receiving video data from a video content provider, and / or a computer graphics system that generates video data, or a combination of these sources. The video data can include one or more pictures. The video encoder 114 encodes the video data from the video source 112 to generate a bitstream. The bitstream can include a sequence of bits that form a codec representation of the video data. The bitstream can include coded pictures and associated data. The coded pictures are codec representations of the pictures. The associated data can include sequence parameter sets, picture parameter sets, and other syntax elements. The I / O interface 116 includes a modulator / demodulator (modem) and / or a transmitter. The encoded video data can be directly sent to the destination device 120 via the I / O interface 116 over a network 130a. The encoded video data can also be stored on a storage medium / server 130b for access by the destination device 120.
[0440] The destination device 120 can include an I / O interface 126, a video decoder 124, and a display device 122.
[0441] The I / O interface 126 may include a receiver and / or a modem. The I / O interface 126 may obtain encoded video data from the source device 110 or the storage medium / server 130b. The video decoder 124 may decode the encoded video data. The display device 122 may display the decoded video data to the user. The display device 122 may be integrated with the target device 120 or may be external to the target device 120 configured to interface with an external display device.
[0442] The video encoder 114 and the video decoder 124 may operate according to video compression standards such as, for example, the High Efficiency Video Coding (HEVC) standard, the Versatile Video Coding (VVC) standard, and other current and / or other standards.
[0443] Figure 9 is a block diagram illustrating an example of a video encoder 200, which may be the video encoder 114 in the system 100 illustrated in Figure 8 FIG.
[0444] The video encoder 200 may be configured to perform any or all of the techniques of the present disclosure. In an example of Figure 9 the video encoder 200 includes a plurality of functional components. The techniques described in the present disclosure may be shared among the various components of the video encoder 200. In some examples, a processor may be configured to perform any or all of the techniques described in the present disclosure.
[0445] The functional components of the video encoder 200 may include a splitting unit 201, a prediction unit 202 (which may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra prediction unit 206), a residual generation unit 207, a transformation unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transformation unit 211, a reconstruction unit 212, a buffer 213, and an entropy coding unit 214.
[0446] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an Intra Block Copy (IBC) unit. The IBC unit may perform prediction in an IBC mode where at least one reference picture is the picture in which the current video block is located.
[0447] In addition, some components such as the motion estimation unit 204 and the motion compensation unit 205 may be highly integrated but are shown separately in the example of Figure 9 for purposes of explanation.
[0448] The splitting unit 201 may split a picture into one or more video blocks. The video encoder 200 and the video decoder 300 may support various video block sizes.
[0449] The mode selection unit 203 may select, for example, one of the intra or inter coding modes based on an error result, and provide the resulting intra or inter coded block to the residual generation unit 207 to generate residual block data and to the reconstruction unit 212 to reconstruct the coded block for use as a reference picture. In some examples, the mode selection unit 203 may select a combined intra and inter prediction (CIIP) mode, where the prediction is based on an inter prediction signal and an intra prediction signal. The mode selection unit 203 may also select the resolution of the motion vector (e.g., sub-pixel or full-pixel accuracy) for a block in the case of inter prediction.
[0450] To perform inter prediction on a current video block, the motion estimation unit 204 may generate motion information for the current video block by comparing one or more reference frames from the buffer 213 with the current video block. The motion compensation unit 205 may determine a predicted video block for the current video block based on the motion information and decoded samples of a picture from the buffer 213 (rather than the picture associated with the current video block).
[0451] The motion estimation unit 204 and the motion compensation unit 205 may perform different operations for the current video block, e.g., perform different operations depending on whether the current video block is in an I-slice, a P-slice, or a B-slice.
[0452] In some examples, the motion estimation unit 204 may perform uni-directional prediction of the current video block, and the motion estimation unit 204 may search for a reference video block of the current video block in the reference pictures of list 0 or list 1. The motion estimation unit 204 may then generate a reference index indicating the reference picture of list 0 or list 1 that contains the reference video block and a motion vector indicating the spatial displacement between the current video block and the reference video block. The motion estimation unit 204 may output the reference index, a prediction direction indicator, and the motion vector as the motion information of the current video block. The motion compensation unit 205 may generate a predicted video block of the current block based on the reference video block indicated by the motion information of the current video block.
[0453] In other examples, the motion estimation unit 204 may perform bidirectional prediction of the current video block. The motion estimation unit 204 may search for a reference video block of the current video block in the reference pictures of list 0 and may also search for another reference video block of the current video block in the reference pictures of list 1. The motion estimation unit 204 may then generate a reference index indicating that the reference picture of list 0 or list 1 contains the reference video block and a motion vector indicating the spatial displacement between the reference video block and the current video block. The motion estimation unit 204 may output the reference index and the motion vector of the current video block as the motion information of the current video block. The motion compensation unit 205 may generate a predicted video block of the current video block based on the reference video block indicated by the motion information of the current video block.
[0454] In some examples, the motion estimation unit 204 may output the entire set of motion information for the decoding process of the decoder.
[0455] In some examples, the motion estimation unit 204 may not output the entire set of motion information of the current video. Instead, the motion estimation unit 204 may signal the motion information of the current video block by referring to the motion information of another video block. For example, the motion estimation unit 204 may determine that the motion information of the current video block is similar enough to the motion information of the neighboring video block.
[0456] In one example, the motion estimation unit 204 may indicate in the syntax structure associated with the current video block: indicating to the video decoder 300 that the current video block has a value of the same motion information as another video block.
[0457] In another example, the motion estimation unit 204 may identify another video block and a motion vector difference (MVD) in the syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 300 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[0458] As discussed above, the video encoder 200 may predictively signal the motion vector. Two examples of predictive signaling techniques that may be implemented by the video encoder 200 include advanced motion vector prediction (AMVP) and merge mode signaling.
[0459] The intra prediction unit 206 may perform intra prediction on the current video block. When the intra prediction unit 206 performs intra prediction on the current video block, the intra prediction unit 206 may generate prediction data of the current video block based on the decoded samples of other video blocks in the same picture. The prediction data of the current video block may include a predicted video block and various syntax elements.
[0460] The residual generation unit 207 may generate residual data for a current video block by subtracting (e.g., denoted by a minus sign) a (plurality of) predicted video blocks of the current video block from the current video block. The residual data for the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.
[0461] In other examples, such as in the skip mode, there may be no residual data for the current video block, and the residual generation unit 207 may not perform the subtraction operation.
[0462] The transform processing unit 208 may generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video block associated with the current video block.
[0463] After the transform processing unit 208 generates the transform coefficient video block associated with the current video block, the quantization unit 209 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0464] The inverse quantization unit 210 and the inverse transform unit 211 may apply inverse quantization and inverse transform to the transform coefficient video block, respectively, to reconstruct the residual video block from the transform coefficient video block. The reconstruction unit 212 may add the reconstructed residual video block to the corresponding samples of one or more predicted video blocks generated by the prediction unit 202 to generate a reconstructed video block associated with the current block for storage in the buffer 213.
[0465] After the reconstruction unit 212 reconstructs the video block, a loop filtering operation may be performed to reduce the blockiness artifacts in the video block.
[0466] The entropy coding unit 214 may receive data from other functional components of the video encoder 200. When the entropy coding unit 214 receives the data, the entropy coding unit 214 may perform one or more entropy coding operations to generate entropy coded data and output a bitstream including the entropy coded data.
[0467] Some embodiments of the technology disclosed herein include determining or enabling a video processing tool or mode. In an example, when a video processing tool or mode is enabled, the encoder will use or implement the tool or mode during the processing of video blocks, but does not necessarily modify the resulting bitstream based on the use of the tool or mode. In other words, when a video processing tool or mode is determined or enabled, the conversion from video blocks to the bitstream (or bitstream representation) of the video will use the video processing tool or mode. In another example, when a video processing tool or mode is enabled, the decoder will utilize the knowledge that the bitstream has been modified based on the video processing tool or mode to process the bitstream. In other words, using the video processing tool or mode determined or enabled, the conversion from the bitstream of the video to video blocks will be performed.
[0468] Figure 10 is a block diagram illustrating an example of a video decoder 300, which may be the video decoder 114 in the system 100 shown in Figure 8 the system 100 shown in
[0469] The video decoder 300 may be configured to perform any or all of the technologies of the present disclosure. In Figure 10 an example, the video decoder 300 includes a plurality of functional components. The technologies described in the present disclosure may be shared among various components of the video decoder 300. In some examples, a processor may be configured to perform any or all of the technologies described in the present disclosure.
[0470] In Figure 10 an example, the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, and a reconstruction unit 306 and a buffer 307. In some examples, the video decoder 300 may perform a decoding process that is generally inverse to the encoding process described with respect to the video encoder 200 ( Figure 9 ).
[0471] The entropy decoding unit 301 may retrieve the encoded bitstream. The encoded bitstream may include entropy-coded video data (e.g., encoded blocks of video data). The entropy decoding unit 301 may decode the entropy-coded video, and based on the entropy-decoded video data, the motion compensation unit 302 may determine motion information including a motion vector, motion vector precision, reference picture list index, and other motion information. The motion compensation unit 302 may determine such information, for example, by performing AMVP and merge mode.
[0472] The motion compensation unit 302 may generate a motion-compensated block, possibly performing interpolation based on an interpolation filter. The identifier of the interpolation filter to be used with sub-pixel precision may be included in a syntax element.
[0473] The motion compensation unit 302 may use the interpolation filter used by the video encoder 200 during encoding of a video block to calculate interpolated values for sub-integer pixels of a reference block. The motion compensation unit 302 may determine the interpolation filter used by the video encoder 200 according to received syntax information and use the interpolation filter to generate a prediction block.
[0474] The motion compensation unit 302 may use some syntax information to determine: the size of the blocks used to encode the (multiple) frames and / or (multiple) slices of the encoded video sequence, the partitioning information that describes how each macroblock of a picture of the encoded video sequence is partitioned, the mode indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-coded block, and other information for decoding the encoded video sequence.
[0475] The intra prediction unit 303 may form a prediction block from spatially adjacent blocks using, for example, an intra prediction mode received in the bitstream. The inverse quantization unit 303 inverse quantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 301. The inverse transform unit 303 applies an inverse transform.
[0476] The reconstruction unit 306 may sum a residual block with a corresponding prediction block generated by the motion compensation unit 202 or the intra prediction unit 303 to form a decoded block. As desired, a deblocking filter may also be applied to filter the decoded block to remove blocking artifacts. The decoded video block is then stored in a buffer 307 that provides reference blocks for subsequent motion compensation / intra prediction and also produces a decoded video for presentation on a display device.
[0477] Next, a list of preferred solutions for some embodiments is provided.
[0478] The following solutions illustrate example embodiments of the techniques discussed in the previous section (e.g., item 1).
[0479] 1. A video processing method (e.g., Figure 7 method 900 depicted in), comprising: performing (902) a conversion between a video including one or more video pictures and a codec representation of the video, wherein each video picture includes one or more slices, and wherein the codec representation conforms to format rules; wherein the format rules specify first information signaled in the codec representation and second information derived from the codec representation, wherein at least the first information or the second information is related to a row index or a column index of one or more slices.
[0480] 2. The method according to solution 1, wherein the format rules specify deriving a slice column index for each codec tree unit column of each video picture.
[0481] 3. The method according to Solution 1, wherein the format rule stipulates deriving a slice row index for each coding tree unit row of each video picture.
[0482] The following solutions illustrate example embodiments of the techniques discussed in the previous section (e.g., Item 2). In these solutions, the video region may be a video picture, and the video unit may be a video block or a coding tree unit or a video strip.
[0483] 4. A video processing method, comprising: performing a conversion between a video unit of a video region of a video and a coded representation of the video, wherein the coded representation complies with a format rule; wherein the format rule stipulates that first control information at the video region controls whether second control information is included at the video unit level; wherein the first control information and / or the second control information includes information regarding luminance mapping and chrominance scaling (LMCS) or chrominance residual scaling (CRS) or retrofit process (RP) for the conversion.
[0484] 5. The method according to Solution 4, wherein the first control information includes an indicator indicating whether the second control information is included in the coded representation.
[0485] 6. The method according to Solutions 4 - 5, wherein a specific value of the first control information indicates disabling LMCS for all video units in the video region.
[0486] 7. The method according to any one of Solutions 4 - 6, wherein the second control information controls enabling of LMCS at the video unit.
[0487] 8. The method according to Solution 4, wherein the first control information includes a plurality of indicators.
[0488] 9. The method according to any one of Solutions 1 to 8, wherein the conversion includes encoding the video into a coded representation.
[0489] 10. The method according to any one of Solutions 1 to 8, wherein the conversion includes decoding the coded representation to generate pixel values of the video.
[0490] 11. A video decoding apparatus, comprising a processor configured to implement the method described in one or more of Solutions 1 to 10.
[0491] 12. A video encoding apparatus, comprising a processor configured to implement the method described in one or more of Solutions 1 to 10.
[0492] 13. A computer program product having computer code stored thereon which, when executed by a processor, causes the processor to implement the method according to any one of Solutions 1 to 10.
[0493] 14. The method, apparatus or system described in this document.
[0494] In the solution described herein, an encoder may conform to format rules by generating an encoded / decoded representation according to the format rules. In the solution described herein, a decoder may use the format rules to parse syntax elements in the encoded / decoded representation, wherein the presence and absence of the syntax elements are known according to the format rules to generate a decoded video.
[0495] Figure 11 Flowchart of an example method 1100 for video processing. Operation 1102 includes: performing a conversion between a video including one or more video pictures and a bitstream of the video, wherein each video picture includes one or more slices, the one or more slices include one or more slice columns, wherein the bitstream conforms to format rules, and wherein the format rules specify deriving a slice column index for each coding tree unit (CTU) column of the slices of the video picture.
[0496] In some embodiments of method 1100, the slice column index of the ctbAddrX-th slice column, denoted as ctbToTileColIdx[ctbAddrX], is derived as follows: tileX = 0
[0497]
[0498] where PicWidthInCtbsY represents the width of the video picture in coding tree blocks (CTBs), and where tileColBd[i] represents the position of the i-th slice column boundary in CTBs.
[0499] In some embodiments of method 1100, each video picture further includes one or more sub-pictures, each sub-picture includes one or more stripes that jointly form a rectangular subset of the video picture, and the format rules further specify deriving the width of the sub-picture in terms of the included slice columns based on the slice column index of the leftmost CTU and / or the rightmost CTU included in the sub-picture.
[0500] In some embodiments of method 1100, the width of the i-th sub-picture in terms of slice columns, denoted as SubpicWidthInTiles[i], is derived as follows:
[0501]
[0502]
[0503] Where sps_num_subpics_minus1 represents the number of sub - pictures in a video picture, where sps_subpic_ctu_top_left_x[i] represents the horizontal position of the top - left CTU of the i - th sub - picture, where sps_subpic_width_minus1[i]+1 defines the width of the i - th sub - picture, and where ctbToTileColIdx[rightX] and ctbToTileColIdx[leftX] respectively represent the tile column indices of the left - most CTU and the right - most CTU included in the sub - picture.
[0504] In some embodiments of method 1100, in response to a slice being segmented into multiple rectangular stripes and the sub - picture including only a subset of the rectangular stripes of the slice, the slice is counted as one slice in the value of the width of the sub - picture.
[0505] Figure 12 A flowchart of an example method 1200 for video processing. Operation 1202 includes: performing a conversion between a video including one or more video pictures and a bitstream of the video, where each video picture includes one or more slices, the one or more slices include one or more slice rows, the bitstream conforms to format rules, and the format rules specify deriving a slice row index for each coding - tree unit (CTU) row of the slices of the video picture.
[0506] In some embodiments of method 1200, the slice row index of the ctbAddrY - th slice row, denoted as ctbToTileRowIdx[ctbAddrY], is derived as follows:
[0507]
[0508] Where PicHeightInCtbsY represents the height of the video picture in terms of coding - tree blocks (CTBs), and where tileRowBd[i] represents the position of the i - th slice row boundary in terms of CTBs.
[0509] In some embodiments of method 1200, each video picture further includes one or more sub - pictures, each sub - picture includes one or more stripes that jointly form a rectangular subset of the video picture, and the format rules also specify that the height of the sub - picture in terms of slice rows is derived based on the slice row indices of the top CTU and / or the bottom CTU included in the sub - picture.
[0510] In some embodiments of method 1200, the height of the i - th sub - picture in terms of slice rows, denoted as SubpicHeightInTiles[i], is derived as follows:
[0511]
[0512]
[0513] Where sps_num_subpics_minus1 represents the number of sub - pictures in a video picture, sps_subpic_ctu_top_left_y[i] represents the vertical position of the top - left CTU of the i - th sub - picture, sps_subpic_height_minus1[i]+1 specifies the height of the i - th sub - picture, and ctbToTileRowIdx[botY] and ctbToTileRowIdx[topY] represent the slice row indices of the bottom CTU and the top CTU included in the sub - picture, respectively.
[0514] In some embodiments of method 1200, in response to a slice being split into multiple rectangular stripes and only a subset of the rectangular stripes of the slice being included in a sub - picture, the slice is counted as one slice in the value of the height of the sub - picture.
[0515] Figure 13 A flowchart of an example method 1300 for video processing. Operation 1302 includes: performing a conversion between a video including at least one video picture and a bitstream of the video according to a rule, where the at least one video picture includes one or more stripes and one or more sub - pictures, and where the rule specifies an order of stripe indices of one or more stripes in the at least one video picture in response to a syntax element associated with the at least one video picture indicating whether each sub - picture of the at least one video picture includes a single stripe.
[0516] In some embodiments of method 1300, the rule further specifies indicating a stripe index in response to each stripe in the at least one video picture being a rectangular stripe. In some embodiments of method 1300, the rule specifies that in the case where the syntax element indicates that each of one or more sub - pictures includes a single rectangular stripe, the order corresponds to an increasing value of the sub - picture index of one or more sub - pictures in the video picture, and the sub - picture indices of the one or more sub - pictures are indicated in a sequence parameter set (SPS) referenced by the at least one video picture. In some embodiments of method 1300, the rule specifies that in the case where the syntax element indicates that each sub - picture includes one or more rectangular stripes, the order corresponds to the order in which one or more stripes are included in a picture parameter set (PPS) referenced by the at least one video picture. In some embodiments of method 1300, the syntax element is included in a picture parameter set (PPS) referenced by the at least one video picture.
[0517] Figure 14Flowchart of an example method 1400 for video processing. Operation 1402 includes: performing a conversion between a video unit of a video region of a video and a bitstream of the video, where the bitstream conforms to format rules, where the format rules specify that first control information of a first level of the video region in the bitstream controls whether second control information is included in a second level of the video unit in the bitstream, where the second level is less than the first level, where the first control information and the second control information include information on whether or how a luminance mapping and chrominance scaling (LMCS) tool is applied to the video unit, and where the LMCS tool includes using chrominance residual scaling (CRS) or a luminance remodelling process (RP) for the conversion.
[0518] In some embodiments of method 1400, the first control information selectively includes a first indicator that indicates whether the LMCS tool is enabled for one or more stripes of a first level of the video region to specify whether the LMCS tool is enabled at the second level of the video unit, and the first indicator is a non-binary value. In some embodiments of method 1400, the first level of the video region includes a picture header. In some embodiments of method 1400, the first level of the video region includes a picture header, the first control information includes a first indicator, when the first indicator is equal to a first value, the LMCS tool is enabled for all stripes of the picture header, when the first indicator is equal to a second value, the LMCS tool is enabled for less than all stripes of the picture header, when the first indicator is equal to a third value, the LMCS tool is disabled for all stripes of the picture header, and the first value, the second value, and the third value are different from each other. In some embodiments of method 1400, when the first control information does not include the first indicator, the value of the first indicator is inferred as a default value.
[0519] In some embodiments of method 1400, the first level of the video region includes a picture header, the first control information includes a first indicator, when the first indicator is equal to a first value, the LMCS tool is disabled for all stripes of the picture header, when the first indicator is equal to a second value, the LMCS tool is disabled for less than all stripes of the picture header, when the first indicator is equal to a third value, the LMCS tool is enabled for all stripes of the picture header, and the first value, the second value, and the third value are different from each other. In some embodiments of method 1400, whether the first indicator is selectively included in the first control information is based on the value of a syntax element in the bitstream that indicates whether the LMCS tool is enabled at the sequence level. In some embodiments of method 1400, the first indicator is coded or decoded using u(v) or u(2) or ue(v). In some embodiments of method 1400, the first indicator is coded or decoded using a truncated unary code.
[0520] In some embodiments of method 1400, based on the value of a first indicator indicating whether the LMCS tool is enabled for one or more strips at a first level of a video region, adaptive parameter set (APS) information of the LMCS tool used by the one or more strips and / or chroma scaling syntax elements are included in the bitstream. In some embodiments of method 1400, second control information selectively includes a second indicator that indicates whether the LMCS tool is enabled or disabled for one or more strips at a second level of a video unit, and the second indicator is included in the bitstream based on the value of the first indicator included in the first control information, and the first indicator indicates whether the LMCS tool is enabled or disabled for one or more strips at a second level of a video unit. In some embodiments of method 1400, the second control information includes a strip header. In some embodiments of method 1400, the second indicator is included in the second control information in response to the first indicator being equal to a first value. In some embodiments of method 1400, the second indicator is included in the second control information in response to performing a conditional check of: first indicator >> 1, or first indicator / 2, or first indicator & 0x01, where >> describes a right shift operation, and where & describes a bitwise logical AND operation.
[0521] In some embodiments of method 1400, in response to the first indicator being equal to a first value, the second indicator is inferred to indicate that the LMCS tool is enabled for one or more strips at a second level of a video unit, or in response to the first indicator being equal to a third value, the second indicator is inferred to indicate that the LMCS tool is disabled for one or more strips at a second level of a video unit, and the first value, the second value of the first indicator, and the third value are different from each other. In some embodiments of method 1400, the first control information includes a plurality of indicators that indicate whether the LMCS tool is enabled for one or more strips at a first level of a video region to specify whether the LMCS tool is enabled at a second level of a video unit, and the plurality of indicators have non-binary values. In some embodiments of method 1400, the plurality of indicators includes at least two indicators included in a picture header. In some embodiments of method 1400, the at least two indicators include a first indicator that specifies whether the LMCS tool is enabled for at least one strip associated with the picture header, and the at least two indicators selectively include a second indicator that specifies whether the LMCS tool is enabled for all strips associated with the picture header. In some embodiments of method 1400, based on the value of the first indicator, the second indicator selectively exists in the plurality of indicators.
[0522] In some embodiments of method 1400, the value of the first indicator indicates that the LMCS tool is enabled for at least one stripe. In some embodiments of method 1400, in response to the absence of a second indicator in the bitstream, it is inferred that the LMCS tool is enabled for all stripes associated with the picture header. In some embodiments of method 1400, the at least two indicators include a third indicator that is selectively included in the stripe header based on a second value of the second indicator. In some embodiments of method 1400, the second value of the second indicator indicates that the LMCS tool is disabled for all stripes. In some embodiments of method 1400, in response to the absence of the third indicator in the bitstream, the value of the third indicator is inferred based on the first value of the first indicator and / or the second value of the second indicator. In some embodiments of method 1400, the at least two indicators include a first indicator that specifies whether the LMCS tool is disabled for at least one stripe associated with the picture header, and the at least two indicators selectively include a second indicator that specifies whether the LMCS tool is disabled for all stripes associated with the picture header. In some embodiments of method 1400, based on the value of the first indicator, the second indicator is present in the plurality of indicators. In some embodiments of method 1400, the value of the first indicator specifies that the LMCS tool is disabled for at least one stripe. In some embodiments of method 1400, in response to the absence of the second indicator in the bitstream, it is inferred that the LMCS tool is disabled for all stripes associated with the picture header. In some embodiments of method 1400, the at least two indicators selectively include the third indicator in the stripe header based on the second value of the second indicator.
[0523] In some embodiments of method 1400, the second value of the second indicator specifies that the LMCS tool is enabled for all stripes. In some embodiments of method 1400, in response to the absence of the third indicator in the bitstream, the value of the third indicator is inferred based on the first value of the first indicator and / or the second value of the second indicator. In some embodiments of method 1400, the plurality of indicators selectively include the first indicator based on the value of a syntax element that indicates whether the LMCS tool is enabled at the sequence level. In some embodiments of method 1400, the plurality of indicators selectively include a third indicator that indicates whether the LMCS tool is enabled or disabled at a second level of the video unit, and the third indicator selectively exists based on the first value of the first indicator and / or the second value of the second indicator. In some embodiments of method 1400, the third indicator selectively exists based on the second indicator that indicates that the LMCS tool is not enabled for all stripes or not disabled for all stripes. In some embodiments of method 1400, the first indicator, the second indicator, and / or the third indicator controls the use of CRS or the luminance RP.
[0524] Figure 15 Flowchart of an example method 1500 for video processing. Operation 1502 includes: performing a conversion between a video and a bitstream of the video according to rules, where the rules specify that a luminance mapping and chrominance scaling (LMCS) tool is enabled when a first syntax element in a reference sequence parameter set indicates that the LMCS tool is enabled, where the rules specify that the LMCS tool is not used when the first syntax element indicates that the LMCS tool is disabled, where the rules specify that when a second syntax element in the bitstream indicates that the LMCS tool is enabled at the picture header level of the video, the LMCS tool is enabled for all slices associated with the picture header of the video picture, where the rules specify that when the second syntax element indicates that the LMCS tool is disabled at the picture header level of the video, the LMCS tool is not used for all slices associated with the picture header, where the rules specify that when a third syntax element selectively included in the bitstream indicates that the LMCS tool is enabled at the slice header level of the video, the LMCS tool is used for the current slice associated with the slice header of the video picture, and where the rules specify that when the third syntax element indicates that the LMCS tool is disabled at the slice header level of the video, the LMCS tool is not used for the current slice.
[0525] In some embodiments of method 1500, the rules specify that when the LMCS tool is used for all slices of a video picture, the third syntax element is not included in the slice header of the bitstream. In some embodiments of method 1500, whether the LMCS tool is enabled or disabled is based on the second syntax element. In some embodiments of method 1500, the LMCS tool is enabled when it is used for all slices of a video picture, and the LMCS tool is disabled when it is not used for all slices of a video picture.
[0526] Figure 16 Flowchart of an example method 1600 for video processing. Operation 1602 includes: performing a conversion between a video including one or more video pictures and a bitstream of the video according to rules, where the rules specify that whether adaptive motion vector difference precision (AMVR) is used in the motion vector coding / decoding of an affine inter prediction mode is based on a syntax element selectively included in a reference sequence parameter set (SPS) indicating whether AMVR is enabled, where the rules specify that when the syntax element indicates that AMVR is disabled, AMVR is not used in the motion vector coding / decoding of the affine inter prediction mode, and where the rules specify that when the syntax element is not included in the SPS, AMVR is inferred not to be used in the motion vector coding / decoding of the affine inter prediction mode.
[0527] Figure 17Flowchart of an example method 1700 for video processing. Operation 1702 includes: performing a conversion between a video including video pictures and a bitstream of the video according to a rule, where the video pictures include sub-pictures, slices, and stripes, and where the rule specifies that since the sub-pictures include stripes segmented from slices, the conversion is performed by avoiding using the number of slices of the video pictures to calculate the height of the sub-pictures.
[0528] In some embodiments of method 1700, the height of the sub-pictures is calculated based on the number of coding tree units (CTUs). In some embodiments of method 1700, the height of the sub-pictures is less than one slice row.
[0529] Figure 18 Flowchart of an example method 1800 for video processing. Operation 1802 includes: performing a conversion between a video including video pictures and a bitstream of the video, where the bitstream indicates the height of the sub-pictures of the video pictures calculated based on the number of coding tree units (CTUs) of the video pictures.
[0530] In some embodiments of method 1800, the height of the sub-pictures is not based on the number of slices of the video pictures. In some embodiments of method 1800, the height of the sub-pictures is less than one slice row.
[0531] Figure 19 Flowchart of an example method 1900 for video processing. Operation 1902 includes: making a determination according to a rule, the determination being about whether the height of the sub-pictures of the video pictures of the video is less than the height of a slice row of the video pictures. Operation 1904 includes: using the determination to perform a conversion between the video and a bitstream of the video.
[0532] In some embodiments of method 1900, the rule specifies that the height of the sub-pictures is less than one slice row when: the sub-pictures include only coding tree units (CTUs) of one slice row, and a first group of CTUs located at the top of the sub-pictures is different from a second group of CTUs located at the top of one slice row, or a third group of CTUs located at the bottom of the sub-pictures is different from a fourth group of CTUs located at the bottom of the one slice row. In some embodiments of method 1900, when each sub-picture of the video pictures includes only one stripe and the height of the sub-pictures is less than one slice row, for each stripe of the video pictures having a picture-level stripe index i, the value of CtbAddrInSlice[i][j] is derived from the picture raster scan CTU address of the j-th CTU in the CTU raster scan of the sub-pictures, and j ranges from 0 to the number of CTUs in the stripe minus 1 (including the end values).
[0533] In some embodiments of method 1900, the rule stipulates that when the distance between a first set of CTUs at the top of a sub-picture and a second set of CTUs at the bottom of the sub-picture is less than a second height of a slice of the video picture, the height of the sub-picture is less than one slice row, where the second height of the slice is based on the number of CTUs of the sub-picture. In some embodiments of method 1900, when each sub-picture of the video picture includes only one strip and the height of the sub-picture is greater than or equal to one slice row, for each strip of the video picture having a picture-level strip index i, the value of CtbAddrInSlice[i][j] is derived from the picture raster scan CTU address of the j-th CTU in the order of CTUs in the sub-picture, and j ranges from 0 to the number of CTUs in the strip minus 1 (including the end values).
[0534] In some embodiments of method 1900, the order of CTUs in the sub-picture is such that the first CTU in the first slice having a first slice index is placed before the second CTU in the second slice having a second slice index, and the value of the first slice index is less than the value of the second slice index. In some embodiments of method 1900, the order of CTUs in the sub-picture is such that the CTUs within a slice in the sub-picture are sorted in the raster scan of CTUs in a slice.
[0535] In some embodiments of (multiple) methods 1100 - 1900, performing the conversion includes encoding the video into a bitstream. In some embodiments of (multiple) methods 1100 - 1900, performing the conversion includes encoding the video into a bitstream, and the method further includes storing the bitstream in a non-transitory computer-readable recording medium. In some embodiments of (multiple) methods 1100 - 1900, performing the conversion includes decoding the video from the bitstream.
[0536] In some embodiments, a video decoding device includes a processor configured to implement operations described for any one or more of methods 1100 to 1900. In some embodiments, a video encoding device includes a processor configured to implement operations described for any one or more of methods 1100 to 1900. In some embodiments, a computer program product having computer instructions stored thereon, which when executed by a processor cause the processor to implement operations described for any one or more of methods 1100 to 1900. In some embodiments, a non-transitory computer-readable storage medium stores a bitstream generated according to operations described for any one or more of methods 1100 to 1900. In some embodiments, a non-transitory computer-readable storage medium stores instructions that cause a processor to implement operations described for any one or more of methods 1100 to 1900. In some embodiments, a bitstream generation method includes: generating a bitstream of a video according to operations described for any one or more of methods 1100 to 1900, and storing the bitstream on a computer-readable program medium. In some embodiments, a method, a device, and a bitstream generated according to the disclosed methods or systems described in this document.
[0537] In this document, the term "video processing" may refer to video encoding, video decoding, video compression, or video decompression. For example, during the conversion from a pixel representation of a video to a corresponding bitstream representation, a video compression algorithm may be applied, and vice versa. As defined by the syntax, the bitstream representation of the current video block may, for example, correspond to bits that are co-located or scattered at different positions within the bitstream. For example, a macroblock may be encoded based on transform and codec error residual values and also using bits in the header and other fields in the bitstream. Further, during the conversion, the decoder may parse the bitstream based on this determination, knowing that some fields may or may not be present, as described in the above solutions. Similarly, the encoder may determine to include or exclude certain syntax fields and accordingly generate an encoded representation by including or excluding the syntax fields from the encoded representation.
[0538] The disclosures and other scenarios, examples, embodiments, modules, and functional operations described in this document can be implemented in digital electronic circuits or in computer software, firmware, or hardware, including the structures disclosed in this document and their structural equivalents, or a combination of one or more of them. The disclosed and other embodiments can be implemented as one or more computer program products encoded on a computer-readable medium, such as one or more computer program instruction modules, for execution by, or to control the operation of, a data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition affecting a machine-readable propagated signal, or a combination of one or more of them. The term "data processing apparatus" encompasses all apparatuses, devices, and machines for processing data, including, for example, programmable processors, computers, or multiple processors or computers. In addition to hardware, the apparatus can also include code that creates an execution environment for the computer programs being discussed, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them. A propagated signal is an artificially generated signal, for example, a machine-generated electrical, optical, or electromagnetic signal, which is generated to encode information for transmission to a suitable receiver device.
[0539] A computer program (also referred to as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, and can be deployed in any form, including as a stand-alone program or as modules, components, subroutines, or other units suitable for use in a computing environment. A computer program does not have to correspond to a file in a file system. The program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program being discussed, or in multiple cooperating files (e.g., files that store one or more modules, subroutines, or portions of code). A computer program can be deployed to execute on one computer or on multiple computers, which are located at one site or distributed across multiple sites and interconnected by a communication network.
[0540] The processes and logical flows described in this document can be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logical flows can also be performed by special-purpose logic circuitry, such as a field-programmable gate array (FPGA) or an application-specific integrated circuit (ASIC), and the apparatus can be implemented as special-purpose logic circuitry (e.g., an FPGA or an ASIC).
[0541] Processors suitable for the execution of a computer program include, for example, both general and special purpose microprocessors, as well as any one or more processors of any type of digital computer. Generally, a processor will receive instructions and data from a read-only memory or a random access memory or both. The basic elements of a computer are a processor for performing instructions and one or more memory devices for storing the instructions and data. Generally, a computer will also include one or more mass storage devices (such as magnetic, magneto-optical or optical disks) for storing data, or operatively coupled to receive data from a mass storage device (such as magnetic, magneto-optical or optical disks) or to transfer data to a mass storage device (such as magnetic, magneto-optical or optical disks), or both. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media and memory devices, including, for example, semiconductor memory devices (such as EPROM, EEPROM and flash memory devices); magnetic disks (such as internal hard disks or removable disks); magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and the memory may be supplemented by, or incorporated in, special purpose logic circuitry.
[0542] Although this patent document contains many details, these details should not be construed as limitations on any subject or the scope of what may be claimed, but rather as descriptions of features specific to particular embodiments of a particular technology. In this patent document, certain features described in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, the various features described in the context of a single embodiment may also be implemented separately in multiple embodiments or in various suitable sub-combinations. Additionally, although features may be described as acting in certain combinations and even initially claimed as such, in some cases one or more features from a claimed combination may be excluded from the combination, and the claimed combination may be directed to a sub-combination or variation of a sub-combination.
[0543] Similarly, although operations are depicted in the drawings in a particular order, this should not be understood to require that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, to achieve the desired result. Additionally, the separation of various system components in the embodiments described in this patent document should not be understood to require such separation in all embodiments.
[0544] Only a few implementations and examples have been described, and other implementations, enhancements and variations may be made based on what is described and illustrated in this patent document.
Claims
1. A video processing method, comprising: performing a conversion between a video including video pictures and a bitstream of the video, wherein the video pictures include one or more slices, and the one or more slices form one or more slice rows and one or more slice columns, wherein the bitstream complies with format rules, and wherein the format rules stipulate deriving a slice column index for each coding tree block (CTB) column of the slices of the video pictures, wherein the video pictures further include one or more sub-pictures, wherein each of the sub-pictures includes one or more stripes that jointly form a rectangular subset of the video pictures, and wherein the format rules further stipulate deriving the width of a sub-picture in terms of slice columns based on the slice column index of the leftmost CTB and / or the rightmost CTB included in the sub-picture, wherein the width of the i-th sub-picture in terms of slice columns, denoted as SubpicWidthInTiles[i], is derived as follows: for(i = 0; i <= sps_num_subpics_minus1; i++){ leftX = sps_subpic_ctu_top_left_x[i] rightX = leftX + sps_subpic_width_minus1[i] SubpicWidthInTiles[i] = ctbToTileColIdx[rightX] + 1 - ctbToTileColIdx[leftX] }, where sps_num_subpics_minus1 plus 1 represents the number of sub-pictures in the video pictures, where sps_subpic_ctu_top_left_x[i] represents the horizontal position of the top-left CTB of the i-th sub-picture, where sps_subpic_width_minus1[i] plus 1 specifies the width of the i-th sub-picture, and where ctbToTileColIdx[rightX] and ctbToTileColIdx[leftX] respectively represent the slice column indices of the leftmost CTB and the rightmost CTB included in the sub-picture.
2. The method according to claim 1, wherein, the slice column index of the ctbAddrX-th slice column, denoted as ctbToTileColIdx[ctbAddrX], is derived as follows: tileX = 0 for(ctbAddrX = 0; ctbAddrX <= PicWidthInCtbsY; ctbAddrX++){ if(ctbAddrX == tileColBd[tileX + 1]) tileX++ ctbToTileColIdx[ctbAddrX] = tileX }, where PicWidthInCtbsY represents the width of the video pictures in terms of CTBs, and where tileColBd[i] represents the position of the i-th slice column boundary in terms of CTBs.
3. The method according to claim 1, wherein, in response to the slice being divided into a plurality of rectangular strips and only a subset of the rectangular strips of the slice being included in the sub-picture, the slice is counted as one slice in the value of the width of the sub-picture.
4. The method according to claim 1, wherein, the format rule further specifies deriving a slice row index for each CTB row of the slices of the video picture.
5. The method according to claim 4, wherein, the slice row index of the ctbAddrY-th slice row, denoted as ctbToTileRowIdx[ctbAddrY], is derived as follows: tileY = 0 for (ctbAddrY = 0; ctbAddrY <= PicHeightInCtbsY; ctbAddrY++) { if (ctbAddrY == tileRowBd[tileY + 1]) tileY++ ctbToTileRowIdx[ctbAddrY] = tileY }, where PicHeightInCtbsY represents the height of the video picture in units of CTB, and where tileRowBd[i] represents the position of the i-th slice row boundary in units of CTB.
6. The method according to claim 4, wherein, the format rule further specifies that the height of the sub-picture in units of slice rows is derived based on the slice row indices of the top CTB and / or bottom CTB included in the sub-picture.
7. The method according to claim 6, wherein, the height of the i-th sub-picture in units of slice rows, denoted as SubpicHeightInTiles[i], is derived as follows: for (i = 0; i <= sps_num_subpics_minus1; i++) { topY = sps_subpic_ctu_top_left_y[i] bottomY = topY + sps_subpic_height_minus1[i] SubpicHeightInTiles[i] = ctbToTileRowIdx[botY] + 1 - ctbToTileRowIdx[topY] }, where sps_subpic_ctu_top_left_y[i] represents the vertical position of the top-left CTB of the i-th sub-picture, where sps_subpic_height_minus1[i] plus 1 specifies the height of the i-th sub-picture, and where ctbToTileRowIdx[botY] and ctbToTileRowIdx[topY] represent the slice row indices of the bottom CTB and top CTB included in the sub-picture, respectively.
8. The method according to claim 6, wherein, in response to the slice being divided into a plurality of rectangular strips and only a subset of the rectangular strips of the slice being included in the sub-picture, the slice is counted as one slice in the value of the height of the sub-picture.
9. The method according to claim 1, wherein, performing the conversion includes encoding the video into the bitstream.
10. The method according to claim 1, wherein, performing the conversion includes encoding the video into the bitstream, and the method further includes storing the bitstream in a non-transitory computer-readable recording medium.
11. The method according to claim 1, wherein, performing the conversion includes decoding the video from the bitstream.
12. A video data processing apparatus, comprising a processor and a non-transitory memory having instructions thereon, wherein, the instructions, when executed by the processor, cause the processor to: perform a conversion between a video including video pictures and a bitstream of the video, wherein the video pictures include one or more slices, and the one or more slices form one or more slice rows and one or more slice columns, wherein the bitstream conforms to format rules, and wherein the format rules specify deriving a slice column index for each coding tree block (CTB) column of the slices of the video pictures, the video pictures further include one or more sub-pictures, wherein each of the sub-pictures includes one or more stripes that together form a rectangular subset of the video picture, and wherein the format rules further specify deriving the width of the sub-picture in terms of slice columns based on the slice column index of the leftmost CTB and / or the rightmost CTB included in the sub-picture; wherein the width of the i-th sub-picture in terms of slice columns, denoted as SubpicWidthInTiles[i], is derived as follows: for(i = 0; i <= sps_num_subpics_minus1; i++){ leftX = sps_subpic_ctu_top_left_x[i] rightX = leftX + sps_subpic_width_minus1[i] SubpicWidthInTiles[i] = ctbToTileColIdx[rightX] + 1 - ctbToTileColIdx[leftX] }, where sps_num_subpics_minus1 plus 1 represents the number of sub-pictures in the video picture, where sps_subpic_ctu_top_left_x[i] represents the horizontal position of the top-left CTB of the i-th sub-picture, where sps_subpic_width_minus1[i] plus 1 specifies the width of the i-th sub-picture, and where ctbToTileColIdx[rightX] and ctbToTileColIdx[leftX] respectively represent the slice column indices of the leftmost CTB and the rightmost CTB included in the sub-picture.
13. The apparatus according to claim 12, wherein, the slice column index of the ctbAddrX-th slice column, denoted as ctbToTileColIdx[ctbAddrX], is derived as follows: tileX = 0 for (ctbAddrX = 0; ctbAddrX <= PicWidthInCtbsY; ctbAddrX++) { if (ctbAddrX == tileColBd[tileX + 1]) tileX++ ctbToTileColIdx[ctbAddrX] = tileX }, where PicWidthInCtbsY represents the width of the video picture in terms of CTBs, and where tileColBd[i] represents the position of the i-th tile column boundary in terms of CTBs; where the format rule also stipulates deriving the tile row index for each CTB row of the tiles of the video picture; where the tile row index of the ctbAddrY-th tile row, denoted as ctbToTileRowIdx[ctbAddrY], is derived as follows: tileY = 0 for (ctbAddrY = 0; ctbAddrY <= PicHeightInCtbsY; ctbAddrY++) { if (ctbAddrY == tileRowBd[tileY + 1]) tileY++ ctbToTileRowIdx[ctbAddrY] = tileY }, where PicHeightInCtbsY represents the height of the video picture in terms of CTBs, and where tileRowBd[i] represents the position of the i-th tile row boundary in terms of CTBs.
14. The apparatus according to claim 12, where the format rule also stipulates that the height of the sub-picture in terms of tile rows is derived based on the tile row indices of the top CTB and / or bottom CTB included in the sub-picture; where the height of the i-th sub-picture in terms of tile rows, denoted as SubpicHeightInTiles[i], is derived as follows: for (i = 0; i <= sps_num_subpics_minus1; i++) { topY = sps_subpic_ctu_top_left_y[i] bottomY = topY + sps_subpic_height_minus1[i] SubpicHeightInTiles[i] = ctbToTileRowIdx[botY] + 1 - ctbToTileRowIdx[topY] }, where sps_subpic_ctu_top_left_y[i] represents the vertical position of the top-left CTB of the i-th sub-picture, where adding 1 to sps_subpic_height_minus1[i] stipulates the height of the i-th sub-picture, and where ctbToTileRowIdx[botY] and ctbToTileRowIdx[topY] respectively represent the tile row indices of the bottom CTB and top CTB included in the sub-picture; Wherein, in response to the slice being divided into a plurality of rectangular strips and only a subset of the rectangular strips of the slice being included in the sub-picture, the slice is counted as one slice in the value of the height of the sub-picture.
15. A non-transitory computer-readable storage medium storing instructions that cause a processor to: perform a conversion between a video including video pictures and a bitstream of the video, wherein, the video pictures include one or more slices, and the one or more slices form one or more slice rows and one or more slice columns, wherein the bitstream conforms to format rules, and wherein the format rules specify deriving a slice column index for each coding tree block (CTB) column of the slices of the video pictures, wherein the video pictures further include one or more sub-pictures, wherein each of the sub-pictures includes one or more strips that jointly form a rectangular subset of the video pictures, and wherein the format rules further specify deriving the width of the sub-picture in terms of slice columns based on the slice column indices of the leftmost CTB and / or the rightmost CTB included in the sub-picture, wherein the width of the i-th sub-picture in terms of slice columns, denoted as SubpicWidthInTiles[i], is derived as follows: for(i = 0; i <= sps_num_subpics_minus1; i++){ leftX = sps_subpic_ctu_top_left_x[i] rightX = leftX + sps_subpic_width_minus1[i] SubpicWidthInTiles[i] = ctbToTileColIdx[rightX] + 1 - ctbToTileColIdx[leftX] }, wherein sps_num_subpics_minus1 plus 1 represents the number of sub-pictures in the video pictures, wherein sps_subpic_ctu_top_left_x[i] represents the horizontal position of the top-left CTB of the i-th sub-picture, wherein sps_subpic_width_minus1[i] plus 1 specifies the width of the i-th sub-picture, and wherein ctbToTileColIdx[rightX] and ctbToTileColIdx[leftX] respectively represent the slice column indices of the leftmost CTB and the rightmost CTB included in the sub-picture.
16. The non-transitory computer-readable storage medium according to claim 15, wherein, the slice column index of the ctbAddrX-th slice column, denoted as ctbToTileColIdx[ctbAddrX], is derived as follows: tileX = 0 for(ctbAddrX = 0; ctbAddrX <= PicWidthInCtbsY; ctbAddrX++){ if(ctbAddrX == tileColBd[tileX + 1]) tileX++ ctbToTileColIdx[ctbAddrX] = tileX }, where PicWidthInCtbsY represents the width of the video picture in terms of CTBs, and where tileColBd[i] represents the position of the i-th tile column boundary in terms of CTBs; where the format rule further stipulates deriving a tile column index for each CTB column of the tiles of the video picture; where the tile row index of the ctbAddrY-th tile row, denoted as ctbToTileRowIdx[ctbAddrY], is derived as follows: tileY = 0 for (ctbAddrY = 0; ctbAddrY <= PicHeightInCtbsY; ctbAddrY++) { if (ctbAddrY == tileRowBd[tileY + 1]) tileY++ ctbToTileRowIdx[ctbAddrY] = tileY }, where PicHeightInCtbsY represents the height of the video picture in terms of CTBs, and where tileRowBd[i] represents the position of the i-th tile row boundary in terms of CTBs.
17. A non-transitory computer-readable recording medium storing a bitstream of a video generated by a method executed by a video processing device, wherein, the method includes: generating the bitstream of the video including the video picture according to a format rule, where the video picture includes one or more tiles, and the one or more tiles form one or more tile rows and one or more tile columns, where the bitstream conforms to the format rule, and where the format rule stipulates deriving a tile column index for each coding tree block CTB column of the tiles of the video picture, where the video picture further includes one or more sub-pictures, where each sub-picture includes one or more stripes that jointly form a rectangular subset of the video picture, and where the format rule further stipulates deriving the width of the sub-picture in terms of tile columns based on the tile column indices of the leftmost CTB and / or the rightmost CTB included in the sub-picture, where the width of the i-th sub-picture in terms of tile columns, denoted as SubpicWidthInTiles[i], is derived as follows: for (i = 0; i <= sps_num_subpics_minus1; i++) { leftX = sps_subpic_ctu_top_left_x[i] rightX = leftX + sps_subpic_width_minus1[i] SubpicWidthInTiles[i] = ctbToTileColIdx[rightX] + 1 - ctbToTileColIdx[leftX] }, where sps_num_subpics_minus1 plus 1 represents the number of sub-pictures in the video picture, where sps_subpic_ctu_top_left_x[i] represents the horizontal position of the top - left CTB of the i - th sub - picture, where sps_subpic_width_minus1[i]+1 defines the width of the i - th sub - picture, and where ctbToTileColIdx[rightX] and ctbToTileColIdx[leftX] respectively represent the tile column indices of the left - most CTB and the right - most CTB included in the sub - picture.
18. The non - transitory computer - readable recording medium according to claim 17, where, the tile column index of the ctbAddrX - th tile column, denoted as ctbToTileColIdx[ctbAddrX], is derived as follows: tileX = 0 for(ctbAddrX = 0; ctbAddrX <= PicWidthInCtbsY; ctbAddrX++){ if(ctbAddrX == tileColBd[tileX + 1]) tileX++ ctbToTileColIdx[ctbAddrX] = tileX }, where PicWidthInCtbsY represents the width of the video picture in terms of CTBs, and where tileColBd[i] represents the position of the i - th tile column boundary in terms of CTBs; where the format rule also specifies deriving the tile row index for each CTB row of the slices of the video picture; where the tile row index of the ctbAddrY - th tile row, denoted as ctbToTileRowIdx[ctbAddrY], is derived as follows: tileY = 0 for(ctbAddrY = 0; ctbAddrY <= PicHeightInCtbsY; ctbAddrY++){ if(ctbAddrY == tileRowBd[tileY + 1]) tileY++ ctbToTileRowIdx[ctbAddrY] = tileY }, where PicHeightInCtbsY represents the height of the video picture in terms of CTBs, and where tileRowBd[i] represents the position of the i - th tile row boundary in terms of CTBs.
19. The non - transitory computer - readable recording medium according to claim 17, where, the format rule also specifies that the height of the sub - picture in terms of tile rows is derived based on the tile row indices of the top CTB and / or bottom CTB included in the sub - picture; where the height of the i - th sub - picture in terms of tile rows, denoted as SubpicHeightInTiles[i], is derived as follows: for(i = 0; i <= sps_num_subpics_minus1; i++){ topY = sps_subpic_ctu_top_left_y[i] bottomY = topY + sps_subpic_height_minus1[i] SubpicHeightInTiles[i] = ctbToTileRowIdx[botY] + 1 - ctbToTileRowIdx[topY] }, where sps_subpic_ctu_top_left_y[i] represents the vertical position of the top - left CTB of the i - th sub - picture, where adding 1 to sps_subpic_height_minus1[i] defines the height of the i - th sub - picture, and where ctbToTileRowIdx[botY] and ctbToTileRowIdx[topY] respectively represent the tile row indices of the bottom CTB and the top CTB included in the sub - picture; where, in response to the slice being divided into a plurality of rectangular strips and only a subset of the rectangular strips of the slice being included in the sub - picture, the slice is counted as one slice in the height value of the sub - picture.
20. A method for storing a bit - stream of a video, comprising: generating the bit - stream of the video including video pictures according to format rules, and storing the bit - stream in a non - transitory computer - readable recording medium, where the video pictures include one or more slices, and the one or more slices form one or more slice rows and one or more slice columns, where the bit - stream conforms to the format rules, and where the format rules specify deriving a slice column index for each coding tree block (CTB) column of the slices of the video pictures, where the video pictures further include one or more sub - pictures, where each sub - picture includes one or more strips that jointly form a rectangular subset of the video picture, and where the format rules also specify deriving the width of the sub - picture in terms of slice columns based on the slice column indices of the left - most CTB and / or the right - most CTB included in the sub - picture, where the width of the i - th sub - picture in terms of slice columns, denoted as SubpicWidthInTiles[i], is derived as follows: for(i = 0; i <= sps_num_subpics_minus1; i++){ leftX = sps_subpic_ctu_top_left_x[i] rightX = leftX + sps_subpic_width_minus1[i] SubpicWidthInTiles[i] = ctbToTileColIdx[rightX] + 1 - ctbToTileColIdx[leftX] }, where adding 1 to sps_num_subpics_minus1 represents the number of sub - pictures in the video picture, where sps_subpic_ctu_top_left_x[i] represents the horizontal position of the top - left CTB of the i - th sub - picture, wherein, sps_subpic_width_minus1[i] plus 1 defines the width of the i-th sub-picture, and wherein, ctbToTileColIdx[rightX] and ctbToTileColIdx[leftX] respectively represent the tile column indices of the leftmost CTB and the rightmost CTB included in the sub-picture.
21. A video decoding apparatus, comprising a processor configured to implement the method according to any one of claims 1 to 11.
22. A video encoding apparatus, comprising a processor configured to implement the method according to any one of claims 1 to 11.
23. A computer program product having computer instructions stored thereon, the instructions, when executed by a processor, causing the processor to implement the method according to any one of claims 1 to 11.
24. A non-transitory computer-readable storage medium storing a bitstream, the bitstream being generated by a video processing apparatus implementing the method according to any one of claims 1 to 11.
25. A non-transitory computer-readable storage medium storing instructions, the instructions causing a processor to implement the method according to any one of claims 1 to 11.
26. A method for generating a bitstream, comprising: generating a bitstream of a video according to the method according to any one of claims 1 to 11, and storing the bitstream on a computer-readable program medium.
Citation Information
Patent Citations
Signalling of video content including sub-picture bitstreams for video coding
CN110431849A
An apparatus, a method and a computer program for video coding and decoding
WO2019073112A1