Derivation of the height of the sub - picture
By introducing two-level control mechanisms and semantic optimization in the video encoding and decoding standard VVC, the problem of inefficient LMCS signaling notification is solved, and the encoding and decoding efficiency and performance are improved.
Patent Information
- Application Number
- CN202180016671.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-02-24
- Filing Date
- 2021-02-23
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2041-02-23
AI Technical Summary
The existing video codec standard VVC has problems with inefficient signaling notifications and incorrect semantics in sub-picture and brightness mapping chromaticity scaling (LMCS) control, resulting in waste of resources and degradation of performance during codec.
A two-level control mechanism is introduced to optimize the signaling notification process by signaling the activation and disabling of LMCS at the picture level and the stripe level, and adjust the semantic definitions in SPS, PH and SH to improve efficiency.
It improves the efficiency of the encoding and decoding process, reduces unnecessary bitstream overhead, and improves the performance and resource utilization of video encoding and decoding.
Smart Images

Figure CN115152210B_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This application is the U.S. national phase entry of International Patent Application No. PCT / US2021 / 019228, filed on February 23, 2021, which claims the priority of U.S. Provisional Patent Application No. US 62 / 980,963, filed on February 24, 2020. The entire disclosure of the above applications is incorporated by reference as part of the disclosure of this application. Technical Field
[0003] This patent document relates to image and video encoding, decoding, and transcoding. Background Art
[0004] Digital video accounts for the largest bandwidth usage on the Internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth demand for digital video usage is expected to continue to grow. Summary of the Invention
[0005] This document discloses techniques that can be used by video encoders and decoders for processing the encoded representation of video using control information useful for decoding the encoded representation.
[0006] In one example aspect, a video processing method is disclosed. The method includes: performing a conversion between a video including one or more video pictures and a bitstream of the video, where each video picture includes one or more slices, the one or more slices include one or more slice columns, the bitstream conforms to format rules, and the format rules specify deriving a slice column index for each coding tree unit (CTU) column of the slices of the video picture.
[0007] In another example aspect, a video processing method is disclosed. The method includes: performing a conversion between a video including one or more video pictures and a bitstream of the video, where each video picture includes one or more slices, the one or more slices include one or more slice columns, the bitstream conforms to format rules, and the format rules specify deriving a slice row index for each coding tree unit (CTU) row of the slices of the video picture.
[0008] In another example aspect, a video processing method is disclosed. The method includes: performing a conversion between a video including at least one video picture and a bitstream of the video according to rules, where the at least one video picture includes one or more strips and one or more sub - pictures, and the rules specify an order of strip indices of one or more strips in the at least one video picture in response to a syntax element associated with the at least one video picture indicating whether each sub - picture of the at least one video picture includes a single strip.
[0009] In another example aspect, a video processing method is disclosed. The method includes: performing a conversion between a video unit of a video region of a video and a bitstream of the video, where the bitstream conforms to format rules, where the format rules stipulate that first-level first control information of the video region in the bitstream controls whether a second level of the video units in the bitstream includes second control information, where the second level is less than the first level, where the first control information and the second control information include information on whether or how a luminance mapping and chrominance scaling (LMCS) tool is applied to the video unit, and where the LMCS tool includes using chrominance residual scaling (CRS) or a luminance remodelling process (RP) for the conversion.
[0010] In another example aspect, a video processing method is disclosed. The method includes: performing a conversion between a video and a bitstream of the video according to rules, where the rules stipulate that a luminance mapping and chrominance scaling (LMCS) tool is enabled when a first syntax element in a reference sequence parameter set indicates that the LMCS tool is enabled, where the rules stipulate that the LMCS tool is not used when the first syntax element indicates that the LMCS tool is disabled, where the rules stipulate that when a second syntax element in the bitstream indicates that the LMCS tool is enabled at the picture header level of the video, the LMCS tool is enabled for all slices associated with the picture header of the video picture, where the rules stipulate that when the second syntax element indicates that the LMCS tool is disabled at the picture header level of the video, the LMCS tool is not used for all slices associated with the picture header, where the rules stipulate that when a third syntax element selectively included in the bitstream indicates that the LMCS tool is enabled at the slice header level of the video, the LMCS tool is used for the current slice associated with the slice header of the video picture, and where the rules stipulate that when the third syntax element indicates that the LMCS tool is disabled at the slice header level of the video, the LMCS tool is not used for the current slice.
[0011] In another example aspect, a video processing method is disclosed. The method includes: performing a conversion between a video including one or more video pictures and a bitstream of the video according to rules, where the rules stipulate that whether adaptive motion vector difference precision (AMVR) is used in the motion vector encoding and decoding of an affine inter prediction mode is based on a syntax element selectively included in a reference sequence parameter set (SPS) indicating whether AMVR is enabled, where the rules stipulate that when the syntax element indicates that AMVR is disabled, AMVR is not used in the motion vector encoding and decoding of the affine inter prediction mode, and where the rules stipulate that when the syntax element is not included in the SPS, AMVR is inferred not to be used in the motion vector encoding and decoding of the affine inter prediction mode.
[0012] In another example aspect, a video processing method is disclosed. The method includes: performing a conversion between a video including video pictures and a bitstream of the video according to a rule, where the video pictures include sub-pictures, slices, and stripes, and where the rule specifies that since the sub-pictures include stripes segmented from slices, the conversion is performed by avoiding using the number of slices of the video pictures to calculate the height of the sub-pictures.
[0013] In another example aspect, a video processing method is disclosed. The method includes: performing a conversion between a video including video pictures and a bitstream of the video, where the bitstream indicates the height of sub-pictures of the video pictures calculated based on the number of coding tree units (CTUs) of the video pictures.
[0014] In another example aspect, a video processing method is disclosed. The method includes: making a determination according to a rule, the determination being about whether the height of sub-pictures of video pictures of a video is less than the height of slice rows of the video pictures; and using the determination to perform a conversion between the video and the bitstream of the video.
[0015] In another example aspect, a video processing method is disclosed. The method includes: performing a conversion between a video including one or more video pictures and a coded representation of the video, where each video picture includes one or more slices, and where the coded representation conforms to format rules; where the format rules specify first information signaled in the coded representation and second information derived from the coded representation, where at least the first information or the second information is related to the row index or column index of one or more slices.
[0016] In another example aspect, a video processing method is disclosed. The method includes: performing a conversion between video units of a video region of a video and a coded representation of the video, where the coded representation conforms to format rules; where the format rules specify that first control information at the video region controls whether second control information is included at the video unit level; where the first control information and / or the second control information includes information about luminance mapping and chrominance scaling (LMCS) or chrominance residual scaling (CRS) or remodelling process (RP) for the conversion.
[0017] In yet another example aspect, a video encoder device is disclosed. The video encoder includes a processor configured to implement the above method.
[0018] In yet another example aspect, a video decoder device is disclosed. The video decoder includes a processor configured to implement the above method.
[0019] In yet another example aspect, a computer-readable medium having code stored thereon is disclosed. The code embodies one of the methods described herein in the form of processor-executable code.
[0020] These and other features are described in this document. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 An example of raster scan strip segmentation of a picture is shown, where the picture is divided into 12 slices and 3 raster scan strips.
[0022] Figure 2 An example of rectangular strip segmentation of a picture is shown, where the picture is divided into 24 slices (6 slice columns and 4 slice rows) and 9 rectangular strips.
[0023] Figure 3 An example of a picture segmented into slices and rectangular strips is shown, where the picture is divided into 4 slices (2 slice columns and 2 slice rows) and 4 rectangular strips.
[0024] Figure 4 A picture segmented into 15 slices, 24 strips, and 24 sub - pictures is shown.
[0025] Figure 5 Is a block diagram of an example video processing system.
[0026] Figure 6 Is a block diagram of a video processing device.
[0027] Figure 7 Is a flowchart of an example method for video processing.
[0028] Figure 8 Is a block diagram showing a video codec system according to some embodiments of the present disclosure.
[0029] Figure 9 Is a block diagram showing an encoder according to some embodiments of the present disclosure.
[0030] Figure 10 Is a block diagram showing a decoder according to some embodiments of the present disclosure.
[0031] Figures 11 to 19 Is a flowchart of an example method for video processing. DETAILED DESCRIPTION
[0032] In this document, chapter headings are used for easy understanding, and the applicability of the technologies and embodiments disclosed in each chapter is not limited to that chapter only. In addition, in some descriptions, H.266 technical terms are used only for easy understanding and not to limit the scope of the disclosed technologies. Therefore, the technologies described herein are also applicable to other video codec protocols and designs. In this document, with respect to the current draft of the VVC specification, deleted text between double brackets (e.g., [[]]) and added text in bold italic are used to indicate editorial changes to the text.
[0033] 1. Overview
[0034] This document relates to video coding and decoding technology. Specifically, it is about the support for sub - pictures, LMCS, and AMVR. Regarding the aspect of sub - pictures, it includes deriving the number of slice rows and slice columns included in a sub - picture, and deriving a list of raster - scanned CTU addresses of CTUs included in a strip when each sub - picture contains only one strip. The LMCS aspect is about signaling to enable LMCS at different levels. The AMVR aspect is about the semantics of the sps_affine_amvr_enabled_flag. These ideas can be applied individually or in various combinations to any video coding standard or non - standard video codec that supports single - layer and / or multi - layer video coding and decoding, such as the Versatile Video Coding (VVC) being developed.
[0035] 2. Abbreviations
[0036] ALF Adaptive Loop Filter
[0037] AMVR Adaptive Motion Vector Resolution Difference
[0038] APS Adaptive Parameter Set
[0039] AU Access Unit
[0040] AUD Access Unit Delimiter
[0041] AVC Advanced Video Coding
[0042] CLVS Coding - Layer Video Sequence
[0043] CPB Coding - Picture Buffer
[0044] CRA Clean Random Access
[0045] CTU Coding - Tree Unit
[0046] CVS Coding - Video Sequence
[0047] DPB Decoded - Picture Buffer
[0048] DPS Decoding Parameter Set
[0049] EOB End of Bit - stream
[0050] EOS End of Sequence
[0051] GDR Gradual Decoding Refresh
[0052] HEVC High - Efficiency Video Coding
[0053] HRD Hypothetical Reference Decoder
[0054] IDR Instantaneous Decoding Refresh
[0055] JEM Joint Exploration Model
[0056] LMCS Luma Mapping with Chroma Scaling
[0057] MCTS Motion Constrained Tile Set
[0058] NAL Network Abstraction Layer
[0059] OLS Output Layer Set
[0060] PH Picture Header
[0061] PPS Picture Parameter Set
[0062] PTL Profile, Tier and Level
[0063] PU Picture Unit
[0064] RBSP Raw Byte Sequence Payload
[0065] SEI Supplementary Enhancement Information
[0066] SPS Sequence Parameter Set
[0067] SVC Scalable Video Coding
[0068] VCL Video Coding Layer
[0069] VPS Video Parameter Set
[0070] VTM VVC Test Model
[0071] VUI Video Usability Information
[0072] VVC Versatile Video Coding
[0073] 3. Preliminary Discussion
[0074] Video coding standards have mainly evolved through the well-known ITU-T and ISO / IEC standards. ITU-T developed H.261 and H.263, ISO / IEC developed MPEG-1 and MPEG-4 Visual, and the two organizations jointly developed the H.262 / MPEG-2 video standard, the H.264 / MPEG-4 Advanced Video Coding (AVC) standard, and the H.265 / HEVC standard. Starting from H.262, video coding standards are based on a hybrid video coding structure, where temporal prediction plus transform coding is used. To explore future video coding technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Team (JVET) in 2015. Since then, JVET has adopted many new methods and incorporated them into a reference software called the Joint Exploration Model (JEM). JVET meetings are held quarterly simultaneously. The goal of the new coding standard is to reduce the bitrate by 50% compared to HEVC. The new video coding standard was officially named Versatile Video Coding (VVC) at the JVET meeting in April 2018, and the first version of the VVC Test Model (VTM) was released at that time. Due to the continuous efforts on VVC standardization, new coding technologies are incorporated into the VVC standard at each JVET meeting. The working draft of VVC and the test model VTM are updated after each meeting. The VVC project now aims to be technically completed (FDIS) at the meeting in July 2020.
[0075] 3.1. Picture segmentation schemes in HEVC
[0076] HEVC includes four different picture segmentation schemes, namely regular slices, non-independent slices, tiles, and wavefront parallel processing (WPP), which can be applied to maximum transmission unit (MTU) size matching, parallel processing, and reducing end-to-end latency.
[0077] Regular slices are similar to those in H.264 / AVC. Each regular slice is encapsulated in its own NAL unit, and intra-picture prediction (intra-sample prediction, motion information prediction, coding mode prediction) and entropy coding dependencies across slice boundaries are disabled. Therefore, a regular slice can be reconstructed independently of other regular slices within the same picture (although there may still be dependencies due to loop filter operations).
[0078] Regular slices are the only tool available for parallelization, and this tool is also available in almost the same form in H.264 / AVC. Parallelization based on regular slices does not require much inter-processor or inter-core communication (except for the inter-processor or inter-core data sharing for motion compensation when decoding predicted coded pictures, which is usually much more than the inter-processor or inter-core data sharing due to intra-picture prediction). However, for the same reason, using regular slices may incur a large amount of coding and decoding overhead due to the bit cost of slice headers and the lack of prediction across slice boundaries. In addition, due to the intra-picture independence of regular slices and each regular slice being encapsulated in its own NAL unit, regular slices (compared with other tools mentioned below) can also serve as a key mechanism for bitstream segmentation to match the MTU size requirements. In many cases, the requirements for the slice layout in the picture for the goals of parallelization and MTU size matching are contradictory. The implementation of this situation has led to the development of the parallelization tools mentioned below.
[0079] Non-independent slices have short slice headers and allow bitstream segmentation at tree block boundaries without breaking any intra-picture prediction. Basically, non-independent slices divide a regular slice into multiple NAL units, reducing the end-to-end latency by allowing a part of the regular slice to be sent before the entire regular slice is encoded.
[0080] In WPP, a picture is divided into single rows of coding tree blocks (CTBs). Entropy decoding and prediction are allowed to use data from CTBs in other partitions. Parallel processing can be performed through the parallel decoding of CTB rows, where the decoding of a CTB row starts with a delay of two CTBs to ensure that the data related to the CTBs above and to the right of the main CTB can be obtained before the main CTB being decoded. Using this staggered start (which looks like a wavefront when represented graphically), as many processors / cores as there are CTB rows in the picture can be parallelized. Since intra-picture prediction between adjacent tree block rows within the picture is allowed, the inter-processor / inter-core communication required to enable intra-picture prediction can be substantial. Compared with not applying WPP partitioning, WPP partitioning does not result in the generation of additional NAL units, so WPP is not a tool for MTU size matching. However, if MTU size matching is required, regular slices can be used together with WPP, but with a certain amount of coding and decoding overhead.
[0081] Slice definitions divide a picture into horizontal and vertical boundaries of slice columns and slice rows. Slice columns extend from the top to the bottom of the picture. Similarly, slice rows extend from the left to the right of the picture. The number of slices in a picture can be simply derived by multiplying the number of slice columns by the number of slice rows.
[0082] Before decoding the top-left CTB of the next slice in the order of the slice raster scan of the picture, the scan order of the CTBs is changed to the local scan order within the slice (in the order of the CTB raster scan of the slice). Similar to a regular strip, a slice breaks the prediction dependency and the entropy decoding dependency within the picture. However, they do not need to be included in separate NAL units (the same as WPP in this regard); thus, slices cannot be used for MTU size matching. Each slice can be processed by one processor / core, and in the case where a strip spans multiple slices, the inter-processor / inter-core communication required for intra-picture prediction between the processing units decoding adjacent slices is limited to transmitting the shared strip header and loop filtering related to the sharing of reconstructed samples and metadata. When a strip contains more than one slice or WPP segment, the entry point byte offset of each slice or WPP segment except the first one in the strip is signaled in the strip header.
[0083] For simplicity, restrictions on the application of four different picture partitioning schemes are specified in HEVC. For most profiles specified in HEVC, a given coded video sequence cannot contain both slices and wavefronts simultaneously. For each strip and slice, one or both of the following conditions must be met: 1) all coded tree blocks in the strip belong to the same slice; 2) all coded tree blocks in the slice belong to the same strip. Finally, a wavefront segment exactly contains one CTB row, and when using WPP, if a strip starts within a CTB row, the strip must end in the same CTB row.
[0084] The most recent revision of HEVC is specified in the JCT-VC output document JCTVC-AC1005, J. Boyce, A. Ramasubramonian, R. Skupin, G. J. Sullivan, A. Tourapis, Y.-Kwang (editors). "HEVC Additional Supplemental Enhancement Information (Draft 4)", October 24, 2017 is publicly available here: http: / / phenix.intevry.fr / jct / doc_end_user / documents- / 29_Macau / wg11 / JCTVC-AC1005-v2.zip. Included within this revision, HEVC specifies three SEI messages related to MCT, namely the Temporal MCTS SEI message, the MCTS Extraction Information Set SEI message, and the MCTS Extraction Information Nested SEI message.
[0085] The time-domain MCTS SEI message indicates the presence of MCTS in the bitstream and signals the MCTS. For each MCTS, the motion vectors are restricted to point to full-sample positions within the MCTS and fractional-sample positions that only require full-sample positions within the MCTS for interpolation, and motion vector candidates predicted from time-domain motion vectors derived from blocks outside the MCTS are not allowed. In this way, each MCTS can be independently decoded without the presence of slices not included in the MCTS.
[0086] The MCTS extraction information set SEI message provides supplementary information that can be used for MCTS sub-bitstream extraction (the part specified as the semantics of the SEI message) to generate a compliant bitstream for the MCTS set. This information consists of multiple extraction information sets, each extraction information set defining multiple MCTS sets and containing the RBSP bytes to replace the VPS, SPS, and PPS to be used during the MCTS sub-bitstream extraction process. When extracting the sub-bitstream according to the MCTS sub-bitstream extraction process, the parameter sets (VPS, SPS, and PPS) need to be rewritten or replaced, and the slice header needs to be slightly updated because the strip addresses related to one or all of the syntax elements (including first_slice_segment_in_pic_flag and slice_segment_address) usually need to have different values.
[0087] 3.2. Segmentation of Pictures in VVC
[0088] In VVC, a picture is divided into one or more slice rows and one or more slice columns. A slice is a sequence of CTUs that cover a rectangular region of the picture. The CTUs within a slice are scanned in raster scan order within that slice.
[0089] A strip consists of an integer number of complete slices or an integer number of consecutive complete CTU rows within a slice of the picture.
[0090] Two strip modes are supported, namely the raster scan strip mode and the rectangular strip mode. In the raster scan strip mode, a strip contains a sequence of complete slices in the slice raster scan of the picture. In the rectangular strip mode, a strip contains multiple complete slices that together form a rectangular region of the picture or multiple consecutive complete CTU rows of a single slice that together form a rectangular region of the picture. The slices within a rectangular strip are scanned in slice raster scan order within the rectangular region corresponding to that strip.
[0091] A sub-picture contains one or more strips that together cover a rectangular region of the picture.
[0092] Figure 1 An example of the raster scan strip segmentation of a picture is shown, where the picture is divided into 12 slices and 3 raster scan strips.
[0093] Figure 2 An example of rectangular strip segmentation of a picture is shown, where the picture is segmented into 24 slices (6 slice columns and 4 slice rows) and 9 rectangular slices.
[0094] Figure 3 An example of a picture segmented into slices and rectangular strips is shown, where the picture is divided into 4 slices (2 slice columns and 2 slice rows) and 4 rectangular strips.
[0095] Figure 4 An example of sub - picture segmentation of a picture is shown, where the picture is segmented into 18 slices, 12 slices on the left (each covering a strip with 4x4 CTUs) and 6 slices on the right (each covering 2 vertically stacked strips with 2x2 CTUs), resulting in a total of 24 strips and 24 sub - pictures of different dimensions (each strip is a sub - picture).
[0096] 3.3. Signaling of SPS / PPS / Picture Header / Strip Header in VVC (such as JVET - Q2001 - vC) 7.3.2.3 Sequence Parameter Set RBSP Syntax
[0097]
[0098]
[0099]
[0100]
[0101]
[0102]
[0103]
[0104]
[0105]
[0106] 7.3.2.4 Picture Parameter Set RBSP Syntax
[0107]
[0108]
[0109]
[0110]
[0111]
[0112]
[0113] 7.3.2.7 Picture Header Structure Syntax
[0114]
[0115]
[0116]
[0117]
[0118]
[0119]
[0120] 7.3.7.1 Regular Strip Header Syntax
[0121]
[0122]
[0123]
[0124]
[0125]
[0126] 3.4. Specifications for slices, strips, and sub - pictures in JVET - Q2001 - vC
[0127] 3 Definitions
[0128] Picture - level strip index: Index of the strip into the list of strips in the picture, in the order they are signaled in the PPS when rect_slice_flag equals 1.
[0129] Sub - picture - level strip index: Index of the strip into the list of strips in the sub - picture, in the order they are signaled in the PPS when rect_slice_flag equals 1.
[0130] 6.5.1 CTB Raster Scan, Slice Scan, and Sub - picture Scan Procedures
[0131] The variable NumTileColumns specifies the number of tile columns, and the list colWidth[i] for i ranging from 0 to NumTileColumn - 1 (inclusive) specifies the width of the i - th tile column in CTBs, derived as follows:
[0132]
[0133] The variable NumTileRows specifies the number of tile rows, and a list RowHeight[j] for j ranging from 0 to NumTileRows-1 (including the end values) specifies the height of the j-th tile row in units of CTB, and is derived as follows:
[0134]
[0135]
[0136] The variable NumTilesInPic is set to be equal to NumTileColumns * NumTileRows.
[0137] A list tileColBd[i] for i ranging from 0 to NumTileColumns (including the end values) specifies the position of the i-th tile column boundary in units of CTB, and is derived as follows:
[0138] for (tileColBd[0] = 0, i = 0; i < NumTileColumns; i++)
[0139] tileColBd[i + 1] = tileColBd[i] + colWidth[i] (25)
[0140] Note 1 – The size of the array tileColBd[] is 1 larger than the actual number of tile columns in the derivation of CtbToTileColBd[].
[0141] A list tileRowBd[j] for j ranging from 0 to NumTileRows (including the end values) specifies the position of the j-th tile row boundary in units of CTB, and is derived as follows:
[0142] for (tileRowBd[0] = 0, j = 0; j < NumTileRows; j++)
[0143] tileRowBd[j + 1] = tileRowBd[j] + RowHeight[j] (26)
[0144] Note 2 – The size of the array tileRowBd[] in the above derivation is 1 larger than the actual number of tile rows in the derivation of CtbToTileRowBd[].
[0145] The list CtbToTileColBd[ctbAddrX] of ctbAddrX ranging from 0 to PicWidthInCtbsY (including the end values) specifies the conversion from the horizontal CTB address to the left tile column boundary in terms of CTBs, and is derived as follows:
[0146]
[0147] Note 3 – The size of the array CtbToTileColBd[] in the above derivation is 1 larger than the actual number of picture widths in CTBs in the derivation of slice_data() signaling.
[0148] The list CtbToTileRowBd[ctbAddrY] of ctbAddrY ranging from 0 to PicHeightInCtbsY (including the end values) specifies the conversion from the vertical CTB address to the top tile column boundary in terms of CTBs, and is derived as follows:
[0149]
[0150] Note 4 – The size of the array CtbToTileRowBd[] in the above derivation is 1 larger than the actual number of picture heights in CTBs in the slice_data() signaling notification.
[0151] For rectangular stripes, the list NumCtusInSlice[i] of i ranging from 0 to num_slices_in_pic_minus1 (including the end values) specifies the number of CTUs in the i-th stripe, the list SliceTopLeftTileIdx[i] of i ranging from 0 to num_slices_in_pic_minus1 (including the end values) specifies the index of the left-top tile of the stripe, and the matrix of i ranging from 0 to num_slices_in_pic_minus1 (including the end values) and j ranging from 0 to NumCtusInSlice[i] – 1 (including the end values) specifies the picture raster scan address of the j-th CTB within the i-th stripe, and is derived as follows:
[0152]
[0153]
[0154] where the function AddCtbsToSlice(sliceIdx,startX,stopX,startY,stopY) is specified as follows:
[0155]
[0156] The requirement for bitstream consistency is that the value of NumCtusInSlice[i] for i ranging from 0 to num_slices_in_pic_minus1 (including the end values) should be greater than 0. Additionally, the requirement for bitstream consistency is that the matrix CtbAddrInSlice[i][j] for i ranging from 0 to num_slices_in_pic_minus1 (including the end values) and j ranging from 0 to NumCtusInSlice[i]-1 (including the end values) should include all CTB addresses in the range from 0 to PicSizeInCtbsY-1 exactly once and only once.
[0157] The list CtbToSubpicIdx[ctbAddrRs] for ctbAddrRs ranging from 0 to PicSizeInCtbsY-1 (including the end values) specifies the conversion from CTB addresses in the picture raster scan to sub-picture indices, which is derived as follows:
[0158]
[0159] The list NumSlicesInSubpic[i] specifies the number of rectangular stripes in the i-th sub-picture, which is derived as follows:
[0160]
[0161] 7.3.4.3 Picture Parameter Set RBSP Semantics
[0162] …
[0163] Equal to 1 specifies signaling the sub-picture ID mapping in the pps. subpic_id_mapping_in_pps_flag equal to 0 specifies not signaling the sub-picture ID mapping in the PPS. If subpic_id_mapping_explicitly_signaled_flag is 0 or subpic_id_mapping_in_sps_flag is equal to 1, the value of subpic_id_mapping_in_pps_flag should be equal to 0. Otherwise (sub pic_id_mapping_explicitly_signaled_flag is equal to 1, and subpic_id_mapping_in_sps_flag is equal to 0), the value of subpic_id_mapping_in_pps_flag should be equal to 1.
[0164] Should be equal to sps_num_subpics_minus1.
[0165] shall be equal to sps_subpic_id_len_minus1.
[0166] pps_subpic_id[i] specifies the sub-picture ID of the i-th sub-picture. The length of the pps_subpic_id[i] syntax element is pps_subpic_id_len_minus1 + 1 bits.
[0167] For each i value in the range from 0 to sps_num_subpics_minus1, the derivation of the variable SubpicIdVal[i] is as follows:
[0168]
[0169] The requirements for bitstream conformance are to apply the following two constraints:
[0170] -- For any two different values of i and j in the range from 0 to sps_num_subpics_minus1 (including the end values), SubpicIdVal[i] shall not be equal to SubpicIdVal[j].
[0171] -- When the current picture is not the first picture of CLVS, for each i value in the range from 0 to sps_num_subpics_minus1 (including the end values), if the value of SubpicIdVal[i] is not equal to the value of SubpicIdVal[i] of the previous picture in decoding order in the same layer, the nal_unit_type of all coded and decoded slice NAL units of the sub-picture in the current picture with sub-picture index i shall be equal to a specific value in the range from IDR_W_RADL to CRA_NUT (including the end values).
[0172] Equal to 1 specifies that picture partitioning is not applied to each picture of the reference PPS. no_pic_partition_flag equal to 0 specifies that each picture of the reference PPS can be partitioned into more than one slice or strip.
[0173] The requirements for bitstream conformance are that for all PPSs referred to by coded pictures within CLVS, the value of no_pic_partition_flag shall be the same.
[0174] The requirements for bitstream conformance are that when the value of sps_num_subpics_minus1 + 1 is greater than 1, the value of no_pic_partition_flag shall not be equal to 1.
[0175] Adding 5 specifies the luma coding / decoding tree block size for each CTU. pps_log2_ctu_size_minus5 shall be equal to sps_log2_ctu_size_minus5.
[0176] Adding 1 specifies the number of explicitly provided tile column widths. The value of num_exp_tile_columns_minus1 shall be in the range of 0 to PicWidthInCtbsY–1 (inclusive of the end values). When no_pic_partition_flag is equal to 1, the value of num_exp_tile_columns_minus1 is inferred to be equal to 0.
[0177] Adding 1 specifies the number of explicitly provided tile row heights. The value of num_exp_tile_rows_minus1 shall be in the range of 0 to PicHeightInCtbsY–1 (inclusive of the end values). When no_pic_partition_flag is equal to 1, the value of num_tile_rows_minus1 is inferred to be equal to 0.
[0178] [i] Adding 1 specifies the width of the i-th tile column in units of CTB, where i ranges from 0 to num_exp_tile_columns_minus1 - 1 (inclusive of the end values).
[0179] tile_column_width_minus1[num_exp_tile_columns_minus1] is used to derive the width of tile columns with indices greater than or equal to num_exp_tile_columns_minus1, as specified in Clause 6.5.1. The value of tile_column_width_minus1[i] shall be in the range of 0 to PicWidthInCtbsY–1 (inclusive of the end values). When not present, the value of tile_column_width_minus1[0] is inferred to be equal to PicWidthInCtbsY - 1.
[0180] [i] Adding 1 specifies the height of the i-th tile row in units of CTB, where i ranges from 0 to num_exp_tile_rows_minus1 – 1 (inclusive of the end values).
[0181] tile_row_height_minus1[num_exp_tile_rows_minus1] is used to derive the height of tile rows with indices greater than or equal to num_exp_tile_rows_minus1, as specified in Clause 6.5.1. The value of tile_row_height_minus1[i] shall be in the range of 0 to PicHeightInCtbsY–1 (inclusive). When not present, the value of tile_row_height_minus1[0] is inferred to be equal to PicHeightInCtbsY-1.
[0182] Equal to 0 specifies that the slices within each slice segment are in raster scan order and the slice segment information is not signaled in the PPS. rect_slice_flag equal to 1 specifies that the slices within each slice segment cover a rectangular region of the picture and the slice segment information is signaled in the PPS. When not present, rect_slice_flag is inferred to be equal to 1. When subpic_info_present_flag is equal to 1, the value of rect_slice_flag shall be equal to 1.
[0183] Equal to 1 specifies that each sub-picture consists of one and only one rectangular slice segment. single_slice_per_subpic_flag equal to 0 specifies that each sub-picture may consist of one or more rectangular slice segments. When single_slice_per_subpic_flag is equal to 1, num_slices_in_pic_minus1 is inferred to be equal to sps_num_subpics_minus1. When not present, the value of single_slice_per_subpic_flag is inferred to be equal to 0.
[0184] Plus 1 specifies the number of rectangular slice segments in each picture with reference to the PPS. The value of num_slices_in_pic_minus1 shall be in the range of 0 to MaxSlicesPerPicture–1 (inclusive), where MaxSlicesPerPicture is specified in Annex A. When no_pic_partition_flag is equal to 1, the value of num_slices_in_pic_minus1 is inferred to be equal to 0.
[0185] A tile_idx_delta value equal to 0 does not exist in the PPS, and all rectangular stripes in the picture of the PPS are specified in raster order according to the process defined in Clause 6.5.1.
[0186] A tile_idx_delta_present_flag equal to 1 specifies that a tile_idx_delta value may exist in the PPS, and all rectangular stripes in the picture of the PPS are specified in the order indicated by the value of tile_idx_delta. When it does not exist, the value of tile_idx_delta_present_flag is inferred to be equal to 0.
[0187] [i] plus 1 specifies the width of the i-th rectangular stripe in units of tiles. The value of slice_width_in_tiles_minus1[i] shall be in the range of 0 to NumTileColumns–1 (including the end values).
[0188] When slice_width_in_tiles_minus1[i] does not exist, the following applies:
[0189] -- If NumTileColumns is equal to 1, the value of slice_width_in_tiles_minus1[i] is inferred to be equal to 0.
[0190] -- Otherwise, the value of slice_width_in_tiles_minus1[i] is inferred as specified in Clause 6.5.1.
[0191] [i] plus 1 specifies the height of the i-th rectangular stripe in units of tiles. The value of slice_height_in_tiles_minus1[i] shall be in the range of 0 to NumTileRows–1 (including the end values).
[0192] When slice_height_in_tiles_minus1[i] does not exist, the following applies:
[0193] -- If NumTileRows is equal to 1, or tile_idx_delta_present_flag is equal to 0, and tileIdx % NumTileColumns is greater than 0, the value of slice_height_in_tiles_minus1[i] is inferred to be equal to 0.
[0194] -- Otherwise (NumTileRows is not equal to 1, and tile_idx_delta_present_flag is equal to 1 or tileIdx % NumTileColumns is equal to 0), when tile_idx_delta_present_flag is equal to 1 or tileIdx % NumTileColumns is equal to 0, the value of slice_height_in_tiles_minus1[i] is inferred to be equal to slice_height_in_tiles_minus1[i - 1].
[0195] [i] specifies the number of explicitly provided slice heights in the current slice that contains multiple rectangular stripes. The value of num_exp_slices_in_tile[i] shall be in the range of 0 to RowHeight[tileY] – 1 (including the end values), where tileY is the stripe row index that contains the i-th stripe. When it does not exist, the value of num_exp_slices_in_tile[i] is inferred to be equal to 0. When num_exp_slices_in_tile[i] is equal to 0, the value of the variable NumSlicesInTile[i] is derived to be equal to 1.
[0196] [j] plus 1 specifies the height of the j-th rectangular stripe in the current slice in terms of CTU rows. The value of exp_slice_height_in_ctus_minus1[j] shall be in the range of 0 to RowHeight[tileY] – 1 (including the end values), where tileY is the slice row index of the current slice.
[0197] When num_exp_slices_in_tile[i] is greater than 0, the variables NumSlicesInTile[i] and SliceHeightInCtusMinus1[i + k] for k in the range of 0 to NumSlicesInTile[i] – 1 are derived as follows:
[0198]
[0199] Specify the difference between the slice index of the first slice in the i-th rectangular strip and the slice index of the first slice in the (i + 1)-th rectangular strip. The value of tile_idx_delta[i] shall be in the range of -NumTilesInPic + 1 to NumTilesInPic – 1 (inclusive). When not present, the value of tile_idx_delta[i] is inferred to be equal to 0. When present, the value of tile_idx_delta[i] shall not be equal to 0.
[0200] …
[0201] 7.4.2.4.5 Order of VCL NAL units and their association with coded pictures
[0202] The order of VCL NAL units within a coded picture is constrained as follows:
[0203] -- For any two coded strip NAL units A and B of a coded picture, let subpicIdxA and subpicIdxB be their sub-picture level index values, and sliceAddrA and sliceAddrB be their slice_address values.
[0204] -- The coded strip NAL unit A shall be before the coded strip NAL unit B when any of the following conditions is true:
[0205] – subpicIdxA is less than subpicIdxB.
[0206] -- subpicidxa is equal to subpicIdxB, and sliceAddrA is less than sliceAddrB.
[0207] 7.4.8.1 General strip header semantics
[0208] The variable CuQpDeltaVal that specifies the difference between the luma quantization parameter of the coding unit containing cu_qp_delta_abs and its prediction is set to be equal to 0. The variables CuQpOffset Cb 、CuQpOffset Cr and CuQpOffset CbCr that specify the values to be used when determining the corresponding values of the quantization parameters of Qp′ Cb 、CuQpOffset Cr and CuQpOffset CbCr are all set to be equal to 0.
[0209] A value of 1 indicates that the PH syntax structure is present in the slice header. A value of 0 for picture_header_in_slice_header_flag indicates that the PH syntax structure is not present in the slice header.
[0210] The requirement for bitstream consistency is that the value of picture_header_in_slice_header_flag should be the same for all coded slices in the CLVS.
[0211] When the picture_header_in_slice_header_flag of a coded slice is equal to 1, the requirement for bitstream consistency is that no VCL NAL unit with nal_unit_type equal to PH_NUT should appear in the CLVS.
[0212] When picture_header_in_slice_header_flag is equal to 0, all coded slices in the current picture should have a picture_header_in_slice_header_flag equal to 0, and the current PU should have a PH NAL unit.
[0213] Specifies the sub-picture ID of the sub-picture containing the slice. If slice_subpic_id is present, the value of the variable CurrSubpicIdx is derived such that SubpicIdVal[CurrSubpicIdx] is equal to slice_subpic_id. Otherwise (slice_subpic_id is not present), CurrSubpicIdx is derived to be equal to 0. The length of slice_subpic_id is sps_subpic_id_len_minus1 + 1 bits.
[0214] Specifies the slice address of the slice. When not present, the value of slice_address is inferred to be equal to 0. When rect_slice_flag is equal to 1 and NumSlicesInSubpic[CurrSubpicIdx] is equal to 1, the value of slice_address is inferred to be equal to 0.
[0215] If rect_slice_flag is equal to 0, the following applies:
[0216] - The slice address is the raster scan slice index.
[0217] - The length of slice_address is Ceil(Log2(NumTilesInPic)) bits.
[0218] The value of slice_address shall be in the range of 0 to NumTilesInPic–1 (inclusive).
[0219] Otherwise (when rect_slice_flag equals 1), the following applies:
[0220] - The slice address is the sub-picture level slice index of the slice.
[0221] - The length of slice_address is Ceil(Log2(NumSlicesInSubpic[CurrSubpicIdx])) bits.
[0222] - The value of slice_address shall be in the range of 0 to NumSlicesInSubpic[CurrSubpicIdx]–1 (inclusive).
[0223] The requirements for bitstream conformance are to apply the following constraints:
[0224] - If rect_slice_flag equals 0 or subpic_info_present_flag equals 0, the value of slice_address shall not be equal to the value of slice_address of any other decoded slice NAL unit of the same decoded picture.
[0225] - Otherwise, a pair of slice_subpic_id and slice_address values shall not be equal to a pair of slice_subpic_id and slice_address values of any other decoded slice NAL unit of the same decoded picture.
[0226] - The shape of the slices of a picture shall be such that when each CTU is decoded, its entire left boundary and entire top boundary are composed of the picture boundary or the boundary of the previously decoded CTU(s).
[0227] [i] can be equal to 1 or 0. Decoders compliant with this version of the specification shall ignore the value of sh_extra_bit[i]. Its value does not affect the decoder's conformance to the profiles specified in this version of the specification.
[0228] Plus 1 (if present) specifies the number of slices in the slice. The value of num_tiles_in_slice_minus1 shall be in the range of 0 to NumTilesInPic–1 (inclusive).
[0229] The variable NumCtusInCurrSlice specifies the number of CTUs in the current slice, and the list CtbAddrInCurrSlice[i] for i ranging from 0 to NumCtusInCurrSlice–1 (inclusive) specifies the picture raster scan address of the ith CTB within the slice, and is derived as follows:
[0230]
[0231]
[0232] The derivation of the variables SubpicLeftBoundaryPos, SubpicTopBoundaryPos, SubpicRightBoundaryPos, and SubpicBotBoundaryPos is as follows:
[0233]
[0234] 3.5. Luminance Mapping with Chroma Scaling (LMCS)
[0235] LMCS consists of two aspects: luminance mapping (the transformation process, denoted as RP) and luminance-dependent chroma residual scaling (CRS). For the luminance signal, the LMCS mode operates based on two domains, including the first domain as the original domain and the second domain as the shaping domain that maps luminance samples to specific values according to the shaping model. In addition, for the chroma signal, residual scaling can be applied, where the scaling factor is derived from the luminance samples.
[0236] The relevant syntax elements and semantic descriptions in the SPS, Picture Header (PH), and Slice Header (SH) are as follows:
[0237] Syntax table
[0238] 7.3.2.3 Sequence Parameter Set RBSP Syntax
[0239]
[0240] 7.3.2.7 Picture Header Structure Syntax
[0241]
[0242]
[0243] 7.3.7 Slice Header Syntax
[0244] 7.3.7.1 General Slice Header Syntax
[0245]
[0246] Semantics
[0247] Equal to 1 specifies the use of a luminance mapping with chroma scaling in CLVS. sps_lmcs_enabled_flag equal to 0 specifies that a luminance mapping with chroma scaling is not used in CLVS.
[0248] Equal to 1 specifies that the luminance mapping with chroma scaling is enabled for all slices related to PH. ph_lmcs_enabled_flag equal to 0 specifies that the luminance mapping with chroma scaling can be disabled for one, more than one, or all slices related to PH. When absent, the value of ph_lmcs_enabled_flag is inferred to be equal to 0.
[0249] Specifies the adaptation_parameter_set_id of the LMCS APS referred to by the slices associated with PH. The TemporalId of the APS NAL unit with aps_params_type equal to LMCS_APS and adaptation_parameter_set_id equal to ph_lmcs_aps_id should be less than or equal to the TemporalId of the picture related to PH.
[0250] Equal to 1 specifies that chroma residual scaling is enabled for all slices associated with PH. ph_chroma_residual_scale_flag equal to 0 specifies that chroma residual scaling can be disabled for one, more than one, or all slices associated with PH. When ph_chroma_residual_scale_flag is absent, it is inferred to be equal to 0.
[0251] Equal to 1 specifies that the luminance mapping with chroma scaling is enabled for the current slice. slice_lmcs_enabled_flag equal to 0 specifies that the luminance mapping with chroma scaling is not enabled for the current slice. When slice_lmcs_enabled_flag is absent, it is inferred to be equal to 0.
[0252] 3.6. Adaptive Motion Vector Difference Precision (AMVR) of Affine Coding / Decoding Blocks
[0253] Affine AMVR is a coding / decoding tool that allows affine inter-coded blocks to send MV differences with different precisions, such as with a precision of 1 / 4 luminance samples (default, amvr_flag set to 0), 1 / 16 luminance samples, 1 luminance sample.
[0254] The relevant syntax elements and semantic descriptions in SPS are as follows:
[0255] Syntax table
[0256] 7.3.2.3 Sequence Parameter Set RBSP Syntax
[0257]
[0258]
[0259] Semantics
[0260] Equal to 1 specifies the use of adaptive motion vector difference resolution in the motion vector coding and decoding of the affine inter prediction mode. sps_affine_amvr_enabled_flag equal to 0 specifies that the adaptive motion vector difference resolution is not used in the motion vector coding and decoding of the affine inter prediction mode. When absent, the value of sps_affine_amvr_enabled_flag is inferred to be equal to 0.
[0261] D.7 Sub-picture Level Information SEI Message
[0262] D.7.1 Sub-picture Level Information SEI Message Syntax
[0263]
[0264] D.7.2 Sub-picture Level Information SEI Message Semantics
[0265] The sub-picture level information SEI message contains information about the level to which the sub-picture sequence in the bitstream conforms when testing the compliance of the bitstream containing the sub-picture sequence extracted according to Annex A.
[0266] When the sub-picture level information SEI message exists in any picture of the CLVS, the sub-picture level information SEI message will exist in the first picture of the CLVS. The sub-picture level information SEI message continues from the current picture to the current layer in decoding order until the end of the CLVS. All sub-picture level information SEI messages applicable to the same CLVS shall have the same content. The sub-picture sequence consists of all sub-pictures within the CLVS that have the same sub-picture index value.
[0267] The requirement for bitstream consistency is that when the sub-picture level information SEI message of CLVS exists, for each i value in the range from 0 to sps_num_subpics_minus1 (including the end values), the value of subpic_treated_as_pic_flag[i] shall be equal to 1.
[0268] Plus 1 specifies the number of reference levels signaled for each sps_num_subpics_minus1 + 1 sub-pictures.
[0269] Equal to 0 specifies that for decoding the sub-bitstream generated for any sub-picture of the extracted bitstream by using the HRD of any CPB specification in the extracted sub-bitstream, the hypothetical stream scheduler (HSS) operates in the intermittent bitrate mode. sli_cbr_constraint_flag equal to 1 specifies that the HSS operates in the constant bitrate (CBR) mode.
[0270] Equal to 1 specifies the existence of the syntax element ref_level_fraction_minus1[i]. explicit_fraction_present_flag equal to 0 specifies the non-existence of the syntax element ref_level_fraction_minus1[i].
[0271] Plus 1 specifies the number of sub-pictures in the picture of CLVS. When it exists, the value of sli_num_subpics_minus1 shall be equal to the value of sps_num_subpics_minus1 in the SPS referred to by the picture in CLVS.
[0272] Shall be equal to 0.
[0273] [i] indicates the level that each sub-picture conforms to as specified in Annex A. Except for the values specified in Annex A, the bitstream shall not contain the value of ref_level_idc. Other values of ref_level_idc[i] are reserved for future use by ITU-T|ISO / IEC. The requirement for bitstream consistency is that for any k value greater than i, the value of ref_level_idc[i] shall be less than or equal to ref_level_idc[k].
[0274] [i][j] plus 1 specifies the fraction of the level limit associated with ref_level_idc[i] that the j-th sub-picture conforms to, as specified in Clause A.4.1.
[0275] The variable SubpicSizeY[j] is set to be equal to (subpic_width_minus1[j] + 1) * CtbSizeY * (subpic_height_minus1[j] + 1) * CtbSizeY.
[0276] When it does not exist, the value of ref_level_fraction_minus1[i][j] is inferred to be equal to Ceil(256 * SubpicSizeY[j] ÷ PicSizeInSamplesY * MaxLumaPs(general_level_idc) ÷ MaxLumaPs(ref_level_idc[i]) - 1.
[0277] The variable RefLevelFraction[i][j] is set to be equal to ref_level_fraction_minus1[i][j] + 1.
[0278] The derivation of the variables SubpicNumTileCols[j] and SubpicNumTileRows[j] is as follows:
[0279]
[0280]
[0281] The derivation of the variables SubpicCpbSizeVcl[i][j] and SubpicCpbSizeNal[i][j] is as follows:
[0282] SubpicCpbSizeVcl[i][j] =
[0283] Floor(CpbVclFactor * MaxCPB * RefLevelFraction[i][j] ÷ 256) (D.6)
[0284] SubpicCpbSizeNal[i][j] =
[0285] Floor(CpbNalFactor * MaxCPB * RefLevelFraction[i][j] ÷ 256) (D.7)
[0286] Wherein, MaxCPB is derived from ref_level_idc[i] as specified in Clause A.4.2.
[0287] The derivation of the variables SubpicBitRateVcl[i][j] and SubpicBitRateNal[i][j] is as follows:
[0288] SubpicBitRateVcl[i][j] =
[0289] Floor(CpbVclFactor * MaxBR * RefLevelFraction[i][j] ÷ 256) (D.8)
[0290] SubpicBitRateNal[i][j] =
[0291] Floor(CpbNalFactor * MaxBR * RefLevelFraction[i][j] ÷ 256) (D.9)
[0292] where, as specified in Clause A.4.2, MaxBR is derived from ref_level_idc[i].
[0293] Note 1 – When extracting sub - pictures, the resulting bit - stream has a CpbSize (indicated or inferred in the SPS) greater than or equal to SubpicCpbSizeVcl[i][j] and SubpicCpbSizeNal[i][j], and a bit - rate (indicated or inferred in the SPS) greater than or equal to SubpicBitRateVcl[i][j] and SubpicBitRateNal[i][j].
[0294] The bit - stream consistency requirement is that the bit - stream generated by extracting the jth sub - picture for j ranging from 0 to sps_num_subpics_minus1 (inclusive), and conforming to the profile with general_tier_flag equal to 0 and level equal to ref_level_idc[i] for i ranging from 0 to num_ref_level_minus1 (inclusive), shall comply with the following constraints for each bit - stream consistency test specified in Appendix C:
[0295] - Ceil(256 * SubpicSizeY[j] ÷ RefLevelFraction[i][j]) shall be less than or equal to MaxLumaPs, where MaxLumaPs is specified in Table A.1 for level ref_level_idc[i].
[0296] - The value of Ceil(256 * (subpic_width_minus1[j] + 1) * CtbSizeY ÷ RefLevelFraction[i][j]) should be less than or equal to Sqrt(MaxLumaPs * 8).
[0297] - The value of Ceil(256 * (subpic_height_minus1[j] + 1) * CtbSizeY ÷ RefLevelFraction[i][j]) should be less than or equal to Sqrt(MaxLumaPs * 8).
[0298] - The value of SubpicNumTileCols[j] should be less than or equal to MaxTileCols, and the value of SubpicNumTileRows[j] should be less than or equal to MaxTileRows, where MaxTileCols and MaxTileRows are specified in Table A.1 for level ref_level_idc[i].
[0299] - The value of SubpicNumTileCols[j] * SubpicNumTileRows[j] should be less than or equal to MaxTileCols * MaxTileRows * RefLevelFraction[i][j], where MaxTileCols and MaxTileRows are specified in Table A.1 for level ref_level_idc[i].
[0300] - For the value of SubpicSizeInSamplesY of AU 0, the sum of the NumBytesInNalUnit variables of AU 0 corresponding to the j-th sub-picture should be less than or equal to FormatCapabilityFactor * (Max(SubpicSizeY[j], fR * MaxLumaSr * RefLevelFraction[i][j] ÷ 256) + MaxLumaSr * (AuCpbRemovalTime[0] - AuNominalRemovalTime[0]) * RefLevelFraction[i][j]) ÷ (256 * MinCr), where MaxLumaSr and FormatCapabilityFactor are the values specified in Tables A.2 and A.3 respectively, applicable to AU 0 of ref_level_idc[i] level, and MinCr is derived as shown in A.4.2.
[0301] - The sum of the NumBytesInNalUnit variables for AU n (where n is greater than 0) corresponding to the j-th sub-picture shall be less than or equal to FormatCapabilityFactor * MaxLumaSr * (AuCpbRemovalTime[n] - AuCpbRemovalTime[n - 1]) * RefLevelFraction[i][j] ÷ (256 * MinCr), where MaxLumaSr and FormatCapabilityFactor are the values specified in Tables A.2 and A.3 respectively, applicable to AU n at the ref_level_idc[i] level, and MinCr is derived as shown in A.4.2.
[0302] For any sub-picture set that contains one or more sub-pictures and consists of multiple sub-pictures in the SubpicSetIndices list of sub-picture indices and NumSubpicsInSet of sub-picture sets, derive the level information of the sub-picture set.
[0303] The variable SubpicSetAccLevelFraction[i] for the total level fraction relative to the reference level ref_level_idc[i], and the variables SubpicSetCpbSizeVcl[i], SubpicSetCpbSizeNal[i], SubpicSetBitRateVcl[i], and SubpicSetBitRateNal[i] of the sub-picture set are derived as follows: [[ID=,8]]
[0304]
[0305] The value of the sub-picture set sequence level indicator SubpicSetLevelIdc is derived as follows:
[0306]
[0307] Where MaxTileCols and MaxTileRows are specified for ref_level_idc[i] in Table A.1.
[0308] The sub-picture set bitstream that conforms to the configuration document with general_tier_flag equal to 0 and level equal to SubpicSetLevelIdc shall comply with the following constraints for each bitstream consistency test specified in Appendix C:
[0309] - For VCL HRD parameters, SubpicSetCpbSizeVcl[i] shall be less than or equal to CpbVclFactor * MaxCPB, where CpbVclFactor is specified in Table A.3 and MaxCPB is specified in Table A.1 in units of CpbVclFactor bits.
[0310] - For NAL HRD parameters, SubpicSetCpbSizeNal[i] shall be less than or equal to CpbNalFactor * MaxCPB, where CpbNalFactor is specified in Table A.3 and MaxCPB is specified in Table A.1 in units of CpbNalFactor bits.
[0311] - For VCL HRD parameters, SubpicSetBitRateVcl[i] shall be less than or equal to CpbVclFactor * MaxBR, where CpbVclFactor is specified in Table A.3 and MaxBR is specified in Table A.1 in units of CpbVclFactor bits.
[0312] - For NAL HRD parameters, SubpicSetBitRateNal[i] shall be less than or equal to CpbNalFactor * MaxCR, where CpbNalFactor is specified in Table A.3 and MaxBR is specified in Table A.1 in units of CpbNalFactor bits.
[0313] Note 2 – When extracting a sub-picture set, the resulting bitstream has a CpbSize (indicated or inferred in the SPS) greater than or equal to SubpicSetCpbSizeVcl[i][j] and SubpicSetCpbSizeNal[i][j], and a bit rate (indicated or inferred in the SPS) greater than or equal to SubpicSetBitRateVcl[i][j] and SubpicSetBitRateNal[i][j].
[0314] 4. Technical problems solved by the disclosed technical solution
[0315] The existing designs of sub-pictures and LMCS in VVC have the following problems:
[0316] 1) The derivation of the list SubpicNumTileRows[] (specifying the number of tile rows included in a sub-picture) is incorrect in Equation D.5 because the index value idx in CtbToTileRowBd[idx] in the equation may be greater than the maximum allowed value. Additionally, the deviation for both SubpicNumTileRows[] and SubpicNumTileCols[] (specifying the number of tile columns included in a sub-picture) uses CTU-based operations, which are unnecessarily complex.
[0317] 2) When single_slice_per_subpic_flag equals 1, the derivation of the array CtbAddrInSlice in Equation 29 is incorrect because the raster scan CTB address values in the array for each slice need to be in the decoding order of CTUs, rather than the raster scan order of CTUs.
[0318] 3) The LMCS signaling notification is inefficient. When ph_lmcs_enabled_flag equals 1, in most cases, LMCS will be enabled for all slices of a picture. However, in the current VVC design, for the case where LMCS is enabled for all slices of a picture, not only does ph_lmcs_enabled_flag equal 1, but also the slice_lmcs_enabled_flag with a value of 1 needs to be signaled for each slice.
[0319] a. When is true, the semantics conflict with the motivation of signaling the slice-level LMCS flag. In the current VVC, when is true, it means that all slices should have LMCS enabled. Therefore, there is no need to also signal the LMCS enable flag in the slice header.
[0320] b. Additionally, when the picture header indicates that LMCS is enabled, generally, LMCS is enabled for all slices. The control of LCMS in the slice header is mainly for handling extreme cases. Therefore, if the PH LMCS flag is true and the SH LMCS flag is always signaled, this may result in signaling unnecessary bits for common user cases.
[0321] 4) The semantics of the SPS affine AMVR flag is incorrect because for each CU with affine inter-coding, affine AMVR can be enabled or disabled.
[0322] 5. Examples of Techniques and Embodiments
[0323] To solve the above problems and some other problems not mentioned, the following summarized methods are disclosed. These items should be considered as examples to explain general concepts and should not be interpreted in a narrow way. In addition, these items can be applied individually or in any combination.
[0324] Related to the sub - pictures for solving the first and second problems
[0325] 1. One or more of the following methods are disclosed:
[0326] a. Derive the slice column index for each CTU column of the picture.
[0327] b. The derivation of the number of slice columns included in the sub-picture is based on the slice column index of the leftmost and / or rightmost CTU included in the sub-picture.
[0328] c. Derive the slice row index for each CTU row of the picture.
[0329] d. The derivation of the number of slice rows included in the sub-picture is based on the slice row index of the top and / or bottom CTU included in the sub-picture.
[0330] e. The term picture-level stripe index is defined as follows:
[0331] The index from the stripe to the list of stripes in the picture defined when rect_slice_flag is equal to 1, in the order in which the stripes are signaled in the PPS when single_slice_per_subpic_flag is equal to 0, or in the order of increasing sub-picture index of the sub-picture corresponding to the stripe when single_slice_per_subpic_flag is equal to 1.
[0332] f. In one example, when the sub-picture contains a stripe split from a slice, the height of the sub-picture cannot be calculated based on the slice.
[0333] g. In one example, the height of the sub-picture can be calculated based on the CTU instead of the slice.
[0334] h. Derive whether the height of the sub-picture is less than one slice row.
[0335] i. In one example, when the sub-picture includes only CTUs from one slice row and when the top CTU in the sub-picture is not the top CTU of the slice row or the bottom CTU in the sub-picture is not the bottom CTU of the slice row, it is derived as true that the height of the sub-picture is less than one slice row.
[0336] ii. When it is indicated that each sub - picture contains only one stripe and the height of the sub - picture is less than one slice row, for each stripe of the picture with picture - level stripe index i, the value of CtbAddrInSlice[i][j] for j in the range from 0 to the number of CTUs in the stripe minus 1 (inclusive) is derived as the picture raster - scan CTU address of the j - th CTU in the CTU raster - scan of the sub - picture.
[0337] iii. In one example, when the distance between the top CTU and the bottom CTU in the sub - picture is less than the height of the slice according to the CTU, whether the height of the sub - picture is less than one slice row is derived as true.
[0338] iv. When it is indicated that each sub - picture contains only one stripe and the height of the sub - picture is greater than or equal to one slice row, for each stripe of the picture with picture - level stripe index i, the value of CtbAddrInSlice[i][j] for j in the range from 0 to the number of CTUs in the stripe minus 1 (inclusive) is derived as the picture raster - scan CTU address of the j - th CTU, and the order of the CTUs is as follows:
[0339] 1) The CTUs in different slices in the sub - picture are sorted such that the first CTU in the first slice with a smaller slice - index value is before the second CTU in the second slice with a larger slice - index value.
[0340] 2) The CTUs within a slice in the sub - picture are sorted in the CTU raster - scan of the slice.
[0341] Related to the LMCS for solving the third problem (including sub - problems)
[0342] 2. Two - level control of LMCS is introduced (which includes two aspects: luminance mapping (the transformation process, represented by RP) and luminance - dependent chrominance - residual scaling (CRS)), where higher - level (e.g., picture - level) and lower - level (e.g., stripe - level) control are used, and the existence of lower - level control information depends on higher - level control information. In addition, the following also applies:
[0343] a. In the first example, one or more of the following sub - items are applied:
[0344] i. The first indicator (e.g., ) can be signaled at a higher level (e.g., in the picture header (PH)) to specify how to enable LMCS at a lower level for non - binary values.
[0345] 1) In one example, when the first indicator is equal to X (e.g., X = 2), it specifies that LMCS is enabled for all the stripes associated with PH; when the first indicator is equal to Y (Y != X) (e.g., Y = 1), it specifies that LMCS is enabled for one or more but not all of the stripes associated with PH; when the first indicator is equal to Z (Z != X and Z != Y) (e.g., Z = 0), it specifies that LMCS is disabled for all the stripes associated with PH.
[0346] a) Alternatively, in addition, when the first indicator is absent, the value of the indicator is inferred to be equal to a default value, e.g., Z.
[0347] 2) In one example, when the first indicator is equal to X (e.g., X = 2), it specifies that LMCS is disabled for all the stripes associated with PH; when the first indicator is equal to Y (Y != X) (e.g., Y = 1), it specifies that LMCS is disabled for one or more but not all of the stripes associated with PH; when the first indicator is equal to Z (Z != X and Z != Y) (e.g., Z = 0), it specifies that LMCS is enabled for all the stripes associated with PH.
[0348] a) Alternatively, in addition, when the first indicator is absent, the value of the indicator is inferred to be equal to a default value, e.g., X.
[0349] 3) Alternatively, in addition, the first indicator may be conditionally signaled according to the value of the LMCS enable flag in the sequence level (e.g., )
[0350] 4) Alternatively, in addition, the first indicator may be coded / decoded using u(v), or u(2) or ue(v).
[0351] 5) Alternatively, in addition, the first indicator may be coded / decoded using a truncated unary code.
[0352] 6) Alternatively, in addition, the LMCS APS information (e.g., ) used by the stripe and / or CS enable flag (e.g., ) may be signaled under the conditional check of the value of the first indicator.
[0353] ii. A second indicator (e.g., ) for enabling / disabling LMCS at a lower level may be signaled at a lower level (e.g., in the stripe header), and this second indicator may be conditionally signaled by checking the value of the first indicator.
[0354] 1) In one example, the second indicator may be signaled under the conditional check of "the first indicator is equal to Y".
[0355] a) Alternatively, under the condition check of "the value of the first indicator >> 1" or "the value of the first indicator / 2" or "the value of the first indicator & 0x01", the second indicator can be signaled.
[0356] b) Alternatively, in addition, when the first indicator is equal to X, it can be inferred that it is enabled; or when the first indicator is equal to z, it can be inferred that it is disabled.
[0357] b. In the second example, one or more of the following sub-items are applied:
[0358] i. One or more indicators can be signaled at a higher level (e.g., in the picture header (PH)) to specify how to enable LMCS at a lower level for non-binary values.
[0359] 1) In one example, two indicators can be signaled in the PH.
[0360] a) In one example, the first indicator specifies whether there is at least one stripe associated with the PH that enables LMCS. And the second indicator specifies whether all stripes associated with the PH enable LMCS.
[0361] i. Alternatively, in addition, the second indicator can be signaled conditionally according to the value of the first indicator. For example, when the first indicator specifies that there is at least one stripe that enables LMCS.
[0362] i. Alternatively, in addition, when the second indicator does not exist, it is inferred that all stripes enable LMCS.
[0363] ii. Alternatively, in addition, the third indicator can be signaled conditionally in the SH according to the value of the second indicator. For example, when the second indicator specifies that not all stripes enable LMCS.
[0364] i. Alternatively, in addition, when the third indicator does not exist, it can be inferred according to the value of the first and / or second indicator (e.g., inferred to be equal to the value of the first indicator).
[0365] b) Alternatively, the first indicator specifies whether there is at least one stripe associated with the PH that disables LMCS. And the second indicator specifies whether all stripes associated with the PH disable LMCS.
[0366] i. Alternatively, in addition, the second indicator can be signaled conditionally according to the value of the first indicator. For example, when the first indicator specifies that there is at least one stripe that disables LMCS.
[0367] i. Alternatively, in addition, when the second indicator is absent, it is inferred that all stripes related to the PH disable the LMCS.
[0368] ii. Alternatively, in addition, according to the value of the second indicator, the third indicator can be signaled conditionally in the SH, for example, when the second indicator specifies that not all stripes disable the LMCS.
[0369] i. Alternatively, in addition, when the third indicator is absent, it can be inferred according to the value of the first and / or second indicator (for example, inferred to be equal to the value of the first indicator).
[0370] 2) Alternatively, in addition, the first indicator can be signaled conditionally according to the LMCS enable flag in the sequence level (for example, ) value.
[0371] ii. The third indicator for enabling / disabling the LMCS at a lower level (for example, in the slice header) can be signaled (for example, ), and this third indicator can be signaled conditionally by checking the value of the first indicator and / or the second indicator.
[0372] 1) In one example, under the condition check of "not all stripes enable the LMCS" or "not all stripes disable the LMCS", the third indicator can be signaled.
[0373] c. In another example, the first and / or second and / or third indicators mentioned in the first / second examples can be used to control the use of RP or CRS instead of the LMCS.
[0374] 3. The semantic update of the three LMCS flags in SPS / PH / SH is as follows:
[0375] Equal to 1 specifies that in CLVS [[used]] There is a luminance mapping with chroma scaling. sps_lmcs_enabled_flag equal to 0 specifies that there is no luminance mapping with chroma scaling used in CLVS.
[0376] Equal to 1 specifies that the luminance mapping with chroma scaling [[enabled]] For all stripes related to the PH, ph_lmcs_enabled_flag equal to 0 specifies that the luminance mapping with chroma scaling [[can disable one or more, or]] All the stripes related to PH. When it does not exist, the value of ph_lmcs_enabled_flag is inferred to be equal to 0.
[0377] Equal to 1 specifies that the luminance mapping with chroma scaling is [[enabled]] The current stripe. slice_lmcs_enabled_flag equal to 0 specifies that the luminance mapping with chroma scaling is not [[enabled for]] The current stripe. When slice_lmcs_enabled_flag does not exist, it is inferred to be equal to 0.
[0378] a. Change the PH and / or SH LMCS signaling notification so that when LMCS is used for all stripes of a picture, there is no LMCS signaling notification in SH.
[0379] i. Alternatively, in addition, how LMCS is inferred depends on the PH LMCS signaling notification.
[0380] 1) In one example, when LMCS is used for all stripes of a picture, it is inferred to be enabled; and when LMCS is not used for all stripes of a picture, it is inferred to be disabled.
[0381] Related to the affine AMVR
[0382] 4. The semantics of the affine AMVR flag in SPS is updated as follows:
[0383] Equal to 1 specifies that the adaptive motion vector difference precision [[is]] Used in the motion vector coding and decoding for the affine inter prediction mode. sps_affine_amvr_enabled_flag equal to 0 specifies that the adaptive motion vector difference precision is not used in the motion vector coding and decoding for the affine inter prediction mode. When it does not exist, the value of sps_affine_amvr_enabled_flag is inferred to be equal to 0.
[0384] 6. Examples
[0385] 6.1. Example 1: Support for sub - pictures
[0386] This example is for item 1 and its sub - items.
[0387] 3 definitions
[0388]
[0389]
[0390] Index of the list of stripes from the stripe to the picture, in the order in which they are signaled in the PPS when rect_slice_flag is equal to 1.]]
[0391] 6.5.1 CTB raster scan, slice scan, and sub-picture scan processes
[0392] …
[0393] For ctbAddrX in the range from 0 to PicWidthInCtbsY (including the end values) CtbToTileColBd[ctbAddrX] Specifies the conversion from the horizontal CTB address to the left tile column boundary in CTB units as derived as follows:
[0394]
[0395] Note 3 – In the above derivation CtbToTileColBd is one larger than the actual picture width in the CTB.
[0396] For ctbAddrY in the range from 0 to PicHeightInCtbsY (including the end values) CtbToTileRowBd[ctbAddrY] Specifies the conversion from the vertical CTB address to the top tile column boundary in CTB units as derived as follows:
[0397]
[0398] Note 4 – In the above derivation CtbToTileRowBd is one larger than the actual picture height in the CTB.
[0399]
[0400]
[0401]
[0402] When rect_slice_flag is equal to 1, the list NumCtusInSlice[i] for i in the range from 0 to num_slices_in_pic_minus1 (inclusive) specifies the number of CTUs in the i-th slice, the list SliceTopLeftTileIdx[i] for i in the range from 0 to num_slices_in_pic_minus1 (inclusive) specifies the slice index of the slice containing the first CTU in the slice, and the matrix CtbAddrInSlice[i][j] for i in the range from 0 to num_slices_in_pic_minus1 (inclusive) and j in the range from 0 to NumCtusInSlice[i] – 1 (inclusive) specifies the picture raster scan address of the j-th CTB within the i-th slice, and the variable NumSlicesInTile[i] specifies the number of slices in the slice containing the i-th slice, as derived below:
[0403]
[0404]
[0405]
[0406] D.7.2 Sub-picture level information SEI message semantics
[0407] …
[0408] [i][j] plus 1 specifies the level restriction score associated with ref_level_idc[i] corresponding to the j-th sub-picture, as specified in Clause A.4.1.
[0409] The variable SubpicSizeY[j] is set to be equal to (subpic_width_minus1[j] + 1) * CtbSizeY * (subpic_height_minus1[j] + 1) * CtbSizeY.
[0410] When not present, the value of ref_level_fraction_minus1[i][j] is inferred to be equal to Ceil(256 * SubpicSizeY[j] ÷ PicSizeInSamplesY * MaxLumaPs(general_level_idc) ÷ MaxLumaPs(ref_level_idc[i]) - 1.
[0411] The variable RefLevelFraction[i][j] is set to be equal to ref_level_fraction_minus1[i][j]+1.
[0412] [[The derivation of the variables SubpicNumTileCols[j] and SubpicNumTileRows[j] is as follows:
[0413]
[0414] - [j] shall be less than or equal to MaxTileCols, and [j] shall be less than or equal to MaxTileRows, where MaxTileCols and MaxTileRows are specified in Table A.1 for level ref_level_idc[i].
[0415] - [j]* [j] shall be less than or equal to MaxTileCols*MaxTileRows*RefLevelFraction[i][j], where MaxTileCols and MaxTileRows are specified in Table A.1 for level ref_level_idc[i]. ...
[0417] For the variable SubpicSetAccLevelFraction[i] of the total level fraction relative to the reference level ref_level_idc[i], and the derivation of the variables SubpicSetCpbSizeVcl[i], SubpicSetCpbSizeNal[i], SubpicSetBitRateVcl[i] and SubpicSetBitRateNal[i] of the sub-picture set is as follows:
[0418]
[0419] 6.2. Example 2: Support for LMCS
[0420] In this example, the syntax and semantics of the LMCS-related syntax elements in the picture header are modified such that when LMCS is used for all stripes of a picture, there is no LMCS signaling notification in the SH.
[0421] 7.3.2.7 Picture Header Structure Syntax
[0422]
[0423]
[0424]
[0425] 7.3.7.1 General strip header syntax
[0426]
[0427] equal to [[1]]Specify that luminance mapping with chroma scaling is enabled for all strips associated with PH. ph_lmcs_enabled_flag equal to 0 specifies that luminance mapping with chroma scaling can be disabled for one or more, or for all strips associated with PH. When not present, it is inferred that the value of ph_lmcs_enabled_flag is equal to 0.
[0428] equal to 1 specifies that luminance mapping with chroma scaling is enabled for the current strip. slice_lmcs_enabled_flag equal to 0 specifies that luminance mapping with chroma scaling is not enabled for the current strip. When slice_lmcs_enabled_flag is not present, it is inferred to be equal to [[0]]
[0429] In the above example, the values of M and N can be set to 1 and 2 respectively. Alternatively, the values of M and N can be set to 2 and 1 respectively.
[0430] Figure 5 is a block diagram of an example video processing system 1900 that can implement the various techniques disclosed herein. Various implementations may include some or all of the components in system 1900. System 1900 may include an input 1902 for receiving video content. The video content may be received in a raw or uncompressed format (such as 8 or 10-bit multi-component pixel values), or it may be received in a compressed or encoded format. Input 1902 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces (such as Ethernet, passive optical network (PON), etc.) and wireless interfaces (such as Wi-Fi or cellular interfaces).
[0431] System 1900 may include a codec component 1904 that may implement various codec or encoding methods described in this document. The codec component 1904 may reduce the average bit rate of the video from the input 1902 to the output of the codec component 1904 to produce a coded representation of the video. Thus, codec techniques are sometimes referred to as video compression or video transcoding techniques. The output of the codec component 1904 may be stored or transmitted via the connected communication, as represented by component 1906. The stored or communicated bitstream (or coded) representation of the video received at the input 1902 may be used by component 1908 to generate pixel values or a displayable video that is sent to the display interface 1910. The process of generating a user-visible video from the bitstream representation is sometimes referred to as video decompression. Additionally, although certain video processing operations are referred to as "codec" operations or tools, it should be understood that codec tools or operations are used at the encoder and corresponding decoding tools or operations that invert the results of the codec will be performed by the decoder.
[0432] Examples of a peripheral bus interface or a display interface may include a Universal Serial Bus (USB), a High-Definition Multimedia Interface (HDMI), a Displayport, etc. Examples of a storage interface include Serial Advanced Technology Attachment (SATA), PCI, an IDE interface, etc. The techniques described in this document may be implemented in various electronic devices, such as a mobile phone, a laptop computer, a smartphone, or other devices capable of performing digital data processing and / or video display.
[0433] Figure 6 is a block diagram of a video processing apparatus 3600. The apparatus 3600 may be used to implement one or more of the methods described herein. The apparatus 3600 may be implemented in a smartphone, a tablet computer, a computer, an Internet of Things (IoT) receiver, etc. The apparatus 3600 may include one or more processors 3602, one or more memories 3604, and video processing hardware 3606. The (multiple) processors 3602 may be configured to implement one or more of the methods described in this document. The (multiple) memories 3604 may be used to store data and code for implementing the methods and techniques described herein. The video processing hardware 3606 may be used to implement some of the techniques described in this document in hardware circuitry.
[0434] Figure 8 is a block diagram showing an example video codec system 100 that may utilize the techniques of the present disclosure.
[0435] As Figure 8As shown, the video encoding and decoding system 100 may include a source device 110 and a destination device 120. The source device 110 generates encoded video data, which may be referred to as a video encoding device. The destination device 120 may decode the encoded video data generated by the source device 110, and the destination device 120 may be referred to as a video decoding device.
[0436] The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.
[0437] The video source 112 may include sources such as a video capture device, an interface for receiving video data from a video content provider, and / or a computer graphics system that generates video data, or a combination of these sources. The video data may include one or more pictures. The video encoder 114 encodes the video data from the video source 112 to generate a bitstream. The bitstream may include a sequence of bits that form a coded representation of the video data. The bitstream may include coded pictures and associated data. A coded picture is a coded representation of a picture. The associated data may include a sequence parameter set, a picture parameter set, and other syntax elements. The I / O interface 116 includes a modulator / demodulator (modem) and / or a transmitter. The encoded video data may be directly sent to the destination device 120 via the I / O interface 116 over a network 130a. The encoded video data may also be stored on a storage medium / server 130b for access by the destination device 120.
[0438] The destination device 120 may include an I / O interface 126, a video decoder 124, and a display device 122.
[0439] The I / O interface 126 may include a receiver and / or a modem. The I / O interface 126 may obtain the encoded video data from the source device 110 or the storage medium / server 130b. The video decoder 124 may decode the encoded video data. The display device 122 may display the decoded video data to a user. The display device 122 may be integrated with the destination device 120, or may be external to the destination device 120 configured to interface with an external display device.
[0440] The video encoder 114 and the video decoder 124 may operate according to video compression standards such as the High Efficiency Video Coding (HEVC) standard, the Versatile Video Coding (VVC) standard, and other current and / or other standards.
[0441] Figure 9 is a block diagram showing an example of a video encoder 200, which may be the video encoder 114 in the system 100 illustrated in Figure 8 FIG.
[0442] Video encoder 200 may be configured to perform any or all of the techniques of the present disclosure. In Figure 9 In an example of Figure 9 , video encoder 200 includes multiple functional components. The techniques described in the present disclosure may be shared among various components of video encoder 200. In some examples, a processor may be configured to perform any or all of the techniques described in the present disclosure.
[0443] The functional components of video encoder 200 may include a splitting unit 201, a prediction unit 202 (which may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra prediction unit 206), a residual generation unit 207, a transformation unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transformation unit 211, a reconstruction unit 212, a buffer 213, and an entropy coding unit 214.
[0444] In other examples, video encoder 200 may include more, fewer, or different functional components. In one example, prediction unit 202 may include an intra block copy (IBC) unit. The IBC unit may perform prediction in IBC mode, where at least one reference picture is the picture in which the current video block is located.
[0445] In addition, some components such as motion estimation unit 204 and motion compensation unit 205 may be highly integrated, but are shown separately in the Figure 9 example for explanatory purposes.
[0446] Splitting unit 201 may split a picture into one or more video blocks. Video encoder 200 and video decoder 300 may support various video block sizes.
[0447] Mode selection unit 203 may select, for example, one of an intra or inter coding mode based on error results, and provide the resulting intra or inter coded block to residual generation unit 207 to generate residual block data and to reconstruction unit 212 to reconstruct the coded block for use as a reference picture. In some examples, mode selection unit 203 may select a combination of intra and inter prediction (CIIP) mode, where the prediction is based on an inter prediction signal and an intra prediction signal. Mode selection unit 203 may also select the resolution of the motion vector (e.g., sub-pixel or full-pixel accuracy) for a block in the case of inter prediction.
[0448] To perform inter prediction for a current video block, motion estimation unit 204 may generate motion information for the current video block by comparing one or more reference frames from buffer 213 with the current video block. Motion compensation unit 205 may determine a predicted video block for the current video block based on motion information from pictures in buffer 213 (rather than the picture associated with the current video block) and decoded samples.
[0449] The motion estimation unit 204 and the motion compensation unit 205 can perform different operations for the current video block. For example, the operations performed depend on whether the current video block is in an I-slice, a P-slice, or a B-slice.
[0450] In some examples, the motion estimation unit 204 can perform uni-directional prediction of the current video block, and the motion estimation unit 204 can search for a reference video block of the current video block in the reference pictures of list 0 or list 1. The motion estimation unit 204 can then generate a reference index indicating that the reference picture of list 0 or list 1 contains the reference video block and a motion vector indicating the spatial displacement between the current video block and the reference video block. The motion estimation unit 204 can output the reference index, the prediction direction indicator, and the motion vector as the motion information of the current video block. The motion compensation unit 205 can generate a predicted video block of the current block based on the reference video block indicated by the motion information of the current video block.
[0451] In other examples, the motion estimation unit 204 can perform bi-directional prediction of the current video block. The motion estimation unit 204 can search for a reference video block of the current video block in the reference pictures of list 0 and can also search for another reference video block of the current video block in the reference pictures of list 1. The motion estimation unit 204 can then generate a reference index indicating that the reference picture of list 0 or list 1 contains the reference video block and a motion vector indicating the spatial displacement between the reference video block and the current video block. The motion estimation unit 204 can output the reference index and the motion vector of the current video block as the motion information of the current video block. The motion compensation unit 205 can generate a predicted video block of the current video block based on the reference video block indicated by the motion information of the current video block.
[0452] In some examples, the motion estimation unit 204 can output the entire set of motion information for the decoding process of the decoder.
[0453] In some examples, the motion estimation unit 204 may not output the entire set of motion information of the current video. Instead, the motion estimation unit 204 can signal the motion information of the current video block by referring to the motion information of another video block. For example, the motion estimation unit 204 can determine that the motion information of the current video block is similar enough to the motion information of a neighboring video block.
[0454] In one example, the motion estimation unit 204 can indicate in the syntax structure associated with the current video block: indicating to the video decoder 300 that the current video block has a value of the same motion information as another video block.
[0455] In another example, the motion estimation unit 204 may identify another video block and a motion vector difference (MVD) in a syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector indicating the video block. The video decoder 300 may use the motion vector indicating the video block and the motion vector difference to determine the motion vector of the current video block.
[0456] As discussed above, the video encoder 200 may predictively signal motion vectors. Two examples of predictive signaling techniques that may be implemented by the video encoder 200 include advanced motion vector prediction (AMVP) and merge mode signaling.
[0457] The intra prediction unit 206 may perform intra prediction on the current video block. When the intra prediction unit 206 performs intra prediction on the current video block, the intra prediction unit 206 may generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block may include a predicted video block and various syntax elements.
[0458] The residual generation unit 207 may generate residual data for the current video block by subtracting (e.g., denoted by a minus sign) the (multiple) predicted video blocks of the current video block from the current video block. The residual data for the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.
[0459] In other examples, such as in the skip mode, there may be no residual data for the current video block, and the residual generation unit 207 may not perform the subtraction operation.
[0460] The transform processing unit 208 may generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video block associated with the current video block.
[0461] After the transform processing unit 208 generates the transform coefficient video block associated with the current video block, the quantization unit 209 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0462] The inverse quantization unit 210 and the inverse transform unit 211 may apply inverse quantization and inverse transform to the transform coefficient video block, respectively, to reconstruct the residual video block from the transform coefficient video block. The reconstruction unit 212 may add the reconstructed residual video block to the corresponding samples of one or more predicted video blocks generated by the prediction unit 202 to produce a reconstructed video block associated with the current block for storage in the buffer 213.
[0463] After reconstructing a video block in the reconstruction unit 212, a loop filtering operation may be performed to reduce blockiness artifacts in the video block.
[0464] The entropy coding unit 214 may receive data from other functional components of the video encoder 200. When the entropy coding unit 214 receives data, the entropy coding unit 214 may perform one or more entropy coding operations to generate entropy coded data and output a bitstream including the entropy coded data.
[0465] Some embodiments of the techniques of the present disclosure include determining or deciding to enable a video processing tool or mode. In an example, when a video processing tool or mode is enabled, the encoder will use or implement the tool or mode during the processing of a video block, but does not necessarily modify the resulting bitstream based on the use of the tool or mode. In other words, when a video processing tool or mode is enabled based on the determination or decision, the conversion from the video block to the bitstream (or bitstream representation) of the video will use the video processing tool or mode. In another example, when a video processing tool or mode is enabled, the decoder will utilize the knowledge that the bitstream has been modified based on the video processing tool or mode to process the bitstream. In other words, using the video processing tool or mode enabled based on the determination or decision, the conversion from the bitstream of the video to the video block will be performed.
[0466] Figure 10 is a block diagram illustrating an example of a video decoder 300, which may be Figure 8 the video decoder 114 in the system 100 shown in
[0467] The video decoder 300 may be configured to perform any or all of the techniques of the present disclosure. In Figure 10 an example, the video decoder 300 includes a plurality of functional components. The techniques described in the present disclosure may be shared among various components of the video decoder 300. In some examples, a processor may be configured to perform any or all of the techniques described in the present disclosure.
[0468] In Figure 10 an example, the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, and a reconstruction unit 306 and a buffer 307. In some examples, the video decoder 300 may perform a decoding process that is generally inverse to the encoding process described with respect to the video encoder 200 ( Figure 9 ).
[0469] The entropy decoding unit 301 may retrieve an encoded bitstream. The encoded bitstream may include entropy-coded video data (e.g., encoded blocks of video data). The entropy decoding unit 301 may decode the entropy-coded video, and based on the entropy-decoded video data, the motion compensation unit 302 may determine motion information including a motion vector, a motion vector precision, a reference picture list index, and other motion information. The motion compensation unit 302 may determine such information, for example, by performing AMVP and merge mode.
[0470] The motion compensation unit 302 may generate a motion-compensated block, possibly with interpolation based on an interpolation filter. An identifier of the interpolation filter to be used with sub-pixel precision may be included in a syntax element.
[0471] The motion compensation unit 302 may use the interpolation filter used by the video encoder 200 during the encoding of a video block to calculate interpolated values of sub-integer pixels of a reference block. The motion compensation unit 302 may determine the interpolation filter used by the video encoder 200 according to received syntax information and use the interpolation filter to generate a prediction block.
[0472] The motion compensation unit 302 may use some syntax information to determine: the size of the blocks used to encode the (multiple) frames and / or (multiple) slices of an encoded video sequence, the partitioning information describing how each macroblock of a picture of the encoded video sequence is partitioned, the mode indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-coded block, and other information for decoding the encoded video sequence.
[0473] The intra prediction unit 303 may form a prediction block from spatially adjacent blocks using, for example, an intra prediction mode received in the bitstream. The inverse quantization unit 303 inverse quantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 3C1. The inverse transform unit 303 applies an inverse transform.
[0474] The reconstruction unit 306 may sum a corresponding prediction block generated by the motion compensation unit 202 or the intra prediction unit 303 and a residual block to form a decoded block. As desired, a deblocking filter may also be applied to filter the decoded block to remove blocking artifacts. The decoded video block is then stored in a buffer 307, which provides reference blocks for subsequent motion compensation / intra prediction and also generates a decoded video for presentation on a display device.
[0475] Next, a list of preferred solutions of some embodiments is provided.
[0476] The following solutions illustrate example embodiments of the techniques discussed in the previous section (e.g., item 1).
[0477] 1. A video processing method (e.g., the method 900 depicted in Figure 7 ), comprising: performing (902) a conversion between a video including one or more video pictures and a codec representation of the video, wherein each video picture includes one or more slices, and wherein the codec representation conforms to format rules; wherein the format rules specify first information signaled in the codec representation and second information derived from the codec representation, wherein at least the first information or the second information is related to the row index or column index of one or more slices.
[0478] 2. The method according to solution 1, wherein the format rules specify deriving the slice column index of each codec tree unit column of each video picture.
[0479] 3. The method according to solution 1, wherein the format rules specify deriving the slice row index of each codec tree unit row of each video picture.
[0480] The following solutions illustrate example embodiments of the techniques discussed in the previous section (e.g., item 2). In these solutions, the video region may be a video picture, and the video unit may be a video block or a codec tree unit or a video strip.
[0481] 4. A video processing method, comprising: performing a conversion between a video unit of a video region of a video and a codec representation of the video, wherein the codec representation conforms to format rules; wherein the format rules specify that first control information at the video region controls whether second control information is included at the video unit level; wherein the first control information and / or the second control information includes information about luminance mapping and chrominance scaling (LMCS) or chrominance residual scaling (CRS) or retrofit process (RP) for the conversion.
[0482] 5. The method according to solution 4, wherein the first control information includes an indicator indicating whether the second control information is included in the codec representation.
[0483] 6. The method according to solutions 4 - 5, wherein a specific value of the first control information indicates disabling LMCS for all video units in the video region.
[0484] 7. The method according to any one of solutions 4 - 6, wherein the second control information controls the enabling of LMCS at the video unit.
[0485] 8. The method according to solution 4, wherein the first control information includes a plurality of indicators.
[0486] 9. The method according to any one of solutions 1 to 8, wherein the conversion includes encoding the video into a codec representation.
[0487] 10. The method according to any one of Solutions 1 to 8, wherein the conversion includes decoding a coded representation to generate pixel values of a video.
[0488] 11. A video decoding apparatus, comprising a processor configured to implement the method described in one or more of Solutions 1 to 10.
[0489] 12. A video encoding apparatus, comprising a processor configured to implement the method described in one or more of Solutions 1 to 10.
[0490] 13. A computer program product having computer code stored thereon, which when executed by a processor causes the processor to implement the method according to any one of Solutions 1 to 10.
[0491] 14. The method, apparatus or system described in this document.
[0492] In the solutions described herein, an encoder can conform to format rules by generating a coded representation according to the format rules. In the solutions described herein, a decoder can use the format rules to parse syntax elements in the coded representation, where the presence and absence of the syntax elements are known according to the format rules to generate a decoded video.
[0493] Figure 11 A flowchart of an example method 1100 for video processing. Operation 1102 includes: performing a conversion between a video including one or more video pictures and a bitstream of the video, where each video picture includes one or more slices, and the one or more slices include one or more slice columns, where the bitstream conforms to format rules, and where the format rules specify a derivation of a slice column index for each coding tree unit (CTU) column of the slices of the video picture.
[0494] In some embodiments of method 1100, the slice column index of the ctbAddrX-th slice column, denoted as ctbToTileColIdx[ctbAddrX], is derived as follows: tileX = 0
[0495]
[0496] where PicWidthInCtbsY represents the width of the video picture in coding tree blocks (CTBs), and where tileColBd[i] represents the position of the i-th slice column boundary in CTBs.
[0497] In some embodiments of method 1100, each video picture further includes one or more sub - pictures, each sub - picture includes one or more stripes that jointly form a rectangular subset of the video picture, and the formatting rules also stipulate that based on the slice column index of the left - most CTU and / or the right - most CTU included in the sub - picture, the width of the sub - picture in terms of the included slices is derived.
[0498] In some embodiments of method 1100, the width of the i - th sub - picture in terms of slices, denoted as SubpicWidthInTiles[i], is derived as follows:
[0499]
[0500]
[0501] where sps_num_subpics_minus1 represents the number of sub - pictures in the video picture, where sps_subpic_ctu_top_left_x[i] represents the horizontal position of the top - left CTU of the i - th sub - picture, where adding 1 to sps_subpic_width_minus1[i] specifies the width of the i - th sub - picture, and where ctbToTileColIdx[rightX] and ctbToTileColIdx[leftX] respectively represent the slice column indices of the left - most CTU and the right - most CTU included in the sub - picture.
[0502] In some embodiments of method 1100, in response to a slice being divided into multiple rectangular stripes and only a subset of the rectangular stripes of the slice being included in the sub - picture, the slice is counted as one slice in the value of the width of the sub - picture.
[0503] Figure 12 FIG. 1200 is a flowchart of an example method for video processing. Operation 1202 includes: performing a conversion between a video including one or more video pictures and a bit - stream of the video, where each video picture includes one or more slices, the one or more slices include one or more slice rows, the bit - stream conforms to formatting rules, and the formatting rules stipulate deriving a slice row index for each coding tree unit (CTU) row of the slices of the video picture.
[0504] In some embodiments of method 1200, the slice row index of the ctbAddrY - th slice row, denoted as ctbToTileRowIdx[ctbAddrY], is derived as follows:
[0505]
[0506] Where PicHeightInCtbsY represents the height of a video picture in terms of Coding Tree Blocks (CTBs), and where tileRowBd[i] represents the position of the i-th tile row boundary in terms of CTBs.
[0507] In some embodiments of method 1200, each video picture further includes one or more sub-pictures, each sub-picture includes one or more stripes that jointly form a rectangular subset of the video picture, and the formatting rules also specify that the height of the sub-picture in terms of tile rows is derived based on the tile row indices of the top CTU and / or bottom CTU included in the sub-picture.
[0508] In some embodiments of method 1200, the height of the i-th sub-picture in terms of tile rows, denoted as SubpicHeightInTiles[i], is derived as follows:
[0509]
[0510]
[0511] Where sps_num_subpics_minus1 represents the number of sub-pictures in the video picture, where sps_subpic_ctu_top_left_y[i] represents the vertical position of the top-left CTU of the i-th sub-picture, where adding 1 to sps_subpic_height_minus1[i] specifies the height of the i-th sub-picture, and where ctbToTileRowIdx[botY] and ctbToTileRowIdx[topY] respectively represent the tile row indices of the bottom CTU and top CTU included in the sub-picture.
[0512] In some embodiments of method 1200, in response to a tile being split into multiple rectangular stripes and only a subset of the rectangular stripes of the tile being included in the sub-picture, the tile is counted as one tile in the value of the height of the sub-picture.
[0513] Figure 13 A flowchart of an example method 1300 for video processing. Operation 1302 includes: performing a conversion between a video including at least one video picture and a bitstream of the video according to rules, where the at least one video picture includes one or more stripes and one or more sub-pictures, and where the rules specify an order of stripe indices of one or more stripes in the at least one video picture in response to a syntax element associated with the at least one video picture indicating whether each sub-picture of the at least one video picture includes a single stripe.
[0514] In some embodiments of method 1300, the rule further specifies that, in response to each stripe in at least one video picture being a rectangular stripe, a stripe index is indicated. In some embodiments of method 1300, the rule specifies that, in the case where a syntax element indicates that each of one or more sub-pictures includes a single rectangular stripe, the order corresponds to an increasing value of a sub-picture index of one or more sub-pictures in the video picture, and the sub-picture index of the one or more sub-pictures is indicated in a sequence parameter set (SPS) referred to by at least one video picture. In some embodiments of method 1300, the rule specifies that, in the case where a syntax element indicates that each sub-picture includes one or more rectangular stripes, the order corresponds to the order in which the one or more stripes are included in a picture parameter set (PPS) referred to by at least one video picture. In some embodiments of method 1300, the syntax element is included in a picture parameter set (PPS) referred to by at least one video picture.
[0515] Figure 14 FIG. 1400 is a flowchart of an example method for video processing. Operation 1402 includes: performing a conversion between a video unit of a video region of a video and a bitstream of the video, wherein the bitstream conforms to format rules, wherein the format rules specify that first control information of a first level of the video region in the bitstream controls whether second control information is included in a second level of the video unit in the bitstream, wherein the second level is less than the first level, wherein the first control information and the second control information include information about whether or how a luminance mapping and chrominance scaling (LMCS) tool is applied to the video unit, and wherein the LMCS tool includes using chrominance residual scaling (CRS) or a luminance remodelling process (RP) for the conversion.
[0516] In some embodiments of method 1400, the first control information selectively includes a first indicator that indicates whether the LMCS tool is enabled for one or more stripes of the first level of the video region to specify whether the LMCS tool is enabled at the second level of the video unit, and the first indicator is a non-binary value. In some embodiments of method 1400, the first level of the video region includes a picture header. In some embodiments of method 1400, the first level of the video region includes a picture header, the first control information includes a first indicator, when the first indicator is equal to a first value, the LMCS tool is enabled for all stripes of the picture header, when the first indicator is equal to a second value, the LMCS tool is enabled for less than all stripes of the picture header, when the first indicator is equal to a third value, the LMCS tool is disabled for all stripes of the picture header, and the first value, the second value, and the third value are different from each other. In some embodiments of method 1400, when the first control information does not include the first indicator, the value of the first indicator is inferred to be a default value.
[0517] In some embodiments of method 1400, the first level of the video region includes a picture header, and the first control information includes a first indicator. When the first indicator is equal to a first value, the LMCS tool is disabled for all slices of the picture header. When the first indicator is equal to a second value, the LMCS tool is disabled for less than all slices of the picture header. When the first indicator is equal to a third value, the LMCS tool is enabled for all slices of the picture header, and the first value, the second value, and the third value are different from each other. In some embodiments of method 1400, whether the first indicator is selectively included in the first control information is based on the value of a syntax element in the bitstream indicating whether the LMCS tool is enabled at the sequence level. In some embodiments of method 1400, the first indicator is encoded and decoded using u(v) or u(2) or ue(v). In some embodiments of method 1400, the first indicator is encoded and decoded using a truncated unary code.
[0518] In some embodiments of method 1400, based on the value of a first indicator indicating whether the LMCS tool is enabled for one or more slices of the first level of the video region, adaptive parameter set (APS) information and / or chroma scaling syntax elements of the LMCS tool used for the one or more slices are included in the bitstream. In some embodiments of method 1400, the second control information selectively includes a second indicator indicating whether the LMCS tool is enabled or disabled for one or more slices of the second level of the video unit, and the second indicator is included in the bitstream based on the value of the first indicator included in the first control information, and the first indicator indicates whether the LMCS tool is enabled or disabled for one or more slices of the second level of the video unit. In some embodiments of method 1400, the second control information includes a slice header. In some embodiments of method 1400, in response to the first indicator being equal to the first value, the second indicator is included in the second control information. In some embodiments of method 1400, in response to performing the following conditional check: first indicator >> 1, or first indicator / 2, or first indicator & 0x01, where >> describes a right shift operation, and where & describes a bitwise logical AND operation, the second indicator is included in the second control information.
[0519] In some embodiments of method 1400, in response to a first indicator being equal to a first value, a second indicator is inferred to indicate enabling of the LMCS tool for one or more stripes at a second level of a video unit, or in response to the first indicator being equal to a third value, the second indicator is inferred to indicate disabling of the LMCS tool for one or more stripes at the second level of the video unit, and the first value, a second value of the first indicator, and the third value are different from each other. In some embodiments of method 1400, the first control information includes a plurality of indicators that indicate whether the LMCS tool is enabled for one or more stripes at a first level of a video region to specify whether the LMCS tool is enabled at a second level of the video unit, and the plurality of indicators have non-binary values. In some embodiments of method 1400, the plurality of indicators include at least two indicators included in a picture header. In some embodiments of method 1400, the at least two indicators include a first indicator that specifies whether the LMCS tool is enabled for at least one stripe associated with the picture header, and the at least two indicators optionally include a second indicator that specifies whether the LMCS tool is enabled for all stripes associated with the picture header. In some embodiments of method 1400, based on the value of the first indicator, the second indicator is optionally present in the plurality of indicators.
[0520] In some embodiments of method 1400, the value of the first indicator indicates that the LMCS tool is enabled for at least one stripe. In some embodiments of method 1400, in response to the absence of a second indicator in the bitstream, it is inferred that the LMCS tool is enabled for all stripes associated with the picture header. In some embodiments of method 1400, the at least two indicators include a third indicator that is selectively included in the stripe header based on a second value of the second indicator. In some embodiments of method 1400, the second value of the second indicator indicates that the LMCS tool is disabled for all stripes. In some embodiments of method 1400, in response to the absence of the third indicator in the bitstream, the value of the third indicator is inferred based on the first value of the first indicator and / or the second value of the second indicator. In some embodiments of method 1400, the at least two indicators include a first indicator that specifies whether the LMCS tool is disabled for at least one stripe associated with the picture header, and the at least two indicators selectively include a second indicator that specifies whether the LMCS tool is disabled for all stripes associated with the picture header. In some embodiments of method 1400, based on the value of the first indicator, the second indicator is present in the plurality of indicators. In some embodiments of method 1400, the value of the first indicator specifies that the LMCS tool is disabled for at least one stripe. In some embodiments of method 1400, in response to the absence of the second indicator in the bitstream, it is inferred that the LMCS tool is disabled for all stripes associated with the picture header. In some embodiments of method 1400, the at least two indicators selectively include the third indicator in the stripe header based on the second value of the second indicator.
[0521] In some embodiments of method 1400, the second value of the second indicator specifies that the LMCS tool is enabled for all stripes. In some embodiments of method 1400, in response to the absence of the third indicator in the bitstream, the value of the third indicator is inferred based on the first value of the first indicator and / or the second value of the second indicator. In some embodiments of method 1400, the plurality of indicators selectively include the first indicator based on the value of a syntax element that indicates whether the LMCS tool is enabled at the sequence level. In some embodiments of method 1400, the plurality of indicators selectively include a third indicator that indicates whether the LMCS tool is enabled or disabled at a second level of the video unit, and the third indicator selectively exists based on the first value of the first indicator and / or the second value of the second indicator. In some embodiments of method 1400, the third indicator selectively exists based on the second indicator that indicates that the LMCS tool is not enabled for all stripes or not disabled for all stripes. In some embodiments of method 1400, the first indicator, the second indicator, and / or the third indicator controls the use of CRS or the luminance RP.
[0522] Figure 15 A flowchart of an example method 1500 for video processing. Operation 1502 includes: performing a conversion between a video and a bitstream of the video according to rules, where the rules specify that a luminance mapping and chrominance scaling (LMCS) tool is enabled when a first syntax element in a reference sequence parameter set indicates that the LMCS tool is enabled, where the rules specify that the LMCS tool is not used when the first syntax element indicates that the LMCS tool is disabled, where the rules specify that when a second syntax element in the bitstream indicates that the LMCS tool is enabled at the picture header level of the video, the LMCS tool is enabled for all slices associated with the picture header of the video picture, where the rules specify that when the second syntax element indicates that the LMCS tool is disabled at the picture header level of the video, the LMCS tool is not used for all slices associated with the picture header, and where the rules specify that when a third syntax element selectively included in the bitstream indicates that the LMCS tool is enabled at the slice header level of the video, the LMCS tool is used for the current slice associated with the slice header of the video picture, and where the rules specify that when the third syntax element indicates that the LMCS tool is disabled at the slice header level of the video, the LMCS tool is not used for the current slice.
[0523] In some embodiments of method 1500, the rules specify that when the LMCS tool is used for all slices of a video picture, the third syntax element is not included in the slice header of the bitstream. In some embodiments of method 1500, whether the LMCS tool is enabled or disabled is based on the second syntax element. In some embodiments of method 1500, the LMCS tool is enabled when it is used for all slices of a video picture, and the LMCS tool is disabled when it is not used for all slices of a video picture.
[0524] Figure 16 A flowchart of an example method 1600 for video processing. Operation 1602 includes: performing a conversion between a video including one or more video pictures and a bitstream of the video according to rules, where the rules specify that whether adaptive motion vector difference precision (AMVR) is used in the motion vector coding / decoding for an affine inter prediction mode is based on a syntax element selectively included in a reference sequence parameter set (SPS) indicating whether AMVR is enabled, where the rules specify that when the syntax element indicates that AMVR is disabled, AMVR is not used in the motion vector coding / decoding for the affine inter prediction mode, and where the rules specify that when the syntax element is not included in the SPS, AMVR is inferred not to be used in the motion vector coding / decoding for the affine inter prediction mode.
[0525] Figure 17Flowchart of an example method 1700 for video processing. Operation 1702 includes: performing a conversion between a video including video pictures and a bitstream of the video according to a rule, where the video pictures include sub-pictures, slices, and strips, and where the rule provides that since the sub-pictures include strips segmented from slices, the conversion is performed by avoiding using the number of slices of the video pictures to calculate the height of the sub-pictures.
[0526] In some embodiments of method 1700, the height of the sub-pictures is calculated based on the number of coding tree units (CTUs). In some embodiments of method 1700, the height of the sub-pictures is less than one slice row.
[0527] Figure 18 Flowchart of an example method 1800 for video processing. Operation 1802 includes: performing a conversion between a video including video pictures and a bitstream of the video, where the bitstream indicates the height of the sub-pictures of the video pictures calculated based on the number of coding tree units (CTUs) of the video pictures.
[0528] In some embodiments of method 1800, the height of the sub-pictures is not based on the number of slices of the video pictures. In some embodiments of method 1800, the height of the sub-pictures is less than one slice row.
[0529] Figure 19 Flowchart of an example method 1900 for video processing. Operation 1902 includes: making a determination according to a rule, the determination being about whether the height of the sub-pictures of the video pictures of the video is less than the height of a slice row of the video pictures. Operation 1904 includes: using the determination to perform a conversion between the video and the bitstream of the video.
[0530] In some embodiments of method 1900, the rule provides that the height of the sub-pictures is less than one slice row when: the sub-pictures include only coding tree units (CTUs) of one slice row, and a first set of CTUs at the top of the sub-pictures is different from a second set of CTUs at the top of one slice row, or a third set of CTUs at the bottom of the sub-pictures is different from a fourth set of CTUs at the bottom of the one slice row. In some embodiments of method 1900, when each sub-picture of the video pictures includes only one strip and the height of the sub-pictures is less than one slice row, for each strip of the video pictures having a picture-level strip index i, the value of CtbAddrInSlice[i][j] is derived from the picture raster scan CTU address of the j-th CTU in the CTU raster scan of the sub-picture, and j ranges from 0 to the number of CTUs in the strip minus 1 (including the end values).
[0531] In some embodiments of method 1900, the rule stipulates that when the distance between a first set of CTUs at the top of a sub-picture and a second set of CTUs at the bottom of the sub-picture is less than a second height of a slice of the video picture, the height of the sub-picture is less than one slice row, where the second height of the slice is based on the number of CTUs of the sub-picture. In some embodiments of method 1900, when each sub-picture of the video picture includes only one strip and the height of the sub-picture is greater than or equal to one slice row, for each strip of the video picture having a picture-level strip index i, the value of CtbAddrInSlice[i][j] is derived from the picture raster scan CTU address of the j-th CTU in the order of CTUs in the sub-picture, and j ranges from 0 to the number of CTUs in the strip minus 1 (including the end values).
[0532] In some embodiments of method 1900, the order of CTUs in the sub-picture is such that the first CTU in the first slice having a first slice index is placed before the second CTU in the second slice having a second slice index, and the value of the first slice index is less than the value of the second slice index. In some embodiments of method 1900, the order of CTUs in the sub-picture is such that the CTUs within a slice in the sub-picture are sorted in the raster scan of the CTUs in a slice.
[0533] In some embodiments of (multiple) methods 1100 - 1900, performing the conversion includes encoding the video into a bitstream. In some embodiments of (multiple) methods 1100 - 1900, performing the conversion includes encoding the video into a bitstream, and the method further includes storing the bitstream in a non-transitory computer-readable recording medium. In some embodiments of (multiple) methods 1100 - 1900, performing the conversion includes decoding the video from the bitstream.
[0534] In some embodiments, a video decoding device includes a processor configured to implement the operations described for any one or more of methods 1100 to 1900. In some embodiments, a video encoding device includes a processor configured to implement the operations described for any one or more of methods 1100 to 1900. In some embodiments, a computer program product having computer instructions stored thereon, which when executed by a processor cause the processor to implement the operations described for any one or more of methods 1100 to 1900. In some embodiments, a non-transitory computer-readable storage medium stores a bitstream generated according to the operations described for any one or more of methods 1100 to 1900. In some embodiments, a non-transitory computer-readable storage medium stores instructions that cause a processor to implement the operations described for any one or more of methods 1100 to 1900. In some embodiments, a bitstream generation method includes: generating a bitstream of a video according to the operations described for any one or more of methods 1100 to 1900, and storing the bitstream on a computer-readable program medium. In some embodiments, a method, a device, and a bitstream generated according to the disclosed methods or systems described in this document.
[0535] In this document, the term "video processing" may refer to video encoding, video decoding, video compression, or video decompression. For example, during the conversion from a pixel representation of a video to a corresponding bitstream representation, a video compression algorithm may be applied, and vice versa. As defined by the syntax, the bitstream representation of the current video block may, for example, correspond to bits co-located or scattered at different positions within the bitstream. For example, a macroblock may be encoded based on the transformed and coded error residual values and also using bits in the header and other fields in the bitstream. Additionally, during the conversion, the decoder may parse the bitstream based on this determination, knowing that some fields may or may not be present, as described in the above solutions. Similarly, the encoder may determine to include or not include certain syntax fields and accordingly generate a coded representation by including or excluding the syntax fields from the coded representation.
[0536] The disclosures and other scenarios, examples, embodiments, modules, and functional operations described in this document can be implemented in digital electronic circuits or in computer software, firmware, or hardware, including the structures disclosed in this document and their structural equivalents, or a combination of one or more of them. The disclosed and other embodiments can be implemented as one or more computer program products encoded on a computer-readable medium, such as one or more computer program instruction modules, for execution by, or to control the operation of, a data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition affecting a machine-readable propagated signal, or a combination of one or more of them. The term "data processing apparatus" encompasses all apparatuses, devices, and machines for processing data, including, for example, programmable processors, computers, or multiple processors or computers. In addition to hardware, the apparatus may also include code for creating an execution environment for the computer programs being discussed, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them. A propagated signal is an artificially generated signal, for example, a machine-generated electrical, optical, or electromagnetic signal, which is generated to encode information for transmission to a suitable receiver device.
[0537] A computer program (also referred to as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, and can be deployed in any form, including as a stand-alone program or as modules, components, subroutines, or other units suitable for use in a computing environment. A computer program does not have to correspond to a file in a file system. The program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program being discussed, or in multiple cooperating files (e.g., files that store one or more modules, subroutines, or portions of code). A computer program can be deployed to execute on one computer or on multiple computers, which are located at one site or distributed across multiple sites and interconnected by a communication network.
[0538] The processes and logical flows described in this document can be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logical flows can also be performed by special-purpose logic circuitry, such as a field-programmable gate array (FPGA) or an application-specific integrated circuit (ASIC), and the apparatus can be implemented as special-purpose logic circuitry (such as an FPGA or an ASIC).
[0539] Processors suitable for the execution of a computer program include, for example, both general and special purpose microprocessors, as well as any one or more processors of any type of digital computer. Generally, a processor will receive instructions and data from a read-only memory or a random access memory or both. The basic elements of a computer are a processor for performing the instructions and one or more memory devices for storing the instructions and data. Generally, a computer will also include one or more mass storage devices (such as magnetic, magneto-optical or optical disks) for storing data, or operatively coupled to receive data from a mass storage device (such as magnetic, magneto-optical or optical disks) or to transfer data to a mass storage device (such as magnetic, magneto-optical or optical disks), or both. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media and memory devices, including, for example, semiconductor memory devices (such as EPROM, EEPROM and flash memory devices); magnetic disks (such as internal hard disks or removable disks); magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and the memory may be supplemented by, or incorporated in, special purpose logic circuitry.
[0540] Although this patent document contains many details, these details should not be construed as limitations on any subject or the scope of what can be claimed, but rather as descriptions of features specific to particular embodiments of a particular technology. In this patent document, certain features described in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, the various features described in the context of a single embodiment may also be implemented separately in multiple embodiments or in various suitable sub-combinations. Additionally, although features may be described as acting in certain combinations and even initially claimed as such, in some cases one or more features from a claimed combination may be excluded from the combination, and the claimed combination may be directed to a sub-combination or variation of a sub-combination.
[0541] Similarly, although operations are depicted in the drawings in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all of the illustrated operations be performed, to achieve a desired result. Additionally, the separation of various system components in the embodiments described in this patent document should not be understood as required in all embodiments.
[0542] Only a few implementations and examples have been described, and other implementations, enhancements and variations may be made based on what is described and illustrated in this patent document.
Claims
1. A video processing method, comprising: Performing a conversion between a video including video pictures and a bitstream of the video, wherein the video pictures include one or more slices, and the one or more slices form one or more slice rows and one or more slice columns, wherein the bitstream complies with format rules, and wherein the format rules specify deriving a slice row index for each coding tree block (CTB) row of the slices of the video pictures; wherein the video pictures further include one or more sub-pictures, wherein each of the one or more sub-pictures includes one or more stripes that together form a rectangular subset of the video picture, and wherein the format rules further specify deriving the height of a sub-picture in terms of slice rows based on the slice row indexes of the top CTB and / or bottom CTB included in the sub-picture; wherein the height of the i-th sub-picture in terms of slice rows, denoted as SubpicHeightInTiles[i], is derived as follows: for(i = 0; i <= sps_num_subpics_minus1; i++){ topY = sps_subpic_ctu_top_left_y[i] bottomY = topY + sps_subpic_height_minus1[i] SubpicHeightInTiles[i] = ctbToTileRowIdx[botY] + 1 - ctbToTileRowIdx[topY] }, where sps_num_subpics_minus1 plus 1 represents the number of sub-pictures in the video picture, where sps_subpic_ctu_top_left_y[i] represents the vertical position of the top-left CTB of the i-th sub-picture, where sps_subpic_height_minus1[i] plus 1 specifies the height of the i-th sub-picture, and where ctbToTileRowIdx[botY] and ctbToTileRowIdx[topY] respectively represent the slice row indexes of the bottom CTB and top CTB included in the sub-picture.
2. The method according to claim 1, wherein, The slice row index of the ctbAddrY-th slice row, denoted as ctbToTileRowIdx[ctbAddrY], is derived as follows: tileY = 0 for(ctbAddrY = 0; ctbAddrY <= PicHeightInCtbsY; ctbAddrY++){ if(ctbAddrY == tileRowBd[tileY + 1]) tileY++ ctbToTileRowIdx[ctbAddrY] = tileY }, where PicHeightInCtbsY represents the height of the video picture in terms of CTBs, and where tileRowBd[i] represents the position of the i-th slice row boundary in terms of CTBs.
3. The method according to claim 1, wherein In response to the slice being divided into a plurality of rectangular strips and only a subset of the rectangular strips of the slice being included in the sub-picture, the slice is counted as one slice in the derivation of the value of the height of the sub-picture.
4. The method according to claim 1, wherein, The format rule also specifies deriving a slice column index for each CTB column of the slices of the video picture; wherein, the slice column index of the ctbAddrX-th slice column, denoted as ctbToTileColIdx[ctbAddrX], is derived as follows: tileX = 0 for (ctbAddrX = 0; ctbAddrX <= PicWidthInCtbsY; ctbAddrX++) { if (ctbAddrX == tileColBd[tileX + 1]) tileX++ ctbToTileColIdx[ctbAddrX] = tileX }, wherein, PicWidthInCtbsY represents the width of the video picture in units of CTB, and wherein, tileColBd[i] represents the position of the i-th slice column boundary in units of CTB.
5. The method according to claim 4, Among them, the format rule also specifies deriving the width of the sub-picture in units of slice columns based on the slice column indices of the leftmost CTB and / or the rightmost CTB included in the sub-picture; wherein, the width of the i-th sub-picture in units of slice columns, denoted as SubpicWidthInTiles[i], is derived as follows: for (i = 0; i <= sps_num_subpics_minus1; i++) { leftX = sps_subpic_ctu_top_left_x[i] rightX = leftX + sps_subpic_width_minus1[i] SubpicWidthInTiles[i] = ctbToTileColIdx[rightX] + 1 - ctbToTileColIdx[leftX] }, wherein, sps_subpic_ctu_top_left_x[i] represents the horizontal position of the top-left CTB of the i-th sub-picture, wherein, adding 1 to sps_subpic_width_minus1[i] specifies the width of the i-th sub-picture, and wherein, ctbToTileColIdx[rightX] and ctbToTileColIdx[leftX] respectively represent the slice column indices of the leftmost CTB and the rightmost CTB included in the sub-picture.
6. The method according to claim 1, wherein, The video picture further includes one or more strips, wherein, the format rule also specifies that in response to a syntax element associated with the video picture indicating the order of the strip indices of the one or more strips in the video picture, the syntax element indicates whether each sub-picture in the video picture includes a single strip; wherein, in response to each of the strips in the video picture being a rectangular strip, the strip index is indicated.
7. The method according to claim 6, Among them, wherein the formatting rule further specifies that, in the case where the syntax element indicates that each of the one or more sub - pictures includes a single rectangular strip, the order corresponds to an increasing value of the sub - picture index of the one or more sub - pictures in the video picture.
8. The method according to claim 6, wherein The formatting rule further specifies that, in the case where the syntax element indicates that each of the one or more sub - pictures may include one or more rectangular strips, the order corresponds to the order in which the one or more strips are included in a picture parameter set (PPS) referenced by the video picture.
9. The method according to claim 1, a determination is made according to the formatting rule, the determination being about whether the height of a sub - picture of the video picture is less than the height of a slice row of the video picture; Among them, the conversion is performed using the determination.
10. The method according to claim 9, Among them, when each sub - picture of the video picture includes only one strip and the height of the sub - picture is less than a slice row, for each strip of the video picture having a picture - level strip index i, it is specified that the value of CtbAddrInSlice[i][j] of the picture raster - scan CTB address of the jth CTB in the strip is derived from the picture raster - scan CTB address of the jth CTB in the CTB raster - scan of the sub - picture, and where i ranges from 0 to the number of strips in the video picture minus 1, inclusive of the end values, and j ranges from 0 to the number of coding tree units (CTUs) in the strip minus 1, inclusive of the end values.
11. The method according to claim 9, Among them, the formatting rule further specifies that the height of the sub - picture is less than a slice row when: the distance between a first set of CTUs at the top of the sub - picture and a second set of CTUs at the bottom of the sub - picture is less than a second height of a slice of the video picture, where the second height of the slice is based on the number of CTUs in the slice.
12. The method according to claim 9, Among them, when each sub - picture of the video picture includes only one strip and the height of the sub - picture is greater than or equal to a slice row, for each strip of the video picture having a picture - level strip index i, it is specified that the value of CtbAddrInSlice[i][j] of the picture raster - scan CTB address of the jth CTB in the strip is derived from the picture raster - scan CTB address of the jth CTB in the order of CTBs in the sub - picture, where i ranges from 0 to the number of strips in the video picture minus 1, inclusive of the end values, and j ranges from 0 to the number of coding tree units (CTUs) in the strip minus 1, inclusive of the end values; where the order of the CTBs in the sub - picture is such that a first CTB in a first slice having a first slice index is placed before a second CTB in a second slice having a second slice index; where the value of the first slice index is less than the value of the second slice index; and Among them, according to the raster scan order of CTUs in a slice, the order of CTBs within a slice in the sub-picture is sorted.
13. The method according to any one of claims 1 to 12, wherein Performing the conversion includes encoding the video into the bitstream.
14. The method according to any one of claims 1 to 12, wherein Performing the conversion includes decoding the video from the bitstream.
15. A video data processing device includes a processor and a non-transitory memory having instructions thereon, wherein, The instructions, when executed by the processor, cause the processor to: Perform a conversion between a video including video pictures and the bitstream of the video, wherein the video pictures include one or more slices, and the one or more slices form one or more slice rows and one or more slice columns, wherein the bitstream conforms to format rules, and wherein the format rules stipulate the derivation of a slice row index for each coding tree block (CTB) row of the slices of the video picture; wherein the video picture further includes one or more sub-pictures, wherein each of the one or more sub-pictures includes one or more stripes that jointly form a rectangular subset of the video picture, and wherein the format rules also stipulate the derivation of the height of the sub-picture in terms of slice rows based on the slice row indices of the top CTB and / or bottom CTB included in the sub-picture; wherein the height of the i-th sub-picture in terms of slice rows, denoted as SubpicHeightInTiles[i], is derived as follows: for(i = 0; i <= sps_num_subpics_minus1; i++){ topY = sps_subpic_ctu_top_left_y[i] bottomY = topY + sps_subpic_height_minus1[i] SubpicHeightInTiles[i] = ctbToTileRowIdx[botY] + 1 - ctbToTileRowIdx[topY] }, where sps_num_subpics_minus1 plus 1 represents the number of sub-pictures in the video picture, where sps_subpic_ctu_top_left_y[i] represents the vertical position of the top-left CTB of the i-th sub-picture, where sps_subpic_height_minus1[i] plus 1 stipulates the height of the i-th sub-picture, and where ctbToTileRowIdx[botY] and ctbToTileRowIdx[topY] respectively represent the slice row indices of the bottom CTB and top CTB included in the sub-picture.
16. The device according to claim 15, wherein, The slice row index of the ctbAddrY-th slice row, denoted as ctbToTileRowIdx[ctbAddrY], is derived as follows: tileY = 0 for(ctbAddrY = 0; ctbAddrY <= PicHeightInCtbsY; ctbAddrY++){ if(ctbAddrY == tileRowBd[tileY + 1]) tileY++ ctbToTileRowIdx[ctbAddrY] = tileY }, where PicHeightInCtbsY represents the height of the video picture in terms of CTBs, and where tileRowBd[i] represents the position of the i-th tile row boundary in terms of CTBs.
17. A non-transitory computer-readable storage medium storing instructions that cause a processor to: perform a conversion between a video including a video picture and a bitstream of the video, Among them, where the video picture includes one or more tiles that form one or more tile rows and one or more tile columns, where the bitstream complies with format rules, and where the format rules specify deriving a tile row index for each coding tree block (CTB) row of the tiles of the video picture; where the video picture further includes one or more sub-pictures, where each of the one or more sub-pictures includes one or more stripes that together form a rectangular subset of the video picture, and where the format rules further specify deriving the height of a sub-picture in terms of tile rows based on the tile row indices of the top CTB and / or bottom CTB included in the sub-picture; where the height of the i-th sub-picture in terms of tile rows, denoted as SubpicHeightInTiles[i], is derived as follows: for (i = 0; i <= sps_num_subpics_minus1; i++) { topY = sps_subpic_ctu_top_left_y[i] bottomY = topY + sps_subpic_height_minus1[i] SubpicHeightInTiles[i] = ctbToTileRowIdx[botY] + 1 - ctbToTileRowIdx[topY] }, where sps_num_subpics_minus1 plus 1 represents the number of sub-pictures in the video picture, where sps_subpic_ctu_top_left_y[i] represents the vertical position of the top-left CTB of the i-th sub-picture, where sps_subpic_height_minus1[i] plus 1 specifies the height of the i-th sub-picture, and where ctbToTileRowIdx[botY] and ctbToTileRowIdx[topY] respectively represent the tile row indices of the bottom CTB and top CTB included in the sub-picture.
18. A non-transitory computer-readable recording medium storing a bitstream of a video generated by a method executed by a video processing device, wherein, The method includes: generating, according to the format rules, a bitstream of the video including the video picture, where the video picture includes one or more tiles that form one or more tile rows and one or more tile columns, where the bitstream complies with the format rules, and where the format rules specify deriving a tile row index for each coding tree block (CTB) row of the tiles of the video picture; Among them, the video picture further includes one or more sub - pictures, wherein each of the one or more sub - pictures includes one or more strips that jointly form a rectangular subset of the video picture, and wherein the format rule also stipulates deriving the height of the sub - picture in terms of tile rows based on the slice row indices of the top CTB and / or bottom CTB included in the sub - picture; wherein the height of the i - th sub - picture in terms of tile rows, denoted as SubpicHeightInTiles[i], is derived as follows: for(i = 0; i <= sps_num_subpics_minus1; i++){ topY = sps_subpic_ctu_top_left_y[i] bottomY = topY + sps_subpic_height_minus1[i] SubpicHeightInTiles[i] = ctbToTileRowIdx[botY]+1 - ctbToTileRowIdx[topY] }, wherein sps_num_subpics_minus1 plus 1 represents the number of sub - pictures in the video picture, wherein sps_subpic_ctu_top_left_y[i] represents the vertical position of the top - left CTB of the i - th sub - picture, wherein sps_subpic_height_minus1[i] plus 1 stipulates the height of the i - th sub - picture, and wherein ctbToTileRowIdx[botY] and ctbToTileRowIdx[topY] respectively represent the slice row indices of the bottom CTB and top CTB included in the sub - picture.
19. A method for storing a bit - stream of a video, comprising: generating, according to the format rule, a bit - stream of the video including the video picture, and storing the bit - stream in a non - transitory computer - readable recording medium, wherein the video picture includes one or more slices, and the one or more slices form one or more slice rows and one or more slice columns, wherein the bit - stream conforms to the format rule, and wherein the format rule stipulates deriving a slice row index for each coding tree block (CTB) row of the slices of the video picture; wherein the video picture further includes one or more sub - pictures, wherein each of the one or more sub - pictures includes one or more strips that jointly form a rectangular subset of the video picture, and wherein the format rule also stipulates deriving the height of the sub - picture in terms of tile rows based on the slice row indices of the top CTB and / or bottom CTB included in the sub - picture; wherein the height of the i - th sub - picture in terms of tile rows, denoted as SubpicHeightInTiles[i], is derived as follows: for(i = 0; i <= sps_num_subpics_minus1; i++){ topY = sps_subpic_ctu_top_left_y[i] bottomY = topY + sps_subpic_height_minus1[i] SubpicHeightInTiles[i] = ctbToTileRowIdx[botY] + 1 - ctbToTileRowIdx[topY] }, Among them, adding 1 to sps_num_subpics_minus1 represents the number of sub - pictures in the video picture, wherein, sps_subpic_ctu_top_left_y[i] represents the vertical position of the top - left CTB of the i - th sub - picture, wherein, adding 1 to sps_subpic_height_minus1[i] specifies the height of the i - th sub - picture, and wherein, ctbToTileRowIdx[botY] and ctbToTileRowIdx[topY] respectively represent the tile row indices of the bottom CTB and the top CTB included in the sub - picture.
Citation Information
Patent Citations
Sub-streams for wavefront parallel processing in video coding
CN104054348A
An apparatus, a method and a computer program for video coding and decoding
WO2019073112A1