Video Processing Method, Apparatus, and Readable Recording Medium

By using control information and rules to specify the use of scaling tools in video encoders and decoders, the problem of inefficiency in video bandwidth occupancy in the prior art is solved, and more efficient video encoding and decoding is achieved.

CN115462085BActive Publication Date: 2025-06-20DOUYIN CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202180026199.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-04-05
Filing Date
2021-04-05
Publication Date
2025-06-20
Estimated Expiration
2041-04-05

AI Technical Summary

Technical Problem

Existing video encoding technologies are inefficient in dealing with the problem of video bandwidth occupation, especially in the Internet and digital communication networks, where the bandwidth occupied by digital video continues to grow.

Method used

By using control information useful for decoding of the codec representation in the video encoder and decoder, the conversion between the video picture and the bitstream of the video picture is performed, rules specify whether and/or how to indicate the use of the scaling tool, including whether the use of the brightness mapping (LMCS) tool with chromatic scaling and the number of explicit scaling list mode types are allowed.

Benefits of technology

It improves the efficiency of video encoding and decoding, reduces the use of video bandwidth, and adapts to the growth of video bandwidth demand.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115462085B_ABST
    Figure CN115462085B_ABST
Patent Text Reader

Abstract

Describes several techniques for video encoding and video decoding. An example method includes: performing a conversion between a video picture of a video including one or more strips and a bitstream of the video picture according to a rule. The rule specifies whether a first syntax element indicating the use of an encoding / decoding tool exists at a first level based on a syntax flag indicating whether a syntax structure at a second level does not exist at the first level, where the second level is higher than the first level, and where the second level is the video picture level or higher than the video picture level.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - reference to related applications

[0002] Under the applicable patent laws and / or rules of the Paris Convention, this application timely claims the priority and benefit of U.S. Provisional Application No. 63 / 005,413, filed on April 5, 2020. For all purposes as provided by law, the entire disclosure of the above application is incorporated by reference as part of the disclosure of this application. Technical field

[0003] This patent document relates to picture and video encoding and decoding. Background art

[0004] In the Internet and other digital communication networks, digital video occupies the largest bandwidth. With the increase in the number of connected user devices capable of receiving and displaying video, the bandwidth demand for digital video use is expected to continue to grow. Summary of the invention

[0005] This document discloses techniques that can be used by video encoders and decoders for processing encoded and decoded representations of video using control information useful for decoding the encoded and decoded representations.

[0006] In one exemplary aspect, a video processing method is disclosed. The method includes: performing a conversion between a video picture of a video including one or more slices and a bitstream of the video picture according to a rule. The rule specifies that whether a first syntax element indicating the use of an encoding and decoding tool exists at a first level is based on a syntax flag indicating whether a syntax structure at a second level does not exist at the first level, where the second level is higher than the first level. The second level is the video picture level or higher than the video picture level.

[0007] In another exemplary aspect, a video processing method is disclosed. The method includes performing a conversion between a video and a bitstream of the video according to a rule. The rule specifies that a syntax element in the sequence parameter set of the video indicates whether a coded layer video sequence (CLVS) of a reference sequence parameter set enables a luminance mapping with chroma scaling (LMCS) tool.

[0008] In another exemplary aspect, a video processing method is disclosed. The method includes performing a conversion between a video and a bitstream of the video according to a rule. The rule specifies that a syntax element in the sequence parameter set of the video indicates whether a reference sequence parameter set enables a sample adaptive offset (SAO) tool for a coded layer video sequence (CLVS).

[0009] In another example aspect, a video processing method is disclosed. The method includes performing a conversion between a video picture of a video including one or more slices and a bitstream of the video picture according to a rule. The rule specifies whether and / or how to indicate the use of a scaling tool is determined based on whether the video picture includes a single slice, where the use of the scaling tool includes whether to allow the use of a luminance mapping with chroma scaling (LMCS) tool for the conversion, and also includes the number of types of scaling modes allowed for the conversion.

[0010] In another example aspect, a video processing method is disclosed. The method includes performing a conversion between a video picture of a video including one or more slices and a bitstream of the video picture according to a rule. The rule specifies whether and / or how to indicate the allowed slice types in a picture header and / or a slice header is determined based on whether the video picture includes a single slice.

[0011] In another example aspect, a video processing method is disclosed. The method includes performing a conversion between a video picture of a video including one or more slices and a bitstream of the video picture according to a rule. The rule specifies the mode type of an encoding / decoding tool is indicated at the video unit level for the conversion.

[0012] In another example aspect, a video processing method is disclosed. The method includes performing a conversion between a video picture of a video including one or more slices and a bitstream of the video picture according to a rule. The rule specifies a non-binary syntax element or a plurality of syntax flags at a first video unit level are used to indicate the use of an encoding / decoding tool at a second video unit level lower than the first video unit level.

[0013] In another example aspect, a video processing method is disclosed. The method includes: performing a conversion between a video including one or more video pictures each including one or more slices and an encoded / decoded representation of the video; where the encoded / decoded representation conforms to a format rule, and the format rule specifies which luminance mapping with a chroma scaling mode or a scaling list mode type can be applied to the conversion of a slice is indicated by a picture header syntax structure in the slice header or a picture header in a picture including a single slice.

[0014] In another example aspect, another video processing method is disclosed. The method includes: performing a conversion between a video including one or more video pictures each including one or more slices and an encoded / decoded representation of the video; where the encoded / decoded representation conforms to a format rule, and the format rule specifies including an indicator that enables a luminance mapping with chroma scaling (LMCS) mode at a first video level depends on a non-binary LMCS-related syntax element at a higher level and whether the picture consists of only one slice.

[0015] In another example aspect, another video processing method is disclosed. The method includes: performing a conversion between a video including one or more video pictures each including one or more stripes and an encoded / decoded representation of the video; wherein the encoded / decoded representation conforms to format rules, and wherein the format rules specify an indicator including an indication of enabling an Explicit Scaling List (ESL) mode at a first video level depending on a higher-level non-binary LMCS-related syntax element and whether a picture consists of only one stripe.

[0016] In another example aspect, another video processing method is disclosed. The method includes: performing a conversion between a video including one or more pictures each including one or more stripes and an encoded / decoded representation of the video, wherein the encoded / decoded representation conforms to format rules, and wherein the format rules specify that whether a picture exactly includes one stripe controls a stripe type or a stripe type flag in the stripe header of the exactly one stripe.

[0017] In another example aspect, another video processing method is disclosed. The method includes: performing a conversion between a video including one or more pictures each including one or more video regions and an encoded / decoded representation of the video, wherein the encoded / decoded representation conforms to format rules, and wherein the format rules specify a two-level signaling including the applicability of a filtering encoding / decoding tool (TX) to the video region.

[0018] In yet another example aspect, a video encoder device is disclosed. The video encoder includes a processor configured to implement the above method.

[0019] In yet another example aspect, a video decoder device is disclosed. The video decoder includes a processor configured to implement the above method.

[0020] In yet another example aspect, a computer-readable medium storing code is disclosed. The code implements one of the methods described herein in the form of processor-executable code.

[0021] These and other features will be described in this document. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 An example of raster scan stripe segmentation of a picture is shown, where the picture is divided into 12 tiles and 3 raster scan stripes.

[0023] Figure 2 An example of rectangular stripe segmentation of a picture is shown, where the picture is divided into 24 tiles (6 tile columns and 4 tile rows) and 9 rectangular stripes.

[0024] Figure 3 An example of a picture segmented into tiles and rectangular stripes is shown, where the picture is divided into 4 tiles (2 tile columns and 2 tile rows) and 4 rectangular stripes.

[0025] Figure 4 Shows a picture divided into 15 slices, 24 strips, and 24 sub - pictures.

[0026] Figure 5 Is a block diagram of an example video processing system.

[0027] Figure 6 Is a block diagram of a video processing device.

[0028] Figure 7 Is a flowchart of an example method for video processing.

[0029] Figure 8 Is a block diagram showing a video coding and decoding system according to some embodiments of the present disclosure.

[0030] Figure 9 Is a block diagram showing an encoder according to some embodiments of the present disclosure.

[0031] Figure 10 Is a block diagram showing a decoder according to some embodiments of the present disclosure.

[0032] Figure 11 Shows an example of the ALF filter shape (chrominance: 5×5 rhombus, luminance: 7×7 rhombus).

[0033] Figure 12 Shows examples of ALF and CC - ALF.

[0034] Figure 13 Is a flowchart representation of a method for video processing according to the present technology.

[0035] Figure 14 Is a flowchart representation of another method for video processing according to the present technology.

[0036] Figure 15 Is a flowchart representation of another method for video processing according to the present technology.

[0037] Figure 16 Is a flowchart representation of another method for video processing according to the present technology.

[0038] Figure 17 Is a flowchart representation of another method for video processing according to the present technology.

[0039] Figure 18 Is a flowchart representation of another method for video processing according to the present technology.

[0040] Figure 19 Is a flowchart representation of another method for video processing according to the present technology. Detailed Implementation Modes

[0041] In this document, chapter headings are used for easy understanding, rather than restricting the applicability of the technologies and embodiments disclosed in each chapter only to that chapter. Additionally, in some descriptions, the use of H.266 terms is merely for easy understanding and not for limiting the scope of the disclosed technologies. Therefore, the technologies described herein are also applicable to other video codec protocols and designs. In this document, with respect to the current draft of the VVC specification, deleted text is indicated by a strikethrough, and added text (including bold italics) is highlighted to show editorial changes in the text.

[0042] 1. Overview

[0043] This document relates to video codec technologies. Specifically, it is about improvements to the Adaptive Loop Filter (ALF), Sample Adaptive Offset (SAO), Luminance Mapping with Chroma Scaling (LMCS), and signaling of scaling lists. These concepts can be applied individually or in various combinations to any video codec standard or non-standard video codec that supports multi-layer video coding, such as the Versatile Video Coding (VVC) being developed.

[0044] 2. Abbreviations

[0045] ALF Adaptive Loop Filter

[0046] APS Adaptive Parameter Set

[0047] AU Access Unit

[0048] AUD Access Unit Delimiter

[0049] AVC Advanced Video Coding

[0050] CLVS Coded Layer Video Sequence

[0051] CPB Coding Picture Buffer

[0052] CRA Clean Random Access

[0053] CTU Coding Tree Unit

[0054] CVS Coded Video Sequence

[0055] DCI Decoding Capability Information

[0056] DPB Decoding Picture Buffer

[0057] EOB End of Bitstream

[0058] EOS End of Sequence

[0059] GDR Gradual Decoding Refresh

[0060] HEVC High Efficiency Video Coding

[0061] HRD Hypothetical Reference Decoder

[0062] IDR Instantaneous Decoding Refresh

[0063] JEM Joint Exploration Model

[0064] LMCS Luminance Mapping with Chroma Scaling

[0065] MCTS Motion Constrained Tile Set

[0066] NAL Network Abstraction Layer

[0067] OLS Output Layer Set

[0068] PH Picture Header

[0069] PPS Picture Parameter Set

[0070] PTL Profile, Tier and Level

[0071] PU Picture Unit

[0072] RADL Random Access Decodable Leading (Picture)

[0073] RAP Random Access Point

[0074] RASL Random Access Skip Leading (Picture)

[0075] RBSP Raw Byte Sequence Payload

[0076] RPL Reference Picture List

[0077] SAO Sample Adaptive Offset

[0078] SEI Supplemental Enhancement Information

[0079] SPS Sequence Parameter Set

[0080] STSA Stepwise Temporal Sub-layer Access

[0081] SVC Scalable Video Coding

[0082] VCL Video Coding Layer

[0083] VPS Video Parameter Set

[0084] VTM VVC Test Model

[0085] VUI Video Usability Information

[0086] Versatile Video Coding (VVC)

[0087] 3.1. Preliminary Discussion

[0088] Video coding standards have evolved mainly through the development of the well-known ITU-T and ISO / IEC standards. ITU-T developed H.261 and H.263, ISO / IEC developed MPEG-1 and MPEG-4 Visual, and the two organizations jointly developed the H.262 / MPEG-2 video, H.264 / MPEG-4 Advanced Video Coding (AVC), and H.265 / HEVC standards. Since H.262, video coding standards have been based on a hybrid video coding structure that utilizes temporal prediction plus transform coding. To explore future video coding technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Team (JVET) in 2015. Since then, JVET has adopted many new methods and incorporated them into a reference software called the Joint Exploration Model (JEM). JVET meetings are held quarterly, and compared with HEVC, the goal of the new coding standard is to reduce the bit rate by 50%. The new video coding standard was officially named Versatile Video Coding (VVC) at the JVET meeting in April 2018, and the first version of the VVC Test Model (VTM) was also released at that time. Due to the continuous efforts in VVC standardization, new coding technologies have been incorporated into the VVC standard at each JVET meeting. The working draft of VVC and the test model VTM are updated after each meeting. The VVC project now aims to be technically completed (FDIS) at the meeting in July 2020.

[0089] 3.1. Picture Segmentation Schemes of HEVC

[0090] HEVC includes four different picture segmentation schemes, namely, regular slices, dependent slices, tiles, and wavefront parallel processing (WPP), which can be applied to maximum transmission unit (MTU) size matching, parallel processing, and reduced end-to-end latency.

[0091] Regular slices are similar to those in H.264 / AVC. Each regular slice is encapsulated in its own NAL unit, and intra-picture prediction (intra-sample prediction, motion information prediction, coding mode prediction) and entropy coding dependencies across slice boundaries are disabled. Therefore, a regular slice can be reconstructed independently of other regular slices within the same picture (although there may still be dependencies due to loop filter operations).

[0092] Regular stripes are the only tool that can be used for parallelization and are also available in H.264 / AVC in almost the same form. Parallelization based on regular stripes does not require much inter-processor or inter-core communication (except for the inter-processor or inter-core data sharing for motion compensation when decoding predictive coded pictures, which is usually much heavier than inter-processor or inter-core data sharing due to intra-picture prediction). However, for the same reason, using regular stripes results in a large amount of encoding and decoding overhead due to the bit cost of stripe headers and the lack of prediction across stripe boundaries. Additionally, due to the intra-picture independence of regular stripes and each regular stripe being encapsulated in its own NAL unit, regular stripes (compared to other tools mentioned below) also serve as a key mechanism for bitstream segmentation to match MTU size requirements. In many cases, the goals of parallelization and MTU size matching pose conflicting requirements for the stripe layout in a picture. Recognition of this situation led to the development of the parallelization tools mentioned below.

[0093] Dependency stripes have short stripe headers and allow for bitstream segmentation at tree block boundaries without breaking any intra-picture prediction. Basically, dependency stripes provide fragmentation of regular stripes into multiple NAL units to provide reduced end-to-end latency by allowing a part of a regular stripe to be sent out before the encoding of the entire regular stripe is completed.

[0094] In WPP, a picture is partitioned into single-row coding tree blocks (CTBs). Entropy decoding and prediction are allowed to use data from CTBs in other partitions. Parallel processing is possible through parallel decoding of CTB rows, where the start of decoding of a CTB row is delayed by two CTBs to ensure that data related to the CTB above and to the right of the subject CTB is available before the subject CTB is decoded. Using this staggered start (which looks like a wavefront when graphically represented), parallelization can use as many processors / cores as there are CTB rows in the picture. Since intra-picture prediction between adjacent tree block rows within a picture is allowed, the inter-processor / inter-core communication required to implement intra-picture prediction can be substantial. Compared to when not applied, WPP partitioning does not produce additional NAL units, so WPP is not a tool for MTU size matching. However, if MTU size matching is required, regular stripes can be used with WPP but with some encoding and decoding overhead.

[0095] Slices define the horizontal and vertical boundaries that divide a picture into slice columns and rows. Slice columns extend from the top to the bottom of the picture. Similarly, slice rows extend from the left to the right of the picture. The number of slices in a picture can be simply derived by multiplying the number of slice columns by the number of slice rows.

[0096] Before decoding the top-left CTB of the next slice in the raster scan order of the picture's slices, the scan order of CTBs is changed to be local within the slice (in the raster scan order of the slice's CTBs). Similar to regular stripes, slices break the intra-picture prediction dependencies as well as the entropy decoding dependencies. However, they do not need to be included in a single NAL unit (the same as WPP in this regard); thus, slices cannot be used for MTU size matching. Each slice can be processed by a single processor / core, and the inter-processor / inter-core communication required for intra-picture prediction between processing units decoding adjacent slices is limited to transmitting the shared stripe header in cases where a stripe spans more than one slice, and loop filtering related to the sharing of reconstructed samples and metadata. When a stripe contains more than one slice or WPP segment, the byte offset of the entry point of each slice or WPP segment in the stripe other than the first one is signaled in the stripe header.

[0097] For simplicity, HEVC specifies restrictions applied to four different picture partitioning schemes. A given coded video sequence cannot simultaneously include both slices and wavefronts for most profiles specified in the HEVC standard. For each stripe and slice, one or both of the following conditions must be satisfied: 1) all coding tree blocks in the stripe belong to the same slice; 2) all coding tree blocks in the slice belong to the same stripe. Finally, a wavefront segment exactly contains one CTB row, and when using WPP, if a stripe starts at a CTB row, it must end at the same CTB row.

[0098] In some embodiments, HEVC specifies three MCTS-related SEI messages, namely, the Temporal MCTS SEI message, the MCTS Extraction Information Set SEI message, and the MCTS Extraction Information Nested SEI message.

[0099] The Temporal MCTS SEI message indicates the presence of MCTS in the bitstream and signals the MCTS. For each MCTS, the motion vectors are restricted to point to full-sampling positions within the MCTS and fractional sampling positions that only require full-sampling positions within the MCTS for interpolation, and motion vector candidates for temporal motion vector prediction derived from blocks outside the MCTS are not allowed. In this way, each MCTS can be decoded independently, without slices that are not included in the MCTS.

[0100] The MCTS extraction information set SEI message provides supplementary information (specified as part of the SEI message semantics) that can be used in MCTS sub-bitstream extraction to generate a bitstream conforming to the MCTS set. This information consists of multiple extraction information sets, each of which defines multiple MCTS sets and contains the RBSP bytes that will replace the VPS, SPS, and PPS to be used during the MCTS sub-bitstream extraction process. When extracting the sub-bitstream according to the MCTS sub-bitstream extraction process, the parameter sets (VPS, SPS, and PPS) need to be rewritten or replaced, and the slice header needs to be updated slightly because one or all of the syntax elements related to the slice address (including first_slice_segment_in_pic_flag and slice_segment_address) usually need to have different values.

[0101] 3.2. Segmentation of VVC Pictures

[0102] In VVC, a picture is divided into one or more slice rows and one or more slice columns. A slice is a sequence of CTUs that cover a rectangular area of the picture. The CTUs in a slice are scanned in raster scan order within that slice.

[0103] A strip consists of an integer number of complete slices or an integer number of consecutive complete CTU rows in the slices of a picture.

[0104] Two strip modes are supported, namely, the raster scan strip mode and the rectangular strip mode. In the raster scan strip mode, a strip contains the complete strip sequence in the strip raster scan of the picture. In the rectangular strip mode, a strip contains multiple complete slices that together form a rectangular area of the picture, or multiple consecutive complete CTU rows of a single slice that together form a rectangular area of the picture. The slices within a rectangular strip are scanned in slice raster scan order within the rectangular area corresponding to that strip.

[0105] A sub-picture contains one or more strips that together cover a rectangular area of the picture.

[0106] Figure 1 An example of the raster scan strip segmentation of a picture is shown, where the picture is divided into 12 slices and 3 raster scan strips.

[0107] Figure 2 An example of the rectangular strip segmentation of a picture is shown, where the picture is divided into 24 strips (6 strip columns and 4 strip rows) and 9 rectangular strips.

[0108] Figure 3 An example of a picture segmented into slices and rectangular strips is shown, where the picture is divided into 4 slices (2 slice columns and 2 slice rows) and 4 rectangular strips.

[0109] Figure 4 An example of sub - picture segmentation of a picture is shown, where the picture is segmented into 18 slices. Each of the 12 on the left - hand side covers a strip of 4×4 CTUs, and each of the 6 slices on the right - hand side covers 2 vertically - stacked strips of 2×2 CTUs, resulting in a total of 24 strips and 24 sub - pictures of different dimensions (each strip is a sub - picture).

[0110] 3.3. Change of Picture Resolution within a Sequence

[0111] In AVC and HEVC, the spatial resolution of a picture cannot be changed unless a new sequence with a new SPS starts with an IRAP picture. VVC allows changing the picture resolution within a sequence at positions where no IRAP pictures are encoded. IRAP pictures are always intra - coded and decoded. This feature is sometimes referred to as Reference Picture Resampling (RPR) because when the reference picture has a different resolution from the current picture being decoded, this feature requires resampling of the reference pictures used for inter - prediction.

[0112] The scaling ratio is restricted to be greater than or equal to 1 / 2 (2 - fold down - sampling from the reference picture to the current picture) and less than or equal to 8 (8 - fold up - sampling). Three resampling filter sets with different frequency cut - offs are specified to handle various scaling ratios between the reference picture and the current picture. The three resampling filter sets are applied to scaling ratios ranging from 1 / 2 to 1 / 1.75, from 1 / 1.75 to 1 / 1.25, and from 1 / 1.25 to 8 respectively. Each set of resampling filters has 16 phases for luminance and 32 phases for chrominance, the same as in the case of motion - compensated interpolation filters. In fact, the normal MC interpolation process is a special case of the resampling process where the scaling ratio ranges from 1 / 1.25 to 8. The horizontal and vertical scaling ratios are derived based on the picture width and height and the left, right, up, and down scaling offsets specified for the reference picture and the current picture.

[0113] Other aspects in which the VVC design supporting this feature is different from HEVC include: i) The picture resolution and the corresponding consistency window are signaled in the PPS instead of in the SPS, while the maximum picture resolution is signaled in the SPS. ii) For a single - layer bitstream, each picture storage (the slot in the DPB for storing a decoded picture) occupies the buffer size required to store a decoded picture with the maximum picture resolution.

[0114] 3.4. Reference Picture Management and Reference Picture List (RPL)

[0115] Reference picture management is a core function required for any video coding and decoding scheme that uses inter prediction. It manages storing reference pictures in the decoded picture buffer (DPB), removing reference pictures from the decoded picture buffer (DPB), and putting the reference pictures in the RPL in the correct order.

[0116] The reference picture management of HEVC, including reference picture marking and removal from the decoded picture buffer (DPB) and reference picture list construction (RPLC), is different from that of AVC. Instead of the reference picture marking mechanism based on a sliding window plus adaptive memory management control operation (MMCO) in AVC, HEVC specifies a reference picture management and marking mechanism based on the so-called reference picture set (RPS), and thus RPLC is based on the RPS mechanism. The RPS consists of reference picture sets associated with pictures (composed of all reference pictures before the associated picture in decoding order), which can be used for inter prediction of the associated picture or any picture after the associated picture in decoding order. The reference picture set consists of five reference picture lists. The first three lists contain all reference pictures that can be used for inter prediction of the current picture and can be used for inter prediction of one or more pictures after the current picture in decoding order. The other two lists consist of all reference pictures that are not used for inter prediction of the current picture but can be used for inter prediction of one or more pictures after the current picture in decoding order. The RPS provides "intra coding" signaling of the DPB state, rather than "inter coding" signaling as in AVC, mainly to improve error resilience. The RPLC process of HEVC is based on the RPS, by signaling an index to an RPS subset for each reference index; this process is simpler than the RPLC process in AVC.

[0117] The reference picture management of VVC is more similar to HEVC than AVC, but is simpler and more robust. As in those standards, two RPLs, list 0 and list 1, are derived, but they are not conceived based on the reference picture sets used in HEVC or the automatic sliding window process used in AVC; instead, they are signaled more directly. The reference pictures for the RPL are listed as active and non-active entries, and only active entries can be used as reference indices for inter prediction of CTUs of the current picture. The invalid entries indicate other pictures to be saved in the DPB for reference by other pictures arriving later in the bitstream.

[0118] 3.5. Parameter Sets

[0119] AVC, HEVC, and VVC specify parameter sets. The types of parameter sets include SPS, PPS, APS, and VPS. All of AVC, HEVC, and VVC support SPS and PPS. VPS was introduced starting from HEVC and is included in both HEVC and VVC. APS is not included in AVC or HEVC but is included in the latest VVC draft text.

[0120] SPS is designed to carry sequence-level header information, and PPS is designed to carry picture-level header information that does not change frequently. With SPS and PPS, information that does not change frequently does not need to be repeated for each sequence or picture, thus avoiding redundant signaling of this information. In addition, the use of SPS and PPS enables out-of-band transmission of important header information, which not only avoids the need for redundant transmission but also improves the error resilience.

[0121] VPS was introduced to carry sequence-level header information common to all layers in a multi-layer bitstream.

[0122] APS was introduced to carry such picture-level or slice-level information that requires a significant number of bits to encode and decode, can be shared by multiple pictures, and can have a significant number of different variations in the sequence.

[0123] 3.6. Slice Headers and Picture Headers of VVC

[0124] Similar to the case of HEVC, the slice header of VVC conveys information about a specific slice. This includes slice address, slice type, slice QP, least significant bit (LSB) of the picture order count (POC), RPS and RPL information, weighted prediction parameters, loop filter parameters, entries of slice and WPP offset, etc.

[0125] VVC introduces a picture header (PH) that contains header parameters for a specific picture. Each picture must have one or only one PH. The PH basically carries those parameters that would be in the slice header if the PH were not introduced, but each parameter has the same value for all slices of the picture. These include IRAP / GDR picture indication, inter-slice / intra-slice allow flag, POC LSB and optional POC MSB, information about RPL, deblocking, SAO, ALF, QP delta, and weighted prediction, coding block partition information, virtual boundary, co-located picture information, etc. It often happens that each picture in an entire picture sequence contains only one slice. In this case, to allow each picture to not have at least two NAL units, the PH syntax structure is allowed to be included in the PH NAL unit or the slice header.

[0126] In VVC, information of the co-located pictures for temporal motion vector prediction is signaled in the picture header or slice header.

[0127] 3.7. Loop Filtering

[0128] In VVC, the deblocking filter, SAO, and ALF are supported as loop filtering methods.

[0129] 3.7.1. SAO

[0130] The same design as in HEVC is used, where sample adaptive offset (SAO) is called after deblocking filtering and before ALF if needed. The key idea of SAO is to reduce sample distortion by first classifying the reconstructed samples into different classes, obtaining the offsets for each class, and then adding the offset to each sample of that class. The offsets for each class are appropriately calculated at the encoder and signaled explicitly to the decoder to effectively reduce sample distortion, while the classification of each sample is performed at both the encoder and the decoder to significantly save side information. To achieve low latency for only one coding tree unit (CTU), a CTU-based syntax design is specified to adapt the SAO parameters to each CTU.

[0131] 3.7.2. Adaptive Loop Filter

[0132] Two diamond filter shapes (as Figure 11 shown) are used in the block-based ALF. The 7×7 diamond is applied to the luma component, and the 5×5 diamond is applied to the chroma component. One of up to 25 filters is selected for each 4×4 block based on the direction and activity of the local gradient. Each 4×4 block in the picture is classified based on directionality and activity. Before filtering each 4×4 block, a simple geometric transformation, such as rotation or diagonal and vertical flipping, can be applied to the filter coefficients depending on the gradient value calculated for that block. This is equivalent to applying these transformations to the samples in the filter support region. The idea is to make different blocks more similar by aligning the directionality of different blocks to which ALF is applied. Block-based classification is not applicable to the chroma component.

[0133] The ALF filter parameters are signaled in the Adaptive Parameter Set (APS). In one APS, up to 25 sets of luminance filter coefficients and clipping value indices, and up to 8 sets of chrominance filter coefficients and clipping value indices can be signaled. To reduce bit overhead, the filter coefficients of different classifications of the luminance component can be combined. In the picture or slice header, up to 7 APS IDs can be signaled to specify the luminance filter set for the current picture or slice. The filtering process is further controlled at the CTB level. The luminance CTB can select a filter set among 16 fixed filter sets and the filter sets signaled in the APS. For the chrominance component, the APS ID is signaled in the picture or slice header to indicate the chrominance filter set for the current picture or slice. At the CTB level, if there are more than one chrominance filter sets in the APS, a filter index is signaled for each chrominance CTB. When ALF is enabled for a CTB, for each sample within the CTB, a diamond filter with the signaled weights is performed, and a clipping operation is applied to capture the difference between adjacent samples and the current sample. The clipping operation introduces non-linearity to make ALF more effective by reducing the influence of adjacent sample values that are too different from the current sample value.

[0134] The Cross-Component Adaptive Loop Filter (CC-ALF) can further enhance each chrominance component on top of the previously described ALF. The goal of CC-ALF is to use the luminance sample values to refine each chrominance component. This is achieved by applying a diamond high-pass linear filter and then using the output of this filtering operation for chrominance refinement. Figure 12 A system-level schematic of the CC-ALF process relative to other loop filters is provided. As Figure 12 shown, CC-ALF uses the same input as the luminance ALF to avoid additional steps in the overall loop filtering process.

[0135] 3.7.3. Signaling of ALF / SAO

[0136] In VVC Draft 8, ALF and SAO share the same advanced control scheme. Both of these codec tools can be controlled either at the sequence level and one of the picture level or slice level (but not both at the picture level and slice level). First, the SPS enable flag is signaled to control ALF / SAO at the CLVS level. At the PPS level, the PPS flag is signaled to indicate whether ALF / SAO is further controlled at the picture level or slice level. If the PPS flag indicates that ALF / SAO is further controlled at the picture level, the PH ALF / SAO enable flag is signaled, followed by the ALF parameters (if it is enabled); if the PPS flag indicates that ALF / SAO is further controlled at the slice level, the SH ALF / SAO enable flag is signaled, followed by the ALF parameters (if it is enabled).

[0137] Table 1: ALF Syntax in SPS

[0138]

[0139] Table 2: ALF Syntax in PPS

[0140]

[0141] Table 3: ALF Syntax in Picture Header

[0142]

[0143]

[0144] Table 4: ALF Syntax in Slice Header

[0145]

[0146]

[0147] Table 5: SAO Syntax in SPS

[0148]

[0149] Table 6: SAO Syntax in PPS

[0150]

[0151] Table 7: SAO Syntax in Picture Header

[0152]

[0153]

[0154] Table 8: SAO Syntax in Slice Header

[0155]

[0156] Table 9: LMCS Syntax in SPS

[0157]

[0158] Table 10: LMCS Syntax in Picture Header

[0159]

[0160] Table 11: LMCS Syntax in Strip Header

[0161]

[0162]

[0163] Equal to 1 specifies that the luminance mapping with chrominance scaling is used in CLVS. sps_lmcs_enabled_flag equal to 0 specifies that the luminance mapping with chrominance scaling is not used in CLVS.

[0164] Equal to 1 specifies the reconstructed picture to which the sample adaptive offset process is applied after the deblocking filtering process. sps_sao_enabled_flag equal to 0 specifies that the sample adaptive offset process is not applied to the reconstructed picture after the deblocking filtering process.

[0165] 3.8. Luminance Mapping with Chrominance Scaling (LMCS)

[0166] Unlike other loop filters (i.e., deblocking filter, SAO filter, and ALF filter) that typically apply a filtering process to the current sample by using information of spatially adjacent samples of the current sample to reduce codec artifacts, the Luminance Mapping with Chrominance Scaling (LMCS) improves compression efficiency by redistributing codewords over the entire dynamic range to modify the input signal before encoding. LMCS has two main components: (a) loop mapping of the luminance component based on an adaptive piecewise linear model, and (b) luminance-dependent chrominance residual scaling for the chrominance component. The luminance mapping utilizes a forward mapping function FwdMap and a corresponding inverse mapping function InvMap. The FwdMap function is signaled using a piecewise linear model with 16 equal pieces. The InvMap function does not need to be signaled but is derived from the FwdMap function. The luminance mapping model is signaled in the APS. Up to 4 LMCS APSs can be used in the encoded / decoded video sequence. When LMCS is enabled for a picture, the APS ID is signaled in the picture header to identify the APS carrying the luminance mapping parameters. When LMCS is enabled for a slice, the InvMap function is applied to all reconstructed luminance blocks to transform the samples back to the original domain. For inter-coded blocks, an additional mapping process is required, which applies the FwdMap function to map the luminance prediction block in the original domain to the mapped domain after the normal compensation process. The chrominance residual scaling is designed to compensate for the interaction between the luminance signal and its corresponding chrominance signal. When luminance mapping is enabled, an additional flag is signaled indicating whether luminance-dependent chrominance residual scaling is enabled. The chrominance residual scaling factor depends on the average of the reconstructed adjacent luminance samples at the top and / or left of the current CU. Once the scaling factor is determined, forward scaling is applied to the intra- and inter-prediction residuals during the encoding stage, and inverse scaling is applied to the reconstructed residuals.

[0167] 3.8.1. Signaling of LMCS-related Syntax Elements

[0168] In the current VVC specification, LMCS control can be signaled in the SPS, PH, and SH. First, the SPS enable flag controls LMCS at the CLVS level. If the SPS enable flag is equal to 1, the PH enable flag is further signaled to control LMCS at the picture level, and if it is enabled at the picture level, the LMCS parameter information is also signaled in the PH. If the PH enable flag is equal to 1, the SH enable flag is further signaled to control LMCS at the slice level, but even if it is enabled at the slice level, the LMCS parameter information cannot be signaled in the SH.

[0169] The related syntax elements and semantics are as follows:

[0170] 7.3.2.7 Picture Header Structure Syntax

[0171]

[0172] 7.3.7.1 General strip header syntax

[0173]

[0174] Equal to 1 specifies that luminance mapping with chroma scaling is enabled for all strips associated with the PH. ph_lmcs_enabled_flag equal to 0 specifies that luminance mapping with chroma scaling is disabled for one, more, or all strips associated with the PH, and if not present, the value of ph_lmcs_enabled_flag is inferred to be equal to 0.

[0175] Equal to 1 specifies that chroma residual scaling is enabled for all strips associated with the PH. ph_chroma_residual_scale_flag equal to 0 specifies that chroma residual scaling can be disabled for one, more, or all strips associated with the PH. When ph_chroma_residual_scale_flag is not present, it is inferred to be equal to 0.

[0176] Equal to 1 specifies that luminance mapping with chroma scaling is enabled for the current strip. slice_lmcs_enabled_flag equal to 0 specifies that luminance mapping with chroma scaling is not enabled for the current strip. When slice_lmcs_enabled_flag is not present, it is inferred to be equal to 0.

[0177] 3.9. Explicit scaling lists

[0178] Explicit signaling of scaling lists is defined in the APS. And for each picture, whether to use explicit signaling is signaled first as a flag in the PH, and if needed, followed by an APS index. In the strip header, when the PH flag indicates the use of an explicit scaling list, each strip further signals a flag indicating whether the current strip uses explicit signaling.

[0179] In the current VVC draft text, the text most relevant to the scaling lists is as follows:

[0180] Sequence parameter set RBSP syntax and semantics

[0181] ....

[0182] Equal to 1 indicates that the scaling list is used for the scaling process of transform coefficients. sps_scaling_list_enabled_flag equal to 0 indicates that the scaling list is not used for the scaling process of transform coefficients. ...

[0183] Picture header structure syntax and semantics

[0184]

[0185] ...

[0186] Equal to 1 indicates that the scaling list data for the slice associated with PH is derived based on the scaling list data contained in the reference scaling list APS. ph_scaling_list_present_flag equal to 0 indicates that the scaling list data for the slice associated with PH is set to be equal to 16. When not present, the value of ph_scaling_list_present_flag is inferred to be equal to 0. When not present, the value of ph_scaling_list_present_flag is inferred to be equal to 0.

[0187] Specify the adaptation_parameter_set_id of the scaling list APS. The Temporalld of the APS NAL unit with aps_params_type equal to SCALING_APS and adaptation_parameter_set_id equal to ph_scaling_list_aps_id should be less than or equal to the Temporalld of the picture associated with PH. ...

[0188] General slice header syntax and semantics

[0189] ...

[0190] Equal to 1 indicates that the scaling list data specified for the current slice is derived from the scaling list data included in the reference scaling list APS with aps_params_type equal to SCALING_APS and adaptation_parameter_set_id equal to ph_scaling_list_aps_id. slice_scaling_list_present_flag equal to 0 indicates that the scaling list data specified for the current picture is the default scaling list data derived as specified in Clause 7.4.3.21. When not present, the value of slice_scaling_list_present_flag is inferred to be equal to 0. ...

[0191] Scaling process of transform coefficients

[0192] In order to derive the scaled transform coefficients d[x][y], where x = 0..nTbW-1, y = 0..nTbH-1, the following cases apply:

[0193] – The intermediate scaling factor m[x][y] is derived as follows:

[0194] – If one or more of the following conditions hold, then m[x][y] is set to be equal to 16:

[0195] – sps_scaling_list_enabled_flag is equal to 0.

[0196] – ph_scaling_list_present_flag is equal to 0.

[0197] – transform_skip_flag[xTbY][yTbY][cIdx] is equal to 1.

[0198] – scaling_matrix_for_lfnst_disabled_flag is equal to 1, and ApplyLfnstFlag is equal to 1.

[0199] –... ...

[0200] 7.3.2.5 Adaptive parameter set RBSP syntax

[0201]

[0202] 7.3.2.21 Scaling list data syntax

[0203]

[0204]

[0205] Equal to 1 specifies that the scaling matrix shall not be applied to the blocks coded with LFNST. scaling_matrix_for_lfnst_disabled_flag equal to 0 specifies that the scaling matrix may be applied to the blocks coded with LFNST.

[0206] 3.10. Latest developments in LMCS, and explicit scaling lists

[0207] To address all of the above issues, it is proposed to replace the PH flag ph_lmcs_enabled_flag with a 2-bit ph_lmcs_mode_idc, and three modes are specified: disabled (mode 0), for all stripes (mode 1), and enabled (mode 2). In mode 1, LMCS is applied to all stripes of the picture, and no signaling of the LMCS control flag in the SH is required. Accordingly, the semantics of the SH LMCS control flag are modified. In addition, a modification to the semantics of ph_chroma_residual_scale_flag is proposed to reflect the intention of enabling / disabling chroma residual scaling for a picture or a stripe.

[0208] The following are some proposed modifications to the syntax structure. Most of the relevant parts added or modified are marked in bold italic underline, and some deleted parts are indicated by [[ ]].

[0209] 7.3.2.7 Picture header structure syntax

[0210]

[0211] 7.3.7.1 General stripe header syntax

[0212]

[0213]

[0214] Equal to 1 specifies that a luminance mapping with chroma scaling shall be applied to all stripes associated with the PH, Equal to 0 specifies that for all stripes associated with the PH [[which may be disabled for one or more stripes, or]] a luminance mapping with chroma scaling. When not present, the value is inferred to be equal to 0. Equal to 1 specifies that chroma residual scaling is enabled for all stripes associated with and ph_chroma_residual_scale_flag being equal to 0 specifies that the chroma residual scaling associated with PH [[can be disabled for one or more stripes, or]] is applied to all stripes When the ph_chroma_residual_scale_flag is not present, it is inferred to be equal to 0. ...

[0215] being equal to 1 specifies that the luminance mapping with chroma scaling is applied to the current stripe, being equal to 0 specifies that the luminance mapping with chroma scaling is not applicable to the current stripe. When is not present, it is inferred to be equal to ...

[0216] To address several issues, the following modifications are proposed:

[0217] (1) The PH flag ph_explicit_scaling_list_enabled_flag is replaced with the 2-bit ph_explicit_scaling_list_mode_idc, specifying 3 modes: disabled (mode 0), for all stripes (mode 1), and enabled (mode 2). In mode 1, the explicit scaling list is used for all stripes of the picture, and no scaling list signaling is required in SH.

[0218] (2) The flag scaling_matrix_for_lfnst_disabled_flag is moved from the scaling_list_data() syntax to the SPS.

[0219] 7.3.2.3 Sequence Parameter Set RBSP Syntax

[0220]

[0221] 7.4.3.3 Sequence Parameter Set RBSP Semantics being equal to 1 specifies that lfnst_idx may be present in the intra coded unit syntax. sps_lfnst_enabled_flag being equal to 0 specifies that lfnst_idx is not present in the intra coded unit syntax.

[0222] Equal to 1 specifies that when decoding a slice, the use of the explicit scaling list signaled in the scaling list APS during the scaling process of transform coefficients is enabled for CLVS. sps_explicit_scaling_list_enabled_flag equal to 0 specifies that when decoding a slice, the use of the explicit scaling list during the scaling process of transform coefficients is disabled for CLVS.

[0223]

[0224] 7.3.2.7 Picture Header Structure Syntax

[0225]

[0226]

[0227] 7.4.3.7 Picture Header Structure Semantics

[0228] Equal to 1 specifies that when decoding a slice, the use of the explicit scaling list signaled in the reference scaling list APS (i.e., the APS with aps_params_type equal to SCALING_APS and adaptation_parameter_set_id equal to ph_explicit_scaling_list_aps_id) during the scaling process of transform coefficients is enabled for the picture. ph_explicit_scaling_list_enabled_flag equal to 0 specifies that when decoding a slice, the use of the explicit scaling list during the scaling process of transform coefficients is enabled for the picture. When absent, the value of ph_explicit_scaling_list_enabled_flag is inferred to be equal to 0]]

[0229]

[0230] Specifies the adaptation_parameter_set_id of the scaling list APS. The Temporalld of the APS NAL unit with aps_params_type equal to SCALING_APS and adaptation_parameter_set_id equal to ph_scaling_list_aps_id shall be less than or equal to the Temporalld of the picture associated with the PH.

[0231] 7.3.2.21 Scaling List Data Syntax

[0232]

[0233]

[0234] 7.4.3.21 Scaling list data semantics

[0235] Equal to 1 specifies that the scaling matrix is not applied to the blocks coded with LFNST. scaling_matrix_for_lfnst_disabled_flag equal to 0 specifies that the scaling matrix can be applied to the blocks coded with LFNST]]

[0236] Equal to 1 specifies that the chroma scaling list is present in scaling_list_data(). scaling_list_chroma_present_flag equal to 0 specifies that the chroma scaling list is not present in scaling_list_data(). The requirement for bitstream conformance is that when ChromaArrayType is equal to 0, scaling_list_chroma_present_flag shall be equal to 0, and when ChromaArrayType is not equal to 0, scaling_list_chroma_present_flag shall be equal to 1.

[0237] 7.3.7.1 General slice header syntax

[0238]

[0239] 7.4.8.1 General slice header semantics

[0240] Equal to 1 specifies that when decoding the current slice, the explicit scaling list signaled in the reference scaling list APS (where aps_params_type is equal to SCALING_APS and adaptation_parameter_set_id is equal to ph_scaling_list_aps_id) is used during the scaling of the transform coefficients. slice_explicit_scaling_list_used_flag equal to 0 specifies that no explicit scaling list is used during the scaling of the transform coefficients when decoding the current slice. When not present, the value of slice_explicit_scaling_list_used_flag is inferred to be equal to

[0241] 4. Examples of technical problems solved by the disclosed technical solution

[0242] The existing designs and the latest progress of ALF, SAO, the Scaling List, and LMCS have the following problems:

[0243] 1. The design of the Scaling List / LMCS solves many problems in the latest VVC text. However, the following problems are further identified:

[0244] a. If a picture contains only one strip, signaling of the strip-level control flag is not necessary.

[0245] b. If a picture contains only one strip, the allowed mode types (e.g., enabled for all strips; disabled for all strips; and enabled for at least one but not all strips) can be reduced to two modes instead of three.

[0246] 2. SAO / ALF can be controlled in either PH or SH, but not both, which limits flexibility.

[0247] 3. The semantics of sps_lmcs_enabled_flag and sps_sao_enabled_flag are inaccurate. Even when the SPS flag is true, each strip or block can choose to apply LMCS / SAO or not.

[0248] 4. The strip type in SH and / or the allowed inter / intra / B strip type flags in PH are signaled without considering the case where a picture contains only one strip.

[0249] 5. List of solutions and embodiments

[0250] To solve the above problems and other problems, the methods summarized below are disclosed. These items should be considered as examples explaining the general concept and should not be construed in a narrow manner. In addition, these items can be used alone or in any combination.

[0251] Related to the LMCS / Scaling List

[0252] 1. The allowed LMCS and / or Scaling List mode types (e.g., enabled for all strips; disabled for all strips; and enabled for at least one but not all strips) can depend on whether the PH syntax structure exists in the strip header (or whether the current picture contains only one strip).

[0253] a. In one example, when the PH syntax structure exists in the strip header (or whether the current picture contains only one strip), only two mode types are allowed.

[0254] b. In one example, how to signal the pattern type may depend on whether the PH syntax structure exists in the slice header (or whether the current picture contains only one slice).

[0255] i. Alternatively, when the PH syntax structure exists in the slice header (or whether the current picture contains only one slice), the signaled pattern type should be 0 or 1 (or the signaled pattern should not be equal to one of the three patterns).

[0256] 2. Whether to signal an indicator for using / enabling LMCS for a lower level (e.g., slice / tile / sub - picture) depends on the non - binary LMCS - related syntax elements signaled at a higher level (e.g., picture, PH / PPS) (such as the LMCS mode index) and whether the PH syntax structure does not exist in the slice header (or whether the current picture contains more than one slice or depends on the pps_one_slice_per_picture_flag).

[0257] a. In one example, whether to signal the lower - level indicator can be based on a conditional check of whether both of the following two conditions are met:

[0258] i. The non - binary LMCS - related syntax elements at the higher level indicate that LMCS is enabled for at least one slice but not for all slices (or indicate that LMCS is enabled at the sequence level or picture level and whether LMCS is used for each slice is controlled at the slice level).

[0259] ii. The current picture consists of more than one slice, or

[0260] the pps_one_slice_per_picture_flag is false.

[0261] b. In one example, if one or both of the following two conditions are true, signaling of the lower - level indicator can be skipped:

[0262] i. The non - binary LMCS - related syntax elements at the higher level indicate that LMCS is enabled for all slices or disabled for all slices (or indicate that LMCS is used by all slices or disabled for all slices).

[0263] ii. The current picture contains only one slice.

[0264] c. In one example, the conditional check changes from the following:

[0265]

[0266] to the following:

[0267]

[0268] d. In one example, the condition check is from the following:

[0269]

[0270] modified to the following:

[0271]

[0272]

[0273] e. Alternatively, in addition, when the lower-level indicator of LMCS (e.g., slice_lmcs_

[0274] used_flag / slice_lmcs_enabled_flag) is not signaled, it is inferred according to the non-binary LMCS-related syntax elements.

[0275] i. In one example, the non-binary LMCS-related syntax element is

[0276] ph_lmcs_mode_idc.

[0277] ii. Alternatively, in addition, the inference using the slice-level LMCS (e.g., slice_lmcs_used_

[0278] flag / slice_lmcs_enabled_flag) is (ph_lmcs_mode_idc == 0? 0 : 1).

[0279] 3. Whether to signal the indicator for using / enabling the Explicit Scaling List (ESL) at the lower level (e.g., slice / tile / sub-picture) depends on the non-binary LMCS-related syntax element (e.g., ESL mode index) signaled at the higher level (e.g., picture, in PH / PPS) and whether the PH syntax structure does not exist in the tile header (or whether the current picture contains more than one tile or depends on pps_one_slice_per_picture_flag).

[0280] a. In one example, whether to signal the lower-level indicator can be based on a condition check of whether both of the following two conditions are met:

[0281] i. The non-binary ESL-related syntax element at the higher level indicates that ESL is enabled for at least one tile rather than for all tiles (or indicates that ESL is enabled at the sequence level or picture level, and whether ESL is used for each tile is controlled at the tile level).

[0282] ii. The current picture consists of more than one strip, or

[0283] the pps_one_slice_per_picture_flag is false.

[0284] b. In one example, if both of the following two conditions are true, signaling of the lower-level indicator can be skipped:

[0285] i. The higher-level non-binary ESL-related syntax element indicates that ESL is enabled for all strips

[0286] or ESL is disabled for all strips (or indicates that ESL is used by all strips or disabled for all strips).

[0287] ii. The current picture contains only one strip.

[0288] c. In one example, the condition check changes from the following:

[0289]

[0290] to the following:

[0291]

[0292] d. In one example, the condition check changes from the following:

[0293]

[0294] to the following:

[0295]

[0296] e. Alternatively, in addition, when the lower-level indicator of ESL (e.g., slice_explicit_

[0297] scaling_list_used_flag / slice_lmcs_enabled_flag) is not signaled, it is inferred according to the non-binary ESL-related syntax element.

[0298] i. In one example, the non-binary ESL-related syntax element is ph_lmcs_mode_idc.

[0299] ii. Alternatively, in addition, the inference using strip-level ESL (e.g., slice_lmcs_used_

[0300] flag / slice_lmcs_enabled_flag) is (ph_lmcs_mode_idc == 0?

[0301] 0:1).

[0302] 4. The semantics of the LMCS SPS flag is modified as follows:

[0303] Equal to 1 specifies to [[Use]] the luminance mapping with chroma scaling. sps_lmcs_enabled_flag equal to 0 specifies to [[Not use]] the luminance mapping with chroma scaling.

[0304] Or as follows:

[0305] Equal to 1 specifies the luminance mapping with chroma scaling, in CLVS of sps_lmcs_enabled_flag equal to 0 specifies the luminance mapping with chroma scaling, and in CLVS of

[0306] Related to the indication of the slice type

[0307] 5. Whether and / or how to signal the slice type in SH (e.g., slice_type) and / or the allowed inter / intra / B slice type flags in PH (ph_inter_slice_allowed_flag, ph_intra_slice_allowed_flag, ph_b_slice_allowed_flag) can depend on whether the picture is only allowed to have one slice.

[0308] a. In one example, whether the picture is only allowed to have one slice can be indicated by pps_one_slice_per_picture_flag being true.

[0309] b. In one example, whether the picture is only allowed to have one slice can be indicated by the presence of the PH syntax structure in the slice header.

[0310] c. In one example, if for the current picture, each picture is only allowed one slice, then the following can be further applied:

[0311] i. If ph_inter_slice_allowed_flag is true, then no signaling is sent.

[0312] ph_intra_slice_allowed_flag

[0313] ii. If ph_B_slice_allowed_flag is true, then no signaling is sent.

[0314] ph_intra_slice_allowed_flag

[0315] iii. If ph_intra_slice_allowed_flag is true, then no signaling is sent.

[0316] ph_b_slice_allowed_flag

[0317] iv. slice_type is not signaled and inferred.

[0318] Related to the Loop filtering techniques represented (e.g., Deblocking filter, ALF, SAO)

[0319] 6. The semantics of the SAO SPS flag is modified as follows:

[0320] Equal to 1 specifies the sample adaptive offset process After the deblocking filtering process [[Applied]] to the reconstructed picture. sps_sao_enabled_flag equal to 0 specifies that the sample adaptive offset process is not applied to the reconstructed picture after the deblocking filtering process.

[0321] 7. Coding and decoding tools The indicator of the enabled mode type of

[0322] can be signaled at the first video unit level.

[0323] i. In one example, the first video unit can be a picture.

[0324] ii. In one example, the sub-video unit can be a strip / slice / sub-picture.

[0325] b. The enabled mode type can be signaled in the PH / PPS.

[0326] c. The allowed mode types can depend on whether the PH syntax structure is present in the slice header (or whether the current picture only contains one slice).

[0327] i. In one example, when the PH syntax structure is present in the slice header (or whether the current picture only contains one slice or depends on the pps_one_slice_per_picture_flag), only two mode types are allowed.

[0328] ii. In one example, how the mode type is signaled can depend on whether the PH syntax structure is present in the slice header (or whether the current picture only contains one slice or depends on the pps_one_slice_per_picture_flag).

[0329] 1. Alternatively, when the PH syntax structure is present in the slice header (or the current picture only contains one slice or the pps_one_slice_per_picture_flag is true), the signaled mode type should be 0 or 1 (or the signaled mode should not be equal to one of the three modes).

[0330] 8. Whether to signal the use / enable for lower levels (e.g., slice / tile / sub - picture) of the indicator depends on whether the PH syntax structure is not present in the slice header (or whether the current picture contains more than one slice or depends on the value of the pps_one_slice_per_picture_flag).

[0331] b. In one example, if the current picture only includes one picture, or the PH syntax structure is present in the slice header, or the pps_one_slice_per_picture_flag is true, then signaling of the lower - level indicator is skipped.

[0332] c. Additionally, alternatively, when the lower - level indicator is not signaled, it is inferred as the enable / use value signaled at the higher level (e.g., in the PH / PPS).

[0333] 9. The use can be indicated at two levels and two - level control of the codec tool is introduced, where higher - level control (e.g., picture - level) and lower - level (e.g., slice - level) control are used, and the presence of the lower - level control information depends on the higher - level control information. Additionally, the following also applies:

[0334] a. In the first example, apply one or more of the following sub-bullets:

[0335] i. The first non-binary value indicator (e.g., ) can be signaled at a higher level (e.g., in the picture header (PH)) to specify how to enable at a lower level

[0336] 1) In one example, when the first indicator is equal to X (e.g., X = 1), it specifies that X is enabled for all stripes associated with the PH; when the first indicator is equal to Y (Y!= X) (e.g., Y = 2), it specifies that one or more but not all stripes associated with the PH are enabled When the first indicator is equal to Z (Z!= X and Z!= Y) (e.g., Z = 0), it specifies that all stripes associated with the PH are disabled

[0337] a) Additionally, alternatively, when the first indicator is absent, the value of the indicator is inferred to be equal to a default value, such as Z.

[0338] b) Alternatively, when the first indicator is equal to Y (Y!= X) (e.g., Y = 1), it specifies that the use of is enabled for the picture during the scaling process of the transform and / or non-transform coefficients when decoding the stripe.

[0339] c) Alternatively, when the first indicator is equal to Y (Y!= X) (e.g., Y = 1), it specifies that the use of can be enabled for the picture.

[0340] 2) In one example, when the first indicator is equal to X (e.g., X = 2), it specifies that all stripes associated with the PH are disabled When the first indicator is equal to Y (Y!= X) (e.g., Y = 1), it specifies that one or more but not all stripes associated with the PH are disabled When the first indicator is equal to Z (Z!= X and Z!= Y) (e.g., Z = 0), it specifies that all stripes associated with the PH are enabled

[0341] a) Additionally, alternatively, when the first indicator is absent, the value of the indicator is inferred to be equal to a default value, such as X.

[0342] b) Alternatively, when the first indicator is equal to Y (Y!=X) (eg, Y=1), it specifies that when decoding a slice, the transform and / or non-transform coefficients are scaled. The usage of images can be disabled.

[0343] c) Alternatively, when the first indicator is equal to Y (Y!=X) (eg, Y=1), it specifies that when decoding a slice, the transform and / or non-transform coefficients are scaled. The use of images is disabled.

[0344] 3) In addition, alternatively, the Enable flags (for example, ) conditionally signals the first indicator.

[0345] 4) In addition, alternatively, the first indicator can be encoded and decoded with u(v), or u(2) or ue(v).

[0346] 5) In addition, alternatively, the first indicator may be encoded or decoded with a truncated unary code. 6) In addition, alternatively, under the condition of checking the value of the first indicator, the corresponding APS information used by the slice (e.g., ALF APS ).

[0347] ii. Enable / disable for lower levels can be signaled at lower levels (e.g. in the slice header) A second indicator (e.g., slice_TX_present_flag) may be conditionally signaled by checking the value of the first indicator.

[0348] 1) In one example, the second indicator may be signaled under a conditional check of "first indicator equals Y".

[0349] a) Alternatively, the second indicator may be signaled under a conditional check of "value of first indicator >> 1" or "value of first indicator / 2" or "value of first indicator & 0x01".

[0350] b) Furthermore, alternatively, the second indicator may be absent and inferred to be enabled when the first indicator is equal to Y; or inferred to be disabled when the first indicator is equal to Z.

[0351] 2) Whether to signal the second indicator may depend on the first indicator and / or whether the current picture consists of more than one slice (or whether the PH syntax is not present in the SH).

[0352] a) If the first indicator indicates enabling for at least one strip rather than all strips and the PH syntax does not exist in the SH (or pps_one_slice_per_picture_flag is false), then the second indicator can be signaled.

[0353] b) If the first indicator indicates whether enabling or disabling is for all strips or the PH syntax exists in the SH (or pps_one_slice_per_picture_flag is true), then signaling of the second indicator can be skipped.

[0354] i. Additionally, alternatively, when not signaled, it is inferred according to the value of the first indicator, for example, set to the value of the first indicator or set to (first indicator == 0? 0 : 1).

[0355] b. In the second example, apply one or more of the following sub - bullet points:

[0356] i. More than one indicator can be signaled at a higher level (e.g., in the picture header (PH)) to specify how to enable at a lower level

[0357] 1) In one example, two indicators (e.g., two 1 - bit flags) can be signaled in the PH

[0358] a) In one example, the first indicator specifies whether there is at least one strip associated with the PH that is enabled The second indicator specifies whether all strips associated with the PH are enabled

[0359] ii. Additionally, alternatively, the second indicator can be signaled conditionally according to the value of the first indicator, for example, when the first indicator specifies that there is at least one enabled strip I. Additionally, alternatively, when the second indicator does not exist, it is inferred that all strips are enabled

[0360] I. Additionally, alternatively, if the first indicator in the compliant bitstream is false, then the second indicator must be false.

[0361] II. Alternatively, if the first indicator in the compliant bitstream is false, then the second indicator must be false.

[0362] iii. Additionally, alternatively, according to at least one of the value of the first indicator and the value of the second indicator, for example, when the first indicator specifies that at least one strip is enabled and the second indicator specifies that not all strips are enabled When it is possible, the third indicator can be signaled conditionally in the SH.

[0363] I. Additionally, alternatively, when the third indicator does not exist, it can be inferred based on the value of the first and / or second indicator (e.g., inferred to be equal to the value of the first indicator).

[0364] b) Alternatively, the first indicator specifies whether there is at least one strip associated with the PH that is disabled. The second indicator specifies whether all strips associated with the PH are disabled

[0365] i. Additionally, alternatively, the second indicator can be signaled conditionally based on the value of the first indicator, e.g., when the first indicator specifies that there is at least one strip.

[0366] I. Additionally, alternatively, when the second indicator does not exist, it is inferred that all strips associated with the PH are disabled

[0367] ii. Additionally, alternatively, based on at least one of the value of the first indicator and the value of the second indicator, e.g., when the first indicator specifies that at least one strip is enabled and the second indicator specifies that not all strips are disabled When it is possible, the third indicator can be signaled conditionally in the SH.

[0368] I. Additionally, alternatively, when the third indicator does not exist, it can be inferred based on the value of the first and / or second indicator (e.g., inferred to be equal to the value of the first indicator).

[0369] 2) Additionally, alternatively, the first indicator can be signaled conditionally based on the enable flag in the sequence level (e.g., ).

[0370] ii. The third indicator for enabling / disabling at a lower level (e.g., in the strip header) can be signaled (e.g., ), and it can be signaled conditionally by checking the value of the first indicator and / or the second indicator. ), and it can be signaled conditionally by checking the value of the first indicator and / or the second indicator.

[0371] 3) In one example, the third indicator can be signaled under the condition check of "not all strips are enabled " or "not all strips are disabled ".

[0372] c. In the third example, two 1-bit flags can be signaled at a higher level (e.g., in the picture header (PH)) to specify how to enable at a lower level (e.g., in the slice header (SH)).

[0373] i. The first PH flag (e.g., named ) being equal to 1 specifies that all slices of the picture use The first PH flag being equal to 0 specifies that each slice of the picture can use or not use

[0374] ii. The second PH flag (e.g., named ph_no_slice_uses_TX_flag) being equal to 1 specifies that no slice in the picture uses The second PH flag being equal to 0 specifies that each slice of the picture can use or not use

[0375] iii. When the first PH flag is equal to 0, only the second PH flag is signaled.

[0376] iv. When the first PH flag is equal to 1 or when (the first PH flag is equal to 0 and the second PH flag is equal to 0), the scaled list APS ID is signaled in the PH.

[0377] v. When the first PH flag is equal to 0 and the second flag is equal to 0, the SH flag (e.g., named ) is signaled in the SH.

[0378] vi. When the first PH flag is equal to 1, the value of the SH flag is inferred to be equal to 0.

[0379] vii. When the first PH flag is equal to 0 and the second PH flag is equal to 0, the value of the SH flag is inferred to be equal to 0.

[0380] viii. If the value of the SH flag is equal to 1, then is used to decode the slice. Otherwise, is used to decode the slice.

[0381] d. In the fourth example, one or more indicators of whether one or more or each of the scaled list related aspects (e.g., enable / disable, APS ID) exist in the PH or SH can be signaled.

[0382] i. Alternatively, one indicator is used, and moreover, this indicator is a 1-bit flag.

[0383] 1) In one example, when the indicator specifies that the relevant aspects are present in the PH, all strips infer the values present in the PH, and the signals of those relevant aspects are skipped in the SH.

[0384] 2) In one example, when the indicator specifies that the relevant aspects are present in the SH, and the signaling of those relevant aspects is skipped in the PH.

[0385] ii. In one example, one or more indicators signal in the PH value.

[0386] iii. In another example, one or more indicators signal in the PPS.

[0387] iv. In another example, one or more indicators signal in the SPS.

[0388] Generally

[0389] 10. The method proposed above can be extended to other codec tools. For example, non-binary value indicators are used to indicate the mode type.

[0390] 11. Whether (e.g., in the PH / SH) to signal the relevant information can be further controlled by some higher-level syntax elements, such as in the PPS / SPS.

[0391] d. Alternatively, whether (e.g., in the PPS / PH / SH) to signal the relevant information can be further controlled by a certain syntax element in the higher level, such as in the SPS.

[0392] e. In one example, the syntax element can exist in the higher level to indicate whether the on / off control can be different in the picture set, picture, or strip.

[0393] Figure 5 FIG. shows a block diagram of an exemplary video processing system 500 in which various techniques disclosed herein can be implemented. Various embodiments may include some or all components of system 500. System 500 may include an input 502 for receiving video content. The video content may be received in a raw or uncompressed format, e.g., 8- or 10-bit multi-component pixel values, or may be in a compressed or encoded format. Input 502 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces (such as Ethernet, passive optical network (PON), etc.), and wireless interfaces (such as Wi-Fi or cellular interfaces).

[0394] System 500 may include an encoding / decoding component 504, which may implement various encoding / decoding or coding methods described in this document. The encoding / decoding component 504 may reduce the average bit rate of the video from the input 502 to the output of the encoding / decoding component 504 to produce an encoded / decoded representation of the video. Thus, encoding / decoding techniques are sometimes referred to as video compression or video transcoding techniques. As represented by component 506, the output of the encoding / decoding component 504 may be stored or transmitted via a connected communication. Component 508 may use the stored or communicatively transmitted bitstream (or encoded / decoded) representation of the video received at the input 502 to generate pixel values or a displayable video to be sent to the display interface 510. The process of generating a user-visible video from the bitstream representation is sometimes referred to as video decompression. Additionally, although certain video processing operations are referred to as "encoding / decoding" operations or tools, it should be understood that encoding / decoding tools or operations are used at the encoder, and the corresponding decoding tools or operations opposite to the encoding result will be performed by the decoder.

[0395] Examples of a peripheral bus interface or a display interface may include Universal Serial Bus (USB), High-Definition Multimedia Interface (HDMI), or DisplayPort, etc. Examples of a storage interface include SATA (Serial Advanced Technology Attachment), PCI, IDE interface, etc. The techniques described in this document may be implemented in various electronic devices, such as mobile phones, laptops, smartphones, or other devices capable of performing digital data processing and / or video display.

[0396] Figure 6 is a block diagram of a video processing apparatus 3600. The apparatus 3600 may be used to implement one or more methods described herein. The apparatus 3600 may be implemented in a smartphone, a tablet computer, a computer, an Internet of Things (IoT) receiver, etc. The apparatus 3600 may include one or more processors 3602, one or more memories 3604, and video processing hardware 3606. The processor 3602 may be configured to implement one or more methods described in this document. The memory(ies) 3604 may be used to store data and code for implementing the methods and techniques described herein. The video processing hardware 3606 may be used to implement some of the techniques described in this document in hardware circuitry.

[0397] Figure 8 is a block diagram showing an example video encoding / decoding system 100 that may utilize the techniques of the present disclosure.

[0398] As Figure 8 shown, the video encoding / decoding system 100 may include a source device 110 and a destination device 120. The source device 110 generates encoded video data, which may be referred to as a video encoding device. The destination device 120 may decode the encoded video data generated by the source device 110, which may be referred to as a video decoding device.

[0399] The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.

[0400] The video source 112 may include sources such as a video capture device, an interface for receiving video data from a video content provider, and / or a computer graphics system for generating video data, or a combination of these sources. The video data may include one or more pictures. The video encoder 114 encodes the video data from the video source 112 to generate a bitstream. The bitstream may include a sequence of bits that forms a coded representation of the video data. The bitstream may include coded pictures and associated data. A coded picture is a coded representation of a picture. The associated data may include a sequence parameter set, a picture parameter set, and other syntax structures. The I / O interface 116 may include a modulator / demodulator (modem) and / or a transmitter. The encoded video data may be sent directly to the destination device 120 via the I / O interface 116 over a network 130a. The encoded video data may also be stored on a storage medium / server 130b for access by the destination device 120.

[0401] The destination device 120 may include an I / O interface 126, a video decoder 124, and a display device 122.

[0402] The I / O interface 126 may include a receiver and / or a modem. The I / O interface 126 may obtain the encoded video data from the source device 110 or the storage medium / server 130b. The video decoder 124 may decode the encoded video data. The display device 122 may display the decoded video data to a user. The display device 122 may be integrated with the destination device 120, or may be external to the destination device 120, which is configured to interface with an external display device.

[0403] The video encoder 114 and the video decoder 124 may operate according to a video compression standard, such as the High Efficiency Video Coding (HEVC) standard, the Versatile Video Coding (VVC) standard, and other current and / or future standards.

[0404] Figure 9 is a block diagram showing an example of a video encoder 200, which may be Figure 8 the video encoder 114 in the system 100 shown.

[0405] The video encoder 200 may be configured to perform any or all of the techniques of the present disclosure. In Figure 9In the example, video encoder 200 includes multiple functional components. The techniques described in this disclosure can be shared among various components of video encoder 200. In some examples, the processor can be configured to execute any or all of the techniques described in this disclosure.

[0406] The functional components of video encoder 200 can include a splitting unit 201, a prediction unit 202 that can include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra prediction unit 206, a residual generation unit 207, a transformation unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transformation unit 211, a reconstruction unit 212, a buffer 213, and an entropy coding unit 214.

[0407] In other examples, video encoder 200 can include more, fewer, or different functional components. In an example, prediction unit 202 can include an intra block copy (IBC) unit. The IBC unit can perform prediction in IBC mode, where at least one reference picture is the picture in which the current video block is located.

[0408] In addition, some components such as motion estimation unit 204 and motion compensation unit 205 can be highly integrated, but for purposes of explanation, are shown separately in the Figure 9 example.

[0409] Splitting unit 201 can split a picture into one or more video blocks. Video encoder 200 and video decoder 300 can support various video block sizes.

[0410] Mode selection unit 203 can select, for example, one of the coding modes (intra or inter) based on error results, and provide the resulting intra or inter coded block to residual generation unit 207 to generate residual block data, and to reconstruction unit 212 to reconstruct the coded block for use as a reference picture. In some examples, mode selection unit 203 can select a combination of intra and inter prediction (CIIP) mode, where the prediction is based on an inter prediction signal and an intra prediction signal. In the case of inter prediction, mode selection unit 203 can also select the resolution of the motion vector for the block (e.g., sub-pixel or integer pixel accuracy).

[0411] To perform inter prediction on the current video block, motion estimation unit 204 can generate motion information for the current video block by comparing one or more reference frames from buffer 213 with the current video block. Motion compensation unit 205 can determine the predicted video block for the current video block based on the motion information of pictures other than the picture associated with the current video block from buffer 213 and decoded samples.

[0412] The motion estimation unit 204 and the motion compensation unit 205 can perform different operations on a current video block, e.g., depending on whether the current video block is in an I-slice, a P-slice, or a B-slice.

[0413] In some examples, the motion estimation unit 204 can perform uni-directional prediction on a current video block, and the motion estimation unit 204 can search for a reference video block of the current video block in the reference pictures of list 0 or list 1. The motion estimation unit 204 can then generate a reference index and a motion vector, where the reference index indicates the reference picture in list 0 or list 1 that contains the reference video block, and the motion vector indicates the spatial displacement between the current video block and the reference video block. The motion estimation unit 204 can output the reference index, a prediction direction indicator, and the motion vector as the motion information of the current video block. The motion compensation unit 205 can generate a predicted video block of the current block based on the reference video block indicated by the motion information of the current video block.

[0414] In other examples, the motion estimation unit 204 can perform bi-directional prediction on a current video block. The motion estimation unit 204 can search for a reference video block of the current video block in the reference pictures of list 0 and can also search for another reference video block of the current video block in the reference pictures of list 1. The motion estimation unit 204 can then generate a reference index and a motion vector, where the reference index indicates the reference picture in list 0 or list 1 that contains the reference video block, and the motion vector indicates the spatial displacement between the reference video block and the current video block. The motion estimation unit 204 can output the reference index and the motion vector of the current video block as the motion information of the current video block. The motion compensation unit 205 can generate a predicted video block of the current video block based on the reference video block indicated by the motion information of the current video block.

[0415] In some examples, the motion estimation unit 204 can output the entire set of motion information for the decoding process of the decoder.

[0416] In some examples, the motion estimation unit 204 may not output the entire set of motion information of the current video. Instead, the motion estimation unit 204 can signal the motion information of the current video block by referring to the motion information of another video block. For example, the motion estimation unit 204 can determine that the motion information of the current video block is similar enough to the motion information of an adjacent video block.

[0417] In one example, the motion estimation unit 204 can indicate a value in the syntax structure associated with the current video block, and this value indicates to the video decoder 300 that the current video block has the same motion information as another video block.

[0418] In another example, the motion estimation unit 204 may identify another video block and a motion vector difference (MVD) in the syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 300 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.

[0419] As described above, the video encoder 200 may predictively signal motion vectors. Two examples of predictive signaling techniques that may be implemented by the video encoder 200 include advanced motion vector prediction (AMVP) and merge mode signaling.

[0420] The intra prediction unit 206 may perform intra prediction on the current video block. When the intra prediction unit 206 performs intra prediction on the current video block, the intra prediction unit 206 may generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block may include a predicted video block and various syntax elements.

[0421] The residual generation unit 207 may generate residual data for the current video block by subtracting (e.g., indicated by a negative sign) the predicted video block of the current video block from the current video block. The residual data for the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.

[0422] In other examples, there may be no residual data for the current video block, e.g., in the skip mode, and the residual generation unit 207 may not perform the subtraction operation.

[0423] The transform processing unit 208 may generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video block associated with the current video block.

[0424] After the transform processing unit 208 generates the transform coefficient video block associated with the current video block, the quantization unit 209 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.

[0425] The inverse quantization unit 210 and the inverse transform unit 211 may apply inverse quantization and inverse transform to the transform coefficient video block, respectively, to reconstruct the residual video block from the transform coefficient video block. The reconstruction unit 212 may add the reconstructed residual video block to the corresponding samples of one or more predicted video blocks generated by the prediction unit 202 to produce a reconstructed video block associated with the current block for storage in the buffer 213.

[0426] After reconstructing a video block in the reconstruction unit 212, a loop filtering operation may be performed to reduce the video block effect in the video block.

[0427] The entropy coding unit 214 may receive data from other functional components of the video encoder 200. When the entropy coding unit 214 receives data, the entropy coding unit 214 may perform one or more entropy coding operations to generate entropy-coded data and output a bitstream including the entropy-coded data.

[0428] Figure 10 is a block diagram illustrating an example of a video decoder 300, which may be Figure 8 the video decoder 114 in the system 100 shown.

[0429] The video decoder 300 may be configured to perform any or all of the techniques of the present disclosure. In Figure 10 an example, the video decoder 300 includes a plurality of functional components. The techniques described in the present disclosure may be shared among various components of the video decoder 300. In some examples, a processor may be configured to perform any or all of the techniques described in the present disclosure.

[0430] In Figure 10 an example, the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, the video decoder 300 may perform a decoding process that is generally inverse to the encoding process described for the video encoder 200 ( Figure 9 ).

[0431] The entropy decoding unit 301 may retrieve the encoded bitstream. The encoded bitstream may include entropy-coded video data (e.g., encoded blocks of video data). The entropy decoding unit 301 may decode the entropy-coded video data, and the motion compensation unit 302 may determine motion information from the entropy-decoded video data, including motion vectors, motion vector precision, reference picture list indices, and other motion information. The motion compensation unit 302 may determine this information (e.g.) by performing AMVP and merge modes.

[0432] The motion compensation unit 302 may generate a motion-compensated block, possibly performing interpolation based on an interpolation filter. The identifier of the interpolation filter to be used with sub-pixel precision may be included in the syntax elements.

[0433] The motion compensation unit 302 may use an interpolation filter such as that used by the video encoder 200 during the encoding of video blocks to calculate the interpolated values of sub-integer pixels of a reference block. The motion compensation unit 302 may determine the interpolation filter used by the video encoder 200 according to the received syntax information, and use the interpolation filter to generate a prediction block.

[0434] The motion compensation unit 302 may use some syntax information to determine the size of the blocks for encoding the (multiple) frames and / or (multiple) slices of the encoded video sequence, the partitioning information that describes how each macroblock of a picture of the encoded video sequence is partitioned, the mode indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-coded block, and other information for decoding the encoded video sequence.

[0435] The intra prediction unit 303 may form a prediction block from spatially adjacent blocks using, for example, an intra prediction mode received in the bitstream. The inverse quantization unit 303 inverse quantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 301. The inverse transform unit 303 applies an inverse transform.

[0436] The reconstruction unit 306 may add a residual block to the corresponding prediction block generated by the motion compensation unit 202 or the intra prediction unit 303 to form a decoded block. If necessary, a deblocking filter may also be applied to filter the decoded block to remove blocking artifact. The decoded video block is then stored in the buffer 307, which provides reference blocks for subsequent motion compensation / intra prediction and also produces the decoded video for presentation on a display device.

[0437] Next, a list of preferred solutions for some embodiments is provided.

[0438] The following solutions illustrate example embodiments of the techniques discussed in the previous section (e.g., item 1).

[0439] 1. A video processing method (e.g., Figure 7 method 700 depicted in), performing (702) a conversion between a video including one or more video pictures including one or more slices and an encoded / decoded representation of the video; wherein the encoded / decoded representation conforms to format rules, and wherein the format rules specify which luminance mapping having a chroma scaling mode or a scaling list mode type can be applied to the conversion of a slice is indicated by a picture header syntax structure in the slice header or a picture header in a picture including a single slice.

[0440] 2. The method of solution 1, wherein the format rules specify that the picture header syntax structure in the group header indicates that only two mode types are allowed.

[0441] The following solutions illustrate example embodiments of the techniques discussed in the previous section (e.g., item 2).

[0442] 3. A video processing method, comprising: performing a conversion between a video including one or more video pictures each including one or more strips and an encoded / decoded representation of the video; wherein the encoded / decoded representation conforms to format rules, and wherein the format rules specify an indicator including indicating enabling a luminance mapping with chroma scaling (LMCS) mode at a first video level depends on a higher-level non-binary LMCS-related syntax element and an indicator of whether the picture consists of only one strip.

[0443] 4. The method according to solution 3, wherein the first video level is the strip level.

[0444] 5. The method according to any one of solutions 3-4, wherein the higher level corresponds to the picture or sequence or picture parameter set level.

[0445] The following solutions illustrate example embodiments of the techniques discussed in the previous section (e.g., items 3-4).

[0446] 6. A video processing method, comprising: performing a conversion between a video including one or more video pictures each including one or more strips and an encoded / decoded representation of the video; wherein the encoded / decoded representation conforms to format rules, and wherein the format rules specify an indicator including indicating enabling an explicit scaling list (ESL) mode at a first video level depends on a higher-level non-binary LMCS-related syntax element and an indicator of whether the picture consists of only one strip.

[0447] 7. The method according to solution 6, wherein the first video level is the strip level.

[0448] 8. The method according to any one of solutions 6-7, wherein the higher level corresponds to the picture or sequence or picture parameter set level.

[0449] The following solutions illustrate example embodiments of the techniques discussed in the previous section (e.g., items 5-8).

[0450] 9. A video encoding method, comprising: performing a conversion between a video including one or more pictures each including one or more strips and an encoded / decoded representation of the video, wherein the encoded / decoded representation conforms to format rules, and wherein the format rules specify that whether the picture exactly includes one strip controls the strip type or strip type flag in the strip header of the exactly one strip.

[0451] 10. The method according to solution 9, wherein the format rules specify that for a picture having exactly one strip, the corresponding picture header syntax structure must be included in the strip header.

[0452] The following solution shows an example embodiment of the technology discussed in the previous section (e.g., item 9).

[0453] 11. A video encoding method, comprising: performing a conversion between a video including one or more pictures each including one or more video regions and an encoded / decoded representation of the video, wherein the encoded / decoded representation conforms to format rules, and wherein the format rules specify two-level signaling including the applicability of a filtering encoding / decoding tool (TX) to the video regions.

[0454] 12. The method according to solution 11, wherein the two-level signaling includes higher-level signaling at the video picture level or a higher level and lower-level signaling at the slice level or a lower level.

[0455] 13. The method according to any one of solutions 11-12, wherein the higher-level signaling includes a non-binary value indicator.

[0456] 14. The method according to any one of solutions 11-13, wherein the lower-level signaling includes a binary value indicator.

[0457] 15. The method according to solution 12, wherein the higher-level signaling includes two 1-bit flags that indicate whether all, some, or conditionally some of the lower-level video regions enable the TX mode.

[0458] 16. The method according to any one of solutions 11-15, wherein the filtering encoding / decoding tool includes the use of a scaling list.

[0459] 17. The method according to any one of solutions 1 to 16, wherein the conversion includes encoding the video into an encoded / decoded representation.

[0460] 18. The method according to any one of solutions 1 to 16, wherein the conversion includes decoding the encoded / decoded representation to generate pixel values of the video.

[0461] 19. A video decoding apparatus, comprising a processor configured to implement the method described in one or more of solutions 1 to 18.

[0462] 20. A video encoding apparatus, comprising a processor configured to implement the method described in one or more of solutions 1 to 18.

[0463] 21. A computer program product having computer code stored thereon, which when executed by a processor causes the processor to implement the method according to any one of solutions 1 to 18.

[0464] 22. The method, apparatus, or system described in this document.

[0465] Figure 13 It is a flowchart representation of method 1300 for video processing according to the present technology. Method 1300 includes, at operation 1310, performing a conversion between a video picture of a video including one or more strips and a bitstream of the video picture according to a rule. The rule specifies whether a first syntax element indicating the use of a codec tool exists at a first level based on a syntax flag indicating whether a syntax structure at a second level does not exist at the first level. The second level is higher than the first level, and the second level is a video picture level or higher than the video picture level.

[0466] In some embodiments, the first level includes a strip header, and the second level includes a picture header. In some embodiments, whether the video picture includes a single strip is indicated by the syntax flag.

[0467] In some embodiments, the rule further specifies whether the first syntax element exists at the first level based on a value of a second syntax element indicating the use of the codec tool at the second level. In some embodiments, the use of the codec tool at the first level is based on (1) a syntax flag indicating whether a syntax structure at the second level does not exist at the first level, and (2) the value of the second syntax element.

[0468] In some embodiments, the codec tool includes a tool that maps luminance samples to specific values and optionally applies a scaling operation to the values of chrominance samples. In some embodiments, the codec tool includes a luminance mapping tool with chrominance scaling. In some embodiments, in response to (1) the second syntax element indicating in the picture header that the luminance mapping tool with chrominance scaling is enabled for at least one strip, and (2) the syntax structure of the picture header not existing in the strip header, a first syntax element indicating the use of the luminance mapping tool with chrominance scaling exists. In some embodiments, in response to (1) the syntax element indicating in the picture header that the luminance mapping tool with chrominance scaling is disabled, or (2) the syntax structure of the picture header existing in the strip header, the first syntax element indicating the use of the luminance mapping tool with chrominance scaling is omitted. In some embodiments, in the case where the use of the luminance mapping tool with chrominance scaling is omitted in the strip header, the use is inferred based on the second syntax element indicated in the picture header.

[0469] In some embodiments, the codec tool includes an explicit scaling list. In some embodiments, the explicit scaling list is used in the scaling process of transform coefficients. In some embodiments, in response to (1) a second syntax element indicating in a picture header that the explicit scaling list is enabled for at least one slice, and (2) the picture header syntax structure not being present in the slice header, there is a first syntax element indicating the use of the explicit scaling list. In some embodiments, in response to (1) the second syntax element indicating in the picture header that the explicit scaling list is disabled, or (2) the picture header syntax structure being present in the slice header, the first syntax element indicating the use of the explicit scaling list is omitted. In some embodiments, in the case where the use of the explicit scaling list is omitted in the slice header, the use is inferred based on the second syntax element indicated in the picture header.

[0470] Figure 14 is a flowchart representation of a method 1400 for video processing according to the present technology. Method 1400 includes, at operation 1410, performing a conversion between a video and a bitstream of the video according to a rule. The rule specifies that a syntax element in the sequence parameter set of the video indicates whether a luminance mapping with chroma scaling (LMCS) tool is enabled for a coded layer video sequence (CLVS) of a reference sequence parameter set.

[0471] In some embodiments, the syntax element being equal to 1 specifies that the LMCS tool is enabled for the CLVS, and wherein the syntax element being equal to 0 specifies that the LMCS tool is disabled for the CLVS.

[0472] Figure 15 is a flowchart representation of a method 1500 for video processing according to the present technology. Method 1500 includes, at operation 1510, performing a conversion between a video and a bitstream of the video according to a rule. The rule specifies that a syntax element in the sequence parameter set of the video indicates whether a sample adaptive offset (SAO) tool is enabled for a coded layer video sequence (CLVS) of a reference sequence parameter set.

[0473] In some embodiments, the syntax element being equal to 1 specifies that the SAO tool is enabled for the CLVS, and wherein the syntax element being equal to 0 specifies that the SAO tool is disabled for the CLVS.

[0474] Figure 16is a flowchart representation of a method 1600 for video processing according to the present technology. Method 1600 includes, at operation 1610, performing a conversion between a video picture of a video including one or more stripes and a bitstream of the video picture according to a rule. Whether and / or how the rule indicates the use of a scaling tool is determined based on whether the video picture includes a single stripe. The use of the scaling tool includes whether to allow a luminance mapping with chroma scaling (LMCS) tool for the conversion. The use also includes the number of types of scaling modes allowed for the conversion.

[0475] In some embodiments, the type of scaling mode indicates whether to enable or disable the scaling tool for all or part of one or more stripes of a video picture. In some embodiments, the rule specifies that whether the video picture includes a single stripe is indicated by whether a picture header syntax structure of the video picture exists in a stripe header of the stripe. In some embodiments, the rule specifies that in the case where the video picture includes a single stripe, only two types of scaling modes are allowed for the conversion. In some embodiments, the two types of scaling modes are indicated by values 0 or 1.

[0476] Figure 17 is a flowchart representation of a method 1700 for video processing according to the present technology. Method 1700 includes, at operation 1710, performing a conversion between a video picture of a video including one or more stripes and a bitstream of the video picture according to a rule. The rule specifies that whether and / or how the allowed stripe types are indicated in a picture header and / or a stripe header is determined based on whether the video picture includes a single stripe.

[0477] In some embodiments, the rule specifies that whether the video picture contains a single stripe is indicated by whether a picture header syntax structure of the video picture exists in a stripe header of the stripe. In some embodiments, in the case where the video picture includes a single stripe, when a second syntax flag in the picture header specifies that one or more stripes in the video are allowed to have a specific stripe type, a first syntax flag in the picture header that specifies whether all stripes of the video picture have the specific stripe type is omitted. In some embodiments, in the case where the video picture includes a single stripe, when a second syntax flag in the picture header specifies that one or more stripes allowing codec type B are allowed, a first syntax flag in the picture header that specifies whether all stripes of the video picture have a specific stripe type is omitted. In some embodiments, in the case where the video picture includes a single stripe, when a second syntax flag in the picture header specifies that all stripes of the video picture have a specific stripe type, a first syntax flag in the picture header that specifies whether one or more stripes allowing codec type B are allowed is omitted. In some embodiments, the type of the stripes in the video picture is omitted and inferred.

[0478] Figure 18is a flowchart representation of a method 1800 for video processing according to the present technology. Method 1800 includes, at operation 1810, performing a conversion between a video picture of a video including one or more strips and a bitstream of the video picture according to a rule. The rule specifies a mode type indicating an encoding / decoding tool at a video unit level for the conversion.

[0479] In some embodiments, the mode type includes: (1) a first type indicating that the encoding / decoding tool is enabled for all sub-units of a video unit, (2) a second type indicating that the encoding / decoding tool is disabled for all sub-units of a video unit, and (3) a third type indicating that the encoding / decoding tool is enabled for at least one but not all sub-units of a video unit. In some embodiments, the video unit includes a picture. In some embodiments, the sub-units include strips, slices, or sub-pictures of a picture. In some embodiments, the mode type is indicated in a picture header or a picture parameter set.

[0480] In some embodiments, the number of allowed mode types is determined based on whether the video picture includes a single strip. In some embodiments, when the video picture includes a single strip, only two mode types are allowed. In some embodiments, how the mode type is indicated is determined based on whether the video picture includes a single strip. In some embodiments, when the video picture includes a single strip, the mode type is limited to a subset of the allowed mode types. In some embodiments, the rule specifies that whether the video picture includes a single strip is indicated by the presence or absence of a picture header syntax structure of the video picture in a strip header of the strip.

[0481] Figure 19 is a flowchart representation of a method 1900 for video processing according to the present technology. Method 1900 includes, at operation 1910, performing a conversion between a video picture of a video including one or more strips and a bitstream of the video picture according to a rule. The rule specifies that a non-binary syntax element or a plurality of syntax flags at a first video unit level are used to indicate the use of an encoding / decoding tool at a second video unit level lower than the first video unit level.

[0482] In some embodiments, the first video unit level includes a picture level, and the second video unit level includes a strip level. In some embodiments, the non-binary syntax element is encoded using at least more than two bits. The non-binary syntax element being equal to X specifies that the encoding / decoding tool is enabled for all strips associated with the video picture, the non-binary syntax element being equal to Y specifies that the encoding / decoding tool is enabled for at least one but not all strips associated with the video picture, and the non-binary syntax element being equal to Z specifies that the encoding / decoding tool is disabled for all strips associated with the video picture, where X!= Y, Y!= Z, and X!= Z.

[0483] In some embodiments, in the case of omitting a non-binary syntax element, the non-binary syntax element is inferred to have a default value. In some embodiments, the non-binary syntax element being equal to Y indicates that the codec tool can be applied to the scaling process of the transformed and / or non-transformed coefficients for this transformation. In some embodiments, X = 1, Y = 2, and Z = 0. In some embodiments, X = 2, Y = 1, and Z = 0.

[0484] In some embodiments, the non-binary syntax element is conditionally indicated based on a corresponding syntax flag at the sequence level. In some embodiments, the non-binary syntax element is encoded as an unsigned integer, an unsigned integer with a left bit first, a syntax element decoded by a 0th order Exp-Golomb code, or a truncated unary value. In some embodiments, the corresponding adaptive parameter set information used by one or more slices of the video picture is indicated based on the non-binary syntax element.

[0485] In some embodiments, the plurality of syntax flags includes a first syntax flag indicating whether a codec tool is enabled or disabled for at least one slice of the video picture, and a second syntax flag indicating whether the codec tool is enabled or disabled for all slices of the video picture. In some embodiments, the second syntax flag is conditionally indicated based on the value of the first syntax flag. In some embodiments, in the case of omitting the second syntax flag, the second syntax flag is inferred to indicate whether the codec tool is enabled or disabled for all slices. In some embodiments, the rule specifies that in the case where the first syntax flag indicates that the codec tool is not enabled for at least one slice of the video picture, the second syntax flag has a value indicating that the codec tool is disabled for all slices of the video picture.

[0486] In some embodiments, the plurality of syntax flags further includes a third syntax flag conditionally indicated in the slice header according to the first syntax flag or the second syntax flag. In some embodiments, the third syntax flag is inferred based on the value of the first syntax flag and / or the second syntax flag. In some embodiments, the third syntax flag is conditionally indicated according to a corresponding syntax flag at the sequence level.

[0487] In some embodiments, the plurality of syntax flags include a first syntax flag indicating whether all slices of a video picture use a codec tool, and a second syntax flag indicating whether not all slices of a video picture use a codec tool. In some embodiments, the second syntax flag is indicated only when the first syntax flag indicates that not all slices of the video picture use the codec tool. In some embodiments, the scaling list adaptive parameter set identifier is included in the picture header when the first syntax flag indicates that all slices of the video picture use the codec tool, or when the first syntax flag and the second syntax flag indicate that at least one slice of the video picture uses the codec tool. In some embodiments, a third syntax flag is indicated in the slice header when the first syntax flag indicates that not all slices of the video picture use the codec tool and the second syntax flag indicates that at least one slice of the video picture uses the codec tool. In some embodiments, the value of the third syntax flag is based on the first syntax flag and / or the second syntax flag.

[0488] In some embodiments, the codec tool is associated with the use or information of a scaling list. In some embodiments, the use or information of the scaling list is omitted in a second video unit level when a non-binary syntax element or a plurality of syntax flags are present in a first video unit level. In some embodiments, a plurality of syntax flags are indicated in the first video unit level, and the first video unit level includes a picture header, a picture parameter set, or a sequence parameter set. In some embodiments, the syntax elements of the second video unit level are conditionally indicated based on the non-binary syntax element or the plurality of syntax flags.

[0489] In some embodiments, the indication of the non-binary syntax element or the plurality of syntax flags is determined based on information in a third video unit level higher than the first video unit level. In some embodiments, the third video unit level includes a picture parameter set or a sequence parameter set. In some embodiments, the syntax elements in the third video unit level indicate that the use of the codec tool varies within a picture set, a picture, or a slice.

[0490] In some embodiments, the conversion includes encoding a video into a bitstream. In some embodiments, the conversion includes decoding the bitstream to generate a video.

[0491] In the solutions described herein, an encoder can conform to format rules by generating a codec representation according to the format rules. In the solutions described herein, a decoder can use the format rules to utilize the knowledge of the presence and absence of syntax elements according to the format rules to parse the syntax elements of the codec representation to generate a decoded video.

[0492] In this document, the term "video processing" may refer to video encoding, video decoding, video compression, or video decompression. For example, during the conversion from the pixel representation of a video to the corresponding bitstream representation, a video compression algorithm may be applied, and vice versa. As defined by the syntax, the bitstream representation of the current video block may (e.g.) correspond to bits that are co-located or scattered at different positions within the bitstream. For example, a macroblock may be encoded based on the transform and coding error residual values, and also using bits in the headers and other fields in the bitstream. Additionally, during the conversion, the decoder may parse the bitstream based on that determination, knowing that some fields may or may not be present, as described in the above solutions. Similarly, the encoder may determine to include or exclude certain syntax fields and may accordingly generate the encoded / decoded representation by including or excluding the syntax fields from the encoded / decoded representation.

[0493] The disclosed and other solutions, examples, embodiments, modules, and functional operations described in this document may be implemented in digital electronic circuitry, or in computer software, firmware, or hardware, including the structures disclosed in this document and their structural equivalents, or in a combination of one or more of them. The disclosed and other embodiments may be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer-readable medium for execution by, or to control the operation of, a data processing apparatus. The computer-readable medium may be a machine-readable storage device, a machine-readable storage substrate, a memory device, a combination of substances implementing a machine-readable propagated signal, or a combination of one or more of them. The term "data processing apparatus" encompasses all apparatuses, devices, and machines for processing data, including, for example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, the apparatus may include code that creates an execution environment for the computer programs being discussed, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them. A propagated signal is an artificially generated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal that is generated to encode information for transmission to a suitable receiver device.

[0494] A computer program (also known as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for a computing environment. A computer program does not necessarily correspond to a file in a file system. The program can be stored in a part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), stored in a single file dedicated to the program in question, or stored in multiple cooperating files (e.g., files that hold one or more modules, subroutines, or portions of code). A computer program can be deployed to execute on one computer or on multiple computers located at one site or distributed across multiple sites and interconnected by a communication network.

[0495] The processes and logical flows described in this document can be performed by one or more programmable processors that execute one or more computer programs to perform functions by operating on input data and generating output. The processes and logical flows can also be performed by special-purpose logic circuitry, and apparatuses can also be implemented as special-purpose logic circuitry, such as an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit).

[0496] By way of example, processors suitable for the execution of a computer program include both general and special purpose microprocessors, as well as any one or more processors of any kind of digital computer. Generally, a processor will receive instructions and data from a read only memory or a random access memory or both. The essential elements of a computer are a processor for executing the instructions and one or more memory devices for storing the instructions and data. Generally, a computer will also include or be operatively coupled to one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, to receive data from or transfer data to the mass storage device, or both. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including by way of example semiconductor memory devices, such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special-purpose logic circuitry.

[0497] Although this patent document contains many details, these details should not be construed as limiting the scope of any subject matter or of what is claimed, but rather as descriptions of features that are specific to particular embodiments of a particular technology. Certain features that are described in this patent document in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, the various features that are described in the context of a single embodiment can also be implemented separately or in any suitable sub-combination in multiple embodiments. Moreover, although the features may be described above as acting in certain combinations and even initially claimed as such, one or more features from a claimed combination can in some cases be deleted from the combination, and the claimed combination can be directed to a sub-combination or variant of a sub-combination.

[0498] Similarly, although operations are depicted in the drawings in a particular order, this should not be understood as requiring that the operations be performed in the particular order shown or in sequential order, or that all of the illustrated operations be performed, to achieve desirable results. Moreover, the separation of various system components in the embodiments described in this patent document should not be understood as required in all embodiments.

[0499] Only some embodiments and examples have been described, and other embodiments, enhancements, and variations can be made based on what is described and illustrated in this patent document.

Claims

1. A video processing method, comprising: Perform the conversion between a video picture of a video including one or more strips and a bitstream of the video picture according to rules, wherein the rules specify whether a first syntax element indicating the use of a codec tool exists at a first level based on a syntax flag indicating whether a syntax structure at a second level does not exist at the first level, wherein the second level is higher than the first level, and wherein the second level is a video picture level or higher than the video picture level.

2. The method according to claim 1, wherein, The first level includes a strip header, and wherein the second level includes a picture header.

3. The method according to claim 1, wherein, The rules further specify whether the first syntax element exists at the first level is also based on the value of a second syntax element indicating the use of the codec tool at the second level.

4. The method according to claim 3, wherein, The use of the codec tool at the first level is based on (1) a syntax flag indicating whether a syntax structure at the second level does not exist at the first level, and (2) the value of the second syntax element.

5. The method according to claim 1, wherein, The codec tool includes an explicit scaling list.

6. The method according to claim 5, wherein, The explicit scaling list is used in the scaling process of transform coefficients.

7. The method according to claim 5, wherein, In response to (1) the second syntax element in the picture header indicating that the explicit scaling list is enabled for at least one strip, and (2) the picture header syntax structure does not exist in the strip header, there is a first syntax element indicating the use of the explicit scaling list.

8. The method according to claim 5, wherein, In response to (1) the second syntax element in the picture header indicating that the explicit scaling list is disabled, or (2) the picture header syntax structure exists in the strip header, the first syntax element indicating the use of the explicit scaling list is omitted.

9. The method according to claim 8, wherein, In the case where the use of the explicit scaling list is omitted in the strip header, the use is inferred according to the second syntax element indicated in the picture header.

10. The method according to claim 1, wherein, Whether the video picture includes a single strip is indicated by the syntax flag.

11. The method according to claim 1, wherein, The rules further specify that a syntax element in the sequence parameter set of the video indicates whether to enable a luminance mapping with chroma scaling (LMCS) tool for a codec layer video sequence (CLVS) that refers to the sequence parameter set.

12. The method according to claim 11, wherein, The syntax element equal to 1 specifies that the LMCS tool is enabled for the CLVS, and wherein the syntax element equal to 0 specifies that the LMCS tool is disabled for the CLVS.

13. The method according to claim 1, wherein, The rules further specify that a syntax element in the sequence parameter set of the video indicates whether to enable a sample adaptive offset (SAO) tool for a codec layer video sequence (CLVS) that refers to the sequence parameter set.

14. The method according to claim 13, wherein, The syntax element equal to 1 specifies that the SAO tool is enabled for the CLVS, and wherein the syntax element equal to 0 specifies that the SAO tool is disabled for the CLVS.

15. The method according to claim 1, wherein, The rules further specify that whether and / or how to indicate the use of a scaling tool is determined based on whether the video picture includes a single strip, wherein the use of the scaling tool includes whether to allow a luminance mapping with chroma scaling (LMCS) tool to be used for the conversion, and the use also includes the number of scaling mode types allowed for the conversion.

16. The method according to claim 15, wherein, The scaling mode type indicates whether to enable or disable the scaling tool for all or part of one or more strips of the video picture.

17. The method according to claim 15 or 16, wherein, The rule specifies that whether the video picture includes a single strip is indicated by whether the picture header syntax structure of the video picture exists in the strip header of the strip.

18. The method according to claim 15, wherein, The rule specifies that in the case where the video picture includes a single strip, only two scaling mode types are allowed for the conversion.

19. The method according to claim 18, wherein, The two scaling mode types are indicated by the values 0 or 1.

20. The method according to claim 1, wherein, The rule also specifies that whether and / or how the allowed strip types are indicated in the picture header and / or strip header is determined based on whether the video picture includes a single strip.

21. The method according to claim 20, wherein, The rule specifies that whether the video picture includes a single strip is indicated by whether the picture header syntax structure of the video picture exists in the strip header of the strip.

22. The method according to claim 20 or 21, wherein, In the case where the video picture includes a single strip, when the second syntax flag in the picture header specifies that one or more strips in the video are allowed to have a specific strip type, the first syntax flag in the picture header that specifies whether all strips of the video picture have the specific strip type is omitted.

23. The method according to claim 20 or 21, wherein, In the case where the video picture includes a single strip, when the second syntax flag in the picture header specifies that one or more strips allowing codec type B are allowed, the first syntax flag in the picture header that specifies whether all strips of the video picture have a specific strip type is omitted.

24. The method according to claim 20 or 21, wherein, In the case where the video picture includes a single strip, when the second syntax flag in the picture header specifies that all strips of the video picture have a specific strip type, the first syntax flag in the picture header that specifies whether one or more strips allowing codec type B are allowed is omitted.

25. The method according to claim 20 or 21, wherein, Omit and infer the type of the strips in the video picture.

26. The method according to claim 1, wherein, The rule also specifies indicating the mode type of the codec tool at the video unit level for the conversion.

27. The method according to claim 26, wherein, The mode type includes: (1) the first type, indicating that the codec tool is enabled for all sub-units of the video unit; (2) the second type, indicating that the codec tool is disabled for all sub-units of the video unit; (3) the third type, indicating that the codec tool is enabled for at least one sub-unit rather than all sub-units of the video unit.

28. The method according to claim 26 or 27, wherein, The video unit includes a picture.

29. The method according to claim 28, wherein, The sub-unit includes a strip, slice or sub-picture of the picture.

30. The method according to claim 26, wherein, The mode type is indicated in the picture header or picture parameter set.

31. The method according to claim 26, wherein, The number of allowed mode types is determined based on whether the video picture includes a single strip.

32. The method according to claim 31, wherein, In the case where the video picture includes a single strip, only two mode types are allowed.

33. The method according to claim 26, wherein, How to indicate the mode type is determined based on whether the video picture includes a single strip.

34. The method according to claim 33, wherein, In the case where the video picture includes a single strip, the mode type is limited to a subset of the multiple allowed mode types.

35. The method according to claim 26, wherein, The rule specifies that whether the video picture includes a single strip is indicated by whether the picture header syntax structure of the video picture exists in the strip header of the strip.

36. The method according to claim 1, wherein, The rule also specifies that non-binary syntax elements or multiple syntax flags at the first video unit level are used to indicate the use of the codec tool at the second video unit level lower than the first video unit level.

37. The method according to claim 36, wherein, The first video unit level includes a picture level, and the second video unit level includes a slice level.

38. The method according to claim 36 or 37, wherein, The non-binary syntax element is encoded and decoded using at least more than two bits, where the non-binary syntax element equal to X specifies enabling the codec tool for all slices associated with the video picture, the non-binary syntax element equal to Y specifies enabling the codec tool for at least one but not all slices associated with the video picture, and the non-binary syntax element equal to Z specifies disabling the codec tool for all slices associated with the video picture, where X != Y, Y != Z, and X != Z.

39. The method according to claim 38, wherein, In the case where the non-binary syntax element is omitted, the non-binary syntax element is inferred to have a default value.

40. The method according to claim 38, wherein, The non-binary syntax element equal to Y indicates that the codec tool is applicable to the scaling process of transform and / or non-transform coefficients for the conversion.

41. The method according to claim 38, wherein, X = 1, Y = 2, and Z = 0.

42. The method according to claim 38, wherein, X = 2, Y = 1, and Z = 0.

43. The method according to claim 38, wherein, The non-binary syntax element is conditionally indicated based on a corresponding syntax flag at the sequence level.

44. The method according to claim 38, wherein, The non-binary syntax element is encoded and decoded as an unsigned integer, a left-most first unsigned integer of 0th order Exp-Golomb encoded syntax element, or a truncated unary value.

45. The method according to claim 38, wherein, Based on the non-binary syntax element, the corresponding adaptive parameter set information used by one or more slices of the video picture is indicated.

46. The method according to claim 36 or 37, wherein, The plurality of syntax flags includes a first syntax flag indicating whether the codec tool is enabled or disabled for at least one slice of the video picture, and a second syntax flag indicating whether the codec tool is enabled or disabled for all slices of the video picture.

47. The method according to claim 46, wherein, The second syntax flag is conditionally indicated based on the value of the first syntax flag.

48. The method according to claim 47, wherein, In the case where the second syntax flag is omitted, the second syntax flag is inferred to indicate enabling or disabling the codec tool for all slices.

49. The method according to claim 47, wherein, The rule specifies that in the case where the first syntax flag indicates enabling the codec tool for at least one slice of the video picture, the second syntax flag has a value indicating disabling the codec tool for all slices of the video picture.

50. The method according to claim 46, wherein, The plurality of syntax flags further includes a third syntax flag conditionally indicated in the slice header according to the first syntax flag or the second syntax flag.

51. The method according to claim 50, wherein, The third syntax flag is inferred based on the value of the first syntax flag and / or the second syntax flag.

52. The method according to claim 50, wherein, The third syntax flag is conditionally indicated according to a corresponding syntax flag at the sequence level.

53. The method according to claim 36 or 37, wherein, The plurality of syntax flags includes a first syntax flag indicating whether all slices of the video picture use the codec tool, and a second syntax flag indicating whether not all slices of the video picture use the codec tool.

54. The method according to claim 53, wherein, The second syntax flag is indicated only when the first syntax flag indicates that not all slices of the video picture use the codec tool.

55. The method according to claim 53, wherein, When the first syntax flag indicates that all slices of the video picture use the codec tool, or when the first syntax flag and the second syntax flag indicate that at least one slice of the video picture uses the codec tool, the scaling list adaptive parameter set identifier is included in the picture header.

56. The method according to claim 53, wherein, When the first syntax flag indicates that not all slices of the video picture use the codec tool and the second syntax flag indicates that at least one slice of the video picture uses the codec tool, a third syntax flag in the slice header is indicated.

57. The method according to claim 56, wherein, The value of the third syntax flag is based on the first syntax flag and / or the second syntax flag.

58. The method according to claim 36 or 37, wherein, The codec tool is associated with the use or information of the scaling list.

59. The method according to claim 58, wherein, When the non-binary syntax element or the plurality of syntax flags are present in the first video unit level, the use or information of the scaling list is omitted in the second video unit level.

60. The method according to claim 58, wherein, The plurality of syntax flags are indicated in the first video unit level, and the first video unit level includes a picture header, a picture parameter set, or a sequence parameter set.

61. The method according to claim 36, wherein, The syntax elements of the second video unit level are conditionally indicated based on the non-binary syntax element or the plurality of syntax flags.

62. The method according to claim 36, wherein, The indication of the non-binary syntax element or the plurality of syntax flags is determined based on information in a third video unit level higher than the first video unit level.

63. The method according to claim 62, wherein, The third video unit level includes a picture parameter set or a sequence parameter set.

64. The method according to claim 62 or 63, wherein,The syntax elements in the third video unit level indicate that the use of the codec tool is different within a picture set, a picture, or a slice.

65. The method according to claim 1, wherein, The conversion includes encoding the video into the bitstream.

66. The method according to claim 1, wherein, The conversion includes decoding the video from the bitstream.

67. A video decoding device, comprising a processor configured to implement the method according to any one of claims 11 to 66.

68. A video encoding device, comprising a processor configured to implement the method according to any one of claims 11 to 66.

69. A non-transitory computer-readable medium storing instructions, which when executed by a processor, cause the processor to implement the method according to any one of claims 11 to 66.

70. A device for processing video data, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to: perform a conversion between a video picture of a video including one or more stripes and a bitstream of the video picture according to rules, wherein, The rule specifies that whether a first syntax element indicating the use of the codec tool exists in a first level is based on a syntax flag indicating whether a syntax structure of a second level does not exist in the first level, where the second level is higher than the first level, and where the second level is a video picture level or higher than the video picture level.

71. A non-transitory computer-readable storage medium storing instructions, which cause a processor to: perform a conversion between a video picture of a video including one or more stripes and a bitstream of the video picture according to rules, wherein, The rule specifies that whether a first syntax element indicating the use of the codec tool exists in a first level is based on a syntax flag indicating whether a syntax structure of a second level does not exist in the first level, where the second level is higher than the first level, and where the second level is a video picture level or higher than the video picture level.

72. A non-transitory computer-readable recording medium storing a bitstream of a video picture including one or more stripes, the bitstream being generated by a method executed by a video processing device, wherein the method includes: Generate the bitstream of the video picture according to the rule, where the rule specifies that whether a first syntax element indicating the use of the codec tool exists in a first level is based on a syntax flag indicating whether a syntax structure of a second level does not exist in the first level, where the second level is higher than the first level, and where the second level is a video picture level or higher than the video picture level.

Citation Information

Patent Citations

  • Unified intra block copy and inter prediction modes

    CN105493505A

  • High Level syntax for video coding and decoding

    GB201919033D0