Advanced control of filtering in video coding
By using control information to process video encoding and decoder in video encoder and decoder, the inefficiency problem in the prior art is solved, and more efficient video processing and bandwidth optimization are achieved.
Patent Information
- Application Number
- CN202510326289.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2020-04-05
- Filing Date
- 2021-04-05
- Publication Date
- 2025-06-06
AI Technical Summary
When existing video encoding technology processes the encoding and decoding representation of video, it lacks effective control information processing, resulting in inefficiency and increased bandwidth usage.
By using control information useful for decoding of the codec representation in the video encoder and the decoder, rules including conversion between the video picture and the bitstream of the video picture are performed, for example, indicating whether to enable the brightness mapping with chromatic scaling or sample adaptive offset tool according to the syntax element.
Improves the efficiency of video encoding and decoding, optimizes bandwidth usage, and enhances the flexibility and control of video processing.
Smart Images

Figure CN120111220A_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application is a divisional application of Chinese invention patent application No. 202180026199.3 filed on September 29, 2022, which is the national phase of international patent application No. PCT / US2021 / 025734 filed on April 5, 2021, and is intended to timely claim priority and benefits of U.S. Provisional Application No. 63 / 005,413 filed on April 5, 2020. For all purposes prescribed by law, the entire disclosure of the aforementioned application is incorporated by reference as part of the disclosure of this patent. Technical Field
[0003] The patent document relates to image and video encoding and decoding. Background Art
[0004] Digital video consumes the largest amount of bandwidth on the Internet and other digital communications networks. As the number of connected user devices capable of receiving and displaying video increases, bandwidth requirements for digital video usage are expected to continue to grow. Summary of the invention
[0005] This document discloses techniques that may be used by a video encoder and decoder for processing a codec representation of a video using control information useful for decoding the codec representation.
[0006] In one example aspect, a video processing method is disclosed. The method includes: performing conversion between a video picture of a video including one or more slices and a bitstream of the video picture according to a rule. The rule specifies whether a first syntax element indicating use of a codec tool is present at a first level based on a syntax flag indicating whether a syntax structure of a second level is not present at the first level, wherein the second level is higher than the first level. The second level is at or above the video picture level.
[0007] In another example aspect, a video processing method is disclosed. The method includes performing conversion between a video and a bitstream of the video according to a rule. The rule specifies that a syntax element in a sequence parameter set of the video indicates whether a luma mapping with chroma scaling (LMCS) tool is enabled for a codec layer video sequence (CLVS) of a reference sequence parameter set.
[0008] In another example aspect, a video processing method is disclosed. The method includes performing conversion between a video and a bitstream of the video according to a rule. The rule specifies that a syntax element in a sequence parameter set of the video indicates whether a sample adaptive offset (SAO) tool is enabled for a codec layer video sequence (CLVS) with respect to a reference sequence parameter set.
[0009] In another example aspect, a video processing method is disclosed. The method includes performing conversion between a video picture of a video including one or more stripes and a bitstream of the video picture according to a rule. The rule specifies whether and / or how to indicate the use of a scaling tool is determined based on whether the video picture includes a single stripe, wherein the use of the scaling tool includes whether to allow the use of a luminance mapping with chroma scaling (LMCS) tool for the conversion, and also includes the number of scaling mode types allowed for the conversion.
[0010] In another example aspect, a video processing method is disclosed. The method includes performing conversion between a video picture of a video including one or more slices and a bitstream of the video picture according to a rule. The rule specifies whether and / or how to indicate in a picture header and / or a slice header that an allowed slice type is determined based on whether the video picture includes a single slice.
[0011] In another example aspect, a video processing method is disclosed. The method includes performing conversion between a video picture of a video including one or more slices and a bitstream of the video picture according to a rule. The rule specifies that a mode type of a codec tool is indicated at a video unit level for conversion.
[0012] In another example aspect, a video processing method is disclosed. The method includes performing conversion between a video picture of a video including one or more slices and a bitstream of the video picture according to a rule. The rule specifies that a non-binary syntax element or a plurality of syntax flags at a first video unit level is used to indicate use of a codec tool at a second video unit level lower than the first video unit level.
[0013] In another example aspect, a video processing method is disclosed. The method includes: performing a conversion between a video including one or more video pictures including one or more slices and a codec representation of the video; wherein the codec representation conforms to a format rule, wherein the format rule specifies which type of luma mapping with a chroma scaling mode or a scaling list mode type is applicable to the conversion of the slices is indicated by a picture header syntax structure in a slice header or a picture header in a picture including a single slice.
[0014] In another example aspect, another video processing method is disclosed. The method includes: performing conversion between a video including one or more video pictures including one or more slices and a codec representation of the video; wherein the codec representation conforms to a format rule, wherein the format rule specifies an indicator including an indication that enabling a luma mapping with chroma scaling (LMCS) mode at a first video level depends on a non-binary LMCS-related syntax element at a higher level and whether the picture consists of only one slice.
[0015] In another example aspect, another video processing method is disclosed. The method includes: performing conversion between a video including one or more video pictures including one or more slices and a codec representation of the video; wherein the codec representation conforms to a format rule, wherein the format rule specifies including an indicator indicating that enabling an explicit scaling list (ESL) mode at a first video level depends on a non-binary LMCS-related syntax element at a higher level and whether the picture consists of only one slice.
[0016] In another example aspect, another video processing method is disclosed. The method includes: performing conversion between a video including one or more pictures including one or more slices and a codec representation of the video, wherein the codec representation conforms to a format rule, wherein the format rule specifies whether a picture includes exactly one slice and a slice type or a slice type flag in a slice header controlling the exactly one slice.
[0017] In another example aspect, another video processing method is disclosed. The method includes: performing conversion between a video including one or more pictures including one or more video regions and a codec representation of the video, wherein the codec representation conforms to a format rule, wherein the format rule specifies two-level signaling including applicability of a filter codec tool (TX) to the video region.
[0018] In yet another example aspect, a video encoder apparatus is disclosed. The video encoder includes a processor configured to implement the above method.
[0019] In yet another example aspect, a video decoder apparatus is disclosed. The video decoder includes a processor configured to implement the above method.
[0020] In yet another example aspect, a computer readable medium having code stored thereon is disclosed. The code is in the form of processor executable code to implement one of the methods described herein.
[0021] These and other features are described throughout this document. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 An example of raster scan striping partitioning of a picture is shown, where the picture is divided into 12 tiles and 3 raster scan strips.
[0023] Figure 2 An example of rectangular strip partitioning of a picture is shown, where the picture is divided into 24 slices (6 slice columns and 4 slice rows) and 9 rectangular strips.
[0024] Figure 3An example of a picture partitioned into slices and rectangular strips is shown, where the picture is divided into 4 slices (2 slice columns and 2 slice rows) and 4 rectangular strips.
[0025] Figure 4 A picture partitioned into 15 slices, 24 slices and 24 sub-pictures is shown.
[0026] Figure 5 is a block diagram of an example video processing system.
[0027] Figure 6 is a block diagram of a video processing device.
[0028] Figure 7 is a flow chart of an example method of video processing.
[0029] Figure 8 is a block diagram illustrating a video encoding and decoding system according to some embodiments of the present disclosure.
[0030] Fig. 9 is a block diagram illustrating an encoder according to some embodiments of the present disclosure.
[0031] Fig.10 is a block diagram illustrating a decoder according to some embodiments of the present disclosure.
[0032] Fig.11 Examples of ALF filter shapes are shown (chroma: 5×5 diamond, luma: 7×7 diamond).
[0033] Fig.12 Examples of ALF and CC-ALF are shown.
[0034] Fig.13 is a flowchart representation of a method of video processing according to the present technology.
[0035] Fig.14 is a flowchart representation of another method for video processing according to the present technology.
[0036] Fig.15 is a flowchart representation of another method for video processing according to the present technology.
[0037] Fig.16 is a flowchart representation of another method for video processing according to the present technology.
[0038] Fig.17 is a flowchart representation of another method of video processing according to the present technology.
[0039] Fig.18 is a flowchart representation of another method for video processing according to the present technology.
[0040] Fig.19 is a flowchart representation of another method for video processing according to the present technology. DETAILED DESCRIPTION
[0041] The section headings used in this document are for ease of understanding and do not limit the applicability of the techniques and embodiments disclosed in each section to only that section. In addition, the use of H.266 terminology in some descriptions is merely for ease of understanding and is not intended to limit the scope of the disclosed technology. Therefore, the techniques described herein are also applicable to other video codec protocols and designs. In this document, editorial changes are displayed in the text relative to the current draft of the VVC specification by indicating deleted text with a strikethrough, highlighting indicating added text (including bold italics).
[0042] 1. Overview
[0043] This document relates to video codec technology. Specifically, it is about improvements in signaling of adaptive loop filter (ALF), sample adaptive offset (SAO), luma mapping with chroma scaling (LMCS), scaling lists. These concepts can be applied alone or in various combinations to any video codec standard or non-standard video codec that supports multi-layer video codec, such as the Versatile Video Codec (VVC) under development.
[0044] 2. Abbreviations
[0045] ALF adaptive loop filter
[0046] APS Adaptive Parameter Set
[0047] AU Access Unit
[0048] AUD access unit delimiter
[0049] AVC advanced video codec
[0050] CLVS codec layer video sequence
[0051] CPB Coded Picture Buffer
[0052] CRA Clean Random Access
[0053] CTU codec tree unit
[0054] CVS codec video sequence
[0055] DCI decoding capability information
[0056] DPB decoded picture buffer
[0057] EOB end of bitstream
[0058] EOS sequence ends
[0059] GDR gradually decodes and refreshes
[0060] HEVC high-efficiency video codec
[0061] HRD Hypothetical Reference Decoder
[0062] IDR instant decoding refresh
[0063] JEM Joint Exploration Mode
[0064] LMCS Luminance Mapping with Chroma Scaling
[0065] MCTS Motion Constraint Episodes
[0066] NAL Network Abstraction Layer
[0067] OLS output layer set
[0068] PH Image Header
[0069] PPS Picture Parameter Set
[0070] PTL profiles, tiers and levels
[0071] PU picture unit
[0072] RADL random access decodable boot (picture)
[0073] RAP Random Access Point
[0074] RASL random access skip boot (image)
[0075] RBSP raw byte sequence payload
[0076] RPL Reference Image List
[0077] SAO Sample Adaptive Offset
[0078] SEI Supplemental Enhancement Information
[0079] SPS sequence parameter set
[0080] STSA Step-by-step temporal sublayer access
[0081] SVC Scalable Video Codec
[0082] VCL video codec layer
[0083] VPS Video Parameter Set
[0084] VTM VVC test model
[0085] VUI Video Availability Information
[0086] VVC multi-function video codec
[0087] 3.1. Preliminary discussion
[0088] Video coding standards have evolved primarily through the development of the well-known ITU-T and ISO / IEC standards. ITU-T developed H.261 and H.263, ISO / IEC developed MPEG-1 and MPEG-4 Visual, and the two organizations jointly developed the H.262 / MPEG-2 Video and H.264 / MPEG-4 Advanced Video Coding (AVC) and H.265 / HEVC standards. Since H.262, video codec standards have been based on a hybrid video codec structure, where temporal prediction plus transform codec is utilized. In order to explore future video codec technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Team (JVET) in 2015. Since then, JVET has adopted many new methods and input them into a reference software called the Joint Exploration Model (JEM). JVET meetings are also held quarterly, and the goal for the new codec standard is to reduce the bit rate by 50% compared to HEVC. The new video coding standard was officially named Versatile Video Codec (VVC) at the JVET meeting in April 2018, and the first version of the VVC Test Model (VTM) was released at that time. Due to the continuous efforts on VVC standardization, new codec technologies are adopted into the VVC standard at each JVET meeting. The working draft of VVC and the test model VTM are updated after each meeting. The VVC project now aims to be technically completed (FDIS) at the July 2020 meeting.
[0089] 3.1.HEVC Image Segmentation Scheme
[0090] HEVC includes four different picture partitioning schemes, namely, regular slices, dependent slices, tiles, and wavefront parallel processing (WPP), which can be applied to maximum transmission unit (MTU) size matching, parallel processing, and reduced end-to-end delay.
[0091] Regular slices are similar to those in H.264 / AVC. Each regular slice is encapsulated in its own NAL unit, and intra-picture prediction (intra-sample prediction, motion information prediction, coding mode prediction) and entropy coding dependencies across slice boundaries are disabled. Therefore, regular slices can be reconstructed independently of other regular slices within the same picture (although there may still be interdependencies due to loop filtering operations).
[0092] Regular slices are the only tool that can be used for parallelization, which is also available in H.264 / AVC in almost the same form. Parallelization based on regular slices does not require much inter-processor or inter-core communication (except for inter-processor or inter-core data sharing for motion compensation when decoding predictively coded pictures, which is usually much heavier than inter-processor or inter-core data sharing due to intra-picture prediction). However, for the same reason, the use of regular slices incurs a large codec overhead due to the bit cost of the slice header and the lack of prediction across slice boundaries. In addition, due to the intra-picture independence of regular slices and the fact that each regular slice is encapsulated in its own NAL unit, regular slices (in contrast to the other tools mentioned below) also serve as a key mechanism for bitstream segmentation to match MTU size requirements. In many cases, the goals of parallelization and MTU size matching place conflicting requirements on the layout of slices within a picture. Recognition of this situation led to the development of the parallelization tools mentioned below.
[0093] Dependent slices have short slice headers and allow the bitstream to be split at treeblock boundaries without breaking any intra-picture prediction. Basically, dependent slices provide fragmentation of regular slices into multiple NAL units to provide reduced end-to-end delay by allowing part of a regular slice to be sent out before the encoding of the entire regular slice is completed.
[0094] In WPP, a picture is partitioned into a single row of codec tree blocks (CTBs). Entropy decoding and prediction are allowed to use data from CTBs in other partitions. Parallel processing is possible through parallel decoding of CTB rows, where the start of decoding of CTB rows is delayed by two CTBs, ensuring that data related to the CTB above and to the right of the subject CTB is available before the subject CTB is decoded. Using this staggered start (when represented graphically, it looks like a wavefront), parallelization can use as many processors / cores as the picture contains CTB rows. Because intra-picture prediction between adjacent tree block rows within a picture is allowed, the inter-processor / inter-core communication required to implement intra-picture prediction can be substantial. WPP partitioning does not generate additional NAL units compared to when it is not applied, so WPP is not a tool for MTU size matching. However, if MTU size matching is required, regular stripes can be used with WPP, but there will be some codec overhead.
[0095] Slices define the horizontal and vertical boundaries that divide the image into slice columns and rows. Slice columns extend from the top of the image to the bottom of the image. Similarly, slice rows extend from the left side of the image to the right side of the image. The number of slices in an image can be simply derived by multiplying the number of slice columns by the number of slice rows.
[0096] Before decoding the top left CTB of the next slice in the order of the slice raster scan of the picture, the scan order of the CTBs is changed to be local within the slice (in the order of the CTB raster scan of the slice). Similar to regular slices, slices break intra-picture prediction dependencies and entropy decoding dependencies. However, they do not need to be included in a single NAL unit (the same as WPP in this respect); therefore, slices cannot be used for MTU size matching. Each slice can be processed by one processor / core, and the inter-processor / inter-core communication required for intra-picture prediction between processing units decoding adjacent slices is limited to transmitting a shared slice header when the slice spans more than one slice, and loop filtering associated with sharing of reconstructed samples and metadata. When more than one slice or WPP segment is included in a slice, the entry point byte offset of each slice or WPP segment in the slice except the first slice or WPP segment is signaled in the slice header.
[0097] For simplicity, HEVC specifies restrictions on the application of four different picture partitioning schemes. A given codec video sequence cannot include both slices and wavefronts from most of the profiles specified in the HEVC standard. For each slice and slice, one or both of the following conditions must be met: 1) all coding tree blocks in a slice belong to the same slice; 2) all coding tree blocks in a slice belong to the same slice. Finally, a wavefront segment contains exactly one CTB row, and when using WPP, if a slice starts on a CTB row, it must end on the same CTB row.
[0098] In some embodiments, HEVC specifies three MCTS-related SEI messages, namely, a temporal MCTS SEI message, an MCTS extraction information set SEI message, and an MCTS extraction information nesting SEI message.
[0099] The temporal MCTS SEI message indicates the presence of MCTS in the bitstream and signals the MCTS. For each MCTS, motion vectors are restricted to point to full sample positions within the MCTS and fractional sample positions that only require full sample positions within the MCTS for interpolation, and motion vector candidates derived from blocks outside the MCTS for temporal motion vector prediction are not allowed. In this way, each MCTS can be decoded independently without the presence of slices not included in the MCTS.
[0100] The MCTS extraction information set SEI message provides supplementary information (specified as part of the SEI message semantics) that can be used in MCTS sub-bitstream extraction to generate a bitstream that conforms to the MCTS set. The information consists of multiple extraction information sets, each of which defines multiple MCTS sets and contains RBSP bytes that replace VPS, SPS, and PPS to be used in the MCTS sub-bitstream extraction process. When extracting a sub-bitstream according to the MCTS sub-bitstream extraction process, the parameter sets (VPS, SPS, and PPS) need to be rewritten or replaced, and the slice header needs to be slightly updated because one or all slice address-related syntax elements (including first_slice_segment_in_pic_flag and slice_segment_address) usually need to have different values.
[0101] 3.2. Segmentation of VVC Images
[0102] In VVC, a picture is divided into one or more slice rows and one or more slice columns. A slice is a sequence of CTUs covering a rectangular area of a picture. The CTUs in a slice are scanned in the slice in raster scan order.
[0103] A slice consists of an integer number of complete slices or an integer number of consecutive complete CTU rows in a slice of a picture.
[0104] Two stripe modes are supported, namely, raster scan stripe mode and rectangular stripe mode. In raster scan stripe mode, a stripe contains a complete sequence of stripes in a stripe raster scan of a picture. In rectangular stripe mode, a stripe contains multiple complete slices that together form a rectangular area of a picture, or multiple consecutive complete CTU rows that together form a slice of a rectangular area of a picture. The slices within a rectangular stripe are scanned in a slice raster scan order within the rectangular area corresponding to the stripe.
[0105] A sub-picture consists of one or more strips that together cover a rectangular area of the picture.
[0106] Figure 1 An example of raster scan striping partitioning of a picture is shown, where the picture is divided into 12 tiles and 3 raster scan strips.
[0107] Figure 2 An example of rectangular slice partitioning of a picture is shown, where the picture is divided into 24 slices (6 slice columns and 4 slice rows) and 9 rectangular slices.
[0108] Figure 3 An example of a picture partitioned into slices and rectangular strips is shown, where the picture is divided into 4 slices (2 slice columns and 2 slice rows) and 4 rectangular strips.
[0109] Figure 4 An example of sub-picture partitioning of a picture is shown, where the picture is partitioned into 18 slices, the 12 on the left hand side each covering one stripe of a 4×4 CTU, and the 6 slices on the right hand side each covering 2 vertically stacked stripes of a 2×2 CTU, resulting in a total of 24 stripes and 24 sub-pictures of different dimensions (each stripe is a sub-picture). 3.3. Change of picture resolution within a sequence
[0110] In AVC and HEVC, the spatial resolution of a picture cannot be changed unless a new sequence using a new SPS starts with an IRAP picture. VVC allows changing the resolution of pictures within a sequence without encoding an IRAP picture, which is always intra-coded. This feature is sometimes called reference picture resampling (RPR) because it requires resampling the reference pictures used for inter prediction when the reference picture has a different resolution than the current picture being decoded.
[0111] The scaling ratio is restricted to be greater than or equal to 1 / 2 (2x downsampling from the reference picture to the current picture) and less than or equal to 8 (8x upsampling). Three resampling filter sets with different frequency cutoffs are specified to handle various scaling ratios between the reference picture and the current picture. The three resampling filter sets are applied to scaling ratios ranging from 1 / 2 to 1 / 1.75, from 1 / 1.75 to 1 / 1.25, and from 1 / 1.25 to 8, respectively. Each set of resampling filters has 16 phases for luma and 32 phases for chroma, which is the same as the case of motion compensated interpolation filters. In fact, the normal MC interpolation process is a special case of the resampling process, where the scaling ratio ranges from 1 / 1.25 to 8. The horizontal and vertical scaling ratios are derived based on the picture width and height and the left, right, top, and bottom scaling offsets specified for the reference picture and the current picture.
[0112] Other aspects of the VVC design that support this feature that differ from HEVC include: i) the picture resolution and corresponding consistency window are signaled in the PPS instead of the SPS, where the maximum picture resolution is signaled. ii) For a single-layer bitstream, each picture store (a slot in the DPB used to store one decoded picture) occupies the buffer size required to store a decoded picture with the maximum picture resolution.
[0113] 3.4. Reference Picture Management and Reference Picture List (RPL)
[0114] Reference picture management is a core functionality required for any video codec using inter prediction. It manages the storage and removal of reference pictures in the decoded picture buffer (DPB) and places the reference pictures in the RPL in the correct order.
[0115] The reference picture management of HEVC, including reference picture marking and removal from the decoded picture buffer (DPB) and reference picture list construction (RPLC), is different from that of AVC. Instead of the reference picture marking mechanism based on sliding window plus adaptive memory management control operation (MMCO) in AVC, HEVC specifies a reference picture management and marking mechanism based on the so-called reference picture set (RPS), and therefore RPLC is based on the RPS mechanism. The RPS consists of a reference picture set associated with a picture (consisting of all reference pictures before the associated picture in decoding order), which can be used for inter-frame prediction of the associated picture or any picture after the associated picture in decoding order. The reference picture set consists of five reference picture lists. The first three lists contain all reference pictures that can be used for inter-frame prediction of the current picture and for inter-frame prediction of one or more pictures after the current picture in decoding order. The other two lists consist of all reference pictures that are not used for inter-frame prediction of the current picture, but can be used for inter-frame prediction of one or more pictures after the current picture in decoding order. RPS provides "intra-frame codec" signaling of DPB status, rather than "inter-frame codec" signaling as in AVC, mainly to improve error resistance. HEVC's RPLC process is based on RPS, by signaling the index to the RPS subset for each reference index; this process is simpler than the RPLC process in AVC.
[0116] VVC's reference picture management is more similar to HEVC than AVC, but simpler and more robust. As in those standards, two RPLs are derived, List 0 and List 1, but they are not based on the reference picture set concept used in HEVC or the automatic sliding window process used in AVC; instead, they are signaled more directly. Reference pictures used for RPLs are listed as active and inactive entries, and only active entries can be used as reference indexes for inter-frame prediction of CTUs of the current picture. Invalid entries indicate other pictures to be saved in the DPB for reference by other pictures that arrive later in the bitstream.
[0117] 3.5. Parameter Set
[0118] AVC, HEVC, and VVC specify parameter sets. Types of parameter sets include SPS, PPS, APS, and VPS. All of AVC, HEVC, and VVC support SPS and PPS. VPS was introduced from HEVC and is included in both HEVC and VVC. APS is not included in AVC or HEVC, but is included in the latest VVC draft text.
[0119] SPS is designed to carry sequence-level header information, and PPS is designed to carry picture-level header information that does not change frequently. With SPS and PPS, information that does not change frequently does not need to be repeated for each sequence or picture, so redundant signaling of that information can be avoided. In addition, the use of SPS and PPS enables out-of-band transmission of important header information, which not only avoids the need for redundant transmission but also improves error resilience.
[0120] VPS is introduced to carry sequence-level header information common to all layers in a multi-layer bitstream.
[0121] APS is introduced to carry such picture-level or slice-level information, which requires quite a lot of bits to encode and decode, can be shared by multiple pictures, and can have quite a lot of different variations in a sequence.
[0122] 3.6.VVC slice header and picture header
[0123] Similar to the case of HEVC, the slice header of VVC conveys information about a specific slice, including slice address, slice type, slice QP, picture order count (POC) least significant bit (LSB), RPS and RPL information, weighted prediction parameters, loop filter parameters, entry offset of slices and WPP, etc.
[0124] VVC introduces a picture header (PH), which contains header parameters for a specific picture. Each picture must have one or only one PH. The PH basically carries those parameters that are in the slice header if the PH is not introduced, but each parameter has the same value for all slices of the picture. These include IRAP / GDR picture indication, inter-slice / intra-slice permission flags, POC LSB and optional POC MSB, information about RPL, deblocking, SAO, ALF, QP delta and weighted prediction, codec block partitioning information, virtual boundaries, co-located picture information, etc. It is often the case that each picture in the entire picture sequence contains only one slice. In this case, in order to allow each picture not to have at least two NAL units, the PH syntax structure is allowed to be included in the PH NAL unit or in the slice header.
[0125] In VVC, information of co-located pictures used for temporal motion vector prediction is signaled in the picture header or slice header.
[0126] 3.7. Loop Filter
[0127] In VVC, deblocking filter, SAO, and ALF are supported as loop filtering methods.
[0128] 3.7.1.SAO
[0129] The same design as HEVC is used, where Sample Adaptive Offset (SAO) is called after deblocking filtering and before ALF if needed. The key concept of SAO is to reduce sample distortion by first classifying the reconstructed samples into different categories, obtaining an offset for each category, and then adding the offset to each sample of that category. The offset for each category is appropriately calculated at the encoder and explicitly signaled to the decoder to effectively reduce sample distortion, while the classification of each sample is performed at both the encoder and decoder to significantly save side information. In order to achieve low latency with only one codec tree unit (CTU), a CTU-based syntax design is specified to adapt the SAO parameters to each CTU.
[0130] 3.7.2. Adaptive Loop Filter
[0131] Two diamond filter shapes are used in block-based ALF (e.g. Fig.11 ). The 7×7 diamonds are applied to the luma component and the 5×5 diamonds are applied to the chroma components. One of up to 25 filters is selected for each 4×4 block based on the direction and activity of the local gradient. Each 4×4 block in the picture is classified based on directionality and activity. Before filtering each 4×4 block, simple geometric transformations such as rotations or diagonal and vertical flips can be applied to the filter coefficients, depending on the gradient values calculated for the block. This is equivalent to applying these transformations to the samples in the filter support area. The idea is to make different blocks more similar by aligning the directionality of different blocks to which the ALF is applied. Block-based classification is not applicable to chroma components.
[0132] ALF filter parameters are signaled in an adaptive parameter set (APS). In one APS, up to 25 luma filter coefficients and clipping value index sets, and up to 8 chroma filter coefficients and clipping value index sets can be signaled. In order to reduce bit overhead, filter coefficients of different classifications of luma components can be merged. In the picture or slice header, up to 7 APS IDs can be signaled to specify the luma filter set for the current picture or slice. The filtering process is further controlled at the CTB level. The luma CTB can select a filter set from 16 fixed filter sets and the filter set signaled in the APS. For chroma components, the APS ID is signaled in the picture or slice header to indicate the chroma filter set used for the current picture or slice. At the CTB level, if there is more than one chroma filter set in the APS, the filter index is signaled for each chroma CTB. When ALF is enabled for a CTB, for each sample within the CTB, a diamond filter with signaled weights is performed and a clipping operation is applied to capture the difference between neighboring samples and the current sample. The clipping operation introduces non-linearity to make the ALF more effective by reducing the impact of neighboring sample values that differ too much from the current sample value.
[0133] The cross-component adaptive loop filter (CC-ALF) can further enhance each chroma component on top of the previously described ALF. The goal of CC-ALF is to refine each chroma component using luma sample values. This is achieved by applying a diamond high-pass linear filter and then using the output of this filtering operation for chroma refinement. Fig.12 A system-level schematic diagram of the CC-ALF process relative to other loop filters is provided. Fig.12 As shown, CC-ALF uses the same input as the luma ALF to avoid an extra step in the overall loop filtering process.
[0134] 3.7.3.ALF / SAO Signaling
[0135] In VVC draft 8, ALF and SAO share the same high-level control scheme. Both codecs can be controlled at the sequence level and one of the picture level or slice level (but not at both the picture level and the slice level). First, the SPS enable flag is signaled to control ALF / SAO at the CLVS level. At the PPS level, the PPS flag is signaled to indicate whether ALF / SAO is further controlled at the picture level or the slice level. If the PPS flag indicates that ALF / SAO is further controlled in the picture level, the PH ALF / SAO enable flag is signaled, followed by the ALF parameters (if it is enabled); if the PPS flag indicates that ALF / SAO is further controlled at the slice level, the SH ALF / SAO enable flag is signaled, followed by the ALF parameters (if it is enabled).
[0136] Table 1 ALF syntax in SPS
[0137]
[0138] Table 2 ALF syntax in PPS
[0139]
[0140] Table 3: ALF syntax in picture header
[0141]
[0142]
[0143] Table 4: ALF syntax in slice header
[0144]
[0145]
[0146] Table 5 SAO syntax in SPS
[0147]
[0148] Table 6: SAO syntax in PPS
[0149]
[0150] Table 7: SAO syntax in picture header
[0151]
[0152] Table 8: SAO syntax in slice header
[0153]
[0154]
[0155] Table 9: LMCS syntax in SPS
[0156]
[0157] Table 10: LMCS syntax in picture header
[0158]
[0159] Table 11: LMCS syntax in stripe header
[0160]
[0161] sps_lmcs_enabled_flag equal to 0 specifies that luma mapping with chroma scaling is not used in CLVS.
[0162] sps_sao_enabled_flag equal to 0 specifies that the sample adaptive offset process is not applied to the reconstructed picture after the deblocking filtering process.
[0163] 3.8. Luminance Mapping with Chroma Scaling (LMCS)
[0164] Unlike other loop filters (i.e., deblocking filters, SAO filters, and ALF filters) that typically apply a filtering process to the current sample using information from its spatial neighbors to reduce coding artifacts, luma mapping with chroma scaling (LMCS) improves compression efficiency by redistributing codewords across the entire dynamic range to modify the input signal before encoding. LMCS has two main components: (a) loop mapping of the luma component based on an adaptive piecewise linear model, and (b) luma-dependent chroma residual scaling for chroma components. Luma mapping utilizes a forward mapping function FwdMap and a corresponding inverse mapping function InvMap. The FwdMap function is signaled using a piecewise linear model with 16 equal pieces. The InvMap function does not need to be signaled, but is derived from the FwdMap function. The luma mapping model is signaled in the APS. Up to 4 LMCS APSs can be used in a coded video sequence. When LMCS is enabled for a picture, the APS ID is signaled in the picture header to identify the APS carrying the luma mapping parameters. When LMCS is enabled for a slice, the InvMap function is applied to all reconstructed luma blocks to convert the samples back to the original domain. For inter-coded blocks, an additional mapping process is required, which applies the FwdMap function to map the luma prediction block in the original domain to the mapped domain after the normal compensation process. Chroma residual scaling is designed to compensate for the interaction between the luma signal and its corresponding chroma signal. When luma mapping is enabled, an additional flag is signaled indicating whether luma-dependent chroma residual scaling is enabled. The chroma residual scaling factor depends on the average of the neighboring luma samples reconstructed at the top and / or left of the current CU. Once the scaling factor is determined, forward scaling is applied to intra and inter prediction residuals during the encoding stage, and inverse scaling is applied to the reconstructed residuals.
[0165] 3.8.1. Signaling of LMCS-related syntax elements
[0166] In the current VVC specification, LMCS control can be signaled in SPS, PH, and SH. First, the SPS enable flag controls LMCS at the CLVS level. If the SPS enable flag is equal to 1, the PH enable flag is further signaled to control LMCS at the picture level, and if it is enabled at the picture level, LMCS parameter information is also signaled in the PH. If the PH enable flag is equal to 1, the SH enable flag is further signaled to control LMCS at the slice level, but even if it is enabled at the slice level, LMCS parameter information cannot be signaled in the SH.
[0167] The relevant syntax elements and semantics are as follows:
[0168] 7.3.2.7 Picture header structure syntax
[0169]
[0170]
[0171] 7.3.7.1 Generic Strip Header Syntax
[0172]
[0173] Equal to 1 specifies that luma mapping with chroma scaling is enabled for all slices associated with the PH. ph_lmcs_enabled_flag equal to 0 specifies that luma mapping with chroma scaling is disabled for one, more than one, or all slices associated with the PH. If not present, the value of ph_lmcs_enabled_flag is inferred to be equal to 0.
[0174] =1 specifies that chroma residual scaling is enabled for all slices associated with the PH. ph_chroma_residual_scale_flag equal to 0 specifies that chroma residual scaling may be disabled for one, multiple or all slices associated with the PH. When ph_chroma_residual_scale_flag is not present, it is inferred to be equal to 0.
[0175] slice_lmcs_enabled_flag is equal to 0. When slice_lmcs_enabled_flag is not present, it is inferred to be equal to 0.
[0176] 3.9. Explicit Scaling Lists
[0177] Explicit signaling of scaling lists is defined in APS. And for each picture, whether explicit signaling is used is first signaled as a flag in PH, followed by the APS index if necessary. In the slice header, when the PH flag tells to use explicit scaling lists, each slice further signals a flag indicating whether the current slice uses explicit signaling.
[0178] In the current VVC draft text, the text most relevant to the scaling list is as follows:
[0179] Sequence Parameter Set RBSP Syntax and Semantics
[0180]
[0181] ....
[0183] sps_scaling_list_enabled_flag equal to 0 specifies that the scaling list is not used for the scaling process of transform coefficients. ...
[0185] Image header structure syntax and semantics
[0186] ...
[0188] equal to 1 specifies that the scaling list data for the slices associated with the PH is derived based on the scaling list data contained in the reference scaling list APS. ph_scaling_list_present_flag equal to 0 specifies that the scaling list data for the slices associated with the PH is set equal to 16. When not present, the value of ph_scaling_list_present_flag is inferred to be equal to 0. When not present, the value of ph_scaling_list_present_flag is inferred to be equal to 0.
[0189] Specifies the adaptation_parameter_set_id of the scaling list APS. The Temporalld of the APS NAL unit with aps_params_type equal to SCALING_APS and adaptation_parameter_set_id equal to ph_scaling_list_aps_id shall be less than or equal to the Temporalld of the picture associated with the PH. ...
[0191] Common Strip Header Syntax and Semantics
[0192]
[0193] ...
[0195] equal to 1 specifies that the scaling list data for the current slice is derived based on the scaling list data contained in the reference scaling list APS with aps_params_type equal to SCALING_APS and adaptation_parameter_set_id equal to ph_scaling_list_aps_id. slice_scaling_list_present_flag equal to 0 specifies that the scaling list data for the current picture is the derived default scaling list data specified in clause 7.4.3.21. When not present, the value of slice_scaling_list_present_flag is inferred to be equal to 0. ...
[0197] Scaling process of transform coefficients
[0198] To derive the scaled transform coefficients d[x][y], where x = 0..nTbW-1, y = 0..nTbH-1, the following applies:
[0199] – The intermediate scaling factor m[x][y] is derived as follows:
[0200] – m[x][y] is set equal to 16 if one or more of the following conditions are true:
[0201] –sps_scaling_list_enabled_flag is equal to 0.
[0202] –ph_scaling_list_present_flag is equal to 0.
[0203] –transform_skip_flag[xTbY][yTbY][cIdx] is equal to 1.
[0204] –scaling_matrix_for_lfnst_disabled_flag is equal to 1, and ApplyLfnstFlag is equal to 1.
[0205] –... ...
[0207] 7.3.2.5 Adaptation Parameter Set RBSP Syntax
[0208]
[0209]
[0210] 7.3.2.21 Scaling List Data Syntax
[0211]
[0212]
[0213] equal to 1 specifies that the scaling matrix should not be applied to blocks coded with LFNST. scaling_matrix_for_lfnst_disabled_flag equal to 0 specifies that the scaling matrix can be applied to blocks coded with LFNST.
[0214] 3.10. Recent Advances in LMCS and Explicit Scaling Lists
[0215] To address all of the above issues, it is proposed to replace the PH flag ph_lmcs_enabled_flag with a 2-bit ph_lmcs_mode_idc, and specify 3 modes: disabled (mode 0), for all slices (mode 1), and enabled (mode 2). In mode 1, LMCS is used for all slices of the picture, and signaling of the LMCS control flag is not required in the SH. The semantics of the SH LMCS control flag are modified accordingly. In addition, a modification of the semantics of ph_chroma_residual_scale_flag is proposed to reflect the intention to enable / disable chroma residual scaling for a picture or slice.
[0216] Below are some proposed changes to the grammatical structure. Most relevant additions or changes are marked with bold, italic, and underlined, and some deletions are indicated with [[ ]].
[0217] 7.3.2.7 Picture header structure syntax
[0218]
[0219] 7.3.7.1 Generic Strip Header Syntax
[0220]
[0221]
[0222] equal to 1 specifies that luma mapping with chroma scaling is applied to all stripes associated with the PH.
[0223] A value of 0 specifies that the PH associated with [[may be disabled for one or more stripes, or]] all stripes Disable Luma mapping with chroma scaling. When it does not exist, The value of is inferred to be equal to 0.
[0224] Equal to 1 specifies the pair All associated slices have chroma residual scaling enabled, and ph_chroma_residual_scale_flag equal to 0 specifies that the residual scale flag associated with PH [[may be disabled for one or more slices, or]] for all slices Chroma residual scaling. When ph_chroma_residual_scale_flag is not present, it is inferred to be equal to 0. ...
[0226] A value equal to 1 specifies that a luma mapping with chroma scaling is applied to the current stripe. A value of 0 specifies chroma scaling. The brightness mapping of is not applicable to the current strip. When absent, it is inferred to be equal to ...
[0228] To address several issues, the following modifications are proposed:
[0229] (1) The PH flag ph_explicit_scaling_list_enabled_flag is replaced with a 2-bit ph_explicit_scaling_list_mode_idc, specifying 3 modes: disabled (mode 0), used for all slices (mode 1), and enabled (mode 2). In mode 1, the explicit scaling list is used for all slices of the picture, and scaling list signaling is not required in SH.
[0230] (2) Move the flag scaling_matrix_for_lfnst_disabled_flag from the scaling_list_data() syntax to the SPS. 7.3.2.3 Sequence Parameter Set RBSP Syntax
[0231]
[0232]
[0233] 7.4.3.3 Sequence parameter set RBSP semantics
[0234] sps_lfnst_enabled_flag is equal to 0 and specifies that lfnst_idx may be present in the syntax of an intra codec unit.
[0235] sps_explicit_scaling_list_enabled_flag equal to 0 specifies that the use of explicit scaling lists signaled in the scaling list APS is enabled for CLVS in the scaling process of transform coefficients when decoding slices. sps_explicit_scaling_list_enabled_flag equal to 0 specifies that the use of explicit scaling lists in the scaling process of transform coefficients when decoding slices is disabled for CLVS.
[0236]
[0237] 7.3.2.7 Picture header structure syntax
[0238]
[0239] 7.4.3.7 Image header structure semantics
[0240] equal to 1 specifies that the use of explicit scaling lists signaled in the reference scaling list APS (i.e., the APS with aps_params_type equal to SCALING_APS and adaptation_parameter_set_id equal to ph_explicit_scaling_list_aps_id) is enabled for the picture during scaling of transform coefficients when decoding slices. ph_explicit_scaling_list_enabled_flag equal to 0 specifies that the use of explicit scaling lists in the scaling of transform coefficients is enabled for the picture when decoding slices. When not present, the value of ph_explicit_scaling_list_enabled_flag is inferred to be equal to 0]]
[0241]
[0242] Specifies the adaptation_parameter_set_id of the scaling list APS. The Temporalld of the APS NAL unit with aps_params_type equal to SCALING_APS and adaptation_parameter_set_id equal to ph_scaling_list_aps_id shall be less than or equal to the Temporalld of the picture associated with the PH.
[0243] 7.3.2.21 Scaling List Data Syntax
[0244]
[0245] 7.4.3.21 Scaling List Data Semantics
[0246] equal to 1 specifies that the scaling matrix is not applied to blocks decoded with LFNST. scaling_matrix_for_lfnst_disabled_flag equal to 0 specifies that the scaling matrix can be applied to blocks decoded with LFNST]]
[0247] equal to 1 specifies that the chroma scaling list is present in scaling_list_data(). scaling_list_chroma_present_flag equal to 0 specifies that the chroma scaling list is not present in scaling_list_data(). A bitstream conformance requirement is that when ChromaArrayType is equal to 0, scaling_list_chroma_present_flag shall be equal to 0 and when ChromaArrayType is not equal to 0, scaling_list_chroma_present_flag shall be equal to 1.
[0248] 7.3.7.1 Generic Strip Header Syntax
[0249]
[0250] 7.4.8.1 Generic Strip Header Semantics
[0251] equal to 1 specifies that the explicit scaling list signaled in the reference scaling list APS (where aps_params_type is equal to SCALING_APS and adaptation_parameter_set_id is equal to ph_scaling_list_aps_id) is used in the scaling process of transform coefficients when decoding the current slice. slice_explicit_scaling_list_used_flag equal to 0 specifies that no explicit scaling list is used in the scaling process of transform coefficients when decoding the current slice. When not present, the value of slice_explicit_scaling_list_used_flag is inferred to be equal to
[0252] 4. Examples of technical problems solved by the disclosed technical solutions
[0253] The existing designs and recent advances of ALF, SAO, scaling lists, and LMCS suffer from the following problems:
[0254] 1. The design of scaled lists / LMCS solves many problems in the state-of-the-art VVC text, however, further issues are identified:
[0255] a. If a picture contains only one slice, signaling of slice level control flags is not necessary.
[0256] b. If a picture contains only one slice, the allowed mode types (e.g., enabled for all slices;
[0257] disabled for all stripes; and enabled for at least one but not all stripes) can be reduced to two modes instead of three.
[0258] 2. SAO / ALF can be controlled in either PH or SH, but not both, which limits flexibility.
[0259] 3. The semantics of sps_lmcs_enabled_flag and sps_sao_enabled_flag are not accurate, even when the SPS flag is true, each slice or block can choose to apply LMCS / SAO or not.
[0260] 4. The slice type in SH and / or the allowed inter / intra / B-slice type flags in PH are signaled regardless of the case where a picture contains only one slice.
[0261] 5. List of solutions and implementation examples
[0262] To solve the above problems and other problems, the following summarized methods are disclosed. These items should be considered as examples to explain the general concept and should not be interpreted in a narrow way. In addition, these items can be used alone or in combination in any way.
[0263] Related to LMCS / zoom lists
[0264] 1. The allowed LMCS and / or scaling list mode types (e.g., enabled for all slices; disabled for all slices; and enabled for at least one but not all slices) may depend on whether the PH syntax structure is present in the slice header (or whether the current picture contains only one slice).
[0265] a. In one example, when the PH syntax structure is present in the slice header (or if the current picture contains only one slice), only two mode types are allowed.
[0266] b. In one example, how to signal the mode type may depend on whether the PH syntax structure is present in the slice header (or whether the current picture contains only one slice).
[0267] i. Alternatively, when the PH syntax structure is present in the slice header (or if the current picture contains only one slice), the signaled mode type should be 0 or 1 (or the signaled mode should not be equal to one of the three modes).
[0268] 2. Whether to signal an indicator to use / enable LMCS for a lower level (e.g., slice / slice / sub-picture) depends on the non-binary LMCS-related syntax elements (e.g., LMCS mode index) signaled in a higher level (e.g., picture, PH / PPS) and whether the PH syntax structure is not present in the slice header (or whether the current picture contains more than one slice or depends on pps_one_slice_per_picture_flag).
[0269] a. In one example, whether to signal a lower-level indicator may be based on a conditional check of whether both of the following conditions are met:
[0270] i. Higher-level non-binary LMCS-related syntax elements indicate that LMCS is enabled for at least one slice but not for all slices (or indicate that LMCS is enabled at the sequence level or picture level, and whether LMCS is used for each slice is controlled at the slice level).
[0271] ii. The current picture consists of more than one slice, or pps_one_slice_per_picture_flag is false.
[0272] b. In one example, the signaling of the lower-level indicator may be skipped if one or both of the following two conditions are true:
[0273] i. Higher-level non-binary LMCS-related syntax elements indicate that LMCS is enabled for all slices or that LMCS is disabled for all slices (or that LMCS is used by all slices or that LMCS is disabled for all slices).
[0274] ii. The current image contains only one strip.
[0275] c. In one example, the conditional checks from:
[0276]
[0277] Modify to the following:
[0278]
[0279] d. In one example, the conditional check is as follows:
[0280]
[0281] Modify to the following:
[0282]
[0283] e. Alternatively, in addition, when the lower level indicator of LMCS (e.g., slice_lmcs_used_flag / slice
[0284] _lmcs_enabled_flag) is not signaled, it is inferred from the non-binary LMCS related syntax elements.
[0285] i. In one example, the non-binary LMCS related syntax element is ph_lmcs_mode_idc.
[0286] ii. Alternatively, in addition, use the slice-level LMCS (e.g., slice_lmcs_used_flag / slice_
[0287] The inference of ph_lmcs_enabled_flag) is (ph_lmcs_mode_idc==0?0:1).
[0288] 3. Whether to signal an indicator that explicit scaling lists (ESLs) are used / enabled for a lower level (e.g., slice / slice / sub-picture) depends on the non-binary LMCS-related syntax elements (e.g., ESL mode index) signaled in a higher level (e.g., picture, in PH / PPS) and whether the PH syntax structure is not present in the slice header (or whether the current picture contains more than one slice or depends on pps_one_slice_per_picture_flag).
[0289] a. In one example, whether to signal a lower-level indicator may be based on a conditional check of whether both of the following conditions are met:
[0290] i. A higher-level non-binary ESL-related syntax element indicates that ESL is enabled for at least one slice but not for all slices (or indicates that ESL is enabled at the sequence level or picture level, and whether ESL is used for each slice is controlled at the slice level).
[0291] ii. The current picture consists of more than one slice, or pps_one_slice_per_picture_flag is false.
[0292] b. In one example, signaling of the lower level indicator may be skipped if both of the following conditions are true:
[0293] i. Higher-level non-binary ESL-related syntax elements indicate that ESL is enabled for all stripes or that ESL is disabled for all stripes (or that ESL is used by all stripes or is disabled for all stripes).
[0294] ii. The current image contains only one strip.
[0295] c. In one example, the conditional checks from:
[0296]
[0297] Modify to the following:
[0298]
[0299]
[0300] d. In one example, the conditional check is as follows:
[0301]
[0302] Modify to the following:
[0303]
[0304] e. Alternatively, in addition, when the lower-level indicators of the ESL (e.g., slice_explicit_scaling_list
[0305] Inferred from non-binary ESL-related syntax elements when the _used_flag / slice_lmcs_enabled_flag) is not signaled.
[0306] i. In one example, the non-binary ESL related syntax element is ph_lmcs_mode_idc.
[0307] ii. Furthermore, alternatively, the inference of using the slice level ESL (eg, slice_lmcs_used_flag / slice_lmcs_enabled_flag) is (ph_lmcs_mode_idc==0?0:1).
[0308] 4. The semantics of the LMCS SPS flag is modified as follows:
[0309] Equal to 1 specified in CLVS [[Use]] Luma mapping with chroma scaling. sps_lmcs_enabled_flag equal to 0 specifies in CLVS [[Do not use]] Luma mapping with chroma scaling.
[0310] Or as follows:
[0311] Equal to 1 specifies Luma mapping with chroma scaling, CLVS of sps_lmcs_enabled_flag is equal to 0 to specify a luminance map with chroma scaling, and In CLVS ]]of Related to the indication of the strip type
[0312] 5. Whether and / or how to signal the slice type (e.g., slice_type) in SH and / or the allowed inter / intra / B slice type flags (ph_inter_slice_allowed_flag, ph_intra_slice_allowed_flag, ph_b_slice_allowed_flag) in PH may depend on whether a picture is allowed to have only one slice.
[0313] a. In one example, whether a picture is only allowed to have one slice may be indicated by pps_one_slice_per_picture_flag being true.
[0314] b. In one example, whether a picture is allowed to have only one slice may be indicated by the presence of a PH syntax structure in the slice header.
[0315] c. In one example, if only one stripe per picture is allowed for the current picture, the following may further apply:
[0316] i. If ph_inter_slice_allowed_flag is true, ph_intra_slice_allowed_flag is not signaled
[0317] ii. If ph_B_slice_allowed_flag is true, then ph_intra_slice_allowed_flag is not signaled
[0318] iii. If ph_intra_slice_allowed_flag is true, ph_b_slice_allowed_flag is not signaled
[0319] iv.slice_type is not signaled and inferred. with by The loop filtering technique represented by Deblocking filter, ALF, SAO)
[0320] 6. The semantics of the SAO SPS flag is modified as follows:
[0321] Equal to 1 specifies sample adaptive offset process exist After the deblocking filtering process [Apply]] to reconstruct the picture. sps_sao_enabled_flag equal to 0 specifies the sample adaptive offset process The deblocking filtering process is not applied to the reconstructed picture afterwards.
[0322] 7. Codec Tools An indicator of the enabled mode type may be signaled at the first video unit level.
[0323] a. In one example, the allowed mode types may include: enabled for all video sub-units; disabled for all video sub-units; enabled for at least one video sub-unit but not all video sub-units.
[0324] i. In one example, the first video unit may be a picture.
[0325] ii. In one example, the sub-video unit may be a slice / slice / sub-picture.
[0326] b. The activation mode type may be signaled in PH / PPS.
[0327] c. The allowed mode types may depend on whether the PH syntax structure is present in the slice header (or whether the current picture contains only one slice).
[0328] i. In one example, when the PH syntax structure is present in the slice header (either the current picture contains only one slice or not, depending on pps_one_slice_per_picture_flag), only two mode types are allowed.
[0329] ii. In one example, how to signal the mode type may depend on whether the PH syntax structure is present in the slice header (or whether the current picture contains only one slice or depends on pps_one_slice_per_picture_flag).
[0330] 1. Alternatively, when the PH syntax structure is present in the slice header (or the current picture contains only one slice or pps_one_slice_per_picture_flag is true), the signaled mode type should be 0 or 1 (or the signaled mode should not be equal to one of the three modes).
[0331] 8. Whether to signal use / enable for lower levels (e.g., slices / tiles / sub-pictures) The indicator depends on whether the PH syntax structure is not present in the slice header (or whether the current picture contains more than one slice or on the value of pps_one_slice_per_picture_flag).
[0332] b. In one example, if the current picture includes only one picture, or the PH syntax structure is present in the slice header, or pps_one_slice_per_picture_flag is true, signaling of the lower-level indicators is skipped.
[0333] c. Furthermore, alternatively, when a lower level indicator is not signaled, it is inferred as an enabled / used value signaled in a higher level (eg, in PH / PPS).
[0334] 9. The use of can be indicated at two levels and the codec tool is introduced Two levels of control, where a higher level control (e.g., picture level) and a lower level (e.g., slice level) control are used, and how / whether the lower level control information exists depends on the higher level control information. In addition, the following also applies:
[0335] a. In the first example, apply one or more of the following sub-bullets:
[0336] i. A first non-binary value indicator (eg, ) can be signaled at a higher level (e.g., in the picture header (PH)) to specify how to enable
[0337] 1) In one example, when the first indicator is equal to X (e.g., X=1), it specifies that X is enabled for all slices associated with the PH; when the first indicator is equal to Y (Y!=X) (e.g., Y=2), it specifies that one or more but not all slices associated with the PH are enabled. When the first indicator is equal to Z (Z!=X and Z!=Y) (eg, Z=0), it specifies that all stripes associated with the PH are disabled.
[0338] a) Furthermore, alternatively, when the first indicator is not present, the value of the indicator is inferred to be equal to a default value, such as Z.
[0339] b) Alternatively, when the first indicator is equal to Y (Y!=X) (eg, Y=1), it specifies that when decoding a slice, the transform and / or non-transform coefficients are scaled. The use of for pictures is enabled.
[0340] c) Alternatively, when the first indicator is equal to Y (Y!=X) (eg, Y=1), it specifies that when decoding the slice, the transform and / or non-transform coefficients are scaled. The use of for pictures can be enabled.
[0341] 2) In one example, when the first indicator is equal to X (eg, X=2), it specifies that all stripes associated with the PH are disabled. When the first indicator is equal to Y (Y!=X) (eg,
[0342] Y = 1), which specifies that one or more but not all stripes associated with the PH are disabled.
[0343] When the first indicator is equal to Z (Z!=X and Z!=Y) (eg, Z=0), it specifies that all slices associated with the PH are enabled.
[0344] a) Furthermore, alternatively, when the first indicator is not present, the value of the indicator is inferred to be equal to a default value, such as X.
[0345] b) Alternatively, when the first indicator is equal to Y (Y!=X) (eg, Y=1), it specifies that when decoding a slice, the transform and / or non-transform coefficients are scaled. The usage of images can be disabled.
[0346] c) Alternatively, when the first indicator is equal to Y (Y!=X) (eg, Y=1), it specifies that when decoding the slice, the transform and / or non-transform coefficients are scaled. The use of images is disabled.
[0347] 3) In addition, alternatively, the Enable flags (for example,
[0348] ) conditionally signals the first indicator.
[0349] 4) In addition, alternatively, the first indicator can be encoded and decoded with u(v), or u(2) or ue(v).
[0350] 5) Furthermore, alternatively, the first indicator may be encoded or decoded using a truncated unary code.
[0351] 6) In addition, alternatively, under the condition of checking the value of the first indicator, the corresponding APS information used by the slice (e.g., ALF APS ).
[0352] ii. Enable / disable for lower levels can be signaled at lower levels (e.g. in the slice header) A second indicator (e.g., slice_TX_present_flag) may be conditionally signaled by checking the value of the first indicator.
[0353] 1) In one example, the second indicator may be signaled under a conditional check of "first indicator equals Y".
[0354] a) Alternatively, the second indicator may be signaled under a conditional check of "value of first indicator >> 1" or "value of first indicator / 2" or "value of first indicator & 0x01".
[0355] b) Furthermore, alternatively, the second indicator may be absent and inferred to be enabled when the first indicator is equal to Y; or inferred to be disabled when the first indicator is equal to Z.
[0356] 2) Whether to signal the second indicator may depend on the first indicator and / or whether the current picture consists of more than one slice (or whether the PH syntax is not present in the SH).
[0357] a) If the first indicator tells (tell) to enable at least one stripe but not all stripes And the PH syntax is not present in the SH (or pps_one_slice_per_picture_flag is false), the second indicator may be signaled.
[0358] b) If the first indicator indicates whether to enable or disable for all stripes Either the PH syntax is present in the SH (or pps_one_slice_per_picture_flag is true), then the signaling of the second indicator may be skipped.
[0359] i. Furthermore, alternatively, when not signaled, it is inferred from the value of the first indicator, for example, set to the value of the first indicator or set to (first indicator == 0? 0:1).
[0360] b. In the second example, apply one or more of the following sub-bullets:
[0361] i. More than one indicator may be signaled at a higher level (e.g., in a picture header (PH)) to specify how to enable
[0362] 1) In one example, two indicators (e.g., two 1-bit flags) may be signaled in the PH
[0363] a) In one example, the first indicator specifies whether an enable The second indicator specifies whether all the slices associated with the PH are enabled.
[0364] ii. In addition, alternatively, the value of the first indicator may be used, for example, when the first indicator specifies that there is an enable A second indicator is conditionally signaled when at least one stripe of the at least one stripe is detected.
[0365] I. Additionally, alternatively, when the second indicator is not present, all stripes are inferred to be enabled
[0366]
[0367] II. Alternatively, if the first indicator in the conforming bitstream is false, then the second indicator is required to be false.
[0368] iii. In addition, alternatively, according to at least one of the value of the first indicator and the value of the second indicator, for example, when the first indicator specifies that at least one stripe is enabled And the second indicator specifies that not all stripes are enabled The third indicator may be conditionally signaled in the SH when .
[0369] I. Furthermore, alternatively, when the third indicator is not present, it may be inferred based on the value of the first and / or second indicator (eg, inferred to be equal to the value of the first indicator).
[0370] b) Alternatively, the first indicator specifies whether there is a disabled At least one slice associated with the PH. The second indicator specifies whether all slices associated with the PH are disabled.
[0371]
[0372] i. In addition, alternatively, the value of the first indicator may be used, for example, when the first indicator specifies the presence of A second indicator is conditionally signaled when at least one stripe of the at least one stripe is detected.
[0373] I. Additionally, alternatively, when the second indicator is not present, it is inferred that all stripes associated with the PH are disabled
[0374] ii. In addition, alternatively, according to at least one of the value of the first indicator and the value of the second indicator, for example, when the first indicator specifies that at least one stripe is enabled And the second indicator specifies that not all stripes are disabled The third indicator may be conditionally signaled in the SH when .
[0375] I. Furthermore, alternatively, when the third indicator is not present, it may be inferred from the value of the first and / or second indicator (eg, inferred to be equal to the value of the first indicator).
[0376] 2) In addition, alternatively, the Enable flags (for example,
[0377] ) value, conditionally signaling the first indicator.
[0378] ii. Enable / disable for lower levels can be signaled at lower levels (e.g. in slice header)
[0379] The third indicator (e.g., ), and it can signal conditionally by checking the value of the first indicator and / or the second indicator.
[0380] 3) In one example, in the "Not all stripes are enabled " or "Not all stripes are disabled In a third example, two 1-bit flags may be signaled at a higher level (e.g., in a picture header (PH)) to specify how to enable the 1-bit flag at a lower level (e.g., in a slice header (SH)).
[0381] i. The first PH flag (e.g., named ph_all_slices_use_TX_flag) equal to 1 specifies that all slices of the picture use The first PH flag equal to 0 specifies that each slice of the picture may or may not be used.
[0382] ii. The second PH flag (e.g., named ph_no_slice_uses_TX_flag) equal to 1 specifies that no slices are used in the picture The second PH flag is equal to 0 to specify that each slice of the picture can be used or not
[0383] iii. When the first PH flag is equal to 0, only the second PH flag is signaled.
[0384] iv. When the first PH flag is equal to 1 or when (the first PH flag is equal to 0 and the second PH flag is equal to 0)
[0385] When , the zoom list APS ID is signaled in PH.
[0386] v. When the first PH flag is equal to 0 and the second flag is equal to 0, a SH flag (eg named slice_use_TX_flag) is signaled in the SH.
[0387] vi. When the first PH flag is equal to 1, the value of the SH flag is inferred to be equal to 0.
[0388] vii. When the first PH flag is equal to 0 and the second PH flag is equal to 0, the value of the SH flag is inferred to be equal to 0.
[0389] viii. If the value of the SH flag is equal to 1, use Decode the strip. Otherwise, use Decode the stripe.
[0390] d. In a fourth example, one or more indicators of the presence or absence of one or more or each of the zoom list related aspects (eg, enabled / disabled, APS ID) in the PH or SH may be signaled.
[0391] i. Alternatively, an indicator is used, again, the indicator is a 1-bit flag.
[0392] 1) In one example, when the indicator specifies that relevant aspects are present in the PH, all stripes infer the values present in the PH and skip the signaling of those relevant aspects in the SH.
[0393] 2) In one example, when the indicator specifies that relevant aspects are present in the SH, and signaling of those relevant aspects is skipped in the PH.
[0394] ii. In one example, one or more indicators are signaled in the pH value.
[0395] iii. In another example, one or more indicators are signaled in the PPS.
[0396] iv. In another example, one or more indicators are signaled in the SPS.
[0397] Generally
[0398] 10. The proposed method can be extended to other codecs, for example, by using non-binary value indicators to indicate the mode type.
[0399] 11. Whether to signal relevant information (eg, in PH / SH) may be further controlled by some syntax elements at a higher level, such as in PPS / SPS.
[0400] d. Alternatively, whether to signal the relevant information (eg, in PPS / PH / SH) may be further controlled by some syntax element in a higher level, such as in SPS.
[0401] e. In one example, a syntax element may be present at a higher level to indicate whether the on / off control may be different in a picture set, picture, or slice.
[0402] Figure 5 5 is a block diagram illustrating an example video processing system 500 in which various techniques disclosed herein may be implemented. Various embodiments may include some or all of the components of system 500. System 500 may include an input 502 for receiving video content. The video content may be received in a raw or uncompressed format, for example, 8 or 10 bit multi-component pixel values, or may be in a compressed or encoded format. Input 502 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, passive optical networks (PONs), etc., and wireless interfaces such as Wi-Fi or cellular interfaces.
[0403] System 500 may include a codec component 504, which may implement various codecs or encoding methods described in this document. Codec component 504 may reduce the average bit rate of the video from input 502 to the output of codec component 504 to generate a codec representation of the video. Therefore, codec technology is sometimes referred to as video compression or video transcoding technology. As represented by component 506, the output of codec component 504 may be stored or sent via a connected communication. Component 508 may use a bitstream (or codec) representation of the storage or communication transmission of the video received at input 502 to generate pixel values or displayable video sent to display interface 510. The process of generating user-visible video from a bitstream representation is sometimes referred to as video decompression. In addition, although some video processing operations are referred to as "codec" operations or tools, it should be understood that codec tools or operations are used at encoders, and corresponding decoding tools or operations that are opposite to the encoding results will be performed by decoders.
[0404] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB) or High Definition Multimedia Interface (HDMI) or Display Port, etc. Examples of storage interfaces include SATA (Serial Advanced Technology Attachment), PCI, IDE interfaces, etc. The techniques described in this document may be implemented in various electronic devices, such as mobile phones, laptops, smartphones, or other devices capable of performing digital data processing and / or video display.
[0405] Figure 6 3600 is a block diagram of a video processing device 3600. Device 3600 can be used to implement one or more methods described herein. Device 3600 can be implemented in a smartphone, a tablet computer, a computer, an Internet of Things (IoT) receiver, etc. Device 3600 may include one or more processors 3602, one or more memories 3604, and video processing hardware 3606. Processor 3602 can be configured to implement one or more methods described in this document. Memory (multiple memories) 3604 can be used to store data and code for implementing the methods and techniques described herein. Video processing hardware 3606 can be used to implement some of the techniques described in this document in hardware circuits.
[0406] Figure 8 is a block diagram illustrating an example video coding system 100 that may utilize the techniques of this disclosure.
[0407] like Figure 8 As shown, the video coding system 100 may include a source device 110 and a destination device 120. The source device 110 generates encoded video data, which may be referred to as a video encoding device. The destination device 120 may decode the encoded video data generated by the source device 110, which may be referred to as a video decoding device.
[0408] Source device 110 may include a video source 112 , a video encoder 114 , and an input / output (I / O) interface 116 .
[0409] The video source 112 may include a source such as a video capture device, an interface for receiving video data from a video content provider, and / or a computer graphics system for generating video data, or a combination of these sources. The video data may include one or more pictures. The video encoder 114 encodes the video data from the video source 112 to generate a bit stream. The bit stream may include a bit sequence that forms a codec representation of the video data. The bit stream may include a codec picture and associated data. The codec picture is a codec representation of the picture. The associated data may include a sequence parameter set, a picture parameter set, and other grammatical structures. The I / O interface 116 may include a modulator / demodulator (modem) and / or a transmitter. The encoded video data may be directly sent to the destination device 120 via the I / O interface 116 through the network 130a. The encoded video data may also be stored on a storage medium / server 130b for access by the destination device 120.
[0410] Destination device 120 may include an I / O interface 126 , a video decoder 124 , and a display device 122 .
[0411] The I / O interface 126 may include a receiver and / or a modem. The I / O interface 126 may obtain encoded video data from the source device 110 or the storage medium / server 130b. The video decoder 124 may decode the encoded video data. The display device 122 may display the decoded video data to the user. The display device 122 may be integrated with the destination device 120, or may be external to the destination device 120, and the destination device 120 may be configured to interface with an external display device.
[0412] The video encoder 114 and the video decoder 124 may operate according to a video compression standard, such as the High Efficiency Video Coding (HEVC) standard, the Versatile Video Coding (VVM) standard, and other current and / or future standards.
[0413] Fig. 9 is a block diagram showing an example of a video encoder 200, which may be Figure 8 Video encoder 114 in system 100 is shown.
[0414] Video encoder 200 may be configured to perform any or all of the techniques of this disclosure. Fig. 9In the example of , video encoder 200 includes multiple functional components. The techniques described in this disclosure can be shared between various components of video encoder 200. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.
[0415] The functional components of the video encoder 200 may include a segmentation unit 201, a prediction unit 202 which may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205 and an intra-frame prediction unit 206, a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213 and an entropy coding unit 214.
[0416] In other examples, the video encoder 200 may include more, fewer, or different functional components. In an example, the prediction unit 202 may include an intra block copy (IBC) unit. The IBC unit may perform prediction in an IBC mode, where at least one reference picture is a picture in which the current video block is located.
[0417] Furthermore, some components such as the motion estimation unit 204 and the motion compensation unit 205 may be highly integrated, but for the purpose of explanation, they are described in detail in the following sections. Fig. 9 are represented separately in the example.
[0418] The partitioning unit 201 may partition a picture into one or more video blocks. The video encoder 200 and the video decoder 300 may support various video block sizes.
[0419] The mode selection unit 203 may select one of the coding modes (intra or inter), for example, based on the error result, and provide the resulting intra or inter coding block to the residual generation unit 207 to generate residual block data, and to the reconstruction unit 212 to reconstruct the coding block for use as a reference picture. In some examples, the mode selection unit 203 may select a combination of intra and inter prediction (CIIP) modes, where the prediction is based on an inter prediction signal and an intra prediction signal. In the case of inter prediction, the mode selection unit 203 may also select a resolution of motion vectors for the block (e.g., sub-pixel or integer pixel precision).
[0420] To perform inter-frame prediction on the current video block, the motion estimation unit 204 may generate motion information of the current video block by comparing the current video block with one or more reference frames from the buffer 213. The motion compensation unit 205 may determine a predicted video block of the current video block based on the motion information of pictures other than the picture associated with the current video block from the buffer 213 and decoded samples.
[0421] Motion estimation unit 204 and motion compensation unit 205 may perform different operations on the current video block, eg, depending on whether the current video block is in an I slice, a P slice, or a B slice.
[0422] In some examples, the motion estimation unit 204 may perform unidirectional prediction on the current video block, and the motion estimation unit 204 may search for a reference video block of the current video block in the reference pictures of list 0 or list 1. The motion estimation unit 204 may then generate a reference index indicating a reference picture in list 0 or list 1 containing the reference video block and a motion vector indicating a spatial displacement between the current video block and the reference video block. The motion estimation unit 204 may output the reference index, the prediction direction indicator, and the motion vector as motion information of the current video block. The motion compensation unit 205 may generate a predicted video block of the current block based on the reference video block indicated by the motion information of the current video block.
[0423] In other examples, the motion estimation unit 204 may perform bidirectional prediction on the current video block, and the motion estimation unit 204 may search for a reference video block of the current video block in the reference pictures in list 0, and may also search for another reference video block of the current video block in the reference pictures in list 1. The motion estimation unit 204 may then generate a reference index and a motion vector, the reference index indicating the reference picture in list 0 or list 1 containing the reference video block, and the motion vector indicating the spatial displacement between the reference video block and the current video block. The motion estimation unit 204 may output the reference index and the motion vector of the current video block as the motion information of the current video block. The motion compensation unit 205 may generate a predicted video block of the current video block based on the reference video block indicated by the motion information of the current video block.
[0424] In some examples, motion estimation unit 204 may output the entire motion information set for use in the decoding process of a decoder.
[0425] In some examples, motion estimation unit 204 may not output the entire set of motion information for the current video. Instead, motion estimation unit 204 may reference motion information of another video block to signal motion information for the current video block. For example, motion estimation unit 204 may determine that the motion information of the current video block is sufficiently similar to the motion information of a neighboring video block.
[0426] In one example, motion estimation unit 204 may indicate a value in a syntax structure associated with the current video block that indicates to video decoder 300 that the current video block has the same motion information as another video block.
[0427] In another example, the motion estimation unit 204 may identify another video block and a motion vector difference (MVD) in a syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 300 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[0428] As described above, the video encoder 200 may predictively signal motion vectors.Two examples of predictive signaling techniques that may be implemented by the video encoder 200 include Advanced Motion Vector Prediction (AMVP) and merge mode signaling.
[0429] The intra prediction unit 206 may perform intra prediction on the current video block. When the intra prediction unit 206 performs intra prediction on the current video block, the intra prediction unit 206 may generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block may include a prediction video block and various syntax elements.
[0430] Residual generation unit 207 may generate residual data for the current video block by subtracting (eg, indicated by a minus sign) the predicted video block of the current video block from the current video block. The residual data for the current video block may include residual video blocks corresponding to different sample components of samples in the current video block.
[0431] In other examples, there may be no residual data for the current video block, for example, in skip mode, and the residual generation unit 207 may not perform a subtraction operation.
[0432] Transform processing unit 208 may generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video block associated with the current video block.
[0433] After transform processing unit 208 generates a transform coefficient video block associated with the current video block, quantization unit 209 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0434] The inverse quantization unit 210 and the inverse transform unit 211 may apply inverse quantization and inverse transform to the transform coefficient video block, respectively, to reconstruct a residual video block from the transform coefficient video block. The reconstruction unit 212 may add the reconstructed residual video block to corresponding samples from one or more prediction video blocks generated by the prediction unit 202 to generate a reconstructed video block associated with the current block for storage in the buffer 213.
[0435] After the reconstruction unit 212 reconstructs the video block, a loop filtering operation may be performed to reduce video blocking artifacts in the video block.
[0436] The entropy coding unit 214 may receive data from other functional components of the video encoder 200. When the entropy coding unit 214 receives the data, the entropy coding unit 214 may perform one or more entropy coding operations to generate entropy-coded data and output a bitstream including the entropy-coded data.
[0437] Fig.10 is a block diagram showing an example of a video decoder 300, which may be Figure 8 The video decoder 114 in the system 100 is shown.
[0438] Video decoder 300 may be configured to perform any or all of the techniques of this disclosure. Fig.10 In the example of , video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared between various components of video decoder 300. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.
[0439] exist Fig.10 In the example of FIG. 3 , the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra-frame prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, the video decoder 300 may perform operations generally similar to those for the video encoder 200 ( Fig. 9 ) is a decoding process that is the inverse of the encoding process described.
[0440] The entropy decoding unit 301 may retrieve an encoded bitstream. The encoded bitstream may include entropy encoded video data (e.g., encoded blocks of video data). The entropy decoding unit 301 may decode the entropy encoded video data, and the motion compensation unit 302 may determine motion information from the entropy decoded video data, including motion vectors, motion vector precision, reference picture list indexes, and other motion information. The motion compensation unit 302 may determine this information, for example, by performing AMVP and merge mode.
[0441] The motion compensation unit 302 may generate a motion compensated block, possibly performing interpolation based on an interpolation filter. An identifier of an interpolation filter to be used with sub-pixel precision may be included in the syntax element.
[0442] The motion compensation unit 302 may calculate interpolated values of sub-integer pixels of the reference block using interpolation filters as used by the video encoder 200 during encoding of the video block. The motion compensation unit 302 may determine the interpolation filters used by the video encoder 200 based on received syntax information and use the interpolation filters to generate a prediction block.
[0443] The motion compensation unit 302 may use some syntax information to determine the size of blocks used to encode (multiple) frames and / or (multiple) slices of the encoded video sequence, partitioning information describing how each macroblock of a picture of the encoded video sequence is partitioned, a mode indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-frame coded block, and other information for decoding the encoded video sequence.
[0444] The intra prediction unit 303 may form a prediction block from spatially neighboring blocks using, for example, an intra prediction mode received in the bitstream. The inverse quantization unit 303 inversely quantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 301. The inverse transform unit 303 applies an inverse transform.
[0445] The reconstruction unit 306 may add the residual block to the corresponding prediction block generated by the motion compensation unit 202 or the intra prediction unit 303 to form a decoded block. If necessary, a deblocking filter may also be applied to filter the decoded block to remove blocking artifacts. The decoded video block is then stored in a buffer 307, which provides reference blocks for subsequent motion compensation / intra prediction and also produces a decoded video for presentation on a display device.
[0446] Next, a list of preferred solutions for some embodiments is provided.
[0447] The following solution illustrates an example embodiment of the technique discussed in the previous section (eg, item 1).
[0448] 1. A video processing method (eg, Figure 7 ), performing (702) a conversion between a video including one or more video pictures including one or more slices and a codec representation of the video; wherein the codec representation conforms to a format rule, wherein the format rule specifies which type of luma mapping with a chroma scaling mode or a scaling list mode type is applicable to the slices, and the conversion is indicated by a picture header syntax structure in a slice header or a picture header in a picture containing a single slice.
[0449] 2. The method of solution 1, wherein the format rule specifies that the picture header syntax structure in the packet header indicates that only two mode types are allowed.
[0450] The following solution illustrates an example embodiment of the technique discussed in the previous section (eg, item 2).
[0451] 3. A video processing method, comprising: performing conversion between a video including one or more video pictures including one or more slices and a codec representation of the video; wherein the codec representation complies with a format rule, wherein the format rule specifies an indicator including an indication that enabling a luma mapping with chroma scaling (LMCS) mode at a first video level depends on a higher-level non-binary LMCS-related syntax element and whether the picture consists of only one slice.
[0452] 4. The method according to solution 3, wherein the first video level is a slice level.
[0453] 5. A method according to any of solutions 3-4, wherein the higher level corresponds to a picture or sequence or picture parameter set level.
[0454] The following solutions illustrate example embodiments of the techniques discussed in the previous section (eg, items 3-4).
[0455] 6. A video processing method, comprising: performing conversion between a video including one or more video pictures including one or more slices and a codec representation of the video; wherein the codec representation complies with a format rule, wherein the format rule specifies an indicator including an indication that enabling an explicit scaling list (ESL) mode at a first video level depends on a higher-level non-binary LMCS-related syntax element and whether the picture consists of only one slice.
[0456] 7. The method according to solution 6, wherein the first video level is a slice level.
[0457] 8. A method according to any of solutions 6-7, wherein the higher level corresponds to a picture or sequence or picture parameter set level.
[0458] The following solutions illustrate example embodiments of the techniques discussed in the previous section (eg, items 5-8).
[0459] 9. A video encoding method, comprising: performing conversion between a video including one or more pictures including one or more slices and a codec representation of the video, wherein the codec representation conforms to a format rule, wherein the format rule specifies whether a picture includes exactly one slice and a slice type or a slice type flag in a slice header controlling the exactly one slice.
[0460] 10. The method of solution 9, wherein the format rule specifies that for a picture with exactly one slice, the corresponding picture header syntax structure must be included in the slice header.
[0461] The following solution illustrates an example embodiment of the technique discussed in the previous section (eg, item 9).
[0462] 11. A video encoding method, comprising: performing a conversion between a video including one or more pictures containing one or more video regions and a codec representation of the video, wherein the codec representation complies with a format rule, wherein the format rule specifies two-level signaling including the applicability of a filtering codec tool (TX) to the video region.
[0463] 12. The method according to solution 11, wherein the two-level signaling includes a higher-level signaling at a video picture level or higher and a lower-level signaling at a slice level or lower.
[0464] 13. A method according to any of solutions 11-12, wherein the higher-level signaling includes a non-binary value indicator.
[0465] 14. A method according to any of solutions 11-13, wherein the lower-level signaling includes a binary value indicator.
[0466] 15. The method of solution 12, wherein the higher level signaling includes two 1-bit flags indicating whether all, some, or conditionally some video regions of the lower level have TX mode enabled.
[0467] 16. A method according to any of solutions 11-15, wherein the filtering codec tool includes the use of a scaling list.
[0468] 17. A method according to any one of solutions 1 to 16, wherein the conversion includes encoding the video into a codec representation.
[0469] 18. A method according to any one of solutions 1 to 16, wherein the conversion includes decoding the codec representation to generate pixel values of the video.
[0470] 19. A video decoding device, comprising a processor configured to implement the method described in one or more of solutions 1 to 18.
[0471] 20. A video encoding device comprising a processor configured to implement the method described in one or more of solutions 1 to 18.
[0472] 21. A computer program product having computer code stored thereon, which, when executed by a processor, causes the processor to implement the method of any one of solutions 1 to 18.
[0473] 22. The methods, apparatus, or systems described in this document.
[0474] Fig.13 1 is a flowchart representation of a method 1300 for video processing according to the present technology. The method 1300 includes, at operation 1310, performing conversion between a video picture of a video including one or more slices and a bitstream of the video picture according to a rule. The rule specifies that whether a first syntax element indicating use of a codec tool is present at a first level is based on a syntax flag indicating whether a syntax structure of a second level is not present at the first level. The second level is higher than the first level, and the second level is a video picture level or higher than the video picture level.
[0475] In some embodiments, the first level comprises a slice header and the second level comprises a picture header.In some embodiments, whether a video picture comprises a single slice is indicated by the syntax flag.
[0476] In some embodiments, the rule further specifies that whether the first syntax element is present at the first level is also based on a value of a second syntax element indicating usage of the codec tool at the second level. In some embodiments, usage of the codec tool at the first level is based on (1) a syntax flag indicating whether the syntax structure of the second level is not present at the first level, and (2) a value of the second syntax element.
[0477] In some embodiments, the codec tool includes a tool that maps luma samples to specific values and optionally applies a scaling operation to the values of chroma samples. In some embodiments, the codec tool includes a luma mapping tool with chroma scaling. In some embodiments, in response to (1) a second syntax element indicating in a picture header that the luma mapping tool with chroma scaling is enabled for at least one slice, and (2) a syntax structure of the picture header is not present in a slice header, a first syntax element indicating use of the luma mapping tool with chroma scaling is present. In some embodiments, in response to (1) the syntax element indicating in a picture header that the luma mapping tool with chroma scaling is disabled, or (2) a syntax structure of the picture header is present in a slice header, the first syntax element indicating use of the luma mapping tool with chroma scaling is omitted. In some embodiments, in the event that use of the luma mapping tool with chroma scaling is omitted in the slice header, the use is inferred from the second syntax element indicated in the picture header.
[0478] In some embodiments, the codec tool includes an explicit scaling list. In some embodiments, the explicit scaling list is used in a scaling process of transform coefficients. In some embodiments, in response to (1) a second syntax element indicating in a picture header that the explicit scaling list is enabled for at least one slice, and (2) a picture header syntax structure is not present in a slice header, a first syntax element indicating the use of the explicit scaling list is present. In some embodiments, in response to (1) a second syntax element indicating in a picture header that the explicit scaling list is disabled, or (2) a picture header syntax structure is present in a slice header, the first syntax element indicating the use of the explicit scaling list is omitted. In some embodiments, in the event that the use of the explicit scaling list is omitted in the slice header, the use is inferred from the second syntax element indicated in the picture header.
[0479] Fig.14 14 is a flowchart representation of a method 1400 for video processing according to the present technology. The method 1400 includes, at operation 1410, performing conversion between a video and a bitstream of the video according to a rule. The rule specifies that a syntax element in a sequence parameter set of the video indicates whether a luma mapping with chroma scaling (LMCS) tool is enabled for a codec layer video sequence (CLVS) of a reference sequence parameter set.
[0480] In some embodiments, the syntax element equal to 1 specifies that the LMCS tool is enabled for CLVS, and wherein the syntax element equal to 0 specifies that the LMCS tool is disabled for CLVS.
[0481] Fig.15 1 is a flowchart representation of a method 1500 for video processing according to the present technology. The method 1500 includes, at operation 1510, performing conversion between a video and a bitstream of the video according to a rule. The rule specifies that a syntax element in a sequence parameter set of the video indicates whether a sample adaptive offset (SAO) tool is enabled for a codec layer video sequence (CLVS) of a reference sequence parameter set.
[0482] In some embodiments, the syntax element equal to 1 specifies that the SAO tool is enabled for CLVS, and wherein the syntax element equal to 0 specifies that the SAO tool is disabled for CLVS.
[0483] Fig.161 is a flowchart representation of a method 1600 for video processing according to the present technology. The method 1600 includes, at operation 1610, performing a conversion between a video picture of a video including one or more strips and a bitstream of the video picture according to a rule. Whether and / or how the rule indicates the use of a scaling tool is determined based on whether the video picture includes a single strip. The use of the scaling tool includes whether a luminance mapping with chroma scaling (LMCS) tool is allowed to be used for the conversion. The use also includes the number of scaling mode types allowed for the conversion.
[0484] In some embodiments, the scaling mode type indicates whether the scaling tool is enabled or disabled for all or part of one or more strips of the video picture. In some embodiments, the rule specifies whether a video picture includes a single stripe by whether a picture header syntax structure of the video picture is present in a stripe header of the stripe. In some embodiments, the rule specifies that in the case where a video picture includes a single stripe, only two scaling mode types are allowed for the conversion. In some embodiments, the two scaling mode types are indicated by a value of 0 or 1.
[0485] Fig.17 17 is a flowchart representation of a method 1700 for video processing according to the present technology. The method 1700 includes, at operation 1710, performing conversion between a video picture of a video including one or more slices and a bitstream of the video picture according to a rule. The rule specifies whether and / or how to indicate in a picture header and / or a slice header that the allowed slice types are determined based on whether the video picture includes a single slice.
[0486] In some embodiments, the rule specifies that whether a video picture includes a single slice is indicated by whether a picture header syntax structure of the video picture is present in a slice header of the slice. In some embodiments, in the case where the video picture includes a single slice, a first syntax flag in the picture header that specifies whether all slices of the video picture have a specific slice type is omitted if a second syntax flag in the picture header specifies that one or more slices in the video are allowed to have a specific slice type. In some embodiments, in the case where the video picture includes a single slice, a first syntax flag in the picture header that specifies whether all slices of the video picture have a specific slice type is omitted if a second syntax flag in the picture header specifies that one or more slices of codec type B are allowed. In some embodiments, in the case where the video picture includes a single slice, a first syntax flag in the picture header that specifies whether one or more slices of codec type B are allowed is omitted if a second syntax flag in the picture header specifies that all slices of the video picture have a specific slice type. In some embodiments, the types of slices in the video picture are omitted and inferred.
[0487] Fig.181 is a flowchart representation of a method 1800 for video processing according to the present technology. The method 1800 includes, at operation 1810, performing conversion between a video picture of a video including one or more slices and a bitstream of the video picture according to a rule. The rule specifies a mode type that indicates a codec tool at a video unit level to use for the conversion.
[0488] In some embodiments, the mode type includes: (1) a first type indicating that the codec tool is enabled for all sub-units of the video unit, (2) a second type indicating that the codec tool is disabled for all sub-units of the video unit, (3) a third type indicating that the codec tool is enabled for at least one sub-unit of the video unit but not all sub-units. In some embodiments, the video unit includes a picture. In some embodiments, the sub-unit includes a slice, a slice, or a sub-picture of a picture. In some embodiments, the mode type is indicated in a picture header or a picture parameter set.
[0489] In some embodiments, the number of allowed mode types is determined based on whether the video picture includes a single slice. In some embodiments, in the case where the video picture includes a single slice, only two mode types are allowed. In some embodiments, how the mode type is indicated is determined based on whether the video picture includes a single slice. In some embodiments, in the case where the video picture includes a single slice, the mode type is limited to a subset of a plurality of allowed mode types. In some embodiments, the rule specifies whether a video picture includes a single slice by whether a picture header syntax structure of the video picture is present in a slice header of the slice.
[0490] Fig.19 1 is a flowchart representation of a method 1900 for video processing according to the present technology. The method 1900 includes, at operation 1910, performing conversion between a video picture of a video including one or more slices and a bitstream of the video picture according to a rule. The rule specifies that a non-binary syntax element or a plurality of syntax flags at a first video unit level is used to indicate the use of a codec tool at a second video unit level lower than the first video unit level.
[0491] In some embodiments, the first video unit level comprises a picture level, and the second video unit level comprises a slice level. In some embodiments, the non-binary syntax element uses at least more than two bits for encoding and decoding. The non-binary syntax element equal to X specifies that the codec tool is enabled for all slices associated with the video picture, the non-binary syntax element equal to Y specifies that the codec tool is enabled for at least one but not all slices associated with the video picture, and the non-binary syntax element equal to Z specifies that the codec tool is disabled for all slices associated with the video picture, where X!=Y, Y!=Z, and X!=Z.
[0492] In some embodiments, in the event that a non-binary syntax element is omitted, the non-binary syntax element is inferred to have a default value. In some embodiments, a non-binary syntax element equal to Y indicates that a codec tool may be applied in a scaling process of transform and / or non-transform coefficients for that conversion. In some embodiments, X=1, Y=2, and Z=0. In some embodiments, X=2, Y=1, and Z=0.
[0493] In some embodiments, the non-binary syntax element is conditionally indicated based on a corresponding syntax flag at the sequence level. In some embodiments, the non-binary syntax element is encoded as an unsigned integer, an unsigned integer 0-order Exp-Golomb codec syntax element with a leftbit first, or a truncated unary value. In some embodiments, corresponding adaptation parameter set information used by one or more slices of a video picture is indicated based on the non-binary syntax element.
[0494] In some embodiments, the plurality of syntax flags include a first syntax flag indicating whether a codec tool is enabled or disabled for at least one slice of a video picture, and a second syntax flag indicating whether a codec tool is enabled or disabled for all slices of the video picture. In some embodiments, the second syntax flag is conditionally indicated based on a value of the first syntax flag. In some embodiments, in the event that the second syntax flag is omitted, the second syntax flag is inferred to indicate that the codec tool is enabled or disabled for all slices. In some embodiments, the rule specifies that in the event that the first syntax flag indicates that the codec tool is not enabled for at least one slice of the video picture, the second syntax flag has a value indicating that the codec tool is disabled for all slices of the video picture.
[0495] In some embodiments, the plurality of syntax flags further comprises a third syntax flag conditionally indicated in the slice header according to the first syntax flag or the second syntax flag. In some embodiments, the third syntax flag is inferred based on the value of the first syntax flag and / or the second syntax flag. In some embodiments, the third syntax flag is conditionally indicated according to a corresponding syntax flag at the sequence level.
[0496] In some embodiments, the plurality of syntax flags include a first syntax flag indicating whether all slices of the video picture use the codec tool, and a second syntax flag indicating whether none of the slices of the video picture use the codec tool. In some embodiments, the second syntax flag is indicated only when the first syntax flag indicates that none of the slices of the video picture use the codec tool. In some embodiments, the scaling list adaptation parameter set identifier is included in the picture header in the case where the first syntax flag indicates that all slices of the video picture use the codec tool, or the first syntax flag and the second syntax flag indicate that at least one slice of the video picture uses the codec tool. In some embodiments, the third syntax flag in the slice header is indicated in the case where the first syntax flag indicates that none of the slices of the video picture use the codec tool and the second syntax flag indicates that at least one slice of the video picture uses the codec tool. In some embodiments, the value of the third syntax flag is based on the first syntax flag and / or the second syntax flag.
[0497] In some embodiments, the codec tool is associated with the use of or information of scaling lists. In some embodiments, the use of or information of scaling lists is omitted in a second video unit level in the event that a non-binary syntax element or a plurality of syntax flags are present in a first video unit level. In some embodiments, a plurality of syntax flags are indicated in a first video unit level, the first video unit level comprising a picture header, a picture parameter set, or a sequence parameter set. In some embodiments, a syntax element at a second video unit level is conditionally indicated based on the non-binary syntax element or the plurality of syntax flags.
[0498] In some embodiments, the non-binary syntax element or the indication of the plurality of syntax flags is determined based on information in a third video unit level higher than the first video unit level. In some embodiments, the third video unit level comprises a picture parameter set or a sequence parameter set. In some embodiments, the syntax element in the third video unit level indicates that usage of the codec tool is different within a picture set, a picture, or a slice.
[0499] In some embodiments, the conversion includes encoding the video into a bitstream. In some embodiments, the conversion includes decoding the bitstream to generate the video.
[0500] In the solution described herein, an encoder can comply with the format rules by generating a codec representation according to the format rules. In the solution described herein, a decoder can use the format rules to parse the syntax elements of the codec representation using the presence and absence of the syntax elements known according to the format rules to produce a decoded video.
[0501] In this document, the term "video processing" may refer to video encoding, video decoding, video compression, or video decompression. For example, during the conversion from a pixel representation of a video to a corresponding bitstream representation, a video compression algorithm may be applied, or vice versa. As defined by the syntax, the bitstream representation of the current video block may correspond to bits that are co-located or scattered at different locations within the bitstream, for example. For example, a macroblock may be encoded based on a transformed and encoded error residual value, and also using bits in a header and other fields in the bitstream. In addition, during conversion, the decoder may parse the bitstream based on this determination, knowing that some fields may or may not be present, as described in the above solution. Similarly, the encoder may determine whether to include or not include certain syntax fields, and generate the codec representation accordingly by including or excluding the syntax fields from the codec representation.
[0502] The disclosed and other solutions, examples, embodiments, modules and functional operations described in this document may be implemented in digital electronic circuits, or in computer software, firmware or hardware, including the structures disclosed in this document and their structural equivalents, or in a combination of one or more of them. The disclosed and other embodiments may be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer-readable medium, for execution by a data processing device or for controlling the operation of the data processing device. The computer-readable medium may be a machine-readable storage device, a machine-readable storage substrate, a memory device, a combination of substances that implement a machine-readable propagation signal, or a combination of one or more of them. The term "data processing device" encompasses all devices, equipment and machines for processing data, including, for example, a programmable processor, a computer or multiple processors or computers. In addition to hardware, the device may include code that creates an execution environment for the computer program in question, for example, code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them. A propagation signal is an artificially generated signal, for example, a machine-generated electrical, optical or electromagnetic signal, which is generated to encode information for transmission to a suitable receiver device.
[0503] A computer program (also referred to as a program, software, software application, script, or code) may be written in any form of programming language, including compiled or interpreted languages, and it may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for a computing environment. A computer program does not necessarily correspond to a file in a file system. A program may be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple cooperating files (e.g., files that store portions of one or more modules, subroutines, or code). A computer program may be deployed to execute on one computer or on multiple computers located in one location or distributed across multiple locations and interconnected by a communications network.
[0504] The processes and logic flows described in this document may be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows may also be performed by, and apparatus may also be implemented as, special purpose logic circuitry, such as an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit).
[0505] For example, processors suitable for executing computer programs include both general-purpose and special-purpose microprocessors, and any one or more processors of any kind of digital computer. Typically, the processor will receive instructions and data from a read-only memory or a random access memory or both. The basic elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include or be operably coupled to one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, to receive data from or transfer data to the mass storage device, or both. However, a computer does not need to have such a device. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices, such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD ROM and DVD-ROM disks. The processor and memory may be supplemented or incorporated into a dedicated logic circuit.
[0506] Although this patent document contains many details, these details should not be interpreted as limitations on the scope of any subject matter or the content claimed, but rather as descriptions of features peculiar to specific embodiments of specific technologies. Certain features described in the context of separate embodiments in this patent document may also be implemented in combination in a single embodiment. Conversely, the various features described in the context of a single embodiment may also be implemented in multiple embodiments individually or in any suitable sub-combination. In addition, although features may be described above as working in certain combinations, and even initially claimed as such, one or more features from the claimed combination may be deleted from the combination in some cases, and the claimed combination may be directed to a variant of a sub-combination or a sub-combination.
[0507] Similarly, although operations are described in a particular order in the drawings, this should not be understood as requiring that the operations be performed in the particular order or order shown, or that all of the operations shown be performed, in order to achieve the desired results. In addition, the separation of various system components in the embodiments described in this patent document should not be understood as requiring such separation in all embodiments.
[0508] Only a few implementations and examples are described, and other implementations, enhancements, and variations may be made based on what is described and shown in this patent document.
Claims
1. A video processing method, include: performing conversion between a video picture of a video comprising one or more slices and a bitstream of the video picture according to a rule, wherein the rule specifies that whether a first syntax element indicating use of a codec tool is present at a first level is based on whether a syntax flag indicating a syntax structure of a second level is not present at the first level, wherein the second level is higher than the first level, and wherein the second level is at or above the video picture level, and The encoding and decoding tools include tools that map luma samples to specific values and optionally apply scaling operations to the values of chroma samples.
2. The method according to claim 1, in, The codec tools include a luma mapping tool with chroma scaling.
3. The method according to claim 2, in, In response to (1) the second syntax element indicating in the picture header that the luma mapping tool with chroma scaling is enabled for at least one slice, and (2) the syntax structure of the picture header is not present in the slice header, there is a first syntax element indicating the use of the luma mapping tool with chroma scaling.
4. The method according to claim 2, in, In response to (1) the syntax element indicating disabling of the luma mapping tool with chroma scaling in the picture header, or (2) the syntax structure of the picture header is present in the slice header, omitting a first syntax element indicating use of the luma mapping tool with chroma scaling.
5. The method according to claim 4, in, In case usage of the luma mapping with chroma scaling tool is omitted in the slice header, the usage is inferred from a second syntax element indicated in the picture header.
6. The method according to claim 1, in, The first level includes a slice header, and wherein the second level includes a picture header.
7. The method according to claim 1, in, The rule also specifies that whether the first syntax element is present at the first level is also based on a value of a second syntax element that indicates usage of the codec tool at the second level.
8. The method according to claim 7, in, The use of the codec tool at the first level is based on (1) a syntax flag indicating whether the syntax structure of the second level is not present at the first level, and (2) a value of the second syntax element.
9. The method according to claim 1, in, The converting includes encoding the video into the bitstream.
10. The method according to claim 1, in, The converting includes decoding the video from the bitstream.
11. An apparatus for processing video data, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to: performing conversion between a video picture of a video comprising one or more slices and a bitstream of the video picture according to a rule, in, The rule specifies that whether a first syntax element indicating use of a codec tool is present at a first level is based on whether a syntax flag indicating a syntax structure of a second level is not present at the first level, wherein the second level is higher than the first level, and wherein the second level is at or above the video picture level, and The encoding and decoding tools include tools that map luma samples to specific values and optionally apply scaling operations to the values of chroma samples.
12. A non-transitory computer-readable storage medium storing instructions that cause a processor to: performing conversion between a video picture of a video comprising one or more slices and a bitstream of the video picture according to a rule, in, The rule specifies that whether a first syntax element indicating use of a codec tool is present at a first level is based on whether a syntax flag indicating a syntax structure of a second level is not present at the first level, wherein the second level is higher than the first level, and wherein the second level is at or above the video picture level, and The encoding and decoding tools include tools that map luma samples to specific values and optionally apply scaling operations to the values of chroma samples.
13. A non-transitory computer-readable recording medium storing a bit stream of a video picture including one or more slices, the bit stream being generated by a method performed by a video processing device, wherein the method include: Generate a bit stream of the video picture according to a rule, wherein the rule specifies that whether a first syntax element indicating use of a codec tool is present at a first level is based on whether a syntax flag indicating a syntax structure of a second level is not present at the first level, wherein the second level is higher than the first level, and wherein the second level is at or above the video picture level, and The encoding and decoding tools include tools that map luma samples to specific values and optionally apply scaling operations to the values of chroma samples.
14. A method for storing a bit stream of a video, include: Generate a bitstream of a video picture including one or more slices according to a rule; as well as storing the bitstream in a non-transitory computer-readable recording medium, wherein the rule specifies that whether a first syntax element indicating use of a codec tool is present at a first level is based on whether a syntax flag indicating a syntax structure of a second level is not present at the first level, wherein the second level is higher than the first level, and wherein the second level is at or above the video picture level, and The encoding and decoding tools include tools that map luma samples to specific values and optionally apply scaling operations to the values of chroma samples.