Signaling notification of encoding and decoding information in video bitstream
By performing grid scanning and rectangular strip segmentation on video units, combined with adaptive loop filters and transform skip encoding/decoding schemes, the problems of high encoding/decoding overhead and low parallel processing efficiency in existing video encoding/decoding technologies are solved, achieving more efficient video processing and resource utilization.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-02-09
- Publication Date
- 2026-03-13
AI Technical Summary
Existing video encoding and decoding technologies suffer from high encoding and decoding overhead, low parallel processing efficiency, and difficulty in MTU size matching when handling video unit segmentation and loop filtering operations. This is especially true in multi-layer video encoding and decoding standards, where it is difficult to efficiently utilize processor resources.
A novel video processing method is adopted, which divides video units into grid scanning strips and rectangular strips, uses an adaptive loop filter (ALF) and transform skip (TS) encoding and decoding scheme, controls the encoding and decoding block size and in-loop filtering operations, optimizes the format rules of the video bitstream to reduce the dependence on cross-slice boundaries and stripe boundaries, and uses a dual-tree encoding and decoding scheme to distinguish the minimum allowable block size of luminance and chrominance components.
It improves the parallel processing efficiency of video encoding and decoding, reduces encoding and decoding overhead, enhances MTU size matching capability, optimizes resource utilization in video processing, and improves the overall performance of video encoding and decoding.
Smart Images

Figure CN115362682B_ABST
Abstract
Description
[0001] Cross-reference to related applications
[0002] In accordance with the applicable Patent Law and / or the Paris Convention, this application promptly claims priority and interest in International Patent Application No. PCT / CN2020 / 075216, filed on February 14, 2020. For all legal purposes, the entire disclosure of the foregoing application is incorporated herein by reference as a part of this application disclosure. Technical Field
[0003] This patent document relates to image and video encoding and decoding. Background Technology
[0004] Digital video accounts for the largest share of bandwidth usage on the Internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth demand for digital video is expected to continue to grow. Summary of the Invention
[0005] This document discloses a technique for processing the codec representation of video using control information useful for the decoded representation, which is processed by the video encoder and decoder.
[0006] In one example aspect, a video processing method is disclosed. The method includes: performing a conversion between a video comprising video units and a bitstream of the video, wherein the bitstream conforms to a format rule, and wherein the format rule specifies that, in response to whether or how the video units are segmented, a first syntax element indicating whether a loop filtering operation is performed across segment boundaries is selectively included in the bitstream.
[0007] In another example, a video processing method is disclosed. The method includes: performing a conversion between video units of a video and a bitstream of the video, wherein the bitstream conforms to a format rule, and wherein the format rule specifies that, in response to whether or how the video units are segmented into stripes, syntax elements indicating whether a loop filtering operation is performed across stripe boundaries are selectively included in the bitstream.
[0008] In another example, a video processing method is disclosed. The method includes: performing a conversion between a video region of a video and a bitstream of the video, wherein the bitstream conforms to a format rule, and wherein the format rule specifies information about the suitability of a tool for the conversion in the bitstream at the video stripe level and / or at the video picture level, and wherein the tool maps luminance samples to specific values and selectively applies scaling operations to the values of chrominance samples.
[0009] In another example, a video processing method is disclosed. The method includes: performing a conversion between a video region of a video and a bitstream of the video, wherein the conversion conforms to a size rule, and wherein the size rule specifies the maximum size of a video region encoded using a Transform Skip (TS) coding scheme or a Block-Based Incremental Pulse Codec Modulation (BDPCM) coding scheme, or the maximum size of a transform block for a video region based on the coding characteristics of the video region.
[0010] In another example, a video processing method is disclosed. The method includes: performing a conversion between a video region of a video and a bitstream of the video, wherein the bitstream conforms to format rules that specify whether and / or how the minimum allowed codec block size used during the conversion controls whether and / or how the maximum allowed transform block size is included in the bitstream.
[0011] In another example, a video processing method is disclosed. The method includes: performing a conversion between a video region of a video and a bitstream of the video, wherein the bitstream conforms to a format rule specifying whether and / or how the minimum permissible codec block size used during the conversion controls whether and / or how the bitstream includes an indication of the maximum size of the video region used for encoding or decoding using a Transform Skip (TS) scheme or a Block-Based Incremental Pulse Codec Modulation (BDPCM) scheme.
[0012] In another example, a video processing method is disclosed. The method includes: performing a conversion between video units of a video and a bitstream of a video, wherein the bitstream conforms to a format rule specifying whether and / or how the minimum permissible codec block size used during the conversion controls whether to include a field in the bitstream indicating whether a transform skip (TS) scheme or a block-based incremental pulse codec modulation (BDPCM) scheme is enabled or disabled.
[0013] In another example, a video processing method is disclosed. The method includes: performing a conversion between video units and a bitstream of video according to format rules, wherein the bitstream conforms to the format rules, which specify whether and / or how the minimum permissible codec block size used during the conversion controls whether to include fields carrying information about the suitability of the codec tools at the video region level of the bitstream.
[0014] In another example, a video processing method is disclosed. The method includes: for a conversion between a video region and a video bitstream, determining that the partitioning schemes for the luminance and chrominance components of the video have different minimum allowed block sizes, since a dual-tree encoding / decoding scheme is enabled for the video region, and performing the conversion based on the determination.
[0015] In another example, a video processing method is disclosed. The method includes: for a conversion between a video region and a bitstream of the video, determining, based on rules, a maximum number of allowed sub-block-based merging candidates for the video region; and performing the conversion based on the determination, wherein the rules specify that the maximum number of sub-block-based merging candidates used during the conversion is derived as the sum of a first variable and a second variable, wherein the first variable is equal to zero in response to affine prediction being disabled, and wherein the second variable is based on whether sub-block-based temporal motion vector prediction (sbTMVP) is enabled.
[0016] In another example, a video processing method is disclosed. The method includes: performing a conversion between a video region and a bitstream of the video region by conforming to processing rules, wherein, since the video is a 4:2:2 video or a 4:4:4 video, the processing rules are applicable to the conversion, wherein the processing rules define chroma and luma alignment of one or more of the following: (a) the number of pixel rows between the virtual boundary of an Adaptive Loop Filter (ALF) operation and the bottom boundary of a Codec Tree Block (CTB); or (b) the filter strength of the filter used for rows between the virtual boundary of the ALF operation and the bottom boundary of the CTB; or (c) a padding method for filling luma and chroma samples in the same row.
[0017] In another example aspect, a video processing method is disclosed. The method includes: a conversion between video units and a codec representation of the video; determining whether an indication of the suitability of in-loop filtering across video regions of the video units is included in the codec representation; and performing the conversion based on the determination.
[0018] In another example, a video processing method is disclosed. The method includes: performing a conversion between a video region of a video and a codec representation of the video, wherein the codec representation conforms to a format rule specifying that information regarding the applicability of a Luminosity Mapping and Chroma Scaling (LMCS) tool to the conversion is indicated at the video stripe level in the codec representation; wherein the LMCS tool includes constructing a current video block based on a first domain and a second domain during the conversion, and / or scaling the chroma residual in a luminosity-dependent manner.
[0019] In another example, a video processing method is disclosed. The method includes: performing a conversion between a video and a codec representation of the video, wherein the conversion conforms to a size rule specifying that: during encoding, the size rule is enforced on the maximum size of a video region encoded using a transform-skip codec scheme or an incremental pulse codec modulation codec scheme; or during decoding, the size rule is enforced on the maximum size of a video region decoded using a transform-skip codec scheme or an incremental pulse codec modulation codec scheme; and parsing and decoding the codec representation.
[0020] In another example, a video processing method is disclosed. The method includes: performing a conversion between a video region of a video and a codec representation of the video, wherein the codec representation conforms to format rules specifying whether or how a maximum permissible transform block size is included in the codec representation, and the format rules specifying whether the minimum permissible transform block size used during the conversion controls an indication of whether or how the maximum permissible transform block size is included in the codec representation.
[0021] In another example, a video processing method is disclosed. The method includes: performing a conversion between video units of a video and a codec representation of the video according to format rules, wherein the codec representation conforms to the format rules, which specify whether the minimum permissible codec block size used during the conversion controls whether a field carrying information about the suitability of the codec tool is included at the video region level.
[0022] In another example, a video processing method is disclosed. The method includes: a conversion between video regions and a codec representation of the video; determining, by using a dual-tree codec for the conversion, a partitioning scheme for the video's luma and chroma components that has different minimum allowable block sizes for the luma and chroma components; and performing the conversion based on the determination.
[0023] In another example, a video processing method is disclosed. The method includes: performing a conversion between a video region and its codec representation by conforming to format rules for codec representation; wherein the format rules specify that the maximum number of sub-block-based merging candidates used during the conversion is derived as the sum of a first variable and a second variable, wherein affine prediction controls the value of the first variable, and wherein sub-block-based temporal motion vector prediction controls the value of the second variable.
[0024] In another example, a video processing method is disclosed. The method includes: since the video is 4:2:2 or 4:4:4, performing a conversion between a video region and a codec representation of the video region by conforming to processing rules applicable to the conversion, wherein the processing rules define chroma and luminance aligned with one or more of the following: (a) the number of pixel rows between the virtual boundary of an adaptive loop filter operation and the bottom boundary of a codec tree block (CTB); or (b) the filter strength of the filter used for rows between the virtual boundary of the adaptive loop filter operation and the bottom boundary of the codec tree block; or (c) a padding method for filling video samples in the same row.
[0025] In yet another example, a video encoder apparatus is disclosed. The video encoder includes a processor configured to implement the methods described above.
[0026] In yet another example, a video decoder apparatus is disclosed. The video decoder includes a processor configured to implement the methods described above.
[0027] In yet another example, a computer-readable medium is disclosed on which code is stored. The code implements one of the methods described herein in the form of processor-executable code.
[0028] This paper describes these and other features. Attached Figure Description
[0029] Figure 1 An example of grid scan strip segmentation of an image is shown, where the image is segmented into 12 pieces and 3 grid scan strips.
[0030] Figure 2 An example of rectangular strip segmentation of an image is shown, where the image is divided into 24 slices (6 slice columns and 4 slice rows) and 9 rectangular strips.
[0031] Figure 3 An example of an image divided into slices and rectangular strips is shown, where the image is divided into 4 slices (2 slice columns and 2 slice rows) and 4 rectangular strips.
[0032] Figure 4 The image is shown as being divided into 15 slices, 24 strips, and 24 sub-images.
[0033] Figure 5 The nominal vertical and horizontal positions of the 4:2:2 luminance and chrominance samples in the image are shown.
[0034] Figure 6 An example of the shape of an ALF filter is shown.
[0035] Figures 7A-7D The subsampled Laplace calculation is shown. Figure 7A The subsampling locations of the vertical gradient are shown. Figure 7B The subsampling locations of the horizontal gradient are shown. Figure 7C The subsampling locations of the diagonal gradient are shown. Figure 7D The subsampling locations of the diagonal gradient are shown.
[0036] Figure 8 An example of the loop filter line buffer requirements for the luminance component in VTM-4.0 is shown.
[0037] Figure 9 An example of the loop filter line buffer requirements for the chroma component in VTM-4.0 is shown.
[0038] Figure 10 An improved block classification at virtual boundaries is shown.
[0039] Figure 11 An example of an improved ALF filter for brightness classification at virtual boundaries is shown.
[0040] Figures 12A-12C An improved luminance ALF filter at the virtual boundary is shown.
[0041] Figure 13 An example of repeated padding of the luminance ALF filter at the boundaries of an image, sub-image, strip, or patch is shown.
[0042] Figures 14A-14D An example of ALF mirror filling is shown.
[0043] Figure 15 A block diagram of an example video processing system is shown.
[0044] Figure 16 A block diagram of a video processing device is shown.
[0045] Figure 17 A flowchart of an example method for video processing is shown.
[0046] Figure 18 This is a block diagram illustrating a video encoding / decoding system according to some embodiments of the present disclosure.
[0047] Figure 19 This is a block diagram illustrating an encoder according to some embodiments of the present disclosure.
[0048] Figure 20 This is a block diagram illustrating a decoder according to some embodiments of the present disclosure.
[0049] Figure 21-31 A flowchart of an example method for video processing is shown. Detailed Implementation
[0050] In this document, chapter headings are used for ease of understanding and do not limit the applicability of the techniques and embodiments disclosed in each chapter to that chapter. Furthermore, the use of H.266 terminology in some descriptions is merely for ease of understanding and not to limit the scope of the disclosed techniques. Therefore, the techniques described herein are also applicable to other video codec protocols and designs.
[0051] 1. Overview
[0052] This article relates to video codec technology. Specifically, it concerns signaling for subpictures, slices, and stripes. These concepts can be applied individually or in various combinations to any video codec standard or non-standard video codec that supports multi-layer video codecs, such as the Multi-Functional Video Codec (VVC) currently under development.
[0053] 2. Abbreviation
[0054] APS Adaptive Parameter Set
[0055] AU Access Unit
[0056] AUD Access Unit Separator
[0057] AVC Advanced Video Codec
[0058] CLVS codec layer video sequence
[0059] CPB image buffer
[0060] CRA Clear Random Access
[0061] CTU (Codec Tree Unit)
[0062] CVS codec video sequence
[0063] DPB Decoding Image Buffer
[0064] DPS Decoding Parameter Set
[0065] EOB bitstream end
[0066] EOS sequence ends
[0067] GDR Progressive Decoding Refresh
[0068] HEVC High-Efficiency Video Encoding and Decoding
[0069] HRD Virtual Reference Decoder
[0070] IDR Instant Decoding and Refresh
[0071] JEM Joint Exploration Model
[0072] MCTS Motion Restraint Piece Set
[0073] NAL Network Abstraction Layer
[0074] OLS Output Layer Set
[0075] PH image header
[0076] PPS Image Parameter Set
[0077] PTL configuration files, hierarchies, and levels
[0078] PU Image Unit
[0079] RBSP raw byte sequence payload
[0080] SEI Supplemental Enhancement Information
[0081] SPS Sequence Parameter Set
[0082] SVC Scalable Video Codec
[0083] VCL (Video Codec Layer)
[0084] VPS Video Parameter Set
[0085] VTM VVC Test Model
[0086] VUI Video Availability Information
[0087] VVC Multi-Functional Video Encoding and Decoding
[0088] 3. Preliminary Discussion
[0089] Video codec standards have primarily evolved through the development of known ITU-T and ISO / IEC standards. ITU-T developed H.261 and H.263, while ISO / IEC developed MPEG-1 and MPEG-4 video. The two organizations jointly developed the H.262 / MPEG-2 video standard, the H.264 / MPEG-4 Advanced Video Codec (AVC) standard, and the H.265 / HEVC standard. Since H.262, video codec standards have been based on a hybrid video codec architecture, employing temporal prediction plus transform coding. To explore future video codec technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Team (JVET) in 2015. Since then, JVET has adopted many new methods and applied them to reference software called the Joint Exploration Model (JEM). JVET meetings are held quarterly, and the goal of the new codec standard is to reduce the bitrate by 50% compared to HEVC. The new video codec standard was officially named Versatile Video Coding (VVC) at the JVET meeting in April 2018, at which time the first version of the VVC Test Model (VTM) was released. With ongoing efforts to standardize VVC, new codec technologies have been adopted into the VVC standard at each JVET meeting. The VVC working draft and test model VTM are updated after each meeting. The current goal of the VVC project is Technical Completion (FDIS) at the July 2020 meeting.
[0090] 3.1. Image Segmentation Schemes in HEVC
[0091] HEVC includes four different image segmentation schemes: regular striping, non-independent striping, slice, and wavefront parallel processing (WPP), which can be applied to maximum transmission unit (MTU) size matching, parallel processing, and reduction of end-to-end latency.
[0092] Regular slices are similar to those in H.264 / AVC. Each regular slice is encapsulated in its own NAL unit, and intra-frame prediction (intra-sample prediction, motion information prediction, and encoding / decoding mode prediction) and entropy coding dependencies across slice boundaries are disabled. Therefore, regular slices can be reconstructed independently of other regular slices within the same image (although interdependencies may still exist due to loop filtering operations).
[0093] Regular striping is the only tool available for parallelization, and it is available in almost the same form in H.264 / AVC. Parallelization based on regular striping requires minimal inter-processor or inter-core communication (except for inter-processor or inter-core data sharing for motion compensation during decoding prediction of encode-decode images, which is typically much heavier than inter-processor or inter-core data sharing due to intra-frame image prediction). However, for the same reason, using regular striping can incur significant encoding / decoding overhead due to the bit cost of the stripe header and the lack of prediction across stripe boundaries. Furthermore, due to the intra-image independence of regular striping and the fact that each regular stripe is encapsulated in its own NAL, regular striping (compared to other tools mentioned below) can also serve as a key mechanism for bitstream segmentation to match MTU size requirements. In many cases, the goals of parallelization and MTU size matching are contradictory in terms of stripe layout requirements within the image. The implementation of such situations led to the development of the parallelization tools mentioned below.
[0094] Non-independent striping has a short stripe header and allows partitioning of the bitstream at tree block boundaries without disrupting any intra-picture predictions. Essentially, non-independent striping divides a regular stripe into multiple NAL units, reducing end-to-end latency by allowing a portion of the regular stripe to be sent before the entire regular stripe's encoding is complete.
[0095] In WPP, images are segmented into single-row codec tree blocks (CTBs). Entropy decoding and prediction are allowed to use data from CTBs in other segments. Parallel processing is possible through parallel decoding of CTB rows, where the decoding of a CTB row begins with a delay of two CTBs to ensure that data associated with CTBs above and to the right of the main CTB is available before the main CTB is decoded. Using this staggered start (which looks like a wavefront when represented graphically), as many processors / cores as images containing CTB rows can be parallelized. Because intra-image prediction is allowed between neighboring tree block rows within an image, the inter-processor / inter-core communication required to achieve intra-image prediction can be substantial. WPP segmentation does not result in additional NAL units compared to segmentation without WPP application, therefore WPP is not a tool for MTU size matching. However, if MTU size matching is required, regular striping can be used with WPP, but with some encoding / decoding overhead.
[0096] A slice is defined by the horizontal and vertical boundaries that divide an image into slice columns and slice rows. Slice columns extend from the top to the bottom of the image. Similarly, slice rows extend from the left to the right of the image. The number of slices in an image can be simply obtained by multiplying the number of slice columns by the number of slice rows.
[0097] Before decoding the top-left CTB of the next slice in the order of slice grid scans of an image, the CTB scan order is changed to be local within the slice (in the order of slice CTB grid scans). Similar to regular stripes, slices break the intra-image prediction dependency and the entropy decoding dependency. However, they do not need to be contained in separate NAL units (the same as WPP in this respect); therefore, slices cannot be used for MTU size matching. Each slice can be processed by one processor / core, and in the case of a stripe spanning multiple slices, the inter-processor / inter-core communication required for intra-image prediction between processing units decoding neighboring slices is limited to transmitting the shared stripe header and loop filtering associated with reconstructed samples and metadata sharing. When a stripe contains more than one slice or WPP segment, the entry point byte offset of each slice or WPP segment in the stripe, except for the first one, is signaled in the stripe header.
[0098] For simplicity, restrictions on the application of four different image segmentation schemes are specified in HEVC. For most profiles specified in HEVC, a given codec video sequence cannot contain both slices and wavefronts simultaneously. For each strip and slice, one or both of the following conditions must be met: 1) All codec tree blocks in a strip belong to the same slice; 2) All codec tree blocks in a slice belong to the same strip. Finally, a wavefront segment contains exactly one CTB line, and when using WPP, if a strip begins within a CTB line, then that strip must end within the same CTB line.
[0099] The most recent revision of HEVC is specified in the JCT-VC output document JCTVC-AC1005, edited by J. Boyce, A. Ramasubramonian, R. Skupin, G. J. Sullivan, A. Tourapis, and Y.-K. Wang. "HEVC Additional Supplemental Enhancement Information (Draft 4)" was released on October 24, 2017 at http: / / phenix.intevry.fr / jct / doc_end_user / documents- / 29_Macau / wg11 / JCTVC-AC1005-v2.zip. Including this revision, HEVC specifies three SEI messages related to MCT: the i.e., the domain MCTS SEI message, the MCTS extracted information set SEI message, and the MCTS extracted information nested SEI message.
[0100] The temporal MCTS SEI message indicates the presence of an MCTS in the bitstream and signals this to the MCTS. For each MCTS, motion vectors are restricted to pointing to full-sample locations within the MCTS and fractional-sample locations that require interpolation only from full-sample locations within the MCTS. Motion vector candidates predicted from temporal motion vectors derived from blocks outside the MCTS are not allowed. This allows each MCTS to be decoded independently even when no slices are not included in the MCTS.
[0101] The MCTS Extraction Information Set (SEI) message provides supplementary information that can be used for MCTS sub-bitstream extraction (specified as part of the semantics of the SEI message) to generate a bitstream conforming to the MCTS set. This information consists of multiple extraction information sets, each defining multiple MCTS sets and containing RBSP bytes for replacing the VPS, SPS, and PPS to be used during the MCTS sub-bitstream extraction process. When extracting the sub-bitstream according to the MCTS sub-bitstream extraction process, the parameter sets (VPS, SPS, and PPS) need to be rewritten or replaced, and the slice header needs to be slightly updated because one or all of the slice address-related syntax elements (including first_slice_segment_in_pic_flag and slice_segment_address) typically need to have different values.
[0102] 3.2. Image Segmentation in VVC
[0103] In VVC, an image is divided into one or more slice rows and one or more slice columns. A slice is a sequence of CTUs that covers a rectangular area of the image. The CTUs in a slice are scanned within that slice in a grid scan order.
[0104] A stripe consists of an integer number of complete slices or an integer number of consecutive complete CTU lines within an image.
[0105] Two stripe modes are supported: grid scan stripe mode and rectangular stripe mode. In grid scan stripe mode, the stripe contains a complete sequence of sheets within a grid scan of the image. In rectangular stripe mode, the stripe contains multiple complete sheets that together form a rectangular region of the image, or multiple consecutive complete CTU rows of a single sheet that together forms a rectangular region of the image. Sheets within a rectangular stripe are scanned in the grid scan sequence within the rectangular region corresponding to that stripe.
[0106] A sub-image contains one or more stripes that collectively cover a rectangular area of the image.
[0107] Figure 1 An example of grid scan strip segmentation of an image is shown, where the image is divided into 12 pieces and 3 grid scan strips.
[0108] Figure 2 An example of rectangular strip segmentation of an image is shown, where the image is divided into 24 slices (6 slice columns and 4 slice rows) and 9 rectangular strips.
[0109] Figure 3 An example of an image divided into slices and rectangular strips is shown, where the image is divided into 4 slices (2 slice columns and 2 slice rows) and 4 rectangular strips.
[0110] Figure 4 An example of sub-image segmentation of an image is shown, where the image is divided into 18 slices: 12 slices on the left, each covering a strip of 4x4 CTUs, and 6 slices on the right, each covering two vertically stacked strips of 2x2 CTUs, resulting in a total of 24 strips and 24 sub-images of different dimensions (each strip being a sub-image).
[0111] 3.3. Signaling notification of SPS / PPS / image header / strip header in VVC (JVET-Q2001-vB)
[0112] 7.3.2.3 Sequence Parameter Set (RBSP) Syntax
[0113]
[0114]
[0115]
[0116]
[0117]
[0118]
[0119]
[0120]
[0121]
[0122]
[0123]
[0124] 7.3.2.4 Image Parameter Set RBSP Syntax
[0125]
[0126]
[0127]
[0128]
[0129]
[0130]
[0131] 7.3.2.7 Image Header Structure Syntax
[0132]
[0133]
[0134]
[0135]
[0136]
[0137]
[0138]
[0139]
[0140] 7.3.7.1 General Strip Header Syntax
[0141]
[0142]
[0143]
[0144]
[0145]
[0146] 3.4. Color Spaces and Chroma Subsampling
[0147] A color space, also known as a color model (or color system), is an abstract mathematical model that simply describes a range of colors as tuples of numbers, typically 3 or 4 values or color components (e.g., RGB). Essentially, a color space is a refinement of a coordinate system and its subspaces.
[0148] For video compression, YCbCr and RGB are the most commonly used.
[0149] YCbCr, Y′CbCr, or YPb / CbPr / Cr, also written as YCBCR or Y′CBCR, is a series of color spaces used as part of the color image pipeline in video and digital photography systems. Y′ is the luminance component, while CB and CR are the blue and red difference chromaticity components. Y′ (with a superscript) is different from Y (where Y is luminance), meaning that the light intensity is based on a gamma-corrected, non-linear encoding of the RGB primary colors.
[0150] Chromaticity subsampling is a technique that utilizes the fact that the human visual system is less sensitive to color differences than to brightness, and encodes images by applying a lower resolution to chromaticity information than to brightness information. 3.4.1 4:4:4
[0152] Each of the three Y′CbCr components has the same sampling rate, therefore there is no chromatic subsampling. This scheme is sometimes used in high-end film scanners and film post-production. 3.4.2 4:2:2
[0154] Both chroma components are sampled at half the luminance sampling rate: the horizontal chroma resolution is halved, while the vertical chroma resolution remains unchanged. This reduces the bandwidth of the uncompressed video signal by one-third with almost no visual difference. An example of the nominal vertical and horizontal positions for the 4:2:2 color format in the VVC working draft is as follows: Figure 5 As shown.
[0155] Figure 5 The image shows the nominal vertical and horizontal positions of the 4:2:2 luminance and chrominance samples. 3.4.3 4:2:0
[0157] In 4:2:0, the horizontal sampling is doubled compared to 4:1:1, but the vertical resolution is halved because the Cb and Cr channels are sampled only on each alternating line in this scheme. Therefore, the data rate is the same. Cb and Cr are subsampled horizontally and vertically by a factor of 2, respectively. There are three variants of the 4:2:0 scheme with different horizontal and vertical positioning.
[0158] In MPEG-2, Cb and Cr are horizontally co-located. Cb and Cr are located between pixels in the vertical direction (positioned alternately).
[0159] • In JPEG / JFIF, H.261, and MPEG-1, Cb and Cr are positioned alternately between alternating luminance samples.
[0160] In a 4:2:0 DV configuration, Cb and Cr are both located in the horizontal direction. In the vertical direction, they are both located on alternating lines.
[0161] Table 3: SubWidthC and SubHeightC values derived from chroma_format_idc and separate_colour_plane_flag
[0162] chroma_format_idc separate_colour_plane_flag Color format SubWidthC SubHeightC 0 0 Monochromaticity 1 1 1 0 4:2:0 2 2 2 0 4:2:2 2 1 3 0 4:4:4 1 1 3 1 4:4:4 1 1
[0163] 3.4. Adaptive Loop Filter (ALF)
[0164] In VVC, an adaptive loop filter (ALF) with block-based filter adaptation is applied. For the luminance component, one of 25 filters is selected for each 4×4 block based on the direction and activity of the local gradient.
[0165] 3.5.1 Filter Shape
[0166] Two diamond filter shapes were used (e.g.) Figure 6 (As shown). The luminance component uses a 7×7 rhombus, and the chrominance component uses a 5×5 rhombus.
[0167] Figure 6 An example of the ALF filter shape is shown (chroma: 5×5 rhombus, luminance: 7×7 rhombus).
[0168] 3.5.2 Block Classification
[0169] For the luminance component, each 4×4 block is divided into one of 25 classes. The classification index C is based on its directionality D and activity. The quantization values are exported as follows:
[0170]
[0171] To calculate D and First, the gradients in the horizontal, vertical, and two diagonal directions are calculated using a one-dimensional Laplace function:
[0172]
[0173]
[0174]
[0175]
[0176] Where indices i and j represent the coordinates of the top-left sample point within the 4×4 block, and R(i, j) represents the reconstructed sample point at coordinates (i, j).
[0177] To reduce the complexity of block classification, a one-dimensional Laplace subsampling computation is applied. For example... Figures 7A-7D As shown, the same subsampling position is used for gradient calculation in all directions.
[0178] Figures 7A-7D Display subsampled Laplace calculation.
[0179] Then, the maximum and minimum values of the gradients in the horizontal and vertical directions are set as follows:
[0180]
[0181] The maximum and minimum values of the gradients in the two diagonal directions are set as follows:
[0182]
[0183] To derive the value of the directionality D, these values are compared with each other and with two thresholds t1 and t2:
[0184] Step 1. If and All are true, D is set to 0.
[0185] Step 2. If Continue from step 3; otherwise, continue from step 4.
[0186] Step 3. If D is set to 2; otherwise, D is set to 1.
[0187] Step 4. If D is set to 4; otherwise, D is set to 3.
[0188] The activity value A is calculated as follows:
[0189]
[0190] A is further quantized to the range of 0 to 4, and the quantized value is represented as
[0191] For the chromaticity components in the image, no classification method is applied; that is, a set of ALF coefficients is applied to each chromaticity component.
[0192] 3.5.3 Geometric Transformation of Filter Coefficients and Clipping Values
[0193] Before filtering each 4×4 luma block, geometric transformations such as rotation or diagonal and vertical flipping are applied to the filter coefficients f(k, l) and the corresponding filter clipping values c(k, l) based on the gradient values calculated for that block. This is equivalent to applying these transformations to samples in the filter's support region. The idea is to make different blocks to which ALF is applied more similar by aligning their orientation.
[0194] Three geometric transformations are introduced: diagonal transformation, vertical flip, and rotation.
[0195] Diagonal: f D (k, l) = f(l, k), c D (k, l) = c(l, k), (3-9)
[0196] Vertical flip: f V (k, l) = f(k, Kl-1), c V (k, l) = c(k, Kl-1) (3-10)
[0197] Rotation: f R (k, l) = f(Kl-1, k), c R (k, l) = c(Kl-1, k) (3-11)
[0198] Where K is the size of the filter, and 0≤k, l≤K-1 are the coefficient coordinates. Therefore, position (0,0) is located in the upper left corner, and position (K-1,K-1) is located in the lower right corner. Based on the gradient values calculated for this block, the transform is applied to the filter coefficients f(k,l) and the clipping value c(k,l). The table below summarizes the relationship between the transform and the four gradients in four directions.
[0199] Table 3.2 Mapping between gradient calculation and transformation for a block
[0200] gradient value Transformation <![CDATA[g d2 <g d1 and g h <g v ]]> No change <![CDATA[g d2 <g d1 and g v <g h ]]> diagonal <![CDATA[g d1 <g d2 and g h <g v ]]> Vertical flip <![CDATA[g d1 <g d2 and g v <g h ]]> Rotation
[0201] 3.5.4 Filter Parameter Signals
[0202] ALF filter parameters are signaled in the Adaptive Parameter Set (APS). Within an APS, up to 25 sets of luma filter coefficients and clipping value indices, and up to 8 sets of chroma filter coefficients and clipping value indices can be signaled. To reduce bit overhead, filter coefficients from different categories of luma components can be merged. In the strip header, the signaling informs the index of the APS used for the current stripe.
[0203] The clipping index decoded from the APS allows the clipping value to be determined using a clipping value table for the luma and chroma components. These clipping values depend on the internal bit depth. More precisely, the clipping value is obtained using the following formula:
[0204] AlfClip = {round(2 B-α*n (3-12) for n∈[0..N-1]}
[0205] B equals the internal bit depth, α is a predefined constant value equal to 2.35, and N equals 4, which is the number of clipped values allowed in VVC.
[0206] The stripe header can signal up to seven APS indices to specify the chroma filter bank used for the current stripe. The filtering process can be further controlled at the CTB level. An always signaling flag indicates whether the ALF is applied to the chroma CTB. The chroma CTB can select a filter bank from 16 fixed filters and filter banks from the APS. A signaling filter bank index is provided for the chroma CTB to indicate which filter bank is applied. The 16 fixed filter banks are predefined and hard-coded in both the encoder and decoder.
[0207] For chroma components, the signaling notification APS index in the stripe header indicates the chroma filter bank used by the current stripe. At the CTB level, if more than one set of chroma filters is set in the APS, a filter index is signaled for each chroma CTB.
[0208] The filter coefficients are quantized using a norm equal to 128. To limit multiplication complexity, bitstream consistency is applied so that coefficient values at non-center positions should be within the range of -2. 7 Up to 2 7 -1, including end values. The coefficient at the center position is not signaled in the bitstream and is considered equal to 128.
[0209] 3.5.5 Filtering Process
[0210] On the decoder side, when ALF is enabled for CTB, each sample R(i,j) within the CU is filtered, producing sample values R'(i,j) as shown below.
[0211] R′(i,j)=R(i,j)+((∑ k≠o ∑ l≠o f(k,l)×K(R(i+k,j+l)-R(i,j),c(k,l))+64)>>7) (3-9)
[0212] Where f(k, l) represents the decoding filter coefficients, K(x, y) represents the pruning function, and c(k, l) represents the pruning parameters for decoding. The variables k and l are... and The values vary between L and L, where L represents the filter length. The clipping function K(x, y) = min(y, max(-y, x)) corresponds to the function Clip3(-y, y, x).
[0213] 3.5.6 Virtual boundary filtering process for reducing row buffer size
[0214] In both hardware and embedded software, image-based processing is practically unacceptable due to the high demands placed on image buffers. Using on-chip image buffers is very expensive, while using off-chip image buffers significantly increases external memory access, power consumption, and data access latency. Therefore, in practical products, DF, SAO, and ALF have switched from image-based decoding to LCU-based decoding. When LCU-based processing is used for DF, SAO, and ALF, the entire decoding process can be completed by the LCUs in a pipelined manner during grid scanning, allowing for parallel processing of multiple LCUs. In this case, DF, SAO, and ALF require row buffers because processing one LCU row requires pixels from the LCU rows above it. Using off-chip row buffers (e.g., DRAM) increases external memory bandwidth and power consumption; using on-chip row buffers (e.g., SRAM) increases chip area. Therefore, although row buffers are already much smaller than image buffers, their size still needs to be reduced.
[0215] In VTM-4.0, such as Figure 8 As shown, the total number of line buffers required for the chroma components is 11.25 lines. The explanation for the line buffer requirements is as follows: Deblocking of horizontal edges overlapping with CTU edges cannot be performed because decision-making and filtering require lines K, L, M, M from the first CTU and lines O, P from the bottom CTU. Therefore, deblocking of horizontal edges overlapping CTU boundaries is postponed until the lower CTUs are present. Thus, for lines K, L, M, N, the reconstructed chroma samples must be stored in the line buffer (4 lines). SAO filtering can then be performed on lines A through J. SAO filtering can be performed on line J because deblocking does not change the samples in line K. For SAO filtering of line K, the edge offset classification decision is only stored in the line buffer (i.e., 0.25 chroma lines). ALF filtering can only be performed on line AF. Figure 8 As shown, ALF classification is performed on each 4×4 block. Each 4×4 block classification requires an 8×8 active window, which in turn requires a 9×9 window to compute the one-dimensional (1D) Laplace function to determine the gradient.
[0216] Therefore, for the block classification requirement of 4×4 blocks overlapping with rows G, H, I, and J, SAO filtering is applied to samples below the virtual boundary. Furthermore, ALF classification requires SAO-filtered samples from rows D, E, and F. Additionally, ALF filtering of row G requires three SAO-filtered samples from the rows above it (rows D, E, and F). Therefore, the total row buffer requirements are as follows:
[0217] – Row KN (Horizontal DF Pixels): 4 rows
[0218] – Row DJ (SAO filtered pixels): 7 rows
[0219] – SAO edge offset classifier value between rows J and K: 0.25 rows
[0220] Therefore, the total number of brightness rows required is 7 + 4 + 0.25 = 11.25.
[0221] Similarly, the line buffer for the chroma component requires, as Figure 9 As shown. The line buffer for evaluating chroma components requires 6.25 lines.
[0222] Figure 8 An example of the loop filter line buffer requirements for the luminance component in VTM-4.0 is shown.
[0223] Figure 9 An example of the loop filter line buffer requirements for the chroma component in VTM-4.0 is shown.
[0224] To eliminate the line buffer requirements of SAO and ALF, the concept of Virtual Boundary (VB) was introduced in the latest VVC to reduce the line buffer requirements of ALF. Modified block classification and filtering are applied to samples near the horizontal CTU boundary. Figure 8 As shown, VB is the horizontal LCU boundary shifted upwards by N pixels. For each LCU, SAO and ALF can process pixels above VB before the arrival of the lower LCU, but cannot process pixels below VB before the arrival of the lower LCU, due to DF. Considering hardware implementation costs, the proposed space between VB and the horizontal LCU boundary is set to four pixels for the luma component (e.g., ...). Figure 8 or Figure 10 In the case of N=4), the chroma component is set to two pixels (e.g., N=2).
[0225] Figure 10 The modified block classification at the virtual boundary is displayed.
[0226] The modified block classification is applicable Figure 11 The brightness components are shown. For the 1D Laplacian gradient calculation of the 4×4 block above the virtual boundary, only samples above the virtual boundary are used. Similarly, for the 1D Laplacian gradient calculation of the 4×4 block below the virtual boundary, only samples below the virtual boundary are used. The quantization of the active value A is scaled accordingly by taking into account the reduced number of samples used in the 1D Laplacian gradient calculation.
[0227] For filtering, both the luminance and chrominance components use a mirror (symmetric) fill operation at the virtual boundaries. For example... Figure 11 As shown, when the filtered sample point is below the virtual boundary, the neighboring sample points above the virtual boundary are filled. At the same time, the corresponding sample points on the other side are also symmetrically filled.
[0228] For another example, if we fill a sample point located at (i, j) (e.g., Figure 12B If the dashed line P0A is used, then even if the sample point is available, the area located at (m, n) will be filled (for example, Figure 12B The corresponding samples with the same filter coefficients (P3B, as shown by the dashed line in the image) are as follows: Figures 12A-12C As shown.
[0229] Figure 12A This displays the 1 required row (on each side) that needs to be filled above / below VB.
[0230] Figure 12B This shows the two required rows (on each side) that need to be filled above / below VB.
[0231] Figure 12C This shows the 3 required rows (on each side) that need to be filled above / below VB.
[0232] Unlike the mirror (symmetric) fill method used at horizontal CTU boundaries, a repeated (one-sided) fill process is applied to strip, patch, and sub-image boundaries when cross-boundary filtering is disabled. The repeated (one-sided) fill process also applies to image boundaries. Filled samples are used for both classification and filtering processes. Figure 13 An example of a repeating padding method for luminance-ALF filtering at image / sub-image / strip / piece boundaries is shown. Adaptive Loop Filter Procedure in Specification 3.5.7
[0233] 8.8.5.2 Encoding and Decoding Tree Block Filtering Process for Luminance Samples
[0234] The input to this process is:
[0235] – Before the adaptive loop filtering process, reconstruct the luminance image sample array recPicture.
[0236] –Filtered and reconstructed luminance image sample array alfPicture L ,
[0237] – Luminance position (xCtb, yCtb) specifies the top-left sample of the current luminance codec block relative to the top-left sample of the current image.
[0238] The output of this process is a modified filtered and reconstructed luminance image sample array, alfPicture. L .
[0239] Call the derivation process of filter index in section 8.8.5.3, take the position (xCtb, yCtb) and the reconstructed brightness image sample array recPicture as input, and take filtIdx[x][y] and transposeIdx[x][y]-1 as output, where x, y = 0.CtbSizeY.
[0240] To derive the reconstructed luminance sample alfPicture of the filter L [x][y], using x, y = 0..CtbSizeY-1, each reconstructed luminance sample within the current luminance codec block recPicture[x][y] is filtered as follows:
[0241] The derivation of the array of luminance filter coefficients f[j] corresponding to the filter specified by filtIdx[x][y] and the array of luminance clipping values c[j] is as follows, where j = 0..11:
[0242] – If AlfCtbFiltSetIdxY[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] is less than 16, then apply the following:
[0243] i=AlfCtbFiltSetIdxY[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] (1453)
[0244] f[j]=AlfFixFiltCoeff[AlfClassToFiltMap[i][filtIdx[x][y]]][j] (1454)
[0245] c[j]=2 BitDepth (1455)
[0246] – Otherwise, if (AlfCtbFiltSetIdxY[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY]) is greater than or equal to 16, then apply the following:
[0247] i=slice_alf_aps_id_luma[AlfCtbFiltSetIdxY[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY]-16] (1456)
[0249] f[j]=AlfCoeff L [i][filtIdx[x][y]][j] (1457)
[0250] c[j] = AlfClip L [i][filtIdx[x][y]][j] (1458)
[0251] –The brightness filter coefficients and clipping value index idx are derived from transposeIdx[x][y] as follows:
[0252] – If transposeIndex[x][y] equals 1, then apply the following:
[0253] idx[]={9,4,10,8,1,5,11,7,3,0,2,6} (1459)
[0254] Otherwise, if transposeIndex[x][y] equals 2, then apply the following:
[0255] idx[]={0,3,2,1,8,7,6,5,4,9,10,11} (1460)
[0256] Otherwise, if transposeIndex[x][y] equals 3, then apply the following:
[0257] idx[]={9,8,10,4,3,7,11,5,1,0,2,6} (1461)
[0258] – Otherwise, apply the following:
[0259] idx[]={0,1,2,3,4,5,6,7,8,9,10,11} (1462)
[0260] – The position (h) of each corresponding luminance sample (x, y) within a given array recPicture of chrominance samples. x+i v y+j The derivation is as follows, where i, j = -3..3:
[0261] h x+i =Clip3(0,pic_width_in_luma_samples-1,xCtb+x+i) (1463)
[0262] v y+j =Clip3(0,pic_height_in_luma_samples-1,yCtb+y+j) (1464)
[0263] The variables clipLeftPos, clipRightPos, clipTopPos, clipBottomPos, clipTopLeftFlag, and clipBotRightFlag are derived by calling the ALF boundary position derivation procedure specified in Section 8.8.5.5, where (xCtb, yCtb) and (x, y) are used as inputs.
[0264] – Modify variable h by calling the ALF sample filling procedure specified in section 8.8.5.6. x+i and v y+j And (xCtb, yCtb), (h x+i v y+j ), 0, clipLeftPos, clipRightPos, clipTopPos, clipBottomPos, clipTopLeftFlag, and clipBotRightFlag are used as inputs.
[0265] The derivation of the variable applyAlfLineBufBoundary is as follows:
[0266] – If the bottom boundary of the current codec block is the bottom boundary of the current image and pic_height_in_luma_samples-yCtb<=CtbSizeY-4, then applyAlfLineBufBoundary is set to 0:
[0267] Otherwise, set applyAlfLineBufBoundary to 1.
[0268] – In Table 45, the vertical sample position offsets y1, y2, y3 and the variable alfShiftY are specified based on the vertical luminance sample position y and applyAlfLineBufBoundary.
[0269] – The derivation of the variable curr is as follows:
[0270] curr = recPicture[h x ][v y (1465)
[0271] – The derivation of variable sum is as follows:
[0272]
[0273]
[0274] sum=curr+((sum+64)>>alfShiftY) (1467)
[0275] – The modified filtered and reconstructed brightness image sample alfPictureL[xCtb+x][yCtb+y] is derived as follows:
[0276] alfPictureL[xCtb+x][yCtb+y]=Clip3(0,(1< <BitDepth)-1,sum) (1468)
[0277] Table 45 - y1, y2, y3, and alfShiftY are based on the vertical luminance sample position y and the specifications of applyAlfLineBufBoundary.
[0278]
[0279] 8.8.5.4 Encoding and Decoding Tree Block Filtering Process for Chroma Samples
[0280] The input to this process is:
[0281] – Before the adaptive loop filtering process, reconstruct the chroma image sample array recPicture.
[0282] –Filter and reconstruct the chroma image sample array alfPicture.
[0283] – Chroma position (xCtbC, yCtbC), specifies the top-left sample of the current chroma codec block relative to the top-left sample of the current image.
[0284] – Replace the chroma filter index altIdx.
[0285] The output of this process is a modified filtered reconstructed chroma image sample array, alfPicture.
[0286] The current width and height of the chroma codec tree block, ctbWidthC and ctbHeightC, are derived as follows:
[0287] ctbWidthC=CtbSizeY / SubWidthC (1500)
[0288] ctbHeightC=CtbSizeY / SubHeightC (1501)
[0289] To derive the filtered reconstructed chroma sample alfPicture[x][y], each reconstructed chroma sample within the current chroma codec block recPicture[x][y] is filtered as follows, where x = 0..ctbWidthC-1, y = 0..ctbHeightC-1:
[0290] –Given the position (h) of each corresponding chromaticity sample point (x, y) within the array. x+i v y+j The derivation is as follows, where i, j = -2..2:
[0291] h x+i =Clip3(0,pic_width_in_luma_samples / SubWidthC-1,xCtbC+x+i) (1502)
[0292] v y+j =Clip3(0,pic_height_in_luma_samples / SubHeightC-1,yCtbC+y+j)(1503)
[0293] The variables clipLeftPos, clipRightPos, clipTopPos, clipBottomPos, clipTopLeftFlag, and clipBotRightFlag are derived by calling the ALF boundary position derivation procedure specified in Section 8.8.5.5, where (xCtbC*SubWidthC, yCtbC*SubHeightC) and (x*SubWidthC, y*SubHeightC) are used as inputs.
[0294] – Modify variable h by calling the ALF sample filling procedure specified in section 8.8.5.6. x+i and v y+j And (xCtb, yCtb), (h x+i v y+j The variables isChroma set, which are equal to 1, clipLeftPos, clipRightPos, clipTopPos, clipBottomPos, clipTopLeftFlag, and clipBotRightFlag are used as inputs.
[0295] The derivation of the variable applyAlfLineBufBoundary is as follows:
[0296] – If the bottom boundary of the current coding tree block is the bottom boundary of the picture and the bottom boundary of pic_height_in_luma_samples-(yCtbC*SubHeightC)<CtbSizeY–4, applyAlfLineBufBoundary is set to 0.
[0297] – Otherwise, applyAlfLineBufBoundary is set to 1.
[0298] – In Table 45, the vertical sampling position offsets y1, y2, and the variable alfShiftC are specified according to the vertical chroma sampling position y and applyAlfLineBufBoundary.
[0299] – The variable curr is derived as follows:
[0300] curr = recPicture[h x [v y (1504)
[0301] – The array f[j] of chroma filter coefficients and the array c[j] of chroma clip values are derived as follows, where j = 0.5:
[0302] f[j] = AlfCoeff C [slice_alf_aps_id_chroma][altIdx][j] (1505)
[0303] c[j] = AlfClipC[slice_alf_aps_id_chroma][altIdx][j] (1506)
[0304] – The variable sum is derived as follows:
[0305]
[0306] sum = curr + ((sum + 64) >> alfShiftC) (l508)
[0307] – The modified filtered reconstructed chroma picture sample alfPicture[xCtbC + x][yCtbC + y] is derived as follows:
[0308] alfPicture[xCtbC + x][yCtbC + y] = Clip3(0, (1 << BitDepth)-1, sum) (1509)
[0309] Table 46-y1, y2, and alfShiftC are based on the vertical chromaticity sample position y and the specification of applyAlfLineBufBoundary.
[0310]
[0311] 4. Examples of technical problems solved by publicly available solutions
[0312] The existing design of signaling notifications for SPS / PPS / image headers / strip headers in VVC has the following problems:
[0313] 1) Even if there is only one tile, the loop_filter_across_tiles_enabled_flag is signaled.
[0314] 2) Even if there is only one slice, signal loop_filter_across_slices_enabled_flag.
[0315] 3) LMCS information is signaled in the image header, not in the strip header.
[0316] 4) The maximum permissible TS block size may be larger than the maximum CU size.
[0317] 5) MaxNumSubblockMergeCand is derived differently depending on whether affine mapping is enabled.
[0318] 6) Apply mirror fill to the ALF virtual boundary to obtain unavailable samples and their corresponding ALF samples (e.g., luma ALF and chroma ALF). The position of the ALF virtual boundary is used to determine which samples are unavailable and need to be filled. However, in the current design, for 4:2:2 / 4:4:4 chroma format videos, the position of the ALF virtual boundary in the chroma ALF is inconsistent with the position in the luma ALF.
[0319] 7) In the current design, ALF filtering of the virtual horizontal CTU boundary neighborhood rows reduces the filtering strength depending on the location of the virtual boundary. However, incorrect locations of the ALF virtual boundary cause the filter strength to decrease at unexpected sample rows.
[0320] 5. Examples of technologies and implementation methods
[0321] To address the aforementioned and other issues, the methods outlined below are disclosed. The listed items should be considered as examples to explain general concepts and should not be interpreted narrowly. Furthermore, these items can be applied individually or in combination in any way.
[0322] In this disclosure, if a neighboring (adjacent or non-adjacent) sample (or line, or row) is not allowed to be located in a different video processing unit (e.g., outside the current picture, or the current subpicture, or the current slice, or the current strip, or the current brick, or the current CTU, or the current processing unit (e.g., the ALF processing unit or the narrow ALF processing unit), or any other current video unit), or if an unreconstructed or cross-filtered video processing unit is not allowed, then the neighboring (adjacent or non-adjacent) sample is "unavailable".
[0323] The padding method used for ALF virtual boundaries can be represented as "mirror padding," where the first unavailable sample located at (i, j) (or the first unavailable line j, or the first unavailable row i) needs to be padded, and even if the second sample is available, the second sample is padded by the "corresponding sample of the first sample" (or "corresponding line of the first line", or "corresponding row of the first row") defined by the filter support in ALF (e.g., the corresponding sample is located at (m, n) (or the corresponding line n, or the corresponding row m), and the distance of (m, n) to the current sample (or the current line, or the current row) is the same).
[0324] 1. Signaling notifications indicating whether to perform loop filtering operations across tile boundaries (e.g., loop_filter_across_tiles_enabled_flag) may depend on whether and / or how video units (e.g., pictures) are segmented into tiles.
[0325] a. In one example, the loop_filter_across_tiles_enabled_flag is signaled only when a video unit is divided into more than one piece.
[0326] b. Alternatively, when there is only one tile in the video unit, skip the signaling notification for loop_filter_across_tiles_enabled_flag.
[0327] c. In one example, the loop_filter_across_tiles_enabled_flag is signaled only when the general constraint flag one_tile_per_pic_constraint_flag is equal to 0.
[0328] i. Alternatively, if the general constraint flag one_tile_per_pic_constraint_flag is equal to 1, then loop_filter_across_tiles_enabled_flag is inferred (or required) to be equal to 0.
[0329] 2. Signaling notifications indicating whether to perform loop filtering operations across slice boundaries (e.g., loop_filter_across_slices_enabled_flag) may depend on whether and / or how video units (e.g., pictures and / or subpictures) are segmented into slices.
[0330] a. In one example, if the video unit is only divided into a single slice, no signaling is sent to loop_filter_across_slices_enabled_flag.
[0331] b. In one example, if each subpicture is only segmented into one strip (e.g., single_slice_per_subpic_flag equals 1), no signaling is sent to loop_filter_across_slices_enabled_flag.
[0332] c. In one example, if each subpicture is segmented into only one strip (e.g., single_slice_per_subpic_flag equals 1), then the signaling notifies loop_filter_across_slices_enabled_flag.
[0333] d. In one example, if the image is segmented into strips in a non-rectangular manner (e.g., rect_slice_flag equals 0), the signaling notifies loop_filter_across_slices_enabled_flag.
[0334] e. In one example, if the image is segmented into strips in a rectangular manner (e.g., rect_slice_flag equals 0) and the number of strips is equal to 1 (e.g., num_slices_in_pic_minus1 equals 0), then no signaling is sent to loop_filter_across_slices_enabled_flag.
[0335] f. In one example, the loop_filter_across_slices_enabled_flag is signaled only when the general constraint flag one_slice_per_pic_constraint_flag is equal to 0.
[0336] i. Alternatively, if the general constraint flag one_slice_per_pic_constraint_flag is equal to 1, then loop_filter_across_slices_enabled_flag is inferred (or required) to be equal to 0.
[0337] 3. LMCS information can be communicated in the strip header and / or picture header signaling (e.g., LMCS usage indication, and / or luminance shaping usage, and / or the adaptation_parameter_set_id of the LMCS APS to be used and / or chroma residual scaling usage).
[0338] a. LMCS information may include a primary indication of whether LMCS is enabled, such as ph_lmcs_enabled_flag or slice_lmcs_enabled_flag.
[0339] b. LMCS information may include a second indication of LMCS parameters, such as ph_lmcs_aps_id and / or ph_chroma_residual_scale_flag.
[0340] c. LMCS information can be signaled in the strip header and image header.
[0341] d. LMCS information can be signaled in the picture header and strip header. When it exists in the strip header, the LMCS information in the picture header (if any) may be overwritten.
[0342] e.LMCS information can be signaled in the strip header or the image header, but not in both simultaneously.
[0343] f. In one example, a signaling notification syntax element can be included in the stripe header to indicate which LMCS APS is used for the current stripe.
[0344] g. In one example, syntax elements (such as lmcs_info_in_ph_flag) are signaled in higher-level video units such as SPS or PPS to indicate whether LMCS information is signaled in the picture header or the strip header.
[0345] i. A syntax element can control only the first instruction.
[0346] ii. Syntax elements can control only the second instruction.
[0347] iii. Syntax elements can control the first and second directives.
[0348] h. If the syntax element for LMCS information in the image header does not exist, it can be set to the default value.
[0349] i. If the syntax element for LMCS information in the strip header does not exist, it can be set to the default value.
[0350] i. If the syntax element for LMCS information in the strip header does not exist, it can be set to the value of the corresponding syntax element for LMCS information in the image header.
[0351] j. In one example, more than one adaptation_parameter_set_id (e.g., a list of ph_lmcs_aps_id[]) of the LMCS APS referred to by the stripe associated with PH can be signaled in the picture header.
[0352] i. In one example, the length of the ph_lmcs_aps_id[] list can depend on the number of stripes in the image.
[0353] ii. In one example, a signaling notification syntax element can be included in the stripe header to indicate which LMCS APS in the ph_lmcs_aps_id[] list is used for the current stripe.
[0354] 4. The maximum permissible size of Transform Skip (TS) and / or BDPCM should be less than or equal to the size of the Codec Tree Block (CTB).
[0355] a. For example, the maximum permissible width and height of the TS and / or BDPCM of the luma block should be less than or equal to CtbSizeY.
[0356] b. For example, the maximum permissible width and height of the TS and / or BDPCM of the chroma block should be less than or equal to CtbSizeY / subWidthC and / or CtbSizeY / subHeightC.
[0357] i. Alternatively, the maximum permissible width of the TS and / or BDPCM of the chroma block should be less than or equal to CtbSizeY / subWidthC and / or CtbSizeY / subHeightC.
[0358] ii. Alternatively, the maximum permissible height of the TS and / or BDPCM of the chroma block should be less than or equal to CtbSizeY / subWidthC and / or CtbSizeY / subHeightC.
[0359] c. For example, it is required that log2_transform_skip_max_size_minus2 plus 2 should be less than or equal to CtbLog2SizeY.
[0360] d. For example, the maximum value of log2_transform_skip_max_size_minus2 is equal to CtbLog2SizeY–2.
[0361] e. For example, MaxTsSize is derived as
[0362] MaxTsSize=Min(CtbSizeY,1<<(log2_transform_skip_max_size_minus2+2)).
[0363] f. For example, MaxTsSize is derived as
[0364] MaxTsSize=1< <Min(log2_transform_skip_max_size_minus2+2,CtbLog2SizeY)。
[0365] g. The maximum permissible block size of the TS and / or BDPCM of the chroma block should be less than or equal to the maximum transform block size of the chroma block.
[0366] i. In one example, assuming MaxTbSizeY represents the maximum transform size of the luma block, the maximum permissible width and height of the TS and / or BDPCM of the chroma block can be less than or equal to MaxTbSizeY / SubWidthC.
[0367] ii. In one example, assuming MaxTbSizeY represents the maximum transform size of the luma block, the maximum permissible width of the TS and / or BDPCM of the chroma block can be less than or equal to MaxTbSizeY / SubWidthC.
[0368] iii. In one example, assuming MaxTbSizeY represents the maximum transform size of the luma block, the maximum permissible height of the TS and / or BDPCM of the chroma block can be less than or equal to MaxTbSizeY / SubHeightC.
[0369] h. The maximum allowed transform skip (TS) block size can be specified using binary syntax elements (e.g.,
[0370] “0” represents 16, “1” represents 32) Signaling notification.
[0371] 5. The maximum allowed transform skip (TS) block size and / or the maximum allowed transform block size should not be less than the minimum decode block size.
[0372] 6. Whether and / or how signaling informs or interprets or limits the maximum permissible transform block size (e.g., MaxTbSizeY in JVET-Q2001-vB) may depend on the minimum permissible codec block size (e.g., MinCbSizeY in JVET-Q2001-vB).
[0373] a. In one example, it is required that MaxTbSizeY must be greater than or equal to MinCbSizeY.
[0374] i. In one example, when MinCbSizeY equals 64, the value of sps_max_luma_transform_size_64_flag should be equal to 1.
[0375] ii. In one example, when MinCbSizeY equals 64, sps_max_luma_transform_size_64_flag is not signaled and is inferred to be 1.
[0376] 7. Whether and / or how signaling informs or interprets or limits the maximum permissible size of TS and / or BDPCM codecs (e.g., MaxTsSize in JVET-Q2001-vB) may depend on the minimum permissible codec block size (e.g., MinCbSizeY in JVET-Q2001-vB).
[0377] a. In one example, it is required that MaxTsSize must be greater than or equal to MinCbSizeY.
[0378] b. In one example, MaxTsSize is required to be less than or equal to W, where W is an integer, such as 32.
[0379] i. For example, it is required that MaxTsSize must satisfy MinCbSizeY <= MaxTsSize <= W.
[0380] ii. For example, when TS and / or BDPCM encoding / decoding is enabled, MinCbSizeY should be less than or equal to X.
[0381] c. In one example, when TS and / or BDPCM are enabled (e.g., sps_transform_skip_enabled_flag equals 1), MaxTsSize must be greater than or equal to MinCbSizeY.
[0382] i. For example, when sps_transform_skip_enabled_flag equals 1, log2_transform_skip_max_size_minus2 should be greater than or equal to log2_min_luma_coding_block_size_minus2.
[0383] d. In one example, MaxTsSize = max(MaxTsSize, MinCbSizeY).
[0384] i. In another example, MaxTsSize = min(W, max(MaxTsSize, MinCbSizeY)), where W is an integer, such as 32.
[0385] e. In one example, when TS and / or BDPCM are enabled (e.g., sps_transform_skip_enabled_flag equals 1), MaxTsSize = max(MaxTsSize, MinCbSizeY).
[0386] i. In one example, when TS and / or BDPCM are enabled (e.g., sps_transform_skip_enabled_flag equals 1), MaxTsSize = min(W, max(MaxTsSize, MinCbSizeY)), where W is an integer, such as 32.
[0387] f. In one example, the signaling notification for MaxTsSize (e.g., log2_transform_skip_max_size_minus2 in JVET-Q2001-vB) may depend on MinCbSizeY.
[0388] i. In one example, the difference between log2_transform_skip_max_size_minus2 and log2_min_luma_coding_block_size_minus2 (represented as log2_diff_max_transform_skip_min_coding_block) can be signaled to indicate MaxTsSize.
[0389] 1) For example, MaxTsSize=1<<(MinCbLog2SizeY+log2_diff_max_trasform_skip_min_coding_block).
[0390] 2) For example, MaxTsSize = min(W, 1 << (MinCbLog2SizeY + log2_diff_max_trasform_skip_min_coding_block)), where W is an integer, such as 32.
[0391] 8. Whether and / or how signaling notifications or interpretations or restrictions on TS and / or BDPCM codecs (e.g., denoted as sps_transform_skip_enabled_flag in JVET-Q2001-vB) may depend on the minimum allowed codec block size (e.g., denoted as MinCbSizeY in JVET-Q2001-vB).
[0392] a. In one example, when MinCbSizeY equals 64, sps_transform_skip_enabled_flag is not signaled and is inferred to be 0.
[0393] b. In one example, when MinCbSizeY is greater than the maximum allowed size of TS and / or BDPCM (e.g., the maximum allowed size of TS and / or BDPCM is 32), sps_transform_skip_enabled_flag is not signaled and is inferred to be 0.
[0394] 9. Whether and / or how the signaling notification or interpretation of the codec tool X in the SPS / PPS / Picture header / strip header may depend on the minimum allowed codec block size (e.g., MinCbSizeY in JVET-Q2001-vB).
[0395] a. If the minimum allowed codec block size is greater than T, then the signaling instruction for the codec tool X in the SPS / PPS / Picture Header / Strip Header can be omitted and it is inferred that it is not used, where T is an integer such as 32.
[0396] b. If the minimum allowed codec block size is greater than T, then the codec tool X in the SPS / PPS / Picture Header / Strip Header must be indicated not to be used, where T is an integer such as 32.
[0397] c. The encoding / decoding tool X can be a combined inter-frame intra-frame prediction (CIIP).
[0398] d. The codec tool X can be a Multiple Transform Selection (MTS).
[0399] e. The encoding / decoding tool X can be Segment-Block Transform (SBT).
[0400] f. The encoding / decoding tool X can be Symmetric Motion Vector Difference (SMVD).
[0401] g. The encoding / decoding tool X can be BDOF.
[0402] h. The encoding / decoding tool X can be affine prediction.
[0403] i. The codec tool X can be the optical flow prediction refinement (PROF).
[0404] j. The codec tool X can be the decoder-side motion vector refinement (DMVR).
[0405] k. The codec tool X can be the bi-directional prediction of CU-level weights (BCW).
[0406] l. The codec tool X can be the Merge with motion vector difference (MMVD).
[0407] m. The codec tool X can be the geometric partition mode (GPM).
[0408] n. The coding tool X can be the intra-block copy (IBC).
[0409] o. The codec tool X can be the palette coding.
[0410] p. The codec tool X can be the adaptive color transform (ACT).
[0411] q. The codec tool X can be the joint Cb-Cr residual coding (JCCR).
[0412] r. The codec tool X can be the cross-component linear model prediction (CCLM).
[0413] s. The codec tool X can be the multi-reference line (MRL).
[0414] t. The codec tool X can be the matrix-based intra prediction (MIP).
[0415] u. The codec tool X can be the intra prediction within sub-partitions (ISP).
[0416] 10. When applying the dual-tree coding, the minimum allowed block size of the binary tree partition (e.g., MinBtSizeY in JVET-Q2001-vB) may be different for the luminance component and the chrominance component.
[0417] a. In one example, MinBtSizeY = 1 << MinBtLog2SizeY is the minimum allowed block size of the binary tree partition for the luminance component, and MinBtSizeC = 1 << MinBtLog2SizeC is the minimum allowed block size of the binary tree partition for the chrominance component, where MinBtLog2SizeY may not be equal to MinBtLog2SizeC.
[0418] i. MinBtLog2SizeY can be predicted by signaling MinCbLog2SizeY. For example, the difference between MinBtLog2SizeY and MinCbLog2SizeY can be signaled.
[0419] ii. MinBtLog2SizeC can be predicted by signaling MinCbLog2SizeY. For example, the difference between MinBtLog2SizeC and MinCbLog2SizeY can be signaled.
[0420] 11. When applying dual-tree coding / decoding, the minimum allowed block size for the ternary tree partition (e.g., MinTtSizeY in JVET-Q2001-vB) may be different for the luma component and the chroma component. MinCbSizeY = 1 << MinCbLog2SizeY represents the minimum allowed coding / decoding block size.
[0421] a. In one example, MinTtSizeY = 1 << MinTtLog2SizeY is the minimum allowed block size for the ternary tree partition of the luma component, and MinTtSizeC = 1 << MinTtLog2SizeC is the minimum allowed block size for the ternary tree partition of the chroma component, where MinTtLog2SizeY may not be equal to MinTtLog2SizeC.
[0422] i. MinTtLog2SizeY can be predicted from MinCbLog2SizeY. For example, the difference between MinTtLog2SizeY and MinCbLog2SizeY can be signaled.
[0423] ii. MinTtLog2SizeC can be predicted from MinCbLog2SizeY. For example, the difference between MinTtLog2SizeC and MinCbLog2SizeY can be signaled.
[0424] 12. The maximum number of sub-block based merge candidates (e.g., MaxNumSubblockMergeCand) is derived as the sum of a first variable and a second variable, where the first variable is equal to zero if affine prediction is disabled (e.g., sps_affine_enabled_flag is equal to 0), and the second variable depends on whether sub-block based TMVP (sbTMVP) is enabled.
[0425] a. In one example, the first variable represents the number of allowed affine merge candidates.
[0426] b. In one example, the second variable can be set to (sps_sbtmvp_enabled_flag&&ph_temporal_mvp_enable_flag).
[0427] c. In one example, the first variable is deduced as KS, where S is the value signaled by a syntax element (e.g., five_minus_num_affine_merge_cand), and K is a fixed value such as 4 or 5.
[0428] d. In one example, MaxNumSubblockMergeCand = 5 - Five_minus_max_num_affine_merge_cand + (sps_sbtmvp_enabled_flag && ph_temporal_mvp_enable_flag).
[0429] e. In one example, MaxNumSubblockMergeCand = 4 - four_minus_max_num_affine_merge_cand + (sps_sbtmvp_enabled_flag && ph_temporal_mvp_enable_flag).
[0430] f. In one example, MaxNumSubblockMergeCand = Min(W, MaxNumSubblockMergeCand), where W is a fixed value, such as 5.
[0431] g. In one example, the indication of the first variable (e.g., five_minus_max_num_affine_merge_cand or four_minus_max_num_affine_merge_cand) can be conditionally signaled.
[0432] i. In one example, signaling notification is only performed when sps_affine_enabled_flag equals 1.
[0433] ii. When it does not exist, five_minus_max_num_affine_merge_cand is inferred to be K (e.g., K = 4 or 5).
[0434] 13. The number of rows (lines) between the ALF virtual boundary and the CTB bottom boundary of the luminance and chrominance components; and / or the filtering intensity of the rows between the ALF virtual boundary and the CTB bottom boundary of the luminance component and the corresponding rows of the chrominance component; and / or the filling method of the luminance and chrominance samples in the same row; alignment in 4:2:2 and 4:4:4 cases.
[0435] a. In one example, for 4:2:2 / 4:4:4 chroma format video, the vertical (and / or horizontal) position of the ALF virtual boundary (VB) in the chroma ALF should be aligned with the vertical (and / or horizontal) position of the ALF virtual boundary in the luma ALF. Let vbPosY represent the vertical (and / or horizontal) position of the ALF VB in the luma ALF, and vbPosC represent the vertical (and / or horizontal) position of the ALF VB in the chroma ALF.
[0436] i. In one example, when the vertical position (vbPosY) of the ALF VB in the luminance ALF is equal to CtbSizeY-S, the vertical position vbPosC of the ALF VB of the chrominance component can be set to be equal to (CtbSizeY-S) / SubHeightC.
[0437] ii. In one example, when the horizontal position vbPosY of the ALF VB in the luminance ALF is equal to CtbSizeY – S, the horizontal position of the ALF VB of the chromaticity component is vbPosC = (CtbSizeY – S) / SubWidthC.
[0438] iii. In the example above, CtbSizeY specifies the luma codec tree block size for each CTU and SubHeightC and SubWidthC, as defined in Table 3-1. S is an integer, such as 4.
[0439] b. In one example, the fill method in the chroma ALF for row K (and / or column H) near the vertical (and / or horizontal) position of the ALF VB should be aligned with the fill method in the luminance ALF. Let Yc represent the vertical (or horizontal) chroma sample point position.
[0440] i. In one example, when Yc equals vbPosC (e.g., in...) Figure 14B When (in Chinese), the following can be applied:
[0441] 1) Rows that are unavailable above (or to the left) (e.g., K=2) can be filled. Alternatively, even if the corresponding K rows are available, the corresponding K rows below (or to the right) of these unavailable rows can also be filled.
[0442] a. In one example, the current row can be used to populate the K unavailable rows above it.
[0443] i. In one example, in Figure 14B In this context, C1, C2, and C3 above the current line can be set to equal C5, C6, and C5, respectively. Furthermore, C0 above the current line can be set to equal C6.
[0444] b. In one example, the current row can be used to fill the corresponding K rows.
[0445] i. In one example, in Figure 14B In this context, C3, C2, and C1 below the current line can be set to equal C5, C6, and C5, respectively. Furthermore, C0 below the current line can be set to equal C6.
[0446] 2) You can fill the K unavailable rows below (or to the right). Alternatively, you can also fill the K unavailable rows above (or to the left) the current row, even if the corresponding K rows are available.
[0447] ii. In one example, when Yc equals vbPosC-M (e.g., M = 1, 2) (e.g., in Figure 14A and Figure 14D (in Chinese), the following can be applied:
[0448] 1) You can fill the K unavailable rows below (or to the right) (e.g., K=1, 2). And even if the corresponding K rows are available, you can also fill the corresponding K rows above (or to the left) these unavailable rows.
[0449] a. In one example, when M equals 1, the current row can be used to fill the K unavailable rows below.
[0450] i. In one example, in Figure 14A In this context, C3, C2, and C1 below the current line can be set to equal C5, C6, and C5, respectively. Furthermore, C0 below the current line can be set to equal C6.
[0451] b. In one example, when M equals 1, the current row can be used to fill the corresponding K rows.
[0452] i. In one example, in Figure 14A In this context, C1, C2, and C3 above the current line can be set to equal C5, C6, and C5, respectively. Furthermore, C0 above the current line can be set to equal C6.
[0453] c. In one example, when M is greater than or equal to 2, the bottommost row above the ALF virtual boundary (e.g., vbPosC-1) can be used to fill the K unavailable rows below.
[0454] i. In one example, in Figure 14DIn the text, C0 below the current line can be set to equal C6.
[0455] d. In one example, when M is greater than or equal to 2, the corresponding K rows can be filled with the corresponding row of the bottom row above the ALF virtual boundary (e.g., vbPos–2*M+1).
[0456] i. In one example, in Figure 14D In the text, C0 above the current line can be set to equal C6.
[0457] iii. In one example, when Yc equals vbPosC+N (e.g., N=1) (e.g., in...) Figure 14C When (in Chinese), the following can be applied:
[0458] 1) You can fill the K unavailable rows above (or to the left) (e.g., K=1).
[0459] Even if the corresponding K rows are available, the corresponding K rows of these unavailable rows below (or to the right) the current row can be filled.
[0460] a. In one example, the current row can be used to populate the K unavailable rows above it.
[0461] i. In one example, in Figure 14C In the text, C0 above the current line can be set to equal C6.
[0462] b. In one example, the current row can be used to fill the corresponding row K.
[0463] i. In one example, in Figure 14C In the text, C0 below the current line can be set to equal C6.
[0464] c. In one example, the ALF filter strength of the M rows (or N columns) near the vertical (or horizontal) position of the ALF VB in the chroma ALF should be aligned with the filter strength in the luma ALF of the 4:2:2 / 4:4:4 chroma format video. The ALF filter strength of the luma ALF and chroma ALF is controlled by alfShiftY (e.g., in Table 45) and alfShiftC (e.g., in Table 46).
[0465] i. In one example, when Yc ≠ vbPosC - M (e.g., M = 0, 1), alfShiftC = T1; when Yc ≠ vbPosC - M, alfShiftC = T2.
[0466] 1) In one example, T1 = 10, T2 = 7.
[0467] ii. In one example, when Yc ≠ vbPosC + M (e.g., M = 0, 1), alfShiftC = T1; when Yc ≠ vbPosC + M, alfShiftC = T2.
[0468] 1) In one example, T1 = 10, T2 = 7.
[0469] 6. Example
[0470] 6.1. Example of loop_filter_across_tiles_enabled_flag signaling notification
[0471] 7.3.2.4 Image Parameter Set RBSP Syntax
[0472]
[0473] 6.2. Example 1 of loop_filter_across_slices_enabled_flag signaling notification
[0474] 7.3.2.4 Image Parameter Set RBSP Syntax
[0475]
[0476] 6.3. Example 2 of loop_filter_across_slices_enabled_flag signaling notification
[0477] 7.3.2.4 Image Parameter Set RBSP Syntax
[0478]
[0479] 6.4. Example #3 of loop_filter_across_slices_enabled_flag signaling notification
[0480] 7.3.2.4 Image Parameter Set RBSP Syntax
[0481]
[0482]
[0483] 6.5. Example 4 of loop_filter_across_slices_enabled_flag signaling notification
[0484] 7.3.2.4 Image Parameter Set RBSP Syntax
[0485]
[0486] 6.6. Implementation Example #1 of LMCS Information Signalling Notification
[0487] 7.3.2.4 Image Parameter Set RBSP Syntax
[0488]
[0489] 6.7. Example 2 of LMCS Information Signaling Notification
[0490] 7.3.2.7 Image Parameter Set (RBSP) Syntax
[0491]
[0492]
[0493] 6.8. Implementation Example #3 of LMCS Information Signalling Notification
[0494] 7.3.7.1 General Strip Header Syntax
[0495] 6.9. Example 4 of LMCS Information Signalling Notification
[0496] 7.3.2.7 Image Parameter Set (RBSP) Syntax
[0497]
[0498] 6.10. LMCS Information Signaling Notification Example #5
[0499] 7.3.7.1 General Strip Header Syntax
[0500]
[0501] 6.11. Example 6 of LMCS Information Signaling Notification
[0502] Example 1: When slice_lmcs_enabled_flag does not exist, it is inferred that it is equal to ph_lmcs_enabled_flag.
[0503] Example 2: When slice_lmcs_aps_id does not exist, it is inferred that it is equal to ph_lmcs_aps_id.
[0504] Example 3: When slice_chroma_residual_scale_flag does not exist, it is inferred that it is equal to ph_chroma_residual_scale_flag.
[0505] 6.12. Example 1 of Striped Information Signalling Notification
[0506] 7.3.2.4 Image Parameter Set RBSP Syntax
[0507]
[0508]
[0509] 6.13. Implementation Examples of Handling ALF Virtual Boundaries
[0510] 8.8.5.4 Encoding and Decoding Tree Block Filtering Process for Chroma Samples
[0511] …
[0512] – In Table 46 [[Table 45]], the vertical chromaticity sample point position offsets y1, y2 and the variable alfShiftC are specified based on the vertical chromaticity sample point position y and applyAlfLineBufBoundary.
[0513] – The derivation of variable curr is as follows:
[0514] curr = recPicture[h x ][v y (1504)
[0515] The derivation of the chroma filter coefficient array f[j] and the chroma clipping value array c[j] is as follows, where
[0516] j = 0..5:
[0517] f[j]=AlfCoeff C [slice_alf_aps_id_chroma][altIdx][j] (1505)
[0518] c[j] = AlfClip C [slice_alf_aps_id_chroma][altIdx][j] (1506)
[0519] The derivation of the variable sum is as follows:
[0520]
[0521] sum=curr+((sum+64)>>alfShiftC) (1508)
[0522] – The modified filter-reconstructed chroma image sample alfPicture[xCtbC+x][yCtbC+y] is derived as follows:
[0523] alfPicture[xCtbC+x][yCtbC+y]=Clip3(0,(1< <BitDepth)-1,sum) (1509)
[0524] Table 46 – Based on the vertical chromaticity sample point position y and applyAlfLineBufBoundary specifying y1, y2, and alfShiftC
[0525]
[0526] Figure 15 A block diagram of an example video processing system 1900 that can implement various technologies of this disclosure is shown. Various implementations may include some or all of the components of system 1900. System 1900 may include an input 1902 for receiving video content. The video content may be received in a raw or uncompressed format (e.g., 8 or 10-bit multi-component pixel values), or in a compressed or encoded format. Input 1902 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, Passive Optical Network (PON), etc., and wireless interfaces such as Wi-Fi or cellular interfaces.
[0527] System 1900 may include codec component 1904, which may implement various codec or encoding methods described in this disclosure. Codec component 1904 may reduce the average bit rate of the video from input 1902 to the output of codec component 1904 to produce a codec representation of the video. Therefore, codec techniques are sometimes referred to as video compression or video transcoding techniques. The output of codec component 1904 may be stored or transmitted via a communication connection as indicated by component 1906. The stored or communicated bitstream (or codec) representation of the video received at input 1902 may be used by component 1908 to generate pixel values or displayable video to be sent to display interface 1910. The process of generating a user-viewable video from the bitstream representation is sometimes referred to as video decompression. Furthermore, although some video processing operations are referred to as “codec” operations or tools, it should be understood that codec tools or operations are used at the encoder, and the corresponding decoding tools or operations that reverse the encoded result will be performed by the decoder.
[0528] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB), High Definition Multimedia Interface (HDMI), or DisplayPort. Examples of storage interfaces include SATA (Serial Advanced Technology Accessory), PCI, IDE, etc. The techniques described in this disclosure can be implemented in a variety of electronic devices, such as mobile phones, laptops, smartphones, or other devices capable of performing digital data processing and / or video display.
[0529] Figure 16 This is a block diagram of a video processing apparatus 3600. Apparatus 3600 can be used to implement one or more methods described in this disclosure. Apparatus 3600 can be located in a smartphone, tablet, computer, Internet of Things (IoT) receiver, etc. Apparatus 3600 may include one or more processors 3602, one or more memories 3604, and video processing hardware 3606. The processors (multiple) 3602 can be configured to implement one or more methods described in this disclosure. The memories (multiple memories) 3604 can be used to store data and code for implementing the methods and techniques described in this disclosure. The video processing hardware 3606 can be used in hardware circuitry to implement some of the techniques described in this disclosure.
[0530] Figure 18 This is a block diagram illustrating an example video codec system 100 that can utilize the technology disclosed herein.
[0531] like Figure 18 As shown, the video encoding / decoding system 100 may include a source device 110 and a target device 120. The source device 110 can generate encoded video data and may be referred to as a video encoding device. The target device 120 can decode the encoded video data generated by the source device 110 and may be referred to as a video decoding device.
[0532] The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.
[0533] Video source 112 may include, for example, a source such as a video capture device, an interface for receiving video data from a video content provider, and / or a computer graphics system for generating video data, or a combination of these sources. Video data may include one or more images. Video encoder 114 encodes the video data from video source 112 to generate a bitstream. The bitstream may include a sequence of bits forming a codec representation of the video data. The bitstream may include codec images and associated data. A codec image is a codec representation of an image. Associated data may include sequence parameter sets, image parameter sets, and other syntax structures. I / O interface 116 may include a modulator / demodulator (modem) and / or a transmitter. Encoded video data may be transmitted directly to target device 120 via network 130a through I / O interface 116. Encoded video data may also be stored on storage medium / server 130b for access by target device 120.
[0534] The target device 120 may include an I / O interface 126, a video decoder 124, and a display device 122.
[0535] I / O interface 126 may include a receiver and / or a modem. I / O interface 126 may acquire encoded video data from source device 110 or storage medium / server 130b. Video decoder 124 may decode the encoded video data. Display device 122 may display the decoded video data to a user. Display device 122 may be integrated with target device 120, or may be external to target device 120, which is configured to interface with an external display device.
[0536] The video encoder 114 and the video decoder 124 can operate according to video compression standards, such as the High Efficiency Video Coding (HEVC) standard, the Multi-Function Video Coding (VVM) standard, and other current and / or other standards.
[0537] Figure 19 This is a block diagram illustrating an example of a video encoder 200. The video encoder 200 can be... Figure 18 The video encoder 114 in the system 100 described herein.
[0538] The video encoder 200 can be configured to perform any or all of the technologies disclosed herein. Figure 19 In the example shown, the video encoder 200 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video encoder 200. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.
[0539] The functional components of the video encoder 200 may include a segmentation unit 201, a prediction unit 202 that may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205 and an intra-frame prediction unit 206, a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213 and an entropy coding unit 214.
[0540] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an intra-block copy (IBC) unit. The IBC unit can perform prediction in IBC mode, where at least one reference picture is the picture containing the current video block.
[0541] Furthermore, some components, such as the motion estimation unit 204 and the motion compensation unit 205, can be highly integrated, but for descriptive purposes... Figure 19 The examples are shown separately.
[0542] The segmentation unit 201 can segment an image into one or more video blocks. The video encoder 200 and the video decoder 300 can support various video block sizes.
[0543] The mode selection unit 203 can select one of the encoding / decoding modes (intra-frame or inter-frame) based, for example, on the error result, and provide the resulting intra-frame or inter-frame codec block to the residual generation unit 207 to generate residual block data and to the reconstruction unit 212 to reconstruct the codec block for use as a reference image. In some examples, the mode selection unit 203 can select an intra-frame and inter-frame prediction combined (CIIP) mode, where prediction is based on inter-frame prediction signals and intra-frame prediction signals. In the case of inter-frame prediction, the mode selection unit 203 can also select the resolution of the motion vector for the block (e.g., sub-pixel or integer pixel resolution).
[0544] To perform inter-frame prediction on the current video block, motion estimation unit 204 can generate motion information for the current video block by comparing one or more reference frames from buffer 213 with the current video block. Motion compensation unit 205 can determine the predicted video block for the current video block based on motion information and decoded samples from images other than those associated with the current video block from buffer 213.
[0545] The motion estimation unit 204 and the motion compensation unit 205 can perform different operations on the current video block, for example, depending on whether the current video block is in an I-band, P-band, or B-band.
[0546] In some examples, motion estimation unit 204 can perform unidirectional prediction on the current video block, and can search for reference images in list 0 or list 1 of reference video blocks for the current video block. Then, motion estimation unit 204 can generate a reference index indicating that the reference image in list 0 or list 1 contains the reference video block, and a motion vector indicating the spatial displacement between the current video block and the reference video block. Motion estimation unit 204 can output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current block based on the reference video block indicated by the motion information of the current video block.
[0547] In other examples, motion estimation unit 204 can perform bidirectional prediction on the current video block. Motion estimation unit 204 can search for reference images in list 0 of the current video block's reference video blocks, and can also search for reference images in list 1 of another reference video block of the current video block. Then, motion estimation unit 204 can generate reference indices indicating the reference images containing the reference video blocks in lists 0 and 1, and motion vectors indicating the spatial displacement between the reference video blocks and the current video block. Motion estimation unit 204 can output the reference index and motion vector of the current video block as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current video block based on the reference video blocks indicated by the motion information of the current video block.
[0548] In some examples, the motion estimation unit 204 can output a complete set of motion information for the decoder to use in the decoding process.
[0549] In some examples, motion estimation unit 204 may not output the complete set of motion information for the current video. Instead, motion estimation unit 204 may signal the motion information of the current video block by referring to the motion information of another video block. For example, motion estimation unit 204 may determine that the motion information of the current video block is sufficiently similar to the motion information of neighboring video blocks.
[0550] In one example, the motion estimation unit 204 may instruct the video decoder 300 in the syntax structure associated with the current video block to indicate that the current video block has the same motion information as another video block.
[0551] In another example, motion estimation unit 204 can identify another video block and motion vector difference (MVD) within the syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. Video decoder 300 can use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[0552] As described above, the video encoder 200 can predictively signal motion vectors. Two examples of predictive signaling techniques that can be implemented by the video encoder 200 include Advanced Motion Vector Prediction (AMVP) and Combined Mode Signaling.
[0553] Intra-prediction unit 206 can perform intra-prediction on the current video block. When intra-prediction unit 206 performs intra-prediction on the current video block, it can generate prediction data for the current video block based on decoded samples from other video blocks in the same frame. The prediction data for the current video block may include the predicted video block and various syntax elements.
[0554] The residual generation unit 207 can generate residual data for the current video block by subtracting (e.g., indicated by a negative sign) multiple predicted video blocks from the current video block. The residual data for the current video block can include residual video blocks corresponding to different sample components of the samples in the current video block.
[0555] In other examples, for the current video block, such as in skip mode, there may not be residual data for the current video block, and the residual generation unit 207 may not perform the subtraction operation.
[0556] The transform processing unit 208 can generate one or more transform coefficient video blocks of the current video block by applying one or more transforms to the residual video blocks associated with the current video block.
[0557] After the transform processing unit 208 generates a transform coefficient video block associated with the current video block, the quantization unit 209 can quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0558] Inverse quantization unit 210 and inverse transform unit 211 can apply inverse quantization and inverse transform to the transform coefficient video block, respectively, to reconstruct the residual video block from the transform coefficient video block. Reconstruction unit 212 can add the reconstructed residual video block to the corresponding sample points of one or more predicted video blocks generated by prediction unit 202 to produce a reconstructed video block associated with the current block stored in buffer 213.
[0559] After the video block is reconstructed by the reconstruction unit 212, a loop filtering operation can be performed to reduce video block artifacts in the video block.
[0560] Entropy encoding unit 214 can receive data from other functional components of video encoder 200. When entropy encoding unit 214 receives data, it can perform one or more entropy encoding operations to generate entropy encoded data and output a bit stream including the entropy encoded data.
[0561] Some embodiments of the disclosed technology include making a decision or determination to enable a video processing tool or mode. In one example, when a video processing tool or mode is enabled, the encoder will use the tool or implement the mode in the processing of video blocks, but not necessarily based on the modification of the generated bitstream using the tool or mode. That is, when a decision or determination is made to enable a video processing tool or mode, the conversion from a video block to a bitstream representation of the video will use the video processing tool or mode. In another example, when a video processing tool or mode is enabled, the decoder will process the bitstream knowing that it has been modified based on the video processing tool or mode. That is, the conversion from the bitstream representation of the video to a video block will be performed using the video processing tool or mode enabled by the decision or determination.
[0562] Figure 20 This is a block diagram illustrating an example of a video decoder 300, which can be... Figure 18 The video decoder 114 in the system 100 shown.
[0563] The video decoder 300 can be configured to perform any or all of the technologies disclosed herein. Figure 20 In the example shown, the video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video decoder 300. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.
[0564] In such Figure 20 The example shown includes an entropy decoding unit 301, a motion compensation unit 302, an intra-frame prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, the video decoder 300 can perform functions typically associated with the video encoder 200. Figure 19 The encoding channel is the opposite of the decoding channel described.
[0565] The entropy decoding unit 301 can obtain the encoded bitstream. The encoded bitstream may include entropy-coded video data (e.g., encoded video data blocks). The entropy decoding unit 301 can decode the entropy-coded video data, and based on the entropy-decoded video data, the motion compensation unit 302 can determine motion information, including motion vectors, motion vector precision, reference image list index, and other motion information. For example, the motion compensation unit 302 can determine such information by performing AMVP and merging modes.
[0566] The motion compensation unit 302 can generate motion compensation blocks, possibly performing interpolation based on an interpolation filter. Identifiers for the interpolation filter at sub-pixel resolution can be included in the syntax elements.
[0567] The motion compensation unit 302 can use the interpolation filter used by the video encoder 20 during the encoding of the video block to calculate the interpolated values of the sub-integer pixels of the reference block. The motion compensation unit 302 can determine the interpolation filter used by the video encoder 200 based on the received syntax information and use the interpolation filter to generate the prediction block.
[0568] The motion compensation unit 302 may use some syntax information to determine the size of the blocks used to encode (multiple) frames and / or (multiple) stripes of the encoded video sequence, segmentation information describing how each macroblock of the image of the encoded video sequence is segmented, a mode indicating how to encode each segment, one or more reference frames (and a list of reference frames) for each inter-frame coded block, and other information for decoding the encoded video sequence.
[0569] Intra-prediction unit 303 can use, for example, an intra-prediction mode received in the bitstream to form prediction blocks from spatially adjacent blocks. Inverse quantization unit 303 performs inverse quantization, i.e., dequantization, on the quantized video block coefficients provided in the bitstream and decoded by entropy decoding unit 301. Inverse transform unit 303 applies an inverse transform.
[0570] The reconstruction unit 306 can add the residual block to the corresponding predicted block generated by the motion compensation unit 202 or the intra-frame prediction unit 303 to form a decoded block. If necessary, a deblocking filter can also be applied to filter the decoded block to remove block artifacts. The decoded video block is then stored in a buffer 307, which provides a reference block for subsequent motion compensation / intra-frame prediction and also generates the decoded video for presentation on the display device.
[0571] The following is a list of preferred solutions for some embodiments.
[0572] The following solutions show example embodiments of the techniques discussed in the preceding sections (e.g., items 1 and 2).
[0573] 1. A video processing method (e.g., Figure 17 The method 1700 shown includes: for the conversion between video units and codec representations of a video, determining (1702) whether an indication of the applicability of in-loop filtering across video regions of a video unit is included in the codec representation; and performing (1704) the conversion based on the determination.
[0574] 2. As in Solution 1, wherein the video unit comprises an image.
[0575] 3. The method of any one of solutions 1-2, wherein the video region includes slices.
[0576] 4. The method of any of solutions 1-2, wherein the video region includes stripes.
[0577] The following solutions show example embodiments of the techniques discussed in the preceding sections (e.g., item 3).
[0578] 5. A video processing method, comprising: performing a conversion between a video region of a video and a codec representation of the video, wherein the codec representation conforms to a format rule, wherein the format rule specifies that information regarding the applicability of a Luminance Mapping and Chroma Scaling (LMCS) tool to the conversion is indicated in the codec representation at the video strip level; wherein the LMCS tool includes constructing a current video block during the conversion based on a first domain and a second domain and / or scaling chroma residuals in a luminance-dependent manner.
[0579] The following solutions show example embodiments of the techniques discussed in the preceding sections (e.g., items 4, 5, 7, and 8).
[0580] 6. A video processing method comprising: performing a conversion between a video and a video codec representation, wherein the conversion conforms to a size rule specifying that: during encoding, the size rule is enforced on the maximum size of a video region encoded using a transform-skip codec scheme or an incremental pulse codec modulation codec scheme; or during decoding, the codec representation is parsed and decoded by enforcing the size rule on the maximum size of a video region decoded using a transform-skip codec scheme or an incremental pulse codec modulation codec scheme.
[0581] 7. As in Solution 6, where the size rule specifies that the maximum size of the video region as a transform skip block is less than or equal to the codec tree block size.
[0582] 8. As in Solution 6, where the size rule specifies that the maximum size of the video region processed using block-based incremental pulse mode encoding and decoding is less than or equal to the codec tree block size.
[0583] The following solutions show example embodiments of the techniques discussed in the preceding sections (e.g., item 5).
[0584] 9. As in Solution 6, where the size rule stipulates that the maximum size of the video region that serves as the transform skip block is less than or equal to the smallest decoding block size.
[0585] 10. As in Solution 6, where the size rule specifies that the maximum size of the video region processed using block-based incremental pulse mode encoding and decoding is less than or equal to the smallest decoding block size.
[0586] The following solutions show example embodiments of the techniques discussed in the preceding sections (e.g., item 7).
[0587] 11. The method of any one of solutions 6 to 10, wherein the size is indicated in a field in the codec representation, and wherein the minimum allowed codec block size for conversion controls the position of the field in the codec representation and / or how the size is interpreted from the field.
[0588] The following solutions show example embodiments of the techniques discussed in the preceding sections (e.g., item 8).
[0589] 12. The method of any one of solutions 6 to 10, wherein the minimum allowed codec block size control for conversion indicates whether a field of the size rule exists in the codec representation and / or how the size is interpreted from the field.
[0590] The following solutions show example embodiments of the techniques discussed in the preceding sections (e.g., item 6).
[0591] 13. A video processing method comprising: performing a conversion between a video region of a video and a codec representation of the video, wherein the codec representation conforms to a format rule specifying whether or how an indication of a minimum permissible transform block size controlled during the conversion is included in the codec representation.
[0592] 14. The method of Solution 13, wherein the format rules specify that the minimum allowed transform block size is greater than or equal to the minimum allowed coding block size.
[0593] The following solutions show example embodiments of the techniques discussed in the preceding sections (e.g., item 9).
[0594] 15. A video processing method comprising: performing a conversion between video units of a video and a codec representation of the video according to format rules, wherein the codec representation conforms to the format rules, the format rules specifying whether a minimum permissible codec block size control used during the conversion includes a field carrying information about the suitability of the codec tool in the conversion at the video region level.
[0595] 16. The method of Solution 15, wherein the video region corresponds to a sequence parameter set or a picture parameter set or a picture header or a strip header.
[0596] 17. The method of any one of solutions 15-16, wherein the encoding / decoding tools include combined inter-frame and intra-frame prediction tools.
[0597] 18. The method of any one of solutions 15-16, wherein the encoding / decoding tool includes a multi-transformation selection encoding / decoding tool.
[0598] The following solutions show example embodiments of the techniques discussed in the preceding sections (e.g., items 10 and 11).
[0599] 19. A video processing method comprising: a conversion between a video region of a video and a codec representation of the video, determining, by using a dual-tree codec for the conversion, a partitioning scheme for the luminance and chrominance components of the video having different minimum allowable block sizes for the luminance and chrominance components, and performing the conversion based on the determination.
[0600] 20. The method of Solution 19, wherein the partitioning scheme includes binary tree partitioning.
[0601] 21. The method of Solution 19, wherein the partitioning scheme includes ternary tree partitioning.
[0602] The following solutions show example embodiments of the techniques discussed in the preceding sections (e.g., item 12).
[0603] 22. A video processing method, comprising: performing a conversion between a video region and a codec representation of the video region by conforming to a format rule for codec representation; wherein the format rule specifies that the maximum number of sub-block-based merging candidates used during the conversion is derived as the sum of a first variable and a second variable, wherein the use of affine prediction controls the value of the first variable, and wherein the use of sub-block-based temporal motion vector prediction controls the value of the second variable.
[0604] 23. The approach of solution 22, where the first variable represents the number of affine merge candidates allowed.
[0605] The following solutions illustrate example embodiments of the techniques discussed in the previous section (e.g., item 13).
[0606] 24. A video processing method comprising: since the video is 4:2:2 or 4:4:4, performing a conversion between a video region of the video and a codec representation region of the video by conforming to processing rules applicable to the conversion, wherein the processing rules define chroma and luminance alignment for one or more of the following: (a) the number of pixel rows between the virtual boundary of an adaptive loop filter operation and the bottom boundary of a codec tree block; or (b) the filtering strength of a filter for rows between the virtual boundary of an adaptive loop filter operation and the bottom boundary of a codec tree block; or (c) a padding method for filling video samples in the same row.
[0607] 25. The method of solution 24, wherein the processing rules further define the vertical alignment of the virtual boundary between the chroma component and the luminance component.
[0608] 26. The method of Solution 24, wherein the processing rules define that the fill method for filling the chromaticity of row K and / or column H is aligned with the fill method for luminance.
[0609] 27. The method of any of the above solutions, wherein the video region includes a video encoding / decoding unit.
[0610] 28. The method of any of the above solutions, wherein the video region includes video images.
[0611] 29. The method of any one of solutions 1 to 28, wherein the conversion includes encoding the video into a codec representation.
[0612] 30. The method of any one of solutions 1 to 28, wherein the conversion includes decoding the encoding / decoding representation to generate pixel values of the video.
[0613] 31. A video decoding apparatus, comprising a processor configured to implement the methods of one or more of solutions 1 to 30.
[0614] 32. A video encoding apparatus, comprising a processor configured to implement the methods of one or more of solutions 1 to 30.
[0615] 33. A computer program product having computer code stored thereon, which, when executed by a processor, causes the processor to implement the method of any one of solutions 1 to 30.
[0616] 34. The methods, apparatus or systems described herein.
[0617] Figure 21 This is a flowchart of an example method (2100) for video processing. Operation 2102 includes performing a conversion between a video comprising video units and a bitstream of the video, wherein the bitstream conforms to a format rule, and wherein the format rule specifies that, in response to whether or how the video units are segmented, a first syntax element indicating whether a loop filtering operation is performed across segment boundaries is selectively included in the bitstream.
[0618] In some embodiments of method 2100, the video unit includes a picture. In some embodiments of method 2100, the first syntax element is in a picture parameter set. In some embodiments of method 2100, a format rule specifies that, in response to a video unit being segmented into more than one slice, the first syntax element is included in the bitstream. In some embodiments of method 2100, a format rule specifies that, in response to a video unit comprising only one slice, the first syntax element is not included in the bitstream. In some embodiments of method 2100, a format rule specifies that, in response to a bitstream including a zero value for a second syntax element indicating the presence of a slice per picture as a constraint, the first syntax element is included in the bitstream.
[0619] In some embodiments of method 2100, the format rule specifies that, in response to a bitstream including a value of a second syntax element indicating the presence of a tile per pic as a constraint, the first syntax element is required to be equal to zero and included in the bitstream. In some embodiments of method 2100, the second syntax element is one_tile_per_pic_constraint_flag. In some embodiments of method 2100, the loop filtering operation includes at least one of a deblocking filtering operation, a sample adaptive offset operation, or an adaptive loop filtering operation. In some embodiments of method 2100, the first syntax element is loop_filter_across_tiles_enabled_flag. In some embodiments of method 2100, performing the conversion includes encoding the video into a bitstream.
[0620] In some embodiments of method 2100, performing the conversion includes generating a bitstream from video, and the method further includes storing the bitstream in a non-transitory computer-readable recording medium. In some embodiments of method 2100, performing the conversion includes decoding video from the bitstream. In some embodiments, a video decoding apparatus includes a processor configured to implement the operations of method 2100 and related embodiments. In some embodiments, a video encoding apparatus includes a processor configured to implement the operations of method 2100 and related embodiments.
[0621] In some embodiments, a computer program product having computer instructions thereon, when executed by a processor, causes the processor to perform the method 2100 and operations related to the relevant embodiments. In some embodiments, a non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by a video processing apparatus, wherein the method includes: generating a bitstream from a video comprising video units; wherein the bitstream conforms to a format rule, and wherein the format rule specifies that, in response to whether or how the video units are fragmented, a first syntax element indicating whether a loop filtering operation is performed at fragment boundaries is selectively included in the bitstream. In some embodiments, a non-transitory computer-readable storage medium stores instructions causing a processor to implement the method 2100 and the operations of the relevant embodiments. In some embodiments, a bitstream generation method includes: generating a bitstream of video according to the operations of the method 2100 and the relevant embodiments, and storing the bitstream in a computer-readable program medium.
[0622] Figure 22 This is a flowchart of an example method (2200) for video processing. Operation 2202 includes performing a conversion between video units of the video and a bitstream of the video, wherein the bitstream conforms to a format rule, and wherein the format rule specifies that, in response to whether or how the video units are segmented into stripes, syntax elements indicating whether a loop filtering operation is performed across stripe boundaries are selectively included in the bitstream.
[0623] In some embodiments of method 2200, the video unit includes an image. In some embodiments of method 2200, the video unit includes sub-images. In some embodiments of method 2200, syntax elements are in an image parameter set. In some embodiments of method 2200, a rule specifies that, in response to a video unit that is segmented into only one strip, syntax elements are not included in the bitstream. In some embodiments of method 2200, a rule specifies that, in response to each sub-image of a video unit that is segmented into only one strip, syntax elements are not included in the bitstream. In some embodiments of method 2200, a rule specifies that, in response to each sub-image of a video unit that is segmented into only one strip, syntax elements are included in the bitstream.
[0624] In some embodiments of method 2200, the rule specifies that, in response to a first flag equal to 1 indicating that each subpicture consists of one and only one rectangular stripe, the syntax element is included in the bitstream. In some embodiments of method 2200, the first flag is single_slice_per_subpic_flag. In some embodiments of method 2200, the first flag is in the picture parameter set.
[0625] In some embodiments of method 2200, a rule specifies that, in response to video units comprising images segmented into non-rectangular shapes, syntax elements are included in the bitstream. In some embodiments of method 2200, a rule specifies that, in response to a second flag equal to 0 specifying that each image adopts a grid scan stripe pattern, syntax elements are included in the bitstream. In some embodiments of method 2200, the grid scan stripe pattern is a non-rectangular stripe pattern. In some embodiments of method 2200, the second flag is a rect_slice_flag. In some embodiments of method 2200, the second flag is in the image parameter set. In some embodiments of method 2200, a rule specifies that, in response to video units comprising images segmented into rectangular shapes and the number of stripes of the video units being equal to 1, syntax elements are not included in the bitstream.
[0626] In some embodiments of method 2200, a rule specifies that, in response to a bitstream including a zero value for a syntax element indicating whether each image has a stripe as a constraint, the syntax element is included in the bitstream. In some embodiments of method 2200, a rule specifies that, in response to a bitstream including a one value for a syntax element indicating whether each image has a stripe as a constraint, the syntax element is set to 0 and included in the bitstream. In some embodiments of method 2200, the loop filtering operation includes at least one of a deblocking filtering operation, a sample adaptive offset operation, or an adaptive loop filtering operation. In some embodiments of method 2200, performing the conversion includes encoding video into a bitstream. In some embodiments of method 2200, performing the conversion includes generating a bitstream from video, and the method further includes storing the bitstream in a non-transitory computer-readable recording medium.
[0627] In some embodiments of method 2200, performing the conversion includes decoding video from the bitstream. In some embodiments, the video decoding apparatus includes a processor configured to implement the operations of method 2200 and related embodiments. In some embodiments, the video encoding apparatus includes a processor configured to implement the operations of method 2200 and related embodiments. In some embodiments, a computer program product having computer instructions thereon, which, when executed by a processor, cause the processor to implement the operations of method 2200 and related embodiments. In some embodiments, a non-transitory computer-readable recording medium stores a bitstream of video generated by a method performed by a video processing apparatus, wherein the method includes: generating a bitstream from video units of the video; wherein the bitstream conforms to a format rule, and wherein the format rule specifies that, in response to whether or how the video units are segmented into stripes, syntax elements indicating whether a loop filtering operation is performed across stripe boundaries are selectively included in the bitstream.
[0628] In some embodiments, a non-transitory computer-readable storage medium stores instructions that cause a processor to implement the operations of method 2200 and related embodiments. In some embodiments, a bitstream generation method includes: generating a bitstream of video according to the operations of method 2200 and related embodiments, and storing the bitstream on a computer-readable program medium.
[0629] Figure 23 This is a flowchart of an example method (2300) for video processing. Operation 2302 includes performing a conversion between a video region of the video and a bitstream of the video, wherein the bitstream conforms to a format rule, and wherein the format rule specifies information about the suitability of a tool for the conversion in the bitstream at the video stripe level and / or at the video picture level, and wherein the tool maps luminance samples to specific values and selectively applies scaling operations to the values of chroma samples.
[0630] In some embodiments of method 2300, the formatting rule specifies that information regarding the applicability of the tool includes a first indication of whether the tool is enabled. In some embodiments of method 2300, the formatting rule specifies that information regarding the applicability of the tool includes a second indication of parameters, wherein the parameters include an identifier for an adaptive parameter set (APS) for the tool at the video picture level, and / or a value indicating whether chroma residual scaling is enabled at the video picture level. In some embodiments of method 2300, the formatting rule specifies that information regarding the applicability of the tool is indicated at both the video stripe level and the video picture level.
[0631] In some embodiments of method 2300, the format rules specify that information about the suitability of the tool is indicated at both the video strip level and the video picture level, and wherein prior information about the suitability of the tool included at the video picture level is overridden by information about the suitability of the tool included at the video strip level.
[0632] In some embodiments of method 2300, the formatting rules specify that information regarding the suitability of the tool is indicated at the video strip level or the video picture level. In some embodiments of method 2300, the formatting rules specify that information regarding the suitability of the tool is indicated at the video strip level to indicate the adaptive parameter set (APS) for the current strip used for the video region. In some embodiments of method 2300, the formatting rules specify that the bitstream includes syntax elements in the sequence parameter set (SPS) or picture parameter set (PPS) to indicate whether information regarding the suitability of the tool is indicated at the video strip level or the video picture level.
[0633] In some embodiments of method 2300, the syntax element controls only the first indication. In some embodiments of method 2300, the syntax element controls only the second indication. In some embodiments of method 2300, the syntax element controls both the first and second indications. In some embodiments of method 2300, the formatting rule specifies that, in response to information not present in the image header, information regarding the applicability of the tool is indicated in the syntax element of the image header with a default value. In some embodiments of method 2300, the formatting rule specifies that, in response to information not present in the strip header, information regarding the applicability of the tool is indicated in the syntax element of the strip header with a default value. In some embodiments of method 2300, wherein the formatting rule specifies that information regarding the applicability of the tool is indicated in the syntax element of the strip header, wherein the formatting rule specifies that the syntax element has a value for a corresponding syntax element that indicates information regarding the applicability of the tool in the image header, and wherein the formatting rule specifies that, in response to information not present in the strip header, the syntax element has a value for a corresponding syntax element.
[0634] In some embodiments of method 2300, a formatting rule specifies that, in response to a stripe associated with an identifier of a plurality of APSs, the identifiers of a plurality of APSs for which an adaptive parameter set (APS) is enabled are indicated in the image header. In some embodiments of method 2300, the length of the list of identifiers of the plurality of APSs depends on the number of stripes in the image. In some embodiments of method 2300, a formatting rule specifies that a syntax element is included in the stripe header, and wherein the syntax element indicates the APS enabled by the tool from the plurality of APSs for the current stripe. In some embodiments of method 2300, when the tool is enabled, if the video region originates from the luminance component, a switching of samples between the reconstructed domain and the original domain of the video region is performed, or wherein, when enabled, if the video region originates from the chroma component, scaling of the chroma residual of the video region is performed.
[0635] Figure 24 This is a flowchart of an example method (2400) for video processing. Operation 2402 includes performing a conversion between a video region of the video and the bitstream of the video, wherein the conversion conforms to a size rule, and wherein the size rule specifies the maximum size of the video region encoded using a Transform Skip (TS) codec scheme or a Block-Based Incremental Pulse Codec Modulation (BDPCM) codec scheme, or the maximum size of the transform block of the video region based on the codec characteristics of the video region.
[0636] In some embodiments of method 2400, the sizing rule specifies that the maximum size of the video region is less than or equal to the size of the codec block (CTB). In some embodiments of method 2400, the sizing rule specifies that the maximum permissible width and height for the luma block, used for a TS encoding or decoding scheme and / or for a BDPCM encoding or decoding scheme, is less than or equal to the CTB size. In some embodiments of method 2400, the sizing rule specifies that the maximum permissible width and height for the chroma block, used for a TS encoding or decoding scheme and / or for a BDPCM encoding or decoding scheme, is less than or equal to the CTB size divided by subWidthC and / or the CTB size divided by subHeightC, where subWidthC and subHeightC depend on the chroma format of the video. In some embodiments of method 2400, the sizing rule specifies that the maximum permissible width for the chroma block, used for a TS encoding or decoding scheme and / or for a BDPCM encoding or decoding scheme, is less than or equal to the CTB size divided by subWidthC and / or the CTB size divided by subHeightC.
[0637] In some embodiments of method 2400, the size rule specifies that the maximum permissible height for a chroma block in a TS encoding or decoding scheme and / or a BDPCM encoding or decoding scheme is less than or equal to the CTB size divided by subWidthC and / or the CTB size divided by subHeightC. In some embodiments of method 2400, a first value is equal to log2_transform_skip_max_size_minus2 plus 2, wherein the first value is less than or equal to a second value for CtbLog2SizeY, and wherein log2_transform_skip_max_size_minus2 plus 2 is equal to log2 of the maximum block size for the TS encoding scheme or the TS decoding scheme.
[0638] In some embodiments of method 2400, a first value describes that the maximum value for log2_transform_skip_max_size_minus2 is equal to a second value for CtbLog2SizeY minus 2, where log2_transform_skip_max_size_minus2 plus 2 is equal to the log2 of the maximum block size for a TS encoding scheme or for a TS decoding scheme. In some embodiments of method 2400, where a size rule specifies that the maximum size of a video region that is a transform skip block is the maximum of (CtbSizeY, 1<<(log2_transform_skip_max_size_minus2+2)), where << represents a left shift operation, where CtbSizeY is the CTB size, and where log2_transform_skip_max_size_minus2 plus 2 is equal to the log2 of the maximum block size for a TS encoding scheme or for a TS decoding scheme. In some embodiments of method 2400, the size rule specifies that the maximum size of a video region that is a transform skip block is 1<<Min(log2_transform_skip_max_size_minus2+2, CtbLog2SizeY), where << represents a left shift operation, where CtbSizeY is the CTB size, and where log2_transform_skip_max_size_minus2 plus 2 is equal to the log2 of the maximum block size for a TS encoding scheme or for a TS decoding scheme. In some embodiments of method 2400, the size rule specifies that the maximum size for a chrominance block for a TS encoding or decoding scheme and / or for a BDPCM encoding or decoding scheme is less than or equal to the size of the maximum transform block for a chrominance block.
[0639] In some embodiments of method 2400, the maximum permissible width and height for the TS encoding or decoding scheme and / or the BDPCM encoding or decoding scheme for the chroma block are less than or equal to the size of the maximum transform block for the luma block divided by SubWidthC. In some embodiments of method 2400, the maximum permissible width for the TS encoding or decoding scheme and / or the BDPCM encoding or decoding scheme for the chroma block are less than or equal to the size of the maximum transform block for the luma block divided by SubWidthC. In some embodiments of method 2400, the maximum permissible height for the TS encoding or decoding scheme and / or the BDPCM encoding or decoding scheme for the chroma block is less than or equal to the size of the maximum transform block for the luma block divided by SubHeightC. In some embodiments of method 2400, the sizing rule specifies that the maximum size of the video region serving as a transform skip block is indicated in the bitstream using binary syntax elements. In some embodiments of method 2400, the sizing rule specifies that the maximum size of the video region serving as a transform skip block and / or the maximum permissible transform block size is greater than or equal to the size of the smallest decoded block.
[0640] Figure 25 This is a flowchart of an example method (2500) for video processing. Operation 2502 includes performing a conversion between a video region of the video and a bitstream of the video, wherein the bitstream conforms to a format rule that specifies whether and / or how the minimum allowed codec block size used during the conversion controls whether the maximum allowed transform block size is included in the bitstream.
[0641] In some embodiments of method 2500, the format rule specifies that the size of the minimum allowed transform block is greater than or equal to the minimum allowed codec block size. In some embodiments of method 2500, in response to the minimum allowed codec block size being equal to 64, the value of sps_max_luma_transform_size_64_flag included in the bitstream is equal to 1. In some embodiments of method 2500, in response to the minimum allowed codec block size being equal to 64, the value of sps_max_luma_transform_size_64_flag is not included in the bitstream and is inferred to be equal to 1.
[0642] Figure 26 This is a flowchart of an example method (2600) for video processing. Operation 2602 includes performing a conversion between a video region of the video and a bitstream of the video, wherein the bitstream conforms to a format rule that specifies whether and / or how the minimum permissible codec block size used during the conversion controls whether and / or how the maximum size of the video region to be included in the bitstream for encoding or decoding using a Transform Skip (TS) scheme or a Block-Based Incremental Pulse Codec Modulation (BDPCM) scheme is indicated.
[0643] In some embodiments of method 2600, the format rule specifies that the maximum size of the video region is greater than or equal to the minimum allowed codec block size. In some embodiments of method 2600, the format rule specifies that the maximum size of the video region is less than or equal to W, where W is an integer. In some embodiments of method 2600, W is 32. In some embodiments of method 2600, the format rule specifies that the maximum size of the video region is greater than or equal to the minimum allowed codec block size, wherein the format rule specifies that the maximum size of the video region is less than or equal to W. In some embodiments of method 2600, the format rule specifies that when the TS scheme and / or BDPCM scheme are enabled, the minimum allowed codec block size is less than or equal to X. In some embodiments of method 2600, the format rule specifies that when the TS scheme and / or BDPCM scheme are enabled, the maximum size of the video region is greater than or equal to the minimum allowed codec block size.
[0644] In some embodiments of method 2600, a format rule specifies that when the TS scheme is enabled, log2 of the maximum size is greater than or equal to log2 of the size of the minimum luma codec block. In some embodiments of method 2600, a format rule specifies that the maximum size of the video region is the maximum of the maximum size or the minimum allowed codec block size. In some embodiments of method 2600, a format rule specifies that the maximum size of the video region is the minimum of a first value and a second value, wherein the first value is an integer W, and wherein the second value is the maximum of the maximum size or the minimum allowed codec block size. In some embodiments of method 2600, a format rule specifies that when the TS scheme and / or the BDPCM scheme are enabled, the maximum size of the video region is the maximum of the maximum size or the minimum allowed codec block size. In some embodiments of method 2600, a format rule specifies that when the TS scheme and / or the BDPCM scheme are enabled, the maximum size of the video region is the minimum of a first value and a second value, wherein the first value is an integer W, and wherein the second value is the maximum of the maximum size or the minimum allowed codec block size.
[0645] In some embodiments of method 2600, the format rule specifies that the maximum size of the video region is included in the bitstream based on the minimum allowed codec block size. In some embodiments of method 2600, the format rule specifies that the maximum size of the video region is indicated by including in the bitstream the difference between log2 of the maximum size of the video region and log2 of the minimum allowed codec block size. In some embodiments of method 2600, the format rule specifies: MaxTsSize = 1 << (MinCbLog2SizeY + log2_diff_max_trasform_skip_min_coding_block), where << denotes a left shift operation, where MaxTsSize is the maximum size of the video region, where MinCbLog2SizeY is log2 of the minimum codec block size, and where log2_diff_max_trasform_skip_min_coding_block is log2 of the difference between the maximum size of the video region and the minimum allowed codec block size in the bitstream.
[0646] In some embodiments of method 2600, the format rules specify:
[0647] MaxTsSize = min(W, 1 << (MinCbLog2SizeY + log2_diff_max_trasform_skip_min_coding_block)), where << indicates a left shift operation, W is an integer, MaxTsSize is the maximum size of the video region, MinCbLog2SizeY is log 2 of the smallest decoding unit size, and log2_diff_max_trasform_skip_min_coding_block is log2 of the difference between the maximum size of the video region in the bitstream and the minimum allowed codec block size.
[0648] Figure 27 This is a flowchart of an example method (2700) for video processing. Operation 2702 includes performing a conversion between video units of the video and a bitstream of the video according to format rules, wherein the bitstream conforms to the format rules, which specify whether and / or how to include fields in the bitstream indicating whether to enable or disable the Transform Skip (TS) scheme or the Block-Based Incremental Pulse Codec Modulation (BDPCM) scheme during the conversion.
[0649] In some embodiments of method 2700, the format rule specifies that, in response to a minimum allowed codec block size equal to 64, the field is inferred to be equal to 0 and is not included in the bitstream. In some embodiments of method 2700, the format rule specifies that, in response to a minimum allowed codec block size greater than the maximum allowed size of the TS codec scheme and / or the BDPCM codec scheme, the field is inferred to be equal to 0 and is not included in the bitstream.
[0650] Figure 28 This is a flowchart of an example method (2800) for video processing. Operation 2802 includes performing a conversion between video units and a video bitstream according to format rules, wherein the bitstream conforms to the format rules, which specify whether and / or how the minimum allowed codec block size used during the conversion controls whether and / or how the video region level in the bitstream includes fields carrying information about the suitability of the codec tools in the conversion.
[0651] In some embodiments of method 2800, the video region level corresponds to a sequence parameter set, a picture parameter set, a picture header, or a stripe header. In some embodiments of method 2800, a format rule specifies that, in response to a minimum allowed codec block size greater than an integer T, fields for codec tools are inferred to be unused or not included in the bitstream. In some embodiments of method 2800, a format rule specifies that, in response to a minimum allowed codec block size greater than an integer T, fields for codec tools indicate that the codec tool is not used or not included in the bitstream. In some embodiments of method 2800, the codec tool includes a combined inter-frame intra-frame prediction (CIIP) tool. In some embodiments of method 2800, the codec tool includes a multiple transform selection (MTS) codec tool. In some embodiments of method 2800, the codec tool includes a segment-block transform (SBT) codec tool. In some embodiments of method 2800, the codec tool includes a symmetric motion vector difference (SMVD) codec tool. In some embodiments of method 2800, the codec tool includes a bidirectional optical flow (BDOF) codec tool. In some embodiments of method 2800, the encoding / decoding tools include affine prediction encoding / decoding tools.
[0652] In some embodiments of method 2800, the encoding / decoding tool includes an optical flow prediction refinement (PROF) encoding / decoding tool. In some embodiments of method 2800, the encoding / decoding tool includes a decoder-side motion vector refinement (DMVR) encoding / decoding tool. In some embodiments of method 2800, the encoding / decoding tool includes a CU-level weighted bidirectional prediction (BCW) encoding / decoding tool. In some embodiments of method 2800, the encoding / decoding tool includes a merge encoding / decoding tool with motion vector difference (MMVD). In some embodiments of method 2800, the encoding / decoding tool includes a geometric partitioning mode (GPM) encoding / decoding tool. In some embodiments of method 2800, the encoding / decoding tool includes an intra-block copy (IBC) encoding / decoding tool. In some embodiments of method 2800, the encoding / decoding tool includes a palette encoding / decoding tool. In some embodiments of method 2800, the encoding / decoding tool includes an adaptive color transformation (ACT) encoding / decoding tool. In some embodiments of method 2800, the encoding / decoding tool includes a joint Cb-Cr residual encoding / decoding (JCCR) encoding / decoding tool. In some embodiments of method 2800, the encoding / decoding tools include cross-component linear prediction (CCLM) encoding / decoding tools. In some embodiments of method 2800, the encoding / decoding tools include multi-reference line (MRL) encoding / decoding tools. In some embodiments of method 2800, the encoding / decoding tools include matrix-based intra-frame prediction (MIP) encoding / decoding tools. In some embodiments of method 2800, the encoding / decoding tools include sub-segment intra-prediction (ISP) encoding / decoding tools.
[0653] Figure 29 This is a flowchart of an example method (2900) for video processing. Operation 2902 includes, for the conversion between a video region of the video and the bitstream of the video, determining that the partitioning schemes for the luminance and chrominance components of the video have different minimum allowed block sizes, since a dual-tree encoding / decoding scheme is enabled for the video region. Operation 2904 includes performing the conversion based on the determination.
[0654] In some embodiments of method 2900, the partitioning scheme includes binary tree partitioning. In some embodiments of method 2900, the first minimum allowable block size for binary tree partitioning of the luminance component is MinBtSizeY = 1 << MinBtLog2SizeY, where the second minimum allowable block size for binary tree partitioning of the chrominance component is MinBtSizeC = 1 << MinBtLog2SizeC, where << represents a left shift operation, where MinBtSizeY is the first minimum allowable block size, where MinBtLog2SizeY is the log2 of MinBtSizeY, where MinBtSizeC is the first minimum allowable block size, where MinBtLog2SizeC is the log2 of MinBtSizeC, and where MinBtLog2SizeY is not equal to MinBtLog2SizeC. In some embodiments of method 2900, MinBtLog2SizeY is signaled based on the minimum decoding unit size. In some embodiments of method 2900, the difference between MinBtLog2SizeY and the log2 of the minimum decoding unit size is included in the bitstream. In some embodiments of method 2900, MinBtLog2SizeC is signaled based on the minimum decoding unit size.
[0655] In some embodiments of method 2900, the difference between MinBtLog2SizeC and the log2 of the minimum decoding unit size is included in the bitstream. In some embodiments of method 2900, the partitioning scheme includes ternary tree partitioning. In some embodiments of method 2900, the minimum allowable coding / decoding block size MinCbSizeY is equal to 1 << MinCbLog2SizeY, where << represents a left shift operation, and where MinCbLog2SizeY is the log2 of the minimum decoding unit size. In some embodiments of method 2900, the first minimum allowable block size for ternary tree partitioning of the luminance component is MinTtSizeY = 1 << MinTtLog2SizeY, where the second minimum allowable block size for ternary tree partitioning of the chrominance component is MinTtSizeC = 1 << MinTtLog2SizeC, where << represents a left shift operation, where MinTtSizeY is the first minimum allowable block size, where MinTtLog2SizeY is the log2 of MinTtSizeY, where MinTtSizeC is the first minimum allowable block size, where MinTtLog2SizeC is the log2 of MinTtSizeC, and where MinTtLog2SizeY is not equal to MinTtLog2SizeC.
[0656] In some embodiments of method 2900, the signaling notification MinTtLog2SizeY is predicted based on the smallest decoding unit size. In some embodiments of method 2900, the difference between MinTtLog2SizeY and log2 of the smallest decoding unit size is included in the bitstream. In some embodiments of method 2900, the signaling notification MinTtLog2SizeC is predicted based on the smallest decoding unit size. In some embodiments of method 2900, the difference between MinTtLog2SizeC and log2 of the smallest decoding unit size is included in the bitstream.
[0657] Figure 30 This is a flowchart of an example method (3000) for video processing. Operation 3002 includes a conversion between a video region of the video and the video bitstream, determining the maximum number of allowed sub-block-based merging candidates for the video region based on rules. Operation 3004 includes performing the conversion based on the determination, wherein the rules specify that the maximum number of sub-block-based merging candidates used during the conversion can be derived as the sum of a first variable and a second variable, wherein the first variable is equal to zero in response to affine prediction being disabled, and wherein the second variable is based on whether sub-block-based temporal motion vector prediction (sbTMVP) is enabled.
[0658] In some embodiments of method 3000, the first variable represents the number of allowed affine merge candidates. In some embodiments of method 3000, the second variable is set to (sps_sbtmvp_enabled_flag&&ph_temporal_mvp_enable_flag). In some embodiments of method 3000, the first variable is deduced as KS, where S is the value of the syntax element signaling notification, and K is a fixed value. In some embodiments of method 3000, the maximum number of merge candidates based on subblocks is MaxNumSubblockMergeCand, where the first variable is five_minus_max_num_affine_merge_cand, and MaxNumSubblockMergeCand = 5 - five_minus_max_num_affine_merge_cand + (sps_sbtmvp_enabled_flag&&ph_temporal_mvp_enable_flag). In some embodiments of method 3000, the maximum number of merge candidates based on sub-blocks is MaxNumSubblockMergeCand, where the first variable is four_minus_max_num_affine_merge_cand, and where MaxNumSubblockMergeCand = 4 - four_minus_max_num_affine_merge_cand + (sps_sbtmvp_enabled_flag && ph_temporal_mvp_enable_flag). In some embodiments of method 3000, the maximum number of merge candidates based on sub-blocks is MaxNumSubblockMergeCand, and MaxNumSubblockMergeCand = Min(W, MaxNumSubblockMergeCand), where W is a fixed value. In some embodiments of method 3000, the first variable is conditionally signaled in the bitstream. In some embodiments of method 3000, the first variable is signaled in the bitstream in response to sps_affine_enabled_flag equaling 1. In some embodiments of method 3000, in response to the first variable not existing in the bitstream, five_minus_max_num_affine_merge_cand is inferred to be K.
[0659] Figure 31This is a flowchart of an example method (3100) for video processing. Operation 3102 includes performing a conversion between a video region of the video and a bitstream of the video region by conforming to processing rules, wherein, since the video is a 4:2:2 video or a 4:4:4 video, the processing rules are applicable to the conversion, wherein the processing rules define chroma and luma alignment of one or more of the following: (a) the number of pixel rows between the virtual boundary of the Adaptive Loop Filter (ALF) operation and the bottom boundary of the Codec Tree Block (CTB); or (b) the filter strength of the filter used for the rows between the virtual boundary of the ALF operation and the bottom boundary of the CTB; or (c) the padding method used to fill luma and chroma samples in the same row.
[0660] In some embodiments of method 3100, the processing rules specify that the vertical and / or horizontal positions are aligned with the virtual boundaries between the chroma and luma components used for ALF operations. In some embodiments of method 3100, the processing rules specify that, in response to the vertical position (vbPosY) of the virtual boundary for the luma component being equal to (CtbSizeY-S), the vertical position vbPosC of the virtual boundary for the chroma component being equal to (CtbSizeY-S) / SubHeightC. In some embodiments of method 3100, the processing rules specify that, in response to the horizontal position (vbPosY) of the virtual boundary for the luma component being equal to (CtbSizeY-S), the horizontal position (vbPosC) of the virtual boundary for the chroma component being equal to (CtbSizeY-S) / SubWidthC. In some embodiments of method 3100, CtbSizeY is the luma CTB size of each code-decode tree unit (CTU), S is an integer, and SubHeightC and / or SubWidthC have values of 1 or 2. In some embodiments of method 3100, the processing rule specifies that the filling method for the K rows and / or H columns of chroma components near the vertical and / or horizontal positions used to fill the virtual boundary is aligned with the filling method used for the luma component, where Yc is the vertical or horizontal chroma sample position. In some embodiments of method 3100, Yc is equal to the vertical or horizontal position of the virtual boundary of the chroma component.
[0661] In some embodiments of method 3100, the top or left K unavailable rows are filled, or K rows of the corresponding K unavailable rows below or to the right of the virtual boundary are filled, or K rows of the corresponding K available rows below or to the right of the virtual boundary are filled. In some embodiments of method 3100, the top K unavailable rows are filled using a virtual boundary. In some embodiments of method 3100, the sample points of the first row are located immediately above the virtual boundary, the sample points of the second row immediately above the first row are set to be equal to the sample points of the first row, and the sample points of the third row, two rows above the first row, are set to be equal to the sample points of the first row. In some embodiments of method 3100, K rows are filled using a virtual boundary. In some embodiments of method 3100, the sample points of the first row are located immediately below the virtual boundary, the sample points of the second row immediately below the first row are set to be equal to the sample points of the first row, and the sample points of the third row, which is two rows below the first row, are set to be equal to the sample points of the first row.
[0662] In some embodiments of method 3100, the K unavailable rows below or to the right are filled, or K rows of the corresponding K unavailable rows above or to the left of the virtual boundary are filled, or K rows of the corresponding K available rows above or to the left of the virtual boundary are filled. In some embodiments of method 3100, Yc is equal to the vertical or horizontal position of the virtual boundary for the chroma component minus M, where M is an integer. In some embodiments of method 3100, the K unavailable rows below or to the right are filled, and K rows of the corresponding K unavailable rows above or to the left of the virtual boundary are filled, or K rows of the corresponding K available rows above or to the left of the virtual boundary are filled. In some embodiments of method 3100, in response to M equaling 1, the K unavailable rows below are filled with the virtual boundary. In some embodiments of method 3100, the sample points of the first row are located immediately below the virtual boundary, the sample points of the second row immediately below the first row are set to be equal to the sample points of the first row, and the sample points of the third row, two rows below the first row, are set to be equal to the sample points of the first row. In some embodiments of method 3100, in response to M equaling 1, the virtual boundary is used to fill either the corresponding K unavailable rows or the corresponding K available rows. In some embodiments of method 3100, the sample points of the first row are located immediately above the virtual boundary, the sample points of the second row immediately above the first row are set to be equal to the sample points of the first row, and the sample points of the third row, two rows above the first row, are set to be equal to the sample points of the first row.
[0663] In some embodiments of method 3100, in response to M being greater than or equal to 2, the bottom row above the virtual boundary is used to fill the K unavailable rows below. In some embodiments of method 3100, the first row of samples is located two rows above the virtual boundary, and the sample located immediately below the virtual boundary in the second row is equal to the sample from the first row. In some embodiments of method 3100, in response to M being greater than or equal to 2, the corresponding row of the bottom row above the virtual boundary is used to fill either the corresponding K unavailable rows or the corresponding K available rows. In some embodiments of method 3100, the first row of samples is located two rows above the virtual boundary, and the sample located two rows above the first row in the second row is equal to the sample from the first row. In some embodiments of method 3100, Yc is equal to the vertical or horizontal position of the virtual boundary of the chroma component plus N, where N is an integer. In some embodiments of method 3100, the top or left K unavailable rows are filled, and K rows of the corresponding K unavailable rows below or to the right of the virtual boundary are filled, or K rows of the corresponding K available rows below or to the right of the virtual boundary are filled. In some embodiments of method 3100, the top K unavailable rows are filled using a virtual boundary.
[0664] In some embodiments of method 3100, the samples in the first row are located two rows below the virtual boundary, and the samples in the second row immediately above the virtual boundary are equal to the samples from the first row. In some embodiments of method 3100, the virtual boundary is used to fill either the corresponding K unavailable rows or the corresponding K available rows. In some embodiments of method 3100, the samples in the first row are located two rows below the virtual boundary, and the samples in the second row two rows below the first row are equal to the samples from the first row. In some embodiments of method 3100, the processing rule specifies that the first filter intensity of M rows or N columns for the chroma components near the vertical or horizontal position of the virtual boundary is aligned with the second filter intensity for the luminance component, wherein the second filter intensity for the luminance component and the first filter intensity for the chroma component are controlled by alfShiftY and alfShiftC, respectively, and where Yc is the vertical or horizontal chroma sample position. In some embodiments of method 3100, when Yc == vbPosC – M, alfShiftC = T1, and when Yc… When Yc ≠ vbPosC – M, alfShiftC = T2, where vbPosC is the horizontal position of the virtual boundary for the chromaticity component. In some embodiments of method 3100, T1 = 10 and T2 = 7. In some embodiments of method 3100, when Yc ≠ vbPosC + M (e.g., M = 0, 1), alfShiftC = T1, and when Yc ≠ vbPosC + M, alfShiftC = T2, where vbPosC is the horizontal position of the virtual boundary for the chromaticity component. In some embodiments of method 3100, T1 = 10 and T2 = 7.
[0665] In some embodiments of methods 2300-3100, performing the conversion includes encoding video into a bitstream. In some embodiments of methods 2300-3100, performing the conversion includes generating a bitstream from video, and the method further includes storing the bitstream in a non-transitory computer-readable recording medium. In some embodiments of methods 2300-3100, performing the conversion includes decoding video from the bitstream. In some embodiments, a video decoding apparatus includes a processor configured to implement the operations of methods 2300-3100 and related embodiments. In some embodiments, a video encoding apparatus includes a processor configured to implement the operations of methods 2300-3100 and related embodiments. In some embodiments, a computer program product having computer instructions thereon causes the processor to implement the operations of methods 2300-3100 and related embodiments when the instructions are executed by the processor. In some embodiments, a non-transitory computer-readable recording medium stores a bitstream of video generated by a method performed by a video processing apparatus, the method comprising: generating a bitstream from video regions of the video, wherein the bitstream conforms to a format rule specifying that information regarding the suitability of a tool for conversion is indicated in the bitstream at the video stripe level and / or at the video picture level, and wherein the tool maps luminance samples to specific values and selectively applies scaling operations to the values of chrominance samples. In some embodiments, a non-transitory computer-readable storage medium storing instructions enables a processor to implement the operations of methods 2300-3100 and related embodiments. In some embodiments, a bitstream generation method includes: generating a bitstream of video according to the operations of methods 2300-3100 and related embodiments, and storing the bitstream on a computer-readable program medium. In some embodiments, a method, apparatus, and bitstream generated according to the methods disclosed herein or the systems described herein.
[0666] In this paper, the term "video processing" can refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm can be applied during the conversion from a pixel representation of a video to a corresponding bitstream representation, and vice versa. For example, as defined by the syntax, the bitstream representation of the current video block can correspond to bits that are co-located or distributed at different positions within the bitstream. For example, macroblocks can be encoded based on the error residual values of the transform and encoding / decoding, and bits can also be used in the header and other fields of the bitstream. Furthermore, as described in the solutions above, during the conversion process, the decoder can parse the bitstream based on this determination, knowing that certain fields may or may not be present. Similarly, the encoder can determine whether certain syntax fields are included or excluded, and generate the encoding / decoding representation accordingly by including or excluding syntax fields from the encoding / decoding representation.
[0667] The disclosures and other solutions, examples, embodiments, modules, and functional operations described in this patent document can be implemented in digital electronic circuits, or computer software, firmware, or hardware, including the structures disclosed herein and their structural equivalents, or combinations thereof. The disclosures and other embodiments herein can be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a tangible and non-volatile computer-readable medium for execution by a data processing apparatus or for controlling the operation of the data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a storage device, a material composition affecting machine-readable propagated signals, or one or more of these. The terms "data processing unit" or "data processing apparatus" include all means, devices, and machines for processing data, including, for example, programmable processors, computers, or multiprocessors or computer groups. In addition to hardware, the apparatus may also include code that creates an execution environment for a computer program, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, or combinations thereof. The propagated signals are artificially generated signals, such as machine-generated electrical, optical, or electromagnetic signals, which are generated to encode information for transmission to a suitable receiver device.
[0668] Computer programs (also known as programs, software, software applications, scripts, or code) can be written in any programming language (including compiled or interpreted languages) and can be deployed in any form, including as standalone programs or as modules, components, subroutines, or other units suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to that program, or in multiple coordinating files (e.g., a file storing one or more modules, subroutines, or portions of code). Computer programs can be deployed and executed on one or more computers located at a single site or distributed across multiple sites interconnected by a communication network.
[0669] The processing and logic flows described in this specification can be executed by one or more programmable processors that execute one or more computer programs to perform functions by manipulating input data and generating outputs. The processing and logic flows can also be executed by special-purpose logic circuitry, and the device can be implemented as special-purpose logic circuitry, such as an FPGA (Field-Programmable Gate Array) or an ASIC (Application-Specific Integrated Circuit).
[0670] For example, processors suitable for executing computer programs include general-purpose and special-purpose microprocessors, as well as one or more of any type of digital computer. Typically, the processor receives instructions and data from read-only memory or random access memory, or both. The basic components of a computer are a processor that executes instructions and one or more storage devices that store the instructions and data. Typically, a computer will also include one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, or receive data from or transfer data to one or more mass storage devices via operative coupling, or both. However, a computer does not necessarily have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable hard disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and memory may be supplemented by or incorporated into special-purpose logic circuitry.
[0671] While this patent document contains numerous details, it should not be construed as limiting the scope of any invention or claim, but rather as a description of features of specific embodiments of a particular invention. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various functions described in the context of a single embodiment may also be implemented individually in multiple embodiments, or in any suitable sub-combination. Furthermore, although the foregoing features may be described as functioning in certain combinations, or even initially claimed to be so, in certain circumstances, one or more features from a combination of claims may be removed from the combination, and a combination of claims may refer to a sub-combination or a variation of a sub-combination.
[0672] Similarly, although the operations are described in a specific order in the accompanying drawings, this should not be construed as requiring the specific order or sequence shown to perform such operations, or all the described operations, in order to obtain the desired result. Furthermore, the separation of various system components in the embodiments of this patent document should not be construed as requiring such separation in all embodiments.
[0673] Only some implementations and examples are described; other implementations, enhancements, and variations can be made based on the content described and illustrated in this patent document.
Claims
1. A video processing method, comprising: Perform the conversion between the video region of the video and the bitstream of the video. The bitstream conforms to the format rules, and The format rules specify that, in the bitstream, at the video stripe level and / or at the video picture level, Luminance Mapping and Chroma Scaling (LMCS) information should be indicated regarding the applicability of the tool to the conversion. The tool maps luminance samples to specific values and selectively applies scaling operations to the values of chrominance samples.
2. The method according to claim 1, wherein, The format rules specify that the information regarding the applicability of the tool includes a first indication of whether the tool is enabled.
3. The method according to claim 1, in, The format rules specify that the information regarding the applicability of the tool includes a second indication of the parameters, and The parameters include an identifier for the adaptive parameter set APS of the tool at the video image level, and / or a value indicating whether chroma residual scaling is enabled at the video image level.
4. The method according to claim 1, wherein, The formatting rules stipulate that information regarding the applicability of the tool is indicated at both the video strip level and the video picture level.
5. The method according to claim 1, in, The formatting rules specify that both the video strip level and the video picture level indicate the information regarding the applicability of the tool, and The prior information regarding the applicability of the tool when included at the video image level is overridden by the information regarding the applicability of the tool when included at the video strip level.
6. The method according to claim 1, wherein, The formatting rules specify that information regarding the applicability of the tool should be indicated at the video strip level or the video picture level.
7. The method according to claim 1, wherein, The format rules specify that information regarding the applicability of the tool is indicated at the video strip level to indicate which adaptive parameter set (APS) is used for the current strip of the video region.
8. The method according to claim 1, wherein, The format rules specify that the bitstream includes syntax elements in a sequence parameter set (SPS) or a picture parameter set (PPS) to indicate whether information regarding the applicability of the tool is indicated at the video strip level or the video picture level.
9. The method according to claim 8, wherein, The syntax element controls only the first instruction.
10. The method according to claim 8, wherein, The syntax element only controls the second instruction.
11. The method according to claim 8, wherein, The syntax element controls the first instruction and the second instruction.
12. The method according to claim 1, wherein, The formatting rules stipulate that, in response to the information not being present in the image header, the information regarding the applicability of the tool is indicated in the syntax elements of the image header having a default value.
13. The method according to claim 1, wherein, The formatting rules stipulate that, in response to the information not being present in the strip header, the information regarding the applicability of the tool is indicated in the syntax elements of the strip header having a default value.
14. The method according to claim 1, in, The formatting rules specify that the information regarding the applicability of the tool should be indicated in the syntax elements of the bar header. The formatting rules specify that the syntax element has a value for a corresponding syntax element that indicates the applicability of the tool in the image header, and The formatting rule specifies that, in response to the information not existing in the strip header, the syntax element has the value of the corresponding syntax element.
15. The method according to claim 1, wherein, The formatting rules specify that, in response to the strip reference multiple adaptive parameter set APS identifiers associated with the image header, multiple APS identifiers in the image header indicate the APS of the enabled tool.
16. The method according to claim 15, wherein, The length of the list of multiple APS identifiers depends on the number of stripes in the image.
17. The method according to claim 15, wherein, The formatting rules specify that a syntax element is included in the strip header, and wherein the syntax element indicates which enabled tool's APS is to be used for the current strip.
18. The method according to any one of claims 1-17, wherein, When the tool is enabled, if the video region is from the luminance component, the tool performs a conversion between samples in the reshaped domain and samples in the original domain of the video region, or if the tool is enabled, if the video region is from the chroma component, the tool performs scaling of the chroma residual of the video region.
19. The method according to claim 1, wherein, The format rules also specify the maximum size of the video region encoded using the Transform Skip TS encoding / decoding scheme or the Block-Based Incremental Pulse Codec Modulation (BDPCM) encoding / decoding scheme, or the maximum size of the transform block of the video region based on the encoding / decoding characteristics of the video region.
20. The method according to claim 19, wherein, The format rules stipulate that the maximum size of the video region is less than or equal to the size of the codec blocks (CTB).
21. The method according to claim 20, wherein, The format rules stipulate that the maximum permissible width and height of the luma block for the TS codec scheme and / or for the BDPCM codec scheme are less than or equal to the CTB size.
22. The method according to claim 20, wherein, The format rules stipulate that the maximum permissible width and height of the chroma block for the TS codec scheme and / or for the BDPCM codec scheme are less than or equal to the CTB size divided by subWidthC and / or the CTB size divided by subHeightC, wherein subWidthC and subHeightC depend on the chroma format of the video.
23. The method of claim 20, wherein, The format rules stipulate that the maximum permissible width for the chroma block used in the TS codec scheme and / or the BDPCM codec scheme is less than or equal to the CTB size divided by subWidthC and / or the CTB size divided by subHeightC.
24. The method of claim 20, wherein, The format rules stipulate that the maximum permissible height for the chroma block used in the TS codec scheme and / or the BDPCM codec scheme is less than or equal to the CTB size divided by subWidthC and / or the CTB size divided by subHeightC.
25. The method according to claim 20, wherein, The first value is equal to log2_transform_skip_max_size_minus2 plus 2, where the first value is less than or equal to the second value for CtbLog2SizeY, and where log2_transform_skip_max_size_minus2 plus 2 is equal to log2 of the maximum block size for the TS encoding scheme or the TS decoding scheme.
26. The method of claim 20, wherein, The first value describing the maximum value of log2_transform_skip_max_size_minus2 is equal to the second value of CtbLog2SizeY minus 2, where log2_transform_skip_max_size_minus2 plus 2 is equal to log2 of the maximum block size used for the TS encoding scheme or the TS decoding scheme.
27. The method according to claim 20, in, The format rules stipulate that the maximum size of the video region that serves as the transform skip block is the minimum of (CtbSizeY, 1 << (log2_transform_skip_max_size_minus2 + 2)). Where << represents a left shift operation. Where CtbSizeY is the CTB size, and Wherein, log2_transform_skip_max_size_minus2 plus 2 equals log2 of the maximum block size used for the TS encoding scheme or the TS decoding scheme.
28. The method according to claim 20, in, The format rules specify that the maximum size of the video region serving as a transform skip block is 1 < <Min(log2_transform_skip_max_size_minus2 + 2 , CtbLog2SizeY), Where << represents a left shift operation. Where CtbSizeY is the CTB size, and Wherein, log2_transform_skip_max_size_minus2 plus 2 equals log2 of the maximum block size used for the TS encoding scheme or the TS decoding scheme.
29. The method according to claim 20, wherein, The format rules stipulate that the maximum size of the chroma block used for the TS encoding / decoding scheme and / or the BDPCM encoding / decoding scheme is less than or equal to the size of the maximum transform block used for the chroma block.
30. The method according to claim 29, wherein, The maximum permissible width and height for the chroma block in the TS encoding / decoding scheme and / or the BDPCM encoding / decoding scheme is less than or equal to the size of the maximum transform block for the luma block divided by SubWidthC.
31. The method according to claim 29, wherein, The maximum permissible width for the chroma block in the TS encoding / decoding scheme and / or the BDPCM encoding / decoding scheme is less than or equal to the size of the maximum transform block for the luma block divided by SubWidthC.
32. The method according to claim 29, wherein, The maximum permissible height for the chroma block in the TS encoding / decoding scheme and / or the BDPCM encoding / decoding scheme is less than or equal to the size of the maximum transform block for the luma block divided by subHeightC.
33. The method according to claim 20, wherein, The format rules specify that the maximum size of the video region, which serves as a transform skip block, is indicated in the bitstream using binary syntax elements.
34. The method of claim 19, wherein the format rule specifies that the maximum size and / or the maximum allowed transform block size of the video region as a transform skip block is greater than or equal to the size of the smallest decoding block.
35. The method according to claim 1, in, The format rules also specify whether and / or how the minimum allowed codec block size used during the conversion controls whether and / or how the maximum allowed transform block size is included in the bitstream.
36. The method according to claim 35, wherein, The format rules stipulate that the minimum allowed transform block size is greater than or equal to the minimum allowed codec block size.
37. The method of claim 35, wherein, In response to the minimum allowed codec block size being equal to 64, the value of sps_max_luma_transform_size_64_flag in the bitstream is equal to 1.
38. The method according to claim 35, wherein, In response to the minimum allowed codec block size being equal to 64, the value of sps_max_luma_transform_size_64_flag is not included in the bitstream and is inferred to be equal to 1.
39. The method according to claim 1, wherein, The format rules also specify whether and / or how the minimum permissible codec block size control used during the conversion includes an indication of the maximum size of the video region to be encoded or decoded using the Transform Skip TS scheme or the Block-Based Incremental Pulse Codec Modulation (BDPCM) scheme in the bitstream.
40. The method according to claim 39, wherein, The format rules stipulate that the maximum size of the video region is greater than or equal to the minimum allowed codec block size.
41. The method according to claim 39, wherein, The format rules stipulate that the maximum size of the video region is less than or equal to W, where W is an integer.
42. The method according to claim 41, wherein, W is 32.
43. The method according to claim 41, wherein, The format rule specifies that the maximum size of the video region is greater than or equal to the minimum allowed codec block size, wherein the format rule specifies that the maximum size of the video region is less than or equal to W.
44. The method according to claim 41, wherein, The format rules stipulate that when the TS scheme and / or the BDPCM scheme are enabled, the minimum allowed codec block size is less than or equal to X.
45. The method according to claim 39, wherein, The format rules stipulate that when the TS scheme and / or the BDPCM scheme are enabled, the maximum size of the video region is greater than or equal to the minimum allowed codec block size.
46. The method according to claim 39, wherein, The format rule stipulates that when the TS scheme is enabled, the maximum size log2 is greater than or equal to the minimum luminance codec block size log2.
47. The method according to claim 39, wherein, The format rules specify that the maximum size of the video region is the maximum value between the maximum size and the minimum allowed codec block size.
48. The method according to claim 47, in, The format rules stipulate that the maximum size of the video region is the minimum of the first and second values. Wherein, the first value is an integer W, and Wherein, the second value is the maximum value of the maximum size or the minimum allowed codec block size.
49. The method according to claim 39, wherein, The format rules stipulate that when the TS scheme and / or the BDPCM scheme are enabled, the maximum size of the video region is the maximum value between the maximum size and the minimum allowed codec block size.
50. The method according to claim 49, in, The format rules stipulate that when the TS scheme and / or the BDPCM scheme are enabled, the maximum size of the video region is the minimum of the first and second values. Wherein, the first value is an integer W, and Wherein, the second value is the maximum value of the maximum size or the minimum allowed codec block size.
51. The method according to claim 39, wherein, The format rules stipulate that, based on the minimum allowed codec block size, the maximum size of the video region is included in the bitstream.
52. The method according to claim 51, wherein, The format rules specify that the maximum size of the video region is indicated by including the difference between log2 of the maximum size of the video region and log2 of the minimum allowed codec block size in the bitstream.
53. The method according to claim 52, wherein, The format rules stipulate that: MaxTsSize=1<<(MinCbLog2SizeY+log2_diff_max_trasform_skip_min_coding_block), Where << represents a left shift operation. Where MaxTsSize is the maximum size of the video region. Among them, MinCbLog2SizeY is the log 2 of the smallest decoding unit size, and Wherein, log2_diff_max_trasform_skip_min_coding_block is log2 of the difference between the maximum size of the video region in the bitstream and the minimum allowed codec block size.
54. The method according to claim 52, wherein, The format rules stipulate that: MaxTsSize=min(W,1<<(MinCbLog2SizeY+log2_diff_max_trasform_skip_min_coding_block)), Where << represents a left shift operation. Where W is an integer, Where MaxTsSize is the maximum size of the video region. Among them, MinCbLog2SizeY is the log 2 of the smallest decoding unit size, and Wherein, log2_diff_max_trasform_skip_min_coding_block is log2 of the difference between the maximum size of the video region in the bitstream and the minimum allowed codec block size.
55. The method according to claim 1, wherein, The format rules also specify whether and / or how the minimum allowed codec block size control used during the conversion includes a field in the bitstream indicating whether to enable or disable the transform skip TS scheme or the block-based incremental pulse codec modulation (BDPCM) scheme.
56. The method according to claim 55, wherein, The format rules stipulate that, in response to the minimum allowed codec block size being equal to 64, the field is inferred to be zero and is not included in the bitstream.
57. The method of claim 55, wherein, The format rules stipulate that, in response to the minimum allowed codec block size being greater than the maximum allowed size of the TS scheme and / or the BDPCM scheme, the field is inferred to be zero and is not included in the bitstream.
58. The method according to claim 1, wherein, The format rules also specify whether and / or how the minimum allowed codec block size used during the conversion controls whether to include a field carrying information about the suitability of the codec tool in the conversion at the video region level of the bitstream.
59. The method according to claim 58, wherein, The video region level corresponds to a sequence parameter set, an image parameter set, an image header, or a strip header.
60. The method according to claim 59, wherein, The format rules stipulate that, in response to the minimum allowed codec block size being greater than an integer T, the field used by the codec tool is inferred to be unused and not included in the bitstream.
61. The method according to claim 59, wherein, The format rules stipulate that, in response to the minimum allowed codec block size being greater than an integer T, the field for the codec tool indicates that the codec tool is not used and is included in the bitstream.
62. The method according to any one of claims 58-61, wherein, The encoding and decoding tools include the combined inter-frame and intra-frame prediction (CIIP) tool.
63. The method according to any one of claims 58-61, wherein, The encoding / decoding tools include the Multi-Transform Selection (MTS) encoding / decoding tool.
64. The method according to any one of claims 58-61, wherein, The encoding and decoding tools include the Segment-Block Transform (SBT) encoding and decoding tool.
65. The method according to any one of claims 58-61, wherein, The encoding and decoding tools include the Symmetric Motion Vector Difference (SMVD) encoding and decoding tool.
66. The method according to any one of claims 58-61, wherein, The encoding and decoding tools include bidirectional optical flow (BDOF) encoding and decoding tools.
67. The method according to any one of claims 58-61, wherein, The encoding and decoding tools include affine prediction encoding and decoding tools.
68. The method according to any one of claims 58-61, wherein, The encoding and decoding tools include the Optical Flow Prediction Refinement PROF encoding and decoding tool.
69. The method according to any one of claims 58-61, wherein, The encoding / decoding tools include a decoder-side motion vector refinement DMVR encoding / decoding tool.
70. The method according to any one of claims 58-61, wherein, The encoding and decoding tools include a CU-level weighted bidirectional prediction (BCW) encoding and decoding tool.
71. The method according to any one of claims 58-61, wherein, The encoding / decoding tools include the Merge encoding / decoding tool with motion vector difference.
72. The method according to any one of claims 58-61, wherein, The encoding and decoding tools include the Geometric Partition Mode (GPM) encoding and decoding tool.
73. The method according to any one of claims 58-61, wherein, The encoding / decoding tools include an intra-block copy IBC encoding / decoding tool.
74. The method according to any one of claims 58-61, wherein, The encoding / decoding tools include a palette encoding / decoding tool.
75. The method according to any one of claims 58-61, wherein, The encoding and decoding tools include the Adaptive Color Transformation (ACT) encoding and decoding tool.
76. The method according to any one of claims 58-61, wherein, The encoding and decoding tools include the Joint Cb-Cr Residual Encoding and Decoding Tool (JCCR).
77. The method according to any one of claims 58-61, wherein, The encoding and decoding tools include a cross-component linear model prediction (CCLM) encoding and decoding tool.
78. The method according to any one of claims 58-61, wherein, The encoding / decoding tools include the Multi-Reference Line (MRL) encoding / decoding tool.
79. The method according to any one of claims 58-61, wherein, The encoding and decoding tools include matrix-based intra-predictive MIP encoding and decoding tools.
80. The method according to any one of claims 58-61, wherein, The encoding / decoding tools include a sub-segment intra-predictive ISP encoding / decoding tool.
81. The method according to claim 1, further comprising: For the conversion between the video region of the video and the bitstream of the video, since a dual-tree encoding and decoding scheme is enabled for the video region, the partitioning schemes for the luminance component and chrominance component of the video are determined to have different minimum allowed block sizes.
82. The method according to claim 81, wherein, The partitioning scheme includes binary tree partitioning.
83. The method according to claim 82, in, The first minimum allowable block size for the binary tree partitioning of the luminance component is MinBtSizeY = 1 << MinBtLog2SizeY, where the second minimum allowable block size for the binary tree partitioning of the chrominance component is MinBtSizeC = 1 << MinBtLog2SizeC, where << represents a left shift operation, where MinBtSizeY is the first minimum allowable block size, where MinBtLog2SizeY is the log2 of MinBtSizeY, where MinBtSizeC is the first minimum allowable block size, where MinBtLog2SizeC is the log2 of MinBtSizeC, and where MinBtLog2SizeY is not equal to MinBtLog2SizeC.
84. The method according to claim 83, wherein, MinBtLog2SizeY is signaled based on a prediction of the minimum decoding unit size.
85. The method according to claim 84, wherein, The difference between MinBtLog2SizeY and the log2 of the minimum decoding unit size is included in the bitstream.
86. The method according to claim 83, wherein, MinBtLog2SizeC is signaled based on a prediction of the minimum decoding unit size.
87. The method according to claim 86, wherein, The difference between MinBtLog2SizeC and the log2 of the minimum decoding unit size is included in the bitstream.
88. The method according to claim 81, wherein, The partitioning scheme includes ternary tree partitioning.
89. The method according to claim 88, wherein, The minimum allowable coding / decoding block size MinCbSizeY is equal to 1 << MinCbLog2SizeY, where << represents a left shift operation, and where MinCbLog2SizeY is the log2 of the minimum decoding unit size.
90. The method according to claim 88, in, The first minimum allowable block size for the ternary tree partitioning of the luminance component is MinTtSizeY = 1 << MinTtLog2SizeY, where the second minimum allowable block size for the ternary tree partitioning of the chrominance component is MinTtSizeC = 1 << MinTtLog2SizeC, where << represents a left shift operation, where MinTtSizeY is the first minimum allowable block size, where MinTtLog2SizeY is the log2 of MinTtSizeY, where MinTtSizeC is the first minimum allowable block size, where MinTtLog2SizeC is the log2 of MinTtSizeC, and where MinTtLog2SizeY is not equal to MinTtLog2SizeC.
91. The method according to claim 90, wherein, MinTtLog2SizeY is signaled based on a prediction of the minimum decoding unit size.
92. The method according to claim 91, wherein, The difference between MinTtLog2SizeY and the log2 of the minimum decoding unit size is included in the bitstream.
93. The method according to claim 90, wherein, MinTtLog2SizeC is signaled based on a prediction of the minimum decoding unit size.
94. The method according to claim 93, wherein, The difference between MinTtLog2SizeC and log2 of the smallest decoding unit size is included in the bitstream.
95. The method according to claim 1, further comprising: For the conversion between the video region of the video and the bitstream of the video, the maximum number of sub-block-based merging candidates allowed for the video region is determined based on rules; The rule stipulates that the maximum number of sub-block-based merge candidates used in the transformation is derived as the sum of the first and second variables. Wherein, in response to disabling affine prediction, the first variable equals 0, and The second variable is based on whether sub-block-based temporal motion vector prediction (sbTMVP) is enabled.
96. The method according to claim 95, wherein, The first variable represents the number of affine merge candidates allowed.
97. The method according to claim 95, wherein, The second variable is set to (sps_sbtmvp_enabled_flag && ph_temporal_mvp_enable_flag).
98. The method according to claim 95, wherein, The first variable is derived as KS, where S is the value of the syntax element signaling notification, and K is a fixed number.
99. The method according to claim 95, in, The maximum number of merge candidates based on sub-blocks is MaxNumSubblockMergeCand. Wherein, the first variable is five_minus_max_num_affine_merge_cand, and Among them, MaxNumSubblockMergeCand = 5 − five_minus_max_num_affine_merge_cand +( sps_sbtmvp_enabled_flag && ph_temporal_mvp_enable_flag ).
100. The method according to claim 95, in, The maximum number of merge candidates based on sub-blocks is MaxNumSubblockMergeCand. Wherein, the first variable is four_minus_max_num_affine_merge_cand, and Among them, MaxNumSubblockMergeCand = 4 − four_minus_max_num_affine_merge_cand +( sps_sbtmvp_enabled_flag && ph_temporal_mvp_enable_flag ).
101. The method according to claim 95, in, The maximum number of merge candidates based on subblocks is MaxNumSubblockMergeCand, and MaxNumSubblockMergeCand = Min( W, MaxNumSubblockMergeCand ), where W is a fixed number.
102. The method according to claim 95, wherein, The first variable is conditionally signaled in the bit stream.
103. The method according to claim 102, wherein, In response to sps_affine_enabled_flag equaling 1, the first variable is signaled in the bitstream.
104. The method according to claim 102, wherein, In response to the absence of the first variable in the bitstream, five_minus_max_num_affine_merge_cand is inferred to be K.
105. The method according to claim 1, further comprising: The conversion between video regions and their bitstreams is performed using processing rules, wherein the video is a 4:2:2 video or a 4:4:4 video. The processing rules define chroma and luminance to conform to one or more of the following: (a) The number of pixel rows between the virtual boundary of the Adaptive Loop Filtering (ALF) operation and the bottom boundary of the codec tree block (CTB); or (b) The filter strength of the filter used for the row between the virtual boundary of the ALF operation and the bottom boundary of the CTB; or (c) Filling method for filling luminance and chrominance samples in the same row.
106. The method according to claim 105, wherein, The processing rules specify that the vertical and / or horizontal positions are aligned with the virtual boundaries between the chroma and luminance components of the ALF operation.
107. The method according to claim 106, in, The processing rule stipulates that, in response to the vertical position vbPosY of the virtual boundary for the luminance component being equal to (CtbSizeY − S), the vertical position vbPosC of the virtual boundary for the chrominance component being equal to (CtbSizeY − S) / SubHeightC.
108. The method according to claim 106, in, The processing rule stipulates that, in response to the horizontal position vbPosY of the virtual boundary for the luminance component being equal to (CtbSizeY − S), the horizontal position vbPosC of the virtual boundary for the chrominance component being equal to (CtbSizeY − S) / SubWeightC.
109. The method according to claim 107 or 108, in, CtbSizeY is the luminance CTB size of each codec tree unit (CTU). Where S is an integer, The values of SubHeightC and / or SubWidthC are 1 or 2.
110. The method according to claim 105, in, The processing rules stipulate that the filling method for the K rows and / or H columns of the chromaticity components near the vertical and / or horizontal positions of the virtual boundary is aligned with the filling method for the luminance components, and Where Yc is the position of the vertical or horizontal chromaticity sample point.
111. The method according to claim 110, in, Yc is equal to the vertical or horizontal position of the virtual boundary of the chromaticity component.
112. The method according to claim 111, in, The top or left K unavailable rows are filled, or Wherein, the K rows corresponding to the K unavailable rows below or to the right of the virtual boundary are filled, or Among them, the K rows corresponding to the K available rows below or to the right of the virtual boundary are filled.
113. The method according to claim 112, wherein, Use the virtual boundary to fill the K unavailable rows above.
114. The method according to claim 113, in, The sample points in the first row are located immediately above and adjacent to the virtual boundary. Wherein, the sample points in the second row immediately above the first row are set to be equal to the sample points in the first row, and Specifically, the sample points in the third row, which is two rows above the first row, are set to be equal to the sample points in the first row.
115. The method according to claim 112, wherein, The K rows are filled using the virtual boundaries.
116. The method according to claim 115, in, The sample points in the first row are located immediately below and adjacent to the virtual boundary. Wherein, the sample points in the second row immediately below the first row are set to be equal to the sample points in the first row, and Specifically, the sample points in the third row, which is two rows below the first row, are set to be equal to the sample points in the first row.
117. The method according to claim 110, in, The K unavailable rows below or to the right are filled, or Among them, the K rows corresponding to the K unavailable rows above or to the left of the virtual boundary are filled, or Among them, the K rows corresponding to the K available rows above or to the left of the virtual boundary are filled.
118. The method according to claim 110, wherein, Yc is equal to the vertical or horizontal position of the virtual boundary used for the chromaticity component minus M, where M is an integer.
119. The method according to claim 118, in, The K unavailable rows below or to the right are filled, and Among them, the K rows corresponding to the K unavailable rows above or to the left of the virtual boundary are filled, or Among them, the K rows corresponding to the K available rows above or to the left of the virtual boundary are filled.
120. The method according to claim 119, wherein, In response to M equaling 1, the K unavailable rows below are filled with the virtual boundary.
121. The method according to claim 120, in, The sample points in the first row are located immediately below and adjacent to the virtual boundary. Wherein, the sample points in the second row immediately below the first row are set to be equal to the sample points in the first row, and Specifically, the sample points in the third row, which is two rows below the first row, are set to be equal to the sample points in the first row.
122. The method according to claim 119, wherein, In response to M equaling 1, the corresponding K unavailable rows or the corresponding K available rows are filled with the virtual boundary.
123. The method according to claim 122, in, The sample points in the first row are located immediately above and adjacent to the virtual boundary. Wherein, the sample points in the second row immediately above the first row are set to be equal to the sample points in the first row, and Specifically, the sample points in the third row, which is two rows above the first row, are set to be equal to the sample points in the first row.
124. The method according to claim 119, wherein, In response to M being greater than or equal to 2, the K unavailable rows below are filled with the bottommost row above the virtual boundary.
125. The method according to claim 124, in, The first row of sample points is located two rows above the virtual boundary. Among them, the sample points located immediately below the virtual boundary in the second row are equal to the sample points from the first row.
126. The method according to claim 119, wherein, In response to M being greater than or equal to 2, the corresponding K unavailable rows or the corresponding K available rows are filled with the corresponding row of the bottom row above the virtual boundary.
127. The method according to claim 126, in, The first row of sample points is located two rows above the virtual boundary. Wherein, the sample points in the second row located above the first row are equal to the sample points from the first row.
128. The method of claim 110, wherein, Yc is equal to the vertical or horizontal position of the virtual boundary of the chromaticity component plus N, where N is an integer.
129. The method according to claim 128, in, The top or left K unavailable rows are filled, and Wherein, the K rows corresponding to the K unavailable rows below or to the right of the virtual boundary are filled, or Among them, the K rows corresponding to the K available rows below or to the right of the virtual boundary are filled.
130. The method according to claim 129, wherein, Use the virtual boundary to fill the K unavailable rows above.
131. The method according to claim 130, in, The sample points in the first row are located two rows below the virtual boundary, and Among them, the sample points located immediately above the virtual boundary in the second row are equal to the sample points from the first row.
132. The method according to claim 129, wherein, The virtual boundaries are used to fill the corresponding K unavailable rows or the corresponding K available rows.
133. The method according to claim 132, in, The first row of sample points is located two rows below the virtual boundary, and Among them, the sample points in the second row, which is located two rows below the first row, are equal to the sample points from the first row.
134. The method according to claim 105, in, The processing rule stipulates that the first filter intensity of the M rows or N columns used for the chroma components near the vertical or horizontal position of the virtual boundary is aligned with the second filter intensity used for the luminance component. Wherein, the second filter intensity for the luminance component and the first filter intensity for the chrominance component are controlled by alfShiftY and alfShiftC, respectively, and Where Yc is the position of the vertical or horizontal chromaticity sample point.
135. The method according to claim 134, in, When Yc == vbPosC – M, alfShiftC = T1; when Yc != vbPosC – M, alfShiftC = T2. Wherein, vbPosC is the horizontal position of the virtual boundary for the chromaticity component.
136. The method of claim 135, wherein T1 = 10 and T2 = 7.
137. The method according to claim 134, in, When Yc == vbPosC + M, alfShiftC = T1; when Yc != vbPosC + M, alfShiftC = T2. Wherein, vbPosC is the horizontal position of the virtual boundary for the chromaticity component.
138. The method of claim 137, wherein T1 = 10 and T2 = 7.
139. The method according to claim 1, wherein, Performing the conversion includes encoding the video into the bitstream.
140. The method according to claim 1, wherein, Performing the conversion includes generating the bitstream from the video, and the method further includes storing the bitstream in a non-transitory computer-readable recording medium.
141. The method according to claim 1, wherein, Performing the conversion includes decoding the video into the bitstream.
142. A video decoding device, comprising a processor configured to implement the method according to any one of claims 1-139 and 141.
143. A video encoding device comprising a processor configured to implement the method according to any one of claims 1-140.
144. A computer program product having computer instructions thereon, wherein, When executed by a processor, the instructions cause the processor to perform the method according to any one of claims 1-141.
145. A non-transitory computer-readable storage medium having instructions stored thereon, wherein, The instructions cause the processor to implement the method according to any one of claims 1-141.
Citation Information
Patent Citations
Flexible range reduction
US20050063471A1