Using prediction mode information for adaptive loop filter in video coding
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-13
- Publication Date
- 2026-08-11
AI Technical Summary
随着能够接收和显示视频的连接用户设备的数量增加,对数字视频使用的带宽需求可能继续增长
[0114] For clarity, any of the embodiments described above may be combined with any one or more of the other embodiments described above to create new embodiments within the scope of this disclosure.
Smart Images

Figure CN122556087A_ABST
Abstract
Description
[0001] Cross-references to related applications
[0002] This patent application claims the benefits of International Patent Application No. PCT / CN2024 / 071764, filed on January 11, 2024; International Patent Application No. PCT / CN2024 / 075689, filed on February 4, 2024; and International Patent Application No. PCT / CN2024 / 079843, filed on March 4, 2024, each of which is incorporated herein by reference. Technical Field
[0003] This disclosure relates to the generation, storage, and use of digital audio and video media information in file formats. Background Technology
[0004] Digital video accounts for the largest share of bandwidth used on the internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth demand for digital video is likely to continue to grow. Summary of the Invention
[0005] The first aspect relates to a method for processing video data, comprising: determining a prediction pattern as side information for a loop filter; and performing a conversion between visual media data and a bitstream based on the loop filter.
[0006] Alternatively, in any of the above aspects, another embodiment of the aspect provides that the loop filter is a sample adaptive compensation (SAO) filter, a bilateral filter (BF), an adaptive loop filter (ALF), or a cross-component ALF (CCALF).
[0007] Alternatively, in any of the above aspects, another embodiment of the aspect provides that the prediction pattern of at least one location is stored before the loop filter is applied during the loop filtering process.
[0008] Alternatively, in any of the above aspects, another implementation of that aspect provides that a prediction pattern for at least one location is stored.
[0009] Alternatively, in any of the above aspects, another embodiment of that aspect provides that a prediction pattern of brightness position is stored.
[0010] Alternatively, in any of the above aspects, another implementation of that aspect provides that a prediction pattern of chromaticity position is stored.
[0011] Alternatively, in any of the above aspects, another implementation of that aspect provides that the stored prediction pattern is obtained by a loop filter.
[0012] Alternatively, in any of the foregoing aspects, another implementation of that aspect provides that the prediction pattern for each location is stored or used in the mapping domain.
[0013] Alternatively, in any of the foregoing aspects, another implementation of that aspect provides that the prediction pattern for each location is stored or used using mapping values representing different prediction patterns.
[0014] Optionally, in any of the above aspects, another implementation of that aspect provides pattern information including predictive patterns.
[0015] Optionally, in any of the above aspects, another implementation of that aspect provides mode information including codec block flags (CBF).
[0016] Optionally, in any of the above aspects, another implementation of the aspect provides mode information including a root codec block flag (rootCBF).
[0017] Alternatively, in any of the above aspects, another implementation of that aspect provides that the mapping value is predefined.
[0018] Alternatively, in any of the above aspects, another implementation of that aspect provides that the mapping value is derived.
[0019] Alternatively, in any of the foregoing aspects, another implementation of that aspect provides that the mapped value is included in the bitstream.
[0020] Alternatively, in any of the foregoing aspects, another implementation of that aspect provides that the location prediction pattern can be classified into one or more categories.
[0021] Alternatively, in any of the above aspects, another implementation of that aspect provides that the predicted pattern is classified into N categories, where N is a positive integer.
[0022] Alternatively, in any of the above aspects, another implementation of that aspect provides that the predicted pattern is classified according to predefined rules.
[0023] Alternatively, in any of the above aspects, another implementation of that aspect provides that locations with INTER modes are classified into a category.
[0024] Alternatively, in any of the foregoing aspects, another implementation of that aspect provides that locations having INTER mode and intra-block copy (IBC) mode are classified into a category.
[0025] Alternatively, in any of the above aspects, another implementation of that aspect provides that locations having INTER mode, intra-block copy (IBC) mode, and intra-template matching prediction (IntraTMP) are classified into a category.
[0026] Alternatively, in any of the above aspects, another implementation of that aspect provides that locations with INTRA patterns are classified into a category.
[0027] Alternatively, in any of the above aspects, another implementation of that aspect provides that locations with non-INTRA patterns are classified into a category.
[0028] Alternatively, in any of the above aspects, another implementation of that aspect provides that positions where the codec block flag (CBF) is equal to 0 are classified as a category.
[0029] Alternatively, in any of the above aspects, another implementation of that aspect provides that positions where the codec block flag (CBF) is not equal to 0 are classified into a category.
[0030] Alternatively, in any of the above aspects, another implementation of that aspect provides that block boundary locations are classified into different categories from other locations within a block.
[0031] Alternatively, in any of the above aspects, another implementation of that aspect provides that the number of merging rules or merging categories is predefined.
[0032] Alternatively, in any of the above aspects, another implementation of that aspect provides that the number of merging rules or merging categories is derived.
[0033] Alternatively, in any of the foregoing aspects, another implementation of that aspect provides that the number of merging rules or merging categories is included in the bitstream.
[0034] Alternatively, in any of the above aspects, another implementation of that aspect provides the use of the mapping value of the prediction mode as side information in other encoding and decoding stages.
[0035] Alternatively, in any of the foregoing aspects, another implementation of that aspect provides a mapping value of the prediction mode used in a bilateral filter.
[0036] Alternatively, in any of the foregoing aspects, another implementation of that aspect provides a mapping value of the prediction mode used in the classification of the bilateral filter.
[0037] Alternatively, in any of the above aspects, another implementation of that aspect provides mapping values for different prediction modes to indicate different filter strengths of the bilateral filter.
[0038] Alternatively, in any of the foregoing aspects, another implementation of that aspect provides a mapping value for the prediction mode used in a Hadamard transform domain filter (HTDF).
[0039] Alternatively, in any of the foregoing aspects, another implementation of that aspect provides a mapping value using a prediction pattern in the classification of HTDF.
[0040] Alternatively, in any of the above aspects, another implementation of that aspect provides mapping values for different prediction modes to indicate different filter strengths of the HTDF.
[0041] Alternatively, in any of the foregoing aspects, another implementation of that aspect provides a mapping value of the prediction mode used in a sample adaptive compensation (SAO) filter.
[0042] Alternatively, in any of the foregoing aspects, another implementation of that aspect provides a mapping value of the predicted mode used in the classification of the SAO filter.
[0043] Alternatively, in any of the foregoing aspects, another implementation of that aspect provides that the mapping value of the predicted pattern is used as an additional pattern for classification.
[0044] Alternatively, in any of the above aspects, another implementation of that aspect provides side information of existing patterns whose predicted pattern mapping values are used as classifications of SOA filters.
[0045] Alternatively, in any of the foregoing aspects, another implementation of that aspect provides a mapping value of the prediction mode when generating compensation for the SAO filter.
[0046] Alternatively, in any of the foregoing aspects, another implementation of that aspect provides a mapping value for the prediction mode used in a cross-component sample adaptive compensation (CCSAO) filter.
[0047] Alternatively, in any of the foregoing aspects, another implementation of that aspect provides a mapping value of the prediction mode used in the classification of the CCSAO filter.
[0048] Alternatively, in any of the foregoing aspects, another implementation of that aspect provides that the mapping value of the prediction mode is used as an additional mode for classification by the CCSAO filter.
[0049] Alternatively, in any of the foregoing aspects, another implementation of that aspect provides side information of existing patterns whose predicted pattern mapping values are used as classifications of the CCSAO filter.
[0050] Alternatively, in any of the foregoing aspects, another implementation of that aspect provides a mapping value of the prediction mode when generating compensation for the CCSAO filter.
[0051] Optionally, in any of the foregoing aspects, another implementation of that aspect provides that the mapping value of the prediction mode is used in other codec tools or other codec stages.
[0052] Optionally, in any of the foregoing aspects, another implementation of that aspect provides whether or how the mapping values of the prediction mode are prepared or modified are included in the bitstream.
[0053] Alternatively, in any of the foregoing aspects, another implementation of that aspect provides that a first syntax element is included in the bitstream to indicate whether a prepared or modified mapping value is used.
[0054] Alternatively, in any of the foregoing aspects, another implementation of that aspect provides that the first syntax element is encoded or decoded via arithmetic encoding / decoding.
[0055] Alternatively, in any of the foregoing aspects, another implementation of that aspect provides that the first syntax element is encoded or decoded using at least one context.
[0056] Alternatively, in any of the foregoing aspects, another implementation of that aspect provides at least one encoding / decoding information that depends on the current block or neighboring blocks.
[0057] Alternatively, in any of the foregoing aspects, another implementation of that aspect provides at least one context-dependent filtering shape of at least one neighboring block.
[0058] Alternatively, in any of the above aspects, another implementation of that aspect provides that the first syntax element is encoded and decoded using bypass encoding and decoding.
[0059] Optionally, in any of the above aspects, another implementation of the aspect provides that the first syntax element is binarized by unary code, rounded unary code, fixed-length code, exponential Golomb code, or rounded exponential Golomb code.
[0060] Alternatively, in any of the foregoing aspects, another implementation of that aspect provides that the first syntax element is conditionally included in the bitstream.
[0061] Alternatively, in any of the above aspects, another implementation of that aspect provides that the first syntax element is included in the bitstream only when the mapping value of the prediction mode is available.
[0062] Alternatively, in any of the above aspects, another implementation of that aspect provides that the first syntax element is encoded and decoded in a predictive manner.
[0063] Alternatively, in any of the above aspects, another implementation of the aspect provides that the first syntax element is predicted based on the on / off decision of the mapping value of the prediction mode of at least one neighboring block.
[0064] Alternatively, in any of the above aspects, another implementation of that aspect provides that, for different color components, the first syntax element is independently included in the bitstream.
[0065] Alternatively, in any of the above aspects, another implementation of that aspect provides that the first syntax element is shared for different color components.
[0066] Alternatively, in any of the foregoing aspects, another implementation of that aspect provides that, for the first color component but not for the second color component, the first syntax element is included in the bitstream.
[0067] Optionally, in any of the foregoing aspects, another implementation of that aspect provides that the first syntax element is included in a sequence parameter set (SPS), picture parameter set (PPS), picture header, strip header, adaptive parameter set (APS), codec tree unit (CTU), or codec unit (CU).
[0068] Alternatively, in any of the foregoing aspects, another implementation of that aspect provides that a first syntax element is included in the bitstream to indicate how the mapping value or boundary strength of the prediction mode is used.
[0069] Alternatively, in any of the foregoing aspects, another implementation of that aspect provides that a first syntax element is included in the bitstream to indicate the number of merged categories of the predicted pattern information or boundary strength information before being combined with existing classifications.
[0070] Alternatively, in any of the above aspects, another implementation of that aspect provides a number of merged categories of N, where N is a positive integer.
[0071] Alternatively, in any of the above aspects, another implementation of that aspect provides a number of merged categories of M, where M is a positive integer.
[0072] Alternatively, in any of the foregoing aspects, another implementation of that aspect provides that a first syntax element is included in the bitstream to indicate whether prediction mode information or boundary strength information is used in combination.
[0073] Alternatively, in any of the above aspects, another implementation of that aspect provides that the first syntax element is arithmetically encoded or decoded.
[0074] Alternatively, in any of the foregoing aspects, another implementation of that aspect provides that the first syntax element is encoded or decoded using at least one context.
[0075] Alternatively, in any of the foregoing aspects, another implementation of that aspect provides at least one encoding / decoding information that depends on the current block or neighboring blocks.
[0076] Alternatively, in any of the foregoing aspects, another implementation of that aspect provides at least one context-dependent filtering shape of at least one neighboring block.
[0077] Alternatively, in any of the above aspects, another implementation of that aspect provides that the first syntax element is encoded and decoded using bypass encoding and decoding.
[0078] Optionally, in any of the above aspects, another implementation of the aspect provides that the first syntax element is binarized by unary code, rounded unary code, fixed-length code, exponential Golomb code, or rounded exponential Golomb code.
[0079] Alternatively, in any of the foregoing aspects, another implementation of that aspect provides that the first syntax element is conditionally included in the bitstream.
[0080] Alternatively, in any of the above aspects, another implementation of that aspect provides that the first syntax element is included in the bitstream only when the mapping value of the prediction mode is available.
[0081] Alternatively, in any of the above aspects, another implementation of that aspect provides that the first syntax element is encoded and decoded in a predictive manner.
[0082] Alternatively, in any of the above aspects, another implementation of the aspect provides that the first syntax element is predicted based on the on / off decision of the mapping value of the prediction mode of at least one neighboring block.
[0083] Alternatively, in any of the above aspects, another implementation of that aspect provides that, for different color components, the first syntax element is independently included in the bitstream.
[0084] Alternatively, in any of the above aspects, another implementation of that aspect provides that the first syntax element is shared for different color components.
[0085] Alternatively, in any of the foregoing aspects, another implementation of that aspect provides that, for the first color component but not for the second color component, the first syntax element is included in the bitstream.
[0086] Optionally, in any of the foregoing aspects, another implementation of that aspect provides that the first syntax element is included in a sequence parameter set (SPS), picture parameter set (PPS), picture header, strip header, adaptive parameter set (APS), codec tree unit (CTU), or codec unit (CU).
[0087] Alternatively, in any of the foregoing aspects, another implementation of that aspect provides that the first syntax element is included in the bitstream at the SPS level or the PPS level.
[0088] Alternatively, in any of the foregoing aspects, another implementation of that aspect provides that the first syntax element is included in the bitstream at the CTU level.
[0089] Alternatively, in any of the foregoing aspects, another implementation of that aspect provides that the first syntax element is included in the bitstream at the APS level.
[0090] Alternatively, in any of the foregoing aspects, another implementation of that aspect provides that for all luminance and chrominance filters within all alternative filter sets, a first syntax element is included in the bitstream.
[0091] Alternatively, in any of the foregoing aspects, another implementation of that aspect provides that, for all luminance and chrominance filters within an alternative filter set, a first syntax element is included in the bitstream.
[0092] Alternatively, in any of the foregoing aspects, another implementation of that aspect provides that the method is used in post-processing or pre-processing.
[0093] Alternatively, in any of the foregoing aspects, another implementation of that aspect provides that the method is used in combination.
[0094] Alternatively, in any of the foregoing aspects, another implementation of that aspect provides that the method can be used alone.
[0095] Optionally, in any of the foregoing aspects, another embodiment of that aspect provides the use of the mapping values of the prediction mode for any loop filtering tool, preprocessing or postprocessing filtering method applied to video encoding and decoding, including ALF, cross-component ALF (CCALF), bilateral filter (BF), SAO, CCSAO or any other filtering method.
[0096] Optionally, in any of the foregoing aspects, another embodiment of that aspect provides a method for applying the mapping values of the prediction mode to a loop filter, or wherein the use of the mapping values of the prediction mode is applied to an ALF, CCALF, any other loop filter, preprocessing filter, or postprocessing filter.
[0097] Optionally, in any of the foregoing aspects, another embodiment of that aspect provides a loop filter applied to a video unit, wherein the video unit is a sequence, picture, sub-picture, strip, slice, CTU, CTU line, CTU group, CU, prediction unit (PU), transform unit (TU), codec tree block (CTB), codec block (CB), prediction block (PB), transform block (TB), or any other region comprising more than one luminance or chrominance sample or pixel.
[0098] Alternatively, in any of the foregoing aspects, another implementation of that aspect provides the use of the method in a bitstream via signal transmission.
[0099] Optionally, in any of the foregoing aspects, another embodiment of that aspect provides that the use of the method is transmitted via signaling at the sequence level, picture group level, picture level, strip level, or slice group level, including in the sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), decoding parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), strip header, or slice group header, or wherein the use of the method is transmitted via signaling at the prediction block (PB), transform block (TB), codec block (CB), prediction unit (PU), transform unit (TU), codec unit (CU), virtual pipeline decoding unit (VPDU), codec tree unit (CTU), CTU row, strip, slice, sub-picture, or other area including more than one sample or pixel.
[0100] Alternatively, in any of the foregoing aspects, another embodiment of that aspect provides that the application of the method depends on the encoded / decoded information, including block size, color format, single-tree segmentation, dual-tree segmentation, color components, stripe type, or image type.
[0101] Optionally, in any of the foregoing aspects, another implementation of that aspect provides that the first syntax element is binarized into a flag, a fixed-length code, an exponential Golomb code, a unary code, a rounded unary code, or a rounded binary code, wherein the first syntax element is signed or unsigned.
[0102] Alternatively, in any of the foregoing aspects, another implementation of that aspect provides that the first syntax element is encoded or decoded using a context model or is bypassed.
[0103] Optionally, in any of the above aspects, another embodiment of the aspect provides conditional transmission of the first syntax element via signal, transmission of the first syntax element via signal only when the corresponding function applies, transmission of the first syntax element via signal when the height or width of the block meets a condition, or transmission via signal when a combination of both is met.
[0104] Optionally, in any of the foregoing aspects, another embodiment of that aspect provides that syntax elements are transmitted via signals at the sequence level, picture group level, picture level, stripe level, or slice group level, in the sequence header, picture header, SPS, VPS, DPS, decoding capability information (DCI), PPS, APS, stripe header, or slice group header.
[0105] Optionally, in any of the foregoing aspects, another embodiment of that aspect provides that method is used in combination with affine, multiple transform selection (MTS), LFNST, Merge with motion vector difference (MMVD), matrix-based intra-prediction (MIP), ISP, cross-component linear model (CCLM), convolutional cross-component model (CCCM), symmetric motion vector difference (SMVD), bidirectional optical flow (BDOF), decoder-side motion vector refinement (DMVR), history-based motion vector prediction (HMVP), template matching, intra-block copy (IBC), or palette, or not in combination with affine, multiple transform selection (MTS), LFNST, Merge with motion vector difference (MMVD), matrix-based intra-prediction (MIP), ISP, cross-component linear model (CCLM), convolutional cross-component model (CCCM), symmetric motion vector difference (SMVD), bidirectional optical flow (BDOF), decoder-side motion vector refinement (DMVR), history-based motion vector prediction (HMVP), template matching, intra-block copy (IBC), or palette.
[0106] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the excluded tool is implicitly disabled without signaling, or that the method is implicitly disabled without signaling when the excluded tool is used.
[0107] Alternatively, in any of the foregoing aspects, another implementation of that aspect provides a conversion that includes encoding visual media data into a bitstream.
[0108] Alternatively, in any of the foregoing aspects, another implementation of that aspect provides a conversion that includes decoding visual media data from a bitstream.
[0109] The second aspect relates to an apparatus for processing video data, comprising: a processor; and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the methods of any of the disclosed embodiments.
[0110] The third aspect relates to a non-transitory computer-readable medium including a computer program product for use by a video codec apparatus, the computer program product including computer-executable instructions stored on the non-transitory computer-readable medium, causing the video codec apparatus to perform any of the methods of the disclosed embodiments when executed by a processor.
[0111] The fourth aspect relates to a non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by a video processing apparatus, wherein the method includes: determining a prediction mode as side information for a loop filter; and generating a bitstream based on the loop filter.
[0112] The fifth aspect relates to a method for storing a bitstream of video, comprising: determining a prediction mode as side information for a loop filter; generating a bitstream based on the determination; and storing the bitstream in a non-transitory computer-readable recording medium.
[0113] The sixth aspect relates to the methods, apparatus, or systems described in this disclosure.
[0114] For clarity, any of the embodiments described above may be combined with any one or more of the other embodiments described above to create new embodiments within the scope of this disclosure.
[0115] These and other features will become clearer from the following detailed description by referring to the accompanying drawings and claims. Attached Figure Description
[0116] To gain a more complete understanding of this disclosure, reference is now made to the following brief description, which is taken in conjunction with the accompanying drawings and detailed description, wherein the same reference numerals denote the same parts.
[0117] Figure 1 Examples of the nominal vertical and horizontal positions of 4:2:2 luminance and chrominance samples in an image are shown.
[0118] Figure 2A sample encoder block diagram is shown.
[0119] Figure 3 An example image is shown, segmented into raster scan strips.
[0120] Figure 4 An example image is shown, segmented into rectangular scan strips.
[0121] Figure 5 An example image showing the structure divided into bricks is shown.
[0122] Figures 6A-6C An example of a Coding Tree Block (CTB) spanning the boundaries of an image is shown.
[0123] Figure 7 An example of an intra-frame prediction mode is shown.
[0124] Figure 8 An example of a block boundary is shown in the image.
[0125] Figure 9 An example of pixels used in a filter is shown.
[0126] Figure 10 An example of the filter shape for an Adaptive Loop Filter (ALF) is shown.
[0127] Figure 11 An example of the transform coefficients supported by a 5×5 rhombus filter is shown.
[0128] Figure 12 An example of the relative coordinates supported by a 5×5 rhombus filter is shown.
[0129] Figure 13 This is a block diagram illustrating an example video processing system.
[0130] Figure 14 This is a block diagram of an example video processing device.
[0131] Figure 15 This is a flowchart of an example method for video processing.
[0132] Figure 16 This is a block diagram illustrating an example video codec system.
[0133] Figure 17 This is a block diagram showing an example encoder.
[0134] Figure 18 This is a block diagram showing an example decoder.
[0135] Figure 19 This is a schematic diagram of an example encoder. Detailed Implementation
[0136] First, it should be understood that although illustrative implementations of one or more embodiments are provided below, the disclosed systems and / or methods can be implemented using any number of techniques, whether currently known or yet to be developed. This disclosure should not be limited in any way to the illustrative implementations, drawings, and techniques shown below, including the exemplary designs and implementations shown and described herein, but can be modified within the scope of the appended claims and their equivalents.
[0137] Chapter headings are used in this disclosure for ease of understanding and not to limit the applicability of the techniques and embodiments disclosed in each chapter to that chapter only. Furthermore, the techniques described herein are applicable to other video codec protocols and designs.
[0138] 1. Preliminary Discussion
[0139] This disclosure relates to video encoding and decoding technologies. Specifically, this disclosure relates to loop filters and other encoding and decoding tools in image / video encoding and decoding. These ideas can be applied individually or in various combinations to video codecs, such as High Efficiency Video Coding (HEVC), Versatile Video Coding (VVC), or other video encoding and decoding technologies.
[0140] 2. Abbreviation
[0141] This disclosure includes the following abbreviations. Advanced video coding (AVC) (ITU-T H.264 | ISO / IEC 14496-10), Coded Picture Buffer (CPB), Clean Random Access (CRA), Coding Tree Unit (CTU), Coded Video Sequence (CVS), Decoded Picture Buffer (DPB), Decoding Parameter Set (DPS), General Constraints Information (GCI), High Efficiency Video Coding (HEVC, also known as ITU-T H.265 | ISO / IEC 23008-2), Joint Exploration Model (JEM), Motion Constrained Tile Set (MCTS), Network Abstraction Layer (NAL), Output Layer Set (OLS), Picture Header The following are included in the ITU-T H recommendations: Picture Parameter Set (PPS), Profile, Tier, and Level (PTL), Picture Unit (PU), Reference Picture Resampling (RPR), Raw Byte Sequence Payload (RBSP), Supplemental Enhancement Information (SEI), Slice Header (SH), Sequence Parameter Set (SPS), Video Coding Layer (VCL), Video Parameter Set (VPS), and Versatile Video Coding (VVC).266 | ISO / IEC 23090-3), VVC Test Model (VTM), Video Usability Information (VUI), Transform Unit (TU), Coding Unit (CU), Deblocking Filter (DF), Sample Adaptive Offset (SAO), Adaptive Loop Filter (ALF), Coding Block Flag (CBF), Quantization Parameter (QP), Rate Distortion Optimization (RDO), and Bilateral Filter (BF).
[0142] 3. Video codec standards
[0143] Video coding standards have evolved primarily through the development of standards by ITU-T and the International Organization for Standardization (ISO) / International Electrotechnical Commission (IEC). ITU-T developed the H.261 and H.263 standards, ISO / IEC developed the Moving Picture Experts Group (MPEG)-1 and MPEG-4 Vision, and the two organizations jointly developed the H.262 / MPEG-2 video standard and the H.264 / MPEG-4 Advanced Video Coding (AVC) standard and the H.265 / HEVC[1]. Starting with H.262, video coding standards are based on a hybrid video coding architecture, which utilizes temporal prediction plus transform coding. In order to explore future video coding technologies beyond HEVC, the Joint Video Exploration Team (JVET) was established by VCEG and MPEG. JVET adopted many methods and incorporated them into a reference software called the Joint Exploration Model (JEM)[2]. When the Multi-Functional Video Codec (VVC) project was officially launched, JVET was renamed the Joint Video Experts Team (JVET). VVC is a codec standard that aims to reduce the bitrate by 50% compared to HEVC. The VVC working draft and VVC Test Model (VTM) are constantly being updated.
[0144] A sample version of the VVC draft, namely the Multi-Functional Video Codec (Draft 10), can be found at: https: / / jvet-experts.org / doc_end_user / documents / 19_Teleconference / wg11 / JVET-S2001-v17.zip. A sample version of the VVC reference software, named VTM, can be found at: https: / / vcgit.hhi.fraunhofer.de / jvet-u-ee2 / VVCSoftware_VTM / - / tree / VTM-11.2.
[0145] The International Telecommunication Union Telecommunication Standardization Sector (ITU-T), the Video Coding Experts Group (VCEG), and the Joint Technical Committee (JTC) 1 / Subcommittee (SC) 29 / Working Group (WG) 11 of the International Organization for Standardization and the International Electrotechnical Commission (ISO / IEC) Moving Picture Experts Group (MPEG) are studying the potential need for standardization of future video coding and decoding technologies with compression capabilities significantly exceeding the current VVC standard. Such future standardization efforts could take the form of multiple extensions to VVC or entirely new standards. These groups are conducting this exploration through a joint collaborative effort called the Joint Video Exploration Team (JVET) to evaluate compression technology designs proposed by experts in the field. JVET has established the first Exploration Experiments (EE) and is using reference software called the Enhanced Compression Model (ECM). The ECM test model is continuously updated.
[0146] 3.1 Color Space and Chromaticity Downsampling
[0147] A color space, also known as a color model (or color system), is a mathematical model that describes a range of colors as tuples of numbers, such as 3 or 4 values or color components (e.g., RGB). Generally, a color space is a refinement of a coordinate system and its subspaces. For video compression, the most commonly used color spaces are Luminance, Blue Difference, and Red Difference (YCbCr) and Red, Green, and Blue (RGB).
[0148] YCbCr, Y'CbCr, or Y Pb / Cb Pr / Cr, also written as YCBCR or Y'CBCR, is a family of color spaces used as part of the color image pipeline in video and digital photography systems. Y' is the luminance component, and CB and CR are the blue and red difference chromaticity components. Y' (with an apostrophe) is distinguished from Y, where Y is luminance; Y' means that light intensity is encoded non-linearly based on gamma-corrected RGB primary colors.
[0149] Chromaticity downsampling is a practice of encoding images by applying a lower resolution to chromaticity information compared to luminance information, taking advantage of the fact that the human visual system is less sensitive to color differences than to luminance differences. 3.1.1 4:4:4
[0151] In a 4:4:4 scheme, each of the three Y'CbCr components has the same sample rate. Therefore, there is no chromaticity downsampling. This scheme is sometimes used in high-end film scanners and film post-production. 3.1.2 4:2:2
[0153] In a 4:3:2 ratio, both chroma components are sampled at half the sampling rate of the luminance. The horizontal chroma resolution is halved, while the vertical chroma resolution remains unchanged. This reduces the bandwidth of the uncompressed video signal by one-third with almost no visual difference. Figure 1 Examples of the nominal vertical and horizontal positions of 4:2:2 luminance and chrominance samples in an image are shown. 3.1.3 4:2:0
[0155] In 4:2:0, the horizontal sampling is doubled compared to 4:1:1, but the vertical resolution is halved because the blue chromatic aberration (Cb) and red chromatic aberration (Cr) channels are sampled only on each alternating row. Therefore, the data rate remains the same. Cb and Cr are downsampled by a factor of 2 in both the horizontal and vertical directions. There are three variations of the 4:2:0 scheme with different horizontal and vertical positions.
[0156] In MPEG-2, Cb and Cr are co-located horizontally. Cb and Cr are located between pixels vertically (at gap positions). In Joint Picture Experts Group (JPEG) / JPEG File Interchange Format (JFIF), H.261, and MPEG-1, Cb and Cr are located at gap positions, in the middle of alternating luminance samples. In 4:2:0 DV, Cb and Cr are co-located horizontally. Vertically, Cb and Cr are co-located on alternating lines.
[0157] Table 1. SubWidthC and SubHeightC values derived from chroma_format_idc and separate_colour_plane_flag.
[0158]
[0159] 3.2 Example codec stream for video codecs
[0160] Figure 2An example encoder block diagram, such as in VVC, is shown. The encoder contains three loop filtering blocks: a deblocking filter (DF), a sample adaptive compensation (SAO), and an ALF. Unlike the DF, which uses predefined filters, the SAO and ALF utilize the original samples of the current image, reducing the mean square error between the original and reconstructed samples by adding compensation and by applying a finite impulse response (FIR) filter, respectively, and by utilizing the encoded and decoded side information through signal transmission compensation and filter coefficients. The ALF is located in the final processing stage of each image and can be viewed as a tool attempting to capture and repair artifacts caused by previous stages.
[0161] 3.3 Definition of Video / Encoding / Decoding Unit
[0162] An image is divided into one or more slice rows and one or more slice columns. A slice is a sequence of CTUs covering a rectangular area of the image. A slice can be divided into one or more bricks, each brick comprising a certain number of CTU rows within the slice. A slice that is not divided into multiple bricks can also be called a brick. However, a brick that is a proper subset of a slice cannot be called a slice. A strip contains multiple slices of an image or multiple bricks of a slice.
[0163] Two stripe modes are supported: raster scan stripe mode and rectangular stripe mode. In raster scan stripe mode, the stripe contains a sequence of slices from a raster scan of the image. In rectangular stripe mode, the stripe contains a certain number of tiles that collectively form a rectangular area of the image. The tiles within the rectangular stripe are arranged in the order of the stripe's raster scan. Figure 3 An example image segmented into raster scan strips is shown. In the example, the image is segmented according to a raster scan strip segmentation pattern, where the image comprises 18×12 luminance CTUs and is divided into 12 patches and 3 raster scan strips.
[0164] Figure 4 Example images segmented into rectangular scan strips are shown. For example, Figure 4 An example of rectangular strip segmentation of an image with 18×12 luminance CTUs is shown, where the image is divided into 24 slices (6 slice columns and 4 slice rows) and 9 rectangular strips.
[0165] Figure 5 Example images showing the structure divided into bricks are shown. For example, Figure 5 An example of an image divided into slices, bricks, and rectangular strips is shown, where the image is divided into 4 slices (2 slice columns and 2 slice rows), 11 bricks (the top left slice contains 1 brick, the top right slice contains 5 bricks, the bottom left slice contains 2 bricks, and the bottom right slice contains 3 bricks) and 4 rectangular strips.
[0166] 3.3.1 CTU / CTB Dimensions
[0167] In VVC, the CTU size for signal transmission in the Sequence Parameter Set (SPS) can be as small as 4×4 via the syntax element log2_ctu_size_minus2.
[0168] 7.3.2.3 Sequence Parameter Set (RBSP) Syntax
[0169]
[0170] `log2_ctu_size_minus2` plus 2 specifies the luma codec block size for each CTU. `log2_min_luma_coding_block_size_minus2` plus 2 specifies the minimum luma codec block size. The variables `CtbLog2SizeY`, `CtbSizeY`, `MinCbLog2SizeY`, `MinCbSizeY`, `MinTbLog2SizeY`, `MaxTbLog2SizeY`, `MinTbSizeY`, `MaxTbSizeY`, `PicWidthInCtbsY`, `PicHeightInCtbsY`, `PicSizeInCtbsY`, `PicWidthInMinCbsY`, `PicHeightInMinCbsY`, `PicSizeInMinCbsY`, `PicSizeInSamplesY`, `PicWidthInSamplesC`, and `PicHeightInSamplesC` are derived as follows:
[0171] CtbLog2SizeY = log2_ctu_size_minus2 + 2 (7-9)
[0172] CtbSizeY = 1 << CtbLog2SizeY (7-10)
[0173] MinCbLog2SizeY = log2_min_luma_coding_block_size_minus2 + 2 (7-11)
[0174] MinCbSizeY = 1 << MinCbLog2SizeY (7-12)
[0175] MinTbLog2SizeY = 2 (7-13)
[0176] MaxTbLog2SizeY = 6 (7-14)
[0177] MinTbSizeY = 1 << MinTbLog2SizeY (7-15)
[0178] MaxTbSizeY = 1 << MaxTbLog2SizeY (7-16)
[0179] PicWidthInCtbsY = Ceil( pic_width_in_luma_samples ÷ CtbSizeY ) (7-17)
[0180] PicHeightInCtbsY = Ceil( pic_height_in_luma_samples ÷ CtbSizeY ) (7-18)
[0181] PicSizeInCtbsY = PicWidthInCtbsY PicHeightInCtbsY (7-19)
[0182] PicWidthInMinCbsY = pic_width_in_luma_samples / MinCbSizeY (7-20)
[0183] PicHeightInMinCbsY = pic_height_in_luma_samples / MinCbSizeY (7-21)
[0184] PicSizeInMinCbsY = PicWidthInMinCbsY PicHeightInMinCbsY (7-22)
[0185] PicSizeInSamplesY = pic_width_in_luma_samples pic_height_in_luma_samples (7-23)
[0186] PicWidthInSamplesC = pic_width_in_luma_samples / SubWidthC (7-24)
[0187] PicHeightInSamplesC = pic_height_in_luma_samples / SubHeightC (7-25)
[0188] 3.3.2 CTUs in a picture
[0189] Figures 6A-6C Examples of CTBs across picture boundaries are shown. Figure 6A CTBs across the bottom picture boundary are shown. Figure 6B CTBs across the right picture boundary are shown. Figure 6C CTBs across the bottom-right picture boundary are shown. Let the Coding Tree Block (CTB) / Largest Coding Unit (LCU) size be indicated by M×N (usually M equals N), and for a CTB located at a picture boundary (or slice or strip or other type of boundary, with the picture boundary as an example), K×L samples are within the picture boundary, where K < M or L < N. For those CTBs as Figures 6A-6C depicted in [reference], the CTB size still equals M×N. However, the bottom boundary / right boundary of the CTB is outside the picture.
[0190] 3.4 Intra prediction
[0191] Figure 7 Examples of intra prediction modes are shown. To capture any edge direction presented in natural videos, the number of directional intra modes is extended from 33 used in HEVC to 65. The extended directional modes are as Figure 7 shown, and the planar and Direct Current (DC) modes remain unchanged. These denser directional intra prediction modes apply to all block sizes, as well as both luma and chroma intra prediction.
[0192] As Figure 7 shown, the angular intra prediction direction can be defined as clockwise from 45 degrees to -135 degrees. In VTM, for non-square blocks, several angular intra prediction modes are adaptively replaced by wide-angle intra prediction modes. The replaced modes are signaled and remapped to the indices of the wide-angle modes after parsing. The total number of intra prediction modes remains unchanged, e.g., 67, and the intra mode coding and decoding remain unchanged.
[0193] In HEVC, each intra-coded block has a square shape, and the length of each side of the block is a power of 2. Therefore, no division operation is needed to generate intra-prediction values using DC mode. In VVC, blocks can have a rectangular shape, which generally requires division for each block. To avoid division for DC prediction, only the longer side is used to calculate the average of non-square blocks.
[0194] 3.5 Inter-frame prediction
[0195] For each inter-frame prediction CU, motion parameters include motion vectors, reference picture indices, reference picture list usage indices, and extended information for new encoding / decoding features used in VVC, which will be used for inter-frame prediction sample generation. Motion parameters can be transmitted via signaling in an explicit or implicit manner. When a CU is encoded / decoded in skip mode, the CU is associated with a PU and has no significant residual coefficients, no encoded / decoded motion vector increments, and / or reference picture indices. A Merge mode is specified, whereby the motion parameters of the current CU are obtained from neighboring CUs, including spatial and temporal candidates and extended scheduling introduced in VVC. The Merge mode can be applied to any inter-frame prediction CU, not just skip mode. An alternative to the Merge mode is explicit transmission of motion parameters, where the motion vectors, the corresponding reference picture index for each reference picture list, the reference picture list usage flag, and other useful information are explicitly transmitted via signaling for each CU.
[0196] 3.6 Deblocking Filter
[0197] Deblocking filtering is an example of a loop filter in a video codec. In VVC, the deblocking filtering process is applied to CU boundaries, transform subblock boundaries, and predictive subblock boundaries. Predictive subblock boundaries include prediction unit boundaries introduced by Subblock Based Temporal Motion Vector Prediction (SbTMVP) and affine modes. Transform subblock boundaries include transform unit boundaries introduced by Subblock Transform (SBT) and Intra Sub-Partitions (ISP) modes, and transforms introduced by implicit partitioning of large CUs. The processing order of the deblocking filter is defined as first performing horizontal filtering on the vertical edges of the entire image, and then performing vertical filtering on the horizontal edges. This specific order allows multiple horizontal or vertical filtering processes to be applied in parallel threads. The filtering process can also be implemented on a CTB-by-CTB basis with only a small processing latency.
[0198] The vertical edges in the image are first filtered. Then, the horizontal edges in the image are filtered using samples modified by the vertical edge filtering process as input. Vertical and horizontal edges in the CTB of each CTU are processed separately on a codec unit basis. The vertical edges of the codec blocks in the codec unit are filtered, starting from the edge on the left-hand side of the codec block and proceeding geometrically towards the right-hand side of the codec block. The horizontal edges of the codec blocks in the codec unit are filtered, starting from the edge on the top of the codec block and proceeding geometrically towards the bottom of the codec block.
[0199] Figure 8 An example of a block boundary in an image is shown. For example, Figure 8 The image samples and horizontal and vertical block boundaries on an 8×8 grid are shown, as well as non-overlapping blocks of the 8×8 samples that can be de-blocked in parallel.
[0200] 3.6.1 Boundary Decision
[0201] The filter is applied to 8×8 block boundaries. Furthermore, such boundaries must be transform block boundaries or codec sub-block boundaries, such as those resulting from the use of Affine Motion Prediction (ATMVP). For other boundaries, deblocking filtering is disabled.
[0202] 3.6.2 Boundary Strength Calculation
[0203] For transform block boundaries / encoder / decoder sub-block boundaries, if the boundary is located in an 8×8 grid, the boundary can be filtered, and the settings of bS[xDi][yDj] (where [xDi][yDj] represents the coordinates) of the edge are defined as in Tables 2 and 3, respectively.
[0204] Table 2 Boundary Strength (when SPS Intra Block Copy (IBC) is disabled)
[0205]
[0206] Table 3 Boundary Strength (when SPS IBC is enabled)
[0207]
[0208] 3.6.3 Deblocking decision for the luminance component
[0209] Figure 9 An example involving the pixels used by the filter is shown. For example, Figure 9The image shows pixels involved in filter on / off decisions and strong / weak filter selection. A wider and stronger luminance filter is used only if conditions 1, 2, and 3 are all true. Condition 1 is the "bulk condition." This condition detects whether samples on the P-side and Q-side belong to a bulk, represented by the variables bSidePisLargeBlk and bSideQisLargeBlk, respectively. bSidePisLargeBlk and bSideQisLargeBlk are defined as follows.
[0210] bSidePisLargeBlk = ((Edge type is vertical and p0 belongs to CU with width >= 32) || (Edge type is horizontal and p0 belongs to CU with height >= 32)) ? TRUE : FALSE
[0211] bSideQisLargeBlk = ((Edge type is vertical and q0 belongs to CU with width >= 32) || (Edge type is horizontal and q0 belongs to CU with height >= 32)) ? TRUE : FALSE
[0212] Based on bSidePisLargeBlk and bSideQisLargeBlk, condition 1 is defined as follows:
[0213] Condition 1 = (bSidePisLargeBlk || bSidePisLargeBlk) ? TRUE : FALSE
[0214] Next, if condition 1 is true, condition 2 will be further examined. First, the following variables are derived:
[0215] First, derive dp0, dp3, dq0, and dq3 using the HEVC method.
[0216] if (p side is greater than or equal to 32)
[0217] dp0 = (dp0 + Abs(p50 - 2)) p40 + p30 + 1) >> 1
[0218] dp3 = (dp3 + Abs(p53 - 2)) p43 + p33 + 1) >> 1
[0219] if (q side is greater than or equal to 32)
[0220] dq0 = (dq0 + Abs(q50 - 2)) q40 + q30 + 1) >> 1
[0221] dq3 = (dq3 + Abs(q53 - 2)) q43 + q33 + 1) >> 1
[0222] Condition 2 = (d < β) ? TRUE : FALSE
[0223] Where d = dp0 + dq0 + dp3 + dq3.
[0224] If conditions 1 and 2 are valid, then further check whether any of the blocks uses a sub-block:
[0225] If (bSidePisLargeBlk)
[0226] {
[0227] If (block P's mode == SUBBLOCKMODE)
[0228] Sp = 5
[0229] else
[0230] Sp = 7
[0231] }
[0232] else
[0233] Sp = 3
[0234] If (bSideQisLargeBlk)
[0235] {
[0236] If (block Q's mode == SUBBLOCKMODE)
[0237] Sq = 5
[0238] else
[0239] Sq = 7
[0240] }
[0241] else
[0242] Sq = 3
[0243] Finally, if both conditions 1 and 2 are valid, the deblocking method will check condition 3 (the strong filter condition), which is defined as follows. In condition 3, StrongFilterCondition, the following variables are derived:
[0244] Derive dpq using the HEVC method.
[0245] Derive sp3 = Abs( p3 - p0 ) using the HEVC method
[0246] if (p side is greater than or equal to 32)
[0247] if (Sp == 5)
[0248] sp3 = ( sp3 + Abs( p5 - p3 ) + 1) >> 1
[0249] else
[0250] sp3 = ( sp3 + Abs( p7 - p3 ) + 1) >> 1
[0251] Derive sq3 = Abs( q0 - q3 ) using the HEVC method.
[0252] if (q side is greater than or equal to 32)
[0253] If (Sq == 5)
[0254] sq3 = ( sq3 + Abs( q5 - q3 ) + 1) >> 1
[0255] else
[0256] sq3 = ( sq3 + Abs( q7 - q3 ) + 1) >> 1
[0257] According to HEVC, StrongFilterCondition = (dpq < (β >> 2), sp3 + sq3 < (3)). β >> 5), and Abs(p0 - q0) is less than (5). tC + 1 ) >> 1) ? TRUE : FALSE.
[0258] 3.6.4 A more robust deblocking filter for luminance
[0259] When a sample on either side of the boundary belongs to a large block, a bilinear filter is used. A sample belonging to a large block is defined as one whose vertical edge width is >= 32 and whose horizontal edge height is >= 32. The bilinear filter is listed below. Then, the block boundary samples pi (i = 0 to Sp-1) and qi (j = 0 to Sq-1) in the HEVC deblocking (as described above) (pi and qi are the i-th sample in the row used for filtering the vertical edge, or the i-th sample in the column used for filtering the horizontal edge) are replaced with the following linear interpolation:
[0260]
[0261]
[0262] in and The item is the location-related limiting as described above, and , , , and The following is given.
[0263] 3.6.5 Color Deblocking Decision
[0264] A strong chroma filter is used on both sides of the block boundary. A chroma filter is selected when the chroma edge on both sides is greater than or equal to 8 (chroma position), and the following decision is satisfied with three conditions: The first is the decision for boundary strength and the size of the block. A chroma filter can be applied when the block width or height orthogonally crossing the block edge in the chroma sample domain is equal to or greater than 8. The second and third are essentially the same as the HEVC luminance deblocking decision, namely the on / off decision and the strong filter decision, respectively.
[0265] In the first decision, for chroma filtering, the boundary strength (bS) is modified, and conditions are checked sequentially. If a condition is met, the remaining conditions with lower priority are skipped. Chroma deblocking is performed when bS equals 2, or when bS equals 1 when a large block boundary is detected. The second and third conditions are essentially the same as the HEVC luma strong filter decision below.
[0266] In the second condition, d is derived using the HEVC luminance deblocking method. The second condition will be true when d is less than β. In the third condition, StrongFilterCondition is derived as follows:
[0267] Derive dpq using the HEVC method.
[0268] Derive sp3 = Abs( p3 - p0 ) using the HEVC method
[0269] Derive sq3 = Abs( q0 - q3 ) using the HEVC method.
[0270] According to the HEVC design, StrongFilterCondition = (dpq < (β >> 2), sp3 + sq3 < (β >> 3), and Abs(p0 - q0) < (5). tC + 1) >> 1)
[0271] 3.6.6 Strong Deblocking Filter for Chroma
[0272] The following strong deblocking filter for chroma is defined:
[0273] p2′= (3 p3+2 p2+p1+p0+q0+4) >> 3
[0274] p1′= (2 p3+p2+2 p1+p0+q0+q1+4) >> 3
[0275] p0′= (p3+p2+p1+2 p0+q0+q1+q2+4) >> 3
[0276] The example chroma filter performs deblocking on a 4x4 chroma sample grid.
[0277] 3.6.7 Location-related amplitude limiting
[0278] Position-dependent limiting (tcPD) is applied to the output samples of the brightness filtering process involving strong and long filters, which modify 7, 5, and 3 samples at the boundaries, respectively. Assuming a quantization error distribution, the limiting value can be increased for samples expected to have higher quantization noise, thus anticipating a larger deviation between the reconstructed sample value and the true sample value.
[0279] For each P or Q boundary filtered by an asymmetric filter, based on the result of the decision process, the position-related threshold table is selected from two tables (e.g., Tc7 and Tc3 listed below) that are provided to the decoder as edge information:
[0280] Tc7 = {6, 5, 4, 3, 2, 1, 1}; Tc3 = {6, 4, 2};
[0281] tcPD = (Sp == 3) ? Tc3 : Tc7;
[0282] tcQD = (Sq == 3) ? Tc3 : Tc7;
[0283] For P or Q boundaries filtered by a short symmetric filter, apply a lower-amplitude position correlation threshold:
[0284] Tc3 = { 3, 2, 1};
[0285] After defining the threshold, the filtered p'i and q'i sample values are clipped according to the tcP and tcQ clipping values:
[0286] p''i = Clip3(p'i + tcPi, p'i – tcPi, p'i );
[0287] q''j = Clip3(q'j + tcQj, q'j – tcQ j, q'j );
[0288] Where p'i and q'i are the filtered sample values, p''i and q''j are the output sample values after clipping, and tcPitcPi is the clipping threshold derived from the VVC tc parameters and tcPD and tcQD. The function Clip3 is the clipping function as specified in VVC.
[0289] 3.6.8 Sub-block Removal and Adjustment
[0290] To enable parallel-friendly deblocking using both long filters and sub-block deblocking, the long filter is restricted to modifying a maximum of 5 samples on the side using sub-block deblocking (AFFINE or ATMVP or Decoder-side Motion Vector Refinement, DMVR), as shown in the long filter's brightness control. Extending this, sub-block deblocking is adjusted such that sub-block boundaries on the 8×8 mesh near the CU boundary or implicit TU boundary are restricted to modifying a maximum of two samples on each side.
[0291] The following applies to sub-block boundaries that are not aligned with the CU boundary.
[0292] If (block Q's mode == SUBBLOCKMODE && edge != 0) {
[0293] if (!(implicitTU && (edge == (64 / 4))))
[0294] if (edge == 2 || edge == (orthogonalLength - 2) || edge== (56 / 4) || edge == (72 / 4))
[0295] Sp = Sq = 2;
[0296] else
[0297] Sp = Sq = 3;
[0298] else
[0299] Sp = Sq = bSideQisLargeBlk ? 5:3
[0300] }
[0301] Where edge = 0 corresponds to the CU boundary, edge = 2 or orthogonalLength-2 corresponds to the sub-block boundary 8 samples away from the CU boundary, and so on. If implicit partitioning of TU is used, then implicitTU is true.
[0302] 3.7 Sample point adaptive compensation
[0303] Sample Adaptive Compensation (SAO) is applied to the reconstructed signal after the deblocking filter, using compensation specified by the encoder for each CTB. The video encoder first decides whether to apply the SAO process to the current slice. If SAO is applied to the current slice, each CTB is classified into one of five SAO types as shown in Table 4. The concept of SAO is to classify pixels into multiple categories and reduce distortion by adding compensation to pixels in each category. SAO operations include Edge Offset (EO) and Band Offset (BO), where EO uses edge attributes for pixel classification in SAO types 1 through 4, and BO uses pixel intensity for pixel classification in SAO type 5. Each applicable CTB has SAO parameters including sao_merge_left_flag, sao_merge_up_flag, SAO type, and four compensations. If sao_merge_left_flag equals 1, the current CTB will reuse the SAO type and compensation of the left CTB. If sao_merge_up_flag equals 1, the current CTB will reuse the SAO type and compensation of the CTB above.
[0304] Table 4. Regulations for SAO Types
[0305]
[0306] 3.8 Adaptive Loop Filter
[0307] Adaptive Loop Filtering (ALF) for video encoding and decoding minimizes the mean square error between the original and decoded samples using Wiener-based adaptive filters. ALF is located at the final processing stage of each picture and can be considered a tool for capturing and repairing artifacts from previous stages. Appropriate filter coefficients are determined by the encoder and explicitly transmitted to the decoder via the signal. To achieve better encoding and decoding efficiency, especially for high-resolution video, local adaptation is used for the luminance signal by applying different filters to different regions or blocks in the picture. In addition to filter adaptation, filter on / off control at the codec tree unit (CTU) level also contributes to improved encoding and decoding efficiency. Syntactically, filter coefficients are sent in a picture-level header called the adaptive parameter set, and the filter on / off flags of the CTUs are interleaved at the CTU level in the striped data. This syntax design not only supports picture-level optimization but also achieves low encoding latency.
[0308] 3.8.1 Signaling of Parameters
[0309] According to the ALF design in VTM, filter coefficients and clipping indices are carried in the ALF Adaptation Parameter Set (APS). An ALF APS can include up to eight chroma filters and a luma filter set with up to 25 filters. Each of the 25 luma categories includes an index. Categories with the same index share the same filters. By merging different categories, the number of bits required to represent the filter coefficients is reduced. The absolute values of the filter coefficients are represented using 0th-order Exp-Golomb code, followed by the sign bit for the non-zero coefficients. When clipping is enabled, a two-bit fixed-length code is also used to transmit the clipping index through the signal for each filter coefficient. The decoder can use up to eight ALF APSs simultaneously.
[0310] The filter control syntax elements for ALF in VTM include two types of information. First, the ALF on / off flag is signaled at the sequence, picture, strip, and CTB levels. Chroma ALF can only be enabled at the corresponding level if Luminance ALF is enabled at the picture and strip levels. Second, if ALF is enabled at the picture, strip, and CTB levels, filter usage information is signaled at that level. If all strips within a picture use the same APS, the referenced ALF APS ID is encoded / decoded at the strip or picture level. A luminance component can reference up to 7 ALF APSs, and a chroma component can reference 1 ALF APS. For luminance CTB, an index is signaled to indicate which ALF APS or offline-trained luminance filter set is used. For chroma CTB, the index indicates which filter in the referenced APS is used.
[0311] The data syntax elements of the ALF associated with the luminance component in VTM are listed below:
[0312]
[0313] `alf_luma_filter_signal_flag` equal to 1 specifies the set of luminance filters transmitted via signaling. `alf_luma_filter_signal_flag` equal to 0 specifies that no luminance filter set is transmitted via signaling. `alf_luma_clip_flag` equal to 0 specifies that linear adaptive loop filtering is applied to the luminance component. `alf_luma_clip_flag` equal to 1 specifies that nonlinear adaptive loop filtering can be applied to the luminance component. `alf_luma_num_filters_signalled_minus1` plus 1 specifies the number of adaptive loop filter classes through which the luminance coefficients can be transmitted via signaling. The value of `alf_luma_num_filters_signalled_minus1` must be in the range of 0 to NumAlfFilters - 1 (inclusive). `alf_luma_coeff_delta_idx[filtIdx]` specifies the index of the adaptive loop filter luminance coefficient increment transmitted via signaling, which is the index of the filter class indicated by `filtIdx` in the range of 0 to NumAlfFilters - 1. When alf_luma_coeff_delta_idx[filtIdx] does not exist, it is presumed to be equal to 0. The length of alf_luma_coeff_delta_idx[filtIdx] is Ceil(Log2(alf_luma_num_filters_signalled_minus1 + 1)) bits. The value of alf_luma_coeff_delta_idx[filtIdx] must be in the range from 0 to alf_luma_num_filters_signalled_minus1 (inclusive).
[0314] `alf_luma_coeff_abs[ sfIdx ][ j ]` specifies the absolute value of the j-th coefficient of the luminance filter transmitted through the signal, indicated by `sfIdx`. When `alf_luma_coeff_abs[ sfIdx ][ j ]` does not exist, it is presumed to be equal to 0. The value of `alf_luma_coeff_abs[ sfIdx ][ j ]` must be in the range of 0 to 128 (inclusive). `alf_luma_coeff_sign[ sfIdx ][ j ]` specifies the sign of the j-th luminance coefficient of the filter indicated by `sfIdx`, as follows:
[0315] If alf_luma_coeff_sign[ sfIdx ][ j ] equals 0, the corresponding luminance filter coefficient has a positive value.
[0316] Otherwise (alf_luma_coeff_sign[ sfIdx ][ j ] equals 1), the corresponding luminance filter coefficient has a negative value.
[0317] When alf_luma_coeff_sign[ sfIdx ][ j ] does not exist, it is presumed to be equal to 0.
[0318] `alf_luma_clip_idx[ sfIdx ][ j ]` specifies the limiting index of the limiting value to be used before multiplying by the j-th coefficient of the luminance filter transmitted through the signal, indicated by `sfIdx`. When `alf_luma_clip_idx[ sfIdx][ j ]` does not exist, it is presumed to be equal to 0. The codec tree unit syntax elements of the ALF associated with the luminance component in the VTM are listed below:
[0319]
[0320] `alf_ctb_flag[ cIdx ][ xCtb >> CtbLog2SizeY ][ yCtb >> CtbLog2SizeY ]` equal to 1 indicates that the adaptive loop filter is applied to the codec tree block of the codec tree unit at the luma position (xCtb, yCtb), for the color component indicated by cIdx. `alf_ctb_flag[ cIdx ][ xCtb >> CtbLog2SizeY ][ yCtb >> CtbLog2SizeY ]` equal to 0 indicates that the adaptive loop filter is not applied to the codec tree block of the codec tree unit at the luma position (xCtb, yCtb), for the color component indicated by cIdx.
[0321] When `alf_ctb_flag[ cIdx ][ xCtb >> CtbLog2SizeY ][ yCtb >> CtbLog2SizeY]` does not exist, it is presumed to be equal to 0. `alf_use_aps_flag` equal to 0 specifies that one of the fixed filter sets is applied to the luma CTB. `alf_use_aps_flag` equal to 1 specifies that a filter set from the APS is applied to the luma CTB. When `alf_use_aps_flag` does not exist, it is presumed to be equal to 0. `alf_luma_prev_filter_idx` specifies the previous filter applied to the luma CTB. The value of `alf_luma_prev_filter_idx` must be in the range of 0 to `sh_num_alf_aps_ids_luma - 1` (inclusive). When `alf_luma_prev_filter_idx` does not exist, it is presumed to be equal to 0.
[0322] The variable AlfCtbFiltSetIdxY[xCtb >> CtbLog2SizeY ][ yCtb >> CtbLog2SizeY ], which represents the filter set index of the luminance CTB at a specified location (xCtb, yCtb), is derived as follows:
[0323] If alf_use_aps_flag equals 0, then AlfCtbFiltSetIdxY[ xCtb >> CtbLog2SizeY ][ yCtb >> CtbLog2SizeY ] is set to equal alf_luma_fixed_filter_idx.
[0324] Otherwise, AlfCtbFiltSetIdxY[ xCtb >> CtbLog2SizeY ][ yCtb >> CtbLog2SizeY] is set to equal to 16 + alf_luma_prev_filter_idx.
[0325] alf_luma_fixed_filter_idx specifies the fixed filter applied to the lumen CTB. The value of alf_luma_fixed_filter_idx must be in the range of 0 to 15 (inclusive).
[0326] Based on the VTM-based ALF design, the ECM-based ALF design further introduces the concept of candidate filter sets into the luminance filter. The luminance filter is trained with multiple candidates / rounds based on the updated luminance CTB ALF on / off decision for each candidate / round. This results in multiple filter sets associated with each trained candidate, and the class merging results for each filter set can be different. Each CTU can select the optimal filter set via RDO, and the relevant candidate information is transmitted via signal transmission. The data syntax elements of the ALF associated with the luminance component in the ECM are listed below:
[0327]
[0328] `alf_luma_num_alts_minus1` incremented by 1 specifies the number of candidate filter sets for the luminance component. The value of `alf_luma_num_alts_minus1` must be in the range of 0 to 3 (inclusive). `alf_luma_clip_flag[altIdx]` equal to 0 specifies that linear adaptive loop filtering is applied to the candidate luminance filter set for the luminance component with index `altIdx`. `alf_luma_clip_flag[altIdx]` equal to 1 specifies that nonlinear adaptive loop filtering can be applied to the candidate luminance filter set for the luminance component with index `altIdx`. `alf_luma_num_filters_signalled_minus1[altIdx]` incremented by 1 specifies the number of adaptive loop filter categories for the candidate luminance filter set with index `altIdx` that the luminance coefficient can be transmitted via signal transmission. The value of alf_luma_num_filters_signalled_minus1[altIdx] must be in the range of 0 to NumAlfFilters - 1 (inclusive).
[0329] `alf_luma_coeff_delta_idx[altIdx][filtIdx]` specifies the index of the adaptive loop filter luminance coefficient increment transmitted via signal transmission. This index is the index of the filter class indicated by `filtIdx`, ranging from 0 to `NumAlfFilters – 1`, in the set of candidate luminance filters with index `altIdx`. `alf_luma_coeff_delta_idx[filtIdx][altIdx]` is presumed to be 0 if it does not exist. The length of `alf_luma_coeff_delta_idx[altIdx][filtIdx]` is Ceil(Log2(alf_luma_num_filters_signalled_minus1[altIdx] + 1)) bits. The value of `alf_luma_coeff_delta_idx[altIdx][filtIdx]` must be in the range of 0 to `alf_luma_num_filters_signalled_minus1[altIdx]` (inclusive). `alf_luma_coeff_abs[altIdx][sfIdx][j]` specifies the absolute value of the j-th coefficient of the luminance filter transmitted through the signal, indicated by `sfIdx`, from the set of candidate luminance filters with index `altIdx`. When `alf_luma_coeff_abs[altIdx][sfIdx][j]` does not exist, it is presumed to be equal to 0. The value of `alf_luma_coeff_abs[altIdx][sfIdx][j]` must be in the range of 0 to 128 (inclusive).
[0330] `alf_luma_coeff_sign[altIdx][sfIdx][j]` specifies the sign of the j-th luminance coefficient of the filter indicated by `sfIdx` in the set of candidate luminance filters with index `altIdx`, as follows:
[0331] If alf_luma_coeff_sign[altIdx][sfIdx][j] equals 0, then the corresponding luminance filter coefficient has a positive value.
[0332] Otherwise (alf_luma_coeff_sign[altIdx][sfIdx][j] equals 1), the corresponding luminance filter coefficient has a negative value.
[0333] When alf_luma_coeff_sign[altIdx][sfIdx][j] does not exist, it is presumed to be equal to 0.
[0334] `alf_luma_clip_idx[altIdx][sfIdx][j]` specifies the limiting index to be used before multiplying the j-th coefficient of the luminance filter transmitted through the signal, indicated by `sfIdx`, in the set of candidate luminance filters with index `altIdx`. When `alf_luma_clip_idx[altIdx][sfIdx][j]` does not exist, it is presumed to be equal to 0. The codec tree unit syntax elements of the ALF associated with the luminance component in the ECM are listed below:
[0335]
[0336] `alf_ctb_luma_filter_alt_idx[ xCtb >> CtbLog2SizeY ][ yCtb >> CtbLog2SizeY ]` specifies the index of the candidate luma filter, which is applied to the luma component's codec tree block of the codec tree unit at the luma location (xCtb, yCtb). If `alf_ctb_luma_filter_alt_idx[ xCtb >> CtbLog2SizeY ][ yCtb >> CtbLog2SizeY ]` does not exist, it is presumed to be equal to 0.
[0337] 3.8.2 Filter Shape
[0338] Figure 10 An example of ALF filter shape is shown. In JEM, there are a maximum of three diamond filter shapes (such as...). Figure 10 (As shown) can be selected for the luma component. At the image level, the filter shape used for the luma component is indicated by a signal transmission index. Each square represents a sample, and Ci (i = 0~6 (left), 0~12 (middle), 0~20 (right)) represents the coefficient to be applied to the sample. For the chroma component in the image, a 5×5 rhombus shape is always used. In VVC, a 7×7 rhombus shape is always used for luma, while a 5×5 rhombus shape is always used for chroma.
[0339] 3.8.3 Classification of ALF
[0340] Each 2×2 (or 4×4) block is classified into one of 25 categories. The classification index C is based on its directionality. and activity The quantization value is derived as follows:
[0341]
[0342] In order to calculate and First, the gradients in the horizontal, vertical, and two diagonal directions are calculated using 1-D Laplacian:
[0343]
[0344]
[0345]
[0346]
[0347] index and Referencing the coordinates of the top left sample point in the 2×2 block, and Indicator coordinates The reconstructed sample points are then used. Then, the gradients in the horizontal and vertical directions are... The maximum and minimum values are set as follows:
[0348] , ,
[0349] Furthermore, the maximum and minimum values of the gradients in the two diagonal directions are set as follows:
[0350] , ,
[0351] In order to derive directionality The values are compared with each other and then compared with two thresholds. and Compare:
[0352] Step 1. If and If both are true, then Set as .
[0353] Step 2. If If yes, continue from step 3; otherwise, continue from step 4.
[0354] Step 3. If ,but Set as ;otherwise Set as .
[0355] Step 4. If ,but Set as ;otherwise Set as .
[0356] Activity value Calculated as:
[0357]
[0358] It is further quantized to the range of 0 to 4 (inclusive), and the quantized value is represented as No classification method was applied to the two chromaticity components in the image; instead, a single set of ALF coefficients was applied to each chromaticity component.
[0359] 3.8.4 Geometric Transformation of Filter Coefficients
[0360] Before filtering each 2×2 block, geometric transformations (such as rotation or diagonal and vertical flips) are applied to the filter coefficients associated with coordinates (k, l), based on the gradient values calculated for that block. This is equivalent to applying these transformations to samples in the filter's support region. The idea is to make the different blocks more similar by aligning the directions of the blocks to which the ALF is applied.
[0361] Three geometric transformations are introduced: diagonal, vertical flip, and rotation.
[0362] diagonal:
[0363] Vertical Flip: ,
[0364] Rotation:
[0365] in It is the size of the filter, and These are coefficient coordinates, which make the position... In the top left corner, and in position In the bottom right corner. Based on the gradient values calculated for this block, the transform is applied to the filter coefficients f(k, l). The relationship between the transform and the four gradients in the four directions is summarized in Table 5.
[0366] Figure 11 An example of the transform coefficients of a relative coordinator acting as a support for a 5×5 rhombus filter is shown. For example, Figure 11 The transformation coefficients for each position based on a 5×5 rhombus are shown.
[0367] Table 5 shows the mapping between gradients and transformations for a single block of computation.
[0368]
[0369] 3.8.5 Filtering Process
[0370] On the decoder side, when ALF is enabled for a block, each sample within the block... The filtering results in sample values as shown below. Where L represents the filter length, Represents the filter coefficients, and This represents the filter coefficients used in decoding.
[0371]
[0372] Figure 12 This example illustrates the relative coordinates used to support a 5×5 diamond filter, where the coordinates of the current sample point are (i, j) and (0, 0). Samples at different coordinates filled with the same color are multiplied by the same filter coefficients.
[0373] 3.8.6 Restating Nonlinear Filtering
[0374] Linear filtering can be reformulated as follows without affecting encoding / decoding efficiency:
[0375]
[0376] in They are the same filter coefficients.
[0377] VVC introduces nonlinearity by using a simple limiting function to measure the value at neighboring sample points ( ) and the current sample value being filtered ( When the difference is too large, the influence of neighboring sample values is reduced, thus making ALF more efficient. More specifically, the ALF filter is modified as follows:
[0378]
[0379] in It is a limiting function, and It is the limiting parameter, which depends on Filter coefficients. The encoder performs optimization to find the optimal values. .
[0380] A limiting parameter is specified for each ALF filter. Each filter coefficient transmits a limiting value via signal transmission. This means that each luminance filter can have a maximum of 12 limiting values transmitted via signal transmission in the bitstream, and each chroma filter can have a maximum of 6 limiting values transmitted via signal transmission in the bitstream. To limit signaling costs and encoder complexity, only 4 fixed values, the same as those used in INTER and INTRA stripes, are used.
[0381] Because the variance of local differences in luminance is typically higher than that in chrominance, two different sets are applied to the luminance and chrominance filters. The maximum sample value in each set (here for a 10-bit bit depth of 1024) is also introduced so that clipping can be disabled if unnecessary. The four values are chosen by dividing the full range of luminance sample values (encoded and decoded on 10 bits) and the chrominance range from 4 to 1024 into approximately equal parts in the logarithmic domain. More precisely, the luminance table of clipping values has been obtained using the following formula:
[0382] AlfClip L Where M=2 10 And N=4
[0383] Similarly, the colorimetric table for the limiting values is obtained according to the following formula:
[0384] AlfClip C Where M=2 10 N=4, A=4
[0385] 3.9 Bilateral Loop Filter
[0386] 3.9.1 Bilateral Image Filter
[0387] Bilateral image filtering is a nonlinear filter that smooths noise while preserving edge structure. Bilateral filtering is a technique where the filter weights decrease not only with increasing distance between samples but also with increasing intensity difference. This improves the smoothing of overly smoothed edges. The weights are defined as follows:
[0388]
[0389] in and It refers to both vertical and horizontal distances, and It is the intensity difference between sample points.
[0390] Edge-preserving denoising bilateral filters employ low-pass Gaussian filters for both the domain and range filters. The domain low-pass Gaussian filter assigns higher weights to pixels spatially closer to the center pixel. The range low-pass Gaussian filter assigns higher weights to pixels similar to the center pixel. Combining the range and domain filters, the bilateral filter at edge pixels becomes a thin Gaussian filter oriented along the edge and significantly reduced in the gradient direction. This is why bilateral filters can smooth noise while preserving edge structure.
[0391] 3.9.2 Bilateral Filters in Video Encoding and Decoding
[0392] Bilateral filters in video encoding and decoding are the encoding and decoding tools of VVC[2]. The filter acts as a loop filter in parallel with the Sample Adaptive Compensation (SAO) filter. Both the bilateral filter and the SAO act on the same input sample, each filter produces compensation, and these compensations are then added to the input sample to produce an output sample, which is then clipped to the next stage. Spatial filtering intensity Determined by the block size, smaller blocks are filtered more strongly, and the intensity of the filter is equal to the strength. Determined by the quantization parameters, stronger filtering is used for higher QP. Only the four nearest samples are used, therefore the filtered sample intensity... It can be calculated as
[0393]
[0394] in Indicates the intensity of the central sample point. This indicates the intensity difference between the center sample point and the sample point above it. and These represent the intensity differences between the central sample point and the sample points below, to the left, and to the right, respectively.
[0395] 4. The technical problem solved by the disclosed technical solution
[0396] The example design of the Adaptive Loop Filter (ALF) in video encoding and decoding has the following problems.
[0397] In the example ALF design, only spatial reconstruction / residual samples are used as input for classification. However, other valuable information can potentially be utilized, such as predictive pattern information.
[0398] 5. List of solutions and implementation examples
[0399] To address the aforementioned problems, the following summarized methods are disclosed. The embodiments should be considered as examples for explaining general concepts and should not be interpreted in a narrow sense. Furthermore, these embodiments can be applied individually or in any combination.
[0400] It should be noted that the disclosed method can be used as a loop filter or post-processing.
[0401] In this disclosure, a video unit can refer to a sequence, picture, subpicture, strip, CTU, block, and / or region. A video unit may include one color component or multiple color components.
[0402] In this disclosure, an ALF processing unit can refer to a sequence, image, sub-image, stripe, CTU, block, region, or sample. An ALF processing unit can include one or more color components. Furthermore, side information can include encoding / decoding information other than the samples to be filtered by the current filter, and such side information can be used to determine the current filter's use of the samples to be filtered.
[0403] 1) It is proposed to use the prediction pattern as side information for loop filters such as deblocking filters / SAO / Bilateral Filter (BIF) / ALF / CCALF.
[0404] a. In one example, the prediction pattern for at least one location can be stored before the loop filter process.
[0405] a) In one example, the prediction pattern for at least one location can be stored.
[0406] 1. In one example, the prediction pattern of brightness location can be stored.
[0407] 2. In one example, the prediction pattern for chromaticity location can be stored.
[0408] 3. The stored prediction patterns can be obtained from the loop filter.
[0409] b) In one example, the prediction pattern for each location can be stored / used in the mapping domain.
[0410] 1. In one example, the prediction mode for each location can be stored / used using mapping values representing different prediction modes (e.g., MODE_INTRA=128, MODE_INTER=256, MODE_IBC=512).
[0411] 2. In one example, the mode information can be a prediction mode (e.g., inter-frame prediction, intra-frame prediction, IBC, or IntraTMP).
[0412] 3. In one example, the pattern information could be the CBF flag.
[0413] 4. In one example, the pattern information could be the rootCBF flag.
[0414] 5. In one example, the mapping value can be predefined.
[0415] 6. In one example, the mapping value can be derived.
[0416] 7. In one example, the mapped value can be transmitted to the decoder via a signal.
[0417] c) In one example, the predicted pattern of a location can be classified into one or more categories.
[0418] 1. In one example, the predicted pattern can be classified into N categories (e.g., N=2).
[0419] 2. In one example, the predicted pattern can be classified according to predefined rules.
[0420] a. In one example, a location with the INTER pattern can be classified into a category.
[0421] b. In one example, locations with both INTER and IBC patterns can be classified into one category.
[0422] c. In one example, locations with INTER, IBC, and IntraTMP modes can be categorized into one class.
[0423] d. In one example, locations with the INTRA pattern can be categorized into a single category.
[0424] e. In one example, a location with a NON-INTRA pattern can be classified into a category.
[0425] f. In one example, a location with a CBF equal to 0 can be classified into a category.
[0426] g. In one example, a location with a CFB that is not equal to 0 can be classified into a category.
[0427] h. In one example, block boundary locations can be classified into a different category than other locations within a block.
[0428] 3. In one example, the number of merge rules or merge categories can be predefined.
[0429] 4. In one example, the number of merge rules or merge categories can be derived.
[0430] 5. In one example, the number of merge rules or merge categories can be transmitted to the decoder via a signal.
[0431] 2) It is proposed to use the prediction pattern as the side information in ALF / CCALF filtering.
[0432] a. For example, f(pm) can be used as side information in ALF / CCALF filtering, where “m” is the mapping value of the prediction mode and “f” is any function, such as a linear function, a limiting function, or any other function.
[0433] b. In one example, the mapping value of the prediction pattern can be used as input to ALF / CCALF.
[0434] a) In one example, the mapping value of the prediction pattern can be used as input to at least one extended tap of ALF / CCALF.
[0435] b) In one example, the mapping value of the prediction pattern can be used as at least one existing tap of the ALF.
[0436] c) In one example, the mapping value of the luminance prediction mode can be used as input to ALF luminance / CCALF.
[0437] d) In one example, the mapping value of the chromaticity prediction mode can be used as input for ALF chromaticity.
[0438] c. In one example, the mapping value of the prediction pattern can be used as input to CCALF.
[0439] a) In one example, the mapping value of the prediction pattern can be used as input to at least one extended tap of CCALF.
[0440] b) In one example, the mapping value of the prediction pattern can be used as input to at least one existing tap of CCALF.
[0441] c) In one example, the mapping value of the luminance prediction mode can be used as input to the extended tap of CCALF.
[0442] d. In one example, the mapped values of the predicted mode can be used after the ALF / CCALF filtering function.
[0443] a) In one example, the difference between the predicted pattern's mapped value and the currently filtered sample may need to be calculated first.
[0444] b) In one example, the filter coefficients can be multiplied by the difference to obtain the final offset.
[0445] e. In one example, the mapped values of the prediction mode can be used in a way that differs from the ALF / CCALF filtering function.
[0446] a) In one example, the mapping value of the predicted pattern value can be used directly.
[0447] b) In one example, the filter coefficients can be multiplied by the mapping value of the prediction mode to obtain the final offset.
[0448] 3) It is proposed to use the prediction pattern as the side information in ALF / CCALF classification.
[0449] a. For example, f(mp) can be used as side information in ALF / CCALF classification, where “mp” is the mapping value of the predicted pattern and “f” is any function, such as a linear function, a limiting function, or any other function.
[0450] b. In one example, the mapping value of the predicted mode can be used in conjunction with the deblocking filter boundary strength (DBF-BS).
[0451] a) In one example, the prediction pattern and DBF-BS can be directly combined. The prediction pattern can be classified into N (e.g., N=2) categories, and the DBF-BS can be classified into M (e.g., M=2) categories. The total number of categories generated by the side information can be equal to M×N (e.g., M×N=4).
[0452] (b) In one example, the prediction pattern and DBF-BS can be conditionally combined. The prediction pattern can be applied only to locations where DBF-BS equals the value K (e.g., K=0).
[0453] 1. In one example, K can be predefined.
[0454] 2. In one example, K can be transmitted via a signal.
[0455] 3. In one example, K can be derived at the decoder.
[0456] c. In one example, the mapping value of the predicted pattern can be used as an additional classifier.
[0457] a) In one example, the classifier unit size of the classifier based on the mapping values of the predicted pattern can be N. M (e.g., N=M=2).
[0458] b) In one example, locations within a classification unit can share the same classification result.
[0459] c) In one example, the classification results can be generated / derived using a specific method.
[0460] 1. In one example, the method may be predefined.
[0461] 2. In one example, the method can be derived.
[0462] 3. In one example, this method can be transmitted to the decoder via a signal.
[0463] 4. In one example, the method could be an averaging function.
[0464] 5. In one example, the method could be a maximum function.
[0465] 6. In one example, the method could be a minimal function.
[0466] 7. In one example, the method could be a linear mapping function.
[0467] 8. In one example, the method could be a non-linear mapping function.
[0468] 9. Alternatively, the method can be any other function.
[0469] d) In one example, whether or not a classifier based on the predicted pattern mapping value is used can be predefined / derived / transmitted via signaling.
[0470] d. In one example, the mapping value of the predicted pattern can be used as side information for an existing classification pattern.
[0471] a) In one example, locations within a classification unit can share the same mapping value for the predicted pattern value.
[0472] b) In one example, the mapping values of the predicted pattern values within a classification unit can be generated / derived using a specific method.
[0473] 1. In one example, the method may be predefined.
[0474] 2. In one example, the method can be derived.
[0475] 3. In one example, this method can be transmitted to the decoder via a signal.
[0476] 4. In one example, the method could be an averaging function.
[0477] 5. In one example, the method could be a maximum function.
[0478] 6. In one example, the method could be a minimal function.
[0479] 7. In one example, the method could be a linear mapping function.
[0480] 8. In one example, the method could be a non-linear mapping function.
[0481] 9. Alternatively, the method can be any other function.
[0482] c) In one example, the mapping values of the predicted pattern can be merged into K categories (e.g., K=2).
[0483] d) In one example, an existing classifier may have N classes (e.g., N=25).
[0484] e) In one example, the mapping value of the prediction pattern can be used to directly expand the number of categories.
[0485] 1. In one example, the number of classes in an existing classifier can be extended to N. K (e.g., N) K=50).
[0486] 2. In one example, the classification result can be calculated using the following formula:
[0487]
[0488] Where c is the final category index. It is a category index generated by an existing classifier.
[0489] 3. In one example, the classification result can be calculated using the following formula:
[0490]
[0491] Where c is the final category index. It is a category index generated by an existing classifier.
[0492] f) In one example, the mapping value of the prediction pattern can be used to modify the number of categories.
[0493] 1. In one example, the number of classes in an existing classifier can be modified to M (e.g., M=12).
[0494] a. In one example, for a band-based classifier, the total number of classes can be modified.
[0495] b. In one example, for a residual-based classifier, the total number of classes can be modified.
[0496] c. In one example, for a texture-based classifier, the number of texture directions can be modified.
[0497] d. In one example, for a texture-based classifier, the number of texture activities can be modified.
[0498] 2. In one example, the number of classes in the existing classifier can be modified to M. K (e.g., M) K=24).
[0499] 3. In one example, the classification result can be calculated using the following formula:
[0500]
[0501] Where c is the final category index. It is a category index generated by an existing classifier.
[0502] 4. In one example, the classification result can be calculated using the following formula:
[0503]
[0504] Where c is the final category index. It is a category index generated by an existing classifier.
[0505] 4) The use of prediction mode information or DBF-BS information in ALF fixed filter classification is proposed.
[0506] a. In one example, prediction mode information or DBF-BS information can be used for fixed filter classification of the luminance ALF.
[0507] b. In one example, the predicted mode information, namely DBF-BS information, can be used for fixed filter classification of chroma ALF.
[0508] c. In one example, prediction pattern information or DBF-BS information can be used for fixed filter classification in CCALF.
[0509] d. For example, f(mp) can be used as side information in ALF / CCALF classification with a fixed filter, where “mp” is the mapping value of the predicted mode or DBF-BS, and “f” is any function, such as a linear function or a limiting function or any other function.
[0510] e. In one example, the mapping values of the prediction pattern can be used in conjunction with DBF-BS.
[0511] a) In one example, the prediction pattern and DBF-BS can be directly combined. The prediction pattern can be classified into N (e.g., N=2) categories, and the DBF-BS can be classified into M (e.g., M=2) categories. The total number of categories generated from the side information can be equal to M×N (e.g., M×N=4).
[0512] (b) In one example, the prediction pattern and DBF-BS can be conditionally combined. The prediction pattern can be applied only to locations where DBF-BS equals the value K (e.g., K=0).
[0513] 1. In one example, K can be predefined.
[0514] 2. In one example, K can be transmitted via a signal.
[0515] 3. In one example, K can be derived at the decoder.
[0516] f. In one example, the predicted pattern mapping value or DBF-BS can be used as an additional classifier for the fixed filter.
[0517] a) In one example, the classifier unit size of the prediction pattern-based mapping value or the DBF-BS classifier can be N. M (e.g., N=M=2).
[0518] b) In one example, locations within a classification unit can share the same classification result.
[0519] c) In one example, the classification results can be generated / derived using a specific method.
[0520] 1. In one example, the method may be predefined.
[0521] 2. In one example, the method can be derived.
[0522] 3. In one example, this method can be transmitted to the decoder via a signal.
[0523] 4. In one example, the method could be an averaging function.
[0524] 5. In one example, the method could be a maximum function.
[0525] 6. In one example, the method could be a minimal function.
[0526] 7. In one example, the method could be a linear mapping function.
[0527] 8. In one example, the method could be a non-linear mapping function.
[0528] 9. Alternatively, the method can be any other function.
[0529] d) In one example, whether a prediction pattern-based mapping value or a DBF-BS classifier is used can be predefined / derived / transmitted via signaling.
[0530] g. In one example, the mapping value of the predicted pattern or DBF-BS can be used as the side information of the existing classification pattern with a fixed filter.
[0531] a) In one example, locations within a classification unit can share the same predicted pattern value mapping or DBF-BS.
[0532] b) In one example, the mapping value or DBF-BS of the predicted pattern value within the classification unit can be generated / derived using a specific method.
[0533] 1. In one example, the method may be predefined.
[0534] 2. In one example, the method can be derived.
[0535] 3. In one example, this method can be transmitted to the decoder via a signal.
[0536] 4. In one example, the method could be an averaging function.
[0537] 5. In one example, the method could be a maximum function.
[0538] 6. In one example, the method could be a minimal function.
[0539] 7. In one example, the method could be a linear mapping function.
[0540] 8. In one example, the method could be a non-linear mapping function.
[0541] 9. Alternatively, the method can be any other function.
[0542] c) In one example, the predicted pattern mapping value or DBF-BS can be merged into K categories (e.g., K=2).
[0543] d) In one example, an existing classifier may have N classes (e.g., N=25).
[0544] e) In one example, the mapping value of the prediction pattern or DBF-BS can be used to directly expand the number of categories.
[0545] 1. In one example, the number of classes in an existing classifier can be extended to N. K (e.g., N) K=50).
[0546] 2. In one example, the classification result can be calculated using the following formula:
[0547]
[0548] Where c is the final category index. It is a category index generated by an existing classifier.
[0549] 3. In one example, the classification result can be calculated using the following formula:
[0550]
[0551] Where c is the final category index. It is a category index generated by an existing classifier.
[0552] f) In one example, the mapping value of the prediction pattern or DBF-BS can be used to modify the number of categories.
[0553] 1. In one example, the number of classes in an existing classifier can be modified to M (e.g., M=12).
[0554] a. In one example, for a band-based classifier, the total number of classes can be modified.
[0555] b. In one example, for a residual-based classifier, the total number of classes can be modified.
[0556] c. In one example, for a texture-based classifier, the number of texture directions can be modified.
[0557] d. In one example, for a texture-based classifier, the number of texture activities can be modified.
[0558] 2. In one example, the number of classes in the existing classifier can be modified to M. K (e.g., M) K=24).
[0559] 3. In one example, the classification result can be calculated using the following formula:
[0560]
[0561] Where c is the final category index. It is a category index generated by an existing classifier.
[0562] 4. In one example, the classification result can be calculated using the following formula:
[0563]
[0564] Where c is the final category index. It is a category index generated by an existing classifier.
[0565] 5) It is proposed to use the mapping value of the prediction mode as side information in other encoding and decoding stages.
[0566] a. In one example, the mapping value of the prediction mode can be used in a bilateral filter.
[0567] a) In one example, the mapping value of the predicted pattern can be used in the classification of the bilateral filter.
[0568] b) In one example, different mapping values of the prediction mode can lead to different filter strengths for the bilateral filter.
[0569] b. In one example, the mapping value of the prediction mode can be used in the HTDF filter.
[0570] a) In one example, the mapping value of the predicted pattern can be used in the classification of the HTDF filter.
[0571] b) In one example, different mapping values of the prediction mode can lead to different filter strengths of the HTDF filter.
[0572] c. In one example, the mapping value of the prediction pattern can be used in SAO.
[0573] a) In one example, the mapping value of the predicted pattern can be used in the classification of SAO.
[0574] 1. In one example, the mapping value of the predicted pattern can be used as an additional classification pattern.
[0575] 2. In one example, the mapping value of the predicted pattern can be used as side information of the existing classification pattern.
[0576] b) In one example, the mapping value of the prediction pattern can be used to generate compensation for the SAO.
[0577] d. In one example, the mapping value of the prediction pattern can be used in CCSAO.
[0578] a) In one example, the mapping value of the predicted pattern can be used in the classification of CCSAO.
[0579] 1. In one example, the mapping value of the predicted pattern can be used as an additional classification pattern.
[0580] 2. In one example, the mapping value of the predicted pattern can be used as side information of the existing classification pattern.
[0581] b) In one example, the mapping value of the prediction pattern can be used to generate compensation for CCSAO.
[0582] e. In one example, the mapping value of the prediction mode can be used in any other encoding / decoding tool / stage.
[0583] 6) In one example, whether / how to modify / prepare the mapping value of the prediction mode can be transmitted via signaling in the bitstream.
[0584] a. In one example, the first syntax element can be signaled to indicate whether the mapping value of the modified / prepared prediction pattern is enabled / used.
[0585] a) In one example, the first syntax element can be encoded or decoded using arithmetic encoding / decoding.
[0586] 1. In one example, the first syntax element can be encoded or decoded using at least one context.
[0587] a. The context can depend on the encoding / decoding information of the current block or neighboring blocks.
[0588] b. The context may depend on the filter shape of at least one neighboring block.
[0589] 2. In one example, the first syntax element can be encoded or decoded via bypass encoding / decoding.
[0590] b) In one example, the first syntax element can be binarized using unary code, rounded unary code, fixed-length code, exponential Golomb code, rounded exponential Golomb code, etc.
[0591] c) In one example, the first syntax element can be conditionally transmitted via a signal.
[0592] 1. For example, the first syntax element can only be signaled if the mapping value of the prediction mode is available.
[0593] d) The first syntax element can be encoded and decoded in a predictive manner.
[0594] 1. The first syntax element can be predicted by the on / off decision of the mapping value of the prediction mode of at least one neighboring block.
[0595] e) For different color components, the first syntax element can be transmitted independently via signal.
[0596] 1. Alternatively, for different color components, the first syntax element can be transmitted and shared via signal transmission.
[0597] 2. Alternatively, for the first color component but not for the second color component, the first syntax element can be transmitted via signal.
[0598] f) Syntax elements can be transmitted via signals in SPS / PPS / image header / strip header / APS / CTU / CU / etc.
[0599] b. In one example, the first syntax element can be signaled to indicate how the mapped values / boundary strengths of the modified / prepared prediction mode are used.
[0600] a) In one example, syntax elements can be signaled to indicate the number of categories to be combined with the predicted pattern information / boundary strength information before combining with existing categories.
[0601] 1. In one example, the number of categories to be merged can be N (e.g., N=2).
[0602] 2. In one example, the number of categories to be merged can be M (e.g., M=4).
[0603] b) In one example, syntax elements can be signaled to indicate whether prediction mode information and boundary strength information are used together.
[0604] c) In one example, the first syntax element can be encoded or decoded using arithmetic encoding / decoding.
[0605] 1. In one example, the first syntax element can be encoded or decoded using at least one context.
[0606] a. The context can depend on the encoding / decoding information of the current block or neighboring blocks.
[0607] b. The context may depend on the filter shape of at least one neighboring block.
[0608] 2. In one example, the first syntax element can be encoded or decoded via bypass encoding / decoding.
[0609] d) In one example, the first syntax element can be binarized using unary code, rounded unary code, fixed-length code, exponential Golomb code, rounded exponential Golomb code, etc.
[0610] e) In one example, the first syntax element can be conditionally transmitted via a signal.
[0611] 1. For example, the first syntax element can only be signaled if the mapping value of the prediction mode is available.
[0612] f) The first syntax element can be encoded and decoded in a predictive manner.
[0613] 1. The first syntax element can be predicted by the on / off decision of the mapping value of the prediction mode of at least one neighboring block.
[0614] g) For different color components, the first syntax element can be transmitted independently via signal.
[0615] 1. Alternatively, for different color components, the first syntax element can be transmitted and shared via signal transmission.
[0616] 2. Alternatively, for the first color component but not for the second color component, the first syntax element can be transmitted via signal.
[0617] h) Syntax elements can be transmitted via signals in SPS / PPS / image header / strip header / APS / CTU / CU / etc.
[0618] 1. In one example, syntax elements can be transmitted via signals at the SPS / PPS level.
[0619] 2. In one example, syntax elements can be transmitted via signals at the CTU level.
[0620] 3. In one example, syntax elements can be transmitted via signals at the APS level.
[0621] a. In one example, syntax elements can be transmitted via signals for all luminance / chrominance filters within all alternative filter sets.
[0622] b. In one example, syntax elements can be signaled for all luminance / chrominance filters within an alternative filter set.
[0623] 7) In one example, the disclosed methods may be used in post-processing and / or pre-processing.
[0624] 8) In one example, the above methods can be used in combination.
[0625] 9) Alternatively, the above methods can be used alone.
[0626] 10) In one example, the proposed method of using the mapping value of the prediction mode for ALF can be applied to any loop filtering tool, preprocessing or postprocessing filtering method in video encoding and decoding (including but not limited to ALF / CCALF / BF / SAO / CCSAO or any other filtering method).
[0627] a. In one example, the proposed method of using the mapping value of the prediction pattern can be applied to the loop filtering method.
[0628] a) In one example, the proposed method of using the mapping values of the prediction pattern can be applied to ALF.
[0629] b) In one example, the proposed method of using the mapping values of the prediction pattern can be applied to CCALF.
[0630] c) Alternatively, the proposed method of using the mapping values of the prediction mode can be applied to other loop filtering methods.
[0631] b. In one example, the proposed method for mapping the predicted pattern values can be applied to a preprocessing filtering method.
[0632] c. In one example, the proposed method for mapping the predicted pattern values can be applied to post-processing filtering methods.
[0633] 11) In the above examples, a video unit can refer to a sequence / picture / subpicture / strip / piece / code-decode tree unit (CTU) / CTU line / CTU group / code-decode unit (CU) / prediction unit (PU) / transform unit (TU) / code-decode tree block (CTB) / code-decode block (CB) / prediction block (PB) / transform block (TB) / any other region containing more than one luminance or chrominance sample / pixel.
[0634] 12) Whether and / or how the methods disclosed above can be applied to transmit signals in a bitstream.
[0635] a. In one example, whether and / or how the methods disclosed above can be applied can be signaled at the sequence level / picture group level / picture level / strip level / piece group level, such as in the sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / strip header / piece group header.
[0636] b. In one example, they can be transmitted via signal at PB / TB / CB / PU / TU / CU / VPDU / CTU / CTU lines / strips / pieces / sub-pictures / other types of areas containing more than one sample point or pixel.
[0637] 13) Whether and / or how the methods disclosed above are applied may depend on the encoded / decoded information, such as block size, color format, single / dual tree segmentation, color components, and stripe / picture type.
[0638] 14) The syntax elements disclosed above can be binarized into flags, fixed-length codes, EG(x) codes, unary codes, rounded unary codes, rounded binary codes, etc. The syntax element can be signed or unsigned.
[0639] 15) The syntax elements disclosed above can be encoded or decoded using at least one context model. Alternatively, the syntax element can be encoded or decoded in a bypass manner.
[0640] 16) The syntax elements disclosed above can be transmitted via signals in a conditional manner.
[0641] a. The SE will only be transmitted via signal if the corresponding function is applicable.
[0642] b. The SE will only be transmitted via signal if the dimensions (width and / or height) of the block meet the conditions.
[0643] 17) The syntax elements disclosed above can be transmitted via signaling at the block level / sequence level / picture group level / picture level / strip level / piece group level, such as in the codec structure of CTU / CU / TU / PU / CTB / CB / TB / PB, or in the sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / strip header / piece group header.
[0644] 18) The proposed methods (multiple methods) can be combined with another codec tool, such as affine / MTS / LFNST / MMVD / MIP / ISP / CCLM / CCCM / SMVD / BDOF / DMVR / HMVP / template matching / IBC / palette / etc.
[0645] 19) The proposed methods (multiple methods) can be excluded by another codec tool, such as affine / MTS / LFNST / MMVD / MIP / ISP / CCLM / CCCM / SMVD / BDOF / DMVR / HMVP / template matching / IBC / palette / etc.
[0646] a. In one example, if the proposed method(s) are used, the excluded codec tools are implicitly disabled without signaling.
[0647] b. In one example, if the excluded codec tool is used, the proposed method(s) are implicitly disabled without signaling.
[0648] 6. Further list of solutions and implementation examples
[0649] 1) The use of DBF-BS as side information for loop filters such as SAO / BIF / ALF / CCALF is proposed.
[0650] a. In one example, the DBF-BS for at least one location may be stored during or after the DBF process.
[0651] a) In one example, DBF-BS can be stored at at least one location.
[0652] 1. In one example, the DBF-BS of the luminance location can be stored.
[0653] 2. In one example, the DBF-BS of the chromaticity location can be stored.
[0654] b) In one example, the DBF-BS of the location can be initialized with a specific value.
[0655] 1. In one example, the DBF-BS of the location can be initialized with 0.
[0656] 2. In one example, the DBF-BS of a location can be initialized using N (e.g., N=32).
[0657] 3. In one example, the initial value may be predefined.
[0658] 4. In one example, the initial value can be derived.
[0659] 5. In one example, the initial value can be transmitted to the decoder via a signal.
[0660] c) In one example, the location's DBF-BS can be stored / used in the original domain.
[0661] 1. In one example, the DBF-BS of a location can be stored / used using default values representing different boundary strengths (e.g., BS=0,1,2).
[0662] d) In one example, the DBF-BS for each a can be stored / used in the mapping domain.
[0663] 1. In one example, the DBF-BS at each location can be stored / used using mapping values representing different boundary strengths (e.g., BS=128, 256, 512).
[0664] 2. In one example, the mapping value can be predefined.
[0665] 3. In one example, the mapping value can be derived.
[0666] 4. In one example, the mapped value can be transmitted to the decoder via a signal.
[0667] e) In one example, the DBF-BS of a location can be merged into one or more categories.
[0668] 1. In one example, the BDF-BS of a location can be merged into N categories (e.g., N=2).
[0669] 2. In one example, the DBF-BS of a location can be merged according to predefined rules.
[0670] a. In one example, the predefined rule could be a threshold-based limiting / mapping function.
[0671] b. In one example, DBF-BS can be set to V0 when DBF-BS is greater than or equal to the threshold, and can be set to V1 when DBF-BS is less than the threshold.
[0672] 3. In one example, the number of merge rules or merge categories can be predefined.
[0673] 4. In one example, the number of merge rules or merge categories can be derived.
[0674] 5. In one example, the number of merge rules or merge categories can be transmitted to the decoder via a signal.
[0675] f) In one example, the DBF-BS of the location can be filtered before use.
[0676] 1. In one example, the DBF-BS of a location can be filtered by a predefined filter.
[0677] 2. In one example, the DBF-BS of a location can be filtered by an offline-trained filter of an ALF.
[0678] 3. In one example, the DBF-BS of a location can be filtered by an online-trained filter.
[0679] a. In one example, an online-trained filter can be transmitted to the decoder via a signal.
[0680] 4. In one example, the DBF-BS of a location can be filtered by any other filter.
[0681] a. In one example, the DBF-BS of a location can be filtered by a Gaussian filter.
[0682] b. In one example, the DBF-BS of a location can be filtered by a Sobel / Prewitt / Roberts / Canny filter.
[0683] c. In one example, the DBF-BS of the location can be filtered by a Hadamard transform domain filter.
[0684] d. In one example, the DBF-BS of a location can be filtered by a bilateral filter.
[0685] e. In one example, the DBF-BS of a location can be filtered by a low-pass filter.
[0686] f. In one example, the DBF-BS of a location can be filtered by a high-pass filter.
[0687] 5. In one example, whether to filter DBF-BS or which filter to apply to DBF-BS can be predefined.
[0688] 6. In one example, whether to filter the DBF-BS or which filter to apply to the DBF-BS can be derived.
[0689] 7. In one example, whether to filter the DBF-BS or which filter to apply to the DBF-BS can be transmitted to the decoder via a signal.
[0690] g) In one example, DBF-BS can be clipped to range / bit depth.
[0691] 1. In one example, the DBF-BS can be clipped to a predefined / transmitted / derived clipping range.
[0692] 2. In one example, the DBF-BS can be clipped to a predefined / transmitted / derived N-bit depth (e.g., N=10).
[0693] h) In one example, DBF-BS can be scaled to range / bit depth.
[0694] 1. In one example, DBF-BS can be scaled to a predefined / signal-transmitted / derived range.
[0695] 2. In one example, DBF-BS can be scaled to a predefined / transmitted / derived N-bit depth (e.g., N=10).
[0696] i) In one example, DBF-BS can be transformed to range / domain / bit depth.
[0697] 1. In one example, DBF-BS can be transformed into a predefined / transmitted / derived range.
[0698] 2. In one example, DBF-BS can be transformed into a predefined / signal-transmitted / derived domain.
[0699] 3. In one example, DBF-BS can be transformed to a predefined / transmitted / derived N-bit depth (e.g., N=10).
[0700] j) In one example, the DBF-BS for at least one location can be derived without invoking the DBF procedure.
[0701] 1. For example, the DBF-BS of a location can be derived and stored, but the location may not be filtered by DBF.
[0702] 2) The use of DBF-BS as side information in ALF / CCALF filtering is proposed.
[0703] a. For example, f(Bs) can be used as side information in ALF / CCALF filtering, where “Bs” is the boundary strength and “f” is any function, such as a linear function, a limiting function, or any other function.
[0704] b. In one example, DBF-BS can be used as input to ALF.
[0705] a) In one example, DBF-BS can be used as input to at least one extended tap of ALF.
[0706] b) In one example, DBF-BS can be used as at least one existing tap of ALF.
[0707] c) In one example, the DBF-BS of the luminance can be used as the input for the ALF luminance.
[0708] d) In one example, the DBF-BS of the chromaticity can be used as input for the ALF chromaticity.
[0709] e) In one example, the filtered DBF-BS of the luminance can be used as the input to the ALF luminance.
[0710] f) In one example, the filtered DBF-BS of the chroma can be used as the input to the ALF chroma.
[0711] c. In one example, DBF-BS can be used as input to CCALF.
[0712] a) In one example, DBF-BS can be used as input to at least one extended tap of CCALF.
[0713] b) In one example, DBF-BS can be used as input to at least one existing tap of CCALF.
[0714] c) In one example, the luminance DBF-BS can be used as input to the extended tap of CCALF.
[0715] d) In one example, the filtered DBF-BS of the luminance can be used as input to CCALF.
[0716] d. In one example, DBF-BS can be used after the ALF / CCALF filtering function.
[0717] a) In one example, the difference between the DBF-BS or the filtered DBF-BS and the currently filtered sample may need to be calculated first.
[0718] b) In one example, the filter coefficients can be multiplied by the difference to obtain the final offset.
[0719] e. In one example, DBF-BS can be used in a different way than the ALF / CCALF filtering function.
[0720] a) In one example, the DBF-BS or filtered DBF-BS value can be used directly.
[0721] b) In one example, the filter coefficients can be multiplied by DBF-BS to obtain the final offset.
[0722] 3) It is proposed to use DBF-BS as edge information in ALF classification.
[0723] a. For example, f(Bs) can be used as edge information in ALF classification, where “Bs” is the boundary strength and “f” is any function, such as a linear function, a limiting function, or any other function.
[0724] b. In one example, DBF-BS / filtered DBF-BS can be used as an additional classification pattern.
[0725] a) In one example, the classification unit size of a DBF-BS-based classifier can be N. M (e.g., N=M=2).
[0726] b) In one example, locations within a classification unit can share the same classification result.
[0727] c) In one example, the classification results can be generated / derived using a specific method.
[0728] 1. In one example, the method may be predefined.
[0729] 2. In one example, the method can be derived.
[0730] 3. In one example, this method can be transmitted to the decoder via a signal.
[0731] 4. In one example, the method could be an averaging function.
[0732] 5. In one example, the method could be a maximum function.
[0733] 6. In one example, the method could be a minimal function.
[0734] 7. In one example, the method could be a linear mapping function.
[0735] 8. In one example, the method could be a non-linear mapping function.
[0736] 9. Alternatively, the method can be any other function.
[0737] d) In one example, whether a DBF-BS-based classifier is used can be predefined / derived / transmitted via signaling.
[0738] c. In one example, DBF-BS / filtered DBF-BS can be used as side information for existing classification patterns.
[0739] a) In one example, locations within a classifier unit can share the same DBF-BS value.
[0740] b) In one example, the DBF-BS value within a classification unit can be generated / derived using a specific method.
[0741] 1. In one example, the method may be predefined.
[0742] 2. In one example, the method can be derived.
[0743] 3. In one example, this method can be transmitted to the decoder via a signal.
[0744] 4. In one example, the method could be an averaging function.
[0745] 5. In one example, the method could be a maximum function.
[0746] 6. In one example, the method could be a minimal function.
[0747] 7. In one example, the method could be a linear mapping function.
[0748] 8. In one example, the method could be a non-linear mapping function.
[0749] 9. Alternatively, the method can be any other function.
[0750] c) In one example, DBF-BS can be merged into K categories (e.g., K=2).
[0751] d) In one example, an existing classifier may have N classes (e.g., N=25).
[0752] e) In one example, the merged DBF-BS can be used to directly expand the number of categories.
[0753] 1. In one example, the number of classes in an existing classifier can be extended to N. K (e.g., N) K=50).
[0754] 2. In one example, the classification result can be calculated using the following formula:
[0755]
[0756] Where c is the final category index. It is a category index generated by an existing classifier.
[0757] 3. In one example, the classification result can be calculated using the following formula:
[0758]
[0759] Where c is the final category index. It is a category index generated by an existing classifier.
[0760] f) In one example, the merged DBF-BS can be used to modify the number of categories.
[0761] 1. In one example, the number of classes in an existing classifier can be modified to M (e.g., M=12).
[0762] a. In one example, for a band-based classifier, the total number of bands can be modified.
[0763] b. In one example, for a texture-based classifier, the number of texture directions can be modified.
[0764] c. In one example, for a texture-based classifier, the number of texture activities can be modified.
[0765] 2. In one example, the number of classes in the existing classifier can be modified to M. K (e.g., M) K=24).
[0766] 3. In one example, the classification result can be calculated using the following formula:
[0767]
[0768] Where c is the final category index. It is a category index generated by an existing classifier.
[0769] 4. In one example, the classification result can be calculated using the following formula:
[0770]
[0771] Where c is the final category index. It is a category index generated by an existing classifier.
[0772] 4) It is proposed to use DBF-BS as side information in other encoding and decoding stages.
[0773] a. In one example, DBF-BS can be used in a bilateral filter.
[0774] a) In one example, DBF-BS can be used in the classification of bilateral filters.
[0775] b) In one example, different DBF-BS values can result in different filter strengths for the bilateral filter.
[0776] b. In one example, DBF-BS can be used in an HTDF filter.
[0777] a) In one example, DBF-BS can be used in the classification of HTDF filters.
[0778] b) In one example, different DBF-BS values can result in different filter strengths for the HTDF filter.
[0779] c. In one example, DBF-BS can be used in SAO.
[0780] a) In one example, DBF-BS can be used in the classification of SAO.
[0781] 1. In one example, DBF-BS can be used as an additional classification pattern.
[0782] 2. In one example, DBF-BS can be used as side information for an existing classification pattern.
[0783] b) In one example, DBF-BS can be used to generate compensation for SAO.
[0784] d. In one example, DBF-BS can be used in CCSAO.
[0785] a) In one example, DBF-BS can be used in CCSAO classification.
[0786] 1. In one example, DBF-BS can be used as an additional classification pattern.
[0787] 2. In one example, DBF-BS can be used as side information for an existing classification pattern.
[0788] b) In one example, DBF-BS can be used to generate compensation for CCSAO.
[0789] e. In one example, DBF-BS can be used in any other codec tool / stage.
[0790] 5) In one example, whether / how to modify / prepare DBF-BS for transmission via signaling in a bitstream.
[0791] a. In one example, the first syntax element can be signaled to indicate whether the modified / prepared DBF-BS is enabled / used.
[0792] a) In one example, the first syntax element can be encoded or decoded using arithmetic encoding / decoding.
[0793] 1. In one example, the first syntax element can be encoded or decoded using at least one context.
[0794] a. The context can depend on the encoding / decoding information of the current block or neighboring blocks.
[0795] b. The context may depend on the filter shape of at least one neighboring block.
[0796] 2. In one example, the first syntax element can be encoded or decoded via bypass encoding / decoding.
[0797] b) In one example, the first syntax element can be binarized using unary code, rounded unary code, fixed-length code, exponential Golomb code, rounded exponential Golomb code, etc.
[0798] c) In one example, the first syntax element can be conditionally transmitted via a signal.
[0799] 1. For example, the first syntax element can only be transmitted via signaling if DBF-BS is available.
[0800] d) The first syntax element can be encoded and decoded in a predictive manner.
[0801] 1. The first syntax element can be predicted by the on / off decision of the DBF-BS of at least one neighboring block.
[0802] e) For different color components, the first syntax element can be transmitted independently via signal.
[0803] 1. Alternatively, for different color components, the first syntax element can be transmitted and shared via signal transmission.
[0804] 2. Alternatively, for the first color component but not for the second color component, the first syntax element can be transmitted via signal.
[0805] f) Syntax elements can be transmitted via signals in SPS / PPS / image header / strip header / APS / CTU / CU / etc.
[0806] 6) In one example, whether DBF-BS can be applied to ALF and / or CCALF may depend on the control information of DBF.
[0807] a. For example, if DBF is turned off, DBF-BS cannot be applied to ALF and / or CCALF.
[0808] 7) In one example, whether samples after / before DBF can be applied to ALF and / or CCALF may depend on the control information of DBF.
[0809] a. For example, if DBF is turned off, samples after DBF cannot be applied to ALF and / or CCALF.
[0810] b. For example, if DBF is turned off, samples prior to DBF cannot be applied to ALF and / or CCALF.
[0811] 8) In one example, the disclosed methods may be used in post-processing and / or pre-processing.
[0812] 9) DBF-BS can be located in the same color component as the sample to be filtered, or in a different color component.
[0813] 10) DBF-BS can be located at the same location as the sample to be filtered, or within the range surrounding the sample to be filtered.
[0814] Figure 13 This is a block diagram illustrating an example video processing system 4000 in which various techniques disclosed herein may be implemented. Various implementations may include some or all of the components of system 4000. System 4000 may include an input 4002 for receiving video content. The video content may be received in a raw or uncompressed format, such as 8 or 10-bit multi-component pixel values, or it may be in a compressed or encoded format. Input 4002 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, Passive Optical Networking (PON), etc., and wireless interfaces such as Wi-Fi or cellular interfaces.
[0815] System 4000 may include an encoding / decoding component 4004 capable of implementing the various encoding / decoding or coding methods described in this disclosure. Encoding / decoding component 4004 may reduce the average bit rate from the video input 4002 to the output of encoding / decoding component 4004 to produce an encoded / decoded representation of the video. Encoding / decoding techniques are therefore sometimes referred to as video compression or video transcoding techniques. The output of encoding / decoding component 4004 may be stored or transmitted via a communication connection, as represented by component 4006. The stored or communicatively transmitted bitstream (or encoded / decoded) representation of the video received at input 4002 may be used by component 4008 to generate pixel values or displayable video that is sent to display interface 4010. The process of generating user-visible video from the bitstream representation is sometimes referred to as video decompression. Furthermore, although certain video processing operations are referred to as “encoding / decoding” operations or tools, it should be understood that encoding / decoding tools or operations are used by the encoder, and the corresponding decoding tools or operations that inversely convert the encoding / decoding results will be performed by the decoder.
[0816] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB), High Definition Multimedia Interface (HDMI), or DisplayPort, etc. Examples of storage interfaces include Serial Advanced Technology Attachment (SATA), Peripheral Component Interconnect (PCI), Integrated Drive Electronics (IDE), etc. The technologies described in this disclosure can be embodied in a variety of electronic devices, such as mobile phones, laptops, smartphones, or other devices capable of performing digital data processing and / or video display.
[0817] Figure 14 This is a block diagram of an example video processing apparatus 4100. Apparatus 4100 can be used to implement one or more methods described herein. Apparatus 4100 can be embodied in a smartphone, tablet, computer, Internet of Things (IoT) receiver, etc. Apparatus 4100 may include one or more processors 4102, one or more memories 4104, and video processing circuitry 4106. The processors(multiple) 4102 can be configured to implement one or more methods described herein. The memories(multiple) 4104 can be used to store data and code used to implement the methods and techniques described herein. The video processing circuitry 4106 can be used to implement some of the techniques described herein in hardware circuitry. In some embodiments, the video processing circuitry 4106 may be at least partially included in the processor 4102, for example, a graphics coprocessor.
[0818] Figure 15 This is a flowchart of an example method 4200 for video processing. Method 4200 includes using a prediction pattern as side information for a loop filter in step 4202. In step 4204, a conversion is performed between visual media data and a bitstream based on the loop filter. According to the example, the conversion in step 4204 may include encoding at an encoder or decoding at a decoder.
[0819] It should be noted that method 4200 can be implemented in an apparatus for processing video data, including a processor and a non-transitory memory having instructions thereon, such as a video encoder 4400, a video decoder 4500, and / or an encoder 4600. In this case, the instructions, when executed by the processor, cause the processor to perform method 4200. Furthermore, method 4200 can be executed by a non-transitory computer-readable medium, which includes a computer program product for use by a video codec device. The computer program product includes computer-executable instructions stored on the non-transitory computer-readable medium, causing the video codec device to perform method 4200 when executed by a processor.
[0820] Figure 16This is a block diagram illustrating an example video encoding / decoding system 4300 that can utilize the techniques disclosed herein. The video encoding / decoding system 4300 may include a source device 4310 and a destination device 4320. The source device 4310 generates encoded video data, and may be referred to as a video encoding device. The destination device 4320 can decode the encoded video data generated by the source device 4310, and may be referred to as a video decoding device.
[0821] Source device 4310 may include video source 4312, video encoder 4314, and input / output (I / O) interface 4316. Video source 4312 may include sources such as video capture devices, interfaces for receiving video data from video content providers, and / or computer graphics systems for generating video data, or combinations of these sources. Video data may include one or more pictures. Video encoder 4314 encodes the video data from video source 4312 to generate a bitstream. The bitstream may include a sequence of bits forming a codec representation of the video data. The bitstream may include codec pictures and associated data. Codec pictures are codec representations of pictures. Associated data may include sequence parameter sets, picture parameter sets, and other syntax structures. I / O interface 4316 may include a modulator / demodulator (modem) and / or transmitter. Encoded video data may be transmitted directly to destination device 4320 via network 4330 through I / O interface 4316. Encoded video data may also be stored on storage medium / server 4340 for access by destination device 4320.
[0822] The destination device 4320 may include an I / O interface 4326, a video decoder 4324, and a display device 4322. The I / O interface 4326 may include a receiver and / or a modem. The I / O interface 4326 may acquire encoded video data from the source device 4310 or the storage medium / server 4340. The video decoder 4324 may decode the encoded video data. The display device 4322 may display the decoded video data to a user. The display device 4322 may be integrated with the destination device 4320 or may be external to the destination device 4320, wherein the destination device 4320 may be configured to interface with an external display device.
[0823] The video encoder 4314 and the video decoder 4324 can operate according to video compression standards, such as the High Efficiency Video Codec (HEVC) standard, the Multi-Functional Video Codec (VVC) standard, and other existing and / or further standards.
[0824] Figure 17 This is a block diagram illustrating an example of a video encoder 4400, which can be... Figure 16The system 4300 shown includes a video encoder 4314. The video encoder 4400 can be configured to perform any or all of the techniques disclosed herein. The video encoder 4400 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video encoder 4400. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.
[0825] The functional components of the video encoder 4400 may include a segmentation unit 4401, a prediction unit 4402, a residual generation unit 4407, a transform processing unit 4408, a quantization unit 4409, an inverse quantization unit 4410, an inverse transform unit 4411, a reconstruction unit 4412, a buffer 4413, and an entropy coding unit 4414. The prediction unit 4402 may include a mode selection unit 4403, a motion estimation unit 4404, a motion compensation unit 4405, and an intra-frame prediction unit 4406.
[0826] In other examples, the video encoder 4400 may include more, fewer, or different functional components. In one example, the prediction unit 4402 may include an intra-block copy (IBC) unit. The IBC unit can perform prediction in an IBC mode, where at least one reference picture is the picture in which the current video block is located.
[0827] Furthermore, some components such as the motion estimation unit 4404 and the motion compensation unit 4405 can be highly integrated, but for illustrative purposes, they are shown separately in the example of the video encoder 4400.
[0828] The segmentation unit 4401 can segment an image into one or more video blocks. The video encoder 4400 and the video decoder 4500 can support various video block sizes.
[0829] The mode selection unit 4403 can, for example, select a codec mode (intra-frame codec or inter-frame codec) based on the error result, and provide the resulting intra-frame or inter-frame codec block to the residual generation unit 4407 to generate residual block data, and provide it to the reconstruction unit 4412 to reconstruct the coded block for use as a reference image. In some examples, the mode selection unit 4403 can select the Combination of Intra-and Inter-Prediction (CIIP) mode, where prediction is based on inter-frame prediction signals and intra-frame prediction signals. In the case of inter-frame prediction, the mode selection unit 4403 can also select the resolution of the motion vector for the block (e.g., sub-pixel precision or integer pixel precision).
[0830] To perform inter-frame prediction on the current video block, motion estimation unit 4404 can generate motion information for the current video block by comparing one or more reference frames from buffer 4413 with the current video block. Motion compensation unit 4405 can determine the predicted video block for the current video block based on the motion information and decoded samples of images from buffer 4413 other than the image associated with the current video block.
[0831] The motion estimation unit 4404 and the motion compensation unit 4405 can perform different operations on the current video block, for example, depending on whether the current video block is in an I-band, P-band, or B-band.
[0832] In some examples, motion estimation unit 4404 can perform unidirectional prediction on the current video block, and can search for a reference video block for the current video block in the reference images of list 0 or list 1. Motion estimation unit 4404 can then generate a reference index indicating the reference image (which contains the reference video block) in list 0 or list 1 and a motion vector indicating the spatial displacement between the current video block and the reference video block. Motion estimation unit 4404 can output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 4405 can generate a predicted video block for the current block based on the reference video block indicated by the motion information of the current video block.
[0833] In other examples, motion estimation unit 4404 can perform bidirectional prediction on the current video block. Motion estimation unit 4404 can search for a reference video block for the current video block in the reference images in list 0, and can also search for another reference video block for the current video block in the reference images in list 1. Motion estimation unit 4404 can then generate reference indices indicating the reference images containing the reference video blocks in lists 0 and 1, and motion vectors indicating the spatial displacement between the reference video blocks and the current video block. Motion estimation unit 4404 can output the reference index and motion vector of the current video block as motion information for the current video block. Motion compensation unit 4405 can generate a predicted video block for the current video block based on the reference video blocks indicated by the motion information of the current video block.
[0834] In some examples, the motion estimation unit 4404 can output a complete set of motion information for use in the decoder's decoding process. In some examples, the motion estimation unit 4404 may not output a complete set of motion information for the current video. Instead, the motion estimation unit 4404 can reference the motion information of another video block to transmit the motion information of the current video block via a signal. For example, the motion estimation unit 4404 may determine that the motion information of the current video block is sufficiently similar to the motion information of neighboring video blocks.
[0835] In one example, the motion estimation unit 4404 may indicate a value to the video decoder 4500 in the syntax structure associated with the current video block, the value indicating that the current video block has the same motion information as another video block.
[0836] In another example, motion estimation unit 4404 may identify another video block and a motion vector difference (MVD) within the syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. Video decoder 4500 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[0837] As discussed above, the video encoder 4400 can transmit motion vectors via signals in a predictive manner. Two examples of predictive signaling techniques that can be implemented by the video encoder 4400 include Advanced Motion Vector Prediction (AMVP) and Merge Pattern Signaling.
[0838] Intra-prediction unit 4406 can perform intra-prediction on the current video block. When intra-prediction unit 4406 performs intra-prediction on the current video block, it can generate prediction data for the current video block based on decoded samples from other video blocks in the same frame. The prediction data for the current video block may include the predicted video block and various syntax elements.
[0839] The residual generation unit 4407 can generate residual data for the current video block by subtracting (or more) predicted video blocks from the current video block. The residual data for the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.
[0840] In other examples, such as in skip mode, there may be no residual data for the current video block, and the residual generation unit 4407 may not perform subtraction operations.
[0841] The transform processing unit 4408 can generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video blocks associated with the current video block.
[0842] After the transform processing unit 4408 generates a transform coefficient video block associated with the current video block, the quantization unit 4409 can quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0843] The inverse quantization unit 4410 and the inverse transform unit 4411 can apply inverse quantization and inverse transform to the transform coefficient video block respectively to reconstruct the residual video block from the transform coefficient video block. The reconstruction unit 4412 can add the reconstructed residual video block to the corresponding samples of one or more predicted video blocks generated by the prediction unit 4402 to generate a reconstructed video block associated with the current block and store it in the buffer 4413.
[0844] After the video block is reconstructed by reconstruction unit 4412, a loop filtering operation can be performed to reduce video block artifacts in the video block.
[0845] The entropy coding unit 4414 can receive data from other functional components of the video encoder 4400. When the entropy coding unit 4414 receives data, it can perform one or more entropy coding operations to generate entropy-coded data and output a bitstream including the entropy-coded data.
[0846] Figure 18 This is a block diagram illustrating an example of a video decoder 4500, which can be... Figure 16 The system 4300 shown includes a video decoder 4324. The video decoder 4500 can be configured to perform any or all of the techniques disclosed herein. In the example shown, the video decoder 4500 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video decoder 4500. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.
[0847] In the example shown, the video decoder 4500 includes an entropy decoding unit 4501, a motion compensation unit 4502, an intra-frame prediction unit 4503, an inverse quantization unit 4504, an inverse transform unit 4505, a reconstruction unit 4506, and a buffer 4507. In some examples, the video decoder 4500 can perform a decoding process that is generally contrasted with the encoding process described with respect to the video encoder 4400.
[0848] The entropy decoding unit 4501 can retrieve the encoded bitstream. The encoded bitstream may include entropy-encoded video data (e.g., encoded video data blocks). The entropy decoding unit 4501 can decode the entropy-encoded video data, and the motion compensation unit 4502 can determine motion information from the entropy-decoded video data. Motion information includes motion vectors, motion vector precision, reference image list index, and other motion information. The motion compensation unit 4502 can determine this information, for example, by executing AMVP and Merge modes.
[0849] The motion compensation unit 4502 can generate motion compensation blocks, possibly performing interpolation based on an interpolation filter. The identifier of the interpolation filter to be used, with sub-pixel accuracy, can be included in the syntax element.
[0850] The motion compensation unit 4502 can use the interpolation filter used by the video encoder 4400 during the encoding of the video block to calculate the sub-integer pixel interpolation of the reference block. The motion compensation unit 4502 can determine the interpolation filter used by the video encoder 4400 based on the received syntax information, and the motion compensation unit 4502 can use the interpolation filter to generate the prediction block.
[0851] The motion compensation unit 4502 may use some syntax information to determine the size of the blocks used to encode (multiple) frames and / or (multiple) stripes of the encoded video sequence, segmentation information describing how each macroblock of the image of the encoded video sequence is segmented, a mode indicating how each segment is encoded, one or more reference frames (and a list of reference frames) for each inter-frame codec block, and other information for decoding the encoded video sequence.
[0852] Intra-prediction unit 4503 can use, for example, an intra-prediction mode received in the bitstream to form prediction blocks from spatially adjacent blocks. Inverse quantization unit 4504 inverse quantizes the quantized video block coefficients provided in the bitstream and decoded by entropy decoding unit 4501. Inverse transform unit 4505 applies the inverse transform.
[0853] The reconstruction unit 4506 can add the residual block to the corresponding predicted block generated by the motion compensation unit 4502 or the intra-frame prediction unit 4503 to form a decoded block. If necessary, a deblocking filter can also be used to filter the decoded block to remove block artifacts. The decoded video block is then stored in the buffer 4507 to provide a reference block for subsequent motion compensation / intra-frame prediction, and also generates decoded video for presentation on the display device.
[0854] Figure 19This is a schematic diagram of an example encoder 4600. Encoder 4600 is suitable for implementing VVC techniques. Encoder 4600 includes three loop filters: a deblocking filter (DF) 4602, a sample adaptive compensation (SAO) 4604, and an adaptive loop filter (ALF) 4606. Unlike DF 4602, which uses predefined filters, SAO 4604 and ALF 4606 utilize the original samples of the current image, respectively, by adding compensation and by applying a finite impulse response (FIR) filter, and by utilizing the encoded / decoded side information through signal transmission compensation and filter coefficients to reduce the mean square error between the original and reconstructed samples. ALF 4606 is located in the final processing stage of each image and can be considered as a tool for attempting to capture and repair artifacts caused by previous stages.
[0855] The encoder 4600 also includes an intra-frame prediction component 4608 and a motion estimation / compensation (ME / MC) component 4610 configured to receive input video. The intra-frame prediction component 4608 is configured to perform intra-frame prediction, while the ME / MC component 4610 is configured to perform inter-frame prediction using a reference image obtained from a reference image buffer 4612. Residual blocks from inter-frame or intra-frame prediction are fed into a transform (T) component 4614 and a quantization (Q) component 4616 to generate quantized residual transform coefficients, which are then fed into an entropy encoder / decoder component 4618. The entropy encoder / decoder component 4618 entropy-encodes and decodes the prediction results and quantized transform coefficients and transmits them toward a video decoder (not shown). The quantization component output from the quantization component 4616 can be fed into an inverse quantization (IQ) component 4620, an inverse transform component 4622, and a reconstruction (REC) component 4624. REC component 4624 is able to output images to DF 4602, SAO 4604 and ALF 4606 for filtering before these images are stored in reference image buffer 4612.
[0856] The following is a list of some preferred solutions.
[0857] The following solutions illustrate examples of the techniques discussed in this article.
[0858] 1. A method for processing video data, comprising: determining a loop filter that employs a receive prediction mode as side information for use as input; and performing a conversion between visual media data and a bitstream based on the loop filter.
[0859] 2. The method according to Solution 1, wherein the loop filter is a sample adaptive compensation (SAO) filter, a bilateral filter (BF), an adaptive loop filter (ALF), or a cross-component ALF (CCALF).
[0860] 3. The method according to solution 1 or 2, wherein the prediction pattern of at least one location is stored prior to the loop filter process.
[0861] 4. The method according to any one of solutions 1-3, wherein the stored prediction mode is used for luminance position or chromaticity position, or wherein the stored prediction mode is acquired by the loop filter.
[0862] 5. The method according to any one of solutions 1-4, wherein the prediction pattern for each location is stored or used in the mapping domain.
[0863] 6. The method according to any one of solutions 1-5, wherein the prediction mode of each position is stored or used using a mapping value representing a different prediction mode, or wherein the mode information is a prediction mode, a codec block flag (CBF) flag, or a root CBF flag, or wherein the mapping value is predefined, derived, or transmitted via signaling in the bit stream.
[0864] 7. The method according to any one of solutions 1-6, wherein the predicted pattern of location is classified into one or more categories.
[0865] 8. The method according to any one of solutions 1-7, wherein the prediction mode is classified into N categories, or wherein the prediction mode is classified according to predefined rules, or wherein locations with inter-frame prediction (INTER) modes are classified into one category, or wherein locations with INTER and intra-block copy (IBC) modes are classified into one category, or wherein locations with INTER, IBC, and intra-template matching prediction (IntraTMP) modes are classified into one category, or wherein locations with intra-frame prediction (INTRA) modes... Locations are classified into a category, or, where locations with a NON-INTRA mode are classified into a category, or, where locations with a CBF equal to 0 are classified into a category, or, where locations with a CBF not equal to 0 are classified into a category, or, where block boundary locations are classified into a category different from other locations within a block, or, where the number of merge rules or merge categories is predefined, or, where the number of merge rules or merge categories is derived, or, where the number of merge rules or merge categories is transmitted to the decoder via a signal.
[0866] 9. The method according to any one of solutions 1-8, wherein the prediction pattern is used as side information in ALF or CCALF filtering.
[0867] 10. The method according to any one of solutions 1-9, wherein f(pm) is used as side information in ALF or CCALF filtering, wherein “m” is the mapping value of the prediction mode, and “f” is any function including a linear function, a limiting function, or any other function; or wherein the mapping value of the prediction mode is used as the input of ALF or CCALF, the input of at least one extended tap of ALF or CCALF, or the input of at least one tap of ALF; or wherein the mapping value of the luminance prediction mode is used as the input of ALF luminance or CCALF; or wherein the mapping value of the chrominance prediction mode is used as the input of ALF chrominance; or wherein the mapping value of the prediction mode is used as the input of CCALF. Alternatively, the mapped value of the prediction mode may be used as the input to CCALF, the input to at least one extended tap of CCALF, or at least one tap of CCALF; or, the mapped value of the prediction mode may be used after the ALF or CCALF filtering function; or, the difference between the mapped value of the prediction mode and the currently filtered sample may be calculated first; or, the filter coefficients may be multiplied by the difference to obtain the final offset; or, the mapped value of the prediction mode may be used in a manner different from the ALF or CCALF filtering function; or, the mapped value of the prediction mode may be used directly; or, the filter coefficients may be multiplied by the mapped value of the prediction mode to obtain the final offset.
[0868] 11. The method according to any one of solutions 1-10, wherein the prediction pattern is used as side information in ALF or CCALF classification.
[0869] 12. The method according to any one of solutions 1-11, wherein f(mp) is used as side information in ALF or CCALF classification, wherein “mp” is the mapping value of the prediction mode, and “f” is any function including a linear function, a limiting function, or any other function; or, wherein the mapping value of the prediction mode is used in conjunction with a deblocking filter (DBF)-Boundary Strength (BS); or, wherein the prediction mode and DBF-BS are directly combined, wherein the prediction mode is classified into N categories, the DBF-BS is classified into M categories, and the total number of categories generated by the side information is equal to M×N; or, wherein the prediction mode and DBF-BS are conditionally combined, wherein the prediction mode is applied only to locations where the DBF-BS is equal to a value K; or, wherein K is predefined, transmitted by signaling, or derived; or, wherein the mapping value of the prediction mode is used as an additional classifier; or, wherein the classification unit size of the classifier based on the mapping value of the prediction mode is N. M, where N=M=2, or, where positions within a classification unit share the same classification result, or, where the classification result is generated or derived using a specific method, or, where the method is predefined, derived, or transmitted via signaling, or, where the method is an average function, a maximum function, a minimum function, a linear mapping function, a nonlinear mapping function, or any other function, or, where whether or not a classifier based on the mapping value of the prediction pattern is used is predefined, derived, or transmitted via signaling.
[0870] 13. The method according to any one of solutions 1-12, wherein the mapping value of the prediction pattern is used as side information of the classification pattern, or wherein positions within a classification unit share the same mapping value of the prediction pattern value, or wherein the mapping value of the prediction pattern value within a classification unit is generated or derived using a specific method, or wherein the method is predefined, derived, or transmitted via signaling, or wherein the method is an averaging function, a maximizing function, a minimizing function, a linear mapping function, a nonlinear mapping function, or any other function, or wherein the mapping values of the prediction pattern are merged into K categories, or wherein the classifier has N categories, or wherein the mapping values of the prediction pattern are used to directly expand the number of categories, or wherein the number of categories of the classifier is expanded to N. K, where N K=50, or, where the classification result is calculated using the following formula:
[0871]
[0872] Where c is the final category index. It is a category index generated by the classifier, or, where,
[0873] The classification result is calculated using the following formula:
[0874]
[0875] Where c is the final category index. It is a category index generated by a classifier.
[0876] 14. The method according to any one of solutions 1-13, wherein the mapping value of the predicted pattern is used to modify the number of categories, or wherein the number of categories of the classifier is modified to M, where M=12, or wherein, for a band-based classifier, the total number of categories is modified, or wherein, for a residual-based classifier, the total number of categories is modified, or wherein, for a texture-based classifier, the number of texture directions is modified, or wherein, for a texture-based classifier, the number of texture activities is modified, or wherein, the number of categories of the classifier is modified to M. K, where M K=24, or, where the classification result is calculated using the following formula:
[0877]
[0878] Where c is the final category index. It is the category index generated by the classifier, or, where the classification result is calculated using the following formula:
[0879]
[0880] Where c is the final category index. It is a category index generated by a classifier.
[0881] 15. The method according to any one of solutions 1-14, wherein the mapping value of the prediction mode is used as side information in other encoding / decoding stages.
[0882] 16. The method according to any one of solutions 1-15, wherein the mapping value of the prediction mode is used in a bilateral filter, or wherein the mapping value of the prediction mode is used in the classification of the bilateral filter, or wherein different mapping values of the prediction mode result in different filter strengths of the bilateral filter, or wherein the mapping value of the prediction mode is used in an HTDF filter, or wherein the mapping value of the prediction mode is used in the classification of the HTDF filter, or wherein different mapping values of the prediction mode result in different filter strengths of the HTDF filter, or wherein the mapping value of the prediction mode is used in SAO, or wherein the mapping value of the prediction mode is used in the classification of SAO, or In some cases, the mapping value of the prediction mode is used as an additional classification mode, or, the mapping value of the prediction mode is used as side information of the classification mode, or, the mapping value of the prediction mode is used to generate compensation for SAO, or, the mapping value of the prediction mode is used in CCSAO, or, the mapping value of the prediction mode is used in the classification of CCSAO, or, the mapping value of the prediction mode is used as an additional classification mode, or, the mapping value of the prediction mode is used as side information of the classification mode, or, the mapping value of the prediction mode is used to generate compensation for CCSAO, or, the mapping value of the prediction mode is used in any other encoding / decoding tool or stage.
[0883] 17. The method according to any one of solutions 1-16, wherein whether or how the mapping value of the prediction mode is modified or prepared is transmitted via signal in the bit stream.
[0884] 18. The method according to any one of solutions 1-17, wherein the first syntax element is signaled to indicate whether a modified or prepared prediction mode mapping value is enabled or used, or wherein the first syntax element is encoded / decoded by arithmetic encoding / decoding, or wherein the first syntax element is encoded / decoded using at least one context, or wherein the context depends on encoding / decoding information of the current block or neighboring blocks, or wherein the context depends on the filter shape of at least one neighboring block, or wherein the first syntax element is encoded / decoded by bypass encoding / decoding, or wherein the first syntax element is binarized by unary code, rounded unary code, fixed-length code, exponential Golomb code, or rounded exponential Golomb code, or wherein the first syntax element is conditionally signaled, or wherein only when the first syntax element is signaled... The first syntax element is transmitted via signaling only when the mapping value of the prediction mode is available, or, wherein the first syntax element is encoded / decoded in a predictive manner, or, wherein the first syntax element is predicted by an on / off decision of the mapping value of the prediction mode of at least one neighboring block, or, wherein the first syntax element is transmitted via signaling independently for different color components, or, wherein the first syntax element is transmitted and shared via signaling for different color components, or, wherein the first syntax element is transmitted via signaling for the first color component but not for the second color component, or, wherein the syntax element is transmitted via signaling in a Sequence Parameter Set (SPS), Picture Parameter Set (PPS), Picture Header, Strip Header, Adaptive Parameter Set (APS), Codec Tree Unit (CTU), or Codec Unit (CU).
[0885] 19. The method according to any one of solutions 1-18, wherein the method is used in post-processing or pre-processing.
[0886] 20. The method according to any one of solutions 1-19, wherein the methods are used in combination or individually.
[0887] 21. The method according to any one of solutions 1-20, wherein the mapping value of the prediction mode for ALF is applied to any loop filtering tool, preprocessing or postprocessing filtering method in video encoding and decoding, including ALF, CCALF, BF, SAO, CCSAO or any other filtering method, or wherein the mapping value of the prediction mode is applied to a loop filter, or wherein the mapping value of the prediction mode is applied to ALF, CCALF, any other loop filter, preprocessing filter or postprocessing filter.
[0888] 22. The method according to any one of solutions 1-21, wherein the loop filter is applied to a video unit, wherein the video unit is a sequence, picture, sub-picture, strip, slice, CTU, CTU line, CTU group, CU, prediction unit (PU), transform unit (TU), codec tree block (CTB), codec block (CB), prediction block (PB), transform block (TB), or any other region containing more than one luminance or chrominance sample or pixel.
[0889] 23. The method according to any one of solutions 1-22, wherein the use of the method is transmitted via signal in the bit stream.
[0890] 24. The method according to any one of solutions 1-23, wherein the use of the method is transmitted via signal at the sequence level, picture group level, picture level, strip level, or slice group level, including in the sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), decoding parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), strip header, or slice group header; or wherein the use of the method is transmitted via signal at the prediction block (PB), transform block (TB), codec block (CB), prediction unit (PU), transform unit (TU), codec unit (CU), virtual pipeline decoding unit (VPDU), codec tree unit (CTU), CTU row, strip, slice, sub-picture, or other region containing more than one sample or pixel.
[0891] 25. The method according to any one of solutions 1-24, wherein the application of the method depends on encoded / decoded information, the encoded / decoded information including block size, color format, single-tree segmentation, dual-tree segmentation, color components, stripe type, or image type.
[0892] 26. The method according to any one of solutions 1-25, wherein the syntax element is: binarized as a flag, fixed-length code, exponential Golomb code, unary code, rounded unary code or rounded binary code, signed or unsigned, encoded or decoded using a context model, encoded or decoded by bypass, conditionally transmitted by signaling, transmitted by signaling only when the corresponding function applies, transmitted by signaling when the height or width of the block satisfies a condition, or transmitted by signaling when both conditions are satisfied.
[0893] 27. The method according to any one of solutions 1-26, wherein syntax elements are transmitted via signals at the sequence level, picture group level, picture level, stripe level or slice group level, in the sequence header, picture header, SPS, VPS, DPS, decoding capability information (DCI), PPS, APS, stripe header or slice group header.
[0894] 28. The method according to any one of solutions 1-27, wherein the method is used in combination with affine, multiple transform selection (MTS), LFNST, Merge with motion vector difference (MMVD), matrix-based intra-prediction (MIP), ISP, cross-component linear model (CCLM), convolutional cross-component model (CCCM), symmetric motion vector difference (SMVD), bidirectional optical flow (BDOF), decoder-side motion vector refinement (DMVR), history-based motion vector prediction (HMVP), template matching, intra-block copy (IBC), or palette, or not in combination with affine, MTS, LFNST, MMVD, MIP, ISP, CCLM, CCCM, SMVD, BDOF, DMVR, HMVP, template matching, IBC, or palette.
[0895] 29. The method according to any one of solutions 1-28, wherein the excluded tool is implicitly disabled without signaling, or wherein the method is implicitly disabled without signaling when the excluded tool is used.
[0896] 30. The method according to any one of solutions 1-29, wherein the conversion includes encoding the visual media data into the bitstream.
[0897] 31. The method according to any one of solutions 1-30, wherein the conversion includes decoding the visual media data from the bitstream.
[0898] 32. An apparatus for processing video data, comprising: a processor; and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of solutions 1-31.
[0899] 33. A non-transitory computer-readable medium comprising a computer program product for use by a video codec device, the computer program product comprising computer-executable instructions stored on the non-transitory computer-readable medium such that, when executed by a processor, the video codec device performs the method according to any one of solutions 1-31.
[0900] 34. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by a video processing apparatus, wherein the method includes: determining a loop filter employing a receive prediction mode as side information for input; and generating the bitstream based on the determination.
[0901] 35. A method for storing a bitstream of video, comprising: determining a loop filter employing a receive prediction mode as input to side information; generating the bitstream based on the determination; and storing the bitstream in a non-transitory computer-readable recording medium.
[0902] 36. A method, apparatus or system described in this disclosure.
[0903] In the described solution, the encoder conforms to the format rules by generating a codec representation based on those rules. In the described solution, the decoder parses the syntax elements in the codec representation using known information about their presence or absence, based on the format rules, to produce the decoded video.
[0904] In this disclosure, the term "video processing" can refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm can be applied during the conversion from a pixel representation of a video to a corresponding bitstream representation, and vice versa. For example, the bitstream representation of a current video block can correspond to bits at co-positions or propagated at different positions in a bitstream defined by a syntax. For example, a macroblock can be encoded based on the error residual values after transformation and encoding / decoding, and can also use bits from the header and other fields in the bitstream. Furthermore, during the conversion, the decoder can parse the bitstream based on the determination, knowing whether certain fields may or may not be present, as described in the solutions above. Similarly, the encoder can determine whether to include or exclude specific syntax fields, and generate the codec representation accordingly by including or excluding syntax fields from the codec representation.
[0905] The disclosed and other solutions, examples, embodiments, modules, and functional operations described herein can be implemented in digital electronic circuits, or in computer software, firmware, or hardware, including the structures disclosed herein and their structural equivalents, or in combinations thereof. The disclosed and other embodiments can be implemented as one or more computer program products, i.e., one or more computer program instruction modules encoded on a computer-readable medium for execution by or control of the operation of a data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a storage device, a material composition affecting machine-readable propagation signals, or a combination thereof. The term "data processing apparatus" includes all means, devices, and machines for processing data, including, for example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, the apparatus may also include code that creates an execution environment for an associated computer program, for example, code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination thereof. Propagation signals are artificially generated signals, such as machine-generated electrical signals, optical signals, or electromagnetic signals, which are generated to encode information to be transmitted to a suitable receiver device.
[0906] Computer programs (also known as programs, software, software applications, scripts, or code) can be written in any programming language, including compiled or interpreted languages, and can be deployed in any form, including standalone programs or modules, components, subroutines, or other units suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple co-located files (e.g., a file storing one or more modules, subroutines, or code portions). A computer program can be deployed to execute on one computer or on multiple computers located at a single site or distributed across multiple sites and interconnected by a communications network.
[0907] The processing and logic flows described in this disclosure can be executed by one or more programmable processors that execute one or more computer programs to perform functions by manipulating input data and generating outputs. The processing and logic flows can also be executed by special-purpose logic circuitry, and the apparatus can also be implemented as special-purpose logic circuitry, such as field-programmable gate arrays (FPGAs) or application-specific integrated circuits (ASICs).
[0908] Processors suitable for executing computer programs include, for example, general-purpose and special-purpose microprocessors, and any one or more processors in any type of digital computer. Typically, the processor receives instructions and data from read-only memory or random access memory, or both. The basic components of a computer are a processor that executes instructions and one or more storage devices that store the instructions and data. Typically, a computer will also include one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks. However, a computer does not necessarily have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and storage devices, including, for example, semiconductor storage devices such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable hard disks; magneto-optical disks; and CD-ROMs and DVD-ROMs. The processor and memory may be supplemented by or incorporated into special-purpose logic circuitry.
[0909] While this disclosure contains numerous details, these details should not be construed as limiting any subject matter or the scope of the claims, but rather as descriptions of features specific to particular embodiments of a particular art. Certain features described in the context of individual embodiments in this disclosure may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented separately in multiple embodiments, or in any suitable sub-combination. Furthermore, although features may function in certain combinations as described above, and even were originally claimed in this manner, in some cases one or more features in the claimed combination may be removed from that combination, and the claimed combination may be for sub-combinations or variations thereof.
[0910] Similarly, although operations are depicted in a specific order in the accompanying drawings, this should not be construed as requiring such operations to be performed sequentially in the particular order or sequence shown, or requiring all shown operations to be performed in order to achieve the desired result. Furthermore, the partitioning of various system components in the embodiments described in this disclosure should not be construed as requiring such partitioning in all embodiments.
[0911] Only a few implementations and examples are described, and other implementations, improvements and variations may be made based on what is described and shown in this disclosure.
[0912] When there is no intermediary component other than a line, trace, or other medium between the first and second components, the first component is directly coupled to the second component. When there is an intermediary component other than a line, trace, or other medium between the first and second components, the first component is indirectly coupled to the second component. The term "coupled" and its variations include direct coupling and indirect coupling. The use of the term "about" means including a range of ±10% of the following figures, unless otherwise specified.
[0913] While several embodiments have been provided in this disclosure, it should be understood that the disclosed systems and methods may be embodied in many other specific forms without departing from the spirit or scope of this disclosure. The present examples are to be considered illustrative rather than restrictive and are not intended to be limited to the details set forth herein. For example, various elements or components may be combined or integrated into another system, or certain features may be omitted or not implemented.
[0914] Furthermore, the technologies, systems, subsystems, and methods described and illustrated as discrete or separate in the various embodiments may be combined or integrated with other systems, modules, technologies, or methods without departing from the scope of this disclosure. Other items shown or discussed as coupled may be directly connected or indirectly coupled or communicated through some interface, device, or intermediate component, whether electrical, mechanical, or otherwise. Other changes, substitutions, and modifications that can be determined by those skilled in the art may be made without departing from the spirit and scope of this disclosure.
Claims
1. A method for processing video data, comprising: The prediction mode is determined to be used as the side information for the loop filter; as well as The conversion between visual media data and bitstream is performed based on the loop filter.
2. The method according to claim 1, wherein, The loop filter is a sample adaptive compensation (SAO) filter, a bilateral filter (BF), an adaptive loop filter (ALF), or a cross-component ALF (CCALF).
3. The method according to claim 1 or 2, wherein, The prediction pattern for at least one location is stored before the loop filter is applied during the loop filtering process.
4. The method according to claim 1, wherein, The prediction pattern for at least one location is stored.
5. The method according to claim 1, wherein, The prediction pattern for brightness location is stored.
6. The method according to claim 1, wherein, The prediction pattern for chromaticity location is stored.
7. The method according to any one of claims 4-6, further comprising obtaining the stored prediction pattern by the loop filter.
8. The method according to any one of claims 1-7, wherein, The prediction pattern for each location is stored or used in the mapping domain.
9. The method according to claim 8, wherein, The prediction pattern for each location is stored or used using mapping values representing different prediction patterns.
10. The method according to claim 8, wherein, The pattern information includes the prediction pattern.
11. The method according to claim 8, wherein, The mode information includes the codec block flag (CBF).
12. The method according to claim 8, wherein, The mode information includes the root codec block flag (rootCBF).
13. The method according to claim 8, wherein, The mapping values are predefined.
14. The method according to claim 8, wherein, The mapping value is derived.
15. The method according to claim 8, wherein, The mapping values are included in the bitstream.
16. The method according to any one of claims 1-15, wherein, The predicted patterns of location can be classified into one or more categories.
17. The method according to claim 16, wherein, The prediction patterns are classified into N categories, where N is a positive integer.
18. The method according to claim 16, wherein, The prediction patterns are classified according to predefined rules.
19. The method of claim 16, wherein, Locations with the INTER pattern are classified into a category.
20. The method of claim 16, wherein, Locations with both INTER mode and Intra-Block Copy (IBC) mode are classified into one category.
21. The method according to claim 16, wherein, Locations with INTER mode, Intra-Block Copy (IBC) mode, and Intra-Template Match Prediction (IntraTMP) are classified into one category.
22. The method according to claim 16, wherein, Locations with the INTRA pattern are classified into a category.
23. The method according to claim 16, wherein, Locations with non-INTRA patterns are classified into one category.
24. The method of claim 16, wherein, The location where the code block flag (CBF) is equal to 0 is classified into a category.
25. The method according to claim 16, wherein, The location where the code block flag (CBF) is not equal to 0 is classified into a category.
26. The method of claim 16, wherein, Block boundary locations are classified into different categories than other locations within a block.
27. The method according to claim 16, wherein, The number of merge rules or merge categories is predefined.
28. The method according to claim 16, wherein, The number of merge rules or merge categories is derived.
29. The method according to claim 16, wherein, The number of merging rules or merging categories is included in the bitstream.
30. The method according to any one of claims 1-29, further comprising using the mapping value of the prediction mode as side information in other encoding / decoding stages.
31. The method of claim 30, further comprising using the mapping value of the prediction mode in a bilateral filter.
32. The method of claim 31, further comprising using the mapping value of the prediction mode in the classification of the bilateral filter.
33. The method according to claim 31 or 32, wherein, The mapping values for different prediction modes indicate the different filter strengths of the bilateral filter.
34. The method of claim 30, further comprising using the mapping value of the prediction mode in a Hadamard transform domain filter (HTDF).
35. The method of claim 34, further comprising using the mapping value of the prediction mode in the classification of the HTDF.
36. The method according to claim 34 or 35, wherein, The mapping values for different prediction modes indicate the different filter strengths of the HTDF.
37. The method of claim 30, further comprising using the mapping value of the prediction mode in a sample adaptive compensation (SAO) filter.
38. The method of claim 34, further comprising using the mapping value of the prediction mode in the classification of the SAO filter.
39. The method according to claim 38, wherein, The mapping value of the predicted pattern is used as an additional pattern for the classification.
40. The method according to claim 38, wherein, The mapping value of the predicted pattern is used as the side information of the existing pattern for classification by the SOA filter.
41. The method of claim 37, further comprising using the mapping value of the prediction mode when generating compensation for the SAO filter.
42. The method of claim 30, further comprising using the mapping value of the prediction mode in a cross-component sample adaptive compensation (CCSAO) filter.
43. The method of claim 34, further comprising using the mapping value of the prediction mode in the classification of the CCSAO filter.
44. The method according to claim 43, wherein, The mapping value of the predicted pattern is used as an additional pattern for classification by the CCSAO filter.
45. The method according to claim 43, wherein, The mapping value of the predicted pattern is used as the side information of the existing pattern for classification by the CCSAO filter.
46. The method of claim 42, further comprising using the mapping value of the prediction mode when generating compensation for the CCSAO filter.
47. The method according to any one of claims 31-46, wherein, The mapping values of the prediction mode are used in other encoding / decoding tools or other encoding / decoding stages.
48. The method according to any one of claims 1-47, wherein, Whether or how the mapping values of the prediction mode are prepared or modified are included in the bitstream.
49. The method according to claim 48, wherein, The first syntax element is included in the bitstream to indicate whether the prepared or modified mapping value is used.
50. The method according to claim 49, wherein, The first syntax element is encoded and decoded using arithmetic encoding and decoding.
51. The method according to claim 50, wherein, The first syntax element is encoded or decoded using at least one context.
52. The method according to claim 51, wherein, The at least one context depends on the encoding / decoding information of the current block or neighboring blocks.
53. The method according to claim 51, wherein, The at least one context depends on the filtering shape of at least one neighboring block.
54. The method according to claim 50, wherein, The first syntax element is encoded and decoded using bypass encoding / decoding.
55. The method according to claim 49, wherein, The first syntax element is binarized using unary code, rounded unary code, fixed-length code, exponential Golomb code, or rounded exponential Golomb code.
56. The method of claim 50, wherein, The first syntax element is conditionally included in the bitstream.
57. The method according to claim 56, wherein, The first syntax element is included in the bitstream only when the mapping value of the prediction mode is available.
58. The method according to claim 50, wherein, The first syntax element is encoded and decoded in a predictive manner.
59. The method according to claim 58, wherein, The first syntax element is predicted based on the on / off decision of the mapping value of the prediction mode of at least one neighboring block.
60. The method of claim 50, wherein, For different color components, the first syntax element is included independently in the bitstream.
61. The method according to claim 60, wherein, The first syntax element is shared for different color components.
62. The method according to claim 60, wherein, For the first color component but not for the second color component, the first syntax element is included in the bitstream.
63. The method according to claim 50, wherein, The first syntax element is included in a sequence parameter set (SPS), picture parameter set (PPS), picture header, strip header, adaptive parameter set (APS), codec tree unit (CTU), or codec unit (CU).
64. The method according to claim 49, wherein, The first syntax element is included in the bitstream to indicate how the mapping value or boundary strength of the prediction mode is used.
65. The method according to claim 64, wherein, The first syntax element is included in the bitstream to indicate the number of categories to be merged for predicting pattern information or boundary strength information before being combined with existing classifications.
66. The method according to claim 65, wherein, The number of merged categories is N, where N is a positive integer.
67. The method according to claim 65, wherein, The number of merged categories is M, where M is a positive integer.
68. The method according to claim 64, wherein, The first syntax element is included in the bitstream to indicate whether prediction mode information or boundary strength information is used in combination.
69. The method according to claim 64, wherein, The first syntax element is arithmetic encoded and decoded.
70. The method according to claim 69, wherein, The first syntax element is encoded or decoded using at least one context.
71. The method according to claim 70, wherein, The at least one context depends on the encoding / decoding information of the current block or neighboring blocks.
72. The method according to claim 70, wherein, The at least one context depends on the filtering shape of at least one neighboring block.
73. The method according to claim 69, wherein, The first syntax element is encoded and decoded using bypass encoding / decoding.
74. The method according to claim 64, wherein, The first syntax element is binarized using unary code, rounded unary code, fixed-length code, exponential Golomb code, or rounded exponential Golomb code.
75. The method according to claim 64, wherein, The first syntax element is conditionally included in the bitstream.
76. The method according to claim 75, wherein, The first syntax element is included in the bitstream only when the mapping value of the prediction mode is available.
77. The method of claim 64, wherein, The first syntax element is encoded and decoded in a predictive manner.
78. The method according to claim 77, wherein, The first syntax element is predicted based on the on / off decision of the mapping value of the prediction mode of at least one neighboring block.
79. The method according to claim 64, wherein, For different color components, the first syntax element is independently included in the bitstream.
80. The method according to claim 79, wherein, The first syntax element is shared for different color components.
81. The method according to claim 79, wherein, For the first color component but not for the second color component, the first syntax element is included in the bitstream.
82. The method according to claim 64, wherein, The first syntax element is included in a sequence parameter set (SPS), picture parameter set (PPS), picture header, strip header, adaptive parameter set (APS), codec tree unit (CTU), or codec unit (CU).
83. The method according to claim 82, wherein, The first syntax element is included in the bitstream at the SPS or PPS level.
84. The method according to claim 82, wherein, The first syntax element is included in the bitstream at the CTU level.
85. The method according to claim 82, wherein, The first syntax element is included in the bitstream at the APS level.
86. The method according to claim 85, wherein, For all luminance and chrominance filters within all candidate filter sets, the first syntax element is included in the bitstream.
87. The method according to claim 85, wherein, For all luminance and chrominance filters within a set of candidate filters, the first syntax element is included in the bitstream.
88. The method according to any one of claims 1-87, wherein, The method is used in post-processing or pre-processing.
89. The method according to any one of claims 1-88, wherein, The methods are used in combination.
90. The method according to any one of claims 1-88, wherein, The method is used alone.
91. The method according to any one of claims 1-90, wherein, The mapping value of the prediction mode is used for an adaptive loop filter (ALF) that is applied to any loop filtering tool, preprocessing or postprocessing filtering method in video encoding and decoding, including ALF, cross-component ALF (CCALF), bilateral filter (BF), SAO, CCSAO or any other filtering method.
92. The method according to claim 91, wherein, The method of using the mapping value of the prediction mode is applied to the loop filter, or the use of the mapping value of the prediction mode is applied to ALF, CCALF, any other loop filter, preprocessing filter, or postprocessing filter.
93. The method according to any one of claims 1-92, wherein, The loop filter is applied to a video unit, wherein the video unit is a sequence, picture, sub-picture, strip, slice, CTU, CTU line, CTU group, CU, prediction unit (PU), transform unit (TU), codec tree block (CTB), codec block (CB), prediction block (PB), transform block (TB), or any other region including more than one luminance or chrominance sample or pixel.
94. The method according to any one of claims 1-93, wherein, The method is used in the bit stream via signal transmission.
95. The method according to any one of claims 1-94, wherein, The method is used at the sequence level, picture group level, picture level, strip level, or slice group level via signal transmission, including in the sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), decoding parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), strip header, or slice group header, or wherein the method is used at the prediction block (PB), transform block (TB), codec block (CB), prediction unit (PU), transform unit (TU), codec unit (CU), virtual pipeline decoding unit (VPDU), codec tree unit (CTU), CTU row, strip, slice, sub-picture, or other area including more than one sample or pixel via signal transmission.
96. The method according to any one of claims 1-95, wherein, The application of the method depends on the encoded and decoded information, which includes block size, color format, single-tree segmentation, double-tree segmentation, color components, stripe type, or image type.
97. The method according to any one of claims 1-96, wherein, The first syntax element is binarized into a flag, a fixed-length code, an exponential Golomb code, a unary code, a rounded unary code, or a rounded binary code, wherein the first syntax element is signed or unsigned.
98. The method according to any one of claims 1-97, wherein, The first syntax element is encoded or decoded using a context model or by bypass encoding / decoding.
99. The method according to any one of claims 1-98, wherein, The first syntax element is conditionally transmitted via a signal, either when the corresponding function applies, or when the height or width of the block meets a condition, or when both conditions are met.
100. The method according to any one of claims 1-99, wherein, Syntax elements are transmitted via signals at the sequence level, picture group level, picture level, stripe level, or slice group level, in the sequence header, picture header, SPS, VPS, DPS, decoding capability information (DCI), PPS, APS, stripe header, or slice group header.
101. The method according to any one of claims 1-100, wherein, The method may be used in combination with or without affine, multiple transform selection (MTS), LFNST, Merge with motion vector difference (MMVD), matrix-based intra-prediction (MIP), ISP, cross-component linear model (CCLM), convolutional cross-component model (CCCM), symmetric motion vector difference (SMVD), bidirectional optical flow (BDOF), decoder-side motion vector refinement (DMVR), history-based motion vector prediction (HMVP), template matching, intra-block copying (IBC), or a palette.
102. The method according to any one of claims 1-101, wherein, The excluded tool is implicitly disabled without signaling, or the method is implicitly disabled without signaling when the excluded tool is used.
103. The method according to any one of claims 1-102, wherein, The conversion includes encoding the visual media data into the bitstream.
104. The method according to any one of claims 1-103, wherein, The conversion includes decoding the visual media data from the bitstream.
105. An apparatus for processing video data, comprising: processor; and a non-transitory memory thereon having instructions, wherein, when executed by the processor, the instructions cause the processor to perform the method according to any one of claims 1-104.
106. A non-transitory computer-readable medium comprising a computer program product for use by a video codec apparatus, the computer program product comprising computer-executable instructions stored on the non-transitory computer-readable medium such that, when executed by a processor, the video codec apparatus performs the method according to any one of claims 1-104.
107. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by a video processing apparatus, wherein, The method includes: Determine to use the prediction mode as side information for the loop filter; and The bit stream is generated based on the loop filter.
108. A method for storing a video bitstream, comprising: The prediction mode is determined to be used as the side information for the loop filter; Based on the determination, a bit stream is generated; as well as The bit stream is stored in a non-transitory computer-readable recording medium.
109. A method, apparatus or system described in this disclosure.