Use of side information for adaptive loop filters in video coding
By taking the reconstruction sample points, predicted sample points, residual values, segmentation information and QP information before DBF/SAO/BF as inputs of ALF, the problem that ALF in the prior art fails to make full use of edge information is solved, and the encoding efficiency and quality of video encoding and decoding are improved.
Patent Information
- Application Number
- CN202380077217.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-11-01
- Filing Date
- 2023-11-01
- Publication Date
- 2025-06-27
AI Technical Summary
In the existing video encoding and decoding technology, the adaptive loop filter (ALF) only uses the reconstruction sample points before the current stage and fails to fully utilize other valuable edge information, such as the sample points before the deblocking filter (DBF), the sample point adaptive compensation filter (SAO) and the bilateral filter (BF), the predicted sample points, residual information, segmentation information and quantization parameter (QP) information, etc., resulting in a lack of encoding and decoding efficiency and quality.
Edge information, such as reconstruction sample points, predicted sample points, residual values, segmentation information and QP information before DBF/SAO/BF, is used as input to ALF to train the filter and enhance the adaptability and accuracy of the filter.
By utilizing more edge information, the encoding and decoding efficiency and quality of video encoding and decoding are improved, especially in high-resolution video processing, better artifact repair and coding efficiency are achieved.
Smart Images

Figure CN120226362A_ABST
Abstract
Description
[0001] Cross - reference to related applications
[0002] This application claims the priority and benefits of International Patent Application No. PCT / CN2022 / 128884, filed on November 1, 2022. The content of the foregoing patent application is incorporated herein by reference in its entirety. Technical field
[0003] This disclosure relates to the generation, storage, and consumption of digital audio - visual media information in a file format. Background art
[0004] Digital video occupies the largest bandwidth used on the Internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth demand for digital video is likely to continue to grow. Summary of the invention
[0005] The first aspect relates to a method for processing video data, including: using side information as an input to an adaptive loop filter (ALF); and performing a conversion between visual media data and a bitstream based on the ALF.
[0006] Optionally, in any of the foregoing aspects, another implementation of this aspect stipulates that the side information includes reconstructed samples obtained before a de - blocking filter (DBF), a sample - adaptive offset (SAO) filter, or a bilateral filter (BF).
[0007] Optionally, in any of the foregoing aspects, another implementation of this aspect stipulates that the reconstructed samples include luminance reconstructed samples, and the luminance reconstructed samples are used as an input source for at least one extended tap of a luminance and chrominance online training filter.
[0008] Optionally, in any of the foregoing aspects, another implementation of this aspect stipulates that the reconstructed samples include luminance reconstructed samples, and the luminance reconstructed samples are used as an input source for at least one spatial domain tap of a luminance and chrominance online training filter.
[0009] Optionally, in any of the foregoing aspects, another implementation of this aspect stipulates that the reconstructed samples include luminance reconstructed samples, and the luminance reconstructed samples are used as an input source for a luminance offline training filter.
[0010] Optionally, in any of the foregoing aspects, another implementation of this aspect stipulates that the reconstructed samples include luminance reconstructed samples, and the luminance reconstructed samples are used as an input source for classification.
[0011] Optionally, in any of the foregoing aspects, another implementation of this aspect stipulates that the classification includes classification of a luminance online training filter.
[0012] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the classification includes classification of a luminance offline training filter.
[0013] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the reconstructed samples include chrominance reconstructed samples, and the chrominance reconstructed samples are used as an input source for at least one extended tap of luminance and chrominance online training filters.
[0014] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the reconstructed samples include chrominance reconstructed samples, and the chrominance reconstructed samples are used as an input source for at least one spatial tap of luminance and chrominance online training filters.
[0015] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the reconstructed samples include chrominance reconstructed samples, and the chrominance reconstructed samples are used as an input source for classification.
[0016] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the classification includes classification of a luminance online training filter.
[0017] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the classification includes classification of a luminance offline training filter.
[0018] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the side information includes predicted samples.
[0019] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the predicted samples include luminance predicted samples, and the luminance predicted samples are used as an input source for at least one extended tap of luminance and chrominance online training filters.
[0020] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the predicted samples include luminance predicted samples, and the luminance predicted samples are used as an input source for at least one spatial tap of luminance and chrominance online training filters.
[0021] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the predicted samples include luminance predicted samples, and the luminance predicted samples are used as an input source for a luminance offline training filter.
[0022] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the predicted samples include luminance predicted samples, and the luminance predicted samples are used as an input source for classification.
[0023] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the classification includes classification of a luminance online training filter.
[0024] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the classification includes the classification of a luminance offline training filter.
[0025] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the reconstructed samples include chrominance reconstructed samples, and the chrominance reconstructed samples are used as an input source for at least one extended tap of the luminance and chrominance online training filters.
[0026] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the reconstructed samples include chrominance reconstructed samples, and the chrominance reconstructed samples are used as an input source for at least one spatial domain tap of the luminance and chrominance online training filters.
[0027] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the reconstructed samples include chrominance reconstructed samples, and the chrominance reconstructed samples are used as an input source for the classification.
[0028] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the classification includes the classification of a luminance online training filter.
[0029] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the classification includes the classification of a luminance offline training filter.
[0030] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the predicted samples are modified before being used as an input to the ALF.
[0031] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the modification includes filtering.
[0032] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the modification includes downsampling or upsampling.
[0033] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the modification includes clipping or shifting.
[0034] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the side information includes a residual value.
[0035] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the residual value includes a luminance residual value, and the luminance residual value is used as an input source for at least one extended tap of the luminance and chrominance online training filters.
[0036] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the residual value includes a luminance residual value, and the luminance residual value is used as an input source for at least one spatial tap of the luminance and chrominance online training filters.
[0037] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the residual value includes a luminance residual value, and the luminance residual value is used as an input source for the luminance offline training filter.
[0038] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the residual value includes a luminance residual value, and the luminance residual value is used as an input source for classification.
[0039] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the classification includes the classification of the luminance online training filter.
[0040] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the classification includes the classification of the luminance offline training filter.
[0041] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the residual value includes a chrominance residual value, and the chrominance residual value is used as an input source for at least one extended tap of the luminance and chrominance online training filters.
[0042] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the residual value includes a chrominance residual value, and the chrominance residual value is used as an input source for at least one spatial tap of the luminance and chrominance online training filters.
[0043] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the residual value includes a chrominance residual value, and the chrominance residual value is used as an input source for classification.
[0044] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the classification includes the classification of the luminance online training filter.
[0045] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the classification includes the classification of the luminance offline training filter.
[0046] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the residual value is modified before being used as an input to the ALF.
[0047] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the modification includes filtering.
[0048] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the modification includes downsampling or upsampling.
[0049] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the modification includes clipping or shifting.
[0050] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the side information includes segmentation information.
[0051] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the segmentation information includes one or more of block size, block shape, and block position.
[0052] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the segmentation information includes luminance segmentation information for classification.
[0053] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the classification includes classification of a luminance online training filter.
[0054] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the classification includes classification of a luminance offline training filter.
[0055] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the segmentation information includes chrominance segmentation information for classification.
[0056] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the classification includes classification of a luminance online training filter.
[0057] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the classification includes classification of a luminance offline training filter.
[0058] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the segmentation information includes luminance segmentation information for filter training.
[0059] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the filter training includes filter training of a luminance online training filter.
[0060] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the filter training includes filter training of a chrominance online training filter.
[0061] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the segmentation information includes chrominance segmentation information for filter training.
[0062] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the filter training includes filter training of a luminance online training filter.
[0063] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the filter training includes filter training of a chrominance online training filter.
[0064] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the segmentation information includes luminance segmentation information for filter training and chrominance segmentation information for the filter training.
[0065] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the side information includes quantization parameter (QP) information.
[0066] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the QP information includes one or more of picture QP information, slice QP information, and block-level QP information.
[0067] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the QP information includes luminance QP information for classification.
[0068] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the classification includes classification of a luminance online training filter.
[0069] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the classification includes classification of a luminance offline training filter.
[0070] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the QP information includes chrominance QP information for classification.
[0071] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the classification includes classification of a luminance online training filter.
[0072] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the classification includes classification of a luminance offline training filter.
[0073] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the QP information includes luminance QP information for filter training.
[0074] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the filter training includes filter training of a luminance online training filter.
[0075] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the filter training includes filter training of a chrominance online training filter.
[0076] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the QP information includes chrominance QP information for filter training.
[0077] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the filter training includes filter training of a luminance online training filter.
[0078] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the filter training includes filter training of a chrominance online training filter.
[0079] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the QP information includes luminance QP information for filter training and chrominance QP information for the filter training.
[0080] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the side information includes boundary strength information.
[0081] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the boundary strength information is generated by a deblocking filter.
[0082] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the boundary strength information includes luminance boundary strength information for classification.
[0083] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the classification includes classification of a luminance online training filter.
[0084] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the classification includes classification of a luminance offline training filter.
[0085] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the boundary strength information includes chrominance boundary strength for classification.
[0086] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the classification includes classification of a luminance online training filter.
[0087] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the classification includes classification of a luminance offline training filter.
[0088] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the boundary strength information includes luminance boundary strength information for filter training.
[0089] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the filter training includes filter training for an in - line luminance training filter.
[0090] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the filter training includes filter training for an in - line chrominance training filter.
[0091] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the boundary strength information includes chrominance boundary strength information for filter training.
[0092] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the filter training includes filter training for an in - line luminance training filter.
[0093] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the filter training includes filter training for an in - line chrominance training filter.
[0094] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the boundary strength information includes luminance boundary strength information for filter training and chrominance boundary strength information for the filter training.
[0095] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the ALF includes a cross - component ALF (CCALF).
[0096] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the side information includes reconstructed samples obtained before a de - blocking filter (DBF), a sample - adaptive offset (SAO) filter, or a bilateral filter (BF).
[0097] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the reconstructed samples include luminance reconstructed samples, and the luminance reconstructed samples are used as an input source for at least one extended tap of the CCALF.
[0098] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the reconstructed samples include luminance reconstructed samples, and the luminance reconstructed samples are used as an input source for at least one spatial domain tap of the CCALF.
[0099] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the reconstructed samples include chrominance reconstructed samples, and the chrominance reconstructed samples are used as an input source for at least one extended tap of the CCALF.
[0100] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the reconstructed samples include chrominance reconstructed samples, and the chrominance reconstructed samples are used as an input source for at least one spatial domain tap of the CCALF.
[0101] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the side information includes predicted samples.
[0102] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the predicted samples include luminance predicted samples, and the luminance predicted samples are used as an input source for at least one extended tap of the CCALF.
[0103] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the predicted samples include luminance predicted samples, and the luminance predicted samples are used as an input source for at least one spatial domain tap of the CCALF.
[0104] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the predicted samples include chrominance predicted samples, and the chrominance predicted samples are used as an input source for at least one extended tap of the CCALF.
[0105] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the predicted samples include chrominance predicted samples, and the chrominance predicted samples are used as an input source for at least one spatial domain tap of the CCALF.
[0106] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the predicted samples are modified before being used as an input to the CCALF.
[0107] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the modification includes filtering.
[0108] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the modification includes downsampling or upsampling.
[0109] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the modification includes clipping or shifting.
[0110] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the side information includes residual values.
[0111] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the residual value includes a luminance residual value, and the luminance residual value is used as an input source for at least one extended tap of the CCALF.
[0112] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the residual value includes a luminance residual value, and the luminance residual value is used as an input source for at least one spatial tap of the CCALF.
[0113] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the residual value includes a chrominance residual value, and the chrominance residual value is used as an input source for at least one extended tap of the CCALF.
[0114] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the residual value includes a chrominance residual value, and the chrominance residual value is used as an input source for at least one spatial tap of the CCALF.
[0115] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the residual value is modified before being used as an input to the CCALF.
[0116] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the modification includes filtering.
[0117] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the modification includes downsampling or upsampling.
[0118] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the modification includes clipping or shifting.
[0119] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the side information includes segmentation information.
[0120] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the segmentation information includes one or more of block size, block shape, and block position.
[0121] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the segmentation information includes luminance segmentation information for an online training filter for training the CCALF.
[0122] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the segmentation information includes chrominance segmentation information for an online training filter for training the CCALF.
[0123] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the segmentation information includes luminance segmentation information for filter training of the CCALF and chrominance segmentation information for filter training of the CCALF.
[0124] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the side information includes quantization parameter (QP) information.
[0125] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the QP information includes one or more of block size, block shape, and block position.
[0126] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the QP information includes luminance QP information for an online training filter for training the CCALF.
[0127] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the QP information includes chrominance QP information for an online training filter for training the CCALF.
[0128] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the QP information includes luminance QP information for filter training of the CCALF and chrominance QP information for filter training of the CCALF.
[0129] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the side information includes boundary strength information.
[0130] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the boundary strength information is generated by a deblocking filter.
[0131] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the boundary strength information includes luminance boundary strength information for an online training filter for training the CCALF.
[0132] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the boundary strength information includes chrominance boundary strength information for an online training filter for training the CCALF.
[0133] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the boundary strength information includes luminance boundary strength information for filter training of the CCALF and chrominance boundary strength information for filter training of the CCALF.
[0134] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the method is implemented in post - processing or pre - processing.
[0135] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that one or more of the methods are used in combination.
[0136] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that one or more of the methods are used individually.
[0137] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the method is applied to any loop - filtering tool, pre - processing filtering method, or post - processing filtering method in video coding and decoding, including but not limited to ALF / CCALF or any other filtering method.
[0138] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the method is applied to a loop - filtering method.
[0139] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the method is applied to ALF, CCAPF, sample - adaptive offset (SAO) filter, or another loop - filtering method.
[0140] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the method is applied to cross - component SAO (CCSAO).
[0141] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the method is applied to a bilateral filter (BF).
[0142] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the method is applied to a Hadamard - transform - domain filter (HTDF).
[0143] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the method is applied to a loop - filtering method.
[0144] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the method is applied to a pre - processing filtering method.
[0145] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the method is applied to a post - processing filtering method.
[0146] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the video units processed by the ALF or the CCALF include one of a sequence, a picture, a sub-picture, a slice, a tile, a coding tree unit (CTU), a CTU row, a group of CTUs, a coding unit (CU), a prediction unit (PU), a transform unit (TU), a coding tree block (CTB), a coding block (CB), a prediction block (PB), a transform block (TB), and any other region containing more than one luma or chroma sample or pixel.
[0147] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the indication of whether and / or how to apply the method is included in the bitstream.
[0148] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the indication is included in the bitstream at sequence level, group of pictures level, picture level, slice level, tile group level, sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), dependent parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), slice header or tile group header.
[0149] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the indication is included in the bitstream at prediction block (PB), transform block (TB), coding block (CB), prediction unit (PU), transform unit (TU), coding unit (CU), virtual pipeline data unit (VPDU), coding tree unit (CTU), CTU row, slice, tile, sub-picture or any other region containing more than one sample or pixel.
[0150] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that whether and / or how to apply any one of the methods depends on coding information, where the coding information includes block size, color format, single-tree or dual-tree segmentation, color component, slice type or picture type.
[0151] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the conversion includes encoding the visual media data into the bitstream.
[0152] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the conversion includes decoding the visual media data from the bitstream.
[0153] A second aspect relates to an apparatus for processing media data, including: one or more processors; and a non-transitory memory having instructions thereon, where the instructions, when executed by the processor, cause the apparatus to perform the method according to any one of the disclosed embodiments.
[0154] A third aspect relates to a non - transitory computer - readable medium, including a computer program product for use by a video codec device. The computer program product includes computer - executable instructions stored on the non - transitory computer - readable medium, such that when the processor executes the computer - executable instructions, the video codec device performs the method described in any one of the disclosed embodiments.
[0155] A fourth aspect relates to a non - transitory computer - readable recording medium that stores a bitstream of a video generated by a method executed by a video processing device, where the method includes the method described in any one of the disclosed embodiments.
[0156] A fifth aspect relates to a method for storing a bitstream of a video, including: the method described in any one of the disclosed embodiments.
[0157] A sixth aspect relates to a method, apparatus, or system described in the present disclosure.
[0158] For clarity, any one of the foregoing embodiments may be combined with one or more of the other foregoing embodiments to produce new embodiments within the scope of the present disclosure.
[0159] These and other features can be more clearly understood from the following detailed description in conjunction with the accompanying drawings and claims. Description of the Drawings
[0160] To more fully understand the present disclosure, reference is now made to the following brief description in conjunction with the accompanying drawings and specific embodiments, in which like reference numerals represent like components.
[0161] Figure 1 An example of the nominal vertical and horizontal positions of luma and chroma samples in a 4:2:2 format in a picture is shown.
[0162] Figure 2 An example encoder block diagram is shown.
[0163] Figure 3 An example picture divided into raster - scan stripes is shown.
[0164] Figure 4 An example picture divided into rectangular scan stripes is shown.
[0165] Figure 5 An example picture divided into bricks is shown.
[0166] Figures 6A to 6C An example of a coding - tree block (CTB) across a picture boundary is shown.
[0167] Figure 7 Shows an example of an intra prediction mode.
[0168] Figure 8 Shows an example of a block boundary in a picture.
[0169] Figure 9 Shows an example of pixels involved in filter use.
[0170] Figure 10 Shows an example of the filter shape of an adaptive loop filter (ALF).
[0171] Figure 11 Shows an example of transform coefficients supported by a 5×5 diamond filter.
[0172] Figure 12 Shows an example of relative coordinates supported by a 5×5 diamond filter.
[0173] Figure 13 Is a block diagram showing an example video processing system.
[0174] Figure 14 Is a block diagram of an example video processing device.
[0175] Figure 15 Is a flowchart of an example method of video processing.
[0176] Figure 16 Is a block diagram showing an example video encoding / decoding system.
[0177] Figure 17 Is a block diagram showing an example encoder.
[0178] Figure 18 Is a block diagram showing an example decoder.
[0179] Figure 19 Is a schematic diagram of an example encoder.
[0180] Figure 20 Is a method for processing video data according to an embodiment of the present disclosure. Detailed implementation
[0181] First, it should be understood that although illustrative implementations of one or more embodiments are provided below, any number of techniques may be used to implement the disclosed systems and / or methods, whether currently known or to be developed. The present disclosure should not in any way be limited to the exemplary implementations, drawings, and techniques shown below, including the exemplary designs and implementations shown and described herein, but may be modified within the full scope of the appended claims and their equivalents.
[0182] The chapter titles used in this disclosure are for ease of understanding and do not limit the applicability of the technologies and embodiments disclosed in each chapter to only that chapter. Additionally, the technologies described herein are applicable to other video codec protocols and designs.
[0183] 1. Initial Discussion
[0184] This disclosure relates to video coding and decoding technologies. Specifically, this document relates to loop filters and other coding and decoding tools in image / video coding. These ideas can be applied, either alone or in various combinations, to video codecs such as High Efficiency Video Coding (HEVC), Versatile Video Coding (VVC), or other video coding and decoding technologies.
[0185] 2. Abbreviations
[0186] This disclosure includes the following abbreviations. Advanced Video Coding (ITU-T Recommendation H.264 | ISO / IEC 14496-10) (AVC), Coding Picture Buffer (CPB), Clean Random Access (CRA), Coding Tree Unit (CTU), Coding Video Sequence (CVS), Decoded Picture Buffer (DPB), Decoding Parameter Set (DPS), General Constraint Information (GCI), High Efficiency Video Coding (also known as ITU-T Recommendation H.265 | ISO / IEC 23008-2) (HEVC), Joint Exploration Model (JEM), Motion Constraint Tile Set (MCTS), Network Abstraction Layer (NAL), Output Layer Set (OLS), Picture Header (PH), Picture Parameter Set (PPS), Profile, Tier and Level (PTL), Picture Unit (PU), Reference Picture Resampling (RPR), Raw Byte Sequence Payload (RBSP), Supplemental Enhancement Information (SEI), Slice Header (SH), Sequence Parameter Set (SPS), Video Coding Layer (VCL), Video Parameter Set (VPS), Versatile Video Coding (also known as ITU-T Recommendation H.266 | ISO / IEC 23090-3) (VVC), VVC Test Model (VTM), Video Usability Information (VUI), Transform Unit (TU), Coding Unit (CU), Deblocking Filter (DF), Sample Adaptive Offset (SAO), Adaptive Loop Filter (ALF), Coding Block Flag (CBF), Quantization Parameter (QP), Rate-Distortion Optimization (RDO), and Bilateral Filter (BF).
[0187] 3. Video Coding and Decoding Standards
[0188] Video coding standards have evolved mainly through the development of ITU-T and ISO / IEC standards. ITU-T developed H.261 and H.263, ISO / IEC developed the Moving Picture Experts Group (MPEG-1) and MPEG-4 Visual, and the two organizations jointly developed the H.262 / MPEG-2 Video and H.264 / MPEG-4 Advanced Video Coding (AVC) and H.265 / HEVC [1] standards. Since H.262, video coding standards have been based on a hybrid video coding structure that utilizes temporal prediction plus transform coding. To explore future video coding technologies beyond HEVC, the Video Coding Experts Group (VCEG) and MPEG jointly established the Joint Video Exploration Team (JVET). JVET adopted many methods and placed them in a reference software called the Joint Exploration Model (JEM) [2]. When the Versatile Video Coding (VVC) project was officially launched, the Joint Video Exploration Team (JVET) was renamed the Joint Video Experts Team (JVET). VVC is a coding standard that aims to reduce the bitrate by 50% compared to HEVC. The VVC working draft and the VVC Test Model (VTM) are continuously updated.
[0189] An example version of the VVC draft (i.e., Versatile Video Coding (Draft 10)) can be found at: https: / / jvet-experts.org / doc_end_user / documents / 19_Teleconference / wg11 / JVET-S2001-v17.zip. An example version of the reference software for VVC (called VTM) can be found at: https: / / vcgit.hhi.fraunhofer.de / jvet-u-ee2 / VVCSoftware_VTM / - / tree / VTM-11.2.
[0190] The Video Coding Experts Group (VCEG) of the International Telecommunication Union Telecommunication Standardization Sector (ITU-T) and the Joint Technical Committee (JTC) 1 / Subcommittee (SC) 29 / Working Group (WG) 11 of the International Organization for Standardization and the International Electrotechnical Commission (ISO / IEC) Moving Picture Experts Group (MPEG) are studying the potential need to standardize future video coding technologies that will significantly exceed the current VVC standard in terms of compression capabilities. Such future standardization actions could take the form of an extension of VVC or a completely new standard. These groups are jointly carrying out this development activity in a joint collaborative effort called JVET to evaluate the compression technology designs proposed by their experts in this area. The first Exploration Experiment (EE) was established by the Joint Video Exploration Team (JVET), and a reference software called the Enhanced Compression Model (ECM) is in use. The test model ECM is continuously updated.
[0191] 3.1 Color Space and Chroma Subsampling
[0192] A color space (also known as a color model (or color system)) is a mathematical model that describes a range of colors as digital tuples, e.g., as 3 or 4 values or color components (e.g., RGB). Generally, a color space is a concretization of a coordinate system and a subspace. For video compression, the most commonly used color spaces are luminance, blue-difference chrominance, and red-difference chrominance (YCbCr) and red, green, and blue (RGB).
[0193] YCbCr, Y'CbCr, or YPb / CbPr / Cr (also written as YCBCR or Y'CBCR) is a collective term for a series of color spaces that are used in video systems and digital photography as part of the color image processing pipeline. Y' is the luminance component, and CB and CR are the blue-difference chrominance component and the red-difference chrominance component. Y' (with an apostrophe) is different from Y, which means that the light intensity is non-linearly encoded based on gamma-corrected RGB primaries.
[0194] Chroma subsampling is the practice of encoding an image by achieving a lower resolution for chrominance information than for luminance information, taking advantage of the fact that the human visual system is less sensitive to chromatic differences than to luminance differences. 3.1.1. 4:4:4
[0196] In 4:4:4, each of the three components of Y'CbCr has the same sampling rate. Thus, there is no chroma subsampling. This scheme is sometimes used in high-end film scanners and film post-production. 3.1.2. 4:2:2
[0198] In 4:2:2, the two chrominance components are sampled at half the sampling rate of luminance. The horizontal chrominance resolution is halved while the vertical chrominance resolution remains the same. This reduces the bandwidth of the uncompressed video signal by one-third with little visual difference. Figure 1 Examples of nominal vertical and horizontal positions in the 4:2:2 color format are shown. 3.1.3 4:2:0
[0200] In 4:2:0, compared with 4:1:1, the horizontal sampling is doubled. However, since the Cb and Cr channels are sampled only on alternate rows in this scheme, the vertical resolution is halved, and thus the data rate remains unchanged. Cb and Cr are each downsampled by a factor of 2 in both the horizontal and vertical directions. There are three variants of the 4:2:0 scheme, each with different horizontal and vertical sampling positions. In MPEG-2, Cb and Cr are horizontally co-located. Cb and Cr are located between pixels vertically (i.e., interleaved). In Joint Photographic Experts Group (JPEG) / JPEG File Interchange Format (JFIF), H.261, and MPEG-1, Cb and Cr are arranged in an interleaved manner, located at the middle position between alternate luminance samples. In 4:2:0 DV, Cb and Cr are horizontally co-located. Vertically, they are co-located on alternate rows.
[0201] chroma_format_idc separate_colour_plane_flag chrominance format SubWidthC SubHeightC 0 0 monochrome 1 1 1 0 4:2:0 2 2 2 0 4:2:2 2 1 3 0 4:4:4 1 1 3 1 4:4:4 1 1
[0202] Table 1. SubWidthC and SubHeightC values derived from chroma_format_idc and Separate_colour_plane_flag
[0203] 3.2 Example encoding / decoding processes of video codecs
[0204] Figure 2 An example of the encoder block diagram of VVC is shown, which includes three loop filter modules: Deblocking Filter (DF), Sample Adaptive Offset (SAO), and Adaptive Loop Filter (ALF). Different from the DF that uses predefined filters, SAO and ALF utilize the original samples of the current picture to reduce the mean square error between the original samples and the reconstructed samples by adding offset values respectively and by applying a Finite Impulse Response (FIR) filter. The encoding / decoding side information is used to signal this offset value and filter coefficients. ALF is located at the last processing stage of each picture and can be regarded as a tool to try to capture and fix the artifacts generated by the previous stages.
[0205] 3.3 Definition of video / encoding / decoding units
[0206] A picture is divided into one or more slice rows and one or more slice columns. A slice is a sequence of Coding Tree Units (CTUs) that cover a rectangular area of the picture. A slice can be divided into one or more tiles, each tile including a certain number of CTU rows within the slice. A slice that is not divided into multiple tiles can also be called a tile. However, a tile that is a proper subset of a slice cannot be called a slice. A strip contains several slices of a picture or several tiles of a slice.
[0207] Two stripe modes are supported, namely, the raster scan stripe mode and the rectangular stripe mode. In the raster scan stripe mode, a stripe contains a series of slices in the slice raster scan of a picture. In the rectangular stripe mode, a stripe contains a certain number of tiles of a picture, and these tiles together form a rectangular area of the picture. The tiles within a rectangular stripe are arranged in the order of the tile raster scan of the stripe. Figure 3 An example of the raster scan stripe segmentation of a picture is shown, where the picture is divided into 12 slices and 3 raster scan stripes.
[0208] Figure 4 An example of the rectangular stripe segmentation of a picture is shown, where the picture is divided into 24 slices (6 slice columns and 4 slice rows) and 9 rectangular stripes.
[0209] Figure 5 An example of a picture segmented into slices, tiles, and rectangular stripes is shown, where the picture is divided into 4 slices (2 slice columns and 2 slice rows), 11 tiles (the upper left slice contains 1 tile, the upper right slice contains 5 tiles, the lower left slice contains 2 tiles, and the lower right slice contains 3 tiles), and 4 rectangular stripes.
[0210] 3.3.1 CTU / CTB Sizes
[0211] In VVC, the CTU size (signaled in the sequence parameter set (SPS) by the syntax element log2_ctu_size_minus2) can be as small as 4×4.
[0212] 7.3.2.3 Sequence Parameter Set RBSP Syntax
[0213]
[0214]
[0215]
[0216] log2_ctu_size_minus2 plus 2 specifies the luma coding tree block size of each CTU. log2_min_luma_coding_block_size_minus2 plus 2 specifies the minimum luma coding block size. The variables CtbLog2SizeY, CtbSizeY, MinCbLog2SizeY, MinCbSizeY, MinTbLog2SizeY, MaxTbLog2SizeY, MinTbSizeY, MaxTbSizeY, PicWidthInCtbsY, PicHeightInCtbsY, PicSizeInCtbsY, PicWidthInMinCbsY, PicHeightInMinCbsY, PicSizeInMinCbsY, PicSizeInSamplesY, PicWidthInSamplesC, and PicHeightInSamplesC are derived as follows:
[0217] CtbLog2SizeY = log2_ctu_size_minus2 + 2 (7-9)
[0218] CtbSizeY = 1 << CtbLog2SizeY (7-10)
[0219] MinCbLog2SizeY = log2_min_luma_coding_block_size_minus2 + 2 (7-11)
[0220] MinCbSizeY = 1 << MinCbLog2SizeY (7-12)
[0221] MinTbLog2SizeY = 2 (7-13)
[0222] MaxTbLog2SizeY = 6 (7-14)
[0223] MinTbSizeY = 1 << MinTbLog2SizeY (7-15)
[0224] MaxTbSizeY = 1 << MaxTbLog2SizeY (7-16)
[0225] PicWidthInCtbsY = Ceil( pic_width_in_luma_samples ÷ CtbSizeY ) (7-17)
[0226] PicHeightInCtbsY = Ceil( pic_height_in_luma_samples ÷ CtbSizeY )(7-18)
[0227] PicSizeInCtbsY = PicWidthInCtbsY * PicHeightInCtbsY (7-19)
[0228] PicWidthInMinCbsY = pic_width_in_luma_samples / MinCbSizeY (7-20)
[0229] PicHeightInMinCbsY = pic_height_in_luma_samples / MinCbSizeY (7-21)
[0230] PicSizeInMinCbsY = PicWidthInMinCbsY * PicHeightInMinCbsY (7-22)
[0231] PicSizeInSamplesY=pic_width_in_luma_samples * pic_height_in_luma_samples (7-23)
[0232] PicWidthInSamplesC = pic_width_in_luma_samples / SubWidthC (7-24)
[0233] PicHeightInSamplesC = pic_height_in_luma_samples / SubHeightC (7-25)
[0234] 3.3.2 CTUs in a Picture
[0235] Assume that the CTB / largest coding unit (LCU) size is represented by M×N (usually M equals N), and for a CTB located at the picture boundary (or slice or strip or other type of boundary, taking the picture boundary as an example), K×L samples are within the picture boundary, where K < M or L < N. Figures 6A to 6C Examples of CTBs crossing the picture boundary are shown. For Figures 6A to 6C those CTBs shown in, the CTB size is still equal to M×N. However, the lower / right boundary of the CTB is outside the picture.
[0236] 3.4 Intra Prediction
[0237] To capture any edge direction presented in natural videos, the number of intra prediction modes in the direction frame is extended from 33 used in HEVC to 65. Figure 7 An example of 67 intra prediction modes is shown. Figure 7 The extended direction modes are shown in [reference], and the planar and direct current (DC) modes remain unchanged. These denser intra prediction modes in the direction frame are applicable to all block sizes as well as for intra prediction of luminance and chrominance.
[0238] The angular intra prediction directions can be defined as from 45 degrees to -135 degrees in the clockwise direction, as Figure 7 shown. In VTM, for non-square blocks, several angular intra prediction modes are adaptively replaced by wide-angle intra prediction modes. The replaced modes are signaled and remapped to the indices of the wide-angle modes after parsing. The total number of intra prediction modes remains unchanged, e.g., 67, and the intra mode coding and decoding remain the same.
[0239] In HEVC, each intra-coded block has a square shape and the length of each side of the block is a power of 2. Thus, no division operation is required to generate the intra predictor using the DC mode. In VVC, the blocks can have a rectangular shape and in general a division operation must be used for each block. To avoid the division operation for DC prediction, only the longer side is used to calculate the average value of non-square blocks.
[0240] 3.5 Inter Prediction
[0241] For each inter-predicted CU, the motion parameters include the motion vector, the reference picture index, the reference picture list use index, and the extended information of the new coding features of VVC for sample generation for inter prediction. The motion parameters can be signaled in an explicit or implicit way. When a CU is coded using the skip mode, the CU is associated with a PU and there are no significant residual coefficients, no coded motion vector differences, and / or reference picture indices. The Merge mode is defined as obtaining the motion parameters of the current CU from neighboring CUs, including spatial and temporal candidates and the extended schedule introduced in VVC. The Merge mode can be applied to any inter-predicted CU, not just for the skip mode. An alternative to the Merge mode is to explicitly send the motion parameters, where for each CU, the motion vector, the corresponding reference picture index for each reference picture list, the reference picture list use flag, and other useful information are explicitly signaled.
[0242] 3.6 Deblocking Filter
[0243] Deblocking filtering is an example loop filter in video codecs. In VVC, the deblocking filtering process is applied to CU boundaries, transform sub-block boundaries, and prediction sub-block boundaries. Prediction sub-block boundaries include prediction unit boundaries introduced by sub-block based temporal motion vector prediction (SbTMVP) and affine modes. Transform sub-block boundaries include transform unit boundaries introduced by sub-block transform (SBT) and intra-subdivision (ISP) modes, as well as transforms due to implicit partitioning of large CUs. The processing order of the deblocking filter is defined as horizontally filtering the vertical edges of the entire picture first, and then vertically filtering the horizontal edges. This specific order enables multiple horizontal filtering or vertical filtering processes to be applied in parallel threads. The filtering process can also be implemented on a per-CTB basis with a very small processing delay.
[0244] First, filter the vertical edges in the picture. Then, use the samples modified by the vertical edge filtering process as input to filter the horizontal edges in the picture. The vertical and horizontal edges in each CTB of a CTU are processed separately according to the coding and decoding units. Filter the vertical edges of the coding and decoding blocks in the coding and decoding units in the geometric order of the coding and decoding blocks, starting from the edge on the left side of the coding and decoding block and proceeding to the edge on the right side of the coding and decoding block. Filter the horizontal edges of the coding and decoding blocks in the coding and decoding units in the geometric order of the coding and decoding blocks, starting from the edge at the top of the coding and decoding block and proceeding to the edge at the bottom of the coding and decoding block. Figure 8 Shows picture samples on an 8×8 grid, as well as horizontal block boundaries and vertical block boundaries, and non-overlapping blocks of 8×8 samples that can be deblocked in parallel. 3.6.1 Boundary Decision
[0245] Filtering is applied to 8×8 block boundaries. Additionally, such a boundary must be a transform block boundary or a coding and decoding sub-block boundary, for example, a transform block boundary or a coding and decoding sub-block boundary caused by using affine motion prediction (ATMVP). For other boundaries, the deblocking filter is disabled.
[0246] 3.6.2 Boundary Strength Calculation
[0247] For transform block boundaries / coding and decoding sub-block boundaries, if the boundary is within the 8×8 grid, the boundary can be filtered, and the setting of bS[xDi][yDj] for that edge (where [xDi][yDj] represents coordinates) is defined as in Tables 2 and 3 respectively.
[0248]
[0249]
[0250] Table 2. Boundary Strength (when SPSIBC is disabled)
[0251]
[0252] Table 3. Boundary Strength (when SPSIBC is enabled)
[0253] 3.6.3 Deblocking Decision for Luma Component
[0254] Figure 9 The pixels involved in the filter on / off decision and strong / weak filter selection are shown. A wider and stronger luma filter is used only when Conditions 1, 2, and 3 are all true. Condition 1 is the "large block condition". This condition detects whether the samples on the P side and Q side belong to large blocks, represented by the variables bSidePisLargeBlk and bSideQisLargeBlk respectively. bSidePisLargeBlk and bSideQisLargeBlk are defined as follows.
[0255] bSidePisLargeBlk = ((the edge type is vertical and p0 belongs to a CU with width >= 32) || (the edge type is horizontal and p0 belongs to a CU with height >= 32))? true : false
[0256] bSideQisLargeBlk = ((the edge type is vertical and q0 belongs to a CU with width >= 32) || (the edge type is horizontal and q0 belongs to a CU with height >= 32))? true : false
[0257] Based on bSidePisLargeBlk and bSideQisLargeBlk, Condition 1 is defined as follows:[[]]
[0258] Condition 1 = (bSidePisLargeBlk || bSidePisLargeBlk)? true : false
[0259] Next, if Condition 1 is true, then Condition 2 is further checked. First, the following variables are derived:[[]]
[0260] First, dp0, dp3, dq0, and dq3 are derived in the HEVC manner
[0261] if (the p side is greater than or equal to 32)[[]]
[0262] dp0 = (dp0 + Abs(p50 - 2 * p40 + p30) + 1) >> 1
[0263] dp3 = (dp3 + Abs(p53 - 2 * p43 + p33) + 1) >> 1
[0264] if (the q side is greater than or equal to 32)[[]]
[0265] dq0 = (dq0 + Abs(q50 - 2 * q40 + q30) + 1) >> 1
[0266] dq3 = (dq3 + Abs(q53 - 2 * q43 + q33) + 1) >> 1
[0267] Condition 2 = (d < β)? true : false
[0268] where d = dp0 + dq0 + dp3 + dq3.
[0269] If both Condition 1 and Condition 2 are satisfied, further check whether any block uses a sub - block:
[0270]
[0271] Finally, if both Condition 1 and Condition 2 are satisfied, the de - block method will check Condition 3 (the strong filtering condition for large blocks), which is defined as follows. In Condition 3 StrongFilterCondition, derive the following variables:
[0272] Derive dpq in the HEVC way.
[0273] Derive sp3 = Abs(p3 - p0) in the HEVC way
[0274]
[0275]
[0276] Derive sq3 = Abs(q0 - q3) in the HEVC way
[0277]
[0278] According to HEVC, StrongFilterCondition = (dpq is less than (β >> 2), sp3 + sq3 is less than (3 * β >> 5), and Abs(p0 - q0) is less than (5 * tC + 1) >> 1)? true : false.
[0279] 3.6.4 Stronger Deblocking Filter for Luminance
[0280] When the samples on either side of the boundary belong to a large block, a bilinear filter is used. When the width of the vertical edge >= 32 and when the height of the horizontal edge >= 32, the samples are defined as belonging to a large block. The bilinear filter is listed as follows. Block boundary samples pi (i = 0 to Sp - 1) and qi (i = 0 to Sq - 1), in the above HEVC deblocking, pi and qi are the i-th samples in the row for filtering the vertical edge, or the i-th samples in the column for filtering the horizontal edge, and then are replaced by linear interpolation as follows:
[0281] p i ′ = (fi i * Middle s,t + (64 - fi i ) * P s + 32) >> 6), clipped to p i ± tcPD i
[0282] q j ′ = (g j * Middle s,t + (64 - g j ) * Q s + 32) >> 6), clipped to q j ± tcPD j
[0283] where the tcPD i and tcPD j terms are the position - dependent clipping described above, and g j 、f i 、Middle s,t 、P s and Q s are given as follows.
[0284] 3.6.5 Chrominance Deblocking Decision
[0285] A chrominance strong filter is used on both sides of the block boundary. Here, when both sides of the chrominance edge are greater than or equal to 8 (chrominance position), the chrominance filter is selected, and the decision satisfies the following three conditions: The first decision is the boundary strength and large block decision. When the block width or height orthogonal to the block edge is equal to or greater than 8 in the chrominance sample domain, this filter can be applied. The second decision and the third decision are basically the same as the HEVC luma deblocking decisions, which are the on / off decision and the strong filter decision respectively.
[0286] In the first decision, for chrominance filtering, the boundary strength (bS) is modified and conditions are checked sequentially. If the condition is satisfied, the remaining conditions with lower priority are skipped. Chrominance deblocking is performed when bS equals 2, or bS equals 1 when a large block boundary is detected. The second and third conditions are basically the same as the HEVC luma strong filter decision, as follows.
[0287] Under the second condition, d is derived according to HEVC luma deblocking. The second condition will be true when d is less than β. Under the third condition, StrongFilterCondition is derived as follows:
[0288] dpq is derived in the HEVC manner.
[0289] sp3 = Abs(p3 - p0) is derived in the HEVC manner
[0290] sq3 = Abs(q0 - q3) is derived in the HEVC manner
[0291] According to the HEVC design, StrongFilterCondition = (dpq is less than (β >> 2), sp3 + sq3 is less than (β >> 3), and Abs(p0 - q0) is less than (5 * tC + 1) >> 1).
[0292] 3.6.6 Strong Deblocking Filter for Chrominance
[0293] The strong deblocking filter for chrominance is defined as follows:
[0294] p2′ = (3 * p3 + 2 * p2 + p1 + p0 + q0 + 4) >> 3
[0295] p1′ = (2 * p3 + p2 + 2 * p1 + p0 + q0 + q1 + 4) >> 3
[0296] p0′ = (p3 + p2 + p1 + 2 * p0 + q0 + q1 + q2 + 4) >> 3
[0297] An example chrominance filter performs deblocking on a 4×4 chrominance sample grid.
[0298] 3.6.7 Position-Dependent Clipping
[0299] Position-dependent clipping tcPD is applied to the output samples of the luma filtering process, involving strong and long filters that modify 7, 5, and 3 samples at the boundary. Assuming the quantization error distribution, the clipping value of the samples can be increased, which are expected to have higher quantization noise and thus a greater deviation between the expected reconstructed sample value and the true sample value.
[0300] For each P or Q boundary filtered with an asymmetric filter, a position-dependent threshold table is selected from two tables (e.g., Tc7 and Tc3 listed below) according to the result of the decision-making process, and these two tables are provided to the decoder as side information:
[0301] Tc7 = {6, 5, 4, 3, 2, 1, 1}; Tc3 = {6, 4, 2};
[0302] tcPD = (Sp == 3)? Tc3 : Tc7;
[0303] tcQD = (Sq == 3)? Tc3 : Tc7;
[0304] For a P or Q boundary filtered with a short symmetric filter, a lower-amplitude position-dependent threshold is applied:
[0305] Tc3 = {3, 2, 1};
[0306] After defining the thresholds, the filtered p'i and q'i sample values are clamped according to the tcP and tcQ clamping values:
[0307] p”i = Clip3(p’i + tcPi, p’i – tcPi, p’i);
[0308] q”j = Clip3(q’j + tcQj, q’j – tcQ j, q’j);
[0309] where p'i and q'j are the filtered sample values, p”i and q”j are the clamped output sample values, and tcPi is the clamping threshold derived from the VVC tc parameters and tcPD and tcQD. The function Clip3 is the clamping function as specified in VVC.
[0310] 3.6.8. Sub-block Deblocking Adjustment
[0311] To perform parallel-friendly deblocking using both a long filter and sub-block deblocking, the long filter is restricted to modifying at most 5 samples on the side where sub-block deblocking (AFFINE or ATMVP or decoder-side motion vector refinement (DMVR)) is used, as shown in the luminance control of the long filter. Extending this, the sub-block deblocking is adjusted such that the sub-block boundaries on the 8×8 grid near the CU or implicit TU boundary are restricted to modifying at most 2 samples on each side.
[0312] The following applies to sub-block boundaries that are not aligned with the CU boundary.
[0313]
[0314]
[0315] Where the edge equal to 0 corresponds to the CU boundary, the edge equal to 2 or equal to orthogonalLength - 2 corresponds to the sub-block boundary 8 samples away from the CU boundary, etc. If the implicit partitioning of the TU is used, the implicit TU is true.
[0316] 3.7. Sample Adaptive Offset
[0317] Sample Adaptive Offset (SAO) is applied to the reconstructed signal after the deblocking filter by using the offset values specified by the encoder for each CTB. The video encoder first decides whether to apply SAO processing to the current slice. If SAO is applied to the slice, each CTB will be classified into one of five SAO types as shown in Table 4. The concept of SAO is to classify pixels into multiple categories and reduce distortion by adding offset values to the pixels in each category. The SAO operations include: Edge Offset (EO), which classifies pixels using edge attributes in SAO types 1 to 4; and Band Offset (BO), which classifies pixels using pixel intensity in SAO type 5. Each applicable CTB has SAO parameters, including sao_merge_left_flag, sao_merge_up_flag, SAO type, and four offset values. If sao_merge_left_flag is equal to 1, the current CTB will reuse the SAO type and offset values of the left CTB. If sao_merge_up_flag is equal to 1, the current CTB will reuse the SAO type and offset values of the upper CTB.
[0318] SAO type Sample Adaptive Offset type to be used number of classes 0 none 0 1 1-D 0-degree mode edge compensation 4 2 1-D 90-degree mode edge compensation 4 3 1-D 135-degree mode edge compensation 4 4 1-D 45-degree mode edge compensation 4 5 band compensation 4
[0319] Table 4. Specification of SAO Types
[0320] 3.8. Adaptive Loop Filter
[0321] Adaptive loop filtering for video coding and decoding minimizes the mean square error between the original samples and the decoded samples by using a Wiener-based adaptive filter. The ALF is located at the last processing stage of each picture and can be regarded as a tool for capturing and fixing artifacts from the previous stage. The appropriate filter coefficients are determined by the encoder and signaled explicitly to the decoder. To achieve better coding and decoding efficiency, especially for high-resolution videos, local adaptivity is applied to the luminance signal by applying different filters to different regions or blocks in the picture. In addition to filter adaptivity, filter on / off control at the coding tree unit (CTU) level also helps to improve coding and decoding efficiency. Syntactically, the filter coefficients are sent in the picture-level header information called the adaptive parameter set, and the filter on / off flags of the CTUs are interleaved at the CTU level in the slice data. This syntax design not only supports picture-level optimization but also achieves a lower coding delay.
[0322] 3.8.1. Signaling of Parameters
[0323] According to the ALF design in VTM, the filter coefficients and the clipping index are carried in the ALF APS. The ALF APS can include up to 8 chroma filters and one luminance filter set, for a total of up to 25 filters. Each of the 25 luminance classes also includes an index. Classes with the same index share the same filter. By combining different classes, the number of bits required to represent the filter coefficients is reduced. The absolute value of the filter coefficients is represented using a 0th order Exp-Golomb code, followed by the sign bit of the non-zero coefficients. When clipping is enabled, a two-bit fixed-length code is also used to signal the clipping index for each filter coefficient. The decoder can use up to 8 ALF APSs simultaneously.
[0324] The filter control syntax elements of the ALF in VTM include two types of information. First, the ALF on / off flags are signaled at the sequence, picture, slice, and CTB levels. The chroma ALF can be enabled at the picture and slice levels only if the luminance ALF is enabled at the corresponding level. Second, if the ALF is enabled at the picture, slice, and CTB levels, the filter usage information is signaled at that level. If all slices within a picture use the same APS, the referenced ALF APS ID is coded at the slice level or the picture level. The luminance component can reference up to 7 ALF APSs, while the chroma component can reference up to 1 ALF APS. For the luminance CTB, an index is signaled to indicate which ALF APS or the offline-trained luminance filter set is used. For the chroma CTB, the index indicates which filter in the referenced APS is used.
[0325] The ALF data syntax elements associated with the luminance (LUMA) component in VTM are listed below:
[0326]
[0327]
[0328] When alf_luma_filter_signal_flag equals 1, it specifies that the luminance filter set is signaled. When alf_luma_filter_signal_flag equals 0, it specifies that the luminance filter set is not signaled. When alf_luma_clip_flag equals 0, it specifies that linear adaptive loop filtering is applied to the luminance component. When alf_luma_clip_flag equals 1, it specifies that non-linear adaptive loop filtering can be applied to the luminance component. alf_luma_num_filters_signalled_minus1 plus 1 specifies the number of classes of adaptive loop filters for luminance coefficients that can be signaled. The value of alf_luma_num_filters_signalled_minus1 shall be in the range of 0 to NumAlfFilters - 1 (including the endpoints). alf_luma_coeff_delta_idx[filtIdx] indicates the index of the signaled adaptive loop filter luminance coefficient increment for the filter class indicated by filtIdx, where filtIdx ranges from 0 to NumAlfFilters - 1. When alf_luma_coeff_delta_idx[filtIdx] does not exist, it is inferred to be equal to 0. The length of alf_luma_coeff_delta_idx[filtIdx] is Ceil(Log2(alf_luma_num_filters_signalled_minus1 + 1)) bits. The value of alf_luma_coeff_delta_idx[filtIdx] shall be in the range of 0 to alf_luma_num_filters_signalled_minus1 (including the endpoints).
[0329] alf_luma_coeff_abs[sfIdx][j] specifies the absolute value of the j-th coefficient of the luminance filter transmitted by the signal indicated by sfIdx. When alf_luma_coeff_abs[sfIdx][j] does not exist, it is inferred to be equal to 0. The value of alf_luma_coeff_abs[sfIdx][j] shall be in the range from 0 to 128 (including the endpoints). alf_luma_coeff_sign[sfIdx][j] specifies the sign of the j-th luminance coefficient of the filter indicated by sfIdx, as follows:
[0330] If alf_luma_coeff_sign[sfIdx][j] is equal to 0, the corresponding luminance filter coefficient has a positive value.
[0331] Otherwise (alf_luma_coeff_sign[sfIdx][j] is equal to 1), the corresponding luminance filter coefficient has a negative value.
[0332] When alf_luma_coeff_sign[sfIdx][j] does not exist, it is inferred to be equal to 0.
[0333] alf_luma_clip_idx[sfIdx][j] specifies the clipping index of the clipping value to be used before multiplying the j-th coefficient of the luminance filter transmitted by the signal indicated by sfIdx. When alf_luma_clip_idx[sfIdx][j] does not exist, it is inferred to be equal to 0. The codec tree syntax elements associated with the luminance component in VTM are listed as follows:
[0334]
[0335]
[0336] alf_ctb_flag[cIdx][xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] being equal to 1 specifies that the adaptive loop filter is applied to the codec tree block of the color component indicated by cIdx of the codec tree unit at the luminance position (xCtb, yCtb). alf_ctb_flag[cIdx][xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] being equal to 0 specifies that the adaptive loop filter is not applied to the codec tree block of the color component indicated by cIdx of the codec tree unit at the luminance position (xCtb, yCtb).
[0337] When alf_ctb_flag[cIdx][xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] does not exist, it is inferred to be equal to 0. alf_use_aps_flag being equal to 0 specifies that one of the fixed filter sets is applied to the luma CTB. alf_use_aps_flag being equal to 1 specifies that the filter set from APS is applied to the luma CTB. When alf_use_aps_flag does not exist, it is inferred to be equal to 0. alf_luma_prev_filter_idx specifies the previous filter applied to the luma CTB. The value of alf_luma_prev_filter_idx shall be in the range of 0 to sh_num_alf_aps_ids_luma - 1 (including the endpoints). When alf_luma_prev_filter_idx does not exist, it is inferred to be equal to 0.
[0338] The variable AlfCtbFiltSetIdxY[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] specifies the filter set index for the luma CTB at position (xCtb, yCtb) and is derived as follows:
[0339] If alf_use_aps_flag is equal to 0, then AlfCtbFiltSetIdxY[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] is set to be equal to alf_luma_fixed_filter_idx.
[0340] Otherwise, AlfCtbFiltSetIdxY[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] is set to be equal to 16 + alf_luma_prev_filter_idx.
[0341] alf_luma_fixed_filter_idx specifies the fixed filter applied to the luma CTB. The value of alf_luma_fixed_filter_idx shall be in the range of 0 to 15 (including the endpoints).
[0342] Based on the VTM-based ALF design, the ECM's ALF design further introduces the concept of alternative filter sets into the luma filters. Based on the updated luma CTU ALF on / off decision for each alternative / round, multiple alternatives / rounds of training are performed on the luma filters. In this way, multiple filter sets will be associated with each training alternative, and the class combination results of each filter set may be different. Each CTU can select the best filter set determined by RDO, and the alternative information related to signaling will be transmitted. The data syntax elements of ALF associated with the luma component in ECM are listed as follows:
[0343]
[0344] alf_luma_num_alts_minus1 plus 1 specifies the number of alternative filter sets for the luma component. The value of alf_luma_num_alts_minus1 should be in the range of 0 to 3 (including the endpoints). alf_luma_clip_flag[altIdx] being equal to 0 specifies that linear adaptive loop filtering is applied to the alternative luma filter set for the luma component with index altIdx. alf_luma_clip_flag[altIdx] being equal to 1 specifies that non-linear adaptive loop filtering may be applied to the alternative luma filter set for the luma component with index altIdx. alf_luma_num_filters_signalled_minus1[altIdx] plus 1 specifies the number of adaptive loop filter classes for which the luma coefficients can be signaled to the alternative luma filter set with index altIdx. The value of alf_luma_num_filters_signalled_minus1[altIdx] should be in the range of 0 to NumAlfFilters - 1 (including the endpoints).
[0345] alf_luma_coeff_delta_idx[altIdx][filtIdx] specifies the index of the adaptive loop filter luma coefficient increment for signal transmission for the filter class represented by filtIdx for the alternative luma filter set with index altIdx, where filtIdx ranges from 0 to NumAlfFilters–1. When alf_luma_coeff_delta_idx[filtIdx][altIdx] does not exist, it is inferred to be equal to 0. The length of alf_luma_coeff_delta_idx[altIdx][filtIdx] is Ceil(Log2(alf_luma_num_filters_signalled_minus1[altIdx]+1)) bits. The value of alf_luma_coeff_delta_idx[altIdx][filtIdx] shall be in the range of 0 to alf_luma_num_filters_signalled_minus1[altIdx] (including the endpoints). alf_luma_coeff_abs[altIdx][sfIdx][j] specifies the absolute value of the j-th coefficient of the luma filter signalled by sfIdx of the alternative luma filter set with index altIdx. When alf_luma_coeff_abs[altIdx][sfIdx][j] does not exist, it is inferred to be equal to 0. The value of alf_luma_coeff_abs[altIdx][sfIdx][j] shall be in the range of 0 to 128 (including the endpoints).
[0346] alf_luma_coeff_sign[altIdx][sfIdx][j] specifies the sign of the j-th luma coefficient of the filter signalled by sfIdx of the alternative luma filter set with index altIdx, as follows:
[0347] If alf_luma_coeff_sign[altIdx][sfIdx][j] is equal to 0, the corresponding luma filter coefficient has a positive value.
[0348] Otherwise (alf_luma_coeff_sign[altIdx][sfIdx][j] is equal to 1), the corresponding luma filter coefficient has a negative value.
[0349] When alf_luma_coeff_sign[altIdx][sfIdx][j] does not exist, it is inferred to be equal to 0.
[0350] alf_luma_clip_idx[altIdx][sfIdx][j] specifies the clipping index of the clipping value to be used before multiplying the j-th coefficient of the luminance filter transmitted through the signal represented by sfIdx of the alternative luminance filter set with index altIdx. When alf_luma_clip_idx[altIdx][sfIdx][j] does not exist, it is inferred to be equal to 0. The codec tree syntax elements associated with the luminance component in the ECM are listed as follows:
[0351]
[0352]
[0353] alf_ctb_luma_filter_alt_idx[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] specifies the index of the alternative luminance filter for the codec tree block of the luminance component of the codec tree unit at the luminance position (xCtb, yCtb). When alf_ctb_luma_filter_alt_idx[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] does not exist, it is inferred to be equal to zero.
[0354] 3.8.2. Filter Shape
[0355] Figure 10 Examples of the filter shapes of the ALF are shown. In JEM, up to three diamond filter shapes can be selected for the luminance component (as Figure 10 shown). The filter shape used for the luminance component is indicated by transmitting an index at the picture level. Each square represents a sample point, and Ci (i is 0 to 6 (left), 0 to 12 (middle), 0 to 20 (right)) represents the coefficient to be applied to that sample point. For the chrominance component in the picture, a 5×5 diamond shape is always used. In VVC, a 7×7 diamond shape is always used for luminance, and a 5×5 diamond shape is always used for chrominance.
[0356] 3.8.3 Classification of ALF
[0357] Each 2×2 (or 4×4) block is classified as one of 25 classes. The classification index C is derived based on its directionality D and the quantized value of the activity as follows:
[0358]
[0359] To calculate D and First, the gradients in the horizontal, vertical, and two diagonal directions are calculated using a 1D Laplacian operator:
[0360]
[0361] The indices i and j refer to the coordinates of the upper-left sample point in the 2×2 block, and R(i, j) indicates the reconstructed sample point at the coordinates (i, j). The maximum and minimum values of the gradients in the horizontal and vertical directions are set as:
[0362]
[0363] And the maximum and minimum values of the gradients in the two diagonal directions are set as:
[0364]
[0365] To derive the value of the directionality D, these values are compared with each other and with two thresholds t1 and t2:
[0366] Step 1. If and are both true, then set D to 0.
[0367] Step 2. If then proceed to Step 3; otherwise, proceed to Step 4;
[0368] Step 3. If then set D to 2; otherwise, set D to 1.
[0369] Step 4. If then set D to 4; otherwise, set D to 3.
[0370] The activity value A is calculated as:
[0371]
[0372] A is further quantized to the range from 0 to 4 (including the endpoints), and the quantized value is denoted as For the two chrominance components in the picture, no classification method is adopted (i.e., a set of ALF coefficients is applied to each chrominance component).
[0373] 3.8.4. Geometric Transformation of Filter Coefficients
[0374] Figure 11An example of the relative coordinates supported by a 5×5 rhombic filter is shown. Before filtering each 2×2 block, geometric transformations such as rotation or diagonal and vertical flipping are applied to the filter coefficients f(k, l) associated with the coordinates (k, l) according to the gradient value calculated for that block. This is equivalent to applying these transformations to the samples in the filter support region. The idea is to make the blocks more similar by aligning the directions of the different blocks to which the ALF is applied.
[0375] Three geometric transformations are introduced, including diagonal transformation, vertical flipping, and rotation:
[0376] Diagonal: f D (k, l) = f(l, k),
[0377] Vertical flip: f V (k, l) = f(k, K - l - 1),
[0378] Rotation: f R (k, l) = f(K - l - 1, k).
[0379] Where K is the size of the filter, and 0 ≤ k, l ≤ K - 1 are the coefficient coordinates, such that the position (0, 0) is at the upper left corner and the position (K - 1, K - 1) is at the lower right corner. The transformation is applied to the filter coefficient f(k, l) according to the gradient value calculated for that block. Table 5 summarizes the relationship between the transformation and the four gradients in four directions. Figure 11 The transformation coefficients for each position based on a 5×5 rhombus are shown.
[0380]
[0381]
[0382] Table 5. Mapping of gradients and transformations calculated for a block.
[0383] 3.8.5. Filtering process
[0384] On the decoder side, when ALF is enabled for a block, each sample R(i, j) within the block is filtered to obtain the sample value R′(i, j) as shown below, where L represents the filter length, f m,n represents the filter coefficient, and f(k, l) represents the decoded filter coefficient.
[0385]
[0386] Figure 12 An example of the relative coordinates supported by a 5×5 rhombic filter is shown, assuming the coordinates (i, j) of the current sample are (0, 0). Samples in different coordinates filled with the same color are multiplied by the same filter coefficient.
[0387] 3.8.6. Nonlinear Filtering Reformulation
[0388] Linear filtering can be reformulated into the following expressions without affecting the encoding and decoding efficiency:
[0389]
[0390] where w(i,j) are the same filter coefficients.
[0391] VVC introduces non - linearity to make ALF more efficient by using a simple clipping function to reduce the impact of these neighboring sample values when the neighboring sample values (I(x + i, y + j)) differ too much from the currently filtered sample value (I(x, y)). More specifically, the ALF filter is modified as follows:
[0392]
[0393] where K(d,b)=min(b,max( - b,d)) is the clipping function, and k(i,j) are the clipping parameters, which depend on the (i,j) filter coefficients. The encoder performs optimization to find the best k(i,j).
[0394] For each ALF filter, a clipping parameter k(i,j) is specified, and each filter coefficient signals a clipping value. This means that at most 12 clipping values can be signaled in the bitstream for each luminance filter, and at most 6 clipping values can be signaled in the bitstream for each chrominance filter. To limit the signaling cost and encoder complexity, only 4 fixed values are used, and these fixed values are the same for inter - frame and intra - frame stripes.
[0395] Since the variance of local differences in luminance is generally higher than that in chrominance, two different sets are applied for the luminance filter and the chrominance filter. The maximum sample value in each set (here 1024 for a 10 - bit bit - depth) is also introduced so that clipping can be disabled when not needed. These 4 values are selected by roughly evenly dividing the entire range of sample values of luminance (for 10 - bit encoding and decoding) in the logarithmic domain and the range of 4 to 1024 for chrominance. More precisely, the luminance table of clipping values is obtained by the following formula:
[0396] where M = 2 10 and N = 4
[0397] Similarly, the chrominance table of clipping values can be obtained according to the following formula:
[0398] where M = 210 , N = 4 and A = 4
[0399] 3.9. Bilateral Loop Filter
[0400] 3.9.1. Bilateral Image Filter
[0401] The bilateral image filter is a non - linear filter that smooths noise while preserving edge structures. Bilateral filtering is a technique where the filter weights decrease not only with the distance between samples but also with the increase in intensity difference. In this way, over - smoothing of edges can be improved. The weights are defined as
[0402]
[0403] where Δx and Δy are the distances in the vertical and horizontal directions respectively, and ΔI is the intensity difference between samples.
[0404] The edge - preserving denoising bilateral filter uses low - pass Gaussian filters for both the domain filter and the range filter. The domain low - pass Gaussian filter gives higher weights to pixels spatially closer to the central pixel. The range low - pass Gaussian filter gives higher weights to pixels similar to the central pixel. Combining the range filter and the domain filter, the bilateral filter at the edge pixels becomes an elongated Gaussian filter that is oriented along the edge and significantly reduced in the gradient direction. This is why the bilateral filter can smooth noise while preserving edge structures.
[0405] 3.9.2. Bilateral Filter in Video Coding and Decoding
[0406] The bilateral filter in video coding and decoding is a coding - decoding tool for VVC[2]. The filter acts as a loop filter in parallel with the sample adaptive offset (SAO) filter. Both the bilateral filter and SAO act on the same input samples, and each filter produces an offset value, which is then added to the input samples to produce the output samples, and the output samples enter the next stage after clipping. The spatial filtering strength σ d is determined by the block size, where the smaller the block, the greater the filtering strength, and the intensity filtering strength σ r is determined by the quantization parameter, where stronger filtering is used for higher QPs. Only the four closest samples are used, so the intensity I of the filtered sample F can be calculated as
[0407]
[0408] where I C represents the intensity of the central sample, ΔI A = I A - I CRepresents the intensity difference between the central sample and the upper sample. ΔI B , ΔI L and ΔI R represent the intensity differences between the central sample and the lower, left, and right samples, respectively.
[0409] 4. Technical problems solved by the disclosed technical solutions
[0410] Example designs of the adaptive loop filter (ALF) in video coding and decoding have the following problems:
[0411] In the example ALF design, only the reconstruction before the current stage is used. However, there is other valuable side information that can potentially be exploited, such as samples before the deblocking filter (DBF), sample adaptive offset filter (SAO), and / or bilateral filter (BF), predicted samples, residual information, segmentation information, QP information, etc.
[0412] 5. List of solutions and embodiments
[0413] To solve the above problems, the methods summarized below are disclosed. The embodiments should be regarded as examples explaining the general concept and should not be interpreted narrowly. In addition, these embodiments can be applied individually or combined in any way.
[0414] It should be noted that the disclosed methods can be used as a loop filter or post - processing.
[0415] In the present disclosure, a video unit may refer to a sequence, picture, sub - picture, slice, CTU, block, and / or region. A video unit may include one color component or multiple color components.
[0416] In the present disclosure, a processing unit may refer to a sequence, picture, sub - picture, slice, CTU, block, region, or sample. A processing unit may include one color component or multiple color components.
[0417] In the present disclosure, side information may refer to any coding and decoding information not input into SAO / ALF / CCALF in VVC.
[0418] Example 1
[0419] In one example, side information can be used for ALF.
[0420] Example 2
[0421] In one example, the reconstruction before DBF / SAO / BF can be used as side information for ALF. In one example, the luminance reconstruction before DBF / SAO / BF can be used as an input source for at least one extended tap of the luminance / chrominance online training filter. In one example, the luminance reconstruction before DBF / SAO / BF can be used as an input source for at least one spatial domain tap of the luminance / chrominance online training filter. In one example, the luminance reconstruction before DBF / SAO / BF can be used as an input source for the luminance offline training filter.
[0422] In one example, the luminance reconstruction before DBF / SAO / BF can be used as an input source for classification. In one example, the luminance reconstruction before DBF / SAO / BF can be used as an input source for the classification of the luminance online training filter. In one example, the luminance reconstruction before DBF / SAO / BF can be used as an input source for the classification of the luminance offline training filter.
[0423] In one example, the chrominance reconstruction before DBF / SAO / BF can be used as an input source for at least one extended tap of the luminance / chrominance online training filter. In one example, the chrominance reconstruction before DBF / SAO / BF can be used as an input source for at least one spatial domain tap of the luminance / chrominance online training filter.
[0424] In one example, the chrominance reconstruction before DBF / SAO / BF can be used as an input source for classification. In one example, the chrominance reconstruction before DBF / SAO / BF can be used as an input source for the classification of the luminance online training filter. In one example, the chrominance reconstruction before DBF / SAO / BF can be used as an input source for the classification of the luminance offline training filter.
[0425] Example 3
[0426] In one example, the predicted sample can be used as side information for ALF.
[0427] In one example, the luminance predicted sample can be used as an input source for at least one extended tap of the luminance / chrominance online training filter. In one example, the luminance predicted sample can be used as an input source for at least one spatial domain tap of the luminance / chrominance online training filter. In one example, the luminance predicted sample can be used as an input source for the luminance offline training filter.
[0428] In one example, the luminance predicted sample can be used as an input source for classification. In one example, the luminance predicted sample can be used as an input source for the classification of the luminance online training filter. In one example, the luminance predicted sample can be used as an input source for the classification of the luminance offline training filter.
[0429] In one example, the chrominance prediction sample can be used as an input source for at least one extended tap of the luminance / chrominance online training filter. In one example, the chrominance prediction sample can be used as an input source for at least one spatial tap of the luminance / chrominance online training filter.
[0430] In one example, the chrominance prediction sample can be used as an input source for classification. In one example, the chrominance prediction sample can be used as an input source for classification of the luminance online training filter. In one example, the chrominance prediction sample can be used as an input source for classification of the luminance offline training filter.
[0431] The prediction sample can be modified before being put into the ALF. The modification can be filtering. The modification can be downsampling / upsampling. The modification can be clipping / shifting.
[0432] Example 4
[0433] Using the residual value as side information for the ALF is proposed.
[0434] In one example, the luminance residual value can be used as an input source for at least one extended tap of the luminance / chrominance online training filter. In one example, the luminance residual value can be used as an input source for at least one spatial tap of the luminance / chrominance online training filter. In one example, the luminance residual value can be used as an input source for the luminance offline training filter.
[0435] In one example, the luminance residual value can be used as an input source for classification. In one example, the luminance residual value can be used as an input source for classification of the luminance online training filter. In one example, the luminance residual value can be used as an input source for classification of the luminance offline training filter.
[0436] In one example, the chrominance residual value can be used as an input source for at least one extended tap of the luminance / chrominance online training filter. In one example, the chrominance residual value can be used as an input source for at least one spatial tap of the luminance / chrominance online training filter.
[0437] In one example, the chrominance residual value can be used as an input source for classification. In one example, the chrominance residual value can be used as an input source for classification of the luminance online training filter. In one example, the chrominance residual value can be used as an input source for classification of the luminance offline training filter.
[0438] The residual value can be modified before being put into the ALF. The modification can be filtering. The modification can be downsampling / upsampling. The modification can be clipping / shifting.
[0439] Example 5
[0440] In one example, the segmentation information can be used as side information for the ALF.
[0441] In one example, the segmentation information can represent block size / shape / position or other information.
[0442] In one example, the luminance segmentation information can be used for classification. In one example, the luminance segmentation information can be used for the classification of the luminance online training filter. In one example, the luminance segmentation information can be used for the classification of the luminance offline training filter.
[0443] In one example, the chrominance segmentation information can be used for classification. In one example, the chrominance segmentation information can be used for the classification of the luminance online training filter. In one example, the chrominance segmentation information can be used for the classification of the luminance offline training filter.
[0444] In one example, the luminance segmentation information can be used for filter training. In one example, the luminance segmentation information can be used for training the luminance online training filter. In one example, the luminance segmentation information can be used for training the chrominance online training filter.
[0445] In one example, the chrominance segmentation information can be used for filter training. In one example, the chrominance segmentation information can be used for training the luminance online training filter. In one example, the chrominance segmentation information can be used for training the chrominance online training filter.
[0446] In one example, the luminance / chrominance segmentation information can participate in the filtering process.
[0447] Example 6
[0448] In one example, the QP information can be used as side information for the ALF.
[0449] In one example, the QP information can represent picture / strip / block-level QP information.
[0450] In one example, the luminance QP information can be used for classification. In one example, the luminance QP information can be used for the classification of the luminance online training filter. In one example, the luminance QP information can be used for the classification of the luminance offline training filter.
[0451] In one example, the chrominance QP information can be used for classification. In one example, the chrominance QP information can be used for the classification of the luminance online training filter. In one example, the chrominance QP information can be used for the classification of the luminance offline training filter.
[0452] In one example, the luma QP information can be used for filter training. In one example, the luma QP information can be used to train the luma online training filter. In one example, the luma QP information can be used to train the chroma online training filter.
[0453] In one example, the chroma QP information can be used for filter training. In one example, the chroma QP information can be used to train the luma online training filter. In one example, the chroma QP information can be used to train the chroma online training filter.
[0454] In one example, the luma / chroma QP information can participate in the filtering process.
[0455] Example 7
[0456] In one example, the boundary strength information can be used as side information for ALF.
[0457] In one example, the boundary strength information can be generated by DBF or other methods.
[0458] In one example, the luma boundary strength information can be used for classification. In one example, the luma boundary strength information can be used for the classification of the luma online training filter. In one example, the luma boundary strength information can be used for the classification of the luma offline training filter.
[0459] In one example, the chroma boundary strength information can be used for classification. In one example, the chroma boundary strength information can be used for the classification of the luma online training filter. In one example, the chroma boundary strength information can be used for the classification of the luma offline training filter.
[0460] In one example, the luma boundary strength information can be used for filter training. In one example, the luma boundary strength information can be used to train the luma online training filter. In one example, the luma boundary strength information can be used to train the chroma online training filter.
[0461] In one example, the chroma boundary strength information can be used for filter training. In one example, the chroma boundary strength information can be used to train the luma online training filter. In one example, the chroma boundary strength information can be used to train the chroma online training filter.
[0462] In one example, the luma / chroma boundary strength information can participate in the filtering process of the luma / chroma online training filter of ALF.
[0463] Example 8
[0464] In one example, the side information of CCALF can be used.
[0465] Example 9
[0466] In one example, the reconstruction before DBF / SAO / BF can be used as side information for CCALF. In one example, the luminance reconstruction before DBF / SAO / BF can be used as an input source for at least one extended tap of CCALF. In one example, the luminance reconstruction before DBF / SAO / BF can be used as an input source for at least one spatial domain tap of CCALF. In one example, the chrominance reconstruction before DBF / SAO / BF can be used as an input source for at least one extended tap of CCALF. In one example, the chrominance reconstruction before DBF / SAO / BF can be used as an input source for at least one spatial domain tap of CCALF.
[0467] Example 10
[0468] In one example, the predicted sample points can be used as side information for CCALF.
[0469] In one example, the luminance predicted sample points can be used as an input source for at least one extended tap of CCALF. In one example, the luminance predicted sample points can be used as an input source for at least one spatial domain tap of CCALF. In one example, the chrominance predicted sample points can be used as an input source for at least one extended tap of CCALF. In one example, the chrominance predicted sample points can be used as an input source for at least one spatial domain tap of CCALF.
[0470] The predicted sample points can be modified before being fed into CCALF. The modification can be filtering. The modification can be downsampling / upsampling. The modification can be clipping / shifting.
[0471] Example 11
[0472] In one example, the residual values can be used as side information for CCALF.
[0473] In one example, the luminance residual values can be used as an input source for at least one extended tap of CCALF. In one example, the luminance residual values can be used as an input source for at least one spatial domain tap of CCALF. In one example, the chrominance residual values can be used as an input source for at least one extended tap of CCALF. In one example, the chrominance residual values can be used as an input source for at least one spatial domain tap of CCALF.
[0474] The residual values can be modified before being fed into CCALF. The modification can be filtering. The modification can be downsampling / upsampling. The modification can be clipping / shifting.
[0475] Example 12
[0476] In one example, the segmentation information can be used as side information for CCALF. In one example, the segmentation information can represent block size / shape / position or other information. In one example, the luminance segmentation information can be used to train the online training filter of CCALF. In one example, the chrominance segmentation information can be used to train the online training filter of CCALF. In one example, the luminance / chrominance segmentation information can participate in the filtering process of CCALF.
[0477] Example 13
[0478] In one example, the QP information can be used as side information for CCALF. In one example, the QP information can represent picture / strip / block-level QP information. In one example, the luminance QP information can be used to train the online training filter of CCALF. In one example, the chrominance QP information can be used to train the online training filter of CCALF. In one example, the luminance / chrominance QP information can participate in the filtering process of CCALF.
[0479] Example 14
[0480] In one example, the boundary strength information can be used as side information for CCALF. In one example, the boundary strength information can be generated by DBF or other methods. In one example, the luminance boundary strength information can be used to train the online training filter of CCALF. In one example, the chrominance boundary strength information can be used to train the online training filter of CCALF. In one example, the luminance / chrominance boundary strength information can participate in the filtering process of CCALF.
[0481] Example 15
[0482] In one example, the disclosed method can be used for post-processing and / or pre-processing.
[0483] Example 16
[0484] In one example, the methods mentioned above can be used jointly.
[0485] Example 17
[0486] In one example, the methods mentioned above can be used separately.
[0487] Example 18
[0488] In one example, the method for using side information described above can be applied to any loop filtering tool, pre-processing or post-processing filtering method in video coding and decoding (including but not limited to ALF / CCALF or any other filtering method).
[0489] Example 19
[0490] In one example, the proposed side information utilization method can be applied to loop filtering methods. In one example, the proposed side information utilization method can be applied to ALF. In one example, the proposed side information utilization method can be applied to CCALF. In one example, the proposed side information utilization method can be applied to SAO. In one example, the proposed side information utilization method can be applied to CCSAO. In one example, the proposed side information utilization method can be applied to bilateral filter (BF). In one example, the proposed side information utilization method can be applied to Hadamard transform domain filter (HTDF). Alternatively, the proposed side information utilization method can be applied to other loop filtering methods.
[0491] Example 20
[0492] In one example, the side information utilization method can be applied to pre-filtering methods. In one example, the side information utilization method can be applied to post-filtering methods.
[0493] Example 21
[0494] In the above examples, a video unit may refer to a sequence / picture / sub-picture / strip / slice / Codec Tree Unit (CTU) / CTU row / CTU group / Codec Unit (CU) / Prediction Unit (PU) / Transformation Unit (TU) / Codec Tree Block (CTB) / Codec Block (CB) / Prediction Block (PB) / Transformation Block (TB) / any other region containing more than one luma or chroma sample / pixel.
[0495] Example 22
[0496] Whether and / or how to apply the methods disclosed above can be signaled in the bitstream.
[0497] In one example, they can be signaled at the sequence level / group of pictures level / picture level / strip level / slice group level, such as in the sequence header, picture header, SPS, VPS, DPS, DCI, PPS, APS, strip header, and slice group header.
[0498] In one example, they can be signaled in PB, TB, CB, PU, TU, CU, VPDU, CTU, CTU row, strip, slice, sub-picture, and other types of regions containing more than one sample or pixel.
[0499] Example 23
[0500] Whether and / or how to apply the methods disclosed above can depend on codec information, such as block size, color format, single / double tree partitioning, color component, strip / picture type.
[0501] 6. References [1] J. Strom, P. Wennersten, J. Enhorn, D. Liu, K. Andersson and R. Sjoberg, “Bilateral Loop Filter in Combination with SAO,” in proceeding of IEEE Picture Coding Symposium (PCS), Nov. 2019.
[0502] Figure 13 FIG. 4000 is a block diagram illustrating an example video processing system 4000 in which various techniques disclosed herein may be implemented. Various embodiments may include some or all of the components of system 4000. System 4000 may include an input 4002 for receiving video content. The video content may be received in a raw or uncompressed format (e.g., 8-bit or 10-bit multi-component pixel values), or may be received in a compressed format or an encoded format. Input 4002 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, passive optical network (PON), etc., and wireless interfaces such as Wi-Fi or cellular interfaces.
[0503] System 4000 may include a codec component 4004 that may implement various codec or encoding methods described in the present disclosure. The codec component 4004 may reduce the average bit rate of the video from the input 4002 to the output of the codec component 4004 to produce a coded representation of the video. Thus, codec techniques are sometimes referred to as video compression or video transcoding techniques. The output of the codec component 4004 may be stored or transmitted via the connected communication as shown by component 4006. The stored or transmitted bitstream representation (or coded representation) of the video received at the input 4002 may be used by component 4008 to generate pixel values or a displayable video, which is sent to the display interface 4010. The process of generating a user-visible video from the bitstream representation is sometimes referred to as video decompression. Additionally, although certain video processing operations are referred to as “encoding” operations or tools, it should be understood that encoding tools or operations are used in an encoder, and the decoder will perform the corresponding decoding tools or operations that are the reverse of the encoding process.
[0504] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB), High-Definition Multimedia Interface (HDMI), DisplayPort, etc. Examples of storage interfaces include Serial Advanced Technology Attachment (SATA), Peripheral Component Interconnect (PCI), Integrated Drive Electronics (IDE) interfaces, etc. The techniques described in the present disclosure may be embodied in various electronic devices such as mobile phones, laptop computers, smartphones, or other devices capable of performing digital data processing and / or video display.
[0505] Figure 14 is a block diagram of an example video processing device 4100. The device 4100 can be used to implement one or more of the methods described herein. The device 4100 can be implemented as a smart phone, a tablet computer, a computer, an Internet of Things (IoT) receiver, etc. The device 4100 can include one or more processors 4102, one or more memories 4104, and video processing circuitry 4106. The processor 4102 can be configured to implement one or more of the methods described in this disclosure. The memory(ies) 4104 can be used to store data and code for implementing the methods and techniques described herein. The video processing circuitry 4106 can be used to implement some of the techniques described in this disclosure in hardware circuitry. In some embodiments, the video processing circuitry 4106 can be at least partially included in the processor 4102 (e.g., a graphics coprocessor).
[0506] Figure 15 is a flowchart of an example method 4200 for video processing. The method 4200 includes determining at step 4202 to use side information as an input to the ALF. Performing a conversion between visual media data and a bitstream based on the ALF at step 4204. According to an example, the conversion of step 4204 can include encoding at an encoder or decoding at a decoder.
[0507] It should be noted that the method 4200 can be implemented in a device for processing video data, the device including a processor and a non-transitory memory having instructions thereon, the device such as, a video encoder 4400, a video decoder 4500, and / or an encoder 4600. In such a case, the instructions, when executed by the processor, cause the processor to execute the method 4200. Additionally, the method 4200 can be performed by a non-transitory computer-readable medium including a computer program product for use by a video codec device. The computer program product includes computer-executable instructions stored on a non-transitory computer-readable medium such that when executed by a processor cause the video codec device to execute the method 4200.
[0508] Figure 16 is a block diagram showing an example video codec system 4300 that can utilize the techniques of this disclosure. The video codec system 4300 can include a source device 4310 and a destination device 4320. The source device 4310 generates encoded video data, and the source device can be referred to as a video encoding device. The destination device 4320 can decode the encoded video data generated by the source device 4310, and the destination device can be referred to as a video decoding device.
[0509] The source device 4310 may include a video source 4312, a video encoder 4314, and an input / output (I / O) interface 4316. The video source 4312 may include sources such as a video capture device, an interface for receiving video data from a video content provider, and / or a computer graphics system for generating video data, or a combination of such sources. The video data may include one or more pictures. The video encoder 4314 encodes the video data from the video source 4312 to generate a bitstream. The bitstream may include a series of bits that form an encoded representation of the video data. The bitstream may include encoded pictures and associated data. The encoded pictures are the encoded representations of the pictures. The associated data may include sequence parameter sets, picture parameter sets, and other syntax structures. The I / O interface 4316 may include a modulator / demodulator (modem) and / or a transmitter. The encoded video data may be directly sent to the destination device 4320 via the I / O interface 4316 over the network 4330. The encoded video data may also be stored on a storage medium / server 4340 for access by the destination device 4320.
[0510] The destination device 4320 may include an I / O interface 4326, a video decoder 4324, and a display device 4322. The I / O interface 4326 may include a receiver and / or a modem. The I / O interface 4326 may obtain the encoded video data from the source device 4310 or the storage medium / server 4340. The video decoder 4324 may decode the encoded video data. The display device 4322 may display the decoded video data to a user. The display device 4322 may be integrated with the destination device 4320 or may be external to the destination device 4320, which may be configured to be connected to an external display device through an interface.
[0511] The video encoder 4314 and the video decoder 4324 may operate according to a video compression standard, such as the High Efficiency Video Coding (HEVC) standard, the Versatile Video Coding (VVC) standard, and other current and / or future standards.
[0512] Figure 17 is a block diagram showing an example of a video encoder 4400, which may be the Figure 16 video encoder 4314 in the system 4300 shown. The video encoder 4400 may be configured to perform any or all of the techniques of the present disclosure. The video encoder 4400 includes a plurality of functional components. The techniques described in the present disclosure may be shared among the various components of the video encoder 4400. In some examples, a processor may be configured to perform any or all of the techniques described in the present disclosure.
[0513] The functional components of the video encoder 4400 may include: a splitting unit 4401; a prediction unit 4402, which may include a mode selection unit 4403, a motion estimation unit 4404, a motion compensation unit 4405, and an intra prediction unit 4406; a residual generation unit 4407; a transformation processing unit 4408; a quantization unit 4409; an inverse quantization unit 4410; an inverse transformation unit 4411; a reconstruction unit 4412; a cache 4413; and an entropy encoding unit 4414.
[0514] In other examples, the video encoder 4400 may include more, fewer, or different functional components. In one example, the prediction unit 4402 may include an intra block copy (IBC) unit. The IBC unit may perform prediction in the IBC mode, where at least one reference picture is the picture in which the current video block is located.
[0515] Furthermore, some components, such as the motion estimation unit 4404 and the motion compensation unit 4405, may be highly integrated, but are shown separately for explanatory purposes in the example of the video encoder 4400.
[0516] The splitting unit 4401 may split a picture into one or more video blocks. The video encoder 4400 and the video decoder 4500 may support various video block sizes.
[0517] The mode selection unit 4403 may, for example, select one of the encoding / decoding modes based on an error result, and provide the resulting intra- or inter-encoded / decoded block to the residual generation unit 4407 to generate residual block data and provide it to the reconstruction unit 4412 to reconstruct the encoded / decoded block for use as a reference picture. In some examples, the mode selection unit 4403 may select an intra and inter prediction combination (CIIP) mode, where the prediction is based on an inter prediction signal and an intra prediction signal. The mode selection unit 4403 may also select the resolution of the motion vector for a block in the case of inter prediction (e.g., sub-pixel or integer pixel accuracy).
[0518] To perform inter prediction on the current video block, the motion estimation unit 4404 may generate motion information for the current video block by comparing one or more reference frames from the buffer 4413 with the current video block. The motion compensation unit 4405 may determine a predicted video block for the current video block based on the motion information and decoded samples of a picture from the cache 4413 (rather than the picture associated with the current video block).
[0519] The motion estimation unit 4404 and the motion compensation unit 4405 may perform different operations on the current video block, for example, depending on whether the current video block is in an I-slice, a P-slice, or a B-slice.
[0520] In some examples, the motion estimation unit 4404 may perform uni - directional prediction on a current video block, and the motion estimation unit 4404 may search for a reference video block for the current video block in the reference pictures of list 0 or list 1. Then, the motion estimation unit 4404 may generate a reference index that indicates the reference picture in list 0 or list 1 that contains the reference video block and a motion vector that indicates the spatial displacement between the current video block and the reference video block. The motion estimation unit 4404 may output the reference index, a prediction direction indicator, and the motion vector as the motion information of the current video block. The motion compensation unit 4405 may generate a predicted video block of the current block based on the reference video block indicated by the motion information of the current video block.
[0521] In other examples, the motion estimation unit 4404 may perform bi - directional prediction on a current video block. The motion estimation unit 4404 may search for a reference video block for the current video block in the reference pictures of list 0 and may also search for another reference video block for the current video block in the reference pictures of list 1. Then, the motion estimation unit 4404 may generate a reference index that indicates the reference pictures in list 0 and list 1 that contain the reference video blocks and a motion vector that indicates the spatial displacement between the reference video block and the current video block. The motion estimation unit 4404 may output the reference index and the motion vector of the current video block as the motion information of the current video block. The motion compensation unit 4405 may generate a predicted video block of the current video block based on the reference video block indicated by the motion information of the current video block.
[0522] In some examples, the motion estimation unit 4404 may output a complete set of motion information for the decoder's decoding process. In some examples, the motion estimation unit 4404 may not output a complete set of motion information for the current video. More precisely, the motion estimation unit 4404 may signal the motion information of the current video block by referring to the motion information of another video block. For example, the motion estimation unit 4404 may determine that the motion information of the current video block is sufficiently similar to the motion information of a neighboring video block.
[0523] In one example, the motion estimation unit 4404 may indicate a certain value in the syntax structure associated with the current video block, and this value indicates to the video decoder 4500 that the current video block has the same motion information as another video block.
[0524] In another example, the motion estimation unit 4404 may identify another video block and a motion vector difference (MVD) in the syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 4500 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[0525] As discussed above, the video encoder 4400 may signal motion vectors in a predictive manner. Two examples of predictive signaling techniques that may be implemented by the video encoder 4400 include advanced motion vector prediction (AMVP) and Merge mode signaling.
[0526] The intra prediction unit 4406 may perform intra prediction on the current video block. When the intra prediction unit 4406 performs intra prediction on the current video block, the intra prediction unit 4406 may generate prediction data for the current video block based on the decoded samples of other video blocks in the same picture. The prediction data for the current video block may include a predicted video block and various syntax elements.
[0527] The residual generation unit 4407 may generate residual data for the current video block by subtracting the predicted video block of the current video block from the current video block. The residual data for the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.
[0528] In other examples, such as in the skip mode, there may be no residual data for the current video block of the current video block, and the residual generation unit 4407 may not perform the subtraction operation.
[0529] The transform processing unit 4408 may generate a transform coefficient video block for the current video block by applying one or more transforms to the residual video block associated with the current video block.
[0530] After the transform processing unit 4408 generates the transform coefficient video block associated with the current video block, the quantization unit 4409 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0531] The dequantization unit 4410 and the inverse transform unit 4411 may apply dequantization and inverse transform to the transform coefficient video block, respectively, to reconstruct the residual video block from the transform coefficient video block. The reconstruction unit 4412 may add the reconstructed residual video block to the corresponding samples of one or more predicted video blocks generated by the prediction unit 4402 to produce a reconstructed video block associated with the current block for storage in the buffer 4413.
[0532] After reconstructing a video block in the reconstruction unit 4412, a loop filtering operation may be performed to reduce video block artifacts in the video block.
[0533] The entropy coding unit 4414 may receive data from other functional components of the video encoder 4400. When the entropy coding unit 4414 receives data, the entropy coding unit 4414 may perform one or more entropy coding operations to generate entropy coded data and output a bitstream including the entropy coded data.
[0534] Figure 18 is a block diagram showing an example of a video decoder 4500, which may be Figure 16 the video decoder 4324 in the system 4300 shown. The video decoder 4500 may be configured to perform any or all of the techniques of the present disclosure. In the example shown, the video decoder 4500 includes a plurality of functional components. The techniques described in the present disclosure may be shared among the various components of the video decoder 4500. In some examples, a processor may be configured to perform any or all of the techniques described in the present disclosure.
[0535] In the example shown, the video decoder 4500 includes an entropy decoding unit 4501, a motion compensation unit 4502, an intra prediction unit 4503, an inverse quantization unit 4504, an inverse transform unit 4505, a reconstruction unit 4506, and a cache 4507. In some examples, the video decoder 4500 may perform a decoding process that is substantially inverse to the encoding process described with respect to the video encoder 4400.
[0536] The entropy decoding unit 4501 may extract the encoded bitstream. The encoded bitstream may include entropy encoded video data (e.g., encoded video data blocks). The entropy decoding unit 4501 may decode the entropy encoded video data, and the motion compensation unit 4502 may determine motion information based on the entropy decoded video data, the motion information including motion vectors, motion vector precision, reference picture list indices, and other motion information. The motion compensation unit 4502 may determine such information, for example, by performing AMVP and Merge modes.
[0537] The motion compensation unit 4502 may generate a motion compensated block, and thus may perform interpolation based on an interpolation filter. The identity of the interpolation filter used with sub-pixel precision may be included in a syntax element.
[0538] The motion compensation unit 4502 may use the interpolation filter used by the video encoder 4400 during the encoding of a video block to calculate the interpolation of sub-integer pixels of a reference block. The motion compensation unit 4502 may determine the interpolation filter used by the video encoder 4400 according to the received syntax information, and use the interpolation filter to generate a prediction block.
[0539] The motion compensation unit 4502 may use some syntax information to determine the size of the blocks used to encode frames and / or slices of an encoded video sequence, the partitioning information that describes how each macroblock of a picture of the encoded video sequence is partitioned, the mode indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-coded block, and other information for decoding the encoded video sequence.
[0540] The intra prediction unit 4503 may form a prediction block from spatially adjacent blocks using, for example, the intra prediction mode received in the bitstream. The inverse quantization unit 4504 inverse quantizes the video block coefficients that are provided in the bitstream and decoded by the entropy decoding unit 4501, i.e., de-quantizes them. The inverse transform unit 4505 applies an inverse transform.
[0541] The reconstruction unit 4506 may add a residual block to the corresponding prediction block generated by the motion compensation unit 4502 or the intra prediction unit 4503 to form a decoded block. If necessary, a de-blocking filter may also be applied to filter the decoded block to remove blocking artifacts. Then, the decoded video blocks are stored in the buffer 4507, which provides reference blocks for subsequent motion compensation / intra prediction and also produces decoded video for presentation on a display device.
[0542] Figure 19 is a schematic diagram of an example encoder 4600. The encoder 4600 is suitable for implementing VVC technology. The encoder 4600 includes three loop filters, namely, a de-blocking filter (DF) 4602, sample adaptive offset (SAO) 4604, and an adaptive loop filter (ALF) 4606. Different from the DF 4602 that uses a predefined filter, the SAO 4604 and the ALF 4606 utilize the original samples of the current picture to reduce the mean squared error between the original samples and the reconstructed samples by respectively adding offset values and applying a finite impulse response (FIR) filter and signaling the offset values and filter coefficients using coded side information. The ALF 4606 is located at the last processing stage of each picture and can be regarded as a tool that attempts to capture and fix the artifacts created by the previous stages.
[0543] The encoder 4600 also includes an intra prediction component 4608 configured to receive an input video and a motion estimation / compensation (ME / MC) component 4610. The intra prediction component 4608 is configured to perform intra prediction, while the ME / MC component 4610 is configured to perform inter prediction using reference pictures obtained from a reference picture buffer 4612. Residual blocks from inter prediction or intra prediction are fed into a transform (T) component 4614 and a quantization (Q) component 4616 to generate quantized residual transform coefficients, which are fed into an entropy coding / decoding component 4618. The entropy coding / decoding component 4618 performs entropy coding / decoding on the prediction results and the quantized transform coefficients and sends them to a video decoder (not shown). The quantized components output from the quantization component 4616 can be fed into an inverse quantization (IQ) component 4620, an inverse transform component 4622, and a reconstruction (REC) component 4624. The REC component 4624 is capable of outputting an image to the DF 4602, SAO 4604, and ALF 4606 for filtering before these pictures are stored in the reference picture buffer 4612.
[0544] Figure 20 is a method for processing video data according to an embodiment of the present disclosure. In block 2002, the method includes using side information as an input to an adaptive loop filter (ALF) and reusing the filter for a current ALF processing unit. In block 2004, the method includes performing a conversion between visual media data and a bitstream based on the ALF.
[0545] Next, a list of some example preferred solutions is provided.
[0546] The following solutions illustrate examples of the techniques discussed herein.
[0547] 1. A method for processing video data, comprising: determining to use side information as an input to an adaptive loop filter (ALF); and performing a conversion between visual media data and a bitstream based on the ALF.
[0548] 2. The method according to solution 1, wherein the side information includes luminance reconstruction samples obtained before applying a deblocking filter (DBF), a sample adaptive offset (SAO) filter, or a bilateral filter (BF).
[0549] 3. The method according to any one of solutions 1 to 2, wherein the side information includes chrominance reconstruction samples obtained before applying the DBF, SAO, or BF.
[0550] 4. The method according to any one of solutions 1 to 3, wherein the side information is used as an input source for extended taps or spatial domain taps of an online training filter.
[0551] 5. The method according to any one of Solutions 1 to 4, wherein the side information is used as an input source for offline training of a filter.
[0552] 6. The method according to any one of Solutions 1 to 5, wherein the side information is used as an input source for classification.
[0553] 7. The method according to any one of Solutions 1 to 6, wherein the side information is used as an input source for classification of a luminance online training filter or a luminance offline training filter.
[0554] 8. The method according to any one of Solutions 1 to 7, wherein the side information includes prediction samples, and wherein the prediction samples are luminance prediction samples or chrominance prediction samples.
[0555] 9. The method according to any one of Solutions 1 to 8, wherein the prediction samples are modified by filtering, downsampling, upsampling, clipping, shifting, or a combination thereof before being used as side information.
[0556] 10. The method according to any one of Solutions 1 to 9, wherein the side information includes residual values, and wherein the residual values are luminance residual values or chrominance residual values.
[0557] 11. The method according to any one of Solutions 1 to 10, wherein the residual values are modified by filtering, downsampling, upsampling, clipping, shifting, or a combination thereof before being used as side information.
[0558] 12. The method according to any one of Solutions 1 to 11, wherein the side information includes segmentation information, and wherein the segmentation information is luminance segmentation information or chrominance segmentation information.
[0559] 13. The method according to any one of Solutions 1 to 12, wherein the segmentation information includes block size, block shape, block position, or a combination thereof.
[0560] 14. The method according to any one of Solutions 1 to 13, wherein the side information is used to train a luminance online training filter or a chrominance online training filter.
[0561] 15. The method according to any one of Solutions 1 to 14, wherein the side information is used as part of a filtering process in the ALF.
[0562] 16. The method according to any one of Solutions 1 to 15, wherein the side information includes quantization parameter (QP) information, and wherein the QP information includes luminance QP information or chrominance QP information.
[0563] 17. The method according to any one of Solutions 1 to 16, wherein the QP information includes picture QP information, slice QP information, block QP information, or a combination thereof.
[0564] 18. The method according to any one of Solutions 1 to 17, wherein the side information includes boundary strength information, and wherein the boundary strength information includes luminance boundary strength information or chrominance boundary strength information.
[0565] 19. The method according to any one of Solutions 1 to 18, wherein the boundary strength information is generated by DBF.
[0566] 20. The method according to any one of Solutions 1 to 19, wherein the ALF is a cross-component ALF (CCALF).
[0567] 21. The method according to any one of Solutions 1 to 20, wherein the method is for preprocessing or postprocessing of video.
[0568] 22. The method according to any one of Solutions 1 to 21, wherein the method is applied as part of a loop filter, and wherein the method is applied to an ALF, CCALF, sample adaptive offset (SAO) filter, cross-component SAO (CCSAO), bilateral filter (BF), Hadamard transform domain filter (HTDF), or a combination thereof.
[0569] 23. The method according to any one of Solutions 1 to 22, wherein the ALF is applied to a video unit, and wherein the video unit is a sequence, picture, sub-picture, slice, tile, coding tree unit (CTU), CTU row, CTU group, coding unit (CU), prediction unit (PU), transform unit (TU), coding tree block (CTB), coding block (CB), prediction block (PB), transform block (TB), or any other region containing more than one luminance or chrominance sample or pixel.
[0570] 24. The method according to any one of Solutions 1 to 23, wherein the application of the method is signaled in the bitstream.
[0571] 25. The method according to any one of Solutions 1 to 24, wherein the application of the method depends on coding information.
[0572] 26. An apparatus for processing video data, comprising: a processor; and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of Solutions 1 to 25.
[0573] 27. A non-transitory computer-readable medium includes a computer program product for use by a video codec device. The computer program product includes computer-executable instructions stored on the non-transitory computer-readable medium that, when executed by a processor, cause the video codec device to perform the method according to any one of Solutions 1 to 25.
[0574] 28. A non-transitory computer-readable recording medium stores a bitstream of a video generated by a method executed by a video processing device. The method includes: determining to use side information as an input to an adaptive loop filter (ALF); and generating the bitstream based on the determination.
[0575] 29. A method for storing a bitstream of a video includes: determining to use side information as an input to an adaptive loop filter (ALF); generating the bitstream based on the determination; and storing the bitstream in a non-transitory computer-readable recording medium.
[0576] 30. A method, apparatus, or system described in the present disclosure.
[0577] In the solutions described herein, an encoder can conform to format rules by generating a codec representation according to the format rules. In the solutions described herein, a decoder can use the format rules to parse syntax elements in the codec representation and, based on the presence and absence of known syntax elements according to the format rules, to generate a decoded video.
[0578] In the present disclosure, the term "video processing" can refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm can be applied during the conversion from a pixel representation of a video to a corresponding bitstream representation (and vice versa). For example, the bitstream representation of a current video block can correspond to bits that are co-located or distributed at different positions in the bitstream, as defined by the syntax. For example, a macroblock can be encoded based on transformed and encoded error residual values and also encoded using bits in the header and other fields in the bitstream. Additionally, during the conversion, a decoder can parse the bitstream based on a determination and know that certain fields may be present or absent, as described in the above solutions. Similarly, an encoder can determine whether to include certain syntax fields and generate an encoded representation accordingly by including or excluding the syntax fields in the encoded representation.
[0579] The disclosed and other solutions, examples, embodiments, modules, and functional operations described in this disclosure can be implemented in digital electronic circuitry or in computer software, firmware, hardware, or a combination of one or more of them, which includes the structures disclosed in this disclosure and their equivalent structures. The disclosed and other embodiments can be implemented as one or more computer program products encoded on a computer-readable medium for execution by, or to control the operation of, a data processing apparatus, i.e., one or more modules of computer program instructions. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition of matter affecting a machine-readable propagated signal, or a combination of one or more of them. The term “data processing apparatus” encompasses all apparatus, devices, and machines for processing data, including, by way of example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, the apparatus can also include code that creates an execution environment for the computer programs being discussed, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them. A propagated signal is an artificially generated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, which has been generated to encode information for transmission to a suitable receiver apparatus.
[0580] A computer program (also known as a program, software, software application, script, or code) can be written in any form of programming language (including a compiled or interpreted language), and it can be deployed in any form (including as a stand-alone program or as a module, component, subroutine, or any other unit suitable for use in a computing environment). A computer program does not necessarily correspond to a file in a file system. The program can be stored in a part of a file that holds other programs or data (e.g., one or more scripts in a markup language document), in a single file dedicated to the program being discussed, or in multiple coordinated files (e.g., files that store one or more modules, subroutines, or portions of code). A computer program can be deployed to be executed on one computer or on multiple computers, which are located at one site or distributed across multiple sites and interconnected by a communication network.
[0581] The processes or logical flows described in this disclosure can be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logical flows can also be performed by, and the apparatus can also be implemented as, special purpose logic circuitry, e.g., a field programmable gate array (FPGA) or an application specific integrated circuit (ASIC).
[0582] Processors suitable for executing computer programs include, for example, any one or more processors of both general and special purpose microprocessors and any kind of digital computer. Generally, a processor will receive instructions and data from a read only memory or a random access memory or both. The basic elements of a computer are a processor for executing the instructions and one or more memory devices for storing the instructions and data. Generally, a computer will also include one or more mass storage devices (such as, for example, magnetic disks, magneto-optical disks, or optical disks) for storing data, or be operatively coupled to receive data from one or more mass storage devices or transfer data to one or more mass storage devices or both. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices such as erasable programmable read only memory (EPROM), electrically erasable programmable read only memory (EEPROM), and flash memory devices; magnetic disks such as internal hard disks or removable disks; magneto-optical disks; and compact disc read only memory (CD ROM) and digital versatile disc read only memory (DVD-ROM) disks. The processor and the memory may be supplemented by, or incorporated in, special purpose logic circuitry.
[0583] Although this disclosure contains many details, these details should not be construed as limiting the scope of any subject matter or of what may be claimed, but rather as descriptions of features that may be specific to particular embodiments of a particular technology. Certain features that are described in this disclosure in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments. Moreover, although the features may be described above as acting in certain combinations and even initially claimed as such, in some cases, one or more features from a claimed combination may be excluded from the combination, and the claimed combination may cover a sub-combination or a variation of a sub-combination.
[0584] Similarly, although the operations are shown in the figures in a particular order, this should not be understood as requiring that the operations be performed in the particular order shown or in a sequential order, or that all of the illustrated operations be performed to achieve the desired result. Additionally, the separation of various system components described in this disclosure should not be understood as requiring such separation in all embodiments.
[0585] Only a few embodiments and examples have been described, and other embodiments, enhancements, and changes may be made based on what is described and shown in this disclosure.
[0586] When there is no intermediate component other than a wire, trace, or another medium between a first component and a second component, the first component is directly coupled to the second component. When there is an intermediate component between the first component and the second component other than a wire, trace, or another medium, the first component is indirectly coupled to the second component. The term "coupled" and its variants include both direct coupling and indirect coupling. Unless otherwise specified, the use of the term "about" means a range of ±10% of the subsequent numerical value.
[0587] Although several embodiments are provided in this disclosure, it should be understood that the disclosed systems and methods may be embodied in many other specific forms without departing from the spirit or scope of the disclosure. This example is considered illustrative rather than restrictive and is not intended to be limited to the details given herein. For example, various elements or components may be combined or integrated into another system, or certain features may be omitted or not implemented.
[0588] In addition, without departing from the scope of the disclosure, the techniques, systems, subsystems, and methods described and shown as discrete or separate in various embodiments may be combined or integrated with other systems, modules, techniques, or methods. Other items shown or discussed as being coupled may be directly connected, or may be indirectly coupled or communicate through some interface, device, or intermediate component, whether electrically, mechanically, or otherwise coupled or communicated. Those skilled in the art will recognize other examples of changes, substitutions, and alterations, and may make these changes, substitutions, and alterations without departing from the spirit and scope disclosed herein.
Claims
1. A method for processing video data, comprising: using side information as an input to an adaptive loop filter (ALF); and performing a conversion between visual media data and a bitstream based on the ALF.
2. The method according to claim 1, wherein the side information includes reconstructed samples obtained before a deblocking filter (DBF), a sample adaptive offset (SAO) filter, or a bilateral filter (BF).
3. The method according to claim 2, wherein the reconstructed samples include luminance reconstructed samples, and the luminance reconstructed samples are used as an input source for at least one extended tap of a luminance or chrominance online training filter.
4. The method according to claim 2, wherein the reconstructed samples include luminance reconstructed samples, and the luminance reconstructed samples are used as an input source for at least one spatial domain tap of a luminance or chrominance online training filter.
5. The method according to claim 2, wherein the reconstructed samples include luminance reconstructed samples, and the luminance reconstructed samples are used as an input source for a luminance offline training filter.
6. The method according to claim 2, wherein the reconstructed samples include luminance reconstructed samples, and the luminance reconstructed samples are used as an input source for classification.
7. The method according to claim 6, wherein the classification includes classification of a luminance online training filter.
8. The method according to claim 6, wherein the classification includes classification of a luminance offline training filter.
9. The method according to claim 2, wherein the reconstructed samples include chrominance reconstructed samples, and the chrominance reconstructed samples are used as an input source for at least one extended tap of a luminance or chrominance online training filter.
10. The method according to claim 2, wherein the reconstructed samples include chrominance reconstructed samples, and the chrominance reconstructed samples are used as an input source for at least one spatial domain tap of a luminance or chrominance online training filter.
11. The method according to claim 2, wherein the reconstructed samples include chrominance reconstructed samples, and the chrominance reconstructed samples are used as an input source for classification.
12. The method according to claim 11, wherein the classification includes classification of a luminance online training filter.
13. The method according to claim 11, wherein the classification includes classification of a luminance offline training filter.
14. The method according to claim 1, wherein the side information includes predicted samples.
15. The method according to claim 14, wherein the predicted samples include luminance predicted samples, and the luminance predicted samples are used as an input source for at least one extended tap of a luminance or chrominance online training filter.
16. The method according to claim 14, wherein the predicted samples include luminance predicted samples, and the luminance predicted samples are used as an input source for at least one spatial domain tap of a luminance or chrominance online training filter.
17. The method according to claim 14, wherein the predicted samples include luminance predicted samples, and the luminance predicted samples are used as an input source for a luminance offline training filter.
18. The method according to claim 14, wherein the predicted samples include luminance predicted samples, and the luminance predicted samples are used as an input source for classification.
19. The method according to claim 18, wherein the classification includes the classification of the luminance online training filter.
20. The method according to claim 18, wherein the classification includes the classification of the luminance offline training filter.
21. The method according to claim 14, wherein the reconstructed samples include chrominance reconstructed samples, and the chrominance reconstructed samples are used as an input source for at least one extended tap of the luminance or chrominance online training filter.
22. The method according to claim 14, wherein the reconstructed samples include chrominance reconstructed samples, and the chrominance reconstructed samples are used as an input source for at least one spatial domain tap of the luminance or chrominance online training filter.
23. The method according to claim 14, wherein the reconstructed samples include chrominance reconstructed samples, and the chrominance reconstructed samples are used as an input source for classification.
24. The method according to claim 23, wherein the classification includes the classification of the luminance online training filter.
25. The method according to claim 23, wherein the classification includes the classification of the luminance offline training filter.
26. The method according to claim 14, further comprising modifying the predicted samples before using the predicted samples as the input to the ALF.
27. The method according to claim 26, wherein the modification includes filtering.
28. The method according to claim 26, wherein the modification includes downsampling or upsampling.
29. The method according to claim 26, wherein the modification includes clipping or shifting.
30. The method according to claim 1, wherein the side information includes residual values.
31. The method according to claim 30, wherein the residual values include luminance residual values, and the luminance residual values are used as an input source for at least one extended tap of the luminance or chrominance online training filter.
32. The method according to claim 30, wherein the residual values include luminance residual values, and the luminance residual values are used as an input source for at least one spatial domain tap of the luminance or chrominance online training filter.
33. The method according to claim 30, wherein the residual values include luminance residual values, and the luminance residual values are used as an input source for the luminance offline training filter.
34. The method according to claim 30, wherein the residual values include luminance residual values, and the luminance residual values are used as an input source for classification.
35. The method according to claim 34, wherein the classification includes the classification of the luminance online training filter.
36. The method according to claim 34, wherein the classification includes the classification of the luminance offline training filter.
37. The method according to claim 30, wherein the residual values include chrominance residual values, and the chrominance residual values are used as an input source for at least one extended tap of the luminance or chrominance online training filter.
38. The method according to claim 30, wherein the residual values include chrominance residual values, and the chrominance residual values are used as an input source for at least one spatial domain tap of the luminance or chrominance online training filter.
39. The method according to claim 30, wherein the residual values include chrominance residual values, and the chrominance residual values are used as an input source for classification.
40. The method according to claim 39, wherein the classification includes the classification of a luminance online training filter.
41. The method according to claim 39, wherein the classification includes the classification of a luminance offline training filter.
42. The method according to claim 14, further comprising modifying the residual value before using the residual value as the input of the ALF.
43. The method according to claim 42, wherein the modification includes filtering.
44. The method according to claim 42, wherein the modification includes downsampling or upsampling.
45. The method according to claim 42, wherein the modification includes clipping or shifting.
46. The method according to claim 1, wherein the side information includes segmentation information.
47. The method according to claim 46, wherein the segmentation information includes one or more of a block size, a block shape, and a block position.
48. The method according to claim 46, wherein the segmentation information includes luminance segmentation information for classification.
49. The method according to claim 48, wherein the classification includes the classification of a luminance online training filter.
50. The method according to claim 48, wherein the classification includes the classification of a luminance offline training filter.
51. The method according to claim 46, wherein the segmentation information includes chrominance segmentation information for classification.
52. The method according to claim 48, wherein the classification includes the classification of a luminance online training filter.
53. The method according to claim 48, wherein the classification includes the classification of a luminance offline training filter.
54. The method according to claim 46, wherein the segmentation information includes luminance segmentation information for filter training.
55. The method according to claim 54, wherein the filter training includes the filter training of a luminance online training filter.
56. The method according to claim 54, wherein the filter training includes the filter training of a chrominance online training filter.
57. The method according to claim 46, wherein the segmentation information includes chrominance segmentation information for filter training.
58. The method according to claim 54, wherein the filter training includes the filter training of a luminance online training filter.
59. The method according to claim 54, wherein the filter training includes the filter training of a chrominance online training filter.
60. The method according to claim 46, wherein the segmentation information includes luminance segmentation information for filter training and chrominance segmentation information for the filter training.
61. The method according to claim 1, wherein the side information includes quantization parameter (QP) information.
62. The method according to claim 61, wherein the QP information includes one or more of picture QP information, slice QP information, and block-level QP information.
63. The method according to claim 61, wherein the QP information includes luminance QP information for classification.
64. The method according to claim 63, wherein the classification includes the classification of a luminance online training filter.
65. The method according to claim 63, wherein the classification includes classification of a luminance offline training filter.
66. The method according to claim 61, wherein the QP information includes chrominance QP information for classification.
67. The method according to claim 66, wherein the classification includes classification of a luminance online training filter.
68. The method according to claim 66, wherein the classification includes classification of a luminance offline training filter.
69. The method according to claim 61, wherein the QP information includes luminance QP information for filter training.
70. The method according to claim 69, wherein the filter training includes filter training of a luminance online training filter.
71. The method according to claim 69, wherein the filter training includes filter training of a chrominance online training filter.
72. The method according to claim 61, wherein the QP information includes chrominance QP information for filter training.
73. The method according to claim 71, wherein the filter training includes filter training of a luminance online training filter.
74. The method according to claim 71, wherein the filter training includes filter training of a chrominance online training filter.
75. The method according to claim 61, wherein the QP information includes luminance QP information for filter training and chrominance QP information for the filter training.
76. The method according to claim 1, wherein the side information includes boundary strength information.
77. The method according to claim 76, wherein the boundary strength information is generated by a deblocking filter.
78. The method according to claim 76, wherein the boundary strength information includes luminance boundary strength information for classification.
79. The method according to claim 78, wherein the classification includes classification of a luminance online training filter.
80. The method according to claim 78, wherein the classification includes classification of a luminance offline training filter.
81. The method according to claim 76, wherein the boundary strength information includes chrominance boundary strength for classification.
82. The method according to claim 81, wherein the classification includes classification of a luminance online training filter.
83. The method according to claim 81, wherein the classification includes classification of a luminance offline training filter.
84. The method according to claim 76, wherein the boundary strength information includes luminance boundary strength information for filter training.
85. The method according to claim 84, wherein the filter training includes filter training of a luminance online training filter.
86. The method according to claim 84, wherein the filter training includes filter training of a chrominance online training filter.
87. The method according to claim 76, wherein the boundary strength information includes chrominance boundary strength information for filter training.
88. The method according to claim 87, wherein the filter training includes filter training of a luminance online training filter.
89. The method according to claim 87, wherein the filter training includes filter training of chrominance online training filters.
90. The method according to claim 76, wherein the boundary strength information includes luminance boundary strength information for filter training and chrominance boundary strength information for the filter training.
91. The method according to claim 1, wherein the ALF includes cross-component ALF (CCALF).
92. The method according to claim 91, wherein the side information includes reconstructed samples obtained before a deblocking filter (DBF), a sample adaptive offset (SAO) filter, or a bilateral filter (BF).
93. The method according to claim 91, wherein the reconstructed samples include luminance reconstructed samples, and the luminance reconstructed samples are used as an input source for at least one extended tap of the CCALF.
94. The method according to claim 91, wherein the reconstructed samples include luminance reconstructed samples, and the luminance reconstructed samples are used as an input source for at least one spatial domain tap of the CCALF.
95. The method according to claim 91, wherein the reconstructed samples include chrominance reconstructed samples, and the chrominance reconstructed samples are used as an input source for at least one extended tap of the CCALF.
96. The method according to claim 91, wherein the reconstructed samples include chrominance reconstructed samples, and the chrominance reconstructed samples are used as an input source for at least one spatial domain tap of the CCALF.
97. The method according to claim 91, wherein the side information includes predicted samples.
98. The method according to claim 97, wherein the predicted samples include luminance predicted samples, and the luminance predicted samples are used as an input source for at least one extended tap of the CCALF.
99. The method according to claim 97, wherein the predicted samples include luminance predicted samples, and the luminance predicted samples are used as an input source for at least one spatial domain tap of the CCALF.
100. The method according to claim 97, wherein the predicted samples include chrominance predicted samples, and the chrominance predicted samples are used as an input source for at least one extended tap of the CCALF.
101. The method according to claim 97, wherein the predicted samples include chrominance predicted samples, and the chrominance predicted samples are used as an input source for at least one spatial domain tap of the CCALF.
102. The method according to claim 97, further comprising modifying the predicted samples before using the predicted samples as the input for the CCALF.
103. The method according to claim 102, wherein the modification includes filtering.
104. The method according to claim 102, wherein the modification includes downsampling or upsampling.
105. The method according to claim 102, wherein the modification includes clipping or shifting.
106. The method according to claim 97, wherein the side information includes residual values.
107. The method according to claim 106, wherein the residual value includes a luminance residual value, and the luminance residual value is used as an input source for at least one extended tap of the CCALF.
108. The method according to claim 106, wherein the residual value includes a luminance residual value, and the luminance residual value is used as an input source for at least one spatial domain tap of the CCALF.
109. The method according to claim 106, wherein the residual value includes a chrominance residual value, and the chrominance residual value is used as an input source for at least one extended tap of the CCALF.
110. The method according to claim 106, wherein the residual value includes a chrominance residual value, and the chrominance residual value is used as an input source for at least one spatial domain tap of the CCALF.
111. The method according to claim 106, further comprising modifying the residual value before using the residual value as the input of the CCALF.
112. The method according to claim 111, wherein the modification includes filtering.
113. The method according to claim 111, wherein the modification includes downsampling or upsampling.
114. The method according to claim 111, wherein the modification includes clipping or shifting.
115. The method according to claim 91, wherein the side information includes segmentation information.
116. The method according to claim 115, wherein the segmentation information includes one or more of block size, block shape, and block position.
117. The method according to claim 115, wherein the segmentation information includes luminance segmentation information for an online training filter for training the CCALF.
118. The method according to claim 115, wherein the segmentation information includes chrominance segmentation information for an online training filter for training the CCALF.
119. The method according to claim 115, wherein the segmentation information includes luminance segmentation information for filter training of the CCALF and chrominance segmentation information for filter training of the CCALF.
120. The method according to claim 91, wherein the side information includes quantization parameter (QP) information.
121. The method according to claim 120, wherein the QP information includes one or more of block size, block shape, and block position.
122. The method according to claim 120, wherein the QP information includes luminance QP information for an online training filter for training the CCALF.
123. The method according to claim 120, wherein the QP information includes chrominance QP information for an online training filter for training the CCALF.
124. The method according to claim 120, wherein the QP information includes luminance QP information for filter training of the CCALF and chrominance QP information for filter training of the CCALF.
125. The method according to claim 91, wherein the side information includes boundary strength information.
126. The method according to claim 120, wherein the boundary strength information is generated by a deblocking filter.
127. The method according to claim 120, wherein the boundary strength information includes the luminance boundary strength information of an online training filter for training the CCALF.
128. The method according to claim 120, wherein the boundary strength information includes the chrominance boundary strength information of an online training filter for training the CCALF.
129. The method according to claim 120, wherein the boundary strength information includes the luminance boundary strength information for filter training of the CCALF and the chrominance boundary strength information for filter training of the CCALF.
130. The method according to any one of claims 1 to 129, wherein the method is performed in post - processing or pre - processing.
131. The method according to any one of claims 1 to 129, wherein one or more of the methods are used jointly.
132. The method according to any one of claims 1 to 129, wherein one or more of the methods are used separately.
133. The method according to any one of claims 1 to 132, wherein the method is applied to any loop filtering tool, pre - processing filtering method or post - processing filtering method in video coding and decoding, including but not limited to ALF / CCALF or any other filtering method.
134. The method according to claim 133, wherein the method is applied to a loop filtering method.
135. The method according to claim 134, wherein the method is applied to ALF, CCAPF, sample adaptive offset (SAO) filter or other loop filtering methods.
136. The method according to claim 134, wherein the method is applied to cross - component SAO (CCSAO).
137. The method according to claim 134, wherein the method is applied to a bilateral filter (BF).
138. The method according to claim 134, wherein the method is applied to a Hadamard transform domain filter (HTDF).
139. The method according to claim 134, wherein the method is applied to a loop filtering method.
140. The method according to claim 134, wherein the method is applied to a pre - processing filtering method.
141. The method according to claim 134, wherein the method is applied to a post - processing filtering method.
142. The method according to any one of claims 1 to 141, wherein the video unit processed by the ALF or the CCALF includes one of a sequence, a picture, a sub - picture, a slice, a strip, a tile, a coding tree unit (CTU), a CTU row, a CTU group, a coding unit (CU), a prediction unit (PU), a transform unit (TU), a coding tree block (CTB), a coding block (CB), a prediction block (PB), a transform block (TB), and any other region containing more than one luminance or chrominance sample or pixel.
143. The method according to any one of claims 1 to 142, wherein an indication of whether and / or how to apply the method is included in the bitstream.
144. The method according to claim 143, wherein the indication is included in a sequence level, group of pictures level, picture level, slice level, slice group level, sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), dependent parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), slice header or slice group header in the bitstream.
145. The method according to claim 143, wherein the indication is included in a prediction block (PB), transform block (TB), codec block (CB), prediction unit (PU), transform unit (TU), codec unit (CU), virtual pipeline data unit (VPDU), codec tree unit (CTU), CTU row, slice, tile, sub-picture or any other region containing more than one sample or pixel in the bitstream.
146. The method according to any one of claims 1 to 145, wherein whether and / or how to apply any one of the methods depends on codec information, and the codec information includes block size, color format, single-tree or dual-tree partitioning, color component, slice type or picture type.
147. The method according to any one of claims 1 to 121, wherein the conversion includes encoding the visual media data into the bitstream.
148. The method according to any one of claims 1 to 121, wherein the conversion includes decoding the visual media data from the bitstream.
149. An apparatus for processing media data, comprising: one or more processors; and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processors, cause the apparatus to perform the method according to any one of claims 1 to 148.
150. A non-transitory computer-readable medium, comprising a computer program product for use by a video codec device, the computer program product including computer-executable instructions stored on the non-transitory computer-readable medium, such that when executed by a processor, cause the video codec device to perform the method according to any one of claims 1 to 148.
151. A non-transitory computer-readable recording medium that stores a bitstream of a video generated by a method executed by a video processing device, wherein the method includes the method according to any one of claims 1 to 148.
152. A method for storing a bitstream of a video, including the method according to any one of claims 1 to 148.
153. A method, apparatus or system described in the present disclosure.