Application of side information to bilateral filter in video coding and decoding
By using more edge information in the bilateral filter, the problem of inefficiency in the prior art is solved, and a more efficient video encoding and decoding effect is achieved.
Patent Information
- Application Number
- CN202380077266.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-11-01
- Filing Date
- 2023-11-01
- Publication Date
- 2025-06-17
AI Technical Summary
In the existing video encoding and decoding technology, bilateral filters fail to fully utilize the edge information of other loop filters during design, resulting in inefficiency.
Use edge information in a bilateral filter, including de-blocking filter, sample point adaptive compensation and sample loop filter before sample point, predicted sample point, residual information, segmentation information and quantization parameter information to improve the performance of the filter.
By utilizing more edge information, the bilateral filter can more effectively reduce the mean square error between the original sample point and the reconstructed sample point, and improve the efficiency and quality of video encoding and decoding.
Smart Images

Figure CN120167111A_ABST
Abstract
Description
[0001] Cross - reference to related patent applications
[0002] This patent application claims the priority of International Patent Application No. PCT / CN2022 / 128887, filed on November 1, 2022. The disclosure and teachings of the aforementioned patent application are incorporated herein by reference in their entirety. Technical Field
[0003] This patent document relates to the generation, storage, and consumption of digital audio - visual media information in a file format. Background Art
[0004] Digital video accounts for the largest bandwidth used on the Internet and other digital communication networks. With the increase in the number of connected user devices capable of receiving and displaying video, the bandwidth demand for digital video is likely to continue to grow. Summary of the Invention
[0005] The first aspect relates to a method for processing video data, including: determining to use side information as an input to a filter; and performing a conversion between visual media data and a bitstream based on the filter.
[0006] The second aspect relates to a device for processing video data, including: a processor; and a non - transitory memory having instructions thereon, where the instructions, when executed by the processor, cause the processor to perform any of the foregoing aspects.
[0007] The third aspect relates to a non - transitory computer - readable medium, including a computer program product for use in a video codec device, the computer program product including computer - executable instructions stored on the non - transitory computer - readable medium, such that when the processor executes the computer - executable instructions, the video codec device performs any of the foregoing aspects.
[0008] The fourth aspect relates to a non - transitory computer - readable recording medium that stores a bitstream of a video generated by a method executed by a video processing device, where the method includes: determining to use side information as an input to a filter; and generating the bitstream based on the determination.
[0009] The fifth aspect relates to a method for storing a bitstream of a video, including: determining to use side information as an input to a filter; generating the bitstream based on the determination; and storing the bitstream in a non - transitory computer - readable recording medium.
[0010] For clarity, any of the foregoing embodiments may be combined with one or more of the other foregoing embodiments to produce new embodiments within the scope of the present disclosure.
[0011] These and other features will be more clearly understood from the following detailed description taken in conjunction with the accompanying drawings and the claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] To more fully understand the present disclosure, reference is now made to the following brief description taken in conjunction with the accompanying drawings and specific embodiments, in which like reference numerals represent like components.
[0013] Figure 1 An example of the nominal vertical and horizontal positions of the luminance and chrominance samples in the 4:2:2 format in the picture is shown.
[0014] Figure 2 An example encoder block diagram is shown.
[0015] Figure 3 An example picture divided into raster scan strips is shown.
[0016] Figure 4 An example picture divided into rectangular scan strips is shown.
[0017] Figure 5 An example picture divided into bricks is shown.
[0018] Figures 6A to 6C An example of a coding tree block (CTB) across picture boundaries is shown.
[0019] Figure 7 An example of an intra prediction mode is shown.
[0020] Figure 8 An example of a block boundary in the picture is shown.
[0021] Figure 9 An example of the pixels involved in filter use is shown.
[0022] Figure 10 An example of the filter shape of an adaptive loop filter (ALF) is shown.
[0023] Figure 11 An example of the transform coefficients supported by a 5×5 diamond filter is shown.
[0024] Figure 12 An example of the relative coordinates supported by a 5×5 diamond filter is shown.
[0025] Figure 13 A block diagram showing an example video processing system is shown.
[0026] Figure 14 A block diagram of an example video processing device is shown.
[0027] Figure 15A flowchart of an example method for video processing.
[0028] Figure 16 A block diagram showing an example video codec system.
[0029] Figure 17 A block diagram showing an example encoder.
[0030] Figure 18 A block diagram showing an example decoder.
[0031] Figure 19 A schematic diagram of an example encoder. Detailed implementation
[0032] First, it should be understood that although illustrative implementations of one or more embodiments are provided below, any number of techniques can be used to implement the disclosed systems and / or methods, whether currently known or to be developed. The present disclosure should not in any way be limited to the exemplary implementations, drawings, and techniques shown below, including the exemplary designs and implementations shown and described herein, but can be modified within the full scope of the appended claims and their equivalents.
[0033] The section headings used in this document are for ease of understanding and do not limit the applicability of the techniques and embodiments disclosed in each section to only that section. Additionally, the techniques described herein are applicable to other video codec protocols and designs.
[0034] 1. Initial discussion
[0035] This document relates to video codec technology. Specifically, this document relates to loop filters and other codec tools in image / video coding. These ideas can be applied alone or in various combinations to video codecs such as High Efficiency Video Coding (HEVC), Versatile Video Coding (VVC), or other video codec technologies.
[0036] 2. Abbreviations
[0037] The present disclosure includes the following abbreviations. Advanced Video Coding (ITU-T Recommendation H.264 | ISO / IEC 14496-10) (AVC), Coding Picture Buffer (CPB), Clean Random Access (CRA), Coding Tree Unit (CTU), Coding Video Sequence (CVS), Decoded Picture Buffer (DPB), Decoding Parameter Set (DPS), General Constraint Information (GCI), High Efficiency Video Coding (also known as ITU-T Recommendation H.265 | ISO / IEC 23008-2) (HEVC), Joint Exploration Model (JEM), Motion Constraint Tile Set (MCTS), Network Abstraction Layer (NAL), Output Layer Set (OLS), Picture Header (PH), Picture Parameter Set (PPS), Profile, Tier and Level (PTL), Picture Unit (PU), Reference Picture Resampling (RPR), Raw Byte Sequence Payload (RBSP), Supplemental Enhancement Information (SEI), Slice Header (SH), Sequence Parameter Set (SPS), Video Coding Layer (VCL), Video Parameter Set (VPS), Versatile Video Coding (also known as ITU-T Recommendation H.266 | ISO / IEC 23090-3) (VVC), VVC Test Model (VTM), Video Usability Information (VUI), Transform Unit (TU), Coding Unit (CU), Deblocking Filter (DF), Sample Adaptive Offset (SAO), Adaptive Loop Filter (ALF), Coding Block Flag (CBF), Quantization Parameter (QP), Rate-Distortion Optimization (RDO), and Bilateral Filter (BF).
[0038] 3. Video Coding Standards
[0039] Video coding standards have evolved mainly through the development of ITU-T and ISO / IEC standards. ITU-T developed H.261 and H.263, ISO / IEC developed the Moving Picture Experts Group (MPEG-1) and MPEG-4 Visual, and the two organizations jointly developed the H.262 / MPEG-2 Video and H.264 / MPEG-4 Advanced Video Coding (AVC) and H.265 / HEVC [1] standards. Since H.262, video coding standards have been based on a hybrid video coding structure that utilizes temporal prediction plus transform coding. To explore future video coding technologies beyond HEVC, the Video Coding Experts Group (VCEG) and MPEG jointly established the Joint Video Exploration Team (JVET). JVET adopted many methods and incorporated them into a reference software called the Joint Exploration Model (JEM) [2]. When the Versatile Video Coding (VVC) project was officially launched, the Joint Video Exploration Team (JVET) was renamed the Joint Video Experts Team (JVET). VVC is a coding standard that aims to reduce the bitrate by 50% compared to HEVC. The VVC working draft and the VVC Test Model (VTM) are continuously updated.
[0040] An example version of the VVC draft (i.e., Versatile Video Coding (Draft 10)) can be found at: https: / / jvet-experts.org / doc_end_user / documents / 19_Teleconference / wg11 / JVET-S2001-v17.zip. An example version of the reference software for VVC (called VTM) can be found at: https: / / vcgit.hhi.fraunhofer.de / jvet-u-ee2 / VVCSoftware_VTM / - / tree / VTM-11.2.
[0041] The Video Coding Experts Group (VCEG) of the International Telecommunication Union Telecommunication Standardization Sector (ITU-T) and the Joint Technical Committee (JTC) 1 / Subcommittee (SC) 29 / Working Group (WG) 11 of the International Organization for Standardization and the International Electrotechnical Commission (ISO / IEC) Moving Picture Experts Group (MPEG) are studying the potential need to standardize future video coding technologies that will significantly exceed the current VVC standard in terms of compression capabilities. Such future standardization actions could take the form of an extension of an extension of VVC or a completely new standard. These groups are jointly conducting this development activity in a joint collaborative effort called JVET to evaluate the compression technology designs proposed by their experts in this area. The first Exploration Experiment (EE) was established by the Joint Video Exploration Team (JVET), and a reference software called the Enhanced Compression Model (ECM) is in use. The test model ECM is continuously updated.
[0042] 3.1 Color Space and Chroma Subsampling
[0043] A color space (also known as a color model (or color system)) is a mathematical model that describes a range of colors as digital tuples, e.g., as 3 or 4 values or color components (e.g., RGB). Generally, a color space is a concretization of a coordinate system and a subspace. For video compression, the most commonly used color spaces are luminance, blue-difference chrominance, and red-difference chrominance (YCbCr) and red, green, and blue (RGB).
[0044] YCbCr, Y'CbCr, or YPb / CbPr / Cr (also written as YCBCR or Y'CBCR) is a collective term for a series of color spaces that are used in video systems and digital photography as part of the color image processing pipeline. Y' is the luminance component, and CB and CR are the blue-difference chrominance component and the red-difference chrominance component. Y' (with an apostrophe) is different from Y, which means that the light intensity is non-linearly encoded based on the gamma-corrected RGB primaries.
[0045] Chroma subsampling is the practice of encoding an image by achieving a lower resolution for chrominance information than for luminance information, taking advantage of the fact that the human visual system is less sensitive to color differences than to luminance. 3.1.1. 4:4:4
[0047] In 4:4:4, each of the three components of Y'CbCr has the same sampling rate. Thus, there is no chroma subsampling. This scheme is sometimes used in high-end film scanners and film post-production. 3.1.2. 4:2:2
[0049] In 4:2:2, the two chrominance components are sampled at half the sampling rate of luminance. The horizontal chrominance resolution is halved while the vertical chrominance resolution remains the same. This reduces the bandwidth of the uncompressed video signal by one-third with little visual difference. Figure 1 Examples of nominal vertical and horizontal positions in the 4:2:2 color format are shown. 3.1.3 4:2:0
[0051] In 4:2:0, compared to 4:1:1, the horizontal sampling is doubled. However, since the Cb and Cr channels are sampled only on alternate rows in this scheme, the vertical resolution is halved, and the data rate thus remains unchanged. Cb and Cr are each downsampled by a factor of 2 in both the horizontal and vertical directions. There are three variants of the 4:2:0 scheme, each with different horizontal and vertical sampling positions. In MPEG-2, Cb and Cr are horizontally co-located. Cb and Cr are located between pixels in the vertical direction (i.e., interleaved). In Joint Photographic Experts Group (JPEG) / PEG File Interchange Format (JFIF), H.261, and MPEG-1, Cb and Cr are arranged in an interleaved manner, at the midpoint between alternate luminance samples. In 4:2:0 DV, Cb and Cr are horizontally co-located. In the vertical direction, they are co-located on alternate rows.
[0052] chroma_format_idc separate_colour_plane_flag chroma format SubWidthC SubHeightC 0 0 monochrome 1 1 1 0 4:2:0 2 2 2 0 4:2:2 2 1 3 0 4:4:4 1 1 3 1 4:4:4 1 1
[0053] Table 1. SubWidthC and SubHeightC values derived from chroma_format_idc and Separate_colour_plane_flag
[0054] 3.2 Example encoding / decoding processes of video codecs
[0055] Figure 2 An example of the encoder block diagram of VVC is shown, which includes three loop filter modules: Deblocking Filter (DF), Sample Adaptive Offset (SAO), and Adaptive Loop Filter (ALF). Different from DF that uses predefined filters, SAO and ALF utilize the original samples of the current picture to reduce the mean square error between the original samples and the reconstructed samples by adding offset values respectively and by applying a Finite Impulse Response (FIR) filter. Encoding / decoding side information is used to signal this offset value and filter coefficients. ALF is at the last processing stage of each picture and can be regarded as a tool to try to capture and fix the artifacts generated by the previous stages.
[0056] 3.3 Definition of video / encoding / decoding units
[0057] A picture is divided into one or more slice rows and one or more slice columns. A slice is a sequence of Coding Tree Units (CTUs) that cover a rectangular area of the picture. A slice can be divided into one or more tiles, each tile including a certain number of CTU rows within the slice. A slice that is not divided into multiple tiles can also be called a tile. However, a tile that is a proper subset of a slice cannot be called a slice. A strip contains several slices of a picture or several tiles of a slice.
[0058] Two stripe modes are supported, namely, the raster scan stripe mode and the rectangular stripe mode. In the raster scan stripe mode, a stripe contains a series of slices in the slice raster scan of a picture. In the rectangular stripe mode, a stripe contains a certain number of tiles of a picture, and these tiles together form a rectangular area of the picture. The tiles within a rectangular stripe are arranged in the order of the tile raster scan of the stripe. Figure 3 An example of a picture divided into raster scan stripes is shown. The example includes a picture with 18×12 luminance CTUs, and the picture is divided into 12 slices and 3 raster scan stripes.
[0059] Figure 4 An example of the rectangular stripe segmentation of a picture with 18×12 luminance CTUs is shown, where the picture is divided into 24 slices (6 slice columns and 4 slice rows) and 9 rectangular stripes.
[0060] Figure 5 An example of a picture segmented into slices, tiles, and rectangular stripes is shown, where the picture is divided into 4 slices (2 slice columns and 2 slice rows), 11 tiles (the upper left slice contains 1 tile, the upper right slice contains 5 tiles, the lower left slice contains 2 tiles, and the lower right slice contains 3 tiles), and 4 rectangular stripes.
[0061] 3.3.1 CTU / CTB Sizes
[0062] In VVC, the CTU size (signaled in the sequence parameter set (SPS) by the syntax element log2_ctu_size_minus2) can be as small as 4×4.
[0063] 7.3.2.3 Sequence Parameter Set RBSP Syntax
[0064]
[0065]
[0066]
[0067] log2_ctu_size_minus2 plus 2 specifies the luma coding tree block size of each CTU. log2_min_luma_coding_block_size_minus2 plus 2 specifies the minimum luma coding block size. The variables CtbLog2SizeY, CtbSizeY, MinCbLog2SizeY, MinCbSizeY, MinTbLog2SizeY, MaxTbLog2SizeY, MinTbSizeY, MaxTbSizeY, PicWidthInCtbsY, PicHeightInCtbsY, PicSizeInCtbsY, PicWidthInMinCbsY, PicHeightInMinCbsY, PicSizeInMinCbsY, PicSizeInSamplesY, PicWidthInSamplesC, and PicHeightInSamplesC are derived as follows:
[0068] CtbLog2SizeY = log2_ctu_size_minus2 + 2 (7-9)
[0069] CtbSizeY = 1 << CtbLog2SizeY (7-10)
[0070] MinCbLog2SizeY = log2_min_luma_coding_block_size_minus2 + 2 (7-11)
[0071] MinCbSizeY = 1 << MinCbLog2SizeY (7-12)
[0072] MinTbLog2SizeY = 2 (7-13)
[0073] MaxTbLog2SizeY = 6 (7-14)
[0074] MinTbSizeY = 1 << MinTbLog2SizeY (7-15)
[0075] MaxTbSizeY = 1 << MaxTbLog2SizeY (7-16)
[0076] PicWidthInCtbsY = Ceil( pic_width_in_luma_samples ÷ CtbSizeY ) (7-17)
[0077] PicHeightInCtbsY = Ceil( pic_height_in_luma_samples ÷ CtbSizeY )(7-18)
[0078] PicSizeInCtbsY = PicWidthInCtbsY * PicHeightInCtbsY (7-19)
[0079] PicWidthInMinCbsY = pic_width_in_luma_samples / MinCbSizeY (7-20)
[0080] PicHeightInMinCbsY = pic_height_in_luma_samples / MinCbSizeY (7-21)
[0081] PicSizeInMinCbsY = PicWidthInMinCbsY * PicHeightInMinCbsY (7-22)
[0082] PicSizeInSamplesY=pic_width_in_luma_samples * pic_height_in_luma_samples (7-23)
[0083] PicWidthInSamplesC = pic_width_in_luma_samples / SubWidthC (7-24)
[0084] PicHeightInSamplesC = pic_height_in_luma_samples / SubHeightC (7-25)
[0085] 3.3.2 CTUs in a Picture
[0086] Figures 6A to 6C Examples of CTBs spanning picture boundaries are shown. Figure 6A A CTB spanning the bottom picture boundary is shown. Figure 6B A CTB spanning the right picture boundary is shown. Figure 6C A CTB spanning the bottom-right picture boundary is shown. Assume that the CTB / largest coding unit (LCU) size is represented as M×N (usually M equals N), and for a CTB located at a picture boundary (or slice or strip or other type of boundary, taking the picture boundary as an example), K×L samples are within the picture boundary, where K < M or L < N.Figures 6A to 6C An example of a CTB across a picture boundary is shown. For Figures 6A to 6C those CTBs shown in , the CTB size remains equal to M×N. However, the lower / right boundary of the CTB is outside the picture.
[0087] 3.4 Intra prediction
[0088] Figure 7 Examples of 67 intra prediction modes are shown. To capture any edge direction presented in natural videos, the number of directional intra modes is extended from 33 used in HEVC to 65. Figure 7 The extended directional modes are shown in , and the planar and direct current (DC) modes remain unchanged. These denser directional intra prediction modes apply to all block sizes as well as luminance and chrominance intra prediction.
[0089] The angular intra prediction direction can be defined as from 45 degrees to -135 degrees in the clockwise direction, as Figure 7 shown. In VTM, for non-square blocks, several angular intra prediction modes are adaptively replaced with wide-angle intra prediction modes. The replaced modes are signaled and remapped to the indices of the wide-angle modes after parsing. The total number of intra prediction modes remains unchanged, e.g., 67, and the intra mode coding and decoding remain unchanged.
[0090] In HEVC, each intra-coded block has a square shape, and the length of each side of the block is a power of 2. Therefore, no division operation is required to generate an intra predictor using the DC mode. In VVC, blocks can have a rectangular shape, and in general, a division operation must be used for each block. To avoid the division operation for DC prediction, only the longer side is used to calculate the average value of non-square blocks.
[0091] 3.5 Inter prediction
[0092] For each inter - predicted CU, the motion parameters include a motion vector, a reference picture index, a reference picture list use index, and extension information of the new coding - decoding features of VVC for sample generation for inter - prediction. The motion parameters can be signaled in an explicit or implicit manner. When a CU is coded or decoded using the skip mode, the CU is associated with a PU and there are no significant residual coefficients, no coded motion vector differences, and / or no reference picture indices. The Merge mode is defined as obtaining the motion parameters of the current CU from neighboring CUs, including spatial and temporal candidates and an extended schedule introduced in VVC. The Merge mode can be applied to any inter - predicted CU, not just to the skip mode. An alternative to the Merge mode is to explicitly send the motion parameters, where for each CU, the motion vector, the corresponding reference picture index for each reference picture list, the reference picture list use flag, and other useful information are signaled explicitly.
[0093] 3.6 Deblocking filter
[0094] Deblocking filtering is an example loop filter in a video codec. In VVC, the deblocking filtering process is applied to CU boundaries, transform sub - block boundaries, and prediction sub - block boundaries. Prediction sub - block boundaries include prediction unit boundaries introduced by sub - block - based temporal motion vector prediction (SbTMVP) and affine modes. Transform sub - block boundaries include transform unit boundaries introduced by sub - block transform (SBT) and intra - sub - partitioning (ISP) modes, as well as transforms due to implicit partitioning of large CUs. The processing order of the deblocking filter is defined as first horizontally filtering the vertical edges of the entire picture and then vertically filtering the horizontal edges. This specific order enables multiple horizontal or vertical filtering processes to be applied in parallel threads. The filtering process can also be implemented on a CTB - by - CTB basis with a very small processing delay.
[0095] First, the vertical edges in the picture are filtered. Then, the samples modified by the vertical - edge filtering process are used as input to filter the horizontal edges in the picture. The vertical and horizontal edges in the CTB of each CTU are processed separately according to the coding - decoding unit. Filtering of the vertical edges of the coding - decoding blocks in the coding - decoding unit starts from the edge on the left - hand side of the coding - decoding block and proceeds towards the edge on the right - hand side of the coding - decoding block in the geometric order of the coding - decoding blocks. Filtering of the horizontal edges of the coding - decoding blocks in the coding - decoding unit starts from the edge on the top - hand side of the coding - decoding block and proceeds towards the edge on the bottom - hand side of the coding - decoding block in the geometric order of the coding - decoding blocks.
[0096] Figure 8 An example of the block boundaries in the picture is shown. This includes picture samples on an 8×8 grid, as well as horizontal and vertical block boundaries, and an illustration of non - overlapping blocks of 8×8 samples that can be deblocked in parallel.
[0097] 3.6.1 Boundary Decision
[0098] Filtering is applied to the 8×8 block boundary. Additionally, such a boundary shall be a transform block boundary or a coding / decoding sub-block boundary, e.g., a transform block boundary or a coding / decoding sub-block boundary resulting from the use of affine motion vector prediction (ATMVP). For other boundaries, the deblocking filter is disabled.
[0099] 3.6.2 Boundary Strength Calculation
[0100] For a transform block boundary / coding / decoding sub-block boundary, if the boundary lies within the 8×8 grid, the boundary can be filtered, and the setting of bS[xDi][yDj] for this edge (where [xDi][yDj] represents coordinates) is defined in Tables 2 and 3 respectively.
[0101]
[0102] Table 2. Boundary Strength (when SPS Intra Block Copy (IBC) is disabled)
[0103]
[0104] Table 3. Boundary Strength (when SPS IBC is enabled)
[0105] 3.6.3 Deblocking Decision for Luma Component
[0106] Figure 9 Examples of pixel participation in filter use are shown. This includes pixel participation in filter on / off decision and strong / weak filter selection. A wider and stronger luma filter is used only when Conditions 1, 2, and 3 are all TRUE. Condition 1 is the "large block condition". This condition detects whether the samples on the P side and the Q side belong to large blocks, represented by the variables bSidePisLargeBlk and bSideQisLargeBlk respectively. bSidePisLargeBlk and bSideQisLargeBlk are defined as follows.
[0107] bSidePisLargeBlk = ((the edge type is vertical and p0 belongs to a CU with width >= 32) || (the edge type is horizontal and p0 belongs to a CU with height >= 32))? true : false
[0108] bSideQisLargeBlk = ((the edge type is vertical and q0 belongs to a CU with width >= 32) || (the edge type is horizontal and q0 belongs to a CU with height >= 32))? true : false
[0109] Based on bSidePisLargeBlk and bSideQisLargeBlk, Condition 1 is defined as follows:
[0110] Condition 1 = (bSidePisLargeBlk || bSidePisLargeBlk)? true : false
[0111] Next, if Condition 1 is true, then Condition 2 is further checked. First, the following variables are derived:
[0112] First, dp0, dp3, dq0, and dq3 are derived in the HEVC manner
[0113] if (the p side is greater than or equal to 32)
[0114] dp0 = (dp0 + Abs(p50 - 2 * p40 + p30) + 1) >> 1
[0115] dp3 = (dp3 + Abs(p53 - 2 * p43 + p33) + 1) >> 1
[0116] if (the q side is greater than or equal to 32)
[0117] dq0 = (dq0 + Abs(q50 - 2 * q40 + q30) + 1) >> 1
[0118] dq3 = (dq3 + Abs(q53 - 2 * q43 + q33) + 1) >> 1
[0119] Condition 2 = (d < β)? true : false
[0120] where d = dp0 + dq0 + dp3 + dq3.
[0121] If both Condition 1 and Condition 2 are satisfied, then it is further checked whether any block uses sub - blocks:
[0122]
[0123]
[0124] Finally, if both Condition 1 and Condition 2 are satisfied, the de - block method will check Condition 3 (the large - block strong filtering condition), which is defined as follows. In Condition 3 StrongFilterCondition, the following variables are derived:
[0125] dpq is derived in the HEVC manner
[0126] sp3 = Abs(p3 - p0) is derived in the HEVC manner
[0127] if (the p side is greater than or equal to 32)
[0128] if(Sp == 5)
[0129] sp3 = (sp3 + Abs(p5 - p3) + 1) >> 1
[0130] else
[0131] sp3 = (sp3 + Abs(p7 - p3) + 1) >> 1
[0132] Derive sq3 = Abs(q0 - q3) in the HEVC way
[0133] if (q side is greater than or equal to 32)
[0134] If (Sq == 5)
[0135] sq3 = (sq3 + Abs(q5 - q3) + 1) >> 1
[0136] else
[0137] sq3 = (sq3 + Abs(q7 - q3) + 1) >> 1
[0138] According to HEVC, StrongFilterCondition = (dpq is less than (β >> 2), sp3 + sq3 is less than (3 * β >> 5), and Abs(p0 - q0) is less than (5 * tC + 1) >> 1)? true : false.
[0139] 3.6.4 Stronger Deblocking Filter for Luminance
[0140] When the samples on either side of the boundary belong to a large block, a bilinear filter is used. When the width of the vertical edge >= 32 and when the height of the horizontal edge >= 32, the samples are defined as belonging to a large block. The bilinear filter is listed as follows. Block boundary samples pi (i = 0 to Sp - 1) and qi (i = 0 to Sq - 1), in the above HEVC deblocking, pi and qi are the i-th samples in the row for filtering the vertical edge, or the i-th samples in the column for filtering the horizontal edge, and then are replaced by linear interpolation, as follows:
[0141] p i ' = (f i * Middle s,t + (64 - f i ) * P s + 32) >> 6), clipped to p i ± tcPD i
[0142] q j ' = (gj *Middle s,t +(64 - g j ) * Q s + 32) >> 6), clipped to q j ±tcPD j
[0143] where tcPD i and tcPD j terms are the position - related clipping described above, and g j 、f i 、Middle s,t 、P s and Q s are given as follows.
[0144] 3.6.5 Chroma Deblocking Decision
[0145] Chroma strong filters are used on both sides of the block boundary. Here, when both sides of the chroma edge are greater than or equal to 8 (chroma position), the chroma filter is selected, and the decision satisfies the following three conditions: The first decision is the boundary strength and the decision for large blocks. When the block width or height orthogonal to the block edge is equal to or greater than 8 in the chroma sample domain, this filter can be applied. The second and third decisions are basically the same as the HEVC luma deblocking decision, which are the on - off decision and the strong filter decision respectively.
[0146] In the first decision, for chroma filtering, the boundary strength (bS) is modified, and the conditions are checked sequentially. If the condition is satisfied, the remaining conditions with lower priority are skipped. Chroma deblocking is performed when bS is equal to 2, or bS is equal to 1 when a large block boundary is detected. The second and third conditions are basically the same as the HEVC luma strong filter decision, as follows.
[0147] In the second condition, d is derived according to the HEVC luma deblocking. When d is less than β, the second condition will be true. In the third condition, StrongFilterCondition is derived as follows:
[0148] dpq is derived in the HEVC manner.
[0149] sp3 = Abs(p3 - p0) is derived in the HEVC manner
[0150] sq3 = Abs(q0 - q3) is derived in the HEVC manner
[0151] According to the HEVC design, StrongFilterCondition = (dpq is less than (β >> 2), sp3 + sq3 is less than (β >> 3), and Abs(p0 - q0) is less than (5 * tC + 1) >> 1).
[0152] 3.6.6 Strong Deblocking Filter for Chrominance
[0153] The strong deblocking filter for chrominance is defined as follows:
[0154] p2′ = (3 * p3 + 2 * p2 + p1 + p0 + q0 + 4) >> 3
[0155] p1′ = (2 * p3 + p2 + 2 * p1 + p0 + q0 + q1 + 4) >> 3
[0156] p0′ = (p3 + p2 + p1 + 2 * p0 + q0 + q1 + q2 + 4) >> 3
[0157] An example chrominance filter performs deblocking on a 4×4 chrominance sample grid.
[0158] 3.6.7 Position-Dependent Clipping
[0159] The position-dependent clipping tcPD is applied to the output samples of the luminance filtering process, involving strong and long filters that modify 7, 5, and 3 samples at the boundaries. Assuming the quantization error distribution, the clipping values of the samples can be increased, which are expected to have higher quantization noise and thus a larger deviation between the expected reconstructed sample values and the true sample values.
[0160] For each P or Q boundary filtered with an asymmetric filter, according to the result of the decision-making process, a position-dependent threshold table is selected from two tables (e.g., Tc7 and Tc3 listed below), which are provided to the decoder as side information:
[0161] Tc7 = {6, 5, 4, 3, 2, 1, 1}; Tc3 = {6, 4, 2};
[0162] tcPD = (Sp == 3)? Tc3 : Tc7;
[0163] tcQD = (Sq == 3)? Tc3 : Tc7;
[0164] For a P or Q boundary filtered with a short symmetric filter, a lower magnitude of position-dependent threshold is applied:
[0165] Tc3 = {3, 2, 1};
[0166] After defining the thresholds, the filtered p'i and q'i sample values are clipped according to the tcP and tcQ clipping values:
[0167] p”i = Clip3(p’i + tcPi, p’i – tcPi, p’i);
[0168] q”j = Clip3(q’j + tcQj, q’j – tcQj, q’j);
[0169] where p'i and q'j are the filtered sample values, p”i and q”j are the output sample values after clipping, and tcPi is the clipping threshold derived from the VVC tc parameter and tcPD and tcQD. The function Clip3 is the clipping function as defined in VVC.
[0170] 3.6.8. Sub - block Deblocking Adjustment
[0171] To perform parallel - friendly deblocking using both the long filter and sub - block deblocking, the long filter is restricted to modifying at most 5 samples on the side where sub - block deblocking (AFFINE or ATMVP or decoder - side motion vector refinement (DMVR)) is used, as shown in the luminance control of the long filter. Extending this, the sub - block deblocking is adjusted such that the sub - block boundaries on the 8×8 grid near the CU or implicit TU boundary are restricted to modifying at most 2 samples on each side.
[0172] The following applies to sub - block boundaries that are not aligned with the CU boundary.
[0173]
[0174] where an edge equal to 0 corresponds to the CU boundary, an edge equal to 2 or equal to orthogonalLength - 2 corresponds to the sub - block boundary 8 samples away from the CU boundary, etc. If the implicit partitioning of the TU is used, implicitTU is true.
[0175] 3.7. Sample Adaptive Compensation
[0176] Sample Adaptive Offset (SAO) is applied to the reconstructed signal after deblocking filtering by using the offset values specified for each Coding Tree Block (CTB) by the encoder. The video encoder first decides whether to apply SAO processing to the current slice. If SAO is applied to the slice, each CTB will be classified into one of five SAO types as shown in Table 4. The concept of SAO is to classify pixels into multiple categories and reduce distortion by adding offset values to pixels in each category. SAO operations include: Edge Offset (EO), which classifies pixels using edge attributes in SAO types 1 to 4; and Band Offset (BO), which classifies pixels using pixel intensity in SAO type 5. Each applicable CTB has SAO parameters, including sao_merge_left_flag, sao_merge_up_flag, SAO type, and four offset values. If sao_merge_left_flag is equal to 1, the current CTB will reuse the SAO type and offset values of the left CTB. If sao_merge_up_flag is equal to 1, the current CTB will reuse the SAO type and offset values of the upper CTB.
[0177]
[0178]
[0179] Table 4. Specification of SAO Types
[0180] 3.8. Adaptive Loop Filter
[0181] Adaptive loop filtering for video coding and decoding minimizes the mean square error between the original samples and the decoded samples by using a Wiener-based adaptive filter. ALF is at the final processing stage of each picture and can be regarded as a tool to capture and fix artifacts from previous stages. Appropriate filter coefficients are determined by the encoder and signaled explicitly to the decoder. To achieve better coding and decoding efficiency, especially for high-resolution video, local adaptivity is used for the luminance signal by applying different filters to different regions or blocks in the picture. In addition to filter adaptivity, filter on / off control at the Coding Tree Unit (CTU) level also helps to improve coding and decoding efficiency. Syntactically, the filter coefficients are sent in picture-level header information called the Adaptive Parameter Set, and the filter on / off flag for a CTU is interleaved at the CTU level in the slice data. This syntax design not only supports picture-level optimization but also achieves a lower coding delay.
[0182] 3.8.1. Signaling of Parameters
[0183] According to the ALF design in VTM, filter coefficients and clipping indices are carried in the ALF APS. The ALF APS can include up to 8 chroma filters and one luma filter set, for a total of up to 25 filters. Each of the 25 luma classes also includes an index. Classes with the same index share the same filter. By combining different classes, the number of bits required to represent the filter coefficients is reduced. The absolute value of the filter coefficient is represented using a 0th order Exp-Golomb code followed by the sign bit of the non-zero coefficient. When clipping is enabled, a two-bit fixed-length code is also used to signal the clipping index for each filter coefficient. The decoder can use up to 8 ALF APSs simultaneously.
[0184] The filter control syntax elements for ALF in VTM include two types of information. First, the ALF on / off flag is signaled at the sequence, picture, slice, and CTB levels. The chroma ALF can be enabled at the picture and slice levels only if the luma ALF is enabled at the corresponding level. Second, if ALF is enabled at the picture, slice, and CTB levels, the filter usage information is signaled at that level. If all slices within a picture use the same APS, the referenced ALF APS ID is coded / decoded at the slice level or picture level. The luma component can reference up to 7 ALF APSs, while the chroma component can reference up to 1 ALF APS. For a luma CTB, an index is signaled to indicate which ALF APS or offline-trained luma filter set is used. For a chroma CTB, the index indicates which filter in the referenced APS is used.
[0185] The ALF data syntax elements associated with the luma component in VTM are listed below:
[0186]
[0187] When alf_luma_filter_signal_flag equals 1, it specifies that the luma filter set is signaled. When alf_luma_filter_signal_flag equals 0, it specifies that the luma filter set is not signaled. When alf_luma_clip_flag equals 0, it specifies that linear adaptive loop filtering is applied to the luma component. When alf_luma_clip_flag equals 1, it specifies that non-linear adaptive loop filtering can be applied to the luma component. alf_luma_num_filters_signalled_minus1 plus 1 specifies the number of classes of adaptive loop filters for luma coefficients that can be signaled. The value of alf_luma_num_filters_signalled_minus1 shall be in the range of 0 to NumAlfFilters - 1 (including the endpoints). alf_luma_coeff_delta_idx[filtIdx] indicates the index of the signaled adaptive loop filter luma coefficient increment for the filter class indicated by filtIdx, where filtIdx ranges from 0 to NumAlfFilters - 1. When alf_luma_coeff_delta_idx[filtIdx] does not exist, it is inferred to be equal to 0. The length of alf_luma_coeff_delta_idx[filtIdx] is Ceil(Log2(alf_luma_num_filters_signalled_minus1 + 1)) bits. The value of alf_luma_coeff_delta_idx[filtIdx] shall be in the range of 0 to alf_luma_num_filters_signalled_minus1 (including the endpoints).
[0188] alf_luma_coeff_abs[sfIdx][j] specifies the absolute value of the j-th coefficient of the signaled luma filter indicated by sfIdx. When alf_luma_coeff_abs[sfIdx][j] does not exist, it is inferred to be equal to 0. The value of alf_luma_coeff_abs[sfIdx][j] shall be in the range of 0 to 128 (including the endpoints). alf_luma_coeff_sign[sfIdx][j] specifies the sign of the j-th luma coefficient of the filter indicated by sfIdx, as follows:
[0189] If alf_luma_coeff_sign[sfIdx][j] equals 0, the corresponding luma filter coefficient has a positive value.
[0190] Otherwise (alf_luma_coeff_sign[sfIdx][j] equals 1), the corresponding luma filter coefficient has a negative value.
[0191] When alf_luma_coeff_sign[sfIdx][j] does not exist, it is inferred to be equal to 0.
[0192] alf_luma_clip_idx[sfIdx][j] specifies the clipping index of the clipping value to be used before multiplying the j-th coefficient of the luma filter transmitted through the signal indicated by sfIdx. When alf_luma_clip_idx[sfIdx][j] does not exist, it is inferred to be equal to 0. The codec tree syntax elements associated with the luma component in VTM are listed below:
[0193]
[0194] alf_ctb_flag[cIdx][xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] being equal to 1 specifies that the adaptive loop filter is applied to the codec tree block of the color component indicated by cIdx of the codec tree unit at the luma position (xCtb, yCtb). alf_ctb_flag[cIdx][xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] being equal to 0 specifies that the adaptive loop filter is not applied to the codec tree block of the color component indicated by cIdx of the codec tree unit at the luma position (xCtb, yCtb).
[0195] When alf_ctb_flag[cIdx][xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] does not exist, it is inferred to be equal to 0. alf_use_aps_flag being equal to 0 specifies that one of the fixed filter sets is applied to the luma CTB. alf_use_aps_flag being equal to 1 specifies that the filter set from APS is applied to the luma CTB. When alf_use_aps_flag does not exist, it is inferred to be equal to 0. alf_luma_prev_filter_idx specifies the previous filter applied to the luma CTB. The value of alf_luma_prev_filter_idx should be in the range of 0 to sh_num_alf_aps_ids_luma - 1 (including the endpoints). When alf_luma_prev_filter_idx does not exist, it is inferred to be equal to 0.
[0196] The variable AlfCtbFiltSetIdxY[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY], which specifies the filter set index of the luma CTB at the position (xCtb, yCtb), is derived as follows:
[0197] If alf_use_aps_flag is equal to 0, then AlfCtbFiltSetIdxY[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] is set to be equal to alf_luma_fixed_filter_idx.
[0198] Otherwise, AlfCtbFiltSetIdxY[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] is set to be equal to 16 + alf_luma_prev_filter_idx.
[0199] alf_luma_fixed_filter_idx specifies the fixed filter applied to the luma CTB. The value of alf_luma_fixed_filter_idx shall be in the range of 0 to 15 (including the endpoints).
[0200] Based on the ALF design of VTM, the ALF design of ECM further introduces the concept of alternative filter sets into the luma filter. Based on the updated luma CTU ALF on / off decision for each alternative / round, multiple alternatives / rounds of training are performed on the luma filter. In this way, multiple filter sets are associated with each training alternative, and the class merging results of each filter set may be different. Each CTU can select the best filter set by RDO and signal the relevant alternative information. The data syntax elements of ALF associated with the luma component in ECM are listed as follows:
[0201]
[0202]
[0203] alf_luma_num_alts_minus1 plus 1 specifies the number of alternative filter sets for the luma component. The value of alf_luma_num_alts_minus1 shall be in the range of 0 to 3 (inclusive of the endpoints). alf_luma_clip_flag[altIdx] being equal to 0 specifies that the linear adaptive loop filter is applied to the alternative luma filter set for the luma component with index altIdx. alf_luma_clip_flag[altIdx] being equal to 1 specifies that the non-linear adaptive loop filter may be applied to the alternative luma filter set for the luma component with index altIdx. alf_luma_num_filters_signalled_minus1[altIdx] plus 1 specifies the number of adaptive loop filter classes for which the luma coefficients can be signalled to the alternative luma filter set with index altIdx. The value of alf_luma_num_filters_signalled_minus1[altIdx] shall be in the range of 0 to NumAlfFilters-1 (inclusive of the endpoints).
[0204] alf_luma_coeff_delta_idx[altIdx][filtIdx] specifies the index of the adaptive loop filter luminance coefficient increment for the filter class signalled through the signal for the alternative luminance filter set with index altIdx, represented by filtIdx, where filtIdx ranges from 0 to NumAlfFilters–1. When alf_luma_coeff_delta_idx[filtIdx][altIdx] does not exist, it is inferred to be equal to 0. The length of alf_luma_coeff_delta_idx[altIdx][filtIdx] is Ceil(Log2(alf_luma_num_filters_signalled_minus1[altIdx]+1)) bits. The value of alf_luma_coeff_delta_idx[altIdx][filtIdx] shall be in the range of 0 to alf_luma_num_filters_signalled_minus1[altIdx] (including the endpoints). alf_luma_coeff_abs[altIdx][sfIdx][j] specifies the absolute value of the j-th coefficient of the luminance filter signalled through the signal indicated by sfIdx of the alternative luminance filter set with index altIdx. When alf_luma_coeff_abs[altIdx][sfIdx][j] does not exist, it is inferred to be equal to 0. The value of alf_luma_coeff_abs[altIdx][sfIdx][j] shall be in the range of 0 to 128 (including the endpoints).
[0205] alf_luma_coeff_sign[altIdx][sfIdx][j] specifies the sign of the j-th luminance coefficient of the filter indicated by sfIdx of the alternative luminance filter set with index altIdx, as follows:
[0206] If alf_luma_coeff_sign[altIdx][sfIdx][j] is equal to 0, the corresponding luminance filter coefficient has a positive value.
[0207] Otherwise (alf_luma_coeff_sign[altIdx][sfIdx][j] is equal to 1), the corresponding luminance filter coefficient has a negative value.
[0208] When alf_luma_coeff_sign[altIdx][sfIdx][j] does not exist, it is inferred to be equal to 0.
[0209] alf_luma_clip_idx[altIdx][sfIdx][j] specifies the clipping index of the clipping value to be used before multiplying the j-th coefficient of the luma filter transmitted through the signal represented by sfIdx of the alternative luma filter set with index altIdx. When alf_luma_clip_idx[altIdx][sfIdx][j] does not exist, it is inferred to be equal to 0. The codec tree syntax elements associated with the luma component in the ECM are listed as follows:
[0210]
[0211] alf_ctb_luma_filter_alt_idx[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] specifies the index of the alternative luma filter of the codec tree block for the luma component of the codec tree unit at the luma position (xCtb, yCtb). When alf_ctb_luma_filter_alt_idx[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] does not exist, it is inferred to be equal to zero.
[0212] 3.8.2. Filter Shape
[0213] Figure 10 Examples of the filter shapes of the ALF are shown. In JEM, up to three diamond filter shapes can be selected for the luma component (as Figure 10 shown). The filter shape used for the luma component is indicated by transmitting an index at the picture level. Each square represents a sample point, and Ci (i is 0 to 6 (left), 0 to 12 (middle), 0 to 20 (right)) represents the coefficient to be applied to that sample point. For the chroma component in the picture, a 5×5 diamond shape is always used. In VVC, a 7×7 diamond shape is always used for luma, and a 5×5 diamond shape is always used for chroma.
[0214] 3.8.3 Classification of ALF
[0215] Each 2×2 (or 4×4) block is classified as one of 25 classes. The classification index C is derived based on its directionality D and the quantized value of activity as follows:
[0216]
[0217] To calculate D and First, the gradients in the horizontal, vertical, and two diagonal directions are calculated using a 1D Laplacian operator:
[0218]
[0219] The indices i and j refer to the coordinates of the upper-left sample point in the 2×2 block, and R(i,j) indicates the reconstructed sample point at the coordinates (i,j). The maximum and minimum values of D for the gradients in the horizontal and vertical directions are set to:
[0220]
[0221] And the maximum and minimum values of the gradients in the two diagonal directions are set to:
[0222]
[0223] To derive the value of the directionality D, these values are compared with each other and with two thresholds t1 and t2:
[0224] Step 1. If and are both true, then set D to 0.
[0225] Step 2. If then continue from Step 3; otherwise continue from Step 4;
[0226] Step 3. If then set D to 2; otherwise, set D to 1.
[0227] Step 4. If then set D to 4; otherwise, set D to 3.
[0228] The activity value A is calculated as:
[0229]
[0230] A is further quantized to the range from 0 to 4 (including the endpoints), and the quantized value is denoted as For the two chrominance components in the picture, no classification method is adopted (i.e., a set of ALF coefficients is applied to each chrominance component).
[0231] 3.8.4. Geometric Transformations of Filter Coefficients
[0232] Before filtering each 2×2 block, geometric transformations such as rotation or diagonal and vertical flipping are applied to the filter coefficients f(k,l) associated with the coordinates (k,l) according to the gradient values calculated for the block. This is equivalent to applying these transformations to the sample points in the filter support region. The idea is to make the different blocks more similar by aligning the directionality of the ALF-applied blocks.
[0233] Three geometric transformations are introduced, including diagonal transformation, vertical flipping, and rotation:
[0234] Diagonal: f D (k, l) = f(l, k),
[0235] Vertical flip: f V (k, l) = f(k, K - l - 1),
[0236] Rotation: f R (k, l) = f(K - l - 1, k).
[0237] Where K is the size of the filter, and 0 ≤ k, l ≤ K - 1 are the coefficient coordinates, such that the position (0, 0) is at the upper left corner and the position (K - 1, K - 1) is at the lower right corner. The transformation is applied to the filter coefficient f(k, l) according to the gradient value calculated for the block. Table 5 summarizes the relationship between the transformation and the four gradients in four directions. Figure 11 Shows the transformation coefficients for each position based on a 5×5 rhombus. This includes the relative coordinates supported by the 5×5 rhombus filter.
[0238] gradient value transformation <![CDATA[g d2 <g d1 and g h <g v > no transformation <![CDATA[g d2 <g d1 and g v <g h > diagonal <![CDATA[g d1 <g d2 and g h <g v > vertical flip <![CDATA[g d1 <g d2 and g v <g h > rotation
[0239] Table 5. Mapping of gradients and transformations calculated for a block.
[0240] 3.8.5. Filtering process
[0241] On the decoder side, when ALF is enabled for a block, each sample point R(i, j) within the block is filtered to obtain the sample point value R′(i, j) as shown below, where L represents the filter length, f m,n represents the filter coefficient, and f(k, l) represents the decoded filter coefficient.
[0242]
[0243] Figure 12 Shows an example of the relative coordinates for a 5×5 rhombus filter support, assuming the coordinates (i, j) of the current sample point are (0, 0). The sample points in different coordinates filled with the same color are multiplied by the same filter coefficient.
[0244] 3.8.6. Non - linear filtering reformulation
[0245] Linear filtering can be reformulated into the following expression without affecting the coding and decoding efficiency:
[0246]
[0247] where w(i, j) are the same filter coefficients.
[0248] VVC introduces non - linearity to make the ALF more efficient by reducing the impact of these neighboring sample values when the neighboring sample values (I(x + i, y + j)) differ too much from the filtered current sample value (I(x, y)) through the use of a simple clipping function. More specifically, the ALF filter is modified as follows:
[0249]
[0250] where K(d, b)=min(b, max( - b, d)) is the clipping function, and k(i, j) is the clipping parameter, which depends on the (i, j) filter coefficients. The encoder performs optimization to find the best k(i, j).
[0251] For each ALF filter, the clipping parameter k(i, j) is specified, and each filter coefficient signals a clipping value. This means that at most 12 clipping values can be signaled in the bit - stream for each luminance filter, and at most 6 clipping values can be signaled in the bit - stream for each chrominance filter. To limit the signaling cost and encoder complexity, only 4 fixed values are used, and these fixed values are the same for inter - frame and intra - frame stripes.
[0252] Since the variance of local differences in luminance is usually higher than that in chrominance, two different sets are applied for the luminance filter and the chrominance filter. The maximum sample value in each set (here 1024 for a 10 - bit bit - depth) is also introduced so that clipping can be disabled when not needed. These 4 values are selected by roughly evenly dividing the entire range of sample values of luminance (encoded and decoded in 10 bits) in the logarithmic domain and the range of 4 to 1024 for chrominance. More precisely, the luminance table of clipping values is obtained by the following formula:
[0253] where M = 2 10 and N = 4
[0254] Similarly, the chrominance table of clipping values can be obtained according to the following formula:
[0255] where M = 2 10 , N = 4 and A = 4
[0256] 3.9. Bilateral Loop Filter
[0257] 3.9.1. Bilateral Image Filter
[0258] The bilateral image filter is a non - linear filter that smooths noise while preserving the edge structure. Bilateral filtering is a technique that makes the filter weights decrease not only with the distance between samples but also with the increase in intensity difference. In this way, over - smoothing of edges can be improved. The weights are defined as
[0259]
[0260] where Δx and Δy are the distances in the vertical and horizontal directions respectively, and ΔI is the intensity difference between the samples.
[0261] The edge-preserving denoising bilateral filter uses low-pass Gaussian filters for both the domain filter and the range filter. The domain low-pass Gaussian filter assigns higher weights to pixels spatially closer to the central pixel. The range low-pass Gaussian filter assigns higher weights to pixels similar to the central pixel. Combining the range filter and the domain filter, the bilateral filter at the edge pixels becomes an elongated Gaussian filter that aligns along the edge orientation and is significantly reduced in the gradient direction. This is why the bilateral filter can smooth the noise while preserving the edge structure.
[0262] 3.9.2. Bilateral Filter in Video Coding and Decoding
[0263] The bilateral filter in video coding and decoding is a coding and decoding tool for VVC[2]. The filter acts as a loop filter in parallel with the sample adaptive offset (SAO) filter. Both the bilateral filter and SAO act on the same input samples, and each filter produces an offset value, which is then added to the input samples to produce the output samples, which enter the next stage after clipping. The spatial filtering strength σ d is determined by the block size, where the smaller the block, the greater the filtering strength, and the intensity filtering strength σ r is determined by the quantization parameter, where stronger filtering is used for higher QPs. Only the four closest samples are used, so the intensity I of the filtered samples F can be calculated as
[0264]
[0265] where I C represents the intensity of the central sample, ΔI A = I A - I C represents the intensity difference between the central sample and the upper sample. ΔI B , ΔI L and ΔI R represent the intensity differences between the central sample and the lower, left, and right samples respectively.
[0266] 4. Technical Problems Solved by the Disclosed Technical Solutions
[0267] The example design of the bilateral filter (BF) in video coding and decoding has the following problems:
[0268] In an example bilateral filter design, only the reconstruction before the current stage is used. However, there is other valuable side information that can potentially be exploited, such as samples before a deblocking filter (DBF), sample adaptive offset filter (SAO), and / or adaptive loop filter (ALF), prediction samples, residual information, segmentation information, QP information, etc.
[0269] 5. List of solutions and embodiments
[0270] To solve the above problems, the methods summarized below are disclosed. The embodiments should be regarded as examples for explaining general concepts and should not be interpreted narrowly. In addition, these embodiments can be applied individually or combined in any way.
[0271] It should be noted that the disclosed methods can be used as loop filters or post-processing.
[0272] In the present disclosure, a processing unit may refer to a sequence, picture, sub-picture, slice, CTU, block, and / or region. The processing unit may include one color component or multiple color components.
[0273] In the present disclosure, side information may refer to any codec information not input into SAO / ALF / cross-component ALF (CCALF) in VVC. For example, side information may include codec information other than the samples to be filtered by the current filter, and such side information can be used to determine the usage of the current filter for the samples to be filtered.
[0274] Example 1
[0275] In one example, side information is used for a bilateral filter (BF).
[0276] Example 2
[0277] In one example, the reconstruction before DBF is used as side information for BF.
[0278] In one example, the luminance reconstruction before DBF can be used to generate the luminance offset value for BF. In one example, the luminance reconstruction before DBF can be used to generate the chrominance offset value for BF.
[0279] In one example, the luminance reconstruction before DBF can be used for the classification of luminance samples in BF. In one example, the classification can be based on a texture-based method. In one example, the classification can be based on a band-based method.
[0280] In one example, the luminance reconstruction before DBF can be used for the classification of chrominance samples in BF. In one example, the classification can be based on a texture-based method. In one example, the classification can be based on a band-based method.
[0281] In one example, chrominance reconstruction prior to DBF can be used to generate the luminance offset value of BF. In one example, chrominance reconstruction prior to DBF can be used to generate the chrominance offset value of BF.
[0282] In one example, chrominance reconstruction prior to DBF can be used for the classification of luminance samples in BF. In one example, the classification can be based on a texture-based method. In one example, the classification can be based on a band-based method.
[0283] In one example, chrominance reconstruction prior to DBF can be used for the classification of chrominance samples in BF. In one example, the classification can be based on a texture-based method. In one example, the classification can be based on a band-based method.
[0284] Example 3
[0285] In one example, the predicted samples are used as side information for BF.
[0286] In one example, the luminance predicted samples can be used to generate the luminance offset value of BF. In one example, the luminance predicted samples can be used to generate the chrominance offset value of BF.
[0287] In one example, the luminance predicted samples can be used for the classification of luminance samples in BF. In one example, the classification can be based on a texture-based method. In one example, the classification can be based on a band-based method.
[0288] In one example, the luminance predicted samples can be used for the classification of chrominance samples in BF. In one example, the classification can be based on a texture-based method. In one example, the classification can be based on a band-based method.
[0289] In one example, the chrominance predicted samples can be used to generate the luminance offset value of BF. In one example, the chrominance predicted samples can be used to generate the chrominance offset value of BF.
[0290] In one example, the chrominance predicted samples can be used for the classification of luminance samples in BF. In one example, the classification can be based on a texture-based method. In one example, the classification can be based on a band-based method.
[0291] In one example, the chrominance predicted samples can be used for the classification of chrominance samples in BF. In one example, the classification can be based on a texture-based method. In one example, the classification can be based on a band-based method.
[0292] The predicted samples can be modified before being put into BF. The modification can be filtering. The modification can be downsampling / upsampling. The modification can be clipping / shifting.
[0293] Example 4
[0294] In one example, the residual value can be used as side information for BF.
[0295] In one example, the luminance residual value can be used to generate a luminance offset value for BF. In one example, the luminance residual value can be used to generate a chrominance offset value for BF.
[0296] In one example, the luminance residual value can be used for the classification of luminance samples in BF. In one example, the classification can be based on a texture-based method. In one example, the classification can be based on a band-based method.
[0297] In one example, the luminance residual value can be used for the classification of chrominance samples in BF. In one example, the classification can be based on a texture-based method. In one example, the classification can be based on a band-based method.
[0298] In one example, the chrominance residual value can be used to generate a luminance offset value for BF. In one example, the chrominance residual value can be used to generate a chrominance offset value for BF.
[0299] In one example, the chrominance residual value can be used for the classification of luminance samples in BF. In one example, the classification can be based on a texture-based method. In one example, the classification can be based on a band-based method.
[0300] In one example, the chrominance residual value can be used for the classification of chrominance samples in BF. In one example, the classification can be based on a texture-based method. In one example, the classification can be based on a band-based method.
[0301] The residual value can be modified before being put into BF. The modification can be filtering. The modification can be downsampling / upsampling. The modification can be clipping / shifting.
[0302] Example 5
[0303] In one example, the segmentation information can be used as side information for BF.
[0304] In one example, the segmentation information can represent block size / shape / position or other information. In one example, the luminance segmentation information can be used to generate a luminance offset value for BF. In one example, the luminance segmentation information can be used to generate a chrominance offset value for BF.
[0305] In one example, the luminance segmentation information can be used for the classification of luminance samples in BF. In one example, the classification can be based on a texture-based method. In one example, the classification can be based on a band-based method.
[0306] In one example, the luminance segmentation information can be used for the classification of chrominance samples in the BF. In one example, the classification can be based on an edge-based method. In one example, the classification can be based on a band-based method.
[0307] In one example, the chrominance segmentation information can be used to generate the luminance offset value of the BF. In one example, the chrominance segmentation information can be used to generate the chrominance offset value of the BF.
[0308] In one example, the chrominance segmentation information can be used for the classification of luminance samples in the BF. In one example, the classification can be based on a texture-based method. In one example, the classification can be based on a band-based method.
[0309] In one example, the chrominance segmentation information can be used for the classification of chrominance samples in the BF. In one example, the classification can be based on a texture-based method. In one example, the classification can be based on a band-based method.
[0310] Example 6
[0311] In one example, the QP information can be used as the edge information of the BF.
[0312] In one example, the QP information can represent the picture / strip / block QP information. In one example, the luminance QP information can be used to generate the luminance offset value of the BF. In one example, the luminance QP information can be used to generate the chrominance offset value of the BF.
[0313] In one example, the luminance QP information can be used for the classification of luminance samples in the BF. In one example, the classification can be based on a texture-based method. In one example, the classification can be based on a band-based method.
[0314] In one example, the luminance QP information can be used for the classification of chrominance samples in the BF. In one example, the classification can be based on a texture-based method. In one example, the classification can be based on a band-based method.
[0315] In one example, the chrominance QP information can be used to generate the luminance offset value of the BF. In one example, the chrominance QP information can be used to generate the chrominance offset value of the BF.
[0316] In one example, the chrominance QP information can be used for the classification of luminance samples in the BF. In one example, the classification can be based on a texture-based method. In one example, the classification can be based on a band-based method.
[0317] In one example, the chrominance QP information can be used for the classification of chrominance samples in the BF. In one example, the classification can be based on a texture-based method. In one example, the classification can be based on a band-based method.
[0318] Example 7
[0319] In one example, the boundary strength information is used as the side information for the BF.
[0320] In one example, the boundary strength information can be generated by DBF or other methods. In one example, the luminance boundary strength information can be used to generate the luminance offset value of the BF. In one example, the luminance boundary strength information can be used to generate the chrominance offset value of the BF.
[0321] In one example, the luminance boundary strength information can be used for the classification of luminance samples in the BF. In one example, the classification can be based on a texture-based method. In one example, the classification can be based on a band-based method.
[0322] In one example, the luminance boundary strength information can be used for the classification of chrominance samples in the BF. In one example, the classification can be based on a texture-based method. In one example, the classification can be based on a band-based method.
[0323] In one example, the chrominance boundary strength information can be used to generate the luminance offset value of the BF. In one example, the chrominance boundary strength information can be used to generate the chrominance offset value of the BF.
[0324] In one example, the chrominance boundary strength information can be used for the classification of luminance samples in the BF. In one example, the classification can be based on a texture-based method. In one example, the classification can be based on a band-based method.
[0325] In one example, the chrominance boundary strength information can be used for the classification of chrominance samples in the BF. In one example, the classification can be based on a texture-based method. In one example, the classification can be based on a band-based method.
[0326] Example 8
[0327] In one example, the side information can be used for the Hadamard transform domain filter (HTDF).
[0328] Example 9
[0329] In one example, the reconstruction before DBF can be used as the side information for the HTDF.
[0330] In one example, the luminance reconstruction before DBF can be used to generate the luminance offset value of the HTDF. In one example, the luminance reconstruction before DBF can be used to generate the chrominance offset value of the HTDF.
[0331] In one example, the luminance reconstruction before DBF can be used for the classification of luminance samples in HTDF. In one example, the classification can be based on a texture-based method. In one example, the classification can be based on a band-based method.
[0332] In one example, the luminance reconstruction before DBF can be used for the classification of chrominance samples in HTDF. In one example, the classification can be based on a texture-based method. In one example, the classification can be based on a band-based method.
[0333] In one example, the chrominance reconstruction before DBF can be used to generate the luminance offset value of HTDF. In one example, the chrominance reconstruction before DBF can be used to generate the chrominance offset value of HTDF.
[0334] In one example, the chrominance reconstruction before DBF can be used for the classification of luminance samples in HTDF. In one example, the classification can be based on a texture-based method. In one example, the classification can be based on a band-based method.
[0335] In one example, the chrominance reconstruction before DBF can be used for the classification of chrominance samples in HTDF. In one example, the classification can be based on a texture-based method. In one example, the classification can be based on a band-based method.
[0336] Example 10
[0337] In one example, the predicted samples are used as side information for HTDF.
[0338] In one example, the luminance predicted samples can be used to generate the luminance offset value of HTDF. In one example, the luminance predicted samples can be used to generate the chrominance offset value of HTDF.
[0339] In one example, the luminance predicted samples can be used for the classification of luminance samples in HTDF. In one example, the classification can be based on a texture-based method. In one example, the classification can be based on a band-based method.
[0340] In one example, the luminance predicted samples can be used for the classification of chrominance samples in HTDF. In one example, the classification can be based on a texture-based method. In one example, the classification can be based on a band-based method.
[0341] In one example, the chrominance predicted samples can be used to generate the luminance offset value of HTDF. In one example, the chrominance predicted samples can be used to generate the chrominance offset value of HTDF.
[0342] In one example, chrominance prediction samples can be used for the classification of luminance samples in the HTDF. In one example, the classification can be based on a texture-based method. In one example, the classification can be based on a band-based method.
[0343] In one example, chrominance prediction samples can be used for the classification of chrominance samples in the HTDF. In one example, the classification can be based on a texture-based method. In one example, the classification can be based on a band-based method.
[0344] The prediction samples can be modified before being put into the HTDF. The modification can be filtering. The modification can be downsampling / upsampling. The modification can be clipping / shifting.
[0345] Example 11
[0346] In one example, the residual values are used as side information for the HTDF.
[0347] In one example, the luminance residual values can be used to generate the luminance offset values of the HTDF. In one example, the luminance residual values can be used to generate the chrominance offset values of the HTDF.
[0348] In one example, the luminance residual values can be used for the classification of luminance samples in the HTDF. In one example, the classification can be based on a texture-based method. In one example, the classification can be based on a band-based method.
[0349] In one example, the luminance residual values can be used for the classification of chrominance samples in the HTDF. In one example, the classification can be based on a texture-based method. In one example, the classification can be based on a band-based method.
[0350] In one example, the chrominance residual values can be used to generate the luminance offset values of the HTDF. In one example, the chrominance residual values can be used to generate the chrominance offset values of the HTDF.
[0351] In one example, the chrominance residual values can be used for the classification of luminance samples in the HTDF. In one example, the classification can be based on a texture-based method. In one example, the classification can be based on a band-based method.
[0352] In one example, the chrominance residual values can be used for the classification of chrominance samples in the HTDF. In one example, the classification can be based on a texture-based method. In one example, the classification can be based on a band-based method.
[0353] The residual values can be modified before being put into the HTDF. The modification can be filtering. The modification can be downsampling / upsampling. The modification can be clipping / shifting.
[0354] Example 12
[0355] In one example, the segmentation information is used as side information for the HTDF.
[0356] In one example, the segmentation information may represent block size / shape / position or other information. In one example, the luminance segmentation information can be used to generate the luminance offset value of the HTDF. In one example, the luminance segmentation information can be used to generate the chrominance offset value of the HTDF.
[0357] In one example, the luminance segmentation information can be used for the classification of luminance samples in the HTDF. In one example, the classification can be based on a texture-based method. In one example, the classification can be based on a band-based method.
[0358] In one example, the luminance segmentation information can be used for the classification of chrominance samples in the HTDF. In one example, the classification can be based on an edge-based method. In one example, the classification can be based on a band-based method.
[0359] In one example, the chrominance segmentation information can be used to generate the luminance offset value of the HTDF. In one example, the chrominance segmentation information can be used to generate the chrominance offset value of the HTDF.
[0360] In one example, the chrominance segmentation information can be used for the classification of luminance samples in the HTDF. In one example, the classification can be based on a texture-based method. In one example, the classification can be based on a band-based method.
[0361] In one example, the chrominance segmentation information can be used for the classification of chrominance samples in the HTDF. In one example, the classification can be based on a texture-based method. In one example, the classification can be based on a band-based method.
[0362] Example 13
[0363] In one example, the QP information is used as side information for the HTDF.
[0364] In one example, the QP information may represent picture / strip / block QP information. In one example, the luminance QP information can be used to generate the luminance offset value of the HTDF. In one example, the luminance QP information can be used to generate the chrominance offset value of the HTDF.
[0365] In one example, the luminance QP information can be used for the classification of luminance samples in the HTDF. In one example, the classification can be based on a texture-based method. In one example, the classification can be based on a band-based method.
[0366] In one example, the luminance QP information can be used for the classification of chrominance samples in BF. In one example, the classification can be based on a texture-based method. In one example, the classification can be based on a band-based method.
[0367] In one example, the chrominance QP information can be used to generate the luminance offset value of HTDF. In one example, the chrominance QP information can be used to generate the chrominance offset value of HTDF.
[0368] In one example, the chrominance QP information can be used for the classification of luminance samples in HTDF. In one example, the classification can be based on a texture-based method. In one example, the classification can be based on a band-based method.
[0369] In one example, the chrominance QP information can be used for the classification of chrominance samples in HTDF. In one example, the classification can be based on a texture-based method. In one example, the classification can be based on a band-based method.
[0370] Example 14
[0371] In one example, the boundary strength information can be used as the side information of HTDF.
[0372] In one example, the boundary strength information can be generated by DBF or other methods. In one example, the luminance boundary strength information can be used to generate the luminance offset value of HTDF. In one example, the luminance boundary strength information can be used to generate the chrominance offset value of HTDF.
[0373] In one example, the luminance boundary strength information can be used for the classification of luminance samples in HTDF. In one example, the classification can be based on a texture-based method. In one example, the classification can be based on a band-based method.
[0374] In one example, the luminance boundary strength information can be used for the classification of chrominance samples in HTDF. In one example, the classification can be based on a texture-based method. In one example, the classification can be based on a band-based method.
[0375] In one example, the chrominance boundary strength information can be used to generate the luminance offset value of HTDF. In one example, the chrominance boundary strength information can be used to generate the chrominance offset value of HTDF.
[0376] In one example, the chrominance boundary strength information can be used for the classification of luminance samples in HTDF. In one example, the classification can be based on a texture-based method. In one example, the classification can be based on a band-based method.
[0377] In one example, the chrominance boundary strength information can be used for the classification of chrominance samples in the HTDF. In one example, the classification can be based on a texture-based method. In one example, the classification can be based on a band-based method.
[0378] Example 15
[0379] In one example, the disclosed method can be used for post-processing and / or pre-processing.
[0380] Example 16
[0381] In one example, the methods mentioned above can be used in combination.
[0382] Example 17
[0383] In one example, the methods mentioned above can be used individually.
[0384] Example 18
[0385] In one example, the described edge information utilization method can be applied to any loop filtering tool, pre-processing or post-processing filtering method in video coding and decoding (including but not limited to ALF / CCALF or any other filtering method).
[0386] Example 19
[0387] In one example, the proposed edge information utilization method can be applied to loop filtering methods. In one example, the proposed edge information utilization method can be applied to ALF. In one example, the proposed edge information utilization method can be applied to CCALF. In one example, the proposed edge information utilization method can be applied to SAO. In one example, the proposed edge information utilization method can be applied to CCSAO. In one example, the proposed edge information utilization method can be applied to the bilateral filter (BF). In one example, the proposed edge information utilization method can be applied to the Hadamard transform domain filter (HTDF). Alternatively, the proposed edge information utilization method can be applied to other loop filtering methods.
[0388] Example 20
[0389] In one example, the edge information utilization method can be applied to pre-processing filtering methods. In one example, the edge information utilization method can be applied to post-processing filtering methods.
[0390] Example 21
[0391] In the above example, a video unit may refer to a sequence / picture / sub-picture / strip / slice / coding tree unit (CTU) / CTU row / group of CTUs / coding unit (CU) / prediction unit (PU) / transformation unit (TU) / coding tree block (CTB) / coding block (CB) / prediction block (PB) / transformation block (TB) / any other region containing more than one luma or chroma sample / pixel.
[0392] Example 22
[0393] Whether and / or how to apply the methods disclosed above can be signaled in the bitstream.
[0394] In one example, they can be signaled at the sequence level / group of pictures level / picture level / strip level / slice group level, such as in the sequence header, picture header, SPS, VPS, DPS, decoding capability information (DCI), PPS, adaptive parameter set (APS), strip header, and slice group header.
[0395] In one example, they can be signaled in PB, TB, CB, PU, TU, CU, virtual pipeline decoding unit (VPDU), CTU, CTU row, strip, slice, sub-picture, and other types of regions containing more than one sample or pixel.
[0396] Example 23
[0397] Whether and / or how to apply the methods disclosed above may depend on coding information, such as block size, color format, single / double tree partitioning, color component, strip / picture type.
[0398] 6. References
[0399] [1] J. Strom, P. Wennersten, J. Enhorn, D. Liu, K. Andersson and R. Sjoberg, “Bilateral Loop Filter in Combination with SAO,” in proceeding of IEEE Picture Coding Symposium (PCS), Nov. 2019.
[0400] Figure 13is a block diagram showing an example video processing system 4000 in which various techniques disclosed herein may be implemented. Various embodiments may include some or all of the components of system 4000. System 4000 may include an input 4002 for receiving video content. The video content may be received in a raw or uncompressed format (e.g., 8-bit or 10-bit multi-component pixel values), or may be received in a compressed format or an encoded format. Input 4002 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, passive optical network (PON), etc., and wireless interfaces such as Wi-Fi or cellular interfaces.
[0401] System 4000 may include a codec component 4004, which may implement various codec or encoding methods described in this document. The codec component 4004 may reduce the average bit rate of the video from the input 4002 to the output of the codec component 4004 to produce a codec representation of the video. Thus, codec techniques are sometimes referred to as video compression or video transcoding techniques. The output of the codec component 4004 may be stored or transmitted via the connected communication, as shown by component 4006. The stored or transmitted bitstream representation (or codec representation) of the video received at input 4002 may be used by component 4008 to generate pixel values or a displayable video, which is sent to a display interface 4010. The process of generating a user-visible video from the bitstream representation is sometimes referred to as video decompression. Additionally, although certain video processing operations are called "encoding" operations or tools, it should be understood that encoding tools or operations are used in an encoder, and a decoder will perform the corresponding decoding tools or operations for the reverse process of encoding.
[0402] Examples of a peripheral bus interface or a display interface may include a universal serial bus (USB), a high-definition multimedia interface (HDMI), or a DisplayPort, etc. Examples of a storage interface include a serial advanced technology attachment (SATA), a peripheral component interconnect (PCI), an integrated drive electronics (IDE) interface, etc. The techniques described in this document may be embodied in various electronic devices (such as a mobile phone, a laptop computer, a smartphone, or other devices capable of performing digital data processing and / or video display).
[0403] Figure 14It is a block diagram of an example video processing device 4100. The device 4100 can be used to implement one or more methods described herein. The device 4100 can be implemented as a smart phone, a tablet computer, a computer, an Internet of Things (IoT) receiver, etc. The device 4100 can include one or more processors 4102, one or more memories 4104, and video processing circuitry 4106. The processor 4102 can be configured to implement one or more methods described in this document. The memory(ies) 4104 can be used to store data and code for implementing the methods and techniques described herein. The video processing circuitry 4106 can be used to implement some of the techniques described in this document in hardware circuitry. In some embodiments, the video processing circuitry 4106 can be at least partially included in the processor 4102 (e.g., a graphics coprocessor).
[0404] Figure 15 It is a flowchart of an example method 4200 for video processing. The method 4200 includes determining to use side information as an input to a filter at step 4202. A conversion between visual media data and a bitstream is performed based on the filter at step 4204. According to an example, the conversion of step 4204 can include encoding at an encoder or decoding at a decoder.
[0405] It should be noted that the method 4200 can be implemented in a device for processing video data, the device including a processor and a non-transitory memory having instructions thereon, the device such as, a video encoder 4400, a video decoder 4500, and / or an encoder 4600. In such a case, the instructions, when executed by the processor, cause the processor to execute the method 4200. Additionally, the method 4200 can be performed by a non-transitory computer-readable medium including a computer program product for use by a video codec device. The computer program product includes computer-executable instructions stored on a non-transitory computer-readable medium such that when executed by a processor, cause the video codec device to execute the method 4200.
[0406] Figure 16 It is a block diagram showing an example video codec system 4300 that can utilize the techniques of the present disclosure. The video codec system 4300 can include a source device 4310 and a destination device 4320. The source device 4310 generates encoded video data, and the source device can be referred to as a video encoding device. The destination device 4320 can decode the encoded video data generated by the source device 4310, and the destination device can be referred to as a video decoding device.
[0407] The source device 4310 may include a video source 4312, a video encoder 4314, and an input / output (I / O) interface 4316. The video source 4312 may include sources such as a video capture device, an interface for receiving video data from a video content provider, and / or a computer graphics system for generating video data, or a combination of such sources. The video data may include one or more pictures. The video encoder 4314 encodes the video data from the video source 4312 to generate a bitstream. The bitstream may include a series of bits that form an encoded representation of the video data. The bitstream may include encoded pictures and associated data. The encoded pictures are the encoded representations of the pictures. The associated data may include sequence parameter sets, picture parameter sets, and other syntax structures. The I / O interface 4316 may include a modulator / demodulator (modem) and / or a transmitter. The encoded video data may be directly sent to the destination device 4320 via the I / O interface 4316 over the network 4330. The encoded video data may also be stored on a storage medium / server 4340 for access by the destination device 4320.
[0408] The destination device 4320 may include an I / O interface 4326, a video decoder 4324, and a display device 4322. The I / O interface 4326 may include a receiver and / or a modem. The I / O interface 4326 may obtain the encoded video data from the source device 4310 or the storage medium / server 4340. The video decoder 4324 may decode the encoded video data. The display device 4322 may display the decoded video data to a user. The display device 4322 may be integrated with the destination device 4320, or may be external to the destination device 4320, which may be configured to be interfaced with an external display device.
[0409] The video encoder 4314 and the video decoder 4324 may operate according to a video compression standard, such as the High Efficiency Video Coding (HEVC) standard, the Versatile Video Coding (VVC) standard, and other current and / or future standards.
[0410] Figure 17 is a block diagram showing an example of a video encoder 4400, which may be the Figure 16 video encoder 4314 in the system 4300 shown. The video encoder 4400 may be configured to perform any or all of the techniques of the present disclosure. The video encoder 4400 includes a plurality of functional components. The techniques described in the present disclosure may be shared among the various components of the video encoder 4400. In some examples, a processor may be configured to perform any or all of the techniques described in the present disclosure.
[0411] The functional components of video encoder 4400 may include: a splitting unit 4401; a prediction unit 4402, which may include a mode selection unit 4403, a motion estimation unit 4404, a motion compensation unit 4405, and an intra prediction unit 4406; a residual generation unit 4407; a transform processing unit 4408; a quantization unit 4409; an inverse quantization unit 4410; an inverse transform unit 4411; a reconstruction unit 4412; a buffer 4413, and an entropy encoding unit 4414.
[0412] In other examples, video encoder 4400 may include more, fewer, or different functional components. In one example, prediction unit 4402 may include an Intra Block Copy (IBC) unit. The IBC unit may perform prediction in the IBC mode, where at least one reference picture is the picture in which the current video block is located.
[0413] In addition, some components, such as motion estimation unit 4404 and motion compensation unit 4405, may be highly integrated, but are shown separately for explanatory purposes in the example of video encoder 4400.
[0414] Splitting unit 4401 may split a picture into one or more video blocks. Video encoder 4400 and video decoder 4500 may support various video block sizes.
[0415] Mode selection unit 4403 may, for example, select one of the encoding / decoding modes based on error results, and provide the resulting intra or inter encoded / decoded block to residual generation unit 4407 to generate residual block data and to reconstruction unit 4412 to reconstruct the encoded / decoded block for use as a reference picture. In some examples, mode selection unit 4403 may select an Intra and Inter Prediction combination (CIIP) mode, where the prediction is based on an inter prediction signal and an intra prediction signal. Mode selection unit 4403 may also select the resolution of the motion vector for a block in the case of inter prediction (e.g., sub-pixel or integer pixel precision).
[0416] To perform inter prediction on the current video block, motion estimation unit 4404 may generate motion information for the current video block by comparing one or more reference frames from buffer 4413 with the current video block. Motion compensation unit 4405 may determine a predicted video block for the current video block based on the motion information and decoded samples of a picture from buffer 4413 (rather than the picture associated with the current video block).
[0417] Motion estimation unit 4404 and motion compensation unit 4405 may perform different operations on the current video block, for example, depending on whether the current video block is in an I-slice, a P-slice, or a B-slice.
[0418] In some examples, the motion estimation unit 4404 may perform uni - directional prediction on a current video block, and the motion estimation unit 4404 may search for a reference video block for the current video block in the reference pictures of list 0 or list 1. Then, the motion estimation unit 4404 may generate a reference index that indicates the reference picture in list 0 or list 1 that contains the reference video block and a motion vector that indicates the spatial displacement between the current video block and the reference video block. The motion estimation unit 4404 may output the reference index, a prediction direction indicator, and the motion vector as the motion information of the current video block. The motion compensation unit 4405 may generate a predicted video block of the current block based on the reference video block indicated by the motion information of the current video block.
[0419] In other examples, the motion estimation unit 4404 may perform bi - directional prediction on a current video block. The motion estimation unit 4404 may search for a reference video block for the current video block in the reference pictures of list 0 and may also search for another reference video block for the current video block in the reference pictures of list 1. Then, the motion estimation unit 4404 may generate a reference index that indicates the reference pictures in list 0 and list 1 that contain the reference video blocks and a motion vector that indicates the spatial displacement between the reference video block and the current video block. The motion estimation unit 4404 may output the reference index and the motion vector of the current video block as the motion information of the current video block. The motion compensation unit 4405 may generate a predicted video block of the current video block based on the reference video block indicated by the motion information of the current video block.
[0420] In some examples, the motion estimation unit 4404 may output a complete set of motion information for the decoder's decoding process. In some examples, the motion estimation unit 4404 may not output a complete set of motion information for the current video. More precisely, the motion estimation unit 4404 may signal the motion information of the current video block by referring to the motion information of another video block. For example, the motion estimation unit 4404 may determine that the motion information of the current video block is sufficiently similar to the motion information of a neighboring video block.
[0421] In one example, the motion estimation unit 4404 may indicate a certain value in the syntax structure associated with the current video block, and this value indicates to the video decoder 4500 that the current video block has the same motion information as another video block.
[0422] In another example, the motion estimation unit 4404 may identify another video block and a motion vector difference (MVD) in a syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 4500 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[0423] As discussed above, the video encoder 4400 may signal motion vectors in a predictive manner. Two examples of predictive signaling techniques that may be implemented by the video encoder 4400 include advanced motion vector prediction (AMVP) and Merge mode signaling.
[0424] The intra prediction unit 4406 may perform intra prediction on the current video block. When the intra prediction unit 4406 performs intra prediction on the current video block, the intra prediction unit 4406 may generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block may include a predicted video block and various syntax elements.
[0425] The residual generation unit 4407 may generate residual data for the current video block by subtracting the predicted video block of the current video block from the current video block. The residual data for the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.
[0426] In other examples, such as in the skip mode, there may be no residual data for the current video block of the current video block, and the residual generation unit 4407 may not perform a subtraction operation.
[0427] The transform processing unit 4408 may generate a transform coefficient video block for the current video block by applying one or more transforms to the residual video block associated with the current video block.
[0428] After the transform processing unit 4408 generates a transform coefficient video block associated with the current video block, the quantization unit 4409 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0429] The inverse quantization unit 4410 and the inverse transform unit 4411 may apply inverse quantization and inverse transform to the transform coefficient video block, respectively, to reconstruct the residual video block from the transform coefficient video block. The reconstruction unit 4412 may add the reconstructed residual video block to the corresponding samples of one or more predicted video blocks generated by the prediction unit 4402 to produce a reconstructed video block associated with the current block for storage in the buffer 4413.
[0430] After reconstructing a video block in the reconstruction unit 4412, a loop filtering operation may be performed to reduce video block artifacts in the video block.
[0431] The entropy coding unit 4414 may receive data from other functional components of the video encoder 4400. When the entropy coding unit 4414 receives data, the entropy coding unit 4414 may perform one or more entropy coding operations to generate entropy-coded data and output a bitstream including the entropy-coded data.
[0432] Figure 18 is a block diagram illustrating an example of a video decoder 4500, which may be Figure 16 the video decoder 4324 in the system 4300 shown. The video decoder 4500 may be configured to perform any or all of the techniques of the present disclosure. In the example shown, the video decoder 4500 includes a plurality of functional components. The techniques described in the present disclosure may be shared among the various components of the video decoder 4500. In some examples, a processor may be configured to perform any or all of the techniques described in the present disclosure.
[0433] In the example shown, the video decoder 4500 includes an entropy decoding unit 4501, a motion compensation unit 4502, an intra prediction unit 4503, an inverse quantization unit 4504, an inverse transform unit 4505, a reconstruction unit 4506, and a cache 4507. In some examples, the video decoder 4500 may perform a decoding process that is substantially inverse to the encoding process described with respect to the video encoder 4400.
[0434] The entropy decoding unit 4501 may extract the encoded bitstream. The encoded bitstream may include entropy-coded video data (e.g., encoded video data blocks). The entropy decoding unit 4501 may decode the entropy-coded video data, and the motion compensation unit 4502 may determine motion information based on the entropy-decoded video data, the motion information including a motion vector, a motion vector precision, a reference picture list index, and other motion information. The motion compensation unit 4502 may determine such information, for example, by performing AMVP and Merge modes.
[0435] The motion compensation unit 4502 may generate a motion-compensated block, and thus may perform interpolation based on an interpolation filter. The identity of the interpolation filter used with sub-pixel precision may be included in a syntax element.
[0436] The motion compensation unit 4502 may use the interpolation filter used by the video encoder 4400 during the encoding of a video block to calculate the interpolation of sub-integer pixels of a reference block. The motion compensation unit 4502 may determine the interpolation filter used by the video encoder 4400 according to the received syntax information, and use the interpolation filter to generate a prediction block.
[0437] The motion compensation unit 4502 can use some syntax information to determine the size of the blocks used to encode the frames and / or slices of the encoded video sequence, the partitioning information describing how each macroblock of the pictures of the encoded video sequence is partitioned, the modes indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-coded block, and other information for decoding the encoded video sequence.
[0438] The intra prediction unit 4503 can form a prediction block from spatially adjacent blocks using, for example, the intra prediction mode received in the bitstream. The inverse quantization unit 4504 inverse quantizes the video block coefficients that are provided in the bitstream and decoded by the entropy decoding unit 4501, i.e., dequantizes them. The inverse transform unit 4505 applies an inverse transform.
[0439] The reconstruction unit 4506 can add the residual block to the corresponding prediction block generated by the motion compensation unit 4502 or the intra prediction unit 4503 to form a decoded block. If necessary, a deblocking filter can also be applied to filter the decoded block to eliminate blocking artifacts. Then, the decoded video block is stored in the cache 4507, which provides reference blocks for subsequent motion compensation / intra prediction and also produces the decoded video for presentation on a display device.
[0440] Figure 19 is a schematic diagram of an example encoder 4600. The encoder 4600 is suitable for implementing VVC technology. The encoder 4600 includes three loop filters, namely, the deblocking filter (DF) 4602, the sample adaptive offset (SAO) 4604, and the adaptive loop filter (ALF) 4606. Different from the DF 4602 that uses predefined filters, the SAO 4604 and the ALF 4606 utilize the original samples of the current picture to reduce the mean square error between the original samples and the reconstructed samples by respectively adding offset values and applying a finite impulse response (FIR) filter and signaling the offset values and filter coefficients using the coded side information. The ALF 4606 is located at the last processing stage of each picture and can be regarded as a tool for trying to capture and fix the artifacts created by the previous stages.
[0441] The encoder 4600 also includes an intra prediction component 4608 configured to receive an input video and a motion estimation / compensation (ME / MC) component 4610. The intra prediction component 4608 is configured to perform intra prediction, while the ME / MC component 4610 is configured to perform inter prediction using reference pictures obtained from a reference picture buffer 4612. Residual blocks from the inter prediction or intra prediction are fed into a transform (T) component 4614 and a quantization (Q) component 4616 to generate quantized residual transform coefficients, which are fed into an entropy coding / decoding component 4618. The entropy coding / decoding component 4618 performs entropy coding / decoding on the prediction result and the quantized transform coefficients and sends them to a video decoder (not shown). The quantized components output from the quantization component 4616 can be fed into an inverse quantization (IQ) component 4620, an inverse transform component 4622, and a reconstruction (REC) component 4624. The REC component 4624 is capable of outputting an image to the DF 4602, SAO 4604, and ALF 4606 for filtering before these pictures are stored in the reference picture buffer 4612.
[0442] Next, a list of some example preferred solutions is provided.
[0443] The following solutions illustrate examples of the techniques discussed herein.
[0444] 1. A method for processing video data, comprising: using side information as an input to a filter; and performing a conversion between visual media data and a bitstream based on the filter.
[0445] 2. The method according to solution 1, wherein the filter is a bilateral filter (BF).
[0446] 3. The method according to any one of solutions 1 to 2, wherein the filter is a Hadamard transform domain filter (HTDF).
[0447] 4. The method according to any one of solutions 1 to 3, wherein the side information includes reconstructed samples obtained before applying a deblocking filter (DBF), and wherein the reconstructed samples are luminance reconstructed samples or chrominance reconstructed samples.
[0448] 5. The method according to any one of solutions 1 to 4, wherein the side information is used to generate a luminance offset value or a chrominance offset value of the filter.
[0449] 6. The method according to any one of solutions 1 to 5, wherein the side information is used for classification of luminance samples or chrominance samples in the filter.
[0450] 7. The method according to any one of solutions 1 to 6, wherein the classification is texture-based or band-based.
[0451] 8. The method according to any one of Solutions 1 to 7, wherein the side information includes predicted samples, and wherein the predicted samples are luminance predicted samples or chrominance predicted samples.
[0452] 9. The method according to any one of Solutions 1 to 8, wherein the predicted samples are modified by filtering, downsampling, upsampling, clipping, shifting, or a combination thereof before being used as side information.
[0453] 10. The method according to any one of Solutions 1 to 9, wherein the side information includes residual values, and wherein the residual values are luminance residual values or chrominance residual values.
[0454] 11. The method according to any one of Solutions 1 to 10, wherein the residual values are modified by filtering, downsampling, upsampling, clipping, shifting, or a combination thereof before being used as side information.
[0455] 12. The method according to any one of Solutions 1 to 11, wherein the side information includes segmentation information, and wherein the segmentation information is luminance segmentation information or chrominance segmentation information.
[0456] 13. The method according to any one of Solutions 1 to 12, wherein the segmentation information includes block size, block shape, block position, or a combination thereof.
[0457] 14. The method according to any one of Solutions 1 to 13, wherein the side information includes quantization parameter (QP) information, and wherein the QP information includes luminance QP information or chrominance QP information.
[0458] 15. The method according to any one of Solutions 1 to 14, wherein the QP information includes picture QP information, slice QP information, block QP information, or a combination thereof.
[0459] 16. The method according to any one of Solutions 1 to 15, wherein the side information includes boundary strength information, and wherein the boundary strength information includes luminance boundary strength information or chrominance boundary strength information.
[0460] 17. The method according to any one of Solutions 1 to 16, wherein the boundary strength information is generated by DBF.
[0461] 18. The method according to any one of Solutions 1 to 17, wherein the method is used for preprocessing or postprocessing of video.
[0462] 19. The method according to any one of Solutions 1 to 18, wherein the method is used as part of a loop filter, and wherein the method is applied to an ALF, CCALF, sample adaptive compensation (SAO) filter, cross-component SAO (CCSAO), bilateral filter (BF), Hadamard transform domain filter (HTDF), or a combination thereof.
[0463] 20. The method according to any one of Solutions 1 to 19, wherein the filter is applied to a video unit, and wherein the video unit is a sequence, picture, sub-picture, slice, strip, tile, coding tree unit (CTU), CTU row, CTU group, coding unit (CU), prediction unit (PU), transform unit (TU), coding tree block (CTB), coding block (CB), prediction block (PB), transform block (TB), or any other region containing more than one luminance or chrominance sample or pixel.
[0464] 21. The method according to any one of Solutions 1 to 20, wherein the application of the method is signaled in the bitstream.
[0465] 22. The method according to any one of Solutions 1 to 21, wherein the application of the method depends on coded information.
[0466] 23. An apparatus for processing video data, comprising: a processor; and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of Solutions 1 to 22.
[0467] 24. A non-transitory computer-readable medium, comprising a computer program product for use in a video coding / decoding device, the computer program product comprising computer-executable instructions stored on the non-transitory computer-readable medium, such that when executed by a processor, cause the video coding / decoding device to perform the method according to any one of Solutions 1 to 22.
[0468] 25. A non-transitory computer-readable recording medium storing a bitstream of a video generated by a method performed by a video processing device, the method comprising: determining to use side information as an input to a filter; and generating the bitstream based on the determination.
[0469] 26. A method for storing a bitstream of a video, comprising: determining to use side information as an input to a filter; generating the bitstream based on the determination; and storing the bitstream in a non-transitory computer-readable recording medium.
[0470] 27. A method, apparatus, or system described in this document.
[0471] The following solutions show further examples of the techniques discussed herein.
[0472] 1. A method for processing video data, comprising: determining to use side information as an input to a filter; and performing a conversion between visual media data and a bitstream based on the filter.
[0473] 2. The method according to solution 1, wherein the filter is a bilateral filter (BF) or a Hadamard transform domain filter (HTDF).
[0474] 3. The method according to any one of solutions 1 to 2, wherein the side information includes reconstructed samples before applying a deblocking filter.
[0475] 4. The method according to any one of solutions 1 to 3, wherein the reconstructed luma samples or the reconstructed chroma samples before applying the deblocking filter are used to generate a luma offset value or a chroma offset value used in the BF or the HTDF.
[0476] 5. The method according to solutions 1 to 4, wherein the side information includes predicted samples.
[0477] 6. The method according to any one of solutions 1 to 5, wherein the luma predicted samples or the chroma predicted samples are used to generate a luma offset value or a chroma offset value used in the BF or the HTDF.
[0478] 7. The method according to any one of solutions 1 to 6, wherein the side information includes prediction values.
[0479] 8. The method according to any one of solutions 1 to 7, wherein the luma residual values or the chroma residual values are used to generate a luma offset value or a chroma offset value used in the BF or the HTDF.
[0480] 9. The method according to any one of solutions 1 to 8, wherein the residual values are downsampled or upsampled before being used as side information in the BF or the HTDF.
[0481] 10. The method according to any one of solutions 1 to 9, wherein the side information includes segmentation information.
[0482] 11. The method according to any one of solutions 1 to 10, wherein the segmentation information includes block size, block shape, block position, or block information.
[0483] 12. The method according to any one of solutions 1 to 11, wherein the luma segmentation information or the chroma segmentation information is used to generate a luma offset value or a chroma offset value used in the BF or the HTDF.
[0484] 13. The method according to any one of solutions 1 to 12, wherein the side information includes quantization parameter (QP) information.
[0485] 14. The method according to any one of Solutions 1 to 13, wherein the QP information includes picture QP information, slice QP information, or block QP information.
[0486] 15. The method according to any one of Solutions 1 to 14, wherein the luma QP information or chroma QP information is used to generate a luma offset value or chroma offset value used in the BF or the HTDF.
[0487] 16. The method according to any one of Solutions 1 to 15, wherein the side information includes boundary strength information.
[0488] 17. The method according to any one of Solutions 1 to 16, wherein the boundary strength information is generated by a deblocking filter (DBF).
[0489] 18. The method according to any one of Solutions 1 to 17, wherein the luma boundary strength information or chroma boundary strength information is used to generate a luma offset value or chroma offset value used in the BF or the HTDF.
[0490] 19. The method according to any one of Solutions 1 to 18, wherein the reconstructed luma samples or reconstructed chroma samples before applying the deblocking filter are used for the classification of luma samples or chroma samples in the BF or the HTDF.
[0491] 20. The method according to any one of Solutions 1 to 19, wherein luma prediction samples or chroma prediction samples are used for the classification of luma samples or chroma samples in the BF or the HTDF.
[0492] 21. The method according to any one of Solutions 1 to 20, wherein the prediction samples are filtered, clipped, or shifted before being used for classification in the BF or the HTDF.
[0493] 22. The method according to any one of Solutions 1 to 21, wherein luma residual values or chroma residual values are used for the classification of luma samples or chroma samples in the BF or the HTDF.
[0494] 23. The method according to any one of Solutions 1 to 22, wherein the residual values are filtered, clipped, or shifted before being used for classification in the BF or the HTDF.
[0495] 24. The method according to any one of Solutions 1 to 23, wherein luma segmentation information or chroma segmentation information is used for the classification of luma samples or chroma samples in the BF or the HTDF.
[0496] 25. The method according to any one of Solutions 1 to 24, wherein luma QP information or chroma QP information is used for the classification of luma samples or chroma samples in the BF or the HTDF.
[0497] 26. The method according to any one of Solutions 1 to 25, wherein the luminance boundary strength information or the chrominance boundary strength information is used for classifying the luminance samples or the chrominance samples in the BF or the HTDF.
[0498] 27. The method according to any one of Solutions 1 to 26, wherein the classification employs a texture-based method, a band-based method, or an edge-based method.
[0499] 28. The method according to any one of Solutions 1 to 27, wherein the edge information is applied to a loop filter, a preprocessing filter, or a postprocessing filter of a video.
[0500] 29. The method according to any one of Solutions 1 to 28, wherein the method is used jointly or separately.
[0501] 30. The method according to any one of Solutions 1 to 29, wherein the method is used as part of a loop filter, and wherein the method is applied to an adaptive loop filter (ALF), a cross-component ALF (CCALF), a sample adaptive offset (SAO) filter, a cross-component SAO (CCSAO), a bilateral filter (BF), a Hadamard transform domain filter (HTDF), or a combination thereof.
[0502] 31. The method according to any one of Solutions 1 to 30, wherein the filter is applied to a video unit, and wherein the video unit is a sequence, a picture, a sub-picture, a slice, a tile, a coding tree unit (CTU), a CTU row, a CTU group, a coding unit (CU), a prediction unit (PU), a transform unit (TU), a coding tree block (CTB), a coding block (CB), a prediction block (PB), a transform block (TB), or any other region containing more than one luminance or chrominance sample or pixel.
[0503] 32. The method according to any one of Solutions 1 to 31, wherein the use of the method is signaled in the bitstream.
[0504] 33. The method according to any one of Solutions 1 to 32, wherein the method is signaled at sequence level, group of pictures level, picture level, slice level or tile group level, including being signaled in a sequence header, a picture header, a sequence parameter set (SPS), a video parameter set (VPS), a decoding parameter set (DPS), decoding capability information (DCI), a picture parameter set (PPS), an adaptive parameter set (APS), a slice header or a tile group header, or wherein the use of the method is signaled in a prediction block (PB), a transform block (TB), a codec block (CB), a prediction unit (PU), a transform unit (TU), a codec unit (CU), a virtual pipeline decoding unit (VPDU), a codec tree unit (CTU), a CTU row, a slice, a tile, a sub-picture or other regions containing more than one sample or pixel.
[0505] 34. The method according to any one of Solutions 1 to 33, wherein the application of the method depends on codec information, including block size, color format, single-tree segmentation, double-tree segmentation, color component, slice type or picture type.
[0506] 35. The method according to any one of Solutions 1 to 34, wherein the conversion includes encoding the visual media data into the bitstream.
[0507] 36. The method according to any one of Solutions 1 to 34, wherein the conversion includes decoding the visual media data from the bitstream.
[0508] 37. An apparatus for processing video data, comprising: a processor; and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of Solutions 1 to 36.
[0509] 38. A non-transitory computer-readable medium, including a computer program product for use by a video codec device, the computer program product including computer-executable instructions stored on the non-transitory computer-readable medium, such that when executed by a processor, cause the video codec device to perform the method according to any one of Solutions 1 to 36.
[0510] 39. A non-transitory computer-readable recording medium storing a bitstream of a video generated by a method executed by a video processing device, wherein the method includes: determining to use side information as an input to a filter; and generating the bitstream based on the determination.
[0511] 40. A method for storing a bitstream of a video, including: determining to use side information as an input to a filter; generating the bitstream based on the determination; and storing the bitstream in a non-transitory computer-readable recording medium.
[0512] In the solutions described herein, an encoder can conform to formatting rules by generating an encoded / decoded representation according to the formatting rules. In the solutions described herein, a decoder can use the formatting rules to parse syntax elements in the encoded / decoded representation and, based on the presence and absence of known syntax elements according to the formatting rules, generate a decoded video.
[0513] In this document, the term "video processing" can refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm can be applied during the conversion from a pixel representation of a video to a corresponding bitstream representation (and vice versa). For example, the bitstream representation of a current video block can correspond to bits that are co-located or distributed at different positions in the bitstream, as defined by the syntax. For example, a macroblock can be encoded based on transformed and encoded error residual values and also using bits in a header and other fields in the bitstream. Additionally, during the conversion, a decoder can parse the bitstream based on determinations and know that certain fields may be present or absent, as described in the solutions above. Similarly, an encoder can determine whether to include certain syntax fields and generate an encoded representation accordingly by including or excluding the syntax fields in the encoded representation.
[0514] The disclosed and other solutions, examples, embodiments, modules, and functional operations described in this document can be implemented in digital electronic circuits or in computer software, firmware, hardware, or a combination of one or more of them, the computer software, firmware, or hardware including the structures disclosed in this document and their equivalent structures. The disclosed and other embodiments can be implemented as one or more computer program products encoded on a computer-readable medium for execution by, or to control the operation of, a data processing apparatus, i.e., one or more modules of computer program instructions. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition of matter affecting a machine-readable propagated signal, or a combination of one or more of them. The term "data processing apparatus" encompasses all apparatus, devices, and machines for processing data, including, for example, programmable processors, computers, or multiple processors or computers. In addition to the hardware, the apparatus can also include code that creates an execution environment for the computer programs being discussed, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them. A propagated signal is a man-made signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, which is generated to encode information for transmission to a suitable receiver apparatus.
[0515] A computer program (also known as a program, software, software application, script, or code) can be written in any form of programming language (including compiled or interpreted languages), and it can be deployed in any form (including as a stand-alone program or as a module, component, subroutine, or any other unit suitable for use in a computing environment). A computer program does not necessarily correspond to a file in a file system. The program can be stored as part of a file that holds other programs or data (e.g., one or more scripts in a markup language document), in a single file dedicated to the program being discussed, or in multiple coordinated files (e.g., files that store one or more modules, subroutines, or portions of code). A computer program can be deployed to execute on one computer or on multiple computers, which may be located at one site or distributed across multiple sites and interconnected by a communication network.
[0516] The processes or logical flows described in this document can be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logical flows can also be performed by special-purpose logic circuitry, such as a field-programmable gate array (FPGA) or an application-specific integrated circuit (ASIC), and the apparatus can also be implemented as special-purpose logic circuitry, such as a field-programmable gate array (FPGA) or an application-specific integrated circuit (ASIC).
[0517] Processors suitable for the execution of a computer program include, by way of example, both general and special purpose microprocessors, and any one or more processors of any kind of digital computer. Generally, a processor will receive instructions and data from a read-only memory or a random access memory or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include one or more mass storage devices for storing data (e.g., magnetic disks, magneto-optical disks, or optical disks), or be operatively coupled to receive data from or transfer data to one or more mass storage devices or both. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including by way of example semiconductor memory devices, such as erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and compact disc read-only memory (CD ROM) and digital versatile disc read-only memory (DVD-ROM) disks. The processor and the memory can be supplemented by, or incorporated in, special-purpose logic circuitry.
[0518] Although this document contains many details, these details should not be construed as limitations on any subject matter scope or what may be claimed, but rather as descriptions of features specific to particular embodiments that may be specific to a particular technology. Certain features described in the context of separate embodiments in this patent document may also be implemented in combination in a single embodiment. Conversely, the various features described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments. Additionally, although features may be described above as acting in certain combinations and even initially claimed as such, in some cases, one or more features from a claimed combination may be excluded from the combination, and the claimed combination may relate to a sub-combination or variation of a sub-combination.
[0519] Similarly, although operations are shown in the figures in a particular order, this should not be construed as requiring that such operations be performed in the particular order shown or in sequential order or that all shown operations be performed to achieve the desired result. Additionally, the separation of various system components described in this patent document should not be construed as requiring such separation in all embodiments.
[0520] Only a few embodiments and examples are described, and other embodiments, enhancements, and changes may be made based on what is described and shown in this patent document.
[0521] A first component is directly coupled to a second component when there is no intermediate component other than a line, trace, or another medium between the first and second components. A first component is indirectly coupled to a second component when there is an intermediate component other than a line, trace, or another medium between the first and second components. The term "coupled" and its variants include both direct and indirect coupling. Unless otherwise specified, the use of the term "about" means a range of ±10% of the subsequent numerical value.
[0522] Although several embodiments are provided in this disclosure, it should be understood that the disclosed systems and methods may be embodied in many other specific forms without departing from the spirit or scope of the disclosure. This example is considered illustrative rather than restrictive and is not intended to be limited to the details given herein. For example, various elements or components may be combined or integrated in another system, or certain features may be omitted or not implemented.
[0523] Additionally, without departing from the scope of the present disclosure, the techniques, systems, subsystems, and methods described and illustrated as discrete or separate in various embodiments may be combined or integrated with other systems, modules, techniques, or methods. Other items shown or discussed as being coupled may be directly connected or may be indirectly coupled or communicate through some interface, device, or intermediate component, whether electrically, mechanically, or otherwise. Those skilled in the art can identify other examples of changes, substitutions, and alterations and can make these changes, substitutions, and alterations without departing from the spirit and scope disclosed herein.
Claims
1. A method for processing video data, comprising: Determine to use side information as the input to the filter; and Perform conversion between visual media data and a bitstream based on the filter.
2. The method according to claim 1, wherein the filter is a bilateral filter (BF) or a Hadamard transform domain filter (HTDF).
3. The method according to any one of claims 1 to 2, wherein the side information includes reconstructed samples before applying the deblocking filter.
4. The method according to any one of claims 1 to 3, wherein the reconstructed luminance samples or the reconstructed chrominance samples before applying the deblocking filter are used to generate a luminance offset value or a chrominance offset value used in the BF or the HTDF.
5. The method according to any one of claims 1 to 4, wherein the side information includes predicted samples.
6. The method according to any one of claims 1 to 5, wherein the luminance predicted samples or the chrominance predicted samples are used to generate a luminance offset value or a chrominance offset value used in the BF or the HTDF.
7. The method according to any one of claims 1 to 6, wherein the side information includes prediction values.
8. The method according to any one of claims 1 to 7, wherein the luminance residual values or the chrominance residual values are used to generate a luminance offset value or a chrominance offset value used in the BF or the HTDF.
9. The method according to any one of claims 1 to 8, wherein the residual values are downsampled or upsampled before being used as side information in the BF or the HTDF.
10. The method according to any one of claims 1 to 9, wherein the side information includes segmentation information.
11. The method according to any one of claims 1 to 10, wherein the segmentation information includes block size, block shape, block position or block information.
12. The method according to any one of claims 1 to 11, wherein the luminance segmentation information or the chrominance segmentation information is used to generate a luminance offset value or a chrominance offset value used in the BF or the HTDF.
13. The method according to any one of claims 1 to 12, wherein the side information includes quantization parameter (QP) information.
14. The method according to any one of claims 1 to 13, wherein the QP information includes picture QP information, slice QP information or block QP information.
15. The method according to any one of claims 1 to 14, wherein the luminance QP information or the chrominance QP information is used to generate a luminance offset value or a chrominance offset value used in the BF or the HTDF.
16. The method according to any one of claims 1 to 15, wherein the side information includes boundary strength information.
17. The method according to any one of claims 1 to 16, wherein the boundary strength information is generated by a deblocking filter (DBF).
18. The method according to any one of claims 1 to 17, wherein the luminance boundary strength information or the chrominance boundary strength information is used to generate a luminance offset value or a chrominance offset value used in the BF or the HTDF.
19. The method according to any one of claims 1 to 18, wherein the reconstructed luminance samples or the reconstructed chrominance samples before applying the deblocking filter are used for classifying the luminance samples or the chrominance samples in the BF or the HTDF.
20. The method according to any one of claims 1 to 19, wherein the luminance prediction samples or the chrominance prediction samples are used for classifying the luminance samples or the chrominance samples in the BF or the HTDF.
21. The method according to any one of claims 1 to 20, wherein the prediction samples are filtered, clipped or shifted before being used for classification in the BF or the HTDF.
22. The method according to any one of claims 1 to 21, wherein the luminance residual values or the chrominance residual values are used for classifying the luminance samples or the chrominance samples in the BF or the HTDF.
23. The method according to any one of claims 1 to 22, wherein the residual values are filtered, clipped or shifted before being used for classification in the BF or the HTDF.
24. The method according to any one of claims 1 to 23, wherein the luminance segmentation information or the chrominance segmentation information is used for classifying the luminance samples or the chrominance samples in the BF or the HTDF.
25. The method according to any one of claims 1 to 24, wherein the luminance QP information or the chrominance QP information is used for classifying the luminance samples or the chrominance samples in the BF or the HTDF.
26. The method according to any one of claims 1 to 25, wherein the luminance boundary strength information or the chrominance boundary strength information is used for classifying the luminance samples or the chrominance samples in the BF or the HTDF.
27. The method according to any one of claims 1 to 26, wherein the classification uses a texture-based method, a band-based method or an edge-based method.
28. The method according to any one of claims 1 to 27, wherein the side information is applied in a loop filter, a pre-processing filter, or a post-processing filter of the video.
29. The method according to any one of claims 1 to 28, wherein the method is used jointly or separately.
30. The method according to any one of claims 1 to 29, wherein the method is applied as part of a loop filter, and wherein the method is applied to an adaptive loop filter (ALF), a cross-component ALF (CCALF), a sample adaptive offset (SAO) filter, a cross-component SAO (CCSAO), a bilateral filter (BF), a Hadamard transform domain filter (HTDF), or a combination thereof.
31. The method according to any one of claims 1 to 30, wherein the filter is applied to a video unit, and wherein the video unit is a sequence, a picture, a sub-picture, a slice, a tile, a coding tree unit (CTU), a CTU row, a CTU group, a coding unit (CU), a prediction unit (PU), a transform unit (TU), a coding tree block (CTB), a coding block (CB), a prediction block (PB), a transform block (TB), or any other region containing more than one luma or chroma sample or pixel.
32. The method according to any one of claims 1 to 31, wherein the use of the method is signaled in the bitstream.
33. The method according to any one of claims 1 to 32, wherein the use of the method is signaled at the sequence level, group of pictures level, picture level, slice level, or tile group level, including signaling in a sequence header, a picture header, a sequence parameter set (SPS), a video parameter set (VPS), a decoding parameter set (DPS), decoding capability information (DCI), a picture parameter set (PPS), an adaptive parameter set (APS), a slice header, or a tile group header, or wherein the use of the method is signaled at a prediction block (PB), a transform block (TB), a coding block (CB), a prediction unit (PU), a transform unit (TU), a coding unit (CU), a virtual pipeline decoding unit (VPDU), a coding tree unit (CTU), a CTU row, a slice, a tile, a sub-picture, or any other region containing more than one sample or pixel.
34. The method according to any one of claims 1 to 33, wherein the application of the method depends on coding information, the coding information including block size, color format, single-tree segmentation, dual-tree segmentation, color component, slice type, or picture type.
35. The method according to any one of claims 1 to 34, wherein the conversion includes encoding the visual media data into the bitstream.
36. The method according to any one of claims 1 to 34, wherein the conversion includes decoding the visual media data from the bitstream.
37. An apparatus for processing video data, comprising: A processor; And a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of claims 1 to 36.
38. A non-transitory computer-readable medium, comprising a computer program product for use by a video codec device, the computer program product including computer-executable instructions stored on the non-transitory computer-readable medium such that, when executed by a processor, cause the video codec device to perform the method according to any one of claims 1 to 36.
39. A non-transitory computer-readable recording medium storing a bitstream of a video generated by a method executed by a video processing device, wherein the method includes: Determine to use side information as the input to the filter; and Generate the bitstream based on the determination.
40. A method for storing a bitstream of a video, comprising: Determine to use side information as the input to the filter; Generate the bitstream based on the determination; and Store the bitstream in a non-transitory computer-readable recording medium.