Reconstruction processed by multiple adaptive loop filters in video coding
By using the boundary strength of the deblocking filter as input in video encoding and decoding, the boundary processing of the adaptive loop filter is optimized, which solves the problem of low efficiency in boundary artifact processing in the prior art and achieves more efficient encoding and decoding results.
Patent Information
- Application Number
- CN202480020722.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-03-21
- Filing Date
- 2024-03-21
- Publication Date
- 2025-11-14
AI Technical Summary
Existing video encoding and decoding technologies are inefficient in dealing with boundary artifacts and cannot effectively utilize edge information for filtering, resulting in low encoding and decoding efficiency.
The boundary strength (DBF-BS) of the deblocking filter is used as the edge information input into the adaptive loop filter (ALF) or cross-component ALF (CC-ALF). The boundary processing process is optimized through boundary strength calculation and filter design.
It improves the efficiency and quality of video encoding and decoding, reduces boundary artifacts, and enhances encoding and decoding efficiency and image quality.
Smart Images

Figure CN120958797A_ABST
Abstract
Description
[0001] Cross-references to related applications
[0002] This patent application claims the benefit of International Patent Application No. PCT / CN2023 / 082725, filed on March 21, 2023, which is incorporated herein by reference. Technical Field
[0003] This patent discloses the generation, storage, and use of digital audio and video media information in file formats. Background Technology
[0004] Digital video accounts for the largest share of bandwidth used on the internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth demand for digital video is likely to continue to grow. Summary of the Invention
[0005] The first aspect relates to a method for processing video data, comprising: determining edge information using a deblocking filter boundary strength (DBF-BS) as input to an adaptive loop filter (ALF) or a cross-component ALF (CC-ALF); and performing a conversion between visual media data and a bitstream based on said ALF or CC-ALF.
[0006] The second aspect relates to an apparatus for processing video data, comprising: a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform a method of any of the preceding aspects.
[0007] The third aspect relates to a non-transitory computer-readable medium comprising a computer program product for use by a video codec apparatus, the computer program product including computer-executable instructions stored on the non-transitory computer-readable medium such that, when executed by a processor, the video codec apparatus performs the methods of any of the preceding aspects.
[0008] The fourth aspect relates to a non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by a video processing apparatus, wherein the method includes: determining side information using DBF-BS as input to an ALF or CC-ALF; and generating the bitstream based on the determination.
[0009] The fifth aspect relates to a method for storing a bitstream of video, comprising: determining side information using DBF-BS as input to an ALF or CC-ALF; generating the bitstream based on the determination; and storing the bitstream in a non-transitory computer-readable recording medium.
[0010] The sixth aspect relates to the methods, apparatus, or systems described in this disclosure.
[0011] For clarity, any of the embodiments described above may be combined with one or more other embodiments described above to create new embodiments within the scope of this disclosure.
[0012] These and other features will become clearer through the following detailed description of the embodiments with reference to the accompanying drawings and claims. Attached Figure Description
[0013] For a more complete understanding of this disclosure, reference is now made to the following brief description, along with the accompanying drawings and detailed description, wherein the same reference numerals denote the same parts.
[0014] Figure 1 An example of the nominal vertical and horizontal positions of 4:2:2 luminance and chrominance samples in an image is shown.
[0015] Figure 2 A sample encoder block diagram is shown.
[0016] Figure 3 An example image is shown, segmented into raster scan strips.
[0017] Figure 4 An example image is shown, segmented into rectangular scan strips.
[0018] Figure 5 An example image showing the structure divided into bricks is shown.
[0019] Figure 6A , 6B Figures 6 and 6C show an example of a codec tree block (CTB) spanning image boundaries.
[0020] Figure 7 An example of an intra-frame prediction mode is shown.
[0021] Figure 8 An example of a block boundary is shown in the image.
[0022] Figure 9 An example of pixels used in a filter is shown.
[0023] Figure 10 An example of the filter shape for an adaptive loop filter (ALF) is shown.
[0024] Figure 11 An example of the transformed coefficients for a 5×5 rhombus filter support is shown.
[0025] Figure 12 An example of relative coordinates for a 5×5 rhombus filter support is shown.
[0026] Figure 13 This is a block diagram illustrating an example video processing system.
[0027] Figure 14 This is a block diagram of an example video processing device.
[0028] Figure 15 This is a flowchart of an example method for video processing.
[0029] Figure 16 This is a block diagram illustrating an example video codec system.
[0030] Figure 17 This is a block diagram showing an example encoder.
[0031] Figure 18 This is a block diagram showing an example decoder.
[0032] Figure 19 This is a schematic diagram of an example encoder. Detailed Implementation
[0033] First, it should be understood that although illustrative implementations of one or more embodiments are provided below, the disclosed systems and / or methods can be implemented using any number of techniques, whether currently known or yet to be developed. This disclosure should not be limited in any way to the illustrative implementations, drawings, and techniques shown below, including the exemplary designs and implementations shown and described herein, but rather to modifications and the full scope of their equivalents within the scope of the appended claims.
[0034] Chapter headings are used in this disclosure for ease of understanding and not to limit the applicability of the techniques and embodiments disclosed in each chapter to that chapter only. Furthermore, the techniques described herein are applicable to other video codec protocols and designs.
[0035] 1. Preliminary Discussion
[0036] This disclosure relates to video encoding and decoding techniques. Specifically, it relates to loop filters and other encoding and decoding tools in image / video encoding and decoding. These ideas can be applied individually or in various combinations to video codecs, such as High Efficiency Video Codec (HEVC), Multi-Functional Video Codec (VVC), or other video encoding and decoding techniques.
[0037] 2. Abbreviation
[0038] This disclosure includes the following abbreviations. Advanced Video Coding (ITU-T H.264 Recommendation | ISO / IEC 14496-10) (AVC), Cocoded Picture Buffer (CPB), Pure Random Access (CRA), Codec Tree Unit (CTU), Codec Video Sequence (CVS), Decoded Picture Buffer (DPB), Decoded Parameter Set (DPS), General Constraint Information (GCI), High-Efficiency Video Coding (also known as ITU-T H.265 Recommendation | ISO / IEC 23008-2, HEVC), Joint Exploration Model (JEM), Motion Constraint Piece Set (MCTS), Network Abstraction Layer (NAL), Output Layer Set (OLS), Picture Header (PH), Picture Parameter Set (PPS), Precedence, Layer and Level (PTL), Picture Unit (PU), Reference Picture Resampling (RPR), Raw Byte Sequence Payload (RBSP), Supplementary Enhancement Information (SEI), Strip Header (SH), Sequence Parameter Set (SPS), Video Coding Layer (VCL), Video Parameter Set (VPS), Multi-Functional Video Coding (also known as ITU-T H.266 Recommendation | ISO / IEC 23090-3, (VVC), VVC Test Model (VTM), Video Availability Information (VUI), Transform Unit (TU), Codec Unit (CU), Deblocking Filter (DF), Sample Adaptive Compensation (SAO), Adaptive Loop Filter (ALF), Codec Block Flag (CBF), Quantization Parameter (QP), Rate Distortion Optimization (RDO), and Bilateral Filter (BF).
[0039] 3. Video codec standards
[0040] Video coding standards have evolved primarily through the development of ITU-T and ISO / IEC standards. ITU-T developed the H.261 and H.263 standards, ISO / IEC developed the MPEG-1 and MPEG-4 Vision standards, and the two organizations jointly developed the H.262 / MPEG-2 video standard and the H.264 / MPEG-4 Advanced Video Coding (AVC) standard and the H.265 / HEVC standard[1]. Starting with H.262, video coding standards are based on a hybrid video coding architecture, which utilizes temporal prediction plus transform coding. In order to explore future video coding technologies beyond HEVC, the Joint Video Exploration Team (JVET) was jointly established by VCEG and MPEG. JVET adopted many methods and incorporated them into a reference software called the Joint Exploration Model (JEM)[2]. When the Multifunctional Video Coding (VVC) project was officially launched, JVET was renamed the Joint Video Experts Group (JVET). VVC is a coding standard that aims to reduce the bit rate by 50% compared to HEVC. The working draft of VVC and the VVC Test Model (VTM) are constantly updated.
[0041] A sample version of the VVC draft, namely the Multi-Functional Video Codec (Draft 10), can be found at: https: / / jvet-experts.org / doc_end_user / documents / 19_Teleconference / wg11 / JVET-S2001-v17.zip. A sample version of the VVC reference software, named VTM, can be found at: https: / / vcgit.hhi.fraunhofer.de / jvet-u-ee2 / VVCSoftware_VTM / - / tree / VTM-11.2.
[0042] The International Telecommunication Union Telecommunication Standardization Sector (ITU-T) Video Coding Experts Group (VCEG) and the International Organization for Standardization and the International Electrotechnical Commission (ISO / IEC) Moving Picture Experts Group (MPEG) Joint Technical Committee (JTC) 1 / Subcommittee (SC) 29 / Working Group (WG) 11 are studying the potential need to standardize future video coding and decoding technologies with compression capabilities significantly exceeding the current VVC standard. This future standardization action could take the form of extended VVC (multiple versions) or entirely new standards. These groups are working together on this exploration in a joint collaborative effort known as the Joint Video Exploration Team (JVET) to evaluate compression technology designs proposed by experts in the field. JVET has established the first exploratory experiment (EE), and reference software called the Enhanced Compression Model (ECM) is in use. The test model ECM is continuously being updated.
[0043] 3.1 Color Space and Chromaticity Downsampling
[0044] A color space, also known as a color model (or color system), is a mathematical model that describes a range of colors as tuples of numbers, such as 3 or 4 values or color components (e.g., RGB). Generally, a color space is a refinement of a coordinate system and its subspaces. For video compression, the most commonly used color spaces are Luminance, Blue Difference, and Red Difference (YCbCr) and Red, Green, and Blue (RGB).
[0045] YCbCr, Y'CbCr, or Y Pb / Cb Pr / Cr, also written as YCBCR or Y'CBCR, is a family of color spaces used as part of the color image pipeline in video and digital photography systems. Y' is the luminance component, and CB and CR are the blue and red difference chromaticity components. Y' (with an apostrophe) is distinguished from Y, which is luminance, meaning that light intensity is non-linearly encoded based on gamma-corrected RGB primary colors.
[0046] Chromaticity downsampling is a practice of encoding images by applying a lower resolution to chromaticity information compared to luminance information, taking advantage of the fact that the human visual system is less sensitive to color differences than to luminance differences. 3.1.1 4:4:4
[0048] In a 4:4:4 scheme, each of the three components of Y'CbCr has the same sampling rate. Therefore, there is no chromaticity downsampling. This scheme is sometimes used in high-end film scanners and film post-production. 3.1.2 4:2:2
[0050] In a 4:3:2 format, both chroma components are sampled at half the luminance sampling rate. The horizontal chroma resolution is halved, while the vertical chroma resolution remains unchanged. This reduces the bandwidth of the uncompressed video signal by one-third, with almost no visual difference. Figure 1 The image shows an example of the nominal vertical and horizontal positions for the 4:2:2 color format. 3.1.3 4:2:0
[0052] In 4:2:0, the horizontal sampling is doubled compared to 4:1:1, but the vertical resolution is halved because the Cb and Cr channels are sampled only on each alternating line. Therefore, the data rate is the same. Cb and Cr are downsampled by a factor of 2 in both the horizontal and vertical directions. There are three variants of the 4:2:0 scheme with different horizontal and vertical positions. In MPEG-2, Cb and Cr are co-located horizontally. Cb and Cr are located between pixels in the vertical direction (at the gap position). In JPEG / JFIF, H.261, and MPEG-1, Cb and Cr are located at the gap position, in the middle of the alternating luma samples. In 4:2:0 DV, Cb and Cr are co-located horizontally. In the vertical direction, they are co-located on alternating lines.
[0053] chroma_format_idc separate_colour_plane_flag Color format SubWidthC SubHeightC 0 0 monochrome 1 1 1 0 4:2:0 2 2 2 0 4:2:2 2 1 3 0 4:4:4 1 1 3 1 4:4:4 1 1
[0054] Table 1 shows the SubWidthC and SubHeightC values derived from chroma_format_idc and separate_colour_plane_flag.
[0055] 3.2 Example Encoding / Decoding Flow of Video Codec
[0056] Figure 2An example of a VVC encoder block diagram is shown, which contains three loop filter blocks: DF, SAO, and ALF. Unlike DF, which uses predefined filters, SAO and ALF utilize the original samples of the current image, reducing the mean square error between the original and reconstructed samples by adding an offset and applying a finite impulse response (FIR) filter, respectively, and by using the encoded / decoded side information through signal transmission offset and filter coefficients. ALF is located in the last processing stage of each image and can be viewed as a tool attempting to capture and repair artifacts caused by previous stages.
[0057] 3.3 Definition of Video / Codec Unit
[0058] An image is divided into one or more slice rows and one or more slice columns. A slice is a sequence of CTUs covering a rectangular area of the image. A slice can be divided into one or more bricks, each brick comprising a certain number of CTU rows within the slice. A slice that is not divided into multiple bricks can also be called a brick. However, a brick that is a proper subset of a slice may not be called a slice. A strip contains multiple slices of an image or multiple bricks of a slice.
[0059] Two stripe modes are supported: raster scan stripe mode and rectangular stripe mode. In raster scan stripe mode, the stripe contains a sequence of slices from a raster scan of the image. In rectangular stripe mode, the stripe contains a certain number of tiles that collectively form a rectangular area of the image. The tiles within the rectangular stripe are arranged in the order of the raster scan of the stripe. Figure 3 An example of raster scan strip segmentation of an image is shown, where the image is divided into 12 slices and 3 raster scan strips.
[0060] Figure 4 An example of rectangular strip segmentation of an image is shown, where the image is divided into 24 slices (6 slice columns and 4 slice rows) and 9 rectangular strips.
[0061] Figure 5 An example of an image divided into slices, bricks, and rectangular strips is shown, where the image is divided into 4 slices (2 slice columns and 2 slice rows), 11 bricks (the top left slice contains 1 brick, the top right slice contains 5 bricks, the bottom left slice contains 2 bricks, and the bottom right slice contains 3 bricks) and 4 rectangular strips.
[0062] 3.3.1 CTU / CTB Dimensions
[0063] In VVC, the CTU size for signal transmission in the Sequence Parameter Set (SPS) can be as small as 4×4 using the syntax element log2_ctu_size_minus2. 7.3.2.3 Sequence Parameter Set (RBSP) Syntax
[0064]
[0065]
[0066] `log2_ctu_size_minus2` plus 2 specifies the luma codec block size for each CTU. `log2_min_luma_coding_block_size_minus2` plus 2 specifies the minimum luma codec block size. The variables `CtbLog2SizeY`, `CtbSizeY`, `MinCbLog2SizeY`, `MinCbSizeY`, `MinTbLog2SizeY`, `MaxTbLog2SizeY`, `MinTbSizeY`, `MaxTbSizeY`, `PicWidthInCtbsY`, `PicHeightInCtbsY`, `PicSizeInCtbsY`, `PicWidthInMinCbsY`, `PicHeightInMinCbsY`, `PicSizeInMinCbsY`, `PicSizeInSamplesY`, `PicWidthInSamplesC`, and `PicHeightInSamplesC` are derived as follows:
[0067] CtbLog2SizeY = log2_ctu_size_minus2 + 2 (7-9) CtbSizeY = 1 << CtbLog2SizeY (7-10)
[0068] MinCbLog2SizeY = log2_min_luma_coding_block_size_minus2 + 2 (7-11)
[0069] MinCbSizeY = 1 << MinCbLog2SizeY (7-12)
[0070] MinTbLog2SizeY=2 (7-13)
[0071] MaxTbLog2SizeY = 6 (7-14)
[0072] MinTbSizeY = 1 << MinTbLog2SizeY (7-15)
[0073] MaxTbSizeY = 1 << MaxTbLog2SizeY (7-16)
[0074] PicWidthInCtbsY = Ceil( pic_width_in_luma_samples ÷ CtbSizeY ) (7-17)
[0075] PicHeightInCtbsY = Ceil( pic_height_in_luma_samples ÷ CtbSizeY ) (7-18)
[0076] PicSizeInCtbsY = PicWidthInCtbsY * PicHeightInCtbsY (7-19)
[0077] PicWidthInMinCbsY = pic_width_in_luma_samples / MinCbSizeY (7-20)
[0078] PicHeightInMinCbsY = pic_height_in_luma_samples / MinCbSizeY (7-21)
[0079] PicSizeInMinCbsY = PicWidthInMinCbsY * PicHeightInMinCbsY (7-22)
[0080] PicSizeInSamplesY= pic_width_in_luma_samples*pic_height_in_luma_samples (7-23)
[0081] PicWidthInSamplesC = pic_width_in_luma_samples / SubWidthC (7-24)
[0082] PicHeightInSamplesC = pic_height_in_luma_samples / SubHeightC (7-25)
[0083] 3.3.2 CTUs in a Picture
[0084] Assume that the CTB / largest coding unit (LCU) size is indicated by M×N (usually M equals N), and for a CTB located at the picture boundary (or slice or strip or other type of boundary, taking the picture boundary as an example), K×L samples are within the picture boundary, where K < M or L < N. For example Figure 6A 、6B As shown in 6C, the CTB size is still equal to M×N; however, the lower / right boundary of the CTB is outside the image.
[0085] 3.4 Intra-frame prediction
[0086] To capture arbitrary edge directions presented in natural video, the number of directional intra-frame modes was expanded from 33 used in HEVC to 65. The expanded directional modes are as follows: Figure 7 As shown, the planar and DC modes remain unchanged. These denser directional intra-prediction modes are applicable to all block sizes as well as luma and chroma intra-prediction.
[0087] like Figure 7 As shown, the angular intra-prediction direction can be defined as clockwise from 45 degrees to -135 degrees. In VTM, for non-square blocks, several angular intra-prediction modes are adaptively replaced by wide-angle intra-prediction modes. The replaced modes are transmitted via signaling and remapped to the wide-angle mode index after resolution. The total number of intra-prediction modes remains unchanged, for example, 67, and the intra-mode encoding and decoding remain unchanged.
[0088] In HEVC, each intra-coded block has a square shape, and the length of each side of the block is a power of 2. Therefore, no division operation is needed to generate intra-prediction values using DC mode. In VVC, blocks can have a rectangular shape, which generally requires division for each block. To avoid division for DC prediction, only the longer side is used to calculate the average of non-square blocks.
[0089] 3.5 Inter-frame prediction
[0090] For each inter-frame predicted CU, motion parameters include motion vectors, reference picture indexes, reference picture list usage indexes, and extended information for new encoding / decoding features of the VVC generated from the samples used in the inter-frame prediction. Motion parameters can be transmitted via signaling in an explicit or implicit manner. When a CU is encoded / decoded in skip mode, the CU is associated with a PU and does not have significant residual coefficients, encoded / decoded motion vector increments, and / or reference picture indexes. Merge mode is defined as the motion parameters of the current CU being obtained from neighboring CUs, including spatial and temporal candidates and extended scheduling introduced in the VVC. Merge mode can be applied to any inter-frame predicted CU, not just skip mode. An alternative to Merge mode is explicit transmission of motion parameters, where the motion vector for each reference picture list, the corresponding reference picture index, the reference picture list usage flag, and other useful information are explicitly transmitted via signaling for each CU.
[0091] 3.6 Deblocking Filter
[0092] Deblocking filtering is an example of a loop filter in a video codec. In VVC, the deblocking filtering process is applied to CU boundaries, transform subblock boundaries, and predictive subblock boundaries. Predictive subblock boundaries include prediction unit boundaries introduced by subblock-based temporal motion vector prediction (SbTMVP) and affine modes. Transform subblock boundaries include transform unit boundaries introduced by subblock transform (SBT) and intra-fractional sub-segmentation (ISP) modes, as well as transforms due to the implicit partitioning of large CUs. The processing order of the deblocking filter is defined as first performing horizontal filtering on the vertical edges of the entire image, and then performing vertical filtering on the horizontal edges. This specific order allows multiple horizontal or vertical filtering processes to be applied in parallel threads. The filtering process can also be implemented on a CTB-by-CTB basis with only a small processing latency.
[0093] The vertical edges in the image are first filtered. Then, the horizontal edges in the image are filtered using samples modified by the vertical edge filtering process as input. Vertical and horizontal edges in the CTB of each CTU are processed individually on a per-code / decode unit basis. The vertical edges of the codec blocks in a codec unit are filtered, starting from the edge on the left-hand side of the codec block and proceeding geometrically through the edges towards the right-hand side of the codec block. The horizontal edges of the codec blocks in a codec unit are filtered, starting from the edge on the top of the codec block and proceeding geometrically through the edges towards the bottom of the codec block.
[0094] 3.6.1 Boundary Decision
[0095] Filtering is applied to 8×8 block boundaries. Alternatively, such boundaries are preferably transform block boundaries or codec sub-block boundaries, such as those resulting from the use of affine motion prediction (ATMVP). For other boundaries, deblocking filtering is disabled.
[0096] 3.6.2 Boundary Strength Calculation
[0097] For transform block boundaries / encoder / decoder sub-block boundaries, if the boundary is located in an 8×8 grid, the boundary can be filtered, and the bS[xDi][yDj] (where [xDi][yDj] represents the coordinates) for that edge is set as defined in Tables 2 and 3, respectively.
[0098]
[0099] Table 2 Boundary Strengths (when SPS IBC is disabled)
[0100]
[0101] Table 3 Boundary Strengths (When SPS IBC is Enabled)
[0102] 3.6.3 Deblocking Decision for Luminance Component
[0103] A wider and stronger brightness filter is used only when conditions 1, 2, and 3 are all true. Condition 1 is the "bulk condition." This condition detects whether samples on the P-side and Q-side belong to a bulk, and is represented by the variables bSidePisLargeBlk and bSideQisLargeBlk, respectively. bSidePisLargeBlk and bSideQisLargeBlk are defined as follows.
[0104] bSidePisLargeBlk = ((Edge type is vertical and p0 belongs to CU with width >= 32) || (Edge type is horizontal and p0 belongs to CU with height >= 32)) ? True: False
[0105] bSideQisLargeBlk = ((Edge type is vertical and q0 belongs to CU with width >= 32) || (Edge type is horizontal and q0 belongs to CU with height >= 32)) ? True: False
[0106] Based on bSidePisLargeBlk and bSideQisLargeBlk, condition 1 is defined as follows:
[0107] Condition 1 = (bSidePisLargeBlk || bSidePisLargeBlk) ? True: False
[0108] Next, if condition 1 is true, condition 2 will be further examined. First, the following variables are derived:
[0109] First, derive dp0, dp3, dq0, and dq3 using the HEVC method.
[0110]
[0111] Where d = dp0 + dq0 + dp3 + dq3.
[0112] If conditions 1 and 2 are valid, then further check whether any of the blocks uses a sub-block:
[0113]
[0114]
[0115] Finally, if both conditions 1 and 2 are valid, the deblocking method will check condition 3 (the strong filter condition), which is defined as follows. In condition 3 (StrongFilterCondition), the following variables are derived:
[0116]
[0117] 3.6.4 A more robust deblocking filter for luminance
[0118] A bilinear filter is used when a sample on either side of the boundary belongs to a large block. A sample is defined as belonging to a large block when the width of the vertical edge is >= 32 and when the height of the horizontal edge is >= 32. The bilinear filter is listed below. Then, the block boundary samples pi (i = 0 to Sp-1) and qi (j = 0 to Sq-1) in the above HEVC deblocking (pi and qi are the i-th sample in the row used for filtering the vertical edge, or the i-th sample in the column used for filtering the horizontal edge) are replaced by the following linear interpolation:
[0119] p i ′=(fi i *Middle s,t +(64-fi i )*P s +32)>>6), Limit to p i ±tcPD i
[0120] q j ′=(g j *Middle s,t +(64-g j )*Q s +32)>>6), Limit to q j ±tcPD j
[0121] tcPD i and tcPD j The term is the aforementioned position-related clipping, and g j f i Middle s,t P s and Q s The following is given.
[0122] 3.6.5 Deblocking Decision for Chroma
[0123] A strong chromaticity filter is applied to both sides of the block boundary. Here, the chromaticity filter is selected when the chromaticity position on both sides of the chromaticity edge is greater than or equal to 8, and the following decision is satisfied with three conditions: The first is the decision for boundary strength and the size of the block. The filter can be applied when the width or height of the block orthogonally passing through the block edge is equal to or greater than 8 in the chromaticity sample domain. The second and third decisions are essentially the same as the decisions for deblocking HEVC luminance, which are the on / off decision and the strong filter decision, respectively.
[0124] In the first decision, the boundary strength (bS) is modified for chroma filtering, and conditions are checked sequentially. If a condition is met, the remaining conditions with lower priority are skipped. Chroma deblocking is performed when bS equals 2, or when bS equals 1 when a large block boundary is detected. The second and third conditions are essentially the same as the HEVC luma strong filter decision below.
[0125] In the second condition, d is derived using the HEVC luminance deblocking method. The second condition will be true when d is less than β. In the third condition, StrongFilterCondition is derived as follows:
[0126] Derive dpq using the HEVC method.
[0127] Derive sp3 = Abs(p3-p0) using the HEVC method.
[0128] Derive sq3 = Abs(q0-q3) using the HEVC method.
[0129] According to the HEVC design, StrongFilterCondition = (dpq < (β >> 2), sp3 + sq3 < (β >> 3), and Abs(p0 - q0) < (5 * tC + 1) >> 1).
[0130] 3.6.6 Strong Deblocking Filter for Chroma
[0131] The following strong deblocking filter is defined for chroma:
[0132] p2′=(3*p3+2*p2+p1+p0+q0+4)>>3
[0133] p1′=(2*p3+p2+2*p1+p0+q0+q1+4)>>3
[0134] p0′=(p3+p2+p1+2*p0+q0+q1+q2+4)>>3
[0135] The example chroma filter performs deblocking on a 4×4 chroma sample grid.
[0136] 3.6.7 Location-related amplitude limiting
[0137] Position-dependent limiting (tcPD) is applied to the output samples of strong and long filters that involve modifying 7, 5, and 3 samples at the boundaries. Assuming a quantization error distribution, the limiting value can be increased for samples expected to have higher quantization noise, thus expecting a larger deviation between the reconstructed sample values and the true sample values.
[0138] For each P or Q boundary filtered by an asymmetric filter, the location-related threshold table is selected from two tables (e.g., Tc7 and Tc3 listed below) that serve as edge information, based on the results of the decision process:
[0139] Tc7={6,5,4,3,2,1,1}; Tc3={6,4,2};
[0140] tcPD=(Sp==3)? Tc3:Tc7;
[0141] tcQD=(Sq==3)? Tc3:Tc7;
[0142] For P or Q boundaries filtered by a short symmetric filter, apply a lower-amplitude position correlation threshold:
[0143] Tc3 = {3, 2, 1};
[0144] After defining the threshold, the filtered p'i and q'i sample values are clipped according to the tcP and tcQ clipping values:
[0145] p”i=Clip3(p’i+tcPi,p’i–tcPi,p’i);
[0146] q”j=Clip3(q’j+tcQj,q’j–tcQj,q’j);
[0147] Where p'i and q'j are the filtered sample values, p”i and q”j are the output sample values after clipping, and tcPitcPi is the clipping threshold derived from the VVC tc parameters, tcPD, and tcQD. The function Clip3 is the clipping function as specified by VVC.
[0148] 3.6.8 Sub-block Removal and Adjustment
[0149] To enable parallel, user-friendly deblocking using both long filters and sub-block deblocking, the long filter is restricted to modifying a maximum of 5 samples on the side using sub-block deblocking (AFFINE, ATMVP, or DMVR), as shown in the long filter's brightness control. Extended, sub-block deblocking is adjusted such that sub-block boundaries on the 8×8 grid near the CU or implicit TU boundaries are restricted to modifying a maximum of two samples on each side.
[0150] The following applies to sub-block boundaries that are not aligned with the CU boundary.
[0151]
[0152] Where edge = 0 corresponds to the CU boundary, edge = 2 or orthogonalLength-2 corresponds to the sub-block boundary 8 samples away from the CU boundary, etc. If implicit partitioning of TU is used, then implicit TU is true.
[0153] 3.7 Sample point adaptive compensation
[0154] SAO is applied to the reconstructed signal after the deblocking filter using compensation specified by the encoder for each CTB. The video encoder first decides whether to apply the SAO process to the current slice. If SAO is applied to a slice, each CTB is classified into one of five SAO types as shown in Table 4. The concept of SAO is to classify pixels into multiple categories and reduce distortion by adding compensation to pixels in each category. SAO operations include Edge Compensation (EO) and Band Compensation (BO), where EO uses edge attributes to classify pixels in SAO types 1 through 4, and BO uses pixel intensity to classify pixels in SAO type 5. Each applicable CTB has SAO parameters including sao_merge_left_flag, sao_merge_up_flag, SAO type, and four compensations. If sao_merge_left_flag equals 1, the current CTB will reuse the SAO type and compensation of the left CTB. If sao_merge_up_flag equals 1, the current CTB will reuse the SAO type and compensation of the upper CTB.
[0155] SAO type The type of sample adaptive compensation to be used Number of categories 0 none 0 1 1-D 0-degree pattern edge compensation 4 2 1-D 90-degree pattern edge compensation 4 3 1-D 135-degree pattern edge compensation 4 4 1-D 45-degree pattern edge compensation 4 5 With compensation 4
[0156] Table 4. Specifications for SAO Types
[0157] 3.8 Adaptive Loop Filter
[0158] Adaptive Loop Filtering (ALF) for video encoding and decoding minimizes the mean square error between the original and decoded samples using Wiener-based adaptive filters. ALF is located at the last processing stage of each picture and can be considered a tool for capturing and repairing artifacts from previous stages. Appropriate filter coefficients are determined by the encoder and explicitly transmitted to the decoder via the signal. To achieve better encoding and decoding efficiency, especially for high-resolution video, local adaptation is used for the luminance signal by applying different filters to different regions or blocks in the picture. In addition to filter adaptation, filter on / off control at the codec tree unit (CTU) level also contributes to improved encoding and decoding efficiency. Syntactically, filter coefficients are sent in a picture-level header called the adaptive parameter set, and the filter on / off flags of the CTUs are interleaved at the CTU level in the striped data. This syntax design not only supports picture-level optimization but also achieves low encoding latency.
[0159] 3.8.1 Signaling of Parameters
[0160] According to the ALF design in VTM, filter coefficients and clipping indices are carried in the ALF APS. The ALF APS can include up to eight chroma filters and a luma filter set with up to 25 filters. An index is also included for each of the 25 luma categories. Categories with the same index share the same filters. By merging different categories, the number of bits required to represent the filter coefficients is reduced. The absolute values of the filter coefficients are represented using 0th-order Exp-Golomb code followed by a sign bit for non-zero coefficients. When clipping is enabled, a two-bit fixed-length code is also used for each filter coefficient transmitted via the signal, providing a clipping index. The decoder can use up to eight ALF APSs simultaneously.
[0161] The ALF filter control syntax elements in VTM include two types of information. First, the ALF on / off flag is transmitted via signaling at the sequence, picture, strip, and CTB levels. Chroma ALF can only be enabled at the picture and strip levels if Luminance ALF is enabled at the corresponding level. Second, if ALF is enabled at the picture, strip, and CTB levels, filter usage information is transmitted via signaling at that level. If all strips within a picture use the same APS, the referenced ALF APS ID is encoded / decoded at the strip or picture level. Luminance components can reference up to 7 ALF APSs, and chroma components can reference 1 ALF APS. For Luminance CTB, an index is transmitted via signaling indicating which ALF APS or offline-trained Luminance filter set is used. For Chroma CTB, an index indicates which filter in the referenced APS is used.
[0162] The data syntax elements of the ALF associated with the luminance component in VTM are listed below:
[0163]
[0164] `alf_luma_filter_signal_flag` equal to 1 specifies that the luminance filter set is transmitted via signal. `alf_luma_filter_signal_flag` equal to 0 specifies that the luminance filter set is not transmitted via signal. `alf_luma_clip_flag` equal to 0 specifies that linear adaptive loop filtering is applied to the luminance component. `alf_luma_clip_flag` equal to 1 specifies that nonlinear adaptive loop filtering can be applied to the luminance component. `alf_luma_num_filters_signalled_minus1` plus 1 specifies the number of adaptive loop filter classes whose luminance coefficients can be transmitted via signal. The value of `alf_luma_num_filters_signalled_minus1` must be in the range of 0 to NumAlfFilters-1 (inclusive). `alf_luma_coeff_delta_idx[filtIdx]` specifies the index of the adaptive loop filter luminance coefficient increment of the filter class transmitted via signal, indicated by `filtIdx` in the range of 0 to NumAlfFilters-1. When alf_luma_coeff_delta_idx[filtIdx] does not exist, it is presumed to be equal to 0. The length of alf_luma_coeff_delta_idx[filtIdx] is Ceil(Log2(alf_luma_num_filters_signalled_minus1+1)) bits. The value of alf_luma_coeff_delta_idx[filtIdx] must be in the range from 0 to alf_luma_num_filters_signalled_minus1 (inclusive).
[0165] `alf_luma_coeff_abs[sfIdx][j]` specifies the absolute value of the j-th coefficient of the luminance filter transmitted through the signal, indicated by `sfIdx`. When `alf_luma_coeff_abs[sfIdx][j]` does not exist, it is presumed to be equal to 0. The value of `alf_luma_coeff_abs[sfIdx][j]` must be in the range of 0 to 128 (inclusive). `alf_luma_coeff_sign[sfIdx][j]` specifies the sign of the j-th luminance coefficient of the filter, indicated by `sfIdx`, as follows:
[0166] If alf_luma_coeff_sign[sfIdx][j] equals 0, the corresponding luminance filter coefficient has a positive value.
[0167] Otherwise (alf_luma_coeff_sign[sfIdx][j] equals 1), the corresponding luminance filter coefficient has a negative value.
[0168] When alf_luma_coeff_sign[sfIdx][j] does not exist, it is presumed to be equal to 0.
[0169] `alf_luma_clip_idx[sfIdx][j]` specifies the limiting index to be used before multiplying by the j-th coefficient of the luminance filter transmitted through the signal, indicated by `sfIdx`. When `alf_luma_clip_idx[sfIdx][j]` does not exist, it is presumed to be equal to 0. The codec tree unit syntax elements of the ALF associated with the luminance component in the VTM are listed below:
[0170]
[0171] `alf_ctb_flag[cIdx][xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY]` equal to 1 indicates that the adaptive loop filter is applied to the codec tree block of the codec tree unit at the luma position (xCtb, yCtb) for the color component indicated by `cIdx`. `alf_ctb_flag[cIdx][xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY]` equal to 0 indicates that the adaptive loop filter is not applied to the codec tree block of the codec tree unit at the luma position (xCtb, yCtb) for the color component indicated by `cIdx`.
[0172] When `alf_ctb_flag[cIdx][xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY]` does not exist, it is presumed to be equal to 0. `alf_use_aps_flag` equal to 0 specifies that one of the fixed filter sets is applied to the luma CTB. `alf_use_aps_flag` equal to 1 specifies that a filter set from the APS is applied to the luma CTB. When `alf_use_aps_flag` does not exist, it is presumed to be equal to 0. `alf_luma_prev_filter_idx` specifies the previous filter applied to the luma CTB. The value of `alf_luma_prev_filter_idx` must be in the range of 0 to `sh_num_alf_aps_ids_luma-1` (inclusive). When `alf_luma_prev_filter_idx` does not exist, it is presumed to be equal to 0.
[0173] The variable AlfCtbFiltSetIdxY[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY], which represents the filter set index of the luminance CTB at a specified location (xCtb, yCtb), is derived as follows:
[0174] If alf_use_aps_flag equals 0, AlfCtbFiltSetIdxY[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] is set to equal alf_luma_fixed_filter_idx.
[0175] Otherwise, AlfCtbFiltSetIdxY[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] is set to equal 16 + alf_luma_prev_filter_idx.
[0176] alf_luma_fixed_filter_idx specifies the fixed filter to be applied to the luminance CTB. The value of alf_luma_fixed_filter_idx must be in the range of 0 to 15 (inclusive).
[0177] Building upon the ALF design of VTM, the ALF design of ECM further introduces the concept of candidate filter sets into the luminance filter. The luminance filter is trained with multiple candidates / rounds based on the updated luminance CTU ALF on / off decisions for each candidate / round. This results in multiple filter sets associated with each trained candidate, and the class merging results for each filter set can differ. Each CTU can select the optimal filter set via RDO, and the associated candidate information is transmitted via signal transmission. The data syntax elements of the ALF associated with the luminance component in ECM are listed below:
[0178]
[0179]
[0180] `alf_luma_num_alts_minus1` plus 1 specifies the number of candidate filter sets for the luminance component. The value of `alf_luma_num_alts_minus1` must be in the range of 0 to 3 (inclusive). `alf_luma_clip_flag[altIdx]` equal to 0 specifies that linear adaptive loop filtering is applied to the luminance component of the candidate luminance filter set with index `altIdx`. `alf_luma_clip_flag[altIdx]` equal to 1 specifies that nonlinear adaptive loop filtering can be applied to the luminance component of the candidate luminance filter set with index `altIdx`. `alf_luma_num_filters_signalled_minus1[altIdx]` plus 1 specifies the number of adaptive loop filter classes for which the luminance coefficients can be transmitted via signal transmission, the candidate luminance filter set with index `altIdx`. The value of alf_luma_num_filters_signalled_minus1[altIdx] must be in the range of 0 to NumAlfFilters-1 (inclusive).
[0181] `alf_luma_coeff_delta_idx[altIdx][filtIdx]` specifies an index of the adaptive loop filter luminance coefficient increments transmitted through the signal, indicated by `filtIdx` ranging from 0 to `NumAlfFilters–1`, for the set of candidate luminance filters with index `altIdx`. When `alf_luma_coeff_delta_idx[filtIdx][altIdx]` does not exist, it is presumed to be equal to 0. The length of `alf_luma_coeff_delta_idx[altIdx][filtIdx]` is Ceil(Log2(alf_luma_num_filters_signalled_minus1[altIdx]+1)) bits. The value of `alf_luma_coeff_delta_idx[altIdx][filtIdx]` must be in the range of 0 to `alf_luma_num_filters_signalled_minus1[altIdx]` (inclusive). `alf_luma_coeff_abs[altIdx][sfIdx][j]` specifies the absolute value of the j-th coefficient of the luminance filter transmitted through the signal, indicated by the `sfIdx` of the alternative luminance filter set with index `altIdx`. When `alf_luma_coeff_abs[altIdx][sfIdx][j]` does not exist, it is presumed to be equal to 0. The value of `alf_luma_coeff_abs[altIdx][sfIdx][j]` must be in the range of 0 to 128 (inclusive).
[0182] alf_luma_coeff_sign[altIdx][sfIdx][j] specifies the sign of the j-th luminance coefficient of the filter indicated by the sfIdx of the set of candidate luminance filters with index altIdx, as follows:
[0183] If alf_luma_coeff_sign[altIdx][sfIdx][j] equals 0, the corresponding luminance filter coefficient has a positive value.
[0184] Otherwise (alf_luma_coeff_sign[altIdx][sfIdx][j] equals 1), the corresponding luminance filter coefficient has a negative value.
[0185] When alf_luma_coeff_sign[altIdx][sfIdx][j] does not exist, it is presumed to be equal to 0.
[0186] `alf_luma_clip_idx[altIdx][sfIdx][j]` specifies the limiting index to be used before multiplying the limiting value by the j-th coefficient of the luminance filter transmitted through the signal, indicated by `sfIdx` of the candidate luminance filter set with index `altIdx`. When `alf_luma_clip_idx[altIdx][sfIdx][j]` does not exist, it is presumed to be equal to 0. The codec tree unit syntax elements of the ALF associated with the luminance component in the ECM are listed below:
[0187]
[0188] `alf_ctb_luma_filter_alt_idx[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY]` specifies the index of the alternative luma filter for the codec tree block of the luma component applied to the luma location (xCtb, yCtb). When `alf_ctb_luma_filter_alt_idx[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY]` does not exist, it is presumed to be equal to 0.
[0189] 3.8.2 Filter Shape
[0190] In JEM, there are a maximum of three diamond filter shapes (e.g., Figure 10 The squares (shown) can be selected for the luma component. At the image level, the filter shape used for the luma component is indicated by a signal transmission index. Each square represents a sample, and Ci (i = 0–6 (left), 0–12 (middle), 0–20 (right)) represents the coefficient to be applied to the sample. For the chroma component in the image, a 5×5 rhombus shape is used. In VVC, a 7×7 rhombus shape is used for luma, while a 5×5 rhombus shape is used for chroma.
[0191] 3.8.3 Classification of ALF
[0192] Each 2×2 (or 4×4) block is classified into one of 25 categories. The classification index C is based on its directionality D and activity. The quantization value is derived as follows:
[0193]
[0194] To calculate D and First, the gradients in the horizontal, vertical, and two diagonal directions are calculated using 1-D Laplacian:
[0195]
[0196] Indices i and j refer to the coordinates of the top-left sample point in the 2×2 block, and r(i,j) indicates the reconstructed sample point at coordinates (i,j). Then, the maximum and minimum values of the gradients D in the horizontal and vertical directions are set as follows:
[0197]
[0198] Furthermore, the maximum and minimum values of the gradients in the two diagonal directions are set as follows:
[0199]
[0200] To derive the values of directionality D, these values are compared with each other and with two thresholds t1 and t2:
[0201] Step 1. If and If both are true, then D is set to 0.
[0202] Step 2. If If so, continue from step 3; otherwise, continue from step 4.
[0203] Step 3. If Then D is set to 2; otherwise, D is set to 1.
[0204] Step 4. If Then D is set to 4; otherwise, D is set to 3.
[0205] The activity value A is calculated as:
[0206]
[0207] A is further quantized to the range of 0 to 4 (inclusive), and the quantized value is represented as No classification method was applied to the two chromaticity components in the image; that is, a single set of ALF coefficients was applied to each chromaticity component.
[0208] 3.8.4 Geometric Transformation of Filter Coefficients
[0209] Before filtering each 2×2 block, geometric transformations (such as rotation or diagonal and vertical flips) are applied to the filter coefficients f(k,l) associated with the coordinates (k,l), based on the gradient values calculated for that block. This is equivalent to applying these transformations to samples in the filter's support region. The idea is to make different blocks to which ALF is applied more similar by aligning their orientations.
[0210] Three geometric transformations are introduced, including diagonal flip, vertical flip, and rotation:
[0211] Diagonal: f D (k,l)=f(l,l),
[0212] Vertical flip: f V (k,l)=f(k,Kl-1),
[0213] Rotation: f R (k,l)=f(Kl-1,k).
[0214] Where K is the size of the filter, and 0 ≤ k, l ≤ K⁻¹ are the coefficient coordinates, such that position (0, 0) is in the upper left corner and position (K⁻¹, K⁻¹) is in the lower right corner. The transform is applied to the filter coefficients f(k, l) based on the gradient values calculated for this block. The relationship between the transform and the four gradients in the four directions is summarized in Table 5. Figure 11 The transformed coefficients are shown for each position based on a 5×5 rhombus.
[0215] gradient value Transformation <![CDATA[g d2 <g d1 And g h <g v ]]> No transformation <![CDATA[g d2 <g d1 And g v <g h ]]> diagonal <![CDATA[g d1 <g d2 And g h <g v ]]> Vertical flip <![CDATA[g d1 <g d2 And g v <g h ]]> Rotation
[0216] Table 5 shows the mapping between gradients and transformations for a single block of computation.
[0217] 3.8.5 Filtering Process
[0218] On the decoder side, when ALF is enabled for a block, each sample R(i,j) within the block is filtered, resulting in sample values R′(i,j) as shown below, where L represents the filter length, f m,n Let f(k,l) represent the filter coefficients, and let f(k,l) represent the decoded filter coefficients.
[0219]
[0220] Figure 12 This example illustrates relative coordinates used in a 5×5 diamond filter, assuming the current sample's coordinates (i, j) are (0, 0). Samples at different coordinates, filled with the same color, are multiplied by the same filter coefficients.
[0221] 3.8.6 Nonlinear Filtering Reconstruction
[0222] Linear filtering can be reconstructed into the following expression without affecting encoding / decoding efficiency:
[0223]
[0224] Where w(i,j) are the same filter coefficients.
[0225] VVC introduces nonlinearity by using a simple limiting function to reduce the impact of neighboring sample values when the difference between the neighboring sample value (I(x+i,y+j)) and the current sample value being filtered (I(x,y)) is too large, thus making ALF more efficient. More specifically, the ALF filter is modified as follows:
[0226]
[0227] Where K(d,b) = min(b,max(-b,d)) is the limiting function, and k(i,j) is the limiting parameter, which depends on the (i,j) filter coefficients. The encoder performs optimization to find the optimal k(i,j).
[0228] A limiting parameter k(i,j) is specified for each ALF filter, and a limiting value is transmitted through the signal for each filter coefficient. This means that up to 12 limiting values and up to 6 limiting values for each luminance filter can be transmitted through the signal in the bitstream. To limit signaling costs and encoder complexity, only 4 fixed values are used, which are the same for both inter-frame and intra-frame stripes.
[0229] Because the variance of local differences in luminance is typically higher than that in chrominance, two different sets of filters are applied for luminance and chrominance. The maximum sample value in each set (here for a 10-bit bit depth of 1024) is also introduced so that clipping can be disabled if unnecessary. Four values were selected by dividing the entire range of luminance sample values (encoded and decoded on 10 bits) and the range of chrominance from 4 to 1024 into approximately equal parts in the logarithmic domain. More precisely, the luminance clipping value table has been obtained using the following formula:
[0230] Where M = 2 10 And N = 4
[0231] Similarly, the chromaticity limit table is obtained according to the following formula:
[0232] Where M = 2 10 N=4 and A=4
[0233] 3.9 Bilateral Loop Filter
[0234] 3.9.1 Bilateral Image Filter
[0235] Bilateral image filters are nonlinear filters that smooth noise while preserving edge structure. Bilateral filtering is a technique where the filter weights decrease not only with the distance between samples but also with the intensity difference. This improves the smoothing of overly smoothed edges. The weights are defined as follows:
[0236]
[0237] Where Δx and Δy are the vertical and horizontal distances, and ΔI is the intensity difference between the sample points.
[0238] The edge-preserving denoising bilateral filter employs low-pass Gaussian filters for both the spatial and domain filters. The domain low-pass Gaussian filter assigns higher weights to pixels spatially closer to the center pixel. The range low-pass Gaussian filter assigns higher weights to pixels similar to the center pixel. Combining the range and domain filters, the bilateral filter at edge pixels becomes a thin Gaussian filter, oriented along the edge and significantly reduced in the gradient direction. This is why the bilateral filter can smooth noise while preserving edge structure.
[0239] 3.9.2 Bilateral Filters in Video Encoding and Decoding
[0240] Bilateral filters in video encoding and decoding are encoding and decoding tools used for VVC [2]. The filters act as loop filters in parallel with SAO filters. Both bilateral filters and SAOs act on the same input samples, each filter produces compensation, and these compensations are then added to the input samples to produce output samples, which are then clipped before proceeding to the next stage. Spatial filtering intensity σ d Determined by the block size, smaller blocks are filtered more strongly, and the intensity of the filter is σ. r The strength of the filter is determined by the quantization parameters, where a stronger QP results in a higher filter strength. Only the four nearest samples are used, therefore the filtered sample strength I... F It can be calculated as
[0241]
[0242] Among them I C The intensity of the central sample point, ΔI A =I A -I C This represents the intensity difference between the central sample point and the sample point above it. ΔI B ΔI L andΔI R These represent the intensity differences between the central sample point and the sample points below, to the left, and to the right, respectively.
[0243] 4. The technical problem solved by the disclosed technical solution
[0244] The example design of the Adaptive Loop Filter (ALF) in video encoding and decoding has the following problems.
[0245] First, in the example ALF design, only the spatial reconstruction samples following other filters (such as DBF, SAO, and BF) are used for filter training and filtering. However, there is other valuable information that can potentially be utilized, such as the boundary strength used by DBF (DBF-BS).
[0246] Second, in the example ALF design, only the spatial reconstruction samples after other filters (such as DBF, SAO, and BF) are used as input for classification. However, there is other valuable information that can potentially be utilized, such as the boundary strength used by DBF (DBF-BS).
[0247] 5. List of solutions and implementation examples
[0248] To address the aforementioned problems, methods outlined below are disclosed. These embodiments should be considered as examples for explaining general concepts and should not be interpreted in a narrow sense. Furthermore, these embodiments can be applied individually or in any combination.
[0249] It should be noted that the disclosed method can be used as a loop filter or post-processing.
[0250] In this disclosure, a video unit may refer to a sequence, picture, subpicture, strip, CTU, block, and / or region. A video unit may include one or more color components.
[0251] In this disclosure, an ALF processing unit can refer to a sequence, image, sub-image, strip, CTU, block, region, or sample. An ALF processing unit may include one or more color components.
[0252] 1) The use of DBF-BS as edge information for ALF / CCALF is proposed.
[0253] a. In one example, a DBF-BS for at least one location may be stored during or after the DBF process.
[0254] a) In one example, at least one location of DBF-BS can be stored.
[0255] 1. In one example, the DBF-BS of the brightness position can be stored.
[0256] 2. In one example, the DBF-BS of the chromaticity position can be stored.
[0257] b) In one example, the DBF-BS of the location can be initialized to a specific value.
[0258] 1. In one example, the DBF-BS of the location can be initialized to 0.
[0259] 2. In one example, the DBF-BS of a location can be initialized to N (e.g., N = 32).
[0260] 3. In one example, the initial value can be predefined.
[0261] 4. In one example, the initial value can be derived.
[0262] 5. In one example, the initial value can be transmitted to the decoder via a signal.
[0263] c) In one example, the location's DBF-BS can be stored / used in the original domain.
[0264] 1. In one example, the DBF-BS of a location can be stored / used using default values representing different boundary strengths (e.g., BS = 0, 1, 2).
[0265] d) In one example, the DBF-BS for each location can be stored / used in the mapping domain.
[0266] 1. In one example, the DBF-BS at each location can be stored / used using mapping values representing different boundary strengths (e.g., BS = 128, 256, 512).
[0267] 2. In one example, the mapping value can be predefined.
[0268] 3. In one example, the mapping value can be derived.
[0269] 4. In one example, the mapped value can be transmitted to the decoder via a signal.
[0270] e) In one example, the DBF-BS of a location can be merged into one or more categories.
[0271] 1. In one example, the BDF-BS of a location can be merged into N categories (e.g., N=2).
[0272] 2. In one example, the DBF-BS of a location can be merged according to predefined rules.
[0273] a. In one example, the predefined rule could be a threshold-based limiting / mapping function.
[0274] b. In one example, it can be set to V0 when DBF-BS is greater than or equal to the threshold, and to V1 when it is less than the threshold.
[0275] 3. In one example, the number of merge rules or merge categories can be predefined.
[0276] 4. In one example, the number of merge rules or merge categories can be derived.
[0277] 5. In one example, the number of merge rules or merge categories can be transmitted to the decoder via a signal.
[0278] f) In one example, the DBF-BS of the location can be filtered before use.
[0279] 1. In one example, the DBF-BS of a location can be filtered by a predefined filter.
[0280] 2. In one example, the DBF-BS of a location can be filtered by an offline-trained filter of an ALF.
[0281] 3. In one example, the DBF-BS of a location can be filtered by an online-trained filter.
[0282] a. In one example, an online-trained filter can be transmitted to the decoder via a signal.
[0283] 4. In one example, the DBF-BS of the location can be filtered by any other filter.
[0284] a. In one example, the DBF-BS of a location can be filtered by a Gaussian filter.
[0285] b. In one example, the DBF-BS of a location can be filtered by a Sobel / Prewitt / Roberts / Canny filter.
[0286] c. In one example, the DBF-BS of a location can be filtered by a Hadamard transform domain filter.
[0287] d. In one example, the DBF-BS of a location can be filtered by a bilateral filter.
[0288] e. In one example, the DBF-BS of a location can be filtered by a low-pass filter.
[0289] f. In one example, the DBF-BS of a location can be filtered by a high-pass filter.
[0290] 5. In one example, whether to filter the DBF-BS or which filter to apply to the DBF-BS can be predefined.
[0291] 6. In one example, whether to filter the DBF-BS or which filter to apply to the DBF-BS can be derived.
[0292] 7. In one example, whether the DBF-BS is filtered or which filter is applied to the DBF-BS can be transmitted to the decoder via a signal.
[0293] g) In one example, DBF-BS can be clipped to a certain range / bit depth.
[0294] 1. In one example, the DBF-BS can be clipped to a predefined / transmitted / derived clipping range.
[0295] 2. In one example, the DBF-BS can be clipped to a predefined / transmitted / derived N-bit depth (e.g., N=10).
[0296] h) In one example, DBF-BS can be scaled to a certain range / bit depth.
[0297] 1. In one example, DBF-BS can be scaled to a predefined / signal-transmitted / derived range.
[0298] 2. In one example, DBF-BS can be scaled to a predefined / transmitted / derived N-bit depth (e.g., N=10).
[0299] i) In one example, DBF-BS can be converted to a certain range / domain / bit depth.
[0300] 1. In one example, DBF-BS can be converted to a predefined / signaled / derived range.
[0301] 2. In one example, DBF-BS can be converted to a predefined / signaled / derived domain.
[0302] 3. In one example, DBF-BS can be converted to a predefined / transmitted / derived N-bit depth (e.g., N=10).
[0303] j) In one example, the DBF-BS for at least one location can be derived without invoking the DBF procedure.
[0304] 1. For example, the DBF-BS of a location can be derived and stored, but that location may not be filtered by DBF.
[0305] 2) The use of DBF-BS as side information in ALF / CCALF filtering is proposed.
[0306] a. For example, f(Bs) can be used as side information in ALF / CCALF filtering, where Bs is the boundary strength and f is any function, such as a linear function or a limiting function or any other function.
[0307] b. In one example, DBF-BS can be used as input to ALF.
[0308] a) In one example, DBF-BS can be used as input to at least one extended tap of ALF.
[0309] b) In one example, DBF-BS can be used as at least one existing tap of ALF.
[0310] c) In one example, the DBF-BS of the luminance can be used as input for the ALF luminance.
[0311] d) In one example, the DBF-BS of the chromaticity can be used as input for the ALF chromaticity.
[0312] e) In one example, the filtered DBF-BS of the luminance can be used as input to the ALF luminance.
[0313] f) In one example, the filtered DBF-BS of the chroma can be used as input to the ALF chroma.
[0314] c. In one example, DBF-BS can be used as input for CCALF.
[0315] a) In one example, DBF-BS can be used as input to at least one extended tap of CCALF.
[0316] b) In one example, DBF-BS can be used as input to at least one existing tap of CCALF.
[0317] c) In one example, the luminance DBF-BS can be used as input to the extended tap of CCALF.
[0318] d) In one example, the filtered DBF-BS of the luminance can be used as input to CCALF.
[0319] d. In one example, DBF-BS can be used by following the ALF / CCALF filtering function.
[0320] a) In one example, the difference between the DBF-BS or filtered DBF-BS and the currently filtered sample may need to be calculated first.
[0321] b) In one example, the filter coefficients can be multiplied by the difference mentioned above to obtain the final offset.
[0322] e. In one example, DBF-BS can be used in a different way than the ALF / CCALF filtering function.
[0323] a) In one example, the DBF-BS or filtered DBF-BS value can be used directly.
[0324] b) In one example, the filter coefficients can be multiplied by DBF-BS to obtain the final offset.
[0325] 3) It is proposed to use DBF-BS as edge information in ALF classification.
[0326] a. For example, f(Bs) can be used as edge information in ALF classification, where Bs is the boundary strength and f is any function, such as a linear function or a limiting function or any other function.
[0327] b. In one example, DBF-BS / filtered DBF-BS can be used as an additional classification mode.
[0328] a) In one example, the classification unit size of a DBF-BS-based classifier can be N×M (e.g., N=M=2).
[0329] b) In one example, locations within a classification unit can share the same classification result.
[0330] c) In one example, the classification results can be generated / derived using a specific method.
[0331] 1. In one example, the method can be predefined.
[0332] 2. In one example, the method can be deduced.
[0333] 3. In one example, the method can be transmitted to the decoder via a signal.
[0334] 4. In one example, the method could be an averaging function.
[0335] 5. In one example, the method could be a maximum function.
[0336] 6. In one example, the method can be a minimal function.
[0337] 7. In one example, the method could be a linear mapping function.
[0338] 8. In one example, the method could be a non-linear mapping function.
[0339] 9. Alternatively, the method can be any other function.
[0340] d) In one example, whether or not a DBF-BS-based classifier is used can be predefined / derived / via signaling.
[0341] c. In one example, DBF-BS / filtered DBF-BS can be used as side information for existing classification patterns.
[0342] a) In one example, locations within a classification unit can share the same DBF-BS value.
[0343] b) In one example, the DBF-BS value within a classification unit can be generated / derived using a specific method.
[0344] 1. In one example, the method can be predefined.
[0345] 2. In one example, the method can be deduced.
[0346] 3. In one example, the method can be transmitted to the decoder via a signal.
[0347] 4. In one example, the method could be an averaging function.
[0348] 5. In one example, the method could be a maximum function.
[0349] 6. In one example, the method can be a minimal function.
[0350] 7. In one example, the method could be a linear mapping function.
[0351] 8. In one example, the method could be a non-linear mapping function.
[0352] 9. Alternatively, the method can be any other function.
[0353] c) In one example, DBF-BS can be merged into K categories (e.g., K=2).
[0354] d) In one example, an existing classifier may have N classes (e.g., N = 25).
[0355] e) In one example, the merged DBF-BS can be used directly to expand the number of categories.
[0356] 1. In one example, the number of classes in an existing classifier can be extended to N×K (e.g., N×K = 50).
[0357] 2. In one example, the classification result can be calculated using the following formula:
[0358] c = c0 + BS × N
[0359] Where c is the final category index, and c0 is the category index generated by the existing classifier.
[0360] 3. In one example, the classification result can be calculated using the following formula:
[0361] c = c0 × K + BS
[0362] Where c is the final category index, and c0 is the category index generated by the existing classifier.
[0363] f) In one example, the merged DBF-BS can be used to modify the number of categories.
[0364] 1. In one example, the number of classes in an existing classifier can be modified to M (e.g., M = 12).
[0365] a. In one example, for a band-based classifier, the total number of bands can be modified.
[0366] b. In one example, for a texture-based classifier, the number of texture directions can be modified.
[0367] c. In one example, for a texture-based classifier, the number of texture activities can be modified.
[0368] 2. In one example, the number of classes in an existing classifier can be modified to N×K (e.g., M×K = 24).
[0369] 3. In one example, the classification result can be calculated using the following formula:
[0370] c = c0 + BS × M
[0371] Where c is the final category index, and c0 is the category index generated by the existing classifier.
[0372] 4. In one example, the classification result can be calculated using the following formula:
[0373] c = c0 × K + BS
[0374] Where c is the final category index, and c0 is the category index generated by the existing classifier.
[0375] 4) It is proposed to use DBF-BS as side information in other encoding and decoding stages.
[0376] a. In one example, DBF-BS can be used in a bilateral filter.
[0377] a) In one example, DBF-BS can be used in the classification of bilateral filters.
[0378] b) In one example, different DBF-BS values can result in different filter strengths for the bilateral filter.
[0379] b. In one example, DBF-BS can be used in an HTDF filter.
[0380] a) In one example, DBF-BS can be used in the classification of HTDF filters.
[0381] b) In one example, different DBF-BS values can result in different filter strengths for the HTDF filter.
[0382] c. In one example, DBF-BS can be used in SAO.
[0383] a) In one example, DBF-BS can be used in the SAO classification.
[0384] 1. In one example, DBF-BS can be used as an additional classification pattern.
[0385] 2. In one example, DBF-BS can be used as side information for an existing classification pattern.
[0386] b) In one example, DBF-BS can be used to generate the offset of SAO.
[0387] d. In one example, DBF-BS can be used in CCSAO.
[0388] a) In one example, DBF-BS can be used in the CCSAO classification.
[0389] 1. In one example, DBF-BS can be used as an additional classification pattern.
[0390] 2. In one example, DBF-BS can be used as side information for an existing classification pattern.
[0391] b) In one example, DBF-BS can be used to generate the offset of CCSAO.
[0392] e. In one example, DBF-BS can be used in any other codec tool / stage.
[0393] 5) In one example, whether / how to modify / prepare DBF-BS can be transmitted via signal in a bitstream.
[0394] a. In one example, the first syntax element can be signaled to indicate whether the modified / prepared DBF-BS is enabled / used.
[0395] a) In one example, the first syntax element can be encoded or decoded using arithmetic encoding / decoding.
[0396] 1. In one example, the first syntax element can be encoded or decoded using at least one context.
[0397] a. The context can depend on the encoding / decoding information of the current block or neighboring blocks.
[0398] b. The context may depend on the filtering shape of at least one neighboring block.
[0399] 2. In one example, the first syntax element can be encoded or decoded using bypass encoding / decoding.
[0400] b) In one example, the first syntax element can be binarized using unary code, rounded unary code, fixed-length code, exponential Golomb code, rounded exponential Golomb code, etc.
[0401] c) In one example, the first syntax element can be conditionally transmitted via signal.
[0402] 1. For example, the first syntax element can only be transmitted via signaling if DBF-BS is available.
[0403] d) The first syntax element can be encoded and decoded in a predictive manner.
[0404] 1. The first syntax element can be predicted by the on / off decision of the DBF-BS of at least one neighboring block.
[0405] e) For different color components, the first syntax element can be transmitted independently via signal.
[0406] 1. Alternatively, for different color components, the first syntax element can be transmitted and shared via signal transmission.
[0407] 2. Alternatively, the first syntax element can be transmitted via signal for the first color component, but not for the second color component.
[0408] f) Syntax elements can be transmitted via signals in SPS / PPS / image header / strip header / APS / CTU / CU / etc.
[0409] 6) In one example, whether DBF-BS can be applied to ALF and / or CCALF may depend on the control information of DBF.
[0410] a. For example, if DBF is turned off, DBF-BS cannot be applied to ALF and / or CCALF.
[0411] 7) In one example, whether samples after / before DBF can be applied to ALF and / or CCALF may depend on the control information of DBF.
[0412] a. For example, if DBF is turned off, samples after DBF cannot be applied to ALF and / or CCALF.
[0413] b. For example, if DBF is turned off, samples prior to DBF cannot be applied to ALF and / or CCALF.
[0414] 8) In one example, the disclosed methods may be used in post-processing and / or pre-processing.
[0415] 9) DBF-BS can be in the same color component as the sample to be filtered, or in a different color component.
[0416] 10) DBF-BS can be located at the same position as the sample to be filtered, or within the range around the sample to be filtered.
[0417] 11) In one example, the above methods can be used in combination.
[0418] 12) Alternatively, the above methods can be used alone.
[0419] 13) In one example, the proposed method of using DBF-BS for ALF can be applied to any loop filtering tool, preprocessing or postprocessing filtering method in video encoding and decoding (including but not limited to ALF / CCALF / BF / SAO / CCSAO or any other filtering method).
[0420] a. In one example, the proposed method using DBF-BS can be applied to loop filtering methods.
[0421] a) In one example, the proposed method using DBF-BS can be applied to ALF.
[0422] b) In one example, the proposed method using DBF-BS can be applied to CCALF.
[0423] c) Alternatively, the proposed method using DBF-BS can be applied to other loop filtering methods.
[0424] b. In one example, the proposed DBF-BS method can be applied to a preprocessing filtering method.
[0425] c. In one example, the proposed DBF-BS method can be applied to a post-processing filtering method.
[0426] 14) In the above examples, a video unit can refer to a sequence / picture / subpicture / strip / piece / code-decode tree unit (CTU) / CTU line / CTU group / code-decode unit (CU) / prediction unit (PU) / transform unit (TU) / code-decode tree block (CTB) / code-decode block (CB) / prediction block (PB) / transform block (TB) / any other region containing more than one luminance or chrominance sample / pixel.
[0427] 15) Whether and / or how the methods disclosed above can be applied to transmit signals in a bitstream.
[0428] a. In one example, they can be transmitted via signaling at the sequence level / picture group level / picture level / strip level / piece group level, such as in the sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / strip header / piece group header.
[0429] b. In one example, they can be transmitted via signal in PB / TB / CB / PU / TU / CU / VPDU / CTU / CTU lines / strips / pieces / sub-pictures / other types of areas containing more than one sample point or pixel.
[0430] 16) Whether and / or how the methods disclosed above are applied may depend on the encoded / decoded information, such as block size, color format, single / double tree segmentation, color components, and stripe / picture type.
[0431] 17) The syntax elements disclosed above can be binarized into flags, fixed-length codes, EG(x) codes, unary codes, rounded unary codes, rounded binary codes, etc. They can be signed or unsigned.
[0432] 18) The syntax elements disclosed above can be encoded or decoded using at least one context model. Alternatively, they can be encoded or decoded using a bypass method.
[0433] 19) The syntax elements disclosed above can be transmitted via signals in a conditional manner.
[0434] a. SE is transmitted via signal only when the corresponding function is available.
[0435] b. SE is transmitted via signal only when the size (width and / or height) of the block meets the conditions.
[0436] 20) The syntax elements disclosed above can be transmitted via signaling at the block level / sequence level / picture group level / picture level / strip level / piece group level, such as in the codec structure of CTU / CU / TU / PU / CTB / CB / TB / PB or in the sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / strip header / piece group header.
[0437] The method proposed in 21)(one or more) can be combined with another codec tool, such as affine / MTS / LFNST / MMVD / MIP / ISP / CCLM / CCCM / SMVD / BDOF / DMVR / HMVP / template matching / IBC / palette / etc.
[0438] The method proposed in 22)(one or more) can be used without another codec tool, such as affine / MTS / LFNST / MMVD / MIP / ISP / CCLM / CCCM / SMVD / BDOF / DMVR / HMVP / template matching / IBC / palette / etc.
[0439] a. In one example, if one or more of the proposed methods are used, then the excluded codec tools are implicitly disabled without signaling.
[0440] b. In one example, if the excluded codec tool is used, then the proposed method(s) is implicitly disabled without signaling.
[0441] 6. References
[0442] [1] J.Strom, P.Wennersten, J.Enhorn, D.Liu, K.Andersson and R.Sjoberg, “Bilateral Loop Filter in Combination with SAO”, Proceedings of the IEEE Picture Coding and Decoding Workshop (PCS), November 2019.
[0443] Figure 13 This is a block diagram illustrating an example video processing system 4000 in which various techniques disclosed herein may be implemented. Various implementations may include some or all of the components of system 4000. System 4000 may include an input 4002 for receiving video content. The video content may be received in a raw or uncompressed format, such as 8 or 10-bit multi-component pixel values, or in a compressed or encoded format. Input 4002 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces (such as Ethernet, Passive Optical Networking (PON), etc.) and wireless interfaces (such as Wi-Fi or cellular interfaces).
[0444] System 4000 may include an encoding / decoding component 4004 capable of implementing the various encoding / decoding or coding methods described in this disclosure. Encoding / decoding component 4004 can reduce the average bit rate from the video input 4002 to the output of encoding / decoding component 4004 to produce an encoding / decoding representation of the video. Encoding / decoding techniques are therefore sometimes referred to as video compression or video transcoding techniques. The output of encoding / decoding component 4004 may be stored or transmitted via a communication connection, as represented by component 4006. The stored or communicatively transmitted bitstream (or encoding / decoding) representation of the video received at input 4002 may be used by component 4008 to generate pixel values or displayable video that is sent to display interface 4010. The process of generating user-visible video from the bitstream representation is sometimes referred to as video decompression. Furthermore, although some video processing operations are referred to as “encoding / decoding” operations or tools, it should be understood that encoding / decoding tools or operations used at the encoder will be followed by corresponding decoding tools or operations that inversely reproduce the encoding / decoding results performed by the decoder.
[0445] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB), High Definition Multimedia Interface (HDMI), or DisplayPort, etc. Examples of storage interfaces include SATA, PCI, IDE, etc. The technologies described in this disclosure can be embodied in a variety of electronic devices, such as mobile phones, laptops, smartphones, or other devices capable of performing digital data processing and / or video display.
[0446] Figure 14 This is a block diagram of an example video processing apparatus 4100. Apparatus 4100 can be used to implement one or more methods described herein. Apparatus 4100 can be embodied in a smartphone, tablet, computer, Internet of Things (IoT) receiver, etc. Apparatus 4100 may include one or more processors 4102, one or more memories 4104, and video processing circuitry 4106. The processor(s) 4102 can be configured to implement one or more methods described herein. The memories 4104 can be used to store data and code for implementing the methods and techniques described herein. The video processing circuitry 4106 can be used to implement some of the techniques described herein in hardware circuitry. In some embodiments, the video processing circuitry 4106 may be at least partially included in the processor 4102, for example, a graphics coprocessor.
[0447] Figure 15This is a flowchart of an example method 4200 for video processing. Method 4200 includes: in step 4202, determining the boundary strength (DBF-BS) of a deblocking filter as input to an adaptive loop filter (ALF) or a cross-component ALF (CC-ALF); and in step 4204, performing a conversion between visual media data and a bitstream based on the ALF or CC-ALF. According to the example, the conversion in step 4204 may include encoding at an encoder or decoding at a decoder.
[0448] It should be noted that method 4200 can be implemented in an apparatus for processing video data, including a processor and a non-transitory memory having instructions thereon, such as a video encoder 4400, a video decoder 4500, and / or an encoder 4600. In this case, the instructions, when executed by the processor, cause the processor to perform method 4200. Furthermore, method 4200 can be executed by a non-transitory computer-readable medium including a computer program product for use by a video codec device. The computer program product includes computer-executable instructions stored on the non-transitory computer-readable medium, such that, when executed by a processor, the computer-executable instructions cause the video codec device to perform method 4200.
[0449] Figure 16 This is a block diagram illustrating an example video encoding / decoding system 4300 from which the techniques of this disclosure can be utilized. The video encoding / decoding system 4300 may include a source device 4310 and a target device 4320. The source device 4310 generates encoded video data, and this source device 4310 may be referred to as a video encoding device. The target device 4320 can decode the encoded video data generated by the source device 4310, and this target device 4320 may be referred to as a video decoding device.
[0450] Source device 4310 may include video source 4312, video encoder 4314, and input / output (I / O) interface 4316. Video source 4312 may include sources such as video capture devices, interfaces for receiving video data from video content providers, and / or computer graphics systems for generating video data, or combinations thereof. Video data may include one or more pictures. Video encoder 4314 encodes the video data from video source 4312 to generate a bitstream. The bitstream may include a sequence of bits forming a codec representation of the video data. The bitstream may include codec pictures and associated data. Codec pictures are codec representations of pictures. Associated data may include sequence parameter sets, picture parameter sets, and other syntax structures. I / O interface 4316 may include a modulator / demodulator (modem) and / or transmitter. Encoded video data may be transmitted directly to target device 4320 via network 4330 through I / O interface 4316. Encoded video data may also be stored on storage medium / server 4340 for access by target device 4320.
[0451] Target device 4320 may include I / O interface 4326, video decoder 4324, and display device 4322. I / O interface 4326 may include a receiver and / or a modem. I / O interface 4326 may acquire encoded video data from source device 4310 or storage medium / server 4340. Video decoder 4324 may decode the encoded video data. Display device 4322 may display the decoded video data to a user. Display device 4322 may be integrated with target device 4320 or may be external to target device 4320, wherein target device 4320 may be configured to interface with an external display device.
[0452] The video encoder 4314 and the video decoder 4324 can operate according to video compression standards, such as the High Efficiency Video Codec (HEVC) standard, the Multi-Functional Video Codec (VVC) standard, and other existing and / or further standards.
[0453] Figure 17 This is a block diagram illustrating an example of a video encoder 4400, which can be... Figure 16 The system 4300 shown includes a video encoder 4314. The video encoder 4400 can be configured to perform any or all of the techniques disclosed herein. The video encoder 4400 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video encoder 4400. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.
[0454] The functional components of the video encoder 4400 may include a segmentation unit 4401, a prediction unit 4402, a residual generation unit 4407, a transform processing unit 4408, a quantization unit 4409, an inverse quantization unit 4410, an inverse transform unit 4411, a reconstruction unit 4412, a buffer 4413, and an entropy coding unit 4414. The prediction unit 4402 may include a mode selection unit 4403, a motion estimation unit 4404, a motion compensation unit 4405, and an intra-frame prediction unit 4406.
[0455] In other examples, the video encoder 4400 may include more, fewer, or different functional components. In one example, the prediction unit 4402 may include an intra-block copy (IBC) unit. The IBC unit can perform prediction in an IBC mode, where at least one reference picture is the picture in which the current video block is located.
[0456] Furthermore, some components such as the motion estimation unit 4404 and the motion compensation unit 4405 can be highly integrated, but for illustrative purposes, they are shown separately in the example of the video encoder 4400.
[0457] The segmentation unit 4401 can segment an image into one or more video blocks. The video encoder 4400 and the video decoder 4500 can support various video block sizes.
[0458] The mode selection unit 4403 can select one of several codec modes (intra-frame codec or inter-frame codec) based, for example, on the error result, and provide the resulting intra-frame or inter-frame codec block to the residual generation unit 4407 to generate residual block data, and to the reconstruction unit 4412 to reconstruct the coded block for use as a reference image. In some examples, the mode selection unit 4403 can select an intra-frame / inter-frame joint prediction (CIIP) mode, where prediction is based on inter-frame prediction signals and intra-frame prediction signals. In the case of inter-frame prediction, the mode selection unit 4403 can also select a resolution for the block based on the motion vector (e.g., sub-pixel precision or integer pixel precision).
[0459] To perform inter-frame prediction on the current video block, motion estimation unit 4404 can generate motion information for the current video block by comparing one or more reference frames from buffer 4413 with the current video block. Motion compensation unit 4405 can determine the predicted video block for the current video block based on the motion information and decoded samples from the image in buffer 4413 (rather than the image associated with the current video block).
[0460] The motion estimation unit 4404 and the motion compensation unit 4405 can perform different operations on the current video block, for example, depending on whether the current video block is in an I-band, P-band, or B-band.
[0461] In some examples, motion estimation unit 4404 can perform unidirectional prediction on the current video block, and can search for a reference video block for the current video block in the reference images of list 0 or list 1. Then, motion estimation unit 4404 can generate a reference index indicating the reference image containing the reference video block in list 0 or list 1, and a motion vector indicating the spatial displacement between the current video block and the reference video block. Motion estimation unit 4404 can output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 4405 can generate a predicted video block for the current block based on the reference video block indicated by the motion information of the current video block.
[0462] In other examples, motion estimation unit 4404 can perform bidirectional prediction on the current video block. Motion estimation unit 4404 can search for a reference video block for the current video block in the reference images in list 0, and can also search for another reference video block for the current video block in the reference images in list 1. Motion estimation unit 4404 can then generate reference indices indicating the reference images containing the reference video blocks in lists 0 and 1, and motion vectors indicating the spatial displacement between the reference video blocks and the current video block. Motion estimation unit 4404 can output the reference index and motion vector of the current video block as motion information for the current video block. Motion compensation unit 4405 can generate a predicted video block for the current video block based on the reference video blocks indicated by the motion information of the current video block.
[0463] In some examples, the motion estimation unit 4404 can output a complete set of motion information for use in the decoder's decoding process. In some examples, the motion estimation unit 4404 may not output a complete set of motion information for the current video. Instead, the motion estimation unit 4404 can reference the motion information of another video block to transmit the motion information of the current video block via a signal. For example, the motion estimation unit 4404 may determine that the motion information of the current video block is sufficiently similar to the motion information of neighboring video blocks.
[0464] In one example, the motion estimation unit 4404 may indicate a value to the video decoder 4500 in the syntax structure associated with the current video block, which indicates that the current video block has the same motion information as another video block.
[0465] In another example, motion estimation unit 4404 may identify another video block and motion vector difference (MVD) in the syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. Video decoder 4500 can use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[0466] As discussed above, the video encoder 4400 can transmit motion vectors via signaling in a predictive manner. Two examples of predictive signaling techniques that can be implemented by the video encoder 4400 include Advanced Motion Vector Prediction (AMVP) and Merge Pattern Signaling.
[0467] Intra-prediction unit 4406 can perform intra-prediction on the current video block. When intra-prediction unit 4406 performs intra-prediction on the current video block, it can generate prediction data for the current video block based on decoded samples of other video blocks in the same frame. The prediction data for the current video block can include the predicted video block and various syntax elements.
[0468] The residual generation unit 4407 can generate residual data for the current video block by subtracting one or more predicted video blocks from the current video block. The residual data for the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.
[0469] In other examples, such as in skip mode, there may be no residual data for the current video block, and the residual generation unit 4407 may not perform subtraction operations.
[0470] The transform processing unit 4408 can generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video blocks associated with the current video block.
[0471] After the transform processing unit 4408 generates a transform coefficient video block associated with the current video block, the quantization unit 4409 can quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0472] The inverse quantization unit 4410 and the inverse transform unit 4411 can apply inverse quantization and inverse transform to the transform coefficient video block respectively to reconstruct the residual video block from the transform coefficient video block. The reconstruction unit 4412 can add the reconstructed residual video block to the corresponding samples of one or more predicted video blocks generated by the prediction unit 4402 to generate a reconstructed video block associated with the current block and store it in the buffer 4413.
[0473] After the video block is reconstructed by reconstruction unit 4412, a loop filtering operation can be performed to reduce video block artifacts in the video block.
[0474] Entropy encoding unit 4414 can receive data from other functional components of video encoder 4400. When entropy encoding unit 4414 receives data, it can perform one or more entropy encoding operations to generate entropy-encoded data and output a bitstream including the entropy-encoded data.
[0475] Figure 18 This is a block diagram illustrating an example of a video decoder 4500, which can be... Figure 16 The system 4300 shown includes a video decoder 4324. The video decoder 4500 can be configured to perform any or all of the techniques disclosed herein. In the example shown, the video decoder 4500 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video decoder 4500. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.
[0476] In the example shown, the video decoder 4500 includes an entropy decoding unit 4501, a motion compensation unit 4502, an intra-frame prediction unit 4503, an inverse quantization unit 4504, an inverse transform unit 4505, a reconstruction unit 4506, and a buffer 4507. In some examples, the video decoder 4500 can perform a decoding process that is the overall inverse of the encoding process described with respect to the video encoder 4400.
[0477] The entropy decoding unit 4501 can retrieve the encoded bitstream. The encoded bitstream may include entropy-encoded video data (e.g., encoded video data blocks). The entropy decoding unit 4501 can decode the entropy-encoded video data, and based on the entropy-decoded video data, the motion compensation unit 4502 can determine motion information including motion vectors, motion vector precision, reference image list index, and other motion information. The motion compensation unit 4502 can determine this information, for example, by executing AMVP and Merge modes.
[0478] The motion compensation unit 4502 can generate motion compensation blocks, possibly performing interpolation based on an interpolation filter. The identifier of the interpolation filter to be used, with sub-pixel accuracy, can be included in the syntax element.
[0479] The motion compensation unit 4502 can use interpolation filters, such as those used by the video encoder 4400 during the encoding of a video block, to calculate interpolations for sub-integer pixels of a reference block. The motion compensation unit 4502 can determine the interpolation filter used by the video encoder 4400 based on the received syntax information, and the motion compensation unit 4502 can use the interpolation filter to generate a prediction block.
[0480] The motion compensation unit 4502 may use some syntax information to determine the size of the blocks used to encode one or more frames and / or one or more stripes of the encoded video sequence, segmentation information describing how each macroblock of the picture of the encoded video sequence is segmented, a mode indicating how each segment is encoded, one or more reference frames (and a list of reference frames) for each inter-frame codec block, and other information for decoding the encoded video sequence.
[0481] Intra-prediction unit 4503 can use, for example, an intra-prediction mode received in the bitstream to form prediction blocks from spatially adjacent blocks. Dequantization unit 4504 dequantizes (i.e., de-quantizes) the quantized video block coefficients provided in the bitstream and decoded by entropy decoding unit 4501. Inverse transform unit 4505 applies the inverse transform.
[0482] The reconstruction unit 4506 can add the residual block to the corresponding predicted block generated by the motion compensation unit 4502 or the intra-frame prediction unit 4503 to form a decoded block. If necessary, a deblocking filter can also be used to filter the decoded block to remove block artifacts. The decoded video block is then stored in a buffer 4507, which provides a reference block for subsequent motion compensation / intra-frame prediction and also generates decoded video for presentation on a display device.
[0483] Figure 19 This is a schematic diagram of the example encoder 4600. Encoder 4600 is suitable for implementing VVC techniques. Encoder 4600 includes three loop filters: a deblocking filter (DF) 4602, a SAO 4604, and an adaptive loop filter (ALF) 4606. Unlike the DF 4602, which uses predefined filters, the SAO 4604 and ALF 4606 utilize the original samples of the current image, respectively, by adding an offset and by applying a finite impulse response (FIR) filter, and by utilizing the encoded / decoded side information through signal transmission offset and filter coefficients to reduce the mean square error between the original and reconstructed samples. The ALF 4606 is located in the last processing stage of each image and can be thought of as a tool to attempt to capture and repair artifacts caused by previous stages.
[0484] The encoder 4600 also includes an intra-frame prediction component 4608 and a motion estimation / compensation (ME / MC) component 4610, which are configured to receive input video. The intra-frame prediction component 4608 is configured to perform intra-frame prediction, while the ME / MC component 4610 is configured to perform inter-frame prediction using a reference image obtained from a reference image buffer 4612. Residual blocks from inter-frame or intra-frame prediction are fed into a transform (T) component 4614 and a quantization (Q) component 4616 to generate quantized residual transform coefficients, which are then fed into an entropy codec component 4618. The entropy codec component 4618 entropy codes the prediction results and the quantized transform coefficients and transmits them toward a video decoder (not shown). The quantized components output from the quantization component 4616 can be fed into the inverse quantization (IQ) component 4620, the inverse transform component 4622, and the reconstruction (REC) component 4624. The REC component 4624 can output the image to the DF 4602, SAO 4604, and ALF 4606 for filtering before these images are stored in the reference image buffer 4612.
[0485] The following is a list of some preferred solutions.
[0486] The following solutions illustrate examples of the techniques discussed in this article.
[0487] 1. A method for processing video data, comprising: determining edge information using a deblocking filter boundary strength (DBF-BS) as input to an adaptive loop filter (ALF) or a cross-component ALF (CC-ALF); and performing a conversion between visual media data and a bitstream based on said ALF or CC-ALF.
[0488] 2. The method according to Solution 1, wherein the DBF-BS of the luminance or chromaticity position is used as input.
[0489] 3. The method according to any one of solutions 1-2, wherein the DBF-BS of the position is initialized to zero, an N value, a predefined value, a derived value, a value transmitted by a signal, or a combination thereof.
[0490] 4. The method according to any one of solutions 1-3, wherein the DBF-BS of the location is used in the original domain or the mapped domain, and wherein the DBF-BS for the location is used together with a default value, a predefined value, a derived value or a value transmitted via signal for the corresponding boundary strength.
[0491] 5. The method according to any one of solutions 1-4, wherein the DBF-BS of the location is merged into N categories, merged according to rules, merged according to a threshold function, merged according to the number of predefined, derived or transmitted merged categories, or a combination thereof.
[0492] 6. The method according to any one of solutions 1-5, wherein the DBF-BS at the location is filtered before use, and wherein the DBF-BS is filtered by a predefined filter, an offline-trained filter for ALF, an online-trained filter for signal transmission, a Gaussian filter, a Sobel filter, a Pravet filter, a Roberts filter, a Canney filter, a Hadamard transform domain filter, a bilateral filter, a low-pass filter, a high-pass filter, or a combination thereof.
[0493] 7. The method according to any one of solutions 1-6, wherein the DBF-BS is predefined, derived, or transmitted via signaling.
[0494] 8. The method according to any one of solutions 1-7, wherein the DBF-BS is clipped, scaled, or converted to a clipping range, bit depth, or domain.
[0495] 9. The method according to any one of solutions 1-8, wherein the DBF-BS of the location is derived and stored, but the location is not filtered by a deblocking filter (DBF).
[0496] 10. The method according to any one of solutions 1-9, wherein f(Bs) is used as edge information, where Bs is the boundary strength, and f is any function.
[0497] 11. The method according to any one of solutions 1-10, wherein DBF-BS luminance, DBF-BS chrominance, filtered DBF-BS luminance, or filtered DBF-BS chrominance is used as input to: ALF tap, ALF extended tap, ALF luminance, ALF chrominance, CC-ALF tap, CC-ALF extended tap, or a combination thereof.
[0498] 12. The method according to any one of solutions 1-11, wherein the difference between the DBF-BS or the filtered DBF-BS and the currently filtered sample is calculated, or the filter coefficients are multiplied by the difference to obtain the final offset.
[0499] 13. The method according to any one of solutions 1-12, wherein the DBF-BS value or the filtered DBF-BS value is used directly, or the filter coefficients are multiplied by the DBF-BS to obtain the final offset.
[0500] 14. The method according to any one of solutions 1-13, wherein DBF-BS is used as edge information in ALF classification.
[0501] 15. The method according to any one of solutions 1-14, wherein the classification unit size of the DBF-BS-based classifier can be N×M, where N and M are integers.
[0502] 16. The method according to any one of solutions 1-15, wherein positions within a classification unit share the same classification result, or the classification result is predefined, derived, transmitted by signal, an average function, a maximum function, a minimum function, a linear mapping function, a nonlinear mapping function, or a combination thereof.
[0503] 17. The method according to any one of solutions 1-16, wherein the use of the DBF-BS-based classifier is predefined, derived, or transmitted via signaling.
[0504] 18. The method according to any one of solutions 1-17, wherein DBF-BS is merged into K categories, and the classifier has N categories, where K and N are integers, and wherein DBF-BS expands N according to the following formula: N×K; c=c0+BS×N; or c=c0×K+BS, where c is the final category index, and c0 is the category index generated by the classifier.
[0505] 19. The method according to any one of solutions 1-18, wherein the merged DBF-BS has K categories and the classifier has M categories, where K and M are integers, and wherein the merged DBF-BS modifies M according to the following formula: M×K; c=c0+BS×M; or c=c0×M+BS, where c is the final category index and c0 is the category index generated by the classifier.
[0506] 20. The method according to any one of solutions 1-19, wherein the total number of merged DBF-BS modification bands, the number of texture directions, or the number of texture activities.
[0507] 21. The method according to any one of solutions 1-20, wherein DBF-BS is used for classification, offset generation or filter strength modification in a bilateral filter, a Hadamard transform domain (HTDF) filter, a sample adaptive compensation (SAO), a cross-component SAO (CCSAO) or a combination thereof.
[0508] 22. The method according to any one of solutions 1-21, wherein the syntax elements in the bitstream indicate the use of DBF-BS, and wherein the syntax elements are transmitted via a signal through context encoding / decoding, bypass encoding / decoding, unary code, rounded unary code, fixed-length code, exponential Golomb code, rounded exponential Golomb code, conditional encoding / decoding, predictive encoding / decoding, independently for color components, through a set of parameters or a combination thereof.
[0509] 23. The method according to any one of solutions 1-22, wherein the application of DBF-BS on ALF or CC-ALF depends on the control information for DBF.
[0510] 24. An apparatus for processing video data, comprising: a processor; and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method of any one of solutions 1-23.
[0511] 25. A non-transitory computer-readable medium comprising a computer program product for use by a video codec apparatus, the computer program product comprising computer-executable instructions stored on the non-transitory computer-readable medium such that, when executed by a processor, the video codec apparatus performs the method described in any one of solutions 1-23.
[0512] 26. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by a video processing apparatus, wherein the method includes: determining edge information using a deblocking filter boundary strength (DBF-BS) as input to an adaptive loop filter (ALF) or a cross-component ALF (CC-ALF); and generating the bitstream based on the determination.
[0513] 27. A method for storing a bitstream of video, comprising: determining edge information using a deblocking filter boundary strength (DBF-BS) as input to an adaptive loop filter (ALF) or a cross-component ALF (CC-ALF); generating the bitstream based on the determination; and storing the bitstream in a non-transitory computer-readable recording medium.
[0514] 28. A method, apparatus or system described in this disclosure.
[0515] In the described solution, the encoder conforms to the format rules by generating a codec representation based on those rules. In the described solution, the decoder parses the syntax elements in the codec representation using known information about their presence or absence, based on the format rules, to produce the decoded video.
[0516] In this disclosure, the term "video processing" can refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm can be applied during the conversion from a pixel representation of a video to a corresponding bitstream representation, and vice versa. For example, the bitstream representation of the current video block can correspond to bits at co-positions or propagated at different positions in the bitstream defined by the syntax. For example, a macroblock can be encoded based on the error residual value after transformation and encoding, and can also use bits from the header and other fields in the bitstream. Furthermore, during the conversion, the decoder can parse the bitstream based on the determination, knowing whether some fields may or may not be present, as described in the solutions above. Similarly, the encoder can determine whether to include or exclude specific syntax fields, and generate the codec representation accordingly by including or excluding syntax fields from the codec representation.
[0517] The disclosed and other solutions, examples, embodiments, modules, and functional operations described herein can be implemented in digital electronic circuits, or in computer software, firmware, or hardware, including the structures disclosed herein and their structural equivalents, or in a combination of one or more. The disclosed and other embodiments can be implemented as one or more computer program products, i.e., one or more computer program instruction modules encoded on a computer-readable medium for execution by or control of the operation of a data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a storage device, a material composition affecting machine-readable propagation signals, or a combination of one or more. The term "data processing apparatus" includes all means, devices, and machines for processing data, including, for example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, the apparatus may also include code that creates an execution environment for an associated computer program, for example, code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more. Propagation signals are artificially generated signals, such as machine-generated electrical signals, optical signals, or electromagnetic signals, which are generated to encode information for transmission to a suitable receiver device.
[0518] Computer programs (also known as programs, software, software applications, scripts, or code) can be written in any programming language, including compiled or interpreted languages, and can be deployed in any form, including standalone programs or modules, components, subroutines, or other units suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the related program, or in multiple co-located files (e.g., a file storing one or more modules, subroutines, or code portions). A computer program can be deployed to execute on one computer or on multiple computers located at a single site or distributed across multiple sites and interconnected by a communications network.
[0519] The processing and logic flows described in this disclosure can be executed by one or more programmable processors that execute one or more computer programs to perform functions by manipulating input data and generating outputs. The processing and logic flows can also be executed by dedicated logic circuitry, and the apparatus can be implemented as dedicated logic circuitry, such as a field-programmable gate array (FPGA) or an application-specific integrated circuit (ASIC).
[0520] Processors suitable for executing computer programs include, for example, general-purpose and special-purpose microprocessors, and any one or more processors of any type of digital computer. Typically, a processor receives instructions and data from read-only memory or random access memory, or both. The basic components of a computer are a processor that executes instructions and one or more storage devices that store the instructions and data. Typically, a computer will also include one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, or receive data from or transfer data to one or more mass storage devices via operative coupling, or both. However, a computer does not necessarily have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and storage devices, including, for example, semiconductor storage devices such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable hard disks; magneto-optical disks; and CD-ROMs and DVD-ROMs. Processors and memory may be supplemented by or incorporated into special-purpose logic circuitry.
[0521] While this patent disclosure contains numerous details, these details should not be construed as limiting any subject matter or the scope of the claims, but rather as descriptions of features specific to particular embodiments of a particular technology. In this patent disclosure, certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented separately in multiple embodiments, or in any suitable sub-combination. Furthermore, although features may function in certain combinations as described above, and even were originally claimed in this manner, in some cases one or more features in the claimed combination may be removed from that combination, and the claimed combination may be for sub-combinations or variations thereof.
[0522] Similarly, although operations are depicted in a specific order in the accompanying drawings, this should not be construed as requiring such operations to be performed sequentially in the specific order or sequence shown, or requiring all shown operations to be performed in order to achieve the desired result. Furthermore, the division of various system components in the embodiments described in this patent disclosure should not be construed as requiring such division in all embodiments.
[0523] Only a few implementations and examples are described, and other implementations, improvements and variations can be made based on what is described and shown in this patent disclosure.
[0524] When there is no intervening component other than a line, trace, or other medium between the first and second components, the first component is directly coupled to the second component. When there is an intervening component other than a line, trace, or other medium between the first and second components, the first component is indirectly coupled to the second component. The term "coupled" and its variations include direct coupling and indirect coupling. The use of the term "about" means including a range of ±10% of the following figures, unless otherwise specified.
[0525] While several embodiments have been provided in this disclosure, it should be understood that the disclosed systems and methods may be embodied in many other specific forms without departing from the spirit or scope of this disclosure. The present examples are to be considered illustrative rather than restrictive and are not intended to be limited to the details set forth herein. For example, various elements or components may be combined or integrated into another system, or certain features may be omitted or not implemented.
[0526] Furthermore, the technologies, systems, subsystems, and methods described and illustrated as discrete or separate in the various embodiments can be combined or integrated with other systems, modules, technologies, or methods without departing from the scope of this disclosure. Other items shown or discussed as couplings can be directly connected or indirectly coupled or communicated through some interface, device, or intermediate component, whether electrically, mechanically, or otherwise. Other examples of variations, substitutions, and alterations can be determined by those skilled in the art and can be made without departing from the spirit and scope of this disclosure.
Claims
1. A method for processing video data, comprising: The boundary strength of the deblocking filter (DBF-BS) is determined as the side information input to the adaptive loop filter (ALF) or cross-component ALF (CC-ALF); as well as Perform the conversion between visual media data and bitstream based on the ALF or CC-ALF.
2. The method of claim 1, wherein the DBF-BS at at least one location is stored during or after the deblocking filter (DBF) process.
3. The method according to any one of claims 1-2, wherein the DBF-BS at at least one location is stored.
4. The method according to any one of claims 1-3, wherein the DBF-BS of the luminance position is stored.
5. The method according to any one of claims 1-4, wherein the DBF-BS of the chromaticity position is stored.
6. The method according to any one of claims 1-2, wherein the DBF-BS at the location is initialized to a specific value.
7. The method of claim 6, wherein the DBF-BS at the position is initialized to 0.
8. The method according to any one of claims 6-7, wherein the DBF-BS at the position is initialized to N, including N = 32.
9. The method according to any one of claims 6-8, wherein the initial value is predefined.
10. The method according to any one of claims 6-9, wherein the initial value is derived.
11. The method according to any one of claims 6-9, wherein the initial value is transmitted to the decoder via a signal.
12. The method according to any one of claims 1-2, wherein the DBF-BS of the location is stored or used in the original domain.
13. The method of claim 12, wherein the DBF-BS of the location is stored or used using default values representing different boundary strengths, including BS = 0, 1, 2.
14. The method according to any one of claims 1-2, wherein the DBF-BS at each location is stored or used in a mapping domain.
15. The method of claim 14, wherein the DBF-BS at each location is stored or used using a mapping value representing different boundary strengths, including BS = 128, 256, or 512.
16. The method according to any one of claims 14-15, wherein the mapping value is predefined.
17. The method according to any one of claims 14-16, wherein the mapping value is derived.
18. The method according to any one of claims 14-17, wherein the mapped value is transmitted to the decoder via a signal.
19. The method according to any one of the preceding claims, wherein the DBF-BS at a given location is merged into one or more categories.
20. The method according to any one of the preceding claims, wherein the BDF-BS at a given location is merged into N categories, including N=2.
21. The method according to any one of the preceding claims, wherein the DBF-BS at a position is merged following a predefined rule, and wherein the predefined rule is a threshold-based limiting or mapping function, or further wherein the DBF-BS is set to V0 when the DBF-BS is greater than or equal to the threshold, or is set to V1 when the DBF-BS is less than the threshold.
22. The method according to any one of the preceding claims, wherein the DBF-BS at a location is merged according to a predefined rule, and wherein the number of the merging rules or merging categories is predefined, derived, or transmitted to the decoder via a signal.
23. The method according to any one of the preceding claims, wherein the DBF-BS at the location is filtered by a predefined filter, or by the offline-trained filter of the ALF, or by the online-trained filter transmitted to the decoder via signal transmission, prior to use.
24. The method according to any one of the preceding claims, wherein the DBF-BS at the location is filtered by a Gaussian filter, a Sobel filter, a Prevet filter, a Roberts filter, a Canney filter, a Hadamard transform domain filter, a bilateral filter, a low-pass filter, a high-pass filter, or other filters.
25. The method according to any one of the preceding claims, wherein whether to filter the DBF-BS or which filter to apply to filter the DBF-BS is predefined, derived, or transmitted to the decoder via a signal.
26. The method according to any one of the preceding claims, wherein the DBF-BS is limited to a range or a bit depth, limited to a predefined or signal-transmitted or derived limiting range, or limited to a predefined or signal-transmitted or derived N-bit depth, including N=10.
27. The method according to any one of the preceding claims, wherein the DBF-BS is scaled to a range or bit depth, a predefined or signal-transmitted or derived range, or a predefined or signal-transmitted or derived N-bit depth, including N=10.
28. The method according to any one of the preceding claims, wherein the DBF-BS is converted to a range or domain or bit depth, or a predefined range or domain transmitted or derived by signal transmission, or a predefined or derived N-bit depth, including N=10.
29. The method according to any one of the preceding claims, wherein the DBF-BS at at least one location is derived without invoking the DBF process, or is derived and stored at the location without being filtered by the DBF.
30. The method according to any one of the preceding claims, wherein f(Bs) is used as edge information in ALF or CCALF filtering, wherein Bs is the boundary strength, and f is any function, such as a linear function or a limiting function or any other function.
31. The method according to any one of the preceding claims, wherein the DBF-BS is used as an input to the ALF, the DBF-BS is used as an input to at least one extended tap of the ALF, or as at least one existing tap of the ALF, the luminance DBF-BS is used as an input to the ALF luminance, the chrominance DBF-BS is used as an input to the ALF chrominance, the luminance filtered DBF-BS is used as an input to the ALF luminance, or the chrominance filtered DBF-BS is used as an input to the ALF chrominance.
32. The method according to any one of the preceding claims, wherein the DBF-BS is used as an input to a CCALF, as an input to at least one extended tap of a CCALF, as an input to at least one existing tap of a CCALF, the luminance DBF-BS is used as an input to an extended tap of a CCALF, or the luminance filtered DBF-BS is used as an input to a CCALF.
33. The method according to any one of the preceding claims, wherein the DBF-BS is used after the ALF or CCALF filtering function, and wherein the difference between the DBF-BS or the filtered DBF-BS and the currently filtered sample is first calculated, or the filter coefficients are multiplied by the difference to obtain the final offset.
34. The method according to any one of the preceding claims, wherein the DBF-BS is used in a manner different from the ALF / CCALF filter function, and wherein the DBF-BS value or the filtered DBF-BS value is used directly, or the filter coefficients are multiplied by the DBF-BS to obtain the final offset.
35. The method according to any one of the preceding claims, wherein the DBF-BS is used as edge information in ALF classification, or f(Bs) is used as edge information in ALF classification, wherein Bs is the boundary strength and f is an arbitrary function, such as a linear function, a limiting function, or any other function.
36. The method according to any one of the preceding claims, wherein the DBF-BS or filtered DBF-BS is used as an additional classification mode, wherein the classification unit size of the classifier based on the DBF-BS is N×M, including N=M=2, or the positions within the classification unit share the same classification result.
37. The method according to any one of the preceding claims, wherein the classification result is generated or derived using a specific method, or is predefined, derived, or transmitted to the decoder via a signal, is an average function, a maximum function, a minimum function, a linear mapping function, a nonlinear mapping function, or other functions, or whether a DBF-BS-based classifier is used is predefined, derived, or transmitted via a signal.
38. The method according to any one of the preceding claims, wherein the DBF-BS or filtered DBF-BS is used as side information of an existing classification pattern.
39. The method according to any one of the preceding claims, wherein the locations within the classification unit share the same DBF-BS value.
40. The method according to any one of the preceding claims, wherein the DBF-BS value within the classification unit is generated or derived using a specific method, or is predefined, derived, and transmitted to the decoder via a signal, and is an average function, a maximum function, a minimum function, a linear mapping function, a nonlinear mapping function, or other function.
41. The method according to any one of the preceding claims, wherein the DBF-BS is merged into K categories, including K=2, or the existing classifier has N categories, including N=25.
42. The method according to any one of the preceding claims, wherein the merged DBF-BS is directly used to expand the number of categories, the number of categories of the existing classifier is expanded to N×K, including N×K=50, and the classification result is calculated by the following formula: c=c0+BS×N, where c is the final category index and c0 is the category index generated by the existing classifier, or the classification result is calculated by the following formula: c=c0×K+BS, where c is the final category index and c0 is the category index generated by the existing classifier.
43. The method according to any one of the preceding claims, wherein the merged DBF-BS is used to modify the number of categories, wherein the number of categories of the existing classifier is modified to M, including M=12, wherein for a band-based classifier the total number of bands is modified, for a texture-based classifier the number of texture directions is modified, or for a texture-based classifier the number of texture activities is modified.
44. The method according to any one of the preceding claims, wherein the merged DBF-BS is used to modify the number of categories, the number of categories of the existing classifier is modified to N×K, including M×K=24, and the classification result is calculated by the following formula: c=c0+BS×M, where c is the final category index, and c0 is the category index generated by the existing classifier, or the classification result is calculated by the following formula: c=c0×K+BS, where c is the final category index, and c0 is the category index generated by the existing classifier.
45. The method according to any one of the preceding claims, wherein the DBF-BS is used as side information in the bilateral filter, and wherein the DBF-BS is used for classification of the bilateral filter, or different DBF-BS values result in different filter strengths of the bilateral filter.
46. The method according to any one of the preceding claims, wherein the DBF-BS is used as side information in a Hadamard transform domain filter (HTDF) filter, and wherein the DBF-BS is used for classification of the HTDF filter, or different DBF-BS values result in different filter strengths of the HTDF filter.
47. The method according to any one of the preceding claims, wherein DBF-BS is used as side information in sample adaptive compensation (SAO) classification, and wherein the DBF-BS is used as an additional classification pattern, or the DBF-BS is used as side information of an existing classification pattern.
48. The method according to any one of the preceding claims, wherein the DBF-BS is used as side information in classification in Cross-Component Sample Adaptive Compensation (CCSAO), and wherein the DBF-BS is used as an additional classification pattern, the DBF-BS is used as side information of an existing classification pattern, the DBF-BS is used to generate the offset of the CCSAO, or the DBF-BS is used to generate the offset of the CCSAO.
49. The method according to any one of the preceding claims, wherein whether or how the DBF-BS is modified or prepared is transmitted via signaling, and wherein a first syntax element is transmitted via signaling to indicate whether the modified or prepared DBF-BS is enabled or used.
50. The method according to any one of the preceding claims, wherein the first syntax element is encoded and decoded by arithmetic encoding and decoding, and wherein the first syntax element is encoded and decoded using at least one context, wherein the context depends on the encoding and decoding information of the current block or neighboring blocks, or the context depends on the filter shape of at least one neighboring block, or wherein the first syntax element is encoded and decoded using bypass encoding and decoding.
51. The method according to any one of the preceding claims, wherein the first syntax element is binarized using a unary code, a rounded unary code, a fixed-length code, an exponential Golomb code, or a rounded exponential Golomb code, and wherein the first syntax element is conditionally transmitted via signaling. It is transmitted via signal only when the Deblocking Filter Boundary Strength (DBF-BS) is available, encoded and decoded in a predictive manner, predicted by the on / off decision of the DBF-BS of at least one neighboring block, transmitted via signal independently for different color components, transmitted via signal and shared for different color components, transmitted via signal for the first color component but not for the second color component, or transmitted via signal in the Sequence Parameter Set (SPS), Picture Parameter Set (PPS), Picture Header, Strip Header, Codec Unit (CU), Codec Tree Unit (CTU), or APS.
52. The method according to any one of the preceding claims, wherein whether DBF-BS is applied to ALF and / or CCALF is based on the control information of DBF, and wherein DBF-BS is not applied to ALF and / or CCALF when DBF is turned off.
53. The method according to any one of the preceding claims, wherein whether samples after or before DBF are applied to ALF and / or CCALF is based on the control information of DBF, and wherein when DBF is turned off, samples after DBF are not applied to ALF and / or CCALF, or when DBF is turned off, samples before DBF are not applied to ALF and / or CCALF.
54. The method according to any one of the preceding claims, wherein any one of the steps is used for post-processing and / or pre-processing and / or used in combination or alone.
55. The method according to any one of the preceding claims, wherein the DBF-BS is in the same color component as the sample to be filtered or in a different color component, or the DBF-BS is at the same location as the sample to be filtered or within a range around the sample to be filtered.
56. The method according to any one of the preceding claims, wherein the DBF-BS is used as a loop filtering tool, preprocessing or postprocessing filter in video encoding and decoding.
57. The method according to any one of the preceding claims, wherein the method using DBF-BS is applied to loop filtering, to ALF, to CCALF, to other loop filtering tools, to preprocessing filtering, or to postprocessing filtering.
58. The method according to any one of the preceding claims, wherein the video unit refers to a sequence, picture, sub-picture, strip, slice, codec tree unit (CTU), CTU row or CTU group, codec unit (CU), prediction unit (PU), transform unit (TU), codec tree block (CTB), codec block (CB), prediction block (PB), transform block (TB), or other region containing more than one luminance or chrominance sample point or pixel.
59. The method according to any one of the preceding claims, wherein the application of the steps is transmitted via signaling in a bitstream, and wherein the signaling is transmitted via signaling at the sequence level, picture group level, picture level, strip level, slice group level, or in the sequence header, picture header, SPS, video codec parameters (VPS), decoding parameter set (DPS), decoding capability information (DCI), PPS, APS, strip header, or slice group header, or in the prediction block (PB), transform block (TB), codec block (CB), picture unit (PU), transform unit (TU), CU, virtual pipeline data unit (VPDU), CTU, CTU line or strip or slice or sub-picture or other region containing more than one sample or pixel.
60. The method according to any one of the preceding claims, wherein the application of the steps is based on encoded or decoded information, block size, color format, single-tree segmentation or dual-tree segmentation, color components, or stripe or image type.
61. The method according to any one of the preceding claims, wherein the syntax element is binarized into a flag, a fixed-length code, an EG(x) code, a unary code, a rounded unary code, a rounded binary code, is encoded or decoded using at least one context model, or is bypassed and decoded, and is signed or unsigned.
62. The method according to any one of the preceding claims, wherein the syntax element is conditionally transmitted by signal, transmitted by signal only when the corresponding function applies, or transmitted by signal only when the dimensions (width and / or height) of the block satisfy a condition.
63. The method according to any one of the preceding claims, wherein syntax elements are transmitted via signals at the block level, sequence level, picture group level, picture level, stripe level or slice group level, in the encoding / decoding structure of CTU, CU, TU, PU, CTB, CB, TB, PB, or in the sequence header, picture header, SPS, VPS, DPS, DCI, PPS, APS, stripe header or slice group header.
64. The method according to any one of the preceding claims, wherein the method further implements other encoding / decoding tools, including affine, multiple transform selection (MTS), low-frequency non-separable transform (LFNST), Merge mode with motion vector difference (MMVD), matrix-based intra-frame prediction (MIP), intra-frame sub-segmentation (ISP), cross-component linear model (CCLM), convolutional cross-component model (CCCM), symmetric motion vector difference (SMVD), bidirectional optical flow (BDOF), decoder-side motion vector refinement (DMVR), history-based motion vector prediction (HMVP), template matching, intra-frame block copy (IBC), or palette.
65. The method according to any one of the preceding claims, wherein the method is not used in conjunction with another codec tool, the other codec tool including affine, MTS, LFNST, MMVD, MIP, ISP, CCLM, CCCM, SMVD, BDOF, DMVR, HMVP, stencil matching, IBC, or palette.