Cross-component adaptive filtering and sub-block coding
By introducing the Cross-Component Adaptive Loop Filter (CC-ALF) tool into video coding, the problem of low efficiency in video component filtering in existing technologies is solved, and more efficient video coding results are achieved.
Patent Information
- Application Number
- CN202080082076.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-11-30
- Filing Date
- 2020-11-27
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2040-11-27
AI Technical Summary
Existing video encoding and decoding technologies struggle to effectively utilize cross-component adaptive loop filters when processing video components, resulting in limitations in video coding efficiency and quality.
The Cross-Component Adaptive Loop Filter (CC-ALF) tool is used to predict the sample value of the current video component by converting between video block and bitstream representation, using the sample values of other video components, and applying symmetric or asymmetric filters for filtering. It supports multiple filter sets and position rules to determine filter activation.
It improves the efficiency and quality of video encoding, enhances the filtering effect between video components, and improves the performance of the video encoder.
Smart Images

Figure CN114930832B_ABST
Abstract
Description
[0001] Cross-reference to related applications
[0002] This application is filed in accordance with the patent law and / or rules applicable under the Paris Convention to promptly claim priority and benefit from international patent application PCT / CN2019 / 122237, filed November 30, 2019. For all purposes provided by law, the entire disclosure of the foregoing application is incorporated by reference as part of the disclosure of this application. Technical Field
[0003] This patent document relates to image and video encoding and decoding. Background Technology
[0004] Digital video accounts for the largest share of bandwidth usage on the Internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth demand for digital video is expected to continue to grow. Summary of the Invention
[0005] This document discloses techniques that can be used by video encoders and decoders to perform cross-component adaptive loop filtering during video encoding or decoding.
[0006] In one example aspect, a method for video processing is disclosed. The method includes: performing a conversion between video blocks of video components and a bitstream representation of the video, wherein the video blocks comprise sub-blocks, wherein a filtering tool is used during the conversion according to rules, and wherein the rules specify that the filtering tool is applied by using a single offset for all samples of each sub-block of the video block.
[0007] In another example aspect, a method for video processing is disclosed. The method includes performing a conversion between a video block of a video component and a bitstream representation of the video using the final offset value of a current sample of a sub-block of a video block for which a filtering tool is applied during the conversion, wherein the filtering tool is applied by using the final offset value for the current sample, and wherein the final offset value is based on a first offset value and a second offset value different from the first offset value.
[0008] In another example, a method for video processing is disclosed. This method includes performing a conversion between video blocks of video components and a bitstream representation of the video by using an N-tap symmetric filter for a cross-component adaptive loop filter (CC-ALF) tool during conversion, wherein the CC-ALF tool predicts sample values of the video blocks of the video components from sample values of another video component of the video, and wherein at least two filter coefficients of the two samples within the support of the N-tap symmetric filter share the same value.
[0009] In another example, a method for video processing is disclosed. This method includes performing a conversion between video blocks of video components and a bitstream representation of the video by using an N-tap asymmetric filter for a cross-component adaptive loop filter (CC-ALF) tool during the conversion, where N is a positive integer, and where the CC-ALF tool predicts sample values of video blocks of the video components from sample values of another video component of the video.
[0010] In another example, a method for video processing is disclosed. The method includes: performing a conversion between a video block of a first video component of the video and a bitstream representation of the video, wherein the conversion of samples of the first video component includes applying a cross-component adaptive loop filter (CC-ALF) tool to the sample difference of a second video component of the video, and wherein the CC-ALF tool predicts the sample values of the video block of the first video component of the video from the sample values of another video component of the video.
[0011] In another example, a method for video processing is disclosed. This method includes performing a conversion between sub-blocks of video blocks of a video component and a bitstream representation of the video by using two or more filters from a set of multiple filters of a cross-component adaptive loop filtering (CC-ALF) tool during conversion, wherein the CC-ALF tool predicts sample values of sub-blocks of video blocks of the video component from sample values of another video component of the video.
[0012] In another example, a method for video processing is disclosed. This method includes performing a conversion between sub-blocks of a video block of a first video component and a bitstream representation of the video using filters supported by a Cross-Component Adaptive Loop Filtering (CC-ALF) tool used during the conversion, wherein the CC-ALF tool predicts sample values of sub-blocks of the video block of the first video component from sample values of another video component of the video.
[0013] In another example, a method for video processing is disclosed. The method includes performing a conversion between video blocks of video components and a bitstream representation of the video by refining the set of samples in the current video frame of the video using samples from multiple video frames in a Cross-Component Adaptive Loop Filter (CC-ALF) tool or an Adaptive Loop Filter (ALF) tool applied during the conversion. The CC-ALF tool predicts sample values of the video blocks of the video components from sample values of another video component of the video, and the ALF tool filters the samples of the video blocks of the video components using a loop filter.
[0014] In another example, a method for video processing is disclosed. The method includes: for a conversion between video blocks of video components and a bitstream representation of the video, determining whether to enable a cross-component adaptive loop filter (CC-ALF) tool for the conversion based on a positional rule, wherein the CC-ALF tool predicts sample values of the video blocks of the video components from sample values of another video component of the video; and performing the conversion based on the determination.
[0015] In another example, a method for video processing is disclosed. The method includes: for a conversion between video units of a video component and a bitstream representation of the video, determining a cross-component adaptive loop filtering for all samples of a sub-block of the video unit using an offset value; and performing the conversion based on the determination, wherein the offset is also used for another processing operation in the conversion, the other processing operation including one or more adaptive loop filtering operations.
[0016] In another example, a method for video processing is disclosed. The method includes: for a conversion between video units of a video component and a bitstream representation of the video, determining a final offset value for a current sample of a sub-block of the video unit to be used for cross-component adaptive loop filtering during the conversion; and performing the conversion based on the determination; wherein the final offset value is based on a first offset value and a second offset value different from the first offset value.
[0017] In another example, a method for video processing is disclosed. The method includes: for a conversion between video units of video components and a bitstream representation of the video, determining that an N-tap symmetric filter will be used for cross-component adaptive loop filter calculation during the conversion; and performing the conversion based on this determination; wherein, with the support of the N-tap symmetric filter, at least two filter coefficients of two samples share the same value.
[0018] In another example, a method for video processing is disclosed. The method includes: for a conversion between video units of video components and a bitstream representation of the video, determining that an N-tap asymmetric filter will be used for cross-component adaptive loop filter calculation during the conversion, where N is a positive integer; and performing the conversion based on this determination.
[0019] In another example, a method for video processing is disclosed. The method includes: performing a conversion between video units of a first component of the video and a bitstream representation of the video; wherein the conversion of samples of the first component includes applying a cross-component adaptive loop filter to the sample difference of a second component of the video.
[0020] In another example, a method for video processing is disclosed. The method includes: during the conversion between sub-blocks of video units of video components and a bitstream representation of the video, determining two or more filters from a plurality of filter sets for cross-component adaptive loop filtering; and performing the conversion based on this determination.
[0021] In another example, a method for video processing is disclosed. The method includes: a conversion between a sub-block of a first component of a video unit and a bitstream representation of the video; determining filters that support multiple components across the video or multiple pictures across the video for performing cross-component adaptive loop filtering during the conversion; and performing the conversion based on the determination.
[0022] In another example, a method for video processing is disclosed. This method includes: a conversion between sub-blocks of video units of a video component and a bitstream representation of the video; determining, based on a positional rule, whether to enable a cross-component adaptive loop filter (CC-ALF) for the conversion; and performing the conversion based on that determination.
[0023] In another example, a video encoder apparatus is disclosed. The video encoder includes a processor configured to implement the methods described above.
[0024] In yet another example, a video encoder apparatus is disclosed. The video encoder includes a processor configured to implement the methods described above.
[0025] In yet another example, a computer-readable medium on which code is stored is disclosed. This code embodies one of the methods described herein in the form of processor-executable code.
[0026] These and other features are described throughout this document. Attached Figure Description
[0027] Figure 1 The nominal vertical and horizontal positions of the 4:2:2 luminance and chrominance samples in the image are shown.
[0028] Figure 2 An example of an encoder block diagram is shown.
[0029] Figure 3 67 intra-frame prediction modes are shown.
[0030] Figures 4A-4B Examples of horizontal and vertical transverse scans are shown.
[0031] Figure 5 This shows an example of the positions of the left and top samples involved in LM mode, as well as the sample positions of the current block.
[0032] Figure 6 An example of dividing a 4x8 sample block into two independent decodeable regions is shown.
[0033] Figure 7 An example sequence is shown for processing pixel rows to maximize the throughput of a 4xN block with a vertical predictor.
[0034] Figure 8 This is an example of a low-frequency non-separable transform (LFNST) process.
[0035] Figure 9 An example of the ALF filter shape is shown (chroma: 5×5 rhombus, luminance: 7×7 rhombus).
[0036] Figures 10A-10D An example of subsampling Laplace calculation is shown. Figure 10A The subsampling locations of the vertical gradient are shown. Figure 10B The subsampling locations of the horizontal gradient are shown. Figure 10C The subsampling locations of the diagonal gradient are shown. Figure 10D The subsampling locations of the diagonal gradient are shown.
[0037] Figure 11 An example of block classification at virtual boundaries is shown.
[0038] Figure 12 An example of ALF filtering with modified luminance components at virtual boundaries is shown.
[0039] Figures 13A-13B An example filter is shown. Figure 13A An example arrangement of CC-ALF relative to other loop filters is shown. Figure 13B An example of a diamond filter is shown.
[0040] Figure 14 An example of a 3×4 diamond filter with 8 unique coefficients is shown.
[0041] Figure 15 The shape of the CC-ALF filter with 8 coefficients in JVET-P0106 is shown.
[0042] Figure 16 The shape of the CC-ALF filter with 6 coefficients in JVET-P0173 is shown.
[0043] Figure 17 The shape of the CC-ALF filter with 6 coefficients in JVET-P0251 is shown.
[0044] Figure 18 An example of the JC-CCALF workflow is shown.
[0045] Figures 19 to 34B Example support for filters used in CC-ALF is shown.
[0046] Figure 35 This is a block diagram of an example video processing system that can implement publicly available technologies.
[0047] Figure 36 A block diagram of an example hardware platform used for video processing.
[0048] Figure 37 A flowchart of an example method for video processing.
[0049] Figure 38 A block diagram illustrating an example of a video decoder.
[0050] Figure 39 A block diagram illustrating an example video codec system that can use the techniques disclosed herein.
[0051] Figure 40 A block diagram illustrating an example of a video encoder.
[0052] Figures 41 to 49 A flowchart of an example method for video processing. Detailed Implementation
[0053] The chapter headings used in this document are for ease of understanding and do not limit the applicability of the techniques and embodiments disclosed in each chapter to that chapter. Furthermore, the use of H.266 terminology in some descriptions is merely for ease of understanding and not intended to limit the scope of the disclosed techniques. Therefore, the techniques described herein are also applicable to other video codec designs.
[0054] 1. Introduction
[0055] This document relates to video codec technology. Specifically, it covers cross-component adaptive loop filter (CC-ALF) and other codec tools in picture / video codecs. It can be applied to existing video codec standards (such as HEVC) or the finalized standard (Multi-Functional Video Codec). It can also be applied to future video codec standards or video codecs.
[0056] 2. Brief description
[0057] Video codec standards have primarily evolved through the development of well-known ITU-T and ISO / IEC standards. ITU-T developed the H.261 and H.263 standards, while ISO / IEC developed the MPEG-1 and MPEG-4 Visual standards. The two organizations jointly developed the H.262 / MPEG-2 video standard, the H.264 / MPEG-4 Advanced Video Codec (AVC) standard, and the H.265 / HEVC standard. Starting with H.262, video codec standards are based on a hybrid video codec architecture, using temporal prediction plus transform coding. To explore future video codec technologies beyond HEVC, the Joint Video Exploration Team (JVET) was jointly established by VCEG and MPEG in 2015. Since then, JVET has adopted many new methods and incorporated them into a reference software called the Joint Exploration Model (JEM). In April 2018, the Joint Video Experts Team (JVET) was established between VCEG (Q6 / 16) and ISO / IEC JTC1 SC29 / WG11 (MPEG) to work on the VVC standard, with the goal of reducing the bit rate by 50% compared to HEVC.
[0058] The latest version of the VVC draft, namely Multi-Functional Video Codec (Draft 7), can be found at the following URL:
[0059] http: / / phenix.it-sudparis.eu / jvet / doc_end_user / documents / 16_Geneva / wg11 / JVET-P2001-v14.zip
[0060] The latest reference software for VVC, called VTM, can be found at the following website:
[0061] https: / / vcgit.hhi.fraunhofer.de / jvet / VVCSoftware_VTM / - / tags / VTM-7.Q
[0062] 2.1. Color Space and Chroma Subsampling
[0063] A color space, also known as a color model (or color system), is an abstract mathematical model that simply describes the range of colors as tuples of numbers, typically 3 or 4 values or color components (e.g., RGB). Essentially, a color space is a detailed description of a coordinate system and its subspaces.
[0064] For video compression, the most frequently used color spaces are YCbCr and RGB.
[0065] YCbCr, Y′CbCr, or Y Pb / Cb Pr / Cr, also written as YCBCR or Y′CBCR, is a family of color spaces used as part of the color imaging pipeline in video and digital photography systems. Y′ is the luminance component, and CB and CR are the blue and red chromaticity components, respectively. Y′ (with an apostrophe) is distinguished from Y (luminance), meaning that light intensity is non-linearly encoded based on gamma-corrected RGB primary colors.
[0066] Chromatic subsampling is a practice that encodes images by achieving a lower resolution for chromatic information than for luminance information, taking advantage of the fact that the human visual system is less sensitive to chromatic differences than to luminance differences. 2.1.1. 4:4:4
[0068] Each of the three Y′CbCr components has the same sampling rate, therefore there is no chromaticity subsampling. This scheme is sometimes used in high-end film scanners and film post-production. 2.1.2. 4:2:2
[0070] The two chroma components are sampled at half the luminance sampling rate: the horizontal chroma resolution is halved, while the vertical chroma resolution remains unchanged. This reduces the bandwidth of the uncompressed video signal by one-third with virtually no parallax. (In the VVC working draft...) Figure 1 The image depicts an example of the nominal vertical and horizontal positions of the 4:2:2 color format.
[0071] Figure 1 The nominal vertical and horizontal positions of the 4:2:2 luminance and chrominance samples in the image are shown. 2.1.3. 4:2:0
[0073] In 4:2:0, the horizontal sampling is twice that of 4:1:1, but the vertical resolution is halved because the Cb and Cr channels are sampled only on each alternating line in this scheme. The data rate is therefore the same. Cb and Cr are subsampled horizontally and vertically, each by a factor of 2. There are three variations of the 4:2:0 scheme with different horizontal and vertical addressing.
[0074] In MPEG-2, Cb and Cr are horizontally co-addressed. Cb and Cr are vertically addressed between pixels (interval addressing).
[0075] • In JPEG / JFIF, H.261, and MPEG-1, Cb and Cr are intermittently and intermittently located between alternating luminance samples.
[0076] In a 4:2:0 DV, Cb and Cr are co-located in the horizontal direction. In the vertical direction, they are co-located on alternating lines.
[0077] Table 2-1. SubWidthC and SubHeightC values derived from chroma_format_idc and separate_colour_plane_flag
[0078]
[0079] 2.2. Encoding and decoding process of a typical video codec
[0080] Figure 2 An example of a VVC encoder block diagram is shown, comprising three loop filtering blocks: Deblocking Filter (DF), Sample Adaptive Offset (SAO), and ALF. Unlike DF, which uses predefined filters, SAO and ALF utilize the original samples of the current image to reduce the mean square error between the original and reconstructed samples, respectively, by adding an offset and by applying a Finite Impulse Response (FIR) filter, using information signaling from the encoder / decoder side to inform the offset and filter coefficients. ALF is located in the final processing stage of each image and can be viewed as a tool attempting to capture and repair artifacts created by the preceding stages.
[0081] 2.3. Intra-mode encoding and decoding with 67 intra-prediction modes
[0082] To capture arbitrary edge directions presented in natural video, the number of directional intra-frame modes was expanded from the 33 used by HEVC to 65. Additional directional modes were added in... Figure 3 The image is depicted as a red dashed arrow, and the planar and DC modes remain the same. These denser directional intra-prediction modes are applicable to all block sizes and both luma and chroma intra-prediction.
[0083] The intra-frame prediction direction for standard angles is defined as 45 degrees to -135 degrees clockwise, such as... Figure 3 As shown. In VTM, for non-square blocks, several regular angular intra-prediction modes are adaptively replaced with wide-angle intra-prediction modes. The replaced modes are signaled using the original method and remapped to the wide-angle mode index after parsing. The total number of intra-prediction modes remains unchanged at 67, and the intra-mode encoding and decoding remain unchanged.
[0084] In HEVC, each intra-codec block has a square shape, and the length of each side is a power of 2. Therefore, generating an intra-predictor using DC mode does not require division. In VVC, blocks can have rectangular shapes, making division for each block necessary in general. To avoid division in DC prediction, only the longer side is used to calculate the average of non-square blocks.
[0085] 2.4. Inter-frame prediction
[0086] For each inter-frame predicted CU, motion parameters include motion vectors, reference picture indexes, reference picture list usage indexes, and additional information required for the new encoding / decoding features of the VVC used to generate inter-frame predicted samples. Motion parameters can be signaled explicitly or implicitly. When encoding / decoding a CU in skip mode, the CU is associated with a PU and has no significant residual coefficients, no encoded / decoded motion vector increments, or reference picture indexes. A merge mode is specified, thereby obtaining the motion parameters of the current CU from neighboring CUs, including spatial and temporal candidates, and additional scheduling introduced in the VVC. The merge mode can be applied to any inter-frame predicted CU, not just skip mode. An alternative to the merge mode is explicit transmission of motion parameters, where motion vectors, corresponding reference picture indexes for each reference picture list, reference picture list usage flags, and other necessary information are explicitly signaled to each CU.
[0087] 2.5. Intra-Block Copying (IBC)
[0088] Intra-Block Copy (IBC) is a tool used in the HEVC extension on SCC. It is well-known for significantly improving the encoding and decoding efficiency of screen content material. Since IBC mode is implemented as a block-level encoding / decoding mode, block matching (BM) is performed at the encoder to find the optimal block vector (or motion vector) for each CU. Here, the block vector indicates the displacement from the current block to a reference block that has already been reconstructed within the current image. The luma block vector of the IBC-encoded CU is rounded to integer precision. The chroma block vector is also rounded to integer precision. When used in conjunction with AMVR, IBC mode can switch between 1-pixel and 4-pixel motion vector precision. The CU of the IBC-encoded CU is treated as a third prediction mode in addition to intra-frame or inter-frame prediction modes. IBC mode is suitable for CUs with a width and height of 64 luma samples or less.
[0089] On the encoder side, hash-based motion estimation is performed on the IBC. The encoder performs RD checks on blocks with a width or height not exceeding 16 luminance samples. For non-merge modes, a block vector search is first performed using a hash-based search. If the hash search does not return valid candidates, a local search based on block matching is performed.
[0090] In hash-based search, hash key matching (32-bit CRC) between the current block and reference blocks is extended to all allowed block sizes. Hash key calculation for each location in the current image is based on 4×4 sub-blocks. For larger current blocks, a hash key is determined to be the matching reference block's hash key when all hash keys in all 4×4 sub-blocks match the hash key at the corresponding reference location. If multiple reference blocks are found to match the current block's hash key, the block vector cost for each matching reference is calculated, and the block vector cost with the lowest cost is selected.
[0091] In block matching search, the search scope is set to cover both the previous and current CTUs.
[0092] At the CU level, IBC mode is notified via flag signaling, and it can be signaled as either IBC AMVP mode or IBC skip / merge mode, as follows:
[0093] - IBC Skip / Merge Mode: The merge candidate index is used to indicate which block vector from the list of neighboring candidate IBC encoded blocks is used to predict the current block. The merge list consists of a spatial domain, HMVP, and paired candidates.
[0094] -IBC AMVP mode: Block vector difference is encoded and decoded in the same way as motion vector difference. The block vector prediction method uses two candidates as predictors, one from the left neighbor and one from the upper neighbor (if IBC encoding / decoding is used). When either neighbor is unavailable, the default block vector will be used as the predictor. Signaling notification flags indicate the block vector predictor index.
[0095] 2.6. Palette Mode
[0096] For palette mode signaling, the palette mode is encoded and decoded into the prediction mode of the codec unit; that is, the prediction mode of the codec unit can be MODE_INTRA, MODE_INTER, MODE_IBC, and MODE_PLT. If palette mode is used, pixel values in the CU are represented by a small set of representative color values. This set is called the palette. For pixels with values close to the palette colors, signaling informs the palette index. For pixels with values outside the palette, the pixel is represented by an escape symbol, and the quantized pixel value is directly signaled.
[0097] To decode a palette-encoded block, the decoder needs to decode both the palette colors and indices. Palette colors are described by a palette table and encoded by a palette table encoding / decoding tool. An escape flag is signaled to each CU to indicate the presence of an escape character in the current CU. If an escape character is present, the palette table is incremented by 1, and the last index is assigned to the escape mode. The palette indices of all pixels in the CU form a palette index map and are encoded by a palette index map encoding / decoding tool.
[0098] For the encoding and decoding of the palette table, the palette predictor is maintained. The predictor is initialized at the beginning of each stripe, where it is reset to 0. For each entry in the palette predictor, a reuse flag is signaled to indicate whether it is part of the current palette. The reuse flag is sent using zero-run-length encoding and decoding. Subsequently, the number of new palette entries is signaled using zero-order exponent Golomb codes. Finally, the component values of the new palette entries are signaled. After encoding the current CU, the palette predictor is updated using the current palette, and entries from previous palette predictors that are not reused in the current palette are added to the end of the new palette predictor until the maximum allowed size (palette fill) is reached.
[0099] To encode and decode the palette index map, horizontal and vertical transverse scans are used, as shown in Figure 4. The scan order is explicitly signaled in the bitstream using the `palette_transpose_flag`.
[0100] Palette indexes are encoded and decoded using two primary palette sample modes: "INDEX" and "COPY_ABOVE". A flag is used to signal the mode except when the top row is used in horizontal scans or the first column or preceding column in vertical scans is in "COPY_ABOVE". In "COPY_ABOVE" mode, the palette index of the sample in the previous row is copied. In "Index" mode, the palette index is explicitly signaled. For both "INDEX" and "COPY_ABOVE" modes, a run-length value is signaled, specifying the number of pixels encoded and decoded using the same mode.
[0101] The encoding order of the index graph is as follows: First, the signaling notifies the number of index values for the CU. Next, truncated binary encoding / decoding is used to signal the actual index values for the entire CU. Both the index number and index values are encoded / decoded in bypass mode. This groups the bypass bins associated with the indexes together. Then, the palette mode (INDEX or COPY_ABOVE) and runs are signaled in an interleaved manner. Finally, the component escape values corresponding to the escape samples of the entire CU are grouped together and encoded / decoded in bypass mode. After signaling the index values, the signaling notification appends the syntax element `last_run_type_flag`. This syntax element, combined with the number of indexes, eliminates the need for signaling the run value corresponding to the last run in the block.
[0102] In VTM, dual trees are enabled for I-stripes, separating the codec units for luma and chroma. Therefore, in this recommendation, the palette is applied separately to the luma (Y component) and chroma (Cb and Cr components). If dual trees are disabled, the palette will be applied jointly to the Y, Cb, and Cr components, the same as the HEVC palette.
[0103] 2.7. Prediction using a cross-component linear model
[0104] In VVC, a cross-component linear model (CCLM) prediction mode is used to predict chromaticity samples based on reconstructed luminance samples from the same CU using the following linear model:
[0105] pred C (i, j) = α·rec L ′(i,j)+β (2-1)
[0106] Among them, pred C (i, j) represent the predicted chromaticity samples in the CU, and rec L (i, j) represent downsampled reconstructed luminance samples of the same CU.
[0107] Figure 5 Example locations of sample points used to derive α and β are shown.
[0108] Besides using the top and left templates together in LM mode to calculate linear model coefficients, they can also be used alternately in two other LM modes, called LM_A and LM_L modes. In LM_A mode, only the top template is used to calculate linear model coefficients. To obtain more sample points, the top template is expanded to (W+H). In LM_L mode, only the left template is used to calculate linear model coefficients. To obtain more sample points, the left template is expanded to (H+W). For non-square blocks, the top template is expanded to W+W, and the left template is expanded to H+H.
[0109] The CCLM parameters (α and β) are derived from at most four adjacent chroma samples and their corresponding downsampled luminance samples. Assuming the current chroma block dimension is W×H, then W′ and H′ are set to...
[0110] - When applying the LM pattern, W′=W, H′=H;
[0111] - When applying the LM-A mode, W' = W + H;
[0112] - When applying the LM-L pattern, H' = H + W;
[0113] The adjacent positions above are denoted as S[0, -1]...S[W'-1, -1], and the adjacent positions to the left are denoted as S[-1, 0]...S[-1, H'-1]. These four sample points are then selected as...
[0114] -When LM mode is applied and the upper and left adjacent samples are available, S[W' / 4, -1], S[3W' / 4, -1], S[-1, H' / 4], S[-1, 3H' / 4];
[0115] - When applying LM-A mode or when only the upper adjacent sample points are available, S[W' / 8, -1], S[3W' / 8, -1], S[5W' / 8, -1], S[7W' / 8, -1];
[0116] - When applying LM-L mode or when only the left adjacent sample point is available, S[-1, H' / 8], S[-1, 3H' / 8], S[-1, 5H' / 8], S[-1, 7H' / 8];
[0117] The four adjacent brightness samples at the selected location are downsampled and compared four times to find the two smaller values: x 0 A and x 1 A and two larger values: x 0 B and x 1 B Their corresponding chromaticity sample values are represented as y 0 A y 1 A y 0 B and y 1 B Then x A x B y A and y B Export as:
[0118] X a =(x0 A +x 1 A +1) >> 1; X b =(x 0 B +x 1 B +1)>>1; Y a =(y 0 A +y 1 A +1)>>1; Y b =(y 0 B +y 1 B +1)>>1 (2-1)
[0119] Finally, the linear model parameters α and β are obtained according to the following equation.
[0120]
[0121] β=Y b -α·X b (2-3)
[0122] The division operation for calculating the parameter α is implemented using a lookup table. To reduce the memory required to store the table, the diff value (the difference between the maximum and minimum values) and the parameter α are represented using exponential notation. For example, the diff is approximated using a 4-bit significant part and an exponent. Therefore, for a 16-digit significant value, the table for 1 / diff is reduced to 16 elements, as shown below:
[0123] DivTable[]={0,7,6,5,5,4,4,3,3,2,2,1,1,1,1,0} (2-4)
[0124] This will help reduce the complexity of the calculations and the memory size required to store the necessary tables.
[0125] To match the chroma sample locations in a 4:2:0 video sequence, two types of downsampling filters are applied to the luminance samples to achieve a 2:1 downsampling ratio in both the horizontal and vertical directions. The selection of the downsampling filters is specified by the SPS level flag. The two downsampling filters are as follows, corresponding to "type-0" and "type-2" content, respectively.
[0126]
[0127]
[0128] Note that when the upper reference line is located at the CTU boundary, only one luminance line (the universal line buffer in intra-frame prediction) is used to downsample luminance samples.
[0129] This parameter calculation is performed as part of the decoding process, not just as part of the encoder search operation. Therefore, the α and β values are not passed to the decoder using syntax.
[0130] For chroma intra-mode encoding and decoding, a total of eight intra-modes are allowed. These modes include five traditional intra-modes and three cross-component linear model modes (LM, LM_A, and LM_L). The chroma mode signaling notification and derivation process are shown in Table 2-2. Chroma mode encoding and decoding directly depends on the intra-prediction mode of the corresponding luma block. Since separate block partitioning structures for luma and chroma components are enabled in I-strips, one chroma block can correspond to multiple luma blocks. Therefore, for chroma DM mode, the intra-prediction mode of the corresponding luma block covering the center position of the current chroma block is directly inherited.
[0131] Table 2-2. Exporting Chroma Prediction Mode from Luminance Mode when cclm_is is enabled
[0132]
[0133] 2.8. Block Differential Pulse Code Modulation and Decoding (BDPCM)
[0134] BDPCM was proposed in JVET-M0057. Due to the shape of the horizontal (vertical) predictor, it uses the left (A) (top (B)) pixels for the prediction of the current pixel. The most throughput-efficient way to process this block is to process all pixels in a column (row) in parallel and process these columns (rows) sequentially. To increase throughput, we introduce the following process: when the predictor selected on the block is vertical, the block with a width of 4 is divided into two halves with horizontal boundaries, and when the predictor selected on the block is horizontal, the block with a height of 4 is divided into two halves with vertical boundaries.
[0135] When blocks are divided, samples from one region are not allowed to use pixels from another region to compute predictions: if this happens, the predicted pixel is replaced by a reference pixel in the prediction direction. This applies to different positions of the current pixel X within a 4×8 block for vertical prediction. Figure 6 As shown in the image.
[0136] Figure 6 An example of dividing a 4x8 sample block into two independent decodeable regions is shown.
[0137] Because of this characteristic, it is now possible to process 4×4 blocks in 2 cycles, 4×8 or 8×4 blocks in 4 cycles, and so on. Figure 7 As shown.
[0138] Figure 7 An example sequence is shown for processing pixel rows to maximize the throughput of a 4xN block with a vertical predictor.
[0139] Table 2-3 summarizes the number of cycles required to process a block based on its block size. It is easy to see that any block with two dimensions greater than or equal to 8 can be processed in 8 pixels or more per cycle.
[0140] Table 2-3. Worst-case throughput of blocks with sizes of 4xN and Nx4
[0141]
[0142] 2.9. Quantization Residual Domain BDPCM
[0143] In JVET-N0413, Quantization Residual Domain BDPCM (hereinafter referred to as RBDPCM) was proposed. Similar to intra-frame prediction, intra-frame prediction is performed on the entire block by copying samples in the prediction direction (horizontal or vertical prediction). The residual is quantized, and the increment between the quantized residual and its predictor (horizontal or vertical) quantization value is encoded and decoded.
[0144] For a block of size M (rows) × N (columns), let r i,j , 0≤i≤M-1, 0≤j≤N-1 is the prediction residual after performing intra-frame prediction horizontally (copying the left adjacent pixel values row by row across the prediction block) or vertically (copying the top adjacent row to each row in the prediction block) using unfiltered samples from the upper or left block boundary samples. Let Q(r) be the prediction residual. i,j ), 0≤i≤M-1, 0≤j≤N-1 represent residuals r i,j The quantized version is then used, where the residual is the difference between the original block and the predicted block values. The block DPCM is then applied to the quantized residual samples, producing a result with element-wise... Modified M×N array When the vertical BDPCM is signaled:
[0145]
[0146] For horizontal prediction, similar rules are applied, and residual quantization samples are obtained in the following way.
[0147]
[0148] Residual Quantization Samples It is sent to the decoder.
[0149] On the decoder side, the above calculation is reversed to generate Q(r). i,j ), 0≤i≤M-1, 0≤j≤N-1. For the vertical prediction case,
[0150]
[0151] Regarding the horizontal situation
[0152]
[0153] Inverse quantization residual Q -1 (Q(r i,j Add to the intra-block prediction value to produce the reconstructed sample values.
[0154] The main advantage of this approach is that inverse DPCM can be performed dynamically during coefficient resolution, by simply adding a predictor while resolving the coefficients, or it can be performed after resolution.
[0155] Transform skipping is always used in the quantized residual domain BDPCM.
[0156] 2.10. VVC Multiple Transform Set (MTS)
[0157] In VTM, large block size transforms up to 64×64 are enabled, primarily for higher resolution video such as 1080p and 4K sequences. For transform blocks with a size (width or height, or both) equal to 64, high-frequency transform coefficients are zeroed, thus retaining only low-frequency coefficients. For example, for an M×N transform block, where M is the block width and N is the block height, when M equals 64, only the left 32 columns of transform coefficients are retained. Similarly, when N equals 64, only the first 32 rows of transform coefficients are retained. When transform skip mode is used for large blocks, the entire block is used without zeroing any values. VTM also supports configurable maximum transform sizes in SPS, allowing the encoder to flexibly choose transform sizes up to 16, 32, or 64 in length, depending on the specific implementation requirements.
[0158] In addition to DCT-II already used in HEVC, the Multiple Transform Selection (MTS) scheme is used for residual coding and decoding of both inter-frame and intra-frame codec blocks. It uses multiple transforms selected from DCT8 / DST7. The newly introduced transform matrices are DST-VII and DCT-VIII. Table 2-4 shows the basic functions of the selected DST / DCT.
[0159] Table 2-4. Transform basis functions for DCT-II / VIII and DSTVII for N-point inputs
[0160]
[0161] To maintain the orthogonality of the transformation matrices, the quantization of the transformation matrices is more precise than that in HEVC. To keep the intermediate values of the transformation coefficients within the 16-bit range, all coefficients should be 10 bits after both the horizontal and vertical transformations.
[0162] To control the MTS scheme, separate enable flags are specified for intra-frame and inter-frame use at the SPS level. When MTS is enabled at SPS, signaling is used to notify the CU level flag to indicate whether MTS should be applied. Here, MTS applies only to luminance. Signaling is used to notify the MTS CU level flag when the following conditions are met.
[0163] - Width and height are both less than or equal to 32
[0164] -CBF logo equals one
[0165] If the MTS CU flag is zero, DCT2 is applied in both directions. However, if the MTS CU flag is 1, two additional signaling flags are used to indicate the transform type for the horizontal and vertical directions, respectively. The transform and signaling notification mapping table is shown in Table 2-5. Transform selection for ISP and implicit MTS is unified by removing intra-frame mode and block shape dependencies. If the current block is in ISP mode, or if the current block is an intra-frame block and both intra-frame and inter-frame explicit MTS are enabled, only DST7 is used for both the horizontal and vertical transform cores. Regarding transform matrix precision, an 8-bit primary transform core is used. Therefore, all transform cores used in HEVC remain unchanged, including 4-point DCT-2 and DST-7, 8-point, 16-point, and 32-point DCT-2. Furthermore, other transform cores, including 64-point DCT-2, 4-point DCT-8, 8-point, 16-point, 32-point DST-7, and DCT-8, use an 8-bit primary transform core.
[0166] Table 2-5. Transformation and Signaling Notification Mapping Table
[0167]
[0168] To reduce the complexity of large-sized DST-7 and DCT-8 blocks, the high-frequency transform coefficients are set to zero for DST-7 and DCT-8 blocks with a size (width or height, or both) equal to 32. Only the coefficients in the 16x16 low-frequency region are retained.
[0169] In HEVC, for example, transform skip mode can be used to encode and decode residuals of blocks. To avoid redundancy in syntax encoding and decoding, the transform skip flag is not signaled when the CU-level MTS_CU_flag is not equal to zero. The block size limit for transform skip is the same as the block size limit for MTS in JEM4, indicating that transform skip applies to the CU when both the block width and height are equal to or less than 32. Note that when LFNST or MIP is activated for the current CU, the implicit MTS transform is set to DCT2. Furthermore, implicit MTS can still be enabled when MTS is enabled for inter-frame codec blocks.
[0170] 2.11. Low-Frequency Non-Separable Transform (LFNST)
[0171] In VVC, LFNST (Low Frequency Inseparable Transform) (which is called the Simplified Quadratic Transform) is applied between the positive primary transform and quantization (at the encoder) and between dequantization and the inverse primary transform (at the decoder), such as... Figure 8 As shown. In LFNST, either a 4×4 non-separable transformation or an 8×8 non-separable transformation is applied depending on the block size. For example, 4x4 LFNST is suitable for small blocks (i.e., min(width, height) < 8), and 8x8 LFNST is suitable for larger blocks (i.e., min(width, height) > 4).
[0172] Using the input as an example, the application of the inseparable transformation used in LFNST is described below. To apply 4x4 LFNST, a 4x4 input block X...
[0173]
[0174] First, it is represented as a vector.
[0175]
[0176] Inseparable transformation is calculated as in The vector indicates the transformation coefficients, and T is a 16x16 transformation matrix. (16×1 coefficient vector) The scan order (horizontal, vertical, or diagonal) of this block is then reorganized into 4×4 blocks. Coefficients with smaller indices are placed in the 4×4 coefficient blocks along with their smaller scan indices.
[0177] 2.11.1. Simplified Inseparable Transformation
[0178] LFNST (Low Frequency Inseparable Transform) applies the inseparable transformation based on a direct matrix multiplication method, enabling it to be implemented in a single channel without multiple iterations. However, it is necessary to reduce the dimension of the inseparable transformation matrix to minimize computational complexity and memory space for storing the transformation coefficients. Therefore, a simplified inseparable transformation (or RST) method is used in LFNST. The main idea of the simplified inseparable transformation is to map an N-dimensional vector (for 8×8 NSST, N is usually equal to 64) to an R-dimensional vector in a different space, where N / R (R < N) is the simplification factor. Thus, instead of an N×N matrix, the RST matrix becomes an R×N matrix, as shown below:
[0179]
[0180] The R rows of the transform form R bases in N-dimensional space. The inverse transform matrix of RT is the transpose of its forward transform. For the 8x8 LFNST, a simplification factor of 4 is applied, and the 64×64 direct matrix, which is the size of a regular 8×8 inseparable transform matrix, is reduced to a 16×48 direct matrix. Therefore, a 48×16 inverse RST matrix is used on the decoder side to generate the core (primary) transform coefficients in the upper left region of the 8×8. When the 16×48 matrix is applied instead of the 16×64 matrix with the same transform set configuration, each matrix takes 48 input data from three 4×4 blocks in the upper left 8×8 block, excluding the lower right 4×4 block. With the help of the simplified dimensions, the memory usage for storing all LFNST matrices is reduced from 10KB to 8KB, where the performance degradation is reasonable. To reduce complexity, the LFNST is restricted to only being applied when all coefficients outside the first coefficient subgroup are unimportant. Therefore, when the LFNST is applied, all primary-only transform coefficients must be zero. This allows for adjustment of the LFNST index signaling notification at the last valid position, thus avoiding the additional coefficient scan in the current LFNST design, which is only needed to check valid coefficients at specific positions. The worst-case handling of LFNST (in terms of per-pixel multiplication) restricts the non-separable transformations of 4×4 and 8×8 blocks to 8×16 and 8×48 transformations, respectively. In these cases, the last valid scan position must be less than 8 when LFNST is applied, and for other sizes less than 16. For blocks with shapes of 4xN and Nx4 where N > 8, the proposed restrictions mean that LFNST is now applied only once, and only to the top-left 4x4 region. Since all primary-only coefficients are zero when LFNST is applied, the number of operations required for the primary transformation is reduced in this case. From the encoder's perspective, coefficient quantization is significantly simplified when testing the LFNST transformation. For the first 16 coefficients (in scan order), rate-distortion optimized quantization must be performed to the maximum extent, while the remaining coefficients are forced to zero.
[0181] 2.11.2. LFNST Transform Selection
[0182] There are a total of 4 transform sets in LFNST, and each transform set uses 2 inseparable transform matrices (kernels). The mapping from intra-prediction modes to transform sets is predefined, as shown in Table 2-6. If one of the three CCLM modes (INTRA_LT_CCLM, INTRA_T_CCLM, or INTRA_L_CCLM) is used for the current block (81 <= predModeIntra <= 83), then transform set 0 is selected for the current chroma block. For each transform set, the selected inseparable quadratic transform candidate is further specified by an LFNST index that is explicitly signaled. This index is signaled once in the bitstream after the transform coefficients.
[0183] Table 2-6. Transformation Selection Table
[0184]
[0185] 2.11.3. LFNST Index Signaling Notifications and Interaction with Other Tools
[0186] Because LFNST is restricted to application only when all coefficients outside the first coefficient subgroup are insignificant, LFNST index encoding / decoding depends on the position of the last valid coefficient. Furthermore, LFNST indexing is context-coded but not dependent on the intra-prediction mode, and only the first bin is context-coded. Additionally, LFNST applies to intra-CUs in both intra- and inter-frame stripes, and to both luma and chroma. If dual-tree is enabled, the LFNST indexes for luma and chroma are signaled separately. For inter-frame stripes (where dual-tree is disabled), a single LFNST index is signaled and used for both luma and chroma.
[0187] When ISP mode is selected, LFNST is disabled, and the RST index is not signaled because the performance improvement is limited even if RST is applied to every feasible partition block. Furthermore, disabling RST for ISP prediction residuals reduces coding complexity. When MIP mode is selected, LFNST is also disabled, and no signaling is used to notify the index.
[0188] Considering that due to the existing maximum transform size limit (64x64), large CUs larger than 64x64 are implicitly partitioned (TU tiling), LFNST index search can increase the data buffer size by four times for a certain number of decoding pipeline stages. Therefore, the maximum size allowed by LFNST is limited to 64x64. Note that LFNST is only enabled with DCT2.
[0189] 2.12. Skip chromaticity transformation
[0190] Chroma Transform Skip (TS) was introduced in VVC. The motivation was to unify TS and MTS signaling notifications between luma and chroma by relocating `transform_skip_flag` and `mts_idx` to the `residual_coding` section. A context model was added for chroma TS. For `mts_idx`, no context model or binary representation was changed. Furthermore, TS residual encoding / decoding was also applied when using chroma TS.
[0191] Semantics
[0192] The terms to be deleted are indicated by underlined italic text.
[0193] `transform_skip_flag[x0][y0][cIdx]` specifies whether the transformation is applied to the associated elements. brightness Transform block. Array indices x0 and y0 specify the position (x0, y0) of the top-left luminance sample of the transform block under consideration relative to the top-left luminance sample of the image. `transform_skip_flag[x0][y0][cIdx]` equal to 1 indicates that the current transform block will not be skipped. brightness The transform block applies the transform. The array index cIdx specifies the indicator for the color components; luminance equals 0, Cb equals 1, and Cr equals 2. `transform_skip_flag[x0][y0][cIdx]`, which is 0, indicates whether the transform is applied to the current transform block; the decision depends on other syntax elements. If `transform_skip_flag[x0][y0][cIdx]` does not exist, it is inferred to be 0.
[0194] 2.13. BDPCM of chromaticity
[0195] In addition to chroma TS support, BDPCM is added to the chroma component. If sps_bdpcm_enable_flag is 1, another syntax element, sps_bdpcm_chroma_enable_flag, is added to SPS. These flags have the following behavior, as shown in Table 2-7.
[0196] Table 2-7. SPS Marker for Luminosity and Chromaticity BDPCM
[0197]
[0198] When BDPCM is only used for luma, the current behavior remains unchanged. When BDPCM is also used for chroma, a bdpcm_chroma_flag is sent for each chroma block. This indicates whether BDPCM was used on the chroma block. When used, BDPCM is applied to both chroma components, and an additional bdpcm_dir_chroma flag is encoded and decoded to indicate the prediction direction for both chroma components.
[0199] The deblocking filter is deactivated at the boundary between two Block-DPCM blocks because neither block uses the transform stage that is typically responsible for block artifacts. This deactivation occurs independently for the luma and chroma components.
[0200] 2.14.ALF
[0201] In VVC, an adaptive loop filter (ALF) with block-based filter adaptation is applied. For the luminance component, one of 25 filters is selected for each 4×4 block based on the direction and activity of the local gradient.
[0202] 2.14.1. Filter Shape
[0203] Two diamond filter shapes were used (e.g.) Figure 9 (As shown). A 7×7 rhombus is applied to the luminance component, and a 5×5 rhombus is applied to the chrominance component.
[0204] 2.14.2. Block Classification
[0205] For the luminance component, each 4×4 block is divided into one of 25 classes. The classification index C is based on its orientation D and activity. The quantization values are exported as follows:
[0206]
[0207] To calculate D and First, the gradients in the horizontal, vertical, and two diagonal directions are calculated using the 1-D Laplacian operator:
[0208]
[0209]
[0210]
[0211]
[0212] Here, indices i and j refer to the coordinates of the top left sample point within the 4×4 block, and R(i,j) indicates the reconstructed sample point at coordinates R(i,j).
[0213] To reduce the complexity of block classification, subsampling 1-D Laplace calculation is applied. For example... Figures 10A-10D As shown, the same subsampling position is used for gradient calculation in all directions.
[0214] Figure 10A The subsampling locations of the vertical gradient are shown. Figure 10B The subsampling locations of the horizontal gradient are shown. Figure 10C The subsampling locations of the diagonal gradient are shown. Figure 10D The subsampling locations of the diagonal gradient are shown.
[0215] The maximum and minimum values of the gradient D in the horizontal and vertical directions are then set as follows:
[0216]
[0217] The maximum and minimum values of the gradients in the two diagonal directions are set as follows:
[0218]
[0219] To derive the values of the directionality D, these values are compared with each other and with two thresholds t1 and t2:
[0220] Step 1. If and If both are true, then D is set to 0.
[0221] Step 2. If Continue from step 3; otherwise, continue from step 4.
[0222] Step 3. If D is set to 2; otherwise, D is set to 1.
[0223] Step 4. If D is set to 4; otherwise, D is set to 3.
[0224] Activity value A is calculated as follows:
[0225]
[0226] A is further quantized to the range of 0 to 4, and the quantized value is represented as
[0227] For the chromaticity components in the image, no classification method is applied; that is, a single set of ALF coefficients is applied to each chromaticity component.
[0228] 2.14.3. Geometric Transformation of Filter Coefficients and Trimmed Values
[0229] Before filtering each 4×4 luma block, geometric transformations, such as rotation or diagonal and vertical flipping, are applied to the filter coefficients f(k, l) and the corresponding filter trimming values c(k, l) based on the gradient values calculated for that block. This is equivalent to applying these transformations to samples in the filter support region. The idea is to make different blocks to which ALF is applied more similar by aligning the orientations of the different blocks.
[0230] Three geometric transformations are introduced: diagonal, vertical flip, and rotation.
[0231] Diagonal: f D (k, l) = f(l, k), c D (k, l) = c(l, k), (2-22)
[0232] Vertical flip: f V (k, l) = f(k, Kl-1), c V (k, l) = c(k, Kl-1) (2-23)
[0233] Rotation: f R (k, l) = f(Kl-1, k), c R (k, l) = c(Kl-1, k) (2-24)
[0234] Where K is the size of the filter, and 0 ≤ k, l ≤ K⁻¹ are the coefficient coordinates, so position (0, 0) is at the top left corner, and position (K⁻¹, K⁻¹) is at the bottom right corner. Based on the gradient values calculated for this block, the transform is applied to the filter coefficients f(k, l) and the trimming value c(k, l). The table below summarizes the relationship between the transform and the four gradients in the four directions.
[0235] Table 2-8 shows the gradient and transformation mapping for a block computation.
[0236] gradient value Transform <![CDATA[g d2 <g d1 and g h <g v ]]> No change <![CDATA[g d2 <g d1 and g v <g h ]]> diagonal <![CDATA[g d1 <g d2 and g h <g v ]]> Flip vertically <![CDATA[g d1 <g d2 and g v <g h ]]> Rotation
[0237] 2.14.4. Filter Parameter Signaling Notification
[0238] ALF filter parameters are signaled in the Adaptive Parameter Set (APS). An APS can signal a set of up to 25 luminance filter coefficients and trimmed value indices, and a set of up to 8 chrominance filter coefficients and trimmed value indices. To reduce bit overhead, filter coefficients from different categories of luminance components can be merged. In the stripe header, the signaling informs the index of the APS used for the current stripe.
[0239] The clipping index decoded from the APS allows the clipping value to be determined using a table of clipping values for both the luminance and chrominance components. These clipping values depend on the internal bit depth. More precisely, the clipping value is obtained using the following formula:
[0240] AlfClip = {round(2 B-α*n (2-25) for n∈[0..N-1]}
[0241] Where b equals the internal bit depth, α is a predefined constant of 2.35, and N equals 4, which is the number of trim values allowed in VVC.
[0242] In the stripe header, up to seven APS indices can be signaled to specify the set of luma filters used for the current stripe. Further control over the filtering process can be achieved at the CTB level. A signaling flag is always used to indicate whether an ALF is applied to the luma CTB. The luma CTB can select a filter set from 16 fixed filter sets and filter sets from the APS. A signaling filter set index is used for the luma CTB to indicate which filter set to apply. The 16 fixed filter sets are predefined and hard-coded in both the encoder and decoder.
[0243] For chroma components, the signaling in the stripe header informs the APS index to indicate the set of chroma filters used for the current stripe. At the CTB level, if there is more than one set of chroma filters in the APS, the filter index is notified for each chroma CTB.
[0244] The filter coefficients are quantized using a norm equal to 128. To limit the multiplication complexity, bitstream consistency is applied, ensuring that coefficients at non-center positions are within -2. 7 Up to 2 7 The range is -1, including -2. 7 and 2 7 -1. The center position coefficient is not signaled in the bitstream and is assumed to be equal to 128.
[0245] 2.14.5. Filtering Process
[0246] On the decoder side, when ALF is enabled for CTB, each sample R(i,j) within the CU is filtered, producing sample values R′(i,j) as shown below.
[0247] R′(i,j)=R(i,j)+((∑ k≠0 ∑ l≠0 f(k,l)×K(R(i+k,j+l)-R(i,j),c(k,l))+64)>>7)
[0248] (2-26)
[0249] Where f(k, l) represents the filter coefficients for decoding, K(x, y) is the trimming function, and c(k, l) represents the trimming parameters for decoding. The variables k and l in... and The values vary between L and L, where L represents the filter length. The corresponding trimming function for the function Clip3(-y, y, x) is K(x, y) = min(y, max(-y, x)).
[0250] 2.14.6. Virtual boundary filtering process for reducing line buffer size
[0251] In VVC, to reduce the line buffer requirements of ALF, modified block classification and filtering are used for samples near the horizontal CTU boundary. For this purpose, as follows... Figure 11 As shown, by shifting the horizontal CTU boundary with “N” sample points, the virtual boundary is defined as a line, where N equals 4 for the luminance component and N equals 2 for the chrominance component.
[0252] like Figure 12 As shown, the modified block classification is applied to the luminance component. For the 1D Laplacian gradient calculation of the 4×4 block above the virtual boundary, only samples above the virtual boundary are used. Similarly, for the 1D Laplacian gradient calculation of the 4×4 block below the virtual boundary, only samples below the virtual boundary are used. The quantization of the active value A is scaled accordingly by taking into account the reduced number of samples used in the 1D Laplacian gradient calculation.
[0253] For filtering, symmetrical padding at virtual boundaries is applied to both the luminance and chrominance components. For example... Figure 12 As shown, when the filtered sample point is below the virtual boundary, the adjacent sample points above the virtual boundary are filled. At the same time, the corresponding sample points on the other side are also filled symmetrically.
[0254] Unlike the symmetrical padding method used at horizontal CTU boundaries, a simple padding process is applied to strip, patch, and sub-image boundaries when cross-boundary filters are disabled. The simple padding process is also applied to image boundaries. Padding samples are used for both classification and filtering processes.
[0255] 2.15. JVET-P0080: CE5-2.1, CE5-2.2: Transcomponent Adaptive Loop Filter
[0256] The following Figure 13A The placement of the CC-ALF relative to other loop filters is shown. The CC-ALF applies a linear diamond filter to the luminance channel of each chroma component. Figure 13B To operate, it is represented as
[0257]
[0258] in
[0259] (x, y) is the position of the refined chromaticity component i.
[0260] (x C ,y C () is the brightness position based on (x, y).
[0261] S i It is a filter that supports the chromaticity component i in luminance.
[0262] c i (x0, y0) represent the filter coefficients
[0263] (2-28)
[0264] Supported area around its centered brightness position (x C ,y C The CC-ALF coefficients are calculated based on the spatial scaling factor between the luma and chroma planes. All filter coefficients are sent in the AP and have an 8-bit dynamic range. The APS can be referenced in the strip header. The CC-ALF coefficients for each chroma component of the strip are also stored in a buffer corresponding to the temporal sublayer. Using strip-level flags facilitates the reuse of these temporal sublayer filter coefficient sets. The application of CC-ALF filters is controlled on variable block sizes (i.e., 16×16, 32×32, 64×64, 128×128) and is notified via context codec flag signaling received for each sample block. The strip-level block size and CC-ALF enable flag are received for each chroma component. Boundary padding for horizontal virtual boundaries utilizes repetition. For the remaining boundaries, the same type of padding as regular ALF is used.
[0265] 2.16. JVET-P1008: CE5 Related: Design of CC-ALF
[0266] In JVET-O0636 and CE5-2.1, a cross-component adaptive loop filter (CC-ALF) was introduced and studied. This filter uses a linear filter to filter the luminance sample values and generates residual correction for the chrominance channels from the concatenated filtered output. This filter is designed to operate in parallel with existing luminance ALFs.
[0267] A CC-ALF design is proposed, which is asserted to be simplified and better aligned with existing ALFs. This design uses a 3x4 rhombus with 8 unique coefficients. This reduces the number of multiplications by 43% compared to the 5x6 design studied in CE5-2.1. When chroma ALF or CC-ALF is enabled for the chroma components of the CTU, the multiplier count per pixel is limited to 16 (compared to 15 for current ALF). The dynamic range of the filter coefficients is limited to 6-bit signed. Figure 14 The filter descriptions for both the proposed solution and the CE5-2.1 solution are shown.
[0268] To better align with existing ALF designs, filter coefficients are signaled in the APS. Up to four filters are supported, and filter selection is indicated at the CTU level. Symmetry line selection is used at virtual boundaries for further ALF reconciliation. Finally, to limit the storage required for the corrected output, the CC-ALF residual output is pruned to -2. BitDepthC-1 Up to 2 BitDepthC-1 -1, including -2 BitDepthC-1 and 2 BitDepthC-1 -1.
[0269] 2.17. Simplified method of CC-ALF in JVET-P2025
[0270] 2.17.1. Alternative Filter Shapes
[0271] As described in this article, the shape of the CC-ALF filter is modified to have 8 or 6 coefficients as shown in the figure.
[0272] Figure 15 The shape of the CC-ALF filter with 8 coefficients in JVET-P0106 is shown.
[0273] Figure 16 The shape of the CC-ALF filter with 6 coefficients in JVET-P0173 is shown. Figure 16 , 17 In 19-34B, luminance samples are indicated by solid-color circles, chrominance samples are indicated by circles with patterns inside, and one or more thick circles with patterns inside indicate chrominance samples determined based on luminance samples indicated by squares surrounding the luminance samples.
[0274] Figure 17 The shape of the CC-ALF filter with 6 coefficients in JVET-P0251 is shown.
[0275] 2.17.2. Joint Chromaticity Cross-Component Adaptive Filtering
[0276] The Joint Chromaticity Cross-Component Adaptive Loop Filter (JC-CCALF) uses only a set of CCALF filter coefficients trained at the encoder to generate a filtered output as a refinement signal. This output is directly added to the Cb component and appropriately weighted before being added to the Cr component. The filter is indicated at the CTU level or by the block size, which is communicated with each stripe signaling.
[0277] The supported range of chroma block sizes extends from the minimum chroma CTU size to the current chroma CTU size. The minimum chroma CTU size is the minimum of the minimum possible width and height of the chroma CTU, i.e., Min(32 / SubWidthC, 32 / SubHeightC), while the current chroma CTU size is the minimum of the width and height of the current chroma CTU, i.e., Min(CtbWidthC, CtbHeightC). For example, if the CTU size is set to a maximum of 128×128, then for 4:4:4 video, the JC-CCALF chroma block size for the strip will be one of 32×32, 64×64, and 128×128, or for 4:2:0 and 4:2:2 video, it will be one of 16×16, 32×32, and 64×64.
[0278] Figure 18 An example of the JC-CCALF workflow is shown.
[0279] 3. Example technical problems solved by the technical solutions disclosed in this paper.
[0280] The current design of CC-ALF has the following problems:
[0281] 1. In the current CC-ALF, the offset is calculated for each chromaticity sample point, which greatly increases the computational complexity.
[0282] 2. The current filtering design in CC-ALF may not be effective because it uses only one filter to support all chroma samples of all types of video, which does not take into account local characteristics.
[0283] 3. In the current filter design of CC-ALF, luminance samples are always used to refine the chrominance components, which may be suboptimal.
[0284] 4. In the current design of CC-ALF or ALF, using luminance and / or chrominance samples in the current frame to refine luminance and / or chrominance samples may be inefficient.
[0285] 5. CC-ALF is not necessarily applicable to chroma samples at specific locations, as these samples may be filtered multiple times, for example in deblocking filters, SAO, and chroma ALF.
[0286] 4. Examples of solutions and implementation methods
[0287] The following list should be considered as examples for explaining general concepts. These items should not be interpreted in a narrow way. Furthermore, these items can be combined in any way.
[0288] In this document, the term "CC-ALF" refers to an encoding / decoding tool that uses sample values from a second color component (e.g., Y) or multiple color components (e.g., both Y and Cr) to refine samples in a first color component (e.g., Cb). It is not limited to the CC-ALF technique. In one example, the juxtaposed luminance sample of a chroma sample located at (x, y) is defined as follows: in 4:2:0 chroma format, 4:2:2 chroma format, and 4:4:4 chroma format, the juxtaposed luminance sample is located at (2x, 2y), (2x, y), and (x, y), respectively. A "corresponding filter sample set" can be used to represent those samples included in the filter support; for example, for CC-ALF, a "corresponding filter sample set" can be used to represent the juxtaposed luminance sample of a chroma sample and its adjacent luminance samples, which are used to derive the refinement / offset of the chroma sample.
[0289] Sub-block-based filtering methods
[0290] 1. Instead of performing the CC-ALF filtering process at the sample level, it proposes to apply it at the sub-block level (containing more than 1 sample), and an offset can be shared by all samples of the color components in the sub-block of CC-ALF / chroma ALF / luminance ALF / other kinds of filtering methods.
[0291] a. In one example, a sub-block can be an M×N (M columns by n rows) array of samples (e.g., 1×2, or 2×1, or 2×2, or 2×4, or 4×2, or 4×4).
[0292] b. In one example, the sub-block dimension can be signaled at the sequence level / picture level / strip level / piece group level, such as in the sequence header / picture header / SPS / VPS / DPS / PPS / APS / strip header / piece group header; or at the video region level, such as CTB.
[0293] c. In one example, the offset can be further modified before refining all samples in the sub-block. The offset is represented as o.
[0294] i. In one example, the offset can be trimmed to a given range before each sample point is added to the sub-block.
[0295] 1) In one example, the offset can be trimmed to the range of T1 to T2, for example, T1 = -5 and T2 = 5.
[0296] 2) In one example, signaling can be used to notify whether and / or how to perform pruning.
[0297] ii. In one example, the offset can be set to T when a specific condition is met.
[0298] 1) In one example, if o < T1 && o > T2, for example T = 0, T1 = -5, T2 = 5, then the offset can be set to T.
[0299] iii. In one example, how the offset is modified might depend on the sub-block width and / or sub-block height. Let the sub-block width be represented as W, and the sub-block height as H.
[0300] 1) In one example, the offset can be modified to o / (M×(W×H)), for example, M=1 or M=2.
[0301] iv. In one example, the offset can be multiplied by T or divided by T, added to T or shifted by T before each sample point is added to the sub-block.
[0302] 1) In one example, signaling can be used to notify T in the bit stream.
[0303] 2) In one example, T can depend on the sample values in the sub-block.
[0304] 3) In one example, T can depend on the bit depth of the sample value.
[0305] a) In one example, the offset can be modified to o*T, where T = v / (2^B) max -1), and v represents the value of a sample point, B max Indicates the maximum bit depth of the component.
[0306] v. Whether to modify the offset can be signaled at the sequence level / picture level / strip level / piece group level, such as in the sequence header / picture header / SPS / VPS / DPS / PPS / APS / strip header / piece group header; or at the video region level, such as CTB.
[0307] d. In one example, a representative sample point from a sub-block can be selected, and the corresponding CC-ALF sample point set in that sub-block can be used to calculate the offset shared by that sub-block.
[0308] i. In one example, the determination of representative chromaticity samples may depend on the location of the samples.
[0309] 1) In one example, depending on the scan order (e.g., raster scan), the representative sample point can be the Nth chromaticity sample point in the sub-block. For example, N can be 1 or 2. In another example, depending on the scan order, N can be the last sample point in the sub-block.
[0310] 2) In one example, representative samples can be located at specific positions within a sub-block, such as the center, top left, top right, bottom left, or bottom right.
[0311] ii. In one example, the determination of representative samples may depend on the sample values.
[0312] 1) In one example, a representative sample point can be the sample point with the largest sample point value in the sub-block.
[0313] 2) In one example, a representative sample point can be the sample point with the smallest sample point value in the sub-block.
[0314] 3) In one example, a representative sample point can be a sample point of the median sample point value in a sub-block.
[0315] iii. In one example, a “corresponding filtered sample set” of a representative sample from CC-ALF (e.g., any of Figures 13 to 17) can be used.
[0316] e. In one example, multiple representative samples from a sub-block can be selected, and the offset can be calculated using their corresponding filtered sample sets.
[0317] i. In one example, the “corresponding filtered sample set” for each of the multiple representative samples can first be modified (e.g., by averaging or weighted averaging) to generate a virtual representative filtered sample set, and this virtual representative filtered sample set is used to calculate the offset.
[0318] 2. A dual-offset thinning method is proposed. In order to filter the sample points, two or more offsets can be added to the sample points to obtain the final thinned sample point values.
[0319] a. In one example, two offsets are used to refine a current sample point, where the first offset is a shared offset for the sub-blocks that include the current sample point, and the second offset is a separate offset for the current sample point.
[0320] i. In one example, the shared offset can be calculated using the first filter support, for example, using the method described in bullet point 1.
[0321] 1) In one example, Figure 19 The filter shape defined in [the code] can be used in CC-ALF, where the sub-block size is 1×2. For example, a sub-block can include elements located at (X... c , Yc ) and (X c , Y c +1) chromaticity samples. Luminance samples, denoted by L0, L1, ..., L5, can be used to calculate the shared offset.
[0322] 2) In one example, it can be used Figure 20 The filter shape is defined in [the code], where the sub-block size is 2×1. For example, a sub-block can include elements located at (X... c , Y c ) and (X c +1, Y c The chromaticity samples are represented by L0, L1, ..., L3. The luminance samples can be used to calculate the shared offset.
[0323] ii. In one example, a second filter can be used to calculate the individual offset for each chromaticity sample.
[0324] 1) In one example, Figure 19 The filter shape defined in [the code] can be used in CC-ALF, where the sub-block size is 1×2. For example, the luminance samples represented by L6 and L1 can be used to calculate the luminance located at (X... c , Y c The individual offset of the first chromaticity sample at (X) and the luminance samples represented by L4 and L7 can be used to calculate the luminance sample located at (X). c , Y c Individual offset of the second chromaticity sample at +1).
[0325] 2) In one example, it can be used Figure 20 The filter shape is defined in [the code], where the sub-block size is 2×1. For example, the brightness samples represented by L4, L1, L5, and L2 can be used to calculate the brightness of samples located in (X...). c , Y c The individual offset of the first chromaticity sample at (X) and the luminance samples represented by L6, L7, L8, and L9 can be used to calculate the luminance sample located at (X). c +1, Y c The individual offset of the second chromaticity sample at ().
[0326] iii. In one example, samples involved in the shared offset calculation (e.g., samples included in the first filter support) are excluded from the second filter support.
[0327] 1) Optionally, some samples involved in the first filter support are included in the second filter support.
[0328] iv. In one example, the number of samples involved in the shared offset calculation (e.g., the number of samples included in the first filter support) is equal to the number of samples involved in the second filter support.
[0329] 1) Optionally, the number of samples involved in the shared offset calculation process is not less than the number of samples involved in the second filter support.
[0330] 2) Optionally, the number of samples involved in the shared offset calculation process is no greater than the number of samples involved in the second filter support.
[0331] b. In one example, one or more shared offsets and / or individual offsets may be signaled at the sequence level / picture level / strip level / piece group level, such as in the sequence header / picture header / SPS / VPS / DPS / PPS / APS / strip header / piece group header.
[0332] i. Shared offsets can be signaled in the scan order of sub-blocks.
[0333] c. In one example, the above method can be applied to CC-ALF / chroma ALF / other filtering methods.
[0334] CC-ALF supports multiple filter shapes
[0335] 3. N-tap symmetric filter support used in CC-ALF, wherein at least two filter coefficients of two samples within the filter support share the same value.
[0336] a. In one example, N can be a specific number, such as N = 6 or 8.
[0337] i. In one example, a symmetric N-tap filter shape can be used in CC-ALF.
[0338] 1) In one example, it can be used in CC-ALF. Figures 21A-21C The shape of a symmetric 8-tap filter is defined in any of them, with 4 unique coefficients (denoted by Ci, where i is from 0 to 3).
[0339] 2) In one example, it can be used in CC-ALF. Figures 22A-22C The shape of a symmetric 8-tap filter with 5 unique coefficients is defined in any of them.
[0340] 3) In one example, it can be used in CC-ALF. Figures 23A-23C The shape of a symmetric 8-tap filter with 6 unique coefficients is defined in any of them.
[0341] b. In one example, the same value shared by two or more coefficients can be signaled at the sequence level / picture level / strip level / piece group level, for example in the sequence header / picture header / SPS / VPS / DPS / PPS / APS / strip header / piece group header, and used to derive two or more coefficients.
[0342] 4. In CC-ALF, an asymmetric N-tap filter shape (N < 8) can be used.
[0343] a. In one example, it can be used in CC-ALF. Figures 24A-24F The shape of the asymmetric 6-tap filter defined in any of them.
[0344] b. In one example, it can be used in CC-ALF. Figures 25A-25D The shape of the asymmetric 5-tap filter defined in any of them.
[0345] c. In one example, it can be used in CC-ALF. Figures 26A-26E The shape of the asymmetric 4-tap filter defined in any of them.
[0346] 5. CC-ALF is invoked to apply the filter coefficients to the sample difference, rather than directly to the samples in the filter support.
[0347] a. In one example, the luminance sample difference can be defined as the difference between a representative luminance sample in the filter support (region) and other luminance samples in the filter support region.
[0348] i. In one example, for a chromaticity sample located at (X, Y), the representative luminance sample is located at (2X, 2Y).
[0349] ii. In one example, Figure 27A Alternatively, the filter shapes defined in 27B can be used in CC-ALF.
[0350] b. In one example, identify multiple representative samples and use the difference between a non-representative sample and a representative sample.
[0351] i. In one example, for a chromaticity sample point located at (X, Y), two representative luminance samples are located at (2X, 2Y) and (2X+1, 2Y).
[0352] ii. In one example, for a chromaticity sample located at (X, Y), two representative luminance samples are located at (2X, 2Y) and (2X, 2Y+1).
[0353] iii. In one example, Figure 28The filter shapes defined in the code can be used in CC-ALF.
[0354] c. The selection of representative samples can be predefined or depend on decoding information (e.g., color format) or be notified by signaling.
[0355] 6. For video units (e.g., strips / pictures), a set of multiple filter supports can be used in CC-ALF.
[0356] a. In one example, the filter taps denoted by N (e.g., those used in bullet points 3-5) and / or the indications of filter support can be signaled at the sequence level / picture level / strip level / piece group level, for example in the sequence header / picture header / SPS / VPS / DPS / PPS / APS / strip header / piece group header.
[0357] b. In one example, the “corresponding filter sample set” used to calculate the offset in a sub-block may be included in some or all of the luminance samples used in a non-sub-block-based filtering method for multiple representative chroma samples.
[0358] i. In one example, Figure 29 The filter shape defined in the code can be used for CC-ALF based on 1×2 sub-blocks.
[0359] ii. In one example, Figure 30 The filter shape defined in the code can be used for CC-ALF based on 2×1 sub-blocks.
[0360] iii. In one example, Figure 31 The 8-tap filter shape defined in the code can be used in a 1×2 subblock CC-ALF.
[0361] iv. In one example, Figure 32 The 8-tap filter shape defined in the code can be used in 2×1 sub-block CC-ALF.
[0362] v. In one example, Figure 33 The 10-tap filter shape defined in the code can be used in a 1×2 sub-block CC-ALF.
[0363] vi. In one example, the 10-tap filter shape defined in Figure 34 can be used in a 2×1 sub-block CC-ALF.
[0364] Filters with multiple components support
[0365] 7. When filtering (e.g., in CC-ALF) samples, the filter support may include samples associated with multiple components.
[0366] a. In one example, the filter coefficients applied to the second color component and the filter coefficients applied to the third color component to correct the first color component can be signaled sequentially (the coefficients applied to the third color component are signaled after the coefficients applied to the second color component).
[0367] a) In another example, the filter coefficients applied to the second color component and the filter coefficients applied to the third color component to correct the first color component can be signaled in an interleaved manner.
[0368] b. It is recommended to correct the samples in the first component by filtering the samples in the second component, which eliminates the case where the first component is Cb / Cr and the second component is Y.
[0369] i. In one example, the filter shape defined in bullet point 3 can be used in the second component to correct the samples in the first component.
[0370] c. It is recommended that when exporting the "sample correction" or "refined sample" of the first component in CC-ALF, samples from more than one component can be filtered.
[0371] i. In one example, a filter shape (e.g., the filter shape in bullet point 3) can be used in all multiple components to correct the samples in the first component.
[0372] 1) In one example, the offsets derived from multiple components can be averaged or weighted to correct the samples in the first component.
[0373] 2) In one example, the product of offsets derived from multiple components can be used to correct samples in the first component.
[0374] ii. In one example, different filter shapes (e.g., the filter shapes defined in bullet point 3) can be used in multiple components to correct samples in the first component.
[0375] 1) In one example, the offsets derived from multiple components can be averaged or weighted to correct the samples in the first component.
[0376] 2) In one example, the product of offsets derived from multiple components can be used to correct samples in the first component.
[0377] Multiple frames supported by ALF or / and CC-ALF
[0378] 8. In the current design of ALF or / and CCALF, samples in the current frame are used to refine samples in the same or different components. It is recommended to use samples from multiple frames to refine samples in the current frame.
[0379] a. In one example, whether samples from multiple frames are used for ALF or / and CC-ALF can be signaled at the sequence level / picture level / strip level / piece group level, for example in the sequence header / picture header / SPS / VPS / DPS / PPS / APS / strip header / piece group header.
[0380] b. In one example, multiple frames used for ALF or / and CC-ALF may or may not be reference images of the current frame.
[0381] i. In one example, the nearest picture of the current frame can be used in ALF or / and CC-ALF, where the code-decode picture (POC) can be used to determine the distance.
[0382] ii. In one example, the short-lived image of the current frame can be used for ALF and / or CC-ALF.
[0383] iii. In one example, the long-term picture of the current frame can be used for ALF and / or CC-ALF.
[0384] iv. In one example, a reference image in the same temporal layer of the current frame can be used in ALF or / and CC-ALF.
[0385] v. In one example, reference images from different temporal layers of the current frame can be used in ALF or / and CC-ALF.
[0386] vi. In one example, multiple frames used for ALF or / and CC-ALF may include or exclude the current frame.
[0387] Disable CC-ALF for chroma samples between the virtual boundary and the bottom boundary of the CTU.
[0388] 9. It is recommended to disable CC-ALF for chroma samples between the ALF virtual boundary and the CTB bottom boundary.
[0389] a. In one example, the ALF virtual boundary can be defined in Section 2.14.6.
[0390] 10. It is recommended to disable ALF for luminance samples and / or chrominance samples between the ALF virtual boundary and the bottom boundary of CTU.
[0391] a. In one example, the ALF virtual boundary can be defined in Section 2.14.6.
[0392] 11. It is recommended to disable the filtering process (e.g., CC-ALF) for samples that have already been filtered by another filter (e.g., deblocking filter).
[0393] a. Optionally, the filtering process (e.g., CC-ALF) can be disabled for samples located at transform / CU edges that can be filtered by another filter (e.g., a deblocking filter).
[0394] General requirements
[0395] 12. Whether and / or how the methods disclosed above can be signaled at the sequence level / picture level / strip level / piece group level, for example in the sequence header / picture header / SPS / VPS / DPS / PPS / APS / strip header / piece group header.
[0396] 13. Whether and / or how the methods disclosed above are applied may depend on information such as the encoding / decoding format, single / dual tree segmentation, and sample location (e.g., relative to CU / CTU).
[0397] Figure 19 A CC-ALF filter based on partially shared sub-blocks is shown.
[0398] Figure 20 A CC-ALF filter based on partially shared 2×1 sub-blocks is shown.
[0399] Figures 21A-21C The filter (X) is shown c , Y c Example of a symmetric 8-tap filter with 4 unique coefficients. Figure 21A A type 1 filter is shown. Figure 21B A type 2 filter is shown, and Figure 21C A type 3 filter is shown.
[0400] Figures 22A-22C An example of a symmetric 8-tap filter with 5 unique coefficients is shown when filtering (Xc, Yc). Figure 22A , 22B 22C shows type 1, type 2 and type 3 filters, respectively.
[0401] Figures 23A-23C An example of a symmetric 8-tap filter with 6 unique coefficients is shown when filtering (Xc, Yc). Figure 23A , 23B 23C shows type 1, type 2 and type 3 filters, respectively.
[0402] Figures 24A-24F The filter (X) is shown c , Y c Asymmetric 6-tap filter when ). Figures 24A to 24F Filters of type 1, type 2, type 3, type 4, type 5 and type 6 are shown respectively.
[0403] Figures 25A-25D The filter (X) is shown c , Y c Asymmetric 5-tap filter (X) c , Y c ). Figures 25A to 25D Type 1, Type 2, Type 3, and Type 4 filters are shown respectively.
[0404] Figures 26A-26D The filter (X) is shown c , Y c Asymmetric 4-tap filter when ). Figures 26A to 26E Type 1, Type 2, Type 3, Type 4, and Type 5 filters are shown respectively.
[0405] Figures 27A-27B An example of filter coefficients applied to luminance differences in CC-ALF is shown. Figure 27A , 27B Type 1 and Type 2 filters are shown respectively.
[0406] Figure 28 An example of filter coefficients applied to luminance differences (different center values) in CC-ALF is shown.
[0407] Figure 29 An example of a 14-tap filter for CC-ALF based on 1×2 sub-blocks is shown.
[0408] Figure 30 An example of a 14-tap filter for CC-ALF based on 2×1 sub-blocks is shown.
[0409] Figure 31 An example of an 8-tap filter for CC-ALF based on 1×2 sub-blocks is shown.
[0410] Figure 32 Example of an 8-tap filter for CC-ALF based on 2×1 sub-blocks.
[0411] Figure 33 An example of a 10-tap filter for CC-ALF based on 1×2 sub-blocks is shown.
[0412] Figures 34A-34B The 10-tap filters of type 1 and type 2 CC-ALF based on 2×1 sub-blocks are shown.
[0413] Figure 35This is a block diagram of an example video processing system 1900 that can implement the various techniques disclosed herein. Various implementations may include some or all of the components in system 1900. System 1900 may include an input 1902 for receiving video content. The video content may be received in a raw or uncompressed format (e.g., 8 or 10-bit multi-component pixel values), or in a compressed or encoded format. Input 1902 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces (such as Ethernet, Passive Optical Networking (PON), etc.) and wireless interfaces (such as Wi-Fi or cellular interfaces).
[0414] System 1900 may include a codec component 1904 capable of implementing the various codec or encoding methods described in this document. Codec component 1904 can reduce the average bit rate of the video from input 1902 to the output of codec component 1904 to produce a bitstream representation of the video. Therefore, codec techniques are sometimes referred to as video compression or video transcoding techniques. The output of codec component 1904 can be stored or transmitted via connected communication, as represented by component 1906. The stored or communicated bitstream (or codec) representation of the video received at input 1902 can be used by component 1908 to generate pixel values or displayable video that is sent to display interface 1910. The process of generating user-visible video from the bitstream representation is sometimes referred to as video decompression. Furthermore, although some video processing operations are referred to as "codec" operations or tools, it should be understood that codec tools or operations used at the encoder will be followed by corresponding decoding tools or operations that inversely reproduce the codec results by the decoder.
[0415] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB), High Definition Multimedia Interface (HDMI), or DisplayPort. Examples of storage interfaces include SATA (Serial Advanced Technology Accessory), PCI, IDE, etc. The technologies described in this document can be implemented in a variety of electronic devices, such as mobile phones, laptops, smartphones, or other devices capable of performing digital data processing and / or video display.
[0416] Figure 36This is a block diagram of a video processing apparatus 3600. Apparatus 3600 can be used to implement one or more of the methods described herein. Apparatus 3600 can be implemented in smartphones, tablet computers, Internet of Things (IoT) receivers, etc. Apparatus 3600 may include one or more processors 3602, one or more memories 3604, and video processing hardware 3606. The processors(multiple) 3602 can be configured to implement one or more methods described herein. The memories(multiple) 3604 can be used to store data and code used to implement the methods and techniques described herein. The video processing hardware 3606 can be used to implement some of the techniques described herein in hardware circuitry.
[0417] The following provides a list of preferred solutions for some embodiments.
[0418] The following solutions illustrate example embodiments of the techniques discussed in the previous section (e.g., Project 1).
[0419] 1. A method for video processing (e.g., Figure 37 The method 3700 includes: for the conversion between a video unit and a bitstream representation of a video component, determining (3702) to use an offset value for cross-component adaptive loop filtering of all samples of a sub-block of the video unit; and performing (3704) a conversion based on the determination, wherein the offset is also used for another processing operation in the conversion, the other processing operation including one or more adaptive loop filtering operations.
[0420] 2. The method according to Solution 1, wherein the video component is the color component of the video.
[0421] 3. The method according to any one of solutions 1-2, wherein the video unit comprises a video block.
[0422] 4. The method according to any one of solutions 1-3, wherein the sub-block size corresponds to an MxN array of sample points.
[0423] 5. The method according to Solution 4, wherein the sub-block size is signaled in the bitstream representation.
[0424] The following solutions illustrate example implementations of the techniques discussed in the previous section (e.g., Project 2).
[0425] 6. A video processing method comprising: for a conversion between a video unit of a video component and a bitstream representation of the video, determining a final offset value of a current sample of a sub-block of the video unit for cross-component adaptive loop filtering during the conversion; and performing the conversion based on the determination; wherein the final offset value is based on a first offset value and a second offset value different from the first offset value.
[0426] 7. The method according to Solution 6, wherein the first offset value is shared with adjacent samples, and the second offset value is specific to the current sample.
[0427] 8. The method according to any one of solutions 6-7, wherein the size of the sub-block is 2×1.
[0428] 9. The method according to any one of solutions 6-8, wherein the sample points used to determine the first offset value and the second offset value are different from each other.
[0429] 10. The method according to any one of solutions 6-8, wherein at least some sample points are used to determine the first offset value and the second offset value.
[0430] 11. The method according to any one of solutions 6-10, wherein a field in the bitstream representation indicates a first offset value for conversion.
[0431] The following solutions illustrate example implementations of the techniques discussed in the previous section (e.g., Project 3).
[0432] 12. A video processing method comprising: for a conversion between video units of video components and a bitstream representation of video, determining that an N-tap symmetric filter will be used for cross-component adaptive loop filter calculation during the conversion; and performing the conversion based on the determination; wherein, with the support of the N-tap symmetric filter, at least two filter coefficients of two samples share the same value.
[0433] 13. The method according to solution 12, wherein N = 6 or 8.
[0434] 14. The method according to any one of solutions 12-13, wherein the field signaling notification in the bitstream representation is the same value shared by the two samples.
[0435] The following solutions illustrate example implementations of the techniques discussed in the previous section (e.g., Project 4).
[0436] 15. A video processing method comprising: for a conversion between video units of video components and a bitstream representation of video, determining that an N-tap asymmetric filter will be used for cross-component adaptive loop filter calculation during the conversion, wherein N is a positive integer; and performing the conversion based on the determination.
[0437] 16. The method according to solution 15, wherein N < 8.
[0438] The following solutions illustrate example implementations of the techniques discussed in the previous section (e.g., Project 5).
[0439] 17. A video processing method comprising: performing a conversion between a video unit of a first component of the video and a bitstream representation of the video; wherein the conversion of samples of the first component includes applying a cross-component adaptive loop filter to the sample difference of a second component of the video.
[0440] 18. The method according to solution 17, wherein the first component is a color component and the second component is a luminance component.
[0441] 19. The method according to solution 17 or 18, wherein the sample difference of the second component is determined by differentiating a representative luminance sample in the luminance filter support region with another luminance sample in the luminance filter support region.
[0442] 20. The method according to solution 19, wherein the sample point is located at position (X, Y) and the representative luminance sample point is selected from position (2X, 2Y), where X and Y are integer offsets from the top left corner of the video unit located at (0, 0) to the sample point position.
[0443] 21. The method according to solution 19, wherein the sample point is located at position (X, Y), and a representative luminance sample point is selected from position (2X, 2Y), and another luminance sample point is located at (2X+1, 2Y), where X and Y are integer offsets from the top left corner of the video unit located at (0, 0) to the sample point position.
[0444] The following solutions illustrate example implementations of the techniques discussed in the previous section (e.g., Project 6).
[0445] 22. A video processing method comprising: during a conversion between a sub-block of a video unit of a video component and a bitstream representation of the video, determining to use two or more filters from a plurality of filter sets for cross-component adaptive loop filtering; and performing the conversion based on the determination.
[0446] 23. The method according to solution 22, wherein the fields in the bitstream representation indicate the length N of the set of multiple filters and / or the support of two or more filters.
[0447] 24. The method according to any one of solutions 22-23, wherein the offset used in cross-component loop filtering is determined by using luminance samples corresponding to the luminance components using predefined support sub-blocks.
[0448] The following solutions illustrate example implementations of the techniques discussed in the previous section (e.g., items 7 and 8).
[0449] 25. A video processing method comprising: a conversion between a sub-block of a first component of a video unit and a bitstream representation of the video; determining filters that support multiple components or multiple pictures across the video for performing cross-component adaptive loop filtering during the conversion; and performing the conversion based on the determination.
[0450] 26. The method according to solution 25, wherein one or more syntax elements in the bitstream representation indicate the use of filters and / or multiple components for supporting filters.
[0451] 27. The method according to solution 26, wherein one or more syntax elements are sequential, wherein filter coefficients of the second component are indicated, followed by filter coefficients of the third component.
[0452] 28. The method according to solution 26, wherein one or more syntax elements are in an interleaved order, wherein the filter coefficients of the second component are indicated, followed by the filter coefficients of the third component.
[0453] 29. The method according to solution 25, wherein one or more syntax elements in the bitstream representation indicate the use of multiple images and / or identify multiple images.
[0454] 30. The method described in solution 25, wherein multiple images are not included in the reference images used during conversion.
[0455] The following solutions illustrate example implementations of the techniques discussed in the previous section (e.g., items 9, 10, and 11).
[0456] 31. A video processing method comprising: for a conversion between a sub-block of a video unit of a video component and a bitstream representation of the video, determining, according to a position rule, whether to enable a cross-component adaptive loop filter (CC-ALF) for the conversion; and performing the conversion based on the determination.
[0457] 32. The method according to solution 31, wherein the location rule specifies the CC-ALF between the bottom boundary of the codec tree block and the virtual boundary of the loop filter.
[0458] 33. The method according to solution 31, wherein the position rule specifies the CC-ALF between the bottom boundary of the codec tree unit and the virtual boundary of the loop filter.
[0459] 34. The method according to solution 31, wherein the location rule specifies that CC-ALF is disabled at the location where another filter is applied during the conversion.
[0460] 35. The method according to any one of solutions 1-34, wherein the video unit comprises a video block or video strip or video picture.
[0461] 36. The method according to any one of solutions 1-34, wherein the syntax elements indicating the method used for conversion are included at the sequence level, picture level, strip level, or piece level in the sequence header, picture header, sequence parameter set, video parameter set, picture parameter set, adaptive parameter set, strip header, or piece group header.
[0462] 37. The method according to any one of solutions 1-36, wherein the method is selectively applied due to the encoding and decoding information of the video.
[0463] 38. The method according to solution 37, wherein the features include the color format or segmentation type or the position of the sub-block relative to the codec unit or codec tree unit.
[0464] 39. The method according to any one of solutions 1-38, wherein performing the conversion includes encoding the video to generate a bitstream representation.
[0465] 40. The method according to any one of solutions 1-38, wherein performing the conversion includes parsing and decoding the bitstream representation to generate video.
[0466] 41. A video decoding apparatus, comprising a processor configured to implement the method described in one or more of solutions 1 to 40.
[0467] 42. A video encoding apparatus comprising a processor configured to implement the method described in one or more of solutions 1 to 40.
[0468] 43. A computer program product having computer code stored thereon, which, when executed by a processor, causes the processor to implement the method of any one of solutions 1 to 40.
[0469] 44. The methods, apparatus or systems described in this document.
[0470] In this document, the term "video processing" can refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm can be applied during the conversion from the pixel representation of a video to the corresponding bitstream representation, and vice versa. As defined in the syntax, the bitstream representation of the current video block can, for example, correspond to bits juxtaposed or scattered at different positions within the bitstream. For example, a macroblock can be encoded based on the error residual values after transformation and encoding / decoding, and also using bits from the header and other fields in the bitstream.
[0471] Figure 38This is a block diagram illustrating an example of a video decoder 300, which can be... Figure 39 The video decoder 114 in the system 100 described herein.
[0472] The video decoder 300 can be configured to perform any or all of the techniques disclosed herein. Figure 38 In the example, the video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video decoder 300. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.
[0473] exist Figure 38 In the example, the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra-frame prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, the video decoder 300 can perform operations related to the video encoder 200 ( Figure 40 The encoding process described is generally the opposite of the decoding process.
[0474] The entropy decoding unit 301 can extract the encoded bitstream. The encoded bitstream may include entropy-encoded video data (e.g., encoded blocks of video data). The entropy decoding unit 301 can decode the entropy-encoded video data, and based on the entropy-encoded video data, the motion compensation unit 302 can determine motion information including motion vectors, motion vector precision, reference image list indexes, and other motion information. The motion compensation unit 302 can determine such information, for example, by executing AMVP and Merge modes.
[0475] The motion compensation unit 302 can generate motion compensation blocks, possibly performing interpolation based on an interpolation filter. The identifier of the interpolation filter to be used at sub-pixel precision can be included in the syntax element.
[0476] The motion compensation unit 302 can use the interpolation filter used by the video encoder 20 during the encoding of a video block to calculate the interpolation values of a sub-integer number of pixels of the reference block. The motion compensation unit 302 can determine the interpolation filter used by the video encoder 200 based on the received syntax information and use the interpolation filter to generate the prediction block.
[0477] The motion compensation unit 302 can use some syntactic information to determine: the size of the blocks used to encode (multiple) frames and / or (multiple) stripes of the encoded video sequence, segmentation information describing how each macroblock of the image of the encoded video sequence is segmented, a mode indicating how each segment is encoded, one or more reference frames (and a list of reference frames) for each inter-frame coded block, and other information for decoding the encoded video sequence.
[0478] Intra-prediction unit 303 can use, for example, an intra-prediction mode received in the bitstream to form prediction blocks from spatially adjacent blocks. Inverse quantization unit 303 inverse quantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by entropy decoding unit 301. Inverse transform unit 303 applies an inverse transform.
[0479] The reconstruction unit 306 can sum the residual blocks using the corresponding prediction blocks generated by the motion compensation unit 202 or the intra-frame prediction unit 303 to form a decoded block. As desired, a deblocking filter can also be applied to filter the decoded block to remove blockage artifacts. The decoded video block is then stored in a buffer 307, which provides a reference block for subsequent motion compensation / intra-frame prediction. The decoded video is also generated for presentation on a display device.
[0480] In this document, the term "video processing" can refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm can be applied during the conversion from a pixel representation of a video to a corresponding bitstream representation, and vice versa. As defined in the syntax, the bitstream representation of the current video block, or bitstream representation, can, for example, correspond to bits juxtaposed or scattered at different positions within the bitstream. For example, a video block can be encoded based on the error residual values after transformation and encoding / decoding, and also using bits from the header and other fields in the bitstream. Furthermore, during the conversion, the decoder can parse the bitstream based on this determination, knowing that certain fields may or may not be present, as described in the solutions above. Similarly, the encoder can determine whether to include or exclude certain syntax fields and generate a bitstream representation accordingly by including or excluding syntax fields from the bitstream representation.
[0481] Figure 39 This is a block diagram illustrating an example video codec system 100 that can utilize the techniques disclosed herein. Figure 39 As shown, the video encoding / decoding system 100 may include a source device 110 and a destination device 120. The source device 110 generates encoded video data and may be referred to as a video encoding device. The destination device 120 can decode the encoded video data generated by the source device 110 and may be referred to as a video decoding device. The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.
[0482] Video source 112 may include sources such as video capture devices, interfaces for receiving video data from video content providers, and / or computer graphics systems that generate video data, or combinations of these sources. Video data may include one or more images. Video encoder 114 encodes the video data from video source 112 to generate a bitstream. The bitstream may include a sequence of bits forming a bitstream representation of the video data. The bitstream may include encoded / decoded images and associated data. The encoded / decoded image is a bitstream representation of the image. The associated data may include sequence parameter sets, image parameter sets, and other syntax elements. I / O interface 116 may include a modulator / demodulator (modem) and / or a transmitter. Encoded video data may be transmitted directly to destination device 120 via network 130a through I / O interface 116. Encoded video data may also be stored on storage medium / server 130b for access by destination device 120.
[0483] Destination device 120 may include I / O interface 126, video decoder 124 and display device 122.
[0484] I / O interface 126 may include a receiver and / or a modem. I / O interface 126 may acquire encoded video data from source device 110 or storage medium / server 130b. Video decoder 124 may decode the encoded video data. Display device 122 may display the decoded video data to a user. Display device 122 may be integrated with destination device 120 or may be external to destination device 120 configured to connect to an external display device.
[0485] The video encoder 114 and the video decoder 124 can operate according to video compression standards such as High Efficiency Video Coding (HEVC), Universal Video Coding (VVC), and other current and / or other standards.
[0486] Figure 40 This is a block diagram illustrating an example of a video encoder 200, which can be... Figure 39 The video encoder 114 in the system 100 described herein.
[0487] The video encoder 200 can be configured to perform any or all of the techniques disclosed herein. Figure 40 In the example, the video encoder 200 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video encoder 200. In some examples, the processor can be configured to perform any or all of the techniques of this disclosure.
[0488] The functional components of the video encoder 200 may include a segmentation unit 201, a prediction unit 202 (which may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra-frame prediction unit 206), a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213, and an entropy coding unit 214.
[0489] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an intra-block copy (IBC) unit. The IBC unit can perform prediction in an IBC mode, where at least one reference picture is the picture in which the current video block is located.
[0490] Furthermore, some components, such as the motion estimation unit 204 and the motion compensation unit 205, can be highly integrated, but for interpretive purposes... Figure 40 The examples are shown separately.
[0491] The segmentation unit 201 can segment an image into one or more video blocks. The video encoder 200 and the video decoder 300 can support various video block sizes.
[0492] The mode selection unit 203 can, for example, select one of the intra-frame or inter-frame coding / decoding modes based on the error result, and provide the obtained intra-frame or inter-frame codec block to the residual generation unit 207 to generate residual block data and to the reconstruction unit 212 to reconstruct the coded block for use as a reference picture. In some examples, the mode selection unit 203 can select a combination of intra-frame and inter-frame prediction (CIIP) modes, where prediction is based on inter-frame prediction signals and intra-frame prediction signals. The mode selection unit 203 can also select the resolution of the motion vector (e.g., sub-pixel or integer pixel precision) for the block in the inter-frame prediction case.
[0493] To perform inter-frame prediction on the current video block, motion estimation unit 204 can generate motion information for the current video block by comparing one or more reference frames from buffer 213 with the current video block. Motion compensation unit 205 can determine the predicted video block for the current video block based on the motion information of the image from buffer 213 (rather than the image associated with the current video block) and decoded samples.
[0494] The motion estimation unit 204 and the motion compensation unit 205 can perform different operations for the current video block. For example, the different operations performed depend on whether the current video block is in an I-strip, a P-strip, or a B-strip.
[0495] In some examples, motion estimation unit 204 can perform unidirectional prediction of the current video block, and can search for a reference video block for the current video block in the reference images of list 0 or list 1. Motion estimation unit 204 can then generate a reference index indicating that the reference video block is present in the reference images of list 0 or list 1, and a motion vector indicating the spatial displacement between the current video block and the reference video block. Motion estimation unit 204 can output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current block based on the reference video block indicated by the motion information of the current video block.
[0496] In other examples, motion estimation unit 204 can perform bidirectional prediction of the current video block. Motion estimation unit 204 can search for a reference video block for the current video block in the reference images of list 0 and can also search for another reference video block for the current video block in the reference images of list 1. Motion estimation unit 204 can then generate a reference index indicating that the reference images in list 0 or list 1 contain the reference video block, and a motion vector indicating the spatial displacement between the reference video block and the current video block. Motion estimation unit 204 can output the reference index and the motion vector of the current video block as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current video block based on the reference video block indicated by the motion information of the current video block.
[0497] In some examples, the motion estimation unit 204 can output the complete set of motion information for the decoder's decoding process.
[0498] In some examples, motion estimation unit 204 may not output the complete set of motion information for the current video. Instead, motion estimation unit 204 may signal the motion information of the current video block by referencing the motion information of another video block. For example, motion estimation unit 204 may determine that the motion information of the current video block is sufficiently similar to the motion information of adjacent video blocks.
[0499] In one example, the motion estimation unit 204 may indicate in the syntax structure associated with the current video block that the current video block has the same motion information value as another video block.
[0500] In another example, motion estimation unit 204 can identify another video block and motion vector difference (MVD) in the syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicating video block. Video decoder 300 can use the motion vector of the indicating video block and the motion vector difference to determine the motion vector of the current video block.
[0501] As discussed above, the video encoder 200 can predictively signal motion vectors. Two examples of predictive signaling notification techniques that can be implemented by the video encoder 200 include Advanced Motion Vector Prediction (AMVP) and Merge Pattern Signaling Notification.
[0502] Intra-prediction unit 206 can perform intra-prediction on the current video block. When intra-prediction unit 206 performs intra-prediction on the current video block, it can generate prediction data for the current video block based on decoded samples from other video blocks in the same frame. The prediction data for the current video block can include the predicted video block and various syntax elements.
[0503] The residual generation unit 207 can generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) multiple predicted video blocks from the current video block. The residual data for the current video block can include residual video blocks corresponding to different sample components of the samples in the current video block.
[0504] In other examples, such as in skip mode, residual data for the current video block may not exist, and the residual generation unit 207 may not perform the subtraction operation.
[0505] The transform processing unit 208 can generate one or more transform coefficient video blocks of the current video block by applying one or more transforms to the residual video block associated with the current video block.
[0506] After the transform processing unit 208 generates a transform coefficient video block associated with the current video block, the quantization unit 209 can quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0507] The inverse quantization unit 210 and the inverse transform unit 211 can apply inverse quantization and inverse transform to the transform coefficient video block respectively to reconstruct the residual video block from the transform coefficient video block. The reconstruction unit 212 can add the reconstructed residual video block to the corresponding samples of one or more predicted video blocks generated by the prediction unit 202 to produce a reconstructed video block associated with the current block for storage in the buffer 213.
[0508] After the video block is reconstructed by reconstruction unit 212, a loop filtering operation can be performed to reduce video block artifacts in the video block.
[0509] Entropy encoding unit 214 can receive data from other functional components of video encoder 200. When entropy encoding unit 214 receives data, it can perform one or more entropy encoding operations to generate entropy encoded data and output a bitstream including the entropy encoded data.
[0510] Some embodiments of the technology disclosed herein include determining or enabling a video processing tool or mode. In one example, when a video processing tool or mode is enabled, the encoder will use or implement the tool or mode in the processing of video blocks, but it is not necessary to modify the resulting bitstream based on the use of the tool or mode. In other words, when a video processing tool or mode is determined or enabled, the conversion from video blocks to a video bitstream will use the video processing tool or mode. In another example, when a video processing tool or mode is enabled, the decoder will utilize knowledge that the bitstream has already been modified based on the video processing tool or mode to process the bitstream. In other words, the conversion from a bitstream representation of video to video blocks will be performed using a video processing tool or mode determined or enabled.
[0511] Some embodiments of the technology disclosed herein include deciding or determining to disable video processing tools or modes. In one example, when video processing tools or modes are disabled, the encoder will not use the tools or modes in the conversion of video blocks to a bitstream representation of the video. In another example, when video processing tools or modes are disabled, the decoder will process the bitstream using knowledge that the bitstream has not yet been modified based on the decision or determination to disable video processing tools or modes.
[0512] Figure 41 A flowchart of an example method 4100 for video processing. Operation 4102 includes: performing a conversion between video blocks of video components and a bitstream representation of the video, wherein the video block comprises sub-blocks, wherein a filtering tool is used during the conversion according to rules, and wherein the rules specify that the filtering tool is applied by using a single offset for all samples of each sub-block of the video block.
[0513] In some embodiments of method 4100, the filtering tool includes a cross-component adaptive loop filter (CC-ALF) tool, wherein the sample values of sub-blocks of a video block of a video component are predicted from the sample values of another video component of the video. In some embodiments of method 4100, the filtering tool includes a chroma adaptive loop filter (chroma ALF) tool, wherein a loop filter is used to filter the samples of sub-blocks of a video block of a chroma video component. In some embodiments of method 4100, the filtering tool includes a luma adaptive loop filter (luma ALF) tool, wherein a loop filter is used to filter the samples of sub-blocks of a video block of a luma video component. In some embodiments of method 4100, the size of the sub-block corresponds to an M-column x N-row sample array. In some embodiments of method 4100, the size of the sub-block includes 1×2 samples, or 2×1 samples, or 2×2 samples, or 2×4 samples, or 4×2 samples, or 4×4 samples. In some embodiments of method 4100, the size of the sub-block is signaled in the bitstream representation.
[0514] In some embodiments of method 4100, the size of the sub-block is signaled in the bitstream representation at the sequence level, picture level, stripe level, or piece group level, or the size of the sub-block is signaled in the bitstream representation at the video region level. In some embodiments of method 4100, the sequence level, picture level, stripe level, or piece group level includes a sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), decoder parameter set (DPS), picture parameter set (PPS), adaptive parameter set (APS), stripe header, or piece group header. In some embodiments of method 4100, the video region level includes a codec tree block (CTB). In some embodiments of method 4100, a single offset is modified according to a second rule to obtain a modified single offset before refining all samples of each sub-block of the video block. In some embodiments of method 4100, the second rule specifies that the single offset is trimmed to the range T1 to T2. In some embodiments of method 4100, T1 = -5, and T2 = 5. In some embodiments of method 4100, the second rule specifies whether to modify a single offset and the technique for modifying a single offset.
[0515] In some embodiments of method 4100, the second rule specifies that the modified single offset is set to the value of T when a single offset meets certain conditions. In some embodiments of method 4100, the modified single offset is set to the value of T in response to a single offset being less than T1 and in response to a single offset being greater than T2. In some embodiments of method 4100, T = 0, T1 = -5, and T2 = 5. In some embodiments of method 4100, the second rule specifies that the technique for modifying the single offset is based on the width and / or height of the sub-block. In some embodiments of method 4100, the modified single offset is obtained by dividing the single offset by (M × W × H), where M is an integer, W is the width of the sub-block, and H is the height of the sub-block. In some embodiments of method 4100, M is 1 or 2. In some embodiments of method 4100, the second rule specifies that the modified single offset is obtained by multiplying or dividing the single offset by T, adding the single offset to T, or shifting the single offset by T, where T is an integer. In some embodiments of method 4100, T is signaled in the bitstream representation. In some embodiments of method 4100, T is based on sample values in a sub-block. In some embodiments of method 4100, T is based on the bit depth of the sample points of the video component.
[0516] In some embodiments of method 4100, the modified single offset is obtained by multiplying a single offset by T, where T = v / (2^B) max -1), where v represents the value of a sample point, B max The maximum bit depth of the sample points representing the video components. In some embodiments of method 4100, the second rule specifies whether to modify a single offset in the bitstream representation at the sequence level, picture level, stripe level, or piece group level, or wherein the second rule specifies whether to modify a single offset in the bitstream representation at the video region level. In some embodiments of method 4100, the sequence level, picture level, stripe level, or piece group level includes a sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), decoder parameter set (DPS), picture parameter set (PPS), adaptive parameter set (APS), stripe header, or piece group header. In some embodiments of method 4100, the video region level includes a codec tree block (CTB).
[0517] In some embodiments of method 4100, a representative sample point is selected from a sub-block, and the CC-ALF sample point set in the sub-block corresponding to a representative sample point is used to calculate a single offset shared by the sub-block. In some embodiments of method 4100, a representative sample point is a representative chroma sample point selected based on its position. In some embodiments of method 4100, a representative sample point is the Nth representative chroma sample point in the sub-block according to the scan order. In some embodiments of method 4100, N is 1 or 2, or the last sample point in the sub-block according to the scan order. In some embodiments of method 4100, a representative sample point is located at a specific position within the sub-block.
[0518] In some embodiments of method 4100, the specific location includes at least one of the following: the center location, the upper left location, the upper right location, the lower left location, or the lower right location within the sub-block. In some embodiments of method 4100, a representative sample is selected based on the sample value. In some embodiments of method 4100, a representative sample includes the sample with the largest sample value within the sub-block. In some embodiments of method 4100, a representative sample includes the sample with the smallest sample value within the sub-block. In some embodiments of method 4100, a representative sample includes the sample with the median sample value within the sub-block. In some embodiments of method 4100, a CC-ALF sample set corresponding to a representative sample is used using a CC-ALF tool. In some embodiments of method 4100, multiple representative samples are selected within the sub-block, and multiple CC-ALF sample sets within the sub-block corresponding to the multiple representative samples are used to compute a single offset shared by the sub-block. In some embodiments of method 4100, the corresponding CC-ALF sample set for each of the plurality of representative samples is first modified to generate a virtual representative filter sample set, and one of the virtual representative filter sample sets is used to calculate the offset.
[0519] Figure 42 A flowchart of an example method 4200 for video processing. Operation 4202 includes: performing a conversion between a video block of a video component and a bitstream representation of a video by using the final offset value of a current sample of a sub-block of a video block for a filtering tool applied during the conversion, wherein the filtering tool is applied by using the final offset value for the current sample, and wherein the final offset value is based on a first offset value and a second offset value different from the first offset value.
[0520] In some embodiments of method 4200, the first offset value is shared with the sub-block, and the second offset value is specific to the current sample. In some embodiments of method 4200, the first offset value is shared with neighboring samples of the current sample, and the second offset value is specific to the current sample. In some embodiments of method 4200, the first offset value is calculated using a first filter support region. In some embodiments of method 4200, the sub-block is 1×2 in size, the sub-block includes chroma samples located at (X, Y) and (X, Y+1), the sub-block includes luminance samples located at (2X-1, 2Y+1), (2X, 2Y+1), (2X+1, 2Y+1), (2X-1, 2Y+2), (2X, 2Y+2), and (2X+1, 2Y+2) for calculating the first offset value, and X and Y are integer offsets from the top-left corner of the video block located at (0, 0) to the chroma sample positions.
[0521] In some embodiments of method 4200, the sub-block size is 2×1, the sub-block includes chroma samples located at (X, Y) and (X+1, Y), and wherein the sub-block includes luminance samples located at (2X+1, 2Y-1), (2X+1, 2Y), (2X+1, 2Y+1), and (2X+1, 2Y+2) for calculating the first offset value, and X and Y are integer offsets from the top-left corner of the video block located at (0, 0) to the chroma sample position. In some embodiments of method 4200, the current sample includes chroma samples, and wherein the second offset value is calculated using a second filter support region. In some embodiments of method 4200, the sub-block has a size of 1×2, includes a chroma sample located at (X, Y), and includes luma samples located at (2X, 2Y) and (2X, 2Y+1), the luma samples being used to calculate a second offset value for the chroma samples, and wherein X and Y are integer offsets from the top-left corner of the video block located at (0, 0) to the chroma sample position. In some embodiments of method 4200, the sub-block has a size of 1×2, includes a chroma sample located at (X, Y+1), includes luma samples located at (2X, 2Y+2) and (2X, 2Y+3), the luma samples being used to calculate a second offset value for the chroma samples, and X and Y are integer offsets from the top-left corner of the video block located at (0, 0) to the chroma sample position.
[0522] In some embodiments of method 4200, the sub-block size is 2×1, the sub-block includes a chroma sample located at (X, Y), and a luminance sample located at (2X, 2Y), (2X+1, 2Y), (2X, 2Y+1), and (2X+1, 2Y+1). The luminance sample is used to calculate a second offset value for the chroma sample, and X and Y are integer offsets from the top-left corner of the video block located at (0, 0) to the chroma sample position. In some embodiments of method 4200, the sub-block size is 2×1, the sub-block includes a chroma sample located at (X+1, Y), and a luminance sample located at (2X+2, 2Y), (2X+3, 2Y), (2X+2, 2Y+1), and (2X+3, 2Y+1). The luminance sample is used to calculate a second offset value for the chroma sample, and X and Y are integer offsets from the top-left corner of the video block located at (0, 0) to the chroma sample position. In some embodiments of method 4200, the samples used to determine the first offset value and the second offset value are different from each other. In some embodiments of method 4200, at least some samples are used to determine both the first and second offset values. In some embodiments of method 4200, the number of samples used to determine the first offset value is the same as the number of samples used to determine the second offset value. In some embodiments of method 4200, the number of first samples used to determine the first offset value is greater than or equal to the number of second samples used to determine the second offset value.
[0523] In some embodiments of method 4200, the number of first samples used to determine the first offset value is less than or equal to the number of second samples used to determine the second offset value. In some embodiments of method 4200, fields in the bitstream representation indicate the first and second offset values used for conversion, and fields are indicated in the bitstream representation at the sequence level, picture level, strip level, or piece group level in the sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), decoder parameter set (DPS), picture parameter set (PPS), adaptive parameter set (APS), strip header, or piece group header. In some embodiments of method 4200, the first offset value is indicated in fields in the bitstream representation according to the scan order of the sub-blocks among a plurality of sub-blocks. In some embodiments of method 4200, the filtering tool includes a cross-component adaptive loop filtering (CC-ALF) tool, which predicts the sample values of sub-blocks of a video block of a video component from the sample values of another video component of the video. In some embodiments of method 4200, the filtering tool includes a chroma adaptive loop filter (chroma ALF) tool, in which a loop filter is used to filter samples of sub-blocks of a video block of the chroma video component. In some embodiments of method 4200, the filtering tool includes a luma adaptive loop filter (luma ALF) tool, in which a loop filter is used to filter samples of sub-blocks of a video block of the luma video component.
[0524] Figure 43 A flowchart of an example method 4300 for video processing. Operation 4302 includes: performing a conversion between video blocks of video components and a bitstream representation of video by using an N-tap symmetric filter for a cross-component adaptive loop filter (CC-ALF) tool during conversion, wherein the CC-ALF tool predicts sample values of video blocks of video components from sample values of another video component of the video, and wherein at least two filter coefficients of two samples within the support of the N-tap symmetric filter share the same value.
[0525] In some embodiments of method 4300, N = 6 or 8. In some embodiments of method 4300, the N-tap symmetric filter has a specific shape used in CC-ALF tools. In some embodiments of method 4300, the N-tap symmetric filter includes a symmetric 8-tap filter shape with 4 unique coefficients. In some embodiments of method 4300, the N-tap symmetric filter includes a symmetric 8-tap filter shape with 5 unique coefficients. In some embodiments of method 4300, the N-tap symmetric filter includes an 8-tap filter shape with 6 unique coefficients. In some embodiments of method 4300, the same value is signaled in the sequence-level or picture-level or slice-level or slice-level bitstream representation in the sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), decoder parameter set (DPS), picture parameter set (PPS), adaptive parameter set (APS), strip header, or slice header, and the same value is used to drive at least two filter coefficients.
[0526] Figure 44 A flowchart of an example method 4400 for video processing. Operation 4402 includes: performing a conversion between video blocks of video components and a bitstream representation of video by using an N-tap asymmetric filter for a cross-component adaptive loop filter (CC-ALF) tool during the conversion, where N is a positive integer, and where the CC-ALF tool predicts the sample values of video blocks of video components from sample values of another video component of the video.
[0527] In some embodiments of method 4400, N < 8. In some embodiments of method 4400, the N-tap asymmetric filter has an asymmetric 6-tap filter shape. In some embodiments of method 4400, the N-tap asymmetric filter has an asymmetric 5-tap filter shape. In some embodiments of method 4400, the N-tap asymmetric filter has an asymmetric 4-tap filter shape.
[0528] Figure 45 A flowchart of an example method 4500 for video processing. Operation 4502 includes: performing a conversion between a video block of a first video component of the video and a bitstream representation of the video, wherein the conversion of samples of the first video component includes applying a cross-component adaptive loop filter (CC-ALF) tool to the sample difference of a second video component of the video, and wherein the CC-ALF tool predicts the sample values of the video block of the first video component of the video from the sample values of another video component of the video.
[0529] In some embodiments of method 4500, the first video component is a chroma component, and the second video component is a luminance component. In some embodiments of method 4500, the sample difference of the second video component is determined by obtaining the difference between a representative luminance sample in the luminance filter support region and another luminance sample in the luminance filter support region. In some embodiments of method 4500, the sample is located at position (X, Y), and the representative luminance sample is selected from position (2X, 2Y), where X and Y are integer offsets from the top-left corner of the video block located at (0, 0) to the sample position. In some embodiments of method 4500, the sample point is located at position (X, Y), and the representative luminance sample point is selected from any one of the positions (2X, 2Y), (2X, 2Y+1), (2X+1, 2Y), (2X+1, 2Y-1), (2X, 2Y-1), (2X-1, 2Y-1), (2X-1, 2Y), and X and Y are integer offsets from the top left corner of the video block located at (0, 0) to the sample point position.
[0530] In some embodiments of method 4500, the sample point is located at position (X, Y), and the representative luminance sample point is selected from any one of the positions (2X, 2Y), (2X, 2Y+1), (2X+1, 2Y), (2X, 2Y-1), (2X-1, 2Y), and X and Y are integer offsets from the top left corner of the video block located at (0, 0) to the sample point position.
[0531] In some embodiments of method 4500, the sample point is located at position (X, Y), and a representative luminance sample point is selected from position (2X, 2Y), and another luminance sample point is located at (2X+1, 2Y), and X and Y are integer offsets from the top-left corner of the video block located at (0, 0) to the sample point position. In some embodiments of method 4500, the sample point is located at position (X, Y), and a representative luminance sample point and another luminance sample point are selected from positions (2X, 2Y+2), (2X+1, 2Y+1), (2X+1, 2Y), (2X, 2Y-1), (2X-1, 2Y), and (2X-1, 2Y+1), where X and Y are integer offsets from the top-left corner of the video block located at (0, 0) to the sample point position. In some embodiments of method 4500, the technique for selecting the representative sample point is predefined. In some embodiments of method 4500, the technique for selecting the representative sample point is based on information in the bitstream representation. In some embodiments of method 4500, the information includes the color format.
[0532] Figure 46 A flowchart of an example method 4600 for video processing. Operation 4602 includes: performing a conversion between sub-blocks of video blocks of a video component and a bitstream representation of the video by using two or more filters from a set of multiple filters of a cross-component adaptive loop filtering (CC-ALF) tool during conversion, wherein the CC-ALF tool predicts sample values of sub-blocks of video blocks of the video component from sample values of another video component of the video.
[0533] In some embodiments of method 4600, fields in the bitstream representation indicate the length N of a set of multiple filters and / or the support of two or more filters. In some embodiments of method 4600, fields are included at the sequence level, picture level, slice level, or piece level in the sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), decoder parameter set (DPS), picture parameter set (PPS), adaptive parameter set (APS), slice header, or piece group header. In some embodiments of method 4600, the offset used in the (CC-ALF) tool is determined by using luminance samples of the luminance components corresponding to sub-blocks using predefined support. In some embodiments of method 4600, the size of the sub-block is 1×2, the sub-block includes chroma samples located at (X, Y) and (X, Y+1), luminance samples located at (2X-1, 2Y), (2X-1, 2Y+1), (2X-1, 2Y+2), (2X-1, 2Y+3), (2X, 2Y-1), (2X, 2Y), (2X, 2Y+1), (2X, 2Y+2), (2X, 2Y+3), (2X, 2Y+4), (2X+1, 2Y), (2X+1, 2Y+1), (2X+1, 2Y+2), and (2X+1, 2Y+3), and X and Y are integer offsets from the top left corner of the video block located at (0, 0) to the sample positions. In some embodiments of method 4600, the size of the sub-block is 2×1, wherein the sub-block includes chroma samples located at (X, Y) and (X+1, Y), wherein luminance samples are located at (2X, 2Y-1), (2X, 2Y), (2X, 2Y+1), (2X, 2Y+2), (2X+1, 2Y-2), (2X+1, 2Y-1), (2X+1, 2Y), (2X+1, 2Y+1), (2X+1, 2Y+2), (2X+1, 2Y+3), (2X+2, 2Y-1), (2X+2, 2Y), (2X+2, 2Y+1), and (2X+2, 2Y+2), and wherein X and Y are integer offsets from the top left corner of the video block located at (0, 0) to the sample positions. In some embodiments of method 4600, the size of the sub-block is 1×2, wherein the sub-block includes chroma samples located at (X, Y) and (X, Y+1), wherein luminance samples are located at (2X-1, 2Y+1), (2X-1, 2Y+2), (2X, 2Y), (2X, 2Y+1), (2X, 2Y+2), (2X, 2Y+3), (2X+1, 2Y+1), and (2X+1, 2Y+2), and wherein X and Y are integer offsets from the top left corner of the video block located at (0, 0) to the sample positions.
[0534] In some embodiments of method 4600, the size of the sub-block is 2×1, the sub-block includes chroma samples located at (X, Y) and (X+1, Y), the luminance samples are located at (2X, 2Y), (2X, 2Y+1), (2X+1, 2Y-1), (2X+1, 2Y), (2X+1, 2Y+1), (2X+1, 2Y+2), (2X+2, 2Y) and (2X+2, 2Y+1), and X and Y are integer offsets from the top left corner of the video block located at (0, 0) to the sample positions. In some embodiments of method 4600, the size of the sub-block is 1×2, the sub-block includes chroma samples located at (X, Y) and (X, Y+1), the luminance samples are located at (2X-1, 2Y), (2X-1, 2Y+2), (2X, 2Y-1), (2X, 2Y), (2X, 2Y+1), (2X, 2Y+2), (2X, 2Y+3), (2X, 2Y+4), (2X+1, 2Y+1), and (2X+1, 2Y+3), and X and Y are integer offsets from the top left corner of the video block located at (0, 0) to the sample positions. In some embodiments of method 4600, the size of the sub-block is 2×1, the sub-block includes chroma samples located at (X, Y) and (X+1, Y), the luminance samples are located at (2X-1, 2Y), (2X, 2Y), (2X, 2Y+1), (2X+1, 2Y-1), (2X+1, 2Y), (2X+1, 2Y+1), (2X+1, 2Y+2), (2X+2, 2Y), (2X+2, 2Y+1), and (2X+3, 2Y+1), and X and Y are integer offsets from the top left corner of the video block located at (0, 0) to the sample positions.
[0535] In some embodiments of method 4600, the size of the sub-block is 2×1, wherein the sub-block includes chroma samples located at (X, Y) and (X+1, Y), wherein luminance samples are located at (2X-1, 2Y+1), (2X, 2Y), (2X, 2Y+1), (2X+1, 2Y-1), (2X+1, 2Y), (2X+1, 2Y+1), (2X+1, 2Y+2), (2X+2, 2Y), (2X+2, 2Y+1), and (2X+3, 2Y), and wherein X and Y are integer offsets from the top left corner of the video block located at (0, 0) to the sample positions.
[0536] Figure 47A flowchart of an example method 4700 for video processing. Operation 4702 includes: performing a conversion between a sub-block of a video block of a first video component and a bitstream representation of the video using a filter that supports multiple video components across the video for use during the conversion by a Cross-Component Adaptive Loop Filtering (CC-ALF) tool, wherein the CC-ALF tool predicts the sample values of the sub-block of the video block of the first video component from the sample values of another video component of the video.
[0537] In some embodiments of method 4700, the plurality of video components include a second video component and a third video component, and one or more syntax elements in the bitstream representation sequentially indicate the use of filter coefficients applied to the second and third video components to correct the first video component. In some embodiments of method 4700, the filter coefficients of the second video component are indicated in the bitstream representation before the filter coefficients of the third video component. In some embodiments of method 4700, one or more syntax elements are in an interleaved order, in which the filter coefficients of the second video component are interleaved with the filter coefficients of the third video component. In some embodiments of method 4700, the correction of samples in the first video component is performed by filtering samples in the second video component. In some embodiments of method 4700, the filter includes a symmetrical 8-tap filter shape with 4 unique coefficients. In some embodiments of method 4700, the filter includes a symmetrical 8-tap filter shape with 5 unique coefficients. In some embodiments of method 4700, the filter includes an 8-tap filter shape with 6 unique coefficients. In some embodiments of method 4700, correction of samples in the first video component is performed using samples from multiple video components. In some embodiments of method 4700, correction of samples in the first video component is performed by using the same filter on multiple video components.
[0538] In some embodiments of method 4700, the filter includes a symmetrical 8-tap filter shape with 4 unique coefficients, or the filter includes a symmetrical 8-tap filter shape with 5 unique coefficients, or the filter includes an 8-tap filter shape with 6 unique coefficients.
[0539] In some embodiments of method 4700, the samples of the first video component are corrected using an average or weighted average of offsets derived from multiple video components. In some embodiments of method 4700, the samples of the first video component are corrected by multiplying by offsets derived from multiple video components. In some embodiments of method 4700, the correction of the samples of the first video component is performed by using different filters on the multiple video components. In some embodiments of method 4700, the different filters include any two or more of a symmetric 8-tap filter shape with four unique coefficients, a symmetric 8-tap filter shape with five unique coefficients, and an 8-tap filter shape with six unique coefficients. In some embodiments of method 4700, the samples of the first video component are corrected using an average or weighted average of offsets derived from multiple video components. In some embodiments of method 4700, the samples of the first video component are corrected by multiplying by offsets derived from multiple video components.
[0540] Figure 48 This is a flowchart of an example method 4800 for video processing. Operation 4802 includes: performing a conversion between video blocks of video components and a bitstream representation of video by using samples from multiple video frames to refine the set of samples in the current video frame of the video in a Cross-Component Adaptive Loop Filter (CC-ALF) tool or an Adaptive Loop Filter (ALF) tool applied during the conversion, wherein the CC-ALF tool predicts the sample values of the video blocks of the video components from the sample values of another video component of the video, and wherein the ALF tool filters the samples of the video blocks of the video components using a loop filter.
[0541] In some embodiments of method 4800, whether to use samples from multiple video frames to refine the sample set in the current video frame is indicated in the sequence-level or picture-level or slice-level or slice-group-level bitstream representation in the sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), decoder parameter set (DPS), picture parameter set (PPS), adaptive parameter set (APS), slice header, or slice group header. In some embodiments of method 4800, the multiple video frames include a reference picture of the current video frame. In some embodiments of method 4800, the multiple video frames do not include a reference picture of the current video frame. In some embodiments of method 4800, the picture of the video closest to the current video frame is used in ALF tools and / or CC-ALF tools, where the codec picture (POC) technique is used to determine the distance between the picture and the current video frame.
[0542] In some embodiments of method 4800, a short-term image of the current video frame is used for the ALF tool and / or the CC-ALF tool. In some embodiments of method 4800, a long-term image of the current video frame is used for the ALF tool and / or the CC-ALF tool. In some embodiments of method 4800, a reference image in the same temporal layer of the current video frame is used for the ALF tool and / or the CC-ALF tool. In some embodiments of method 4800, reference images in different temporal layers of the current video frame are used for the ALF tool and / or the CC-ALF tool. In some embodiments of method 4800, multiple video frames include the current video frame.
[0543] Figure 49 A flowchart of example method 4900 for video processing. Operation 4902 includes: for the conversion between video blocks of video components and the bitstream representation of the video, determining whether to enable the Cross-Component Adaptive Loop Filter (CC-ALF) tool for the conversion based on a positional rule, wherein the CC-ALF tool predicts the sample values of the video blocks of the video components from the sample values of another video component of the video. Operation 4904 includes performing the conversion based on the determination.
[0544] In some embodiments of method 4900, the location rule specifies that the CC-ALF tool is disabled for chroma samples between the bottom boundary of the codec tree block and the virtual boundary of the loop filter. In some embodiments of method 4900, the virtual boundary is a line obtained by shifting a horizontal codec tree unit (CTU). In some embodiments of method 4900, the location rule specifies that the CC-ALF tool is disabled for chroma samples and / or luma samples between the bottom boundary of the codec tree unit (CTU) and the virtual boundary of the loop filter. In some embodiments of method 4900, the virtual boundary is a line obtained by shifting a horizontal codec tree unit (CTU). In some embodiments of method 4900, the location rule specifies that the CC-ALF tool is disabled at locations where another filter is applied during conversion. In some embodiments of method 4900, the location rule specifies that the CC-ALF tool is disabled for samples located at transform edges or codec unit (CU) edges, where the samples are filtered by another filter applied during conversion.
[0545] In some embodiments of methods 4100-4900, the syntax elements indicating the method of conversion used are included in the sequence-level, picture-level, strip-level, or slice-level bitstream representation in the sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), decoder parameter set (DPS), picture parameter set (PPS), adaptive parameter set (APS), strip header, or slice group header. In some embodiments of methods 4100-4900, the method is selectively applied based on the characteristics of the video indicated in the bitstream representation. In some embodiments of methods 4100-4900, the characteristics include color format or segmentation type or the position of sub-blocks relative to a codec unit (CU) or codec tree unit (CTU). In some embodiments of methods 4100-4900, performing the conversion includes encoding the video to generate the bitstream representation. In some embodiments of methods 4100-4900, performing the conversion includes parsing and decoding the bitstream representation to generate the video.
[0546] In some embodiments, a video decoding apparatus includes a processor configured to implement one or more of the methods 4100-4900. In some embodiments, a video encoding apparatus includes a processor configured to implement one or more of the methods 4100-4900. In some embodiments, a computer program product having computer code stored thereon, which, when executed by a processor, causes the processor to implement any one of the methods 4100-4900. In some embodiments, a computer-readable medium storing a bitstream representation or a bitstream representation generated according to any one or more of the methods 4100-4900.
[0547] Based on the foregoing, it will be understood that specific embodiments of the technology disclosed herein have been described for illustrative purposes, but various modifications may be made without departing from the scope of the invention. Accordingly, the technology disclosed herein is not limited to what is claimed in the appended claims.
[0548] The disclosed and other solutions, examples, embodiments, modules, and functional operations described in this document can be implemented in digital electronic circuits or in computer software, firmware, or hardware, containing the structures disclosed in this document and their equivalents, or combinations thereof. The disclosed and other embodiments can be implemented as one or more computer program products encoded on a computer-readable medium, such as one or more computer program instruction modules, for operation by or controlling a data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a complex influencing machine-readable propagating signals, or a combination thereof. The term "data processing apparatus" encompasses all means, devices, and machines for processing data, including, for example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, the apparatus may also include code that creates an environment for the computer program in question, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination thereof. Propagating signals are artificially generated signals, such as machine-generated electrical, optical, or electromagnetic signals, which are generated to encode information for transmission to a suitable receiver device.
[0549] Computer programs (also known as programs, software, software applications, scripts, or code) can be written in any programming language, including compiled or interpreted languages, and can be deployed in any form, including standalone programs or modules, components, subroutines, or other units suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple co-located files (e.g., a file storing one or more modules, subroutines, or code portions). A computer program can be deployed to execute on one computer or on multiple computers located at a single site or distributed across multiple sites and interconnected by a communications network.
[0550] The processes and logic flows described in this document can be executed by one or more programmable processors that perform one or more computer programs to perform functions by manipulating input data and generating outputs. The processes and logic flows can also be executed by special-purpose logic circuitry (e.g., field-programmable gate arrays (FPGAs) or application-specific integrated circuits (ASICs)), and the apparatus can be implemented as special-purpose logic circuitry (e.g., FPGAs or ASICs).
[0551] Processors suitable for executing computer programs include, for example, both general-purpose and special-purpose microprocessors, and any one or more processors of any type of digital computer. Typically, a processor receives instructions and data from read-only memory or random access memory, or both. The basic components of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include one or more mass storage devices (e.g., magneto-optical, magneto-optical, or optical disc) for storing data, or operatively coupled to receive data from or transfer data to a mass storage device (e.g., magneto-optical, magneto-optical, or optical disc), or both. However, a computer does not necessarily need to have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices (e.g., EPROM, EEPROM, and flash memory devices); magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD-ROMs and DVD-ROMs. Processors and memory may be supplemented by or incorporated into special-purpose logic circuitry.
[0552] While this patent document contains numerous details, these details should not be construed as limiting any subject matter or the scope of the claims, but rather as descriptions of features specific to particular embodiments of a particular technology. In this patent document, certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented separately in multiple embodiments or in various suitable sub-combinations. Furthermore, although features may be described above as operating in certain combinations and even initially claimed in the same manner, in certain circumstances one or more features from the claimed combination may be removed from the combination, and the claimed combination may be for sub-combinations or variations thereof.
[0553] Similarly, although operations are depicted in a specific order in the accompanying drawings, this should not be construed as requiring such operations to be performed in the specific order or sequence shown, or to perform all of the shown operations to achieve the desired result. Furthermore, the separation of various system components in the embodiments described in this patent document should not be construed as requiring such separation in all embodiments.
[0554] Only a few implementations and examples are described, and other implementations, enhancements and variations can be made based on what is described and shown in this patent document.
Claims
1. A method for processing video data, comprising: Perform the conversion between the video block of the first video component of the video and the bitstream of the video. The conversion of the samples of the first video component includes applying the cross-component adaptive loop filter CC-ALF tool to the sample difference of the second video component of the video. The CC-ALF tool predicts the sample values of the video block of the first video component from the sample values of another video component of the video. The first video component is a chroma component, and the second video component is a luminance component. The sample difference of the second video component is determined by obtaining the difference between a representative luminance sample in the luminance filter support region and another luminance sample in the luminance filter support region. The sample point is located at position (X, Y), and the representative luminance sample point is selected from position (2X, 2Y), and the other luminance sample point is located at (2X+1, 2Y), or the sample point is located at position (X, Y), and the representative luminance sample point is selected from position (2X, 2Y), and the other luminance sample point is located at (2X, 2Y+1). Where X and Y are integer offsets from the top left corner of the video block located at (0, 0) to the sample point position.
2. The method according to claim 1, wherein the sample point is located at position (X, Y), and the representative brightness sample point is selected from any one of the positions (2X, 2Y), (2X, 2Y+1), (2X+1, 2Y), (2X+1, 2Y-1), (2X, 2Y-1), (2X-1, 2Y-1), (2X-1, 2Y-1), (2X-1, 2Y).
3. The method according to claim 1, wherein the selection of the representative luminance sample is based on the color format of the video block.
4. The method of claim 1, wherein filters supporting multiple video components across the video are used in the cross-component adaptive loop filter CC-ALF tool during conversion.
5. The method of claim 4, wherein the plurality of video components includes a second video component and a third video component, and One or more syntax elements in the bitstream sequentially indicate the use of filter coefficients applied to the second and third video components to correct the first video component.
6. The method of claim 5, wherein the filter coefficients of the second video component are indicated in the bitstream before the filter coefficients of the third video component, or the filter coefficients of the second video component are interleaved with the filter coefficients of the third video component.
7. The method of claim 5, wherein the correction of the samples in the first video component is performed by filtering the samples in the second video component, or the correction of the samples in the first video component is performed by using samples from the plurality of video components.
8. The method of claim 7, wherein the correction of the samples in the first video component is performed by using the same filter on the plurality of video components.
9. The method of claim 8, wherein the filter comprises a symmetrical 8-tap filter shape having 4 unique coefficients, or wherein the filter comprises a symmetrical 8-tap filter shape having 5 unique coefficients, or wherein the filter comprises an 8-tap filter shape having 6 unique coefficients.
10. The method according to claim 8, wherein, The samples in the first video component are corrected using the average or weighted average of the offsets derived from the plurality of video components, or the samples in the first video component are corrected by multiplying by the offsets derived from the plurality of video components.
11. The method of claim 1, wherein samples from multiple video frames of the video are used to refine the set of samples in the current video frame of the video in a cross-component adaptive loop filter CC-ALF tool or an adaptive loop filter ALF tool applied during conversion, and The ALF tool uses a loop filter to filter the samples of the video block of the first video component.
12. The method of claim 11, wherein whether or not samples from the plurality of video frames are used to refine the sample set in the current video frame is indicated in the bitstream at the sequence level, picture level, strip level, or slice level in the sequence header, picture header, sequence parameter set SPS, video parameter set VPS, decoder parameter set DPS, picture parameter set PPS, adaptive parameter set APS, strip header, or slice header.
13. The method of claim 11, wherein the plurality of video frames includes a reference image of the current video frame or the current video frame, and The image of the video that is closest to the current video frame is used by the ALF tool and / or the CC-ALF tool. The short image of the current video frame is used by the ALF tool and / or the CC-ALF tool; The long-term image of the current video frame is used in the ALF tool and / or the CC-ALF tool; The reference image in the same temporal layer of the current video frame is used by the ALF tool and / or the CC-ALF tool; or The reference images in different temporal layers of the current video frame are used in the ALF tool and / or the CC-ALF tool.
14. The method of claim 1, wherein the conversion comprises encoding the video block into the bitstream.
15. The method of claim 1, wherein the conversion comprises decoding the video block from the bitstream.
16. An apparatus for processing video data, the apparatus comprising a processor and a non-transitory memory storing instructions thereon, wherein the instructions, when executed by the processor, cause the processor to: Perform the conversion between the video block of the first video component of the video and the bitstream of the video. The conversion of the samples of the first video component includes applying the cross-component adaptive loop filter CC-ALF tool to the sample difference of the second video component of the video, and The CC-ALF tool predicts the sample values of the video block of the first video component from the sample values of another video component of the video. The first video component is a chroma component, and the second video component is a luminance component. The sample difference of the second video component is determined by obtaining the difference between a representative luminance sample in the luminance filter support region and another luminance sample in the luminance filter support region. The sample point is located at position (X, Y), and the representative luminance sample point is selected from position (2X, 2Y), and the other luminance sample point is located at (2X+1, 2Y), or the sample point is located at position (X, Y), and the representative luminance sample point is selected from position (2X, 2Y), and the other luminance sample point is located at (2X, 2Y+1). Where X and Y are integer offsets from the top left corner of the video block located at (0, 0) to the sample point position.
17. A non-transitory computer-readable storage medium storing instructions, said instructions causing a processor to: Perform the conversion between the video block of the first video component of the video and the bitstream of the video. The conversion of the samples of the first video component includes applying the cross-component adaptive loop filter CC-ALF tool to the sample difference of the second video component of the video, and The CC-ALF tool predicts the sample values of the video block of the first video component from the sample values of another video component of the video. The first video component is a chroma component, and the second video component is a luminance component. The sample difference of the second video component is determined by obtaining the difference between a representative luminance sample in the luminance filter support region and another luminance sample in the luminance filter support region. The sample point is located at position (X, Y), and the representative luminance sample point is selected from position (2X, 2Y), and the other luminance sample point is located at (2X+1, 2Y), or the sample point is located at position (X, Y), and the representative luminance sample point is selected from position (2X, 2Y), and the other luminance sample point is located at (2X, 2Y+1). Where X and Y are integer offsets from the top left corner of the video block located at (0, 0) to the sample point position.
18. A method for storing a video bitstream, comprising: For the video blocks of the first video component of the video, generate the bitstream of the video; as well as The bitstream is stored in a non-transitory computer-readable recording medium. The generation of samples for the first video component includes applying a cross-component adaptive loop filter (CC-ALF) tool to the sample difference of the second video component of the video, and The CC-ALF tool predicts the sample values of the video block of the first video component from the sample values of another video component of the video. The first video component is a chroma component, and the second video component is a luminance component. The sample difference of the second video component is determined by obtaining the difference between a representative luminance sample in the luminance filter support region and another luminance sample in the luminance filter support region. The sample point is located at position (X, Y), and the representative luminance sample point is selected from position (2X, 2Y), and the other luminance sample point is located at (2X+1, 2Y), or the sample point is located at position (X, Y), and the representative luminance sample point is selected from position (2X, 2Y), and the other luminance sample point is located at (2X, 2Y+1). Where X and Y are integer offsets from the top left corner of the video block located at (0, 0) to the sample point position.
Citation Information
Patent Citations
A high definition signal decoder
CN101060627A