Method and device for video processing and medium
By using filter-based non-intra-frame or intra-frame prediction methods and leveraging various filter techniques to optimize video encoding and decoding, the problem of improving encoding and decoding efficiency in existing technologies is solved, and more efficient video processing is achieved.
Patent Information
- Application Number
- CN202480036518.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-05-29
- Filing Date
- 2024-05-29
- Publication Date
- 2026-01-02
AI Technical Summary
Existing video encoding and decoding technologies have room for improvement in encoding and decoding efficiency, especially in intra-frame prediction, where it is difficult to further improve encoding and decoding efficiency.
The method employs filter-based non-intra-frame or intra-frame prediction, using Transform Unit (TU), Codec Unit (CU), Prediction Unit (PU), or sub-blocks to determine the prediction of the current video region. It utilizes various filter techniques for video processing, including slope adjustment, gradient PDPC, secondary MPM list, and DIMD chroma mode fusion.
It improves the effectiveness and efficiency of video encoding and decoding, optimizes the intra-frame prediction process, and enhances encoding and decoding performance.
Smart Images

Figure CN121264051A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present disclosure generally relate to video processing technology, and more particularly, to filter-based prediction in video coding. BACKGROUND
[0002] Nowadays, digital video capability is being applied to various aspects of people's life. For video coding / decoding, various types of video compression technologies have been proposed, such as MPEG-2, MPEG-4, ITU-T H.263, ITU-T H.264 / MPEG-4 Part 10 Advanced Video Coding (AVC), ITU-T H.265 High Efficiency Video Coding (HEVC) standard, Versatile Video Coding (VVC) standard. However, it is generally desired to further improve the coding efficiency of video coding technology. SUMMARY
[0003] Embodiments of the present disclosure provide a solution for video processing.
[0004] In a first aspect, a method for video processing is proposed. The method comprises: for conversion between a current video region of a video and a bitstream of the video, determining a filter-based non-intra or intra prediction for the current video region based on one of: a transform unit (TU), a coding unit (CU), a prediction unit (PU), or a sub-block; and performing the conversion based on the filter-based non-intra or intra prediction. The method according to the first aspect of the present disclosure uses the filter-based non-intra or intra prediction, thereby improving the coding effectiveness and coding efficiency.
[0005] In a second aspect, an apparatus for video processing is proposed. The apparatus comprises a processor and a non-transitory memory having instructions thereon. The instructions, when executed by the processor, cause the processor to perform the method according to the first aspect of the present disclosure.
[0006] In a third aspect, a non-transitory computer-readable storage medium is proposed. The non-transitory computer-readable storage medium stores instructions that cause a processor to perform the method according to the first aspect of the present disclosure.
[0007] In a fourth aspect, another non-transitory computer-readable recording medium is proposed. The non-transitory computer-readable recording medium stores a bitstream of a video generated by a method performed by an apparatus for video processing. The method comprises: determining a filter-based non-intra or intra prediction for a current video region of the video based on one of: a transform unit (TU), a coding unit (CU), a prediction unit (PU), or a sub-block; and generating the bitstream based on the filter-based non-intra or intra prediction.
[0008] In a fifth aspect, a method for storing a bitstream of a video is suggested. The method includes determining a filter-based non-intra or intra prediction of a current video region of the video based on one of a transform unit (TU), a coding unit (CU), a prediction unit (PU), or a sub-block; generating the bitstream based on the filter-based non-intra or intra prediction; and storing the bitstream in a non-transitory computer-readable recording medium.
[0009] This summary is provided to introduce a selection of concepts that are further described below in the detailed description. This summary is not intended to identify key or essential features of the claimed subject matter, nor is it used to limit the scope of the claimed subject matter. BRIEF DESCRIPTION OF DRAWINGS
[0010] The above and other objects, features and advantages of the example embodiments of the present disclosure will be more apparent from the following detailed description taken in conjunction with the accompanying drawings, in which like reference characters refer to like elements throughout. In the example embodiments of the present disclosure, the same reference numbers in different drawings identify the same components.
[0011] Figure 1 A block diagram showing an example video coding system is shown in accordance with some embodiments of the present disclosure; Figure 2 A block diagram showing a first example video encoder is shown in accordance with some embodiments of the present disclosure; Figure 3 A block diagram showing an example video decoder is shown in accordance with some embodiments of the present disclosure; Figure 4 A diagram showing the effect of the slope adjustment parameter "u". Left: model created with current CCLM. Right: updated model as suggested; Figure 5 Neighboring blocks (L, A, BL, AR, AL) used in derivation of the general MPM list are shown; Figure 6 Neighboring reconstructed samples used for DIMD chroma mode are shown; Figure 7 Intra template matching search region used is shown; Figure 8 Use of IntraTMP block vector for IBC blocks is shown; Figure 9A And Figure 9B Partitioning methods for angular modes are shown respectively; Figure 10 Extended MRL candidate list is shown; Figure 11 A diagram of the template region is shown; Figure 12The spatial part of the convolution filter is shown; Figure 13 The reference region used to derive the filter coefficients (with its padding) is shown; Figure 14 Four Sobel-based gradient patterns used for GLM are shown; Figure 15 The spatial GPM candidate is shown; Figure 16 The GPM template is shown; Figure 17 The GPM blending is shown; Figure 18 The possible positions of the candidate region are shown; Figure 19 The positions of the neighboring spatial candidates are shown; Figure 20 The transform selection process for the planar mode with direction is shown; Figure 21 The luma block used to derive the direct block vector is shown; Figure 22 The three types of reconstruction regions defined including thirteen columns or rows of reconstructed pixels are shown; Figure 23 The three types of filter shapes defined with fifteen inputs and generating one output are shown; Figure 24 Examples of prediction for different positions in the current block are shown; Figure 25 The spatial part of the filter is shown; Figure 26 The reference region used to derive the filter coefficients is shown; Figure 27 The filter shape is shown; Figure 28 The proposed method on the decoder is shown; Figure 29 The luma samples L0,..., L5 are shown with respect to the chroma sample C (shown in the half-pel luma grid); Figure 30 An example of the filter being modulated by the correlation of the current template (padded with a grid) with the reference template (padded with diagonal lines) in the reference picture is shown; Figure 31 An example of the filter being modulated by the correlation of the current template (padded with a grid) with the reference template (padded with diagonal lines) in the current picture is shown; Figure 32 A flowchart of a method for video processing according to an embodiment of the present disclosure is shown; and Figure 33 A block diagram of a computing device in which various embodiments of the present disclosure can be implemented is shown.
[0012] Throughout the drawings, identical or similar reference numerals can designate identical or similar elements throughout the several views. DETAILED DESCRIPTION
[0013] The principles of the present disclosure will now be described with reference to some embodiments. It should be understood that the embodiments are described for illustrative purposes only and help the person skilled in the art to understand and implement the present disclosure, but do not imply any limitation on the scope of the present disclosure. The disclosure described herein can be implemented in various ways in addition to the ways described below.
[0014] In the following description and claims, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs.
[0015] Reference in the present disclosure to "one embodiment", "an embodiment", "example embodiment" etc. indicates that a described embodiment can include a particular feature, structure, or characteristic, but every embodiment can not necessarily include the particular feature, structure, or characteristic. Moreover, such phrases are not necessarily referring to the same embodiment. Furthermore, when a particular feature, structure, or characteristic is described in connection with an example embodiment, it is submitted that this is within the knowledge of those skilled in the art to effect such feature, structure, or characteristic in connection with other embodiments whether or not explicitly described.
[0016] It should be understood that although the terms "first" and "second" etc. can be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first element could be termed a second element, and, similarly, a second element could be termed a first element without departing from the scope of the example embodiments. As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed terms.
[0017] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of example embodiments. As used herein, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises", "comprising", "includes" and / or "including", when used herein, specify the presence of stated features, elements and / or components etc. but do not preclude the presence or addition of one or more other features, elements, components and / or combinations thereof.
[0018] Example Environment Figure 1is a block diagram illustrating an example video coding system 100 that can utilize the techniques of this disclosure. As shown, video coding system 100 can include a source device 110 and a destination device 120. Source device 110 can also be referred to as a video encoding device, and destination device 120 can also be referred to as a video decoding device. In operation, source device 110 can be configured to generate encoded video data, and destination device 120 can be configured to decode the encoded video data generated by source device 110. Source device 110 can include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.
[0019] Video source 112 can include a source such as a video capture device. Examples of video capture devices include, but are not limited to, an interface to receive video data from a video content provider, a computer graphics system to generate video data, and / or a combination thereof.
[0020] Video data can include one or more pictures. Video encoder 114 encodes video data from video source 112 to generate a bitstream. The bitstream can include a sequence of bits that form an encoded representation of the video data. The bitstream can include encoded pictures and associated data. An encoded picture is an encoded representation of a picture. The associated data can include sequence parameter sets, picture parameter sets, and other syntax structures. I / O interface 116 can include a modulator / demodulator and / or a transmitter. Encoded video data can be transmitted directly to destination device 120 via I / O interface 116, over network 130A. Encoded video data can also be stored onto a storage medium / server 130B for access by destination device 120.
[0021] Destination device 120 can include an I / O interface 126, a video decoder 124, and a display device 122. I / O interface 126 can include a receiver and / or a modem. I / O interface 126 can acquire encoded video data from source device 110 or storage medium / server 130B. Video decoder 124 can decode the encoded video data. Display device 122 can display the decoded video data to a user. Display device 122 can be integrated with destination device 120, or can be external to destination device 120 which is configured to interface with an external display device.
[0022] Video encoder 114 and video decoder 124 can operate according to a video compression standard, such as the High Efficiency Video Coding (HEVC) standard, Versatile Video Coding (VVC) standard, and other existing and / or future standards.
[0023] Figure 2is a block diagram illustrating an example of a video encoder 200 that can be Figure 1 the example system 100.
[0024] The video encoder 200 can be configured to implement any or all of the techniques of this disclosure. In Figure 2 In the example of FIG. 2, the video encoder 200 includes a plurality of functional components. The techniques described in this disclosure can be shared between the components of the video encoder 200. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.
[0025] In some embodiments, the video encoder 200 can include a partitioning unit 201, a prediction unit 202 that can include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra-prediction unit 206, a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213, and an entropy encoding unit 214.
[0026] In other examples, the video encoder 200 can include more, less, or different functional components. In one example, the prediction unit 202 can include an intra block copy (IBC) unit. The IBC unit can perform prediction in an IBC mode in which at least one reference picture is the picture in which the current video block is located.
[0027] Furthermore, although some components, such as the motion estimation unit 204 and the motion compensation unit 205, can be integrated, these components are shown separately in Figure 2 the example of FIG. 2 for illustrative purposes.
[0028] The partitioning unit 201 can partition a picture into one or more video blocks. The video encoder 200 and the video decoder 300 can support various video block sizes.
[0029] The mode selection unit 203 can select one of a plurality of encoding modes (intra- or inter- coding) based on, for example, error results, and provide the resulting intra- or inter- coded block to the residual generation unit 207 to generate residual block data and to the reconstruction unit 212 to reconstruct the encoded block for use as a reference picture. In some examples, the mode selection unit 203 can select a combined intra-inter prediction (CIIP) mode in which prediction is based on both an inter prediction signal and an intra prediction signal. In the case of inter prediction, the mode selection unit 203 can also select a resolution for motion vectors (e.g., sub-pixel precision or integer pixel precision) for the block.
[0030] To perform inter prediction on a current video block, motion estimation unit 204 can generate motion information for the current video block by comparing one or more reference frames from buffer 213 to the current video block. Motion compensation unit 205 can determine a predicted video block for the current video block based on the motion information and decoded samples of pictures from buffer 213 other than the picture associated with the current video block.
[0031] Motion estimation unit 204 and motion compensation unit 205 can perform different operations on a current video block, e.g., depending on whether the current video block is in an I slice, a P slice, or a B slice. As used herein, an "I slice" can refer to a portion of a picture composed of macroblocks that are all based on macroblocks within the same picture. Further, as used herein, a "P slice" and a "B slice" can refer to portions of a picture composed of macroblocks that are independent of macroblocks in the same picture, in some aspects.
[0032] In some examples, motion estimation unit 204 can perform uni-prediction on a current video block, and motion estimation unit 204 can search reference pictures of List 0 or List 1 for a reference video block for the current video block. Motion estimation unit 204 can then generate a reference index that indicates a reference picture of List 0 or List 1 that contains the reference video block, and a motion vector that indicates a spatial displacement between the current video block and the reference video block. Motion estimation unit 204 can output the reference index, a prediction direction indicator, and the motion vector as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current video block based on the reference video block indicated by the motion information for the current video block.
[0033] Alternatively, in other examples, motion estimation unit 204 can perform bi-prediction on a current video block. Motion estimation unit 204 can search reference pictures of List 0 for one reference video block for the current video block, and can also search reference pictures of List 1 for another reference video block for the current video block. Motion estimation unit 204 can then generate multiple reference indices that indicate multiple reference pictures of List 0 and List 1 that contain the multiple reference video blocks, and multiple motion vectors that indicate multiple spatial displacements between the multiple reference video blocks and the current video block. Motion estimation unit 204 can output the multiple reference indices and the multiple motion vectors for the current video block as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current video block based on the multiple reference video blocks indicated by the motion information for the current video block.
[0034] In some examples, the motion estimation unit 204 can output a full set of motion information for use in the decoding process by the decoder. Alternatively, in some embodiments, the motion estimation unit 204 can signal the motion information for the current video block with reference to the motion information of another video block. For example, the motion estimation unit 204 can determine that the motion information for the current video block is sufficiently similar to the motion information of a neighboring video block.
[0035] In one example, the motion estimation unit 204 can indicate a value in a syntax structure associated with the current video block that indicates to the video decoder 300 that the current video block has the same motion information as another video block.
[0036] In another example, the motion estimation unit 204 can identify another video block and a motion vector difference (MVD) in a syntax structure associated with the current video block. The motion vector difference indicates a difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 300 can use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[0037] As discussed above, the video encoder 200 can signal motion vectors in a predictive manner. Two examples of predictive signaling techniques that can be implemented by the video encoder 200 include advanced motion vector prediction (AMVP) and Merge mode signaling.
[0038] The intra prediction unit 206 can perform intra prediction on the current video block. When the intra prediction unit 206 performs intra prediction on the current video block, the intra prediction unit 206 can generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block can include a predicted video block and various syntax elements.
[0039] The residual generation unit 207 can generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) the prediction video block(s) for the current video block from the current video block. The residual data for the current video block can include residual video blocks corresponding to different sample components of the samples in the current video block.
[0040] In other examples, such as in skip mode, there can be no residual data for the current video block for the current video block, and the residual generation unit 207 can not perform the subtraction operation.
[0041] The transform processing unit 208 can generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video blocks associated with the current video block.
[0042] After the transform processing unit 208 generates a transform coefficient video block associated with the current video block, the quantization unit 209 can quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0043] The inverse quantization unit 210 and the inverse transform unit 211 can apply inverse quantization and inverse transform to the transform coefficient video block, respectively, to reconstruct the residual video block from the transform coefficient video block. The reconstruction unit 212 can add the reconstructed residual video block to the corresponding samples of one or more predicted video blocks generated by the prediction unit 202 to produce a reconstructed video block associated with the current video block for storage in the buffer 213.
[0044] After the video block is reconstructed by reconstruction unit 212, a loop filtering operation can be performed to reduce video block artifacts in the video block.
[0045] Entropy encoding unit 214 can receive data from other functional components of video encoder 200. When entropy encoding unit 214 receives data, it can perform one or more entropy encoding operations to generate entropy-encoded data and output a bitstream including the entropy-encoded data.
[0046] Figure 3 This is a block diagram illustrating an example of a video decoder 300 according to some embodiments of the present disclosure. The video decoder 300 may be... Figure 1 An example of video decoder 124 in system 100 is shown.
[0047] The video decoder 300 can be configured to perform any or all of the technologies disclosed herein. Figure 3 In the example, the video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video decoder 300. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.
[0048] exist Figure 3 In the example, the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra-frame prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, the video decoder 300 can perform a decoding process that is generally contrasted with the encoding process described with respect to the video encoder 200.
[0049] Entropy decoding unit 301 can retrieve an encoded bitstream. The encoded bitstream can include entropy encoded video data (e.g., encoded blocks of video data). Entropy decoding unit 301 can decode the entropy encoded video data, and motion compensation unit 302 can determine motion information from the entropy decoded video data, including motion vectors, motion vector precision, reference picture list index, and other motion information. Motion compensation unit 302 can determine such information, for example, by performing AMVP and Merge modes. AMVP is used, including deriving a number of most probable candidates based on data from neighboring PBs and reference pictures. The motion information typically includes a horizontal motion vector displacement value and a vertical motion vector displacement value, one or two reference picture indices, and in the case of a prediction region in a B slice, an identification of which reference picture list is associated with each index. As used herein, in some aspects, "Merge mode" can refer to deriving motion information from a spatially or temporally neighboring block.
[0050] Motion compensation unit 302 can generate a motion compensated block, possibly performing interpolation based on an interpolation filter. An identifier for the interpolation filter used at sub-pixel precision can be included in the syntax elements.
[0051] Motion compensation unit 302 can use the interpolation filter used by video encoder 200 during encoding of the video block to calculate interpolated values for sub-integer pixels of the reference block. Motion compensation unit 302 can determine the interpolation filter used by video encoder 200 from the received syntax information, and motion compensation unit 302 can use the interpolation filter to generate the prediction block.
[0052] Motion compensation unit 302 can use at least some of the syntax information to determine the size of the blocks used to encode frames and / or slices of the encoded video sequence, partitioning information describing how each macroblock of a picture of the encoded video sequence is partitioned, modes indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-coded block, and other information used to decode the encoded video sequence. As used herein, in some aspects, a "slice" can refer to a data structure that can be decoded independently of other slices of the same picture in terms of entropy coding, signal prediction, and residual signal reconstruction. A slice can be an entire picture, or can also be a region of a picture.
[0053] Intra prediction unit 303 can use intra prediction modes, e.g., received in the bitstream, to form a prediction block from spatial neighboring blocks. Dequantization unit 304 dequantizes quantized video block coefficients provided in the bitstream and decoded by entropy decoding unit 301. Inverse transform unit 305 applies an inverse transform.
[0054] The reconstruction unit 306 can obtain the decoded block, e.g., by adding the residual block to the corresponding prediction block generated by the motion compensation unit 302 or the intra prediction unit 303. If needed, a deblocking filter can also be applied to filter the decoded block in order to remove blocking artifacts. The decoded video blocks are then stored in the buffer 307, which provides reference blocks for subsequent motion compensation / intra prediction, and which also produces decoded video for presentation on a display device.
[0055] Some example embodiments of the present disclosure will be described in detail below. It should be noted that the use of section headings in this document is for convenience only and not to be construed as limiting the embodiments disclosed in that section to that section only. Furthermore, although some embodiments are described with reference to a multi-functional video codec or other specific video codec, the disclosed techniques are applicable to other video coding technologies. Moreover, although some embodiments describe video encoding steps in detail, it will be appreciated that corresponding decoding steps will be implemented by a decoder. Furthermore, the term video processing includes video encoding or compression, video decoding or decompression, and video transcoding, in which video pixels are represented from one compressed format to another compressed format or at a different compressed bit rate.
[0056] 1. BRIEF OVERVIEW The present disclosure relates to video coding techniques. In particular, it is about non-intra prediction in image / video coding. It can be applied to existing video coding standards such as HEVC, VVC, etc. It can also be applicable to future video coding standards or video codecs.
[0057] 2. INTRODUCTION Video coding standards have evolved mainly through the development of the well-known ITU-T and ISO / IEC standards. The ITU-T produced H.261 and H.263 standards, ISO / IEC produced MPEG-1 and MPEG-4 Visual, and both organizations produced the H.262 / MPEG-2 Video standard and the H.264 / MPEG-4 Advanced Video Coding (AVC) standard and the H.265 / HEVC standard. From H.262, video coding standards are based on the hybrid video coding structure, where temporal prediction is combined with transform coding. To explore future video coding technologies beyond HEVC, the Joint Video Exploration Team (JVET) was founded by VCEG and MPEG jointly in 2015. The JVET meeting is held once every quarter, and the new video coding standard was officially named as Versatile Video Coding (VVC) in the April 2018 JVET meeting, and the first version of VVC test model (VTM) was released at that time. The working draft of VVC and the test model VTM are updated after each meeting. The VVC project achieved the technical completion (FDIS) in the July 2020 meeting.
[0058] 2.1 Intra prediction In intra prediction, the minimum chroma intra prediction unit (SCIPU) constraint in VVC is removed. In addition, the VPDU constraint for reducing CCLM prediction delay is also removed.
[0059] 2.1.1 Multi-model LM (MMLM) The CCLM included in VVC is extended by adding three multi-model LM (MMLM) modes. In each MMLM mode, the reconstructed neighboring samples are classified into two categories using a threshold that is the average of the luma reconstructed neighboring samples. A linear model for each category is derived using the least mean square (LMS) method. For the CCLM mode, a linear model is also derived using the LMS method. A slope adjustment is applied to the cross-component linear model (CCLM) and multi-model LM prediction. The adjustment is to tilt the linear function that maps luma values to chroma values with respect to a center point determined by the average luma value of the reference samples.
[0060] 2.1.1.1 Slope adjustment for CCLM CCLM maps luma values to chroma values using a model with 2 parameters. The slope parameter “a” and the bias parameter “b” define the mapping as follows:
[0061] The adjustment “u” to the slope parameter is signaled to update the model to the following form:
[0062] where
[0063] With this selection, the mapping function is tilted or rotated around the point with luminance value y r The average value of the reference luma samples used in the model creation is taken as y r in order to provide a meaningful modification of the model. The following picture illustrates this process. Figure 4 A diagram showing the effect of the slope adjustment parameter "u". Figure 4 The left side of the picture corresponds to the model created with the current CCLM. Figure 4 The right side of the picture corresponds to the updated model as proposed.
[0064] Embodiments The slope adjustment parameter is provided as an integer between -4 and 4 (inclusive) and is signaled in the bitstream. The unit of the slope adjustment parameter is 1 / 8 of a chroma sample value per one luma sample value (for 10-bit content).
[0065] For CCLM models using reference samples from both the top and the left of the block ("LM_CHROMA IDX" and "MMLM_CHROMA IDX"), the adjustment is available, but not for "single-sided" modes. This selection is based on a trade-off between coding efficiency and complexity.
[0066] When applying slope adjustment to multi-mode CCLM models, both models can be adjusted, so up to two slope updates are signaled for a single chroma block.
[0067] Encoder scheme The proposed encoder method performs a SATD-based search for the best value of the slope update for Cr and a similar SATD-based search for Cb. If either results in a non-zero slope adjustment parameter, the combined slope adjustment pair (SATD-based update for Cr, SATD-based update for Cb) is included in the list of RD checks for the TU.
[0068] 2.1.2 Gradient PDPC In VVC, for some scenarios, PDPC can not be applied due to the unavailability of the auxiliary reference samples. In these cases, a gradient-based PDPC extended from the horizontal / vertical mode is applied. The PDPC weight (wT / wL) and nScale parameters used to determine the decay of the PDPC weight with respect to the distance from the left / above boundary are set equal to the corresponding parameters in the horizontal / vertical mode. When the auxiliary reference samples are located at fractional sample positions, bilinear interpolation is applied.
[0069] 2.1.3 Secondary MPM A secondary MPM list is introduced. The existing primary MPM (PMPM) list consists of 6 entries, and the secondary MPM (SMPM) list includes 16 entries. A general MPM list with 22 entries is first constructed, then the first 6 entries in the general MPM list are included into the PMPM list, and the remaining entries form the SMPM list. The first entry in the general MPM list is the planar mode. The remaining entries consist of the intra modes of the left (L), above (A), below-left (BL), above-right (AR), and above-left (AL) neighboring blocks, the band direction modes with offsets added to the first two available band directions of the neighboring blocks, and the default mode.
[0070] If the CU block is vertically oriented, the order of the neighboring blocks is A, L, BL, AR, AL; otherwise, the order is L, A, BL, AR, AL. Figure 5 The neighboring blocks (L, A, BL, AR, AL) used in the derivation of the general MPM list are shown.
[0071] The PMPM flag is first parsed, if equal to 1, the PMPM index is parsed to determine which entry of the PMPM list is selected, otherwise the SPMPM flag is parsed to determine whether to parse the SMPM index or the remaining mode.
[0072] 2.1.4 Reference sample interpolation and smoothing for intra prediction The 4-tap cubic interpolation is replaced with a 6-tap cubic interpolation filter for deriving the prediction samples from the reference samples.
[0073] For reference sample filtering, a 6-tap Gaussian filter is applied for larger blocks (W >= 32 and H >= 32), otherwise the existing VVC 4-tap Gaussian interpolation filter is applied. The 4-tap interpolation filter is used instead of nearest-neighbor rounding to derive the extended intra reference samples.
[0074] 2.1.5 Decoder-side intra mode derivation (DIMD) When DIMD is applied, two intra modes are derived from the reconstructed neighboring samples, and these two prediction values are combined with the planar mode prediction value, where the weights are derived from the gradient. The division operation in the weight derivation is performed with the same LUT-based integerization scheme used by CCLM. For example, the division operation in the orientation calculation
[0075] The following LUT-based scheme is used to calculate:
[0076] where DivSigTable[ 16 ] = { 0, 7, 6, 5, 5, 4, 4, 3, 3, 2, 2, 1, 1, 1, 1, 0}.
[0077] The derived intra mode is included into the main list of intra most probable modes (MPM), thus the DIMD process is performed before constructing the MPM list. The main derived intra mode of the DIMD block is stored together with the block and used for the MPM list construction of the neighboring blocks.
[0078] 2.1.5.1 DIMD chroma mode The DIMD chroma mode uses the DIMD derivation method to derive the chroma intra prediction mode of the current block based on the neighboring reconstructed Y, Cb and Cr samples in the second neighboring row and column. Specifically, the horizontal and vertical gradients are calculated for each collocated reconstructed luma sample and the reconstructed Cb and Cr samples of the current chroma block to construct the HoG. Then, the intra prediction mode with the largest histogram amplitude value is used to perform the chroma intra prediction of the current chroma block. Figure 6 The neighboring reconstructed samples used for the DIMD chroma mode are shown.
[0079] When the intra prediction mode derived from the DIMD chroma mode is the same as the intra prediction mode derived from the DM mode, the intra prediction mode with the second largest histogram amplitude value is used as the DIMD chroma mode. A CU level flag is signaled to indicate whether the proposed DIMD chroma mode is applied.
[0080] 2.1.6 Fusion of chroma intra prediction modes The DM mode and the four default modes can be fused with the MMLM LT mode as follows:
[0081] wherein is the prediction value obtained by applying the non-LM mode, is the prediction value obtained by applying the MMLM LT mode, and is the final prediction value of the current chroma block. The two weights and are determined by the intra prediction modes of the neighboring chroma blocks, and is set equal to 2. Specifically, when both the above neighboring block and the left neighboring block are coded with LM mode, = { 1, 3}; when both the above neighboring block and the left neighboring block are coded with non-LM mode, = { 3, 1}; otherwise, = { 2, 2}.
[0082] For syntax design, if non-LM mode is chosen, a flag is signaled to indicate whether fusion is applied or not. This method is only applicable to I slice.
[0083] 2.1.7 Intra Template Matching Intra Template Matching Prediction (IntraTMP) is a special Intra prediction mode that copies the best prediction block from the reconstructed part of the current frame, whose L-shaped template matches the current template. For a predefined search range, the encoder searches for the most similar template to the current template in the reconstructed part of the current frame and uses the corresponding block as the prediction block. The encoder then signals the use of this mode and the same prediction operation is performed at the decoder side.
[0084] The prediction signal is generated by matching the L-shaped causal neighbors of the current block with another block in a predefined search area in Figure 7 which includes: R1: the current CTU; R2: the top-left CTU; R3: the top CTU; R4: the left CTU.
[0085] The sum of absolute differences (SAD) is used as the cost function.
[0086] Within each region, the decoder searches for the template with the smallest SAD with respect to the current template and uses its corresponding block as the prediction block.
[0087] The dimensions of all regions (SearchRange_w, SearchRange_h) are set to be proportional to the block dimensions (BlkW, BlkH) to have a fixed number of SAD comparisons per pixel. That is:
[0088] where “ ” is a constant that controls the gain / complexity trade-off. In practice, “ ” is equal to 5. Figure 7 The Intra Template Matching search regions used are shown.
[0089] To speed up the template matching process, the search range of all search regions is down-sampled by a factor of 2. This results in a 4x reduction of the template matching search. After the best match is found, a refinement process is performed. The refinement is done by a second template matching search around the best match with a reduced range. The reduced range is defined as min(BlkW, BlkH) / 2.
[0090] The Intra Template Matching tool is enabled for CUs with width and height size smaller or equal to 64. This maximum CU size for Intra Template Matching is configurable.
[0091] When DIMD is not used for the current CU, the Intra Template Matching prediction mode is signaled at CU level by a dedicated flag.
[0092] 2.1.7.1 IntraTMP derived block vector candidates for IBC In this method, the block vectors (BV) derived from Intra Template Matching Prediction (IntraTMP) are used for Intra Block Copy (IBC). The stored IntraTMP BVs of neighboring blocks are used together with the IBC BVs as spatial BV candidates in the IBC candidate list construction.
[0093] The IntraTMP block vectors are stored in the IBC block vector buffer and the current IBC block can use both the IBC BV and the IntraTMP BV of the neighboring blocks as BV candidates for the IBC BV candidate list as shown in Figure 8 .
[0094] The IntraTMP block vectors are added to the IBC block vector candidate list as spatial candidates.
[0095] 2.1.8 Fusion for Template based Intra mode derivation (TIMD) For each Intra prediction mode in the MPMs, the SATD between the prediction samples of the template and the reconstructed samples is calculated. The first two Intra prediction modes with the smallest SATD are selected as the TIMD modes. After applying the PDPC process, these two TIMD modes are fused with a weight and this weighted Intra prediction is used to code the current CU. The position dependent Intra prediction combination (PDPC) is included in the derivation of the TIMD modes.
[0096] The cost of the two selected modes is compared with a threshold and the cost factor of 2 is applied in the test as follows:
[0097] If this condition is true, the fusion is applied, otherwise only mode1 is used.
[0098] The weight of the modes is calculated from their SATD cost as follows:
[0099] The division operation is done using the same LUT based integerization scheme used by CCLM.
[0100] 2.1.9 Intra prediction fusion The intra prediction method derives the prediction samples as a weighted combination of multiple predictors generated from different reference lines. In this process, multiple intra prediction values are generated and then fused by a weighted average. The process of deriving the prediction values to be used in the fusion process is described as follows: 1) For angular intra prediction modes including TIMD and DIMD, the proposed method derives the intra prediction by weighting the intra predictions obtained from multiple reference lines represented as where is the intra prediction from the default reference line, and is the prediction from the line above the default reference line. The weights are set to and .
[0101] 2) For TIMD mode with mixing, is used for the first mode ( ), and is used for the second mode ( ).
[0102] 3) For DIMD mode with mixing, the number of selected prediction values for the weighted average is increased from 3 to 6.
[0103] When the angular intra mode has a non-integer slope (required reference sample interpolation) and the block size is larger than 16, the intra prediction fusion method is applied to the luma block which uses MRL together, while not applied to the ISP coded block. In the method studied in sub-test a, PDPC is applied to the intra prediction mode using the reference line closest to the current block.
[0104] 2.1.10 Combination of CIIP with TIMD and TM Merge In the CIIP mode, the prediction samples are generated by weighting the inter prediction signal predicted using the CIIP-TM Merge candidate and the intra prediction signal predicted using the intra prediction mode derived by TIMD. This method is only applied to the coded blocks with an area smaller than or equal to 1024.
[0105] The TIMD derivation method is used to derive the intra prediction mode in CIIP. Specifically, the intra prediction mode with the smallest SATD value in the TIMD mode list is selected and mapped to one of the 67 regular intra prediction modes.
[0106] Furthermore, it is proposed to modify the weights (wIntra, wInter) for both tests if the derived intra prediction mode is an angular mode. For near-horizontal modes (2 <= angle mode index < 34), the current block is vertically split; for near-vertical modes (34 <= angle mode index <= 66), the current block is horizontally split.
[0107] (wIntra, wInter) for different sub-blocks are as shown in Figure 9A and Figure 9B .
[0108] Table 1. Modified weights for angular patterns
[0109] With CIIP-TM, a CIIP-TM Merge candidate list is constructed for the CIIP-TM mode. Merge candidates are refined by template matching. CIIP-TM Merge candidates are also reordered as regular Merge candidates by the ARMC method. The maximum number of CIIP-TM Merge candidates is equal to 2.
[0110] 2.1.11 Extended multi-reference line (MRL) list The MRL list in VVC is extended to include more reference lines for intra prediction. The extended reference line list consists of the line indices {1, 3, 5, 7, 12}. For template-based intra mode derivation (TIMD), only the first two reference line candidates (i.e., {1, 3}) are used, instead of the full MRL candidate list. Figure 10 The extended MRL candidate list is shown.
[0111] 2.1.12 Template-based multi-reference line intra prediction The template-based multi-reference line intra prediction (TMRL) mode combines a reference line and a prediction mode together and uses a template matching method to construct a list of candidate combinations. The index to the candidate combination list is coded to indicate which reference line and prediction mode to use when coding the current block. The regular multi-reference line (MRL) for non-TIMD parts is replaced by the TMRL mode.
[0112] The TMRL mode extends the reference line candidate list and the intra prediction mode candidate list. The extended reference line candidate list is {1, 3, 5, 7, 12}. The restriction on the top CTU row remains unchanged. The size of the intra prediction mode candidate list is 10. The construction of the intra prediction mode candidate list is similar to MPM, except that the planar mode is excluded from the intra prediction mode candidate list, the DC mode is added after the 5 neighboring PU modes and the DIMD mode (if it is not included), and the additional modes with the to an angular mode of an incremental angle (compared to the existing angular modes in the intra prediction mode candidate list).
[0113] The TMRL candidate is constructed as follows. For a block, there are 5x10=50 combinations of the extended reference line and allowed intra prediction modes. Since the extended reference line starts from reference line 1, the area covered by reference line 0 is used for template matching. The SAD cost on the template area (see Figure 11 ) is calculated between the prediction (generated by the 50 combinations) and the reconstruction. The 20 combinations with the smallest SAD cost are selected in ascending order to form the TMRL candidate list.
[0114] For TMR signaling, instead of coding the reference line and the intra mode directly, the index of the TMRL candidate list is coded to indicate which combination of the reference line and the prediction mode is used to code the current block.
[0115] 2.1.13 Convolutional Cross-Component Intra Prediction Model In this method, a convolutional cross-component model (CCCM) is applied in a similar spirit as the current CCLM mode to predict the chroma samples from the reconstructed luma samples. As with CCLM, when chroma downsampling is used, the reconstructed luma samples are downsampled to match the lower resolution chroma grid. Similar to CCLM, the top, left, or both top and left reference samples are used as the template for model derivation.
[0116] In addition, similar to CCLM, there is an option to use a single model or a multi-model variant of the CCCM. The multi-model variant uses two models, one model is derived for samples above the average luma reference value and the other model is derived for the remaining samples (following the spirit of the CCLM design). The multi-model CCCM mode can be selected for PUs with at least 128 available reference samples.
[0117] 2.1.13.1 Convolutional Filter The convolutional 7-tap filter consists of a 5-tap plus sign-shaped spatial component, a non-linear term, and a bias term. The input to the spatial 5-tap component of the filter is composed of the center (C) luma sample co-located with the chroma sample to be predicted and its above / north (N), below / south (S), left / west (W), and right / east (E) neighbors, as follows. Figure 12 The spatial part of the convolutional filter is shown.
[0118] The non-linear term P is expressed as the 2nd power of the center luma sample C and is scaled to the sample value range of the content:
[0119] That is, for 10-bit content, it is computed as:
[0120] The bias term B represents a scalar offset between input and output (similar to the offset term in CCLM) and is set to the middle chroma value (512 for 10-bit content).
[0121] The output of the filter is computed as the convolution of the filter coefficients c i with the input values, and is clipped to the range of valid chroma samples: predChromaVal = c0C + c1N + c2S + c3E + c4W + c5P + c6B.
[0122] 2.1.13.2 Computation of filter coefficients The filter coefficients c i are computed by minimizing the MSE between the predicted chroma samples and the reconstructed chroma samples in the reference region. Figure 13 A reference region consisting of 6 rows of chroma samples above and to the left of the PU is shown. The reference region extends one PU width to the right of the PU boundary and one PU height below the PU boundary. The region is adjusted to only include available samples. The region extension shown in blue is needed to support the "edge samples" of the plus-shaped spatial filter, and is padded when in unavailable regions.
[0123] The MSE minimization is performed by computing the auto-correlation matrix for the luma input and the cross-correlation vector between the luma input and the chroma output. The auto-correlation matrix is LDL-decomposed, and the final filter coefficients are computed using back-substitution. This process roughly follows the computation of the ALF filter coefficients in ECM, however LDL-decomposition is chosen instead of Cholesky decomposition to avoid the use of square root operations.
[0124] The auto-correlation matrix is computed using the reconstructed values of the luma and chroma samples. These samples are full-range (e.g. between 0 and 1023 for 10-bit content), resulting in relatively large values in the auto-correlation matrix. This requires high bit-depth operations during the model parameter computation. It is proposed to remove a fixed offset from the luma and chroma samples in each PU for each model. This reduces the amplitude of the values used in the model creation, and allows to reduce the precision required for fixed-point arithmetic. As a result, it is proposed to use 16-bit decimal precision instead of the 22-bit precision of the original CCCM implementation.
[0125] For simplicity, the reference sample values immediately outside the top-left corner of the PU are used as offsets (offsetLuma, offsetCb, and offsetCr). The sample values used in both the model creation and the final prediction (i.e., the luma and chroma in the reference region and the luma in the current PU) reduce these fixed values as follows:
[0126] And the chroma values are predicted using the following equation, where offsetChroma is equal to offsetCr and offsetCb for the Cr and Cb components, respectively: predChromaVal = c0C' + c1N' + c2S' + c3E' + c4W' + c5P' + c6B +offsetChroma.
[0127] To avoid any additional sample-level operations, the luma offsets are removed during the luma reference sample interpolation. This can be done, for example, by replacing the rounding term used in the luma reference sample interpolation with an updated offset that includes both the rounding term and offsetLuma. The chroma offsets can be removed by subtracting them directly from the reference chroma samples. As an alternative, the effect of the chroma offsets can be removed from the cross-component vector, resulting in the same outcome. To add the chroma offsets back to the output of the convolutional prediction operation, the chroma offsets are added to the bias term of the convolutional model.
[0128] The process of CCCM model parameter calculation requires division operations. Division operations are not always considered friendly to implementations. Division operations are replaced by multiplication (with a scaling factor) and shift operations, where the scaling factor and the number of shifts are computed based on the denominator, similar to the method used in the computation of the CCLM parameters.
[0129] 2.1.13.3 Gradient linear model For YUV 4:2:0 color format, the gradient linear model (GLM) method can be used to predict chroma samples from luma sample gradients. Two modes are supported: two-parameter GLM mode and three-parameter GLM mode.
[0130] Compared to CCLM, two-parameter GLM utilizes luma sample gradients to derive the linear model instead of down-sampled luma values. Specifically, when applying two-parameter GLM, the input of the CCLM process (i.e., down-sampled luma samples ) is replaced by luma sample gradients . Other parts of CCLM (e.g., parameter derivation, prediction sample linear transformation) remain unchanged.
[0131]
[0132] In the three-parameter GLM, the chroma samples can be predicted based on both the luma sample gradient with different parameters and the down-sampled luma values. The model parameters of the three-parameter GLM are derived from the 6 rows and columns of neighboring samples by the MSE minimization method based on LDL decomposition as used in CCCM.
[0133]
[0134] For signaling, when CCLM mode is enabled for the current CU, one flag is signaled to indicate whether GLM is enabled for both Cb and Cr components; if GLM is enabled, another flag is signaled to indicate which of the two GLM modes is selected, and one syntax element is further signaled to select one of the 4 gradient filters for gradient calculation.
[0135] Four gradient filters are used for GLM, as shown in Figure 14 .
[0136] 2.1.13.4 Bitstream signaling The use of the mode is signaled using a PU-level flag that is CABAC coded. A new CABAC context is included to support this. When it comes to signaling, CCCM is considered as a sub-mode of CCLM. That is, the CCCM flag is signaled only when the intra prediction mode is LM_CHROMA.
[0137] 2.1.14 Spatial Geometry Partition Mode (SGPM) SGPM is an intra mode of inter coding tool similar to GPM, where two prediction parts are generated from the intra prediction process. In this mode, a candidate list is constructed, where each entry contains one partition split and two intra prediction modes, as shown in Figure 15 . 26 partition modes and 3 intra prediction modes are used to form the combinations. The length of the candidate list is set to equal 16. The selected candidate index is signaled.
[0138] The list is reordered using template matching ( Figure 16 ), where the SAD between the prediction of the template and the reconstruction is used for the ordering. The template size is fixed to 1.
[0139] For each partition mode, the same intra-inter GPM list derivation is used to derive the IPM list for each part. The IPM list size is set to 3. In the list, the mode derived by TIMD is replaced by 2 derived modes with horizontal and vertical orientation.
[0140] The SGPM pattern is applied with a restricted block size, which is: 4 <= width <= 64, 4 <= height <= 64, and width < height. 8. Height < Width 8. Width Height >= 32.
[0141] Adaptive blending has also been used in spatial GPM, where Figure 17 The mixing depth τ shown is derived as follows:
[0142] 2.1.15 Nonlocal cross-component prediction Cross-component prediction (CCP), including CCLM, CCCM, and their variants, is employed by ECM to leverage cross-component correlations. With CCLM or CCCM, training samples are always adjacent to the current block. However, the cross-component relationships of the current block can be more relevant to cross-component relationships in non-local regions.
[0143] A nonlocal cross-component prediction method is proposed to enhance CCP by gaining more advantages from nonlocal regions.
[0144] Method #1: A Non-Adjacent Cross-Component Prediction (NA-CCP) model is proposed. Using the NA-CCP model, samples from regions that are not adjacent to the current block can be used to derive the CCCM model for the current block. A candidate region list with 6 candidates is constructed by sequentially examining potential 8×8 regions. If an examined region is available, it is added to the candidate region list. The upper-left position of the potential 8×8 regions is pre-determined as {(-xStep, 0), (0, -yStep), (xStep, -yStep), (-xStep, yStep), (-xStep, -yStep), (-2...} xStep, 0), (0, -2) yStep), (-2 xStep, 2 yStep), (2) xStep, -2 yStep), (-2 xStep, yStep), (xStep, -2 yStep), (-2 xStep, -yStep), (-xStep, -2 yStep), (-2 xStep, -2 { (0, 0), (xStep, 0), (xStep, yStep), (0, yStep), (-xStep, yStep), (-xStep, 0), (-xStep, -yStep), (0, -yStep), (xStep, -yStep), (xStep, -xStep / 2), (0, -xStep / 2), (-xStep / 2, -xStep / 2), (-xStep / 2, 0), (0, xStep / 2), (xStep / 2, xStep / 2), (-xStep / 2, yStep / 2), (-xStep / 2, -yStep / 2)} where xStep = Max(width, 16), yStep = Max(height, 16). Figure 18 Some possible positions of the candidate regions are shown.
[0145] A flag is signaled to indicate whether NA-CCP is applied to the chroma block. If NA-CCP is applied, an index is signaled to indicate which candidate in the candidate region list is used to derive the CCCM model.
[0146] Method #2: A history-based cross-component prediction (H-CCP) mode is proposed. With H-CCP, similar to the HMVP table, a H-CCLM table and a H-CCCM table are maintained. After a block coded with CCLM or CCCM is decoded, the corresponding table is updated. In the implementation of H-CCP, the size of the H-CCLM table or the H-CCCM table is 6. If the current block is coded with CCLM or CCCM mode, a flag is signaled to indicate whether H-CCP is applied. If H-CCP is used, an index is further signaled to indicate which candidate model in the H-CCLM table or the H-CCCM table is selected.
[0147] 2.1.16 Cross-component Merge mode for chroma intra coding Cross-component prediction (CCP) including cross-component linear model (CCLM), convolutional cross-component model (CCCM) and gradient linear model (GLM) is adopted by ECM to exploit cross-component correlation. A cross-component Merge (CCMerge) mode is proposed as a new CCP mode. The cross-component model parameters of the current chroma block coded with CCMerge can be inherited from the neighboring blocks coded with CCP. With CCMerge, CCP can be more efficient and has less signaling overhead.
[0148] In CCMerge, the final cross-component model parameters of the current chroma block can be inherited from its spatial neighbors and non-adjacent neighbors or default model. A list is created including the CCP models from the spatial neighbors and non-adjacent neighbors coded with CCLM, MMLM, CCCM, GLM, chroma merge and CCMerge mode. After including the neighboring CCP models, a default model is further included to fill the remaining empty positions in the list. To avoid including redundant CCP models in the list, a de-duplication operation is applied. More details are described as follows. Figure 19The positions of the spatial neighboring candidates are shown.
[0149] Spatial non-adjacent neighboring candidates The positions of the spatial neighboring candidates are shown. Figure 19 The spatial candidates are included in the following order: B1 -> A1 -> B0 -> A0 -> B2.
[0150] Spatial non-adjacent neighboring candidates After checking all spatial adjacent neighboring candidates, spatial non-adjacent neighboring candidates are considered. In the current ECM design, in Inter Merge mode, two sets of spatial non-adjacent neighboring candidates are obtained. In the proposed method, the positions and inclusion order of the spatial non-adjacent neighboring candidates from the first set are used.
[0151] CCLM candidates with default scaling parameters If the list is not full, after including the spatial adjacent and non-adjacent candidates, CCLM candidates with default scaling parameters are considered. The default scaling parameters are {0, 1 / 8, -1 / 8, 2 / 8, -2 / 8, 3 / 8}, and the offset parameters are derived according to the selected default scaling parameter, the average neighboring reconstructed luma sample value (Yavg) and the average neighboring reconstructed Cb / Cr sample value (Cavg).
[0152] 2.1.16.1 Merge model candidates When merging CCLM candidates, only the scaling parameters are inherited. The offset parameters are derived by using the inherited scaling parameters, Yavg and Cavg.
[0153] When merging MMLM candidates, the scaling parameters and the classification threshold are inherited. The offset parameters in each class are derived according to the inherited classification threshold and Yavg and Cavg in each class. If there is no available neighboring reconstructed sample in a certain class, the offset parameter is directly inherited from the candidate.
[0154] When merging CCCM candidates, all the convolution parameters, offsets (i.e., offsetLuma, offsetCb and offsetCr) and the classification threshold are inherited.
[0155] When merging GLM candidates, if the GLM candidate is a 3 -parameter GLM mode, all the gradient pattern indices and model parameters are inherited; otherwise, if the GLM candidate is a 2 -parameter GLM mode, the offset parameters are derived by using the inherited scaling parameters, Yavg and Cavg.
[0156] When merging chroma fusion candidates, the derived MMLM parameters are inherited and used as the merged MMLM candidate.
[0157] For a CCMerge block, if its merged candidate mode is CCLM, MMLM, CCCM or GLM, the merged candidate mode is stored as the propagated mode of the current chroma block; otherwise, if its merged candidate mode is chroma fusion, the propagated mode is set to MMLM. When merging a CCMerge candidate, how to inherit or derive the CCP parameters depends on the propagated mode of the CCMerge candidate, as described in the above five paragraphs.
[0158] 2.1.16.2 Signaling After the cclm_mode_flag syntax element, an additional flag is signaled to indicate whether CCMerge is used or not. If CCMerge is used, the candidate index is additionally signaled. The candidate index signaled is shared for Cb / Cr color components. Currently, the maximum allowed number of candidates is set to the default value 6. If the maximum allowed number of candidates is modified to 1, the candidate index does not need to be signaled. Each bin of the candidate index is context coded with a separate context.
[0159] 2.1.17 Planar mode with direction Two additional planar modes, where only horizontal interpolation or only vertical interpolation is used to obtain the predicted samples.
[0160] For the planar horizontal mode, only horizontal linear interpolation is performed based on the left reference sample and the top-right reference sample to predict the current sample as:
[0161] For the planar vertical mode, only vertical linear interpolation is performed based on the top reference sample and the bottom-left reference sample to predict the current sample as:
[0162] The transform kernel selection for the planar horizontal mode and the planar vertical mode is shown in Figure 20 If the intra prediction mode of the current block is the planar vertical mode, the horizontal intra prediction mode is used to derive the transform kernels in the MTS set and the LFNST set. In addition, if the intra prediction mode of the current block is the planar horizontal mode, the vertical intra prediction mode is used to derive the transform kernels in the MTS set and the LFNST set.
[0163] 2.1.18 Direct block vector for chroma blocks Direct block vector is used for chroma blocks in dual tree slice. When chroma dual tree is activated, a flag is signaled to indicate whether the chroma block is coded using IBC mode or not. If Figure 21 one of the five positions of the luma block is coded with IBC or intraTMP mode, its block vector is scaled and used as the block vector for the chroma block. Template matching is used to perform the block vector scaling.
[0164] 2.1.19 Extrapolation filter based intra prediction mode (EFI mode) The proposed extrapolation filter based intra prediction is processed in two steps. First, extrapolation filter coefficients are obtained from the neighboring reconstructed pixels of the current block with a predetermined template. Second, extrapolation generates the prediction values from top-left to bottom-right position by position within the current block.
[0165] 2.1.19.1 Searching mean, minimum and maximum Similar to CCCM mode, the mean should be removed when feeding the input to the EIP filter. The value of the DC mode of the current block is used as the mean for EIP prediction. The minimum and maximum are searched from the reconstructed pixels in the reconstructed region with thirteen columns and thirteen rows.
[0166] 2.1.19.2 Calculation of filter coefficients Three types of reconstructed regions and three filter shapes are proposed as shown in Figure 22 When the current block is predicted with the proposed EIP mode, the decoder decodes the related syntax elements to determine the selected reconstructed region type and filter shape for the current block.
[0167] Figure 22 Three types of reconstructed regions defined including thirteen columns or rows of reconstructed pixels are shown.
[0168] Figure 23 Three types of filter shapes defined with fifteen inputs and generating one output are shown.
[0169] The selected filter is slid in the selected reconstructed region with one pixel step to collect the input samples and output samples of EIP. While removing the mean from the input samples and output samples, the auto-correlation matrix and cross-correlation vector are constructed. Then, the EIP coefficients are obtained by the same method in CCCM.
[0170] 2.1.19.3 Prediction of the current block EIP mode predicts the current block position by position as shown in Figure 24
[0171] For positions located in the top-left corner of the current block, the input of the EIP filter is a reconstructed sample.
[0172] For positions located along the boundary of the current block, part of the input of the EIP filter is a reference sample and part of the input of the EIP filter is a previously predicted sample.
[0173] For other positions in the current block, the input of the EIP filter is a previously predicted sample.
[0174] To reduce the prediction error, the searched minimum and maximum values are applied to limit the output range of each prediction value,
[0175] is the prediction value at (x, y) in the current block, is the searched minimum and maximum values from the thirteen reconstructed columns and rows, is the i-th coefficient of the derived EIP filter, is the predicted reconstructed value or prediction value for the current position, is the value computed by the DC prediction mode.
[0176] 2.1.20 Intra TMP based on linear filter model The proposed 6-tap filter consists of a 5-tap spatial component plus a sign-shaped bias term. The input of the spatial 5-tap component of the filter consists of the center (C) sample in the reference block and its above / north (N), below / south (S), left / west (W) and right / east (E) neighbors, which are located in corresponding positions in the current block to be predicted, as follows. Figure 25 The spatial part of the filter is shown.
[0177] The bias term B represents a scalar offset between input and output, and is set to the middle luma value (512 for 10-bit content).
[0178] The output of the filter is computed as follows: predLumaVal = c0C + c1N + c2S + c3E + c4W + c5B.
[0179] As shown in Figure 26 The filter coefficients ci are computed by minimizing the MSE between the reference and current templates. The area extension shown in the blue area is needed to support the "edge samples" of the plus-shaped spatial filter, and is padded when in the unavailable area.
[0180] MSE minimization is performed by calculating the autocorrelation matrix for the reference template input and the current template output. The autocorrelation matrix is decomposed using LDL, and the final filter coefficients are calculated using back substitution.
[0181] The Intra TMP-FLM mode is used by transmitting the CU-level flag of the codec via signal transmission. Specifically, IntraTMP-FLM is considered a sub-mode of Intra TMP. That is, the Intra TMP-FLM flag is only transmitted via signal transmission when the Intra TMP flag is true.
[0182] 2.1.21 Filtered Intra-Frame Block Copy (FIBC) A filtered IBC (FIBC) is proposed, which applies a linear filter to the predicted samples of the IBC. For example... Figure 27 As shown, the proposed filter consists of five spatial terms and one bias term. The five spatial terms consist of the center (C) position and its neighbors above / north (N), below / south (S), left / west (W), and right / east (E).
[0183] predVal = C+ N + S + W + E +
[0184] in It is a coefficient. This is the offset. The filter coefficients are derived using up to four rows / columns of samples above and to the left of the current CU. The filter coefficients are derived based on minimizing the difference between the template samples and their corresponding reference samples using the same regression-based minimization technique employed in ECM, which is used by other tools such as CCCM.
[0185] For signaling, an additional indication flag is introduced for FIBC, which is transmitted via signaling after the IBC-LIC flag. Specifically, when the IBC-LIC flag is true, this flag is transmitted via signaling and used to indicate whether FIBC is applied to the current block.
[0186] 2.2 Cross-component residual model (CCRM) for inter-frame prediction 2.2.1 Introduction It is proposed that when blocks use inter-frame prediction or intra-block copy (IBC), the cross-component residual model (CCRM) is applied to predict chrominance samples from reconstructed luminance samples. Figure 28The decoder side of the method is shown. The cross-component filter uses the prediction signal of luma and chroma is derived. The derived filter is applied to the reconstructed luma signal to produce the final chroma prediction.
[0187] 2.2.2 Calculation of the convolution filter and filter coefficients The proposed 8-tap filter consists of 6 spatial luma samples, a non-linear term and a bias term. As shown in Figure 29 the spatial luma samples (L0,..., L5) are obtained from a luma grid that selects the 6 luma samples closest to the chroma position C without downsampling. The predicted chroma value is obtained as:
[0188] where nonlinear is the operator of the non-linearity of the CCCM and B is the bias.
[0189] Figure 29 The luma samples L0,.., L5 are shown with respect to the chroma sample C (shown in the half-pel luma grid).
[0190] The filter coefficients are derived using the ECM's method of extrapolation-free Gaussian elimination and before filter derivation, necessary offsets are applied to the samples.
[0191] When a block has less than 64 chroma samples, the intra reference samples are used as additional input samples in the filter derivation. The CCCM design uses up to 6 rows and columns of intra reference samples.
[0192] Blocks with 256 or more chroma samples are divided into sub-blocks with up to 256 chroma samples. Sub-blocks containing zero luma residuals are skipped.
[0193] 2.2.3 Bitstream signaling The use of the mode is signaled using a TU-level flag that is CABAC coded. A new CABAC context is included to support this. The CCRM flag is signaled only if the luma Cbf of the TU is non-zero and the predMode of the CU is either MODE INTER or MODE IBC.
[0194] 2.2.4 Multi-hypothesis prediction (MHP) In the multi-hypothesis inter prediction mode, one or more additional motion- compensated prediction signals are signaled in addition to the traditional bi-prediction signal. The resulting overall prediction signal is obtained by a sample-wise weighted superposition. With the bi-prediction signal and the first additional inter prediction signal / hypothesis the resulting prediction signal is obtained as follows:
[0195] According to the following mapping, the weighting factor The new syntax element add_hyp_weight_idx specifies the following:
[0196] Similar to the above, more than one additional prediction signal can be used. The resulting overall prediction signal is iteratively accumulated with each additional prediction signal.
[0197]
[0198] The resulting overall prediction signal is obtained as the last one. (i.e., has the largest index) of Within this EE, a maximum of two additional prediction signals can be used (i.e., Limited to 2).
[0199] The motion parameters for each additional prediction hypothesis can be transmitted via signaling by explicitly specifying the reference index, motion vector prediction index, and motion vector difference, or by implicitly specifying the merge index. A separate multi-hypothesis merge flag distinguishes between these two signaling modes.
[0200] For inter-frame AMVP mode, MHP is only applied when unequal weights in BCW are selected in bidirectional prediction mode.
[0201] Combining MHP and BDOF is possible; however, BDOF is only applied to the bidirectional prediction signal portion of the predicted signal (i.e., the ordinary first two assumptions).
[0202] 2.3 On the cross-component model for residual encoding and decoding in image and video encoding and decoding The detailed embodiments described below should be considered as examples for explaining general concepts. These embodiments should not be interpreted in a narrow sense. Furthermore, these embodiments can be combined in any way.
[0203] The term "video unit" or "code-decoder unit" can refer to a picture, strip, slice, code-decoder tree block (CTB), code-decoder tree unit (CTU), code-decoder block (CB), CU, PU, TU, PB, TB.
[0204] The term "block" can refer to code-decode tree block (CTB), code-decode tree unit (CTU), code-decode block (CB), CU, PU, TU, PB, TB.
[0205] The term "motion vector" or "block vector" can refer to a vector of horizontal and vertical displacement between a position of a reference block and a position of a current block. The reference block can be a video unit in a reference picture in the RPL list. The reference block can also be a video unit in the current picture.
[0206] The term "LM" can refer to any linear regression based method, such as CCLM, MMLM, CCCM, GL-CCCM, non-down-sampled CCCM, GLM, GLM with luma values, etc. It can also be referred to as the term "cross component prediction (CCP)".
[0207] The term "CCLM" can refer to a single model LM mode, which can be single model CCLM, single model CCCM, single model GL-CCCM, non-down-sampled single model CCCM, single model GLM, single model GLM with luma values, multi-model CCLM, MMLM, multi-model CCCM, multi-model GL-CCCM, non-down-sampled multi-model CCCM, multi-model GLM, multi-model GLM with luma values, etc.
[0208] The term "MMLM" can refer to a multi-model LM mode, which can be multi-model CCLM, MMLM, multi-model CCCM, multi-model GL-CCCM, non-down-sampled multi-model CCCM, multi-model GLM, multi-model GLM with luma values, etc.
[0209] The term "CCCM" can refer to a regular CCCM mode, or GL-CCCM mode, or non-down-sampled CCCM, CCRM, etc.
[0210] The term "GL-CCCM" can refer to a CCCM mode that takes into account the gradient and position of the involved samples.
[0211] The term "non-down-sampled CCCM" can refer to a CCCM mode that takes into account non-down-sampled luma samples.
[0212] The term "CCRM" can refer to cross component model based residual coding or derivation. It can also imply cross component model based inter / IBC prediction, such as inter / IBC CCCM. It can also imply cross component model based intra prediction, such as intra CCCM.
[0213] In this document, cross component prediction (CCP) can refer to any cross component prediction method, such as any kind of CCLM / CCCM / GLM / GL-CCCM.
[0214] Note that the following terms are not limited to the specific terms defined in the existing standards. Any changes to the coding tools apply.
[0215] 1) The residual (and / or prediction) of a chroma block can be derived based on a cross-component model.
[0216] a. For example, the cross-component model can be a specific extrapolation filter (e.g., EIP, etc.).
[0217] b. For example, the cross-component model can be a specific interpolation filter (e.g., GLM, etc.).
[0218] c. For example, the cross-component model can be a specific convolution filter (CCCM, GL-CCCM, non-down-sampling CCCM, CCRM, inter-CCCM, intra-CCCM, etc.).
[0219] d. For example, the cross-component model can be a specific linear filter (e.g., CCLM, MMLM, etc.).
[0220] 2) The cross-component model (e.g., CCRM) used for residual coding can not contain a non-linear term.
[0221] a. For example, the cross-component model used for residual coding can contain a linear term and / or a bias term, but not a non-linear term.
[0222] 3) The CCRM can be used for intra or IBC blocks.
[0223] a. For example, it can be used for intra or IBC blocks in intra (such as I) slices.
[0224] b. For example, it can be used for intra or IBC blocks in inter (such as B or P) slices.
[0225] c. For example, in addition, it can be used for single tree.
[0226] d. For example, in addition, it can be used for dual tree.
[0227] e. For example, in single tree I slice, both luma and chroma are coded by IBC (or intraTMP), the CCRM can be generated based on the reconstructed luma and chroma samples within the reference block retrieved / guided by the block vector, and the residual model is applied to estimate the reconstructed values of the chroma samples in the current block.
[0228] f. For example, in dual tree, luma is coded by IBC (or intraTMP) while chroma is coded by intra, the CCRM can be generated based on the reconstructed luma samples within the reference luma block retrieved / guided by the block vector, and the reconstructed chroma samples co-located (e.g., in the same position) with the luma block, and the residual model is applied to estimate the reconstructed values of the chroma samples in the current block.
[0229] 4) CCRM can be used for DBV coded chroma blocks.
[0230] a. For example, based on the block vector of a DBV coded chroma block, the reference chroma block and its collocated luma block can be identified. These samples can be used as training samples for CCRM model computation.
[0231] b. For example, the derived CCRM model is applied to the reconstructed luma signal of the DBV chroma block to generate the final chroma prediction.
[0232] 5) CCRM model can be generated based on the correlation between luma and chroma reconstructed values from neighboring samples adjacent / non-adjacent to the current block.
[0233] a. For example, alternatively, CCRM model can be generated based on the correlation between luma and chroma reconstructed values in a reference block in the reference picture.
[0234] b. For example, alternatively, CCRM model can be generated based on the correlation between luma and chroma reconstructed values in a reference block in the current picture.
[0235] 6) For example, CCCM for intra prediction and CCRM (e.g., CCRM) for inter prediction can share the same logic.
[0236] a. For example, both of them can follow the same logic to obtain training samples.
[0237] b. For example, both of them can follow the same logic to determine training region.
[0238] 7) CCRM model can be generated based on non-downsampled luma samples.
[0239] a. For example, CCRM model coefficients can be solved based on non-downsampled luma samples as the training samples of the reference region.
[0240] b. For example, CCRM model can be applied to chroma blocks, where the chroma prediction of the current chroma block is generated based on non-downsampled luma samples of the collocated luma block.
[0241] 8) More than one CCRM model (e.g., MM-CCRM) for a block can be generated.
[0242] a. For example, the training samples of the CCRM can be divided into more than one class (e.g., two classes), and each set of samples can contribute to a unique model. In this way, multiple models can be generated, each model having its own filter coefficients. Each derived filter is applied to its corresponding set of luma reconstructed signals to produce the final prediction values for the current chroma samples belonging to the corresponding class.
[0243] i. For example, according to a multi-model CCRM (e.g., MM-CCRM) mode, the training sample pairs (e.g., these training samples are in the reference frame) on luma and chroma sample pairs of the reference block can be divided into more than one class.
[0244] ii. For example, alternatively, according to a multi-model CCRM (e.g., MM- CCRM) mode, the training sample pairs (e.g., these training samples are in the reference frame) on luma and chroma sample pairs of neighboring samples adjacent / non-adjacent to the reference block can be divided into more than one class.
[0245] iii. For example, alternatively, according to a multi-model CCRM (e.g., MM- CCRM) mode, the training sample pairs (e.g., these training samples are in the current frame) on luma and chroma sample pairs of neighboring samples adjacent / non-adjacent to the current video unit can be divided into more than one class.
[0246] iv. For example, in addition, by following the same criteria (e.g., by a threshold), the luma samples in the current video unit are divided into more than one group, and for the luma samples belonging to each class, a corresponding model can be applied to generate the model estimated chroma samples belonging to the group.
[0247] b. For example, multiple sets of training samples can be utilized to derive multiple models.
[0248] i. In one example, there are two sets, where the distance between the training samples and the current samples is different.
[0249] c. For example, the threshold (e.g., class threshold) to separate the samples into different classes can depend on the values of the samples within or neighboring the training region.
[0250] i. For example, the training region can be a reference block of the current video unit (e.g., these training samples are in the reference frame).
[0251] 1. For example, the reference block can be derived based on a block vector.
[0252] 2. For example, the reference block can be derived based on a motion vector.
[0253] ii. For example, the threshold value can be derived based on samples neighboring / non- neighboring the reference block of the current video unit (e.g., these training samples are in the reference frame).
[0254] iii. For example, the threshold value can be derived based on samples neighboring / non- neighboring the current video unit (e.g., these training samples are in the current frame).
[0255] iv. For example, the threshold value can be derived based on an average / median / intermediate operation on more than one sample within or neighboring the training region.
[0256] v. For example, the category threshold value can be derived based on non-downsampled luma sample values.
[0257] 1. Alternatively, the category threshold value can be derived based on downsampled luma sample values.
[0258] a. For example, a K-tap (such as K = 6) downsampling filter can be used to downsize K surrounding luma samples to one downsampled luma sample value.
[0259] vi. For example, the category threshold value can be derived based on an offset removal scheme.
[0260] 1. For example, the offset can be derived based on a luma sample located at a fixed position (such as top-left or center, etc.) in the reference video unit.
[0261] 2. For example, the offset value for category threshold derivation and CCRM model calculation can be the same.
[0262] vii. For example, the category threshold value can be derived at sub-block level.
[0263] viii. For example, the category threshold value can be derived at CU / PU / TU level.
[0264] ix. For example, the category threshold value can be calculated based on (downsampled or non-downsampled) luma prediction samples.
[0265] x. For example, the category threshold value can be derived based on luma residual sample values.
[0266] 1. For example, prediction samples of a second video unit (e.g., sub-block) that do not have non-zero residual can not be counted in the calculation of the category threshold value for a first video unit.
[0267] a. For example, the second video unit can be a subset of the first video unit.
[0268] b. For example, the second video unit can be equal to the first video unit.
[0269] d. For example, MM-CCRM can be applied on a sub-block level.
[0270] i. For example, the sub-block size can be pre-defined.
[0271] 1. For example, the pre-defined sub-block size can be 16x16 or 32x32, etc.
[0272] 2. For example, a pre-defined rule can be used to determine the sub-block size of MM-CCRM for a particular video unit.
[0273] a. For example, the sub-block size can be adaptive to the block dimension (width and / or height) of the current video block.
[0274] b. For example, a minimum number of chroma samples can be guaranteed for a sub-block of a video unit that is MM-CCRM coded.
[0275] ii. For example, if a video unit is larger than a pre-defined sub-block size, the video unit can be divided into more than one sub-block and MM-CCRM is performed.
[0276] iii. For example, at least one sub-block of a video unit can have more than one CCRM model.
[0277] iv. For example, each sub-block (and its associated training region) can have its own class threshold.
[0278] 1. For example, the class threshold of a particular sub-block can be calculated based on the training sample values belonging to such sub-block.
[0279] a. For example, the luma training samples in the reference block can be used to calculate the class threshold.
[0280] v. For example, all sub-blocks (and their associated training regions) can share the same class threshold.
[0281] 1. For example, one class threshold can be calculated and used for all sub-blocks.
[0282] 2. For example, the class threshold of all sub-blocks in the current video unit can be calculated based on the training sample values of the current video unit.
[0283] 3. For example, the class threshold of all applicable sub-blocks in the current video unit can be calculated based on the training sample values of the current video unit.
[0284] a. For example, sub-blocks that do not contain non-zero residuals can not be counted.
[0285] vi. For example, each sub-block of a video unit can have its own training samples, and the training samples of a particular sub-block can be partitioned into more than one class.
[0286] 1. For example, the training samples in a reference video unit of a reference picture can be classified based on sub-blocks.
[0287] vii. For example, the training samples from a current picture can be classified into more than one group, but can not be partitioned into sub-blocks.
[0288] e. For example, MM-CCRM can be applied based on TU level (or PU / CU level).
[0289] i. For example, MM-CCRM can be applied based on TU / CU / PU (e.g., for the application of MM-CCRM, the TU / CU / PU can not be partitioned into sub-blocks).
[0290] ii. For example, whether to use TU / PU / CU based multi-model CCRM (i.e., MM-CCRM) can be determined at TU / PU / CU level.
[0291] 1. For example, a video unit (e.g., TU / PU / CU) can choose to use sub-block based single model CCRM or TU / PU / CU based MM-CCRM.
[0292] a. For example, the decision can be made at TU / PU / CU level.
[0293] f. For example, whether and / or how to apply MM-CCRM (and / or CCRM) can be derived based on codec information at both encoder side and decoder side (e.g., without signaling).
[0294] i. In one example, it can be derived instantaneously, e.g., using information of previously coded / reconstructed samples.
[0295] ii. For example, the determination of whether to use sub-block based CCRM or TU / CU / PU level CCRM can be implicitly derived based on codec information (e.g., without signaling).
[0296] iii. For example, the determination of whether to use M1xM2 sub-block based CCRM or N1xN2 sub-block based CCRM can be implicitly derived based on codec information (e.g., without signaling).
[0297] 1. For example, M1 = 16 or 8 or 32 or TU / CU / PU.
[0298] 2. For example, M2 = 16 or 8 or 32 or TU / CU / PU.
[0299] 3. For example, N1 = 16 or 8 or 32 or TU / CU / PU.
[0300] 4. For example, N2 = 16 or 8 or 32 or TU / CU / PU.
[0301] 5. For example, M1!= N1 and / or M2!= N2.
[0302] iv. For example, the determination of whether to use sub-block based MM-CCRM or TU / CU / PU level MM-CCRM can be implicitly derived based on coding information (e.g., without signaling).
[0303] v. For example, the determination of whether to use M1 x M2 sub-block based MM-CCRM or N1 x N2 sub-block based MM-CCRM can be implicitly derived based on coding information (e.g., without signaling).
[0304] 1. For example, M1 = 16 or 8 or 32 or TU / CU / PU.
[0305] 2. For example, M2 = 16 or 8 or 32 or TU / CU / PU.
[0306] 3. For example, N1 = 16 or 8 or 32 or TU / CU / PU.
[0307] 4. For example, N2 = 16 or 8 or 32 or TU / CU / PU.
[0308] 5. For example, M1!= N1 and / or M2!= N2.
[0309] vi. For example, the determination of whether to use single model CCRM or MM-CCRM can be implicitly derived based on coding information (e.g., without signaling).
[0310] vii. For example, the determination can be according to a method based on decoder-derived cost.
[0311] 1. For example, the decoder-derived cost can be calculated based on minimizing the SAD / SATD / SSE / MSE between model-estimated sample values and true reconstructed sample values, where the sample can refer to at least one of the training samples.
[0312] 2. For example, the method with lower cost can be selected as the final method applied to the current video unit.
[0313] viii. For example, the determination can be based on reference picture information.
[0314] 1. For example, the determination can be based on the POC distance of the current picture and its reference picture.
[0315] 2. For example, the determination can be based on reference index.
[0316] g. For example, whether and / or how to apply MM-CCRM (and / or CCRM) can be signaled in the bitstream.
[0317] i. For example, a syntax element (e.g., flag, index, etc.) can be signaled based on whether the current block is CCRM coded or not.
[0318] 1. For example, if the video unit is CCRM coded, a syntax element (e.g., flag, index, etc.) can be further signaled to indicate whether it is MM-CCRM.
[0319] ii. For example, a syntax element (e.g., flag, index, etc.) can be signaled to indicate whether it is subblock-based MM-CCRM or TU / CP / PU-based MM-CCRM.
[0320] iii. For example, a syntax element (e.g., flag, index, etc.) can be signaled to indicate whether it is subblock-based CCRM or TU / CP / PU-based CCRM.
[0321] iv. For example, a syntax element can be conditionally signaled based on block dimension (width and / or height).
[0322] h. For example, a block restriction can be applied to indicate the allowance of MM-CCRM mode.
[0323] i. In one example, assuming the width and height of a chroma CU / PU / TU are denoted as W and H, MM-CCRM can be allowed in case at least one of the following conditions is met: 1. W H>T0 or W H>= T0 (e.g., T0 = 16 or 32 or 64 or 128).
[0324] 2. W>T1, or, W>= T1.
[0325] 3. H>T2, or, H>= T2.
[0326] 4. Min (W,H)>T3, or, Min (W,H)>= T3.
[0327] 5. Max (W,H)<T4, or, Max (W,H)<= T4.
[0328] 6. W<T5 H, or, W<= T5 H.
[0329] 7. W>T6 H, or, W>= T6 H.
[0330] 8. H<T7 W, or, H<= T7 W.
[0331] 9. H>T8 W, or, H>= T8 W.
[0332] 10. W H<T9, or W H<= T9.
[0333] ii. In one example, blocks enabled for a particular tool (e.g., affine motion compensation enabled) can be prohibited from MM-CCRM.
[0334] i. For example, CCRM coded video units can always use multi-model CCRM.
[0335] i. Alternatively, CCRM coded video units can use single model CCRM or multi-model CCRM.
[0336] 9) Chroma Cb and Cr can share one CCRM.
[0337] a. Alternatively, chroma Cb and Cr can build their own CCRM.
[0338] 10) For CCRM model filter design, sample values and / or gradients and / or location information can be considered.
[0339] a. For example, at least one K-tap filter can be used for CCRM model, which consists of K1 sample terms, K2 gradient terms, K3 positioning / location terms, K4 non-linear terms, K5 bias terms, etc.
[0340] i. For example, K1 = 0 or 1 or 2 or 5 or 6.
[0341] ii. For example, K2 = 0 or 1 or 2 or 4.
[0342] iii. For example, K3 = 0 or 1 or 2 or 4.
[0343] iv. For example, K4 = 0 or 1 or 2 or 4.
[0344] v. For example, K5 = 0 or 1.
[0345] vi. For example, K = K1 + K2 + K3 + K4 + K5.
[0346] vii. For example, the sample term can be calculated based on the luma sample value.
[0347] viii. For example, the gradient term can be calculated based on more than one sample neighboring the particular luma sample.
[0348] ix. For example, the positioning / location term can be calculated based on the horizontal coordinate and / or the vertical coordinate of the particular luma sample, wherein the coordinates can be relative to the top-left position of the particular reference region.
[0349] x. For example, the non-linear term can be the square of a particular value (e.g., a bit-depth dependent intermediate value such as 512 or 256, or a particular luma value).
[0350] xi. For example, the non-linear term can be the square of a gradient value based on the particular gradient term.
[0351] xii. For example, the offset can be subtracted from the terms of the K-tap filter.
[0352] 1. For example, the offset can be derived based on a pre-defined rule (such as the value of the top-left training sample in the training region, or the average / intermediate value of more than one sample in the training region).
[0353] xiii. For example, the coefficients of the K-tap filter can be solved by a Gaussian elimination solver.
[0354] xiv. For example, the coefficients of the K-tap filter can be solved by an LDL decomposition method.
[0355] i. For example, the coefficients of the K-tap filter can be solved by linear regression.
[0356] ii. For example, the coefficients of the K-tap filter can be solved by linear equations.
[0357] b. For example, more than one filter can be used, and the final prediction can be derived based on fusing together the filtered outputs of the multiple filters.
[0358] i. For example, the weights of fusing the multiple filtered values can be solved by a Gaussian elimination solver.
[0359] ii. For example, the weights that fuse multiple filter values can be solved by an LDL decomposition method.
[0360] 11) For example, more than one filter can be allowed for a CCRM coded video unit, and which filter is finally selected can be signaled or partitioned.
[0361] a. For example, a syntax element can be signaled to indicate which filter (e.g., CCLM or CCCM) is used for CCRM mode.
[0362] b. For example, indicating which filter (e.g., CCLM or CCCM) is used for CCRM mode can be determined based on template costs from both the encoder and the decoder.
[0363] c. For example, indicating which filter (e.g., CCLM or CCCM) is used for CCRM mode can be determined based on decoder-derived costs from both the encoder and the decoder.
[0364] 12) The filter output can be clipped to a certain value.
[0365] a. For example, it can be clipped based on the reconstructed values in the training region.
[0366] i. For example, the training region can be derived based on the block vector (or motion vector).
[0367] ii. For example, the training region can be adjacent to the current block.
[0368] iii. For example, the training region can be the reference region of the current block.
[0369] iv. For example, the filter output can be clipped within the minimum and maximum of the reconstructed (or predicted) luma sample values in the training region.
[0370] b. For example, it can be clipped based on the reconstructed (or predicted) values in the co-located luma block of the current chroma block.
[0371] i. For example, it can be clipped within the minimum and maximum of the current block luma reconstructed (or predicted) values.
[0372] c. For example, if the value is outside the valid range, it can be ignored / dropped / not used.
[0373] 13) The CCRM parameters can be stored in a buffer and used for coding of future blocks.
[0374] a. For example, CCRM parameters for a video unit (e.g., CU, PU, color component, Cb, Cr, etc.) can include model type, model coefficients, whether it is a single model or multiple models, threshold to separate samples into multiple models, etc.
[0375] b. For example, it can be stored in a local cache for coding of future blocks in the current picture.
[0376] c. For example, it can be stored in a temporal / picture / frame buffer for coding of future blocks in future decoded pictures.
[0377] i. For example, CCRM parameters of a current frame / picture can be stored, which can be referenced for the CCP process of future frames / pictures.
[0378] ii. For example, it can be stored in association with motion and mode information of a video unit.
[0379] 14) A video block can inherit model parameters from a previous CCRM coded block.
[0380] a. For example, a video block can be coded by a CCM inheritance mode.
[0381] b. For example, a video block can be coded by a CCMerge (e.g., CCMerge) mode.
[0382] c. For example, model parameters of a previous CCRM coded block can be stored in a buffer (e.g., local buffer, picture buffer, temporal buffer, history-based LUT, etc.).
[0383] d. In one example, the parameters can refer to filter information, linear or non-linear parameters of a model, model index, etc.
[0384] 15) A final prediction of a block can be generated based on multiple prediction candidates from different CCRM.
[0385] a. For example, more than one CCRM prediction can be fused together.
[0386] b. For example, weights / coefficients of different fusion terms can be solved based on a Gaussian elimination method.
[0387] c. For example, weights / coefficients of different fusion terms can be solved based on an LDL decomposition method.
[0388] d. For example, a bias term can be involved for fusion.
[0389] e. For example, a non-linear term can be involved for fusion.
[0390] 16) The enabling of CCRM mode can depend on at least one of the following aspects: a. The prediction mode of the video unit (e.g., MODE_INTRA, MODE_INTER, MODE_IBC, MODE_PLT, etc.).
[0391] b. The transform type of the video unit (e.g., ACT, color transform, etc.).
[0392] c. The number of non-zero coefficients of the video unit.
[0393] d. The partition tree type (e.g., single tree, dual tree).
[0394] e. The slice type (e.g., I, B, P slice).
[0395] f. The color format (e.g., whether 4:0:0).
[0396] g. The availability of chroma components.
[0397] h. For example, CCRM can not be enabled for ACT and / or 4:0:0 color format.
[0398] 17) The disclosed CCRM mode can be based on one of the following filters: a. CCLM and / or its variants.
[0399] b. MMLM and / or its variants.
[0400] c. CCCM and / or its variants (e.g., GL-CCCM, non-down-sampled CCCM, BVG- CCCCM, inter-frame CCCM, intra-frame CCCM, etc.).
[0401] d. GLM and / or its variants.
[0402] e. Any cross-component prediction that uses information in one channel / component to predict information in another channel / component.
[0403] f. Any filter-based prediction where the filter coefficients are solved based on the correlation between the prediction and / or the reconstructed information.
[0404] 18) Block restrictions can be applied to limit the application of certain types of CCP modes.
[0405] a. For example, CCP modes can only be allowed to be used for cases where the block size satisfies a pre-defined rule.
[0406] b. For example, the syntax element can be signaled only if the CCP mode applies.
[0407] c. For example, if the CCP mode is not allowed to be used, the syntax element can be inferred to a particular value indicating that no such CCP mode is used for such a block.
[0408] d. For example, at least one of the following block restrictions can be applied to the CCRM mode (assuming W denotes the block width and H denotes the block height): i. W < T1, or, W <= T1.
[0409] ii. H < T2, or, H <= T2.
[0410] iii. Min (W, H) > T3, or, Min (W, H) >= T3.
[0411] iv. Max (W, H) < T4, or, Max (W, H) <= T4.
[0412] v. W < T5 H, or, W <= T5 H.
[0413] vi. W > T6 H, or, W >= T6 H.
[0414] vii. H < T7 W, or, H <= T7 W.
[0415] viii. H > T8 W, or, H >= T8 W.
[0416] ix. W H < T9, or W H <= T9.
[0417] x. For example, T1, T2,... T9 can be predefined integer constants.
[0418] e. For example, the CCRM mode can be allowed only for small blocks.
[0419] i. For example, it can be allowed for blocks smaller than 4x4 or 8x8 or 16x16 or 32x32.
[0420] ii. For example, it can be allowed for blocks with a number of samples smaller than 32 or 64 or 128.
[0421] iii. For example, it can be used for blocks with fewer than 32, 64, or 128 samples.
[0422] iv. For example, it may not be allowed for 2xN blocks, where N can be greater than 4, 8, or 16.
[0423] v. For example, it may not be allowed for Nx2 blocks, where N can be greater than 4, 8, or 16.
[0424] 19) The disclosed method can be used in a single tree.
[0425] 20) The disclosed method can be used in two trees.
[0426] 21) The disclosed method can be used in inter-frame (such as B or P) stripes.
[0427] 22) The disclosed method can be used in intra-frame (such as I) stripes.
[0428] 23) The “block vector” in the disclosed method can be a “motion vector”.
[0429] 24) The training / reference samples in the disclosed method may refer to the predicted samples and / or reconstructed samples in the training / reference region.
[0430] 25) Whether and / or how the methods disclosed above can be applied can be transmitted via signaling at the sequence level / picture group level / picture level / strip level / piece group level, such as in the sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / strip header / piece group header.
[0431] 26) Whether and / or how the methods disclosed above can be applied to transmit signals at PB / TB / CB / PU / TU / CU / VPDU / CTU / CTU lines / strips / films / sub-images / other types of areas containing more than one sample point or pixel.
[0432] 27) Whether and / or how the methods disclosed above are applied may depend on the encoded / decoded information, such as block size, color format, single / double tree segmentation, color components, and stripe / picture type.
[0433] 3 questions There are several problems with existing video encoding and decoding technologies, and these problems will be further improved in order to achieve higher encoding and decoding gains.
[0434] 1. Inter-frame prediction blocks can be generated based on applying filters to reference samples.
[0435] 2. The filter used for IBC / intraTMP prediction can be improved.
[0436] 3. Linear / nonlinear filter models can be applied on the basis of CU / TU / PU or on the basis of sub-blocks.
[0437] a. Furthermore, when it is based on sub-blocks, another filter can be applied to mitigate discontinuities between sub-blocks.
[0438] b. Furthermore, the deblocking process can be modified based on a sub-block-based filter model.
[0439] 4. MHP can be improved by considering intra-frame / IBC and solver-based hybrid / fusion weight derivations.
[0440] 4 Detailed Solutions The detailed embodiments described below should be considered as examples for explaining general concepts. These embodiments should not be interpreted in a narrow sense. Furthermore, these embodiments can be combined in any way.
[0441] The term "video unit" or "code-decoder unit" can refer to a picture, strip, slice, code-decoder tree block (CTB), code-decoder tree unit (CTU), code-decoder block (CB), CU, PU, TU, PB, TB.
[0442] The term "block" can refer to code-decode tree block (CTB), code-decode tree unit (CTU), code-decode block (CB), CU, PU, TU, PB, TB.
[0443] The term "non-intra-frame" can refer to inter-frame, IBC, PLT, etc. It can also be inferred to be intraTMP.
[0444] The terms "motion vector" or "block vector" can refer to the vector of horizontal and vertical displacements between the position of a reference block and the position of the current block. The reference block can be a video unit in a reference image within the RPL list. Alternatively, the reference block can be a video unit in the current image.
[0445] The term "LM" can refer to any linear regression-based method, such as CCLM, MMLM, CCCM, GL-CCCM, CCCM without downsampling, GLM, GLM with luminance values, etc. It can also be referred to as the term "Cross-Component Prediction (CCP)".
[0446] The term "CCLM" can refer to a single-model LM mode, which can be a single-model CCLM, a single-model CCCM, a single-model GL-CCCM, a single-model CCCM without subsampling, a single-model GLM, a single-model GLM with luminance values, a multi-model CCLM, a MMLM, a multi-model CCCM, a multi-model GL-CCCM, a multi-model CCCM without subsampling, a multi-model GLM, a multi-model GLM with luminance values, etc.
[0447] The term "MMLM" can refer to a multi-model LM mode, which can be multi-model CCLM, MMLM, multi-model CCCM, multi-model GL-CCCM, multi-model CCCM without downsampling, multi-model GLM, multi-model GLM with luminance values, etc.
[0448] The term "CCCM" can refer to regular CCCM mode, GL-CCCM mode, CCCM without downsampling, CCRM, etc.
[0449] The term "GL-CCCM" can refer to a CCCM mode that takes into account the gradient and location of the samples involved.
[0450] The term "CCCM without downsampling" can refer to a CCCM mode that takes into account unsampled luminance samples.
[0451] The term "CCRM" can refer to residual derivation based on a cross-component model. It can also presuppose inter-frame / IBC coding / decoding based on a CCCM model (such as inter-frame / IBC CCCM). It can also presuppose intra-frame coding / decoding based on a CCCM model (such as intra-frame CCCM).
[0452] In this document, cross-component prediction (CCP) can refer to any cross-component prediction method, such as any kind of CCLM / CCCM / GLM / GL-CCCM.
[0453] It should be noted that the following terms are not limited to the specific terms defined in existing standards. Any changes to encoding / decoding tools also apply.
[0454] Figure 30 An example is shown where the filter is modulated by the correlation between the current template (filled with a grid) and a reference template (filled with diagonal lines) in a reference image. Figure 31 An example is shown where the filter is modulated by the correlation between the current template (filled with a grid) and a reference template (filled with diagonal lines) in the current image.
[0455] 1) Inter-frame prediction blocks can be generated based on reference samples in a reference image by applying filters.
[0456] a. For example, filter coefficients can be solved by minimizing the difference between the reference template and the current template, for example, such as... Figure 30 As shown.
[0457] i. For example, the current template may include sample points above and / or to the left of the current block.
[0458] ii. For example, the reference template may include sample points above and / or to the left of the reference block of the current block guided by the motion vector.
[0459] 1. In addition, for example, reference templates can be found within reference images.
[0460] b. For example, the training samples used to solve for filter coefficients may include samples with up to M rows and / or N columns.
[0461] i. For example, M = 2, 4, 6, or 12.
[0462] ii. For example, N = 2, 4, 6, or 12.
[0463] iii. For example, training samples may include samples to the left of the current block / reference block.
[0464] iv. For example, training samples may include samples above the current block / reference block.
[0465] c. For example, a template may include samples with up to X rows and / or Y columns.
[0466] i. For example, X = 1, 2, or 3.
[0467] ii. For example, Y = 1, 2, or 3.
[0468] iii. For example, the template may include samples to the left of the current block / reference block.
[0469] iv. For example, the template may include sample points above the current block / reference block.
[0470] d. For example, the current prediction block can be derived based on applying the filter to the reference block.
[0471] 2) IBC or intraTMP prediction blocks can be generated based on reference samples in the current image by applying filters.
[0472] a. For example, filter coefficients can be solved by minimizing the difference between the reference template and the current template, for example, such as... Figure 31 As shown.
[0473] i. For example, the current template may include sample points above and / or to the left of the current block.
[0474] ii. For example, the reference template may include sample points above and / or to the left of the reference block of the current block guided by the block vector.
[0475] 1. In addition, for example, reference templates can be used within the current image.
[0476] b. For example, the training samples used to solve for filter coefficients may include samples with up to M rows and / or N columns.
[0477] i. For example, M = 2, 4, 6, or 12.
[0478] ii. For example, N = 2, 4, 6, or 12.
[0479] iii. For example, training samples may include samples to the left of the current block / reference block.
[0480] iv. For example, training samples may include samples above the current block / reference block.
[0481] c. For example, a template may include samples with up to X rows and / or Y columns.
[0482] i. For example, X = 1, 2, or 3.
[0483] ii. For example, Y = 1, 2, or 3.
[0484] iii. For example, the template may include samples to the left of the current block / reference block.
[0485] iv. For example, the template may include sample points above the current block / reference block.
[0486] d. For example, the current prediction block can be derived based on applying the filter to the reference block.
[0487] 3) Filter-based non-intra-frame prediction can be applied to luminance samples.
[0488] a. For example, the filter model derivation can be based on a reference luminance template and the current luminance template, and the current luminance prediction can be derived by applying the filter model to the reference luminance block.
[0489] b. For example, both filter derivation and filter application can be based on a color component / channel (such as luminance) information.
[0490] 4) Filter-based non-intra-frame prediction can be applied to chroma samples.
[0491] a. For example, the filter model derivation can be based on a reference chromaticity template and the current chromaticity template, and the current chromaticity prediction can be derived by applying the filter model to the reference chromaticity block.
[0492] b. For example, both filter derivation and filter application can be based on a color component / channel (such as chroma or Cb or Cr) information.
[0493] 5) Different filter coefficients can be calculated separately for luminance and chrominance samples.
[0494] a. For example, luminance and chrominance can share the same filter shape.
[0495] b. For example, luminance and chrominance can use separate filter coefficients.
[0496] 6) For the application of the filter that derives the current prediction, samples within and / or adjacent to the reference block can be considered.
[0497] a. For example, samples adjacent to the reference block and samples within the reference block can both be used as filter inputs.
[0498] i. For example, such a process can be needed to generate boundary samples for the current block.
[0499] b. For example, samples within only the reference block can be used as filter inputs.
[0500] i. For example, such a process may be needed to generate internal samples of the current block.
[0501] c. For example, samples to the right and / or below the reference block can be filled before being fed to the filter.
[0502] i. For example, it can be copied from the nearest available sample point within the reference block.
[0503] ii. For example, it can be mirrored from available sample points within a reference block.
[0504] d. For example, reconstructed samples that are adjacent to and / or within the reference block can be used.
[0505] e. For example, prediction samples that are adjacent to and / or within the reference block can be used.
[0506] f. For example, whether to use referenced reconstructed or predicted samples can depend on whether the reference block is BV-guided or MV-guided.
[0507] g. For example, whether to use referenced reconstructed or predicted samples can depend on whether the sample to be computed is at the boundary of the current block or inside the current block.
[0508] 7) Sample values and / or gradients and / or location information can be considered for filter design.
[0509] a) For example, linear filters can be used.
[0510] b) For example, nonlinear filters can be used.
[0511] c) For example, convolutional filters can be used.
[0512] d) For example, extrapolation filters can be used.
[0513] e) For example, at least one K-tap filter can be used, which consists of K1 sample terms, K2 gradient terms, K3 localization / position terms, K4 nonlinear terms, K5 bias terms, etc.
[0514] i. For example, K1 = 0 or 1 or 2 or 5 or 6.
[0515] ii. For example, K2 = 0, 1, 2, or 4.
[0516] iii. For example, K3 = 0, 1, 2, or 4.
[0517] iv. For example, K4 = 0, 1, 2, or 4.
[0518] v. For example, K5 = 0 or 1.
[0519] vi. For example, K = K1 + K2 + K3 + K4 + K5.
[0520] vii. For example, sample items can be calculated based on reference sample values.
[0521] viii. For example, gradient terms can be computed based on more than one sample point adjacent to a particular reference sample point.
[0522] ix. For example, the location / position item can be calculated based on the horizontal and / or vertical coordinates of a specific reference point, where the coordinates can be relative to the upper left position of a specific reference area.
[0523] x. For example, at least one item can correspond to sample values of different color components.
[0524] 1. For example, at least one item may be the sample value of the corresponding luminance sample (which may be downsampled) of the chrominance sample.
[0525] xi. For example, the nonlinear term can be the square of a specific value (e.g., an intermediate value related to bit depth, such as 512 or 256, or a specific reference value).
[0526] xii. For example, a nonlinear term can be the square of the gradient value based on a particular gradient term.
[0527] xiii. For example, it may not contain nonlinear terms.
[0528] 1. For example, it may contain linear terms and / or bias terms, but not nonlinear terms.
[0529] xiv. For example, the offset can be subtracted from the terms of the K-tap filter.
[0530] 1. For example, the offset can be derived based on predefined rules (such as the value of the top left training sample in the training region, or the average / median value of more than one sample in the training region).
[0531] xv. For example, the coefficients of a K-tap filter can be solved using a Gaussian elimination solver.
[0532] xvi. For example, the coefficients of a K-tap filter can be solved using the LDL decomposition method.
[0533] iii. For example, the coefficients of a K-tap filter can be solved by linear regression.
[0534] iv. For example, the coefficients of a K-tap filter can be solved using linear equations.
[0535] 8) For non-intra-frame blocks, more than one filter model can be generated, and the final prediction can be derived based on applying different models to samples of different classes.
[0536] f) For example, training samples can be divided into more than one class (e.g., two classes), and each set of samples can contribute to a unique filter model. In this way, multiple filter models can be generated, each with its own filter coefficients. Each derived filter is applied to its corresponding set of reference signals to produce a final predicted value for the current sample belonging to the corresponding class.
[0537] g) For example, the threshold for separating samples into different categories can depend on the values of the samples in the training region.
[0538] i. For example, the training region can contain samples that are adjacent to the reference block of the current block.
[0539] 1. For example, a reference block can be derived based on a block vector.
[0540] 2. For example, a reference block can be derived based on motion vectors.
[0541] ii. For example, the threshold can be derived based on the average / median / intermediate operation of more than one sample in the training region.
[0542] 9) Filter-based non-intra-frame (or intra-frame) prediction can be performed based on TU / CU / PU.
[0543] h) For example, TU / CU / PU may not be divided into sub-blocks for filter applications (e.g., M×M sub-blocks, such as M=16).
[0544] i) For example, for TU / CU / PU, the (linear / nonlinear) filter model can be modulated using a set of training samples about the entire TU / CU / PU.
[0545] i. For example, a set of model coefficients for a filter can be solved, and such a filter can be applied to estimate the predicted / residual samples of TU / CU / PU.
[0546] 10) For example, CCRM (or inter-frame / IBC CCCM) can be applied at the TU / PU / CU level (e.g., never dividing the TU / PU / CU into multiple sub-blocks).
[0547] j) For example, the model coefficients of CCRM (or inter-frame / IBC CCCM) filters can be solved based on training samples at the TU / PU / CU level, and such filters can be applied to estimate the predicted / residual samples of TU / PU / CU.
[0548] i. For example, all applicable samples in TU / PU / CU can be estimated from the same model with the same model coefficients.
[0549] k) For example, CCRM (or inter-frame / IBC CCCM) based on TU / PU / CU can be applied to all allowed block sizes.
[0550] i. For example, regardless of the size of the block (e.g., CU / PU / TU), as long as CCRM (or inter-frame / IBC CCCM) is allowed for such a block, the filter can be modulated at the block level without sub-block partitioning / segmentation.
[0551] ii. For example, model derivation can be applied to CU / PU / TU, and model application can be applied to all applicable samples in such CU / PU / TU.
[0552] 11) Filter-based non-intra-frame (or intra-frame) prediction can be performed on a sub-block basis.
[0553] l) For example, CU / PU / TU can be divided into more than one sub-block, and each sub-block can have its own filter model (e.g., coefficients, filter shape, filter terms, etc.).
[0554] For example, the size of a sub-block can be a fixed size (e.g., width and / or height and / or number of sample points).
[0555] n) For example, the size of the sub-block can be predefined (e.g., such as 16x16, or 8x8, or 32x32).
[0556] o) For example, the training samples for each sub-block can be different.
[0557] i. For example, the training samples for each sub-block can be derived based on the corresponding sub-block in the reference CU / PU / TU (e.g., the reference video unit can also be divided into sub-blocks for training sample derivation).
[0558] ii. For example, the training samples for each sub-block can be derived based on the partial samples that are adjacent to the current CU / PU / TU or the reference CU / PU / TU.
[0559] 1. For example, some sample points can correspond to each sub-block.
[0560] iii. Alternatively, the training samples for each sub-block can be derived based on neighboring samples that are adjacent or non-adjacent to the current CU / PU / TU or the reference CU / PU / TU.
[0561] p) For example, the CCCM (or its variant filter modes) for block vector guidance in IBC / intraTMP can be executed on a sub-block basis.
[0562] q) For example, CCCM / CCLM / intraTMP filters / IBC filters (or their variant filter modes) can be executed based on sub-blocks.
[0563] 12) Filter-based non-intra-frame (or intra-frame) prediction for video units can be performed based on adaptive sub-block granularity.
[0564] r) For example, a particular video unit can be divided into multiple sub-blocks, where all sub-blocks have the same size.
[0565] s) For example, different video units can have different sub-block size granularity.
[0566] For example, the width and / or height of sub-blocks can be defined for video units of different sizes.
[0567] i. Alternatively, the number of sub-blocks in width and / or height can be defined.
[0568] u) For example, the rules for obtaining the size of sub-blocks can be predefined.
[0569] i. For example, larger video units can have larger sub-block sizes, while smaller video units can be assigned smaller sub-block sizes.
[0570] v) For example, the rules for obtaining the sub-block size can be transmitted via signals in the bitstream.
[0571] i. For example, whether a video unit is divided into sub-blocks can be transmitted via signal in the bitstream.
[0572] ii. For example, whether the video unit is divided into X1-sized sub-blocks or X2-sized sub-blocks can be transmitted via signal in the bitstream.
[0573] w) For example, adaptive sub-block sizes can be allocated based on the width and / or height of the video unit's TU / PU / CU.
[0574] i. For example, it can depend on the width and / or height of the video unit.
[0575] ii. For example, it can depend on the total number of samples in the video unit.
[0576] 13) For example, whether and how the sub-block size can be determined can be derived based on encoding and decoding information at both the encoder and decoder sides (e.g., without signal transmission).
[0577] x) For example, the determination of whether to use a sub-block-based filter or a TU / CU / PU level filter can be implicitly derived based on the encoding / decoding information (e.g., without signal transmission).
[0578] y) For example, the determination of whether to use a filter based on M1xM2 sub-blocks or a filter based on N1xN2 sub-blocks can be implicitly derived based on encoding and decoding information (e.g., without signal transmission).
[0579] i. For example, M1 = 16 or 8 or 32 or TU / CU / PU.
[0580] ii. For example, M2 = 16 or 8 or 32 or TU / CU / PU.
[0581] iii. For example, N1 = 16 or 8 or 32 or TU / CU / PU.
[0582] iv. For example, N2 = 16 or 8 or 32 or TU / CU / PU.
[0583] v. For example, M1 != N1 and / or M2 != N2.
[0584] z) For example, determination can be based on a template cost-based approach.
[0585] i. For example, the template cost can be derived based on minimizing the filter model to estimate the SAD / SATD / SSE / MSE between the sample values and the true reconstructed sample values, where the sample can be referenced to at least one of the training samples.
[0586] ii. For example, a method with lower template cost can be chosen as the final method to be applied to the current video unit.
[0587] aa) For example, it can be determined based on information from a reference image.
[0588] i. For example, determination can be based on the POC distance between the current image and its reference image.
[0589] ii. For example, determination can be based on a reference index.
[0590] bb) Alternatively, for example, whether and how the sub-block size can be determined can be transmitted via signaling in the bitstream.
[0591] 14) For example, whether and how to apply sub-block-based modeling filters can depend on indicators from one of the following tools: cc) OBMC and / or its variants.
[0592] dd) LIC and / or its variants.
[0593] ee) IBC and / or palette and / or BDPCM.
[0594] ff) intraTMP and / or its variants.
[0595] gg) DMVR and / or its variants.
[0596] Inter-frame TM and / or its variants (hh)
[0597] ii) Affine and / or its variants.
[0598] jj) sbTMVP and / or its variants.
[0599] kk) ISP and / or its variants.
[0600] ll) GPM and / or its variants.
[0601] CIIP and / or its variants (mm)
[0602] nn) SGPM and / or its variants.
[0603] oo) SBT and / or its variants.
[0604] pp) Merge and / or its variants.
[0605] qq) AMVP and / or its variants.
[0606] 15) For example, the second filtering process can be applied to samples of video units processed by the first filter.
[0607] For example, the input to the second filter can be the output of the first filter.
[0608] i. For example, the input to the second filter can be samples filtered by the first filter.
[0609] For example, the first filter can be based on a linear / nonlinear model (e.g., CCRM, intra / inter / IBC / intraTMP CCCM, EIF, CCLM, etc.).
[0610] For example, the first filter can be applied at the sub-block level.
[0611] i. For example, TU / PU / CU can be divided into M×N (such as M=N=16) sub-blocks for filter processing.
[0612] For example, the second filter could be a smoothing filter and / or a boundary filter.
[0613] For example, the second filter can be applied to boundary samples located at the boundaries of sub-blocks.
[0614] i. For example, a K-tap filter can be applied to samples located in the first row and / or first column and / or last row and / or last column of each sub-block.
[0615] ii. For example, K = 2.
[0616] For example, the input to the second filter can be at least one sample point within the sub-block to be processed, and at least one sample point within a neighboring sub-block of the sub-block to be processed.
[0617] i. For example, the output of a sample filtered by a 2-tap boundary filter can be based on y = (a0) x0 + a1 x1 + offset) >> shift is derived, where y represents the output value, a0 and a1 are weights / coefficients, x0 and x1 are the input samples within and next to the sub-block to be processed, shift can be calculated based on the sum of a0 and a1, and offset can be a fixed value.
[0618] ii. For example, the weights / coefficients of the input samples of the second filter can be predefined.
[0619] iii. For example, the weight of input samples within the sub-block to be processed can be greater than the weight of input samples adjacent to the sub-block to be processed.
[0620] For example, the value of at least one sample point within a video unit can be modified through a second filtering process.
[0621] i. For example, the second filter can be applied to internal samples located inside a sub-block.
[0622] ii. For example, the values of boundary samples in the first row and / or first column and / or last row and / or last column of a sub-block can be modified by a second filtering process.
[0623] iii. For example, samples located in the first row and / or first column and / or last row and / or last column of the current video unit may not be modified by the second filtering process.
[0624] For example, whether a sample can be modified by a second filter can be determined based on which sub-block the sample belongs to.
[0625] i. For example, the first row samples of the first row sub-block may not be filtered by the second filter.
[0626] ii. For example, the first column samples of the first column sub-block may not be filtered by the second filter.
[0627] iii. For example, the last row of samples in the last sub-block may not be filtered by the second filter.
[0628] iv. For example, the last column of samples in the last sub-block may not be filtered by the second filter.
[0629] For example, whether a sample to be processed can be further modified by a second filter can be determined based on whether the input sample of the second filter was processed by the first filter.
[0630] i. For example, a second filter can only be applied to modify the value of the sample to be processed if all the input samples of the second filter have been processed by the first filter.
[0631] ii. For example, input samples may include the sample to be processed on one side of the edge / boundary and its (multiple) neighboring samples on the other side of the edge / boundary.
[0632] 1. For example, edges / boundaries can refer to sub-block edges / boundaries.
[0633] 16) For example, the deblocking process can be based on a specific linear / nonlinear filter model.
[0634] For example, whether and / or how to apply a deblocking filter to boundary / edge samples can be determined based on the filter model (e.g., CCRM, intra / inter / IBC / intraTMP CCCM, EIF, CCLM, etc.).
[0635] For example, the deblocking strength can be determined based on the filter model.
[0636] i. For example, whether to apply a strong deblocking filter or a weak deblocking filter to a specific edge / boundary can be determined based on the filter model.
[0637] ii. For example, whether to use a long deblocking filter or a short deblocking filter for a specific edge / boundary can be determined based on the filter model.
[0638] For example, the values of deblocking filter parameters (e.g., tC and / or beta) can be determined based on the filter model.
[0639] For example, it can be determined whether the prediction / residual of samples along the edge / boundary is based on a sub-block filter model.
[0640] For example, the filter model can be one of the following filters.
[0641] i. Sub-block based CCRM.
[0642] ii. Sub-block based intra-frame CCCM.
[0643] iii. Sub-block based IBC filter / CCCM.
[0644] iv. Sub-block based intraTMP filter / CCCM.
[0645] v. Sub-block based EIF.
[0646] vi. Sub-block based CCLM.
[0647] 17) More than one filter can be used for non-intra-frame blocks, and the final prediction can be derived based on fusing the filtered outputs of multiple filters together.
[0648] For example, the weights that fuse multiple filter values can be solved using a Gaussian elimination solver.
[0649] For example, the weights that fuse multiple filter values can be solved using the LDL decomposition method.
[0650] (hhh) For example, the weights for fusing multiple filter values can be solved using a linear regression method.
[0651] iii) For example, the weights for fusing multiple filter values can be solved using linear equations.
[0652] For example, bias terms can be involved in fusion.
[0653] For example, nonlinear terms can be involved in fusion.
[0654] 18) More than one filter may be allowed for use in non-intra-frame blocks, and which filter is ultimately selected may be transmitted or partitioned by the signal.
[0655] For example, syntax elements can be transmitted via signals to indicate which filter (e.g., CCLM or CCCM) is used for non-intra-frame blocks.
[0656] For example, which filter (e.g., CCLM or CCCM) is used for a non-intra-frame block can be determined based on template costs from both the encoder and decoder.
[0657] 19) The filter output can be limited to a certain value.
[0658] For example, it can be amplitude-limited based on the reconstructed / predicted values of samples in the training region and / or template.
[0659] i. For example, the filter output can be limited to the minimum and maximum values of the reconstructed (or predicted) luminance sample values in the training region and / or template.
[0660] For example, it can be amplitude-limited based on the reconstructed (or predicted) values of samples in the reference block.
[0661] i. For example, it can be bounded to the minimum and maximum values of the reconstructed (or predicted) values of the samples in the reference block.
[0662] For example, if the value is outside the valid range, it can be ignored / discarded / not used.
[0663] 20) Filter parameters for non-intra-blocks can be stored in a buffer and used for encoding and decoding of future blocks.
[0664] For example, filter parameters can include model type, model coefficients, whether it is a single model or multiple models, threshold for separating samples into multiple models, etc.
[0665] For example, it can be stored in a local cache for encoding and decoding future blocks in the current image.
[0666] For example, it can be stored in the temporal domain / image / frame buffer for encoding and decoding future blocks in future decoded images.
[0667] i. For example, the filter parameters of a block in the current frame / image can be stored and referenced for the filtering process of blocks in future frames / images.
[0668] ii. For example, it can be stored in association with motion and pattern information of the video unit.
[0669] 21) Video blocks can inherit model parameters from previous filter-based blocks.
[0670] For example, video blocks can be encoded and decoded using a filter inheritance mode.
[0671] For example, the parameters of the filter model of the previous filter-based block can be stored in a cache (e.g., local cache, image cache, temporal cache, history-based LUT, etc.).
[0672] 22) The mixing / fusion weights of different assumptions of MHP can be derived based on the encoding and decoding information at both the encoder and decoder.
[0673] For example, the blending / fusion weights of MHP can be adaptively / instantly calculated for each video unit.
[0674] (www) For example, the mixing / fusion weights of MHP can be calculated based on a linear / nonlinear filter model.
[0675] i. For example, filter coefficients can be mixed / fused weights to fuse different assumptions of MHP.
[0676] ii. For example, filter coefficients can be derived based on a solver based on Gaussian elimination (or LDL).
[0677] iii. For example, the filter coefficients of the model can be derived based on a set / group of training (template) samples.
[0678] 1. For example, training samples can be constructed based on samples in a template.
[0679] 2. For example, a template can be constructed using up to M rows above and up to N columns to the left (e.g., adjacent) of the current video unit.
[0680] 3. For example, a template can be constructed using the assumption that the cells above the MHP assumption are at most M rows and the cells to the left are at most N columns (e.g., adjacent).
[0681] 4. For example, M and / or N can be equal to 6 or 4 of the luminance samples.
[0682] 5. For example, M=N.
[0683] (xxx) In addition, for example, video units encoded and decoded by MHP can be derived based on intra-frame / intraTMP prediction.
[0684] (yyy) In addition, for example, video units encoded and decoded by MHP can be derived based on IBC prediction.
[0685] 23) The permission and / or use of filter-based modes may depend on at least one of the following aspects: The prediction mode of the video unit (e.g., MODE_INTRA, MODE_INTER, MODE_IBC, MODE_PLT, etc.).
[0686] (aaaa) Prediction modes, such as affine or DMVR or bidirectional / unidirectional prediction or sub-block-based prediction or BDOF or OBMC or CIIP or LIC, etc.
[0687] bbbb) The transformation type of the video unit (e.g., ACT, color transformation, etc.).
[0688] cccc) The number of non-zero coefficients in a video unit.
[0689] dddd) Split tree type (e.g., single tree, double tree).
[0690] eeee) Strip type (e.g., I, B, P stripes).
[0691] ffff) Color format (e.g., whether it is 4:0:0).
[0692] gggg) The block dimension (e.g., width / height) of the current block.
[0693] 24) The disclosed filter-based non-intra-frame (or intra-frame) prediction can be based on one of the following filters: hhhh) Any filter-based prediction, where the filter coefficients are solved based on the correlation between the prediction and / or reconstruction information.
[0694] iii) Filters based on CCLM and / or its variants.
[0695] jjjj) Filters based on MMLM and / or its variants.
[0696] kkkk) Filters based on CCCM and / or its variants (e.g., GL-CCCM, non-downsampled CCCM, BVG-CCCM, CCRM, etc.).
[0697] llll) Filters based on GLM and / or its variants.
[0698] mmmm) CCRM (e.g., inter-frame CCCM, IBC CCCM, non-intra CCCM) and / or its variants.
[0699] nnnn) CCCM for intra and / or its variants.
[0700] oooo) CCCM for IBC and / or its variants.
[0701] pppp) CCCM for intraTMP and / or its variants.
[0702] qqqq) Convolution filters and / or its variants.
[0703] rrrr) Filters based on Gaussian elimination solver and / or its variants.
[0704] 25) Block restrictions can be applied to limit the application of specific types of filter-based non-intra prediction modes.
[0705] ssss) For example, filter-based non-intra prediction modes can only be allowed for block sizes that meet predefined rules.
[0706] tttt) For example, syntax elements can only be signaled when filter-based non-intra prediction modes are applicable.
[0707] uuuu) For example, if filter-based non-intra prediction modes are not allowed to be used, syntax elements can be presumed to indicate a specific value that no such CCP mode is used for such blocks.
[0708] vvvv) For example, at least one of the following block restrictions can be applied to filter-based non-intra prediction modes (assuming W represents the block width and H represents the block height): i. W < T1, or, W <= T1.
[0709] ii. H < T2, or, H <= T2.
[0710] iii. Min (W, H) > T3, or Min (W, H) >= T3.
[0711] iv. Max (W, H) < T4, or Max (W, H) <= T4.
[0712] v. W < T5 H, or W <= T5 H.
[0713] vi. W > T6 H, or W >= T6 H.
[0714] vii. H < T7 W, or H <= T7 W.
[0715] viii. H > T8 W, or H >= T8 W.
[0716] ix. W H < T9, or W H <= T9.
[0717] x. For example, T1, T2,... T9 can be predefined integer constants.
[0718] wwww) For example, the filter-based non-intra prediction mode can be allowed only for small blocks.
[0719] i. For example, it can be allowed for blocks smaller than 4x4 or 8x8 or 16x16 or 32x32.
[0720] ii. For example, it can be allowed for blocks with a sample number less than 32 or 64 or 128.
[0721] iii. For example, it can be allowed for blocks with a sample number less than 32 or 64 or 128.
[0722] iv. For example, for a 2xN block, it can be not allowed, where N can be greater than 4 or 8 or 16.
[0723] v. For example, for an Nx2 block, it can be not allowed, where N can be greater than 4 or 8 or 16.
[0724] 26) The disclosed method can be used in a single tree.
[0725] 27) The disclosed method can be used in a dual tree.
[0726] 28) The disclosed method can be used in inter-frame (such as B or P) stripes.
[0727] 29) The disclosed method can be used in intra-frame (such as I) stripes.
[0728] 30) The “block vector” in the disclosed method can be a “motion vector”.
[0729] 31) The training / reference samples in the disclosed method may refer to the predicted samples and / or reconstructed samples in the training / reference region.
[0730] 32) Whether and / or how the methods disclosed above can be applied to be transmitted via signaling at the sequence level / picture group level / picture level / strip level / piece group level, such as in the sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / strip header / piece group header.
[0731] 33) Whether and / or how the methods disclosed above can be applied to transmit signals at PB / TB / CB / PU / TU / CU / VPDU / CTU / CTU lines / strips / films / sub-images / other types of areas containing more than one sample point or pixel.
[0732] 34) Whether and / or how the methods disclosed above are applied may depend on the encoded / decoded information, such as block size, color format, single / double tree segmentation, color components, and stripe / picture type.
[0733] Further embodiments will be described below. Figure 32 A flowchart of a method 3200 for video processing according to an embodiment of the present disclosure is shown. Method 3200 is implemented during the conversion between video regions, such as video units or video blocks, and the bitstream of the video.
[0734] At box 3210, for the conversion between the current video region and the video bitstream, the filter-based non-intra-frame or intra-frame prediction of the current video region is determined based on one of the following: Transform Unit (TU), Codec Unit (CU), Prediction Unit (PU), or Sub-block.
[0735] At box 3220, the conversion based on filter-based non-intra-frame or intra-frame prediction is performed.
[0736] Method 3200 enables the application of filter-based non-intra-frame or intra-frame prediction to video encoding and decoding, thereby improving encoding and decoding efficiency and / or encoding and decoding effectiveness.
[0737] In some embodiments, filter-based non-intra-frame or intra-frame prediction is based on TU, CU, or PU, and TU, CU, or PU are not divided into sub-blocks for filtering.
[0738] In some embodiments, for TU, CU, or PU, the filter model is modulated using a set of training samples with respect to TU, CU, or PU, and the filter model is linear or nonlinear.
[0739] In some embodiments, a set of model coefficients of the filter model is determined, and the filter model is applied to determine the estimation of the predicted or residual samples of TU, CU, or PU.
[0740] In some embodiments, at least one of the following CCRM, CCCM, or IBC CCCM—Residual Codec Recognition Based on Cross-Component Model (CCRM), Inter-Frame Convolutional Cross-Component Model (CCCM), or Intra-Frame Block Copy (IBC) CCCM—is applied at the level of TU, PU, or CU, and TU, PU, or CU are not divided into sub-blocks.
[0741] In some embodiments, the coefficients of at least one of the CCRM, CCCM, or IBC CCCM are determined based on training samples at the TU, PU, or CU level, and at least one of the CCRM, CCCM, or IBC CCCM is applied to determine the estimation of predicted samples or residual samples for the TU, PU, or CU.
[0742] In some embodiments, all applicable samples in TU, PU, or CU are determined based on the same filter model with the same model coefficients.
[0743] In some embodiments, at least one of the following CCRM, inter-frame CCCM, or IBC CCCM based on TU, PU, or CU is suitable for multiple block sizes allowed for the conversion.
[0744] In some embodiments, CCRM or inter-frame CCCM or IBC CCCM is permitted for the current video region, and the filter model is modulated at the level of the current video region, regardless of the size of the current video region, which is one of the following: CU, PU, or TU.
[0745] In some embodiments, the model derivation is applied to the CU, PU, or TU, and the derived model is applied to applicable samples in the CU, PU, or TU.
[0746] In some embodiments, filter-based non-intra-frame or intra-frame prediction is performed based on sub-blocks.
[0747] In some embodiments, CU, PU, or TU is divided into multiple sub-blocks, each sub-block having multiple filter models, wherein the filter models among the multiple filter models have at least one of the following: model coefficients, filter shape, or filter terms.
[0748] In some embodiments, the size of the multiple sub-blocks is fixed or predefined.
[0749] In some embodiments, at least one of the width, height, or number of samples of the plurality of sub-blocks is fixed or predefined.
[0750] In some embodiments, the training samples for each sub-block are different.
[0751] In some embodiments, the training samples of each sub-block are determined based on the corresponding sub-block in the reference CU, reference PU, or reference TU, and the reference CU, reference PU, or reference TU are divided into sub-blocks for training sample derivation.
[0752] In some embodiments, the training samples of each sub-block are determined based on a subset of samples that are adjacent to the current CU, current PU, current TU, or reference CU, reference PU, or reference TU.
[0753] In some embodiments, a subset of sample points corresponds to each sub-block.
[0754] In some embodiments, the training samples of each sub-block are determined based on neighboring samples that are adjacent or not adjacent to the current CU, current PU, current TU, or reference CU, reference PU, or reference TU.
[0755] In some embodiments, block vector-guided convolutional cross-component model (CCCM) for intra-block copying (IBC) or intra-template matching prediction (intraTPM) is performed based on sub-blocks.
[0756] In some embodiments, at least one of the following is performed based on sub-blocks: Convolutional Cross-Component Model (CCCM), Cross-Component Linear Model (CCLM), Intra-Template Matching Prediction (intraTPM) filter, Intra-Block Copy (IBC) filter, or variant filter mode.
[0757] In some embodiments, filter-based non-intra-frame or intra-frame prediction for the current video region is performed at an adaptive sub-block granularity.
[0758] In some embodiments, the video unit is divided into multiple sub-blocks of the same size.
[0759] In some embodiments, multiple video units are divided into multiple groups of sub-blocks, with each group having a different sub-block size granularity.
[0760] In some embodiments, for video units of different sizes, at least one of the sub-block width or sub-block height is defined.
[0761] In some embodiments, the number of sub-blocks in width and / or height is defined.
[0762] In some embodiments, the rules for obtaining the size of the sub-blocks are predefined.
[0763] In some embodiments, larger video units have larger sub-block sizes, and smaller video units have smaller sub-block sizes.
[0764] In some embodiments, rules for obtaining the size of sub-blocks are included in the bitstream.
[0765] In some embodiments, whether video units are divided into sub-blocks is included in the bitstream.
[0766] In some embodiments, whether video units are divided into sub-blocks is included in the bitstream.
[0767] In some embodiments, the adaptive sub-block size is assigned based on the width and / or height of the TU, PU, or CU of the current video region.
[0768] In some embodiments, the adaptive sub-block size is based on the width and / or height of the current video region.
[0769] In some embodiments, the adaptive sub-block size is based on the total number of samples in the current video region.
[0770] In some embodiments, whether and how the sub-block size is determined is based on encoding and decoding information at the encoder and decoder for the conversion.
[0771] In some embodiments, the determination of whether to use a sub-block-based filter or a TU, CU, or PU level filter is implicitly determined based on the encoding / decoding information.
[0772] In some embodiments, the determination of whether to use a filter based on M1xM2 sub-blocks or a filter based on N1xN2 sub-blocks is implicitly determined based on encoding and decoding information, where M1, M2, N1, and N2 are positive integers.
[0773] In some embodiments, M1 is one of the following: 16 or 8 or 32 or TU or CU or PU, M2 is one of the following: 16 or 8 or 32 or TU or CU or PU, N1 is one of the following: 16 or 8 or 32 or TU or CU or PU, N2 is one of the following: 16 or 8 or 32 or TU or CU or PU, M1 is not equal to N1, and / or M2 is not equal to N2.
[0774] In some embodiments, whether and how the sub-block size is determined is based on a template cost-based scheme.
[0775] In some embodiments, the template cost is determined based on minimizing one of the following: the sum of absolute differences (SAD), the sum of absolute transformation differences (SATD), the sum of squared errors (SSE), or the mean squared error (MSE) between the estimated sample values from the filter model and the reconstructed sample values, with the sample references being at least one training sample.
[0776] In some embodiments, a scheme with lower template cost is selected as the scheme to be applied to the current video region.
[0777] In some embodiments, whether and how the sub-block size is determined is based on information from a reference image.
[0778] In some embodiments, the information of the reference image includes at least one of the following: the sequential image (POC) distance between the current image and the current image, or the reference index of the reference image.
[0779] In some embodiments, whether and how the sub-block size is determined is indicated in the bitstream.
[0780] In some embodiments, whether and how a sub-block-based modeling filter is applied is based on one of the following indicators: Overlapping Block Motion Compensation (OBMC), Local Illumination Compensation (LIC), Intra-Block Copying (IBC), Palette Encoder / Decoder Tool, Block-Based Incremental Pulse Code Modulation (BDPCM), Intra-Template Matching Prediction (intraTMP), Decoder-Side Motion Vector Refinement (DMVR), Inter-Frame Template Matching (TM), Affine Tool, Sub-Block-Based Temporal Motion Vector Prediction (sbTMVP), Intra-Segmentation (ISP), Geometric Segmentation Mode (GPM), Intra-Inter-Frame Joint Prediction (CIIP), Spatial GPM (SGPM), Sub-Block Transform (SBT), Merge Tool, or Advanced Motion Vector Prediction (AMVP).
[0781] In some embodiments, method 3200 further includes: processing samples of the current video region with a first filter; and applying a second filter to the processed samples.
[0782] In some embodiments, the input to the second filter includes the output of the first filter.
[0783] In some embodiments, the input to the second filter includes samples filtered by the first filter.
[0784] In some embodiments, the first filter is based on a linear or nonlinear model, which includes one of the following: residual coding-decoding (CCRM) based on a cross-component model, intra-convolutional cross-component model (CCCM), inter-frame CCCM, intra-block copy (IBC) CCCM, intra-template matching prediction (intraTMP) CCCM, EIF, or cross-component linear model (CCLM).
[0785] In some embodiments, the first filter is applied at the sub-block level.
[0786] In some embodiments, the TU, PU, or CU is divided into multiple sub-blocks for filtering, the multiple sub-blocks including 16×16 sub-blocks.
[0787] In some embodiments, the second filter includes at least one of the following: a smoothing filter or a boundary filter.
[0788] In some embodiments, the second filter is applied to boundary samples located at the boundaries of the sub-blocks.
[0789] In some embodiments, a K-tap filter is applied to samples located in the first row and / or first column and / or last row and / or last column of each sub-block, where K is 2.
[0790] In some embodiments, the input to the second filter includes at least one sample point within the sub-block to be processed, the at least one sample point being within a neighboring sub-block of the sub-block to be processed.
[0791] In some embodiments, the output of the samples filtered by the 2-tap boundary filter is based on y = (a0) x0 +a1 x1 + offset)>>shift is determined, where y represents the output value, a0 and a1 are weights or coefficients, x0 and x1 are the input samples within and next to the sub-block to be processed, shift is calculated based on the sum of a0 and a1, and offset is a fixed value.
[0792] In some embodiments, the weights or coefficients of the input samples of the second filter are predefined.
[0793] In some embodiments, the first weight of the input sample within the sub-block to be processed is greater than the second weight of the input sample adjacent to the sub-block to be processed.
[0794] In some embodiments, at least one value of at least one sample point in the current video region is modified by a second filter.
[0795] In some embodiments, the second filter is applied to internal samples located inside the sub-block.
[0796] In some embodiments, the values of boundary samples at one of the following locations are modified by a second filter: the first row of the sub-block, the first column of the sub-block, the last row of the sub-block, or the last column of the sub-block.
[0797] In some embodiments, the values of boundary samples at one of the following locations are not modified by the second filter: the first row of the sub-block, the first column of the sub-block, the last row of the sub-block, or the last column of the sub-block.
[0798] In some embodiments, whether a sample is modified by the second filter is based on the sub-block to which the sample belongs.
[0799] In some embodiments, at least one of the first row samples of the first row sub-block, the first column samples of the first column sub-block, the last row samples of the last row sub-block, or the last column samples of the last column sub-block is not filtered by the second filter.
[0800] In some embodiments, whether the sample to be processed is further modified by the second filter is based on whether the input sample of the second filter was processed by the first filter.
[0801] In some embodiments, the input samples of the second filter are processed by the first filter, and the second filter is applied to modify the values of the samples to be processed.
[0802] In some embodiments, the input samples include at least one of the following: a sample to be processed on one side of an edge or boundary, or a neighboring sample on the other side of an edge or boundary.
[0803] In some embodiments, method 3200 further includes applying a deblocking process to the current video region based on a linear or nonlinear filter model.
[0804] In some embodiments, whether and / or how the deblocking process is applied to boundary or edge samples is determined based on a filter model, which includes one of the following: residual coding-decoding (CCRM) based on a cross-component model, intra-convolutional cross-component model (CCCM), inter-frame CCCM, intra-block copy (IBC) CCCM, intra-template matching prediction (intraTMP) CCCM, EIF, or cross-component linear model (CCLM).
[0805] In some embodiments, the deblocking intensity of the deblocking process is determined based on a filter model.
[0806] In some embodiments, whether strong or weak deblocking is performed on edges or boundaries is determined based on the filter model.
[0807] In some embodiments, whether to perform long deblocking or short deblocking on edges or boundaries is determined based on the filter model.
[0808] In some embodiments, the values of the deblocking filter parameters in the deblocking process are determined based on the filter model.
[0809] In some embodiments, the values of the deblocking filter parameters in the deblocking process are determined based on whether the predictions or residuals of samples along the edges or boundaries are based on a sub-block filter model.
[0810] In some embodiments, the filter model includes one of the following: residual codec based on a sub-block-based cross-component model (CCRM), sub-block-based intra-convolutional cross-component model (CCCM), sub-block-based intra-block copy (IBC) filter or IBC CCCM, sub-block-based intra-template matching prediction (intraTMP) filter or CCCM, sub-block-based EIF, or sub-block-based cross-component linear model (CCLM).
[0811] In some embodiments, the mixed weights or fusion weights of multiple hypotheses in multiple hypothesis prediction (MHP) are determined based on encoding and decoding information at the encoder and decoder for the transformation.
[0812] In some embodiments, the MHP blending weights or fusion weights are adaptively or instantaneously determined for each video unit.
[0813] In some embodiments, the hybrid weights or fusion weights of MHP are determined based on a filter model based on linear or nonlinearity.
[0814] In some embodiments, the filter coefficients include mixed weights or fused weights to fuse different assumptions of the MHP.
[0815] In some embodiments, the filter coefficients are determined based on one of the following: Gaussian elimination or an LDL-based scheme.
[0816] In some embodiments, the filter coefficients of the model are determined based on a set or group of training samples or template samples.
[0817] In some embodiments, training samples are constructed based on samples in a template.
[0818] In some embodiments, the template is constructed using at most M rows above and at most N columns to the left of the current video region, where M and N are positive integers, or the template is constructed using at most M rows above and at most N columns to the left of the hypothetical unit of the MHP hypothesis.
[0819] In some embodiments, M is equal to 6 or 4 of the luminance samples, and / or N is equal to 6 or 4 of the luminance samples.
[0820] In some embodiments, M=N.
[0821] In some embodiments, the video units encoded and decoded by MHP are determined based on intra-frame prediction or intra-frame template matching prediction (intraTMP).
[0822] In some embodiments, video units encoded and decoded by MHP are determined based on intra-block copy (IBC) prediction.
[0823] In some embodiments, at least one of sample values, gradients, or positional information is used to determine a filter, the filter including at least one item corresponding to sample values for different color components.
[0824] In some embodiments, at least one item includes a sample value corresponding to a luminance sample, wherein the luminance sample is a downsampled chrominance sample.
[0825] In some embodiments, the permission and / or use of filter-based modes is based on at least one of the following: the prediction mode of the current video block, the transform type of the current video block, the number of non-zero coefficients of the current video block, the segmentation tree type of the current video block, the stripe type, the color format, or the block dimension of the current video block.
[0826] In some embodiments, the prediction mode includes at least one of the following: affine mode, decoder-size motion vector refinement, bidirectional prediction, unidirectional prediction, sub-block-based prediction, bidirectional optical flow (BDOF), overlapping block motion compensation (OBMC), intra-frame and inter-frame joint prediction (CIIP), or local illumination compensation (LIC).
[0827] In some embodiments, filter-based non-intra-frame or intra-frame prediction is based on one of the following: filter-based prediction, cross-component linear model (CCLM), a variant of CCLM, multi-model linear model (MMLM), a variant of MMLM, convolutional cross-component model (CCCM), a variant of CCCM, gradient linear model (GLM) or a variant of GLM, cross-component model-based residual codec (CCRM), a variant of CCRM, a convolutional filter, a variant of a convolutional filter, a filter based on a Gaussian elimination solver, or a variant thereof.
[0828] In some embodiments, CCCM includes at least one of the following: CCCM for intra-frame, CCCM for intra-block copy (IBC), or CCCM for intra-template matching prediction (intraTMP).
[0829] In some embodiments, CCRM includes at least one of the following: inter-frame CCCM, intra-frame block copy (IBC) CCCM, or non-intra-frame CCCM.
[0830] In some embodiments, the method is used in a single tree or a double tree.
[0831] In some embodiments, the method is used in inter-frame stripes, which include at least one of the following: B stripes or P stripes.
[0832] In some embodiments, the method is used in intra-frame stripes, which include I-stripes.
[0833] In some embodiments, the block vector (BV) involved in the method is replaced by a motion vector (MV).
[0834] In some embodiments, the training samples or reference samples used in the method refer to at least one of the prediction samples or reconstructed samples in the training region or reference region.
[0835] In some embodiments, an indication of whether and / or how to apply the method is given at one of the following levels: sequence level, picture group level, picture level, strip level, or slice group level.
[0836] In some embodiments, the indication of whether and / or how to apply the method is indicated in one of the following: sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), dependency parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), strip header, or slice header.
[0837] In some embodiments, an indication of whether and / or how to apply a method is included in one of the following: prediction block (PB), transform block (TB), codec block (CB), prediction unit (PU), transform unit (TU), codec unit (CU), virtual pipeline data unit (VPDU), codec tree unit (CTU), CTU row, strip, slice, sub-picture, or region containing more than one sample point or pixel.
[0838] In some embodiments, method 3200 further includes: determining whether and / or how to apply the method based on the encoded and decoded information of the current video block, the encoded and decoded information including at least one of the following: block size, color format, single-tree segmentation and / or dual-tree segmentation, color components, stripe type, or picture type.
[0839] In some embodiments, the conversion includes encoding the current video region into a bitstream.
[0840] In some embodiments, the conversion includes decoding the current video region from the bitstream.
[0841] According to another embodiment of this disclosure, a non-transitory computer-readable recording medium is provided. This non-transitory computer-readable recording medium stores a bitstream of video generated by a method performed by means of a video processing apparatus. The method includes: determining a filter-based non-intra-frame or intra-frame prediction of a current video region based on one of: a transform unit (TU), a codec unit (CU), a prediction unit (PU), or a sub-block; and generating a bitstream based on the filter-based non-intra-frame or intra-frame prediction.
[0842] According to further embodiments of this disclosure, a method for storing a bitstream of video is provided. The method includes: determining a filter-based non-intra-frame or intra-frame prediction of a current video region based on one of: a transform unit (TU), a codec unit (CU), a prediction unit (PU), or a sub-block; generating a bitstream based on the filter-based non-intra-frame or intra-frame prediction; and storing the bitstream in a non-transitory computer-readable recording medium.
[0843] The embodiments of this disclosure can be described according to the following entries, and their features can be combined in any reasonable manner.
[0844] Item 1. A method for video processing, comprising: for a conversion between a current video region of a video and a bitstream of the video, determining a filter-based non-intra-frame or intra-frame prediction of the current video region based on one of: a transform unit (TU), a codec unit (CU), a prediction unit (PU), or a sub-block; and performing the conversion based on the filter-based non-intra-frame or intra-frame prediction.
[0845] Item 2. The method according to Item 1, wherein the filter-based non-intra-frame or intra-frame prediction is based on the TU or the CU or the PU, and the TU or the CU or the PU is not divided into sub-blocks for filtering.
[0846] Item 3. The method according to Item 1 or 2, wherein for the TU, the CU, or the PU, the filter model is modulated using a set of training samples with respect to the TU, the CU, or the PU, and the filter model is linear or nonlinear.
[0847] Item 4. The method according to Item 3, wherein a set of model coefficients of the filter model is determined, and the filter model is applied to determine the estimation of the predicted or residual samples of the TU, the CU, or the PU.
[0848] Item 5. The method according to any one of Items 1-4, wherein at least one of Residual Codec Reduction Based on Cross-Component Model (CCRM), Inter-Frame Convolutional Cross-Component Model (CCCM), or Intra-Frame Block Copy (IBC) CCCM is applied at a level based on the TU, the PU, or the CU, and the TU, the PU, or the CU is not divided into sub-blocks.
[0849] Item 6. According to the method of Item 5, the coefficients of at least one of the CCRM, the inter-frame CCCM, or the IBC CCCM are determined based on training samples at the TU, PU, or CU level, and at least one of the CCRM, the inter-frame CCCM, or the IBC CCCM is applied to determine the estimation of the predicted samples or residual samples of the TU, PU, or CU.
[0850] Item 7. According to the method described in Item 6, all applicable samples in the TU, PU, or CU are determined based on the same filter model with the same model coefficients.
[0851] Item 8. The method according to any one of Items 5-7, wherein at least one of the CCRM, the inter-frame CCCM, or the IBC CCCM based on the TU, the PU, or the CU is applicable to multiple block sizes allowed for the conversion.
[0852] Item 9. The method according to Item 8, wherein the CCRM or inter-frame CCCM or IBC CCCM is permitted for the current video region, and the filter model is modulated at the level of the current video region, regardless of the size of the current video region, wherein the current video region is one of the following: the CU, the PU, or the TU.
[0853] Item 10. The method according to Item 8 or 9, wherein the model derivation is applied to the CU or the PU or the TU, and the derived model is applied to applicable samples in the CU or the PU or the TU.
[0854] Item 11. The method according to Item 1, wherein the filter-based non-intra-frame or intra-frame prediction is performed based on the sub-block.
[0855] Item 12. The method according to Item 11, wherein the CU or the PU or the TU is divided into a plurality of sub-blocks, the plurality of sub-blocks having a plurality of filter models, wherein the filter models in the plurality of filter models have at least one of the following: model coefficients, filter shape or filter terms.
[0856] Item 13. The method according to Item 11 or 12, wherein the sub-block size of the plurality of sub-blocks is fixed or predefined.
[0857] Item 14. The method according to Item 13, wherein at least one of the width, height or number of samples of the plurality of sub-blocks is fixed or predefined.
[0858] Item 15. The method according to any one of items 11-14, wherein the training samples of each sub-block are different.
[0859] Item 16. The method according to Item 15, wherein the training samples of each sub-block are determined based on the corresponding sub-block in a reference CU or reference PU or reference TU, the reference CU or the reference PU or the reference TU being divided into sub-blocks for training sample derivation.
[0860] Item 17. The method according to Item 15, wherein the training samples of each sub-block are determined based on a subset of samples that are adjacent to the current CU or the current PU or the current TU or the reference CU or the reference PU or the reference TU.
[0861] Item 18. The method according to Item 17, wherein the partial sample points correspond to each sub-block.
[0862] Item 19. The method according to Item 15, wherein the training samples of each sub-block are determined based on neighboring samples that are adjacent or not adjacent to the current CU, the current PU, the current TU, the reference CU, the reference PU, or the reference TU.
[0863] Item 20. The method according to any one of items 11-19, wherein the block vector-guided convolutional cross-component model (CCCM) for intra-block copying (IBC) or intra-template matching prediction (intraTPM) is performed based on sub-blocks.
[0864] Item 21. The method according to any one of Items 11-20, wherein at least one of the following is performed based on a sub-block: Convolutional Cross-Component Model (CCCM), Cross-Component Linear Model (CCLM), Intra-Template Matching Prediction (intraTPM) filter, Intra-Block Copy (IBC) filter, or variant filter mode.
[0865] Item 22. The method according to any one of items 11-21, wherein the filter-based non-intra-frame or intra-frame prediction for the current video region is performed based on an adaptive sub-block granularity.
[0866] Item 23. The method according to Item 22, wherein the video unit is divided into multiple sub-blocks of the same size.
[0867] Item 24. The method according to Item 22, wherein a plurality of video units are divided into a plurality of sub-blocks, the plurality of sub-blocks having different sub-block size granularities.
[0868] Item 25. The method according to Item 22, wherein at least one of sub-block width or sub-block height is defined for video units of different sizes.
[0869] Item 26. The method according to Item 25, wherein the number of sub-blocks in width and / or height is defined.
[0870] Item 27. The method described in Item 22, wherein the rules for obtaining the size of the sub-blocks are predefined.
[0871] Item 28. The method according to Item 27, wherein the larger video unit has a larger sub-block size and the smaller video unit has a smaller sub-block size.
[0872] Item 29. The method according to Item 22, wherein the rules for obtaining the size of the sub-blocks are included in the bitstream.
[0873] Item 30. The method according to Item 29, wherein whether or not video units are divided into sub-blocks is included in the bitstream.
[0874] Item 31. The method according to Item 30, wherein whether or not video units are divided into sub-blocks is included in the bitstream.
[0875] Item 32. The method according to any one of items 11-31, wherein the adaptive sub-block size is allocated based on the width and / or height of the TU, PU, or CU of the current video region.
[0876] Item 33. The method according to Item 32, wherein the adaptive sub-block size is based on the width and / or height of the current video region.
[0877] Item 34. The method according to Item 32, wherein the adaptive sub-block size is based on the total number of samples in the current video region.
[0878] Item 35. The method according to any one of items 11-34, wherein whether and how the sub-block size is determined is based on encoding / decoding information at the encoder and decoder for the transformation.
[0879] Item 36. The method according to Item 35, wherein the determination of whether to use a sub-block-based filter or a TU or CU or PU level filter is implicitly determined based on encoding / decoding information.
[0880] Item 37. The method according to Item 35, wherein the determination of whether to use a filter based on M1xM2 sub-blocks or a filter based on N1xN2 sub-blocks is implicitly determined based on encoding / decoding information, where M1, M2, N1, and N2 are positive integers.
[0881] Item 38. The method according to Item 37, wherein M1 is one of the following: 16 or 8 or 32 or TU or CU or PU, M2 is one of the following: 16 or 8 or 32 or TU or CU or PU, N1 is one of the following: 16 or 8 or 32 or TU or CU or PU, N2 is one of the following: 16 or 8 or 32 or TU or CU or PU, M1 is not equal to N1, and / or M2 is not equal to N2.
[0882] Item 39. The method according to Item 35, wherein whether and how the sub-block size is determined is based on a template cost-based scheme.
[0883] Item 40. The method according to Item 39, wherein the template cost is determined based on minimizing one of the following: sum of absolute differences (SAD), sum of absolute transform differences (SATD), sum of squared errors (SSE), or mean squared error (MSE) between estimated sample values and reconstructed sample values from the filter model, wherein the sample references at least one training sample among training samples.
[0884] Item 41. The method according to Item 39 or 40, wherein the scheme with lower template cost is selected as the scheme to be applied to the current video region.
[0885] Item 42. The method according to Item 35, wherein whether and how the size of the sub-block is determined is based on information from a reference image.
[0886] Item 43. The method according to Item 42, wherein the information of the reference image includes at least one of the following: the sequential image (POC) distance between the current image and the current image, or the reference index of the reference image.
[0887] Item 44. The method according to any one of items 11-34, wherein whether and how the sub-block size is determined is indicated in the bitstream.
[0888] Item 45. The method according to any one of items 11-44, wherein whether and how the sub-block-based modeling filter is applied is based on an indicator of one of the following: Overlapping Block Motion Compensation (OBMC), Local Illumination Compensation (LIC), Intra-Block Copying (IBC), Palette Encoder / Decoder Tool, Block-based Incremental Pulse Code Modulation (BDPCM), Intra-Template Matching Prediction (intraTMP), Decoder-Side Motion Vector Refinement (DMVR), Inter-Frame Template Matching (TM), Affine Tool, Sub-block-based Temporal Motion Vector Prediction (sbTMVP), Intra-Segmentation (ISP), Geometric Segmentation Mode (GPM), Intra-Inter-Frame Joint Prediction (CIIP), Spatial GPM (SGPM), Sub-block Transform (SBT), Merge Tool, or Advanced Motion Vector Prediction (AMVP).
[0889] Item 46. The method according to any one of items 1-45 further includes: processing samples of the current video region with a first filter; and applying a second filter to the processed samples.
[0890] Item 47. The method according to Item 46, wherein the input of the second filter includes the output of the first filter.
[0891] Item 48. The method according to Item 46, wherein the input of the second filter includes samples filtered by the first filter.
[0892] Item 49. The method according to any one of items 46-48, wherein the first filter is based on a linear or nonlinear model, the model including one of: residual coding-decoding (CCRM) based on a cross-component model, intra-convolutional cross-component model (CCCM), inter-frame CCCM, intra-block copy (IBC) CCCM, intra-template matching prediction (intraTMP) CCCM, EIF, or cross-component linear model (CCLM).
[0893] Item 50. The method according to any one of items 46-49, wherein the first filter is applied at the sub-block level.
[0894] Item 51. The method according to Item 50, wherein the TU or PU or CU is divided into a plurality of sub-blocks for filtering, the plurality of sub-blocks comprising 16×16 sub-blocks.
[0895] Item 52. The method according to any one of items 46-51, wherein the second filter comprises at least one of: a smoothing filter or a boundary filter.
[0896] Item 53. The method according to any one of items 46-52, wherein the second filter is applied to boundary samples located at the boundary of the sub-block.
[0897] Item 54. The method according to Item 53, wherein a K-tap filter is applied to samples located in the first row and / or first column and / or last row and / or last column of each sub-block, where K is 2.
[0898] Item 55. The method according to any one of items 46-53, wherein the input of the second filter includes at least one sample point within the sub-block to be processed, the at least one sample point being within a neighboring sub-block of the sub-block to be processed.
[0899] Item 56. The method according to Item 55, wherein the output of the sample filtered by the 2-tap boundary filter is based on y = (a0) x0 + a1 x1 + offset)>>shift is determined, where y represents the output value, a0 and a1 are weights or coefficients, x0 and x1 are the input samples within and next to the sub-block to be processed, shift is calculated based on the sum of a0 and a1, and offset is a fixed value.
[0900] Item 57. The method according to Item 56, wherein the weights or coefficients of the input samples of the second filter are predefined.
[0901] Item 58. The method according to Item 56, wherein the first weight of the input sample within the sub-block to be processed is greater than the second weight of the input sample adjacent to the sub-block to be processed.
[0902] Item 59. The method according to any one of items 46-58, wherein at least one value of at least one sample point in the current video region is modified by the second filter.
[0903] Item 60. The method according to Item 59, wherein the second filter is applied to internal samples located inside the sub-block.
[0904] Item 61. The method according to Item 59, wherein the value of the boundary sample point is modified by the second filter at one of the following locations: the first row of the sub-block, the first column of the sub-block, the last row of the sub-block, or the last column of the sub-block.
[0905] Item 62. The method according to Item 59, wherein the value of the boundary sample at one of the following locations is not modified by the second filter: the first row of the sub-block, the first column of the sub-block, the last row of the sub-block, or the last column of the sub-block.
[0906] Item 63. The method according to any one of items 46-62, wherein whether a sample is modified by the second filter is based on the sub-block to which the sample belongs.
[0907] Item 64. The method according to Item 63, wherein at least one of the first row sample of the first row sub-block, the first column sample of the first column sub-block, the last row sample of the last row sub-block, or the last column sample of the last column sub-block is not filtered by the second filter.
[0908] Item 65. The method according to any one of items 46-64, wherein whether the sample to be processed is further modified by the second filter is based on whether the input sample of the second filter is processed by the first filter.
[0909] Item 66. The method according to Item 65, wherein the input sample of the second filter is processed by the first filter, and the second filter is applied to modify the value of the sample to be processed.
[0910] Item 67. The method according to Item 65 or 66, wherein the input sample includes at least one of the following: a sample to be processed on one side of the edge or boundary, or a neighboring sample on the other side of the edge or boundary.
[0911] Item 68. The method according to any one of items 1-67 further includes: applying the deblocking process to the current video region based on a linear or nonlinear filter model.
[0912] Item 69. The method according to Item 68, wherein whether and / or how the deblocking process is applied to boundary or edge samples is determined based on a filter model, said filter model including one of: residual coding-decoding (CCRM) based on a cross-component model, intra-convolutional cross-component model (CCCM), inter-frame CCCM, intra-block copy (IBC) CCCM, intra-template matching prediction (intraTMP) CCCM, EIF, or cross-component linear model (CCLM).
[0913] Item 70. The method according to Item 68 or 69, wherein the deblocking intensity of the deblocking process is determined based on the filter model.
[0914] Item 71. The method according to Item 70, wherein whether strong or weak deblocking is performed on the edge or boundary is determined based on the filter model.
[0915] Item 72. The method according to Item 70, wherein whether to perform long deblocking or short deblocking on the edge or boundary is determined based on the filter model.
[0916] Item 73. The method according to any one of items 68-72, wherein the values of the deblocking filter parameters of the deblocking process are determined based on the filter model.
[0917] Item 74. The method according to any one of items 68-72, wherein the values of the deblocking filter parameters of the deblocking process are determined based on whether the prediction or residual of the samples along the edge or boundary is based on a sub-block filter model.
[0918] Item 75. The method according to any one of Items 68-74, wherein the filter model comprises one of the following: residual coding-decoding based on a sub-block-based cross-component model (CCRM), a sub-block-based intra-convolutional cross-component model (CCCM), a sub-block-based intra-block copy (IBC) filter or IBC CCCM, a sub-block-based intra-template matching prediction (intraTMP) filter or CCCM, a sub-block-based EIF, or a sub-block-based cross-component linear model (CCLM).
[0919] Item 76. The method according to any one of items 1-75, wherein the mixed weights or fusion weights of multiple hypotheses of the multiple hypothesis prediction (MHP) are determined based on encoding and decoding information at the encoder and decoder for the transformation.
[0920] Item 77. The method according to Item 76, wherein the mixing weights or fusion weights of the MHP are adaptively or instantaneously determined for each video unit.
[0921] Item 78. The method according to Item 76 or 77, wherein the mixing weights or fusion weights of the MHP are determined based on a linear or nonlinear filter model.
[0922] Item 79. The method according to Item 78, wherein the filter coefficients include the mixing weights or the fusion weights to fuse the different assumptions of the MHP.
[0923] Item 80. The method according to Item 78, wherein the filter coefficients are determined based on one of the following: Gaussian elimination or an LDL-based scheme.
[0924] Item 81. The method according to Item 78, wherein the filter coefficients of the model are determined based on a set or group of training samples or template samples.
[0925] Item 82. The method according to Item 81, wherein the training samples are constructed based on samples in the template.
[0926] Item 83. The method according to Item 81, wherein the template is constructed using at most M rows above and at most N columns to the left adjacent to the current video region, where M and N are positive integers, or wherein the template is constructed using at most M rows above and at most N columns to the left adjacent to the hypothetical unit of the MHP hypothesis.
[0927] Item 84. The method according to Item 83, wherein M is equal to 6 or 4 of the luminance samples, and / or N is equal to 6 or 4 of the luminance samples.
[0928] Item 85. The method described in Item 83, where M=N.
[0929] Item 86. The method according to any one of items 76-85, wherein the video unit encoded and decoded by MHP is determined based on intra-frame prediction or intra-frame template matching prediction (intraTMP).
[0930] Item 87. The method according to any one of items 76-85, wherein the video unit encoded and decoded by MHP is determined based on intra-block copy (IBC) prediction.
[0931] Item 88. The method according to any one of items 1-87, wherein at least one of sample values, gradient, or position information is used to determine a filter, said filter comprising at least one item corresponding to sample values of different color components.
[0932] Item 89. The method according to Item 88, wherein the at least one item includes a sample value corresponding to a luminance sample, the luminance sample being a downsampled chrominance sample.
[0933] Item 90. The method according to any one of items 1-89, wherein the permission and / or use of the filter-based mode is based on at least one of the following: the prediction mode of the current video block, the transform type of the current video block, the number of non-zero coefficients of the current video block, the segmentation tree type of the current video block, the strip type, the color format, or the block dimension of the current video block.
[0934] Item 91. The method according to Item 90, wherein the prediction mode includes at least one of the following: affine mode, decoder-size motion vector refinement, bidirectional prediction, unidirectional prediction, sub-block-based prediction, bidirectional optical flow (BDOF), overlapping block motion compensation (OBMC), intra-frame and inter-frame joint prediction (CIIP), or local illumination compensation (LIC).
[0935] Item 92. The method according to any one of items 1-92, wherein the filter-based non-intra-frame or intra-frame prediction is based on one of the following: filter-based prediction, cross-component linear model (CCLM), a variant of CCLM, multi-model linear model (MMLM), a variant of MMLM, convolutional cross-component model (CCCM), a variant of CCCM, gradient linear model (GLM) or a variant of GLM, cross-component model-based residual encoding / decoding (CCRM), a variant of said CCRM, a convolutional filter, a variant of said convolutional filter, a filter based on a Gaussian elimination solver or a variant thereof.
[0936] Item 93. The method according to Item 92, wherein the CCCM includes at least one of the following: a CCCM for intra-frame, a CCCM for intra-block copy (IBC), or a CCCM for intra-template matching prediction (intraTMP).
[0937] Item 94. The method according to Item 92, wherein the CCRM includes at least one of the following: inter-frame CCCM, intra-frame block copy (IBC) CCCM, or non-intra-frame CCCM.
[0938] Item 95. The method according to any one of items 1-94, wherein the method is used in a single tree or a double tree.
[0939] Item 96. The method according to any one of items 1-95, wherein the method is used in an inter-frame stripe, the inter-frame stripe comprising at least one of: a B stripe or a P stripe.
[0940] Item 97. The method according to any one of items 1-95, wherein the method is used in an intra-frame stripe, the intra-frame stripe comprising an I-strip.
[0941] Item 98. The method according to any one of items 1-97, wherein the block vector (BV) involved in the method is replaced by a motion vector (MV).
[0942] Item 99. The method according to any one of items 1-98, wherein the training sample or reference sample used in the method refers to at least one of the prediction sample or reconstructed sample in the training region or reference region.
[0943] Item 100. The method according to any one of items 1-99, wherein an indication of whether and / or how the method is applied is indicated at one of the following: sequence level, picture group level, picture level, strip level, or slice group level.
[0944] Item 101. The method according to any one of items 1-99, wherein an indication of whether and / or how to apply the method is indicated in one of the following: sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), dependency parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), strip header, or slice header.
[0945] Item 102. The method according to any one of items 1-99, wherein an indication of whether and / or how the method is applied is included in one of the following: prediction block (PB), transform block (TB), codec block (CB), prediction unit (PU), transform unit (TU), codec unit (CU), virtual pipeline data unit (VPDU), codec tree unit (CTU), CTU row, strip, slice, sub-picture, or region containing more than one sample point or pixel.
[0946] Item 103. The method according to any one of items 1-99 further comprises: determining whether and / or how to apply the method based on the encoded and decoded information of the current video block, the encoded and decoded information including at least one of the following: block size, color format, single-tree segmentation and / or dual-tree segmentation, color components, stripe type, or picture type.
[0947] Item 104. The method according to any one of items 1-103, wherein the conversion includes encoding the current video region into the bitstream.
[0948] Item 105. The method according to any one of items 1-103, wherein the conversion includes decoding the current video region from the bitstream.
[0949] Item 106. An apparatus for video processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform a method according to any one of items 1-105.
[0950] Item 107. A non-transitory computer-readable storage medium storing instructions that cause a processor to execute the method according to any one of items 1-105.
[0951] Item 108. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by means of a video processing apparatus, wherein the method comprises: determining a filter-based non-intra-frame or intra-frame prediction of a current video region of the video based on one of: a transform unit (TU), a codec unit (CU), a prediction unit (PU), or a sub-block; and generating the bitstream based on the filter-based non-intra-frame or intra-frame prediction.
[0952] Item 109. A method for storing a bitstream of video, comprising: determining a filter-based non-intra-frame or intra-frame prediction of a current video region of the video based on one of: a transform unit (TU), a codec unit (CU), a prediction unit (PU), or a sub-block; generating the bitstream based on the filter-based non-intra-frame or intra-frame prediction; and storing the bitstream in a non-transitory computer-readable recording medium.
[0953] Example device Figure 33 A block diagram of a computing device 3300 in which various embodiments of the present disclosure may be implemented is shown. The computing device 3300 may be implemented as a source device 110 (or video encoder 114 or 200) or a destination device 120 (or video decoder 124 or 300), or may be included in a source device 110 (or video encoder 114 or 200) or a destination device 120 (or video decoder 124 or 300).
[0954] It should be understood that, Figure 33 The computing device 3300 shown is for illustrative purposes only and is not intended to imply any limitation on the functionality and scope of the embodiments of this disclosure.
[0955] like Figure 33 As shown, computing device 3300 includes general-purpose computing device 3300. Computing device 3300 may include at least one or more processors or processing units 3310, memory 3320, storage unit 3330, one or more communication units 3340, one or more input devices 3350, and one or more output devices 3360.
[0956] In some embodiments, the computing device 3300 can be implemented as any user terminal or server terminal with computing capabilities. The server terminal can be a server, a large computing device, etc., provided by a service provider. The user terminal can be, for example, any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, stations, units, devices, multimedia computers, multimedia tablet computers, internet nodes, communicators, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices, or any combination thereof. It is conceivable that the computing device 3300 can support any type of interface to the user (such as "wearable" circuitry devices, etc.).
[0957] Processing unit 3310 can be a physical processor or a virtual processor, and can perform various processes based on programs stored in memory 3320. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of computing device 3300. Processing unit 3310 may also be referred to as a central processing unit (CPU), microprocessor, controller, or microcontroller.
[0958] Computing device 3300 typically includes various computer storage media. Such media can be any media accessible by computing device 3300, including but not limited to volatile and non-volatile media, or removable and non-removable media. Memory 3320 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), or flash memory) or any combination thereof. Storage cell 3330 can be any removable or non-removable media and may include machine-readable media, such as memory, flash drives, disks, or other media that can be used to store information and / or data and can be accessed within computing device 3300.
[0959] The computing device 3300 may also include additional removable / non-removable storage media, volatile / non-volatile storage media. Although in Figure 33 Not shown, but a disk drive for reading from and / or writing to a removable non-volatile disk, and an optical disc drive for reading from and / or writing to a removable non-volatile optical disc may be provided. In this case, each drive may be connected to a bus (not shown) via one or more data media interfaces.
[0960] Communication unit 3340 communicates with another computing device via a communication medium. Furthermore, the functionality of the components in computing device 3300 can be implemented by a single computing cluster or multiple computing machines that can communicate via communication connections. Therefore, computing device 3300 can operate in a networked environment using logical connections to one or more other servers, networked personal computers (PCs), or other general-purpose network nodes.
[0961] Input device 3350 can be one or more of various input devices, such as a mouse, keyboard, trackball, voice input device, etc. Output device 3360 can be one or more of various output devices, such as a monitor, speaker, printer, etc. With the aid of communication unit 3340, computing device 3300 can also communicate with one or more external devices (not shown), such as storage devices and display devices. Computing device 3300 can also communicate with one or more devices that enable a user to interact with computing device 3300, or, if necessary, with any device (e.g., network card, modem, etc.) that enables computing device 3300 to communicate with one or more other computing devices. Such communication can be performed via an input / output (I / O) interface (not shown).
[0962] In some embodiments, some or all of the components of computing device 3300 may be deployed in a cloud computing architecture, rather than being integrated into a single device. In a cloud computing architecture, components may be remotely provided and work together to achieve the functionality described herein. In some embodiments, cloud computing provides computing, software, data access, and storage services without requiring end users to know the physical location or configuration of the systems or hardware providing these services. In various embodiments, cloud computing provides services via a wide area network (WAN), such as the Internet, using suitable protocols. For example, a cloud computing provider provides applications via a WAN that can be accessed through a web browser or any other computing component. The software or components of the cloud computing architecture, along with the corresponding data, may be stored on servers at remote locations. Computing resources in a cloud computing environment may be consolidated or distributed across remote data center locations. Cloud computing infrastructure may provide services through shared data centers, although they appear as a single access point to users. Therefore, a cloud computing architecture can be used to provide the components and functionality described herein from service providers at remote locations. Alternatively, the components and functionality described herein may be provided by conventional servers or installed directly or otherwise on client devices.
[0963] In embodiments of this disclosure, computing device 3300 may be used to implement video encoding / decoding. Memory 3320 may include one or more video codec modules 3325 having one or more program instructions. These modules are accessible and executable by processing unit 3310 to perform the functions of the various embodiments described herein.
[0964] In an example embodiment of performing video encoding, input device 3350 may receive video data as input 3370 to be encoded. The video data may be processed, for example, by video codec module 3325 to generate an encoded bitstream. The encoded bitstream may be provided as output 3380 via output device 3360.
[0965] In an example embodiment of performing video decoding, input device 3350 may receive an encoded bitstream as input 3370. The encoded bitstream may be processed, for example, by video codec module 3325 to generate decoded video data. The decoded video data may be provided as output 3380 via output device 3360.
[0966] While this disclosure has been specifically shown and described with reference to preferred embodiments, those skilled in the art will understand that various changes in form and detail may be made without departing from the spirit and scope of this application as defined by the appended claims. These variations are intended to be covered by the scope of this application. Therefore, the foregoing description of embodiments of this application is not intended to be limiting.
Claims
1. A method for video processing, comprising: For the conversion between the current video region and the bitstream of the video, a filter-based non-intra-frame or intra-frame prediction of the current video region is determined based on one of the following: Transform Unit (TU), Encoder / Decoder Unit (CU), Prediction Unit (PU), or sub-block; and The transformation is performed based on the filter-based non-intra-frame or intra-frame prediction.
2. The method of claim 1, wherein the filter-based non-intra-frame or intra-frame prediction is based on the TU or the CU or the PU, and the TU or the CU or the PU is not divided into sub-blocks for filtering.
3. The method of claim 1 or 2, wherein for the TU, the CU, or the PU, the filter model is modulated using a set of training samples with respect to the TU, the CU, or the PU, and the filter model is linear or nonlinear.
4. The method of claim 3, wherein a set of model coefficients of the filter model is determined, and the filter model is applied to determine an estimate of the predicted or residual samples of the TU, the CU, or the PU.
5. The method according to any one of claims 1-4, wherein at least one of cross-component model-based residual coding and decoding (CCRM), inter-frame convolutional cross-component model (CCCM), or intra-frame block copy (IBC) CCCM is applied at a level based on the TU, the PU, or the CU, and the TU, the PU, or the CU is not divided into sub-blocks.
6. The method of claim 5, wherein the coefficients of at least one of the CCRM, the inter-frame CCCM, or the IBC CCCM are determined based on training samples at the TU, PU, or CU level, and at least one of the CCRM, the inter-frame CCCM, or the IBC CCCM is applied to determine the estimation of the predicted samples or residual samples of the TU, the PU, or the CU.
7. The method of claim 6, wherein all applicable samples in the TU, PU, or CU are determined based on the same filter model having the same model coefficients.
8. The method according to any one of claims 5-7, wherein at least one of the CCRM, the inter-frame CCCM, or the IBC CCCM based on the TU, the PU, or the CU is suitable for a plurality of block sizes allowed for the conversion.
9. The method of claim 8, wherein the CCRM or inter-frame CCCM or IBC CCCM is permitted for the current video region, and the filter model is modulated at the level of the current video region, regardless of the size of the current video region, wherein the current video region is one of the following: the CU, the PU, or the TU.
10. The method of claim 8 or 9, wherein the model derivation is applied to the CU or the PU or the TU, and the derived model is applied to applicable samples in the CU or the PU or the TU.
11. The method of claim 1, wherein the filter-based non-intra-frame or intra-frame prediction is performed based on the sub-block.
12. The method of claim 11, wherein the CU or the PU or the TU is divided into a plurality of sub-blocks, the plurality of sub-blocks having a plurality of filter models, wherein the filter models in the plurality of filter models have at least one of the following: model coefficients, filter shape, or filter terms.
13. The method according to claim 11 or 12, wherein the size of the plurality of sub-blocks is fixed or predefined.
14. The method of claim 13, wherein at least one of the width, height, or number of samples of the plurality of sub-blocks is fixed or predefined.
15. The method according to any one of claims 11-14, wherein the training samples of each sub-block are different.
16. The method of claim 15, wherein the training samples of each sub-block are determined based on the corresponding sub-block in a reference CU or reference PU or reference TU, wherein the reference CU or the reference PU or the reference TU is divided into sub-blocks for training sample derivation.
17. The method of claim 15, wherein the training samples of each sub-block are determined based on a subset of samples adjacent to the current CU, the current PU, the current TU, the reference CU, the reference PU, or the reference TU.
18. The method of claim 17, wherein the partial sample points correspond to each sub-block.
19. The method of claim 15, wherein the training samples of each sub-block are determined based on neighboring samples that are adjacent or not adjacent to the current CU, the current PU, the current TU, the reference CU, the reference PU, or the reference TU.
20. The method according to any one of claims 11-19, wherein the block vector-guided convolutional cross-component model (CCCM) for intra-block copying (IBC) or intra-template matching prediction (intraTPM) is performed based on sub-blocks.
21. The method according to any one of claims 11-20, wherein at least one of the following is performed based on sub-blocks: Convolutional Cross-Component Model (CCCM), Cross-Component Linear Model (CCLM), Intra-Template Matching Prediction (intraTPM) filter, Intra-Block Copy (IBC) filter, or variant filter mode.
22. The method according to any one of claims 11-21, wherein the filter-based non-intra-frame or intra-frame prediction for the current video region is performed based on adaptive sub-block granularity.
23. The method of claim 22, wherein the video unit is divided into a plurality of sub-blocks having the same size.
24. The method of claim 22, wherein the plurality of video units are divided into a plurality of sub-blocks, the plurality of sub-blocks having different sub-block size granularities.
25. The method of claim 22, wherein at least one of sub-block width or sub-block height is defined for video units of different sizes.
26. The method of claim 25, wherein the number of sub-blocks in width and / or height is defined.
27. The method of claim 22, wherein the rules for obtaining the size of the sub-blocks are predefined.
28. The method of claim 27, wherein the larger video unit has a larger sub-block size, and the smaller video unit has a smaller sub-block size.
29. The method of claim 22, wherein the rules for obtaining the size of the sub-blocks are included in the bitstream.
30. The method of claim 29, wherein whether or not the video unit is divided into sub-blocks is included in the bitstream.
31. The method of claim 30, wherein whether or not the video unit is divided into sub-blocks is included in the bitstream.
32. The method according to any one of claims 11-31, wherein the adaptive sub-block size is allocated based on the width and / or height of the TU, PU, or CU of the current video region.
33. The method of claim 32, wherein the adaptive sub-block size is based on the width and / or height of the current video region.
34. The method of claim 32, wherein the adaptive sub-block size is based on the total number of samples in the current video region.
35. The method according to any one of claims 11-34, wherein whether and how the sub-block size is determined is based on encoding / decoding information at the encoder and decoder for the transformation.
36. The method of claim 35, wherein the determination of whether to use a sub-block-based filter or a TU, CU, or PU level filter is implicitly determined based on encoding / decoding information.
37. The method of claim 35, wherein the determination of whether to use a filter based on M1xM2 sub-blocks or a filter based on N1xN2 sub-blocks is implicitly determined based on encoding / decoding information, and M1, M2, N1, and N2 are positive integers.
38. The method of claim 37, wherein M1 is one of: 16, 8, 32, TU, CU, or PU. M2 is one of the following: 16, 8, 32, TU, CU, or PU. N1 is one of the following: 16, 8, 32, TU, CU, or PU. N2 is one of the following: 16, 8, 32, TU, CU, or PU. M1 is not equal to N1, and / or M2 is not equal to N2.
39. The method of claim 35, wherein whether and how the sub-block size is determined is based on a template cost-based scheme.
40. The method of claim 39, wherein the template cost is determined based on minimizing one of the following: sum of absolute differences (SAD), sum of absolute transform differences (SATD), sum of squared errors (SSE), or mean squared error (MSE) between estimated sample values and reconstructed sample values from the filter model, wherein the sample references at least one training sample among training samples.
41. The method of claim 39 or 40, wherein the scheme with lower template cost is selected as the scheme to be applied to the current video region.
42. The method of claim 35, wherein whether and how the size of the sub-block is determined is based on information from a reference image.
43. The method of claim 42, wherein the information of the reference image includes at least one of the following: the sequential image (POC) distance between the current image and the current image, or the reference index of the reference image.
44. The method according to any one of claims 11-34, wherein whether and how the sub-block size is determined is indicated in the bitstream.
45. The method according to any one of claims 11-44, wherein whether and how the sub-block-based modeling filter is applied is based on an indicator of one of the following: Overlapping Block Motion Compensation (OBMC) Local illumination compensation (LIC) Intra-block copy (IBC) Palette encoding / decoding tools Block-based incremental pulse code modulation (BDPCM) Intra-template matching prediction (intraTMP) Decoder-side motion vector refinement (DMVR) Inter-frame template matching (TM) Affine tools Sub-block-based temporal motion vector prediction (sbTMVP) Intra-frame sub-segmentation (ISP) Geometric Partitioning Pattern (GPM) Intra-frame and inter-frame joint prediction (CIIP) Airspace GPM (SGPM) Subblock Transformation (SBT) Merge tool, or Advanced Motion Vector Prediction (AMVP).
46. The method according to any one of claims 1-45, further comprising: The samples of the current video region are processed by the first filter; as well as The second filter is applied to the processed sample points.
47. The method of claim 46, wherein the input of the second filter includes the output of the first filter.
48. The method of claim 46, wherein the input of the second filter comprises samples filtered by the first filter.
49. The method according to any one of claims 46-48, wherein the first filter is based on a linear or nonlinear model, the model including one of the following: residual coding-decoding (CCRM) based on a cross-component model, intra-convolutional cross-component model (CCCM), inter-frame CCCM, intra-block copy (IBC) CCCM, intra-template matching prediction (intraTMP) CCCM, EIF, or cross-component linear model (CCLM).
50. The method according to any one of claims 46-49, wherein the first filter is applied at the sub-block level.
51. The method of claim 50, wherein the TU or PU or CU is divided into a plurality of sub-blocks for filtering, the plurality of sub-blocks comprising 16×16 sub-blocks.
52. The method according to any one of claims 46-51, wherein the second filter comprises at least one of: a smoothing filter or a boundary filter.
53. The method according to any one of claims 46-52, wherein the second filter is applied to boundary samples located at the boundary of the sub-block.
54. The method of claim 53, wherein a K-tap filter is applied to samples located in the first row and / or first column and / or last row and / or last column of each sub-block, where K is 2.
55. The method according to any one of claims 46-53, wherein the input of the second filter includes at least one sample point within the sub-block to be processed, the at least one sample point being within a neighboring sub-block of the sub-block to be processed.
56. The method of claim 55, wherein the output of the sample filtered by the 2-tap boundary filter is based on y = (a0) x0 + a1 The shift is determined by x1 + offset >> shift, where y represents the output value, a0 and a1 are weights or coefficients, x0 and x1 are the input samples within and next to the sub-block to be processed, shift is calculated based on the sum of a0 and a1, and offset is a fixed value.
57. The method of claim 56, wherein the weights or coefficients of the input samples of the second filter are predefined.
58. The method of claim 56, wherein the first weight of the input sample within the sub-block to be processed is greater than the second weight of the input sample adjacent to the sub-block to be processed.
59. The method according to any one of claims 46-58, wherein at least one value of at least one sample point in the current video region is modified by the second filter.
60. The method of claim 59, wherein the second filter is applied to internal samples located within a sub-block.
61. The method of claim 59, wherein the value of the boundary sample at one of the following locations is modified by the second filter: the first row of the sub-block, the first column of the sub-block, the last row of the sub-block, or the last column of the sub-block.
62. The method of claim 59, wherein the value of the boundary sample at one of the following locations is not modified by the second filter: the first row of the sub-block, the first column of the sub-block, the last row of the sub-block, or the last column of the sub-block.
63. The method according to any one of claims 46-62, wherein whether a sample is modified by the second filter is based on the sub-block to which the sample belongs.
64. The method of claim 63, wherein at least one of the first row sample of the first row sub-block, the first column sample of the first column sub-block, the last row sample of the last row sub-block, or the last column sample of the last column sub-block is not filtered by the second filter.
65. The method according to any one of claims 46-64, wherein whether the sample to be processed is further modified by the second filter is based on whether the input sample of the second filter is processed by the first filter.
66. The method of claim 65, wherein the input sample of the second filter is processed by the first filter, and the second filter is applied to modify the value of the sample to be processed.
67. The method of claim 65 or 66, wherein the input sample includes at least one of the following: a sample to be processed on one side of an edge or boundary, or a neighboring sample on the other side of the edge or boundary.
68. The method according to any one of claims 1-67, further comprising: The deblocking process is applied to the current video region based on a linear or nonlinear filter model.
69. The method of claim 68, wherein whether and / or how the deblocking process is applied to boundary or edge samples is determined based on a filter model, said filter model comprising one of: residual coding-decoding (CCRM) based on a cross-component model, intra-convolutional cross-component model (CCCM), inter-frame CCCM, intra-block copy (IBC) CCCM, intra-template matching prediction (intraTMP) CCCM, EIF, or cross-component linear model (CCLM).
70. The method of claim 68 or 69, wherein the deblocking intensity of the deblocking process is determined based on the filter model.
71. The method of claim 70, wherein whether strong or weak deblocking is performed on the edge or boundary is determined based on the filter model.
72. The method of claim 70, wherein whether to perform long deblocking or short deblocking on the edge or boundary is determined based on the filter model.
73. The method according to any one of claims 68-72, wherein the values of the deblocking filter parameters in the deblocking process are determined based on the filter model.
74. The method according to any one of claims 68-72, wherein the values of the deblocking filter parameters in the deblocking process are determined based on whether the prediction or residual of samples along the edge or boundary is based on a sub-block filter model.
75. The method according to any one of claims 68-74, wherein the filter model comprises one of the following: residual code-decoder based on a sub-block-based cross-component model (CCRM), a sub-block-based intra-convolutional cross-component model (CCCM), a sub-block-based intra-block copy (IBC) filter or IBC CCCM, a sub-block-based intra-template matching prediction (intraTMP) filter or CCCM, a sub-block-based EIF, or a sub-block-based cross-component linear model (CCLM).
76. The method according to any one of claims 1-75, wherein the mixed weights or fusion weights of the multiple hypotheses of the multiple hypothesis prediction (MHP) are determined based on encoding and decoding information at the encoder and decoder for the transformation.
77. The method of claim 76, wherein the mixing weights or fusion weights of the MHP are adaptively or instantaneously determined for each video unit.
78. The method of claim 76 or 77, wherein the mixing weights or fusion weights of the MHP are determined based on a linear or nonlinear filter model.
79. The method of claim 78, wherein the filter coefficients include the mixing weights or the fusion weights to fuse the different assumptions of the MHP.
80. The method of claim 78, wherein the filter coefficients are determined based on one of Gaussian elimination or an LDL-based scheme.
81. The method of claim 78, wherein the filter coefficients of the model are determined based on a set or group of training samples or template samples.
82. The method of claim 81, wherein the training samples are constructed based on samples in the template.
83. The method of claim 81, wherein the template is constructed using at most M rows above and at most N columns to the left adjacent to the current video region, where M and N are positive integers, or The template is constructed using a maximum of M rows above and a maximum of N columns to the left of the assumption cell adjacent to the MHP assumption.
84. The method of claim 83, wherein M is equal to 6 or 4 of the luminance samples, and / or N is equal to 6 or 4 of the luminance samples.
85. The method of claim 83, wherein M=N.
86. The method according to any one of claims 76-85, wherein the video unit encoded and decoded by MHP is determined based on intra-frame prediction or intra-frame template matching prediction (intraTMP).
87. The method according to any one of claims 76-85, wherein the video unit encoded and decoded by MHP is determined based on intra-block copy (IBC) prediction.
88. The method according to any one of claims 1-87, wherein at least one of sample values, gradient, or position information is used to determine a filter, said filter comprising at least one item corresponding to sample values of different color components.
89. The method of claim 88, wherein the at least one item includes a sample value corresponding to a luminance sample, the luminance sample being a downsampled chrominance sample.
90. The method according to any one of claims 1-89, wherein the permission and / or use of the filter-based mode is based on at least one of the following: the prediction mode of the current video block, the transform type of the current video block, the number of non-zero coefficients of the current video block, the segmentation tree type of the current video block, the stripe type, the color format, or the block dimension of the current video block.
91. The method of claim 90, wherein the prediction mode comprises at least one of the following: Affine mode, decoder-size motion vector refinement, bidirectional prediction, unidirectional prediction, sub-block-based prediction, bidirectional optical flow (BDOF), overlapping block motion compensation (OBMC), intra-frame and inter-frame joint prediction (CIIP), or local illumination compensation (LIC).
92. The method according to any one of claims 1-92, wherein the filter-based non-intra-frame or intra-frame prediction is based on one of the following: Filter-based prediction, cross-component linear model (CCLM), variants of CCLM, multi-model linear model (MMLM), variants of MMLM, convolutional cross-component model (CCCM), variants of CCCM, gradient linear model (GLM) or variants of GLM, cross-component model-based residual encoding and decoding (CCRM), variants of said CCRM, convolutional filter, variants of said convolutional filter, Gaussian elimination-based filter or variants thereof.
93. The method of claim 92, wherein the CCCM comprises at least one of the following: an intra-frame CCCM, an intra-frame block copy (IBC) CCCM, or an intra-frame template matching prediction (intraTMP) CCCM.
94. The method of claim 92, wherein the CCRM comprises at least one of the following: inter-frame CCCM, intra-frame block copy (IBC) CCCM, or non-intra-frame CCCM.
95. The method according to any one of claims 1-94, wherein the method is used in a single tree or a dual tree.
96. The method according to any one of claims 1-95, wherein the method is used in an inter-frame stripe, the inter-frame stripe comprising at least one of: a B stripe or a P stripe.
97. The method according to any one of claims 1-95, wherein the method is used in an intra-frame stripe, the intra-frame stripe comprising an I-strip.
98. The method according to any one of claims 1-97, wherein the block vector (BV) involved in the method is replaced by a motion vector (MV).
99. The method according to any one of claims 1-98, wherein the training sample or reference sample used in the method refers to at least one of the prediction sample or reconstructed sample in the training region or reference region.
100. The method according to any one of claims 1-99, wherein an indication of whether and / or how to apply the method is indicated in one of the following places: sequence level, Image group level, Image level, strip level, or Film series level.
101. The method according to any one of claims 1-99, wherein an indication of whether and / or how to apply the method is indicated in one of the following: Sequence header, Image header, Sequence Parameter Set (SPS) Video Parameter Set (VPS) Dependency Parameter Set (DPS) Decoding Capability Information (DCI) Image Parameter Set (PPS) Adaptive Parameter Set (APS) strip head, or The beginning of the film.
102. The method according to any one of claims 1-99, wherein an indication of whether and / or how to apply the method is included in one of the following: Predicted blocks (PB) Transform Block (TB) Code Block (CB) Prediction Unit (PU) Transformer Unit (TU) Codec Unit (CU) Virtual Pipeline Data Unit (VPDU) Code-decode tree unit (CTU) CTU line, strips, piece, Sub-images, or A region containing more than one sample point or pixel.
103. The method according to any one of claims 1-99, further comprising: The method is determined based on the encoded and decoded information of the current video block, wherein the encoded and decoded information includes at least one of the following: Block size, Color format, Single-tree segmentation and / or dual-tree segmentation Color components, Strip type, or Image type.
104. The method according to any one of claims 1-103, wherein the conversion includes encoding the current video region into the bitstream.
105. The method according to any one of claims 1-103, wherein the conversion includes decoding the current video region from the bitstream.
106. An apparatus for video processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of claims 1-105.
107. A non-transitory computer-readable storage medium storing instructions that cause a processor to execute the method according to any one of claims 1-105.
108. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by means of a video processing apparatus, wherein the method includes: The current video region of the video is determined based on filter-based non-intra-frame or intra-frame prediction using one of the following: Transform Unit (TU), Encoder / Decoder Unit (CU), Prediction Unit (PU), or sub-block; and The bitstream is generated based on the filter-based non-intra-frame or intra-frame prediction.
109. A method for storing a bitstream of video, comprising: The current video region of the video is determined based on filter-based non-intra-frame or intra-frame prediction of one of the following: transform unit (TU), codec unit (CU), prediction unit (PU), or sub-block; The bitstream is generated based on the filter-based non-intra-frame or intra-frame prediction; and The bitstream is stored in a non-transitory computer-readable recording medium.