Method and device for video processing and medium
Through the offset-based regression model and local lighting compensation (LIC) video block prediction method, combined with intra-frame fusion of multiple reference lines, the problem of insufficient encoding and decoding efficiency of existing video encoding and decoding technologies is solved, and a more efficient encoding and decoding effect is achieved.
Patent Information
- Application Number
- CN202480007446.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-01-13
- Filing Date
- 2024-01-11
- Publication Date
- 2025-08-19
AI Technical Summary
The encoding and decoding efficiency of existing video encoding and decoding technologies needs to be further improved.
Offset-based regression model and local lighting compensation (LIC) are used to determine the prediction of video blocks, combining intra-frame fusion of more than two reference lines to improve the encoding and decoding efficiency.
The encoding and codec efficiency and effectiveness of video encoding and codec are improved through improved prediction methods.
Smart Images

Figure CN120513626A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate generally to video processing techniques, and more particularly, to regression models for prediction. Background Art
[0002] Digital video capabilities are now being used in every aspect of our lives. Various video compression technologies have been proposed for video encoding and decoding, including MPEG-2, MPEG-4, ITU-T H.263, ITU-T H.264 / MPEG-4 Part 10 Advanced Video Codec (AVC), ITU-T H.265 High Efficiency Video Codec (HEVC), and Versatile Video Codec (VVC). However, further improvements in the encoding and decoding efficiency of video encoding and decoding technologies are often desired. Summary of the Invention
[0003] Embodiments of the present disclosure provide a solution for video processing.
[0004] In a first aspect, a method for video processing is provided. The method includes: for conversion between a current video block of a video and a bitstream of the video, determining a filtered prediction of the current video block based on an offset-based regression model that modulates a relationship between a current region of the current video block and a reference region; and performing conversion based on the filtered prediction. The method according to the first aspect of the present disclosure uses the offset-based regression model to determine the filtered prediction. This improves codec efficiency and codec effectiveness.
[0005] In a second aspect, another method for video processing is provided. The method includes: determining a prediction for a current video block of a video based on local illumination compensation (LIC) for conversion between a current video block and a video bitstream, the LIC being based on a regression model that modulates a relationship between a current region of the current video block and a reference region for the LIC; and performing conversion based on the prediction. The method according to the second aspect of the present disclosure utilizes the LIC based on the regression model. This improves codec efficiency and effectiveness.
[0006] In a third aspect, another method for video processing is provided. The method includes: applying at least one intra-frame blending for at least one color component to the current video block based on more than two reference lines for conversion between a current video block and a video bitstream; and performing the conversion based on the application. The method according to the third aspect of the present disclosure applies the intra-frame blending for the color component based on multiple reference lines. This improves codec efficiency and codec effectiveness.
[0007] In a fourth aspect, a device for video processing is provided. The device includes a processor and a non-volatile memory having instructions thereon. The instructions, when executed by the processor, cause the processor to perform the method according to the first aspect, the second aspect, or the third aspect of the present disclosure.
[0008] In a fifth aspect, a non-transitory computer-readable storage medium is provided, wherein the non-transitory computer-readable storage medium stores instructions for causing a processor to execute the method according to the first aspect, the second aspect, or the third aspect of the present disclosure.
[0009] In a sixth aspect, a non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a bitstream of a video generated by a method performed by an apparatus for video processing. The method includes: determining a filtered prediction of a current video block of the video based on an offset-based regression model, the offset-based regression model modulating a relationship between a current region and a reference region of the current video block; and generating a bitstream based on the filtered prediction.
[0010] In a seventh aspect, a method for storing a bitstream of a video is provided. The method includes determining a filtered prediction of a current video block of the video based on an offset-based regression model that modulates a relationship between a current region and a reference region of the current video block; generating a bitstream based on the filtered prediction; and storing the bitstream in a non-transitory computer-readable recording medium.
[0011] In an eighth aspect, another non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a bitstream of a video generated by a method performed by an apparatus for video processing. The method includes: determining a prediction of a current video block of the video based on local illumination compensation (LIC), the LIC being based on a regression model that modulates a relationship between a current region and a reference region of the current video block for the LIC; and generating a bitstream based on the prediction.
[0012] In a ninth aspect, another method for storing a bitstream of a video is provided. The method includes: determining a prediction of a current video block of the video based on local illumination compensation (LIC), the LIC being based on a regression model that modulates a relationship between a current region and a reference region of the current video block for the LIC; generating a bitstream based on the prediction; and storing the bitstream in a non-transitory computer-readable recording medium.
[0013] In a tenth aspect, another non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a bitstream of a video generated by a method performed by an apparatus for video processing. The method includes: applying at least one intra-frame blending for at least one color component to a current video block of the video based on more than two reference lines; and generating a bitstream based on the application.
[0014] In an eleventh aspect, another method for storing a bitstream of a video is provided. The method includes: applying at least one intra-frame blending for at least one color component to a current video block of the video based on more than two reference lines; generating a bitstream based on the application; and storing the bitstream in a non-transitory computer-readable recording medium.
[0015] This summary is intended to introduce a selection of concepts in a simplified form that are further described below in the detailed description. This summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The above and other objects, features and advantages of the exemplary embodiments of the present disclosure will become more apparent through the following detailed description with reference to the accompanying drawings.In the exemplary embodiments of the present disclosure, the same reference numerals generally refer to the same components.
[0017] Figure 1 A block diagram illustrating an example video encoding and decoding system is shown according to some embodiments of the present disclosure;
[0018] Figure 2 shows a block diagram illustrating a first example video encoder according to some embodiments of the present disclosure;
[0019] Figure 3 shows a block diagram illustrating an example video decoder according to some embodiments of the present disclosure;
[0020] Figure 4A and Figure 4B shows the effect of the slope adjustment parameter "u", where Figure 4A corresponds to a model created using the current CCLM, and Figure 4B corresponds to the updated model as proposed;
[0021] Figure 5 shows the neighboring blocks (L, A, BL, AR, AL) used in the derivation of the general MPM list;
[0022] Figure 6 shows adjacent reconstructed samples used for DIMD chroma mode;
[0023] Figure 7 The intra-frame template matching search area used is shown;
[0024] Figure 8A and Figure 8B The division methods for angle modes are shown respectively;
[0025] Figure 9 The expanded MRL candidate list is shown;
[0026] Figure 10 The spatial portion of the convolution filter is shown;
[0027] Figure 11 shows the reference region (and its filling) used to derive the filter coefficients;
[0028] Figure 12 Four Sobel-based gradient modes for GLM are shown;
[0029] Figure 13 The template area is shown;
[0030] Figure 14 Shows the location of the spatial merge candidate;
[0031] Figure 15 shows the candidate pairs considered for redundancy check of spatial Merge candidates;
[0032] Figure 16 Shows motion vector scaling for temporal Merge candidates;
[0033] Figure 17 The candidate positions for the time domain Merge candidate, C0 and C1, are shown;
[0034] Figure 18 The MMVD search point is shown;
[0035] Figure 19 shows the expanded CU area used in BDOF;
[0036] Figure 20 A schematic diagram for a symmetric MVD mode is shown;
[0037] Figure 21 Decoding side motion vector refinement is shown;
[0038] Figure 22 Shown are the top and left neighboring blocks used in the derivation of CIIP weights;
[0039] Figure 23 An example of GPM partitioning grouped at the same angle is shown;
[0040] Figure 24 shows the unidirectional prediction MV selection for geometric partitioning mode;
[0041] Figure 25 An exemplary generation of blending weights w0 using a geometric partitioning mode is shown;
[0042] Figure 26 A flowchart of a method for video processing according to an embodiment of the present disclosure is shown;
[0043] Figure 27 A flowchart of a method for video processing according to an embodiment of the present disclosure is shown;
[0044] Figure 28 A flowchart showing a method for video processing according to an embodiment of the present disclosure is shown; and
[0045] Figure 29 A block diagram is shown of a computing device in which various embodiments of the present disclosure may be implemented.
[0046] Throughout the drawings, the same or similar reference numbers generally refer to the same or similar elements. DETAILED DESCRIPTION
[0047] The principles of the present disclosure will now be described with reference to some embodiments. It should be understood that these embodiments are described only for the purpose of illustrating and helping those skilled in the art to understand and implement the present disclosure, and do not imply any limitation on the scope of the present disclosure. In addition to the methods described below, the disclosure described herein can also be implemented in various ways.
[0048] In the following description and claims, unless defined otherwise, all scientific and technical terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs.
[0049] References in this disclosure to "one embodiment," "an embodiment," "an example embodiment," and the like indicate that the described embodiment may include a particular feature, structure, or characteristic, but not every embodiment is required to include that particular feature, structure, or characteristic. Furthermore, these phrases do not necessarily refer to the same embodiment. Furthermore, when a particular feature, structure, or characteristic is described in conjunction with an example embodiment, it is intended that such feature, structure, or characteristic, whether or not explicitly described, be applicable to other embodiments and that it is within the knowledge of those skilled in the art to apply such feature, structure, or characteristic.
[0050] It should be understood that although the terms "first" and "second" and the like may be used herein to describe various elements, these elements should not be limited to these terms. These terms are only used to distinguish one element from another. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element without departing from the scope of the example embodiments. As used herein, the term "and / or" includes any and all combinations of one or more of the listed terms.
[0051] The terms used herein are used only for the purpose of describing specific embodiments and are not intended to limit the example embodiments. As used herein, the singular forms "a," "an," and "the" are also intended to include the plural forms, unless the context clearly indicates otherwise. It should also be understood that the terms "comprise," "including," "having," "including," and / or "comprising" when used herein indicate the presence of the features, elements, and / or components, etc., but do not exclude the presence or addition of one or more other features, elements, components, and / or combinations thereof. Sample Environment
[0052] Figure 1 is a block diagram illustrating an example video codec system 100 that can utilize the techniques of the present disclosure. As shown, the video codec system 100 may include a source device 110 and a destination device 120. The source device 110 may also be referred to as a video encoding device, and the destination device 120 may also be referred to as a video decoding device. In operation, the source device 110 may be configured to generate encoded video data, and the destination device 120 may be configured to decode the encoded video data generated by the source device 110. The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.
[0053] The video source 112 may include a source such as a video capture device. Examples of a video capture device include, but are not limited to, an interface for receiving video data from a video content provider, a computer graphics system for generating video data, and / or a combination thereof.
[0054] The video data may include one or more pictures. The video encoder 114 encodes the video data from the video source 112 to generate a bitstream. The bitstream may include a sequence of bits that form a coded representation of the video data. The bitstream may include coded pictures and associated data. The coded pictures are coded representations of the pictures. The associated data may include sequence parameter sets, picture parameter sets, and other syntax structures. The I / O interface 116 may include a modulator / demodulator and / or a transmitter. The coded video data may be transmitted directly to the destination device 120 via the network 130A via the I / O interface 116. The coded video data may also be stored on a storage medium / server 130B for access by the destination device 120.
[0055] Destination device 120 may include an I / O interface 126, a video decoder 124, and a display device 122. The I / O interface 126 may include a receiver and / or a modem. The I / O interface 126 may obtain encoded video data from source device 110 or storage medium / server 130B. The video decoder 124 may decode the encoded video data. The display device 122 may display the decoded video data to a user. The display device 122 may be integrated with the destination device 120, or may be external to the destination device 120, the destination device 120 being configured to interface with an external display device.
[0056] The video encoder 114 and the video decoder 124 may operate according to a video compression standard, such as the High Efficiency Video Codec (HEVC) standard, the Versatile Video Codec (VVC) standard, and other existing and / or future standards.
[0057] Figure 2 is a block diagram illustrating an example of a video encoder 200 according to some embodiments of the present disclosure, which may be Figure 1 An example of the video encoder 114 in the system 100 is shown.
[0058] Video encoder 200 may be configured to implement any or all of the techniques of this disclosure. Figure 2 In the example of , video encoder 200 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of video encoder 200. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.
[0059] In some embodiments, the video encoder 200 may include a segmentation unit 201, a prediction unit 202, a residual generation unit 207, a transformation unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transformation unit 211, a reconstruction unit 212, a cache 213 and an entropy coding unit 214, and the prediction unit 202 may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205 and an intra-frame prediction unit 206.
[0060] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an intra block copy (IBC) unit. The IBC unit may perform prediction in an IBC mode in which at least one reference picture is the picture in which the current video block is located.
[0061] Furthermore, although some components (such as the motion estimation unit 204 and the motion compensation unit 205) may be integrated, for the purpose of explanation, these components are described in detail in the following sections. Figure 2 are shown separately in the example.
[0062] The partitioning unit 201 may partition a picture into one or more video blocks. The video encoder 200 and the video decoder 300 may support various video block sizes.
[0063] The mode selection unit 203 can, for example, select one of a plurality of coding modes (intra-frame coding or inter-frame coding) based on the error result, and provide the resulting intra-frame coded block or inter-frame coded block to the residual generation unit 207 to generate residual block data, and to the reconstruction unit 212 to reconstruct the coded block for use as a reference picture. In some examples, the mode selection unit 203 can select a joint intra-frame and inter-frame prediction (CIIP) mode, in which prediction is based on an inter-frame prediction signal and an intra-frame prediction signal. In the case of inter-frame prediction, the mode selection unit 203 can also select a resolution for the motion vector for the block (e.g., sub-pixel precision or integer pixel precision).
[0064] To perform inter-frame prediction on the current video block, the motion estimation unit 204 may generate motion information for the current video block by comparing the current video block with one or more reference frames from the cache 213. The motion compensation unit 205 may determine a predicted video block for the current video block based on the motion information and decoded samples of pictures from the cache 213 other than the picture associated with the current video block.
[0065] The motion estimation unit 204 and the motion compensation unit 205 may perform different operations on the current video block, for example, depending on whether the current video block is in an I slice, a P slice, or a B slice. As used herein, an "I slice" may refer to a portion of a picture consisting of macroblocks, all of which are based on macroblocks within the same picture. Furthermore, as used herein, in some aspects, "P slices" and "B slices" may refer to portions of a picture consisting of macroblocks that are independent of macroblocks in the same picture.
[0066] In some examples, motion estimation unit 204 may perform unidirectional prediction on the current video block, and motion estimation unit 204 may search the reference pictures in list 0 or list 1 to find a reference video block for the current video block. Motion estimation unit 204 may then generate a reference index indicating the reference picture in list 0 or list 1 that contains the reference video block and a motion vector indicating the spatial displacement between the current video block and the reference video block. Motion estimation unit 204 may output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 205 may generate a predicted video block for the current video block based on the reference video block indicated by the motion information for the current video block.
[0067] Alternatively, in other examples, the motion estimation unit 204 may perform bidirectional prediction on the current video block. The motion estimation unit 204 may search the reference pictures in list 0 for a reference video block for the current video block, and may also search the reference pictures in list 1 for another reference video block for the current video block. The motion estimation unit 204 may then generate reference indexes indicating multiple reference pictures in list 0 and list 1 containing multiple reference video blocks, and motion vectors indicating multiple spatial displacements between the multiple reference video blocks and the current video block. The motion estimation unit 204 may output the multiple reference indexes and multiple motion vectors for the current video block as motion information for the current video block. The motion compensation unit 205 may generate a predicted video block for the current video block based on the multiple reference video blocks indicated by the motion information of the current video block.
[0068] In some examples, motion estimation unit 204 may output a complete set of motion information for use in the decoding process of a decoder. Alternatively, in some embodiments, motion estimation unit 204 may signal the motion information of the current video block with reference to the motion information of another video block. For example, motion estimation unit 204 may determine that the motion information of the current video block is sufficiently similar to the motion information of an adjacent video block.
[0069] In one example, motion estimation unit 204 may indicate to video decoder 300 a value in a syntax structure associated with the current video block that indicates the current video block has the same motion information as another video block.
[0070] In another example, the motion estimation unit 204 may identify another video block and a motion vector difference (MVD) in a syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 300 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[0071] As discussed above, the video encoder 200 may signal motion vectors in a predictive manner.Two examples of prediction signaling techniques that may be implemented by the video encoder 200 include Advanced Motion Vector Prediction (AMVP) and Merge mode signaling.
[0072] The intra-frame prediction unit 206 can perform intra-frame prediction on the current video block. When the intra-frame prediction unit 206 performs intra-frame prediction on the current video block, the intra-frame prediction unit 206 can generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block can include a predicted video block and various syntax elements.
[0073] The residual generation unit 207 can generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) the predicted video block(s) of the current video block from the current video block. The residual data for the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.
[0074] In other examples, such as in skip mode, there may be no residual data for the current video block, and the residual generation unit 207 may not perform a subtraction operation.
[0075] Transform processing unit 208 may generate one or more transform coefficient video blocks for a current video block by applying one or more transforms to a residual video block associated with the current video block.
[0076] After transform processing unit 208 generates a transform coefficient video block associated with the current video block, quantization unit 209 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0077] The inverse quantization unit 210 and the inverse transform unit 211 may apply inverse quantization and inverse transform, respectively, to the transform coefficient video block to reconstruct a residual video block from the transform coefficient video block. The reconstruction unit 212 may add the reconstructed residual video block to corresponding samples from one or more prediction video blocks generated by the prediction unit 202 to generate a reconstructed video block associated with the current video block for storage in the buffer 213.
[0078] After reconstruction unit 212 reconstructs the video block, a loop filtering operation may be performed to reduce video blocking artifacts in the video block.
[0079] The entropy coding unit 214 may receive data from other functional components of the video encoder 200. When the entropy coding unit 214 receives data, the entropy coding unit 214 may perform one or more entropy coding operations to generate entropy-coded data and output a bitstream including the entropy-coded data.
[0080] Figure 3 is a block diagram illustrating an example of a video decoder 300 according to some embodiments of the present disclosure. The video decoder 300 may be Figure 1 An example of the video decoder 124 in the system 100 is shown.
[0081] Video decoder 300 may be configured to perform any or all of the techniques of this disclosure. Figure 3 In the example of FIG, video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of video decoder 300. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.
[0082] exist Figure 3 In the example of FIG. 3 , the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, the video decoder 300 can perform a decoding process that is generally opposite to the encoding process described with respect to the video encoder 200.
[0083] The entropy decoding unit 301 can retrieve an encoded bitstream. The encoded bitstream may include entropy-encoded video data (e.g., encoded blocks of video data). The entropy decoding unit 301 can decode the entropy-encoded video data, and the motion compensation unit 302 can determine motion information from the entropy-decoded video data, which includes motion vectors, motion vector precision, reference picture list indexes, and other motion information. The motion compensation unit 302 can determine such information, for example, by performing AMVP and Merge mode. AMVP is used, which includes deriving several most likely candidates based on data from adjacent PBs and reference pictures. The motion information typically includes horizontal and vertical motion vector displacement values, one or two reference picture indexes, and, in the case of prediction regions in B slices, an identification of which reference picture list is associated with each index. As used herein, in some aspects, "Merge mode" may refer to deriving motion information from adjacent blocks in the spatial or temporal domain.
[0084] The motion compensation unit 302 may generate a motion compensated block, possibly performing interpolation based on an interpolation filter. An identifier for the used interpolation filter with sub-pixel precision may be included in the syntax element.
[0085] Motion compensation unit 302 may calculate interpolated values for sub-integer pixels of a reference block using interpolation filters used by video encoder 200 during encoding of the video block. Motion compensation unit 302 may determine the interpolation filters used by video encoder 200 based on received syntax information, and motion compensation unit 302 may use the interpolation filters to generate a prediction block.
[0086] The motion compensation unit 302 can use at least part of the syntax information to determine the size of the blocks used to encode the (multiple) frames and / or (multiple) slices of the coded video sequence, partition information describing how each macroblock of the picture of the coded video sequence is partitioned, a mode indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-frame coded block, and other information used to decode the coded video sequence. As used herein, in some aspects, a "slice" can refer to a data structure that can be decoded independently of other slices of the same picture in terms of entropy coding and decoding, signal prediction, and residual signal reconstruction. A slice can be an entire picture or a region of a picture.
[0087] The intra prediction unit 303 can form a prediction block from spatially adjacent blocks using, for example, an intra prediction mode received in the bitstream. The inverse quantization unit 304 inversely quantizes, i.e., dequantizes, the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 301. The inverse transform unit 305 applies an inverse transform.
[0088] The reconstruction unit 306 can obtain the decoded block, for example, by adding the residual block to the corresponding prediction block generated by the motion compensation unit 302 or the intra-frame prediction unit 303. If necessary, a deblocking filter can also be used to filter the decoded block to remove blocking artifacts. The decoded video block is then stored in the buffer 307, which provides reference blocks for subsequent motion compensation / intra-frame prediction and also produces the decoded video for presentation on a display device.
[0089] Some exemplary embodiments of the present disclosure are described in detail below. It should be understood that the section titles used in this document are for ease of understanding and do not limit the embodiments disclosed in the section to only that section. In addition, although certain embodiments are described with reference to a multifunctional video codec or other specific video codecs, the disclosed technology is also applicable to other video coding and decoding technologies. In addition, although some embodiments describe the video encoding steps in detail, it should be understood that the corresponding decoding steps corresponding to the de-encoding will be implemented by the decoder. In addition, the term "video processing" includes video encoding or compression, video decoding or decompression, and video transcoding, in which video pixels are represented from one compression format to another compression format or at different compression bit rates. 1. Brief Overview This disclosure relates to video coding technology. Specifically, it relates to linear / nonlinear / polynomial regression model prediction and related offset removal algorithms in image / video coding. It can be applied to existing video coding standards such as HEVC and VVC. It is also applicable to future video coding standards or video codecs. 2. Introduction Video codec standards have evolved primarily through the development of the well-known ITU-T and ISO / IEC standards. ITU-T developed the H.261 and H.263 standards, while ISO / IEC developed MPEG-1 and MPEG-4 Vision. The two organizations jointly developed the H.262 / MPEG-2 Video standard, the H.264 / MPEG-4 Advanced Video Codec (AVC) standard, and the H.265 / HEVC standard. Starting with H.262, video codec standards have been based on a hybrid video codec architecture that utilizes temporal prediction plus transform coding. To explore future video codec technologies beyond HEVC, the Joint Video Exploration Team (JVET) was established in 2015 by VCEG and MPEG. JVET meetings are held quarterly, and the new video codec standard was officially named the Versatile Video Codec (VVC) at the April 2018 JVET meeting. The first version of the VVC Test Model (VTM) was also released at that time. The VVC working draft and test model (VTM) have been updated after each meeting. The VVC project achieved technical completion (FDIS) at the July 2020 meeting. 2.1 Intra-frame Prediction In intra prediction, the minimum chroma intra prediction unit (SCIPU) constraint in VVC is removed. In addition, the VPDU constraint used to reduce CCLM prediction delay is also removed. 2.1.1 Multi-Model LM (MMLM) The CCLM included in VVC is extended by adding three multi-model LM (MMLM) modes. In each MMLM mode, the reconstructed neighboring samples are classified into two categories using a threshold that is the average of the reconstructed neighboring samples of luma. The linear model for each category is derived using the least mean square (LMS) method. For the CCLM mode, the linear model is also derived using the LMS method. Slope adjustment is applied to the cross-component linear model (CCLM) and multi-model LM prediction. The adjustment tilts the linear function that maps luma values to chroma values relative to the center point determined by the average luma value of the reference samples. 2.1.1.1 CCLM Slope Adjustment CCLM uses a 2-parameter model to map luma values to chroma values. The slope parameter "a" and the bias parameter "b" define the mapping as follows: chromaVal=a*lumaVal+b. The adjustment to the slope parameter “u” is signaled to update the model to the following form: chromaVal=a'*lumaVal+b' in a'=a+u, b'=bu*y r . With this choice, the mapping function is centered around the illuminance value y r The average of the reference brightness samples used in model creation is used as y r , in order to provide meaningful modifications to the model. The following figure illustrates this process. Figure 4A and Figure 4B shows the effect of the slope adjustment parameter "u", where Figure 4A corresponds to a model created using the current CCLM, and Figure 4B corresponds to the updated model as proposed. Implementation The slope adjustment parameter is provided as an integer between -4 and 4 (inclusive) and is signaled in the bitstream. The unit of the slope adjustment parameter is 1 / 8 of the chroma sample value for each luma sample value. th (For 10-bit content.) Adjustments apply to CCLM models that use reference samples both above and to the left of the block ("LM_CHROMA_IDX" and "MMLM_CHROMA_IDX"), but not to "one-sided" mode. This choice is based on a trade-off between codec efficiency and complexity. When slope adjustment is applied for a multi-mode CCLM model, both models may be adjusted so that a maximum of two slope updates are signaled for a single chroma block. Encoder Method The proposed encoder method performs a SATD-based search for the optimal value of the slope update for Cr and a similar SATD-based search for Cb. If either term results in a non-zero slope adjustment parameter, the combined slope adjustment pair (SATD-based update for Cr, SATD-based update for Cb) is included in the list of RD checks for the TU. 2.1.2 Gradient PDPC In VVC, for some scenes, PDPC may not be applied due to the unavailability of secondary reference samples. In these cases, a gradient-based PDPC extended from the horizontal / vertical mode is applied. The PDPC weights (wT / wL) and nScale parameters used to determine the attenuation of the PDPC weight relative to the distance from the left / top boundary are set equal to the corresponding parameters in the horizontal / vertical mode, respectively. When the secondary reference samples are located at fractional sample positions, bilinear interpolation is applied. 2.1.3 Secondary MPM A secondary MPM list is introduced. The existing primary MPM (PMPM) list consists of 6 entries, and the secondary MPM (SMPM) list includes 16 entries. First, a general MPM list with 22 entries is constructed, and then the first 6 entries in the general MPM list are included in the PMPM list, and the remaining entries form the SMPM list. The first entry in the general MPM list is the plane mode. The remaining entries are as follows: Figure 5 The shown consists of intra modes for the left (L), above (A), below left (BL), above right (AR), and above left (AL) neighboring blocks, a directional mode with an offset added from the first two available directional modes of the neighboring blocks, and a default mode. If the CU block is vertically oriented, the order of neighboring blocks is A, L, BL, AR, AL; otherwise, it is L, A, BL, AR, AL. Figure 5 Neighboring blocks (L, A, BL, AR, AL) used in the derivation of the general MPM list are shown. First parse the PMPM flag, if it is equal to 1, parse the PMPM index to determine which entry of the PMPM list is selected, otherwise parse the SPMPM flag to determine whether to parse the SMPM index or the remaining mode. 2.1.4 Reference Sample Interpolation and Smoothing for Intra-frame Prediction The 4-tap cubic interpolation is replaced by a 6-tap cubic interpolation filter for the derivation of predicted samples from reference samples. For reference sample filtering, a 6-tap Gaussian filter is applied to larger blocks (W>=32 and H>=32), otherwise the existing VVC 4-tap Gaussian interpolation filter is applied. Extended intra reference samples are derived using a 4-tap interpolation filter instead of nearest neighbor rounding. 2.1.5 Decoder-side Intra Mode Derivation (DIMD) When DIMD is applied, two intra modes are derived from the reconstructed neighboring samples and these two predictions are combined with the planar mode prediction, where the weights are derived from the gradients. The division operation in the weight derivation is performed using the same lookup table (LUT) based integerization scheme used by CCLM. For example, the division operation in the direction calculation Orient=G y / G x Calculated via the following LUT-based scheme: x=Floor(Log2(Gx)) normDiff=((Gx<<4)>>x)&15 x+=(3+(normDiff!=0)?1:0) Orient=(Gy*(DivSigTable[normDiff]|8)+(1<<(x-1)))>>x in DivSigTable
[16] ={0,7,6,5,5,4,4,3,3,2,2,1,1,1,1,0}. The derived intra modes are included into the main list of intra most probable modes (MPMs), so the DIMD process is performed before the MPM list is built. The main derived intra modes of a DIMD block are stored with the block and are used for MPM list construction of neighboring blocks. 2.1.5.1 DIMD Chroma Mode The DIMD colorimetric mode uses the DIMD derivation method based on Figure 6 The chroma intra prediction mode of the current block is derived from the adjacent reconstructed Y, Cb, and Cr samples in the second adjacent row and column shown. Specifically, the horizontal gradient and vertical gradient are calculated for each co-located reconstructed luma sample and reconstructed Cb and Cr samples of the current chroma block to construct the HoG. The intra prediction mode with the largest histogram amplitude value is then used to perform chroma intra prediction for the current chroma block. Figure 6 Neighboring reconstructed samples used for DIMD chroma mode are shown. When the intra prediction mode derived from the DIMD chroma mode is the same as the intra prediction mode derived from the DM mode, the intra prediction mode with the second largest histogram magnitude value is used as the DIMD chroma mode. A CU level flag is signaled to indicate whether the proposed DIMD chroma mode is applied. 2.1.6 Fusion of Chroma Intra Prediction Modes The DM mode and the four default modes can be merged with the MMLM_LT mode as follows: pred=(w0*pred0+w1*pred1+(1<<(displacement-1)))>>displacement Where pred0 is the prediction value obtained by applying the non-LM mode, pred1 is the prediction value obtained by applying the MMLM_LT mode, and pred is the final prediction value of the current chroma block. The two weights w0 and w1 are determined by the intra prediction mode of the adjacent chroma blocks, and the displacement is set to be equal to 2. Specifically, when both the upper and left adjacent blocks are encoded and decoded using the LM mode, {w0, w1} = {1, 3}; when both the upper and left adjacent blocks are encoded and decoded using the non-LM mode, {w0, w1} = {3, 1}; otherwise, {w0, w1} = {2, 2}. For syntax design, if non-LM mode is selected, a flag is signaled to indicate whether fusion is applied. This method is only applicable to I slices. 2.1.7 Intra-frame Template Matching Intra Template Matching Prediction (Intra TMP) is a special intra prediction mode that copies the best prediction block from the reconstructed part of the current frame, whose L-shaped template matches the current template. For a predefined search range, the encoder searches for the template that is most similar to the current template in the reconstructed part of the current frame and uses the corresponding block as the prediction block. The encoder then signals the use of this mode, and the same prediction operation is performed on the decoder side. The prediction signal is obtained by combining the L-shaped causal neighbors of the current block with the Figure 7 generated by matching another block in a predefined search area in the , the predefined search area including: R1: current CTU, R2: Upper left CTU, R3: Upper CTU, R4: left CTU. The sum of absolute differences (SAD) is used as the cost function. In each region, the decoder searches for the template with the smallest SAD relative to the current template and uses its corresponding block as the prediction block. The dimensions of all regions (SearchRange_w, SearchRange_h) are set to be proportional to the block dimensions (BlkW, BlkH) to have a fixed number of SAD comparisons per pixel. That is: SearchRange_w=a*BlkW, SearchRange_h=a*BlkH. Where "a" is a constant that controls the gain / complexity tradeoff. In practice, "a" is equal to 5. Figure 7 The intra template matching search area used is shown. The intra template matching tool is enabled for CUs with width and height dimensions less than or equal to 64. This maximum CU size for intra template matching is configurable. When DIMD is not used for the current CU, the intra template matching prediction mode is signaled at the CU level through a dedicated flag. 2.1.8 Fusion for Template-Based Intra Mode Derivation (TIMD) For each intra prediction mode in the MPM, the SATD between the template's prediction and the reconstructed samples is calculated. The first two intra prediction modes with the smallest SATD are selected as TIMD modes. These two TIMD modes are fused using weights after applying the PDPC process, and such weighted intra prediction is used to encode and decode the current CU. Position-dependent intra prediction combining (PDPC) is included in the derivation of TIMD modes. The costs of the two selected modes are compared with the threshold. In the test, a cost factor of 2 is applied as follows: costMode2<2*costMode1. If this condition is true, then the blend is applied, otherwise only mode1 is used. The weight of a pattern is calculated from its SATD cost as follows: weight1=costMode2 / (costMode1+costMode2), weight2=1-weight1. The division operation is performed using the same lookup table (LUT) based integerization scheme used by CCLM. 2.1.9 Combination of CIIP with TIMD and TM Merge In CIIP mode, prediction samples are generated by weighting the inter prediction signal predicted using CIIP-TM Merge candidates and the intra prediction signal predicted using TIMD-derived intra prediction modes. This method is only applied to codec blocks with an area less than or equal to 1024. The TIMD derivation method is used to derive intra prediction modes in CIIP. Specifically, the intra prediction mode with the smallest SATD value in the TIMD mode list is selected and mapped to one of the 67 conventional intra prediction modes. Figure 8A and Figure 8B The division methods for the angle modes are shown respectively. Furthermore, it is proposed to modify the weights (wIntra, wInter) for both tests if the derived intra prediction mode is an angular mode. For near-horizontal modes (2 <= angular mode index < 34), the current block is split vertically, as Figure 8A As shown; for the near vertical mode (34 <= angle mode index <= 66), the current block is divided horizontally, as Figure 8B shown. (wIntra, wInter) for different sub-blocks are shown in Table 1. Table 1. Modified weights used for angle mode Sub-block index (wIntra,wInter) 0 (6,2) 1 (5,3) 2 (3,5) 3 (2,6) Using CIIP-TM, a CIIP-TM Merge candidate list is constructed for the CIIP-TM mode. Merge candidates are refined by template matching. CIIP-TM Merge candidates are also reordered as regular Merge candidates by the ARMC method. The maximum number of CIIP-TM Merge candidates is equal to 2. 2.1.10 Extended Multiple Reference Line (MRL) List The MRL list in VVC is extended to include more reference lines for intra prediction. The extended reference line list consists of Figure 9 The row indices shown are {1,3,5,7,12}. Figure 9 The extended MRL candidate list is shown.For template-based intra mode derivation (TIMD), only the first two reference row candidates (ie, {1, 3}) are used instead of the complete MRL candidate list. 2.1.11 Convolutional Cross-Component Intra Prediction Model In this method, a convolutional cross-component model (CCCM) is applied to predict chroma samples from reconstructed luma samples, in a similar way to what is done in the current CCLM mode. As with CCLM, when chroma downsampling is used, the reconstructed luma samples are downsampled to match the lower resolution chroma grid. And, similar to CCLM, there is an option to use a single model or a multi-model variant of CCCM. The multi-model variant uses two models, one model is derived for samples above the average luminance reference value, and the other model is for the remaining samples (following the spirit of CCLM design). Multi-model CCCM mode can be selected for PUs with at least 128 available reference samples. 2.1.11.1 Convolutional Filters The convolutional 7-tap filter consists of a 5-tap plus-shaped spatial component, a nonlinear term, and a bias term. The input to the spatial 5-tap component of the filter consists of the center (C) luma sample co-located with the chroma sample to be predicted and its neighbors above / north (N), below / south (S), left / west (W), and right / east (E), as shown below. Figure 10 The spatial portion of the convolution filter is shown. The nonlinear term P is expressed as the square of the center luminance sample C and is scaled to the sample value range of the content: P=(C*C+midVal)>>bitDepth. That is, for 10-bit content, it is calculated as: P=(C*C+512)>>10. The bias term B represents a scalar offset between the input and output (similar to the offset term in CCLM) and is set to the intermediate chrominance value (512 for 10-bit content). The output of the filter is calculated as the filter coefficient c i Convolution with the input value and clipped to the range of valid chroma samples: predChromaVal=c0C+c1N+c2S+c3E+c4W+c5P+c6B. 2.1.11.2 Calculation of filter coefficients Filter coefficient c i is calculated by minimizing the MSE between the predicted chroma samples and the reconstructed chroma samples in the reference region. Figure 11 A reference region consisting of 6 rows of chroma samples above and to the left of the PU is shown. Figure 11 The reference region (and its padding) used to derive the filter coefficients is shown. The reference region extends one PU width to the right of the PU boundary and one PU height below the PU boundary. The region is adjusted to include only available samples. The extension of the region shown in blue is needed to support the "side samples" of the plus-shaped spatial filter and is padded when not in the available region. MSE minimization is performed by calculating the autocorrelation matrix for the luma input and the cross-correlation vector between the luma input and the chroma output. The autocorrelation matrix is LDL decomposed, and the final filter coefficients are calculated using inverse substitution. This process roughly follows the calculation of the ALF filter coefficients in ECM, however, LDL decomposition is chosen instead of Cholesky decomposition to avoid the use of square root operations. 2.1.11.3 Gradient Linear Model Compared to CCLM, GLM uses the gradient of luma samples to derive the linear model instead of the downsampled luma values. Specifically, when GLM is applied, the input of the CCLM process (i.e., the downsampled luma samples L) is replaced by the luma sample gradient G. The other parts of CCLM (e.g., parameter derivation, linear transformation of prediction samples) remain unchanged. C=α·G+β For signaling, when CCLM mode is enabled for the current CU, two flags are separately transmitted through signals for the Cb component and the Cr component to indicate whether GLM is enabled for each component; if GLM is enabled for a component, a syntax element is further transmitted through signals to select one of the four gradient filters for gradient calculation. Enable four gradient filters for GLM, such as Figure 12 As shown, Figure 12 Four Sobel-based gradient modes for the GLM are shown. 2.1.11.4 Gradient Linear Model with Luminance Values In ECM-6.0, GLM uses the gradient of luma samples to predict chroma samples as follows: pred C (i,j)=α·G(i,j)+β, where pred C (i, j) represents the predicted value of the chrominance sample, G(i, j) represents the gradient of the corresponding reconstructed luminance sample, and the linear model parameters α and β are derived from the adjacent reconstructed samples based on the CCLM linear minimum mean square error (LMMSE) method. In the test, a new GLM mode is evaluated, in which the chrominance samples are based on the gradient G(i,j) of the luma samples and the reconstructed value rec′ of the downsampled luma samples. L (i,j) is predicted using different parameters: pred C (i,j)=α0·G(i,j)+α1·rec′ L (i,j)+α2·midValue, The model parameters α0, α1, and α2 are derived from the six rows and columns of adjacent sample points based on the LDL decomposition method of the CCCM model in ECM-6.0. For signaling, a flag is signaled to indicate whether GLM is enabled for both Cb and Cr components, and a syntax element indicating that the gradient mode is encoded by truncated unary code. The original GLM mode is retained and the new GLM mode is signaled as an additional mode by signaling an extra flag in the bitstream. 2.1.11.5 Bitstream Signaling The use of this mode is signaled via a PU-level flag via the CABAC codec. A new CABAC context is included to support this. When it comes to signaling, CCCM is considered a submode of CCLM. That is, the CCCM flag is only signaled if the intra prediction mode is LM_CHROMA. 2.1.12 Template-based Multi-reference Intra Prediction In template-based multi-reference row intra prediction, instead of directly signaling the reference row and intra mode, an index into a candidate list is encoded to indicate which combination of reference row and prediction mode is used to encode the current block, and the combination selected from the combination list is encoded using a truncated Golomb-Rice codec with a divisor of 4. A list of 20 candidates is constructed by combining the MPM with the reference rows {1, 3, 5, 7, 12}. Compared to regular intra-MPM, the MPM list construction is modified as follows: PLANAR mode is not included in the intra prediction mode candidate list, DC mode is added after 5 adjacent modes and DIMD mode. Incremental angles from ±1 to ±4 are added to the angle modes already included in the list. There are 5x10=50, which are Figure 13 The template regions are sorted in ascending order of SAD cost. Since the extended reference rows start from reference row 1, the region covered by reference row 0 is used for template cost calculation. The 20 combinations with the smallest SAD cost form a candidate list. 2.1.13 Intra-frame prediction fusion In this test, the intra prediction is formed by fusion intra prediction derived from different reference lines as follows: For the angular intra prediction mode including the single mode case of TIMD and DIMD, the proposed method is implemented by transforming the angular intra prediction mode from the angular intra prediction mode represented as p fusion =w0p line +w1p line+1 The intra prediction is derived by weighting the intra prediction obtained from multiple reference rows, where p line is intra prediction from the default reference line, and p line+1 is the prediction from the row above the default reference row. The weights are set to w0=3 / 4 and w1=1 / 4. For TIMD mode with mixing, p line Used for 1 st mode (w0=1, w1=0), and p line+1 Used for 2 nd mode (w0=0,w1=1). For DIMD mode with hybrid, the number of prediction values selected for weighted averaging is increased from 3 to 6. When angular intra mode has non-integer slope (needs reference sample interpolation) and the block size is greater than 16, intra prediction fusion is applied to luma blocks, which is used together with MRL, but not applied to ISP-coded blocks. For intra prediction mode, PDPC is applied using the reference line closest to the current block. 2.1.14 IntraTMP Adaptation for Camera-Captured Content In the test, IntraTMP was enabled for the camera captured content and an acceleration method was applied, where the search area was downsampled by a factor of 2 and the template matching search was reduced by a factor of 4. After the best match was found, a second refinement pass was performed, where another template matching search was performed around the best match, with a reduced search range defined as min(width, height) / 2 of the current block. 2.2 Inter-frame prediction For each inter-predicted CU, the motion parameters consist of a motion vector, a reference picture index and a reference picture list usage index, as well as additional information required by the new codec features of VVC in order to be used for inter-predicted sample generation. The motion parameters can be signaled explicitly or implicitly. When a CU is coded in skip mode, the CU is associated with one PU and has no significant residual coefficients, coded motion vector increments or reference picture indices. A Merge mode is specified, whereby the motion parameters for the current CU are obtained from neighboring CUs, including spatial and temporal candidates, as well as the additional scheduling introduced in VVC. Merge mode can be applied to any inter-predicted CU, not just for skip mode. An alternative to Merge mode is the explicit transmission of motion parameters, where the motion vector, the corresponding reference picture index and reference picture list usage flag for each reference picture list, as well as other required information, are explicitly signaled for each CU. In addition to the inter-frame coding features in HEVC, VVC includes many new and refined inter-frame prediction coding tools listed below: – Extended Merge forecast, – Merge mode with MVD (MMVD), – Symmetric MVD (SMVD) signaling, – affine motion compensated prediction, – Sub-block based temporal motion vector prediction (SbTMVP), – Adaptive Motion Vector Resolution (AMVR), – Sports Field Storage: 1 / 16 th Luminance sample MV storage and 8x8 motion field compression, – Bidirectional prediction with CU level weights (BCW), – Bidirectional Optical Flow (BDOF), – Decoder-side motion vector refinement (DMVR), – Geometric Partitioning Mode (GPM), – Joint Intra-frame and Inter-frame Prediction (CIIP). The following text provides details of these inter prediction methods specified in VVC. 2.2.1 Extended Merge Prediction In VVC, the Merge candidate list is constructed by including the following five types of candidates in order: 1) Airspace MVP from airspace adjacent CU, 2) Temporal MVP from the same CU, 3) History-based MVP from FIFO table, 4) Pairwise average MVP, 5) Zero MV. The size of the merge list is signaled in the sequence parameter set header, and the maximum allowed size of the merge list is 6. For each CU coded in merge mode, the index of the best merge candidate is encoded using truncated unary binarization (TU). The first binary bit of the merge index is coded using context, and bypass coding is used for the remaining binary bits. This section provides the derivation process for each category of Merge candidates. As in HEVC, VVC also supports parallel derivation of Merge candidate lists for all CUs in a certain size area. 2.2.1.1 Spatial Candidate Derivation Figure 14 Shows the location of the spatial merge candidate. The derivation of spatial Merge candidates in VVC is the same as that in HEVC, except that the positions of the first two Merge candidates are swapped. Figure 14 Among the candidates at the positions shown, select up to four Merge candidates. The derivation order is B 0, 、A 0, 、B 1, , A1 and B2. Position B2 is considered only when one or more CUs at positions B0, A0, B1 and A1 are not available (for example, because it belongs to another slice or piece) or is intra-coded. After the candidate at position A1 is added, a redundancy check is performed on the addition of the remaining candidates, which ensures that candidates with the same motion information are not included in the list, thereby improving the coding efficiency. In order to reduce computational complexity, not all possible candidate pairs are considered in the mentioned redundancy check. Instead, only the Figure 15 The pairs are linked by arrows in , and a candidate is added to the list only if the corresponding candidates used for redundancy check do not have the same motion information. Figure 15 The candidate pairs considered for redundancy check of spatial merge candidates are shown. Temporal candidate derivation In this step, only one candidate is added to the list. Specifically, in the derivation of the temporal merge candidate, the scaled motion vector is derived based on the co-located CU belonging to the co-located reference picture. The reference picture list to be used for the derivation of the co-located CU is explicitly signaled in the slice header. Figure 16 The motion vector scaling for the time domain Merge candidate is shown. Figure 16 As shown by the dotted line in , the scaled motion vector for the temporal merge candidate is obtained by scaling the motion vector of the co-located CU using the POC distances tb and td, where tb is defined as the POC difference between the reference picture of the current picture and the current picture, and td is defined as the POC difference between the reference picture of the co-located picture and the co-located picture. The reference picture index of the temporal merge candidate is set equal to 0. like Figure 17 As shown, the position for the time domain candidate is selected between candidates C0 and C1, Figure 17 Candidate positions C0 and C1 for the temporal Merge candidate are shown. If the CU at position C0 is unavailable, intra-coded, or outside the current row of the CTU, position C1 is used. Otherwise, position C0 is used for the derivation of the temporal Merge candidate. Merge candidate derivation based on history History-based MVP (HMVP) Merge candidate spatial MVP and TMVP are then added to the Merge list. In this method, the motion information of previously coded blocks is stored in a table and used as the MVP for the current CU. During the encoding / decoding process, a table with multiple HMVP candidates is maintained. When a new CTU row is encountered, the table is reset (cleared). As long as there is a CU with non-sub-block inter-frame coding, the associated motion information is added to the last entry of the table as a new HMVP candidate. The HMVP table size S is set to 6, which indicates that a maximum of 6 history-based MVP (HMVP) candidates can be added to the table. When inserting a new motion candidate into the table, a constrained first-in-first-out (FIFO) rule is utilized, where a redundancy check is first applied to find if there is an identical HMVP in the table. If found, the identical HMVP is removed from the table, and all subsequent HMVP candidates are moved forward. HMVP candidates can be used in the Merge candidate list construction process. The latest HMVP candidates in the table are checked in order and inserted into the candidate list after the TMVP candidate. For spatial or temporal Merge candidates, redundancy check is applied to the HMVP candidates. To reduce the number of redundant checking operations, the following simplifications are introduced: 1. The number of HMPV candidates used for Merge list generation is set to (N<=4)?M:(8-N), where N indicates the number of existing candidates in the Merge list and M indicates the number of available HMVP candidates in the table. 2. Once the total number of available Merge candidates reaches the maximum allowed Merge candidate minus 1, the HMVP is terminated. Merge candidate list building process. Pairwise Average Merge Candidate Derivation Pairwise average candidates are generated by averaging predefined candidate pairs in the existing merge candidate list. The predefined pairs are defined as {(0,1),(0,2),(1,2),(0,3),(1,3),(2,3)}, where the numbers represent the merge index of the merge candidate list. The averaged motion vector is calculated separately for each reference list. If two motion vectors are available in a list, they are averaged even if they point to different reference pictures. If only one motion vector is available, it is used directly. If no motion vector is available, the list remains invalid. When the merge list is not full after adding pairwise average merge candidates, a zero MVP will be inserted at the end until the maximum number of merge candidates is reached. 2.2.1.2Merge Estimation Area Merge Estimation Region (MER) allows independent derivation of Merge candidate lists for CUs in the same Merge Estimation Region (MER). Candidate blocks in the same MER as the current CU are not included in the generation of the Merge candidate list for the current CU. In addition, only when (xCb+cbWidth)>>Log2ParMrgLevel is greater than xCb>> Log2ParMrgLevel and (yCb+cbHeight)>>Log2ParMrgLevel is greater than (yCb>> The update process of the history-based motion vector predictor candidate list is updated only when (log2ParMrgLevel) is set, where (xCb, yCb) is the top left luma sample position of the current CU in the picture, and (cbWidth, cbHeight) is the CU size. The MER size is selected at the encoder side and signaled in the sequence parameter set as log2_parallel_merge_level_minus2. 2.2.1.3 Merge Mode with MVD (MMVD) In addition to the Merge mode (in which the implicitly derived motion information is directly used for prediction sample generation of the current CU), the Merge mode with motion vector difference (MMVD) is introduced in VVC. The MMVD flag is transmitted by signal immediately after the Skip flag and Merge flag are sent to specify whether the MMVD mode is used for the CU. In MMVD, after a merge candidate is selected, it is further refined using signaled MVD information. This further information includes a merge candidate flag, an index specifying the magnitude of motion, and an index indicating the direction of motion. In MMVD mode, one of the first two candidates in the merge list is selected as the MV basis. The merge candidate flag is signaled to specify which candidate to use. The distance index specifies motion magnitude information and indicates a predefined offset from the starting point. Figure 18 The MMVD search points are shown. Figure 18 As shown in Table 2, the offset is added to the horizontal or vertical component of the starting MV. The relationship between the distance index and the predefined offset is specified in Table 2. Table 2 – Relationship between distance index and predefined offsets The direction index indicates the direction of the MVD relative to the starting point. The direction index can represent four directions as shown in Table 3. It should be noted that the meaning of the MVD symbol can change according to the information of the starting MV. When the starting MV is a unidirectionally predicted MV or a bidirectionally predicted MV, and both lists point to the same side of the current picture (that is, the POCs of both references are greater than the POC of the current picture, or both are less than the POC of the current picture), the symbols in Table 3 specify the sign of the MV offset added to the starting MV. When the starting MV is a bidirectionally predicted MV and the two MVs point to different sides of the current picture (that is, the POC of one reference is greater than the POC of the current picture, and the POC of the other reference is less than the POC of the current picture), the symbols in Table 3 specify the sign of the MV offset added to the list 0 MV component of the starting MV, and the signs of list 1 MV have opposite values. Table 3 – Signs of MV offsets specified by direction index Direction IDX 00 01 10 11 x-axis + - N / A N / A y-axis N / A N / A + - 2.2.1.4 Bidirectional Prediction with CU-Level Weights (BCW) In HEVC, the bidirectional prediction signal is generated by averaging two prediction signals obtained from two different reference pictures and / or using two different motion vectors. In VVC, the bidirectional prediction mode is extended beyond simple averaging to allow weighted averaging of the two prediction signals. P bi-pred=((8-w)*P0+w*P1+4)>>3 (2-1) Five weights are allowed in weighted average bidirectional prediction, w∈{-2,3,4,5,10}. For each bidirectionally predicted CU, the weight w is determined in one of two ways: 1) For non-Merge CUs, the weight index is transmitted by signal after the motion vector difference; 2) For Merge CUs, the weight index is inferred from neighboring blocks based on the Merge candidate index. BCW is only applied to CUs with 256 or more luma samples (i.e., CU width multiplied by CU height is greater than or equal to 256). For low-latency pictures, all 5 weights are used. For non-low-latency pictures, only 3 weights (w∈{3,4,5}) are used. At the encoder, a fast search algorithm is applied to find the weight index without significantly increasing the encoder complexity. These algorithms are summarized below. When combined with AMVR, unequal weights for 1-pixel and 4-pixel motion vector precision are only conditionally checked if the current picture is a low-latency picture. When combined with affine, affine ME is performed for unequal weights if and only if the affine mode is selected as the current best mode. – When the two reference pictures in bidirectional prediction are the same, unequal weights are only checked conditionally. – Do not search for unequal weights when certain conditions are met, which depend on the POC distance between the current picture and its reference pictures, the codec QP, and the temporal level. The BCW weight index is encoded using a context-coded bit and a subsequent bypass-coded bit. The first context-coded bit indicates whether equal weights are used; if unequal weights are used, an additional bit is signaled using the bypass codec to indicate which unequal weights are used. Weighted prediction (WP) is a codec tool supported by the H.264 / AVC and HEVC standards for efficiently encoding and decoding video content with gradients. Support for WP has also been added to the VVC standard. WP allows weighting parameters (weights and offsets) to be signaled for each reference picture in each of the reference picture lists L0 and L1. The weights and offsets for the corresponding reference picture(s) are then applied during motion compensation. WP and BCW are designed for different types of video content. To avoid interaction between WP and BCW, which would complicate VVC decoder design, if a CU uses WP, the BCW weight index is not signaled, and w is assumed to be 4 (i.e., equal weights are applied). For a merge CU, the weight index is inferred from neighboring blocks based on the merge candidate index. This can be applied to both normal merge mode and inherited affine merge mode. For constructed affine merge mode, affine motion information is constructed based on the motion information of up to three blocks. The BCW index of a CU using constructed affine merge mode is simply set equal to the BCW index of the first control point MV. In VVC, CIIP and BCW cannot be applied jointly for a CU. When a CU is encoded or decoded using CIIP mode, the BCW index of the current CU is set to 2, i.e., equal weight. 2.2.1.5 Bidirectional Optical Flow (BDOF) The Bidirectional Optical Flow (BDOF) tool is included in VVC. BDOF (formerly known as BIO) is included in JEM. Compared to the JEM version, the BDOF in VVC is a simpler version that requires much less computation, especially in terms of the number of multiplications and the size of the multipliers. BDOF is used to refine the bidirectional prediction signal of a CU at the 4x4 sub-block level. BDOF is applied to a CU if it meets all of the following conditions: - The CU is coded using "true" bi-prediction mode, ie, one of the two reference pictures precedes the current picture in display order, and the other follows the current picture in display order. - The distances from the two reference pictures to the current picture (ie, the POC difference) are the same. - Both reference images are short-term reference images. -CU is not encoded or decoded using Affine mode or ATMVP Merge mode. - The CU has more than 64 luma samples. - CU height and CU width are both greater than or equal to 8 luma samples. -BCW weight index indicates equal weight. -Do not enable WP for the current CU. -CIIP mode is not used for the current CU. BDOF is applied only to the luma component. As its name indicates, BDOF mode is based on the concept of optical flow, which assumes that the motion of objects is smooth. For each 4x4 sub-block, the motion refinement (v) is calculated by minimizing the difference between the L0 prediction samples and the L1 prediction samples. x ,v y ). Motion refinement is then used to adjust the bidirectionally predicted sample values in the 4x4 sub-blocks. The following steps are applied in the BDOF process. First, the horizontal and vertical gradients of the two predicted signals are calculated by directly calculating the difference between two adjacent samples. and k=0,1, that is, Among them I (k) (i, j) is the sample value at coordinate (i, j) of the prediction signal in list k, k=0, 1, and shift1 is calculated based on the luma bit depth bitDepth as shift1=max(6, bitDepth-6). Then, the autocorrelations and cross-correlations of the gradients S1, S2, S3, S5, and S6 are calculated as S1=∑ (i,j)∈Ω Abs(ψ x (i,j)),S3=∑ (i,j)∈Ω θ(i,j)·Sign(ψ x (i,j)) (2-3) S5=∑ (i,j)∈Ω Abs(ψ y (i,j)),S6=∑ (i,j)∈Ω θ(i,j)·Sign(ψ y (i,j)) in θ(i,j)=(I (1) (i,j)>>n b )-(I (0) (i,j)>>nb ) where Ω is a 6x6 window around the 4x4 sub-block, and n a and n b The values of are set equal to min(1, bitDepth-11) and min(4, bitDepth-8) respectively. The motion refinement (v) is then derived using the cross-correlation and autocorrelation terms as follows x ,v y ): in th′ BIO =2 max(5,BD-7) . is a floor function, and Based on the motion refinement and gradients, the following adjustments are calculated for each sample in the 4x4 sub-block: Finally, the BDOF samples of the CU are calculated by adjusting the bidirectional prediction samples as follows: pred BDOF (x,y)=(I (0) (x,y)+I (1) (x,y)+b(x,y)+o offset )>>Displacement(2-7) These values are chosen so that the multipliers in the BDOF process do not exceed 15 bits and the maximum bit width of the intermediate parameters in the BDOF process is kept within 32 bits. In order to derive the gradient value, it is necessary to generate some predicted sample points I in the list k (k = 0, 1) outside the current CU boundary (k) (i,j). Figure 19 FIG4 shows the expanded CU area used in BDOF. Figure 19 As shown, BDOF in VVC uses an extended row / column around the CU boundary. In order to control the computational complexity of generating prediction samples outside the boundary, the prediction samples in the extended area (white positions) are generated by directly obtaining the reference samples at nearby integer positions without interpolation (using the floor() operation on the coordinates), and the normal 8-tap motion compensation interpolation filter is used to generate the prediction samples within the CU (gray positions). These extended sample values are only used in gradient calculations. For the remaining steps in the BDOF process, if any samples and gradient values outside the CU boundary are needed, they are filled (i.e. repeated) from their nearest neighbors. When the width and / or height of a CU is greater than 16 luma samples, it will be divided into sub-blocks with a width and / or height equal to 16 luma samples, and the sub-block boundaries are regarded as CU boundaries in the BDOF process. The maximum unit size for the BDOF process is limited to 16x16. The BDOF process can be skipped for each sub-block. When the SAD between the initial L0 prediction samples and the L1 prediction samples is less than a threshold, the BDOF process is not applied to the sub-block. The threshold is set to be equal to (8*W*(H>>1), where W indicates the sub-block width and H indicates the sub-block height. In order to avoid the additional complexity of the SAD calculation, the SAD between the initial L0 prediction samples and the L1 prediction samples calculated in the DVMR process is reused here. If BCW is enabled for the current block, that is, the BCW weight index indicates unequal weights, then bidirectional optical flow is disabled. Similarly, if WP is enabled for the current block, that is, luma_weight_lx_flag is 1 for either of the two reference pictures, then BDOF is also disabled. BDOF is also disabled when the CU is encoded or decoded using symmetric MVD mode or CIIP mode. 2.2.1.6 Symmetrical MVD Encoding and Decoding In VVC, in addition to the normal unidirectional prediction and bidirectional prediction mode MVD signaling, a symmetric MVD mode for bidirectional prediction MVD signaling is also applied. In symmetric MVD mode, the motion information including the reference picture indexes of both list 0 and list 1 and the MVD of list 1 is not transmitted through the signal but is derived. The decoding process of the symmetric MVD mode is as follows: 1) At the stripe level, the variables BiDirPredFlag, RefIdxSymL0, and RefIdxSymL1 are derived as follows: – If mvd_l1_zero_flag is 1, BiDirPredFlag is set equal to 0. Otherwise, if the most recent reference picture in list 0 and the most recent reference picture in list 1 form a forward and backward reference picture pair or a backward and forward reference picture pair, then BiDirPredFlag is set to 1, and both the list 0 reference picture and the list 1 reference picture are short-term reference pictures. Otherwise, BiDirPredFlag is set to 0. 2) At the CU level, if the CU is bidirectionally predicted and BiDirPredFlag is equal to 1, the symmetric mode flag indicating whether the symmetric mode is used is explicitly signaled. When the symmetric mode flag is true, only mvp_l0_flag, mvp_l1_flag, and MVD0 are explicitly signaled. The reference indexes of list 0 and list 1 are set equal to the reference picture pair, respectively. MVD1 is set equal to (-MVD0). The final motion vector is shown in the following formula. Figure 20 A schematic diagram of a symmetric MVD pattern is shown. In the encoder, symmetric MVD motion estimation starts with initial MV evaluation. A set of initial MV candidates includes MVs obtained from unidirectional prediction search, MVs obtained from bidirectional prediction search, and MVs from the AMVP list. The MV with the lowest rate-distortion cost is selected as the initial MV for symmetric MVD motion search. 2.2.1.7 Decoder-side Motion Vector Refinement (DMVR) To improve the accuracy of MVs in Merge mode, decoder-side motion vector refinement based on bilateral matching is applied in VVC. In bidirectional prediction operations, the refined MV is searched around the initial MV in reference picture lists L0 and L1. The BM method calculates the distortion between two candidate blocks in reference picture lists L0 and L1. Figure 21 Decoding side motion vector refinement is shown. Figure 21 As shown, the SAD between the red blocks of each MV candidate around the initial MV is calculated. The MV candidate with the lowest SAD becomes the refined MV and is used to generate the bidirectionally predicted signal. In VVC, DMVR can be applied to CUs that are coded or decoded using the following modes and features: -CU-level Merge mode with bi-predictive MV. - Relative to the current picture, one reference picture is in the past and the other reference picture is in the future. - The distances from the two reference pictures to the current picture (ie, the POC difference) are the same. - Both reference images are short-term reference images. - The CU has more than 64 luma samples. - CU height and CU width are both greater than or equal to 8 luma samples. – BCW weight index indicates equal weight. – WP is not enabled for the current block. – CIIP mode is not used for the current block. The refined MV derived by the DMVR process is used to generate inter-frame prediction samples and is also used for temporal motion vector prediction in future picture codecs. The original MV is used in the deblocking process and is also used for spatial motion vector prediction in future CU codecs. Additional features of DMVR are mentioned in the following sub-items. Search Solutions In DVMR, the search point is around the initial MV, and the MV offset follows the MV difference mirror rule. In other words, any point examined by DMVR (represented by a candidate MV pair (MV0, MV1)) follows the following two equations: MV0′=MV0+MV_offset (2-9), MV1′=MV1-MV_offset (2-10). Where MV_offset represents the refinement offset between the initial MV and the refined MV in one of the reference pictures. The refinement search range is two integer luma samples away from the initial MV. The search includes an integer sample offset search phase and a fractional sample refinement phase. A 25-point full search is applied for integer sample offset search. The SAD of the initial MV pair is first calculated. If the SAD of the initial MV pair is less than a threshold, the integer sample stage of DMVR is terminated. Otherwise, the SAD of the remaining 24 points is calculated and checked in raster scan order. The point with the smallest SAD is selected as the output of the integer sample offset search stage. To reduce the impact of DMVR refinement uncertainty, it is proposed to bias towards the original MV during the DMVR process. The SAD between the reference blocks referenced by the initial MV candidates is reduced by 1 / 4 of the SAD value. The integer sample search is followed by fractional sample refinement. To save computational complexity, fractional sample refinement is derived using the parametric error surface equation rather than an additional search with SAD comparison. Fractional sample refinement is conditionally invoked based on the output of the integer sample search phase. Fractional sample refinement is further applied when the integer sample search phase terminates with the center having the minimum SAD in either the first or second iteration of the search. In the sub-pixel offset estimation based on the parameter error surface, the center position cost and the costs at the four neighboring positions from the center are used to fit the two-dimensional parabolic error surface equation of the following form E(x,y)=A(xx min ) 2 +B(yy min ) 2 +C (2-11) Where (x min ,y min) corresponds to the fractional position with the minimum cost, and C corresponds to the value of the minimum cost. By solving the above equation using the cost values of the five search points, (x min ,y min ) is calculated as: x min =(E(-1,0)-E(1,0)) / (2(E(-1,0)+E(1,0)-2E(0,0))) (2-12), y min =(E(0,-1)-E(0,1)) / (2((E(0,-1)+E(0,1)-2E(0,0))) (2-13). x min and y min The value of is automatically constrained to be between -8 and 8, since all cost values are positive and the minimum is E(0,0). This corresponds to a half-pixel offset with 1 / 16 pixel MV precision in VVC. The calculated fraction (x min ,y min ) is added to the integer distance refinement MV to get the sub-pixel accurate refinement increment MV. Bilinear interpolation and sample filling In VVC, the resolution of the MV is 1 / 16 luma samples. Samples at fractional positions are interpolated using an 8-tap interpolation filter. In DMVR, the search points are centered around the original fractional pixel MV with integer sample offsets, so samples at those fractional positions need to be interpolated for the DMVR search process. To reduce computational complexity, a bilinear interpolation filter is used to generate fractional samples for the search process in DMVR. Another important effect of using a bilinear filter is that, with a 2-sample search range, DVMR does not access more reference samples than the normal motion compensation process. After obtaining the refined MV through the DMVR search process, a normal 8-tap interpolation filter is applied to generate the final prediction. To avoid accessing more reference samples than the normal MC process, samples that are not required for the interpolation process based on the original MV but are required for the interpolation process based on the refined MV are filled from those available samples. Maximum DMVR Processing Unit When the width and / or height of a CU is greater than 16 luma samples, it will be further divided into sub-blocks with width and / or height equal to 16 luma samples. The maximum unit size for the DMVR search process is limited to 16x16. 2.2.1.8 Joint Intra-Frame and Inter-Frame Prediction (CIIP) In VVC, when a CU is encoded and decoded in Merge mode, if the CU contains at least 64 luma samples (i.e., the CU width multiplied by the CU height is equal to or greater than 64), and if both the CU width and the CU height are less than 128 luma samples, an additional flag is transmitted by signal to indicate whether the intra / inter joint prediction (CIIP) mode is applied to the current CU. As its name indicates, CIIP prediction combines the inter prediction signal with the intra prediction signal. The inter prediction signal P in CIIP mode inter is derived using the same inter-frame prediction process as applied to the conventional Merge mode; and the intra-frame prediction signal P intra The conventional intra prediction process with planar mode is derived. Then, the intra prediction signal and the inter prediction signal are combined using weighted averaging, where the weight values are calculated depending on the codec mode of the top neighboring block and the left neighboring block as follows: – If the top neighboring block is available and is intra-coded, set isIntraTop to 1, otherwise set isIntraTop to 0; – If the left neighboring block is available and is intra-coded, set isIntraLeft to 1, otherwise set isIntraLeft to 0; – If (isIntraLeft + isIntraTop) is equal to 2, then wt is set to 3; – Otherwise, if (isIntraLeft + isIntraTop) is equal to 1, wt is set to 2; – Otherwise, set wt to 1. The CIIP forecast is formed as follows: P CIIP =((4-wt)*P inter +wt*P intra +2)>>2 (2-14). Figure 22 The top and left neighboring blocks used in the derivation of the CIIP weights are shown. 2.2.1.9 Multiple Hypothesis Prediction (MHP) In inter-AMVP mode, normal Merge mode and MMVD mode, up to two additional prediction values are signaled. The resulting overall prediction signal is iteratively accumulated with each additional prediction signal. p n+1 =(1-α n+1 )p n +α n+1 h n+1 The weighting factor α is specified according to the following table. add_hyp_weight_idx α 0 1 / 4 1 -1 / 8 For inter-AMVP mode, MHP is applied only when unequal weights in BCW are selected in bi-prediction mode. 2.2.1.10 Overlapped Sub-Block Motion Compensation (OBMC) When OBMC is applied, top and left boundary pixels of a CU are refined using motion information of neighboring blocks with weighted prediction. The conditions under which OBMC should not be applied are as follows: When OBMC is disabled at the SPS level. When the current block has intra mode or IBC mode. When the current block applies LIC. When the current luminance block area is less than or equal to 32. Sub-block boundary OBMC is performed by applying the same blending to the top, left, bottom and right sub-block boundary pixels using the motion information of the neighboring sub-blocks. It is enabled for sub-block based codecs: Affine AMVP mode; Affine Merge mode and sub-block-based temporal motion vector prediction (SbTMVP); Sub-block based bilateral matching. 2.2.1.11 Local Illumination Compensation (LIC) LIC is an inter-frame prediction technique that models the local illumination variation between the current block and its prediction block as a function of the local illumination variation between the current block template and the reference block template. The parameters of this function can be represented by a scale α and an offset β, which form a linear equation, namely α*p[x]+β, to compensate for illumination variation, where p[x] is the reference sample at position x on the reference picture pointed to by the MV. When surround motion compensation is enabled, the MV must be clipped to account for the surround offset. Since α and β can be derived based on the current block template and the reference block template, no signaling overhead is required for them, except for the signaling of the LIC flag for AMVP mode to indicate the use of LIC. Local illumination compensation is used for uni-directionally predicted inter CUs with the following modifications. Neighboring samples within the frame can be used to derive LIC parameters; Disable LIC for blocks with less than 32 luma samples; For both non-subblock mode and affine mode, LIC parameter derivation is performed based on the template block samples corresponding to the current CU, rather than based on the partial template block samples corresponding to the first top-left 16x16 unit; The samples of the reference block template are generated by using the MC with the block MV without rounding it to integer pixel precision. 2.2.1.12 Geometric Partitioning Mode (GPM) In VVC, geometric partitioning mode for inter prediction is supported. The geometric partitioning mode is signaled using a CU level flag as a Merge mode, where other Merge modes include normal Merge mode, MMVD mode, CIIP mode, and sub-block Merge mode. The geometric partitioning mode is used for each possible CU size w×h= 2 m ×2 n A total of 64 splits are supported, where m,n∈{3…6} excluding x64 and 64x8. When this mode is used, the CU is divided into two parts by a geometrically positioned straight line ( Figure 23 ). Figure 23 An example of a GPM partition grouped at the same angle is shown. The position of the partition line is mathematically derived from the angle and offset parameters of the specific partition. Each part of the geometric partition in the CU is inter-predicted using its own motion; only unidirectional prediction is allowed for each partition, i.e., each part has one motion vector and one reference index. Unidirectional prediction motion constraints are applied to ensure that, as with conventional bidirectional prediction, only two motion-compensated predictions are required for each CU. If geometric partitioning mode is used for the current CU, a geometric partitioning index and two Merge indices (one for each partition) indicating the partitioning mode (angle and offset) of the geometric partitioning are further transmitted by signal. The number of maximum GPM candidate sizes is explicitly transmitted by signal in the SPS, and the syntax binarization for the GPM Merge index is specified. After predicting each part of the geometric partitioning, a hybrid process with adaptive weights is used to adjust the sample values along the geometric partitioning edge. This is a prediction signal for the entire CU, and the transform and quantization process will be applied to the entire CU as in other prediction modes. Finally, the motion field of the CU predicted using the geometric partitioning mode is stored. One-way prediction candidate list construction The unidirectional prediction candidate list is directly derived from the merge candidate list constructed according to the extended merge prediction process. Let n be the index of the unidirectional prediction motion in the geometric unidirectional prediction candidate list. The LX motion vector (where X is equal to the parity of n) of the nth extended merge candidate is used as the nth unidirectional prediction motion vector for the geometric partition mode. These motion vectors are Figure 24 Marked with an "x". Figure 24Uni-prediction MV selection for geometric partitioning mode is shown.If a corresponding LX motion vector of the nth extended Merge candidate does not exist, the L(1-X) motion vector of the same candidate is used as the uni-prediction motion vector for geometric partitioning mode instead. Blending along geometric segmentation edges After predicting each part of the geometric partition using its own motion, blending is applied to the two prediction signals to derive samples around the geometric partition edges. The blending weight for each position of the CU is derived based on the distance between the individual position and the partition edge. The distance from the segmentation edge to the position (x,y) is derived as: where i,j are the indices of the angle and offset for the geometric partition, which depend on the geometric partition index transmitted by the signal. ρ x,j and ρ y,j The sign of depends on the angle index i. The weights of each part of the geometric segmentation are derived as follows: wldxL(x,y)=partIdx? 32+d(x,y):32-d(x,y) (2-19) w1(x,y)=1-w0(x,y)(2-21). partIdx depends on the angle index i. An example of the weight w0 is shown below. Figure 25 An exemplary generation of blending weights w0 using a geometric partitioning mode is shown. Motion field storage for geometric partitioning patterns Mv1 from the first part of the geometric partition, Mv2 from the second part of the geometric partition, and the combined Mv of Mv1 and Mv2 are stored in the motion field of the geometric partition mode coded CU. The type of motion vector stored for each individual position in the motion field is determined as: sType=abs(motionIdx)<32?2:(motionIdx≤0?(1-partIdx):partIdx) (2-22) where motionIdx is equal to d(4x+2,4y+2). partIdx depends on the angle index i. If sType is equal to 0 or 1, then Mv0 or Mv1 is stored in the corresponding motion field, otherwise if sType is equal to 2, the combined Mv from Mv0 and Mv2 is stored. The combined Mv is generated using the following process: 1) If Mv1 and Mv2 are from different reference picture lists (one from L0 and the other from L1), Mv1 and Mv2 are simply combined to form a bi-directional prediction motion vector. 2) Otherwise, if Mv1 and Mv2 are from the same list, only the unidirectional predicted motion Mv2 is stored. GPM with inter and intra prediction (GPM Inter-Intra) With GPM Inter-Intra, in addition to the Merge candidates for each non-rectangular partitioned area in the CU to which GPM is applied, a predefined intra prediction mode for the geometric partition line can also be selected. In the proposed method, for each GPM-separated area, a flag from the encoder is used to determine whether it is intra prediction mode or inter prediction mode. When it is inter prediction mode, the unidirectional prediction signal is generated by the MV from the Merge candidate list. On the other hand, when it is intra prediction mode, the unidirectional prediction signal is generated from the neighboring pixels for the intra prediction mode specified by the index from the encoder. The variation of possible intra prediction modes is limited by the geometric shape. Finally, the two unidirectional prediction signals are mixed in the same way as ordinary GPM. 2.3 About LIC, AMVP-MERGE, and Adaptive DMVR Mode The following detailed embodiments should be considered as examples to explain the general concept. These embodiments should not be interpreted in a narrow sense. In addition, these embodiments can be combined in any way. The term “video unit” or “codec unit” or “block” may refer to a codec tree block (CTB), a codec tree unit (CTU), a codec block (CB), a CU, a PU, a TU, a PB, a TB. In this disclosure, regarding “blocks encoded and decoded in mode N”, “mode N” here can be a prediction mode (e.g., MODE_INTRA, MODE_INTER, MODE_PLT, MODE_IBC, etc.) or a coding and decoding technology (e.g., AMVP, Merge, SMVD, BDOF, PROF, DMVR, AMVR, TM, Affine, CIIP, GPM, GEO, TPM, MMVD, BCW, HMVP, SbTMVP, etc.). In this disclosure, "bidirectional DMVR" may refer to conventional DMVR that refines both L0 and L1 motion vectors, as described in Section 2.1.14. In addition, "unidirectional DMVR" may refer to a DMVR process that refines only L0 or L1 motion vectors, such as the adaptive DMVR described in Section 2.1.23. In the following discussion, LIC parameters may refer to two parameters (such as a slope parameter "a" and a bias parameter "b") derived based on a linear model used to map neighboring samples of a current block to neighboring samples of a temporally co-located block (e.g., the temporally co-located block may be pointed to by a motion vector or a rounded motion vector of the current block). Furthermore, the LIC parameters may be used to estimate prediction values for samples within the current video unit. In the following discussion, the AMVP mode may be a conventional AMVP mode, an affine AMVP mode, and / or an SMVD mode, and / or an AMVP-MERGE mode. It should be noted that the following terms are not limited to the specific terms defined in existing standards, and any variants of codecs are also applicable. 2.3.1 In order to solve the first problem, the following method is proposed: a. Propose support for multiple MVD / MV accuracies for the AMVP-MERGE mode. i. In one example, the supported precision candidates can be the same as the precision candidates used for the conventional AMVP mode, for example, for the non-affine case, half pixel, 1 / 4 pixel, 1 pixel, 4 pixels are applied; for the affine case, 1 / 16 pixel, 1 / 8 pixel, 1 / 4 pixel are applied. ii. In another example, at least one of the supported precision candidates may be different from the precision candidates used for the regular AMVP mode. b. For example, the motion vector difference (eg, MVD) on the AMVP side of the AMVP-MERGE mode may be encoded with other precisions besides 1 / 4 pixel resolution. i. For example, the MVD value can be 4-pixel accurate. ii. For example, the MVD value can be 1 pixel accurate. iii. For example, the MVD value can be half-pixel accurate. iv. For example, the MVD value can be 1 / 8 pixel accurate. v. For example, the MVD value can be 1 / 16 pixel accurate. c. For example, a second interpolation filter (eg, a Y-tap filter, such as Y=6) in addition to the first interpolation filter (eg, an X-tap filter, such as X=8 or 12) may also be used for motion compensation. i. For example, when to use the second interpolation filter for an AMVP-MERGE coded block may depend on the MVD prediction and / or the final MV accuracy. ii. For example, when the MVD transmitted by the signal is 1 / 2 pixel precision and the final MV is also 1 / 2 pixel precision, the second interpolation filter can be used. Otherwise, the first interpolation filter can be used. iii. In one example, the second interpolation filter used in the AMVP-MERGE mode may be the same as the second interpolation filter used for the regular AMVP mode (eg, a second interpolation filter for 1 / 2 pixel precision). iv. Alternatively, the second interpolation filter used in the AMVP-MERGE mode may be different from the second interpolation filter used for the regular AMVP mode (eg, a second interpolation filter for 1 / 2 pixel precision). d. Alternatively, half-pixel MVD precision may not be allowed for AMVP-MERGE mode. e. For example, at least one syntax element (e.g., a flag and / or parameter index) may be signaled at the block level to indicate which motion vector precision is used to encode / signal the MVD value and / or MV of a video unit encoded in AMVP-MERGE mode. i. In addition, the signaled MVD at any resolution can be converted to internal precision (eg, 1 / 16 pixel resolution) for subsequent processes such as motion compensation. 2.3.2 To solve the second problem, the following method is proposed: a. In one example, the prediction unit generated based on the AMVP-MERGE mode can be used as the MHP hypothesis. i. For example, the AMVP-MERGE mode prediction block can be used as the basic hypothesis of the MHP block. ii. For example, syntax elements / structures related to MHP hypothesis data (e.g., whether there are additional hypotheses associated with the video unit coded in AMVP-MERGE mode) may be signaled in the bitstream immediately after the video unit is identified as a video unit coded in AMVP-MERGE mode; If so, additional assumptions based on AMVP or MERGE's MHP are used, etc.). iii. For example, the AMVP-MERGE prediction block can be used as an additional hypothesis for the MHP block. iv. For example, additional hypotheses for MHP blocks can be generated based on AMVP-MERGE motion candidates. 1) For example, in this case, syntax elements related to the AMVP-MERGE motion candidate (e.g., which side is AMVP / MERGE encoded, the reference index of the AMVP side, the MVD value for the AMVP side and / or the MVP index of the AMVP side) can be transmitted by signal in the multi-hypothesis data structure. v. For example, the additional assumption of whether LIC is used for AMVP-MERGE encoding or not can be inherited from the usage of LIC of the base assumption. 1) For example, if the base hypothesis is LIC coded, then the additional hypothesis coded by AMVP-MERGE is LIC coded without signaling the use of LIC for such additional hypothesis. 2) Alternatively, if the base hypothesis is not LIC coded, then the additional hypothesis coded by AMVP-MERGE is not LIC coded, without signaling the use of LIC for such additional hypothesis. vi. For example, whether LIC is used for the hypothesis (basic hypothesis and / or additional hypothesis) encoded and decoded via AMVP-MERGE may depend on the use of LIC on the Merge side of the AMVP-MERGE candidate. 1) For example, suppose an AMVP-MERGE candidate consists of a unidirectional merge candidate on one side (L0 or L1) and a unidirectional AMVP candidate on the other side (L1 or L0). a. For example, the use of LIC for such an AMVP-MERGE candidate can be obtained from its Merge Candidates are inherited. b. For example, if the Merge candidate uses LIC, then LIC can be used for AMVP- Prediction block for MERGE codec. vii. Alternatively, the use of the hypothetical LIC for AMVP-MERGE codecs may be signaled in the bitstream. 2.3.3 To solve the third problem, the following method is proposed: a. For example, an AMVP-MERGE candidate may be used in one or more of the following codec modes. i. CIIP mode (and / or its variants, e.g., conventional CIIP, CIIP-PDPC, CIIP-TM, etc.). ii. MMVD mode (and / or its variants, eg, conventional MMVD, affine MMVD, etc.). iii. MHP model (and / or its variants, such as MHP basic assumptions and / or MHP additional assumptions, etc.). iv. GPM mode (and / or its variants, e.g., conventional GPM, GPM-TM, GPM-MMVD, GPM Inter-Intra, etc.). b. For example, AMVP-MERGE candidates may first be refined by a decoder-side motion vector refinement process (eg, TM- or DMVR-based motion vector refinement) and then used for the second encoding mode (eg, as listed in the above sub-item). c. For example, an AMVP-MERGE candidate may be inserted into another candidate list. i. For example, AMVP-MERGE candidates can be inserted into the regular Merge candidate list. 1) For example, an AMVP-MERGE candidate can be used in the regular Merge mode and / or its variants. 2) For example, AMVP-MERGE candidates may be used in MMVD mode and / or its variants. 3) For example, the AMVP-MERGE candidate can be used in CIIP mode and / or its variants. 4) For example, the AMVP-MERGE candidate may be used in the MHP mode and / or its variants. 5) For example, the AMVP-MERGE candidate may be used in GPM mode and / or its variants. ii. For example, AMVP-MERGE candidates can be inserted into the regular TM Merge candidate list. 1) For example, AMVP-MERGE candidates may be used in conventional TM Merge mode and / or its variants. iii. Additionally, AMVP-MERGE candidates may be inserted into another prediction list after the original candidate of that prediction list. d. For example, AMVP-MERGE candidates can be reordered based on a decoder-derived method (through TM or DMVR-based cost evaluation), and then M of the AMVP-MERGE candidates will be selected to be added to the second candidate list (such as a regular Merge candidate list, a regular TM Merge candidate list). i. Additionally, more than one AMVP-MERGE candidate may be reordered together. ii. Alternatively, the first candidate from the first AMVP-MERGE prediction list and the second candidate from the second prediction list may be reordered together. e. For example, once an AMVP-MERGE candidate is used for a codec block, additional syntax elements may be signaled that specify the prediction direction (L0 or L1) of the AMVP portion and / or the reference picture index of the selected AMVP candidate and / or the motion vector predictor index of the selected AMVP candidate and / or the motion vector difference associated with the AMVP motion vector predictor. i. Alternatively, the motion vector predictor index on the AMVP side of the AMVP-MERGE candidate may not be transmitted through a signal (eg, the motion vector predictor index may be selected by a decoder-side method through a TM or DMVR-based cost evaluation). f. Alternatively, once an AMVP-MERGE candidate is used, additional syntax element(s) may be signaled that specify the predictor index of the Merge candidate. i. Alternatively, the motion vector predictor index on the Merge side of the AMVP-MERGE candidate may not be transmitted through a signal (eg, the motion vector predictor index may be selected by a decoder-side method through a TM or DMVR-based cost evaluation). g. In one example, the Merge part of the AMVP-MERGE mode may be first refined by a decoder-side motion vector refinement process (such as TM or DMVR) before generating AMVP-MERGE candidates. h. In one example, the AMVP portion of the AMVP-MERGE mode may be first refined by a decoder-side motion vector refinement process (such as TM or DMVR) before generating AMVP-MERGE candidates. 2.3.4 To solve the fourth problem, the following method is proposed: a. For example, the adaptive DMVR motion candidate can be used in one or more of the following codec modes. i. CIIP mode (and / or its variants, e.g., conventional CIIP, CIIP-PDPC, CIIP-TM, etc.). ii. MMVD mode (and / or its variants, eg, conventional MMVD, affine MMVD, etc.). iii. MHP model (and / or its variants, such as MHP basic assumptions and / or MHP additional assumptions, etc.). iv. GPM mode (and / or its variants, e.g., conventional GPM, GPM-TM, GPM-MMVD, GPM inter-frame-intraframe, etc.). v. AMVP mode (and / or its variants, e.g., regular AMVP, SMVD, AMVP-MERGE, affine AMVP, etc.). b. For example, adaptive DMVR motion candidates may first be refined by a decoder-side motion vector refinement process (eg, TM- or DMVR-based motion vector refinement) and then used in the second encoding mode (eg, as listed in the above sub-item). c. For example, when adaptive DMVR motion candidates are used in AMVP mode. i. A DMVR motion candidate may refer to a motion vector pair containing both an L0 motion vector and an L1 motion vector, and / or both an L0 reference picture index and an L1 reference picture index. ii. It can be used as an MVP candidate. iii. The candidate index of the DMVR motion candidate (instead of the L0 and L1 motion vector predictor indices and the L0 and L1 reference picture indices) may be signaled in the bitstream for the AMVP mode. iv. The reference picture index for one prediction direction (L0 or L1) may be signaled for AMVP mode, and the reference picture index for the other prediction direction (L1 or L0) may be inferred (eg, according to DMVR conditions). d. For example, the adaptive DMVR motion candidate can be inserted into another candidate list. i. For example, adaptive DMVR motion candidates can be inserted into the regular Merge candidate list. 1) For example, adaptive DMVR motion candidates can be used in the conventional Merge mode and / or its variants. 2) For example, adaptive DMVR motion candidates can be used in MMVD mode and / or its variants. 3) For example, adaptive DMVR motion candidates can be used in CIIP mode and / or its variants. 4) For example, adaptive DMVR motion candidates can be used in MHP mode and / or its variants. 5) For example, adaptive DMVR motion candidates can be used in GPM mode and / or its variants. ii. For example, the adaptive DMVR motion candidate can be inserted into the regular TM Merge candidate list. 1) For example, adaptive DMVR motion candidates can be used in conventional TM Merge mode and / or its variants. iii. Additionally, furthermore, the adaptive DMVR motion candidate may be inserted into another prediction list after the original candidate of that prediction list. e. For example, the adaptive DMVR motion candidates can be reordered based on a decoder-derived method (through TM- or DMVR-based cost evaluation), and then M of the adaptive DMVR motion candidates will be selected to be added to the second candidate list (such as the regular Merge candidate list, the regular TM Merge candidate list). i. Additionally, more than one adaptive DMVR motion candidate can be reordered together. ii. Alternatively, the first candidate from the first adaptive DMVR Merge list and the second candidate from the second prediction list may be reordered together. 2.3.5 To solve the fifth problem, the following method is proposed: a. For example, the enabling / disabling of a first codec tool may be controlled by a second syntax element signaled at a syntax level higher than the codec block level. i. For example, the syntax level higher than the codec block level can indicate the sequence level / picture group level / picture level / slice level / slice group level, such as in the sequence header / picture header / SPS / VPS / DPS / DCI / In PPS / APS / strip header / slice group header. ii. For example, a single codec may be controlled by the second syntax element. iii. For example, more than one codec may be controlled by the second syntax element. iv. For example, the first codec tool may be a prediction mode using a decoder-side motion inference method (Such as AMVP-MERGE mode, etc.). v. For example, the second syntax element may be a parameter that specifies whether a decoder-side motion derivation method (such as DMVR and TM, etc.) vi. For example, the second syntax element may be a parameter that specifies whether decoder-side motion vector refinement is allowed (such as DMVR, etc.) with SPS / PPS / PH / SH marks. vii. For example, the second syntax element may be a parameter that specifies whether decoder-side template matching (such as TM and / or inter-frame TM, etc.) is allowed. 2.3.6 Regarding LIC parameter derivation (e.g., as shown in the sixth question), assuming that the linear model used for the LIC-encoded block is based on at least two parameters: a slope parameter "a" and a bias parameter "b", and the relationship between the neighboring samples of the current block and the neighboring samples of the time-domain co-located block can be expressed by "reconTempNeigh=a*reconCurNeigh+b", where "reconTempNeigh" represents the reconstructed / predicted value of the neighboring samples of the time-domain co-located block, and "reconCurNeigh" represents the reconstructed / predicted value of the neighboring samples of the current block, the following method is proposed: a. For example, at least one adjustment factor may be applied to adjust at least one LIC parameter derived for the LIC model. a. For example, the adjustment factor may be signaled / present in the bitstream. b. For example, the adjustment factors can be derived at both the encoder and the decoder. b. For example, at least one syntax element (eg, syntax parameter, index, variable, offset value, or integer) may be transmitted by signal at the video unit level for use in calculating at least one LIC parameter of at least one LIC model. a. For example, the video unit level may be PU / CU / block level. i. For example, in addition, the video unit level can be sequence / group of pictures / picture / slice / slice group / PB / TB / CB / PU / TU / CU / VPDU / CTU / CTU row / slice / slice / sub-picture level. b. For example, syntax element(s) may be used to adjust the value of at least one LIC parameter of at least one LIC model. i. For example, syntax element(s) may be used as indicator(s) of adjustment factor(s). ii. Alternatively, syntax element(s) may be used to represent / indicate the value of at least one LIC parameter. iii. For example, an indicator can be transmitted by signaling to adjust the parameters of a LIC model. c. For example, the derivation of LIC parameters can be based on both decoder-derived methods and signaled syntax elements. d. For example, a syntax element may be an indicator of an integer. i. For example, the value of the syntax element may be in the range of [-N, +N], for example, N=4. ii. For example, LIC parameters can be directly derived based on integers. e. For example, a syntax element may be an indicator of an index. i. For example, based on the first value of the index, the second value may be derived from a (predefined) lookup table for LIC parameter derivation. f. For example, how many syntax elements are signaled may depend on how many linear / LIC models are used for a video unit (eg, codec block). i. For example, if M LIC models are used for a video unit, M syntax elements may be signaled to be associated with the video unit. c. For example, the predicted sample value derivation of the LIC-encoded video unit can be performed based on the updated model: ValueAfter=a'*ValueBefore+b', where a'=a+Delta, b'=b-Delta*funcD. a. For example, the updated model can be used to estimate / derive prediction samples within the current block. i. Alternatively, the updated model can be used to modulate the relationship between the neighboring samples of the current block and the neighboring samples of the time-domain co-located block. b. For example, "Delta" can be the slope adjustment / offset value of the LIC model. i. For example, at least one indicator of "Delta" may be signaled in the bitstream. ii. Alternatively, the value of "Delta" can be derived based on decoded information (eg, decoded sample values, decoded prediction modes of neighboring / reference blocks). iii. For example, "Delta" can be an integer. iv. For example, “Delta” may be an integer in the range of [-N, +N], for example, N=4. v. For example, "Delta" can be a number / value / integer / constant / variable derived from an index on a lookup table. c. For example, “funcD” can be calculated by averaging the reconstructed / predicted values of all available / appropriate / possible neighboring samples of the current block (or all available / appropriate / possible neighboring samples of the time-domain co-located block). i. For example, "funcD" may be calculated by averaging neighboring / reference samples from both intra-coded blocks and inter-coded blocks. 1. Alternatively, "funcD" can be calculated by averaging only neighboring / reference samples from inter-coded blocks. ii. For example, “funcD” may be calculated by averaging all available neighboring / reference samples located to the left and / or above the current block and / or the temporally co-located block. 1. Alternatively, only a portion of such samples (eg, adjacent to the first upper left M×M unit, such as M=16 or 8 or 4 or 32) may be considered. iii. For example, (neighboring samples of) a temporally co-located block can be retrieved / pointed to by a block motion vector (or its variants). 1. Alternatively, (neighboring samples of) the temporally co-located block can be retrieved / pointed to by a rounded block motion vector (eg, rounded to integer pixel precision). iv. For example, the averaging process can be processed with (or without) a rounding factor. v. Alternatively, the averaging process can be replaced by other functions (such as summation, etc.). d. For example, the updated model (eg, with adjustments) may be used / allowed for all LIC-encoded blocks. i. Alternatively, the updated model may be used / allowed to be used for a certain class of LIC-coded blocks. 1. For example, "a certain category" can be determined based on available / appropriate / possible neighboring samples (such as both left neighboring samples and top neighboring samples are available, or only left neighboring samples are available, or only top neighboring samples are available, etc.). 2. For example, “a certain type” may be determined based on a prediction mode (such as AMVP coded or Merge coded, unidirectional prediction or bidirectional prediction, etc.). ii. Alternatively, the updated model may be used / allowed only if both left and above reference samples are available / appropriate for the video unit. iii. Alternatively, the updated model can be used / allowed only if the video unit is unidirectionally predicted. d. The adjusted information for LIC or CCLM or MM-CCLM can be coded in a predictive manner. e. The adjusted information for LIC or CCLM or MM-CCLM may be encoded and decoded using at least one context model. a. The context model may depend on codec information. b. Alternatively, it can be encoded and decoded in a bypass mode. f. For example, the neighboring / reference samples used to derive the LIC model parameters may not come from all available / appropriate / possible neighboring / reference samples on the left and above the coded block and the time-domain co-located block. a. For example, it may refer to neighboring / reference samples from both intra-coded blocks and inter-coded blocks. i. Alternatively, it may refer to neighboring / reference samples only from inter-coded blocks. b. For example, it may refer to a neighboring / reference sample located to the left (or above) of the current block and / or the time-domain co-located block. i. Alternatively, only a portion of such samples (eg, adjacent to the first upper left MxM unit, such as M=16 or 8 or 4 or 32) may be considered. c. For example, (neighboring samples of) a temporally co-located block can be retrieved / pointed to by a block motion vector (or its variant). i. Alternatively, (neighboring samples of) the temporally co-located block can be retrieved / pointed to by a rounded block motion vector (eg, rounded to integer pixel precision). g. For example, whether to apply / allow adjustments to a video unit (eg, an updated model) may depend on the encoded information. a. For example, both the original model (without adjustments) and the updated model (with adjustments) may be used / allowed. i. Alternatively, only newer models will be used / allowed. b. For example, whether adjustment-based LIC model update is allowed (or applied) can be signaled in the bitstream. i. For example, it can be at (at least) a video unit level (such as sequence / group of pictures / picture / Slice / slice group / PB / TB / CB / PU / TU / CU / VPDU / CTU / CTU line / slice / slice / sub-picture level) are transmitted through the signal. c. For example, whether to allow (or apply) the adjustment-based LIC model update can be derived at both the encoder and decoder sides. 2.3.7 Whether and / or how to apply adjustments for CCLM, MM-CCLM or LIC may depend on codec information such as block dimension, codec mode, (transformed) residual, transform, etc. a. For example, if the width and / or height and / or size of the block is less than a threshold, then the adjustment may not be applied. 2.3.8 Regarding the signaling and determination of LIC (e.g., as shown in the seventh question), the following method is proposed: a. For example, the LIC flag at the video unit level (eg, CU / PU level LIC flag) may not be transmitted through a signal, but derived at both the encoder side and the decoder side. a. For example, the CU / PU level LIC flag for AMVP coded blocks may not be signaled. b. For example, the CU / PU level LIC flag for affine AMVP coded blocks may not be signaled. c. For example, the CU / PU level LIC flag for AMVP-MERGE coded blocks may not be signaled. b. For example, whether to use LIC for a video unit may depend on coded information (eg, a decoder-derived method). a. For example, the LIC flag at the video unit level can be derived implicitly at both the encoder side and the decoder side. i. For example, the CU / PU level LIC flag for a block decoded via non-Merge (such as AMVP and / or Affine AMVP) codec can be derived implicitly. ii. For example, for Merge (and / or its variants, such as TM, BM, MHP, ADMVR, CU / PU level LIC for blocks encoded and decoded by CIIP, GPM, sbTMVP, Affine Merge, etc. Flags can be deduced implicitly. iii. For example, implicit derivation can be based on the decoder derivation method. iv. For example, implicit deduction can be based on template matching. v. For example, implicit deduction can be based on two-sided matching. b. For example, the decoder-derived cost / error / distortion may be calculated for both the non-LIC case and the LIC case, and the one with the smaller cost / error / distortion is determined to be the coding method to be used for the video unit. i. For example, template (and / or bilateral) matching costs may be calculated separately for LIC-coded video units and non-LIC-coded video units. ii. For example, the cost / error / distortion is obtained by comparing the temporal (co-located) neighboring samples in the reference picture and / or Or reference samples are derived. iii. For example, the cost / error / distortion is not derived from the current block samples in the current picture. c. For example, the coded information used for LIC mode may be neighboring / reference samples from both intra-coded blocks and inter-coded blocks. i. Alternatively, the coded information may be neighboring / reference samples only from inter-coded blocks. d. For example, the coded information used for the LIC mode may be all available neighboring / reference samples located to the left and / or above the current block and / or the time-domain co-located block. i. Alternatively, only a portion of such samples (eg, adjacent to the first upper left MxM unit, such as M=16 or 8 or 4 or 32) may be considered. e. For example, the temporally co-located block can be retrieved / pointed to by a block motion vector (or its variant). i. Alternatively, the temporally co-located block can be retrieved / pointed to by a rounded block motion vector (eg, rounded to integer pixel precision). c. For example, the Merge index of the Merge block encoded and decoded by LIC may not be transmitted through a signal (eg, derived at both the encoder and the decoder). a. For example, the motion (eg, motion vector, reference index, prediction direction, etc.) of the LIC-encoded Merge block can be derived at both the encoder side and the decoder side. b. For example, for all (or multiple, or a predefined portion) available / possible / appropriate merge candidates, multiple template (and / or bilateral) matching costs / errors / distortions can be calculated separately. The one with the smallest cost / error / distortion is determined to be the motion used for the video unit. i. For example, the template is constructed from temporally (co-located) neighboring samples and / or reference samples in a reference picture. ii. For example, the template is not constructed from the current block samples in the current picture. iii. For example, the template is constructed using a spot without LIC. iv. For example, the template is constructed using a spot with LIC. d. For example, an optimal set of LIC parameters (such as a and b calculated by a least squares fitting method) can be determined from more than one set of LIC parameters. a. For example, more than one set of LIC parameters may be applicable to a LIC-encoded video unit. b. For example, which set of LIC parameters is used for a video unit can be derived at both the encoder side and the decoder side. i. Alternatively, a syntax element (eg, an index) may be signaled that specifies the LIC parameter set to be used for the video unit. c. For example, for all appropriate LIC parameter sets, multiple template (and / or bilateral) matching costs / errors / The distortion can be calculated separately. The one with the smallest cost / error / distortion is determined as the LIC parameter set to be used for the video unit. i. For example, the template is constructed from temporally (co-located) neighboring samples and / or reference samples in a reference picture. ii. For example, the template is not constructed from the current block samples in the current picture. generally 2.3.9 Whether to apply the above disclosed method and / or how to apply the above disclosed method can be transmitted through a signal at the sequence level / picture group level / picture level / slice level / slice group level, such as in the sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / strip header / slice group header. 2.3.10 Whether to apply the method disclosed above and / or how to apply the method disclosed above can be transmitted by signal at PB / TB / CB / PU / TU / CU / VPDU / CTU / CTU row / slice / slice / sub-picture / other types of areas containing more than one sample or pixel. 2.3.11 Whether to apply the above disclosed method and / or how to apply the above disclosed method may depend on the coded information, such as block size, color format, single / dual tree partitioning, color component, slice / picture type. 3 questions There are some problems in the existing video coding and decoding technology, which will be further improved in order to obtain higher coding and decoding gain. In ECM-7.0, intraTMP is applied to both camera-captured video and screen content video. Template matching is used in intraTMP mode to find the optimal prediction block with the minimum SAD between the current and reference templates. The same process is applied on the decoder side. However, the predictions found on the decoder side may not be accurate enough for high codec efficiency. 2. In ECM-7.0, LIC is applied to compensate for local illumination changes between the current block and the reference block. Build a linear model Luma cur =a*Luma ref +b to modulate the relationship between the brightness samples of the current region and the reference region. However, the linear model can be further improved. 4 Detailed solutions The following detailed embodiments should be considered as examples to explain the general concept. These embodiments should not be interpreted in a narrow sense. In addition, these embodiments can be combined in any way. The term “video unit” or “codec unit” may refer to a picture, a slice, a slice, a codec tree block (CTB), a codec tree unit (CTU), a codec block (CB), a CU, a PU, a TU, a PB, or a TB. The term "block" may refer to a codec tree block (CTB), a codec tree unit (CTU), a codec block (CB), a CU, a PU, a TU, a PB, or a TB. It should be noted that the following terms are not limited to the specific terms defined in existing standards, and any changes in codec tools are also applicable. 4.1 Regarding the derivation of the final prediction of the intraTMP block and related issues (e.g., the first problem), the following method is proposed: a. For example, an offset-based linear / non-linear / polynomial regression model may be used to modulate the relationship between the current region and the reference region, and a filtered prediction block may be generated based on the model. a. Assume that the final sample point in the current area is sampVal = a0Y0 + a1Y1 + a2Y2 + a3Y3 + ... + a i Y i +a i+1 B is calculated, where Y0…Y i represents the value of the sample point in the reference area, B represents the bias term, a0…a i+1 represents the filter coefficient, Y i The value of can be derived by subtracting the offset from the reference sample value. i. For example, the offset can be derived based on (a plurality of) reference samples located at (a plurality of) predefined positions. ii. For example, the offset can be derived based on the average / mean / median / maximum / minimum of the samples within a reference region. iii. For example, the offset can be equal to 0. iv. For example, the bias B can be equal to a variable based on the bitDepth of the video sequence. For example, for 10 bits the capacity is 512, and for 8 bits the content is 256. 1. For example, the bias B can be equal to 1 << (bitDepth - 1), where bitDepth represents the internal bit depth of the samples of the luminance and chrominance arrays. 2. Alternatively, the bias B can be equal to 0. v. For example, Y0 to Y i can represent a plurality of samples adjacent to each other, such as Y0 being in the center and being surrounded by Y1...Y above / below / left / right of Y0. i surrounded. vi. Alternatively, Y0 to Y i can represent samples in different prediction block candidates, where there can be a total of i + 1 prediction candidates. vii. Alternatively, Y0 to Y i can represent samples in different reference rows, where there can be a total of i + 1 reference rows. viii. Alternatively, Y0 to Y i can represent samples in different columns / rows within a template, where there can be a total of i + 1 columns / rows in the template. ix. For example, the filter coefficients a0 to a i+1 can be derived based on LDL decomposition, or Gaussian elimination, or least squares, or LU decomposition, or Cholesky decomposition, etc. b. For example, a list of filtered prediction block candidates can be generated, and which one is used as the final filtered prediction block can be signaled in the bitstream. a. For example, the list can be reordered based on the template cost. b. For example, N candidates out of M candidates (such as N < M) can be selected first, and then the final filtered prediction is selected from these N candidates. i. For example, the selection of the N candidates can be performed based on the template cost. c. For example, syntax parameters (such as candidate indices) can be signaled at the block level. d. Alternatively, the final filtered prediction block can be determined implicitly based on the template cost (eg, SAD, SATD, mean-removed SAD, offset-removed SAD, etc.). c. For example, multiple filtered prediction blocks can be fused together and used as the final prediction for the current block. a. For example, the blending weights between different filtered prediction blocks can be determined based on SAD / SATD costs. b. For example, the mixing weights between different filtered prediction blocks can be predefined fixed values. c. For example, the mixing weights between different filtered prediction blocks can be determined based on LDL decomposition, or LU decomposition, or Gaussian elimination. d. For example, whether to apply blending or filtering to the final prediction block generation can be signaled in the bitstream. 4.2 Regarding the LIC model and related issues (e.g., the second question), the following approach is proposed: a. For example, LIC can be applied based on linear / nonlinear / polynomial regression models. a. For example, linear / nonlinear / polynomial regression models can be used to modulate the current template / reference / block / Relationship between a region and a reference template / reference / block / region. b. For example, the linear / nonlinear / polynomial regression model for LIC can be expressed as: sampVal = a0Y0 + a1Y1+a2Y2+a3Y3+…+a i Y i +a i+1 B, where Y0…Y i represents the value of the sample point based on the reference area, B represents the bias term, a0…a i+1 represents the filter coefficient, Y i The value of can be derived by subtracting the offset from the reference sample value. i. For example, the offset can be derived based on reference sample point(s) located at predefined position(s). ii. For example, the offset can be derived based on the average / mean / median / maximum / minimum value of the samples within the reference area. iii. For example, the offset may be equal to 0. iv. For example, the bias B may be equal to a variable based on the bitDepth of the video sequence, eg, 512 for 10-bit content and 256 for 8-bit content. 1. Alternatively, bias B can be equal to 0. v. For example, Y0 to Yi It can represent multiple samples that are adjacent to each other, such as Y0 being at the center and surrounded by Y1...Y above / below / left / right of Y0. i around. vi. Alternatively, Y0 to Y i Samples in different reference rows may be represented, where there may be a total of i+1 reference rows. vii. Alternatively, Y0 to Y i It can represent the sample points in different columns / rows in the template, where the template can have There are i+1 columns / rows in total. viii. For example, the model coefficients can be solved by minimizing the MSE between the current template and the reference template. 1. For example, filter coefficients a0 to a i+1 It can be based on LDL decomposition, Gaussian elimination, or least squares, Or LU decomposition, or Cholesky decomposition, etc. are derived. b. For example, the template size for LIC can be larger than a row and / or column of samples. a. For example, LIC prediction can be derived based on more than one row / column of neighboring samples adjacent to the current block and / or reference block. 4.3 For example, intra luma fusion can be applied based on more than two reference rows. a. For example, intra-frame luminance fusion can be applied based on linear / non-linear / polynomial regression models. b. For example, the weights of different predictions from different reference rows can be calculated based on LDL decomposition, or Gaussian elimination, or least squares, or LU decomposition, or Cholesky decomposition, etc. 4.4 For example, intra chroma blending can be applied based on more than two reference lines. c. For example, intra-frame chroma fusion can be applied based on linear / non-linear / polynomial regression models. d. For example, the weights of different predictions from different reference rows can be calculated based on LDL decomposition, or Gaussian elimination, or least squares, or LU decomposition, or Cholesky decomposition, etc. 4.5 Whether to apply the above disclosed method and / or how to apply the above disclosed method can be transmitted through a signal at the sequence level / picture group level / picture level / slice level / slice group level, such as in the sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / strip header / slice group header. 4.6 Whether to apply the above-disclosed method and / or how to apply the above-disclosed method can be transmitted by signal at PB / TB / CB / PU / TU / CU / VPDU / CTU / CTU row / slice / slice / sub-picture / other types of areas containing more than one sample or pixel. 4.7 Whether to apply the above disclosed method and / or how to apply the above disclosed method may depend on the coded information, such as block size, color format, single / dual tree partitioning, color component, slice / picture type.
[0090] More details will be discussed further below. Figure 26 FIG2 is a flowchart of a method 2600 for video processing according to an embodiment of the present disclosure. The method 2600 is implemented for converting between a current video block of a video and a bitstream of the video.
[0091] At block 2610, a filtered prediction of the current video block is determined based on an offset-based regression model. The offset-based regression model modulates the relationship between the current region and a reference region of the current video block. That is, an offset-based linear / non-linear / polynomial regression model can be used to modulate the relationship between the current region and the reference region, and a filtered prediction block can be generated based on the model.
[0092] At block 2620, conversion is performed based on the filtered prediction. In some embodiments, conversion includes encoding the current video block into a bitstream. Alternatively or additionally, in some embodiments, conversion includes decoding the current video block from a bitstream.
[0093] Method 2600 enables determining a filtered prediction block based on an offset-based regression model. Coding efficiency and / or coding effectiveness can thus be improved.
[0094] In some embodiments, the shift-based regression model includes at least one of: a shift-based linear regression model, a shift-based nonlinear regression model, or a shift-based polynomial regression model.
[0095] In some embodiments, the current video block includes an intra template matching prediction (intraTMP) block, and the filtered prediction is a final prediction for the current video block.
[0096] In some embodiments, determining the filtered prediction includes: determining a plurality of samples in a reference region; determining a plurality of values based on the plurality of samples and an offset; determining a final sample in a current region based on a weighted sum of the plurality of values and an offset value; and determining the filtered prediction based on the final sample. Assume that the final sample in the current region is obtained by sampVal=a0Y0+a1Y1+a2Y2+a3Y3+…+a i Yi +a i+1 B is calculated, where Y0…Y i represents the value of the sample point in the reference area, B represents the bias term, a0…a i+1 represents the filter coefficient, Y i The value of can be derived by subtracting the offset from the reference sample value.
[0097] In some embodiments, multiple values are determined by subtracting offsets from multiple samples. i The value of can be derived by subtracting the offset from the reference sample value.
[0098] In some embodiments, the weighted sum is determined based on a plurality of filter coefficients for a plurality of values and a bias value.
[0099] In some embodiments, the plurality of filter coefficients are determined based on at least one of: LDL decomposition or a variant of LDL decomposition, LU decomposition or a variant of LU decomposition, Cholesky decomposition or a variant of Cholesky decomposition, Gaussian elimination or a variant of Gaussian elimination, or a least squares tool or a variant of a least squares tool.
[0100] In some embodiments, the offset is determined based on at least one reference sample located at at least one predefined position.
[0101] In some embodiments, the offset is determined based on metric values of samples within the reference region.
[0102] In some embodiments, the metric value comprises one of: an average value, a mean value, a median value, a maximum value, or a minimum value.
[0103] In some embodiments, the offset is zero.
[0104] In some embodiments, the offset value is determined based on the bit depth of the video sequence of the video. In one example, the bit depth is 10 bits and the offset value is 512. In another example, the bit depth is 8 bits and the offset value is 256.
[0105] In some embodiments, the bias value is determined by: 1<<(bitDepth-1), where bitDepth represents the internal bit depth of the sample of at least one of the following: luma array or chroma array, and << represents an arithmetic left shift operation.
[0106] In some embodiments, the bias value is zero. For example, bias B can be equal to 0.
[0107] In some embodiments, the plurality of spots are adjacent to each other.
[0108] In some embodiments, the plurality of sample points include a first sample point at the center of the reference region, a second sample point above the first sample point, a third sample point below the first sample point, a fourth sample point to the left of the first sample point, and a fifth sample point to the right of the first sample point. For example, Y0 to Y i can represent a plurality of sample points adjacent to each other, such as Y0 being at the center and being surrounded by Y1…Y above / below / left / right of Y0 i Surrounded by.
[0109] In some embodiments, the plurality of sample points are among the plurality of prediction candidates of the current video block. The number of the plurality of sample points may be equal to the number of the plurality of prediction candidates.
[0110] In some embodiments, the plurality of sample points are among the plurality of reference lines of the current video block. The number of the plurality of sample points may be equal to the number of the plurality of reference lines.
[0111] In some embodiments, the plurality of sample points are in multiple columns or rows within a template of the current video block. The number of the plurality of sample points may be equal to the number of the multiple columns or rows.
[0112] In some embodiments, a list of filtered prediction candidates for the current video block is determined, and an indication in the bitstream specifies the filtered prediction among the list of filtered prediction candidates. For example, a list of filtered prediction block candidates may be generated, and which one is used as the final filtered prediction block may be signaled in the bitstream.
[0113] In some embodiments, the list of filtered prediction candidates is re-ordered based on a template cost.
[0114] In some embodiments, a first number of filtered prediction candidates are selected from the list of filtered prediction candidates, and a filtered prediction is selected from the first number of filtered prediction candidates, the first number being less than the total number of candidates in the list of filtered prediction candidates. For example, N candidates out of M candidates (such as N<M) may be selected first, and then the final filtered prediction is selected from those N candidates.
[0115] In some embodiments, the first number of filtered prediction candidates are selected based on the template cost of the list of filtered prediction candidates.
[0116] In some embodiments, the indication includes a syntax parameter included at the block level in the bitstream. For example, the syntax parameter may be an index of a candidate.
[0117] In some embodiments, a list of filtered prediction candidates for the current video block is determined, and the filtered prediction is determined based on the template cost of the filtered prediction candidates in the list.
[0118] In some embodiments, the template cost of the filtered prediction candidate comprises one of the following: sum of absolute differences (SAD), sum of absolute transformed differences (SATD), mean removed SAD, or offset removed SAD.
[0119] In some embodiments, determining the filtered prediction includes: determining a plurality of filtered prediction candidates; and determining a filtered prediction for the current video block based on a fusion of the plurality of filtered prediction candidates. For example, the plurality of filtered prediction blocks may be fused together and used as a final prediction for the current block.
[0120] In some embodiments, fusion is determined based on a blending weight of multiple filtered prediction candidates, where the blending weight is determined based on a sum of absolute differences (SAD) cost or a sum of absolute transformed differences (SATD) cost. That is, the blending weight between different filtered prediction blocks can be determined based on the SAD / SATD cost.
[0121] In some embodiments, fusion is determined based on a blending weight of a plurality of filtered prediction candidates, where the blending weight is a predefined fixed value. That is, the blending weight between different filtered prediction blocks may be a predefined fixed value.
[0122] In some embodiments, fusion is determined based on a blending weight of the plurality of filtered prediction candidates, where the blending weight is determined based on at least one of: LDL decomposition, LU decomposition, or Gaussian elimination. That is, the blending weight between different filtered prediction blocks may be determined based on LDL decomposition, LU decomposition, or Gaussian elimination.
[0123] In some embodiments, whether blending or filtering is applied to the final prediction for the current video block is included in the bitstream. For example, whether blending or filtering is applied to the final prediction block generation can be signaled in the bitstream.
[0124] According to another embodiment of the present disclosure, a non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a bitstream of a video generated by a method performed by an apparatus for video processing. In the method, a filtered prediction of a current video block of the video is determined based on an offset-based regression model. The offset-based regression model modulates a relationship between a current region and a reference region of the current video block. The bitstream is generated based on the filtered prediction.
[0125] According to further embodiments of the present disclosure, a method for storing a bitstream of a video is provided. In this method, a filtered prediction of a current video block of the video is determined based on an offset-based regression model. The offset-based regression model modulates a relationship between a current region and a reference region of the current video block. A bitstream is generated based on the filtered prediction. The bitstream is stored in a non-transitory computer-readable recording medium.
[0126] Figure 27 FIG27 is a flowchart of a method 2700 for video processing according to an embodiment of the present disclosure. The method 2700 is implemented for converting between a current video block of a video and a bitstream of the video.
[0127] At block 2710, a prediction for a current video block is determined based on local illumination compensation (LIC). LIC is based on a regression model that modulates a relationship between a current region and a reference region of the current video block for the LIC.
[0128] At block 2720, conversion is performed based on the prediction. In some embodiments, conversion includes encoding the current video block into a bitstream. Alternatively or additionally, in some embodiments, conversion includes decoding the current video block from a bitstream.
[0129] Method 2700 enables determining a prediction of a block by using a LIC based on a regression model. Codec efficiency and / or codec effectiveness can thus be improved.
[0130] In some embodiments, the regression model includes at least one of: a linear regression model, a nonlinear regression model, or a polynomial regression model.
[0131] In some embodiments, determining the prediction includes: determining a plurality of samples in a reference region; determining a plurality of values based on the plurality of samples and an offset; determining a final sample in the current region based on a weighted sum of the plurality of values and the offset value; and determining the prediction based on the final sample.
[0132] In some embodiments, the plurality of values are determined by subtracting an offset from the plurality of samples.
[0133] In some embodiments, the weighted sum is determined based on a plurality of filter coefficients for a plurality of values and a bias value.
[0134] In some embodiments, the plurality of filter coefficients are determined based on at least one of: LDL decomposition or a variant of LDL decomposition, LU decomposition or a variant of LU decomposition, Cholesky decomposition or a variant of Cholesky decomposition, Gaussian elimination or a variant of Gaussian elimination, or a least squares tool or a variant of a least squares tool.
[0135] In some embodiments, the offset is determined based on at least one reference sample located at at least one predefined position.
[0136] In some embodiments, the offset is determined based on metric values of samples within the reference region.
[0137] In some embodiments, the metric value comprises one of: an average value, a mean value, a median value, a maximum value, or a minimum value.
[0138] In some embodiments, the offset is zero.
[0139] In some embodiments, the offset value is determined based on the bit depth of the video sequence of the video. In one example, the bit depth is 10 bits and the offset value is 512. In another example, the bit depth is 8 bits and the offset value is 256.
[0140] In some embodiments, the bias value is zero.
[0141] In some embodiments, the plurality of spots are adjacent to each other.
[0142] In some embodiments, the plurality of sample points include a first sample point at the center of the reference area, a second sample point above the first sample point, a third sample point below the first sample point, a fourth sample point to the left of the first sample point, and a fifth sample point to the right of the first sample point.
[0143] In some embodiments, the plurality of samples are in a plurality of reference rows of the current video block.
[0144] In some embodiments, the number of the plurality of samples is equal to the number of the plurality of reference rows.
[0145] In some embodiments, the multiple samples are in multiple columns or rows within the template of the current video block.
[0146] In some embodiments, the number of the plurality of samples is equal to the number of the plurality of columns or rows.
[0147] In some embodiments, the template size for LIC is larger than at least one of: a row of samples or a column of samples.
[0148] In some embodiments, the prediction by the LIC is determined based on multiple rows or columns of neighboring samples.
[0149] In some embodiments, the plurality of rows or columns of neighboring samples are adjacent to at least one of: the current video block or a reference block of the current video block.
[0150] According to another embodiment of the present disclosure, a non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a bitstream of a video generated by a method performed by an apparatus for video processing. In the method, a prediction of a current video block of the video is determined based on local illumination compensation (LIC). The LIC is based on a regression model that modulates a relationship between a current region and a reference region of the current video block for the LIC. The bitstream is generated based on the prediction.
[0151] According to further embodiments of the present disclosure, a method for storing a bitstream of a video is provided. In this method, a prediction of a current video block of the video is determined based on local illumination compensation (LIC). The LIC is based on a regression model that modulates a relationship between a current region and a reference region of the current video block for the LIC. A bitstream is generated based on the prediction. The bitstream is stored in a non-transitory computer-readable recording medium.
[0152] Figure 28 FIG28 is a flowchart of a method 2800 for video processing according to an embodiment of the present disclosure. The method 2800 is implemented for converting between a current video block of a video and a bitstream of the video.
[0153] At block 2810 , at least one intra blend for at least one color component is applied to a current video block based on more than two reference rows.
[0154] At block 2820, conversion is performed based on the application. In some embodiments, conversion includes encoding the current video block into a bitstream. Alternatively or additionally, in some embodiments, conversion includes decoding the current video block from a bitstream.
[0155] Method 2800 enables intra-frame blending for a color component, such as a luma component or a chroma component. In this way, codec efficiency and / or codec effectiveness can be improved.
[0156] In some embodiments, at least one intra blend for at least one color component includes at least one of: intra luma blending or intra chroma blending. For example, intra luma blending and / or intra chroma blending may be applied based on more than two reference rows.
[0157] In some embodiments, at least one intra-frame fusion is applied based on a regression model. For example, the regression model may include at least one of the following: a linear regression model, a nonlinear regression model, or a polynomial regression model. For example, intra-frame luma fusion and / or intra-frame chroma fusion may be applied based on a linear / nonlinear / polynomial regression model.
[0158] In some embodiments, applying at least one intra blend to the current video block comprises: determining a plurality of predictions of at least one color component based on a plurality of reference rows of the current video block; and determining a blend of the plurality of predictions based on a plurality of weights.
[0159] In some embodiments, the plurality of weights are determined based on at least one of: LDL decomposition or a variant thereof, LU decomposition or a variant thereof, Cholesky decomposition or a variant thereof, Gaussian elimination or a variant thereof, or a least squares tool or a variant thereof. For example, weights for different predictions from different reference rows can be calculated based on LDL decomposition, Gaussian elimination, least squares, LU decomposition, Cholesky decomposition, etc.
[0160] According to another embodiment of the present disclosure, a non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a bitstream of a video generated by a method performed by an apparatus for video processing. In the method, at least one intra-frame blending for at least one color component is applied to a current video block of the video based on more than two reference lines. The bitstream is generated based on the application.
[0161] According to further embodiments of the present disclosure, a method for storing a video bitstream is provided. In this method, at least one intra-frame blending for at least one color component is applied to a current video block of the video based on more than two reference lines. The bitstream is generated based on the application and stored in a non-transitory computer-readable recording medium.
[0162] In some embodiments, information about whether and / or how to apply method 2600, method 2700, and / or method 2800 is included in the bitstream.
[0163] In some embodiments, the information is indicated at one of: sequence level, group of pictures level, picture level, slice level, or slice group level.
[0164] In some embodiments, the information is indicated in a sequence header, a picture header, a sequence parameter set (SPS), a video parameter set (VPS), a decoding parameter set (DPS), decoding capability information (DCI), a picture parameter set (PPS), an adaptation parameter set (APS), a slice header, or a slice group header.
[0165] In some embodiments, information is indicated in an area comprising more than one sample or pixel.
[0166] In some embodiments, the region includes one of the following: a prediction block (PB), a transform block (TB), a codec block (CB), a prediction unit (PU), a transform unit (TU), a codec unit (CU), a virtual pipeline data unit (VPDU), a codec tree unit (CTU), a CTU row, a slice, a slice, or a sub-picture.
[0167] In some embodiments, the method 2800 further includes determining information based on the encoded and decoded information.
[0168] In some embodiments, the encoded information includes at least one of: block size, color format, single-tree partitioning or dual-tree partitioning, color component, slice type, or picture type.
[0169] It should be understood that the methods 2600, 2700, and 2800 can be applied individually or in any combination. By using the methods 2600, 2700, and / or 2800 individually or in combination, codec effectiveness and / or codec efficiency can be improved.
[0170] The embodiments of the present disclosure may be described in view of the following items, features of which may be combined in any reasonable way.
[0171] Item 1. A method for video processing, comprising: for a conversion between a current video block of a video and a bitstream of the video, determining a filtered prediction of the current video block based on an offset-based regression model, wherein the offset-based regression model modulates a relationship between a current region and a reference region of the current video block; and performing the conversion based on the filtered prediction.
[0172] Item 2. The method of Item 1, wherein the offset-based regression model comprises at least one of: an offset-based linear regression model, an offset-based nonlinear regression model, or an offset-based polynomial regression model.
[0173] Item 3. The method of Item 1 or 2, wherein the current video block comprises an intra template matching prediction (intraTMP) block, and the filtered prediction is a final prediction for the current video block.
[0174] Item 4. A method according to any one of Items 1-3, wherein determining the filtered prediction comprises: determining a plurality of samples in the reference region; determining a plurality of values based on the plurality of samples and an offset; determining a final sample in the current region based on a weighted sum of the plurality of values and a bias value; and determining the filtered prediction based on the final sample.
[0175] Clause 5. The method of clause 4, wherein the plurality of values are determined by subtracting the offset from the plurality of samples.
[0176] Clause 6. The method of clause 4 or 5, wherein the weighted sum is determined based on a plurality of filter coefficients for the plurality of values and the bias value.
[0177] Item 7. A method according to Item 6, wherein the multiple filter coefficients are determined based on at least one of: LDL decomposition or a variant of LDL decomposition, LU decomposition or a variant of LU decomposition, Cholesky decomposition or a variant of Cholesky decomposition, Gaussian elimination or a variant of Gaussian elimination, or a least squares tool or a variant of a least squares tool.
[0178] Clause 8. The method according to any of Clauses 4-8, wherein the offset is determined based on at least one reference sample located at at least one predefined position.
[0179] Clause 9. A method according to any one of Clauses 4-8, wherein the offset is determined based on a metric value of a sample point within the reference region.
[0180] Clause 10. The method of clause 9, wherein the metric value comprises one of the following: an average value, a mean value, a median value, a maximum value, or a minimum value.
[0181] Clause 11. The method of any one of clauses 4-10, wherein the offset is zero.
[0182] Clause 12. A method according to any one of clauses 4-11, wherein the bias value is determined based on a bit depth of a video sequence of the video.
[0183] Item 13. The method of Item 12, wherein the bit depth is 10 bits and the offset value is 512.
[0184] Item 14. The method of Item 12, wherein the bit depth is 8 bits and the offset value is 256.
[0185] Item 15. The method of Item 12, wherein the bias value is determined by: 1<< (bitDepth-1), where bitDepth represents the internal bit depth of the samples of at least one item in the luma array or the chroma array, and << represents an arithmetic left shift operation.
[0186] Item 16. The method of any one of Items 4-11, wherein the bias value is zero.
[0187] Item 17. The method according to any one of Items 4-16, wherein the plurality of spots are adjacent to each other.
[0188] Item 18. A method according to Item 17, wherein the multiple sample points include a first sample point at the center of the reference area, a second sample point above the first sample point, a third sample point below the first sample point, a fourth sample point to the left of the first sample point, and a fifth sample point to the right of the first sample point.
[0189] Item 19. The method of any one of Items 4-16, wherein the plurality of samples are among a plurality of prediction candidates for the current video block.
[0190] Clause 20. The method of clause 19, wherein the number of the plurality of samples is equal to the number of the plurality of prediction candidates.
[0191] Item 21. The method of any one of Items 4-16, wherein the plurality of samples are in a plurality of reference rows of the current video block.
[0192] Item 22. The method of Item 21, wherein the number of the plurality of samples is equal to the number of the plurality of reference rows.
[0193] Item 23. The method of any one of Items 4-16, wherein the plurality of samples are in a plurality of columns or rows within a template of the current video block.
[0194] Item 24. The method of Item 23, wherein the number of the plurality of samples is equal to the number of the plurality of columns or rows.
[0195] Item 25. The method of any one of Items 1-24, wherein a list of filtered prediction candidates for the current video block is determined, an indication in the bitstream specifying the filtered prediction in the list of filtered prediction candidates.
[0196] Clause 26. The method of clause 25, wherein the list of filtered prediction candidates is reordered based on template costs.
[0197] Item 27. A method according to item 25 or 26, wherein a first number of filtered prediction candidates are selected from the list of filtered prediction candidates, and the filtered prediction is selected from the first number of filtered prediction candidates, the first number being less than the total number of candidates in the list of filtered prediction candidates.
[0198] Item 28. The method of Item 27, wherein the first number of filtered prediction candidates is selected based on a template cost of the list of filtered prediction candidates.
[0199] Clause 29. A method according to any of Clauses 25-28, wherein the indication comprises syntax parameters included at a block level in the bitstream.
[0200] Clause 30. The method of clause 29, wherein the syntax parameter comprises an index of a candidate.
[0201] Item 31. A method according to any of Items 1-24, wherein a list of filtered prediction candidates for the current video block is determined, and the filtered prediction is determined based on template costs of the filtered prediction candidates in the list.
[0202] Item 32. The method of Item 31, wherein the template cost of the filtered prediction candidate comprises one of the following: sum of absolute differences (SAD), sum of absolute transformed differences (SATD), mean removed SAD, or offset removed SAD.
[0203] Item 33. A method according to any one of Items 1-32, wherein determining the filtered prediction comprises: determining a plurality of filtered prediction candidates; and determining the filtered prediction for the current video block based on a fusion of the plurality of filtered prediction candidates.
[0204] Item 34. A method according to Item 33, wherein the fusion is determined based on a blending weight of the multiple filtered prediction candidates, and the blending weight is determined based on a sum of absolute differences (SAD) cost or a sum of absolute transformed differences (SATD) cost.
[0205] Item 35. The method according to Item 33, wherein the fusion is determined based on a blending weight of the plurality of filtered prediction candidates, the blending weight being a predefined fixed value.
[0206] Item 36. The method of Item 33, wherein the fusion is determined based on a blending weight of the plurality of filtered prediction candidates, the blending weight being determined based on at least one of: LDL decomposition, LU decomposition, or Gaussian elimination.
[0207] Item 37. The method of any one of Items 33-36, wherein whether blending or filtering is applied for the final prediction of the current video block is included in the bitstream.
[0208] Item 38. A method for video processing, comprising: determining, for a conversion between a current video block of a video and a bitstream of the video, a prediction of the current video block based on local illumination compensation (LIC), the LIC being based on a regression model that modulates a relationship between a current region and a reference region of the current video block for the LIC; and performing the conversion based on the prediction.
[0209] Item 39. The method of Item 38, wherein the regression model comprises at least one of: a linear regression model, a nonlinear regression model, or a polynomial regression model.
[0210] Item 40. A method according to Item 38 or 39, wherein determining the prediction includes: determining a plurality of sample points in the reference area; determining a plurality of values based on the plurality of sample points and an offset; determining a final sample point in the current area based on a weighted sum of the plurality of values and a bias value; and determining the prediction based on the final sample point.
[0211] Item 41. The method of Item 40, wherein the plurality of values are determined by subtracting the offset from the plurality of samples.
[0212] Item 42. The method of Item 40 or 41, wherein the weighted sum is determined based on a plurality of filter coefficients for the plurality of values and the bias value.
[0213] Item 43. A method according to Item 42, wherein the plurality of filter coefficients are determined based on at least one of: LDL decomposition or a variant of LDL decomposition, LU decomposition or a variant of LU decomposition, Cholesky decomposition or a variant of Cholesky decomposition, Gaussian elimination or a variant of Gaussian elimination, or a least squares tool or a variant of a least squares tool.
[0214] Item 44. The method of any one of Items 40-43, wherein the offset is determined based on at least one reference sample located at at least one predefined position.
[0215] Item 45. The method of any one of Items 40-43, wherein the offset is determined based on a metric value of a sample point within the reference region.
[0216] Item 46. The method of Item 45, wherein the metric value comprises one of the following: an average value, a mean value, a median value, a maximum value, or a minimum value.
[0217] Item 47. The method of any one of Items 40-46, wherein the offset is zero.
[0218] Item 48. A method according to any of Items 40-47, wherein the bias value is determined based on a bit depth of a video sequence of the video.
[0219] Item 49. The method of Item 48, wherein the bit depth is 10 bits and the offset value is 512.
[0220] Item 50. The method of Item 48, wherein the bit depth is 8 bits and the offset value is 256.
[0221] Item 51. The method of any one of Items 40-47, wherein the bias value is zero.
[0222] Item 52. The method according to any one of Items 40-51, wherein the plurality of spots are adjacent to each other.
[0223] Item 53. A method according to Item 52, wherein the multiple sample points include a first sample point at the center of the reference area, a second sample point above the first sample point, a third sample point below the first sample point, a fourth sample point to the left of the first sample point, and a fifth sample point to the right of the first sample point.
[0224] Item 54. The method of any one of Items 40-51, wherein the plurality of samples are in a plurality of reference rows of the current video block.
[0225] Item 55. The method of Item 54, wherein the number of the plurality of samples is equal to the number of the plurality of reference rows.
[0226] Item 56. The method of any one of Items 40-51, wherein the plurality of samples are in a plurality of columns or rows within a template of the current video block.
[0227] Item 57. The method of Item 56, wherein the number of the plurality of samples is equal to the number of the plurality of columns or rows.
[0228] Item 58. The method according to any one of Items 38-57, wherein the template size for LIC is larger than at least one of: a row of spots or a column of spots.
[0229] Item 59. The method of any one of Items 38-58, wherein the prediction by the LIC is determined based on a plurality of rows or columns of neighboring samples.
[0230] Item 60. The method of Item 59, wherein the plurality of rows or columns of the neighboring samples are adjacent to at least one of: the current video block or a reference block of the current video block.
[0231] Item 61. A method for video processing, comprising: for converting between a current video block of a video and a bitstream of the video, applying at least one intra-frame blending for at least one color component to the current video block based on more than two reference lines; and performing the conversion based on the application.
[0232] Item 62. The method of Item 61, wherein the at least one intra-frame blend for at least one color component comprises at least one of: an intra-frame luma blend or an intra-frame chroma blend.
[0233] Item 63. A method according to Item 61 or 62, wherein the at least one intra-frame fusion is applied based on a regression model.
[0234] Item 64. The method of Item 63, wherein the regression model comprises at least one of: a linear regression model, a nonlinear regression model, or a polynomial regression model.
[0235] Item 65. A method according to any one of Items 61-64, wherein applying the at least one intra-frame fusion to the current video block includes: determining multiple predictions of the at least one color component based on multiple reference rows of the current video block; and determining a fusion of the multiple predictions based on multiple weights.
[0236] Item 66. A method according to Item 65, wherein the multiple weights are determined based on at least one of: LDL decomposition or a variant of LDL decomposition, LU decomposition or a variant of LU decomposition, Cholesky decomposition or a variant of Cholesky decomposition, Gaussian elimination or a variant of Gaussian elimination, or a least squares tool or a variant of a least squares tool.
[0237] Item 67. A method according to any of Items 1-66, wherein information on whether to apply the method and / or how to apply the method is included in the bitstream.
[0238] Item 68. The method of Item 67, wherein the information is indicated at one of: sequence level, group of pictures level, picture level, slice level, or slice group level.
[0239] Item 69. A method according to item 67 or 68, wherein the information is indicated in a sequence header, a picture header, a sequence parameter set (SPS), a video parameter set (VPS), a decoding parameter set (DPS), decoding capability information (DCI), a picture parameter set (PPS), an adaptation parameter set (APS), a slice header or a slice group header.
[0240] Item 70. A method according to any one of Items 67-69, wherein the information is indicated in an area comprising more than one sample or pixel.
[0241] Item 71. A method according to item 70, wherein the region includes one of the following: a prediction block (PB), a transform block (TB), a codec block (CB), a prediction unit (PU), a transform unit (TU), a codec unit (CU), a virtual pipeline data unit (VPDU), a codec tree unit (CTU), a CTU row, a slice, a slice or a sub-picture.
[0242] Item 72. The method according to any one of Items 67-71 further comprises: determining the information based on the encoded and decoded information.
[0243] Item 73. The method of Item 72, wherein the encoded information comprises at least one of: block size, color format, single-tree partitioning or dual-tree partitioning, color component, slice type, or picture type.
[0244] Item 74. The method of any one of Items 1-73, wherein the converting comprises encoding the current video block into the bitstream.
[0245] Item 75. The method of any one of Items 1-73, wherein the converting comprises decoding the current video block from the bitstream.
[0246] Item 76. An apparatus for video processing, comprising a processor and non-volatile memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of items 1-75.
[0247] Item 77. A non-transitory computer-readable storage medium storing instructions for causing a processor to perform the method according to any one of Items 1-75.
[0248] Item 78. A non-transitory computer-readable recording medium storing a bitstream of a video generated by a method performed by an apparatus for video processing, wherein the method comprises: determining a filtered prediction of a current video block of the video based on an offset-based regression model, the offset-based regression model modulating a relationship between a current region and a reference region of the current video block; and generating the bitstream based on the filtered prediction.
[0249] Item 79. A method for storing a bitstream of a video, comprising: determining a filtered prediction of a current video block of the video based on an offset-based regression model, wherein the offset-based regression model modulates a relationship between a current region and a reference region of the current video block; generating the bitstream based on the filtered prediction; and storing the bitstream in a non-transitory computer-readable recording medium.
[0250] Item 80. A non-transitory computer-readable recording medium storing a bitstream of a video generated by a method performed by an apparatus for video processing, wherein the method comprises: determining a prediction of a current video block of the video based on local illumination compensation (LIC), the LIC being based on a regression model that modulates a relationship between a current region and a reference region of the current video block for the LIC; and generating the bitstream based on the prediction.
[0251] Item 81. A method for storing a bitstream of a video, comprising: determining a prediction of a current video block of the video based on local illumination compensation (LIC), the LIC being based on a regression model that modulates a relationship between a current region and a reference region of the current video block for the LIC; generating the bitstream based on the prediction; and storing the bitstream in a non-transitory computer-readable recording medium.
[0252] Item 82. A non-transitory computer-readable recording medium storing a bitstream of a video generated by a method performed by an apparatus for video processing, wherein the method comprises: applying at least one intra-frame blending for at least one color component to a current video block of the video based on more than two reference lines; and generating the bitstream based on the application.
[0253] Item 83. A method for storing a bitstream of a video, comprising: applying at least one intra-frame blending for at least one color component to a current video block of the video based on more than two reference lines; generating the bitstream based on the applying; and storing the bitstream in a non-transitory computer-readable recording medium. Example device
[0254] Figure 29 A block diagram of a computing device 2900 in which various embodiments of the present disclosure may be implemented is shown. The computing device 2900 may be implemented as a source device 110 (or video encoder 114 or 200) or a destination device 120 (or video decoder 124 or 300), or may be included in a source device 110 (or video encoder 114 or 200) or a destination device 120 (or video decoder 124 or 300).
[0255] It should be understood that Figure 29 The computing device 2900 shown in FIG. 2 is for illustrative purposes only and is not intended to in any way imply any limitation on the functionality and scope of the disclosed embodiments.
[0256] like Figure 29 As shown, computing device 2900 comprises a general computing device 2900. Computing device 2900 may include at least one or more processors or processing units 2910, memory 2920, storage unit 2930, one or more communication units 2940, one or more input devices 2950, and one or more output devices 2960.
[0257] In some embodiments, computing device 2900 can be implemented as any user terminal or server terminal with computing capability. A server terminal can be a server, a large computing device, etc. provided by a service provider. A user terminal can be, for example, any type of mobile terminal, fixed terminal, or portable terminal, including a mobile phone, a station, a unit, a device, a multimedia computer, a multimedia tablet computer, an Internet node, a communicator, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an electronic book device, a gaming device, or any combination thereof, and includes accessories and peripherals of these devices, or any combination thereof. It is conceivable that computing device 2900 can support any type of interface to a user (such as a "wearable" circuit device, etc.).
[0258] The processing unit 2910 may be a physical processor or a virtual processor and may implement various processes based on a program stored in the memory 2920. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capabilities of the computing device 2900. The processing unit 2910 may also be referred to as a central processing unit (CPU), a microprocessor, a controller, or a microcontroller.
[0259] The computing device 2900 typically includes various computer storage media. Such media can be any media accessible by the computing device 2900, including but not limited to volatile media and non-volatile media, or removable media and non-removable media. The memory 2920 can be a volatile memory (e.g., registers, cache, random access memory (RAM)), a non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM) or flash memory) or any combination thereof. The storage unit 2930 can be any removable or non-removable medium and can include machine-readable media, such as memory, flash drive, disk or other media that can be used to store information and / or data and can be accessed in the computing device 2900.
[0260] The computing device 2900 may also include additional removable / non-removable storage media, volatile / non-volatile storage media. Figure 29 Although not shown, a magnetic disk drive for reading from and / or writing to a removable nonvolatile magnetic disk, and an optical disk drive for reading from and / or writing to a removable nonvolatile optical disk may be provided. In this case, each drive may be connected to a bus (not shown) via one or more data medium interfaces.
[0261] The communication unit 2940 communicates with another computing device via a communication medium. In addition, the functions of the components in the computing device 2900 can be implemented by a single computing cluster or multiple computing machines communicating via a communication connection. Thus, the computing device 2900 can operate in a networked environment using logical connections to one or more other servers, networked personal computers (PCs), or other general-purpose network nodes.
[0262] Input device 2950 may be one or more of various input devices, such as a mouse, keyboard, trackball, voice input device, etc. Output device 2960 may be one or more of various output devices, such as a display, speaker, printer, etc. Computing device 2900 may also communicate with one or more external devices (not shown) via communication unit 2940, such as storage devices and display devices, one or more devices that enable a user to interact with computing device 2900, or any device that enables computing device 2900 to communicate with one or more other computing devices (e.g., a network card, modem, etc.), if desired. Such communication may be performed via an input / output (I / O) interface (not shown).
[0263] In some embodiments, some or all components of the computing device 2900 may also be arranged in a cloud computing architecture rather than being integrated into a single device. In a cloud computing architecture, components can be provided remotely and can work together to implement the functionality described in this disclosure. In some embodiments, cloud computing provides computing, software, data access, and storage services without requiring the end user to know the physical location or configuration of the systems or hardware providing these services. In various embodiments, cloud computing provides services via a wide area network (such as the Internet) using appropriate protocols. For example, a cloud computing provider provides an application via a wide area network that can be accessed through a web browser or any other computing component. The software or components of the cloud computing architecture and the corresponding data may be stored on servers at a remote location. Computing resources in a cloud computing environment can be consolidated or distributed across locations in remote data centers. Cloud computing infrastructure can provide services through shared data centers, although they appear to be a single access point for users. Therefore, cloud computing architecture can be used to provide the components and functionality described herein from a service provider at a remote location. Alternatively, the components and functionality described herein can be provided from a conventional server or installed directly or otherwise on a client device.
[0264] In an embodiment of the present disclosure, the computing device 2900 may be used to implement video encoding / decoding. The memory 2920 may include one or more video encoding / decoding modules 2925 having one or more program instructions. These modules are accessible and executable by the processing unit 2910 to perform the functions of the various embodiments described herein.
[0265] In an example embodiment performing video encoding, an input device 2950 may receive video data as input 2970 to be encoded. The video data may be processed, for example, by a video codec module 2925 to generate an encoded bitstream. The encoded bitstream may be provided as output 2980 via an output device 2960.
[0266] In an example embodiment performing video decoding, an input device 2950 may receive an encoded bitstream as input 2970. The encoded bitstream may be processed by, for example, a video codec module 2925 to generate decoded video data. The decoded video data may be provided as output 2980 via an output device 2960.
[0267] Although the present disclosure has been specifically shown and described with reference to the preferred embodiments of the present disclosure, it will be understood by those skilled in the art that various changes in form and details may be made without departing from the spirit and scope of the present application as defined by the appended claims. Such variations are intended to be encompassed by the scope of the present application. Therefore, the foregoing description of the embodiments of the present application is not intended to be limiting.
Claims
1. A method for video processing, comprising: determining, for a conversion between a current video block of a video and a bitstream of the video, a filtered prediction of the current video block based on an offset-based regression model that modulates a relationship between a current region and a reference region of the current video block; as well as The converting is performed based on the filtered prediction.
2. The method of claim 1 , wherein the shift-based regression model comprises at least one of: Based on the linear regression model of the shift, A nonlinear regression model based on the shift, or Shift-based polynomial regression model.
3. The method of claim 1 or 2, wherein the current video block comprises an intra template matching prediction (intraTMP) block, and the filtered prediction is a final prediction for the current video block.
4. The method of any one of claims 1 to 3, wherein determining the filtered prediction comprises: determining a plurality of sample points in the reference area; determining a plurality of values based on the plurality of samples and the offset; determining a final sample point in the current area based on a weighted sum of the multiple values and a bias value; as well as The filtered prediction is determined based on the final sample. The method of claim 4 , wherein the plurality of values are determined by subtracting the offset from the plurality of samples.
6. The method according to claim 4 or 5, wherein the weighted sum is determined based on a plurality of filter coefficients for the plurality of values and the bias value.
7. The method of claim 6 , wherein the plurality of filter coefficients are determined based on at least one of: LDL breakdown or variants of LDL breakdown, LU decomposition or a variant of LU decomposition, Cholesky decomposition or a variant of Cholesky decomposition, Gaussian elimination or a variant of Gaussian elimination, or The Least Squares tool or a variation of the Least Squares tool.
8. The method according to any one of claims 4 to 8, wherein the offset is determined based on at least one reference sample located at at least one predefined position.
9. The method according to any one of claims 4 to 8, wherein the offset is determined based on a metric value of a sample point within the reference area.
10. The method of claim 9, wherein the metric value comprises one of the following: an average value, a mean value, a median value, a maximum value, or a minimum value.
11. The method of any one of claims 4-10, wherein the offset is zero.
12. The method according to any one of claims 4 to 11, wherein the offset value is determined based on a bit depth of a video sequence of the video. The method of claim 12 , wherein the bit depth is 10 bits and the offset value is 512. The method of claim 12 , wherein the bit depth is 8 bits and the offset value is 256.
15. The method according to claim 12, wherein the bias value is determined by: 1<< (bitDepth-1), where bitDepth represents the internal bit depth of the samples of at least one item in the luma array or the chroma array, and << represents an arithmetic left shift operation.
16. The method of any one of claims 4-11, wherein the bias value is zero.
17. The method according to any one of claims 4 to 16, wherein the plurality of spots are adjacent to each other.
18. The method of claim 17, wherein the plurality of sample points include a first sample point at the center of the reference area, a second sample point above the first sample point, a third sample point below the first sample point, a fourth sample point to the left of the first sample point, and a fifth sample point to the right of the first sample point.
19. The method according to any one of claims 4 to 16, wherein the plurality of samples are among a plurality of prediction candidates for the current video block.
20. The method according to claim 19, wherein the number of the plurality of samples is equal to the number of the plurality of prediction candidates.
21. The method of any one of claims 4-16, wherein the plurality of samples are in a plurality of reference rows of the current video block.
22. The method according to claim 21, wherein the number of the plurality of samples is equal to the number of the plurality of reference lines.
23. The method of any one of claims 4-16, wherein the plurality of samples are in a plurality of columns or rows within a template of the current video block.
24. The method of claim 23, wherein the number of the plurality of samples is equal to the number of the plurality of columns or rows.
25. The method of any one of claims 1-24, wherein a list of filtered prediction candidates for the current video block is determined, an indication in the bitstream specifying the filtered prediction in the list of filtered prediction candidates.
26. The method of claim 25, wherein the list of filtered prediction candidates is reordered based on template costs.
27. A method according to claim 25 or 26, wherein a first number of filtered prediction candidates are selected from the list of filtered prediction candidates, and the filtered prediction is selected from the first number of filtered prediction candidates, the first number being less than the total number of candidates in the list of filtered prediction candidates.
28. The method of claim 27, wherein the first number of filtered prediction candidates are selected based on template costs of the list of filtered prediction candidates.
29. The method according to any of claims 25-28, wherein the indication comprises syntax parameters included at block level in the bitstream.
30. The method of claim 29, wherein the syntax parameter comprises an index of a candidate.
31. The method of any one of claims 1-24, wherein a list of filtered prediction candidates for the current video block is determined, and the filtered prediction is determined based on template costs of the filtered prediction candidates in the list.
32. The method of claim 31 , wherein the template cost of the filtered prediction candidate comprises one of: Sum of Absolute Differences (SAD), The sum of absolute transformed differences (SATD), SAD with mean removal, or Offset removed SAD.
33. The method of any one of claims 1-32, wherein determining the filtered prediction comprises: determining a plurality of filtered prediction candidates; as well as The filtered prediction for the current video block is determined based on a fusion of the plurality of filtered prediction candidates.
34. The method of claim 33, wherein the fusion is determined based on a blending weight of the plurality of filtered prediction candidates, the blending weight being determined based on a sum of absolute differences (SAD) cost or a sum of absolute transformed differences (SATD) cost. 35 . The method of claim 33 , wherein the fusion is determined based on a blending weight of the plurality of filtered prediction candidates, the blending weight being a predefined fixed value.
36. The method of claim 33, wherein the fusion is determined based on a blending weight of the plurality of filtered prediction candidates, the blending weight being determined based on at least one of: LDL breakdown, LU decomposition, or Gaussian elimination.
37. The method of any one of claims 33-36, wherein whether blending or filtering is applied for final prediction of the current video block is included in the bitstream.
38. A method for video processing, comprising: determining, for conversion between a current video block of a video and a bitstream of the video, a prediction for the current video block based on local illumination compensation (LIC), the LIC being based on a regression model that modulates a relationship between a current region and a reference region of the current video block for the LIC; as well as The converting is performed based on the prediction.
39. The method of claim 38, wherein the regression model comprises at least one of: Linear regression model, nonlinear regression models, or Polynomial regression model.
40. The method of claim 38 or 39, wherein determining the prediction comprises: determining a plurality of sample points in the reference area; determining a plurality of values based on the plurality of samples and the offset; determining a final sample point in the current area based on a weighted sum of the multiple values and a bias value; as well as The prediction is determined based on the final sample point.
41. The method of claim 40, wherein the plurality of values are determined by subtracting the offset from the plurality of samples.
42. The method of claim 40 or 41, wherein the weighted sum is determined based on a plurality of filter coefficients for the plurality of values and the bias value.
43. The method of claim 42, wherein the plurality of filter coefficients are determined based on at least one of: LDL breakdown or variants of LDL breakdown, LU decomposition or a variant of LU decomposition, Cholesky decomposition or a variant of Cholesky decomposition, Gaussian elimination or a variant of Gaussian elimination, or The Least Squares tool or a variation of the Least Squares tool.
44. The method according to any one of claims 40 to 43, wherein the offset is determined based on at least one reference sample located at at least one predefined position.
45. The method according to any one of claims 40 to 43, wherein the offset is determined based on a metric value of samples within the reference area.
46. The method of claim 45, wherein the metric value comprises one of the following: an average value, a mean value, a median value, a maximum value, or a minimum value.
47. The method of any one of claims 40-46, wherein the offset is zero.
48. The method of any one of claims 40-47, wherein the offset value is determined based on a bit depth of a video sequence of the video.
49. The method of claim 48, wherein the bit depth is 10 bits and the offset value is 512.
50. The method of claim 48, wherein the bit depth is 8 bits and the offset value is 256.
51. The method of any one of claims 40-47, wherein the bias value is zero.
52. The method of any one of claims 40-51, wherein the plurality of spots are adjacent to each other.
53. The method of claim 52, wherein the plurality of sample points include a first sample point at the center of the reference area, a second sample point above the first sample point, a third sample point below the first sample point, a fourth sample point to the left of the first sample point, and a fifth sample point to the right of the first sample point.
54. The method of any one of claims 40-51, wherein the plurality of samples are in a plurality of reference rows of the current video block. The method of claim 54 , wherein the number of the plurality of samples is equal to the number of the plurality of reference lines.
56. The method of any one of claims 40-51, wherein the plurality of samples are in a plurality of columns or rows within a template of the current video block.
57. The method of claim 56, wherein the number of the plurality of samples is equal to the number of the plurality of columns or rows.
58. The method of any one of claims 38-57, wherein the template size for LIC is larger than at least one of: a row of spots or a column of spots.
59. The method of any one of claims 38-58, wherein the prediction by the LIC is determined based on a plurality of rows or columns of neighboring samples.
60. The method of claim 59, wherein the plurality of rows or columns of the neighboring samples are adjacent to at least one of: the current video block or a reference block of the current video block.
61. A method for video processing, comprising: For conversion between a current video block of a video and a bitstream of the video, applying at least one intra blend for at least one color component to the current video block based on more than two reference lines; as well as The converting is performed based on the application.
62. The method of claim 61, wherein the at least one intra blend for at least one color component comprises at least one of: an intra luma blend or an intra chroma blend.
63. The method of claim 61 or 62, wherein the at least one intra-frame fusion is applied based on a regression model.
64. The method of claim 63, wherein the regression model comprises at least one of: Linear regression model, nonlinear regression models, or Polynomial regression model.
65. The method of any one of claims 61-64, wherein applying the at least one intra blend to the current video block comprises: determining a plurality of predictions of the at least one color component based on a plurality of reference rows of the current video block; as well as A fusion of the plurality of predictions is determined based on a plurality of weights.
66. The method of claim 65, wherein the plurality of weights are determined based on at least one of: LDL breakdown or variants of LDL breakdown, LU decomposition or a variant of LU decomposition, Cholesky decomposition or a variant of Cholesky decomposition, Gaussian elimination or a variant of Gaussian elimination, or The Least Squares tool or a variation of the Least Squares tool.
67. The method according to any one of claims 1-66, wherein information on whether and / or how to apply the method is included in the bitstream.
68. The method of claim 67, wherein the information is indicated at one of: sequence level, group of pictures level, picture level, slice level, or slice group level.
69. The method of claim 67 or 68, wherein the information is indicated in a sequence header, a picture header, a sequence parameter set (SPS), a video parameter set (VPS), a decoding parameter set (DPS), decoding capability information (DCI), a picture parameter set (PPS), an adaptation parameter set (APS), a slice header, or a slice group header.
70. The method according to any one of claims 67 to 69, wherein the information is indicated in an area comprising more than one sample or pixel.
71. The method of claim 70, wherein the region comprises one of a prediction block (PB), a transform block (TB), a codec block (CB), a prediction unit (PU), a transform unit (TU), a codec unit (CU), a virtual pipeline data unit (VPDU), a codec tree unit (CTU), a CTU row, a slice, a slice, or a sub-picture.
72. The method of any one of claims 67-71, further comprising: The information is determined based on the encoded and decoded information.
73. The method of claim 72, wherein the coded information comprises at least one of: block size, color format, single-tree partitioning or dual-tree partitioning, color component, slice type, or picture type.
74. The method of any one of claims 1-73, wherein the converting comprises encoding the current video block into the bitstream.
75. The method of any one of claims 1-73, wherein the converting comprises decoding the current video block from the bitstream.
76. An apparatus for processing video data, comprising a processor and non-volatile memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method of any one of claims 1-75.
77. A non-transitory computer-readable storage medium storing instructions for causing a processor to execute the method according to any one of claims 1-75.
78. A non-transitory computer-readable recording medium storing a bitstream of a video generated by a method performed by a video processing device, wherein the method comprises: determining a filtered prediction of a current video block of the video according to an offset-based regression model that modulates a relationship between a current region and a reference region of the current video block; as well as The bitstream is generated based on the filtered prediction.
79. A method for storing a bitstream of a video, comprising: determining a filtered prediction of a current video block of the video according to an offset-based regression model that modulates a relationship between a current region and a reference region of the current video block; generating the bitstream based on the filtered prediction; as well as The bitstream is stored in a non-transitory computer-readable recording medium.
80. A non-transitory computer-readable recording medium storing a bitstream of a video generated by a method performed by a video processing device, wherein the method comprises: determining a prediction for a current video block of the video based on local illumination compensation (LIC), the LIC being based on a regression model that modulates a relationship between a current region and a reference region of the current video block for the LIC; as well as The bitstream is generated based on the prediction.
81. A method for storing a bitstream of a video, comprising: determining a prediction for a current video block of the video based on local illumination compensation (LIC), the LIC being based on a regression model that modulates a relationship between a current region and a reference region of the current video block for the LIC; generating the bitstream based on the prediction; as well as The bitstream is stored in a non-transitory computer-readable recording medium.
82. A non-transitory computer-readable recording medium storing a bitstream of a video generated by a method performed by a video processing device, wherein the method comprises: applying at least one intra blend for at least one color component to a current video block of the video based on more than two reference rows; as well as The bitstream is generated based on the application.
83. A method for storing a bitstream of a video, comprising: applying at least one intra blend for at least one color component to a current video block of the video based on more than two reference rows; generating the bitstream based on the application; as well as The bitstream is stored in a non-transitory computer-readable recording medium.