Method and device for video processing and medium

By determining the conversion metric between the video block and the bit stream in the video encoding and decoding technology, and determining intra prediction based on the metric and sample point position, the problem of insufficient accuracy and efficiency of intra prediction in the prior art is solved, and a more efficient encoding and decoding process is achieved.

CN120019643APending Publication Date: 2025-05-16DOUYIN CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380072326.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-10-13
Filing Date
2023-10-11
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The existing video encoding and decoding technology has shortcomings in improving the encoding and decoding efficiency, especially in terms of the accuracy and efficiency of intra prediction.

Method used

By determining the conversion metric between the video block and the bitstream, the intra prediction is determined based on the metric and sample point position, thereby optimizing the encoding and decoding process of the video block.

Benefits of technology

It improves the accuracy and encoding and decoding efficiency of intra-frame prediction, and enhances the effectiveness of video encoding and decoding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120019643A_ABST
    Figure CN120019643A_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide a solution for video processing. A method for video processing is presented. The method comprises: for a conversion between a current video block of a video and a bitstream of the video, determining a metric for intra prediction of the current video block; determining an intra prediction of a sample point at the first position in the current video block based on the metric and the first position; and performing a conversion based on the intra prediction.
Need to check novelty before this filing date? Find Prior Art

Description

CROSS REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to U.S. Provisional Patent Application No. 63 / 415,882 filed on October 13, 2022, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] Embodiments of the present disclosure relate generally to video processing techniques, and more particularly, to intra-prediction determination. Background Art

[0003] Nowadays, digital video capabilities are being applied to all aspects of people's lives. For video encoding / decoding, various types of video compression technologies have been proposed, such as MPEG-2, MPEG-4, ITU-TH.263, ITU-TH.264 / MPEG-4 Part 10 Advanced Video Codec (AVC), ITU-TH.265 High Efficiency Video Codec (HEVC) standard, and Versatile Video Codec (VVC) standard. However, it is generally expected to further improve the encoding and decoding efficiency of video encoding and decoding technologies. Summary of the invention

[0004] Embodiments of the present disclosure provide a solution for video processing.

[0005] In a first aspect, a method for video processing is proposed. The method includes: determining a metric for intra-frame prediction for a current video block of a video and a bitstream of the video for conversion between the current video block; determining an intra-frame prediction for a sample at a first position in the current video block based on the metric and a first position; and performing conversion based on the intra-frame prediction. According to the method of the first aspect of the present disclosure, the intra-frame prediction of the current video block is determined based on the metric, and thus the intra-frame prediction can be improved. In this way, codec efficiency and codec effectiveness can be enhanced.

[0006] In a second aspect, a device for video processing is provided. The device includes a processor and a non-volatile memory having instructions thereon. The instructions, when executed by the processor, cause the processor to perform the method according to the first aspect of the present disclosure.

[0007] In a third aspect, a non-transitory computer-readable storage medium is provided, wherein the non-transitory computer-readable storage medium stores instructions for causing a processor to execute the method according to the first aspect of the present disclosure.

[0008] In a fourth aspect, another non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a bitstream of a video generated by a method performed by an apparatus for video processing. The method includes: determining a metric for intra-frame prediction of a current video block of the video; determining an intra-frame prediction of a sample at a first position in the current video block based on the metric and a first position; and generating a bitstream based on the intra-frame prediction.

[0009] In a fifth aspect, a method for storing a bitstream of a video is provided. The method includes: determining a metric for intra-frame prediction of a current video block of the video; determining an intra-frame prediction of a sample at a first position in the current video block based on the metric and a first position; generating a bitstream based on the intra-frame prediction; and storing the bitstream in a non-transitory computer-readable recording medium.

[0010] This Summary is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The above and other objects, features and advantages of the exemplary embodiments of the present disclosure will become more apparent through the following detailed description with reference to the accompanying drawings. In the exemplary embodiments of the present disclosure, the same reference numerals generally refer to the same components.

[0012] Figure 1 A block diagram illustrating an example video encoding and decoding system is shown according to some embodiments of the present disclosure;

[0013] Figure 2 A block diagram illustrating a first example video encoder is shown according to some embodiments of the present disclosure;

[0014] Figure 3 shows a block diagram illustrating an example video decoder according to some embodiments of the present disclosure;

[0015] Figure 4 An example of an encoder block diagram is shown;

[0016] Figure 5 67 intra prediction modes are shown;

[0017] Fig. 6A and Figure 6B Reference samples for wide-angle intra prediction are shown;

[0018] Figure 7 The discontinuity problem is shown when the orientation exceeds 45°;

[0019] Fig. 8A and Figure 8B The positions of the sample points used for the derivation of α and β are shown respectively;

[0020] Figure 9A-9D The definitions of the samples used by PDPC applied to diagonal and adjacent angle intra modes are shown respectively;

[0021] Fig.10 A gradient approach for non-vertical / non-horizontal patterns is shown;

[0022] Fig.11 The value of nScale is shown in relation to nTbH and the number of modes; for all cases where nScale<0, the gradient method is used;

[0023] Fig.12 Flowcharts are shown: current PDPC (left side) and proposed PDPC (right side);

[0024] Fig.13 Neighboring blocks (L, A, BL, AR, AL) used for derivation of the general MPM list are shown;

[0025] Fig.14 An example of the proposed intra reference mapping is shown;

[0026] Fig.15 An example of four reference rows of a neighboring prediction block is shown;

[0027] Fig.16A and Fig. 16B Examples of sub-divisions depending on the block size are shown respectively;

[0028] Fig.17 An example matrix-weighted intra prediction process is shown;

[0029] Fig.18 Target points, template points, and reference points of a template used for DIMD are shown;

[0030] Fig.19 shows a selected set of pixels on which a gradient analysis is performed;

[0031] Fig. 20 The convolution of a 3x3 Sobel gradient filter with a template is shown;

[0032] Fig.21 The proposed intra-block decoding process is shown;

[0033] Fig. 22 HoG computation from a template of width 3 pixels is shown;

[0034] Fig.23shows the prediction fusion by weighted averaging of two HoG modes and planes;

[0035] Fig.24 The top and left neighboring blocks used for CIIP weight derivation are shown;

[0036] Fig.25 An example of GPM partitioning grouped by the same angle is shown;

[0037] Fig.26 Unidirectional prediction MV selection for geometric partitioning mode is shown;

[0038] Fig. 27 An exemplary generation of warp weights w_0 using a geometric segmentation pattern is shown;

[0039] Fig.28 A conventional angle IPM (indicated by a black arrow) and an extended angle IPM (indicated by a dashed line) are shown;

[0040] Fig.29 The positions of the training samples are shown;

[0041] Fig.30 A flowchart showing a method for video processing according to some embodiments of the present disclosure; and

[0042] Fig.31 A block diagram of a computing device is shown in which various embodiments of the present disclosure may be implemented.

[0043] Same or similar reference numbers generally refer to same or similar elements throughout the drawings. DETAILED DESCRIPTION

[0044] The principle of the present disclosure will now be described with reference to some embodiments. It should be understood that these embodiments are described only for the purpose of illustrating and helping those skilled in the art to understand and implement the present disclosure, without implying any limitation on the scope of the present disclosure. In addition to the methods described below, the disclosure described herein can also be implemented in various ways.

[0045] In the following description and claims, unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs.

[0046] References in this disclosure to "one embodiment," "an embodiment," "an example embodiment," and the like indicate that the described embodiment may include a particular feature, structure, or characteristic, but not every embodiment must include the particular feature, structure, or characteristic. Furthermore, these phrases do not necessarily refer to the same embodiment. Furthermore, when a particular feature, structure, or characteristic is described in conjunction with an example embodiment, it is claimed that such feature, structure, or characteristic, whether or not explicitly described, is within the knowledge of those skilled in the art to affect correlation with other embodiments.

[0047] It should be understood that, although the terms "first" and "second" etc. may be used herein to describe various elements, these elements should not be limited to these terms. These terms are only used to distinguish one element from another element. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element without departing from the scope of the exemplary embodiments. As used herein, the term "and / or" includes any and all combinations of one or more of the listed terms.

[0048] The terms used herein are only used for the purpose of describing specific embodiments and are not intended to limit the exemplary embodiments. As used herein, the singular forms "a", "an" and "the" are also intended to include plural forms unless the context clearly indicates otherwise. It should also be understood that the terms "include", "comprises", "has", "has", "includes" and / or "comprising" are used herein to indicate the presence of the features, elements and / or components, etc., but do not exclude the presence or addition of one or more other features, elements, components and / or combinations thereof. Example Environment

[0049] Figure 1 1 is a block diagram illustrating an example video codec system 100 that may utilize the techniques of the present disclosure. As shown, the video codec system 100 may include a source device 110 and a destination device 120. The source device 110 may also be referred to as a video encoding device, and the destination device 120 may also be referred to as a video decoding device. In operation, the source device 110 may be configured to generate encoded video data, and the destination device 120 may be configured to decode the encoded video data generated by the source device 110. The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.

[0050] The video source 112 may include a source such as a video acquisition device. Examples of a video acquisition device include, but are not limited to, an interface for receiving video data from a video content provider, a computer graphics system for generating video data, and / or a combination thereof.

[0051] The video data may include one or more pictures. The video encoder 114 encodes the video data from the video source 112 to generate a bitstream. The bitstream may include a bit sequence that forms a coded representation of the video data. The bitstream may include coded pictures and associated data. The coded pictures are coded representations of pictures. The associated data may include sequence parameter sets, picture parameter sets, and other grammatical structures. The I / O interface 116 may include a modulator / demodulator and / or a transmitter. The coded video data may be directly transmitted to the destination device 120 via the network 130A via the I / O interface 116. The coded video data may also be stored on a storage medium / server 130B for access by the destination device 120.

[0052] The destination device 120 may include an I / O interface 126, a video decoder 124, and a display device 122. The I / O interface 126 may include a receiver and / or a modem. The I / O interface 126 may obtain encoded video data from the source device 110 or the storage medium / server 130B. The video decoder 124 may decode the encoded video data. The display device 122 may display the decoded video data to the user. The display device 122 may be integrated with the destination device 120, or may be outside the destination device 120, which is configured to be connected to an external display device interface.

[0053] The video encoder 114 and the video decoder 124 may operate according to a video compression standard, such as the High Efficiency Video Codec (HEVC) standard, the Versatile Video Codec (VVC) standard, and other existing and / or future standards.

[0054] Figure 2 is a block diagram showing an example of a video encoder 200 according to some embodiments of the present disclosure, which may be Figure 1 An example of a video encoder 114 in the system 100 is shown.

[0055] Video encoder 200 may be configured to implement any or all of the techniques of this disclosure. Figure 2 In the example of , video encoder 200 includes multiple functional components. The techniques described in this disclosure can be shared between the various components of video encoder 200. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.

[0056] In some embodiments, the video encoder 200 may include a segmentation unit 201, a prediction unit 202, a residual generation unit 207, a transformation unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transformation unit 211, a reconstruction unit 212, a cache 213 and an entropy coding unit 214, and the prediction unit 202 may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205 and an intra-frame prediction unit 206.

[0057] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an intra-block copy (IBC) unit. The IBC unit may perform prediction in an IBC mode in which at least one reference picture is a picture in which the current video block is located.

[0058] Furthermore, although some components (such as the motion estimation unit 204 and the motion compensation unit 205) may be integrated, for the purpose of explanation, these components are described in detail below. Figure 2 are shown separately in the example.

[0059] The partitioning unit 201 may partition a picture into one or more video blocks. The video encoder 200 and the video decoder 300 may support various video block sizes.

[0060] The mode selection unit 203 may select one of a plurality of encoding modes (intra-frame encoding or inter-frame encoding), for example, based on the error result, and provide the generated intra-frame encoded block or inter-frame encoded block to the residual generation unit 207 to generate residual block data, and to the reconstruction unit 212 to reconstruct the encoded block for use as a reference picture. In some examples, the mode selection unit 203 may select a combined intra-frame and inter-frame prediction (CIIP) mode in which the prediction is based on an inter-frame prediction signal and an intra-frame prediction signal. In the case of inter-frame prediction, the mode selection unit 203 may also select a resolution for the motion vector (e.g., sub-pixel precision or integer pixel precision) for the block.

[0061] To perform inter-frame prediction on the current video block, the motion estimation unit 204 may generate motion information for the current video block by comparing the current video block with one or more reference frames from the cache 213. The motion compensation unit 205 may determine a predicted video block for the current video block based on the motion information and decoded samples of pictures from the cache 213 other than the picture associated with the current video block.

[0062] The motion estimation unit 204 and the motion compensation unit 205 may perform different operations on the current video block, for example, depending on whether the current video block is in an I slice, a P slice, or a B slice. As used herein, an "I slice" may refer to a portion of a picture consisting of macroblocks, all of which are based on macroblocks within the same picture. Furthermore, as used herein, in some aspects, a "P slice" and a "B slice" may refer to a portion of a picture consisting of macroblocks that are independent of macroblocks in the same picture.

[0063] In some examples, the motion estimation unit 204 may perform unidirectional prediction on the current video block, and the motion estimation unit 204 may search the reference pictures of list 0 or list 1 to find the reference video block for the current video block. The motion estimation unit 204 may then generate a reference index and a motion vector, the reference index indicating the reference picture in list 0 or list 1 containing the reference video block, and the motion vector indicating the spatial displacement between the current video block and the reference video block. The motion estimation unit 204 may output the reference index, the prediction direction indicator, and the motion vector as the motion information of the current video block. The motion compensation unit 205 may generate a predicted video block for the current video block based on the reference video block indicated by the motion information of the current video block.

[0064] Alternatively, in other examples, the motion estimation unit 204 may perform bidirectional prediction on the current video block. The motion estimation unit 204 may search the reference pictures in list 0 to find a reference video block for the current video block, and may also search the reference pictures in list 1 to find another reference video block for the current video block. The motion estimation unit 204 may then generate a plurality of reference indexes and a plurality of motion vectors, the plurality of reference indexes indicating a plurality of reference pictures in list 0 and list 1 containing a plurality of reference video blocks, and the plurality of motion vectors indicating a plurality of spatial displacements between the plurality of reference video blocks and the current video block. The motion estimation unit 204 may output the plurality of reference indexes and the plurality of motion vectors of the current video block as the motion information of the current video block. The motion compensation unit 205 may generate a predicted video block for the current video block based on the plurality of reference video blocks indicated by the motion information of the current video block.

[0065] In some examples, motion estimation unit 204 may output a complete set of motion information for use in a decoding process by a decoder. Alternatively, in some embodiments, motion estimation unit 204 may signal motion information of a current video block with reference to motion information of another video block. For example, motion estimation unit 204 may determine that motion information of a current video block is sufficiently similar to motion information of a neighboring video block.

[0066] In one example, motion estimation unit 204 may indicate a value in a syntax structure associated with the current video block that indicates to video decoder 300 that the current video block has the same motion information as another video block.

[0067] In another example, the motion estimation unit 204 may identify another video block and a motion vector difference (MVD) in a syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 300 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.

[0068] As discussed above, the video encoder 200 may signal motion vectors in a predictive manner.Two examples of prediction signaling techniques that may be implemented by the video encoder 200 include Advanced Motion Vector Prediction (AMVP) and Merge mode signaling.

[0069] The intra prediction unit 206 may perform intra prediction on the current video block. When the intra prediction unit 206 performs intra prediction on the current video block, the intra prediction unit 206 may generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block may include a prediction video block and various syntax elements.

[0070] The residual generation unit 207 can generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) the predicted video block(s) of the current video block from the current video block. The residual data of the current video block may include residual video blocks corresponding to different sample components of samples in the current video block.

[0071] In other examples, such as in skip mode, there may be no residual data for the current video block, and the residual generation unit 207 may not perform a subtraction operation.

[0072] Transform processing unit 208 may generate one or more transform coefficient video blocks for a current video block by applying one or more transforms to the residual video block associated with the current video block.

[0073] After transform processing unit 208 generates a transform coefficient video block associated with the current video block, quantization unit 209 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values ​​associated with the current video block.

[0074] The inverse quantization unit 210 and the inverse transform unit 211 may apply inverse quantization and inverse transform to the transform coefficient video block, respectively, to reconstruct a residual video block from the transform coefficient video block. The reconstruction unit 212 may add the reconstructed residual video block to corresponding samples from one or more prediction video blocks generated by the prediction unit 202 to generate a reconstructed video block associated with the current video block for storage in the buffer 213.

[0075] After reconstruction unit 212 reconstructs the video block, a loop filtering operation may be performed to reduce video blocking artifacts in the video block.

[0076] The entropy encoding unit 214 may receive data from other functional components of the video encoder 200. When the entropy encoding unit 214 receives the data, the entropy encoding unit 214 may perform one or more entropy encoding operations to generate entropy-encoded data and output a bitstream including the entropy-encoded data.

[0077] Figure 3 is a block diagram showing an example of a video decoder 300 according to some embodiments of the present disclosure, which may be Figure 1 An example of a video decoder 124 in the system 100 is shown.

[0078] Video decoder 300 may be configured to perform any or all of the techniques of this disclosure. Figure 3 In the example of , video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared between the various components of video decoder 300. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.

[0079] exist Figure 3 In the example of , the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, and a reconstruction unit 306 and a buffer 307. In some examples, the video decoder 300 can perform a decoding process that is generally opposite to the encoding process described with respect to the video encoder 200.

[0080] The entropy decoding unit 301 can retrieve the encoded bitstream. The encoded bitstream may include entropy encoded video data (e.g., encoded blocks of video data). The entropy decoding unit 301 can decode the entropy encoded video data, and the motion compensation unit 302 can determine motion information from the entropy decoded video data, the motion information including motion vectors, motion vector precision, reference picture list indexes, and other motion information. The motion compensation unit 302 can determine such information, for example, by performing AMVP and Merge mode. AMVP is used, including deriving several most likely candidates based on data and reference pictures from adjacent PBs. The motion information typically includes horizontal motion vector displacement values ​​and vertical motion vector displacement values, one or two reference picture indexes, and in the case of prediction areas in B strips, also includes an identification of which reference picture list is associated with each index. As used herein, in some aspects, "Merge mode" may refer to deriving motion information from spatial neighboring blocks or temporal neighboring blocks.

[0081] The motion compensation unit 302 may generate a motion compensated block, possibly performing interpolation based on an interpolation filter.Identifiers for the interpolation filters used with sub-pixel precision may be included in the syntax elements.

[0082] The motion compensation unit 302 may calculate interpolated values ​​for sub-integer pixels of a reference block using interpolation filters used by the video encoder 200 during encoding of the video block. The motion compensation unit 302 may determine the interpolation filters used by the video encoder 200 based on received syntax information, and the motion compensation unit 302 may use the interpolation filters to generate a prediction block.

[0083] The motion compensation unit 302 may use at least part of the syntax information to determine the size of blocks used to encode (multiple) frames and / or (multiple) slices of the encoded video sequence, partition information describing how each macroblock of a picture of the encoded video sequence is partitioned, a mode indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-coded block, and other information for decoding the encoded video sequence. As used herein, in some aspects, a "slice" may refer to a data structure that can be decoded independently of other slices of the same picture in terms of entropy coding and decoding, signal prediction, and residual signal reconstruction. A slice may be an entire picture, or it may be a region of a picture.

[0084] The intra prediction unit 303 may use, for example, an intra prediction mode received in the bitstream to form a prediction block from spatially neighboring blocks. The inverse quantization unit 304 inverse quantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 301. The inverse transform unit 305 applies an inverse transform.

[0085] The reconstruction unit 306 may obtain the decoded block, for example, by adding the residual block to the corresponding prediction block generated by the motion compensation unit 302 or the intra prediction unit 303. If necessary, a deblocking filter may also be applied to filter the decoded block to remove blocking artifacts. The decoded video block is then stored in a buffer 307, which provides reference blocks for subsequent motion compensation / intra prediction, and the buffer 307 also generates the decoded video for presentation on a display device.

[0086] Some exemplary embodiments of the present disclosure will be described in detail below. It should be noted that the section titles used in this document are for ease of understanding, and the embodiments disclosed in the section are not limited to that section. In addition, although some embodiments are described with reference to multifunctional video codecs or other specific video codecs, the disclosed technology is also applicable to other video coding and decoding technologies. In addition, although some embodiments describe the video encoding steps in detail, it should be understood that the corresponding decoding steps of de-encoding will be implemented by the decoder. In addition, the term video processing includes video encoding or compression, video decoding or decompression, and video transcoding, in which video pixels are represented from one compression format to another compression format or at different compression bit rates. 1. Brief Overview The present disclosure relates to video coding technology. Specifically, it relates to intra-frame prediction in image / video coding. It can be applied to existing video coding standards such as HEVC or versatile video coding (VVC). It can also be applied to future video coding standards or video codecs. 2. Introduction Video codec standards have evolved primarily through the development of the well-known ITU-T and ISO / IEC standards. ITU-T produced H.261 and H.263, ISO / IEC produced MPEG-1 and MPEG-4 Visual, and the two organizations jointly produced the H.262 / MPEG-2 Video and H.264 / MPEG-4 Advanced Video Codec (AVC) and H.265 / HEVC standards. Since H.262, video codec standards have been based on a hybrid video codec structure in which temporal prediction plus transform codecs are utilized. In order to explore future video codec technologies beyond HEVC, the Joint Video Exploration Team (JVET) was jointly established by VCEG and MPEG in 2015. Since then, many new methods have been adopted by JVET and put into reference software called the Joint Exploration Model (JEM). In April 2018, the Joint Video Experts Group (JVET) between VCEG (Q6 / 16) and ISO / IEC JTC1 SC29 / WG11 (MPEG) was created to work on the VVC standard with the goal of a 50% bitrate reduction compared to HEVC. The latest version of the VVC draft, Versatile Video Codec (Draft 10), can be found at: http: / / phenix.it-sudparis.eu / jvet / doc_end_user / documents / 20_Teleconference / wg11 / JVET-T2001-v1.zip. The latest reference software for VVC is called VTM and can be found at the following URL: https: / / vcgit.hhi.fraunhofer.de / jvet / VVCSoftware_VTM / - / tags / VTM-11.0. 2.1. Encoding and decoding process of typical video codecs Figure 4 An example 400 of an encoder block diagram for VVC is shown, which includes three in-loop filter blocks: deblocking filter (DF), sample adaptive offset (SAO), and ALF. Unlike DF, which uses a predefined filter, SAO and ALF use the original samples of the current picture to reduce the mean square error between the original samples and the reconstructed samples by adding an offset and applying a finite impulse response (FIR) filter, respectively, where the codec's side information transmits the offset and filter coefficients through signaling. ALF is located at the last processing stage of each picture and can be seen as a tool that tries to obtain and repair artifacts produced by previous stages. 2.2. Intra-mode codec with 67 intra-prediction modes Figure 5 A schematic diagram 500 showing 67 intra prediction modes is shown. In order to capture arbitrary edge directions present in natural video, the number of directional intra modes is extended from 33 used in HEVC to 65, as shown in FIG. Figure 5 As shown, planar mode and DC mode remain unchanged. These dense directional intra prediction modes are applicable to all block sizes and to both luma intra prediction and chroma intra prediction. In HEVC, each intra-coded block has a square shape, and the length of each of its sides is a power of 2. Therefore, no division operation is required when generating intra prediction values ​​using DC mode. In VVC, blocks can have a rectangular shape, which generally requires the use of a division operation for each block. To avoid division operations for DC prediction, only the longer sides are used to calculate the average value for non-square blocks. 2.2.1. Wide-angle intra prediction Although 67 modes are defined in VVC, the exact prediction direction for a given intra prediction mode index further depends on the block shape. Conventional angular intra prediction directions are defined as from 45 degrees to -135 degrees in a clockwise direction. In VVC, several conventional angular intra prediction modes are adaptively replaced with wide-angle intra prediction modes for non-square blocks. The replaced mode is transmitted via a signal using the original mode index, which is remapped to the index of the wide-angle mode after parsing. The total number of intra prediction modes remains unchanged, i.e. 67, and the intra mode encoding and decoding method remains unchanged. To support these prediction directions, a top reference of length 2W+1 and a left reference of length 2H+1 are defined, as Fig. 6A and Figure 6B shown. Fig. 6A and Figure 6B Schematic diagram 600 and schematic diagram 650 respectively show reference samples for wide-angle intra prediction. The number of replaced modes in the wide-angle direction mode depends on the aspect ratio of the block. The replaced intra-frame prediction modes are shown in Table 1. Table 1 - Intra prediction modes replaced by wide mode Figure 7 Schematic diagram 700 showing the problem of discontinuities in directions exceeding 45°. Figure 7 As shown, in the case of wide-angle intra prediction, two vertically adjacent prediction samples can use two non-adjacent reference samples. Therefore, a low-pass reference sample filter and side smoothing are applied to wide-angle prediction to reduce the added gap Δp αnegative impact. If the wide angle mode represents a non-fractional offset. There are 8 modes in the wide angle mode that meet this condition, which are [-14, -12, -10, -6, 72, 76, 78, 80]. When a block is predicted by these modes, the samples in the reference cache are directly copied without applying any interpolation. With this modification, the number of samples that need to be smoothed is reduced. In addition, it aligns the design of non-fractional modes in the conventional prediction mode and the wide angle mode. In VVC, in addition to supporting the 4:2:0 chroma format, the 4:2:2 chroma format and the 4:4:4 chroma format are also supported. The chroma derivation mode (DM) derivation table for the 4:2:2 chroma format was originally ported from HEVC, expanding the number of entries from 35 to 67 to align with the expansion of the intra prediction mode. Since the HEVC specification does not support prediction angles below -135 degrees and above 45 degrees, the luma intra prediction mode in the range of 2 to 5 is mapped to 2. Therefore, the chroma DM derivation table for the 4:2:2 chroma format is updated by replacing some values ​​of the entries of the mapping table to more accurately convert the prediction angles of the chroma blocks. 2.3. Inter-frame prediction For each inter-predicted CU, the motion parameters include motion vectors, reference picture indices, and reference picture list usage indices, as well as additional information required by the new coding features of VVC that will be used for sample generation for inter-prediction. Motion parameters can be transmitted through signals in an explicit or implicit manner. When a CU is encoded and decoded in skip mode, the CU is associated with one PU and has no significant residual coefficients, no coded motion vector delta, or reference picture index. Merge mode is specified, in which the motion parameters for the current CU are obtained from neighboring CUs (including spatial and temporal candidates), and additional mechanisms introduced in VVC are adopted. Merge mode can be applied to any inter-predicted CU, not just skip mode. An alternative to Merge mode is explicit transmission of motion parameters, in which motion vectors, corresponding reference picture indices for each reference picture list, reference picture list usage flags, and other required information are explicitly transmitted through signals per CU. 2.4. Intra-block copy (IBC) Intra-block copying (IBC) is a tool adopted in the HEVC extension on SCC. It is well known that it significantly improves the encoding and decoding efficiency of screen content materials. Since the IBC mode is implemented as a block-level codec mode, block matching (BM) is performed at the encoder to find the best block vector (or motion vector) for each CU. Here, the block vector is used to indicate the displacement from the current block to the reference block, which has been reconstructed within the current picture. The luminance block vector of the CU encoded and decoded by IBC is integer precision. The chrominance block vector is also rounded to integer precision. When combined with AMVR, the IBC mode can switch between 1-pixel and 4-pixel motion vector precision. The CU encoded and decoded by IBC is regarded as a third prediction mode in addition to the intra or inter prediction mode. The IBC mode is applicable to CUs with a width and height that are less than or equal to 64 luminance samples. On the encoder side, hash-based motion estimation is performed for IBC. The encoder performs RD checks on blocks with a width or height no greater than 16 luma samples. For non-Merge mode, a block vector search is first performed using a hash-based search. If the hash search does not return a valid candidate, a local search based on block matching is performed. In hash-based search, the hash key matching (32-bit CRC) between the current block and the reference blocks is extended to all allowed block sizes. The hash key calculation for each position in the current picture is based on 4×4 sub-blocks. For current blocks of larger size, the hash key is determined to match the hash key of the reference block when all hash keys in all 4×4 sub-blocks match the hash keys in the corresponding reference positions. If the hash keys of multiple reference blocks are found to match the hash key of the current block, the block vector cost of each matching reference is calculated, and the one with the smallest cost is selected. In the block matching search, the search range is set to cover the previous CTU and the current CTU. At the CU level, the IBC mode is signaled using a flag, which can be signaled in IBC AMVP mode or IBC Skip / Merge mode as shown below: – IBC Skip / Merge mode: The Merge candidate index is used to indicate which block vector from the list of neighboring candidate IBC codec blocks is used to predict the current block. The Merge list includes spatial candidates, HMVP candidates, and pairwise candidates. – IBC AMVP mode: Block vector differences are encoded in the same way as motion vector differences. The block vector prediction method uses two candidates as predictors, one from the left neighbor and one from the top neighbor (in case of IBC codec). When either neighbor is not available, the default block vector will be used as the predictor. A flag is signaled to indicate the block vector predictor index. 2.5. Cross-component linear model prediction To reduce cross-component redundancy, a cross-component linear model (CCLM) prediction mode is used in VVC, for which chroma samples are predicted based on the reconstructed luma samples of the same CU using the following linear model: pred C (i,j)=α·rec L ′(i,j)+β (2-1) where pred C (i, j) represents the predicted chroma sample in the CU, and rec L ′(i,j) represents the downsampled reconstructed luma sample of the same CU. The CCLM parameters (α and β) are derived using up to four neighboring chroma samples and their corresponding downsampled luma samples. Assuming the current chroma block size is W×H, W' and H' are set to - When LM mode is applied, W'=W, H'=H; - When LM-T mode is applied, W'=W+H; - When LM-L mode is applied, H'=H+W. The upper neighboring position is denoted as S[0, -1] ... S[W'-1, -1], and the left neighboring position is denoted as S[-1, 0] ... S[-1, H'-1]. Then, four sample points are selected as: - When LM mode is applied and both the upper neighboring sample points and the left neighboring sample points are available, S[W' / 4,-1], S[3 *W' / 4,-1],S[-1,H' / 4],S[-1,3*H' / 4]; - When LM-T mode is applied or only upper neighbor samples are available, S[W' / 8,-1],S[3*W' / 8, -1],S[5*W' / 8,-1],S[7*W' / 8,-1]; - When LM-L mode is applied or only left neighbor samples are available, S[-1,H' / 8], S[-1,3*H' / 8],S[-1,5*H' / 8],S[-1,7*H' / 8]. The four neighboring brightness samples at the selected position are downsampled and compared four times to find the larger of the two values: x 0 A and x 1 A , and two smaller values: x 0 B and x 1 B Their corresponding chrominance sample values ​​are represented by y 0 A.y 1 A.y 0 B and y 1 B. Then x A 、x B ,y A and B is derived as: X a =(x 0 A +x 1 A +1)>>1;X b =(x 0 B +x 1 B +1)>>1;Y a =(y 0 A +y 1 A +1)>>1;Y b =(y 0 B +y 1 B +1)>>1 (2-2) Finally, the parameters α and β of the linear model are obtained according to the following formula. β=Y b -α·X b (2-4) Fig. 8A and Figure 8B Schematic diagrams 800 and 850 show the locations of sample points used to derive α and β, respectively. Fig. 8A and Figure 8B An example of the positions of the left and upper samples and the samples of the current block involved in CCLM mode is shown. The division operation for calculating the parameter α is implemented using a lookup table. In order to reduce the memory required to store the table, the diff value (the difference between the maximum and minimum values) and the parameter α are represented by exponential notation. For example, diff is approximated by a 4-bit significant part and an exponent. Therefore, the table for 1 / diff is reduced to 16 elements for 16 values ​​of the significant part, as shown below: DivTable[]={0, 7, 6, 5, 5, 4, 4, 3, 3, 2, 2, 1, 1, 1, 1, 0} (2-5). This will help reduce the complexity of the calculations and the memory size required to store the required tables. In addition to the upper template and the left template being used together to calculate the linear model coefficients, they can also be used alternately in the other two LM modes, called LM_T mode and LM_L mode. In LM_T mode, only the upper template is used to calculate the linear model coefficients. To obtain more samples, the upper template is expanded to (W+H) samples. In LM_L mode, only the left template is used to calculate the linear model coefficients. To obtain more samples, the left template is expanded to (H+W) samples. In LM mode, the left template and the upper template are used to calculate the linear model coefficients. To match the chroma sample positions for 4:2:0 video sequences, two types of downsampling filters are applied to the luma samples to achieve a 2 to 1 downsampling ratio in both the horizontal and vertical directions. The choice of downsampling filter is specified by the SPS level flag. The two downsampling filters are as follows, corresponding to "type-0" and "type-2" content, respectively. Note that when the above reference line is at a CTU boundary, only one luma line (common line buffer in intra prediction) is used to generate the downsampled luma samples. This parameter calculation is performed as part of the decoding process, not just as an encoder search operation. Therefore, no syntax is used to pass the α and β values ​​to the decoder. For chroma intra mode coding and decoding, a total of 8 intra modes are allowed for chroma intra mode coding and decoding. These modes include five conventional intra modes and three cross-component linear model modes (LM, LM_T and LM_L). The chroma mode signaling and derivation process are shown in Table 2. The chroma mode coding and decoding depends directly on the intra prediction mode of the corresponding luminance block. Since the independent block partitioning structure for luminance component and chrominance component in the I strip is enabled, one chroma block can correspond to multiple luminance blocks. Therefore, for the chroma DM mode, the intra prediction mode of the corresponding luminance block covering the center position of the current chroma block is directly inherited. Table 2 Derivation of chroma prediction mode from luma mode when CCLM is enabled Regardless of the value of sps_cclm_enabled_flag, a single binarization table is used, as shown in Table 3. Table 3 Unified binarization table for chroma prediction mode Value of intra_chroma_pred_mode Binary String 4 00 0 0100 1 0101 2 0110 3 0111 5 10 6 110 7 111 In Table 3, the first binary bit indicates whether it is normal mode (0) or LM mode (1). If it is LM mode, the next binary bit indicates whether it is LM_CHROMA (0). If it is not LM_CHROMA, the next binary bit indicates whether it is LM_L (0) or LM_T (1). For this case, when sps_cclm_enabled_flag is 0, the first binary bit of the binarization table for the corresponding intra_chroma_pred_mode can be discarded before entropy coding. Or, in other words, the first binary bit is inferred to be 0 and therefore not coded. This single binarization table is used for both cases where sps_cclm_enabled_flag is equal to 0 and equal to 1. The first two binary bits in Table 3 are context-coded using their own context models, and the remaining binary bits are bypass-coded. In addition, to reduce luma-chroma delay in dual trees, when the 64×64 luma codec tree node is split using Not Split (and ISP is not used for 64×64 CU) or QT, the chroma CU in the 32×32 / 32×16 chroma codec tree node is allowed to use CCLM in the following way: – If a 32×32 chroma node is not partitioned or is split by a QT partition, all chroma CUs in the 32×32 node can use CCLM. If a 32×32 chroma node is split using horizontal BT and the 32×16 child node is not split or is split using vertical BT, all chroma CUs in the 32×16 chroma node can use CCLM. Under all other luma and chroma codec tree partitioning conditions, CCLM is not allowed for chroma CUs. 2.6. Position-dependent intra prediction combination In VVC, the results of intra prediction for DC, planar, and several angular modes are further modified by the Position Dependent Intra Prediction Combination (PDPC) method. PDPC is an intra prediction method that calls for a combination of boundary reference samples and HEVC-style intra prediction with filtered boundary reference samples. PDPC is applied to the following intra modes without signaling: planar, DC, intra angles less than or equal to horizontal, and intra angles greater than or equal to vertical and less than or equal to 80. PDPC is not applied if the current block is in BDPCM mode or the MRL index is greater than 0. The prediction sample pred(x',y') is predicted using a linear combination of the intra prediction mode (DC, planar, angular) and the reference sample according to the following equation 2-8: pred(x′,y′)=Clip(0, (1<<BitDepth)-1, (wL×R -1,y′ +wT×R x′,-1 +(64-wL-wT)×pred(x′,y′)+32)>>6) (2-8) Where R x,-1 ,R -1,y Respectively represent the reference sample points located at the top and left boundaries of the current sample point (x, y). If PDPC is applied for DC, planar, horizontal and vertical intra modes, no additional boundary filters are required, as is required in the case of HEVC DC mode boundary filters or horizontal / vertical mode edge filters. The PDPC process is the same for DC mode and planar mode. For angular mode, if the current angular mode is HOR_IDX or VER_IDX, the left or top reference samples, respectively, are not used. The PDPC weights and scaling factors depend on the prediction mode and block size. PDPC is applied to blocks with width and height both greater than or equal to 4. Figure 9A-9D Schematic diagrams 900, 920, 940 and 960 show the definition of the samples used for PDPC applied to diagonal and adjacent angle intra modes, respectively. 9A to 9D The reference samples (R) for PDPC applied in various prediction modes are shown. x,-1and R -1,y ) definition. Fig. 9A The diagonal upper right mode is shown. Fig. 9B The diagonal lower left mode is shown. Fig. 9C The adjacent diagonal upper right mode is shown. Fig.9D The adjacent diagonal lower left mode is shown. The prediction sample pred(x',y') is located at (x',y') in the prediction block. As an example, for the diagonal mode, the reference sample R x,-1 The coordinate x is given by: x = x' + y' + 1, and the reference point R -1,y The coordinate y of is similarly given by: y = x' + y' + 1. For other angle modes, the reference point R x,-1 and R -1,y Can be located at fractional sample positions. In this case, the sample value at the nearest integer sample position is used. 2.7. Gradient PDPC Fig.10 A schematic diagram 1000 of a gradient method for non-vertical / non-horizontal patterns is shown. Fig.10 As shown in Figure 1, the gradient-based approach is extended to non-vertical / non-horizontal modes. Here, the gradient is calculated as r(-1,y)–r(-1+d,-1), where d is the horizontal displacement depending on the angular direction. A few points to note here: The gradient term r(-1,y)–r(-1+d,-1) needs to be calculated once for each row since it does not depend on the x position. The calculation of d is already part of the original intra prediction process which can be reused, so a separate calculation of d is not needed. Therefore, d is 1 / 32 pixel accurate. When d is in fractional position, two-tap (linear) filtering has been used, i.e., if dPos is the displacement with 1 / 32 pixel precision, dInt is the (rounded down) integer part (dPos>>5), and dFract is the fractional part with 1 / 32 pixel precision (dPos&31), then r(-1+d) is calculated as: r(-1+d)=(32–dFrac)*r(-1+dInt)+dFrac*r(-1+dInt+1). As described in a, this 2-tap filtering is performed once per row (if necessary). Finally, the prediction signal is calculated: p(x,y)=Clip(((64–wL(x))*p(x,y)+wL(x)*(r(-1,y)-r(-1+d,-1))+32)>>6). Where wL(x)=32>>((x<<1)>>nScale2), and nScale2=(log2(nTbH)+log2(nTbW)–2)>>2, which are the same as the vertical / horizontal mode. In short, the same process is applied compared to the vertical / horizontal mode (in fact, d=0 indicates the vertical / horizontal mode). Secondly, when (nScale<0) or PDPC cannot be applied due to unavailability of auxiliary reference samples, the gradient-based method is activated for non-vertical / non-horizontal modes. Fig.11 1100 shows the value of nScale as a function of nTbH and the number of modes; for all cases where nScale < 0 the gradient method is used. The value of nScale as a function of TB size and angular mode is Fig.11 is shown in , to better visualize the gradient method being used. Additionally, in Fig.12 , a flowchart 1200 is shown for the current PDPC (left side) and the proposed PDPC (right side). 2.8. Auxiliary MPM Auxiliary MPM lists are introduced. The existing primary MPM (PMPM) list includes 6 entries, while the auxiliary MPM (SMPM) list includes 16 entries. First, a general MPM list with 22 entries is constructed, then the first 6 entries in the general MPM list are included in the PMPM list, and the remaining entries form the SMPM list. The first entry in the general MPM list is the plane mode. Fig.13 A schematic diagram 1300 of neighboring blocks (L, A, BL, AR, AL) used for derivation of a general MPM list is shown. Fig.13 As shown, the remaining entries consist of the intra modes of the left (L), above (A), lower left (BL), upper right (AR), and upper left (AL) neighboring blocks, directional modes with additional offsets from the first two available directional modes of the neighboring blocks, and a default mode. If the CU block is vertically oriented, the order of neighboring blocks is A, L, BL, AR, AL; otherwise, it is L, A, BL, AR, AL. First the PMPM flag is parsed, if it is equal to 1, the PMPM index is parsed to determine which entry of the PMPM list is selected, otherwise the SPMPM flag is parsed to determine whether to parse the SMPM index or the remaining modes. 2.9.6 Tap Intra-frame Interpolation Filter In order to improve the prediction accuracy, it is proposed to replace the 4-tap cubic interpolation filter with a 6-tap interpolation filter. The filter coefficients are derived based on the same polynomial regression model, but the polynomial order is 6. The filter coefficients are listed below, {0,0,256,0,0,0}, / / 0 / 32 position {0,-4,253,9,-2,0}, / / 1 / 32 position {1,-7,249,17,-4,0}, / / 2 / 32 position {1,-10,245,25,-6,1}, / / 3 / 32 position {1,-13,241,34,-8,1}, / / 4 / 32 position {2,-16,235,44,-10,1}, / / 5 / 32 position {2,-18,229,53,-12,2}, / / 6 / 32 position {2,-20,223,63,-14,2}, / / 7 / 32 position {2,-22,217,72,-15,2}, / / 8 / 32 position {3,-23,209,82,-17,2}, / / 9 / 32 position {3,-24,202,92,-19,2}, / / 10 / 32 position {3,-25,194,101,-20,3}, / / 11 / 32 position {3,-25,185,111,-21,3}, / / 12 / 32 position {3,-26,178,121,-23,3}, / / 13 / 32 position {3,-25,168,131,-24,3}, / / 14 / 32 position {3,-25,159,141,-25,3}, / / 15 / 32 position {3,-25,150,150,-25,3}, / / half pixel position The reference samples used for interpolation come from the reconstructed samples or are padded as in HEVC, so a conditional check on the availability of reference samples is not required. It is recommended to use a 4-tap cubic interpolation filter instead of using the nearest integer operation to derive the extended intra-frame reference samples. Fig.14 An exemplary schematic diagram 1400 regarding the proposed intra-frame reference mapping is shown. Fig.14As shown in the example in , in order to derive the value of the reference sample point P, a four-tap interpolation filter is used, while in JEM-3.0 or HM, P is directly set to X1. 2.10. Multiple Reference Line (MRL) Intra Prediction Multiple reference line (MRL) intra prediction uses more reference lines for intra prediction. Fig.15 An exemplary diagram 1500 is shown of four reference rows of neighboring prediction blocks. Fig.15 In , an example of 4 reference lines is depicted, where the samples of segments A and F are not extracted from the reconstructed neighboring samples, but are filled with the nearest samples from segments B and E, respectively. HEVC intra picture prediction uses the nearest reference line (i.e., reference line 0). In MRL, 2 additional lines (reference line 1 and reference line 3) are used. The index of the selected reference row (mrl_idx) is signaled and used to generate intra prediction values. For reference row indices greater than 0, only additional reference row modes are included in the MPM list, and only the MPM index is signaled without the remaining modes. The reference row index is signaled before the intra prediction mode, and in case a non-zero reference row index is signaled, the planar mode is excluded from the intra prediction mode. The first row of MRL for blocks within a CTU is disabled to prevent the use of extended reference samples outside the current CTU row. In addition, PDPC is disabled when additional rows are used. For MRL mode, the derivation of the DC value in the DC intra prediction mode for a non-zero reference row index is aligned with the derivation for reference row index 0. MRL requires storage of 3 neighboring luma reference rows with a CTU to generate predictions. The Cross Component Linear Model (CCLM) tool also requires 3 neighboring luma reference rows for its downsampling filter. The definition of MRL using the same 3 rows is aligned with CCLM to reduce storage requirements for the decoder. 2.11. Intra-frame sub-segmentation (ISP) Intra sub-partitioning (ISP) divides the luma intra prediction block into 2 or 4 sub-partitions vertically or horizontally according to the block size. For example, the minimum block size of ISP is 4×8 (or 8×4). If the block size is larger than 4×8 (or 8×4), the corresponding block is divided into four sub-partitions. We note that M×128 (M≤64) and 128×N (N≤64) ISP blocks may generate potential problems for 64×64 VDPU. For example, an M×128 CU in a single-tree case has an M×128 luma TB and two corresponding Chroma TB. If the CU uses ISP, the luma TB will be divided into 4 M×32TBs (can only be divided horizontally), each of which is smaller than a 64×64 block. However, in the current ISP design, the chroma blocks are indivisible. Therefore, the two chroma components will have a size larger than a 32×32 block. Similarly, a 128×NCU using ISP may also produce similar situations. Therefore, these two situations are problems with the 64×64 decoder pipeline. Therefore, the CU size that can use ISP is limited to a maximum value of 64×64. Fig.16A and 16B Exemplary schematic diagrams 1600 and 1650 respectively show sub-division depending on block size. Fig.16A and 16B Examples of both possibilities are shown. Fig.16A Examples of sub-partitioning for 4x8 and 8x4 CUs are shown. Fig. 16B Examples of sub-partitions for CUs other than 4x8, 8x4, and 4x4 are shown. All sub-partitions satisfy the condition of having at least 16 samples. In ISP, 1×N / 2×N sub-block predictions that rely on reconstructed values ​​of previously decoded 1×N / 2×N sub-blocks in the codec block are not allowed, so the minimum width of the prediction for a sub-block is four samples. For example, when an 8×N (N>4) codec block is encoded using ISP with vertical partitioning, it is divided into two prediction regions of size 4×N and four transforms of size 2×N. Similarly, when a 4×N codec block is encoded using ISP with vertical partitioning, it is predicted using a full 4×N block; four 1×N transforms are used. Although 1×N and 2×N transform sizes are allowed, it can be asserted that the transforms of these blocks within the 4×N region can be performed in parallel. For example, when a 4×N prediction region includes four 1×N transforms, there is no transform in the horizontal direction; the transform in the vertical direction can be performed as a single 4×N transform in the vertical direction. Similarly, when a 4×N prediction region includes two 2×N transform blocks, the transform operations of the two 2×N blocks in each direction (horizontally and vertically) can be performed in parallel. Therefore, there is no added delay in processing these smaller blocks compared to processing intra blocks of a 4×4 conventional codec. Table 4 Entropy coding and decoding coefficient group size Block size Coefficient group size 1×N,N≥16 1×16 N×1,N≥16 16×1 2×N,N≥8 2×8 N×2,N≥8 8×2 All other possible M×N situations 4×4 For each sub-partition, the reconstructed sample is obtained by adding the residual signal to the prediction signal. Here, the residual signal is generated by processes such as entropy decoding, inverse quantization, and inverse transformation. Therefore, the reconstructed sample value of each sub-partition can be used to generate a prediction for the next sub-partition, and each sub-partition is processed repeatedly. In addition, the first sub-partition to be processed is the upper left sample containing the CU, and then continues to the sub-partition downward (horizontal partition) or to the right (vertical partition). Therefore, the reference samples used to generate the sub-partition prediction signal are located only on the left and above the row. All sub-partitions share the same intra-frame mode. The following is a summary of the interaction of ISP with other codec tools. – Multiple Reference Line (MRL): If the MRL index of a block is not 0, the ISP codec mode will be inferred to be 0, so the ISP mode information will not be sent to the decoder. – Entropy coding coefficient group size: As shown in Table 4, the size of the entropy coding sub-block has been modified to have 16 samples in all possible cases. It is worth noting that the new size only affects blocks generated by ISP where one of the dimensions is less than 4 samples. In all other cases, the coefficient group maintains 4×4 dimensions. – CBF codec: It is assumed that at least one subpartition has a non-zero CBF. Thus, if n is the number of subpartitions, and the first n-1 subpartitions have yielded zero CBFs, the CBF of the nth subpartition is assumed to be 1. – Transform size restriction: All ISP transforms with length greater than 16 points use DCT-II. – MTS flag: If the CU uses ISP codec mode, the MTS CU flag will be set to 0 and will not be sent to the decoder. Therefore, the encoder will not perform RD tests on the different available transforms for each resulting sub-partition. The transform selection for ISP mode will instead be fixed and will be selected based on the intra mode used, the processing order and the block size. Therefore, no signaling is required. For example, let t H and t V are the horizontal and vertical transforms selected for the w×h sub-partition, respectively, where w is the width and h is the height. The transforms are then selected according to the following rules: - If w=1 or h=1, there is no horizontal transform or vertical transform, respectively. – If w ≥ 4 and w ≤ 16, t H =DST-VII, otherwise t H =DCT-II. – If h≥4 and h≤16, t V =DST-VII, otherwise t V =DCT-II. In ISP mode, all 67 intra prediction modes are allowed. PDPC is also applied if the corresponding width and height are at least 4 samples long. In addition, the reference sample filtering process (reference smoothing) and the conditions for intra interpolation filter selection no longer exist, and in ISP mode, the cubic (DCT-IF) filter is always applied for fractional position interpolation. 2.12. Matrix Weighted Intra Prediction (MIP) The matrix weighted intra prediction (MIP) method is a new intra prediction technique added to VVC. In order to predict the samples of a rectangular block with a width of W and a height of H, the matrix weighted intra prediction (MIP) takes a row of H reconstructed neighboring boundary samples on the left side of the block and a row of W reconstructed neighboring boundary samples above the block as input. If the reconstructed samples are not available, they are generated as in conventional intra prediction. Fig.17 As shown, the generation of the prediction signal is based on the following three steps, namely averaging, matrix-vector multiplication and linear interpolation. Fig.17 An example 1700 of a matrix-weighted intra prediction process is shown. 2.12.1. Average neighboring samples Among the boundary samples, four samples or eight samples are selected by averaging based on block size and shape. Specifically, input boundary bdry top and bdry left According to a predefined rule that depends on the block size, it is reduced to smaller boundary samples by averaging adjacent boundary samples. and Then, the two shrinking boundaries and Spliced ​​into a reduced boundary vector bdry red , so for blocks of shape 4×4, the size of the reduced boundary vector is 4, while for blocks of all other shapes, the size of the reduced boundary vector is 8. If mode refers to a MIP mode, this stitching is defined as follows: Matrix multiplication Taking the averaged samples as input, a matrix-vector multiplication is performed and then the offset is added. The result is a downscaled prediction signal generated on a subsampled set of samples in the original block. red The reduced prediction signal pred is generated in red , pred red The width is W red And the height is H red Here, W red and H red is defined as: By calculating the matrix-vector product and adding the offset, the downscaled prediction signal pred red Calculated: pred red =A.bdry red +b (2-12) Here, A is a matrix. If W = H = 4, then A has W red ·H red rows and 4 columns. In all other cases, A has 8 columns. b is of size W red ·H red The matrix A and the offset vector b are taken from one of the sets S0, S1, S2. The index idx=idx(W,H) is defined as follows: Here, each coefficient of matrix A is represented with 8 bits of precision. Set S0 consists of 16 matrices and 16 offset vectors Composition, each matrix With 16 rows and 4 columns, each offset vector The size of is 16. The matrices and offset vectors of this set are used for blocks of size 4×4. Set S1 consists of 8 matrices and 8 offset vectors Composition, each matrix With 16 rows and 8 columns, each offset vector The size of is 16. The set S2 consists of 6 matrices and 6 offset vectors Composition, each matrix With 64 rows and 8 columns, each offset vector The size is 64. 2.12.3. Interpolation The prediction signals at the remaining positions are generated from the prediction signals on the subsampled set by linear interpolation, which is a single-step linear interpolation in each direction. Interpolation is performed first in the horizontal direction and then in the vertical direction, and the execution order is not affected by the block shape or block size. 2.12.4.MIP mode signaling and coordination with other codec tools For each codec unit (CU) in intra mode, a flag is sent indicating whether the MIP mode is to be applied. If the MIP mode is to be applied, the MIP mode (predModeIntra) is signaled. For the MIP mode, a transposed flag (isTransposed) that determines whether the mode is transposed and a MIP mode identifier (modeId) that determines which matrix to use for a given MIP mode are derived as follows isTransposed=predModeIntra&1 modeId=predModeIntra>>1 (2-14) The MIP codec mode is coordinated with other codecs by taking into account the following aspects: – For MIPs on larger blocks, LFNST is enabled. Here, the LFNST transform in planar mode is used. - Reference sample derivation for MIP is performed exactly the same as for regular intra prediction mode. – For the upsampling step used in MIP prediction, the original reference samples are used instead of the downsampled samples. – Clipping is performed before upsampling, rather than after upsampling. – Regardless of the maximum transform size, MIP is allowed to be up to 64×64. The number of MIP modes is 32 for sizeId=0, 16 for sizeId=1, and 12 for sizeId=2. 2.13. Decoder-side intra-mode derivation In JEM-2.0, intra modes are extended from 35 to 67 modes in HEVC, and they are derived at the encoder and explicitly signaled to the decoder. In JEM-2.0, a lot of overhead is spent on intra mode encoding and decoding. For example, in all intra codec configurations, the intra mode signaling overhead can be as high as 5-10% of the total bit rate. This paper proposes a decoder-side intra mode derivation method to reduce the intra mode encoding and decoding overhead while maintaining prediction accuracy. In order to reduce the overhead of intra mode signaling, this paper proposes a decoder-side intra mode derivation (DIMD) method. In the proposed method, information is derived from neighboring reconstructed samples of the current block at both the encoder and the decoder, instead of explicitly signaling the intra mode. The intra mode derived by DIMD is used in two ways: 1) For a 2N×2N CU, when the corresponding CU-level DIMD flag is turned on, the DIMD mode is used as the intra mode for intra prediction; 2) For N×N CU, DIMD mode is used to replace a candidate of the existing MPM list to improve the efficiency of intra mode coding and decoding. 2.13.1. Template-based intra-mode derivation Fig.18 Schematic diagram 1800 showing target points, template points, and reference points of a template used in DIMD. Fig.18 As shown, the target represents the current block (block size is N) for which the intra prediction mode is to be estimated. The template (composed of Fig.18 The template size is expressed as the number of samples in the template that extend above and to the left of the target block, i.e., L. In the current implementation, templates of size 2 (i.e., L=2) are used for 4×4 and 8×8 blocks, and templates of size 4 (i.e., L=4) are used for 16×16 and larger blocks. According to the definition of JEM-2.0, the reference of the template (given by Fig.18 The dotted area in the figure refers to a set of neighboring samples from the top and left of the template. Unlike the template samples that are always from the reconstructed area, the reference samples of the template may not have been reconstructed when encoding / decoding the target block. In this case, the existing reference sample replacement algorithm of JEM-2.0 is used to replace the unavailable reference samples with the available reference samples. For each intra prediction mode, DIMD calculates the absolute difference (SAD) between the reconstructed template samples and its predicted samples obtained from the reference samples of the template. The intra prediction mode that generates the minimum SAD is selected as the final intra prediction mode for the target block. 2.13.2. DIMD of Intra-frame 2N×2N CU For intra 2N×2N CUs, DIMD is used as an additional intra mode, which is adaptively selected by comparing the DIMD intra mode with the best normal intra mode (i.e., explicitly signaled). For each intra 2N×2N CU, a flag is signaled to indicate the use of DIMD. If the flag is 1, the CU is predicted using the intra mode derived by DIMD; otherwise, DIMD is not applied and the CU is predicted using the intra mode explicitly signaled in the bitstream. When DIMD is enabled, chroma components always reuse the same intra mode derived for luma components, i.e., DM mode. In addition, for each DIMD-encoded CU, blocks in the CU can be adaptively selected to derive their intra modes at the PU level or the TU level. Specifically, when the DIMD flag is 1, another CU-level DIMD control flag is transmitted by a signal to indicate at which level DIMD is performed. If the flag is 0, it means that DIMD is performed at the PU level and all TUs in the PU use the same derived intra mode for intra prediction; otherwise (i.e., the DIMD control flag is 1), it means that DIMD is performed at the TU level and each TU in the PU derives its own intra mode. Further, when DIMD is enabled, the number of angular directions increases to 129, and DC mode and planar mode remain unchanged. To accommodate the increased granularity of angular intra modes, the precision of intra interpolation filtering for DIMD-encoded CUs is increased from 1 / 32 pixel to 1 / 64 pixel. In addition, in order to use the derived intra modes of DIMD-encoded CUs as MPM candidates for neighboring intra blocks, the 129 directions of DIMD-encoded CUs are converted to "normal" intra modes (i.e., 65 angular intra directions) before they are used as MPMs. 2.13.3. DIMD of Intra-N×N CU In the proposed method, the intra mode of an intra N×N CU is always signaled. However, to improve the efficiency of intra mode encoding and decoding, the intra mode derived from DIMD is used as the MPM candidate for predicting the intra mode of the four PUs in the CU. In order not to increase the overhead of MPM index signaling, the DIMD candidate is always placed first in the MPM list and the last existing MPM candidate is deleted. At the same time, deduplication is performed so that if a DIMD candidate is redundant, it will not be added to the MPM list. 2.13.4.DIMD Intra-frame Pattern Search Algorithm To reduce the complexity of encoding / decoding, a simple fast intra mode search algorithm is used for DIMD. First, an initial estimation process is performed to provide a good starting point for the intra mode search. Specifically, an initial candidate list is created by selecting N fixed modes from the allowed intra modes. Then, the SAD is calculated for all candidate intra modes, and the intra mode that minimizes the SAD is selected as the starting intra mode. In order to achieve a good complexity / performance trade-off, the initial candidate list consists of 11 intra modes, including DC mode, planar mode, and every 4 modes of the 33 angular intra directions defined in HEVC, namely intra modes 0, 1, 2, 6, 10…30, 34. If the starting Intra mode is DC mode or Planar mode, it is used as the DIMD mode. Otherwise, based on the starting Intra mode, a refinement process is then applied, where the best Intra mode is identified through an iterative search. It works by comparing the SAD values ​​of three Intra modes separated by a given search interval at each iteration and maintaining the Intra mode that minimizes the SAD. The search interval is then reduced to half and the Intra mode selected from the previous iteration is used as the center Intra mode for the current iteration. For the current DIMD implementation with 129 angular Intra directions, a maximum of 4 iterations are used in the refinement process to find the best DIMD Intra mode. 2.14. Decoder-side intra-mode derivation In this paper, a method is proposed to avoid transmitting the luma intra prediction mode in the bitstream. This is achieved by using previously encoded / decoded pixels to derive the luma intra mode in the same way at both the encoder and the decoder. The process defines a new codec mode called DIMD, the selection of which is signaled in the bitstream for intra codec blocks using a simple flag. DIMD competes with other codec modes at the encoder, including the classic intra codec mode (where the intra prediction mode is encoded). Note that in this paper, DIMD is applied only to luma. For chroma, the classic intra codec mode is applied. As done for other codec modes (classic intra, inter, merge, etc.), the rate-distortion cost is calculated for the DIMD mode and compared with the codec costs of other modes to decide whether to select the DIMD mode as the final codec mode for the current block. At the decoder side, the DIMD flag is parsed first. If the DIMD flag is true, the intra prediction mode is derived during the reconstruction process using the same previously coded neighboring pixels. Otherwise, the intra prediction mode is parsed from the bitstream as in the classic intra codec mode. 2.14.1. Intra-frame prediction mode derivation 2.14.1.1 Gradient analysis To derive the intra prediction mode for a block, a set of neighboring pixels is first selected on which the gradient analysis is performed. For normative purposes, these pixels should be in the pool of decoded / reconstructed pixels. Fig.19 1600 is shown showing a selected set of pixels on which gradient analysis is performed. Fig.19 As shown, the templates around the current block are selected by T pixels to the left and T pixels above it. In the proposal, T=2 is set. Next, a gradient analysis is performed on the pixels of the template. This determines the dominant angular orientation of the template, assuming (a core premise of our approach) that it is likely to be the same as the angular orientation of the current patch. Therefore, a simple 3×3 Sobel gradient filter is used, defined by the following matrix that will be convolved with the template: For each pixel of the template, each of the two matrices is a 3×3 window centered on the current pixel, each matrix is ​​point-by-point multiplied and combined with its 8 immediate neighbors, and then the results are summed. Therefore, the two values ​​corresponding to the gradient at the current pixel Gx (from the product with Mx) and Gy (from the product with My) are obtained in the horizontal direction and vertical direction, respectively. Fig. 20 The convolution process 2000 is shown, which shows the convolution of a 3×3 Sobel gradient filter with a template. The current pixel is Fig. 20 The template pixels (including the current pixel) are pixels that can be subjected to gradient analysis. Unavailable pixels are pixels that cannot be subjected to gradient analysis due to the lack of some neighbors. Reconstructed pixels are available (reconstructed) pixels that are outside the template under consideration and are used for gradient analysis of template pixels. If a reconstructed pixel is unavailable (for example, because the block is too close to the image boundary), the gradient analysis of all template pixels that use this reconstructed pixel will not be performed. 2.14.1.2 Gradient Histogram and Mode Derivation For each template pixel, use G x and G y The magnitude (G) and direction (O) of the gradient are calculated as follows: It is worth noting that a fast implementation of the atan function is proposed. The direction of the gradient is then converted into an intra-frame angular prediction mode and used to index the histogram (initialized to zero first). The histogram value for this intra-frame angular mode is increased by G. Once all template pixels in the template have been processed, the histogram will contain the cumulative value of the gradient strength for each intra-frame angular mode. The mode showing the highest peak in the histogram is selected as the intra-frame prediction mode for the current block. If the maximum value in the histogram is 0 (meaning that gradient analysis cannot be performed, or the area making up the template is flat), then the DC mode is selected as the intra-frame prediction mode for the current block. For the block at the top of the CTU, the gradient analysis of the pixels at the top of the template is not performed. The DIMD flag is encoded using three possible contexts, depending on the neighboring blocks on the left and above, similar to the Skip flag encoding. Context 0 corresponds to the case where neither the left neighboring block nor the neighboring block above is encoded using DIMD mode, context 1 corresponds to the case where only one neighboring block is encoded using DIMD, and context 2 corresponds to the case where both neighbors are DIMD encoded. The initial symbol probability of each context is set to 0.5. 2.14.2.1 Prediction of 30 Intra-frame Modes One advantage DIMD offers over the classic intra mode codec is that the derived intra mode can have higher precision, allowing more accurate prediction at no additional cost since it is not transmitted in the bitstream. The derived intra modes cross-cover 129 angular modes, so there are 130 modes in total including DC (in this context, a derived intra mode can never be planar). The classic intra codec modes are unchanged, i.e., the prediction and mode codec still use 67 modes. The required changes are performed for wide-angle intra prediction and simplified PDPC to accommodate prediction using 129 modes. Note that only the prediction process uses the extended intra modes, which means that for any other purpose (e.g., deciding whether to filter reference samples), the modes are converted back to 67 mode precision. 2.14.3. Other regulatory changes In DIMD mode, the luma intra mode is derived during the reconstruction process before the block is reconstructed. This is done to avoid dependencies on reconstructed pixels during parsing. However, by doing this, the luma intra mode of a block will be undefined for the chroma components of the block and the luma components of neighboring blocks. This causes a problem because: For chroma, a fixed list of mode candidates is defined. In general, if the luma mode is equal to one of the chroma candidates, the candidate will be replaced by the vertical diagonal (VDIA_IDX) intra mode. Since in DIMD, luma mode is not available, the initial chroma mode candidate list is not modified. In classic intra mode where luma intra prediction mode is parsed from the bitstream, the MPM list is constructed using luma intra modes of neighboring blocks, which may not be available if those blocks are coded using DIMD. In this case, in this paper, DIMD coded blocks are treated as inter blocks during MPM list construction, which means that they are effectively considered unavailable. 2.15.DIMD The three angle modes are selected from the Histogram of Gradients (HoG) calculated from the neighboring pixels of the current block. Once the three modes are selected, their prediction values ​​are calculated normally and then their weighted average is used as the final prediction value for the block. To determine the weights, the corresponding magnitude in the HoG is used for each of the three modes. The DIMD mode is used as an optional prediction mode and is always checked in FullRD mode. The current version of DIMD has modified some aspects in signaling, HoG calculation and prediction fusion. The purpose of this modification is to improve the codec performance and address the complexity issues raised during the last meeting (i.e., throughput of 4×4 blocks). The following sections describe the modifications in each aspect. 2.15.1. Signaling Fig.21 The proposed intra-block decoding process 2100 is shown. Fig.21 The order of parsing flags / indexes integrated with the proposed DIMD in VTM5 is shown. As can be seen, the DIMD flag of a block is first parsed using a single CABAC context, which is initialized to a default value of 154. If flag == 0, parsing will continue normally. Otherwise (if flag == 1), only the ISP index is parsed, and the following flags / indexes are inferred to be zero: BDPCM flag, MIP flag, MRL index. In this case, the entire IPM parsing is also skipped. During the parsing phase, when a regular non-DIMD block queries its DIMD neighbor's IPM, the pattern PLANAR_IDX is used as a virtual IPM for the DIMD block. 2.15.2. Texture analysis Fig. 22 A schematic diagram 2200 showing the HoG calculation from a template with a width of 3 pixels is shown. Texture analysis of DIMD includes the calculation of the Histogram of Gradients (HoG) ( Fig. 22 ). HoG calculation is performed by applying horizontal and vertical Sobel filters to the pixels in a template of width 3 around the block. The exception is that if the upper template pixels belong to a different CTU, they will not be used for texture analysis. Once calculated, the IPMs corresponding to the two highest histogram bins are selected for the block. In previous versions, all pixels in the center line of the template participated in the HoG calculation. However, the current version improves the throughput of this process by applying the Sobel filter more sparsely on 4x4 blocks. For this purpose, only one pixel from the left and one pixel from the top are used. This Fig. 22 is shown in . Besides reducing the number of operations for gradient computation, this feature also simplifies the selection of the best 2 patterns from the HoG, since the resulting HoG cannot have more than two non-zero magnitudes. 2.15.3. Prediction Fusion The current version of the method also uses a fusion of three prediction values ​​for each block, just like the previous version. However, the choice of prediction mode is different and a combined hypothetical intra prediction method is used, where planar mode is considered to be used in combination with other modes when calculating candidates for intra prediction. In the current version, the two IPMs corresponding to the two highest HoG stripes are combined with planar mode. Prediction fusion is applied as a weighted average of the three predictions above. For this purpose, the weight of the plane is fixed to 21 / 64 (~1 / 3). The remaining weight of 43 / 64 (~2 / 3) is then shared between the two HoG IPMs in proportion to the magnitude of their HoG stripes. Fig.23 Visualize this process. Fig.23 The prediction fusion by weighted averaging of two HoG modes and planes is shown. 2.16. Combined Inter and Intra Prediction (CIIP) In VVC, when a CU is encoded and decoded in Merge mode, if the CU contains at least 64 luma samples (i.e., the CU width multiplied by the CU height is equal to or greater than 64), and if both the CU width and the CU height are less than 128 luma samples, an additional flag is transmitted by signal to indicate whether the combined inter / intra prediction (CIIP) mode is applied to the current CU. As its name suggests, CIIP prediction combines the inter prediction signal with the intra prediction signal. The inter prediction signal P in CIIP mode inter The intra prediction signal P is derived using the same inter prediction process applied to the conventional Merge mode; and intra The conventional intra prediction process using planar mode is derived. Fig.24 A schematic diagram 2400 showing the top and left neighboring blocks used for CIIP weight derivation is shown. Then, the intra-frame and inter-frame prediction signals are combined using weighted averaging, where the weight values ​​depend on the coding mode of the top and left neighboring blocks and are calculated as follows (e.g. Fig.24 shown): – If the top neighbor is available and is intra-coded, set isIntraTop to 1, otherwise set isIntraTop to 0; – If the left neighbor is available and is intra-coded, set isIntraLeft to 1, otherwise set isIntralLeft to 0; – If (isIntraLeft+isIntraTop) is equal to 2, wt is set to 3; – Otherwise, if (isIntraLeft+isIntraTop) is equal to 1, wt is set to 2; – Otherwise, set wt to 1. The CIIP forecast is constructed as follows: P CIIP =((4-wt)*P intra +wt*P intra +2)>>2 (2-17) 2.17. Geometric Partitioning Mode (GPM) In VVC, geometric partitioning mode is supported for inter prediction. The geometric partitioning mode is transmitted as a Merge mode through a signal using a CU level flag. Other Merge modes include normal Merge mode, MMVD mode, CIIP mode, and sub-block Merge mode. For each possible CU size w×h=2 m ×2 n Where m,n∈{3…6} and excluding 8x64 and 64x8, the geometric segmentation mode supports a total of 64 segmentations. When this mode is used, the CU is divided into two parts by a geometrically positioned straight line ( Fig.25 ). Fig.25 Schematic 2500 showing an example of a GPM partition grouped by the same angle. The location of the partition line is mathematically derived from the angle and offset parameters of the particular partition. Each part of the geometric partition in the CU is inter-predicted using its own motion; only unidirectional prediction is allowed for each partition, i.e., each part has one motion vector and one reference index. Unidirectional prediction motion constraints are applied to ensure that only two motion compensated predictions are required for each CU, as with conventional bidirectional prediction. The unidirectional prediction motion for each partition is derived using the process described in Section 2.17.1. If geometric partitioning mode is used for the current CU, a geometric partitioning index (angle and offset) indicating the partitioning mode of the geometric partitioning and two Merge indices (one for each partition) are further transmitted through the signal. The number of maximum GPM candidate sizes is explicitly transmitted through the signal in the SPS, and the syntax binarization for the GPM Merge index is specified. After predicting each part of the geometric partitioning, the sample values ​​along the edges of the geometric partitioning are adjusted using the hybrid processing with adaptive weights in Section 2.17.2. This is a prediction signal for the entire CU, and like in other prediction modes, the transform and quantization process will be applied to the entire CU. Finally, as in Section 2.17.3, the motion field of the CU predicted using the geometric partitioning mode is stored. 2.17.1. One-way prediction candidate list construction The unidirectional prediction candidate list is derived directly from the Merge candidate list constructed according to the extended Merge prediction process. Denote n as the index of the unidirectional prediction motion in the geometric unidirectional prediction candidate list. The LX motion vector of the nth extended Merge candidate is used as the nth unidirectional prediction motion vector for the geometric partition mode, where X equals the parity of n. Fig.26 A schematic diagram 2600 is shown showing unidirectional prediction MV selection for geometric partitioning mode. These motion vectors are Fig.26 If the corresponding LX motion vector of the n-th extended Merge candidate does not exist, the L(1-X) motion vector of the same candidate is used as the unidirectional prediction motion vector for the geometric partition mode instead. 2.17.2. Blending Along Geometric Partition Edges After predicting each part of the geometric partition using its own motion, blending is applied to the two prediction signals to derive samples around the geometric partition edge. The blending weight for each position of the CU is derived based on the distance between the respective position and the partition edge. The distance from position (x, y) to the segmentation edge is derived as: where i,j are the angle and offset indices of the geometric partition, which depend on the geometric partition index transmitted by the signal. x,j and ρ y,j The sign of depends on the angle index i. The weights of each part of the geometric segmentation are derived as follows: wIdxL(x,y)=partIdx? 32+d(x,y):32-d(x,y) (2-22) w1(x,y)=1-w0(x,y) (2-24) partIdx depends on the angle index i. An example of weight w0 is Fig. 27 shown. Fig. 27 A schematic diagram 2700 illustrating an exemplary generation of warp weights w_0 using a geometric partitioning mode is shown. 2.17.3. Motion Field Storage for Geometry Partitioning Mode Mv1 from the first part of the geometric partition, Mv2 from the second part of the geometric partition, and Mv which is a combination of Mv1 and Mv2 are stored in the motion field of the CU coded in the geometric partition mode. The type of motion vector stored for each single position in the motion field is determined as: sType=abs(motionIdx)<32?2: (motionIdx≤0?(1-partIdx):partIdx) (2-25) Where motionIdx is equal to d(4x+2,4y+2), which is recalculated according to equation (2-18). partIdx depends on the angle index i. If sType is equal to 0 or 1, then Mv0 or Mv1 is stored in the corresponding motion field, otherwise, if sType is equal to 2, then the combined Mv from Mv0 and Mv2 is stored. The combined Mv is generated using the following process: 1) If Mv1 and Mv2 are from different reference picture lists (one from L0 and the other from L1), then Mv1 and Mv2 are simply combined to form a bi-directional prediction motion vector. Otherwise, if Mv1 and Mv2 are from the same list, only the unidirectional predicted motion Mv2 is stored. 2.18. Multiple Hypothesis Prediction (MHP) In Multiple Hypothesis Prediction (MHP), up to two additional prediction values ​​are signaled over Inter AMVP mode, Normal Merge mode, Affine Merge and MMVD mode. The resulting overall prediction signal is iteratively accumulated with each additional prediction signal. p n+1 =(1-α n+1 ) n +α n+1 h n+1 . The weighting factor α is specified according to the following table: add_hyp_weight_idx α 0 1 / 4 1 -1 / 8 For inter-AMVP mode, MHP is applied only if unequal weights in BCW are selected in bi-prediction mode. The additional assumption may be Merge mode or AMVP mode. In the case of Merge mode, motion information is indicated by a Merge index, and the Merge candidate list is the same as in the geometric partitioning mode. In the case of AMVP mode, reference index, MVP index, and MVD are transmitted by signal. 2.19. Convolutional Cross-Component Model (CCCM) for Intra Prediction In this paper, we propose to apply chroma prediction based on convolutional cross-component model (CCCM) to improve the compression efficiency of ECM. When CCCM prediction mode is activated through PU level flag, the proposed 7-tap convolution filter maps luma values ​​to chroma values. The filter input consists of 5 spatial luma samples, a nonlinear term and a bias term. For each chroma block, the filter coefficients are derived for reference samples in the neighborhood of the PU using regression-based MSE minimization. 3. Question The current design of intra prediction can only use the redundancy between the samples in the current block and the samples in the neighboring current blocks. The non-local texture information is not fully used. 4. Detailed solution The following detailed embodiments should be considered as examples to explain the general concept. These embodiments should not be interpreted in a narrow manner. In addition, these embodiments can be combined in any way. In this disclosure, the term decoder-side intra mode derivation (DIMD) or template-based intra prediction mode (TIMD) refers to codec tools that use previously decoded blocks to derive intra prediction modes. In the present disclosure, a “regular intra prediction mode (IPM) candidate set” is used to indicate allowed IPMs for blocks used for intra coding (e.g., 35 modes in HEVC, 67 modes in VVC), and a “regular intra prediction mode” may refer to an IPM in the regular IPM candidate set. Fig.28 A schematic diagram 2800 showing a conventional angular IPM (indicated by a black arrow) and an extended angular IPM (indicated by a dashed line) is shown. In the present disclosure, an "extended intra prediction mode (IPM) candidate set" includes all conventional IPMs and extended IPMs (such as in Fig.28 in the example). In the following discussion, RightShift(x,n) can be defined as RightShift(x,n) can also be defined as RightShift(x,n)=(x+offset0)>>n. Clip3(x)=max(minV,min(maxV,x)), where minV and maxV are the minimum and maximum values ​​of the samples. Concept of the proposed method 1. In one example, the intra prediction at position (x, y) denoted as P(x, y) can be obtained by the surface equation P(x,y)=f(x,y) is modeled. a. For example, P(x,y) can be described by a linear model. i. For example, P(x,y)=f1(x,y)=ax+by+c. b. For example, P(x,y) can be described by a quadratic model. i. For example, P(x,y)=f2(x,y)=ax 2 +by 2 +cxy+dx+ey+f. c. For example, P(x,y) can be described by a polynomial model. i. For example, Wherein M is an integer, such as 3. d. For example, P(x,y) can be described by a model consisting of at least one non-polynomial calculation. i. For example, non-polynomial calculations can be cos, sin, tan, cos -1 ,sin -1 ,tan -1 ,cosh,sinh,tanh,cosh -1 ,sinh -1 ,tanh -1 ,ax,log a x or any combination between them and / or polynomial functions. 2. In one example, the origin coordinate (0,0) may be at a specific location. a. For example, (0,0) can be the upper left corner of the current block. b. For example, (0,0) can be at the upper right corner of the current block. c. For example, (0,0) can be at the lower left corner of the current block. d. For example, (0,0) can be at the lower right corner of the current block. e. For example, (0,0) may be located at the center of the current block. f. For example, (0,0) may be a position adjacent to the current block. i. For example, (0,0) may be the upper left position adjacent to the current block. 3. In one example, the predicted value may be calculated using integer arithmetic. a. In one example, shifting and / or clipping may be applied. b. In one example, P(x,y)=RightShift(f(x,y),s), where s is an integer. c. In one example, P(x,y)=Clip3(RightShift(f(x,y),s)), where s is an integer. Derivation of parameters 4. In one example, at least one parameter of the model (such as a, b, c, d, e, f, a in item 1) i ) Can be specified as a fixed value. a. In one example, the fixed value may be zero. 5. In one example, at least one parameter of the model (such as a, b, c, d, e, f, a in item 1) i ) can be derived from at least one neighboring reconstruction sample. a. In one example, N neighboring reconstruction samples may be included in a training set. Based on the training set, least mean square (LMS), also known as regression-based MSE minimization or linear regression, may be used to determine parameters. i. For the training sample points (also called reference sample points) in the training process, the coordinates of the training sample points are regarded as inputs, and the reconstructed sample point values ​​are regarded as target values. ii. After going through all training samples, the parameters that produce the minimum mean square error between the reconstructed sample values ​​and the predicted sample values ​​from the model are output. b. In one example, parameter derivation can share the same derivation model with another codec tool (such as CCCM). c. In one example, the training sample points may be adjacent and / or non-adjacent neighboring samples of the current block, such as Fig.29 as shown in . Fig.29 A schematic diagram 2900 showing the locations of training samples is shown. i. For example, the training samples may include K reconstructed lines above the current block (eg, K=1, 2, ...). ii. For example, the training samples may include K reconstructed columns to the left of the current block (eg, K=1, ...). iii. For example, whether a sample point can be included in the training set may depend on whether it is available. iv. For example, whether a sample point can be included in the training set may depend on whether it has been reconstructed. v. The determination of the training set may depend on the size / position of the current block. Signaling 6. In one example, whether to apply the proposed method claimed in item 1 may be signaled in a first syntax element (SE). a. In one example, SE may be signaled in SPS / PPS / sequence header / picture header / slice header / CTU / CU / TU / PU / etc. b. In one example, SE can be predictively encoded and decoded. c. In one example, SE is conditionally transmitted via a signal. i. For example, the SE for a block is signaled only if the block is intra-coded. d. In one example, SE can be encoded and decoded using at least one arithmetic context model. ii. The arithmetic context model may depend on neighboring blocks. e. In one example, SE can be bypassed codec. f. In one example, SE can be encoded and decoded in a layered manner. iii. If a higher level SE (such as in SPS) is signaled to be disabled, a lower level SE (such as in CU) may not be signaled. 7. In one example, how to apply the proposed method claimed in item 1 can be used with the second syntax element (SE) is transmitted via the signal. a. In one example, SE can be in SPS / PPS / sequence header / picture header / slice header It is transmitted through signals in / CTU / CU / TU / PU / , etc. b. In one example, SE can be predictively encoded and decoded. c. In one example, SE is conditionally transmitted via a signal. iv. For example, only in case the proposed method is applied (which can be signaled with the first SE), the SE for the block is signaled. d. In one example, SE can be encoded and decoded using at least one arithmetic context model. v. The arithmetic context model may depend on neighboring blocks. e. In one example, SE can be bypassed codec. f. In one example, SE can be encoded and decoded in a layered manner. 8. In one example, at least one parameter of the model (such as a, b, c, d, e, f, a in item 1) i ) It can be transmitted by signaling in a third SE. a. In one example, SE is signaled only if the model is applied. g. In one example, SE can be encoded and decoded using at least one arithmetic context model. h. In one example, SE can be bypassed codec. i. In one example, SE can be predictively encoded and decoded. i. In one example, the prediction of the parameter can be derived according to Derivation of parameters the chapter of 9. In one example, the proposed method claimed in item 1 can be regarded as an intra prediction mode that can be predictively coded. vi. If the SE at a higher level (such as in the SPS) is signaled to be disabled, the SE at a lower level (such as in the CU) may not be signaled. General requirements 10. Whether and / or how to apply the method disclosed above can be signaled in the bitstream. j. In one example, they can be signaled at the sequence level / group of pictures level / picture level / strip level / slice group level, such as in the sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / strip header / slice group header. k. In one example, they can be signaled at the PB / TB / CB / PU / TU / CU / VPDU / CTU / CTU row / strip / slice / sub-picture / other types of regions containing more than one sample or pixel. 11. Whether and / or how to apply the method disclosed above can depend on coding information, such as block size, color format, single / double tree segmentation, color component, slice / picture type. a. For example, it can be based on whether the width (W) and / or height (H) of the video unit meet predefined conditions, such as one or more combinations of the following: i. W < T1 or W <= T1, where T1 can be 8 or 16 or 32 or 64. ii. W > T2 or W >= T2, where T2 can be 2 or 4 or 8. iii. H < T3 or H <= T3, where T3 can be 8 or 16 or 32 or 64. iv. H > T4 or H >= T4, where T4 can be 2 or 4 or 8. v. W / H < T5 or W / H <= T5, where T5 can be 1 / 8 or 1 / 4 or 1 / 2 or 1 or 2 or 4 or 8 or 16. vi. W / H > T6 or W / H >= T6, where T6 can be 1 / 8 or 1 / 4 or 1 / 2 or 1 or 2 or 4 or 8 or 16. vii. H / W < T7 or H / W <= T7, where T7 can be 1 / 8 or 1 / 4 or 1 / 2 or 1 or 2 or 4 or 8 or 16. viii. H / W>T8 or H / W>=T8, where T8 can be 1 / 8 or 1 / 4 or 1 / 2 or 1 or 2 or 4 or 8 or 16. ix.W == T9, where T9 can be 8 or 16 or 32 or 64. xH == T10, where T10 can be 8 or 16 or 32 or 64.

[0087] Fig.30 A flow chart of a method 3000 for video processing according to an embodiment of the present disclosure is shown. The method 3000 is implemented for conversion between a current video block of a video and a bit stream of the video.

[0088] At block 3010, a metric for intra prediction for a current video block is determined. At block 3020, an intra prediction for a sample at a first position in the current video block is determined based on the metric and the first position.

[0089] At block 3030, conversion is performed based on the intra prediction. In some embodiments, the conversion may include encoding the current video block into a bitstream. Alternatively or additionally, the conversion may include decoding the current video block from a bitstream.

[0090] The method 3000 enables determining the intra prediction of the sample at the position based on the metric and the position of the sample, thereby improving the intra prediction of the current video block. In this way, the coding efficiency and coding effectiveness can be improved.

[0091] In some embodiments, the metric includes a surface metric, such as a surface equation. For example, an intra prediction at a position (x, y) represented as P(x, y) can be modeled by a surface equation P(x, y) = f(x, y). The surface metric includes at least one of the following: a linear metric, a quadratic metric, a polynomial metric, or a non-polynomial metric.

[0092] In some embodiments, the non-polynomial metric includes at least one of the following: a trigonometric metric such as cosine, sin, or tan, such as cosine. -1 ,sin -1 ,tan -1 Inverse trigonometric measures, such as cosh, sinh, tanh Hyperbolic trigonometric measures, such as cosh -1 ,sinh -1 ,tanh -1 Inverse hyperbolic trigonometric measures, exponential measures such as log a The logarithmic metric of x, or a combination of non-polynomial and polynomial metrics.

[0093] In some embodiments, the metric includes one of the following: ax+by+c, ax 2+by 2 +cxy+dx+ey+f, or Where (x, y) represents the coordinates of the sample point in the current video block, and a, b, c, d, e, f and M are the parameters of the measurement.

[0094] In some embodiments, the first position of the sample point is a relative position relative to the second position. For example, the second position can be referred to as an original coordinate. The origin coordinate (0,0) can be at a specific position. As an example, the second position includes one of the following: the upper left corner position of the current video block, the upper right corner position of the current video block, the lower left corner position of the current video block, the lower right corner position of the current video block, the center position of the current video block, or a position adjacent to the current video block.

[0095] In some embodiments, the position adjacent to the current video block includes an upper left position adjacent to the current video block. For example, (0,0) may be an upper left position adjacent to the current block.

[0096] In some embodiments, determining the intra prediction of the sample at the first position comprises: determining a value of a metric based on the first position; and determining the intra prediction of the sample at the first position by applying at least one of the following to the value: an integer operation, a shift operation, or a clipping operation. For example, the prediction value may be calculated using integer operations. Shifting and / or clipping may be applied.

[0097] In some embodiments, the intra prediction of the sample at the first position is determined by using RightShift(f(x,y),s), wherein (x,y) represents the coordinates of the first position, f() represents a metric, f(x,y) represents a value of the metric, s represents a predefined integer, RightShift(f(x,y),s) represents a shift operation, and wherein the shift operation RightShift(f(x,y),s) is defined as: Where offset0 and offset1 are predefined values.

[0098] In some embodiments, intra prediction of a sample at a first position is determined by using Clip3(RightShift(f(x,y),s)), where (x,y) represents the coordinates of the first position, f() represents a metric, f(x,y) represents a value of the metric, s represents a predefined integer, RightShift(f(x,y),s) represents a shift operation, Clip3(RightShift(f(x,y),s)) represents a clipping operation, where the shift operation RightShift(f(x,y),s) is defined as: Wherein offset0 and offset1 are predefined values; and wherein the clipping operation Clip3(z) is defined as: Clip3(z)=max(minV,min(maxV,z)), where minV and maxV represent the minimum and maximum values ​​of the samples of the current video block, and z represents the value determined by the shift operation RightShift(f(x,y),s).

[0099] In some embodiments, at least one parameter of the metric includes at least one predefined value. For example, a, b, c, d, e, f, a, b, c, d, e, f, c, d ... i Parameters such as , may be specified as fixed values. As an example, the fixed value may be zero.

[0100] In some embodiments, determining the metric includes: determining at least one parameter of the metric based on at least one reference sample of the current video block. For example, such as a, b, c, d, e, f, a described above i Parameters such as these may be derived from at least one adjacent reconstruction sample point.

[0101] In some embodiments, at least one parameter is determined based on at least one reference sample using at least one of the following: least mean square (LMS), regression-based mean square error (MSE) minimization, or linear regression. As used herein, the reference sample may be referred to as a training sample. At least one reference sample may be included in a training set. The process for determining at least one parameter may be referred to as a training process.

[0102] In some embodiments, the at least one reference sample comprises at least one neighboring reconstructed sample of the current video block, and determining at least one parameter of the metric comprises: determining at least one reconstructed sample value based on at least one coordinate of the at least one neighboring reconstructed sample; and determining at least one parameter based on at least one difference between the at least one reconstructed sample value and at least one predicted sample value determined according to the metric. For training samples (also referred to as reference samples) in the training process, the coordinates of the training samples are considered as inputs, and the reconstructed sample values ​​are considered as target values. After passing through all training samples, the parameter that can produce the minimum mean square error between the reconstructed sample value and the predicted sample value from the pattern is output.

[0103] In some embodiments, at least one parameter of the metric is determined using a parameter derivation tool used for a codec tool.In some embodiments, the codec tool comprises a convolutional cross-component model (CCCM).

[0104] In some embodiments, the at least one reference sample includes at least one of the following: a first neighboring sample adjacent to the current video block, or a second neighboring sample non-adjacent to the current video block. For example, the reference sample may be a neighboring sample adjacent to and / or non-adjacent to the current video block, such as Fig.29 shown.

[0105] In some embodiments, the at least one reference sample includes at least one of the following: at least one reconstructed sample row above the current video block, or at least one reconstructed sample column to the left of the current video block.

[0106] In some embodiments, at least one reference sample of the current video block is determined based on at least one of: whether the reference sample is available, whether the reference sample is reconstructed, the size of the current video block, or the position of the current video block. For example, whether a sample can be included in the training set can depend on whether it is available. For another example, whether a sample can be included in the training set can depend on whether it has been reconstructed. The determination of the training set can depend on the size / position of the current video block.

[0107] In some embodiments, information about the application of the method is included in at least one syntax element in the bitstream. In some embodiments, the at least one syntax element includes at least one of the following: a first syntax element indicating whether to apply the method, or a second syntax element indicating how to apply the method. For example, whether to apply the proposed method can be signaled in the first syntax element. How to apply the proposed method can be signaled in the second syntax element.

[0108] In some embodiments, if the current video block is intra-coded, the first syntax element is included in the bitstream.

[0109] In some embodiments, if the method is applied to the current video block, the second syntax element is included in the bitstream. For example, the second syntax element for the block is signaled only if the proposed method is applied (which may be signaled using the first SE).

[0110] In some embodiments, at least one syntax element is included in at least one of the following: a sequence parameter set (SPS), a picture parameter set (PPS), a sequence header, a picture header, a slice header, a codec tree unit (CTU), a codec unit (CU), a transform unit (TU) or a prediction unit (PU).

[0111] In some embodiments, at least one syntax element is included in at least one of the following: sequence level, picture group level, picture level, slice level, slice group level, sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), decoding parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptation parameter set (APS), slice header, or slice group header.

[0112] In some embodiments, at least one syntax element is included in a region containing more than one sample or pixel.

[0113] In some embodiments, the region includes one of the following: a prediction block (PB), a transform block (TB), a codec block (CB), a prediction unit (PU), a transform unit (TU), a codec unit (CU), a virtual pipeline data unit (VPDU), a codec tree unit (CTU), a CTU row, a slice, a slice, or a sub-picture.

[0114] In some embodiments, at least one syntax element is predictively coded.

[0115] In some embodiments, at least one syntax element is encoded using at least one arithmetic context model.

[0116] In some embodiments, at least one arithmetic context model is based on at least one neighboring block of the current video block.

[0117] In some embodiments, at least one syntax element is bypass coded.

[0118] In some embodiments, at least one syntax element is encoded and decoded in a hierarchical manner.

[0119] In some embodiments, the information is based on codec information of the current video block.

[0120] In some embodiments, the codec information includes at least one of the following: block size, color format, single tree or dual tree partitioning, color component, slice type, or picture type.

[0121] In some embodiments, the information is based on whether at least one of a width of the current video block or a height of the current video block satisfies at least one condition.

[0122] In some embodiments, at least one condition includes at least one of the following: the width of the current video block is less than a first threshold, the width of the current video block is less than or equal to the first threshold, the width of the current video block is greater than a second threshold, the width of the current video block is greater than or equal to the second threshold, the height of the current video block is less than a third threshold, the height of the current video block is less than or equal to the third threshold, the height of the current video block is greater than a fourth threshold, the height of the current video block is greater than or equal to the fourth threshold, the first ratio of width to height is less than a fifth threshold, the first ratio is less than or equal to the fifth threshold, the first ratio is greater than a sixth threshold, the first ratio is greater than or equal to the sixth threshold, the second ratio of height to width is less than a seventh threshold, the second ratio is less than or equal to the seventh threshold, the second ratio is greater than an eighth threshold, the second ratio is greater than or equal to the eighth threshold, the width is equal to the first value, or the height is equal to the second value.

[0123] In some embodiments, the first threshold value (T1) may be one of: 8, 16, 32, or 64. The second threshold value (T2) may be one of: 2, 4, or 8. The third threshold value (T3) may be one of: 8, 16, 32, or 64. The fourth threshold value (T4) may be one of: 2, 4, or 8. The fifth threshold value (T5) may be one of: 1 / 8, 1 / 4, 1 / 2, 1, 2, 4, 8, or 16. The sixth threshold value (T6) may be one of: 1 / 8, 1 / 4, 1 / 2, 1, 2, 4, 8, or 16. The seventh threshold value (T7) may be one of: 1 / 8, 1 / 4, 1 / 2, 1, 2, 4, 8, or 16. The eighth threshold value (T8) may be one of: 1 / 8, 1 / 4, 1 / 2, 1, 2, 4, 8, or 16. The first value (T9) may be one of: 8, 16, 32, or 64. The second value (T10) can be one of the following: 8, 16, 32 or 64.

[0124] In some embodiments, at least one parameter of the metric is included in a third syntax element in the bitstream.In some embodiments, if the metric is applied to determine intra prediction for the current video block, the third syntax element is included in the bitstream.

[0125] In some embodiments, the third element is encoded using at least one arithmetic context model.

[0126] In some embodiments, the third syntax element is bypass coded.

[0127] In some embodiments, the third syntax element is predictively coded.

[0128] In some embodiments, the method is performed as an intra prediction mode, the intra prediction mode being a predictive codec.

[0129] In some embodiments, if syntax elements of a first level associated with an intra-prediction mode are to be disabled, syntax elements of a second level lower than the first level associated with the intra-prediction mode are not included in the bitstream.

[0130] According to an additional embodiment of the present disclosure, a non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a bitstream of a video generated by a method performed by an apparatus for video processing. In the method, a metric for intra-frame prediction of a current video block of the video is determined. Based on the metric and a first position, an intra-frame prediction of a sample at a first position in the current video block is determined. A bitstream is generated based on the intra-frame prediction.

[0131] According to some other embodiments of the present disclosure, a method for storing a bitstream of a video is provided. In the method, a metric for intra-frame prediction of a current video block of the video is determined. Based on the metric and a first position, an intra-frame prediction of a sample at a first position in the current video block is determined. A bitstream is generated based on the intra-frame prediction. The bitstream is stored in a non-transitory computer-readable recording medium.

[0132] The embodiments of the present disclosure may be described according to the following items, the features of which may be combined in any reasonable way.

[0133] Item 1. A method for video processing, comprising: determining a metric for intra-frame prediction for a current video block of a video and a bitstream of the video; determining, based on the metric and a first position, an intra-frame prediction for a sample at the first position in the current video block; and performing the conversion based on the intra-frame prediction.

[0134] Item 2. A method according to Item 1, wherein the metric includes a surface metric, and the surface metric includes at least one of the following: a linear metric, a quadratic metric, a polynomial metric, or a non-polynomial metric.

[0135] Item 3. A method according to Item 2, wherein the non-polynomial metric includes at least one of the following: a trigonometric function metric, an inverse trigonometric function metric, a hyperbolic trigonometric function metric, an inverse hyperbolic trigonometric function metric, an exponential metric, a logarithmic metric, or a combination of a non-polynomial metric and a polynomial metric.

[0136] Item 4. A method according to item 2 or item 3, wherein the metric comprises one of the following: ax+by+c, ax 2 +by 2 +cxy+dx+ey+f, or Wherein (x, y) represents the coordinates of the sample point in the current video block, and a, b, c, d, e, f and M are parameters of the metric.

[0137] Item 5. The method according to any one of Items 1 to 4, wherein the first position of the sample point is a relative position with respect to a second position, and the second position comprises one of the following: The upper left corner position of the current video block, the upper right corner position of the current video block, the lower left corner position of the current video block, the lower right corner position of the current video block, the center position of the current video block, or a position adjacent to the current video block.

[0138] Item 6. The method of Item 5, wherein the position adjacent to the current video block comprises an upper left position adjacent to the current video block.

[0139] Item 7. A method according to any one of Items 1 to 6, wherein determining the intra-frame prediction of the sample point at the first position includes: determining a value of the metric based on the first position; and determining the intra-frame prediction of the sample point at the first position by applying at least one of the following to the value: an integer operation, a shift operation, or a limiting operation.

[0140] Item 8. The method of Item 7, wherein the intra prediction of the sample at the first position is determined by using RightShift(f(x,y),s), wherein (x,y) represents the coordinates of the first position, f() represents the metric, f(x,y) represents the value of the metric, s represents a predefined integer, RightShift(f(x,y),s) represents the shift operation, and wherein the shift operation RightShift(f(x,y),s) is defined as: Where offset0 and offset1 are predefined values.

[0141] Item 9. The method of Item 7, wherein the intra prediction of the sample at the first position is determined by using Clip3(RightShift(f(x,y),s)), wherein (x,y) represents the coordinates of the first position, f() represents the metric, f(x,y) represents the value of the metric, s represents a predefined integer, RightShift(f(x,y),s) represents the shift operation, Clip3(RightShift(f(x,y),s)) represents the clipping operation, wherein the shift operation RightShift(f(x,y),s) is defined as: Wherein offset0 and offset1 are predefined values; and wherein the clipping operation Clip3(z) is defined as: Clip3(z)=max(minV,min(maxV,z)), wherein minV and maxV represent the minimum and maximum values ​​of the samples of the current video block, and z represents the value determined by the shift operation RightShift(f(x,y),s).

[0142] Item 10. A method according to any one of Items 1 to 9, wherein at least one parameter of the metric comprises at least one predefined value.

[0143] Item 11. A method according to any one of Items 1 to 9, wherein determining the metric comprises: determining at least one parameter of the metric based on at least one reference sample of the current video block.

[0144] Item 12. A method according to Item 11, wherein the at least one parameter is determined based on the at least one reference sample point by using at least one of the following: least mean square (LMS), regression-based mean square error (MSE) minimization, or linear regression.

[0145] Item 13. A method according to Item 11 or Item 12, wherein the at least one reference sample includes at least one neighboring reconstructed sample of the current video block, and determining the at least one parameter of the metric includes: determining at least one reconstructed sample value based on at least one coordinate of the at least one neighboring reconstructed sample; and determining the at least one parameter based on at least one difference between the at least one reconstructed sample value and at least one predicted sample value determined according to the metric.

[0146] Clause 14. The method of clause 11, wherein the at least one parameter of the metric is determined by using a parameter derivation tool used for a codec tool.

[0147] Item 15. A method according to Item 14, wherein the encoding and decoding tool includes a convolutional cross-component model.

[0148] Item 16. The method according to any one of Items 11 to 15, wherein the at least one reference sample comprises at least one of: a first neighboring sample adjacent to the current video block, or a second neighboring sample non-adjacent to the current video block.

[0149] Item 17. A method according to any one of Items 11 to 16, wherein the at least one reference sample comprises at least one of the following: at least one reconstructed sample row above the current video block, or at least one reconstructed sample column to the left of the current video block.

[0150] Item 18. A method according to any one of Items 11 to 16, wherein the at least one reference sample of the current video block is determined based on at least one of: whether the reference sample is available, whether the reference sample is reconstructed, the size of the current video block, or the position of the current video block.

[0151] Item 19. A method according to any one of items 1 to 18, wherein information about the application of the method is included in at least one syntax element in the bitstream.

[0152] Item 20. A method according to Item 19, wherein the at least one syntax element comprises at least one of the following: a first syntax element indicating whether to apply the method, or a second syntax element indicating how to apply the method.

[0153] Item 21. The method of Item 20, wherein the first syntax element is included in the bitstream if the current video block is intra-coded.

[0154] Item 22. The method of item 20 or item 21, wherein if the method is applied to the current video block, the second syntax element is included in the bitstream.

[0155] Item 23. A method according to any one of Items 19 to 22, wherein the at least one syntax element is included in at least one of the following: a sequence parameter set (SPS), a picture parameter set (PPS), a sequence header, a picture header, a slice header, a codec tree unit (CTU), a codec unit (CU), a transform unit (TU), or a prediction unit (PU).

[0156] Item 24. A method according to any one of Items 19 to 22, wherein the at least one syntax element is included in at least one of the following: sequence level, picture group level, picture level, slice level, slice group level, sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), decoding parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptation parameter set (APS), slice header, or slice group header.

[0157] Item 25. A method according to any one of Items 19 to 22, wherein the at least one syntax element is included in a region containing more than one sample or pixel.

[0158] Item 26. A method according to item 25, wherein the area includes one of the following: a prediction block (PB), a transform block (TB), a codec block (CB), a prediction unit (PU), a transform unit (TU), a codec unit (CU), a virtual pipeline data unit (VPDU), a codec tree unit (CTU), a CTU row, a slice, a slice or a sub-picture.

[0159] Item 27. A method according to any one of Items 19 to 26, wherein the at least one syntax element is predictively encoded.

[0160] Item 28. A method according to any one of items 19 to 27, wherein the at least one syntax element is encoded and decoded using at least one arithmetic context model.

[0161] Item 29. The method of Item 28, wherein the at least one arithmetic context model is based on at least one neighboring block of the current video block.

[0162] Item 30. A method according to any one of items 19 to 29, wherein the at least one syntax element is bypass coded.

[0163] Item 31. A method according to any one of Items 19 to 30, wherein the at least one syntax element is encoded and decoded in a hierarchical manner.

[0164] Item 32. A method according to any one of Items 19 to 31, wherein the information is based on codec information of the current video block.

[0165] Item 33. A method according to item 32, wherein the codec information includes at least one of the following: block size, color format, single tree or dual tree partitioning, color component, slice type, or picture type.

[0166] Item 34. A method according to Item 32 or Item 33, wherein the information is based on whether at least one of the width of the current video block or the height of the current video block satisfies at least one condition.

[0167] Item 35. A method according to Item 34, wherein the at least one condition includes at least one of the following: the width of the current video block is less than a first threshold, the width of the current video block is less than or equal to the first threshold, the width of the current video block is greater than a second threshold, the width of the current video block is greater than or equal to the second threshold, the height of the current video block is less than a third threshold, the height of the current video block is less than or equal to the third threshold, the height of the current video block is greater than a fourth threshold, the height of the current video block is greater than or equal to the fourth threshold, a first ratio of the width to the height is less than a fifth threshold, the first ratio is less than or equal to the fifth threshold, the first ratio is greater than a sixth threshold, the first ratio is greater than or equal to the sixth threshold, a second ratio of the height to the width is less than a seventh threshold, the second ratio is less than or equal to the seventh threshold, the second ratio is greater than an eighth threshold, the second ratio is greater than or equal to the eighth threshold, the width is equal to the first value, or the height is equal to the second value.

[0168] Item 36. A method according to Item 35, wherein: the first threshold includes one of the following: 8, 16, 32 or 64, the second threshold includes one of the following: 2, 4 or 8, the third threshold includes one of the following: 8, 16, 32 or 64, the fourth threshold includes one of the following: 2, 4 or 8, the fifth threshold includes one of the following: 1 / 8, 1 / 4, 1 / 2, 1, 2, 4, 8 or 16, the sixth threshold includes one of the following: 1 / 8, 1 / 4, 1 / 2, 1, 2, 4, 8 or 16, the seventh threshold includes one of the following: 1 / 8, 1 / 4, 1 / 2, 1, 2, 4, 8 or 16, the eighth threshold includes one of the following: 1 / 8, 1 / 4, 1 / 2, 1, 2, 4, 8 or 16, the first value includes one of the following: 8, 16, 32 or 64, and / or the second value includes one of the following: 8, 16, 32 or 64.

[0169] Clause 37. A method according to any one of clauses 1 to 31, wherein at least one parameter of the metric is included in a third syntax element in the bitstream.

[0170] Item 38. The method of Item 37, wherein the third syntax element is included in the bitstream if the metric is applied to determine intra-prediction for the current video block.

[0171] Item 39. A method according to Item 37 or Item 38, wherein the third element is encoded and decoded using at least one arithmetic context model.

[0172] Item 40. A method according to any one of items 37 to 39, wherein the third syntax element is bypass coded.

[0173] Item 41. A method according to any one of Items 37 to 40, wherein the third syntax element is predictively encoded.

[0174] Item 42. A method according to any one of items 1 to 41, wherein the method is performed as an intra-frame prediction mode, the intra-frame prediction mode is predictive coding.

[0175] Item 43. A method according to item 42, wherein if a syntax element of a first level associated with the intra-frame prediction mode is to be disabled, a syntax element of a second level lower than the first level associated with the intra-frame prediction mode is not included in the bitstream.

[0176] Item 44. A method according to any one of Items 1 to 43, wherein the converting includes encoding the current video block into the bitstream.

[0177] Item 45. A method according to any one of Items 1 to 43, wherein the converting comprises decoding the current video block from the bitstream.

[0178] Item 46. An apparatus for video processing, comprising a processor and a non-volatile memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform a method according to any one of Items 1 to 45.

[0179] Item 47. A non-transitory computer-readable storage medium storing instructions for causing a processor to perform a method according to any one of Items 1 to 45.

[0180] Item 48. A non-transitory computer-readable recording medium storing a bitstream of a video generated by a method performed by an apparatus for video processing, wherein the method comprises: determining a metric for intra-frame prediction for a current video block of the video; determining, based on the metric and a first position, an intra-frame prediction of a sample at the first position in the current video block; and generating the bitstream based on the intra-frame prediction.

[0181] Item 49. A method for storing a bitstream of a video, comprising: determining a metric for intra-frame prediction for a current video block of the video; determining, based on the metric and a first position, an intra-frame prediction for a sample at the first position in the current video block; generating the bitstream based on the intra-frame prediction; and storing the bitstream in a non-transitory computer-readable recording medium. Example Device

[0182] Fig.31 A block diagram of a computing device 3100 in which various embodiments of the present disclosure may be implemented is shown. The computing device 3100 may be implemented as a source device 110 (or video encoder 114 or 200) or a destination device 120 (or video decoder 124 or 300), or may be included in a source device 110 (or video encoder 114 or 200) or a destination device 120 (or video decoder 124 or 300).

[0183] It should be understood that Fig.31 The computing device 3100 shown in the figure is for illustrative purposes only and is not intended to in any way imply any limitation on the functionality and scope of the embodiments of the present disclosure.

[0184] like Fig.31 As shown, computing device 3100 includes a general computing device 3100. Computing device 3100 may include at least one or more processors or processing units 3110, memory 3120, storage unit 3130, one or more communication units 3140, one or more input devices 3150, and one or more output devices 3160.

[0185] In some embodiments, the computing device 3100 can be implemented as any user terminal or server terminal with computing power. The server terminal can be a server, a large computing device, etc. provided by a service provider. The user terminal can be, for example, any type of mobile terminal, fixed terminal, or portable terminal, including a mobile phone, a station, a unit, a device, a multimedia computer, a multimedia tablet computer, an Internet node, a communicator, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an e-book device, a gaming device, or any combination thereof, including accessories and peripherals of these devices, or any combination thereof. It is conceivable that the computing device 3100 can support any type of interface to the user (such as a "wearable" circuit device, etc.).

[0186] The processing unit 3110 may be a physical processor or a virtual processor and may implement various processes based on a program stored in the memory 3120. In a multi-processor system, multiple processing units execute computer executable instructions in parallel to improve the parallel processing capability of the computing device 3100. The processing unit 3110 may also be referred to as a central processing unit (CPU), a microprocessor, a controller, or a microcontroller.

[0187] The computing device 3100 typically includes various computer storage media. Such media can be any media accessible by the computing device 3100, including but not limited to volatile media and non-volatile media, or removable media and non-removable media. The memory 3120 can be a volatile memory (e.g., a register, a cache, a random access memory (RAM)), a non-volatile memory (such as a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM) or flash memory) or any combination thereof. The storage unit 3130 can be any removable or non-removable medium, and can include machine-readable media, such as a memory, a flash drive, a disk or other media that can be used to store information and / or data and can be accessed in the computing device 3100.

[0188] The computing device 3100 may also include additional removable / non-removable storage media, volatile / non-volatile storage media. Fig.31 Although not shown in the figure, a disk drive for reading from and / or writing to a removable nonvolatile disk and an optical drive for reading from and / or writing to a removable nonvolatile optical disk may be provided. In this case, each drive may be connected to the bus (not shown) via one or more data medium interfaces.

[0189] The communication unit 3140 communicates with another computing device via a communication medium. In addition, the functions of the components in the computing device 3100 can be implemented by a single computing cluster or multiple computing machines, which can communicate via a communication connection. Therefore, the computing device 3100 can operate in a networked environment using logical connections to one or more other servers, networked personal computers (PCs), or other general network nodes.

[0190] The input device 3150 may be one or more of various input devices, such as a mouse, keyboard, trackball, voice input device, and the like. The output device 3160 may be one or more of various output devices, such as a display, a speaker, a printer, and the like. With the aid of the communication unit 3140, the computing device 3100 may also communicate with one or more external devices (not shown), such as storage devices and display devices, and the computing device 3100 may also communicate with one or more devices that enable a user to interact with the computing device 3100, or, if necessary, the computing device 3100 may also communicate with any device (e.g., a network card, a modem, etc.) that enables the computing device 3100 to communicate with one or more other computing devices. Such communication may be performed via an input / output (I / O) interface (not shown).

[0191] In some embodiments, some or all components of the computing device 3100 may also be arranged in a cloud computing architecture rather than being integrated in a single device. In a cloud computing architecture, components may be provided remotely and work together to implement the functions described in the present disclosure. In some embodiments, cloud computing provides computing, software, data access and storage services, which will not require the end user to know the physical location or configuration of the system or hardware that provides these services. In various embodiments, cloud computing provides services via a wide area network (such as the Internet) using a suitable protocol. For example, a cloud computing provider provides an application via a wide area network, which can be accessed through a web browser or any other computing component. The software or components of the cloud computing architecture and the corresponding data may be stored on a server at a remote location. The computing resources in a cloud computing environment may be merged or distributed at the location of a remote data center. Cloud computing infrastructure can provide services through a shared data center, although they appear as a single access point to the user. Therefore, a cloud computing architecture may be used to provide components and functions described herein from a service provider at a remote location. Alternatively, the components and functions described herein may be provided by a conventional server, or may be installed on a client device directly or otherwise.

[0192] In an embodiment of the present disclosure, the computing device 3100 may be used to implement video encoding / decoding. The memory 3120 may include one or more video encoding / decoding modules 3125 having one or more program instructions. These modules are accessible and executable by the processing unit 3110 to perform the functions of the various embodiments described herein.

[0193] In an example embodiment performing video encoding, input device 3150 may receive video data as input 3170 to be encoded. The video data may be processed, for example, by video codec module 3125 to generate an encoded bitstream. The encoded bitstream may be provided as output 3180 via output device 3160.

[0194] In an example embodiment performing video decoding, input device 3150 may receive an encoded bitstream as input 3170. The encoded bitstream may be processed, for example, by video codec module 3125 to generate decoded video data. The decoded video data may be provided as output 3180 via output device 3160.

[0195] Although the present disclosure has been specifically shown and described with reference to the preferred embodiments of the present disclosure, it will be appreciated by those skilled in the art that various changes may be made in form and detail without departing from the spirit and scope of the present application as defined by the appended claims. These modifications are intended to be encompassed by the scope of the present application. Therefore, the foregoing description of the embodiments of the present application is not intended to be limiting.

Claims

1. A method for video processing, comprising: For conversion between a current video block of a video and a bitstream of the video, determining a metric for intra prediction for the current video block; Determining, based on the metric and the first position, an intra prediction for a sample at the first position in the current video block; as well as The conversion is performed based on the intra prediction.

2. The method of claim 1 , wherein the metric comprises a surface metric, the surface metric comprising at least one of: Linear metrics, Secondary metric, Polynomial metric, or Non-polynomial metrics.

3. The method of claim 2, wherein the non-polynomial metric comprises at least one of the following: Trigonometric metrics, Inverse trigonometric metrics, Hyperbolic trigonometric metrics, Inverse hyperbolic trigonometric metrics, Index metrics, Logarithmic metric, or Combination of non-polynomial and polynomial metrics.

4. The method of claim 2 or claim 3, wherein the metric comprises one of the following: ax+by+c, ax 2 +by 2 +cxy + dx + ey + f, or Wherein (x, y) represents the coordinates of the sample point in the current video block, and a, b, c, d, e, f and M are parameters of the metric.

5. The method according to any one of claims 1 to 4, wherein the first position of the sample point is a relative position relative to a second position, and the second position comprises one of the following: the upper left corner position of the current video block, the upper right corner position of the current video block, the lower left corner position of the current video block, the lower right corner position of the current video block, the center position of the current video block, or The position of the adjacent current video block. The method of claim 5 , wherein the position adjacent to the current video block comprises an upper left position adjacent to the current video block.

7. The method according to any one of claims 1 to 6, wherein determining the intra prediction of the sample at the first position comprises: determining a value of the metric based on the first position; as well as The intra prediction of the sample at the first position is determined by applying at least one of the following to the value: Integer operations, Shift operations, or Limiting operation.

8. The method according to claim 7, wherein the intra prediction of the sample at the first position is determined by using RightShift(f(x,y),s), wherein (x,y) represents the coordinates of the first position, f() represents the metric, f(x,y) represents the value of the metric, s represents a predefined integer, RightShift(f(x,y),s) represents the shift operation, and wherein the shift operation RightShift(f(x,y),s) is defined as: Where offset0 and offset1 are predefined values.

9. The method according to claim 7, wherein the intra prediction of the sample at the first position is determined by using Clip3(RightShift(f(x,y),s)), wherein (x,y) represents the coordinates of the first position, f() represents the metric, f(x,y) represents the value of the metric, s represents a predefined integer, RightShift(f(x,y),s) represents the shift operation, Clip3(RightShift(f(x,y),s)) represents the clipping operation, wherein the shift operation RightShift(f(x,y),s) is defined as: Where offset0 and offset1 are predefined values; and The clipping operation Clip3(z) is defined as: Clip3(z)=max(minV,min(maxV,z)), where minV and maxV represent the minimum and maximum values ​​of the samples of the current video block, and z represents the value determined by the shift operation RightShift(f(x,y),s).

10. The method according to any one of claims 1 to 9, wherein at least one parameter of the metric comprises at least one predefined value.

11. The method according to any one of claims 1 to 9, wherein determining the metric comprises: At least one parameter of the metric is determined based on at least one reference sample of the current video block.

12. The method of claim 11, wherein the at least one parameter is determined based on the at least one reference sample by using at least one of: Least Mean Square (LMS), Minimization of mean squared error (MSE) based on regression, or Linear regression.

13. The method of claim 11 or claim 12, wherein the at least one reference sample comprises at least one neighboring reconstructed sample of the current video block, and determining the at least one parameter of the metric comprises: determining at least one reconstruction sample value based on at least one coordinate of the at least one neighboring reconstruction sample; as well as The at least one parameter is determined based on at least one difference between the at least one reconstructed sample value and at least one predicted sample value determined according to the metric.

14. The method of claim 11, wherein the at least one parameter of the metric is determined by using a parameter derivation tool used for a codec tool.

15. The method of claim 14, wherein the encoding and decoding tool comprises a convolutional cross-component model.

16. The method according to any one of claims 11 to 15, wherein the at least one reference sample point comprises at least one of the following: The first neighboring sample adjacent to the current video block, or A second neighboring sample point that is non-adjacent to the current video block.

17. The method according to any one of claims 11 to 16, wherein the at least one reference sample point comprises at least one of the following: at least one reconstructed sample row above the current video block, or At least one reconstructed sample point column on the left side of the current video block.

18. The method according to any one of claims 11 to 16, wherein the at least one reference sample of the current video block is determined based on at least one of the following: Are reference points available? Whether the reference sample is reconstructed, The size of the current video block, or The position of the current video block.

19. The method according to any one of claims 1 to 18, wherein information about the application of the method is included in at least one syntax element in the bitstream.

20. The method of claim 19, wherein the at least one syntax element comprises at least one of: a first syntax element indicating whether to apply the method, or A second syntax element indicating how to apply the method.

21. The method of claim 20, wherein the first syntax element is included in the bitstream if the current video block is intra-coded.

22. The method of claim 20 or claim 21, wherein the second syntax element is included in the bitstream if the method is applied to the current video block.

23. The method according to any one of claims 19 to 22, wherein the at least one syntax element is included in at least one of: Sequence Parameter Set (SPS), Picture Parameter Set (PPS), Sequence header, Picture header, Strip head, Codec Tree Unit (CTU), Codec Unit (CU), Transform Unit (TU), or Prediction Unit (PU).

24. The method according to any one of claims 19 to 22, wherein the at least one syntax element is included in at least one of: Sequence level, Picture group level, Picture level, Stripe level, Film group level, Sequence header, Picture header, Sequence Parameter Set (SPS), Video Parameter Set (VPS), Decoding Parameter Set (DPS), Decoding Capability Information (DCI), Picture Parameter Set (PPS), Adaptive Parameter Set (APS), Strip header, or Film group header.

25. The method according to any one of claims 19 to 22, wherein the at least one syntax element is included in a region containing more than one sample or pixel.

26. The method of claim 25, wherein the region comprises one of a prediction block (PB), a transform block (TB), a codec block (CB), a prediction unit (PU), a transform unit (TU), a codec unit (CU), a virtual pipeline data unit (VPDU), a codec tree unit (CTU), a CTU row, a slice, a slice, or a sub-picture.

27. The method according to any one of claims 19 to 26, wherein the at least one syntax element is predictively coded.

28. The method according to any one of claims 19 to 27, wherein the at least one syntax element is encoded and decoded using at least one arithmetic context model.

29. The method of claim 28, wherein the at least one arithmetic context model is based on at least one neighboring block of the current video block.

30. The method according to any one of claims 19 to 29, wherein the at least one syntax element is bypass coded.

31. The method according to any one of claims 19 to 30, wherein the at least one syntax element is encoded and decoded in a hierarchical manner.

32. The method of any one of claims 19 to 31, wherein the information is based on codec information of the current video block.

33. The method of claim 32, wherein the codec information comprises at least one of: block size, color format, single tree or dual tree partitioning, color component, slice type, or picture type.

34. The method of claim 32 or claim 33, wherein the information is based on whether at least one of a width of the current video block or a height of the current video block satisfies at least one condition.

35. The method of claim 34, wherein the at least one condition comprises at least one of the following: The width of the current video block is smaller than a first threshold, The width of the current video block is less than or equal to the first threshold, The width of the current video block is greater than a second threshold, The width of the current video block is greater than or equal to the second threshold, The height of the current video block is less than a third threshold, The height of the current video block is less than or equal to the third threshold, The height of the current video block is greater than a fourth threshold, The height of the current video block is greater than or equal to the fourth threshold, a first ratio of the width to the height is less than a fifth threshold, the first ratio is less than or equal to the fifth threshold, the first ratio is greater than a sixth threshold, the first ratio is greater than or equal to the sixth threshold, a second ratio of the height to the width is less than a seventh threshold, the second ratio is less than or equal to the seventh threshold, the second ratio is greater than an eighth threshold, the second ratio is greater than or equal to the eighth threshold, The width is equal to the first value, or The height is equal to a second value.

36. The method of claim 35, wherein: The first threshold value includes one of the following: 8, 16, 32 or 64, The second threshold value includes one of the following: 2, 4 or 8, The third threshold value includes one of the following: 8, 16, 32 or 64, The fourth threshold includes one of the following: 2, 4 or 8, The fifth threshold value includes one of the following: 1 / 8, 1 / 4, 1 / 2, 1, 2, 4, 8 or 16, The sixth threshold value includes one of the following: 1 / 8, 1 / 4, 1 / 2, 1, 2, 4, 8 or 16, The seventh threshold value includes one of the following: 1 / 8, 1 / 4, 1 / 2, 1, 2, 4, 8 or 16, The eighth threshold value includes one of the following: 1 / 8, 1 / 4, 1 / 2, 1, 2, 4, 8 or 16, The first value includes one of the following: 8, 16, 32 or 64, and / or The second value includes one of the following: 8, 16, 32 or 64.

37. The method according to any one of claims 1 to 31, wherein at least one parameter of the metric is included in a third syntax element in the bitstream.

38. The method of claim 37, wherein the third syntax element is included in the bitstream if the metric is applied to determine intra-prediction for the current video block.

39. A method according to claim 37 or claim 38, wherein the third element is encoded and decoded using at least one arithmetic context model.

40. The method according to any one of claims 37 to 39, wherein the third syntax element is bypass coded.

41. The method according to any one of claims 37 to 40, wherein the third syntax element is predictively coded.

42. The method according to any one of claims 1 to 41, wherein the method is performed as an intra prediction mode, the intra prediction mode being a predictive codec.

43. The method of claim 42, wherein if syntax elements of a first level associated with the intra-prediction mode are to be disabled, syntax elements of a second level lower than the first level associated with the intra-prediction mode are not included in the bitstream.

44. The method of any one of claims 1 to 43, wherein the converting comprises encoding the current video block into the bitstream.

45. The method of any one of claims 1 to 43, wherein the converting comprises decoding the current video block from the bitstream.

46. ​​An apparatus for video processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method of any one of claims 1 to 45.

47. A non-transitory computer-readable storage medium storing instructions for causing a processor to execute the method according to any one of claims 1 to 45.

48. A non-transitory computer-readable recording medium storing a bitstream of a video generated by a method performed by an apparatus for video processing, wherein the method comprises: determining a metric for intra prediction for a current video block of the video; Determining, based on the metric and the first position, an intra prediction for a sample at the first position in the current video block; as well as The bitstream is generated based on the intra prediction.

49. A method for storing a bitstream of a video, comprising: determining a metric for intra prediction for a current video block of the video; Determining, based on the metric and the first position, an intra prediction for a sample at the first position in the current video block; generating the bitstream based on the intra prediction; as well as The bit stream is stored in a non-transitory computer-readable recording medium.