Apparatus and method for image encoding and decoding using chrominance prediction mode based on reference vector

WO2026206046A1PCT designated stage Publication Date: 2026-10-01HYUNDAI MOTOR CO LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2026/004897
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-27
Filing Date
2026-03-27
Publication Date
2026-10-01

Smart Images

  • Figure KR2026004897_01102026_PF_FP_ABST
    Figure KR2026004897_01102026_PF_FP_ABST
Patent Text Reader

Abstract

Provided is an image decoding method. The method comprises the steps of: determining prediction information of partitions of the current block by decoding a bitstream; generating prediction blocks of the partitions by using the prediction information of the partitions; deriving a weight model on the basis of regression using reference samples of the current block and reference samples of the partitions; and generating a prediction block of the current block by blending the prediction blocks of the partitions by using the weight model.
Need to check novelty before this filing date? Find Prior Art

Description

Device and method for image encoding and decoding using a reference vector-based chrominance prediction mode

[0001] The present disclosure relates to image encoding and decoding using a reference vector-based chrominance prediction mode.

[0002] The following description merely provides background information related to the present embodiment and does not constitute prior art.

[0003] Because video data contains a large amount of data compared to audio or still image data, storing or transmitting it as is without compression processing requires significant hardware resources, including memory.

[0004] Therefore, typically when storing or transmitting video data, the encoder compresses the video data for storage or transmission, and the decoder receives the compressed video data, decompresses it, and plays it. Such video compression technologies include H.264 / AVC, HEVC (High Efficiency Video Coding), and VVC (Versatile Video Coding), which improves coding efficiency by more than 30% compared to HEVC.

[0005] However, as video size, resolution, and frame rates are gradually increasing, and the amount of data that needs to be encoded is also growing accordingly, a new compression technology is required that offers better encoding efficiency and higher image quality improvement effects than existing compression technologies.

[0006] The present disclosure provides an image encoding and decoding method and apparatus for predicting a current color difference block from a reference color difference block using a prediction model in a reference vector-based color difference prediction mode.

[0007] In addition, the present invention provides a method for transmitting or storing a bitstream generated by an image encoding method.

[0008] In addition, the present invention provides a recording medium that stores a bitstream generated by an image encoding method.

[0009] One aspect of the present disclosure provides an image decoding method. The method comprises: a process of deriving a reference vector of a current chrominance block by scaling a reference vector of a luminance region corresponding to a current chrominance block based on a chrominance subsampling format; and a process of predicting a current chrominance block using the reference vector of the current chrominance block.

[0010] Another aspect of the present disclosure provides an image encoding method. The method comprises: a process of deriving a reference vector of a current chrominance block by scaling a reference vector of a luminance region corresponding to a current chrominance block based on a chrominance subsampling format; and a process of predicting a current chrominance block using the reference vector of the current chrominance block.

[0011] Another aspect of the present disclosure provides a method for providing a bitstream generated by an image encoding method to an image decoding device.

[0012] Another aspect of the present disclosure provides a computer-readable recording medium for storing a bitstream generated by an image encoding method.

[0013] FIG. 1 is an exemplary block diagram of an image encoding device capable of implementing the technologies of the present disclosure.

[0014] Figure 2 is a diagram illustrating a method for dividing blocks using a QTBTTT (QuadTree plus BinaryTree TernaryTree) structure.

[0015] FIGS. 3a and 3b are diagrams showing a plurality of intra prediction modes including wide-angle intra prediction modes.

[0016] Figure 4 is an example of the surrounding blocks of the current block.

[0017] FIG. 5 is an exemplary block diagram of an image decoding device capable of implementing the technologies of the present disclosure.

[0018] FIG. 6 is a drawing showing reference regions according to one embodiment of the present disclosure.

[0019] FIG. 7 is a drawing for illustrating a TIMD mode according to one embodiment of the present disclosure.

[0020] FIG. 8 is a flowchart of a reference vector-based color difference prediction method according to one embodiment of the present disclosure.

[0021] FIG. 9 is a method for deriving a color difference reference vector according to one embodiment of the present disclosure.

[0022] This is a flowchart.

[0023] FIG. 10 is an illustrative diagram for explaining a method for deriving a reference vector of a color difference block in a dual tree structure according to one embodiment of the present disclosure.

[0024] FIGS. 11a, FIGS. 11b, FIGS. 11c and FIGS. 11d are drawings illustrating various template regions according to one embodiment of the present disclosure.

[0025] FIG. 12 is a diagram illustrating a model-based prediction method for color difference blocks according to one embodiment of the present disclosure.

[0026] FIG. 13 is a flowchart for color difference block prediction according to one embodiment of the present disclosure.

[0027] FIG. 14 is a flowchart for determining a method for predicting a color difference block according to one embodiment of the present disclosure.

[0028] FIG. 15a is a flowchart for determining a method for predicting a color difference block in a dual tree structure according to one embodiment of the present disclosure. In FIG. 15a, a convolution model based

[0029] FIG. 15b is a flowchart for determining a method for predicting a color difference block in a dual tree structure according to one embodiment of the present disclosure. In FIG. 15b, a convolution model based

[0030] FIG. 16 is a drawing for illustrating sampling of a template area according to one embodiment of the present disclosure.

[0031] FIG. 17 is a diagram illustrating the derivation of a linear model based on statistical values ​​according to one embodiment of the present disclosure.

[0032] FIG. 18 is a diagram showing the positional relationship of input samples of a convolution model according to one embodiment of the present disclosure.

[0033] FIGS. 19a and FIGS. 19b are drawings for explaining the application of a convolution model according to one embodiment of the present disclosure.

[0034] Some embodiments of the present disclosure will be described in detail below with reference to the exemplary drawings. It should be noted that in assigning reference numerals to the components of each drawing, the same components are given the same reference numeral whenever possible, even if they are shown in different drawings. Furthermore, in describing the embodiments, if it is determined that a detailed description of related known components or functions could obscure the essence of the embodiments, such detailed description is omitted.

[0035] FIG. 1 is an exemplary block diagram of an image encoding device capable of implementing the technologies of the present disclosure. Hereinafter, the image encoding device and its sub-components will be described with reference to FIG. 1.

[0036] The video encoding device may be configured to include a picture splitting unit (110), a prediction unit (120), a subtractor (130), a conversion unit (140), a quantization unit (145), a reordering unit (150), an entropy encoding unit (155), an inverse quantization unit (160), an inverse conversion unit (165), an adder (170), a loop filter unit (180), and a memory (190).

[0037] Each component of the video encoding device may be implemented in hardware or software, or as a combination of hardware and software. Additionally, the function of each component may be implemented in software, and a microprocessor may be implemented to execute the software function corresponding to each component.

[0038] A single image (video) consists of one or more sequences containing multiple pictures. Each picture is divided into multiple regions, and encoding is performed for each region. For example, a single picture is divided into one or more tiles and / or slices. Here, one or more tiles can be defined as a tile group. Each tile or slice is divided into one or more Coding Tree Units (CTUs). And each CTU is divided into one or more Coding Units (CUs) by a tree structure. Information applicable to each CU is encoded as the syntax of the CU, and information applicable to all CUs included in a single CTU is encoded as the syntax of the CTU. Additionally, information applicable to all blocks within a single slice is encoded as the syntax of the slice header, and information applicable to all blocks constituting one or more pictures is encoded in the Picture Parameter Set (PPS) or the picture header. Furthermore, information commonly referenced by multiple pictures is encoded in a Sequence Parameter Set (SPS). Also, information commonly referenced by one or more SPSs is encoded in a Video Parameter Set (VPS). Additionally, information commonly applicable to a single tile or tile group may be encoded as the syntax of a tile or tile group header. The syntax included in the SPS, PPS, slice header, and tile or tile group header may be referred to as high-level syntax.

[0039] The picture splitting unit (110) determines the size of the CTU. Information regarding the size of the CTU (CTU size) is encoded as a syntax of SPS or PPS and transmitted to an image decoding device.

[0040] The picture division unit (110) divides each picture constituting the image into multiple CTUs having a predetermined size, and then recursively divides the CTUs using a tree structure. The leaf nodes in the tree structure become the CUs, which are the basic units of encoding.

[0041] The tree structure may be a QuadTree (QT) in which an upper node (or parent node) is divided into four lower nodes (or child nodes) of equal size, a BinaryTree (BT) in which an upper node is divided into two lower nodes, a TernaryTree (TT) in which an upper node is divided into three lower nodes in a 1:2:1 ratio, or a structure that combines two or more of these QT, BT, and TT structures. For example, a QTBT (QuadTree plus BinaryTree) structure may be used, or a QTBTTT (QuadTree plus BinaryTree TernaryTree) structure may be used. Here, BTTT combined may be referred to as an MTT (Multiple-Type Tree).

[0042] Figure 2 is a diagram illustrating a method for dividing blocks using a QTBTTT structure.

[0043] As illustrated in FIG. 2, the CTU can first be split into a QT structure. Quadtree splitting can be repeated until the size of the splitting block reaches the minimum block size of the leaf node allowed in QT (MinQTSize). A first flag (QT_split_flag) indicating whether each node of the QT structure is split into four nodes of the lower layer is encoded by the entropy encoder (155) and signaled to the image decoder. If the leaf node of the QT is not larger than the maximum block size of the root node allowed in BT (MaxBTSize), it can be further split into one or more of the BT structure or TT structure. In the BT structure and / or TT structure, multiple splitting directions may exist. For example, there may be two directions in which the block of the corresponding node is split horizontally and vertically. As shown in Figure 2, when MTT splitting begins, a second flag (mtt_split_flag) indicating whether the nodes have been split, and if splitting has occurred, a flag indicating the splitting direction (vertical or horizontal) and / or the splitting type (binary or ternary) are encoded by the entropy encoding unit (155) and signaled to the image decoding device.

[0044] Alternatively, prior to encoding the first flag (QT_split_flag) indicating whether each node is split into four nodes of the lower layer, the CU split flag (split_cu_flag) indicating whether the node is split may be encoded. If the value of the CU split flag (split_cu_flag) indicates that it is not split, the block of the corresponding node becomes a leaf node in the split tree structure and becomes a coding unit (CU), which is the basic unit of encoding. If the value of the CU split flag (split_cu_flag) indicates that it is split, the video encoding device starts encoding from the first flag in the manner described above.

[0045] When QTBT is used as another example of a tree structure, there may be two types: a type that divides the block of the corresponding node horizontally into two blocks of the same size (i.e., symmetric horizontal splitting) and a type that divides it vertically (i.e., symmetric vertical splitting). A splitting flag (split_flag) indicating whether each node of the BT structure is split into a block of a lower layer and splitting type information indicating the type of splitting are encoded by the entropy encoding unit (155) and transmitted to the image decoding device. Meanwhile, there may also be an additional type that divides the block of the corresponding node into two blocks of an asymmetric shape. The asymmetric shape may include a shape that divides the block of the corresponding node into two rectangular blocks with a size ratio of 1:3, or a shape that divides the block of the corresponding node diagonally.

[0046] A CU can have various sizes depending on the QTBT or QTBTTT partitioning from a CTU. Hereinafter, the block corresponding to the CU to be encoded or decoded (i.e., the leaf node of QTBTTT) is referred to as the 'current block'. Depending on the adoption of QTBTTT partitioning, the shape of the current block may be not only square but also rectangular.

[0047] The prediction unit (120) predicts the current block and generates a prediction block. The prediction unit (120) includes an intra prediction unit (122) and an inter prediction unit (124).

[0048] Generally, current blocks within a picture can each be predictively coded. Typically, the prediction of a current block can be performed using an intra-prediction technique (using data from the picture containing the current block) or an inter-prediction technique (using data from a picture coded prior to the picture containing the current block). Inter-prediction includes both unidirectional and bidirectional prediction.

[0049] The intra prediction unit (122) predicts samples within the current block using samples (reference samples) located around the current block within the current picture containing the current block. “Sample” means “pixel,” and the two terms may be used interchangeably in this specification.

[0050] There are multiple intra prediction modes depending on the prediction direction. For example, as shown in FIG. 3a, the multiple intra prediction modes may include two non-directional modes, including Planar mode and DC mode, and 65 directional modes. The surrounding pixels to be used and the calculation formula are defined differently depending on each prediction mode.

[0051] For efficient directional prediction for a rectangular current block, additional directional modes (intra-prediction modes 67 through 80 and -1 through -14) illustrated by dashed arrows in FIG. 3b may be used. These may be referred to as "wide angle intra-prediction modes." In FIG. 3b, the arrows indicate corresponding reference samples used for prediction and do not indicate the prediction direction. The prediction direction is opposite to the direction indicated by the arrows. Wide angle intra-prediction modes are modes that perform prediction in the opposite direction of a specific directional mode without additional bit transmission when the current block is rectangular. Among the wide angle intra-prediction modes, some wide angle intra-prediction modes available for the current block may be determined by the ratio of the width to the height of the rectangular current block. For example, wide-angle intra prediction modes with an angle less than 45 degrees (intra prediction modes 67 to 80) are available when the current block is a rectangular shape with a height less than the width, and wide-angle intra prediction modes with an angle greater than -135 degrees (intra prediction modes -1 to -14) are available when the current block is a rectangular shape with a width greater than the height.

[0052] The intra prediction unit (122) can determine the intra prediction mode to use for encoding the current block. In some examples, the intra prediction unit (122) may encode the current block using several intra prediction modes and select an appropriate intra prediction mode to use from the tested modes. For example, the intra prediction unit (122) may calculate rate-distortion values ​​using rate-distortion analysis of several tested intra prediction modes and select an intra prediction mode having the best rate-distortion features among the tested modes.

[0053] The intra prediction unit (122) selects one intra prediction mode among a plurality of intra prediction modes and predicts the current block using a calculation formula and surrounding pixels (reference pixels) determined according to the selected intra prediction mode. Information regarding the selected intra prediction mode is encoded by the entropy encoding unit (155) and transmitted to the image decoding device.

[0054] Meanwhile, the intra prediction unit (222) may configure some of the modes that are likely to be the intra prediction mode of the current block among the multiple intra prediction modes as a candidate list in order to efficiently encode intra prediction mode information indicating which of the multiple intra prediction modes is used as the intra prediction mode of the current block. This candidate list may be referred to as the most probable mode (MPM) list.

[0055] A candidate list can be derived based on the intra-prediction modes of the surrounding blocks of the current block. For example, all or part of the left block (L), top block (A), bottom-left block (BL), top-right block (AR), and top-left block (AL) of the current block can be used as surrounding blocks to construct the candidate list.

[0056] Mode information indicating whether the intra prediction mode of the current block is selected from the candidate list is generated and encoded by the entropy encoding unit (155). If the intra prediction mode of the current block is included in the candidate list, a first intra identification information is encoded to indicate which of the intra prediction modes in the candidate list was selected as the intra prediction mode of the current block. On the other hand, if the intra prediction mode of the current block is not included in the candidate list, a second intra identification information is encoded to indicate which of the remaining modes other than MPM was selected as the intra prediction mode of the current block.

[0057] The inter prediction unit (124) generates a prediction block for the current block using a motion compensation process. The inter prediction unit (124) searches for the block most similar to the current block within a reference picture that is encoded and decoded before the current picture, and generates a prediction block for the current block using the searched block. Then, it generates a motion vector (MV) corresponding to the displacement between the current block in the current picture and the prediction block in the reference picture. Generally, motion estimation is performed on the luminance component, and the motion vector calculated based on the luminance component is used for both the luminance component and the chroma component. Motion information including information about the reference picture used to predict the current block and information about the motion vector is encoded by the entropy encoding unit (155) and transmitted to the image decoding device.

[0058] The inter prediction unit (124) may perform interpolation on a reference picture or reference block to increase the accuracy of the prediction. That is, subsamples between two consecutive integer samples are interpolated by applying filter coefficients to a plurality of consecutive integer samples including those two integer samples. When the process of searching for the block most similar to the current block is performed for the interpolated reference picture, the motion vector can be expressed with precision in fractional units rather than precision in integer samples. The precision or resolution of the motion vector can be set differently for each unit of the target area to be encoded, such as slice, tile, CTU, CU, etc. When such Adaptive Motion Vector Resolution (AMVR) is applied, information regarding the motion vector resolution to be applied to each target area must be signaled for each target area. For example, if the target area is a CU, information regarding the motion vector resolution applied to each CU is signaled. The information regarding the motion vector resolution may be information indicating the precision of the difference motion vector described later.

[0059] Meanwhile, the inter prediction unit (124) can perform inter prediction using bi-prediction. In the case of bi-prediction, two reference pictures and two motion vectors representing the block location most similar to the current block within each reference picture are used. The inter prediction unit (124) selects a first reference picture and a second reference picture from the reference picture list 0 (RefPicList0) and reference picture list 1 (RefPicList1), respectively, and generates a first reference block and a second reference block by searching for a block similar to the current block within each reference picture. Then, it generates a prediction block for the current block by averaging or weighting the first reference block and the second reference block. Then, it transmits motion information containing information about the two reference pictures used to predict the current block and information about the two motion vectors to the entropy encoding unit (155). Here, reference picture list 0 consists of restored pictures that are prior to the current picture in the display order, and reference picture list 1 may consist of restored pictures that are prior to the current picture in the display order. However, this is not necessarily limited to this, and restored pictures prior to the current picture in the display order may be additionally included in reference picture list 0, and conversely, restored pictures prior to the current picture may be additionally included in reference picture list 1.

[0060] Various methods can be used to minimize the amount of bits required to encode motion information.

[0061] For example, if the reference picture and motion vector of the current block are identical to the reference picture and motion vector of a neighboring block, the motion information of the current block can be transmitted to an image decoder by encoding information that identifies the neighboring block. In other words, the motion information of the neighboring block is used as the motion information of the current block. This method is called 'merge mode'.

[0062] In merge mode, the inter prediction unit (124) selects a predetermined number of merge candidate blocks (hereinafter referred to as 'merge candidates') from the surrounding blocks of the current block.

[0063] As for the surrounding blocks to induce merge candidates, all or part of the left block (A0), bottom-left block (A1), top block (B0), top-right block (B1), and top-left block (B2) adjacent to the current block within the current picture, as shown in FIG. 4, may be used. Additionally or alternatively, blocks not adjacent to the current block may be used.

[0064] Additionally, blocks located within a reference picture (which may or may not be the same as the reference picture used to predict the current block) other than the current picture where the current block is located may be used as merge candidates. For example, blocks located at the same position as the current block within the reference picture (co-located blocks) or blocks adjacent to such co-located blocks may be additionally used as merge candidates.

[0065] If the number of merge candidates selected by the method described above is less than the preset number, history-based merge candidates, combined merge candidates, and zero merge candidates may be added to the list.

[0066] History-based merge candidates may be movement information within a list containing movement information of other blocks that were encoded / decoded prior to the encoding / decoding of the current block. The list of history-based merge candidates may be initialized on a CTU row basis within a predefined image region. Here, the predefined image region may be a picture, tile, or slice.

[0067] A combined merge candidate can be derived based on the statistical value of two or more pre-inserted merge candidates within the merge candidate list. For example, the statistical value may be an average.

[0068] Zero merge candidates can be zero-vector motion information. Zero-vector motion information refers to motion information that has a zero vector.

[0069] The inter prediction unit (124) constructs a merge list containing a predetermined number of merge candidates using these surrounding blocks. Among the merge candidates included in the merge list, it selects a merge candidate to be used as movement information for the current block and generates merge index information to identify the selected candidate. The generated merge index information is encoded by the entropy encoding unit (155) and transmitted to the image decoding device.

[0070] Merge skip mode is a special case of merge mode in which, after quantization, when all transform coefficients for entropy coding are close to zero, only neighbor block selection information (merge index) is transmitted without transmitting residual signals. Consequently, the predicted block generated by the merge index is restored as the current block. By utilizing skip mode, relatively high coding efficiency can be achieved in images with minimal motion, still images, and screen content.

[0071] Hereinafter, merge mode and skip mode will be collectively referred to as merge / skip mode.

[0072] Another method for encoding motion information is the AMVP (Advanced Motion Vector Prediction) mode.

[0073] In AMVP mode, the inter prediction unit (124) can construct a list of motion vector candidates in a manner similar to merge mode.

[0074] For example, candidate motion vectors for the motion vector of the current block are derived using the surrounding blocks of the current block. As for the surrounding blocks used to derive the candidate motion vectors for the predicted motion vectors, all or part of the left block (A0), bottom-left block (A1), top block (B0), top-right block (B1), and top-left block (B2) adjacent to the current block within the current picture shown in FIG. 4 may be used. Additionally or alternatively, blocks not adjacent to the current block may be used.

[0075] Additionally, blocks located within a reference picture (which may or may not be the same as the reference picture used to predict the current block) other than the current picture where the current block is located may be used as neighboring blocks to derive predicted motion vector candidates. For example, blocks located at the same position as the current block within the reference picture (co-located blocks) or blocks adjacent to such co-located blocks may be used.

[0076] In addition, history-based merge candidates can be included in the motion vector candidate list.

[0077] If the number of motion vector candidates is smaller than the preset number by the method described above, a 0 vector is added to the motion vector candidates.

[0078] The inter prediction unit (124) derives predicted motion vector candidates using the motion vectors of the surrounding blocks and determines a predicted motion vector for the current block's motion vector using the predicted motion vector candidates. Then, it calculates a difference motion vector by subtracting the predicted motion vector from the current block's motion vector.

[0079] Predicted motion vectors can be obtained by applying a predefined function (e.g., median, mean operation, etc.) to the predicted motion vector candidates. In this case, the image decoder is also aware of the predefined function. Furthermore, since the surrounding blocks used to derive the predicted motion vector candidates have already been encoded and decoded, the image decoder is also aware of the motion vectors of those surrounding blocks. Therefore, the image decoder does not need to encode information to identify the predicted motion vector candidates. Consequently, in this case, information regarding the difference motion vector and the reference picture used to predict the current block is encoded.

[0080] Meanwhile, the predicted motion vector may be determined by selecting one of the predicted motion vector candidates. In this case, information for identifying the selected predicted motion vector candidate is additionally encoded, along with information about the difference motion vector and information about the reference picture used to predict the current block.

[0081] The subtractor (130) generates a residual block by subtracting the prediction block generated by the intra prediction unit (122) or the inter prediction unit (124) from the current block.

[0082] The conversion unit (140) converts residual signals within a residual block having pixel values ​​in a spatial domain into conversion coefficients in the frequency domain. The conversion unit (140) can convert the residual signals within the residual block using the entire size of the residual block as the conversion unit, or it can divide the residual block into multiple sub-blocks and use the sub-blocks as the conversion unit to perform the conversion. Alternatively, it can divide the residual signals into two sub-blocks, a conversion area and a non-conversion area, and use only the conversion area sub-block as the conversion unit to convert the residual signals. Here, the conversion area sub-block may be one of two rectangular blocks having a size ratio of 1:1 with respect to the horizontal axis (or vertical axis). In this case, a flag (cu_sbt_flag) indicating that only the sub-block has been converted, direction (vertical / horizontal) information (cu_sbt_horizontal_flag) and / or position information (cu_sbt_pos_flag) are encoded by the entropy encoding unit (155) and signaled to the image decoding device. Additionally, the size of the converted area sub-block may have a size ratio of 1:3 with respect to the horizontal axis (or vertical axis), and in this case, a flag (cu_sbt_quad_flag) distinguishing the corresponding division is additionally encoded by the entropy encoding unit (155) and signaled to the image decoding device.

[0083] Meanwhile, the transformation unit (140) can perform transformations on the residual block individually in the horizontal and vertical directions. For the transformation, various types of transformation functions or transformation matrices may be used. For example, a pair of transformation functions for horizontal transformation and vertical transformation can be defined as a Multiple Transform Set (MTS). The transformation unit (140) can select one pair of transformation functions with the best transformation efficiency among the MTS and transform the residual block in the horizontal and vertical directions, respectively. Information (mts_idx) regarding the selected pair of transformation functions among the MTS is encoded by the entropy encoding unit (155) and signaled to the image decoder.

[0084] The quantization unit (145) quantizes the transformation coefficients output from the transformation unit (140) using quantization parameters and outputs the quantized transformation coefficients to the entropy encoding unit (155). The quantization unit (145) may quantize the associated residual block directly without transformation for any block or frame. The quantization unit (145) may apply different quantization coefficients (scaling values) depending on the position of the transformation coefficients within the transformation block. The quantization matrix applied to the quantized transformation coefficients arranged in two dimensions can be encoded and signaled to an image decoder.

[0085] The reordering unit (150) can perform reordering of coefficient values ​​for quantized residual values.

[0086] The reordering unit (150) can convert a two-dimensional coefficient array into a one-dimensional coefficient sequence using coefficient scanning. For example, the reordering unit (150) can output a one-dimensional coefficient sequence by scanning from DC coefficients to coefficients in the high-frequency range using a zig-zag scan or a diagonal scan. Depending on the size of the conversion unit and the intra-prediction mode, a vertical scan that scans the two-dimensional coefficient array in the column direction and a horizontal scan that scans the two-dimensional block-shaped coefficients in the row direction may be used instead of a zig-zag scan. That is, depending on the size of the conversion unit and the intra-prediction mode, the scanning method to be used among a zig-zag scan, a diagonal scan, a vertical scan, and a horizontal scan may be determined.

[0087] The entropy encoding unit (155) generates a bitstream by encoding a sequence of one-dimensional quantized transformation coefficients output from the reordering unit (150) using various encoding methods such as CABAC (Context-based Adaptive Binary Arithmetic Code) and Exponential Golomb.

[0088] Additionally, the entropy encoding unit (155) encodes information related to block division, such as CTU size, CU division flag, QT division flag, MTT division type, and MTT division direction, so that the video decoder can divide the block in the same way as the video encoding unit. Additionally, the entropy encoding unit (155) encodes information regarding a prediction type indicating whether the current block is encoded by intra prediction or by inter prediction, and encodes intra prediction information or inter prediction information according to the prediction type.

[0089] Intra prediction information is information indicating the intra prediction mode of the current block. For example, as described above, the intra prediction information may include information indicating whether the intra prediction mode of the current block is included in a candidate list. If the intra prediction mode of the current block is included in the candidate list, the intra prediction information may include first intra identification information to indicate the intra prediction mode of the current block among the candidate lists. On the other hand, if the intra prediction mode of the current block is not included in the candidate list, the intra prediction information may include second intra identification information to indicate the intra prediction mode of the current block among the remaining intra prediction modes not included in the candidate list.

[0090] Inter-predicted information may include information indicating whether the encoding mode of the motion information is merge mode or AMVP mode. In the case of merge mode, the inter-predicted information may include a merge index indicating a merge candidate used as the motion information of the current block within the merge candidate list. On the other hand, in the case of AMVP mode, the inter-predicted information may include information regarding the reference picture index, the predicted motion vector, and the difference motion vector.

[0091] Additionally, the entropy encoding unit (155) encodes information related to quantization, namely information about quantization parameters and information about quantization matrices.

[0092] The inverse quantization unit (160) inversely quantizes the quantized transformation coefficients output from the quantization unit (145) to generate transformation coefficients. The inverse transformation unit (165) converts the transformation coefficients output from the inverse quantization unit (160) from the frequency domain to the spatial domain to restore the residual block.

[0093] The adder (170) restores the current block by adding the restored residual block and the prediction block generated by the prediction unit (120). The pixels within the restored current block are used as reference pixels when intra-predicting the next block in sequence.

[0094] The loop filter section (180) performs filtering on the restored pixels to reduce blocking artifacts, ringing artifacts, blurring artifacts, etc. caused by block-based prediction and transformation / quantization. The loop filter section (180) may include all or part of a deblocking filter (182), a SAO (Sample Adaptive Offset) filter (184), and an ALF (Adaptive Loop Filter, 186) as an in-loop filter.

[0095] The deblocking filter (182) filters the boundaries between restored blocks to remove blocking artifacts caused by block-unit encoding / decoding, and the SAO filter (184) and ALF (186) perform additional filtering on the deblocking filtered image. The SAO filter (184) and ALF (186) are filters used to compensate for the difference between restored pixels and original pixels caused by lossy coding. The SAO filter (184) improves not only subjective image quality but also encoding efficiency by applying an offset in CTU units. In contrast, the ALF (186) performs block-unit filtering, and compensates for distortion by applying different filters by distinguishing the degree of edge and change of the corresponding block. Information regarding the filter coefficients to be used in the ALF can be encoded and signaled to an image decoder.

[0096] The restored blocks filtered through the deblocking filter (182), SAO filter (184), and ALF (186) are stored in memory (190). Once all blocks within a picture are restored, the restored picture can be used as a reference picture for inter-predicting blocks within a picture to be encoded later.

[0097] The video encoding device can store the bitstream of encoded video data on a non-transient recording medium or transmit it to a video decoding device using a communication network.

[0098] FIG. 5 is an exemplary block diagram of an image decoding device capable of implementing the technologies of the present disclosure. Hereinafter, the image decoding device and its sub-components will be described with reference to FIG. 5.

[0099] The image decoding device may be configured to include an entropy decoding unit (510), a reordering unit (515), an inverse quantization unit (520), an inverse transformation unit (530), a prediction unit (540), an adder (550), a loop filter unit (560), and a memory (570).

[0100] Similar to the image encoding device of FIG. 1, each component of the image decoding device may be implemented in hardware or software, or in combination of hardware and software. Additionally, the function of each component may be implemented in software, and a microprocessor may be implemented to execute the function of the software corresponding to each component.

[0101] The entropy decoding unit (510) determines the current block to be decoded by decoding the bitstream generated by the video encoding device and extracting information related to block division, and extracts prediction information, information on residual signals, etc., necessary to restore the current block.

[0102] The entropy decoding unit (510) extracts information about the CTU size from the SPS (Sequence Parameter Set) or PPS (Picture Parameter Set) to determine the size of the CTU and divides the picture into CTUs of the determined size. Then, the CTU is determined as the top layer of the tree structure, i.e., the root node, and divides the CTU using the tree structure by extracting division information for the CTU.

[0103] For example, when splitting a CTU using a QTBTTT structure, first, a first flag (QT_split_flag) related to QT splitting is extracted to split each node into four nodes of the lower layer. Then, for the nodes corresponding to the leaf nodes of QT, a second flag (mtt_split_flag) related to MTT splitting and splitting direction (vertical / horizontal) and / or splitting type (binary / ternary) information are extracted to split the corresponding leaf nodes into an MTT structure. Accordingly, each node below the leaf nodes of QT is recursively split into a BT or TT structure.

[0104] As another example, when splitting a CTU using the QTBTTT structure, a CU splitting flag (split_cu_flag) indicating whether to split the CU is first extracted, and if the block is split, a first flag (QT_split_flag) is extracted. During the splitting process, each node may undergo zero or more iterative MTT splitting after zero or more iterative QT splittings. For example, the CTU may undergo MTT splitting immediately, or conversely, only multiple QT splittings may occur.

[0105] As another example, when splitting a CTU using a QTBT structure, a first flag (QT_split_flag) related to the splitting of QT is extracted to split each node into four nodes of the lower layer. Then, for the nodes corresponding to the leaf nodes of QT, a split flag (split_flag) indicating whether to further split into BTs and split direction information are extracted.

[0106] Meanwhile, when the entropy decoding unit (510) determines the current block to be decoded using the division of the tree structure, it extracts information regarding the prediction type indicating whether the current block is intra-predicted or inter-predicted. If the prediction type information indicates intra-predicted, the entropy decoding unit (510) extracts syntax elements for the intra-predicted information (intra-predicted mode) of the current block. If the prediction type information indicates inter-predicted, the entropy decoding unit (510) extracts syntax elements for the inter-predicted information, namely information indicating the motion vector and the reference picture that the motion vector refers to.

[0107] Additionally, the entropy decoder (510) extracts information regarding quantization-related information and information regarding residual signals, as well as information regarding the quantized transformation coefficients of the current block.

[0108] The reordering unit (515) can change the sequence of one-dimensional quantized transformation coefficients entropy-decoded in the entropy decoding unit (510) back into a two-dimensional coefficient array (i.e., block) in the reverse order of the coefficient scanning order performed by the image encoding device.

[0109] The inverse quantization unit (520) inversely quantizes the quantized transformation coefficients and inversely quantizes the quantized transformation coefficients using quantization parameters. The inverse quantization unit (520) may apply different quantization coefficients (scaling values) to the quantized transformation coefficients arranged in two dimensions. The inverse quantization unit (520) may perform inverse quantization by applying a matrix of quantization coefficients (scaling values) from an image encoding device to a two-dimensional array of quantized transformation coefficients.

[0110] The inverse transformation unit (530) generates a residual block for the current block by inversely transforming the inversely quantized transformation coefficients from the frequency domain to the spatial domain and restoring the residual signals.

[0111] Additionally, when the inverse transformation unit (530) inversely transforms only a part of the transformation block (sub-block), it extracts a flag (cu_sbt_flag) indicating that only the sub-block of the transformation block has been transformed, information on the directionality (vertical / horizontal) of the sub-block (cu_sbt_horizontal_flag) and / or information on the position of the sub-block (cu_sbt_pos_flag), restores residual signals by inversely transforming the transformation coefficients of the corresponding sub-block from the frequency domain to the spatial domain, and creates a final residual block for the current block by filling the areas that have not been inversely transformed with "0" values ​​of residual signals.

[0112] Additionally, when MTS is applied, the inverse transformation unit (530) determines a transformation function or transformation matrix to be applied in the horizontal and vertical directions, respectively, using MTS information (mts_idx) signaled from the video encoding device, and performs an inverse transformation on the transformation coefficients within the transformation block in the horizontal and vertical directions using the determined transformation function.

[0113] The prediction unit (540) may include an intra prediction unit (542) and an inter prediction unit (544). The intra prediction unit (542) is activated when the prediction type of the current block is an intra prediction, and the inter prediction unit (544) is activated when the prediction type of the current block is an inter prediction.

[0114] The intra prediction unit (542) determines the intra prediction mode of the current block among a plurality of intra prediction modes from the syntax elements for the intra prediction mode extracted from the entropy decoding unit (510), and predicts the current block using reference pixels around the current block according to the intra prediction mode.

[0115] The inter prediction unit (544) determines the motion vector of the current block and the reference picture that the motion vector refers to using the syntax elements for the inter prediction mode extracted from the entropy decoding unit (510), and predicts the current block using the motion vector and the reference picture.

[0116] The adder (550) restores the current block by adding the residual block output from the inverse transformation unit (530) and the prediction block output from the inter prediction unit (544) or the intra prediction unit (542). The pixels within the restored current block are used as reference pixels when intra-predicting the block to be decoded later.

[0117] The loop filter section (560) may include a deblocking filter (562), an SAO filter (564), and an ALF (566) as an in-loop filter. The deblocking filter (562) deblocks the boundaries between restored blocks to remove blocking artifacts caused by block-unit decoding. The SAO filter (564) and the ALF (566) perform additional filtering on the restored blocks after deblocking filtering to compensate for the difference between the restored pixels and the original pixels caused by lossy coding. The filter coefficients of the ALF are determined using information about the filter coefficients decoded from the non-stream.

[0118] The restored blocks filtered through the deblocking filter (562), SAO filter (564), and ALF (566) are stored in memory (570). When all blocks within a picture are restored, the restored picture is used as a reference picture to inter-predict blocks within the picture to be encoded later.

[0119] An improved coding tool performed by the aforementioned image encoding device or image decoding device is disclosed below.

[0120] Improvements to IBC mode

[0121] Various improvement techniques of the IBC mode described below can be performed by the intra prediction unit (122) of the video encoding device or the intra prediction unit (542) of the video decoding device.

[0122] IBC (Intra Block Copy) refers to a technique that generates a predicted block for a current block using the reference block most similar to the current block within a restored area of ​​the current picture containing the current block. The displacement from the current block to the reference block is called the block vector.

[0123] IBC mode is similar to intra prediction in that it utilizes restored pixels within the current picture. However, except for the fact that it uses restored regions within the current picture rather than a reference picture restored before the current picture, it is also similar to inter prediction technology in that it identifies the restored region of the block most similar to the current block and represents it as a vector.

[0124] Therefore, encoding tools applied to inter-prediction technology can be applied to IBC mode. “Motion vectors” in inter-prediction technology can be replaced with “block vectors” in IBC mode.

[0125] For example, similar to Inter-prediction, IBC mode may also have sub-modes such as Skip mode, Merge mode, and AMVP mode, and the details of each sub-mode described in Inter-prediction can be applied directly to IBC mode. For instance, in the case of IBC Merge mode, IBC-related information or block vector information of the current block may be inherited from the block used to derive each candidate within the block vector candidate list.

[0126] In the case of IBC merge mode, the prediction information may include index information (merge index information) indicating the candidate among the candidates in the block vector candidate list that is used as the block vector for the current block. In the case of IBC AMVP mode, the information may include information on the block vector difference along with index information (AMVP index information) indicating the candidate among the candidates in the block vector candidate list that is used as the block vector for the current block. The final block vector is derived by combining the Block Vector Predictor (BVP) and the Block Vector Difference (BVD) indicated by the AMVP index information.

[0127] Improvement of intra-prediction

[0128] Various improvement techniques for intra prediction described below can be performed by the intra prediction unit (122) of the video encoding device or the intra prediction unit (542) of the video decoding device.

[0129] In one embodiment, a method for inducing at least one intra prediction mode for the current block without separate signaling may be used. At least one intra prediction mode for the current block may be generated based on information about the surrounding area of ​​the current block or surrounding blocks.

[0130] In one embodiment, the current block can be predicted using a Spatial Geometric Partitioning Mode (SGPM). SGPM may be a sub-mode of the intra-frame prediction technique.

[0131] In SGPM, the current block is divided into multiple (e.g., two) partitions or sub-regions through geometric partitioning. Predicted blocks corresponding to each partition can be generated through intra-predicted techniques (e.g., directional prediction mode, planar mode, or DC mode, etc.) or block vector-based prediction techniques (e.g., IntraTMP or IBC mode). A predicted block for the current block can be generated through a weighted sum of predicted blocks corresponding to each sub-region.

[0132] The geometric partition type (or partition mode) of the current block and the prediction information used in each partition can be signaled from the image encoding device to the image decoder.

[0133] The weights of the partitions can be explicitly signaled or implicitly derived. The weight matrix for weighted summing can be determined according to an agreement between the decoding and encoding units. Here, the weight matrix may include weights to be applied to each sample location. The weight for each sample location can be set according to the distance from the partition boundary determined by the partitioning mode.

[0134] In one embodiment, the current block can be predicted using a matrix-based intra-prediction (MIP) mode. Predicted samples of the current block can be generated based on a multiplication operation with a matrix of restored reference samples around the current block.

[0135] For the matrix, an index to indicate the matrix to be used in the current block among multiple defined matrices can be encoded by the image encoding device and signaled to the image decoder. Alternatively, information regarding the coefficient values ​​constituting the matrix may be signaled.

[0136] In one embodiment, the current block can be predicted using an Intra Template Matching Prediction (IntraTMP).

[0137] You can set a previously restored area around the current block as the template for the current block.

[0138] In the restored reference region of the current picture, at least one reference block having the template with the smallest error with the template of the current block can be searched, and a prediction block for the current block can be generated based on the searched reference block. According to an embodiment, block vector information derived by IntraTMP can be stored and used in the subsequent encoding or decoding process of the block.

[0139] As one embodiment for deriving an intra prediction mode of the current block, at least one intra prediction mode for the current block may be derived based on a gradient operation for reference samples within a restored reference region in the current picture.

[0140] Here, the reference region may be a region adjacent to the current block or a non-adjacent region of the current block. The non-adjacent region may be a region indicated by a block vector of a neighboring block of the current block or a block vector derived through intra-template matching.

[0141] FIG. 6 is a drawing showing reference regions according to one embodiment of the present disclosure.

[0142] In FIG. 6(a), an L-shaped region containing sample lines for each of the top, top-left, and left regions of the current block within the restored region around the current block can be used as a reference region. Here, the number of sample lines can be 1 to 3.

[0143] In FIG. 6(b), the lower-left area of ​​the current block within the restored area around the current block can be used as a reference area.

[0144] In FIG. 6(c), the upper-right area of ​​the current block within the restored area surrounding the current block can be used as a reference area.

[0145] In FIG. 6(d), the lower-left and upper-right regions of the current block within the restored area surrounding the current block can be used as reference regions.

[0146] The intra prediction unit can generate information on the direction and amplitude of gradients using horizontal and vertical gradients for reference samples within a reference region. Gradients can be derived using boundary filters such as Sobel filters.

[0147] The direction of the gradient can correspond to the aforementioned directional intra-prediction modes. If the direction of the gradient does not match the direction of a predefined intra-prediction mode, the intra-prediction mode having the directionality closest to the direction of the gradient can be mapped to the direction of the gradient.

[0148] The intra prediction unit can generate a Histogram of Gradient (HoG) for each intra prediction mode by accumulating gradient magnitudes corresponding to the same intra prediction mode. One or more intra prediction modes for the current block can be derived based on the magnitudes of the gradients on the HoG.

[0149] For example, the intra prediction mode having the largest gradient size can be set as the intra prediction mode of the current block.

[0150] As another example, multiple intra prediction modes may be selected as the intra prediction modes of the current block in order of gradient magnitude. In this case, the prediction block of the current block may be generated based on a weighted sum of multiple prediction blocks generated according to the multiple intra prediction modes. The weights for the weighted sum may be determined based on the gradient magnitude for each intra prediction mode. For example, the weights may be proportional to the gradient magnitude.

[0151] As another example, prediction blocks generated by one or more gradient-based intra-prediction modes may be weighted and summed with prediction blocks predicted by non-directional modes (DC or Planar modes) or block vectors.

[0152] The size of the reference area and / or filter for gradient calculation can be adaptively determined based on at least one of the size of the current block, the aspect ratio, or the resolution of the picture.

[0153] The histogram information configured in the above process is stored and can be used for blocks to be encoded or decoded later.

[0154] For example, when constructing a candidate list of intra prediction modes for the current block based on prediction mode information of surrounding blocks, directional intra prediction modes having a low amplitude in the histogram information of the surrounding blocks may be excluded from the candidate list.

[0155] As another embodiment for inducing an intra prediction mode of the current block, at least one intra prediction mode for the current block may be induced based on the frequency of occurrence of intra prediction modes around the current block.

[0156] The intra prediction unit can count the frequency of occurrence of intra prediction modes of surrounding blocks of the current block. Surrounding blocks may include blocks immediately adjacent to the current block or non-adjacent blocks separated from the current block by a certain distance or more.

[0157] For example, the occurrence frequency may refer to a pixel-based occurrence frequency. That is, the occurrence frequency of the intra prediction mode of a specific surrounding block can be set as the area of ​​that surrounding block, i.e., the product of its width and height.

[0158] Additionally or alternatively, the occurrence frequency can be derived based on the distance from the current block to the surrounding blocks under consideration. That is, the occurrence frequency can be set inversely proportional to the distance.

[0159] The intra prediction unit generates a Histogram of Occurrences (HoG) based on the occurrence frequencies of surrounding intra prediction modes and can select one or more intra prediction modes in order of highest frequency. The prediction block for the current block can be generated through a weighted sum of prediction blocks predicted based on the selected intra prediction modes. The weights for the weighted sum can be determined based on the occurrence frequencies of the intra prediction modes. That is, the weights can be set in proportion to the occurrence frequency.

[0160] As another example, prediction blocks generated by one or more intra-prediction modes derived based on occurrence frequency may be weighted and summed with prediction blocks predicted based on non-directional modes (DC or Planar modes) or block vectors.

[0161] As another embodiment for deriving an intra prediction mode of the current block, at least one intra prediction mode for the current block may be derived based on template matching.

[0162] FIG. 7 is a drawing for illustrating a TIMD mode according to one embodiment of the present disclosure.

[0163] In FIG. 7, a restored area around the current block can be defined as a template. The template may be an adjacent area of ​​the current block or a non-adjacent area.

[0164] The intra prediction unit (122, 542) can generate an intra prediction mode candidate list using intra prediction modes of surrounding blocks of the current block. As an example, the intra prediction mode candidate list may be an MPM list. The MPM list may include intra-frame prediction modes of spatially adjacent blocks, multiple intra-frame prediction modes derived using the DIMD method, and intra-frame prediction modes of spatially non-adjacent blocks as candidates. Prediction mode candidates may be sorted based on the matching cost between the predicted template area and the restored template area using each prediction mode candidate in the candidate list. The candidate list may further include prediction modes having a prediction angle adjacent to the prediction angle of the prediction mode candidate as candidates. The candidate list may include a DC mode, a vertical intra prediction mode, and a horizontal intra prediction mode as default modes.

[0165] Predicted samples for a template can be generated from samples within the template's surrounding reference region using each candidate in the candidate list.

[0166] The intra prediction unit (122, 542) calculates the matching cost between the reconstructed samples and the predicted samples within the template, and then can derive one or more intra prediction modes having a small matching cost. When deriving multiple intra prediction modes, the prediction block of the current block can be generated by weighted summing the prediction blocks generated using each intra prediction mode. The weights for the weighted sum can be set to be inversely proportional to the matching cost.

[0167] For example, two intra prediction modes may be selected in order of lowest template cost. By comparing the costs of the two selected intra prediction modes, it may be determined whether to use only the intra prediction mode with the lowest cost or to use both intra prediction modes for the intra prediction of the current block.

[0168] Meanwhile, prediction blocks generated by one or more intra prediction modes derived based on template matching may be weighted and summed with prediction blocks predicted based on non-directed modes (DC or Planar modes) or block vectors.

[0169] In one embodiment, when a directional intra prediction mode is applied to the current block, a prediction block can be generated by weighting reference pixels existing in both directions based on the prediction direction of the current directional mode. In this case, the weights used for the weighting sum may be determined according to at least one of the size of the current block, the prediction mode of the current block, and the location of the predicted pixel.

[0170] In one embodiment, when a DC mode is applied to the current block, a prediction block can be generated by weighting the prediction block generated in DC mode with reference pixels located at the top and left. In this case, the weights used for the weighting sum may be determined based on at least one of the size of the current block or the location of the predicted pixel.

[0171] In one embodiment, cross-component prediction may be used. For example, a model for cross-component prediction may be derived through the correlation between reconstructed color difference samples around the color difference block of the current block and reconstructed luminance samples around the luminance block at a location corresponding to the color difference block. The color difference block may be generated by applying the cross-component prediction model to the luminance block.

[0172] The prediction model between components can be a linear model (e.g., Cross Component Linear Model) or a non-linear model (e.g., Convolutional Cross Component Model).

[0173] In one embodiment, when a prediction is performed for a color difference block of the current block, the prediction block of the color difference block may be generated using a Direct Mode (DM) mode. The DM mode refers to a method of performing a prediction of a color difference component block using a prediction mode of a luminance block corresponding to the color difference block.

[0174] In one embodiment, the current block can be predicted using a directional planar mode. The directional planar mode includes a horizontal planar mode and a vertical planar mode. The horizontal planar mode is a mode that predicts the current sample by interpolating a reference sample located on the same horizontal line as the current sample among the reference samples within the left reference line of the current block, and a reference sample at the top right of the current block. The vertical planar mode is a mode that predicts the current sample by interpolating a reference sample located on the same vertical line as the current sample among the reference samples within the top reference line of the current block, and a reference sample at the bottom left of the current block.

[0175] Improvement of Inter prediction

[0176] Various improvement techniques for inter prediction can be performed by the inter prediction unit (124) of the video encoding device or the inter prediction unit (544) of the video decoding device.

[0177] As described above, the inter-prediction mode may include AMVP mode and merge mode. Information regarding which mode to use (e.g., a 1-bit flag) may be signaled from the video encoding device to the video decoder.

[0178] In one embodiment, the current block can be predicted using a Geometric Partitioning Mode (GPM). The GPM may be a sub-mode of the cross-frame prediction technique.

[0179] In GPM mode, the current block can be divided into multiple partitions or sub-regions through geometric division. The divided sub-regions can be non-rectangular in shape.

[0180] Multiple prediction blocks can be generated based on motion information for each sub-region. A prediction block for the current block can be generated by weighting the multiple prediction blocks.

[0181] Motion information for each sub-region can be encoded by a video encoding device and signaled to a video decoder. For example, a list of motion information candidates can be constructed. Among the candidates in the candidate list, index information indicating the motion information candidate to be applied to each sub-region can be signaled. The index information can be signaled for each sub-region.

[0182] According to an embodiment, each candidate in the candidate list may be composed of a combination of motion information for sub-regions. In this case, one index information may be signaled. Motion information for each sub-region may be determined by the candidate corresponding to the index information.

[0183] In GPM mode, the partitioning mode or partitioning type of the current block can be defined by the distance between the center of the current block and the boundary line and the angle formed by the center of the current block and the boundary line. Information regarding the partitioning mode can be encoded by an image encoding device and transmitted to an image decoder.

[0184] In GPM mode, the weight matrix for weighted summing can be determined according to an agreement between the image encoding device and the image decoder. The weight for each sample position can be set according to the distance from the partition boundary line determined by the partition type.

[0185] In one embodiment, a combined technique of inter-prediction and intra-prediction may be used. A prediction block for the current block may be derived by weighting an inter-prediction block generated based on inter-prediction and an intra-prediction block generated based on intra-prediction.

[0186] Inter-predicted blocks can be generated based on merge mode. Alternatively, motion information for the current block can be derived by defining a reconstructed area (current template) around the current block in the current picture and using a template matching method based on a target template. Depending on the template matching method, motion information for the current block can be derived through a process of finding a reference block that has a template most similar to the current template.

[0187] An intra prediction block can be generated based on a predefined intra prediction mode. For example, the intra prediction mode can be derived by calculating statistical values ​​based on a restored area of ​​the current picture and using said statistical values. For instance, at least one of the aforementioned gradient-based, frequency-based, or template-matching-based intra prediction mode derivation methods may be utilized. For instance, a restored area (current template) surrounding the current block in the current picture is defined, and the current block can be intra-predicted based on a template-matching method using the target template.

[0188] The weights of the inter-prediction block and intra-prediction block may vary depending on each sample location and can be adaptively determined in the decoder using the prediction information of the current block, information on the reconstructed area around the current block, aspect ratio, etc.

[0189] Local illumination compensation (LIC)

[0190] The local illuminance compensation mode is a technique for compensating for changes in illuminance between regions. The local illuminance compensation mode may be a mode that derives a linear model including at least one of weights and offsets by calculating the correlation between the template of the current block and the template of the reference block, and applies the linear model to part or all of the current block. Here, the weights of the linear model refer to parameters used for multiplication operations with the reference block of the target block, and the offset refers to parameters used for addition operations with the reference block of the target block. In the local illuminance compensation mode, the sample value (P) can be modified as follows.

[0191] P' = α*P + β (α is the weight, β is the offset)

[0192] Local illumination compensation can be applied to inter-prediction techniques. However, the present invention is not limited thereto. For example, linear filtering for local illumination compensation can be applied to predicted blocks using reference vectors, such as IntraTMP or IBC. An IBC mode combined with LIC can be referred to as an IBC-LIC (Intra Block Copy-Local illumination Compensation) mode, and an IntraTMP mode combined with LIC can be referred to as an IntraTMP-LIC (Intra Block Copy-Local illumination Compensation) mode.

[0193] In some embodiments of the present disclosure, convolution filtering may be applied to the predicted block using a reference vector. This mode may be referred to as a filtered IBC (Filtered IBC, FIBC) mode.

[0194] Filter coefficients can be derived based on the correlation between the template around the current block (current template) and the template around the reference block at the location pointed to by the current block's block vector (reference template). Predicted samples within the block predicted by IBC mode can be modified using the filter coefficients.

[0195] In the present disclosure, the current color difference block can be predicted using a reference vector-based color difference prediction method. A reference vector-based color difference prediction method is a method that performs color difference prediction using a reference vector derived from a luminance block corresponding to the color difference block. This method may be referred to as a Direct Reference Vector (DRV) mode.

[0196] Here, the corresponding luminance block refers to a block having a collocated position.

[0197] In the following, the reference vector of the chrominance block may refer to either a motion vector or a block vector. A motion vector is information indicating a restored area within a reference picture rather than the current picture, and a block vector is information indicating a restored area within the current picture. The mode for deriving the block vector of the chrominance block from the luminance block may be referred to as the Direct Block Vector (DBV) mode.

[0198] Whether the prediction mode of a color difference block is a reference vector-based color difference prediction method is determined by an agreement between an image encoding device and an image decoder, or the image encoding device may transmit information to the image decoder regarding whether a reference vector-based color difference prediction method is applied. As an example, when the prediction mode of a color difference block is an IBC mode, a reference vector-based color difference prediction method according to the present disclosure may be applied.

[0199] FIG. 8 is a flowchart of a reference vector-based color difference prediction method according to one embodiment of the present disclosure.

[0200] In FIG. 8, the reference vector-based color difference prediction method derives a reference vector of the color difference block from a luminance region corresponding to the color difference block (S810), and performs a prediction for the color difference block using the reference vector of the color difference block (S820).

[0201] In step S810, the method of deriving the reference vector of the color difference block may vary depending on the block division structure of the luminance component and the color difference component of the current block.

[0202] Whether a reference vector-based color difference prediction method is applied to the current color difference block can be derived from the image encoding device and image decoder, or determined through signaling.

[0203] Below, the derivation of the color difference block reference vector of step S810 is described.

[0204] FIG. 9 is a flowchart of a method for deriving a color difference reference vector according to one embodiment of the present disclosure.

[0205] In step S910, it is determined whether the block division structure of the luminance component and chrominance component of the current block is a single tree structure or a dual tree structure.

[0206] In step S920, in a single tree structure indicating that the luminance component block division structure and the chrominance component block division structure are identical, the chrominance block can inherit the reference vector of the corresponding luminance block. That is, the block vector information of the luminance block corresponding to the chrominance block can be used as the block vector of the chrominance block.

[0207] In step S930, in a dual tree structure indicating that the block partitioning structures of the luminance component and the chrominance component are different, luminance subregions having reference vectors within the luminance block corresponding to the chrominance block are searched. Specifically, among the luminance subregions covering predefined sample locations within the luminance block corresponding to the chrominance block, luminance blocks having reference vectors are searched, thereby deriving at least one reference vector of a luminance block.

[0208] FIG. 10 is an illustrative diagram for explaining a method for deriving a reference vector of a color difference block in a dual tree structure according to one embodiment of the present disclosure.

[0209] Referring to FIG. 10, a reference vector can be derived by searching for predicted luminance samples using a reference vector among luminance samples at predetermined locations within a luminance block corresponding to a color difference block in a predetermined sample order.

[0210] The predefined location may be at least one of the center sample, top-left sample, top-right sample, bottom-left sample, and bottom-right sample within the luminance region. The predefined search order may be the order of center sample, top-left sample, top-right sample, bottom-left sample, and bottom-right sample.

[0211] As an example, if a luminance subregion containing luminance samples at a predefined location is predicted using a block vector such as in IntraTMP mode or IBC mode, the block vector of the luminance subregion can be used as a block vector or block vector candidate for a chrominance block.

[0212] As an example, if the prediction mode of a luminance block containing a predefined luminance sample is a prediction mode that blends multiple prediction blocks including a prediction block derived using the reference vector of the luminance block, the reference vector of the luminance block may be used as a reference vector or a candidate reference vector of the chrominance block. Here, the blending prediction mode may be a DIMD mode, an OBIC mode, or a TIMD mode, and the multiple prediction blocks may include a prediction block derived using an intra prediction mode.

[0213] As an example, if a luminance block containing a predefined luminance sample is predicted using a bidirectional reference vector, the two unidirectional reference vectors of the luminance block may be used as reference vectors or reference vector candidates for the chrominance block. The list of reference vector candidates for the chrominance block may include bidirectional reference vectors as candidates.

[0214] In other embodiments, the search location and search order may be changed.

[0215] Meanwhile, in a single tree structure or a dual tree structure, the reference vector of the luminance block can be scaled based on a chrominance subsampling format.

[0216] When the chrominance subsampling format is Y:Cb:Cr = 4:2:0, the resolution of the chrominance component is half the resolution of the luminance component, and the chrominance samples correspond to luminance samples sampled at a ratio of 2:1 in the vertical and horizontal directions. That is, four luminance samples can correspond to one chrominance component. In the 4:2:0 format, the reference vector of the luminance block can be downsampled at a ratio of 2:1 in the vertical and horizontal directions. The reference vector of the chrominance block becomes half the reference vector of the luminance block. The scaled reference vector of the luminance block can be expressed as Equation 1.

[0217] [Mathematical Formula 1]

[0218] scaled RV x = (luma_RV x + 1) >> 1

[0219] scaled RV y = (luma_RV y + 1) >> 1

[0220] In mathematical formula 1, luma_RV x and luma_RV y represent the horizontal and vertical components of the reference vector, respectively, and scaled RV x and scaled RV y and represent the horizontal and vertical components of the scaled reference vector, respectively. The symbol “A>>B” signifies an operator that right-shifts the bit sequence A by B bits.

[0221] When the chrominance subsampling format is Y:Cb:Cr = 4:2:2, the horizontal resolution of the chrominance component is half the horizontal resolution of the luminance component, and the chrominance samples correspond to luminance samples sampled at a ratio of 2:1 in the horizontal direction. That is, four luminance samples can correspond to two chrominance components. In the 4:2:2 format, the reference vector of the luminance block can be downsampled at a ratio of 2:1 in the horizontal direction. The vertical length of the reference vector of the chrominance block is equal to the vertical length of the reference vector of the luminance block, and the horizontal length of the reference vector of the chrominance block is half the horizontal length of the reference vector of the luminance block. The scaled reference vector of the luminance block can be expressed as Equation 2.

[0222] [Mathematical Formula 2]

[0223] scaled RV x = (luma_RV x + 1) >> 1

[0224] scaled RV y = luma_RV y

[0225] When the chrominance subsampling format is Y:Cb:Cr = 4:4:4, the horizontal resolution of the chrominance components is equal to the horizontal resolution of the luminance components, and the chrominance samples correspond directly to the luminance samples. That is, four luminance samples can correspond to four chrominance components. In the 4:4:4 format, the reference vector of the luminance block may not be scaled. The scaled reference vector of the luminance block can be expressed as in Equation 3.

[0226] [Mathematical Formula 3]

[0227] scaled RV x = luma_RV x

[0228] scaled RV y = luma_RV y

[0229] In step S940, a list of reference vector candidates is generated using the reference vector of the luminance region.

[0230] At least one reference vector of the luminance region may be added to the reference vector candidate list as a reference vector candidate. If the reference vector of the luminance region is scaled, the scaled reference vector may be added to the reference vector candidate list.

[0231] In step S950, the reference vector candidate list can be sorted based on the template matching cost. The reference vector candidates within the reference vector candidate list can be sorted in order of lowest or highest template matching cost.

[0232] FIGS. 11a, FIGS. 11b, FIGS. 11c and FIGS. 11d are drawings illustrating various template regions according to one embodiment of the present disclosure.

[0233] In FIG. 11a, an L-shaped region within a restored area surrounding the current color difference block having a size of W x H can be defined as the template region of the current color difference block. The template region may include a left region of size N x H, a top region of size W x M, and a top-left region of size N x M. It may be referred to as the current template region.

[0234] The reference color difference block is a color difference region located at a distance from the current color difference block by a reference vector or a scaled reference vector, and an L-shaped region within the restored area surrounding the reference color difference block can be defined as the template region of the reference color difference block. It may be referred to as the reference template region.

[0235] W, H, N, M are integers greater than or equal to 1 and may be predefined in the image encoding device and image decoding device.

[0236] The current template area and the reference template area can have the same shape and size.

[0237] The template matching cost of a reference vector candidate can be calculated based on the difference between the template region of the current chrominance block and the template region of the reference chrominance block indicated by the reference vector candidate. Specifically, predicted samples within the current template region are derived using the restored samples within the reference template region, and the matching cost between the predicted samples within the current template region and the restored samples can be calculated.

[0238] Functions for calculating matching costs may include SAD (Sum of Absolute Differences), SATD (Sum of Absolute Transformed Differences), MR-SAD (Mean-Removed Sum of Absolute Differences), MSE (Mean Squared Error), or SSE (Sum of Squared Error).

[0239] For example, the template matching cost of a reference vector candidate can be calculated based on SAD, which represents the absolute difference between the predicted samples and the restored samples within the current template region. As another example, the template matching cost of a reference vector candidate can be calculated based on SATD, which represents the result of summing the absolute values ​​after performing a Hadamard transformation on the predicted samples and the restored samples within the current template region.

[0240] In Fig. 11b, N additional The top-right area of ​​size x M can be used as a template area.

[0241] In Fig. 11c, N x M additional The bottom-left area of ​​the size can be used as a template area.

[0242] In Fig. 11d, N additional Top-right area of ​​size x M and N x M additional The bottom-left area of ​​the size can be used as a template area.

[0243] In another embodiment, step S950 may be omitted.

[0244] In step S960, the reference vector of the current color difference block is derived by selecting one of the reference vector candidates in the reference vector candidate list.

[0245] The reference vector of the current color difference block can be determined implicitly or explicitly.

[0246] In the implicit method, the candidate list of reference vectors for the color difference block is sorted based on the template matching cost, and among the sorted reference vector candidates, the candidate with the smallest template matching cost can be selected as the reference vector for the color difference block.

[0247] In an explicit method, candidate index information is transmitted from an image encoding device to an image decoder, and among the reference vector candidates in a reference vector candidate list, a candidate selected by the candidate index information can be selected as the reference vector of a color difference block.

[0248] Below, the color difference block prediction of step S820 is described.

[0249] A color difference block can be predicted from the region indicated by the reference vector of the color difference block using one of a plurality of prediction methods for the color difference block. The prediction method for the color difference block may include a default prediction method and a model-based prediction method.

[0250] The default prediction method is a method that uses a reference region indicated by the reference vector of the color difference block as the prediction block of the color difference block. Recovered samples of the reference region indicated by the reference vector of the color difference block can be used as prediction samples of the color difference block. In a dual tree structure, recovered samples of the reference color difference region indicated by a reference vector candidate selected from a list of reference vector candidates can be used as prediction samples of the color difference block.

[0251] A model-based prediction method is a method for predicting chrominance blocks by applying a prediction model to a reference region indicated by the reference vector of the chrominance block. Model-based prediction methods may include convolutional model-based prediction methods and linear model-based prediction methods.

[0252] FIG. 12 is a diagram illustrating a model-based prediction method for color difference blocks according to one embodiment of the present disclosure.

[0253] In FIG. 12, a reference chrominance block A is derived at a location separated by a reference vector scaled from the current chrominance block B, and the current chrominance block B can be predicted by applying a convolutional model or a linear model to the reference chrominance block A. T() represents a convolutional model or a linear model.

[0254] FIG. 13 is a flowchart for color difference block prediction according to one embodiment of the present disclosure.

[0255] In Fig. 13, a prediction method for the color difference block is determined (S1310). As a prediction method for the color difference block, one of a default prediction method, a linear model-based prediction method, or a convolution model-based prediction method may be selected.

[0256] In one embodiment, the prediction method for the color difference block may be determined based on information decoded from the bitstream. Either a default prediction method or a model-based prediction method may be determined based on a 1-bit flag (DRVmodFlag). As a model-based prediction method, either a linear model-based prediction method or a convolutional model-based prediction method may be determined based on a 1-bit flag (DRVmodeIdx).

[0257] In one embodiment, the prediction method of a color difference block may be determined based on the prediction method of a luminance block corresponding to the color difference block. The prediction method of the color difference block may be determined based on which prediction method the corresponding luminance block was predicted by among a default prediction method, a linear model-based prediction method, or a convolution model-based prediction method. That is, the color difference block may inherit the prediction method of the corresponding luminance block.

[0258] In a single tree structure, a reference vector of a chrominance block is derived from a corresponding luminance block, and if the corresponding luminance block is predicted using a linear model-based prediction method, the prediction method of the chrominance block can be determined to be a linear model-based prediction method.

[0259] In a dual tree structure, when multiple luminance sub-regions within a luminance block corresponding to a chrominance block are predicted using a reference vector, the prediction method for the chrominance block can be determined based on the number of luminance sub-regions predicted by the default prediction method, the number of luminance sub-regions predicted by the linear model-based prediction method, and the number of luminance sub-regions predicted by the convolution model-based prediction method. For example, when multiple sub-regions within a luminance block corresponding to a chrominance block are predicted using a reference vector or by blending a reference vector-based predictor and an intra-predictor, if the number of sub-regions predicted using the linear model-based prediction method is greater than a threshold, the prediction method for the chrominance block can be determined to be the linear model-based prediction method. Here, the threshold may be predefined or signaled between the image encoding device and the image decoder.

[0260] In one embodiment, the model-based prediction method is allowed only in a single-tree structure, and the model-based prediction method may not be allowed in a dual-tree structure. That is, in a dual-tree structure, only the default prediction method may be allowed.

[0261] In one embodiment, a model-based prediction method may be allowed even in a dual tree structure. It is determined that a prediction model is used to predict the current color difference block from a reference color difference block indicated by the reference vector of the current color difference block in a dual tree structure.

[0262] FIG. 14 is a flowchart for determining a method for predicting a color difference block according to one embodiment of the present disclosure.

[0263] The flowchart of FIG. 14 can be performed in an image decoder. The image encoding device can encode the encoding information of FIG. 14 and transmit the encoding information to a decoder.

[0264] A DRVflag indicating whether the current color difference block is predicted using the reference vector of the luminance block is parsed (S1401).

[0265] When DRVflag is 1, the prediction mode of the current chrominance block is determined to be the mode predicted using the reference vector of the luminance block.

[0266] It is determined whether the luminance block corresponding to the current color difference block is encoded based on a convolution model (S1403). In a dual tree structure, it is determined whether the number of sub-regions predicted using a convolution model-based prediction method among the luminance sub-regions within the luminance block is greater than or equal to a threshold.

[0267] If the luminance block corresponding to the current color difference block is encoded based on a convolution model, the prediction method of the current color difference block is determined to be a convolution-based prediction method (S1415).

[0268] If the luminance block corresponding to the current color difference block is not encoded based on a convolution model, it is determined whether the luminance block corresponding to the current color difference block is encoded based on a linear model (S1405). In a dual tree structure, it is determined whether the number of sub-regions predicted using a linear model-based prediction method among the luminance sub-regions within the luminance block is greater than or equal to a threshold.

[0269] If the luminance block corresponding to the current color difference block is encoded based on a linear model, the prediction method for the current color difference block is determined to be a linear model-based prediction method (S1415).

[0270] If the luminance block corresponding to the current color difference block does not use both convolution model-based prediction and linear model-based prediction, the syntax DRVmodeFlag is parsed (S1407).

[0271] It is determined whether the syntax DRVmodeFlag is 0 (S1409).

[0272] When the syntax DRVmodeFlag is 0, the prediction method for the color difference block is determined to be the default prediction method (S1419).

[0273] If the syntax DRVmodeFlag is not 0, the syntax DRVmodeIdx representing one of the convolutional model-based prediction method and the linear model-based prediction method is parsed (S1411).

[0274] It is determined whether the syntax DRVmodeIdx is 0 (S1413).

[0275] If the syntax DRVmodeIdx is 0, the prediction method of the current color difference block is determined to be a linear model-based prediction method (S1415).

[0276] When the syntax DRvmodeIDx is 1, the prediction method of the current color difference block is determined to be a convolution model-based prediction method (S1417).

[0277] FIG. 15a is a flowchart for determining a method for predicting a color difference block in a dual tree structure according to one embodiment of the present disclosure. In FIG. 15a, a convolution model-based prediction method may be implicitly determined.

[0278] FIG. 15b is a flowchart for determining a method for predicting a color difference block in a dual tree structure according to one embodiment of the present disclosure. In FIG. 15b, a convolution model-based prediction method can be explicitly determined.

[0279] In FIGS. 15a and 15b, it is determined whether the number of luminance sub-regions within a luminance block in a dual tree structure, predicted using a linear model-based prediction method or a convolutional model-based prediction method, is greater than or equal to a threshold. For example, if the number of luminance sub-regions predicted using a convolutional model-based prediction method is greater than or equal to a first threshold, the prediction method of the current chrominance block may be determined to be a convolutional model-based prediction method. For example, if the number of luminance sub-regions predicted using a linear model-based prediction method is greater than or equal to a second threshold, the prediction method of the current chrominance block may be determined to be a linear model-based prediction method. Here, the first threshold and the second threshold may be predefined or signaled between the image encoding device and the image decoder.

[0280] Whether a convolution model-based prediction method is explicitly or implicitly derived may be defined in parameter sets such as SPS and PPS.

[0281] If it is determined through the above process that the current color difference block prediction method is the default prediction method, the restored samples of the reference color difference region indicated by the reference vector of the color difference block can be used as prediction samples of the color difference block (S1320).

[0282] Specifically, the default prediction method can be expressed as Equation 4.

[0283] [Mathematical Formula 4]

[0284] pred c (i, j) = rec c (i - scaledRV x , j - scaledRV y )

[0285] In mathematical equation 4, pred c represents the predicted sample of the current color difference block, and i and j represent the horizontal and vertical position coordinates of the sample within the current color difference block. i is included in the range from 0 to the width of the current color difference block, and j is included in the range from 0 to the height of the current color difference block. The coordinate (0, 0) represents the position of the top-left sample within the current color difference block. rec c represents the restored sample within the reference color difference block, and scaledRV x and scaledRV y represents the horizontal and vertical components of the reference vector scaled according to the color format. Coordinates (i - scaledRV x , j - scaledRV y ) indicates the sample location within the reference color difference block.

[0286] As an example, a restored block is determined based on the block vector of a color difference block within a restored area of ​​the current picture, and the restored block can be used as a prediction block for the color difference block.

[0287] As an example, a restored block is determined based on the motion vector of a color difference block within a restored area of ​​a reference picture, and the restored block can be used as a prediction block for the color difference block.

[0288] If it is determined that the current prediction method for the color difference block is a model-based prediction method rather than a default prediction method, a prediction model for the color difference block is created (S1330).

[0289] The prediction model of the color difference block can be inherited from the corresponding luminance block, or derived based on the template of the reference color difference block indicated by the reference vector of the color difference block and the template of the current color difference block.

[0290] Information regarding whether the prediction model of the color difference block is inherited from the luminance block or derived based on a template is included in a parameter set applied to multiple pictures, such as SPS, PPS, etc., and can be signaled from the video encoding device to the video decoder.

[0291] When the model-based prediction method is a linear model-based prediction method, the linear model of the chrominance block can be defined as Equation 5.

[0292] [Mathematical Formula 5]

[0293] pred c (i, j) = α×rec c (i - scaledRV x , j - scaledRv y ) + β

[0294] In mathematical equation 5, pred c represents the predicted samples within the current color difference block as the output of the linear model, and rec c represents the reconstructed samples within the reference chrominance block as input to the linear model. Parameters α and β represent the scaling parameter and offset parameter, respectively, i and j represent the current location coordinates within the chrominance block, and scaledRV x and scaledRV y represents a scaled reference vector to indicate the location of the reference color difference block. i is included in the range from 0 to the width of the current color difference block, and j is included in the range from 0 to the height of the current color difference block. The coordinate (0, 0) indicates the location of the top-left sample within the current color difference block.

[0295] In one embodiment, when a luminance block corresponding to a chrominance block in a single tree structure is predicted by a linear model-based prediction method, the parameters of the linear model of the chrominance block may be inherited from the corresponding luminance block.

[0296] In one embodiment, in a dual tree structure, if a chrominance block inherits a reference vector of a specific sub-region within a luminance block corresponding to the chrominance block, and the specific sub-region is predicted using a linear model-based prediction method, the chrominance block may inherit the linear model of the specific sub-region. For example, if a reference vector-based predictor for a luminance sub-region within a corresponding luminance block is derived using a reference vector and a linear model, and the luminance sub-region is predicted by blending multiple prediction regions including the reference vector-based predictor, and the chrominance block inherits a reference vector from the luminance sub-region, the chrominance block may inherit the parameters of the linear model of the luminance sub-region. In another example, model inheritance in a dual tree structure may be applied equally to a default prediction method and a convolutional model-based prediction method.

[0297] In one embodiment, if the number of sub-regions to which a linear model-based prediction method is applied among a plurality of sub-regions within a luminance block corresponding to a color difference block in a dual tree structure is greater than a threshold, the parameters of the linear model of the color difference block can be inherited from the sub-regions.

[0298] In one embodiment, the parameters of the linear model may be derived based on an error minimization technique between the template region of the current color difference block and the template region of the reference color difference block. As template regions used to derive the prediction model of the color difference block, the regions shown in FIG. 11 may be used.

[0299] Error minimization techniques to derive the parameters α and β of a linear model may refer to Mean Square Error (MSE) optimization or Rate-Distortion (RDO) optimization.

[0300] In MSE optimization, the values ​​of the linear model parameters are set such that the error between the predicted samples and the reconstructed samples in the template region of the current chrominance block is minimized.

[0301] The error minimization technique can be expressed as Equation 6.

[0302] [Mathematical Formula 6]

[0303]

[0304] In mathematical equation 6, ref t (i, j, α, β) represents the result of predicting the sample at position (i, j) within the template region of the current color difference block from the template region of the reference color difference block using a linear model with model parameters α and β, and rec t (i, j) represents a restored sample within the template area of ​​the current color difference block. i and j represent location coordinates within the current template area. In the case of Fig. 11a, (i, j) may be included in the ranges {-N ≤ i < 0, 0 ≤ j < H-1} and {-N ≤ i < W, -M ≤ j < 0}. (0, 0) represents the top-left sample within the current color difference block.

[0305] ref t (i, j, α, β) can be defined as in Equation 7. In Equation 7, rec t (i - scaledRV x , j - scaledRv y ) represents a reference sample at a location located a distance of the reference vector from the current color difference block.

[0306] [Mathematical Formula 7]

[0307] ref t (i, j, α, β) = α×rec t (i - scaledRV x , j - scaledRv y ) + β

[0308] The parameters of the linear model are derived using mathematical equations 6 and 7, and the current color difference block can be predicted from the reference color difference block using mathematical equation 5.

[0309] In one embodiment, downsampling of the template region may be performed to simplify the derivation of the linear model. Some samples are selected from the template region of the current chrominance block and the template region of the reference chrominance block, respectively, and the parameters of the linear model can be derived by applying an error minimization technique to the selected samples.

[0310] FIG. 16 is a drawing for illustrating sampling of a template area according to one embodiment of the present disclosure.

[0311] In FIG. 16, four red samples located at predefined positions within the template area of ​​the current color difference block and four red samples within the template area of ​​the reference color difference block can be used to derive a linear model. The red samples may represent restored samples. Specifically, by applying a linear model to the four red samples within the reference template area, prediction samples of the current template area are derived, and the values ​​of the parameters of the linear model can be determined such that the error between the restored samples of the current template area and the derived prediction samples is minimized.

[0312] The locations and number of samples used to derive the linear model may be predefined or signaled in the image encoding device and image decoder.

[0313] The location and number of samples can be derived based on the current size of the color difference block, the shape of the template, and the area of ​​the template. For example, when the width of the color difference block is greater than the height, the shape of the template is as in FIG. 11b, and the top area of ​​the template is greater than the threshold, more than 4 samples can be used.

[0314] In one embodiment, the parameters of the linear model can be derived based on the statistical values ​​of samples within the current template region and the reference template region.

[0315] The parameters of the linear model can be derived based on Equation 8.

[0316] [Mathematical Formula 8]

[0317]

[0318] FIG. 17 is a diagram illustrating the derivation of a linear model based on statistical values ​​according to one embodiment of the present disclosure.

[0319] In Fig. 17, for the derivation of linear model parameters, the representative values ​​x of u samples with large values ​​within the template region of the reference chrominance block are used. ref_a , the representative value x of u samples with small values ​​within the template area of ​​the reference color difference block ref_b , the representative value x of the u samples with large values ​​within the template area of ​​the current color difference block cur_a , the representative value x of u samples with small values ​​within the current color difference block's template area cur_b This can be used.

[0320] u is an integer greater than or equal to 1, and the locations and number of samples used to derive the linear model may be predefined or signaled in the image encoding device and image decoder. Representative values ​​may be the mean, median, middle, weighted average, etc.

[0321] When the model-based prediction method is a convolution model-based prediction method, the convolution model of the chrominance block can be defined as Equation 9.

[0322] [Mathematical Formula 9]

[0323] pred c (i, j) = c0×C + c1×N + c2×S + c3×E + c4×W + c T ×P + c6×B

[0324] In mathematical equation 9, pred c represents the predicted sample within the current color difference block as the output of the convolution model. c0, c1, c2, c3, c4, c5, and c6 represent the parameters of the convolution model. C represents the reference color difference sample within the reference color difference block. N, S, E, and W represent the adjacent samples above, below, to the right, and to the left of the reference color difference sample C, respectively.

[0325] C can be expressed as in mathematical formula 10.

[0326] [Mathematical Formula 10]

[0327] C(i, j) = rec c (i - scaledRV x , j - scaledRV y )

[0328] In mathematical formula 10, rec c represents the restored sample. i is included in the range from 0 to the width of the current color difference block, and j is included in the range from 0 to the height of the current color difference block.

[0329] FIG. 18 is a diagram showing the positional relationship of input samples of a convolution model according to one embodiment of the present disclosure.

[0330] P is a nonlinear term of the convolution model and can be expressed as the square of C. P can be defined as in Equation 11. In Equation 11, bitdepth represents the bit depth of the video coding. For example, bitdepth can be 8 bits or 10 bits.

[0331] [Mathematical Formula 11]

[0332] P = (C×C + 2bitdepth-1) >> bitdepth

[0333] B is a bias term of the convolution model and can be an intermediate value of the bit depth. B can be defined as in Equation 12.

[0334] [Mathematical Formula 12]

[0335] B = 2bitdepth-1

[0336] In other embodiments, some of the terms regarding C, N, E, W, P, and B may be omitted.

[0337] In one embodiment, when a luminance block corresponding to a chrominance block in a single tree structure is predicted by a convolution model-based prediction method, the parameters of the convolution model of the chrominance block can be inherited from the corresponding luminance block.

[0338] In one embodiment, in a dual tree structure, a chrominance block inherits a reference vector of a specific sub-region within a luminance block corresponding to the chrominance block, and if the specific sub-region is predicted using a convolution model-based prediction method, the chrominance block may inherit the convolution model of the specific sub-region.

[0339] In one embodiment, if the number of sub-regions to which a convolution model-based prediction method is applied among a plurality of sub-regions within a luminance block corresponding to a chrominance block in a dual tree structure is greater than a threshold, the parameters of the convolution model of the chrominance block can be inherited from the sub-regions.

[0340] In one embodiment, the parameters of the convolution model may be derived based on an error minimization technique between the template region of the current color difference block and the template region of the reference color difference block. As template regions used to derive the prediction model of the color difference block, the regions shown in FIG. 11 may be used.

[0341] Error minimization techniques for deriving the parameters of a convolutional model may refer to Mean Square Error (MSE) optimization or Rate-Distortion (RDO) optimization.

[0342] In MSE optimization, the values ​​of the convolution model parameters are set such that the error between the predicted samples and the reconstructed samples in the template region of the current chrominance block is minimized.

[0343] The error minimization technique can be expressed as Equation 13.

[0344] [Mathematical Formula 13]

[0345]

[0346] In mathematical formula 13, ref t_conv (i, j; c0, c1, c2, c3, c4, c5, c6) represents the result of predicting the sample at position (i, j) within the template region of the current color difference block from the template region of the reference color difference block using a convolutional model. rec t (i, j) represents the restored sample at position (i, j) within the template area of ​​the current color difference block. i and j represent location coordinates within the current template area. In the case of Fig. 11a, (i, j) may be included in the ranges {-N ≤ i < 0, 0 ≤ j < H-1} and {-N ≤ i < W, -M ≤ j < 0}. (0, 0) represents the top-left sample within the current color difference block.

[0347] ref t_conv (i, j; c0, c1, c2, c3, c4, c5, c6) can be defined using mathematical formulas 14 and 15.

[0348] [Mathematical Formula 14]

[0349] ref t_conv (i, j; c0, c1, c2, c3, c4, c5, c6) = c0×C t + c1×N t + c2×S t + c3×E t + c4×W t + c T ×P t + c6×B t

[0350] [Mathematical Formula 15]

[0351] C t (i, j) = rec t (i - scaledRV x , j - scaledRv y )

[0352] In mathematical formulas 14 and 15, C t represents a sample within the reference template area, and rec t (i - scaledRV x, j - scaledRv y ) represents a reference sample at a location located a distance of the reference vector from the current color difference block.

[0353] A prediction model of the color difference block is generated through the aforementioned process.

[0354] Afterwards, the current color difference block is predicted by applying the prediction model of the color difference block to the reference color difference block (S1340).

[0355] When a linear model is derived, the current color difference block can be predicted by applying the linear model to reference samples within the reference color difference block indicated by the reference vector, as in Equation 5.

[0356] When a convolution model is derived, the current chrominance block is predicted by applying the convolution model to the reference chrominance block. The convolution model can be applied on a sample basis.

[0357] FIGS. 19a and FIGS. 19b are drawings for explaining the application of a convolution model according to one embodiment of the present disclosure.

[0358] In FIG. 19a, it is illustrated that one sample in the current color difference block is predicted by applying a convolution model to reference samples in the reference color difference block.

[0359] In FIG. 19b, it is shown that all samples of the current color difference block are predicted by applying a convolution model to the reference color difference block.

[0360] Although the flowcharts and timing diagrams in this specification describe each process as being executed sequentially, this is merely an illustrative explanation of the technical concept of one embodiment of the present disclosure. In other words, a person skilled in the art to which one embodiment of the present disclosure belongs may modify and adapt the flowcharts and timing diagrams in various ways, such as changing the order described in the flowcharts and timing diagrams or executing one or more of the processes in parallel, without departing from the essential characteristics of one embodiment of the present disclosure; therefore, the flowcharts and timing diagrams are not limited to a chronological order.

[0361] It should be understood that the exemplary embodiments described above may be implemented in many different ways. The functions or methods described in one or more examples may be implemented in hardware, software, firmware, or any combination thereof. It should be understood that the functional components described herein are labeled as "...unit" to particularly emphasize their implementation independence.

[0362] Meanwhile, the various functions or methods described in the present embodiment may be implemented as instructions stored in a non-transient recording medium that can be read and executed by one or more processors. A non-transient recording medium includes, for example, any type of recording device in which data is stored in a form readable by a computer system. For example, a non-transient recording medium includes storage media such as an EPROM (erasable programmable read-only memory), a flash drive, an optical drive, a magnetic hard drive, and a solid-state drive (SSD).

[0363] The above description is merely an illustrative explanation of the technical concept of the present embodiment, and a person skilled in the art to which the present embodiment belongs would be able to make various modifications and variations within the scope of the essential characteristics of the present embodiment. Accordingly, the present embodiments are intended to explain, not limit, the technical concept of the present embodiment, and the scope of the technical concept of the present embodiment is not limited by these embodiments. The scope of protection of the present embodiment shall be interpreted by the claims below, and all technical concepts within an equivalent scope shall be interpreted as being included within the scope of rights of the present embodiment.

[0364] CROSS-REFERENCE TO RELATED APPLICATION

[0365] This patent application claims priority to Korean patent application No. 10-2025-0039848 filed on March 27, 2025, the entire contents of which are incorporated into this patent application by reference.

Claims

1. In an image decoding method using a reference vector-based chrominance prediction mode, A process of deriving a reference vector of the current color difference block by scaling a reference vector of a luminance region corresponding to the current color difference block based on a color difference subsampling format; and A process of predicting the current color difference block using the reference vector of the current color difference block. Includes, The above-mentioned corresponding luminance area is, An image decoding method for a predicted region by blending a plurality of prediction blocks including a prediction block derived using a reference vector of the above-mentioned luminance region.

2. In Paragraph 1, The above plurality of prediction blocks are, Image decoding method including a prediction block derived using an intra prediction mode.

3. In Paragraph 1, The above-mentioned corresponding luminance area is, An image decoding method comprising a luminance block corresponding to the current chrominance block in a single tree structure.

4. In Paragraph 1, The above-mentioned corresponding luminance area is, An image decoding method comprising a region where a reference vector is stored among a plurality of subregions within a luminance block corresponding to the current color difference block in a dual tree structure.

5. In Paragraph 4, When multiple reference vectors are stored in the multiple sub-regions mentioned above, a process of generating a list of candidate reference vectors based on the multiple reference vectors; A process of sorting the reference vector candidates based on the template matching costs of the reference vector candidates within the above reference vector candidate list; and The process of deriving the reference vector of the current color difference block from the above-mentioned aligned reference vector candidates. A video decoding method including 6. In Paragraph 1, The process of predicting the current color difference block mentioned above is, The process of deriving the prediction model of the current color difference block above; and A process of generating a prediction block of the current color difference block by applying the prediction model to a reference color difference block indicated by the reference vector of the current color difference block. A video decoding method including 7. In Paragraph 6, The process of predicting the current color difference block mentioned above is, A process of obtaining information from a bitstream regarding whether a prediction model of the current color difference block is used to predict the current color difference block; and A process of determining whether to use the prediction model to predict the current color difference block based on information regarding whether the prediction model is used to predict the current color difference block. A video decoding method including additional 8. In Paragraph 6, The process of predicting the current color difference block mentioned above is, A process of determining that a prediction model is used to predict the current color difference block from a reference color difference block indicated by the reference vector of the current color difference block in a dual tree structure. A video decoding method including additional 9. In Paragraph 6, The process of predicting the current color difference block mentioned above is, A process of determining that, in a dual tree structure, if the number of sub-regions predicted using a luminance reference vector and a prediction model among a plurality of sub-regions within a luminance block corresponding to the current color difference block is greater than or equal to a threshold, the prediction model is used to predict the current color difference block from the reference color difference block indicated by the reference vector. A video decoding method including additional 10. In Paragraph 6, The process of deriving the above prediction model is, When the luminance region corresponding to the above current color difference block is predicted using a luminance reference vector and a prediction model, the process of inheriting the prediction model of the said luminance region as the prediction model of the said current color difference block. A video decoding method including 11. In Paragraph 6, The process of deriving the above prediction model is, A process of deriving prediction samples within the template area of ​​the current color difference block by applying the prediction model to the template area of ​​the reference color difference block; and A process of determining the parameters of the prediction model so that the error between the predicted samples and the restored samples within the template area of ​​the current color difference block is minimized. A video decoding method including 12. In Paragraph 1, The above current color difference block prediction model is, An image decoding method comprising at least one of a linear model or a convolutional model.

13. In an image encoding method using a reference vector-based chrominance prediction mode, A process of deriving a reference vector of the current color difference block by scaling a reference vector of a luminance region corresponding to the current color difference block based on a color difference subsampling format; and A process of predicting the current color difference block using the reference vector of the current color difference block. A video decoding method including 14. A method for providing image data to an image decoding device, A process of encoding the above image data to generate a bitstream; and The process includes transmitting the above bitstream to the above video decoder, and The process of generating the above bitstream is, A process of deriving a reference vector of the current color difference block by scaling a reference vector of a luminance region corresponding to the current color difference block based on a color difference subsampling format; and A process of predicting the current color difference block using the reference vector of the current color difference block. A method including