Method and apparatus for video coding with local illumination compensation-based prediction in geometric partition mode

By applying local illumination compensation in geometric partitioning mode, the problem of increased encoded data volume in existing technologies is solved, encoding and decoding efficiency is improved, and video quality is enhanced, especially in the case of local or global illumination changes.

CN121128178APending Publication Date: 2025-12-12HYUNDAI MOTOR CO LTD +2
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202480032789.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-05-07
Filing Date
2024-05-08
Publication Date
2025-12-12

AI Technical Summary

Technical Problem

Existing video encoding and decoding technologies face challenges when image size, resolution, and frame rate increase, leading to a larger amount of encoded data and a need for higher encoding and decoding efficiency and improved image enhancement effects. In particular, in geometric partitioning mode, it is difficult to effectively utilize local lighting compensation to handle local or global lighting changes.

Method used

By applying local illumination compensation in geometric partitioning mode, a reference block is generated using motion information, and local illumination compensation is performed by mixing the linear relationship parameters of adjacent reference pixels to generate the final predicted block of the current block.

Benefits of technology

It improves video encoding and decoding efficiency and enhances video quality by effectively compensating for local or global lighting changes between the current frame and the reference frame.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121128178A_ABST
    Figure CN121128178A_ABST
Patent Text Reader

Abstract

The present embodiment provides a video coding and decoding method and apparatus using local illumination compensation-based prediction in geometric partition mode. In this embodiment, an image decoding device decodes a geometric partition mode index, and partitions a current block into partitioned blocks according to the geometric partition mode index. An image decoding device generates reference blocks of partitioned blocks by using motion information, and generates a prediction block of a current block by mixing the reference blocks. When adjacent reference pixels of a current block are available, and all or some of the adjacent reference pixels of each reference block are available, the image decoding device generates adjacent reference pixels of a mixed prediction block by mixing the adjacent reference pixels of the reference blocks, and deriving a parameter of a linear relationship expression between adjacent reference pixels of the current block and adjacent reference pixels of the hybrid prediction block. The image decoding apparatus generates a final prediction block of the current block by applying local illumination compensation to the mixed prediction block using parameters of the linear relationship expression.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to a video coding method and apparatus using prediction based on local illumination compensation in a geometric partition mode. BACKGROUND

[0002] The statements in this section merely provide background information related to the present disclosure and can not constitute the prior art.

[0003] Since video data has a large amount of data compared to audio data or still image data, a large amount of hardware resources including memory are required to store or transmit uncompressed video data.

[0004] Accordingly, an encoder is generally used to compress and store or transmit video data. A decoder receives compressed video data, decompresses the received compressed video data, and plays the decompressed video data. Video compression techniques include H.264 / Advanced Video Coding (AVC), High Efficiency Video Coding (HEVC), and Versatile Video Coding (VVC) which improves coding efficiency by about 30% or more than HEVC.

[0005] However, as the image size, resolution, and frame rate gradually increase, the amount of data to be encoded also increases. Accordingly, there is a need to provide a new compression technique that has higher coding efficiency and improved image enhancement effects than existing compression techniques.

[0006] In VVC, a geometric partitioning mode (GPM) is adopted as an inter prediction technique to perform prediction according to a more flexible partitioning than square and rectangular partitioning. In the GPM technique, the encoder provides the decoder with a mode index indicating one of the pre-defined geometric partitioning modes and motion information of two partitioned regions. The decoder uses the transmitted information to compensate for the motion of the two partitioned regions and uses a blending matrix according to the geometric partitioning mode to weight-sum the two partitioned regions, thereby generating a final prediction signal of the current block.

[0007] When generating a prediction signal of a current block by using motion compensation based on inter prediction, local or global illumination changes can occur between the current frame and the reference frame. Local illumination compensation (LIC) is a technique that can improve the prediction efficiency of the current block by compensating for the changed illumination value. Therefore, in order to improve video coding efficiency and video quality, a method of effectively using illumination compensation when performing inter prediction based on GMP is required. SUMMARY

[0008] TECHNICAL PROBLEM The present disclosure is directed to a video coding method and apparatus that applies local illumination compensation to a prediction signal of a current block to compensate for local or global illumination changes between a current frame and a reference frame when performing prediction of the current block according to a geometric partition mode.

[0009] TECHNICAL SOLUTION At least one aspect of the present disclosure provides a method of reconstructing a current block by a video decoding apparatus. The method includes decoding a geometric partition mode index from a bitstream, and partitioning the current block into partitioned blocks according to the geometric partition mode index. The method further includes generating reference blocks of the partitioned blocks by utilizing motion information. Here, the reference blocks include a first reference block and a second reference block. The method further includes generating a prediction block of the current block by blending the reference blocks of the partitioned blocks. The method further includes checking availability of neighboring reference pixels of the current block at an upper side and a left side of the current block. When the neighboring reference pixels of the current block are available, the method further includes checking availability of neighboring reference pixels of each of the reference blocks at an upper side and a left side of each of the reference blocks. When the neighboring reference pixels of each of the reference blocks are all or partially available, the method further includes generating blended neighboring reference pixels of the prediction block by blending the neighboring reference pixels of the reference blocks. The method further includes deriving parameters of a linear relationship between the neighboring reference pixels of the current block and the blended neighboring reference pixels of the prediction block. The method further includes generating a final prediction block of the current block by applying local illumination compensation to the prediction block based on the parameters of the linear relationship.

[0010] Another aspect of the disclosure provides a method of encoding a current block by a video encoding device. The method includes determining a geometric partition mode index, and partitioning the current block into partitioned blocks according to the geometric partition mode index. The method also includes generating reference blocks for the partitioned blocks by utilizing motion information. Here, the reference blocks include a first reference block and a second reference block. The method also includes generating a prediction block for the current block by blending the reference blocks for the partitioned blocks. The method also includes checking availability of neighboring reference pixels of the current block at the top and left sides of the current block. When the neighboring reference pixels of the current block are available, the method further includes checking availability of neighboring reference pixels of each of the reference blocks at the top and left sides of each of the reference blocks. When the neighboring reference pixels of each of the reference blocks are all or partially available, the method further includes generating blended neighboring reference pixels of the prediction block by blending the neighboring pixels of the reference blocks. The method further includes deriving parameters of a linear relationship between the neighboring reference pixels of the current block and the blended neighboring reference pixels of the prediction block. The method further includes generating a first final prediction block for the current block by applying local illumination compensation to the prediction block based on the parameters of the linear relationship.

[0011] Yet another aspect of the disclosure provides a method of providing video data to a video decoding device. The method includes encoding the video data into a bitstream, and transmitting the bitstream to the video decoding device. The encoding of the video data includes determining a geometric partition mode index, and partitioning a current block into partitioned blocks according to the geometric partition mode index. The encoding of the video data also includes generating reference blocks for the partitioned blocks by utilizing motion information. Here, the reference blocks include a first reference block and a second reference block. The encoding of the video data also includes generating a prediction block for the current block by blending the reference blocks for the partitioned blocks. The encoding of the video data also checks availability of neighboring reference pixels of the current block at the top and left sides of the current block. When the neighboring reference pixels of the current block are available, the encoding of the video data further includes checking availability of neighboring reference pixels of each of the reference blocks at the top and left sides of each of the reference blocks. When the neighboring reference pixels of each of the reference blocks are all or partially available, the encoding of the video data further includes generating blended neighboring reference pixels of the prediction block by blending the neighboring pixels of the reference blocks. The encoding of the video data further includes deriving parameters of a linear relationship between the neighboring reference pixels of the current block and the blended neighboring reference pixels of the prediction block. The encoding of the video data further includes generating a final prediction block for the current block by applying local illumination compensation to the prediction block based on the parameters of the linear relationship.

[0012] Advantages As described above, the present application provides a video coding method and apparatus that applies local illumination compensation to a prediction signal of a current block to compensate for local or global illumination changes between a current frame and a reference frame when performing prediction on the current block according to a geometric partition mode. Accordingly, the video coding method and apparatus improves video coding efficiency and enhances video quality. BRIEF DESCRIPTION OF DRAWINGS

[0013] FIG. 1 is a block diagram of a video encoding apparatus to which the present technology can be applied.

[0014] FIG. 2 illustrates a method of partitioning a block using a quad-tree plus binary tree ternary tree (QTBTTT) structure.

[0015] FIGS. 3a and 3b illustrate a plurality of intra prediction modes including a wide angle intra prediction mode.

[0016] FIG. 4 illustrates neighboring blocks of a current block.

[0017] FIG. 5 is a block diagram of a video decoding apparatus to which the present technology can be applied.

[0018] FIG. 6 is a diagram illustrating block partitioning according to geometric partitioning.

[0019] FIGS. 7a and 7b are diagrams illustrating a straight line that splits a block into two parts.

[0020] FIG. 8 is a diagram illustrating a geometric partition mode (GPM) merge list for geometric motion prediction.

[0021] FIG. 9 is a block diagram detailing a portion of a video decoding apparatus according to at least one embodiment of the present application.

[0022] FIG. 10 is a diagram illustrating blending of neighboring reference pixels of a reference block according to at least one embodiment of the present application.

[0023] FIG. 11 is a diagram illustrating blending of neighboring reference pixels of a reference block according to another embodiment of the present application.

[0024] FIG. 12 is a diagram illustrating blending of neighboring reference pixels of a reference block according to yet another embodiment of the present application.

[0025] FIGS. 13a and 13b are diagrams illustrating blending masks of reference pixels according to at least one embodiment of the present application.

[0026] FIGS. 14a and 14b are diagrams illustrating blending masks of reference pixels according to another embodiment of the present application.

[0027] FIG. 15 is a flowchart of a method of encoding a current block by a video encoding apparatus according to at least one embodiment of the present application.

[0028] FIG. 16 is a flowchart of a method of reconstructing a current block by a video decoding apparatus according to at least one embodiment of the present application. DETAILED DESCRIPTION

[0029] Hereinafter, some embodiments of the present application will be described in detail with reference to the accompanying drawings. In the following description, the same drawing reference numerals are used for the same elements, even in different drawings. Further, in the following description of some embodiments, detailed descriptions of known components and functions incorporated herein will be omitted for clarity and conciseness, when deemed that the detailed description of related known components and functions obfuscates the subject matter of the present application.

[0030] FIG. 1 is a block diagram of a video encoding apparatus to which the present application is applied. Hereinafter, the video encoding apparatus and components of the apparatus will be described with reference to the drawing.

[0031] The encoding apparatus can include a picture partitioner 110, a predictor 120, a subtractor 130, a transformer 140, a quantizer 145, a rearrangement unit 150, an entropy encoder 155, an inverse quantizer 160, an inverse transformer 165, an adder 170, a loop filter unit 180, and a memory 190.

[0032] Each component of the encoding apparatus can be implemented as hardware or software, or a combination of hardware and software. In addition, the function of each component can be implemented as software, and a microprocessor can also be implemented to perform the function of software corresponding to each component.

[0033] A video is composed of one or more sequences including a plurality of pictures. Each picture is divided into a plurality of regions, and encoding is performed for each region. For example, a picture is divided into one or more tiles or / and slices. Here, one or more tiles can be defined as a tile group. Each tile or / and slice is divided into one or more coding tree units (CTU). In addition, each CTU is divided into one or more coding units (CU) by a tree structure. Information applied to each coding unit (CU) is coded as a syntax of the CU, and information commonly applied to CUs included in one CTU is coded as a syntax of the CTU. In addition, information commonly applied to all blocks in one slice is coded as a syntax of a slice header, and information applied to all blocks constituting one or more pictures is coded as a picture parameter set (PPS) or a picture header. Furthermore, information commonly referred to by a plurality of pictures is coded as a sequence parameter set (SPS). In addition, information commonly referred to by one or more SPSs is coded as a video parameter set (VPS). Furthermore, information commonly applied to one tile or tile group can also be coded as a syntax of a tile or tile group header. Syntaxes included in the SPS, PPS, slice header, tile or tile group header can be referred to as high-level syntaxes.

[0034] The picture partitioner 110 determines the size of a coding tree unit (CTU). Information on the size of the CTU (CTU size) is coded as a syntax of the SPS or PPS, and is transmitted to the video decoding apparatus.

[0035] The picture partitioner 110 divides each picture constituting a video into a plurality of coding tree units (CTUs) having a predetermined size, and then recursively divides the CTUs by using a tree structure. A leaf node in the tree structure becomes a coding unit (CU), which is a basic unit of encoding.

[0036] The tree structure can be a quadtree (QT) in which a higher node (or parent node) is split into four lower nodes (or child nodes) having the same size. The tree structure can also be a binary tree (BT) in which a higher node is split into two lower nodes. The tree structure can also be a ternary tree (TT) in which a higher node is split into three lower nodes in a ratio of 1:2:1. The tree structure can also be a structure in which two or more of the QT structure, the BT structure, and the TT structure are mixed. For example, a quadtree plus binary tree (QTBT) structure can be used, or a quadtree plus binary tree ternary tree (QTBTTT) structure can be used. Here, a binary tree ternary tree (BTTT) is added to the tree structure to be called a multiple-type tree (MTT).

[0037] FIG. 2 is a schematic diagram for describing a method of splitting a block by using a QTBTTT structure.

[0038] As shown in FIG. 2, a CTU can be first split into a QT structure. The quadtree splitting can be recursive until the size of the split block reaches a minimum block size of a leaf node allowed in the QT (MinQTSize). A first flag (QT_split_flag) indicating whether each node of the QT structure is split into four nodes of a lower layer is encoded by the entropy encoder 155 and signaled to the video decoding apparatus. When a leaf node of the QT is not greater than a maximum block size of a root node allowed in the BT (MaxBTSize), the leaf node can be further split by at least one of a BT structure or a TT structure. There can be multiple splitting directions in the BT structure and / or the TT structure. For example, there can be two directions, i.e., a direction of splitting a block of a corresponding node horizontally and a direction of splitting the block of the corresponding node vertically. As shown in FIG. 2, when the MTT splitting starts, a second flag (mtt_split_flag) indicating whether a node is split, and a flag additionally indicating a splitting direction (vertical or horizontal) and / or a flag indicating a splitting type (binary or ternary) in the case where the node is split are encoded by the entropy encoder 155 and signaled to the video decoding apparatus.

[0039] Alternatively, a CU split flag (split_cu_flag) indicating whether a node is split can also be encoded before a first flag (QT_split_flag) indicating whether each node is split into four nodes of a lower layer is encoded. When a value of the CU split flag (split_cu_flag) indicates that each node is not split, blocks of the corresponding node become leaf nodes in the split tree structure and become a CU, which is a basic unit of encoding. When a value of the CU split flag (split_cu_flag) indicates that each node is split, the video encoding apparatus starts encoding the first flag first in the above-described scheme.

[0040] When QTBT is used as another example of a tree structure, there can be two types, i.e., a type in which a block of a corresponding node is horizontally split into two blocks having the same size (i.e., symmetric horizontal split) and a type in which the block of the corresponding node is vertically split into two blocks having the same size (i.e., symmetric vertical split). A split flag (split_flag) indicating whether each node of the BT structure is split into blocks of a lower layer and split type information indicating a split type are encoded by the entropy encoder 155 and are transmitted to the video decoding apparatus. On the other hand, there can additionally be a type in which the block of the corresponding node is split into two blocks that are asymmetric to each other. The asymmetric form can include a form in which the block of the corresponding node is split into two rectangular blocks having a size ratio of 1:3, or can further include a form in which the block of the corresponding node is split in a diagonal direction.

[0041] A CU can have various sizes according to QTBT or QTBTTT splitting from a CTU. Hereinafter, a block corresponding to a CU to be encoded or decoded (i.e., a leaf node of the QTBTTT) is referred to as a "current block". When QTBTTT splitting is employed, a shape of the current block can be a rectangular shape in addition to a square shape.

[0042] The predictor 120 predicts the current block to generate a prediction block. The predictor 120 includes an intra predictor 122 and an inter predictor 124.

[0043] In general, each of the current blocks in a picture can be predictively encoded. In general, the prediction of the current block can be performed by utilizing an intra prediction technique that utilizes data from the picture including the current block or an inter prediction technique that utilizes data from pictures that are encoded before the picture including the current block. The inter prediction includes both uni-prediction and bi-prediction.

[0044] The intra predictor 122 predicts pixels in the current block by utilizing pixels (reference pixels) located adjacent to the current block in the current picture including the current block. There are a plurality of intra prediction modes according to a prediction direction. For example, as shown in FIG. 3a, the plurality of intra prediction modes can include two non-directional modes including a Planar mode and a DC mode, and can include 65 directional modes. The adjacent pixels and algorithmic equations to be used are defined differently according to each prediction mode.

[0045] In order to efficiently perform directional prediction on the current block having a rectangular shape, the directional modes (intra prediction modes #67 to #80, #-1 to #-14) shown by dotted arrows in FIG. 3b can be additionally used. The directional modes can be referred to as "wide angle intra-prediction modes". In FIG. 3b, the arrows indicate respective reference samples used for prediction, not representative of the prediction direction. The prediction direction is opposite to the direction indicated by the arrow. When the current block has a rectangular shape, the wide angle intra-prediction mode is a mode in which prediction is performed in a direction opposite to a specific directional mode without additional bit transmission. In this case, in the wide angle intra-prediction mode, some of the wide angle intra-prediction modes available for the current block can be determined by a ratio of the width to the height of the current block having a rectangular shape. For example, when the current block has a rectangular shape having a height smaller than a width, the wide angle intra-prediction modes (intra prediction modes #67 to #80) having an angle smaller than 45 degrees are available. When the current block has a rectangular shape having a width greater than a height, the wide angle intra-prediction modes (intra prediction modes -1 to -14) having an angle greater than -135 degrees are available.

[0046] The intra predictor 122 can determine an intra prediction to be used for encoding the current block. In some examples, the intra predictor 122 can encode the current block by utilizing a plurality of intra prediction modes, and can also select an appropriate intra prediction mode to be used from among the test modes. For example, the intra predictor 122 can calculate rate-distortion values by utilizing rate-distortion analysis on a plurality of test intra prediction modes, and can also select an intra prediction mode having the best rate-distortion characteristics among the test modes.

[0047] The intra predictor 122 selects one of the plurality of intra prediction modes, and predicts the current block by utilizing the adjacent pixels (reference pixels) and algorithmic equations determined according to the selected intra prediction mode. Information on the selected intra prediction mode is encoded by the entropy encoder 155 and transmitted to the video decoding apparatus.

[0048] The inter predictor 124 generates a prediction block of the current block by using a motion compensation process. The inter predictor 124 searches for a block most similar to the current block among reference pictures that have been encoded and decoded earlier than the current picture, and generates a prediction block of the current block by using the searched block. In addition, a motion vector (MV) corresponding to a displacement between the current block in the current picture and the prediction block in the reference picture is generated. Generally, motion estimation is performed on a luma component, and the motion vector calculated based on the luma component is used for both the luma component and a chroma component. Motion information including information of the reference picture and information on the motion vector used to predict the current block is encoded by the entropy encoder 155 and transmitted to the video decoding apparatus.

[0049] The inter predictor 124 can also perform interpolation of a reference picture or a reference block to increase the accuracy of prediction. In other words, a sub-sample is interpolated between two consecutive integer samples by applying a filter coefficient to a plurality of consecutive integer samples including the two integer samples. When performing a process of searching for a block most similar to the current block with respect to an interpolated reference picture, a decimal unit precision can be expressed for a motion vector instead of an integer sample unit precision. The precision or resolution of the motion vector can be differently set for each target region to be encoded, for example, a unit such as a slice, a tile, a CTU, a CU, or the like. When such adaptive motion vector resolution (AMVR) is applied, information on the motion vector resolution to be applied to each target region should be signaled for each target region. For example, when the target region is a CU, information on the motion vector resolution applied to each CU is signaled. The information on the motion vector resolution can be information representing the precision of a motion vector difference described below.

[0050] On the other hand, the inter predictor 124 can perform inter prediction by using bi-prediction. In the case of bi-prediction, two reference pictures and two motion vectors representing the most similar block position to the current block in each of the reference pictures are used. The inter predictor 124 selects a first reference picture and a second reference picture from a reference picture list 0 (RefPicListO) and a reference picture list 1 (RefPicListl), respectively. The inter predictor 124 also searches for blocks most similar to the current block in the respective reference pictures to generate a first reference block and a second reference block. Further, a prediction block of the current block is generated by averaging or weighted-averaging the first reference block and the second reference block. In addition, motion information including information on the two reference pictures used for predicting the current block and including information on the two motion vectors is transmitted to the entropy encoder 155. Here, the reference picture list 0 can be constituted by pictures in the pre-reconstructed pictures that precede the current picture in display order, and the reference picture list 1 can be constituted by pictures in the pre-reconstructed pictures that follow the current picture in display order. However, although not particularly limited thereto, a pre-reconstructed picture that follows the current picture in display order can be additionally included in the reference picture list 0. Conversely, a pre-reconstructed picture that precedes the current picture can also be additionally included in the reference picture list 1.

[0051] In order to minimize the amount of bits consumed for encoding the motion information, various methods can be used.

[0052] For example, in the case where the reference picture and the motion vector of the current block are the same as those of a neighboring block, the information of the neighboring block can be identified to encode the motion information of the current block to be transmitted to the video decoding apparatus. This method is referred to as a merge mode.

[0053] In the merge mode, the inter predictor 124 selects a predetermined number of merge candidates (hereinafter, referred to as "merge candidates") from neighboring blocks of the current block.

[0054] As the neighboring blocks used to derive the merge candidates, all or some of a left block A0, a lower-left block Al, an upper block B0, an upper-right block Bl, and an upper-left block B2 adjacent to the current block in the current picture can be used, as illustrated in FIG. 4. In addition, blocks located within a reference picture (which can be the same as or different from the reference picture used to predict the current block) other than the current picture in which the current block is located can also be used as the merge candidates. For example, a co-located block of the current block within the reference picture or a block adjacent to the co-located block can be additionally used as the merge candidates. If the number of the merge candidates selected by the above-described method is less than a predetermined number, a zero vector is added to the merge candidates.

[0055] The inter predictor 124 configures a merge list including a predetermined number of merge candidates by using neighboring blocks. A merge candidate to be used as motion information of the current block is selected from among the merge candidates included in the merge list, and merge index information for identifying the selected candidate is generated. The generated merge index information is encoded by the entropy encoder 155 and transmitted to the video decoding apparatus.

[0056] The merge skip mode is a special case of the merge mode. After quantization, when all transform coefficients for entropy encoding are close to zero, only neighboring block selection information is transmitted without transmitting a residual signal. By using the merge skip mode, relatively high encoding efficiency can be achieved for images with slight motion, still images, screen content images, and the like.

[0057] Hereinafter, the merge mode and the merge skip mode are collectively referred to as a merge / skip mode.

[0058] Another method for encoding motion information is an advanced motion vector prediction (AMVP) mode.

[0059] In the AMVP mode, the inter predictor 124 derives a motion vector prediction candidate for a motion vector of the current block by using neighboring blocks of the current block. As the neighboring blocks for deriving the motion vector prediction candidate, all or some of the left block A0, the lower-left block Al, the upper block B0, the upper-right block Bl, and the upper-left block B2 adjacent to the current block in the current picture shown in FIG. 4 can be used. In addition, blocks located within a reference picture (which can be the same as or different from the reference picture used for predicting the current block) other than the current picture in which the current block is located can also be used as the neighboring blocks for deriving the motion vector prediction candidate. For example, a co-located block or a block adjacent to the co-located block within the reference picture of the current block can be used. If the number of motion vector candidates selected by the above-described method is less than a predetermined number, a zero vector is added to the motion vector candidates.

[0060] The inter predictor 124 derives a motion vector prediction candidate by using a motion vector of a neighboring block, and determines a motion vector prediction of a motion vector of the current block by using the motion vector prediction candidate. In addition, a motion vector difference is calculated by subtracting the motion vector prediction from the motion vector of the current block.

[0061] Motion vector prediction can be obtained by applying a predefined function (e.g., median and average value calculation, etc.) to the motion vector prediction candidates. In this case, the video decoding device also knows the predefined function. In addition, since the neighboring blocks used to derive the motion vector prediction candidates are blocks that have already been encoded and decoded, the video decoding device can also already know the motion vectors of the neighboring blocks. Therefore, the video encoding device does not need to encode information for identifying the motion vector prediction candidates. Accordingly, in this case, the information about the motion vector difference and the information about the reference picture used to predict the current block are encoded.

[0062] On the other hand, motion vector prediction can also be determined by a scheme of selecting any one of the motion vector prediction candidates. In this case, information for identifying the selected motion vector prediction candidate is additionally encoded together with the information about the motion vector difference and the information about the reference picture used to predict the current block.

[0063] The subtractor 130 generates a residual block by subtracting the prediction block generated by the intra predictor 122 or the inter predictor 124 from the current block.

[0064] The transformer 140 transforms a residual signal in the residual block having pixel values in a spatial domain into a transform coefficient in a frequency domain. The transformer 140 can transform the residual signal in the residual block by using the entire size of the residual block as a transform unit, or can also divide the residual block into a plurality of sub-blocks and can perform the transformation by using the sub-blocks as transform units. Alternatively, the residual block is divided into two sub-blocks, i.e., a transform region and a non-transform region, to transform the residual signal by using only the transform region sub-block as a transform unit. Here, the transform region sub-block can be one of two rectangular blocks having a size ratio of 1:1 based on a horizontal axis (or a vertical axis). In this case, a flag (cu_sbt_flag) indicating only the transform sub-block, and direction (vertical / horizontal) information (cu_sbt_horizontal_flag) and / or position information (cu_sbt_pos_flag) are encoded by the entropy encoder 155 and signaled to the video decoding device. In addition, the size of the transform region sub-block based on the horizontal axis (or the vertical axis) can have a size ratio of 1:3. In this case, a flag (cu_sbt_quad_flag) indicating division of the corresponding partition is additionally encoded by the entropy encoder 155 and signaled to the video decoding device.

[0065] On the other hand, the transformer 140 can perform a transform of the residual block separately in a horizontal direction and a vertical direction. For the transform, various types of transform functions or transform matrices can be used. For example, a pair of transform functions used for horizontal transform and vertical transform can be defined as a multiple transform set (MTS). The transformer 140 can select one transform function pair having the highest transform efficiency among the MTS, and can transform the residual block in each of the horizontal direction and the vertical direction. Information (mts_idx) about the transform function pair in the MTS is encoded by the entropy encoder 155 and signaled to the video decoding apparatus.

[0066] The quantizer 145 quantizes the transform coefficients output from the transformer 140 with a quantization parameter, and outputs the quantized transform coefficients to the entropy encoder 155. The quantizer 145 can also quantize the relevant residual block immediately without transforming any block or frame. The quantizer 145 can also apply different quantization coefficients (scaling values) according to the positions of the transform coefficients in the transform block. A quantization matrix applied to the quantized transform coefficients arranged in two dimensions can be encoded and signaled to the video decoding apparatus.

[0067] The rearrangement unit 150 can perform a rearrangement of coefficient values on the quantized residual values.

[0068] The rearrangement unit 150 can change a 2D coefficient array into a 1D coefficient sequence by utilizing a coefficient scan. For example, the rearrangement unit 150 can scan a DC coefficient to a coefficient of a high frequency region with a zig-zag scan or a diagonal scan to output a 1D coefficient sequence. According to the size of the transform unit and the intra prediction mode, a vertical scan that scans a 2D coefficient array in a column direction and a horizontal scan that scans a 2D block type coefficient in a row direction can also be utilized instead of the zig-zag scan. In other words, according to the size of the transform unit and the intra prediction mode, a scan method to be used can be determined among the zig-zag scan, the diagonal scan, the vertical scan, and the horizontal scan.

[0069] The entropy encoder 155 encodes the sequence of the 1D quantized transform coefficients output from the rearrangement unit 150 by utilizing various encoding schemes including a Context-based Adaptive Binary Arithmetic Code (CABAC), an Exponential Golomb, etc., to generate a bitstream.

[0070] Further, the entropy encoder 155 encodes information related to block partitioning (e.g., CTU size, CTU partition flags, QT partition flags, MTT partition types, and MTT partition directions, etc.) to enable a video decoding device to partition blocks identically to the video encoding device. Further, the entropy encoder 155 encodes information on a prediction type indicating whether the current block is coded by intra prediction or by inter prediction. The entropy encoder 155 encodes intra prediction information (i.e., information on an intra prediction mode) or inter prediction information (a merge index in the case of merge mode, and information on a reference picture index and a motion vector difference in the case of AMVP mode) according to the prediction type. Further, the entropy encoder 155 encodes information related to quantization (i.e., information on a quantization parameter and information on a quantization matrix).

[0071] The inverse quantizer 160 inverse-quantizes quantized transform coefficients output from the quantizer 145 to generate transform coefficients. The inverse transformer 165 transforms the transform coefficients output from the inverse quantizer 160 from a frequency domain to a spatial domain to reconstruct a residual block.

[0072] The adder 170 adds the reconstructed residual block and the prediction block generated by the predictor 120 to reconstruct the current block. Pixels in the reconstructed current block are used as reference pixels when intra-predicting a next block.

[0073] The loop filter 180 performs filtering on reconstructed pixels to reduce blocking artifacts, ringing artifacts, blurring artifacts, etc. that occur due to block-based prediction and transform / quantization. The loop filter 180 as an in-loop filter can include all or some of a deblocking filter 182, a sample adaptive offset (SAO) filter 184, and an adaptive loop filter (ALF) 186.

[0074] The deblocking filter 182 filters boundaries between reconstructed blocks to remove blocking artifacts that occur due to block unit encoding / decoding, and the SAO filter 184 and the ALF 186 additionally filter the deblocking-filtered video. The SAO filter 184 and the ALF 186 are filters for compensating for a difference between reconstructed pixels and original pixels that occurs due to lossy coding. The SAO filter 184 applies an offset in units of CTU to enhance subjective image quality and coding efficiency. On the other hand, the ALF 186 performs block unit filtering, and applies different filters by dividing a boundary of a corresponding block and a degree of change to compensate for distortion. Information on filter coefficients to be used for the ALF can be encoded and signaled to the video decoding apparatus.

[0075] The reconstructed blocks filtered through the deblocking filter 182, the SAO filter 184, and the ALF 186 are stored in the memory 190. When all blocks in one picture are reconstructed, the reconstructed picture can be used as a reference picture for inter-prediction of blocks within a picture to be encoded subsequently.

[0076] The video encoding apparatus can store a bitstream of the encoded video data in a nonvolatile storage medium or transmit the bitstream to the video decoding apparatus through a communication network.

[0077] FIG. 5 is a functional block diagram of a video decoding apparatus that can implement the present technology. Hereinafter, with reference to FIG. 5, a video decoding apparatus and components of the apparatus are described.

[0078] The video decoding apparatus can include an entropy decoder 510, a rearrangement unit 515, an inverse quantizer 520, an inverse transformer 530, a predictor 540, an adder 550, a loop filter 560, and a memory 570.

[0079] Similar to the video encoding apparatus of FIG. 1, each component of the video decoding apparatus can be implemented as hardware or software, or as a combination of hardware and software. In addition, the function of each component can be implemented as software, and a microprocessor can also be implemented to perform the function corresponding to each component.

[0080] The entropy decoder 510 extracts information related to block partitioning by decoding a bitstream generated by the video encoding apparatus to determine a current block to be decoded, and extracts prediction information required to reconstruct the current block and information on a residual signal.

[0081] The entropy decoder 510 determines the size of a CTU by extracting information on the CTU size from a sequence parameter set (SPS) or a picture parameter set (PPS), and partitions a picture into CTUs having the determined size. Also, the CTU is determined as the highest layer of a tree structure (i.e., a root node), and partitioning information of the CTU can be extracted to partition the CTU by utilizing the tree structure.

[0082] For example, when the CTU is partitioned by utilizing the QTBTTT structure, a first flag (QT_split_flag) related to partitioning of the QT is first extracted to partition each node into four nodes of a lower layer. In addition, a second flag (mtt_split_flag), a partition direction (vertical / horizontal), and / or a partition type (binary / ternary) related to partitioning of the MTT are extracted with respect to a node corresponding to a leaf node of the QT to partition the corresponding leaf node in the MTT structure. As a result, each node below the leaf node of the QT is recursively partitioned in the BT or TT structure.

[0083] As another example, when the CTU is partitioned by utilizing the QTBTTT structure, a CU partition flag (split_cu_flag) indicating whether to partition a CU is extracted. When the corresponding block is partitioned, a first flag (QT_split_flag) can also be extracted. During the partitioning process, for each node, 0 or more times of recursive MTT partitioning can occur after 0 or more times of recursive QT partitioning. For example, for a CTU, MTT partitioning can occur immediately, or vice versa, or only multiple times of QT partitioning can occur.

[0084] As another example, when the CTU is partitioned by utilizing the QTBT structure, a first flag (QT_split_flag) related to partitioning of the QT is extracted to partition each node into four nodes of a lower layer. In addition, a partition flag (split_flag) indicating whether to further partition a node corresponding to a leaf node of the QT in a BT and partition direction information are extracted.

[0085] On the other hand, when the entropy decoder 510 determines a current block to be decoded by utilizing partitioning of a tree structure, the entropy decoder 510 extracts information on prediction type information indicating whether the current block is intra-predicted or inter-predicted. When the prediction type information indicates intra-prediction, the entropy decoder 510 extracts a syntax element of intra-prediction information (intra-prediction mode) for the current block. When the prediction type information indicates inter-prediction, the entropy decoder 510 extracts information of a syntax element representing inter-prediction information, i.e., a motion vector and a reference picture of a motion vector reference.

[0086] Further, the entropy decoder 510 extracts quantization-related information and extracts information on quantized transform coefficients of the current block as information on the residual signal.

[0087] The reordering unit 515 can reorder the sequence of the 1D quantized transform coefficients entropy-decoded by the entropy decoder 510 into a 2D coefficient array (i.e., a block) in the reverse order of the coefficient scan order performed by the video encoding apparatus.

[0088] The inverse quantizer 520 inverse quantizes the quantized transform coefficients and inverse quantizes the quantized transform coefficients by utilizing a quantization parameter. The inverse quantizer 520 can also apply different quantization coefficients (scaling values) to the quantized transform coefficients arranged in 2D. The inverse quantizer 520 can perform inverse quantization by applying a matrix of quantization coefficients (scaling values) from the video encoding apparatus to the 2D array of quantized transform coefficients.

[0089] The inverse transformer 530 reconstructs the residual signal by inverse transforming the inverse quantized transform coefficients from the frequency domain to the spatial domain to generate a residual block of the current block.

[0090] Further, when the inverse transformer 530 inverse transforms a partial region (sub-block) of the transform block, the inverse transformer 530 extracts a flag (cu_sbt_flag) to transform only the sub-block of the transform block, directional (vertical / horizontal) information (cu_sbt_horizontal_flag) of the sub-block, and / or position information (cu_sbt_pos_flag) of the sub-block. The inverse transformer 530 also inverse transforms the transform coefficients of the corresponding sub-block from the frequency domain to the spatial domain to reconstruct the residual signal and fills a region not inverse transformed with a value "0" as the residual signal to generate a final residual block of the current block.

[0091] Further, when the MTS is applied, the inverse transformer 530 determines a transform function or a transform matrix to be applied on each of the horizontal direction and the vertical direction by utilizing MTS information (mts_idx) signaled from the video encoding apparatus. The inverse transformer 530 also performs inverse transformation on the transform coefficients in the transform block in the horizontal direction and the vertical direction by utilizing the determined transform function.

[0092] The predictor 540 can include an intra predictor 542 and an inter predictor 544. The intra predictor 542 is activated when the prediction type of the current block is intra prediction, and the inter predictor 544 is activated when the prediction type of the current block is inter prediction.

[0093] The intra predictor 542 determines an intra prediction mode of the current block from among a plurality of intra prediction modes according to syntax elements of the intra prediction mode extracted from the entropy decoder 510, and predicts the current block by using neighboring reference pixels of the current block according to the intra prediction mode.

[0094] The inter predictor 544 determines a motion vector and a reference picture of a motion vector reference of the current block by using syntax elements of the inter prediction mode extracted from the entropy decoder 510, and predicts the current block by using the motion vector and the reference picture.

[0095] The adder 550 reconstructs the current block by adding the residual block output from the inverse transformer 530 to the prediction block output from the inter predictor 544 or the intra predictor 542. Pixels within the reconstructed current block are used as reference pixels when intra predicting a block to be decoded subsequently.

[0096] The loop filter unit 560, which is an in-loop filter, can include a deblocking filter 562, an SAO filter 564, and an ALF 566. The deblocking filter 562 performs deblocking filtering on boundaries between reconstructed blocks in order to remove blocking artifacts occurring due to block-based decoding. The SAO filter 564 and the ALF 566 perform additional filtering on the reconstructed blocks after deblocking filtering in order to compensate for differences between reconstructed pixels and original pixels occurring due to lossy encoding. Filter coefficients of the ALF are determined by using information about filter coefficients decoded from the bitstream.

[0097] The reconstructed blocks filtered by the deblocking filter 562, the SAO filter 564, and the ALF 566 are stored in the memory 570. When all blocks in one picture are reconstructed, the reconstructed picture can be used as a reference picture for inter predicting blocks within a picture to be encoded subsequently.

[0098] The present invention, in some embodiments thereof, relates to encoding and decoding video pictures as described above. More specifically, the present invention provides a video coding method and apparatus that applies local illumination compensation to a prediction signal of a current block to compensate for local or global illumination changes between the current frame and a reference frame when performing prediction on the current block according to a geometric partition mode.

[0099] The following embodiments can be performed by the inter predictor 124 in a video encoding apparatus. The following embodiments can also be performed by the inter predictor 544 in a video decoding apparatus.

[0100] A video coding device can generate the signaling information associated with the present embodiment from the perspective of optimizing rate-distortion when encoding the current block. The video coding device can encode the signaling information using the entropy encoder 155 and send the encoded signaling information to the video decoding device. The video decoding device can decode the signaling information associated with the decoding of the current block from the bitstream using the entropy decoder 510.

[0101] In the following description, the term "target block" can be used interchangeably with the current block or coding unit (CU), or can refer to some region of a coding unit.

[0102] In addition, a value of true for a flag indicates a case where the flag is set to 1. In addition, a value of false for a flag indicates a case where the flag is set to 0.

[0103] I-1. Merge / skip mode of inter prediction and merge mode with motion vector difference (MMVD) The following describes a method of constructing a merge candidate list of motion information in the merge / skip mode of inter prediction. To support the merge / skip mode, the video coding device can construct the merge candidate list by selecting a preset number (e.g., 6) of merge candidates.

[0104] The video coding device searches for spatial merge candidates. The video coding device searches for spatial merge candidates from neighboring blocks as shown in FIG. 4. Up to four spatial merge candidates can be selected.

[0105] The video coding device searches for temporal merge candidates. The video coding device can add a block co-located with the current block (the block is in a reference picture other than the current picture in which the target block is located) as a temporal merge candidate. Here, the reference picture can be the same as or different from the reference picture used to predict the current block. One temporal merge candidate can be selected.

[0106] The video coding device searches for a history-based motion vector predictor (HMVP) candidate. The video coding device can store motion vectors of previous h CUs (where h is a natural number) in a table and can use the motion vectors as merge candidates. The table has a size of 6 and stores motion vectors of previous CUs in a first-in first-out (FIFO) manner. This indicates that up to 6 HMVP candidates can be stored in the table. The video coding device can set the latest motion vector among the HMVP candidates stored in the table as a merge candidate.

[0107] The video coding device searches for pairwise average MVP (PAMVP) candidates. The video coding device can set an average of motion vectors of first and second candidates in the merge candidate list as a merge candidate.

[0108] If the merge candidate list cannot be filled even after performing all the above search processes (i.e., if a preset number of merge candidates cannot be assembled), the video coding device adds a zero motion vector as a merge candidate.

[0109] From a codec efficiency optimization perspective, the video coding device can determine a merge index indicating one of the candidates in the merge candidate list. The video coding device can use the merge index to derive a motion vector predictor (MVP) from the merge candidate list, and can then determine the MVP as the motion vector of the current block. In addition, the video coding device can signal the merge index to the video decoding device.

[0110] In the skip mode, the video coding device uses the same motion vector transmission method as in the merge mode, but does not transmit a residual block equivalent to a difference between the current block and the predicted block.

[0111] The above method of constructing the merge candidate list can be equivalently performed by the video decoding device. The video decoding device can decode the merge index. The video decoding device can use the merge index to derive an MVP from the merge candidate list, and can then determine the MVP as the motion vector of the current block.

[0112] On the other hand, when using a merge mode with motion vector difference technique (MMVD), the video coding device can derive an MVP from the merge candidate list by exploiting the merge index. For example, the first or second candidate in the merge candidate list can be used as the MVP. In addition, from a codec efficiency optimization perspective, the video coding device determines a distance index and a direction index. The video coding device can derive a motion vector difference (MVD) by exploiting the distance index and the direction index, and can then reconstruct the motion vector of the current block by summing the MVD and the MVP. In addition, the video coding device can signal the merge index, the distance index, and the direction index to the video decoding device.

[0113] The above MMVD technique can be equivalently performed by the inter predictor 544 in the video decoding device. The video decoding device can decode the merge index, the distance index, and the direction index. The video decoding device can construct the merge candidate list, and can then derive an MVP from the merge candidate list by exploiting the merge index. The video decoding device can use the distance index and the direction index to derive a MVD, and can then sum the MVD and the MVP to reconstruct the motion vector of the current block.

[0114] I-2. Advanced motion vector prediction mode (AMVP mode) for inter prediction The following implementations are described in the context of inter predictor 124 within a video encoding device, but as previously mentioned, the following implementations can also be performed by inter prediction unit 544 in a video decoding device.

[0115] A method of constructing a candidate list of motion information in the advanced motion vector prediction mode (AMVP mode) for inter prediction is described below. To support the AMVP mode, inter predictor 124 in a video encoding device can select a pre-determined number (e.g., two) of candidates to form a candidate list.

[0116] The video encoding device searches spatial candidates. The video encoding device searches spatial candidates from neighboring blocks, as shown in FIG. 4. Up to two spatial candidates can be selected.

[0117] The video encoding device searches temporal candidates. The video encoding device can add a block that is collocated with the current block and within a reference picture (which is not the current picture that owns the target block) as a temporal candidate. Here, the reference picture can be the same or different from the reference picture used to predict the current block. One temporal candidate can be selected.

[0118] If the candidate list cannot be filled even after performing all the above search processes (i.e., if a pre-determined number of merge candidates cannot be assembled), the video encoding device adds a zero motion vector as a merge candidate.

[0119] From a codec efficiency optimization perspective, the video encoding device can determine a candidate index that indicates one of the candidates in the candidate list. The video encoding device can use the candidate index to derive a motion vector predictor (MVP) from the candidate list. In addition, from a codec efficiency optimization perspective, the video encoding device determines a motion vector and then subtracts the MVP to calculate a MVD. Furthermore, the video encoding device can signal the candidate index and the MVD to a video decoding device.

[0120] The above method of constructing an AMVP candidate list can be equivalently performed by inter predictor 544 in a video decoding device. The video decoding device can decode the candidate index and the MVD. The video decoding device can use the candidate index to derive a MVP from the candidate list. The video decoding device can sum the MVD and the MVP to reconstruct the motion vector of the current block.

[0121] I-3. Geometric partition mode (GPM) for inter prediction The following implementations are described in the context of a video decoding device, but they can be implemented in the same or similar manner in a video encoding device.

[0122] In the versatile video coding (VVC), a geometric partitioning mode (GPM) uses various shapes of prediction units instead of rectangular shapes.

[0123] FIG. 6 is a diagram illustrating block partitioning according to a geometric partitioning.

[0124] A video decoding device divides a current block into two parts by a straight line that is perpendicular to a line segment having an angle θ and a distance ρ from a center of the block. Hereinafter, the two divided blocks are referred to as a first block partition and a second block partition. Block partition and partitioned block are used interchangeably. The straight line that divides the current block is referred to as a partition boundary.

[0125] Here, the center of the block represents a single virtual position having the current block before partitioning, in which a 1 / 2 position of a height of the block and a 1 / 2 position of a width of the block intersect. The angle θ represents an angle of rotating counterclockwise from a virtual horizontal axis passing through the center of the block to a line segment that is perpendicular to the partition boundary. The distance ρ represents a distance between the center of the block and the partition boundary.

[0126] As described above, the straight line that divides the block into two parts, i.e., the partition boundary, divides the current block into two different block partitions. The video decoding device uses the above-described geometric partitioning to divide the current block into two blocks, i.e., the first block partition and the second block partition, which perform separate prediction based on the partition boundary.

[0127] As an example, in the regular GPM, information about the partition boundary as illustrated in FIG. 6 is signaled from a video encoding device to a video decoding device. The video decoding device can use the parsed partition boundary information to decode the geometric block partition of the current block. Here, the partition boundary information can include the angle θ and the distance ρ based on the center of the block. In addition, the video encoding device can further signal information indicating whether to apply the GPM to the video decoding device.

[0128] Hereinafter, information about the partition boundary according to the geometric partitioning mode is used interchangeably with geometric partitioning mode information, geometric partitioning information, or partition information.

[0129] FIGS. 7a and 7b are diagrams illustrating a straight line that divides a block into two parts.

[0130] As another example, the lookup table can be configured to include a combination of angles and distances that divide the current block into the first block partition and the second block partition, as shown in FIGS. 7a, 7b, and Table 1. Then, an index indicating the combination of angles and distances in the lookup table can be signaled from the video encoding apparatus to the video decoding apparatus. According to a protocol between the video encoding apparatus and the video decoding apparatus, the combination of angles and distances corresponding to each index can be defined based on a fixed lookup table. As a variant, the lookup table can be adaptively reconfigured according to a pre-agreed rule.

[0131] As described above, the geometric partition shape is based on a partition boundary, which is a straight line representing two partitions of a block. Information of such a straight line can include an index distanceIdx representing a distance p from the center of the block to the partition boundary, and an index angleIdx representing an angle q of a line segment perpendicular to the partition boundary. The index indicating the angle of the line segment perpendicular to the partition boundary can be set as shown in FIG. 7a. Further, 64 geometric partition shapes based on these angles and distances can be set as shown in FIG. 7b.

[0132] The 64 geometric partition shapes can be signaled by utilizing a syntax merge_gpm_partition_idx, which is an index indicating a geometric partition shape, as shown in Table 1. That is, a form in which a current block is divided into a first block partition and a second block partition according to various angles and distances can be effectively signaled by utilizing a single index.

[0133] [Table 1] The index distanceIdx derived from the example of FIG. 7b is a value excluding the size of the current block. Therefore, an actual distance between a pixel in the current block and the straight line can be calculated by utilizing the size information of the current block, the index angleIdx representing the angle, and the index distanceIdx representing the distance. Here, the actual distance is a value expressed in pixel units.

[0134] On the other hand, the actual distance can be used to calculate a weight of each pixel in the current block. For example, for a single pixel in the first block partition, as the actual distance between the pixel and the straight line increases, the weight of the predictor of the first block partition, as described above, increases, and the weight of the predictor of the second block partition decreases. For a pixel located on the partition boundary, the two predictors can use weights having the same value. In this case, the sum of the weights of the two predictors for a single pixel is maintained at 1.

[0135] For the two different block partitions of the current block present within the current picture, the video decoding device performs prediction by utilizing each motion vector (mv0 or mv1). The video decoding device applies a weighted-based blending process to the predicted blocks of the first block partition and the predicted blocks of the second block partition to generate a final predicted block for the current block.

[0136] To perform inter prediction for the current block, the video decoding device obtains a first predicted block by utilizing the motion vector of the first block partition and a second predicted block by utilizing the motion vector of the second block partition. At this time, the video decoding device can obtain the predicted blocks in the process of obtaining the predicted blocks for each block partition by multiplying different weights according to the pixel positions as described above. In addition, the video decoding device can use a shift operation and a clipping operation in the process of generating the final predicted block from the predicted blocks multiplied by the weights.

[0137] FIG. 8 is a schematic diagram illustrating a geometric partition mode (GPM) merge list for geometric motion prediction.

[0138] As illustrated in FIG. 8, the video decoding device selects motion information for motion prediction from the merge list and then uses the selected motion information.

[0139] However, unlike the existing block-based motion prediction technique, for geometric motion prediction, the video decoding device performs uni-prediction for a single block partition by restricting the prediction direction, as illustrated in FIG. 8. This is because, in geometric motion prediction, when motion prediction is performed by bi-prediction for each block partition, the memory bandwidth used for prediction can be doubled compared to block-based motion prediction. Therefore, in order to effectively address the above-described memory bandwidth increase problem, a technique of restricting the directionality of prediction for each block partition can be applied.

[0140] In the case of restricting the directionality of prediction for each block partition in geometric motion prediction, a GPM merge list for geometric motion prediction can be generated by utilizing the existing merge list. To generate the GPM merge list, the video decoding device first constructs the merge list as described above. Then, the video decoding device can generate the GPM merge list for geometric motion prediction by the prediction directionality and the order within the list. At this time, the video decoding device adds L0 direction motion information to the GPM merge list to generate a merge candidate for geometric motion prediction of the first block partition. In addition, the video decoding device can add L1 direction motion information to the GPM merge list to generate a merge candidate for geometric motion prediction of the second block partition. In other words, the video decoding device can derive uni-directional motion information from bi-directional motion information and can add the derived motion information to the GPM merge list.

[0141] As shown in FIG. 8, in the existing geometric motion prediction method, the video decoding device can directly use the merge candidates in the GPM merge list as the motion information for the motion prediction of the first block partition and the second block partition.

[0142] On the other hand, even in the AMVP mode of inter prediction, the video decoding device can derive the motion information related to GPM in a similar manner as described above.

[0143] II. Local Illumination Compensation (LIC) When generating the prediction signal of the current block by using motion compensation of inter prediction, local or global illumination change can occur between the current frame and the reference frame. Local illumination compensation (LIC) is a technique that can improve the prediction efficiency for the current block by performing illumination compensation according to the change of the image capturing environment.

[0144] To effectively perform local illumination compensation, the video decoding device derives the parameters of the linear relationship equation between the neighboring L-shaped template of the current block and the neighboring L-shaped template of the reference block. For example, the parameters of the linear relationship equation can be derived based on the linear least square method. At this time, the neighboring L-shaped template of the current block exists in the reconstructed region of the current block. The video decoding device uses the derived parameters of the linear relationship equation to perform local illumination compensation of the current block based on the reference block, as shown in Equation 1.

[0145] [Equation 1] In Equation 1, a and β are the parameters of the linear relationship equation, representing the slope and the offset, respectively. P(x, y) represents the pixel in the reference block, i.e., the predicted pixel, and P'(x, y) represents the predicted pixel in the current block that is subject to local illumination compensation.

[0146] As described above, local illumination compensation is applied to the current block, and the linear relationship equation is derived based on the neighboring reference pixels adjacent to the reference block and the neighboring reference pixels adjacent to the current block. In this case, when there is no available pixel in the neighboring reference pixels of the reference block, local illumination compensation can not be performed on the current block. When the available pixel does not exist on both the left and the top of the reference block, but exists on a part of the left and the top, local illumination compensation can be performed by using some existing pixels.

[0147] In the video encoding device and the video decoding device, the parameters a and β of the linear relationship equation used in local illumination compensation are calculated in the same way. Therefore, the video encoding device can not separately transmit the parameters of the linear relationship equation to the video decoding device.

[0148] When applying LIC in merge mode to a neighboring block selected according to a merge index, the video decoding device performs LIC by applying LIC parameters of the neighboring block to the current block. In AMVP mode, a flag is signaled to indicate whether LIC is to be performed or not, and the video decoding device derives the LIC parameters by utilizing a template of the current block and a template of a reference block. In addition, LIC is applied in the case of uni-prediction, and is not applied when the current block is predicted according to Combined Inter Intra Prediction (CIIP). CIIP is a technique that generates an intra-predicted block after the current block (the block is an intra-predicted block) and an inter-predicted block after the current block (the block is an inter-predicted block), and then performs a weighted sum of the intra-predicted block and the inter-predicted block to generate a final predicted block of the current block.

[0149] The following implementations are described in the context of a video decoding device, but they can be implemented in a video encoding device in the same or similar manner.

[0150] III. Embodiments According to the Present Disclosure A video decoding device according to some embodiments can determine prediction and transform units, and in response to a current block corresponding to the determined units, can perform prediction and inverse transform by utilizing a determined prediction technique and prediction mode to finally generate a reconstructed block of the current block. The operations shown in FIG. 9 can be performed by the inverse transformer 530, the predictor 540, and the adder 550 of the video decoding device. On the other hand, the same operations as shown in FIG. 9 can be performed by the inverse transformer 165, the picture partitioner 110, the predictor 120, and the adder 170 in a video encoding device. In this case, the video decoding device uses encoding information parsed from a bitstream, but the video encoding device can use encoding information set from a high level in terms of minimizing bit rate distortion. Hereinafter, for convenience, the implementations are described with a focus on the video decoding device.

[0151] As shown in FIG. 5, the predictor 540 includes an intra-predictor 542 and an inter-predictor 544 according to a prediction technique, but as shown in FIG. 9, the predictor 540 can include all or part of a prediction unit determiner 902, a prediction technique determiner 904, a prediction mode determiner 906, and a prediction performer 908.

[0152] When the color format of the input video is a YUV format (YUV420, YUV411, YUV422, YUV444, etc.), the video decoding device can perform prediction and reconstruction of the luma components, and then can perform prediction and reconstruction of the chroma components. In this way, the luma components and the chroma components can be reconstructed sequentially by the component order shown in FIG. 9. On the other hand, when the color format of the input video is RGB, the video encoding device can perform color format conversion from RGB to YUV, and then can encode the converted video. Here, under the YUV format, the color format represents the association between pixels in the luma components and pixels in the chroma components.

[0153] A prediction unit determiner 902 determines a prediction unit (PU). A prediction technique determiner 904 determines a prediction technique, such as an intra prediction, an inter prediction, or an intra block copy (IBC) mode, a palette mode, etc., in response to the prediction unit. A prediction mode determiner 906 determines a detailed prediction mode for the prediction technique. A prediction executor 908 generates a prediction block for the current block according to the determined prediction mode.

[0154] An inverse transformer 530 includes a transform unit determiner 910 and an inverse transform executor 912. The transform unit determiner 910 determines a transform unit (TU) in response to the dequantized signal of the current block, and the inverse transform executor 912 performs inverse transform on the transform unit represented by the dequantized signal to generate a residual signal.

[0155] An adder 550 sums the prediction block and the residual signal to generate a reconstructed block. The reconstructed block is stored in a memory, and can be used to predict other blocks in the future.

[0156] The prediction unit determined by the prediction unit determiner 902 can be one of the current block or a sub-block partitioned from the current block. In this case, depending on the color format, the prediction unit of the chroma component can correspond in size to the prediction unit of the luma component. Alternatively, the prediction units of the luma component and the chroma component can be determined separately, and the prediction can be performed in response to the prediction unit of the chroma component.

[0157] The prediction technique determiner 904 determines a prediction technique for the prediction unit. As described above, the prediction technique can be one of an inter prediction, an intra prediction, an IBC mode, and a palette mode. In this case, the prediction technique for the chroma component can be determined to be the same as the prediction technique of the corresponding luma component without signaling and parsing any additional information.

[0158] In one example, when the prediction technique for the current block is not intra prediction, the video decoding device parses a 1-bit flag information. When the parsed flag indicates a skip mode, the video decoding device can determine the prediction mode for the current block as an inter prediction merge mode or an IBC merge mode. In the skip mode, the video decoding device can also use the prediction signal instead of the reconstructed signal while skipping the inverse transform, i.e., without parsing the residual signal.

[0159] On the other hand, when the parsed flag does not indicate the skip mode for the current block, the prediction technique determiner 904 can determine one of the prediction techniques by parsing a series of 1-bit flags as the prediction technique for the current block, such as inter prediction, intra prediction, IBC mode, palette mode, etc.

[0160] For example, when the skip mode is not applied to the current block and the inter prediction or IBC mode is determined as the prediction technique, the video decoding device parses a 1-bit flag. The video decoding device can determine the prediction mode for the current block as a regular merge mode or an advanced motion vector prediction (AMVP) mode based on the parsed flag.

[0161] The prediction mode determiner 906 determines a detailed prediction mode for the prediction technique.

[0162] As an example, when the prediction technique for the current block is the inter prediction, the prediction mode determiner 906 can determine the regular merge mode or the AMVP mode as the prediction mode for the current block. In the regular merge mode or the AMVP mode, the video decoding device generates prediction blocks based on one or more motion compensations based on the parsed motion information, and performs a weighted sum of the plurality of generated prediction blocks to generate a final prediction signal for the current block.

[0163] In another example, when the prediction technique for the current block is the inter prediction, the prediction mode determiner 906 can determine a geometric partitioning mode (GPM) as the prediction mode for the current block. In the GPM, the video decoding device divides the current block into two or more partitioned blocks according to geometric partitioning, generates prediction blocks through one or more motion compensations based on the parsed motion information for the current block, and performs a weighted sum of the plurality of generated prediction blocks to generate a final prediction signal for the current block. For example, as described above, the current block can be divided into two partitioned blocks.

[0164] The prediction executor 908 generates a prediction block for the current block according to the determined prediction technique and prediction mode.

[0165] As an example, the prediction performer 908 generates a prediction block of the current block according to the prediction mode, and the adder 550 sums the prediction block of the current block and the residual signal to generate a reconstructed block.

[0166] Meanwhile, the entropy decoder 510 can reconstruct the quantized secondary transform coefficients when the secondary transform is undergone. On the other hand, the entropy decoder 510 can reconstruct the quantized primary transform coefficients when the secondary transform is not applied. The inverse quantizer 520 can apply inverse quantization to the transform coefficients reconstructed based on the quantization parameter to generate inverse quantized transform coefficients.

[0167] The transform unit determiner 910 of the inverse transformer 530 determines a transform unit (TU) for the inverse quantized transform coefficients. At this time, when a single TU is divided into a plurality of sub-blocks, a single sub-block can be used as a TU.

[0168] The inverse transformer 530 can determine a non-separable secondary inverse transform kernel and a separable primary inverse transform kernel when the secondary transform is undergone. On the other hand, the inverse transformer 530 can determine a separable primary inverse transform kernel or a non-separable primary inverse transform kernel when the secondary transform is not applied.

[0169] Hereinafter, the terms "geometric partition mode", "GPM partition mode", and "partition mode according to GPM" are used interchangeably. Also, as described above, the two partitioned blocks according to GPM are defined as a first partitioned block and a second partitioned block, respectively.

[0170] For example, when the prediction technique for the current block is inter prediction and the current block is predicted according to GPM, the prediction performer 908 in the video decoding apparatus according to the present application can predict the current block as follows.

[0171] When the current block is predicted according to GPM, the video decoding apparatus generates a merge list according to a regular merge mode, and then uses the merge list to generate a GPM (motion vector) merge list. The GPM merge list can be the list shown in FIG. 8. The video decoding apparatus decodes gpm_mode_idx and merge_index. Here, the GPM index gpm_mode_idx can be merge_gpm_partition_idx according to Table 1, and gpm_mode_idx indicates one of the GPM partition modes. The merge index merge_index can indicate a motion information predictor within the GPM merge list. The video decoding apparatus determines a geometric partition mode according to gpm_mode_idx, and uses merge_idx and the GPM merge list to determine motion information, i.e., a motion vector, for the partitioned blocks.

[0172] FIG. 10 is a diagram illustrating blending of neighboring reference pixels of a reference block according to at least one embodiment of the present application.

[0173] The video decoding device calculates a first reference block (B) based on the motion vector of the top-left block (i.e., the first partitioned block) and a second reference block (B) based on the motion vector of the bottom-right block (i.e., the second partitioned block), as shown in the example of FIG. 10. The video decoding device can perform blending processing on the two reference blocks to generate a predicted block (B`), and use the blended predicted block to generate the final prediction signal of the current block.

[0174] At this time, in order to apply local illumination compensation to the current block, the video decoding device selects reference pixels for local illumination compensation, i.e., reference templates of B_refpl and B_refp2, for the two reference blocks (B), respectively. The video decoding device applies filter based on the blending mask to the reference pixels to derive neighboring reference pixels B_refp` of the blended predicted block. The video decoding device uses the neighboring reference pixels B_refp` of the blended predicted block (i.e., the template of the blended predicted block) and the neighboring reference pixels A_refp of the current block (i.e., the template of the current block) as the basis for calculating local illumination compensation (LIC) parameters a and b according to a linear relationship as shown in Equation 2.

[0175] [Equation 2] For example, Equation 2 can derive the LIC parameters based on a least square method such as Equations 3 and 4.

[0176] [Equation 3] [Equation 4] In Equations 3 and 4, C(n) represents pixel values in the template of the current block for deriving the LIC parameters, and R(n) represents pixel values within the template of the blended predicted block for deriving the LIC parameters. N represents the number of pixels for deriving the LIC parameters.

[0177] In order to apply local illumination compensation in the geometric partitioning mode, the video decoding device can apply the calculated LIC parameters to the blended predicted block B, and thereby generate the final prediction signal of the current block.

[0178] In the following, a method performed by a video decoding device according to the present application for applying local illumination compensation in the GPM mode is described in detail.

[0179] The video decoding device uses the decoded syntax of the related syntax to determine whether to apply the prediction based on local illumination compensation in the GPM mode. At this time, the determination of whether to apply the prediction based on LIC can be transmitted from the video encoding device to the video decoding device by using the high-level syntax. Alternatively, it can be transmitted from the video encoding device to the video decoding device by using the syntax of the level of a specific block size (e.g., CTU or CU) or the syntax of the block group level. The video decoding device uses the decoded syntax as a basis for determining whether to predict the current block according to the geometric partition mode and whether to apply the local illumination compensation. Here, the syntax can include a flag indicating whether to apply the GPM mode and a flag indicating whether to apply the prediction based on LIC.

[0180] To apply the LIC to the current block, the video decoding device determines the availability of the neighboring reference pixels for modeling the linear relationship of the local illumination compensation. Here, the availability includes not only the presence or absence of the previously decoded neighboring reference pixels A_refp, B_refp1, and B_refp2 (i.e., the reference template), but also whether the neighboring reference pixels are available for applying the local illumination compensation to the current block. If the template A_refp of the current block is not available for the local illumination compensation, the video decoding device can perform the prediction / decoding of the current block, thereby covering the steps related to the local illumination compensation. Alternatively, the video decoding device can perform the prediction / decoding of the current block by setting a to 1 and b to 0, thereby omitting the step of calculating the LIC parameters.

[0181] When the neighboring reference pixel A_refp of the current block is available, the video decoding device performs the motion vector-based compensation for each of the partitioned blocks according to the geometric partition mode to determine the prediction block B, i.e., the reference block, of the partitioned block. At this time, as described above, the motion information of the partitioned block can be generated from the merge index and the GPM merge list.

[0182] The video decoding device determines the availability of B_refp1 and B_refp2, which are the neighboring reference pixels of the reference block. At this time, even if B exists, their positions can cause all or some of B_refp1 and B_refp2 to be unavailable. If both B_refp1 and B_refp2 are not available, the video decoding device can perform the prediction / decoding of the current block while omitting the steps related to the local illumination compensation. Alternatively, the video decoding device can perform the prediction / decoding of the current block by setting a to 1 and b to 0, thereby omitting the step of calculating the LIC parameters.

[0183] When both A_refp and B_refp1 are available, the video decoding device can use A_refp and B_refp1 to compute LIC parameters a and b. The video decoding device can then apply the computed LIC parameters to the prediction block to generate the final prediction signal of the current block.

[0184] As another example, the video decoding device can not use all the pixels included in A_refp, B_refp1, and B_refp2. The video decoding device can filter A_refp, B_refp1, and B_refp2, or can extract and generate some pixels from A_refp, B_refp1, and B_refp2, and can use the generated some pixels to compute LIC parameters a and b.

[0185] As yet another example, when a certain position of B makes some of B_refp1 and B_refp2 available, the video decoding device can use only those available to compute LIC parameters a and b. For example, as shown in the example of FIG. 11, when only the left-side pixel of B_refp1 and the top-side pixel of B_refp2 are available, the video decoding device can use the left-side pixel of A_refp, the left-side pixel of B_refp1, the top-side pixel of A_refp, and the top-side pixel of B_refp2 instead of the blending of B_refp1 and B_refp2 to compute LIC parameters a and b.

[0186] As yet another example, as shown in FIG. 12, when only the left-side pixel of B_ref1 and the left-side pixel of B_ref2 are available, the video decoding device applies blending to the available pixels to derive B_refp`. The video decoding device can use the left-side pixels of A_refp and B_refp` to compute LIC parameters a and b.

[0187] In yet another example, when only some of B_refp1 and B_refp2 are available, the video decoding device can perform the prediction / decoding of the current block while omitting the steps related to local illumination compensation. Alternatively, the video decoding device can perform the prediction / decoding of the current block by setting a to 1 and b to 0, thereby omitting the step of computing LIC parameters.

[0188] FIGS. 13a and 13b are schematic diagrams illustrating a blending mask of reference pixels, according to at least one embodiment of the present application.

[0189] The video decoding apparatus can use a mask having the same shape as the mask for B as a blending mask for blending B_refp1 and B_refp2, as shown in FIG. 13a. More accurately, as shown in the example of FIG. 13b, the video decoding apparatus can blend the reference pixels by using a mask having the same shape as the mask that has been used in the line of pixels within the B block of the most neighboring reference pixels. In the blending mask shown in FIG. 13b, the weight of the pixel increases as the color of the pixel approaches white.

[0190] FIGS. 14a and 14b are schematic diagrams illustrating a blending mask of reference pixels according to another embodiment of the present application.

[0191] As another example, the video decoding apparatus can determine a mask shape by using a partition structure of a block containing A_refp as a blending mask for B_refp1 and B_refp2, as shown in FIG. 14a. The video decoding apparatus can construct a mask according to the determined mask shape, as shown in the example of FIG. 14b, and then can use the constructed mask to blend the reference pixels. In the blending mask shown in FIG. 14b, the weight of the pixel increases as the color of the pixel approaches white.

[0192] When the local illumination compensation is applied, the video decoding apparatus can perform in-loop filtering by reflecting the local illumination compensation. For example, when determining the filtering strength of the deblocking filter 562 at a boundary between two neighboring blocks, the video decoding apparatus can include the LIC parameters a and b in the parameters for determining the filtering strength. For example, when (i) the local illumination compensation is applied to the blocks containing the target pixel for in-loop filtering, and (ii) the a values of two of the blocks containing the target pixel are not 1 and the b values of the two blocks are not 0, the video decoding apparatus can determine the BS for determining the filtering strength of the deblocking filter 562 as 1 to perform filtering at the boundary between the two blocks, even if the reference pictures of the two blocks are the same. Alternatively, if the a values of the two blocks containing the target pixel for filtering are different, the video decoding apparatus can determine the BS as 1 to determine the filtering strength of the deblocking filter 562 to perform filtering at the boundary between the two blocks. In addition, the foregoing description can also be applied when the video encoding apparatus determines the filtering strength of the deblocking filter 182.

[0193] Hereinafter, a method of applying LIC when predicting a current block according to GPM is described by using the illustrations of FIGS. 15 and 16.

[0194] FIG. 15 is a flowchart of a method of encoding a current block by a video encoding apparatus according to at least one embodiment of the present application.

[0195] The video encoding device determines a GPM index (S1500). From a rate-distortion optimization perspective, the video encoding device can determine the GPM index that indicates the geometric partition shape of the current block.

[0196] The video encoding device partitions the current block into partitioned blocks according to the GPM index (S1502). For example, the current block can be partitioned into a first partitioned block and a second partitioned block.

[0197] The video encoding device generates a merge list according to a regular merge mode, and then uses the merge list to generate a GPM (motion vector) merge list. From a rate-distortion optimization perspective, the video encoding device can determine a merge index. The video encoding device uses the merge index and the GPM merge list to determine the motion information of the partitioned blocks.

[0198] The video encoding device uses the motion information to generate reference blocks of the partitioned blocks (S1504). Here, the reference blocks include a first reference block and a second reference block. The first reference block and the second reference block correspond to the first partitioned block and the second partitioned block, respectively.

[0199] The video encoding device blends the reference blocks of the partitioned blocks to generate a prediction block of the current block (S1506).

[0200] The video encoding device performs the following steps to determine the availability status of neighboring reference pixels for linear modeling of local illumination compensation. Here, the availability includes not only the presence or absence of previously decoded neighboring reference pixels (A_refp, B_refp1, and B_refp2), but also whether the neighboring reference pixels are available for applying the local illumination compensation to the current block.

[0201] The video encoding device checks the availability of neighboring reference pixels of the current block at the top and left side of the current block (S1508).

[0202] When the neighboring reference pixel (A_refp) of the current block is available (S1508 is Yes), the video encoding device performs the following steps.

[0203] The video encoding device checks the availability of neighboring reference pixels of the respective reference blocks at the top and left side of the respective reference blocks (S1510).

[0204] For example, as shown in FIG. 11 or FIG. 12, if some neighboring reference pixels (B_refp1) of the first reference block and some neighboring reference pixels (B_refp2) of the second reference block are available, the video encoding device can use the available B_refp1 and B_refp2.

[0205] If all or some of the neighboring reference pixels (B_refpl and B_refp2) of the corresponding reference block are available (S1510 is Yes), the video encoding device performs the following steps.

[0206] The video encoding device blends the neighboring reference pixels of the reference block to generate the neighboring reference pixels of the blended prediction block (S1512).

[0207] The video encoding device derives parameters of a linear relationship between the neighboring reference pixels of the current block and the neighboring reference pixels of the blended prediction block (S1514).

[0208] For example, the video encoding device selects pixels for linear modeling from the neighboring reference pixels (A_refp) of the current block and the neighboring reference pixels (B_refp') of the blended prediction block, and then uses the selected reference pixels to calculate the parameters a and b of the linear relationship. The pixels for calculating the parameters can be all or some of the neighboring reference pixels (A_refp) of the current block and the neighboring reference pixels (B_refp') of the blended prediction block. In this case, the neighboring reference pixels of the blended prediction block corresponding to the neighboring reference pixels of the current block in position are used for derivation of the parameters.

[0209] The video encoding device uses the parameters of the linear relationship to apply local illumination compensation to the blended prediction block, thereby generating a first final prediction block of the current block (S1516).

[0210] If the neighboring reference pixels (A_refp) of the current block are not available (S1508 is No), or if the neighboring reference pixels (B_refpl and B_refp2) of the reference block are not available (S1510 is No), the video encoding device derives the blended prediction block as a second final prediction block of the current block (S1518). For example, the video encoding device does not perform operations related to local illumination compensation, and generates the second final prediction block of the current block. Alternatively, the video encoding device omits the process of deriving the parameters of the linear relationship. That is, the video encoding device generates the second final prediction block of the current block by assigning 1 to a and 0 to b, thereby omitting local illumination compensation.

[0211] The video encoding device determines a flag indicating whether local illumination compensation is to be applied based on the first final prediction block and the second final prediction block (S1520).

[0212] From the perspective of rate-distortion optimization, the video encoding device can determine the flag indicating whether local illumination compensation is to be applied. For example, if the first final prediction block is optimal, the video encoding device sets the flag to true. In contrast, if the second final prediction block is optimal, the video encoding device sets the flag to false.

[0213] The video encoding device encodes the flag (S1522).

[0214] For example, the video encoding device can determine whether local illumination compensation is to be applied to a picture or slice. A flag indicating whether local illumination compensation is to be applied can be transmitted by utilizing a high-level syntax or by utilizing a block group level (e.g., CTU, CU) or block level syntax.

[0215] Depending on the value of the flag indicating whether local illumination compensation is to be applied, the video encoding device subtracts the first final prediction block or the second final prediction block from the current block to generate a residual block. The video encoding device then transforms / quantizes the residual block to generate and encode quantized transform coefficients.

[0216] FIG. 16 is a flowchart of a method of reconstructing a current block by a video decoding device, according to at least one embodiment of the present disclosure.

[0217] The video decoding device decodes the GPM index from the bitstream (S1600).

[0218] The video decoding device partitions the current block into partitioned blocks according to the GPM index (S1602). For example, the current block can be partitioned into a first partitioned block and a second partitioned block.

[0219] The video decoding device generates a merge list according to a regular merge mode, and then uses the merge list to generate a GPM (motion vector) merge list. The video decoding device decodes a merge index from the bitstream. The video decoding device uses the merge index and the GPM merge list to determine motion information of the partitioned blocks.

[0220] The video decoding device uses the motion information to generate reference blocks of the partitioned blocks (S1604). Here, the reference blocks include a first reference block and a second reference block. At this time, the first reference block and the second reference block correspond to the first partitioned block and the second partitioned block, respectively.

[0221] The video decoding device blends the reference blocks of the partitioned blocks to generate a prediction block of the current block (S1606).

[0222] The video decoding device decodes a flag indicating whether local illumination compensation is to be applied from the bitstream (S1608).

[0223] For example, the video decoding device can determine whether local illumination compensation is to be applied to a picture or slice. A flag indicating whether local illumination compensation is to be applied can be transmitted by utilizing a high-level syntax or by utilizing a block group level syntax (e.g., CTU or CU level) or block level syntax.

[0224] The video decoding device checks the flag (S1610).

[0225] If the flag indicating whether to apply local illumination compensation is true (S1610 is Yes), the video decoding device performs the following steps to determine the availability status of neighboring reference pixels for linear modeling of local illumination compensation. Here, the availability includes not only the presence or absence of previously decoded neighboring reference pixels (A_refp, B_refpl, and B_refp2), but also whether the neighboring reference pixels are available for applying local illumination compensation to the current block.

[0226] The video decoding device checks the availability of neighboring reference pixels of the current block (S1612).

[0227] If the neighboring reference pixel (A_refp) of the current block is available (S1612 is Yes), the video decoding device performs the following steps.

[0228] The video decoding device checks the availability of neighboring reference pixels of the corresponding reference block on the upper side and the left side of the corresponding reference block (S1614).

[0229] For example, as shown in the examples of FIG. 11 or 12, if some neighboring reference pixels (B_refpl) of the first reference block and some neighboring reference pixels (B_refp2) of the second reference block are available, the video decoding device can use the available B_refpl and B_refp2.

[0230] When all or some neighboring reference pixels (B_refpl and B_refp2) of the corresponding reference block are available, the video decoding device performs the following steps.

[0231] The video decoding device blends the neighboring reference pixels of the reference blocks to generate neighboring reference pixels of the blended prediction block (S1616).

[0232] The video decoding device derives parameters of a linear relationship between the neighboring reference pixels of the current block and the neighboring reference pixels of the blended prediction block (S1618).

[0233] For example, the video decoding device selects pixels for linear modeling from the neighboring reference pixels (A_refp) of the current block and the neighboring reference pixels (B_refp') of the blended prediction block, and then uses the selected reference pixels to calculate the parameters a and b of the linear relationship. The pixels for calculating the parameters can be all or some of the neighboring reference pixels (A_refp) of the current block and the neighboring reference pixels (B_refp') of the blended prediction block. In this case, the neighboring reference pixels of the blended prediction block corresponding to the neighboring reference pixels of the current block in position are used to derive the parameters.

[0234] The video decoding device uses the parameters of the linear relationship to apply the local illumination compensation to the blended prediction block to generate a final prediction block of the current block (S1620).

[0235] In contrast, if the flag indicating whether to apply the local illumination compensation is false (S1610 is No), or if the neighboring reference pixels of the current block (A_refp) are not available (S11612 is No), or if the neighboring reference pixels of the reference block (B_refpl and B_refp2) are not available (S11614 is No), the video decoding device derives the blended prediction block as the final prediction block of the current block (S1630).

[0236] As another example, when some of the neighboring reference pixels of the reference block (B_refpl and B_refp2) are available, the video decoding device also derives the blended prediction block as the final prediction block of the current block.

[0237] For example, the video decoding device does not perform the operations related to the local illumination compensation and generates the final prediction block of the current block. Alternatively, the video decoding device omits the process of deriving the parameters of the linear relationship. That is, the video decoding device generates the final prediction block of the current block by assigning 1 to a and 0 to b, thereby omitting the local illumination compensation.

[0238] The video decoding device decodes the quantized transform coefficients from the bitstream, inverse quantizes / inverse transforms the quantized transform coefficients to reconstruct a residual block. The video decoding device sums the final prediction block of the current block and the residual block to generate a reconstructed block of the current block.

[0239] Although the steps in the respective flowcharts are described as being sequentially performed, the steps merely illustrate the technical ideas of some embodiments of the present application. Therefore, a person of ordinary skill in the art to which the present application pertains can perform the steps by changing the order described in the respective figures or by performing two or more steps in parallel. Therefore, the steps in the respective flowcharts are not limited to the order shown in terms of the time of occurrence.

[0240] It should be understood that the above description presents illustrative embodiments that can be implemented in various other manners. The functions described in some embodiments can be implemented by hardware, software, firmware, and / or a combination thereof. It should also be understood that the functional components described in the present application are marked as “... unit” to highlight the possibility of their independent implementation.

[0241] On the other hand, in some embodiments, various methods or functions described can be implemented as instructions stored in a non-volatile recording medium that can be read and executed by one or more processors. The non-volatile recording medium can include various types of recording devices that store data in a form readable by a computer system, for example. For example, the non-volatile recording medium can include storage media such as erasable programmable read-only memory (EPROM), a flash drive, an optical disk drive, a magnetic hard disk drive, and a solid state drive (SSD), and the like.

[0242] Although exemplary embodiments of the present application are described for illustrative purposes, it will be understood by those of ordinary skill in the art to which the present application pertains that various modifications, additions and substitutions are possible, without departing from the idea and scope of the present application. Therefore, the embodiments of the present application are described for the sake of brevity and clarity. The scope of the technical idea of the embodiments of the present application is not limited by the examples illustrated. Accordingly, it will be understood by those of ordinary skill in the art to which the present application pertains that the scope of the present application should not be limited by the explicitly described embodiments above, but by the claims and their equivalents.

[0243] (Reference Signs) 124: Inter predictor 544: Inter predictor 902: Prediction unit determiner 904: Prediction technique determiner 906: Prediction mode determiner 908: Prediction executor.

[0244] Cross Reference to Related Applications This application claims priority to and the benefit of Korean Patent Application No. 10-2023-0063330, filed May 16, 2023, and Korean Patent Application No. 10-2024-0059764, filed May 7, 2024, the entire contents of each of which are incorporated herein by reference.

Claims

1. A method for reconstructing a current block using a video decoding device, the method comprising: Decode geometric partitioning patterns from a bitstream; The current block is partitioned into partitioned blocks based on the geometric partitioning pattern index; Reference blocks for partitioned blocks are generated by utilizing motion information, the reference blocks including a first reference block and a second reference block; The prediction block for the current block is generated by mixing the reference block of the partitioned block; as well as Check the availability of adjacent reference pixels on the top and left sides of the current block. The method further includes: when the adjacent reference pixels of the current block are available: Check the availability of adjacent reference pixels of each reference block on the top and left sides of each reference block, and The method further includes: when all or part of the adjacent reference pixels of each reference block are available: The adjacent reference pixels of the blended prediction block are generated by using the adjacent reference pixels of the blended reference block; Parameters for deriving the linear relationship between the neighboring reference pixels of the current block and the neighboring reference pixels of the blended prediction block; and The final prediction block of the current block is generated by applying local lighting compensation to the blended prediction blocks using parameters based on linear relationships.

2. The method according to claim 1, further comprising: A flag indicating whether local illumination compensation should be applied during bitstream decoding; as well as Check the aforementioned mark. When the flag is true, the method continues to check the availability of neighboring reference pixels of the current block.

3. The method according to claim 1, further comprising: When the flag is false, the mixed prediction block is derived as the final prediction signal for the current block.

4. The method of claim 1, further comprising: when adjacent reference pixels of the current block are unusable, or when adjacent reference pixels of each reference block of the reference block are unusable: The mixed prediction block is derived as the final prediction signal for the current block.

5. The method according to claim 1, wherein, The parameters for deriving linear relationships include: The parameters are derived by applying linear least squares to the neighboring reference pixels of the current block and the neighboring reference pixels of the mixed prediction block.

6. The method according to claim 1, wherein, The parameters for deriving the linear relationship include the following: when only the adjacent reference pixels on the left side of the first reference block and the adjacent reference pixels on the upper side of the second reference block are available: Instead of mixing the adjacent reference pixels on the left side of the first reference block and the adjacent reference pixels on the top side of the second reference block, the adjacent reference pixels on the left side of the current block, the adjacent reference pixels on the left side of the first reference block, the adjacent reference pixels on the top side of the current block, and the adjacent reference pixels on the top side of the second reference block are used to continue to derive the parameters of the linear relationship.

7. The method according to claim 1, wherein, The neighboring reference pixels for generating the blended prediction block include those where only the left-hand neighboring reference pixels of the first reference block and the left-hand neighboring reference pixels of the second reference block are available: The left-side adjacent reference pixels of the mixed prediction block are generated by mixing the left-side adjacent reference pixels of the first reference block and the left-side adjacent reference pixels of the second reference block.

8. The method according to claim 7, wherein, The parameters for deriving linear relationships include: The parameters of the linear relationship are further derived by using the neighboring reference pixels to the left of the current block and the neighboring reference pixels to the left of the mixed prediction block.

9. The method according to claim 1, wherein, The neighboring reference pixels for generating the blended prediction block include: The reference pixels of the reference block are blended by using the same blending mask as the prediction block used to generate the blend.

10. The method according to claim 1, wherein, The neighboring reference pixels for generating the blended prediction block include: The reference pixels of the reference block are blended by using a blending mask determined based on the partitioning structure of the block containing the current block.

11. The method of claim 1, further comprising: The in-loop filter is applied to the boundary between the final predicted block, which has already undergone local lighting compensation, and its neighboring blocks. The application of in-loop filtering includes: The filtering intensity at the boundary is determined by using parameters that represent the linear relationship between the final predicted block and its neighboring blocks.

12. A method for encoding a current block using a video encoding device, the method comprising: Determine the geometric partitioning pattern index; The current block is partitioned into partitioned blocks based on the geometric partitioning pattern index; Reference blocks for partitioned blocks are generated by utilizing motion information, the reference blocks including a first reference block and a second reference block; The prediction block for the current block is generated by mixing the reference block of the partitioned block; as well as Check the availability of adjacent reference pixels on the top and left sides of the current block. The method further includes, when the adjacent reference pixels of the current block are available, Check the availability of adjacent reference pixels of each reference block on the top and left sides of each reference block, and The method further includes, when all or part of the adjacent reference pixels of each reference block are available, The adjacent reference pixels of the blended prediction block are generated by blending the adjacent pixels of the reference block; Parameters for deriving the linear relationship between the neighboring reference pixels of the current block and the neighboring reference pixels of the blended prediction block; and The first final prediction block for the current block is generated by applying local illumination compensation to the blended prediction blocks using parameters based on a linear relationship.

13. The method of claim 12, further comprising: The hybrid prediction block is derived as the second final prediction block of the current block; A flag indicating whether local lighting compensation should be applied is determined based on the first and second final prediction blocks; and The flag is encoded.

14. The method of claim 12, further comprising: when adjacent reference pixels of the current block are unusable, or when adjacent reference pixels of each reference block of the reference block are unusable: Skip generating the first final prediction block.

15. A method for providing video data to a video decoding device, the method comprising: Encode video data into a bitstream; as well as Send the bitstream to the video decoding device. The encoded video data includes: Determine the geometric partitioning pattern index; The current block is partitioned into partitioned blocks based on the geometric partitioning pattern index; Reference blocks for partitioned blocks are generated by utilizing motion information, the reference blocks including a first reference block and a second reference block; The prediction block for the current block is generated by mixing the reference block of the partitioned block; and Check the availability of adjacent reference pixels on the top and left sides of the current block. The encoded video data further includes, when the adjacent reference pixels of the current block are available, Check the availability of adjacent reference pixels of each reference block on the top and left sides of each reference block, and The encoded video data further includes, when all or part of the adjacent reference pixels of each reference block are available, The adjacent reference pixels of the blended prediction block are generated by blending the adjacent pixels of the reference block; Parameters for deriving the linear relationship between the neighboring reference pixels of the current block and the neighboring reference pixels of the blended prediction block; and The final prediction block of the current block is generated by applying local lighting compensation to the blended prediction blocks using parameters based on linear relationships.

Citation Information

Patent Citations

  • Housing including metal sheet, electronic device including same, and method for manufacturing same

    KR1020230063330A