Illumination compensation method for prediction based on geometric segmentation mode
By applying local illumination compensation techniques in video coding, the problem of illumination variation in inter-frame prediction is solved, thereby improving coding efficiency and video quality.
Patent Information
- Application Number
- CN202480032042.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-05-03
- Filing Date
- 2024-05-07
- Publication Date
- 2025-12-19
AI Technical Summary
Existing video coding techniques, when using geometric segmentation patterns for inter-frame prediction, fail to effectively compensate for local or global illumination changes between the current frame and the reference frame, resulting in limitations on coding efficiency and video quality.
The Local Illumination Compensation (LIC) technique is used to generate local illumination compensation parameters, which are applied to the reference block of the inter-frame segmentation block to generate the final prediction block, compensating for the illumination changes between the current frame and the reference frame.
It improves video coding efficiency and video quality, and enhances the accuracy and effectiveness of inter-frame prediction.
Smart Images

Figure CN121176010A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to a video encoding method and apparatus using illumination compensation in prediction based on geometric partition mode. BACKGROUND
[0002] The statements in this section merely provide background information related to the present disclosure and can not constitute the prior art.
[0003] Since video data has a larger amount of data than audio data or still image data, if no compression processing is performed, a large amount of hardware resources including a memory are required to store or transmit the video data.
[0004] Therefore, the video data is generally compressed using an encoder before being stored or transmitted. A decoder receives the compressed video data, decompresses the received compressed video data, and plays the decompressed video data. Video compression techniques include H.264 / Advanced Video Coding (AVC), High Efficiency Video Coding (HEVC), and Versatile Video Coding (VVC), which has a coding efficiency of more than about 30% compared to HEVC.
[0005] However, as the image size, resolution, and frame rate gradually increase, the amount of data to be encoded also increases. Therefore, a new compression technique that provides higher coding efficiency and better image enhancement effects than existing compression techniques is required.
[0006] In VVC, a geometric partition mode (GPM) is adopted as an inter prediction technique to perform prediction according to a more flexible partition than square and rectangular partition. In the GPM technique, the encoder provides the decoder with a mode index indicating one of the pre-defined geometric partition modes and motion information of two partition regions. The decoder uses the transmitted information to compensate for the motion of the two partition regions and performs a weighted sum of the two partition regions using a blending matrix according to the geometric partition mode, thereby generating a final prediction signal of the current block.
[0007] When a prediction signal of a current block is generated by using motion compensation based on inter prediction, local or global illumination changes can occur between the current frame and the reference frame. Local illumination compensation (LIC) is a technique that can improve the prediction efficiency of the current block by compensating for changed illumination values. Therefore, in order to improve the video coding efficiency and the video quality, a method of efficiently using illumination compensation when performing inter prediction based on GPM is required. SUMMARY
[0008] TECHNICAL PROBLEM
[0009] The present disclosure aims to provide a video coding method and apparatus which applies illumination compensation to a prediction signal of a current block to compensate for local or global illumination changes between a current frame and a reference frame when performing prediction on the current block according to a geometric partitioning mode.
[0010] Technical solutions
[0011] At least one aspect of the present disclosure provides a method of reconstructing a current block by a video decoding apparatus. The method comprises decoding a geometric partitioning mode (GPM) index and a candidate index from a bitstream. The method further comprises constructing a GPM candidate list, and partitioning the current block into two partition blocks according to the GPM index. The method further comprises deriving motion information from the GPM candidate list by using the candidate index for an inter partition block, at least one of the partition blocks being an inter partition block. The method further comprises generating a reference block of the inter partition block by using the motion information. The method further comprises obtaining local illumination compensation parameters (LIC parameters) defining a linear relationship between the reference block and the inter partition block to which local illumination compensation (LIC) is applied. The method further comprises generating a final prediction block of the inter partition block by applying the LIC to the reference block of the inter partition block based on the LIC parameters.
[0012] Another aspect of the present disclosure provides a method of encoding a current block by a video encoding apparatus. The method comprises determining a geometric partitioning mode (GPM) index and a candidate index. The method further comprises constructing a GPM candidate list, and partitioning the current block into two partition blocks according to the GPM index. The method further comprises deriving motion information from the GPM candidate list by using the candidate index for an inter partition block, at least one of the partition blocks being an inter partition block. The method further comprises generating a reference block of the inter partition block by using the motion information. The method further comprises obtaining local illumination compensation parameters (LIC parameters) defining a linear relationship between the reference block and the inter partition block to which local illumination compensation (LIC) is applied. The method further comprises generating a final prediction block of the inter partition block by applying the LIC to the reference block of the inter partition block based on the LIC parameters.
[0013] Yet another aspect of the disclosure provides a method of providing video data to a video decoding device, the method comprising: encoding the video data into a bitstream; and transmitting the bitstream to the video decoding device. Here, the encoding of the video data comprises determining a geometric partition mode (GPM) index and a candidate index. The encoding of the video data further comprises constructing a GPM candidate list; and partitioning a current block into two partition blocks according to the geometric partition mode index. The encoding of the video data further comprises deriving, for at least one inter partition block among the partition blocks, motion information from the GPM candidate list by using the candidate index, the inter partition block being partitioned with motion compensation. The encoding of the video data further comprises generating a reference block of the inter partition block by using the motion information. The encoding of the video data further comprises obtaining local illumination compensation parameters (LIC parameters) defining a linear relationship between the reference block and the inter partition block to which local illumination compensation (LIC) is applied. The encoding of the video data further comprises generating a final prediction block of the inter partition block by applying the LIC to the reference block of the inter partition block based on the LIC parameters.
[0014] Advantages
[0015] As described above, the disclosure provides a video encoding method and device that applies illumination compensation to a prediction signal of a current block to compensate for local or global illumination changes between a current frame and a reference frame when performing prediction on the current block according to a geometric partition mode. Accordingly, the video encoding method and device improve video encoding efficiency and enhance video quality. BRIEF DESCRIPTION OF DRAWINGS
[0016] Figure 1 is a block diagram of a video encoding device that can implement the techniques of the disclosure;
[0017] Figure 2 shows a method of partitioning a block using a quad-tree plus binary-tree ternary-tree (QTBTTT) structure;
[0018] Figure 3A and Figure 3B shows a plurality of intra prediction modes including a wide-angle intra prediction mode;
[0019] Figure 4 shows a surrounding block of a current block;
[0020] Figure 5 is a block diagram of a video decoding device that can implement the techniques of the disclosure;
[0021] Figure 6 is a diagram showing block partitioning according to a geometric partition;
[0022] Figure 7A and 7Bis a diagram showing a straight line that splits a block into two parts;
[0023] Figure 8 is a diagram showing a geometry partition mode (GPM) merge list for geometry-based motion prediction;
[0024] Figure 9 is a block diagram showing a portion of a video decoding apparatus in accordance with at least one embodiment of the present disclosure;
[0025] Figure 10 is a flowchart of the operation of a prediction executor in accordance with at least one embodiment of the present disclosure;
[0026] Figure 11 is a flowchart of a method of performing local illumination compensation in accordance with at least one embodiment of the present disclosure;
[0027] Figure 12 is a diagram showing template definitions according to angle indices in accordance with at least one embodiment of the present disclosure;
[0028] Figure 13 is a diagram showing template definitions in accordance with another embodiment of the present disclosure;
[0029] Figure 14 is a flowchart of a method of encoding a current block by a video encoding apparatus in accordance with at least one embodiment of the present disclosure;
[0030] Figure 15 is a flowchart of a method of reconstructing a current block by a video decoding apparatus in accordance with at least one embodiment of the present disclosure. DETAILED DESCRIPTION
[0031] Some embodiments of the present disclosure are described in detail below with reference to the attached drawing figures, wherein the same reference numerals are used to denote the same elements throughout the several views. Additionally, in the following description, well-known components are not described in detail since they can obscure the understanding of the present disclosure. Further, the description of some embodiments does not apply to other embodiments.
[0032] Figure 1 is a block diagram of a video encoding apparatus that can implement the techniques of the present disclosure. The video encoding apparatus and its components are described below with reference to Figure 1 .
[0033] The encoding apparatus can include a picture partitioner 110, a predictor 120, a subtractor 130, a transformer 140, a quantizer 145, a rearrangement unit 150, an entropy encoder 155, an inverse quantizer 160, an inverse transformer 165, an adder 170, a loop filter unit 180, and a memory 190.
[0034] Each component of the encoding apparatus can be realized as hardware or software, or as a combination of hardware and software. Furthermore, the functions of each component can be realized as software, and a microprocessor can also be implemented to execute the software functions corresponding to each component.
[0035] A video is composed of one or more sequences including a plurality of pictures. Each picture is divided into a plurality of regions, and encoding is performed for each region. For example, a picture can be divided into one or more tiles or / and slices. Here, one or more tiles can be defined as a tile group. Each tile or / and slice is divided into one or more coding tree units (CTUs). Furthermore, each CTU is divided into one or more coding units (CUs) in a tree structure. Information applied to each coding unit (CU) is encoded as the syntax of the CU, and information commonly applied to the CUs included in one CTU is encoded as the syntax of the CTU. Furthermore, information commonly applied to all blocks in one slice is encoded as the syntax of a slice header, and information applied to all blocks constituting one or more pictures is encoded to a picture parameter set (PPS) or a picture header. Furthermore, information commonly referred to by a plurality of pictures is encoded to a sequence parameter set (SPS). Furthermore, information commonly referred to by one or more SPSs is encoded to a video parameter set (VPS). Furthermore, information commonly applied to one tile or tile group can also be encoded as the syntax of a tile or tile group header. The syntax included in the SPS, PPS, slice header, tile or tile group header can be referred to as high-level syntax.
[0036] The picture partitioner 110 determines the size of a coding tree unit (CTU). Information on the size of the CTU (CTU size) is encoded as the syntax of the SPS or PPS, and transmitted to the video decoding apparatus.
[0037] The picture partitioner 110 divides each picture constituting a video into a plurality of coding tree units (CTUs) having a predetermined size, and then recursively partitions the CTUs by using a tree structure. A leaf node in the tree structure becomes a coding unit (CU), which is a basic unit of encoding.
[0038] The tree structure can be a quad-tree (QT) structure in which an upper node (or parent node) is split into four lower nodes (or child nodes) of equal size. The tree structure can also be a binary tree (BT) structure in which an upper node is split into two lower nodes. The tree structure can also be a ternary tree (TT) structure in which an upper node is split into three lower nodes in a 1:2:1 ratio. The tree structure can also be a structure in which two or more of the QT structure, the BT structure, and the TT structure are mixed. For example, a quad-tree plus binary tree (QTBT) structure can be used, or a quad-tree plus binary tree plus ternary tree (QTBTTT) structure can be used. Here, the tree structure that gathers the binary tree plus ternary tree (BTTT) can be referred to as a multi-type tree (MTT).
[0039] Figure 2 is a diagram for describing a method of splitting a block by using a QTBTTT structure.
[0040] As shown in Figure 2 , a CTU can be first split into a QT structure. The quad-tree splitting can be performed recursively until the size of the split block reaches a minimum block size of a leaf node allowed in the QT (MinQTSize). A first flag (QT_split_flag) indicating whether each node of the QT structure is split into four lower nodes, is encoded by the entropy encoder 155 and signaled to the video decoding apparatus. When the leaf node of the QT has a size not greater than a maximum block size of a root node allowed in the BT (MaxBTSize), the leaf node can be further split into at least one of a BT structure or a TT structure. There can be multiple splitting directions in the BT structure and / or the TT structure. For example, two directions can be included, i.e., a direction in which a block of a corresponding node is split in a horizontal direction, and a direction in which the block of the corresponding node is split in a vertical direction. As shown in Figure 2 , when MTT splitting is started, a second flag (mtt_split_flag) indicating whether to split a node, an additional flag indicating a splitting direction (vertical or horizontal), and / or a flag indicating a splitting type (binary or ternary) if the node is split, are all encoded by the entropy encoder 155 and signaled to the video decoding apparatus.
[0041] Alternatively, a CU split flag (split_cu_flag) indicating whether to split a node can be encoded before a first flag (QT_split_flag) indicating whether to split each node into four lower-level nodes is encoded. When the value of the CU split flag (split_cu_flag) indicates not to split each node, the block of the corresponding node becomes a leaf node in the split tree structure and becomes a CU as a basic unit of encoding. When the value of the CU split flag (split_cu_flag) indicates to split each node, the video encoding apparatus starts to encode the first flag first according to the above-described scheme.
[0042] When QTBT is used as another example of a tree structure, there can be two types, i.e., a type in which a block of a corresponding node is split into two blocks of the same size in a horizontal direction (i.e., symmetric horizontal splitting), and a type in which a block of a corresponding node is split into two blocks of the same size in a vertical direction (i.e., symmetric vertical splitting). A split flag (split_flag) indicating whether each node of the BT binary tree structure is split into lower-level blocks and split type information indicating a split type are encoded by the entropy encoder 155 and transmitted to the video decoding apparatus. In addition, there can be an additional type in which a block of a corresponding node is split into two blocks that are not symmetric to each other. This asymmetric form can include a form in which a block of a corresponding node is split into two rectangular blocks having a size ratio of 1:3, or can include a form in which a block of a corresponding node is split in a diagonal direction.
[0043] According to QTBT or QTBTTT splitting from a CTU, a CU can have various sizes. Hereinafter, a block corresponding to a CU to be encoded or decoded (i.e., a leaf node of QTBTTT) will be referred to as a "current block". Since QTBTTT splitting is adopted, the shape of the current block can be rectangular in addition to a square.
[0044] The predictor 120 predicts the current block to generate a prediction block. The predictor 120 includes an intra predictor 122 and an inter predictor 124.
[0045] In general, each current block in a picture can be predictively encoded. The prediction of the current block can be generally performed by an intra prediction technique (using data from a picture containing the current block) or an inter prediction technique (using data from a picture that has been encoded before a picture containing the current block). The inter prediction includes both uni-prediction and bi-prediction.
[0046] The intra predictor 122 predicts pixels within the current block by using pixels (reference pixels) located at the periphery of the current block in the current picture containing the current block. According to a prediction direction, there are various intra prediction modes. For example, as shown in FIG. 1, there are a planar mode, a DC mode, and a directional mode. Figure 3AAs shown, the plurality of intra prediction modes can include 2 non-directional modes (including the planar mode and the DC mode), and can include 65 directional modes. According to each prediction mode, the surrounding pixels to be used and the calculation formula are defined differently.
[0047] To perform efficient directional prediction for a current block having a rectangular shape, additional directional modes (intra prediction modes #67 to #80, #-1 to #-14) shown in dashed arrows in FIG. 6 can be used. Figure 3B Figure 3B In FIG. 6, the arrows indicate the corresponding reference samples used for prediction, and do not represent the prediction direction. The prediction direction is opposite to the direction shown by the arrows. When the current block has a rectangular shape, the wide-angle intra prediction modes are modes that can perform prediction in the opposite direction of a particular directional mode without additional bit transmission. In this case, among the wide-angle intra prediction modes, some of the wide-angle intra prediction modes that can be used for the current block can be determined by the width-to-height ratio of the rectangular current block. For example, when the current block has a rectangular shape with a height smaller than a width, the wide-angle intra prediction modes having an angle smaller than 45 degrees (intra prediction modes #67 to #80) can be used. When the current block has a rectangular shape with a width greater than a height, the wide-angle intra prediction modes having an angle greater than -135 degrees can be used.
[0048] The intra predictor 122 can determine an intra prediction mode for encoding of the current block. In some examples, the intra predictor 122 can encode the current block by using a plurality of intra prediction modes, and can also select an appropriate intra prediction mode to be used from among the tested modes. For example, the intra predictor 122 can calculate rate-distortion values by using rate-distortion analysis on the plurality of tested intra prediction modes, and can select an intra prediction mode having the best rate-distortion characteristics among the tested modes.
[0049] The intra predictor 122 selects an intra prediction mode among the plurality of intra prediction modes, and predicts the current block by using the surrounding pixels (reference pixels) and the arithmetic formula determined according to the selected intra prediction mode. Information about the selected intra prediction mode is encoded by the entropy encoder 155, and is transmitted to the video decoding apparatus.
[0050] The inter predictor 124 generates a prediction block of the current block by using a motion compensation process. The inter predictor 124 searches for a block most similar to the current block in a reference picture that is earlier in encoding and decoding than the current picture, and generates a prediction block of the current block by using the searched block. In addition, a motion vector (MV) is also generated, which corresponds to a displacement between the current block in the current picture and the prediction block in the reference picture. Generally, motion estimation is performed on the luma component, and the motion vector calculated based on the luma component is used for both the luma component and the chroma component. Motion information including information about the reference picture used to predict the current block and information about the motion vector is encoded by the entropy encoder 155 and transmitted to the video decoding apparatus.
[0051] To improve the prediction accuracy, the inter predictor 124 can also perform interpolation on the reference picture or the reference block. In other words, sub-samples between two consecutive integer samples are interpolated by applying filter coefficients to a plurality of consecutive integer samples including the two integer samples. When performing the process of searching for a block most similar to the current block on the interpolated reference picture, the motion vector can be expressed in decimal unit accuracy, not in integer sample unit accuracy. The accuracy or resolution of the motion vector can be set differently for each target area to be encoded (e.g., slice, tile, CTU, CU, etc.). When such adaptive motion vector resolution (AMVR) is applied, information about the motion vector resolution to be applied to each area should be signaled for each target area. For example, when the target area is a CU, information about the motion vector resolution to be applied to each CU is signaled. The information about the motion vector resolution can be information indicating the accuracy of a motion vector difference, which will be described below.
[0052] Furthermore, the inter-frame predictor 124 can perform inter-frame prediction using bidirectional prediction. In the case of bidirectional prediction, two reference images and two motion vectors representing the position of the block most similar to the current block in each reference image are used. The inter-frame predictor 124 selects a first reference image and a second reference image from reference image list 0 (RefPicList0) and reference image list 1 (RefPicList1), respectively. The inter-frame predictor 124 also searches for the block most similar to the current block in each reference image to generate a first reference block and a second reference block. Furthermore, a predicted block for the current block is generated by averaging or weighted averaging the first reference block and the second reference block. In addition, motion information containing information about the two reference images used to predict the current block and information about the two motion vectors is transmitted to the entropy encoder 155. Here, reference image list 0 may consist of images in the pre-reconstructed images that are preceding the current image in display order, while reference image list 1 may consist of images in the pre-reconstructed images that are following the current image in display order. However, this disclosure is not limited to this, and pre-reconstructed images that are following the current image in display order may also be additionally included in reference image list 0. Conversely, pre-reconstructed images preceding the current image can also be additionally included in the reference image list 1.
[0053] To minimize the number of bits required to encode motion information, various methods can be used.
[0054] For example, when the reference image and motion vector of the current block are the same as those of surrounding blocks, the information that can identify the surrounding blocks is encoded to transmit the motion information of the current block to the video decoding device. This method is called merging mode.
[0055] In merge mode, the inter-frame predictor 124 selects a predetermined number of merge candidate blocks (hereinafter referred to as "merge candidates") from the surrounding blocks of the current block.
[0056] like Figure 4 As shown, the surrounding blocks used to derive merge candidates can be all or some of the left block A0, lower left block A1, upper block B0, upper right block B1, and upper left block B2 adjacent to the current block in the current image. Additionally, blocks within a reference image (which may be the same as or different from the reference image used to predict the current block) outside the current image can also be used as merge candidates. For example, blocks in the reference image that are at the same position as the current block, or blocks adjacent to that block at the same position, can be additionally used as merge candidates. If the number of merge candidates selected using the above method is less than a preset number, a zero vector can be added to the merge candidates.
[0057] The inter-frame predictor 124 configures a merge list containing a predetermined number of merge candidates by using peripheral blocks. It selects merge candidates from the merge list to be used as motion information for the current block and generates merge index information to identify the selected candidate. The generated merge index information is encoded by the entropy encoder 155 and transmitted to the video decoding device.
[0058] Merge skip mode is a special case of merge mode. After quantization, when all transform coefficients used for entropy coding are close to zero, only the peripheral block selection information is sent, and the residual signal is not sent. By using merge skip mode, relatively high coding efficiency can be achieved for images with little motion, still images, and screen content images.
[0059] Hereinafter, merge mode and merge skip mode will be collectively referred to as merge / skip mode.
[0060] Another method for encoding motion information is the Advanced Motion Vector Prediction (AMVP) model.
[0061] In AMVP mode, the inter-frame predictor 124 derives candidate predicted motion vectors for the current block by using the surrounding blocks of the current block. The surrounding blocks used to derive the candidate predicted motion vectors can be, for example, blocks such as... Figure 4 The current image shown contains all or some of the following adjacent blocks: left block A0, lower left block A1, upper block B0, upper right block B1, and upper left block B2. Alternatively, blocks within a reference image (which may be the same as or different from the reference image used to predict the current block) outside the current image can be used as peripheral blocks for deriving the predicted motion vector candidates. For example, blocks within the reference image that are at the same location as the current block, or blocks adjacent to that block at the same location, can be used. If the number of motion vector candidates selected using the above method is less than a preset number, zero vectors can be added to the motion vector candidates.
[0062] The inter-frame predictor 124 derives predicted motion vector candidates using motion vectors from surrounding blocks, and determines the predicted motion vector of the current block's motion vector using these candidates. Furthermore, the motion vector difference is calculated by subtracting the predicted motion vector from the current block's motion vector.
[0063] The prediction motion vector can be obtained by applying a predefined function (e.g., center value and average value calculation, etc.) to the prediction motion vector candidate. In this case, the video decoding apparatus also knows the predefined function. Furthermore, since the surrounding blocks used to derive the prediction motion vector candidate are already encoded and decoded blocks, the video decoding apparatus can also already know the motion vectors of the surrounding blocks. Therefore, the video encoding apparatus does not need to encode the information for identifying the prediction motion vector candidate. Thus, in this case, the information about the motion vector difference and the information about the reference picture used to predict the current block are encoded.
[0064] Furthermore, the prediction motion vector can also be determined by a scheme of selecting any one of the prediction motion vector candidates. In this case, the information for identifying the selected prediction motion vector candidate is jointly encoded together with the information about the motion vector difference and the information about the reference picture used to predict the current block.
[0065] The subtractor 130 generates a residual block by subtracting the prediction block generated by the intra predictor 122 or the inter predictor 124 from the current block.
[0066] The transformer 140 transforms a residual signal in the residual block having pixel values in a spatial domain into a transform coefficient in a frequency domain. The transformer 140 can transform the residual signal in the residual block by using an overall size of the residual block as a transform unit, or can divide the residual block into a plurality of sub-blocks and can perform the transformation by using the sub-blocks as transform units. Alternatively, the residual block is divided into two sub-blocks of a transform region and a non-transform region, and the residual signal is transformed by using only the transform region sub-block as a transform unit. Here, the transform region sub-block can be one of two rectangular blocks having a size ratio of 1:1 with a horizontal axis (or a vertical axis) as a reference. In this case, a flag (cu_sbt_flag) indicating that only the sub-block is transformed, directional (vertical / horizontal) information (cu_sbt_horizontal_flag), and / or position information (cu_sbt_pos_flag) are encoded by the entropy encoder 155 and signaled to the video decoding apparatus. Furthermore, the size of the transform region sub-block can have a size ratio of 1:3 with the horizontal axis (or the vertical axis) as a reference. In this case, a flag (cu_sbt_quad_flag) for distinguishing the corresponding division is additionally encoded by the entropy encoder 155 and signaled to the video decoding apparatus.
[0067] Further, the transformer 140 can perform a transform on the residual block in a horizontal direction and a vertical direction, respectively. For the transform, a plurality of types of transform functions or transform matrices can be used. For example, a pair of transform functions for horizontal transform and vertical transform can be defined as a multiple transform set (MTS). The transformer 140 can select a transform function pair having the highest transform efficiency among the MTS, and can transform the residual block in each of the horizontal direction and the vertical direction. Information (mts_idx) about the transform function pair in the MTS is encoded by the entropy encoder 155 and signaled to the video decoding apparatus.
[0068] The quantizer 145 quantizes the transform coefficients output from the transformer 140 using a quantization parameter, and outputs the quantized transform coefficients to the entropy encoder 155. For some blocks or frames, the quantizer 145 can not perform a transform, but directly quantize the relevant residual block. The quantizer 145 can also apply different quantization coefficients (scaling values) according to the positions of the transform coefficients in the transform block. A quantization matrix applied to the quantized transform coefficients arranged in two dimensions can be encoded and signaled to the video decoding apparatus.
[0069] The rearrangement unit 150 can perform a re-alignment of coefficient values on the quantized residual values.
[0070] The rearrangement unit 150 can change a 2D coefficient array into a 1D coefficient sequence by using a coefficient scan. For example, the rearrangement unit 150 can output a 1D coefficient sequence by scanning from a DC coefficient to a high frequency domain coefficient using a zigzag scan or a diagonal scan. Instead of the zigzag scan, a vertical scan scanning a 2D coefficient array in a column direction and a horizontal scan scanning a 2D block type coefficient in a row direction can be used according to the size of the transform unit and the intra prediction mode. In other words, according to the size of the transform unit and the intra prediction mode, a scan to be used can be determined among the zigzag scan, the diagonal scan, the vertical scan, and the horizontal scan.
[0071] The entropy encoder 155 encodes the 1D quantized transform coefficient sequence output from the rearrangement unit 150 using a plurality of encoding schemes such as context-based adaptive binary arithmetic coding (CABAC), exponential Golomb, etc., to generate a bitstream.
[0072] Furthermore, the entropy encoder 155 encodes information related to block partitioning (e.g., CTU size, CTU partition flags, QT partition flags, MTT partition types, MTT partition directions, etc.) to enable the video decoding device to partition blocks identically as the video encoding device. Furthermore, the entropy encoder 155 encodes information regarding a prediction type indicating whether the current block is coded by intra prediction or inter prediction. Depending on the prediction type, the entropy encoder 155 encodes intra prediction information (i.e., information regarding an intra prediction mode) or inter prediction information (a merge index in the case of merge mode, information regarding a reference picture index and motion vector difference in the case of AMVP mode). Furthermore, the entropy encoder 155 encodes quantization-related information (i.e., information regarding a quantization parameter and information regarding a quantization matrix).
[0073] The inverse quantizer 160 inverse quantizes quantized transform coefficients output from the quantizer 145 to generate transform coefficients. The inverse transformer 165 transforms the transform coefficients output from the inverse quantizer 160 from the frequency domain into the spatial domain to reconstruct a residual block.
[0074] The adder 170 adds the reconstructed residual block to the prediction block generated by the predictor 120 to reconstruct the current block. Pixels in the reconstructed current block can be used as reference pixels when intra-predicting subsequent blocks.
[0075] The loop filter 180 performs filtering processes on reconstructed pixels in order to reduce blocking artifacts, ringing artifacts, blurring artifacts, etc. that are generated due to block-based prediction and transform / quantization. The loop filter 180, which is an in-loop filter, can include all or some of a deblocking filter 182, a sample adaptive offset (SAO) filter 184, and an adaptive loop filter (ALF) 186.
[0076] The deblocking filter 182 filters boundaries between reconstructed blocks in order to remove blocking artifacts generated due to encoding / decoding of block units, and the SAO filter 184 and the ALF 186 perform additional filtering processes on the deblocked video. The SAO filter 184 and the ALF 186 are used to compensate for differences between reconstructed pixels and original pixels caused by lossy encoding. The SAO filter 184 applies offsets in units of CTUs to improve subjective picture quality and coding efficiency. On the other hand, the ALF 186 performs block-unit filtering and compensates for distortion by applying different filters depending on the degree of a boundary and a change amount of a corresponding block. Information regarding filter coefficients to be used for the ALF can be encoded and signaled to the video decoding device.
[0077] The reconstructed blocks filtered by the deblocking filter 182, the SAO filter 184 and the ALF 186 are stored in the memory 190. When all blocks in a picture are reconstructed, the reconstructed picture can be used as a reference picture for inter prediction of blocks in a subsequent picture to be encoded.
[0078] The video encoding apparatus can store the encoded video data bitstream into a non-transitory storage medium, or transmit the bitstream to a video decoding apparatus through a communication network.
[0079] Figure 5 is a functional block diagram of a video decoding apparatus that can implement the techniques of the present disclosure. Hereinafter, the video decoding apparatus will be described with reference to Figure 5 The video decoding apparatus and its components will be described.
[0080] The video decoding apparatus can include an entropy decoder 510, a rearrangement unit 515, an inverse quantizer 520, an inverse transformer 530, a predictor 540, an adder 550, a loop filter 560 and a memory 570.
[0081] Like the video encoding apparatus of Figure 1 The components of the video decoding apparatus can be implemented as hardware or software, or a combination of hardware and software. In addition, the functions of the components can be implemented as software, and a microprocessor can be implemented to perform the software functions corresponding to the components.
[0082] The entropy decoder 510 extracts information related to block partitioning by decoding a bitstream generated by the video encoding apparatus, and determines a current block to be decoded, and extracts prediction information and information about a residual signal required to reconstruct the current block.
[0083] The entropy decoder 510 determines the size of a CTU by extracting information about the CTU size from a sequence parameter set (SPS) or a picture parameter set (PPS), and partitions a picture into CTUs having the determined size. In addition, the CTU is determined as the highest layer of a tree structure (i.e., a root node), and partitioning information for the CTU is extracted to partition the CTU by using the tree structure.
[0084] For example, when the QTBTTT structure is used to partition the CTU, a first flag (QT_split_flag) related to QT partitioning is first extracted to partition each node into four lower nodes. In addition, for a node corresponding to a leaf node of the QT, a second flag (mtt_split_flag) related to MTT partitioning, a partition direction (vertical / horizontal), and / or a partition type (binary / ternary) are extracted to partition the corresponding leaf node into the MTT structure. As a result, each node below the leaf node of the QT is recursively partitioned into the BT or TT structure.
[0085] As another example, when a CTU is partitioned using a QTBT structure, a CU partition flag (split_cu_flag) is extracted indicating whether a CU is partitioned. When a corresponding block is partitioned, a first flag (QT_split_flag) can also be extracted. During partitioning, for each node, 0 or more recursive MTT partitions can be performed after 0 or more recursive QT partitions. For example, for a CTU, MTT partitioning can be performed immediately; or conversely, only multiple QT partitions can be performed.
[0086] As another example, when a CTU is partitioned using a QTBT structure, a first flag (QT_split_flag) related to QT partitioning is extracted to partition each node into four lower nodes. In addition, a split flag (split_flag) and split direction information are extracted indicating whether a node corresponding to a leaf node of QT is further partitioned into BT.
[0087] In addition, when the entropy decoder 510 determines a current block to be decoded by using tree structure partitioning, the entropy decoder 510 extracts information about prediction type indicating whether the current block is intra-predicted or inter-predicted. When the prediction type information indicates intra-prediction, the entropy decoder 510 extracts a syntax element for intra-prediction information (intra-prediction mode) of the current block. When the prediction type information indicates inter-prediction, the entropy decoder 510 extracts information representing a syntax element for inter-prediction information, i.e., a motion vector and a reference picture to which the motion vector refers.
[0088] In addition, the entropy decoder 510 extracts quantization-related information and extracts information about quantized transform coefficients of the current block as information about a residual signal.
[0089] The rearrangement unit 515 can transform the 1D quantized transform coefficient sequence entropy-decoded by the entropy decoder 510 into a 2D coefficient array (i.e., a block) in an order opposite to that of the coefficient scanning order performed by the video encoding apparatus.
[0090] The inverse quantizer 520 inverse-quantizes the quantized transform coefficients and inverse-quantizes the quantized transform coefficients by using a quantization parameter. The inverse quantizer 520 can also apply different quantization coefficients (scaling values) to the 2D-arranged quantized transform coefficients. The inverse quantizer 520 can perform inverse quantization by applying a matrix of quantization coefficients (scaling values) from the video encoding apparatus to the 2D array of quantized transform coefficients.
[0091] The inverse transformer 530 generates a residual block for the current block by inverse-transforming the inverse-quantized transform coefficients from the frequency domain to the spatial domain to reconstruct a residual signal.
[0092] Further, when the inverse transformer 530 inverse-transforms a partial region (sub-block) of the transform block, the inverse transformer 530 extracts a flag (cu_sbt_flag) indicating that only the sub-block of the transform block is transformed, directional (vertical / horizontal) information (cu_sbt_horizontal_flag) of the sub-block, and / or position information (cu_sbt_pos_flag). The inverse transformer 530 inverse-transforms the transform coefficients of the corresponding sub-block from the frequency domain to the spatial domain to reconstruct the residual signal, and fills regions not inverse-transformed with "0" values as the residual signal to generate a final residual block for the current block.
[0093] Further, when MTS is applied, the inverse transformer 530 determines transform functions or transform matrices to be applied on each of the horizontal and vertical directions by using MTS information (mts_idx) signaled from the video encoding apparatus. The inverse transformer 530 also performs inverse-transforms on the transform coefficients in the transform block along the horizontal and vertical directions, respectively, by using the determined transform functions.
[0094] The predictor 540 can include an intra-predictor 542 and an inter-predictor 544. When the prediction type of the current block is intra-prediction, the intra-predictor 542 is activated, and when the prediction type of the current block is inter-prediction, the inter-predictor 544 is activated.
[0095] The intra-predictor 542 determines an intra-prediction mode for the current block among a plurality of intra-prediction modes according to syntax elements for the intra-prediction mode extracted from the entropy decoder 510. The intra-predictor 542 also predicts the current block by using surrounding reference pixels of the current block according to the intra-prediction mode.
[0096] The inter-predictor 544 determines a motion vector of the current block and a reference picture referred by the motion vector by using syntax elements for the inter-prediction mode extracted from the entropy decoder 510, and predicts the current block by using the motion vector and the reference picture.
[0097] The adder 550 reconstructs the current block by adding the residual block output from the inverse transformer 530 to the prediction block output from the inter-predictor 544 or the intra-predictor 542. Pixels within the reconstructed current block are used as reference pixels when intra-predicting a subsequent block to be decoded.
[0098] The in-loop filter 560 as an in-loop filter can include a deblocking filter 562, an SAO filter 564, and an ALF 566. The deblocking filter 562 performs deblocking filtering on boundaries between reconstructed blocks in order to remove blocking artifacts caused by block-based decoding. The SAO filter 564 and the ALF 566 perform additional filtering on the deblocking-filtered reconstructed blocks in order to compensate for differences between reconstructed pixels and original pixels caused by lossy encoding. The filtering coefficients of the ALF are determined by using information about the filtering coefficients decoded from the bitstream.
[0099] The reconstructed blocks filtered by the deblocking filter 562, the SAO filter 564, and the ALF 566 are stored in the memory 570. When all blocks in a picture are reconstructed, the reconstructed picture can be used as a reference picture for inter prediction of blocks in a subsequent picture to be encoded.
[0100] Some embodiments of the disclosure relate to encoding and decoding of video pictures as described above. More specifically, the disclosure provides a video encoding method and apparatus that applies illumination compensation to a prediction signal of a current block to compensate for local or global illumination changes between a current frame and a reference frame when performing prediction on the current block according to a geometric partition mode.
[0101] The following embodiments can be performed by the inter predictor 124 in the video encoding apparatus. The following embodiments can also be performed by the inter predictor 544 in the video decoding apparatus.
[0102] The video encoding apparatus can generate the signaling information associated with the present embodiments from a rate-distortion optimization perspective when encoding the current block. The video encoding apparatus can encode the signaling information using the entropy encoder 155 and transmit the encoded signaling information to the video decoding apparatus. The video decoding apparatus can decode the signaling information associated with the decoding of the current block from the bitstream using the entropy decoder 510.
[0103] In the following description, the term "target block" can be used interchangeably with a current block or a coding unit (CU), or can refer to some region of a coding unit.
[0104] In addition, a value of "true" for a flag indicates that the flag is set to 1. In addition, a value of "false" for a flag indicates that the flag is set to 0.
[0105] I-1. Merge / skip mode of inter prediction and merge mode with motion vector difference (MMVD)
[0106] A method of constructing a merge candidate list of motion information in the inter prediction merge / skip mode is described as follows. To support the merge / skip mode, a video encoding device can construct a merge candidate list by selecting a preset number (e.g., 6) of merge candidates.
[0107] The video encoding device searches for spatial merge candidates. As shown in FIG. 2, the video encoding device searches for spatial merge candidates from the surrounding blocks. At most, four spatial merge candidates can be selected. Figure 4
[0108] The video encoding device searches for temporal merge candidates. The video encoding device can append a block at the same position as the current block in a reference picture (not the current picture where the target block is located) as a temporal merge candidate. Here, the reference picture can be the same or different from the reference picture used to predict the current block. One temporal merge candidate can be selected.
[0109] The video encoding device searches for history-based prediction motion vector (HMVP) candidates. The video encoding device can store motion vectors of previous h CUs (where h is a natural number) in a table and use these motion vectors as merge candidates. The table has a size of 6 and stores motion vectors of previous CUs in a first-in-first-out (FIFO) manner. This indicates that at most 6 HMVP candidates can be stored in the table. The video encoding device can set the latest motion vector among the HMVP candidates stored in the table as a merge candidate.
[0110] The video encoding device searches for pair-wise average MVP (PAMVP) candidates. The video encoding device can set an average of motion vectors of the first and second candidates in the merge candidate list as a merge candidate.
[0111] If the merge candidate list cannot be filled with the preset number of merge candidates even after performing all the above search processes, the video encoding device adds a zero motion vector as a merge candidate.
[0112] From the viewpoint of coding efficiency optimization, the video encoding device can determine a merge index indicating one candidate in the merge candidate list. The video encoding device can use the merge index to derive a motion vector predictor (MVP) from the merge candidate list, and then can determine the MVP as the motion vector of the current block. In addition, the video encoding device can signal the merge index to a video decoding device.
[0113] In the skip mode, the video encoding device uses the same motion vector transmission method as in the merge mode, but does not transmit a residual block equal to the difference between the current block and the predicted block.
[0114] The method of constructing the merge candidate list can also be performed by a video decoding device. The video decoding device can decode the merge index. The video decoding device can derive the MVP from the merge candidate list using the merge index, and then determine the MVP as the motion vector of the current block.
[0115] On the other hand, when using the merge mode technique with motion vector difference (MMVD), the video encoding device can derive the MVP from the merge candidate list using the merge index. For example, the first or second candidate in the merge candidate list can be used as the MVP. In addition, the video encoding device can determine the distance index and the direction index from the perspective of coding efficiency optimization. The video encoding device can derive the motion vector difference (MVD) using the distance index and the direction index, and then reconstruct the motion vector of the current block by summing the MVD and the MVP. In addition, the video encoding device can signal the merge index, the distance index, and the direction index to the video decoding device.
[0116] The MMVD technique described above can also be performed by the inter- predictor 544 in a video decoding device. The video decoding device can decode the merge index, the distance index, and the direction index. The video decoding device can construct the merge candidate list, and then can derive the MVP from the merge candidate list using the merge index. The video decoding device can derive the MVD using the distance index and the direction index, and then can sum the MVD and the MVP to reconstruct the motion vector of the current block.
[0117] I-2. Advanced Motion Vector Prediction Mode (AMVP Mode) for Inter- Prediction
[0118] The following embodiments are described with focus on the inter-predictor 124 within a video encoding device, but as mentioned previously, they can also be performed by the inter-predictor 544 in a video decoding device.
[0119] The following describes a method of constructing a candidate list of motion information in the advanced motion vector prediction mode (AMVP mode) for inter-prediction. To support the AMVP mode, the inter-predictor 124 in a video encoding device can select a pre-set number (e.g., two) of candidates to form a candidate list.
[0120] The video encoding device searches for spatial candidates. As shown in FIG. 1C, the video encoding device searches for spatial candidates from the surrounding blocks. At most two spatial candidates can be selected. Figure 4
[0121] The video encoding device searches for temporal candidates. The video encoding device can append a block at the same position as the current block in a reference picture (which is not the current picture containing the target block) as a temporal candidate. Here, the reference picture can be the same as or different from the reference picture used to predict the current block. One temporal candidate can be selected.
[0122] If the candidate list cannot be filled even after performing all the search processes described above, i.e., if the preset number of merge candidates cannot be recruited, the video encoding device adds a zero motion vector as a merge candidate.
[0123] From the viewpoint of coding efficiency optimization, the video encoding device can determine a candidate index indicating one of the candidates in the candidate list. The video encoding device can use the candidate index to derive a motion vector predictor (MVP) from the candidate list. In addition, from the viewpoint of coding efficiency optimization, the video encoding device determines a motion vector and then subtracts the MVP from the motion vector to calculate an MVD. Furthermore, the video encoding device can signal the candidate index and the MVD to the video decoding device.
[0124] The method of constructing the AMVP candidate list described above can be equally performed by the inter predictor 544 in the video decoding device. The video decoding device can decode the candidate index and the MVD. The video decoding device can use the candidate index to derive the MVP from the candidate list. The video decoding device can sum the MVD and the MVP to reconstruct the motion vector of the current block.
[0125] I-3. Geometric partition mode (GPM) for inter prediction
[0126] The following embodiments are described with a focus on a video decoding device, but they can be implemented in a video encoding device in the same or similar manner.
[0127] In Versatile Video Coding (VVC), geometric partition mode (GPM) uses prediction units of various shapes, not just rectangles.
[0128] Figure 6 is a diagram showing block partitioning according to geometric partitioning.
[0129] The video decoding device partitions the current block into two parts by a straight line that is perpendicular to a line segment having an angle θ and a distance ρ with respect to the center of the block. Hereinafter, the two blocks partitioned are referred to as a first block partition and a second block partition. Block partition and partitioned block are used interchangeably. The straight line that partitions the current block is referred to as a partition boundary.
[0130] Here, the center of the block indicates a virtual position where the 1 / 2 position of the height of the block intersects with the 1 / 2 position of the width of the block for the current block before partitioning. The angle θ indicates the angle of rotation counterclockwise from the virtual horizontal axis passing through the center of the block to the line segment perpendicular to the partition boundary. The distance ρ indicates the distance between the center of the block and the partition boundary.
[0131] As described above, the straight line that divides the block into two parts, i.e., the segmentation boundary, divides the current block into two distinct block partitions. The video decoding device uses the aforementioned geometric segmentation to divide the current block into two blocks, namely the first block partition and the second block partition, which perform separate predictions based on the segmentation boundary.
[0132] As an example, in previous GPMs, such as Figure 6 The segmentation boundary information is signaled from the video encoding device to the video decoding device. The video decoding device can use the resolved segmentation boundary information to decode the geometric block partitions of the current block. Here, the segmentation boundary information may include the angle θ and distance ρ relative to the center of the block. Furthermore, the video encoding device may also signal information indicating whether GPM should be applied to the video decoding device.
[0133] In the following text, information about the segmentation boundary based on the geometric segmentation pattern is used interchangeably with geometric segmentation pattern information, geometric segmentation information, or segmentation information.
[0134] Figure 7A and Figure 7B It is a diagram showing the straight lines that divide the block into two parts.
[0135] As another example, a lookup table can be configured to contain combinations of angles and distances that divide the current block into a first partition and a second partition, such as... Figure 7A , Figure 7B As shown in Table 1. Then, the index indicating the angle and distance combination in the lookup table can be signaled from the video encoding device to the video decoding device. The angle and distance combinations corresponding to each index can be defined based on a fixed lookup table according to an agreement between the video encoding device and the video decoding device. As a variation, the lookup table can be adaptively reconfigured according to pre-agreed rules.
[0136] As described above, the geometric segmentation shape is based on a segmentation boundary, which is a straight line representing the two partitions of the block. This straight line information can include an index `distanceIdx` representing the distance `ρ` from the center of the block to the segmentation boundary, and an index `angleIdx` representing the angle `θ` of the line segment perpendicular to the segmentation boundary. The index indicating the angle of the line segment perpendicular to the segmentation boundary can be as follows: Figure 7A The settings are shown below. Furthermore, 64 geometric segmentation shapes based on these angles and distances can be configured as follows: Figure 7B Configure as shown.
[0137] The 64 geometric partition shapes can be represented by using a syntax merge_gpm_partition_idx, which is an index indicating the geometric partition shape, as shown in Table 1. That is, a form in which the current block is partitioned into a first block partition and a second block partition according to various angles and distances can be efficiently signaled by using a single index.
[0138] [Table 1]
[0139]
[0140] The index distanceIdx derived from the example in Figure 7B is a value not including the size of the current block. Accordingly, the actual distance between the pixels in the current block and the straight line can be calculated by using the size information of the current block, the index angleIdx indicating the angle, and the index distanceIdx indicating the distance. Here, the actual distance is a value indicated in pixels.
[0141] On the other hand, the actual distance can be used to calculate the weight of each pixel in the current block. For example, for a single pixel in the first block partition, as the actual distance between the pixel and the straight line increases, the weight of the prediction value of the first block partition increases, and the weight of the prediction value of the second block partition decreases, as described above. For a pixel located on the partition boundary, the two prediction values can use weights having the same value. In this case, the sum of the weights of the two prediction values for a single pixel remains 1.
[0142] The video decoding apparatus performs prediction on two different block partitions of the current block existing within the current picture by using respective motion vectors (mv0 or mv1). The video decoding apparatus applies weight-based blending processing on the prediction block of the first block partition and the prediction block of the second block partition to generate a final prediction block of the current block.
[0143] To perform inter prediction of the current block, the video decoding apparatus obtains a first prediction block by using the motion vector of the first block partition and a second prediction block by using the motion vector of the second block partition. At this time, as described above, in the process of obtaining the prediction block of each block partition, the video decoding apparatus can obtain the prediction block by multiplying different weights according to the pixel position. In addition, in the process of generating the final prediction block from the prediction blocks multiplied by the weights, the video decoding apparatus can use a shift operation and a clipping operation.
[0144] Figure 8 is a table showing a merge list of geometric partition modes (GPMs) for geometric motion prediction.
[0145] As Figure 8As shown, the video decoding device selects motion information for motion prediction from the merge list, and then uses the selected motion information.
[0146] However, unlike the existing block-based motion prediction technique, for geometric motion prediction, the video decoding device performs uni-prediction for each block partition by restricting the prediction direction, as shown. Figure 8 This is because, compared with the block-based motion prediction, in the geometric motion prediction, the memory bandwidth for prediction is doubled when bi-prediction is performed for each block partition. Therefore, in order to efficiently solve the above-mentioned memory bandwidth increase problem, a technique of restricting the directionality of prediction for each block partition can be applied.
[0147] In the case of restricting the directionality of prediction for each block partition in the geometric motion prediction, a GPM merge list for the geometric motion prediction can be generated by using the existing merge list. In order to generate the GPM merge list, the video decoding device first constructs the merge list as described above. Then, the video decoding device can generate the GPM merge list for the geometric motion prediction in the order of prediction directionality and within the list. At this time, the video decoding device adds L0 direction motion information to the GPM merge list to generate a merge candidate for the geometric motion prediction of the first block partition. In addition, the video decoding device can also add L1 direction motion information to the GPM merge list to generate a merge candidate for the geometric motion prediction of the second block partition. In other words, the video decoding device can derive uni-directional motion information from the existing bi-directional motion information, and add the derived motion information to the GPM merge list.
[0148] As shown in Figure 8 In the existing geometric motion prediction method, the video decoding device can directly use the merge candidate in the GPM merge list as the motion information for the motion prediction of the first block partition and the second block partition.
[0149] On the other hand, even in the AMVP mode of inter prediction, the video decoding device can derive GPM-related motion information in a manner similar to the above.
[0150] In a GPM corresponding to an enhanced compression model (ECM) beyond VVC, a video decoding device generates a prediction signal of one region by using motion compensation among two partition blocks, and generates a prediction signal of another region by using intra prediction, and then performs a weighted sum of the prediction signals of the two partition blocks to generate a final prediction signal. This GPM is referred to as an intra GPM. As described above, the two partition blocks are defined as a first partition block and a second partition block, respectively. The video decoding device parses whether to apply intra prediction to the first partition block, and when not to apply intra prediction to the first partition block, parses whether to apply intra prediction to the second partition block. On the other hand, when applying intra prediction to the first partition block, the video decoding device generates a prediction signal of the second partition block by using motion compensation.
[0151] In the ECM, the GPM can apply MMVD (gpm_mmvd) to each partition block, and apply motion information modification based on template matching (gpm_tm). At this time, the video decoding device can apply gpm_mmvd or gpm_tm to a single coding unit (CU).
[0152] II. Local Illumination Compensation (LIC)
[0153] When a prediction signal of a current block is generated by using motion compensation based on inter prediction, a local or global illumination change can occur between the current frame and a reference frame. Local illumination compensation (LIC) is a technique to improve the prediction efficiency of the current block by performing illumination compensation according to changes in a picture capturing environment.
[0154] In order to effectively perform local illumination compensation, the video decoding device derives parameters representing a linear relationship between a peripheral L-shaped template of the current block and a peripheral L-shaped template of a reference block. For example, the parameters of the linear relationship can be derived based on a linear least squares method. At this time, the peripheral L-shaped template of the current block exists in a reconstructed region of the current block. The video decoding device performs local illumination compensation of the current block based on the reference block using the derived parameters of the linear relationship, as shown in Equation 1.
[0155] [Equation 1]
[0156]
[0157] In Equation 1, α and β are parameters of the linear relationship, and represent a slope and an offset, respectively. P(x, y) represents a pixel in the reference block, i.e., a prediction pixel, and P'(x, y) represents a prediction pixel in the current block that has undergone local illumination compensation.
[0158] As described above, local illumination compensation is applied to the current block, and the linear relationship is derived based on the neighboring reference pixels of the reference block and the neighboring reference pixels of the current block. In this case, when there is no available pixel among the neighboring reference pixels of the reference block, the local illumination compensation can not be performed on the current block. When there are available pixels on the left and the top of the reference block, but not on both, the local illumination compensation can be performed by using the available partial pixels.
[0159] The parameters a and b of the linear relationship used in the local illumination compensation are calculated in the same way in the video encoding apparatus and the video decoding apparatus. Therefore, the video encoding apparatus can not separately transmit the parameters of the linear relationship to the video decoding apparatus.
[0160] In the merge mode, when the LIC is applied to the neighboring block selected according to the merge index, the video decoding apparatus performs the LIC by applying the LIC parameters of the neighboring block to the current block. In the AMVP mode, a flag indicating whether to perform the LIC is signaled, and the video decoding apparatus derives the LIC parameters by using the template of the current block and the template of the reference block. In addition, the LIC is applied in the case of uni-prediction, and is not applied when the current block is predicted according to the combined inter-intra prediction (CIIP). The CIIP is a technique that generates an intra-predicted block (which is an intra-predicted block) and an inter-predicted block (which is an inter-predicted block) for the current block, and then performs a weighted sum of the intra-predicted block and the inter-predicted block to generate a final predicted block of the current block.
[0161] In addition, the LIC can be applied to the bi-predicted block. When the LIC is applied to the neighboring block selected according to the merge index in the merge mode, the video decoding apparatus performs the LIC by applying the LIC parameters of the neighboring block to the current block. In the AMVP mode, a flag indicating whether to perform the LIC is signaled, and the video decoding apparatus derives the LIC parameters by using the template of the current block and the template of the reference block. In addition, when the bi-directional optical flow (BDOF) and the decoder-side motion vector refinement (DMVR) are performed, the LIC is not applied. The BDOF is a method of refining bi-directional motion information by using the OF, and the DMVR is a method of refining motion information based on the bilateral matching (BM). When the overlapping block motion compensation (OBMC) is performed and the LIC is applied to the neighboring block of the current block, the video decoding apparatus performs the OBMC on the current block by using the LIC parameters of the neighboring block. The OBMC is a technique of mixing a predicted block based on original motion information and a predicted block based on motion information of a neighboring block to generate a final predicted block.
[0162] When the bi-directional LIC is applied, the video decoding apparatus generates the final prediction signal P' according to Equation 2.
[0163] [Formula 2]
[0164]
[0165] In Formula 2, w can be calculated by using BCW_idx signaled at the CU level. Here, bi-prediction with CU-level weights (BCW) denotes weights used for summing bi-predicted blocks. P0' denotes a L0 prediction block to which local illumination compensation is applied, and P1' denotes a L1 prediction block to which local illumination compensation is applied. At this time, P0'(x, y) and P1'(x, y) can be generated by using the following iterative method.
[0166] A template of a current block is defined as T, a template of a L0 prediction block is defined as T0, and a template of a L1 prediction block is defined as T1. The video decoding apparatus derives parameters a0 and b0, which represent a relationship between T and T0 as a linear model in Formula 3.
[0167] [Formula 3]
[0168]
[0169] A template of the current block corrected according to Formula 3 is defined as T', and the video decoding apparatus derives parameters a1 and b1, which represent a relationship between T' and T1 as a linear model as Formula 4.
[0170] [Formula 4]
[0171]
[0172] A template of the current block corrected according to Formula 4 is defined as T", and the video decoding apparatus corrects a0 and b0 by representing a relationship between T" and T0 as a linear model as Formula 5.
[0173] [Formula 5]
[0174]
[0175] By using a0', b0', a1, and b1 according to Formula 4 and Formula 5, the video decoding apparatus calculates P0'(x, y) and P1'(x, y), and generates a final prediction signal of the current block according to Formula 2.
[0176] The following embodiments are described with a focus on a video decoding apparatus, but they can be implemented in the same or similar manner in a video encoding apparatus.
[0177] III. Embodiments According to the Disclosure
[0178] According to some embodiments, a video decoding apparatus can determine prediction units and transform units, and in response to the current block corresponding to the determined units, can perform prediction and inverse transform by using the determined prediction techniques and prediction modes to ultimately generate a reconstructed block of the current block. Figure 9 The operations shown can be performed by the inverse converter 530, predictor 540, and adder 550 of the video decoding device. On the other hand, with Figure 9 The same operation shown can also be performed by the inverse transformer 165, image segmenter 110, predictor 120, and adder 170 of the video encoding apparatus. In this case, the video decoding apparatus uses the encoded information parsed from the bitstream, while from the perspective of minimizing bit rate distortion, the video encoding apparatus can use encoded information set from a higher level. For ease of description, the following embodiment will be described with focus on the video decoding apparatus.
[0179] like Figure 5 As shown, predictor 540, based on prediction techniques, includes an intra-frame predictor 542 and an inter-frame predictor 544, but as... Figure 9 As shown, the predictor 540 may include all or some of the following: prediction unit determiner 902, prediction technique determiner 904, prediction mode determiner 906, and prediction executor 908.
[0180] When the input video's color format is YUV (YUV420, YUV411, YUV422, YUV444, etc.), the video decoding device can perform prediction and reconstruction of the luminance component, and then perform prediction and reconstruction of the chrominance component. In this way, the luminance and chrominance components can be... Figure 9 The components shown are reconstructed sequentially. On the other hand, when the input video's color format is RGB, the video encoding device can perform a color format conversion from RGB to YUV, and then encode the converted video. Here, in the YUV format, the color format represents the correspondence between luminance component pixels and chrominance component pixels.
[0181] Prediction Unit Determiner 902 determines the prediction unit (PU). Prediction Technique Determiner 904, in response to the prediction unit, determines the prediction technique (e.g., intra-frame prediction, inter-frame prediction, intra-block copy (IBC) mode, palette mode, etc.). Prediction Mode Determiner 906 determines a detailed prediction mode for the given prediction technique. Prediction Executor 908 generates a prediction block for the current block based on the determined prediction mode.
[0182] The inverse converter 530 includes a converter unit determiner 910 and an inverse converter actuator 912. The converter unit determiner 910 determines the converter unit (TU) in response to the inverse quantization signal of the current block, and the inverse converter actuator 912 performs an inverse converter on the converter unit represented by the inverse quantization signal to generate a residual signal.
[0183] The adder 550 sums the prediction block and the residual signal to generate a reconstructed block. The reconstructed block is stored in a memory and can be used to predict other blocks in the future.
[0184] The prediction unit determined by the prediction unit determiner 902 can be a current block or one of sub-blocks partitioned from the current block. In this case, depending on the color format, the prediction unit of the chroma component can correspond in size to the prediction unit of the luma component. Alternatively, the prediction units of the luma and chroma components can be determined separately and the prediction can be performed in response to the prediction unit of the chroma component.
[0185] The prediction technique determiner 904 determines a prediction technique for the prediction unit. As described above, the prediction technique can be one of inter prediction, intra prediction, IBC mode, or palette mode. In this case, the prediction technique of the chroma component can be determined to be the same as the prediction technique of the corresponding luma component without signaling and parsing additional information.
[0186] In one example, when the prediction technique of the current block is not intra prediction, the video decoding device parses a 1-bit flag information. When the parsed flag indicates a skip mode, the video decoding device can determine the prediction mode of the current block to be an inter prediction merge mode or an IBC merge mode. In the skip mode, the video decoding device can also skip inverse transform (i.e., not parse the residual signal) and use the prediction signal instead of the reconstructed signal.
[0187] On the other hand, when the parsed flag does not indicate the skip mode for the current block, the prediction technique determiner 904 can determine the prediction technique of the current block to be one of inter prediction, intra prediction, IBC mode, or palette mode by parsing a series of 1-bit flags.
[0188] For example, when the skip mode is not applied to the current block and it is determined that the prediction technique is inter prediction or IBC mode, the video decoding device parses a 1-bit flag. Based on the parsed flag, the video decoding device can determine the prediction mode of the current block to be a regular merge mode or an AMVP (advanced motion vector prediction) mode.
[0189] The prediction mode determiner 906 determines a detailed prediction mode for the prediction technique.
[0190] As an example, when the prediction technique of the current block is inter prediction, the prediction mode determiner 906 can determine a general merge mode or an AMVP mode as the prediction mode of the current block. In the general merge mode or the AMVP mode, the video decoding device generates prediction blocks according to one or more motion compensations based on the parsed motion information, and performs a weighted sum of the generated prediction blocks to generate a final prediction block of the current block.
[0191] In another example, when the prediction technique of the current block is inter prediction, the prediction mode determiner 906 can determine a geometric partitioning mode (GPM) as the prediction mode of the current block. In the GPM, the video decoding device partitions the current block into two or more partition blocks according to geometric partitioning, generates prediction blocks through one or more motion compensations based on the parsed motion information of the current block, and performs a weighted sum of the generated prediction blocks to generate a final prediction signal of the current block. For example, as described above, the current block can be partitioned into two partition blocks.
[0192] The prediction executor 908 generates a prediction block of the current block according to the determined prediction technique and prediction mode.
[0193] As an example, the prediction executor 908 generates a prediction block of the current block according to the prediction mode, and the adder 550 sums the prediction block of the current block and the residual signal to generate a reconstructed block.
[0194] In addition, the entropy decoder 510 can reconstruct the quantized secondary transform coefficients when the secondary transform is performed. On the other hand, in the case where the secondary transform is not applied, the entropy decoder 510 can reconstruct the quantized primary transform coefficients. The inverse quantizer 520 can apply inverse quantization to the transform coefficients reconstructed based on the quantization parameter to generate inverse quantized transform coefficients.
[0195] The transform unit determiner 910 of the inverse transformer 530 determines a transform unit (TU) for the inverse quantized transform coefficients. At this time, when a single TU is partitioned into a plurality of sub-blocks, a single sub-block can be used as a TU.
[0196] The inverse transformer 530 can determine a non-separable secondary inverse transform kernel and a separable primary inverse transform kernel when the secondary transform is performed. On the other hand, in the case where the secondary transform is not applied, the inverse transformer 530 can determine a separable primary inverse transform kernel or a non-separable primary inverse transform kernel.
[0197] The following description can interchangeably use a geometric partitioning mode, a GPM partitioning mode, and a GPM-based partitioning mode.
[0198] Figure 10 is a flowchart of operations of a prediction executor according to at least one embodiment of the disclosure.
[0199] The operation of the prediction performer 908 in the video decoding apparatus according to the present disclosure is described in detail below. As an example, when the prediction technique for the current block is inter prediction, and the current block is predicted according to GPM, the current block can be predicted according to Figure 10 the sequence shown in FIG. 6.
[0200] When the current block is predicted according to GPM, the video decoding apparatus generates and uses a general motion vector candidate list to generate a GPM (motion vector) candidate list. The GPM candidate list can be a GPM merge list as shown in FIG. 7. Alternatively, the GPM candidate list can be a list containing candidates with bi-directional motion information. Figure 8
[0201] The video decoding apparatus checks whether gpm_mmvd is applied. When gpm_mmvd is not applied, the video decoding apparatus checks whether gpm_intra, i.e., intra GPM, is applied. When gpm_intra is not applied, the video decoding apparatus checks whether gpm_tm is applied.
[0202] When gpm_mmvd is applied, the video decoding apparatus decodes gpm_mmvd_idx, gpm_mode_idx, and merge_idx. After reordering GPM partition modes based on TM, the video decoding apparatus determines a geometric partition mode according to gpm_mode_idx, and determines a motion information predictor of the partitioned block by using merge_idx and the GPM candidate list. Here, reordering GPM partition modes based on TM involves calculating TM costs of 64 partition modes according to Table 1, and reordering GPM partition modes according to the calculated costs. Therefore, gpm_mode_idx indicates one of the reordered partition modes. The video decoding apparatus generates a motion information difference according to gpm_mmvd_idx, and sums the motion information difference with the motion information predictor to generate motion information. As described above, gpm_mmvd_idx can contain a distance index and a direction index. The video decoding apparatus uses the motion information to generate a prediction block of the partitioned block. The video decoding apparatus decodes a matrix index for blending weighted sum, and then blends the prediction block of the partitioned block using a weighted sum matrix corresponding to the weighted sum matrix index, thereby generating a final prediction signal of the current block. Here, the GPM index gpm_mode_idx can be merge_gpm_partition_idx according to Table 1, and indicates one of the GPM partition modes. The merge index merge_index can indicate a motion information predictor within the GPM candidate list.
[0203] When gpm_intra is applied, the video decoding device decodes gpm_mode_idx and merge_idx. After TM-based reordering of partition modes according to GPM, the video decoding device determines a geometric partition mode according to gpm_mode_idx and uses merge_idx and the GPM candidate list to determine motion information for a partition block that is inter-predicted. The video decoding device uses the motion information to generate a prediction block for the partition block that is inter-predicted. In addition, the video decoding device generates a prediction block for a partition block that is intra-predicted. The video decoding device decodes a weighting matrix index for blending, and then blends the prediction block for the partition block that is inter-predicted and the prediction block for the partition block that is intra-predicted using a weighting matrix according to the weighting matrix index to generate a final prediction signal for the current block.
[0204] When gpm_tm is applied, the video decoding device decodes gpm_mode_idx and merge_idx. After TM-based reordering of partition modes according to GPM, the video decoding device determines a geometric partition mode according to gpm_mode_idx and uses merge_idx and the GPM candidate list to determine motion information for a partition block. The video decoding device modifies the motion information, such as a motion vector (MV), based on TM. The video decoding device uses the modified motion information to generate a prediction block for the partition block. The video decoding device decodes a weighting matrix index for blending, and then blends the prediction block for the partition block using a weighting matrix according to the weighting matrix index to generate a final prediction signal for the current block.
[0205] If neither gpm_mmvd, gpm_intra, nor gpm_tm is applied, the video decoding device decodes gpm_mode_idx and merge_idx. After TM-based reordering of partition modes according to GPM, the video decoding device determines a geometric partition mode according to gpm_mode_idx and uses merge_idx and the GPM candidate list to determine motion information for a partition block. The video decoding device uses the motion information to generate a prediction block for the partition block. The video decoding device decodes a weighting matrix index for blending, and then blends the prediction block for the partition block using a weighting matrix according to the weighting matrix index to generate a final prediction signal for the current block.
[0206] The above describes the case where the prediction mode is merge mode, but it can be similarly applied when GPM is applied in AMVP mode.
[0207] Figure 11 is a flowchart of a method of performing local illumination compensation according to at least one embodiment of the present disclosure.
[0208] As an example, according to the sequence shown in Figure 11 The video decoding device can apply LIC to the current block. When LIC is applied, the video decoding device can derive LIC parameters or inherit LIC parameters from the neighboring blocks, according to whether LIC parameters are to be derived. The video decoding device uses the LIC parameters to perform LIC on the current block.
[0209] In another example, when the current block is predicted according to GPM, the video decoding device can parse gpm_mode_idx and merge_idx, and then apply LIC to the partitioned blocks for motion compensation. By signaling a 1-bit flag for each partitioned block, it can be determined whether LIC is applied or not. Alternatively, when LIC is implicitly applied to the reconstructed blocks neighboring the current block, LIC can be applied to the partitioned blocks of the current block. As mentioned above, gpm_mode_idx is an index indicating the GPM partition mode of the current block, and merge_idx is an index of motion information of the partitioned blocks for motion compensation. For example, merge_idx_0 and merge_idx_1 can be used for the first and second partitioned blocks of the current block, respectively.
[0210] Hereinafter, the partitioned blocks for motion compensation are referred to as inter-partitioned blocks.
[0211] The inheritance and derivation of LIC parameters will be described below.
[0212] An example case is that the current block is predicted according to a geometric partition mode, and only one of the two partitioned blocks generates a prediction block according to motion compensation, i.e., only one inter-partitioned block.
[0213] As an example, when the neighboring block providing motion information according to merge_idx uses LIC in the prediction process, the video decoding device can inherit and use the LIC parameters a and b of the neighboring block.
[0214] For the inheritance / derivation of LIC parameters, the templates of the respective partitioned blocks are defined according to the angle index angleIdx, as shown in Figure 12 and Table 2.
[0215]
Table 2
[0216]
[0217] In the example of Figure 12 , "a" and "b" represent positive integers. A represents the upper template above the current block, L represents the left template of the current block, and LT represents the top-left corner of the current block. In Table 2, L+A represents the top-left corner template of the current block, i.e., the L-shaped template.
[0218] In an example case, the surrounding block providing motion information according to the merge_idx can use LIC in the prediction process, and the surrounding block can be located within the template of the inter-partition block being motion compensated. In this case, the video decoding device can inherit the LIC parameters a and b of the block.
[0219] In another example, the surrounding block providing motion information according to the merge_idx can use LIC in the prediction process, and the surrounding block is located within the template of the partition block being motion compensated. In this case, the video decoding device can derive the LIC parameters a and b using the template of the current partition block (i.e., the inter-partition block) and the template of the reference block according to the motion information. At this time, when the motion information is obtained from the LT position in the example, the video decoding device can derive the LIC parameters a and b using the L+A template. Figure 12
[0220] For example, the video decoding device can derive the LIC parameters according to the least square method, such as Equation 6 and Equation 7.
[0221] [Equation 6]
[0222]
[0223] [Equation 7]
[0224]
[0225] In Equation 6 and Equation 7, C(n) is a pixel value within the template of the current partition block for deriving the LIC parameters, and R(n) is a pixel value within the template of the reference block for deriving the LIC parameters. N denotes the number of pixels for deriving the LIC parameters.
[0226] As yet another example, the video decoding device can derive the LIC parameters according to Equation 8 and Equation 9.
[0227] [Equation 8]
[0228]
[0229] [Equation 9]
[0230]
[0231] In Equation 8 and Equation 9, C max and C min denote the maximum value and the minimum value within the template of the current partition block, respectively, and R max and R min denote the maximum value and the minimum value within the template of the reference block, respectively.
[0232] For example, the video decoding device can derive the LIC parameters by using all pixels within the template of the current partition block or the reference block. Alternatively, the video decoding device can sample s samples (where s is a positive integer) within the template and use a portion of the sampled pixels to derive the LIC parameters. As an example, when performing the sampling, the video decoding device can apply different sampling rates to the top template and the left template of the current partition block.
[0233] As an example, when bi-prediction is applied to the current partition block, the LIC parameters can be derived by using the following iterative method. Here, the current partition block corresponding to the inter-partition block can be the first partition block or the second partition block. Let T denote the template of the current partition block, T0denote the template of the L0reference block of the same shape, and T1denote the template of the L1reference block of the same shape. The video decoding device derives parameters a0and b0that represent a linear model of the relationship between T and T0as shown in Equation 10.
[0234] [Equation 10]
[0235]
[0236] Let T' denote the template of the current partition block enhanced according to Equation 10, and the video decoding device derives parameters a1and b1that represent a linear model of the relationship between T' and T1according to Equation 11.
[0237] [Equation 11]
[0238]
[0239] Let T" denote the template of the current partition block enhanced according to Equation 11, and the video decoding device corrects a0and b0by representing T" and T0as a linear model according to Equation 12.
[0240] [Equation 12]
[0241]
[0242] Using a0', b0', a1, and b1according to Equations 11 and 12, the video decoding device calculates P0'(x, y) and P1'(x, y) to which local illumination compensation is applied, and generates a final prediction signal P' of the current partition block according to Equation 13.
[0243] [Equation 13]
[0244]
[0245] In Equation 13, w can be calculated based on the BCW_idx of the surrounding blocks of the current block. At this time, w can use the aspect ratio of the current block or a fixed value.
[0246] As another example, when bi-prediction is applied to the current partition, the LIC parameters can be derived according to the following iterative method. Let the template of the current partition be defined as T, the template of the L0 prediction block of the same shape be defined as To, and the template of the L1 prediction block of the same shape be defined as Ti. The video decoding device derives parameters a1 and b1 representing a linear model between T and Ti as shown in Equation 14.
[0247] [Equation 14]
[0248]
[0249] Let the template of the current partition enhanced according to Equation 14 be defined as T', and the video decoding device derives parameters a0 and b0 representing a linear model between T' and To as shown in Equation 15.
[0250] [Equation 15]
[0251]
[0252] Let the template of the current partition corrected according to Equation 15 be defined as T", and the video decoding device corrects a1 and b1 by representing T" and Ti as a linear model as shown in Equation 16.
[0253] [Equation 16]
[0254]
[0255] By using a0, b0, a1', and b1' according to Equations 15 and 16, the video decoding device calculates P0'(x, y) and P1'(x, y) with local illumination compensation applied, and generates the final prediction signal of the current partition according to Equation 13. As mentioned above, in Equation 13, w can be calculated based on the BCW_idx of the surrounding blocks of the current block. At this time, w can be the aspect ratio of the current block or a fixed value.
[0256] As another example, when bi-prediction is applied to the current partition, the LIC parameters can be derived as follows. Let the template of the current partition be defined as T, the template of the L0 prediction block of the same shape be defined as To, and the template of the L1 prediction block of the same shape be defined as Ti. Then, the video decoding device performs a weighted sum of To and Ti according to Equation 17 to generate a new template T 01 .
[0257] [Equation 17]
[0258]
[0259] In Formula 17, k can be a predefined value. Alternatively, k can be derived based on a picture order count (POC) difference between the current frame and the reference frame. Alternatively, the video decoding device can apply a weighted sum on a portion of the template instead of applying the weighted sum on the entire template.
[0260] The video decoding device can calculate T 01 and the LIC parameters between T 01 and T. The video decoding device can apply a weighted sum on the prediction blocks to generate a single prediction block and apply the LIC parameters to the weighted-summed prediction block to generate the final prediction signal of the current partition. In this case, the weighted sum of the prediction blocks can be performed in the same way as Formula 17.
[0261] As an example, when the current block is predicted according to GPM and possibly bi-predicted for each partition, the video decoding device can omit the step of constructing the GPM (motion vector) candidate list, as shown in Figure 10 .
[0262] As an example, for the template of the current partition defined in Figure 12 , the video decoding device can not apply LIC to the latter partition when the surrounding blocks do not lie within the template of the partition that is subject to motion compensation.
[0263] As an example, the video decoding device can extend the geometric partition boundary of the current block to the template of the current block, as shown in Figure 13 , and then define the template for each partition according to the extended partition boundary. In the example of Figure 13 , "a" and "b" are positive integers.
[0264] In addition, when the surrounding block providing the motion information according to merge_idx uses LIC in the prediction process, the video decoding device can apply LIC to the current partition as follows.
[0265] When the above-mentioned surrounding block lies within the template of the partition that is subject to motion compensation, the video decoding device can inherit the LIC parameters a and b of the block.
[0266] When the above-mentioned surrounding block lies within the template of the partition that is subject to motion compensation, the video decoding device can use the template of the motion-compensated partition and the same-shape template of the same-shape reference block obtained from the motion information. For example, the video decoding device can derive the LIC parameters between the templates according to Formula 6 and Formula 7. Alternatively, the video decoding device can derive the LIC parameters between the templates according to Formula 8 and Formula 9.
[0267] When the above-mentioned surrounding blocks are not located within the template of the partition block for which motion compensation is performed, the video decoding device can not apply LIC to the partition block.
[0268] As an example, to perform LIC as described above, the video decoding device can perform a redundancy check among the surrounding reconstructed blocks of the current block when constructing a motion vector candidate list of the current block according to Figure 8 As an example, to perform LIC as described above, the video decoding device can perform a redundancy check among the surrounding reconstructed blocks of the current block when constructing a motion vector candidate list of the current block according to
[0269] The above describes the case where, in the geometric partition mode, the current block is predicted and only one of the two partition blocks generates a prediction block according to motion compensation. Another example describes the case where, in the geometric partition mode, the current block is predicted and both of the two partition blocks generate prediction blocks according to motion compensation.
[0270] As an example, when LIC is applied only to the first partition block, the video decoding device can inherit or derive the LIC parameters a0 and b0 of the first partition block according to the above-mentioned process (i.e., inheriting / deriving the LIC parameters when only one of the two partition blocks generates a prediction block according to motion compensation).
[0271] As another example, when LIC is applied only to the second partition block, the video decoding device can inherit or derive the LIC parameters a1 and b1 of the second partition block according to the above-mentioned process (i.e., inheriting / deriving the LIC parameters when only one of the two partition blocks generates a prediction block according to motion compensation).
[0272] In yet another example, when LIC is applied to both the first partition block and the second partition block, the video decoding device can inherit or derive the LIC parameters a0 and b0 of the first partition block and the LIC parameters a1 and b1 of the second partition block according to the above-mentioned process (i.e., inheriting / deriving the LIC parameters when only one of the two partition blocks generates a prediction block according to motion compensation).
[0273] As an example, when LIC is applied to the first partition block or the second partition block, the video decoding device can generate the final prediction signal of the current block as follows.
[0274] For the partition block to be processed by LIC, the video decoding device performs LIC according to Equation 18.
[0275] [Equation 18]
[0276]
[0277] In Equation 18, P(x,y) represents the initial prediction signal, α and β represent the LIC parameters, and P'(x,y) represents the prediction signal with LIC applied. The video decoding device can perform a weighted summation of two prediction blocks (i.e., the prediction block of the segmented block with LIC applied and the prediction block of the remaining segmented block) to generate the final prediction signal of the current block.
[0278] As another example, when LIC is applied to both the first and second segmentation regions, the video decoding device can generate the final prediction signal for the current block as follows.
[0279] For two segmented blocks, the video decoding device performs LIC on each of the two segmented blocks according to Equation 18. Here, P(x,y) represents the initial prediction signal, α and β represent the LIC parameters, and P'(x,y) represents the prediction signal with LIC applied. The video decoding device can perform a weighted summation of the two prediction blocks (i.e., the prediction blocks of the two segmented blocks with LIC applied) to generate the final prediction signal for the current block.
[0280] For example, when the LIC parameters of the first segment and the second segment are both inherited, the video decoding device performs LIC on each of the two segment blocks using the inherited parameters. The video decoding device can then perform a weighted summation of the two prediction blocks (i.e., prediction blocks of the two segment blocks for which LIC has been applied based on the inherited LIC parameters) to generate the final prediction signal for the current block.
[0281] As another example, the video decoding device can first perform a weighted summation of the initial prediction blocks of two segmented blocks to generate the prediction block of the current block, and then perform LIC on the prediction block generated by the weighted summation to generate the final prediction signal of the current block.
[0282] As such Figure 13 In the example shown, based on the reference templates of each segmented block and the surrounding templates around the current block, the video decoding device can derive the LIC parameters according to Equations 6 and 7. Alternatively, the video decoding device can derive the LIC parameters according to Equations 8 and 9. By using the derived LIC parameters, the video decoding device can perform LIC on the prediction blocks generated by weighted summation to generate the final prediction signal for the current block.
[0283] As another example, during the construction of the template and reference template for the current block used to derive LIC parameters, the video decoding device does not perform a weighted summation of the templates used to blend the segmented blocks. For example, the video decoding device could be based on... Figure 13The illustrated split boundary generates a template for the first split block and a template for the second split block. As described above, the video decoding device first performs a weighted sum of the initial prediction blocks of the two split blocks to generate a prediction block for the current block. However, to generate the reference template, the video decoding device does not perform a blending of the template of the initial prediction block of the first split block with the template of the initial prediction block of the second split block.
[0284] The execution of LIC with TM-based GPM split pattern reordering will be described below.
[0285] When predicting the current block according to GPM and performing TM-based GPM split pattern reordering, the video decoding device can use the LIC parameters according to the GPM split pattern in Table 1. Figure 13 The templates defined in the example construct a template for each GPM split pattern. At this time, in the process of constructing the template of the current block and the reference template, the weighted sum for blending between the templates of the split blocks is not performed.
[0286] For example, when a split block of the current block has a split block to be subjected to LIC, the video decoding device inherits the LIC parameters of the latter split block from the neighboring blocks. The video decoding device can perform LIC on the template of the split block using the inherited LIC parameters, thereby constructing the template of the current block.
[0287] For example, when a split block of the current block has a split block to be subjected to LIC, the video decoding device can not perform LIC in the process of TM-based GPM split pattern reordering.
[0288] The execution of LIC with TM-based MV refinement will be described below.
[0289] As an example, when predicting the current block according to GPM and performing TM-based MV refinement, the video decoding device can calculate the LIC parameters and then perform TM-based MV refinement. The video decoding device can inherit the LIC parameters from the neighboring blocks or can derive the LIC parameters using the MV before refinement.
[0290] As an example, when TM-based MV refinement is performed after the LIC parameters are calculated, the video decoding device can perform TM-based MV refinement using the LIC-processed template.
[0291] Below, a method of applying local illumination compensation (LIC) when predicting a current block according to a geometric partition mode (GPM) is described by using Figure 14 and Figure 15
[0292] Figure 14 is a flowchart of a method of a video encoding device encoding a current block according to at least one embodiment of the present disclosure.
[0293] The video coding device determines the GPM index and the candidate index (S1400).
[0294] As shown in Table 1, the GPM index can determine a partition boundary according to a geometric partition mode using an angle index angleldx and a distance index distanceIdx.
[0295] The video coding device constructs the GPM candidate list (S1402).
[0296] The video coding device generates a general motion vector candidate list, and then uses the general motion vector candidate list to generate the GPM (motion vector) candidate list.
[0297] The video coding device partitions the current block into two partition blocks according to the GPM index (S1404).
[0298] The video coding device derives motion information from the GPM candidate list by using the candidate index for the inter partition block that uses motion compensation among the partition blocks (S1406).
[0299] Here, there is at least one inter partition block among the partition blocks. When there are two inter partition blocks, the video coding device can use separate candidate indexes to derive motion information for each inter partition block.
[0300] The video coding device generates a reference block of the inter partition block using the motion information (S1408).
[0301] The video coding device obtains LIC parameters (S1410). Here, the LIC parameters define a linear relationship between the reference block and the inter partition block to which LIC is applied.
[0302] For example, when a surrounding block that provides the motion information uses LIC and the surrounding block is located within a template of the inter partition block, the video coding device can inherit the LIC parameters from the surrounding block.
[0303] In another example, when a surrounding block that provides the motion information uses LIC and the surrounding block is located within a template of the inter partition block, the video coding device can derive the LIC parameters by using a least square method based on pixel values within the template of the inter partition block and pixel values within a template of the reference block.
[0304] As yet another example, when a surrounding block that provides the motion information uses LIC and the surrounding block is located within a template of the inter partition block, the video coding device can derive the LIC parameters by using maximum and minimum pixel values within the template of the inter partition block and maximum and minimum pixel values within a template of the reference block.
[0305] As an example, when bi-prediction is applied to the inter-partition block, the video encoding device can calculate, using an iterative method, a first LIC parameter representing a linear relationship between the template of the inter-partition block and the template of the first reference block, and a second LIC parameter representing a linear relationship between the template of the inter-partition block and the template of the second reference block, based on the template of the inter-partition block, the template of the first reference block, and the template of the second reference block.
[0306] As another example, when bi-prediction is applied to the inter-partition block, the video encoding device can perform a weighted sum of the template of the first reference block and the template of the second reference block to generate a reference template, and can derive the LIC parameter using the template of the inter-partition block and the reference template.
[0307] The video encoding device applies LIC to the reference blocks of the inter-partition block based on the LIC parameter, thereby generating a final prediction block for the inter-partition block (S1412).
[0308] As an example, when bi-prediction is applied to the inter-partition block, the video encoding device applies LIC to the first reference block based on the first LIC parameter to generate a first prediction block for the inter-partition block, and applies LIC to the second reference block based on the second LIC parameter to generate a second prediction block for the inter-partition block. The video encoding device can perform a weighted sum of the first prediction block and the second prediction block to generate the final prediction block.
[0309] As another example, when bi-prediction is applied to the inter-partition block, the video encoding device performs a weighted sum of the first reference block and the second reference block to generate a prediction block for the inter-partition block. The video encoding device can apply LIC to the prediction block for the inter-partition block based on the LIC parameter to generate the final prediction block.
[0310] The video encoding device generates a final prediction signal for the current block based on the final prediction block for the inter-partition block (S1414).
[0311] For example, when one of the partition blocks is an inter-partition block, the video encoding device generates prediction blocks for the remaining partition blocks that do not use motion compensation. The video encoding device can blend the final prediction block for the inter-partition block with the prediction blocks for the remaining partition blocks to generate the final prediction signal for the current block.
[0312] In another example, when there are two inter-partition blocks, the video encoding device can blend the final prediction blocks for the inter-partition blocks to generate the final prediction signal for the current block.
[0313] The video encoding device encodes the GPM index and the candidate index (S1416).
[0314] The video encoding device then subtracts the final prediction signal of the current block from the current block to generate a residual signal. The video encoding device transforms / quantizes the residual signal to generate quantized transform coefficients and encodes the quantized transform coefficients.
[0315] Figure 15 is a flowchart of a method of reconstructing a current block by a video decoding device according to at least one embodiment of the present disclosure.
[0316] The video decoding device decodes the GPM index and the candidate index from the bitstream (S1500).
[0317] As shown in Table 1, the GPM index can determine a partition boundary according to a geometric partition mode using an angle index angleldx and a distance index distanceIdx.
[0318] The video decoding device constructs a GPM candidate list (S1502).
[0319] The video decoding device can generate a general motion vector candidate list and then can generate a GPM (motion vector) candidate list using the general motion vector candidate list.
[0320] The video decoding device partitions the current block into two partition blocks according to the GPM index (S1504).
[0321] The video decoding device derives motion information from the GPM candidate list by using the candidate index for an inter partition block that uses motion compensation among the partition blocks (S1506).
[0322] Here, there is at least one inter partition block among the partition blocks. When there are two inter partition blocks, the video decoding device can use separate candidate indexes to derive motion information of each inter partition block.
[0323] The video decoding device generates a reference block of the inter partition block using the motion information (S1508).
[0324] The video decoding device obtains LIC parameters (S1510). Here, the LIC parameters define a linear relationship between the reference block and the inter partition block to which LIC is applied.
[0325] The video decoding device applies LIC to the reference block of the inter partition block based on the LIC parameters to generate a final prediction block of the inter partition block (S1512).
[0326] The video decoding device generates a final prediction signal of the current block based on the final prediction block of the inter partition block (S1514).
[0327] As an example, when one of the split blocks is an inter split block, the video decoding device generates a prediction block for the remaining split blocks that does not use motion compensation. The video decoding device can blend the final prediction blocks of the inter split block with the prediction blocks of the remaining split blocks to generate a final prediction signal for the current block.
[0328] As another example, when there are two inter split blocks, the video decoding device can blend the final prediction blocks of the inter split blocks to generate a final prediction signal for the current block.
[0329] The video decoding device decodes quantized transform coefficients from the bitstream, inverse quantizes / inverse transforms the quantized transform coefficients to reconstruct a residual block. The video decoding device sums the final prediction signal of the current block and the residual block to generate a reconstructed block of the current block.
[0330] Although the steps in each flowchart are described in a particular order, the steps are merely illustrative of the technical idea of some embodiments of the present disclosure. Therefore, a person of ordinary skill in the art to which the present disclosure pertains can perform the steps by changing the sequence described in each drawing, or performing two or more steps in parallel. Therefore, the steps in each flowchart are not limited to the time sequence shown.
[0331] It should be understood that the above description presents exemplary embodiments that can be implemented in various other ways. The functions described in some embodiments can be implemented by hardware, software, firmware, and / or a combination thereof. It should also be understood that the functional components described in the present disclosure are labeled as “... unit” in order to emphasize the possibility of independent implementation.
[0332] On the other hand, various methods or functions described in some embodiments can be implemented as instructions stored in a non-transitory recording medium that can be read and executed by one or more processors. The non-transitory recording medium can include, for example, various types of recording devices in which data is stored in a form readable by a computer system. For example, the non-transitory recording medium can include storage media such as erasable programmable read-only memory (EPROM), a flash drive, an optical disk drive, a magnetic hard disk drive, and a solid state drive (SSD), etc.
[0333] Although embodiments of the present disclosure are described for illustrative purposes, it should be understood by those of ordinary skill in the art to which the present disclosure pertains that various modifications, additions, and substitutions can be made without departing from the idea and scope of the present disclosure. Therefore, the embodiments of the present disclosure are described for brevity and clarity. The scope of the technical idea of the embodiments of the present disclosure is not limited by these descriptions. Therefore, it should be understood by those of ordinary skill in the art to which the present disclosure pertains that the scope of the present disclosure should not be limited by the embodiments explicitly described above, but should be limited by the claims and their equivalents.
[0334] (ref)
[0335] 124: inter predictor
[0336] 544: inter predictor
[0337] 902: prediction unit determiner
[0338] 904: prediction technique determiner
[0339] 906: prediction mode determiner
[0340] 908: prediction executor
[0341] CROSS-REFERENCE TO RELATED APPLICATIONS
[0342] This application claims priority to and the benefit of Korean Patent Application No. 10-2023-0062654, filed May 15, 2023, and Korean Patent Application No. 10-2024-0059176, filed May 3, 2024, the entire contents of which are incorporated herein by reference.
Claims
1. A method for reconstructing a current block using a video decoding device, the method comprising the following steps: Decode the geometric segmentation pattern (GPM) index and candidate index from the bitstream; Construct a GPM candidate list; Based on the geometric segmentation pattern index, the current block is divided into two segmentation blocks; For the inter-frame segmentation block that utilizes motion compensation within the segmentation block, motion information is derived from the GPM candidate list using the candidate index, wherein there is at least one inter-frame segmentation block within the segmentation block; The reference block for the inter-frame segmentation block is generated by using the motion information; Obtain Local Illumination Compensation (LIC) parameters, which define the linear relationship between the reference block and the inter-frame segmentation block to which Local Illumination Compensation (LIC) is applied; and The final prediction block of the inter-frame segmentation block is generated by applying the LIC parameter to the reference block of the inter-frame segmentation block.
2. The method according to claim 1, wherein, The steps to obtain the LIC parameters include: When the peripheral block providing the motion information uses LIC and is located within the template of the inter-frame segmentation block, the LIC parameters are inherited from the peripheral block.
3. The method according to claim 1, wherein, The steps to obtain the LIC parameters include: When the peripheral block providing the motion information uses LIC and is located within the template of the inter-frame segmentation block, the LIC parameters are derived by using the least squares method based on the pixel values within the template of the inter-frame segmentation block and the pixel values within the template of the reference block.
4. The method according to claim 1, wherein, The steps to obtain the LIC parameters include: When the peripheral block providing the motion information uses LIC and is located within the template of the inter-frame segmentation block, the LIC parameters are derived by using the maximum and minimum pixel values within the template of the inter-frame segmentation block and the maximum and minimum pixel values within the template of the reference block.
5. The method according to claim 1, wherein, The steps to obtain the LIC parameters include: When bidirectional prediction is applied to the inter-frame segmentation block, a first LIC parameter and a second LIC parameter are calculated using an iterative method based on the template of the inter-frame segmentation block, the template of the first reference block, and the template of the second reference block. The first LIC parameter represents the linear relationship between the template of the inter-frame segmentation block and the template of the first reference block, while the second LIC parameter represents the linear relationship between the template of the inter-frame segmentation block and the template of the second reference block.
6. The method according to claim 5, wherein, The steps for generating the final prediction block include: The first prediction block of the inter-frame segmentation block is generated by applying the LIC to the first reference block based on the first LIC parameter. A second predicted block of the inter-frame segmentation block is generated by applying the LIC parameter to the second reference block; and The final prediction block is generated by weighted summation of the first prediction block and the second prediction block.
7. The method according to claim 1, wherein, The steps to obtain the LIC parameters include: When bidirectional prediction is applied to the inter-frame segmentation block, a reference template is generated by weighted summation of the templates of the first reference block and the second reference block, and the LIC parameters are derived by using the template of the inter-frame segmentation block and the reference template.
8. The method according to claim 7, wherein, The steps for generating the final prediction block include: The prediction block of the inter-frame segmentation block is generated by weighted summation of the first reference block and the second reference block; and The final prediction block is generated by applying the LIC parameter to the prediction block of the inter-frame segmentation block.
9. The method according to claim 3, wherein, The template of the inter-frame segmentation block is determined based on the segmentation angle indicated by the geometric segmentation pattern index, while the template of the reference block has the same shape as the template of the inter-frame segmentation block.
10. The method according to claim 3, wherein, The template of the inter-frame segmentation block is determined based on the segmentation boundaries indicated by the geometric segmentation pattern index and extended to the template of the current block, and The template of the reference block has the same shape as the template of the inter-frame segmentation block.
11. The method according to claim 1, further comprising the following step: When one of the segmented blocks is an inter-frame segmentation block: Generate prediction blocks for the remaining segmented blocks without using motion compensation; as well as The final prediction signal of the current block is generated by mixing the final prediction block of the inter-frame segmentation block with the prediction blocks of the remaining segmentation blocks.
12. The method according to claim 1, further comprising the following step: When there are two inter-frame segmentation blocks: The final prediction signal of the current block is generated by mixing the final prediction blocks of the inter-frame segmentation blocks.
13. A method for encoding a current block by a video encoding device, the method comprising the following steps: Determine the geometric partitioning pattern (GPM) index and candidate index; Construct a GPM candidate list; Based on the geometric segmentation pattern index, the current block is divided into two segmentation blocks; For the inter-frame segmentation block that utilizes motion compensation within the segmentation block, motion information is derived from the GPM candidate list using the candidate index, wherein there is at least one inter-frame segmentation block within the segmentation block; The reference block for the inter-frame segmentation block is generated by using the motion information; Obtain Local Illumination Compensation (LIC) parameters, which define the linear relationship between the reference block and the inter-frame segmentation block to which Local Illumination Compensation (LIC) is applied; and The final prediction block of the inter-frame segmentation block is generated by applying the LIC parameter to the reference block of the inter-frame segmentation block.
14. The method of claim 13, further comprising the step of: The geometric segmentation pattern index and the candidate index are encoded.
15. The method of claim 13, further comprising the step of: When one of the segmented blocks is the inter-frame segmented block: Generate prediction blocks for the remaining segmented blocks that do not utilize motion compensation; as well as The final prediction signal of the current block is generated by mixing the final prediction block of the inter-frame segmentation block with the prediction blocks of the remaining segmentation blocks.
16. The method of claim 13, further comprising the step of: When there are two inter-frame segmentation blocks: The final prediction signal of the current block is generated by mixing the final prediction blocks of the inter-frame segmentation blocks.
17. A method for providing video data to a video decoding device, the method comprising the following steps: The video data is encoded into a bitstream; as well as The bitstream is sent to the video decoding device. The step of encoding the video data includes: Determine the geometric partitioning pattern (GPM) index and candidate index; Construct a GPM candidate list; Based on the geometric segmentation pattern index, the current block is divided into two segments; For the inter-frame segmentation block that utilizes motion compensation within the segmentation block, motion information is derived from the GPM candidate list using the candidate index, wherein there is at least one inter-frame segmentation block within the segmentation block; The reference block for the inter-frame segmentation block is generated by using the motion information; Obtain Local Illumination Compensation (LIC) parameters, which define the linear relationship between the reference block and the inter-frame segmentation block to which Local Illumination Compensation (LIC) is applied; and The final prediction block of the inter-frame segmentation block is generated by applying the LIC parameter to the reference block of the inter-frame segmentation block.
Citation Information
Patent Citations
/ Analog / Digital Converter
KR100340908B1
Test system
KR1020230062654A
Shopping mall operation system and shopping mall operation method using shopping basket
KR1020240059176A