Bidirectional prediction method based on geometric partitioning
The video coding method employs bidirectional prediction with geometric segmentation to enhance encoding efficiency and video quality, effectively managing memory usage for high-resolution video data.
Patent Information
- Application Number
- PCT/KR2024/016854
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-10-29
- Filing Date
- 2024-10-31
- Publication Date
- 2025-06-05
AI Technical Summary
Existing video coding technologies face challenges in efficiently encoding and decoding high-resolution, high-frame-rate video data due to increased memory usage and data size, necessitating improved encoding efficiency and image quality.
A video coding method and device that performs bidirectional prediction based on effective memory usage by applying geometric segmentation-based prediction to a current block, dividing it into sub-regions, and generating prediction signals using motion compensation and blending matrices.
This approach enhances video encoding efficiency and improves video quality by effectively utilizing memory resources during bidirectional prediction, addressing the limitations of existing compression technologies.
Smart Images

Figure KR2024016854_05062025_PF_FP_ABST
Abstract
Description
Bidirectional prediction method based on geometric segmentation
[0001] The present disclosure relates to a video coding method and device using bidirectional prediction based on geometric segmentation.
[0002] The content described below merely provides background information related to the present invention and does not constitute prior art.
[0003] Since video data has a large amount of data compared to voice data or still image data, it requires a lot of hardware resources, including memory, to store or transmit it without processing for compression.
[0004] Therefore, when storing or transmitting video data, an encoder is used to compress the video data and store or transmit it, and a decoder receives the compressed video data, decompresses it, and plays it back. Examples of such video compression technologies include H.264 / AVC, HEVC (High Efficiency Video Coding), and VVC (Versatile Video Coding), which improves encoding efficiency by about 30% compared to HEVC.
[0005] However, as the size, resolution, and frame rate of images are gradually increasing, and the amount of data that needs to be encoded is also increasing, a new compression technology that has better encoding efficiency and better image quality improvement than existing compression technologies is required.
[0006] GPM (Geometric Partitioning Mode) determines two sub-regions of the current block based on information parsed from the bitstream, obtains a geometric partitioning mode, and motion vector information for each sub-region. GPM generates prediction signals for the two sub-regions by performing motion compensation based on the partitioning mode and motion vector information. GPM generates the final prediction signal of the current block by weighting and adding the prediction signals of the two sub-regions based on a blending matrix according to the geometric partitioning mode. Bidirectional GPM technology generates prediction signals for each sub-region based on bidirectional prediction when applying GPM to generate the final prediction signal of the current block. Meanwhile, when applying bidirectional GPM technology to improve video encoding efficiency and enhance video quality, a method for effectively using memory may be considered.
[0007] The present disclosure aims to provide a video coding method and device that performs bidirectional prediction based on effective memory usage when applying geometric segmentation-based prediction to a current block.
[0008] According to an embodiment of the present disclosure, a method for restoring a current block, performed by a video decoding device, comprises the steps of: decoding merge indices indicating P0 initial motion information of a P0 sub-region of the current block and P1 initial motion information of a P1 sub-region, wherein the P0 sub-region and the P1 sub-region are generated by dividing the current block into two according to a geometric partitioning mode; constructing a merge candidate list of the current block; obtaining the P0 initial motion information and the P1 initial motion information from the merge candidate list based on the merge indices; changing the P0 initial motion information and / or the P1 initial motion information when bidirectional prediction is applied to both the P0 sub-region and the P1 sub-region; generating a P0 initial prediction signal of the P0 sub-region based on the P0 initial motion information or the changed P0 initial motion information, and generating a P1 initial prediction signal of the P1 sub-region based on the P1 initial motion information or the changed P1 initial motion information; A method is provided, comprising: a step of generating a P0 final prediction signal and a P1 final prediction signal based on the P0 initial prediction signal and the P1 initial prediction signal; and a step of generating a final prediction signal of the current block by weighting and adding the P0 final prediction signal and the P1 final prediction signal based on blending according to the geometric division mode.
[0009] According to another embodiment of the present disclosure, a method for encoding a current block, performed by a video encoding apparatus, comprises the steps of: obtaining merge indices indicating P0 initial motion information of a P0 sub-region of the current block and P1 initial motion information of a P1 sub-region, wherein the P0 sub-region and the P1 sub-region are generated by dividing the current block into two according to a geometric partitioning mode; constructing a merge candidate list of the current block; obtaining the P0 initial motion information and the P1 initial motion information from the merge candidate list based on the merge indices; generating a P0 initial prediction signal of the P0 sub-region based on the P0 initial motion information or the changed P0 initial motion information, and generating a P1 initial prediction signal of the P1 sub-region based on the P1 initial prediction information or the changed P0 initial motion information; generating a P0 final prediction signal and a P1 final prediction signal based on the P0 initial prediction signal and the P1 initial prediction signal; And, the method includes a step of generating a final prediction signal of the current block by weighting and adding the P0 final prediction signal and the P1 final prediction signal based on blending according to the geometric division mode.
[0010] According to another embodiment of the present disclosure, a method for providing video data to a video decoding device comprises: encoding the video data into a bitstream; and transmitting the bitstream to the video decoding device, wherein the encoding the video data comprises: obtaining merge indices indicating P0 initial motion information of a P0 sub-region of a current block and P1 initial motion information of a P1 sub-region, wherein the P0 sub-region and the P1 sub-region are generated by dividing the current block into two according to a geometric partitioning mode; constructing a merge candidate list of the current block; obtaining the P0 initial motion information and the P1 initial motion information from the merge candidate list based on the merge indices; generating a P0 initial prediction signal of the P0 sub-region based on the P0 initial motion information or the changed P0 initial motion information, and generating a P1 initial prediction signal of the P1 sub-region based on the P1 initial prediction information or the changed P0 initial motion information; A method is provided, comprising: a step of generating a P0 final prediction signal and a P1 final prediction signal based on the P0 initial prediction signal and the P1 initial prediction signal; and a step of generating a final prediction signal of the current block by weighting and adding the P0 final prediction signal and the P1 final prediction signal based on blending according to the geometric division mode.
[0011] As described above, according to the present embodiment, when applying geometric segmentation-based prediction to a current block, a video coding method and device for performing bidirectional prediction based on effective memory usage are provided, thereby improving video encoding efficiency and improving video quality.
[0012] FIG. 1 is an exemplary block diagram of an image encoding device capable of implementing the techniques of the present disclosure.
[0013] Figure 2 is a drawing for explaining a method of dividing a block using the QTBTTT (QuadTree plus BinaryTree TernaryTree) structure.
[0014] FIGS. 3A and 3B are diagrams illustrating multiple intra prediction modes, including wide-angle intra prediction modes.
[0015] Figure 4 is an example diagram of the surrounding blocks of the current block.
[0016] FIG. 5 is an exemplary block diagram of an image decoding device capable of implementing the techniques of the present disclosure.
[0017] Figure 6 is an example diagram showing block division according to geometric division.
[0018] Figures 7a and 7b are examples showing straight lines that divide a block into two.
[0019] Figure 8 is an example diagram showing a GPM (Geometric Partition Mode) merge list used for geometric motion prediction.
[0020] FIG. 9 is a block diagram illustrating in detail a portion of an image decoding device according to one embodiment of the present disclosure.
[0021] FIG. 10 is a flowchart illustrating prediction of a current block according to GPM according to one embodiment of the present disclosure.
[0022] FIG. 11 is a flowchart illustrating prediction of sub-regions according to one embodiment of the present disclosure.
[0023] FIG. 12 is an exemplary diagram showing a template of an initial prediction block according to one embodiment of the present disclosure.
[0024] FIG. 13 is a flowchart illustrating a change in prediction information in the LA direction according to one embodiment of the present disclosure.
[0025] FIG. 14 is a flowchart illustrating a change in prediction information in the LB direction according to one embodiment of the present disclosure.
[0026] FIG. 15 is an exemplary diagram showing a combined prediction signal according to one embodiment of the present disclosure.
[0027] FIG. 16 is an exemplary diagram showing a template of an initial prediction block according to one embodiment of the present disclosure.
[0028] FIG. 17 is an exemplary diagram showing a template of an initial prediction block according to another embodiment of the present disclosure.
[0029] FIG. 18 is a flowchart illustrating a change in prediction information in the LA direction according to another embodiment of the present disclosure.
[0030] FIG. 19 is a flowchart illustrating a change in prediction information in the LB direction according to another embodiment of the present disclosure.
[0031] Hereinafter, embodiments of the present invention will be described in detail with reference to exemplary drawings. When designating components in each drawing, it should be noted that, where possible, identical components are given the same reference numerals, even if they appear in different drawings. Furthermore, in describing the present embodiments, detailed descriptions of related known structures or functions will be omitted if they are deemed to obscure the gist of the present embodiments.
[0032] FIG. 1 is an exemplary block diagram of an image encoding device capable of implementing the techniques of the present disclosure. Hereinafter, the image encoding device and its subcomponents will be described with reference to the illustration in FIG. 1.
[0033] The video encoding device may be configured to include a picture segmentation unit (110), a prediction unit (120), a subtractor (130), a transformation unit (140), a quantization unit (145), a reordering unit (150), an entropy encoding unit (155), an inverse quantization unit (160), an inverse transformation unit (165), an adder (170), a loop filter unit (180), and a memory (190).
[0034] Each component of the video encoding device may be implemented in hardware, software, or a combination of hardware and software. Furthermore, the functions of each component may be implemented in software, with a microprocessor executing the software functions corresponding to each component.
[0035] A single image (video) is composed of one or more sequences containing multiple pictures. Each picture is divided into multiple regions, and encoding is performed for each region. For example, a single picture is divided into one or more tiles and / or slices. Here, one or more tiles can be defined as a tile group. Each tile or slice is divided into one or more Coding Tree Units (CTUs). Each CTU is then divided into one or more Coding Units (CUs) by a tree structure. Information applied to each CU is encoded as the syntax of the CU, and information commonly applied to CUs included in a CTU is encoded as the syntax of the CTU. In addition, information commonly applied to all blocks within a single slice is encoded as the syntax of the slice header, and information applied to all blocks constituting one or more pictures is encoded in the Picture Parameter Set (PPS) or the picture header. Furthermore, information commonly referenced by multiple pictures is encoded in a Sequence Parameter Set (SPS). And, information commonly referenced by one or more SPS is encoded in a Video Parameter Set (VPS). In addition, information commonly applied to one tile or tile group may be encoded as syntax of a tile or tile group header. Syntaxes included in an SPS, PPS, slice header, tile or tile group header may be referred to as high level syntax.
[0036] The picture segmentation unit (110) determines the size of the CTU. Information about the size of the CTU (CTU size) is encoded as the syntax of SPS or PPS and transmitted to the image decoding device.
[0037] The picture segmentation unit (110) divides each picture constituting an image into a plurality of CTUs having a predetermined size, and then recursively divides the CTUs using a tree structure. A leaf node in the tree structure becomes a CU, which is a basic unit of encoding.
[0038] The tree structure may be a QuadTree (QT) in which an upper node (or parent node) is divided into four lower nodes (or child nodes) of the same size, a BinaryTree (BT) in which an upper node is divided into two lower nodes, or a TernaryTree (TT) in which an upper node is divided into three lower nodes in a 1:2:1 ratio, or a structure that mixes two or more of the QT structures, BT structures, and TT structures. For example, a QTBT (QuadTree plus BinaryTree) structure may be used, or a QTBTTT (QuadTree plus BinaryTree TernaryTree) structure may be used. Here, BTTT may be combined and referred to as a MTT (Multiple-Type Tree).
[0039] Figure 2 is a drawing for explaining a method of dividing a block using the QTBTTT structure.
[0040] As illustrated in FIG. 2, a CTU may first be split into a QT structure. The quadtree splitting may be repeated until the size of the splitting block reaches the minimum block size (MinQTSize) of the leaf node allowed in the QT. A first flag (QT_split_flag) indicating whether each node of the QT structure is split into four nodes of the lower layer is encoded by the entropy encoding unit (155) and signaled to the image decoding device. If the leaf node of the QT is not larger than the maximum block size (MaxBTSize) of the root node allowed in the BT, it may be further split into one or more of the BT structure or the TT structure. There may be multiple splitting directions in the BT structure and / or the TT structure. For example, there may be two directions in which the block of the corresponding node is split horizontally and two directions in which the block is split vertically. As illustrated in FIG. 2, when MTT splitting begins, a second flag (mtt_split_flag) indicating whether nodes have been split, and if splitting has occurred, a flag indicating the splitting direction (vertical or horizontal) and / or a flag indicating the splitting type (Binary or Ternary) are encoded by the entropy encoding unit (155) and signaled to the image decoding device.
[0041] Alternatively, before encoding the first flag (QT_split_flag) indicating whether each node is split into four nodes of a lower layer, a CU split flag (split_cu_flag) indicating whether the node is split may be encoded. If the CU split flag (split_cu_flag) value indicates that the node is not split, the block of the corresponding node becomes a leaf node in the split tree structure and becomes a CU (coding unit), which is a basic unit of encoding. If the CU split flag (split_cu_flag) value indicates that the node is split, the video encoding device starts encoding from the first flag in the above-described manner.
[0042] As another example of a tree structure, when QTBT is used, there may be two types: a type that horizontally splits the block of the corresponding node into two blocks of the same size (i.e., symmetric horizontal splitting) and a type that vertically splits it (i.e., symmetric vertical splitting). A split flag (split_flag) indicating whether each node of the BT structure is split into blocks of a lower layer and split type information indicating the type of split are encoded by the entropy encoding unit (155) and transmitted to the image decoding device. Meanwhile, there may additionally be a type that splits the block of the corresponding node into two blocks of an asymmetrical shape. The asymmetric shape may include a shape that splits the block of the corresponding node into two rectangular blocks with a size ratio of 1:3, or a shape that splits the block of the corresponding node in a diagonal direction.
[0043] A CU can have various sizes depending on the QTBT or QTBTTT partitioning from the CTU. Hereinafter, the block corresponding to the CU to be encoded or decoded (i.e., the leaf node of the QTBTTT) is referred to as the "current block." Depending on the QTBTTT partitioning employed, the current block may be rectangular as well as square.
[0044] The prediction unit (120) predicts the current block and generates a prediction block. The prediction unit (120) includes an intra prediction unit (122) and an inter prediction unit (124).
[0045] In general, each current block within a picture can be predictively coded. Prediction of the current block can typically be performed using either intra-prediction (using data from the picture containing the current block) or inter-prediction (using data from a picture coded before the picture containing the current block). Inter-prediction encompasses both unidirectional and bidirectional prediction.
[0046] The intra prediction unit (122) predicts pixels within the current block using pixels (reference pixels) located around the current block within the current picture including the current block. There are multiple intra prediction modes depending on the prediction direction. For example, as shown in Fig. 3a, the multiple intra prediction modes may include two non-directional modes including the Planar mode and the DC mode, and 65 directional modes. The surrounding pixels to be used and the calculation formula are defined differently depending on each prediction mode.
[0047] For efficient directional prediction for a rectangular current block, directional modes (intra prediction modes 67 to 80 and -1 to -14) indicated by dotted arrows in Fig. 3b may be additionally used. These may be referred to as "wide-angle intra-prediction modes." In Fig. 3b, the arrows point to corresponding reference samples used for prediction, and do not indicate the prediction direction. The prediction direction is opposite to the direction indicated by the arrows. Wide-angle intra-prediction modes are modes that perform prediction in the opposite direction of a specific directional mode without additional bit transmission when the current block is rectangular. At this time, among the wide-angle intra-prediction modes, some wide-angle intra-prediction modes available for the current block may be determined based on the ratio of the width and height of the rectangular current block. For example, wide-angle intra prediction modes (intra prediction modes 67 to 80) having an angle less than 45 degrees are available when the current block is a rectangular shape whose height is smaller than its width, and wide-angle intra prediction modes (intra prediction modes -1 to -14) having an angle greater than -135 degrees are available when the current block is a rectangular shape whose width is larger than its height.
[0048] The intra prediction unit (122) can determine an intra prediction mode to be used to encode the current block. In some examples, the intra prediction unit (122) can encode the current block using multiple intra prediction modes and select an appropriate intra prediction mode to be used from the tested modes. For example, the intra prediction unit (122) can calculate bit-rate distortion values using rate-distortion analysis for multiple tested intra prediction modes and select an intra prediction mode with the best bit-rate distortion characteristics among the tested modes.
[0049] The intra prediction unit (122) selects one intra prediction mode from among multiple intra prediction modes and predicts the current block using surrounding pixels (reference pixels) and an operation formula determined according to the selected intra prediction mode. Information about the selected intra prediction mode is encoded by the entropy encoding unit (155) and transmitted to the image decoding device.
[0050] The inter prediction unit (124) generates a prediction block for the current block using a motion compensation process. The inter prediction unit (124) searches for a block most similar to the current block within reference pictures that were encoded and decoded before the current picture, and generates a prediction block for the current block using the searched block. Then, a motion vector (MV) corresponding to the displacement between the current block within the current picture and the prediction block within the reference picture is generated. Generally, motion estimation is performed on the luma component, and the motion vector calculated based on the luma component is used for both the luma component and the chroma component. The motion information including information on the reference picture used to predict the current block and information on the motion vector is encoded by the entropy encoding unit (155) and transmitted to the image decoding device.
[0051] The inter prediction unit (124) may perform interpolation on a reference picture or a reference block to improve prediction accuracy. That is, subsamples between two consecutive integer samples are interpolated by applying filter coefficients to a plurality of consecutive integer samples including the two integer samples. When a process of searching for a block most similar to the current block is performed on the interpolated reference picture, the motion vector can be expressed up to a precision in decimal units rather than a precision in integer sample units. The precision or resolution of the motion vector can be set differently for each target region to be encoded, such as a slice, tile, CTU, CU, etc. When such adaptive motion vector resolution (AMVR) is applied, information on the motion vector resolution to be applied to each target region must be signaled for each target region. For example, when the target region is a CU, information on the motion vector resolution applied to each CU is signaled. Information on the motion vector resolution may be information indicating the precision of a differential motion vector, which will be described later.
[0052] Meanwhile, the inter prediction unit (124) can perform inter prediction using bi-prediction. In the case of bi-prediction, two reference pictures and two motion vectors indicating the block position most similar to the current block within each reference picture are used. The inter prediction unit (124) selects a first reference picture and a second reference picture from reference picture list 0 (RefPicList0) and reference picture list 1 (RefPicList1), respectively, and searches for a block similar to the current block within each reference picture to generate a first reference block and a second reference block. Then, the first reference block and the second reference block are averaged or weighted averaged to generate a prediction block for the current block. Then, motion information including information on two reference pictures used to predict the current block and information on two motion vectors is transmitted to the entropy encoding unit (155). Here, reference picture list 0 may be composed of pictures that are before the current picture in display order among the restored pictures, and reference picture list 1 may be composed of pictures that are after the current picture in display order among the restored pictures. However, this is not necessarily limited to this, and restored pictures that are after the current picture in display order may be additionally included in reference picture list 0, and conversely, restored pictures that are before the current picture may be additionally included in reference picture list 1.
[0053] Various methods can be used to minimize the number of bits required to encode motion information.
[0054] For example, if the reference picture and motion vector of the current block are identical to those of a neighboring block, the motion information of the current block can be transmitted to the image decoding device by encoding information that can identify the neighboring block. This method is called 'merge mode'.
[0055] In merge mode, the inter prediction unit (124) selects a predetermined number of merge candidate blocks (hereinafter referred to as 'merge candidates') from the surrounding blocks of the current block.
[0056] As the surrounding blocks for deriving merge candidates, all or part of the left block (A0), the lower left block (A1), the upper block (B0), the upper right block (B1), and the upper left block (B2) adjacent to the current block within the current picture may be used, as illustrated in FIG. 4. In addition, a block located within a reference picture (which may or may not be the same as the reference picture used to predict the current block) other than the current picture in which the current block is located may be used as a merge candidate. For example, a block co-located with the current block within the reference picture or blocks adjacent to the block at the co-located block may be additionally used as a merge candidate. If the number of merge candidates selected by the method described above is less than a preset number, a 0 vector is added to the merge candidates.
[0057] The inter prediction unit (124) uses these surrounding blocks to construct a merge list containing a predetermined number of merge candidates. Among the merge candidates included in the merge list, the merge candidate to be used as motion information of the current block is selected and merge index information for identifying the selected candidate is generated. The generated merge index information is encoded by the entropy encoding unit (155) and transmitted to the video decoding device.
[0058] Merge Skip mode is a special case of merge mode. After quantization, when all transform coefficients for entropy encoding are close to zero, only neighboring block selection information is transmitted without transmitting residual signals. By utilizing merge skip mode, relatively high encoding efficiency can be achieved for low-motion images, still images, and screen content images.
[0059] Hereinafter, merge mode and merge skip mode are collectively referred to as merge / skip mode.
[0060] Another method for encoding motion information is Advanced Motion Vector Prediction (AMVP) mode.
[0061] In AMVP mode, the inter prediction unit (124) derives predicted motion vector candidates for the motion vector of the current block using neighboring blocks of the current block. As neighboring blocks used to derive predicted motion vector candidates, all or some of the left block (A0), the lower left block (A1), the upper block (B0), the upper right block (B1), and the upper left block (B2) adjacent to the current block in the current picture as shown in FIG. 4 may be used. In addition, a block located in a reference picture (which may or may not be the same as the reference picture used to predict the current block) other than the current picture in which the current block is located may be used as the neighboring block used to derive predicted motion vector candidates. For example, a block co-located with the current block in the reference picture or blocks adjacent to the block in the co-located block may be used. If the number of motion vector candidates is less than a preset number by the method described above, a 0 vector is added to the motion vector candidates.
[0062] The inter prediction unit (124) derives predicted motion vector candidates using the motion vectors of these surrounding blocks, and determines a predicted motion vector for the motion vector of the current block using the predicted motion vector candidates. Then, the predicted motion vector is subtracted from the motion vector of the current block to produce a differential motion vector.
[0063] The predicted motion vector can be obtained by applying a predefined function (e.g., median, mean, etc.) to the predicted motion vector candidates. In this case, the image decoding device also knows the predefined function. In addition, since the surrounding blocks used to derive the predicted motion vector candidates are blocks that have already been encoded and decoded, the image decoding device also already knows the motion vectors of the surrounding blocks. Therefore, the image encoding device does not need to encode information to identify the predicted motion vector candidates. Therefore, in this case, information about the differential motion vector and information about the reference picture used to predict the current block are encoded.
[0064] Alternatively, the predicted motion vector can be determined by selecting one of the predicted motion vector candidates. In this case, information for identifying the selected predicted motion vector candidate is additionally encoded, along with information about the differential motion vector and the reference picture used to predict the current block.
[0065] The subtractor (130) subtracts the prediction block generated by the intra prediction unit (122) or inter prediction unit (124) from the current block to generate a residual block.
[0066] The transformation unit (140) transforms residual signals within a residual block having pixel values in a spatial domain into transform coefficients in a frequency domain. The transformation unit (140) may transform the residual signals within the residual block using the entire size of the residual block as a transformation unit, or may divide the residual block into a plurality of sub-blocks and use the sub-blocks as transformation units to perform the transformation. Alternatively, the residual signals may be transformed using only the transformation domain sub-block as a transformation unit by dividing the sub-blocks into two sub-blocks, that is, a transformation domain and a non-transform domain. Here, the transformation domain sub-block may be one of two rectangular blocks having a size ratio of 1:1 with respect to the horizontal axis (or vertical axis). In this case, a flag (cu_sbt_flag) indicating that only a sub-block has been converted, directionality (vertical / horizontal) information (cu_sbt_horizontal_flag), and / or position information (cu_sbt_pos_flag) are encoded by the entropy encoding unit (155) and signaled to the image decoding device. In addition, the size of the conversion area sub-block may have a size ratio of 1:3 with respect to the horizontal axis (or vertical axis), and in this case, a flag (cu_sbt_quad_flag) distinguishing the corresponding division is additionally encoded by the entropy encoding unit (155) and signaled to the image decoding device.
[0067] Meanwhile, the transformation unit (140) can individually perform transformations on the residual block in the horizontal and vertical directions. For the transformation, various types of transformation functions or transformation matrices can be used. For example, a pair of transformation functions for horizontal transformation and vertical transformation can be defined as a Multiple Transform Set (MTS). The transformation unit (140) can select one transformation function pair with the best transformation efficiency among the MTS and transform the residual block in the horizontal and vertical directions, respectively. Information (mts_idx) on the transformation function pair selected among the MTS is encoded by the entropy encoding unit (155) and signaled to the image decoding device.
[0068] The quantization unit (145) quantizes the transform coefficients output from the transform unit (140) using quantization parameters and outputs the quantized transform coefficients to the entropy encoding unit (155). The quantization unit (145) may directly quantize a related residual block without transformation for a certain block or frame. The quantization unit (145) may also apply different quantization coefficients (scaling values) according to the positions of the transform coefficients within the transform block. The quantization matrix applied to the quantized transform coefficients arranged in two dimensions may be encoded and signaled to an image decoding device.
[0069] The rearrangement unit (150) can perform rearrangement of coefficient values for quantized residual values.
[0070] The reordering unit (150) can change a two-dimensional coefficient array into a one-dimensional coefficient sequence by using coefficient scanning. For example, the reordering unit (150) can output a one-dimensional coefficient sequence by scanning from the DC coefficient to the coefficients of the high-frequency region by using a zig-zag scan or a diagonal scan. Depending on the size of the transformation unit and the intra prediction mode, a vertical scan that scans the two-dimensional coefficient array in the column direction or a horizontal scan that scans the two-dimensional block-shaped coefficients in the row direction may be used instead of the zig-zag scan. That is, depending on the size of the transformation unit and the intra prediction mode, the scanning method to be used may be determined among the zig-zag scan, the diagonal scan, the vertical scan, and the horizontal scan.
[0071] The entropy encoding unit (155) generates a bitstream by encoding a sequence of one-dimensional quantized transform coefficients output from the rearrangement unit (150) using various encoding methods such as CABAC (Context-based Adaptive Binary Arithmetic Code) and Exponential Golomb.
[0072] In addition, the entropy encoding unit (155) encodes information related to block division, such as CTU size, CU division flag, QT division flag, MTT division type, and MTT division direction, so that the image decoding device can divide the block in the same manner as the image encoding device. In addition, the entropy encoding unit (155) encodes information about a prediction type indicating whether the current block is encoded by intra prediction or inter prediction, and encodes intra prediction information (i.e., information about an intra prediction mode) or inter prediction information (information about an encoding mode of motion information (merge mode or AMVP mode), a merge index in the case of a merge mode, and a reference picture index and a differential motion vector in the case of an AMVP mode) according to the prediction type. In addition, the entropy encoding unit (155) encodes information related to quantization, that is, information about a quantization parameter and information about a quantization matrix.
[0073] The inverse quantization unit (160) inversely quantizes the quantized transform coefficients output from the quantization unit (145) to generate transform coefficients. The inverse transform unit (165) transforms the transform coefficients output from the inverse quantization unit (160) from the frequency domain to the spatial domain to restore the residual block.
[0074] An adder (170) adds the restored residual block and the predicted block generated by the prediction unit (120) to restore the current block. The pixels within the restored current block are used as reference pixels when intra-predicting the next block.
[0075] The loop filter unit (180) performs filtering on restored pixels to reduce blocking artifacts, ringing artifacts, blurring artifacts, etc. that occur due to block-based prediction and transformation / quantization. The loop filter unit (180) may include all or part of a deblocking filter (182), a sample adaptive offset (SAO) filter (184), and an adaptive loop filter (ALF, 186) as an in-loop filter.
[0076] The deblocking filter (182) filters the boundaries between restored blocks to remove blocking artifacts caused by block-based encoding / decoding, and the SAO filter (184) and the ALF (186) perform additional filtering on the deblocking-filtered image. The SAO filter (184) and the ALF (186) are filters used to compensate for the differences between restored pixels and original pixels caused by lossy coding. The SAO filter (184) improves not only subjective image quality but also encoding efficiency by applying an offset in units of CTUs. In contrast, the ALF (186) performs block-based filtering, and compensates for distortion by applying different filters by distinguishing the edges and degrees of variation of the corresponding block. Information on filter coefficients to be used in the ALF can be encoded and signaled to an image decoding device.
[0077] The restored blocks filtered through the deblocking filter (182), SAO filter (184), and ALF (186) are stored in the memory (190). When all blocks within a picture are restored, the restored picture can be used as a reference picture for inter-predicting blocks within a picture to be encoded later.
[0078] The video encoding device can store the bitstream of encoded video data on a non-transitory storage medium or transmit it to the video decoding device using a communication network.
[0079] FIG. 5 is an exemplary block diagram of an image decoding device capable of implementing the techniques of the present disclosure. Hereinafter, the image decoding device and its subcomponents will be described with reference to FIG. 5.
[0080] The video decoding device may be configured to include an entropy decoding unit (510), a rearrangement unit (515), an inverse quantization unit (520), an inverse transformation unit (530), a prediction unit (540), an adder (550), a loop filter unit (560), and a memory (570).
[0081] Similar to the video encoding device of FIG. 1, each component of the video decoding device may be implemented in hardware, software, or a combination of hardware and software. Furthermore, the functions of each component may be implemented in software, with a microprocessor executing the software functions corresponding to each component.
[0082] The entropy decoding unit (510) decodes the bitstream generated by the image encoding device to extract information related to block division, thereby determining the current block to be decoded, and extracts prediction information, information on residual signals, etc. required to restore the current block.
[0083] The entropy decoding unit (510) extracts information about the CTU size from the Sequence Parameter Set (SPS) or the Picture Parameter Set (PPS), determines the size of the CTU, and divides the picture into CTUs of the determined size. Then, the CTU is determined as the top layer of the tree structure, i.e., the root node, and the CTU is divided using the tree structure by extracting division information about the CTU.
[0084] For example, when splitting a CTU using the QTBTTT structure, first, the first flag (QT_split_flag) related to the splitting of QT is extracted, and each node is split into four nodes of the lower layer. Then, for the nodes corresponding to the leaf nodes of QT, the second flag (mtt_split_flag) related to the splitting of MTT and the split direction (vertical / horizontal) and / or split type (binary / ternary) information are extracted, and the corresponding leaf nodes are split into the MTT structure. Accordingly, each node below the leaf nodes of QT are split recursively into the BT or TT structure.
[0085] As another example, when splitting a CTU using the QTBTTT structure, the CU split flag (split_cu_flag) indicating whether the CU is split is first extracted, and if the block is split, the first flag (QT_split_flag) may be extracted. During the splitting process, each node may undergo zero or more repeated QT splits followed by zero or more repeated MTT splits. For example, a CTU may undergo an MTT split right away, or conversely, may undergo only multiple QT splits.
[0086] As another example, when splitting a CTU using the QTBT structure, the first flag (QT_split_flag) related to the splitting of QT is extracted, and each node is split into four nodes of the lower layer. Furthermore, for nodes corresponding to leaf nodes of QT, a split flag (split_flag) indicating whether to further split into BTs and splitting direction information are extracted.
[0087] Meanwhile, when the entropy decoding unit (510) determines the current block to be decoded by using the division of the tree structure, it extracts information on the prediction type indicating whether the current block is intra-predicted or inter-predicted. If the prediction type information indicates intra-prediction, the entropy decoding unit (510) extracts syntax elements for intra-prediction information (intra-prediction mode) of the current block. If the prediction type information indicates inter-prediction, the entropy decoding unit (510) extracts syntax elements for inter-prediction information, i.e., information indicating a motion vector and a reference picture referenced by the motion vector.
[0088] Additionally, the entropy decoding unit (510) extracts information about the quantized transform coefficients of the current block as information related to quantization and information about residual signals.
[0089] The rearrangement unit (515) can change the sequence of one-dimensional quantized transform coefficients entropy-decoded in the entropy decoding unit (510) back into a two-dimensional coefficient array (i.e., block) in the reverse order of the coefficient scanning performed by the image encoding device.
[0090] The inverse quantization unit (520) inversely quantizes the quantized transform coefficients and inversely quantizes the quantized transform coefficients using the quantization parameters. The inverse quantization unit (520) may also apply different quantization coefficients (scaling values) to the quantized transform coefficients arranged in two dimensions. The inverse quantization unit (520) may perform inverse quantization by applying a matrix of quantized coefficients (scaling values) from an image encoding device to a two-dimensional array of quantized transform coefficients.
[0091] The inverse transform unit (530) inversely transforms the inverse quantized transform coefficients from the frequency domain to the spatial domain to restore residual signals, thereby generating a residual block for the current block.
[0092] In addition, when the inverse transform unit (530) inversely transforms only a portion of a transform block (sub-block), it extracts a flag (cu_sbt_flag) indicating that only a sub-block of the transform block has been transformed, directionality (vertical / horizontal) information (cu_sbt_horizontal_flag) of the sub-block, and / or position information (cu_sbt_pos_flag) of the sub-block, and inversely transforms the transform coefficients of the corresponding sub-block from the frequency domain to the spatial domain to restore residual signals, and fills “0” values with residual signals for areas that have not been inversely transformed, thereby generating a final residual block for the current block.
[0093] In addition, when MTS is applied, the inverse transform unit (530) determines a transform function or a transform matrix to be applied in the horizontal and vertical directions using MTS information (mts_idx) signaled from the image encoding device, and performs inverse transform on the transform coefficients within the transform block in the horizontal and vertical directions using the determined transform function.
[0094] The prediction unit (540) may include an intra prediction unit (542) and an inter prediction unit (544). The intra prediction unit (542) is activated when the prediction type of the current block is intra prediction, and the inter prediction unit (544) is activated when the prediction type of the current block is inter prediction.
[0095] The intra prediction unit (542) determines the intra prediction mode of the current block among a plurality of intra prediction modes from the syntax elements for the intra prediction mode extracted from the entropy decoding unit (510), and predicts the current block using reference pixels around the current block according to the intra prediction mode.
[0096] The inter prediction unit (544) uses the syntax elements for the inter prediction mode extracted from the entropy decoding unit (510) to determine the motion vector of the current block and the reference picture referenced by the motion vector, and predicts the current block using the motion vector and the reference picture.
[0097] An adder (550) adds the residual block output from the inverse transform unit (530) and the predicted block output from the inter prediction unit (544) or the intra prediction unit (542) to restore the current block. The pixels within the restored current block are used as reference pixels when intra-predicting a block to be decoded later.
[0098] The loop filter unit (560) may include a deblocking filter (562), an SAO filter (564), and an ALF (566) as in-loop filters. The deblocking filter (562) deblocks the boundaries between restored blocks to remove blocking artifacts caused by block-by-block decoding. The SAO filter (564) and the ALF (566) perform additional filtering on restored blocks after deblocking filtering to compensate for differences between restored pixels and original pixels caused by lossy coding. The filter coefficients of the ALF are determined using information about filter coefficients decoded from the non-stream.
[0099] The restored blocks filtered through the deblocking filter (562), SAO filter (564), and ALF (566) are stored in the memory (570). When all blocks within a picture are restored, the restored picture is used as a reference picture for inter-predicting blocks within a picture to be encoded later.
[0100] The present embodiment relates to encoding and decoding of images (videos) as described above. More specifically, the present invention provides a video coding method and device that performs bidirectional prediction based on efficient memory usage when applying geometric segmentation-based prediction to a current block.
[0101] The following embodiments may be performed by a prediction unit (120) within a video encoding apparatus. Additionally, they may be performed by a prediction unit (540) within a video decoding apparatus.
[0102] The video encoding device can generate signaling information related to the present embodiment in terms of rate distortion optimization in encoding the current block. The video encoding device can encode the signaling information using the entropy encoding unit (155) and then transmit it to the video decoding device. The video decoding device can decode the signaling information related to the decoding of the current block from the bitstream using the entropy decoding unit (510).
[0103] In the following description, the term "target block" may be used interchangeably with the current block or coding unit (CU). Alternatively, the term "target block" may also refer to a portion of a coding unit.
[0104] Also, a value of a flag being true indicates that the flag is set to 1. Also, a value of a flag being false indicates that the flag is set to 0.
[0105] I-1. Merge / Skip Mode and MMVD in Inter Prediction
[0106] Hereinafter, a method for constructing a merge candidate list of motion information in the merge / skip mode of inter prediction is described. To support the merge / skip mode, a video encoding device can construct a merge candidate list by selecting a preset number of merge candidates (e.g., 6).
[0107] The video encoding device searches for spatial merge candidates. The video encoding device searches for spatial merge candidates from surrounding blocks, as illustrated in FIG. 4. Up to four spatial merge candidates can be selected.
[0108] A video encoding device searches for temporal merge candidates. The video encoding device may add a co-located block as a temporal merge candidate, which is a block located in the same location as the current block within a reference picture (which may or may not be the same as the reference picture used to predict the current block) other than the current picture in which the target block is located. Only one temporal merge candidate may be selected.
[0109] A video encoding device searches for HMVP (History-based Motion Vector Predictor) candidates. The video encoding device can store the motion vectors of the previous h CUs (where h is a natural number) in a table and use them as merge candidates. The table size is 6, and the motion vectors of the previous CUs are stored in a First-in First Out (FIFO) manner. This indicates that up to 6 HMVP candidates are stored in the table. The video encoding device can set the most recent motion vectors among the HMVP candidates stored in the table as merge candidates.
[0110] The video encoding device searches for PAMVP (Pairwise Average MVP) candidates. The video encoding device can set the motion vector average of the first and second candidates in the merge candidate list as the merge candidate.
[0111] If the merge candidate list cannot be filled even after performing all of the above-described search processes (i.e., if the preset number cannot be filled), the video encoding device adds a zero motion vector as a merge candidate.
[0112] In terms of optimizing encoding efficiency, a video encoding device can determine a merge index that indicates one candidate within a merge candidate list. The video encoding device can then derive a motion vector predictor (MVP) from the merge candidate list using the merge index, and then determine the MVP as the motion vector of the current block. Furthermore, the video encoding device can signal the merge index to a video decoding device.
[0113] The video encoding device uses the same motion vector transmission method as the merge mode in the skip mode, but does not transmit the residual block corresponding to the difference between the current block and the predicted block.
[0114] The method for constructing the aforementioned merge candidate list can be performed in the same manner by a video decoding device. The video decoding device can decode the merge index. The video decoding device can then derive an MVP from the merge candidate list using the merge index, and then determine the MVP as the motion vector of the current block.
[0115] Meanwhile, when using MMVD (Merge mode with Motion Vector Difference) technology, the video encoding device can derive an MVP from the merge candidate list using a merge index. For example, the first or second candidate in the merge candidate list can be used as the MVP. In addition, in terms of optimizing encoding efficiency, the video encoding device determines a distance index and a direction index. The video encoding device can derive a motion vector difference (MVD) using the distance index and the direction index, and then reconstruct the motion vector of the current block by adding the MVD and the MVP. In addition, the video encoding device can signal the merge index, the direction index, and the direction index to the video decoding device.
[0116] The aforementioned MMVD technique can be performed in the same manner by the inter prediction unit (544) within the video decoding device. The video decoding device can decode a merge index, a distance index, and a direction index. The video decoding device can construct a merge candidate list and then derive an MVP from the merge candidate list using the merge index. The video decoding device can derive an MVD using the distance index and the direction index, and then reconstruct the motion vector of the current block by adding the MVD and the MVP.
[0117] I-2. GPM (Geometric Partitioning Mode) technology
[0118] The following embodiments are described with a focus on an image decoding device, but can be implemented identically or similarly in an image encoding device.
[0119] In VVC (Versatile Video Coding), GPM (Geometric Partitioning Mode) uses prediction units of various shapes other than square shapes.
[0120] Figure 6 is an example diagram showing block division according to geometric division.
[0121] The video decoding device divides the current block into two parts based on a straight line perpendicular to a line segment having a constant angle (θ) and a constant distance (ρ) from the center of the block. Hereinafter, the two divided blocks are referred to as the first block partition and the second block partition. The term "block partition" is used interchangeably with the term "partitioned block." The straight line that divides the current block into two parts is referred to as the partitioning boundary.
[0122] Here, the center of the block represents a virtual position where the position at half the height of the block and the position at half the width of the block intersect with respect to the current block before division. The angle (θ) represents the angle rotated counterclockwise from the virtual horizontal axis passing through the center of the block to the line segment perpendicular to the division boundary. The distance (ρ) represents the distance between the center of the block and the division boundary.
[0123] As described above, the straight line that divides the block, i.e., the division boundary, divides the current block into two different block divisions. The image decoding device uses the geometric division described above to divide the current block into two blocks, a first block division and a second block division, based on the division boundary, for performing separate predictions.
[0124] As an example, in a conventional GPM, information on segmentation boundaries, such as the example in FIG. 6, is signaled from an image encoding device to an image decoding device. The image decoding device can decode geometric block segments of the current block using the information on the parsed segmentation boundaries. Here, the information on the segmentation boundaries can include an angle (θ) and a distance (ρ) based on the center of the block. In addition, the image encoding device can additionally signal information indicating whether or not to apply the GPM to the image decoding device.
[0125] Hereinafter, information on the segmentation boundary according to the geometric segmentation mode is used interchangeably with geometric segmentation mode information, geometric segmentation information, or segmentation information.
[0126] Figures 7a and 7b are examples showing straight lines that divide a block into two.
[0127] As another example, a lookup table including combinations of angles and distances for dividing a current block into a first block division and a second block division, as shown in FIGS. 7A and 7B and Table 1, may be constructed. Thereafter, an index indicating a combination of angles and distances within the lookup table may be signaled from a video encoding device to a video decoding device. The combination of angles and distances corresponding to each index may be defined based on a fixed lookup table according to an agreement between the video encoding device and the video decoding device. As another example, the lookup table may be adaptively reconfigured according to a pre-agreed rule.
[0128] As described above, the geometric segmentation shape is based on the segmentation boundary, which is a straight line representing the division of the block. Information about this straight line may include an index distanceIdx representing the distance (ρ) from the center of the block to the segmentation boundary, and an index angleIdx representing the angle (θ) of the line segment perpendicular to the segmentation boundary. The index representing the angle of the line segment perpendicular to the segmentation boundary may be set as illustrated in Fig. 7a. In addition, 64 geometric segmentation shapes according to these angles and distances may be set as illustrated in Fig. 7b.
[0129] 64 geometric partition shapes can be signaled using the merge_gpm_partition_idx syntax, which is an index indicating a geometric partition shape, as shown in Table 1. That is, the shapes that divide the current block into the first block partition and the second block partition according to various angles and various distances can be efficiently signaled using a single index.
[0130]
[0131]
[0132] The index distanceIdx derived from the example in Fig. 7b is a value that excludes the size of the current block. Therefore, the actual distance between a pixel in the current block and a straight line can be calculated using the size information of the current block, the index angleIdx indicating the angle, and the index distanceIdx indicating the distance. Here, the actual distance is a value expressed in pixel units.
[0133] Meanwhile, weights for each pixel within the current block can be calculated using the actual distance. For example, for a pixel within the first block partition, as the actual distance between the pixel and the straight line increases, the weight of the predictor of the first block partition, as described above, may increase, and the weight of the predictor of the second block partition may decrease. For pixels located on the partition boundary, the two predictors may use weights with the same value. In this case, the sum of the weights of the two predictors for a single pixel remains 1.
[0134] For two different block divisions of the current block within the current picture, the video decoding device performs prediction using each motion vector (mv0 or mv1). The video decoding device applies a weighted-sum-based blending matrix to the prediction block for the first block division and the prediction block for the second block division to generate a final prediction block of the current block.
[0135] In order to perform inter prediction of the current block, the video decoding device obtains a first prediction block using a motion vector for the first block division, and obtains a second prediction block using a motion vector for the second block division. At this time, the video decoding device, in the process of obtaining the prediction block for each block division, may obtain the prediction block in a form in which different weights are multiplied according to pixel positions as described above. In addition, the video decoding device may use a shift operation and a clipping operation in the process of generating a final prediction block from the prediction blocks in a form in which weights are multiplied.
[0136] Figure 8 is an example diagram showing a GPM (Geometric Partition Mode) merge list used for geometric motion prediction.
[0137] As in the example of Fig. 8, the video decoding device can select motion information for motion prediction from a merge list and then use the selected motion information.
[0138] However, unlike the existing block-based motion prediction technology, for geometric motion prediction, the image decoding device performs unidirectional prediction for one block division by limiting the prediction directionality, as illustrated in FIG. 8. This is because, compared to block-based motion prediction, when motion prediction is performed according to bidirectional prediction for each block division in geometric motion prediction, the memory bandwidth used for prediction doubles. Therefore, to efficiently solve the aforementioned memory bandwidth increase problem, a technology for limiting the prediction directionality for each block division can be applied.
[0139] In the case where the direction of prediction is limited for each block partition in geometric motion prediction, a GPM merge list for geometric motion prediction can be generated using an existing merge list rule. To generate a GPM merge list, the image decoding device first configures the merge list as described above. Thereafter, the image decoding device can generate a GPM merge list for geometric motion prediction from the merge list according to the prediction direction and the order in the list. At this time, the image decoding device adds motion information in the L0 direction to the GPM merge list to generate a merge candidate for geometric motion prediction of the first block partition. In addition, the image decoding device can add motion information in the L1 direction to the GPM merge list to generate a merge candidate for geometric motion prediction of the second block partition. That is, the image decoding device can derive unidirectional motion information in one direction from the existing bidirectional motion information, and then add the derived motion information to the GPM merge list.
[0140] As in the example of Fig. 8, in the existing geometric motion prediction method, the image decoding device can directly use the merge candidates in the GPM merge list as motion information for motion prediction of the first block division and the second block division.
[0141] Beyond VVC, the Enhanced Compression Model (ECM) adds a technology to GPM that generates a prediction signal for at least one of the two sub-regions based on intra prediction. In this case, if both sub-regions are intra-predicted, it is classified as SGPM (Spatial GPM) technology. The two sub-regions are referred to as sub-region 0 and sub-region 1, respectively. GPM extracts information from the bitstream indicating whether sub-region 0 is intra-predicted, and if sub-region 0 is not intra-predicted, extracts information indicating whether sub-region 1 is intra-predicted. If sub-region 0 is intra-predicted, sub-region 1 generates a prediction signal based on motion compensation. In ECM, GPM can perform MMVD (Merge with Motion Vector Difference) on each sub-region and compensate for motion information based on TM (Template Matching). Only one GPM-MMVD and one GPM-TM can be applied to one CU. GPM-MMVD can be applied independently to each sub-domain.
[0142] TM calculates the template matching cost between the template of the predicted block and the template of the current block. Here, motion information can be generated from the initial prediction signal according to compensation. In the case of inter prediction, the predicted block is a reference block within the reference picture, and in the case of intra block copy (IBC), the predicted block can be a reference block within the restoration area of the current picture. TM can compensate for the motion information of the initial prediction signal so that the template matching cost is minimized.
[0143] Bidirectional GPM technology generates prediction signals for each sub-region based on bidirectional prediction when applying GPM to generate the final prediction signal of the current block. Bidirectional GPM technology can be applied to all blocks to which GPM is applied, except for CUs of 8×8, 8×16, and 16×8 sizes. The process of generating a GPM merge list from a general merge list is applied to blocks of the aforementioned small sizes. For blocks of the remaining sizes, bidirectional GPM technology can generate a general merge list and predict each sub-region based on merge index information extracted from the bitstream. Bidirectional GPM-MMVD and bidirectional GPM-TM are supported, and BDOF (Bi-directional Optical Flow)-based motion vector refinement technology can be applied to each sub-region. BDOF additionally compensates for the motion of predicted samples using bidirectional motion prediction based on the assumption that samples or objects constituting an image move at a constant speed and there is little change in sample values.
[0144] Bidirectional GPM technology processes bidirectional motion information across two sub-regions, which can result in inefficient memory usage. Below, we describe a method for performing bidirectional prediction based on efficient memory usage.
[0145] Hereinafter, the terms general merge candidate list, merge candidate list, and merge list are used interchangeably. Furthermore, the terms GPM merge candidate list and GPM merge list are used interchangeably.
[0146] The following embodiments are described with a focus on an image decoding device, but can be implemented identically or similarly in an image encoding device.
[0147] II. Embodiments according to the present disclosure
[0148] FIG. 9 is a block diagram illustrating in detail a portion of an image decoding device according to one embodiment of the present disclosure.
[0149] The video decoding device according to the present embodiment determines prediction and transformation units, and performs prediction and inverse transformation on the current block corresponding to the determined unit using the determined prediction technique and prediction mode, thereby finally generating a restoration block of the current block. The example illustrated in FIG. 9 may be performed by the inverse transformation unit (530), the prediction unit (540), and the adder (550) of the video decoding device. Meanwhile, the same operations as the example illustrated in FIG. 9 may be performed by the inverse transformation unit (165), the picture division unit (110), the prediction unit (120), and the adder (170) of the video encoding device. At this time, the video decoding device uses encoding information parsed from the bitstream, but the video encoding device may use encoding information set from a higher level in terms of minimizing rate distortion. Hereinafter, for convenience, the present embodiment will be described with a focus on the video decoding device.
[0150] As shown in the example of FIG. 5, the prediction unit (540) includes an intra prediction unit (542) and an inter prediction unit (544) depending on the prediction technology, but as shown in FIG. 9, the prediction unit (540) may include all or part of the prediction unit determination unit (902), the prediction technology determination unit (904), the prediction mode determination unit (906), and the prediction execution unit (908).
[0151] If the color format of the input video is a YUV format (such as YUV420, YUV411, YUV422, YUV444), the video decoding device can perform prediction and restoration of the chroma component after performing prediction and restoration of the luma component. That is, the luma component and the chroma component can be sequentially restored by the components illustrated in FIG. 8. Meanwhile, if the color format of the input video is RGB, the video encoding device can perform color format conversion from RGB to YUV and then encode the converted video. Here, in the case of the YUV format, the color format represents the correspondence between the pixels of the luma component and the pixels of the chroma component.
[0152] The prediction unit determination unit (902) determines a prediction unit (PU). The prediction technique determination unit (904) determines a prediction technique (e.g., intra prediction, inter prediction, IBC (Intra Block Copy) mode, palette mode, etc.) for the prediction unit. The prediction mode determination unit (906) determines a detailed prediction mode for the prediction technique. The prediction execution unit (908) generates a prediction block of the current block according to the determined prediction mode.
[0153] The inverse transform unit (530) includes a transform unit determination unit (910) and an inverse transform execution unit (912). The transform unit determination unit (910) determines a transform unit (TU) for the inverse quantization signals of the current block, and the inverse transform execution unit (912) inversely transforms the transform unit expressed as the inverse quantization signals to generate residual signals.
[0154] An adder (550) adds the prediction block and residual signals to generate a restoration block. The restoration block is stored in memory and can be used to predict other blocks.
[0155] The prediction unit determined by the prediction unit determination unit (902) may be the current block or one of the sub-blocks into which the current block is divided. At this time, the prediction unit of the chroma component may have a size corresponding to the prediction unit of the luma component according to the color format. Alternatively, after the prediction units of the luma component and the chroma component are determined separately, prediction may be performed on the prediction unit of the chroma component.
[0156] The prediction technology determination unit (904) determines the prediction technology for each prediction unit. As described above, the prediction technology may be one of inter-prediction, intra-prediction, IBC mode, and palette mode. In this case, the prediction technology for the chroma component can be determined in the same manner as the prediction technology for the corresponding luma component, without separate signaling or parsing of information.
[0157] For example, if the prediction technique of the current block is not intra prediction, the video decoding device parses 1-bit flag information. For example, if the parsed flag indicates Skip mode, the video decoding device determines the prediction mode of the current block as the merge mode of inter prediction or the IBC merge mode. In the case of Skip mode, the video decoding device can use the prediction signals as restored signals without performing the inverse transformation process (i.e., without parsing the residual signals).
[0158] On the other hand, if the parsed flag does not indicate a Skip mode for the current block, the prediction technique determination unit (904) can parse a series of 1-bit flags to determine the prediction technique of the current block as one of techniques such as inter prediction, intra prediction, IBC mode, palette mode, etc.
[0159] For example, if Skip is not applied to the current block and the prediction technique is determined to be Inter-Prediction or IBC mode, the video decoding device parses a 1-bit flag. Depending on the parsed flag, the prediction mode of the current block can be determined as either General Merge mode or Advanced Motion Vector Prediction (AMVP) mode.
[0160] The prediction mode determination unit (906) determines a detailed prediction mode for the prediction technology.
[0161] For example, if the prediction technology of the current block is inter prediction, the prediction mode determining unit (906) can determine the general merge mode or AMVP mode as the prediction mode of the current block. In the general merge mode or AMVP mode, the image decoding device generates prediction blocks according to one or more motion compensations based on parsed motion information, and weights and combines the generated multiple prediction blocks to generate final prediction signals of the current block.
[0162] As another example, if the prediction technology of the current block is inter prediction, the prediction mode determining unit (906) may determine Geometric Partitioning Mode (GPM) as the prediction mode of the current block. In GPM, the image decoding device divides the current block into two or more sub-regions according to geometric partitioning, generates prediction blocks according to one or more motion compensations based on the parsed motion information and prediction mode information of the current block, and weights and combines the generated multiple prediction blocks to generate final prediction signals of the current block.
[0163] The prediction execution unit (908) generates a prediction block of the current block according to the determined prediction technology and prediction mode.
[0164] As an example, the prediction performing unit (908) generates a prediction block of the current block according to the prediction mode, and the adder (550) adds the prediction block of the current block and residual signals to generate a restoration block.
[0165] Meanwhile, the entropy decoding unit (510) can restore the quantized second-order transform coefficients when the second-order transform is applied. On the other hand, when the second-order transform is not applied, the entropy decoding unit (510) can restore the quantized first-order transform coefficients. The inverse quantization unit (520) can apply inverse quantization to the restored transform coefficients based on the quantization parameter to generate inverse quantized transform coefficients.
[0166] The transformation unit determination unit (910) of the inverse transformation unit (530) determines a transformation unit (TU) for the inverse quantized transformation coefficients. At this time, if one TU is divided into multiple sub-blocks, one sub-block can be used as a TU.
[0167] The inverse transform unit (530) can determine a non-separable second-order inverse transform kernel and a separable first-order inverse transform kernel when a second-order transform is applied. On the other hand, when the second-order transform is not applied, the inverse transform unit (530) can determine a separable first-order inverse transform kernel or a non-separable first-order inverse transform kernel. The inverse transform unit (530) can perform an inverse transform on the inverse quantized transform coefficients using the inverse transform kernel. The inverse transform unit (530) can determine whether to perform a non-separable first-order inverse transform and a non-separable second-order inverse transform based on signaling / parsing, or can implicitly determine the same based on information such as the size of the current TU, aspect ratio, etc.
[0168] Hereinafter, with respect to bidirectional GPM, the operation of the prediction performing unit (908) within the image decoding device is described in detail.
[0169] As described above, the image decoding device including the prediction performing unit (908) generates a final prediction signal of the current block according to the determined prediction technique and prediction mode.
[0170] If the prediction technique of the current block is inter prediction, the current block is predicted according to GPM, and is divided into two sub-regions, the sub-regions can be classified into P0 sub-region and P1 sub-region centered on the division boundary. As another example, the current block can be divided into three or more sub-regions.
[0171] If the prediction technology of the current block is inter prediction and the current block is predicted according to GPM, the image decoding device can predict the current block according to the diagram in FIG. 10.
[0172] FIG. 10 is a flowchart illustrating prediction of a current block according to GPM according to one embodiment of the present disclosure.
[0173] The video decoding device checks the GPM flag (S1000).
[0174] If the GPM flag is false (No in S1000), the video decoding device may apply a prediction technique other than GPM. The application of a prediction technique other than GPM is outside the scope of this disclosure, and therefore, a detailed description thereof is omitted.
[0175] If the GPM flag is true (Yes in S1000), the image decoding device performs the following steps to predict the current block according to the geometric segmentation mode.
[0176] The video decoding device constructs a general merge candidate list (S1002). The video decoding device may construct the general merge candidate list using motion information of previously restored areas surrounding the current block and / or previously restored blocks in non-adjacent areas. As an example, the video decoding device may apply template matching (TM) to the general merge candidate list and reorder the candidate list based on the template costs of the merge candidates.
[0177] The video decoding device checks whether the current block is a small block (S1004).
[0178] A small block has a size of N×M (where N and M are integers greater than or equal to 1). N and M can be adaptively determined based on the size (height and width) of the current image. Alternatively, N and M can be fixed values determined by an agreement between the image encoding device and the image decoding device.
[0179] If the current block is a small block (Yes in S1004), the video decoding device constructs a GPM merge candidate list (S1006). The video decoding device can construct the GPM merge candidate list using a general merge candidate list. As described above, the GPM merge candidate list can be constructed based on the unidirectional motion information of each candidate in the general merge list.
[0180] If the current block is not a small block (No in S1004), the video decoding device may omit composing the GPM merge candidate list.
[0181] The video decoding device decodes merge indices that indicate motion information (or, interchangeably, "prediction information") of sub-regions. Meanwhile, the video encoding device may obtain merge indices from a higher level. Alternatively, the video encoding device may select optimal merge indices in terms of rate distortion based on a general merge candidate list or a GPM merge candidate list. The video encoding device may signal the merge indices to the video decoding device.
[0182] The video decoding device obtains motion information of each sub-region from a general merge candidate list or a GPM merge candidate list based on merge indices, and the video decoding device generates P0 and P1 initial prediction signals based on the obtained motion information (S1008).
[0183] The video decoding device performs GPM split mode reordering (S1010).
[0184] The video decoding device can perform GPM segmentation mode reordering and parse an index indicating an optimal segmentation mode. Using the parsed index, the video decoding device can determine the GPM segmentation mode of the current block from the reordered GPM segmentation modes.
[0185] In the case of bidirectional prediction, the video decoding device generates P0 and P1 final prediction signals based on P0 and P1 initial prediction signals (S1012).
[0186] The video decoding device performs a weighted sum of the final prediction signals P0 and P1 (S1014). The video decoding device can generate the final prediction signal of the current block by performing the weighted sum based on the GPM splitting mode and matrix blending.
[0187] Below, steps S1008 to S1012 are described in detail using the flowchart of Fig. 11.
[0188] FIG. 11 is a flowchart illustrating prediction of sub-regions according to one embodiment of the present disclosure.
[0189] In the example of FIG. 11, steps S1100 to S1110 and step S1130 are steps for generating P0 and P1 initial prediction signals, step S1112 is a step for GPM split mode reordering, and steps S1114 to S1122 are steps for generating P0 and P1 final prediction signals.
[0190] The example in Fig. 11 can be applied to both the P0 and P1 sub-regions. Hereinafter, the P0 and P1 sub-regions are expressed as the Pi (i=0, 1) sub-region, or sub-region.
[0191] The video decoding device checks whether the prediction technology of the sub-region is intra prediction (S1100).
[0192] If the prediction technology of the sub-region is intra prediction (Yes in S1100), the image decoding device applies intra prediction to the sub-region to generate a prediction signal (S1130).
[0193] If the prediction technology of the sub-region is not intra prediction (No in S1100), the video decoding device generates an initial prediction block of the sub-region based on the motion information of one of the candidates obtained from the merge list. The video decoding device can decode a TM flag indicating whether template matching is applied from the bitstream.
[0194] The video decoding device checks the TM flag (S1102).
[0195] If the TM flag is true (Yes in S1102), the video decoding device performs motion compensation based on template matching (S1104). The video decoding device can compensate for motion information of a sub-region based on template matching using the template of the current block.
[0196] The video decoding device can decode an MMVD flag indicating whether MMVD is applied from a bitstream.
[0197] If the TM flag is false (No in S1102), the video decoding device checks the MMVD flag (S1106).
[0198] If the MMVD flag is true (Yes in S1106), the video decoding device performs motion compensation based on the MMVD in the sub-region (S1108). The video decoding device can decode the MMVD information and compensate for the motion information in the sub-region based on the MMVD information.
[0199] As mentioned above, only one MMVD and TM can be applied to a sub-domain.
[0200] According to the steps described above, the image decoding device generates an initial prediction signal of the sub-region. The initial prediction signal may be an uncorrected initial prediction block, a corrected initial prediction block, or an intra-prediction signal.
[0201] The video decoding device checks whether an initial prediction signal of the P1 sub-region is generated (S1110).
[0202] If the initial prediction signal of the P1 sub-region is not generated (i.e., if the initial prediction signal of the P0 sub-region is generated, No in S1110), the steps for generating the initial prediction signal of the P1 sub-region are repeatedly performed. On the other hand, if the initial prediction signal of the P1 sub-region is generated (Yes in S1110), the image decoding device performs the following steps.
[0203] The video decoding device performs GPM split mode reordering (S1112).
[0204] As illustrated in FIG. 12, the video decoding device can generate a template for each initial prediction block within a reference picture using the initial prediction signal of each sub-region. The video decoding device can calculate the cost between the weighted template and the template of the current block based on each GPM division mode. The video decoding device can reorder the GPM division modes according to the calculated cost.
[0205] In the example of Fig. 12, Region A and Region L are divided according to the GPM division boundary. Region A (P0), Region A (P1), Region L (P0), and Region L (P1) represent templates of each initial prediction block. Region A (Cand.) and Region L (Cand.) represent weighted templates, and Region A and Region L represent templates of the current block. MV0 represents the motion vector of the P0 sub-region, and MV1 represents the motion vector of the P1 sub-region.
[0206] The video encoding device can signal an index indicating an optimal GPM partitioning mode to the video decoding device after performing the same reordering process, and the video decoding device can parse the index and determine the GPM partitioning mode of the current block from the reordered GPM partitioning modes based on the parsed index.
[0207] The video decoding device checks the DMVR (Decoder-side MV Refinement) condition (S1114). For example, the DMVR condition may indicate that bidirectional prediction is performed on the initial prediction signal of each sub-region, one reference picture has a POC (Picture Order Count) value greater than that of the current picture, the remaining reference pictures have POC values smaller than that of the current picture, and the distance between the current picture and each reference picture is the same.
[0208] If the DMVR condition is satisfied (Yes in S1114), the video decoding device performs sub-block-based motion compensation on the initial prediction signal in units of p×q blocks (p and q are integers greater than or equal to 1) (S1116). For example, the video decoding device can perform sub-block-based BDOF as the sub-block-based motion compensation.
[0209] If the DMVR condition is not satisfied (No in S1114), the video decoding device checks whether another sub-region is an intra-predicted region (S1118). For example, the video decoding device can check whether another sub-region (e.g., P0 sub-region) is an intra-predicted region when generating the final prediction signal of the current sub-region (e.g., P1 sub-region).
[0210] For the initial prediction signal of the current sub-region or the signal on which sub-block-based motion compensation has been performed, if another sub-region is an intra-predicted region (Yes in S1118), the image decoding device performs sub-block-based OBMC (Overlapped Block Motion Compensation) on the current sub-region (S1120). That is, the image decoding device can divide the current block into sub-blocks and compensate the prediction signal of the current sub-region by using the motion of the surrounding sub-blocks in units of sub-blocks.
[0211] The video decoding device checks whether the final prediction signal of the P1 sub-region is generated (S1122).
[0212] If the final prediction signal of the P1 sub-region is not generated (i.e., if the final prediction signal of the P0 sub-region is generated, No in S1122), the steps for generating the final prediction signal of the P1 sub-region are repeatedly performed. On the other hand, if the final prediction signal of the P1 sub-region is generated (Yes in S1122), the image decoding device finishes generating the final prediction signal.
[0213] Below, the weighted sum process for generating the final prediction signal of the current block is described in detail.
[0214] The video decoding device can explicitly or implicitly determine the final blending region on which blending will be performed during the weighted sum process. The video decoding device can calculate the blending matrix to be used in the weighted sum process based on the determined final blending region.
[0215] As an example, depending on the size of the current luma block and the derived geometric partitioning mode, the blending region can be determined as a fixed region.
[0216] As another example, the blending region can be explicitly determined by signaling / parsing index information that can determine the blending region.
[0217] As another example, the blending region can be implicitly determined based on the blending region determination process. For example, after generating the initial or final prediction signal for each sub-region, the blending region can be implicitly determined using the initial or final prediction values around the geometric segmentation boundary.
[0218] The video decoding device determines the blending matrix W based on the final blending region, which is implicitly or explicitly determined. B f can be calculated. The value W of each coefficient in the blending matrix B f (ij) is (0, 2) for w (w is an integer greater than or equal to 0). w) can be an integer value existing in the range. In this case, w can be determined according to the size and / or geometric division mode of the current block.
[0219] Blending matrix W B f And using the final prediction signals P0 and P1 of the two sub-regions, the image decoding device calculates the final prediction signal P of the current block according to the weighted sum as in mathematical expression 1. G You can get it.
[0220]
[0221] Meanwhile, when predicting the current block in geometric segmentation mode, the image decoding device can change the prediction information of each sub-region to generate an initial prediction signal using limited memory based on the prediction information of each sub-region. For example, if the initial prediction information of two sub-regions each indicates bidirectional prediction, or if one sub-region performs bidirectional prediction and the other sub-region performs unidirectional prediction, the image decoding device can change the prediction information.
[0222] Hereinafter, when prediction of both sub-regions is performed by bidirectional prediction, information of reference pictures (e.g., reference picture index) in the motion information of each sub-region is defined as follows. That is, for the P0 sub-region, information of reference pictures is defined as P0_LA_idx and P0_LB_idx, and for the P1 sub-region, information of reference pictures is defined as P1_LA_idx and P1_LB_idx.
[0223] In order to efficiently use memory in bidirectional GPM technology, the video decoding device can change prediction information in the LA direction as shown in Fig. 13 and prediction information in the LB direction as shown in Fig. 14. The LA direction and the LB direction represent two directions used for bidirectional prediction.
[0224] Below, a method for changing the prediction information in the direction of LA using the city of Fig. 13 is described.
[0225] FIG. 13 is a flowchart illustrating a change in prediction information in the LA direction according to one embodiment of the present disclosure.
[0226] The video decoding device checks whether P0_LA_idx and P1_LA_idx are identical (S1300). The identicalness of P0_LA_idx and P1_LA_idx indicates that the reference picture of the LA direction prediction signal of the P0 sub-region and the reference picture of the LA direction prediction signal of the P1 sub-region are identical.
[0227] If P0_LA_idx and P1_LA_idx are the same (Yes in S1300), the video decoding device compares the value of LA_W×LA_H with the threshold value T (S1302).
[0228] LA_W and LA_H can be calculated as in mathematical expression 2.
[0229]
[0230] The size of the current block is W×H, and T may be a value determined according to the size of the current block. As shown in Fig. 15, the threshold value T may be pW×qH (p and q are integers greater than or equal to 1), and may be a predefined value according to an agreement between the image encoding device and the image decoding device.
[0231] LA_W is defined as the sum of W and the difference (a positive integer greater than or equal to 0) between the horizontal components of the motion vectors in the LA direction for the P0, P1 prediction signals. LA_H is defined as the sum of H and the difference (a positive integer greater than or equal to 0) between the vertical components of the motion vectors in the LA direction. Therefore, LA_W×LA_H is a rectangle that includes all of the P0, P1 prediction signals, has a minimum area, and overlaps with the P0, P1 prediction signals at the vertices on the diagonal. LB_W×LB_H can be similarly defined according to Equation 2 and FIG. 15. Hereinafter, LA_W×LA_H is referred to as the combined prediction signal in the LA direction, and LB_W×LB_H is referred to as the combined prediction signal in the LB direction.
[0232] If LA_W×LA_H is less than the threshold T, the same information as the initially parsed value can be used as movement information in the LA direction.
[0233] If P0_LA_idx and P1_LA_idx are the same and LA_W×LA_H < T (Yes in S1302), the image decoding device starts processing information in the LA direction of the P0 prediction signal.
[0234] The video decoding device checks the TM flag for the P0 prediction signal (S1304).
[0235] If the TM flag for the P0 prediction signal is false (No in S1304), the video decoding device checks the MMVD flag for the P0 prediction signal (S1306).
[0236] As mentioned above, TM and MMVD can be performed as one process.
[0237] If the TM flag for the P0 prediction signal is true (Yes in S1304), the image decoding device limits the TM search range for the P0 prediction signal so that the prediction signal after the final motion compensation is performed is within pW×qH shown in FIG. 15 (S1308).
[0238] If the MMVD flag for the P0 prediction signal is true (Yes in S1306), the image decoding device configures new candidates among the distance candidates and direction candidates of the MVD (Motion Vector Difference) and limits the MMVD range for the P0 prediction signal so that the prediction signal after the final motion compensation is within pW×qH shown in FIG. 15 (S1310).
[0239] For example, a video encoding device can propose an MMVD range for a P0 prediction signal according to the same process and signal information about one of the new candidates to a video decoding device. The video decoding device can then calculate an MVD for the P0 prediction signal from the new candidates based on the parsed information.
[0240] If the TM range or MMVD range for the LA direction of the P0 prediction signal is limited, or the MMVD flag for the P0 prediction signal is false (No in S1306), the image decoding device finishes processing the LA direction information of the P0 prediction signal. Thereafter, the same process may be performed for the LA direction information of the P1 prediction signal.
[0241] The video encoding device checks the TM flag for the P1 prediction signal (S1312)
[0242] If the TM flag for the P1 prediction signal is false (No in S1312), the video encoding device checks the MMVD flag for the P1 prediction signal (S1314).
[0243] As mentioned above, TM and MMVD can be performed as one process.
[0244] If the TM flag for the P1 prediction signal is true (Yes in S1312), the image decoding device limits the TM search range for the P1 prediction signal so that the prediction signal after the final motion compensation is performed is within pW×qH shown in FIG. 15 (S1316).
[0245] If the MMVD flag for the P1 prediction signal is true (Yes in S1314), the image decoding device configures new candidates among the distance candidates and direction candidates of the MVD (Motion Vector Difference) and limits the MMVD range for the P1 prediction signal so that the prediction signal after the final motion compensation is within pW×qH shown in FIG. 15 (S1318).
[0246] For example, a video encoding device can propose an MMVD range for a P1 prediction signal using the same process and signal information about one of the new candidates to a video decoding device. The video decoding device can then calculate an MVD for the P1 prediction signal from the new candidates based on the parsed information.
[0247] If the TM range or MMVD range for the LB direction of the P1 prediction signal is limited, or the MMVD flag for the P1 prediction signal is false (No in S1314), the image decoding device finishes processing the LA direction information of the P1 prediction signal.
[0248] P0_LA_idx and P1_LA_idx are different (No. of S1300), or P0_LA_idx and P1_LA_idx are the same and LA_W×LA_H <T를 만족하지 않는 경우(S1302의 No), 영상 복호화 장치는 P0 및 P1 서브 영역의 예측 정보를 변경한다.
[0249] The video decoding device can use only the LA-direction prediction information as the prediction information for the P0 sub-region, as follows. Alternatively, the video decoding device can use only the LB-direction prediction information. In other words, the P0 sub-region uses unidirectional prediction information.
[0250] As shown in Fig. 16, the video decoding device defines an 'L'-shaped template of the P0 prediction signal in the LA and LB directions and an 'L'-shaped template of the current block, and calculates the cost by comparing the template of the current block with the template of each prediction signal. The video decoding device can use only the prediction information of the direction corresponding to the template with the lower cost as the prediction information of the P0 sub-region.
[0251] The template shape of the P0 prediction signal may vary depending on the angle of the segmentation boundary. For example, the P0 template shape according to the segmentation boundary may be defined according to an agreement between the video encoding device and the video decoding device. The template may not be an 'L'-shaped template, but may be the upper template and / or the left template as shown in FIG. 17. The cost may be calculated using only the templates illustrated in FIG. 17 and the upper and / or left templates of the current block. A normalized cost may be used depending on the template area. The cost may be one of values such as MSE (Mean Squared Error), SAD (Sum of Absolute Differences), etc.
[0252] The video decoding device can use only LA-direction prediction information as prediction information for the P1 sub-region, as follows. Alternatively, the video decoding device can use only LB-direction prediction information. In other words, the P1 sub-region uses unidirectional prediction information.
[0253] The video decoding device defines templates for the P1 prediction signals in the LA and LB directions and templates for the current block, and compares the templates for the current block with the templates for each prediction signal to calculate the cost. The video decoding device can only use prediction information in the direction corresponding to the template with the lowest cost as prediction information for the P1 sub-region.
[0254] The shape of the template of the P1 prediction signal and the template of the current block may be an 'L' shape, as shown in Fig. 16. Alternatively, the shape of the template of the P1 prediction signal may vary depending on the angle of the segmentation boundary. For example, the shape of the P0 template according to the segmentation boundary may be defined according to an agreement between the video encoding device and the video decoding device. The template may not be an 'L'-shaped template, but may be the upper template and / or the left template, as shown in Fig. 17.
[0255] Additionally, the video decoding device may have different P0_LB_idx and P1_LB_idx, or the same P0_LB_idx and P1_LB_idx may have LB_W×LB_H. <T를 만족하지 않는지를 확인한다(S1330). 여기서, LB_W와 LB_H는 수학식 2와 같이 산정될 수 있다.
[0256] P0_LB_idx and P1_LB_idx are different, or P0_LB_idx and P1_LB_idx are the same and LB_W×LB_H <T를 만족하지 않는 경우(S1330의 Yes), 영상 복호화 장치는 P0 및 P1 서브 영역의 예측 정보를 변경한다(S1332).
[0257] As described above, the video decoding device can use only the prediction information in the LA direction as the prediction information of the P0 sub-region. Alternatively, the video decoding device can use only the prediction information in the LB direction. That is, the P0 sub-region uses unidirectional prediction information. As described above, the video decoding device can use only the prediction information in the LA direction as the prediction information of the P1 sub-region. Alternatively, the video decoding device can use only the prediction information in the LB direction. That is, the P1 sub-region uses unidirectional prediction information.
[0258] On the other hand, if P0_LB_idx and P1_LB_idx are the same, then LB_W×LB_H <T를 만족하는 경우(S1330의 No), 영상 복호화 장치는 P0 또는 P1 서브 영역의 예측 정보를 변경한다(S1334).
[0259] The video decoding device defines templates of P0 and P1 prediction signals in the LA direction and templates of the current block, as shown in FIG. 16 or FIG. 17, and compares the template of the current block with the template of each prediction signal to calculate the cost. The video decoding device can improve the prediction performance by additionally using prediction information in the LA direction for prediction signals with high costs. For example, if the cost of the P0 sub-region in the LB direction is high, the P0 sub-region can additionally use prediction information in the LA direction. That is, the video decoding device can use bidirectional motion information for the P0 sub-region and unidirectional motion information for the P1 sub-region.
[0260] Below, a method for changing prediction information in the LB direction using the city of Fig. 14 is described.
[0261] FIG. 14 is a flowchart illustrating a change in prediction information in the LB direction according to one embodiment of the present disclosure.
[0262] The video decoding device checks whether P0_LB_idx and P1_LB_idx are identical (S1400). The identicalness of P0_LB_idx and P1_LB_idx indicates that the reference picture of the LB direction prediction signal of the P0 sub-region and the reference picture of the LB direction prediction signal of the P1 sub-region are identical.
[0263] If P0_LB_idx and P1_LB_idx are the same (Yes in S1400), the video decoding device compares the value of LB_W × LB_H with the threshold value T (S1402). Here, LB_W and LB_H can be calculated as in mathematical expression 2.
[0264] If LB_W×LB_H is less than the threshold T, the same information as the initially parsed value can be used as the movement information in the LA direction.
[0265] If P0_LB_idx and P1_LB_idx are the same and LB_W×LB_H < T (Yes in S1302), the image decoding device can process LB direction information of the P0, P1 prediction signals according to steps S1304 to S1318 of FIG. 13.
[0266] P0_LB_idx and P1_LB_idx are different (No. of S1400), or P0_LB_idx and P1_LB_idx are the same and LB_W×LB_H <T를 만족하지 않는 경우(S1402의 No), 영상 복호화 장치는 P0 및 P1 서브 영역의 예측 정보를 변경한다.
[0267] As described above, the video decoding device can use only the prediction information in the LA direction as the prediction information of the P0 sub-region. Alternatively, the video decoding device can use only the prediction information in the LB direction. That is, the P0 sub-region uses unidirectional prediction information. As described above, the video decoding device can use only the prediction information in the LA direction as the prediction information of the P1 sub-region. Alternatively, the video decoding device can use only the prediction information in the LB direction. That is, the P1 sub-region uses unidirectional prediction information.
[0268] Additionally, the video decoding device may have different P0_LA_idx and P1_LA_idx, or the same P0_LA_idx and P1_LA_idx may have LA_W×LA_H. <T를 만족하지 않는지를 확인한다(S1430). 여기서, LA_W와 LA_H는 수학식 2와 같이 산정될 수 있다.
[0269] P0_LA_idx and P1_LA_idx are different, or P0_LA_idx and P1_LA_idx are the same and LA_W×LA_H <T를 만족하지 않는 경우(S1430의 Yes), 영상 복호화 장치는 P0 및 P1 서브 영역의 예측 정보를 변경한다(S1432).
[0270] As described above, the video decoding device can use only the prediction information in the LA direction as the prediction information of the P0 sub-region. Alternatively, the video decoding device can use only the prediction information in the LB direction. That is, the P0 sub-region uses unidirectional prediction information. As described above, the video decoding device can use only the prediction information in the LA direction as the prediction information of the P1 sub-region. Alternatively, the video decoding device can use only the prediction information in the LB direction. That is, the P1 sub-region uses unidirectional prediction information.
[0271] P0_LA_idx and P1_LA_idx are the same and LA_W×LA_H <T를 만족하는 경우(S1430의 No), 영상 복호화 장치는 P0 또는 P1 서브 영역의 예측 정보를 변경한다(S1434).
[0272] The video decoding device defines templates of P0 and P1 prediction signals in the LA direction and templates of the current block, as shown in FIG. 16 or FIG. 17, and compares the template of the current block with the template of each prediction signal to calculate the cost. The video decoding device can improve the prediction performance by additionally using prediction information in the LB direction for prediction signals with high costs. For example, if the cost of the P0 sub-region in the LA direction is high, the P0 sub-region can additionally use prediction information in the LB direction. That is, the video decoding device can use bidirectional motion information for the P0 sub-region and unidirectional motion information for the P1 sub-region.
[0273] Hereinafter, using the illustrations of FIGS. 18 and 19, a method for changing prediction information in order to efficiently use memory in a bidirectional GPM technique is described when, among the P0 and P1 sub-regions, the prediction signal of one sub-region is a bidirectional prediction signal and the prediction signal of the remaining sub-regions is a unidirectional prediction signal.
[0274] Hereinafter, using the city of Fig. 18, a method for changing prediction information in the LA direction when the prediction signals of the P0 sub-region and the P1 sub-region exist in the LA direction is described. As described above, one of the P0 sub-region and the P1 sub-region is a unidirectional prediction signal in the LA direction.
[0275] FIG. 18 is a flowchart illustrating a change in prediction information in the LA direction according to another embodiment of the present disclosure.
[0276] The video decoding device checks whether P0_LA_idx and P1_LA_idx are identical (S1800). The identicalness of P0_LA_idx and P1_LA_idx indicates that the reference picture of the LA direction prediction signal of the P0 sub-region and the reference picture of the LA direction prediction signal of the P1 sub-region are identical.
[0277] If P0_LA_idx and P1_LA_idx are the same (Yes in S1800), the video decoding device compares the value of LA_W × LA_H with the threshold value T (S1802). Here, LA_W and LA_H can be calculated as in mathematical expression 2.
[0278] If LA_W×LA_H is less than the threshold T, the same information as the initially parsed value can be used as movement information in the LA direction.
[0279] If P0_LA_idx and P1_LA_idx are the same and LA_W×LA_H < T (Yes in S1802), the image decoding device can process LA direction information of the P0, P1 prediction signals according to steps S1304 to S1318 of FIG. 13.
[0280] P0_LA_idx and P1_LA_idx are different (No. of S1800), or P0_LA_idx and P1_LA_idx are the same and LA_W×LA_H <T를 만족하지 않는 경우(S1802의 No), 영상 복호화 장치는 P0 및 P1 서브 영역의 예측 정보를 변경한다(S1830).
[0281] As described above, the video decoding device can use only the prediction information in the LA direction as the prediction information of the P0 sub-region. Alternatively, the video decoding device can use only the prediction information in the LB direction. As described above, the video decoding device can use only the prediction information in the LA direction as the prediction information of the P1 sub-region. Alternatively, the video decoding device can use only the prediction information in the LB direction.
[0282] Hereinafter, using the illustration of Fig. 19, a method for changing prediction information in the LB direction when the prediction signals of the P0 sub-region and the P1 sub-region exist in the LB direction is described. As described above, one of the signals in the P0 sub-region and the P1 sub-region is a LB direction, i.e., a unidirectional prediction signal.
[0283] FIG. 19 is a flowchart illustrating a change in prediction information in the LB direction according to another embodiment of the present disclosure.
[0284] The video decoding device checks whether P0_LB_idx and P1_LB_idx are identical (S1900). The identicalness of P0_LB_idx and P1_LB_idx indicates that the reference picture of the LB direction prediction signal of the P0 sub-region and the reference picture of the LB direction prediction signal of the P1 sub-region are identical.
[0285] If P0_LB_idx and P1_LB_idx are the same (Yes in S1900), the video decoding device compares the value of LB_W × LB_H with the threshold value T (S1802). Here, LB_W and LB_H can be calculated as in mathematical expression 2.
[0286] If LB_W×LB_H is less than the threshold T, the same information as the initially parsed value can be used as the movement information in the LB direction.
[0287] If P0_LB_idx and P1_LB_idx are the same and LB_W×LB_H < T (Yes in S1902), the image decoding device can process LB direction information of the P0, P1 prediction signals according to steps S1304 to S1318 of FIG. 13.
[0288] P0_LB_idx and P1_LB_idx are different (No. of S1900), or P0_LB_idx and P1_LB_idx are the same and LB_W×LB_H <T를 만족하지 않는 경우(S1902의 No), 영상 복호화 장치는 P0 및 P1 서브 영역의 예측 정보를 변경한다(S1930).
[0289] As described above, the video decoding device can use only the prediction information in the LA direction as the prediction information of the P0 sub-region. Alternatively, the video decoding device can use only the prediction information in the LB direction. As described above, the video decoding device can use only the prediction information in the LA direction as the prediction information of the P1 sub-region. Alternatively, the video decoding device can use only the prediction information in the LB direction.
[0290] As an example, in the example of Fig. 10, when constructing a GPM merge candidate list regardless of the block size, each candidate in the GPM merge candidate list can only have unidirectional prediction information (e.g., one reference picture list, one reference picture index, one motion vector). When performing bidirectional prediction for each sub-region, a motion vector difference can be additionally transmitted for the unidirectional prediction information.
[0291] Assume that an L0-direction motion vector of a P0 sub-region exists and an L1-direction motion vector does not exist, and that an L0-direction motion vector of a P1 sub-region does not exist and an L1-direction motion vector exists. As an example, for bidirectional prediction of a P1 sub-region, an image decoding device generates an L0-direction prediction signal of a P0 sub-region based on an L0-direction motion vector. The image decoding device can generate an L0-direction motion vector of a P1 sub-region by adding a parsed motion vector difference value to an L0-direction motion vector. The image decoding device can generate an L0-direction prediction signal of a P1 sub-region based on the generated L0-direction motion vector. As another example, for bidirectional prediction of a P0 sub-region, an image decoding device generates an L1-direction prediction signal of a P1 sub-region based on an L1-direction motion vector. The video decoding device can generate an L1-direction motion vector of the P0 sub-region by adding the parsed motion vector difference value to the L1-direction motion vector, and the video decoding device can generate an L1-direction prediction signal of the P0 sub-region based on the generated L1-direction motion vector.
[0292] As described above, when bidirectional prediction is applied to both the P0 sub-region and the P1 sub-region, two motion vector differences can be transmitted for the two unidirectional prediction information. In this case, the two motion vector differences can be transmitted separately. Alternatively, only one motion vector difference can be transmitted, and another motion vector difference can be derived by changing the sign of the transmitted motion vector difference.
[0293] Although the flowchart / timing diagram of this specification describes each process as being executed sequentially, this is merely an illustrative description of the technical idea of one embodiment of the present disclosure. In other words, a person of ordinary skill in the art to which one embodiment of the present disclosure belongs may modify and apply various modifications and variations by changing the order described in the flowchart / timing diagram without departing from the essential characteristics of one embodiment of the present disclosure, or by executing one or more of the processes in parallel. Therefore, the flowchart / timing diagram is not limited to a chronological order.
[0294] It should be understood that the exemplary embodiments described above can be implemented in many different ways. The functions or methods described in one or more examples can be implemented in hardware, software, firmware, or any combination thereof. It should be understood that the functional components described herein are labeled as "units" to further emphasize their implementation independence.
[0295] Meanwhile, the various functions or methods described in this embodiment may also be implemented as instructions stored on a non-transitory storage medium that can be read and executed by one or more processors. Non-transitory storage media include, for example, all types of storage devices that store data in a form readable by a computer system. For example, non-transitory storage media include storage media such as erasable programmable read-only memory (EPROM), flash drives, optical drives, magnetic hard drives, and solid-state drives (SSDs).
[0296] The above description is merely an example of the technical idea of the present embodiment, and those skilled in the art will appreciate that various modifications and variations can be made without departing from the essential characteristics of the present embodiment. Therefore, the present embodiments are not intended to limit the technical idea of the present embodiment, but rather to explain it, and the scope of the technical idea of the present embodiment is not limited by these embodiments. The scope of protection of the present embodiment should be interpreted by the claims below, and all technical ideas within a scope equivalent thereto should be interpreted as being included in the scope of rights of the present embodiment.
[0297]
[0298]
[0299] CROSS-REFERENCE TO RELATED APPLICATION
[0300] This patent application claims priority to Korean patent application No. 10-2023-0172498, filed on December 1, 2023, and Korean patent application No. 10-2024-0149164, filed on October 29, 2024, the entire contents of which are incorporated herein by reference.
Claims
1. A method for restoring a current block performed by a video decoding device, A step of decoding merge indices indicating P0 initial motion information of a P0 sub-region of the current block and P1 initial motion information of a P1 sub-region, wherein the P0 sub-region and the P1 sub-region are generated by dividing the current block into two according to a geometric division mode; A step of constructing a merge candidate list of the current block; A step of obtaining the P0 initial motion information and the P1 initial motion information from the merge candidate list based on the merge indices; A step of changing the P0 initial motion information and / or the P1 initial motion information, when bidirectional prediction is applied to both the P0 sub-region and the P1 sub-zero; A step of generating a P0 initial prediction signal of the P0 sub-region based on the P0 initial motion information or the changed P0 initial motion information, and generating a P1 initial prediction signal of the P1 sub-region based on the P1 initial motion information or the changed P1 initial motion information; A step of generating a P0 final prediction signal and a P1 final prediction signal based on the P0 initial prediction signal and the P1 initial prediction signal; and A step of generating a final prediction signal of the current block by weighting and adding the P0 final prediction signal and the P1 final prediction signal based on blending according to the geometric division mode. A method comprising:
2. In paragraph 1, A step of decoding an index indicating the P0 initial motion information and an index indicating the P1 initial motion information in the above merge candidate list. A method further comprising:
3. In paragraph 1, The step of generating the above P0 initial prediction signal is: A step of correcting the P0 initial movement information based on template matching using the template of the current block; or A step of correcting the P0 initial motion information based on the motion vector difference information in the merge mode. A method comprising:
4. In paragraph 1, The step of generating the above P0 final prediction signal is: A method for generating the P0 final prediction signal by compensating the P0 initial prediction signal based on sub-block-based motion compensation when the motion information compensation condition is satisfied.
5. In paragraph 3, The above motion information correction conditions are: A method including: (i) a case where bidirectional prediction is performed on an initial prediction signal of each sub-region, (ii) a case where one reference picture has a POC (Picture Order Count) value greater than that of the current picture and the remaining reference pictures have a POC value smaller than that of the current picture, and (iii) a case where the distance between the current picture and each reference picture is equal.
6. In paragraph 1, The step of changing the above P0 initial movement information is: A method using P0_LA_index and P0_LB_index indicating reference pictures in the LA direction and the LB direction as bidirectional motion information of the P0 sub-region, and using P1_LA_index and P1_LB_index indicating reference pictures in the LA direction and the LB direction as bidirectional motion information of the P1 sub-region.
7. In Article 6 The step of changing the above P0 initial movement information is: (i) if the P0_LA_index and the P1_LA_index are the same, and (ii) the combined prediction signal including the P0 initial prediction signal and the P1 initial prediction signal in the LA direction is smaller than a preset threshold range, A step of limiting the search range for template matching so that the P0 prediction signal in the LA direction exists within the preset threshold range; or A step of limiting the range of motion vector differential information in merge mode so that the P0 prediction signal in the LA direction exists within the preset threshold range. A method comprising:
8. In Article 6 The step of changing the above P0 initial movement information is: (i) if the P0_LA_index and the P1_LA_index are different, or (ii) if the P0_LA_index and the P1_LA_index are the same, but the combined prediction signal including the P0 initial prediction signal and the P1 initial prediction signal in the LA direction is greater than a preset threshold range, A method for using only initial motion information in a direction corresponding to a smaller cost as motion information of the P0 sub-region based on a cost between templates of P0 prediction signals in the LA direction and the LB direction and the template of the current block.
9. In Article 8 The step of changing the above P0 initial movement information is: (i) the P0_LA_index and the P1_LA_index are different, or (ii) the P0_LA_index and the P1_LA_index are the same, but the combined prediction signal including the P0 initial prediction signal and the P1 initial prediction signal in the LA direction is greater than or equal to a preset threshold range, (iii) if the P0_LB_index and the P1_LB_index are different, or (iv) if the P0_LB_index and the P1_LB_index are the same, but the combined prediction signal including the P0 initial prediction signal and the P1 initial prediction signal in the LB direction is greater than or equal to a preset threshold range, A method for using only initial motion information in a direction corresponding to a smaller cost as motion information of the P0 sub-region based on a cost between templates of P0 prediction signals in the LA direction and the LB direction and the template of the current block.
10. In Article 8 The step of changing the above P0 initial movement information is: (i) the P0_LA_index and the P1_LA_index are different, or (ii) the P0_LA_index and the P1_LA_index are the same, but the combined prediction signal including the P0 initial prediction signal and the P1 initial prediction signal in the LA direction is greater than or equal to a preset threshold range, (iii) if the P0_LB_index and the P1_LB_index are the same, and (iv) the combined prediction signal including the P0 initial prediction signal and the P1 initial prediction signal in the LB direction is smaller than a preset threshold range, A method of additionally using initial motion information in the LA direction for a sub-region corresponding to a large cost based on a cost between templates of the P0 prediction signal and the P1 prediction signal in the LA direction and the template of the current block.
11. In paragraph 1, A step of reordering geometric partitioning modes by applying template matching between templates of the P0 initial prediction signal and the P1 initial prediction signal, and a template of a current block, wherein the P0 initial prediction signal and the P1 initial prediction signal depend on each geometric partitioning mode; A step of decrypting an index indicating the optimal splitting mode; and A step of determining the geometric division mode of the current block from the geometric division modes reordered according to the above index. A method further comprising:
12. A method for encoding a current block performed by a video encoding device, A step of obtaining merge indices indicating P0 initial motion information of a P0 sub-region of the current block and P1 initial motion information of a P1 sub-region, wherein the P0 sub-region and the P1 sub-region are generated by dividing the current block into two according to a geometric division mode; A step of constructing a merge candidate list of the current block; A step of obtaining the P0 initial motion information and the P1 initial motion information from the merge candidate list based on the merge indices; A step of generating a P0 initial prediction signal of the P0 sub-region based on the P0 initial motion information or the changed P0 initial motion information, and generating a P1 initial prediction signal of the P1 sub-region based on the P1 initial prediction information or the changed P0 initial motion information; A step of generating a P0 final prediction signal and a P1 final prediction signal based on the P0 initial prediction signal and the P1 initial prediction signal; and A step of generating a final prediction signal of the current block by weighting and adding the P0 final prediction signal and the P1 final prediction signal based on blending according to the geometric division mode. A method comprising:
13. In paragraph 12, A method further comprising the step of encoding the above merge indices.
14. A method for providing video data to a video decoding device, A step of encoding the above video data into a bitstream; and A step of transmitting the above bitstream to the image decoding device Including, The step of encoding the above video data is: A step of obtaining merge indices indicating P0 initial motion information of a P0 sub-region of a current block and P1 initial motion information of a P1 sub-region, wherein the P0 sub-region and the P1 sub-region are generated by dividing the current block into two according to a geometric division mode; A step of constructing a merge candidate list of the current block; A step of obtaining the P0 initial motion information and the P1 initial motion information from the merge candidate list based on the merge indices; A step of generating a P0 initial prediction signal of the P0 sub-region based on the P0 initial motion information or the changed P0 initial motion information, and generating a P1 initial prediction signal of the P1 sub-region based on the P1 initial prediction information or the changed P0 initial motion information; A step of generating a P0 final prediction signal and a P1 final prediction signal based on the P0 initial prediction signal and the P1 initial prediction signal; and A step of generating a final prediction signal of the current block by weighting and adding the P0 final prediction signal and the P1 final prediction signal based on blending according to the geometric division mode. A method comprising:
Citation Information
Patent Citations
Cosmetic composition floating the ceramide capsules
KR1020250050587A
Prop assembly for article
KR102482813B1
A method of measuring bone mineral density of based on machine learning using radiographic image of the hip joint taken by X-ray
KR102578943B1