Video encoding method and apparatus using inter prediction based on template matching
By refining motion vector candidates in the template area and optimizing inter-frame prediction, the problem of insufficient encoding efficiency of high resolution and high frame rate videos is solved, and more efficient video encoding and quality improvement is achieved.
Patent Information
- Application Number
- CN202380082623.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-11-22
- Filing Date
- 2023-11-24
- Publication Date
- 2025-08-29
AI Technical Summary
When existing video encoding technology processes high-resolution and high frame rate video, the encoding efficiency is insufficient and cannot effectively improve the video quality.
By performing initial refinement, reordering and filtering of motion vector candidates on the template area, the inter prediction process is optimized and video encoding efficiency and quality is improved.
It improves video encoding efficiency, enhances video quality, and adapts to the encoding needs of high resolution and high frame rate videos.
Smart Images

Figure CN120569971A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to a video encoding method and apparatus using inter-frame prediction based on template matching. Background Art
[0002] The statements in this section merely provide background information related to the present disclosure and may not constitute prior art.
[0003] Since video data has a large amount of data compared to audio or still image data, the video data requires a large amount of hardware resources (including memory) to store or transmit the video data without a process for compression.
[0004] Therefore, encoders are typically used to compress and store or transmit video data. Decoders receive the compressed video data, decompress it, and play the decompressed video data. Video compression technologies include H.264 / Advanced Video Coding (AVC), High Efficiency Video Coding (HEVC), and Versatile Video Coding (VVC), which has improved coding efficiency by approximately 30% or more compared to HEVC.
[0005] However, as image size, resolution, and frame rate gradually increase, the amount of data to be encoded also increases. Thus, a new compression technology that provides higher encoding efficiency and improved image enhancement effects than existing compression technologies is needed.
[0006] Outside of VVC, adaptive reordering of mergecandidates (ARMC) reorders motion vector candidates in merge mode based on template matching. When ARMC is applied, the decoder divides the merge candidates into multiple subgroups and reorders the merge candidates in each subgroup according to the template matching cost. The template matching cost is calculated as the sum of absolute differences (SAD) between the samples in the template area of the current block and their corresponding reference samples. The template consists of samples in the neighboring reconstructed area of the current block. Outside of VVC, reordering is applied to regular merge mode, template matching merge mode, or affine merge mode. Figure 6 As described in , ARMC can be applied to the conventional merge mode based on the template area of the current block and the reference block. Figure 7 As shown in , ARMC can be applied to affine merge mode based on the template area of the current block and the template area of the sub-block of the reference block. In order to improve video coding efficiency and enhance video quality when predicting the current block, a method for effectively performing inter-frame prediction based on template matching is needed. Summary of the Invention
[0007] [Technical Issues]
[0008] The present invention seeks to provide a video encoding method and apparatus that compensates for motion of a current block during inter-frame prediction of the current block by performing initial refinement of motion vector candidates, reordering of motion vector candidates, and filtering on a template region.
[0009] [Technical solution]
[0010] At least one aspect of the present invention provides a method for reconstructing a current block by a video decoding device. The method includes decoding a merge index of the current block from a bitstream. The method also includes generating a motion vector candidate list for the current block. The motion vector candidate list includes a preset number of motion vector candidates, and the motion vector candidates are unidirectional motion vectors or bidirectional motion vectors. The method also includes performing initial motion refinement on the motion vector candidates. The method also includes reordering the motion vector candidates. The method also includes selecting motion information of the current block from the motion vector candidates by using the merge index. The method also includes generating a prediction block for the current block by using the selected motion information. When performing the initial motion refinement or reordering the motion vector candidates, the method includes applying filtering to block partition boundaries within a template area of the current block when using template matching.
[0011] Another aspect of the present invention provides a method for encoding a current block by a video encoding device. The method includes generating a motion vector candidate list for the current block. The motion vector candidate list includes a preset number of motion vector candidates, and the motion vector candidates are unidirectional motion vectors or bidirectional motion vectors. The method also includes performing initial motion refinement on the motion vector candidates. The method also includes reordering the motion vector candidates. The method also includes determining a merge index indicating one of the motion vector candidates. The method also includes selecting motion information of the current block from the motion vector candidates by using the merge index. The method also includes generating a prediction block for the current block by using the selected motion information. When performing the initial motion refinement or reordering the motion vector candidates, the method includes applying filtering to block partition boundaries within a template area of the current block when using template matching.
[0012] Another aspect of the present invention provides a computer-readable recording medium storing a bitstream generated by a video encoding method. The video encoding method includes generating a motion vector candidate list for a current block. The motion vector candidate list includes a preset number of motion vector candidates, and the motion vector candidates are unidirectional motion vectors or bidirectional motion vectors. The video encoding method also includes performing initial motion refinement on the motion vector candidates. The video encoding method also includes reordering the motion vector candidates. The video encoding method also includes determining a merge index indicating one of the motion vector candidates. The video encoding method also includes selecting motion information of the current block from the motion vector candidates by using the merge index. The video encoding method also includes generating a prediction block for the current block by using the selected motion information. When performing the initial motion refinement or reordering the motion vector candidates, the video encoding method includes applying filtering to block partition boundaries within a template area of the current block when using template matching.
[0013] [Beneficial Effects]
[0014] As described above, the present invention provides a video encoding method and apparatus that compensates for the motion of a current block during inter-frame prediction of the current block by performing initial refinement of motion vector candidates, reordering of motion vector candidates, and filtering on a template region. Thus, the video encoding method and apparatus improve video encoding efficiency and enhance video quality. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 is a block diagram of a video encoding device that may implement the techniques of this disclosure.
[0016] Figure 2 A method for partitioning blocks using a quadtree plus binary tree ternary tree (QTBTTT) structure is described.
[0017] Figure 3a and Figure 3b A plurality of intra prediction modes including a wide-angle intra prediction mode are shown.
[0018] Figure 4 Shows the neighboring blocks of the current block.
[0019] Figure 5 is a block diagram of a video decoding device that may implement the techniques of this disclosure.
[0020] Figure 6 is a diagram showing template areas of a current block and a reference block.
[0021] Figure 7 is a simplified diagram showing a template area of a current block and template areas of sub-blocks of a reference block.
[0022] Figure 8is a diagram illustrating a method of deriving constructed affine merge candidates for affine motion prediction.
[0023] Figure 9 is a detailed block diagram of a portion of a video decoding device according to at least one embodiment of the present disclosure.
[0024] Figure 10 is a diagram illustrating motion vector refinement using template matching, in accordance with at least one embodiment of the present disclosure.
[0025] Figure 11 is a diagram illustrating templates of neighboring sub-blocks according to at least one embodiment of the present disclosure.
[0026] Figure 12 is a diagram illustrating motion vector refinement using template matching according to another embodiment of the present disclosure.
[0027] Figure 13 is a diagram illustrating filtering of a template region according to at least one embodiment of the present disclosure.
[0028] Figure 14 is a simplified diagram illustrating filtering of a template region according to another embodiment of the present disclosure.
[0029] Figure 15 is a simplified diagram illustrating filtering of a template region according to another embodiment of the present disclosure.
[0030] Figure 16 The present invention is a flowchart of a method for encoding a current block by a video encoding device according to at least one embodiment of the present disclosure.
[0031] Figure 17 The present invention is a flowchart of a method for reconstructing a current block by a video decoding device according to at least one embodiment of the present disclosure. DETAILED DESCRIPTION
[0032] Hereinafter, some embodiments of the present disclosure are described in detail with reference to the accompanying drawings. In the following description, the same reference numerals represent the same elements, even though the elements are shown in different drawings. In addition, in the following description of some embodiments, for the purpose of clarity and brevity, detailed descriptions of related known components and functions may be omitted when it is considered that they would obscure the subject matter of the present disclosure.
[0033] Figure 1 is a block diagram of a video encoding device that can implement the technology of the present disclosure. Figure 1 A diagram illustrating a video encoding device and components of the device.
[0034] The encoding apparatus may include an image divider 110 , a predictor 120 , a subtractor 130 , a transformer 140 , a quantizer 145 , a rearrangement unit 150 , an entropy encoder 155 , an inverse quantizer 160 , an inverse transformer 165 , an adder 170 , a loop filter unit 180 , and a memory 190 .
[0035] Each component of the coding device can be implemented as hardware or software or as a combination of hardware and software. In addition, the function of each component can be implemented as software, and a microprocessor can also be implemented to execute the function of the software corresponding to each component.
[0036] A video consists of one or more sequences of multiple images. Each picture is divided into multiple regions, and encoding is performed on each region. For example, a picture is split into one or more tiles and / or slices. Here, one or more tiles can be defined as tile groups. Each tile and / or slice is split into one or more coding tree units (CTUs). In addition, each CTU is split into one or more coding units (CUs) through a tree structure. Information applied to each coding unit (CU) is encoded as the syntax of the CU, and information commonly applied to the CUs contained in a CTU is encoded as the syntax of the CTU. In addition, information commonly applied to all blocks in a slice is encoded as the syntax of the slice header, and information applied to all blocks constituting one or more pictures is encoded as a picture parameter set (PPS) or picture header. In addition, information commonly referenced by multiple pictures is encoded into a sequence parameter set (SPS). In addition, information commonly referenced by one or more SPSs is encoded into a video parameter set (VPS). Furthermore, information commonly applied to a tile or tile group can also be encoded as the syntax of the tile or tile group header. The syntax included in the SPS, PPS, slice header, tile, or tile group header may be referred to as a high-level syntax.
[0037] The image divider 110 determines the size of a coding tree unit (CTU). Information on the size of the CTU (CTU size) is encoded as a syntax of an SPS or PPS and delivered to a video decoding apparatus.
[0038] The image splitter 110 splits each picture constituting a video into a plurality of coding tree units (CTUs) of a predetermined size, and then recursively splits the CTUs using a tree structure. Leaf nodes in the tree structure become coding units (CUs), which are basic units of encoding.
[0039] The tree structure can be a quadtree (QT), wherein a higher node (or parent node) is divided into four lower nodes (or child nodes) of the same size. The tree structure can also be a binary tree (BT), wherein a higher node is divided into two lower nodes. The tree structure can also be a ternary tree (TT), wherein a higher node is divided into three lower nodes at a ratio of 1:2:1. The tree structure can also be a structure in which two or more of the QT structure, BT structure, and TT structure are mixed. For example, a quadtree plus binary tree (QTBT) structure can be used, or a quadtree plus binary tree ternary tree (QTBTTT) structure can be used. Here, a binary tree ternary tree (BTTT) is added to the tree structure to be referred to as a multi-type tree (MTT).
[0040] Figure 2 is a simplified diagram for describing a method of partitioning a block by using a QTBTTT structure.
[0041] like Figure 2 As shown, the CTU can first be split into a QT structure. The quadtree partitioning can be recursive until the size of the partitioned block reaches the minimum block size (MinQTSize) of the leaf node allowed in QT. The first flag (QT_split_flag) indicating whether each node of the QT structure is split into four nodes of the lower layer is encoded by the entropy encoder 155 and sent to the video decoding device. When the leaf node of QT is not larger than the maximum block size (MaxBTSize) of the root node allowed in BT, the leaf node can be further split into at least one of the BT structure or the TT structure. There may be multiple partitioning directions in the BT structure and / or the TT structure. For example, there may be two directions, namely, the direction in which the blocks of the corresponding node are partitioned horizontally and the direction in which the blocks of the corresponding node are partitioned vertically. As shown in FIG. Figure 2 As shown, when MTT splitting begins, a second flag (mtt_split_flag) indicating whether the node is split and a flag indicating the split direction (vertical or horizontal) and / or a flag indicating the split type (binary or ternary) if the node is split are encoded by the entropy encoder 155 and notified to the video decoding device using a signal.
[0042] Alternatively, before encoding the first flag (QT_split_flag) indicating whether each node is split into four nodes in the lower layer, a CU split flag (split_cu_flag) indicating whether the node is split may also be encoded. When the value of the CU split flag (split_cu_flag) indicates that each node is not split, the block of the corresponding node becomes a leaf node in the split tree structure and becomes a CU as a basic unit of encoding. When the value of the CU split flag (split_cu_flag) indicates that each node is split, the video encoding device first starts encoding the first flag using the above-mentioned scheme.
[0043] When QTBT is used as another embodiment of a tree structure, there may be two types, namely, a type in which the block of the corresponding node is horizontally split into two blocks of the same size (i.e., symmetrical horizontal splitting) and a type in which the block of the corresponding node is vertically split into two blocks of the same size (i.e., symmetrical vertical splitting). A split flag (split_flag) indicating whether each node of the BT structure is split into blocks of the lower layer and split type information indicating the split type are encoded by the entropy encoder 155 and delivered to the video decoding device. At the same time, a type in which the block of the corresponding node is split into two blocks that are asymmetric to each other may be additionally presented. The asymmetric form may include a form in which the block of the corresponding node is split into two rectangular blocks with a size ratio of 1:3, or may also include a form in which the block of the corresponding node is split in a diagonal direction.
[0044] A CU can have different sizes depending on the QTBT or QTBTTT split from a CTU. Hereinafter, the block corresponding to the CU to be encoded or decoded (i.e., a leaf node of a QTBTTT) is referred to as the "current block." Due to the QTBTTT split, the shape of the current block can be rectangular in addition to square.
[0045] The predictor 120 predicts the current block to generate a predicted block. The predictor 120 includes an intra predictor 122 and an inter predictor 124.
[0046] Typically, each of the current blocks in a picture can be predictively coded. Typically, prediction of the current block can be performed using intra-prediction techniques (using data from the picture that includes the current block) or inter-prediction techniques (using data from pictures coded before the picture that includes the current block). Inter-prediction includes both unidirectional prediction and bidirectional prediction.
[0047] The intra-frame predictor 122 predicts pixels in the current block by using pixels (reference pixels) located on neighbors of the current block in the current picture including the current block. There are multiple intra-frame prediction modes according to the prediction direction. For example, Figure 3a As shown in , a plurality of intra prediction modes may include 2 non-directional modes including a planar mode and a DC mode, and may include 65 directional modes. Neighboring pixels and an arithmetic expression to be used are defined differently according to each prediction mode.
[0048] In order to perform efficient directional prediction for a current block with a rectangular shape, it is possible to additionally use Figure 3b The dotted arrows in FIG. 5 show the directional modes (#67 to #80, intra prediction modes #-1 to #-14). The directional mode may be referred to as a "wide-angle intra prediction mode". Figure 3b, the arrows indicate the corresponding reference samples used for prediction and do not represent the prediction direction. The prediction direction is opposite to the direction indicated by the arrow. When the current block has a rectangular shape, the wide-angle intra prediction mode is a mode in which prediction is performed in a direction opposite to the specific direction mode without additional bit transmission. In this case, in the wide-angle intra prediction mode, some wide-angle intra prediction modes available for the current block can be determined by the ratio of the width and height of the current block having a rectangular shape. For example, when the current block has a rectangular shape with a height smaller than the width, a wide-angle intra prediction mode (intra prediction modes #67 to #80) with an angle smaller than 45 degrees is available. When the current block has a rectangular shape with a width greater than the height, a wide-angle intra prediction mode with an angle greater than -135 degrees can be used.
[0049] The intra-frame predictor 122 may determine an intra-frame prediction to be used for encoding the current block. In some embodiments, the intra-frame predictor 122 may encode the current block using multiple intra-frame prediction modes and may further select an appropriate intra-frame prediction mode to be used from a test mode. For example, the intra-frame predictor 122 may calculate a rate-distortion value by performing a rate-distortion analysis on multiple test intra-frame prediction modes and may further select an intra-frame prediction mode having the best rate-distortion characteristics from the test mode.
[0050] The intra-frame predictor 122 selects an intra-frame prediction mode from a plurality of intra-frame prediction modes and predicts the current block by using adjacent pixels (reference pixels) and an arithmetic expression determined according to the selected intra-frame prediction mode. Information about the selected intra-frame prediction mode is encoded by the entropy encoder 155 and delivered to the video decoding device.
[0051] The inter-frame predictor 124 generates a prediction block for the current block using motion compensation processing. The inter-frame predictor 124 searches for a block most similar to the current block in a reference picture that was encoded and decoded earlier than the current picture, and generates a prediction block for the current block using the searched block. Furthermore, a motion vector (MV) is generated, which corresponds to the displacement between the current block in the current picture and the prediction block in the reference picture. Typically, motion estimation is performed on the luma component, and the motion vector calculated based on the luma component is used for both the luma component and the chroma components. Motion information, including information about the reference picture and information about the motion vector used to predict the current block, is encoded by the entropy encoder 155 and delivered to the video decoding device.
[0052] The inter-frame predictor 124 may also perform interpolation on a reference picture or reference block to increase prediction accuracy. In other words, subsamples between two consecutive integer samples are interpolated by applying filter coefficients to a plurality of consecutive integer samples including the two integer samples. When searching for a block most similar to the current block in an interpolated reference picture, the motion vector may be represented with decimal precision rather than integer sample precision. The precision or resolution of the motion vector may be set differently for each target region to be encoded (e.g., a unit such as a slice, tile, CTU, or CU). When this adaptive motion vector resolution (AMVR) is applied, information regarding the motion vector resolution to be applied to each target region should be signaled for each target region. For example, when the target region is a CU, information regarding the motion vector resolution applied to each CU is signaled. The information regarding the motion vector resolution may be information indicating the precision of the motion vector difference, which will be described below.
[0053] Meanwhile, the inter-frame predictor 124 can perform inter-frame prediction using bidirectional prediction. Bidirectional prediction uses two reference pictures and two motion vectors representing the block positions most similar to the current block in each reference picture. The inter-frame predictor 124 selects a first reference picture and a second reference picture from reference picture list 0 (RefPicList0) and reference picture list 1 (RefPicList1), respectively. The inter-frame predictor 124 also searches for blocks most similar to the current block in each reference picture to generate a first reference block and a second reference block. Furthermore, a prediction block for the current block is generated by averaging or weighted averaging the first and second reference blocks. Motion information, including information about the two reference pictures used to predict the current block and information about the two motion vectors, is delivered to the entropy encoder 155. Reference picture list 0 may consist of pictures preceding the current picture in display order among previously reconstructed pictures, and reference picture list 1 may consist of pictures following the current picture in display order among previously reconstructed pictures. However, although not particularly limited thereto, a pre-reconstructed picture after the current picture in display order may be additionally included in reference picture list 0. Conversely, a pre-reconstructed picture before the current picture may also be additionally included in reference picture list 1.
[0054] To minimize the number of bits consumed for encoding motion information, various methods can be used.
[0055] For example, when the reference picture and motion vector of the current block are the same as those of a neighboring block, information that can identify the neighboring block is encoded to deliver the motion information of the current block to the video decoding device. This method is called merge mode.
[0056] In the merge mode, the inter predictor 124 selects a predetermined number of merge candidate blocks (hereinafter, referred to as “merge candidates”) from the neighboring blocks of the current block.
[0057] As a neighboring block for deriving merge candidates, as in Figure 4 As shown in FIG, all or some of the left block A0, lower left block A1, upper block B0, upper right block B1, and upper left block B2 adjacent to the current block in the current picture can be used. Furthermore, blocks located in reference pictures other than the current picture in which the current block is located (which may be the same as or different from the reference picture used to predict the current block) can also be used as merge candidates. For example, a block co-located with the current block or a block adjacent to the co-located block in the reference picture can also be used as a merge candidate. If the number of merge candidates selected by the method described above is less than the preset number, a zero vector is added to the merge candidates.
[0058] The inter-frame predictor 124 configures a merge list including a predetermined number of merge candidates using adjacent blocks. A merge candidate to be used as motion information for the current block is selected from the merge candidates included in the merge list, and merge index information for identifying the selected candidate is generated. The generated merge index information is encoded by the entropy encoder 155 and delivered to the video decoding device.
[0059] Merge skip mode is a special case of merge mode. After quantization, when all transform coefficients used for entropy coding are close to zero, only the neighbor block selection information is transmitted without the residual signal. By using merge skip mode, relatively high coding efficiency can be achieved for images with slight motion, still images, screen content images, etc.
[0060] Hereinafter, the merge mode and the merge skip mode are collectively referred to as the merge / skip mode.
[0061] Another method for encoding motion information is the Advanced Motion Vector Prediction (AMVP) mode.
[0062] In the AMVP mode, the inter-frame predictor 124 derives a motion vector predictor candidate for the motion vector of the current block by using the neighboring blocks of the current block. As the neighboring blocks for deriving the motion vector predictor candidate, the same as in Figure 4All or some of the left block A0, lower left block A1, upper block B0, upper right block B1, and upper left block B2 adjacent to the current block in the current picture shown in FIG. Furthermore, blocks located in a reference picture other than the current picture in which the current block is located (which may be the same as or different from the reference picture used to predict the current block) may also be used as neighboring blocks for deriving motion vector predictor candidates. For example, a block co-located with the current block or a block adjacent to the co-located block within the reference picture may be used. If the number of motion vector candidates selected by the above method is less than a preset number, a zero vector is added to the motion vector candidates.
[0063] The inter-frame predictor 124 derives motion vector predictor candidates by using the motion vectors of the adjacent blocks, and determines a motion vector predictor for the motion vector of the current block by using the motion vector predictor candidates. In addition, a motion vector difference is calculated by subtracting the motion vector predictor from the motion vector of the current block.
[0064] The motion vector predictor can be obtained by applying a predefined function (e.g., center value and average calculation) to the motion vector predictor candidate. In this case, the video decoding device also knows the predefined function. In addition, since the neighboring blocks used to derive the motion vector predictor candidate are blocks that have already been encoded and decoded, the video decoding device may also already know the motion vectors of the neighboring blocks. Therefore, the video encoding device does not need to encode information for identifying the motion vector predictor candidate. Therefore, in this case, information about the motion vector difference and information about the reference picture used to predict the current block are encoded.
[0065] Meanwhile, the motion vector predictor can also be determined by selecting any one of the motion vector predictor candidates. In this case, information for identifying the selected motion vector predictor candidate is jointly encoded with information about the motion vector difference and information about the reference picture used to predict the current block.
[0066] The subtractor 130 generates a residual block by subtracting the prediction block generated by the intra predictor 122 or the inter predictor 124 from the current block.
[0067] The transformer 140 converts the residual signal in the residual block having pixel values in the spatial domain into transform coefficients in the frequency domain. The transformer 140 may transform the residual signal in the residual block by using the total size of the residual block as a transform unit, or may further partition the residual block into multiple sub-blocks and perform the transform using the sub-blocks as transform units. Alternatively, the residual block is partitioned into two sub-blocks, a transform region and a non-transform region, so that the residual signal is transformed using only the transform region sub-block as a transform unit. Here, the transform region sub-block may be one of two rectangular blocks having a size ratio of 1:1 based on the horizontal axis (or vertical axis). In this case, a flag (cu_sbt_flag) indicates that only the sub-block is transformed, and directional (vertical / horizontal) information (cu_sbt_horizontal_flag) and / or position information (cu_sbt_pos_flag) are encoded by the entropy encoder 155 and signaled to the video decoding device. Furthermore, the size of the transform region sub-block may have a size ratio of 1:3 based on the horizontal axis (or vertical axis). In this case, a flag (cu_sbt_quad_flag) for splitting the corresponding split is additionally encoded by the entropy encoder 155 and signaled to the video decoding apparatus.
[0068] At the same time, the transformer 140 may perform transforms on the residual block separately in the horizontal and vertical directions. For the transforms, different types of transform functions or transform matrices may be used. For example, a pair of transform functions for horizontal and vertical transforms may be defined as a multiple transform set (MTS). The transformer 140 may select a transform function pair with the highest transform efficiency in the MTS and transform the residual block in each of the horizontal and vertical directions. Information (mts_idx) of the transform function pair in the MTS is encoded by the entropy encoder 155 and signaled to the video decoding device.
[0069] The quantizer 145 quantizes the transform coefficients output from the transformer 140 using a quantization parameter and outputs the quantized transform coefficients to the entropy encoder 155. The quantizer 145 can also immediately quantize the relevant residual block without requiring a transform for any block or frame. The quantizer 145 can also apply different quantization coefficients (scaling values) depending on the position of the transform coefficient in the transform block. The quantization matrix applied to the quantized transform coefficients arranged in two dimensions can be encoded and signaled to the video decoding device.
[0070] The rearrangement unit 150 may perform rearrangement of coefficient values with respect to the quantized residual value.
[0071] The rearrangement unit 150 can convert the 2D coefficient array into a 1D coefficient sequence by using coefficient scanning. For example, the rearrangement unit 150 can output a 1D coefficient sequence by scanning the DC coefficient into high-frequency domain coefficients using zigzag scanning or diagonal scanning. Depending on the size of the transform unit and the intra-frame prediction mode, vertical scanning, which scans the 2D coefficient array in the column direction, and horizontal scanning, which scans the 2D block type coefficients in the row direction, can also be used instead of zigzag scanning. In other words, depending on the size of the transform unit and the intra-frame prediction mode, the scanning method to be used can be determined from zigzag scanning, diagonal scanning, vertical scanning, and horizontal scanning.
[0072] The entropy encoder 155 generates a bitstream by encoding the sequence of ID-quantized transform coefficients output from the rearrangement unit 150 using various encoding schemes including context-based adaptive binary arithmetic coding (CABAC), Exponential Golomb, and the like.
[0073] In addition, the entropy encoder 155 encodes information related to block partitioning (such as CTU size, CTU partition flag, QT partition flag, MTT partition type, MTT partition direction, etc.) to allow the video decoding device to partition the block equally to the video encoding device. In addition, the entropy encoder 155 encodes information about the prediction type indicating whether the current block is encoded by intra-frame prediction or inter-frame prediction. The entropy encoder 155 encodes intra-frame prediction information (i.e., information about the intra-frame prediction mode) or inter-frame prediction information (in the case of merge mode, merge index, and in the case of AMVP mode, information about the reference picture index and motion vector difference) according to the prediction type. In addition, the entropy encoder 155 encodes information related to quantization (i.e., information about the quantization parameter and information about the quantization matrix).
[0074] The inverse quantizer 160 dequantizes the quantized transform coefficient output from the quantizer 145 to generate a transform coefficient. The inverse transformer 165 transforms the transform coefficient output from the inverse quantizer 160 from the frequency domain to the spatial domain to reconstruct a residual block.
[0075] The adder 170 reconstructs the current block by adding the reconstructed residual block and the prediction block generated by the predictor 120. When intra prediction is performed on a next sequential block, pixels in the reconstructed current block may be used as reference pixels.
[0076] The loop filter unit 180 performs filtering on the reconstructed pixels to reduce blocking artifacts, ringing artifacts, blurring artifacts, etc. that occur due to block-based prediction and transformation / quantization. The loop filter unit 180 as a loop filter may include all or some of the deblocking filter 182, the sample adaptive offset (SAO) filter 184, and the adaptive loop filter (ALF) 186.
[0077] The deblocking filter 182 filters the boundaries between reconstructed blocks to remove blocking artifacts caused by block-based encoding / decoding, while the SAO filter 184 and ALF 186 perform additional filtering for the deblocked filtered video. The SAO filter 184 and ALF 186 are filters used to compensate for differences between reconstructed and original pixels caused by lossy decoding. The SAO filter 184 applies offsets as CTU units to enhance subjective image quality and coding efficiency. Meanwhile, the ALF 186 performs block-based filtering, applying different filters based on the boundaries of the corresponding blocks and the degree of variance to compensate for distortion. Information regarding the filter coefficients used for the ALF can be encoded and signaled to the video decoding device.
[0078] The reconstructed blocks filtered by the deblocking filter 182, the SAO filter 184, and the ALF 186 are stored in the memory 190. When all blocks in one picture are reconstructed, the reconstructed picture can be used as a reference picture for inter-frame prediction of blocks in a picture to be encoded later.
[0079] The video encoding device may store the bitstream of encoded video data in a non-transitory storage medium or transmit the bitstream to a video decoding device over a communication network.
[0080] Figure 5 is a functional block diagram of a video decoding device that can implement the techniques of the present invention. Figure 5 , describes a video decoding device and components of the device.
[0081] The video decoding apparatus may include an entropy decoder 510 , a reordering unit 515 , an inverse quantizer 520 , an inverse transformer 530 , a predictor 540 , an adder 550 , a loop filter unit 560 , and a memory 570 .
[0082] and Figure 1 Similar to the video encoding device of the present invention, each component of the video decoding device can be implemented as hardware or software or a combination of hardware and software. In addition, the function of each component can be implemented as software, and a microprocessor can also be implemented to execute the function of the software corresponding to each component.
[0083] The entropy decoder 510 extracts information related to block partitioning by decoding a bitstream generated by a video encoding apparatus to determine a current block to be decoded, and extracts prediction information required to reconstruct the current block and information about a residual signal.
[0084] The entropy decoder 510 extracts information about the CTU size from a sequence parameter set (SPS) or a picture parameter set (PPS) to determine the size of the CTU and partitions the picture into CTUs of the determined size. Furthermore, the CTU is determined to be the highest level of the tree structure, i.e., the root node, and partition information of the CTU can be extracted to partition the CTU using the tree structure.
[0085] For example, when a CTU is split using the QTBTTT structure, the first flag (QT_split_flag) related to QT splitting is first extracted to split each node into four nodes in the lower layer. In addition, for the nodes corresponding to the leaf nodes of the QT, the second flag (mtt_split_flag) related to MTT splitting, the splitting direction (vertical / horizontal) and / or the splitting type (binary / ternary) are extracted to split the corresponding leaf nodes into the MTT structure. As a result, each node below the leaf node of the QT is recursively split into a BT or TT structure.
[0086] As another embodiment, when a CTU is split using the QTBTTT structure, a CU split flag (split_cu_flag) indicating whether the CU is split is extracted. When the corresponding block is split, a first flag (QT_split_flag) may also be extracted. During the splitting process, for each node, zero or more recursive MTT splits may occur after zero or more recursive QT splits. For example, for a CTU, MTT splitting may occur immediately, or conversely, multiple QT splits may occur.
[0087] For another example, when a CTU is split using a QTBT structure, the first flag related to the QT split (QT_split_flag) is extracted to split each node into four nodes in the lower layer. In addition, a split flag (split_flag) indicating whether the node corresponding to the QT leaf node is further split into BTs and split direction information are extracted.
[0088] Meanwhile, when the entropy decoder 510 determines the current block to be decoded by partitioning using a tree structure, the entropy decoder 510 extracts information about the prediction type indicating whether the current block is intra-predicted or inter-predicted. When the prediction type information indicates intra-prediction, the entropy decoder 510 extracts syntax elements for intra-prediction information (intra-prediction mode) of the current block. When the prediction type information indicates inter-prediction, the entropy decoder 510 extracts information representing syntax elements for inter-prediction information (i.e., motion vectors and reference pictures referenced by the motion vectors).
[0089] Also, the entropy decoder 510 extracts quantization-related information and extracts information on a quantized transform coefficient of the current block as information on a residual signal.
[0090] The rearrangement unit 515 may change the sequence of 1D quantized transform coefficients entropy-decoded by the entropy decoder 510 into a 2D coefficient array (ie, block) again in the reverse order of the coefficient scanning order performed by the video encoding apparatus.
[0091] The inverse quantizer 520 dequantizes the quantized transform coefficients and dequantizes the quantized transform coefficients by using quantization parameters. The inverse quantizer 520 can also apply different quantization coefficients (scaling values) to the quantized transform coefficients arranged in 2D. The inverse quantizer 520 can perform inverse quantization by applying a matrix of quantization coefficients (scaling values) from the video encoding device to the 2D array of quantized transform coefficients.
[0092] The inverse transformer 530 generates a residual block for the current block by reconstructing a residual signal by inversely transforming the dequantized transform coefficients from the frequency domain into the spatial domain.
[0093] In addition, when the inverse transformer 530 inversely transforms a partial region (sub-block) of the transform block, the inverse transformer 530 extracts a flag (cu_sbt_flag) indicating that only the sub-block of the transform block is transformed, direction (vertical / horizontal) information (cu_sbt_horizontal_flag) of the sub-block, and / or position information (cu_sbt_pos_flag) of the sub-block. The inverse transformer 530 also inversely transforms the transform coefficients of the corresponding sub-block from the frequency domain to the spatial domain to reconstruct the residual signal and fills the untransformed region with a value of "0" as the residual signal to generate a final residual block for the current block.
[0094] In addition, when MTS is applied, the inverse transformer 530 determines the transformation function or transformation matrix to be applied in each of the horizontal and vertical directions by using the MTS information (mts_idx) signaled from the video encoding device. The inverse transformer 530 also performs inverse transformation on the transformation coefficients in the transformation block in the horizontal and vertical directions by using the determined transformation function.
[0095] The predictor 540 may include an intra predictor 542 and an inter predictor 544. The intra predictor 542 is activated when the prediction type of the current block is intra prediction, and the inter predictor 544 is activated when the prediction type of the current block is inter prediction.
[0096] The intra predictor 542 determines an intra prediction mode of the current block among a plurality of intra prediction modes according to the syntax elements of the intra prediction mode extracted from the entropy decoder 510. The intra predictor 542 also predicts the current block by using neighboring reference pixels of the current block according to the intra prediction mode.
[0097] The inter predictor 544 determines a motion vector of a current block and a reference picture to which the motion vector refers by using a syntax element for an inter prediction mode extracted from the entropy decoder 510 .
[0098] The adder 550 reconstructs the current block by adding the residual block output from the inverse transformer 530 and the prediction block output from the inter predictor 544 or the intra predictor 542. When intra-predicting a block to be decoded later, pixels within the reconstructed current block are used as reference pixels.
[0099] The loop filter unit 560, which serves as a loop filter, may include a deblocking filter 562, an SAO filter 564, and an ALF 566. The deblocking filter 562 performs deblocking filtering on the boundaries between reconstructed blocks to remove blocking artifacts caused by block-by-block decoding. The SAO filter 564 and the ALF 566 perform additional filtering on the reconstructed blocks after deblocking filtering to compensate for differences between reconstructed and original pixels caused by lossy encoding. The filter coefficients of the ALF are determined using information about the filter coefficients decoded from the bitstream.
[0100] The reconstructed blocks filtered by the deblocking filter 562, the SAO filter 564, and the ALF 566 are stored in the memory 570. When all blocks in one picture are reconstructed, the reconstructed picture can be used as a reference picture for inter-frame prediction of blocks in a picture to be encoded later.
[0101] In some embodiments, the present invention relates to encoding and decoding a video image as described above. More specifically, the present invention provides a video encoding method and apparatus that compensates for motion of a current block during inter-frame prediction of the current block by performing initial refinement of motion vector candidates, reordering of motion vector candidates, and filtering on a template region.
[0102] The following embodiments may be performed by the inter-frame predictor 124 in the video encoding apparatus. The following embodiments may also be performed by the inter-frame predictor 544 in the video decoding apparatus.
[0103] The video encoding device may generate signaling information associated with this embodiment in terms of optimizing rate-distortion when encoding the current block. The video encoding device may use the entropy encoder 155 to encode the signaling information and transmit the encoded signaling information to the video decoding device. The video decoding device may use the entropy decoder 510 to decode the signaling information associated with decoding the current block from the bitstream.
[0104] In the following description, the term "target block" may be used interchangeably with a current block or a coding unit (CU), or may refer to a certain region of a coding unit.
[0105] In addition, a flag value of true indicates when the flag is set to 1. In addition, a flag value of false indicates when the flag is set to 0.
[0106] The following describes inter prediction techniques related to merge mode.
[0107] I-1. Merge / Skip Mode and MMVD
[0108] Merge / skip modes include normal merge mode, merge mode with motion vector difference mode (MMVD), geometric partitioning mode (GPM), and sub-block merge mode. Here, sub-block merge mode is divided into sub-block based temporal motion vector prediction (SbTMVP) mode and affine merge mode.
[0109] The following describes a method of synthesizing a merge candidate list of motion information in a conventional merge / skip mode. To support the merge / skip mode, the inter-frame predictor 124 in the video encoding apparatus may select a preset number (eg, 6) of merge candidates to form a merge candidate list.
[0110] The inter-frame predictor 124 searches for spatial merging candidates. The inter-frame predictor 124 searches for spatial merging candidates from neighboring blocks, such as Figure 4 As shown in . Up to four spatial merging candidates can be selected. Spatial merging candidates are also called spatial MVPs (SMVPs).
[0111] The inter-frame predictor 124 searches for temporal merge candidates. The inter-frame predictor 124 may add blocks that are co-located with the current block and located in a reference picture (in addition to the current picture holding the target block) as temporal merge candidates. Here, the reference picture may be the same as or different from the reference picture used to predict the current block. A single temporal merge candidate may be selected. A temporal merge candidate is also referred to as a temporal motion vector predictor candidate or a temporal MVP (TMVP) candidate.
[0112] The inter-frame predictor 124 searches for history-based motion vector predictor (HMVP) candidates. The inter-frame predictor 124 may store the motion vectors of the previous h CUs in a table (where h is a natural number), and may then use the stored motion vectors as merge candidates. The size of the table is 6, and the inter-frame predictor 124 stores the motion vectors of the previous CUs in a first-in-first-out (FIFO) manner. This indicates that up to six HMVP candidates are stored in the table. The inter-frame predictor 124 may set the most recent motion vector among the HMVP candidates stored in the table as the merge candidate.
[0113] The inter predictor 124 searches for a pairwise average MVP (PAMVP) candidate and may set an average of motion vectors of a first candidate and a second candidate in the merge candidate list as a merge candidate.
[0114] If the merge candidate list cannot be filled up (ie, the preset number of merge candidates is not satisfied) after performing all the above search processes, the inter predictor 124 adds a zero motion vector as a merge candidate.
[0115] In order to optimize coding efficiency, the inter-frame predictor 124 may determine a merge index indicating one of the candidates in the merge candidate list. The inter-frame predictor 124 may use the merge index to derive a motion vector predictor (MVP) from the merge candidate list and then determine the MVP as the motion vector of the current block. In addition, the video encoding apparatus may signal the merge index to the video decoding apparatus.
[0116] When in skip mode, the video encoding apparatus uses the same motion vector transmission method as in merge mode, but does not transmit a residual block corresponding to a difference between a current block and a predicted block.
[0117] The above-described method of synthesizing a merge candidate list can be similarly performed by the inter-frame predictor 544 in the video decoding device. The video decoding device can decode the merge index. The inter-frame predictor 544 can use the merge index to derive the MVP from the merge candidate list and then determine the MVP as the motion vector of the current block.
[0118] On the other hand, when using the MMVD technique, the inter-frame predictor 124 can use a merge index to derive an MVP from the merge candidate list. For example, the first or second candidate in the merge candidate list can be used as the MVP. Furthermore, to optimize coding efficiency, the inter-frame predictor 124 determines a distance index and a direction index. The inter-frame predictor 124 can use the distance index and the direction index to derive a motion vector difference (MVD), and can then sum the MVD with the MVP to reconstruct the motion vector of the current block. Furthermore, the video encoding device can signal the merge index, distance index, and direction index to the video decoding device.
[0119] The above-described MMVD technique can also be performed by the inter-frame predictor 544 in the video decoding device. The video decoding device can decode the merge index, distance index, and direction index. After the inter-frame predictor 544 combines the merge candidate list, it can use the merge index to derive the MVP from the merge candidate list. After the inter-frame predictor 544 derives the MVD using the distance index and direction index, the inter-frame predictor 544 can sum the MVD with the MVP to reconstruct the motion vector of the current block.
[0120] I-2. Affine Merge Mode
[0121] Inter-frame prediction is motion prediction that reflects a translational motion model, that is, it predicts horizontal (x-axis direction) and vertical (y-axis direction) motion. However, in reality, there may be various forms of motion other than translational motion, such as rotation, enlargement, or reduction. Affine motion prediction can reflect these different forms of motion.
[0122] There are two types of models that can be used for affine motion prediction. One is a model that uses four parameters for two control point motion vectors (CPMVs), one for each of the upper left and upper right corners of the target block to be encoded. The other model uses six parameters for three control point motion vectors, one for each of the upper left, upper right, and lower left corners of the target block.
[0123] The four-parameter affine model is expressed as shown in Equation 1. The motion at the sample position (x, y) in the target block can be calculated as shown in Equation 1. Here, the position of the upper left sample of the target block is assumed to be (0, 0).
[0124] [Formula 1]
[0125]
[0126] Furthermore, a 6-parameter affine model is expressed as shown in Equation 2. The motion at the sample position (x, y) in the target block can be calculated as shown in Equation 2.
[0127] [Formula 2]
[0128]
[0129] Here, (mv 0x , mv 0y ) is the motion vector of the upper left control point, (mv 1x , mv 1y ) is the motion vector of the upper right corner control point, and (mv 2x , mv 2y ) is the motion vector of the lower left control point. W is the horizontal length of the target block, and H is the vertical length of the target block.
[0130] Affine motion prediction may be performed on each sample in the target block using the motion vector calculated according to Equation 1 or Equation 2. Alternatively, to reduce computational complexity, affine motion prediction may be performed on a sub-block basis, for example, by partitioning the target block into sub-blocks of size 4×4.
[0131] The motion vector (mv x , mv y ) is set to have 1 / 16 sample accuracy. In this case, the motion vector (mv) calculated according to equation 1 or 2 is x , mv y ) can be rounded to 1 / 16 sample unit.
[0132] The video encoding apparatus performs intra prediction, inter prediction (translational motion prediction), affine motion prediction, etc., and selects the optimal prediction method by calculating the rate-distortion (RD) cost. To perform affine motion prediction, the inter predictor 124 of the video encoding apparatus determines which of two types of models to use, and determines two or three control points according to the determined type. The inter predictor 124 uses the control point motion vector to calculate the motion vector (mv) of each of the subblocks in the target block. x , mv y Then, the motion vector of each sub-block (mv x , mv y ) is used to perform motion compensation in a reference picture on a sub-block-by-sub-block basis to generate a prediction block for each sub-block in a target block.
[0133] A video encoding device encodes an affine-related syntax element and transmits the affine-related syntax element to a video decoding device, wherein the affine-related syntax element includes a flag indicating whether affine motion prediction has been applied to a target block, type information indicating a type of an affine model, and motion information indicating a motion vector for each control point. When affine motion prediction is performed, the type information and motion information of the control point may be signaled, and the motion vector of the control point may be signaled a number of times determined by the type information.
[0134] The video decoding device determines the type of the affine model and the control point motion vector using the signaled syntax, and calculates the motion vector (mv) of each 4×4 sub-block in the target block by using Equation 1 or 2. x , mv y If the motion vector resolution information of the affine motion vector of the target block is signaled, the motion vector (mv) is modified with the precision identified by the motion vector resolution information by using an operation such as rounding. x , mv y ).
[0135] The video decoding device decodes the video by using the motion vector (mv x , mv y ), generates a prediction block for each sub-block by performing motion compensation within the reference picture.
[0136] In order to reduce the amount of bits required to encode the control point motion vectors, the above-described conventional method of inter-frame prediction (translational motion prediction) may be applied.
[0137] In one example, when in affine merge mode, the inter-frame predictor 124 of the video encoding device organizes a predefined number (e.g., 5) of affine merge candidate lists. First, the inter-frame predictor 124 of the video encoding device derives inherited affine merge candidates from neighboring blocks of the target block. For example, by Figure 4 The neighboring samples A0, A1, B0, B1 and B2 of the target block shown in derive a predefined number of inherited affine merge candidates and generate a merge candidate list. Each of the inherited affine merge candidates in the candidate list corresponds to a combination of two or three CPMVs (control point motion vectors).
[0138] The inter-frame predictor 124 derives inherited affine merge candidates from the control point motion vectors of neighboring blocks of the target block predicted in affine mode. Some embodiments may limit the number of merge candidates derived from neighboring blocks predicted in affine mode. For example, the inter-frame predictor 124 may derive two inherited affine merge candidates (one of A0 and A1, and one of B0, B1, and B2) from neighboring blocks predicted in affine mode. The priority order may be in the order of A0, A1, followed by B0, B1, and B2.
[0139] On the other hand, if the total number of merge candidates is greater than three, the inter predictor 124 may derive the required number of supplementary constructed affine merge candidates from the translational motion vectors of the neighboring blocks.
[0140] Figure 8 is a diagram illustrating a method of deriving constructed affine merge candidates for affine motion prediction.
[0141] The inter-frame predictor 124 derives control point motion vectors CPMV1, CPMV2, and CPMV3 from each neighboring block in the group {B2, B3, A2}, each neighboring block in the group {B1, B0}, and each neighboring block in the group {A1, A0}. As an example, the priority within each group of neighboring blocks can be B2, B3, and A2 in order, B1 and B0 in order, and A1 and A0 in order. Another control point motion vector CPMV4 is derived from the co-located block T in the reference picture. The inter-frame predictor 124 combines two or three of the four control point motion vectors to generate the required number of supplementary constructed affine merge candidates. The priority order of the combinations is as follows. The elements within each group are listed in the following order: control point motion vector at the top left corner, top right corner, and then bottom left corner.
[0142] {CPMV1,CPMV2,CPMV3},{CPMV1,CPMV2,CPMV4},{CPMV1,CPMV3,CPMV4},{CPMV2,CPMV3,CPMV4},{CPMV1,CPMV2},{CPMV1,CPMV3}
[0143] If the merge candidate list cannot be filled by using the inherited affine merge candidate and the constructed affine merge candidate, the inter predictor 124 may add a zero motion vector as a candidate.
[0144] The inter-frame predictor 124 selects a merge candidate from the merge candidate list to optimize decoding efficiency and determines a merge index indicating the merge candidate. The inter-frame predictor 124 uses the selected merge candidate to perform affine motion prediction on the target block. When the merge candidate consists of two control point motion vectors, affine motion prediction is performed using a four-parameter model. On the other hand, when the merge candidate consists of three control point motion vectors, affine motion prediction is performed using a six-parameter model. The video encoding device encodes the merge index and signals it to the video decoding device.
[0145] The video decoding apparatus decodes the merge index. The inter-frame predictor 544 of the video decoding apparatus synthesizes a merge candidate list in the same manner as the video encoding apparatus, and performs affine motion prediction by using a control point motion vector corresponding to the merge candidate indicated by the merge index.
[0146] I-3. Geometric Partitioning Mode (GPM)
[0147] In GPM, the inter-frame predictor 124 performs inter-frame prediction based on the sub-regions that have been geometrically partitioned from the current block. The inter-frame predictor 124 performs inter-frame prediction on the two sub-regions by using different motion information items (i.e., motion vectors). The inter-frame predictor 124 generates a final prediction signal by weighted summing the prediction signals from each region to minimize discontinuity at the boundary between the sub-regions.
[0148] When composing the GPM candidate list, the motion information of each sub-region is derived from the regular merge candidate list. If the index of the merge candidate list is even, the motion information in L0 (first reference list) is selected, and if the index is odd, the motion information in L1 (second reference list) is selected.
[0149] I-4. Decoder-side Motion Vector Refinement (DMVR)
[0150] Decoder-side motion vector refinement (DMVR) is a method that refines motion vectors at the decoder side by fine-tuning the motion vectors (MV0 and MV1) in bi-prediction using bilateral matching (BM). Hereinafter, motion vectors in bi-prediction and motion vector pairs are used interchangeably.
[0151] In bidirectional prediction, the video decoding device searches for a refined motion vector that is adjacent to the initial motion vector generated from the reference pictures in the reference lists L0 and L1. Here, the initial motion vectors are the two motion vectors MV0 and MV1 for bidirectional prediction. The BM technique calculates the BM cost, which is the distortion between the two candidate blocks in the reference pictures of L0 and L1. The SAD (sum of absolute differences) or SSE (sum of squared errors) between the two candidate blocks can be calculated as the BM cost. The video decoding device generates a refined motion vector from the motion vector candidate with the minimum BM cost as shown in Equation 3.
[0152] [Formula 3]
[0153] MV_0'=MV_0+MVoffset
[0154] MV_1'=MV_1-MVoffset
[0155] Here, MVoffset is the offset applied to the initial motion vector when motion vector refinement is performed, and is the difference between the candidate motion vector and the initial motion vector. The offset can be formed as the sum of an integer offset in integer samples and a sub-pixel offset in sub-pixel or sub-pixel samples. As shown in Equation 3, the two motion vector candidates follow the mirroring convention of the offset.
[0156] I-5. Sub-block based temporal motion vector prediction (SbTMVP)
[0157] Similar to the temporal merge candidate described above, SbTMVP utilizes motion information within a co-located picture to improve the motion vector prediction of each sub-block in merge mode. Each sub-block is one of the blocks generated by partitioning the current block, and a co-located picture refers to the picture containing the co-located block. Before deriving the motion information of the co-located picture, SbTMVP applies motion shifting.
[0158] SbTMVP determines e.g. Figure 8 , and whether the pixel A1 described in [ 15 ] has a motion vector that uses the co-located picture as a reference picture. If the pixel A1 has a motion vector that uses the co-located picture as a reference picture, then the motion vector of the pixel A1 is selected as the motion shift. On the other hand, if the pixel A1 does not have a motion vector that uses the co-located picture as a reference picture, then the motion shift is set to zero.
[0159] SbTMVP applies motion shifting to co-located blocks within a co-located picture. For each sub-block, SbTMVP extracts motion information (motion vector and reference index) for the corresponding sub-block in the co-located picture. SbTMVP applies temporal motion scaling to the extracted motion information to generate motion information for each sub-block.
[0160] The following embodiments are described with reference to a video decoding apparatus, but may be implemented in a video encoding apparatus in the same or similar manner.
[0161] II. Implementations According to the Present Disclosure
[0162] Figure 9 is a detailed block diagram of a portion of a video decoding device according to at least one embodiment of the present disclosure.
[0163] According to some embodiments, a video decoding apparatus may determine a prediction and transformation unit, and in response to a current block corresponding to the determined unit, perform prediction and inverse transformation by using a determined prediction technique and prediction mode to ultimately generate a reconstructed block of the current block. Figure 9 The operations shown in FIG. 5 can be performed by the inverse transformer 530, the predictor 540, and the adder 550 of the video decoding device. Figure 9 The same operations shown in FIG can be performed by the inverse transformer 165, the image splitter 110, the predictor 120, and the adder 170 of the video encoding device. In this case, the video decoding device uses the encoding information parsed from the bitstream, but the video encoding device can use the encoding information set from a higher level in terms of minimizing rate distortion. Hereinafter, for convenience, the embodiment is described with the video decoding device as the center.
[0164] like Figure 5 As shown in FIG, the predictor 540 includes an intra predictor 542 and an inter predictor 544 depending on the prediction technique, but as shown in FIG. Figure 9 As shown in , the predictor 540 may include all or part of a prediction unit determiner 902 , a prediction technique determiner 904 , a prediction mode determiner 906 , and a prediction executor 908 .
[0165] If the color format of the input video is a YUV format (YUV420, YUV411, YUV422, YUV444, etc.), the video decoding device may perform prediction and reconstruction of the luminance component and then perform prediction and reconstruction of the chrominance component. In this way, the luminance component and the chrominance component may be obtained by Figure 9 On the other hand, if the color format of the input video is RGB, the video encoding device may perform color format conversion from RGB to YUV and then encode the converted video. Here, in the YUV format, the color format represents the correspondence between pixels in the luminance component and pixels in the chrominance component.
[0166] A prediction unit determiner 902 determines a prediction unit (PU). A prediction technique determiner 904 determines a prediction technique for the PU, such as intra prediction, inter prediction, intra block copy (IBC) mode, palette mode, etc. A prediction mode determiner 906 determines a detailed prediction mode for the prediction technique. A prediction executor 908 generates a prediction block for the current block based on the determined prediction mode.
[0167] The inverse transformer 530 includes a transform unit determiner 910 and an inverse transform performer 912. The transform unit determiner 910 determines a transform unit (TU) in response to a dequantized signal of the current block, and the inverse transform performer 912 inversely transforms the transform unit represented by the dequantized signal to generate a residual signal.
[0168] The adder 550 sums the predicted block and the residual signal to generate a reconstructed block. The reconstructed block is stored in a memory and can be used to predict other blocks in the future.
[0169] The prediction unit determined by the prediction unit determiner 902 may be a sub-block of the current block or a sub-block partitioned from the current block. In this case, the prediction unit of the chrominance component may correspond in size to the prediction unit of the luma component, depending on the color format. Alternatively, the prediction units of the luma component and the chrominance component may be determined separately, and prediction may be performed with respect to the prediction unit of the chrominance component.
[0170] The prediction technique determiner 904 determines a prediction technique for the prediction unit. As described above, the prediction technique may be one of inter-frame prediction, intra-frame prediction, IBC mode, and palette mode. In this case, the prediction technique for the chroma component may be determined to be the same as the prediction technique for the corresponding luma component without signaling or parsing any additional information.
[0171] In one embodiment, if the prediction technique for the current block is not intra prediction, the video decoding device parses 1-bit flag information. If the parsed flag indicates skip mode, the video decoding device may determine the prediction mode for the current block as inter-prediction merge mode or IBC merge mode. The video decoding device may also use the prediction signal as a reconstruction signal and skip inverse transform.
[0172] On the other hand, if the parsed flag does not indicate a skip mode with respect to the current block, the prediction technology determiner 904 may determine one of the prediction technologies, such as inter-frame prediction, intra-frame prediction, IBC mode, palette mode, etc., by parsing a series of 1-bit flags into the prediction technology of the current block.
[0173] For example, if skip mode is not applied to the current block and inter prediction or IBC mode is determined as the prediction technique, the video decoding apparatus parses a 1-bit flag and determines the prediction mode of the current block as merge mode or AMVP (Advanced Motion Vector Prediction) mode based on the parsed flag.
[0174] The prediction mode determiner 906 determines a detailed prediction mode for the prediction technique.
[0175] For example, when the prediction mode of the current block is merge mode or skip mode, the prediction mode determiner 906 may determine whether to perform sub-block based prediction by parsing a one-bit flag. If sub-block based prediction is to be performed, affine prediction or SbTMVP prediction may be performed. If sub-block based prediction is not to be performed, prediction may be performed according to techniques such as conventional merge mode, MMVD, DMVR, GPM, etc.
[0176] The prediction executor 908 generates a prediction block for the current block according to the determined prediction technique and prediction mode.
[0177] For example, when the prediction technique for the current block is inter prediction and the prediction mode is merge mode or skip mode, the prediction execution unit 908 parses the merge index and synthesizes a motion vector candidate list. The prediction execution unit 908 reorders the motion vector candidates in the motion vector candidate list and uses the merge index to obtain motion information for the current block from the reordered motion vector candidate list. The prediction execution unit 908 uses the motion information to generate a final prediction signal for the current block.
[0178] After parsing the merge index, the video decoding device synthesizes a motion vector candidate list for the current block according to the above-mentioned method for synthesizing a merge candidate list of motion information in conventional merge / skip mode. The motion vector candidate list includes a preset number (e.g., n) of motion vector candidates. Hereinafter, the terms motion vector candidate list and merge candidate list are used interchangeably. Motion vector candidate and merge candidate are used interchangeably. A motion vector candidate can be a unidirectional motion vector or a bidirectional motion vector.
[0179] When synthesizing a motion vector candidate list, the video decoding apparatus performs initial motion refinement on motion vector candidates included in the motion vector candidate list.
[0180] Figure 10 is a diagram illustrating motion vector refinement using template matching, in accordance with at least one embodiment of the present disclosure.
[0181] In one embodiment, when the prediction method of the current block is not a sub-block based prediction method, the video decoding device refines the unidirectional motion vector among the motion vector candidates by using template matching, such as Figure 10 As shown in . The video decoding device refines the unidirectional motion vector by using template matching between a template in a reconstructed area adjacent to the current block and a template at a corresponding position adjacent to the reference block. Here, the reference block is indicated by the above-mentioned unidirectional motion vector.
[0182] Hereinafter, the template in the reconstructed area adjacent to the current block is used interchangeably with the "template of the current block." The template at the corresponding position adjacent to the reference block is used interchangeably with the "corresponding template of the reference block." The area covered by the template of the current block is denoted by the template area of the current block. The area covered by the corresponding template of the reference block is denoted by the corresponding template area of the reference block.
[0183] In order to calculate the template matching cost between samples in the template area of the current block and samples in the corresponding template area of the reference block, the video decoding device can use one of multiple measurements, such as the sum of absolute differences (SAD), mean square error (MSE), sum of absolute transform differences (SATD), etc.
[0184] like Figure 10 As shown, the video decoding apparatus performs template matching using a template within a template matching search area defined by "a" and "b." The template matching search area may be adaptively determined based on the size and aspect ratio of the current block, etc. The video decoding apparatus may refine the unidirectional motion vector so as to minimize the template matching cost within the template matching search area.
[0185] The template matching search area exists within the reference picture. Hereinafter, the template matching search area and the search area are used interchangeably. The corresponding template of the reference block exists within the search area.
[0186] As another example, when the prediction method of the current block is a sub-block based prediction method, the video decoding apparatus uses template matching to refine the unidirectional motion vector among the motion vector candidates. For the sub-blocks of size p×q at the top and left of the current block, as in Figure 11 In an example of , a video decoding apparatus refines a unidirectional motion vector by using template matching between a template adjacent to each subblock and a template at a corresponding position adjacent to a reference subblock.
[0187] Figure 12 is a diagram illustrating motion vector refinement using template matching according to another embodiment of the present disclosure.
[0188] In one embodiment, when the prediction method of the current block is not a sub-block based prediction method, the video decoding device refines the bidirectional motion vector among the motion vector candidates by using template matching, such as Figure 12 As shown in . The video decoding device refines the bidirectional motion vectors (MV0 and MV1) by using template matching between a template in a reconstructed area adjacent to the current block and a template at a corresponding position adjacent to the reference block. Alternatively, the video decoding device can refine the bidirectional motion vector by using bilateral matching (BM) between reference blocks. Here, the reference blocks in the reference pictures in both directions (L0 reference picture and L1 reference picture) are represented by the above-mentioned bidirectional motion vectors.
[0189] In one embodiment, if the motion vector candidate is a bidirectional motion vector, the video decoding apparatus may refine the bidirectional motion vector by using template matching and then apply further refinement to the primarily refined bidirectional motion vector using BM between reference blocks.
[0190] At this time, in order to calculate the template matching cost between samples in the template region of the current block and samples in the corresponding template region of the reference block, the video decoding apparatus may use one of measurements such as SAD, MSE, SATD, etc.
[0191] The video decoding device performs template matching by using a template in a template matching search area defined by "a" and "b", as shown in FIG. Figure 12 At this time, the template matching search area can be adaptively determined based on the size, aspect ratio, etc. of the current block. The video decoding device can refine the bidirectional motion vectors within the template matching search area in such a way as to minimize the template matching cost.
[0192] The video decoding apparatus may calculate the BM cost between the reference block in the reference picture L0 and the reference block in the reference picture L1 using one of various measurements such as SAD, MSE, SATD, etc.
[0193] In one example, when the prediction method of the current block is a sub-block based prediction method, the video decoding device refines the bidirectional motion vector in the motion vector candidate by using template matching. For the p×q size sub-block at the upper left of the current block, as shown in FIG. Figure 11 As shown, the video decoding apparatus refines the bidirectional motion vector by using template matching between a template adjacent to each subblock and a template at a corresponding position adjacent to a reference subblock.
[0194] For another example, when the prediction method of the current block is a sub-block based prediction method, the video decoding device refines the motion information of the bidirectional motion vector by using the BM. Here, the motion information may be the initial motion information or the motion information after the initial refinement according to a non-BM method. For the sub-block of size p×q at the upper left of the current block, the video decoding device may perform BM by comparing each sub-block with a reference sub-block at the corresponding position, such as Figure 11 shown.
[0195] As yet another embodiment, the video decoding apparatus may omit the initial motion refinement.
[0196] After performing the initial motion refinement, the video decoding apparatus reorders the motion vector candidates in the motion vector candidate list. Alternatively, when the initial motion refinement is omitted, the video decoding apparatus may reorder the motion vector candidates before the initial motion refinement.
[0197] In one embodiment, the video decoding device generates a reference block based on the motion information of each motion vector candidate and then calculates a template matching cost between the template of the current block and the corresponding template of the reference block. The video decoding device compares the template matching costs of the motion vector candidates. The video decoding device reorders the motion vector candidates in the motion vector candidate list in order of increasing template matching cost. In doing so, the video decoding device may use one of the following metrics, such as SAD, MSE, SATD, etc., to calculate the template matching cost between samples in the template area of the current block and samples in the corresponding template area of the reference block.
[0198] As another embodiment, the video decoding apparatus may omit reordering of motion vector candidates.
[0199] In one embodiment, the video decoding device performs template matching within a search area, and the search area may be adaptively determined based on the size, aspect ratio, etc. of the current block.
[0200] In one embodiment, for the corresponding templates of two reference blocks obtained from the bidirectional motion vectors among the motion vector candidates, the video decoding device calculates the template matching cost between the template of the current block and the corresponding template of each reference block, and then averages the two template matching cost values to calculate the final template matching cost for the aforementioned motion vector candidate.
[0201] The video decoding apparatus uses the merge index to select the motion vector of the current block from the motion vector candidates in the reordered motion vector candidate list. Alternatively, when reordering of the motion vector candidates is omitted, the video decoding apparatus uses the merge index to select the motion vector of the current block from the motion vector candidates in the motion vector candidate list before reordering.
[0202] When a bidirectional motion vector is selected according to a merge index, the video decoding device performs final motion refinement by using a reference block obtained by using the bidirectional motion vectors (MV0 and MV1). The video decoding device can perform final motion refinement by performing sub-block-based BM and / or bidirectional optical flow (BDOF). BDOF further compensates for the motion of samples predicted by using bidirectional motion prediction based on optical flow, assuming that the samples or objects constituting the video move at a constant speed and that the sample values hardly change.
[0203] The video decoding device performs motion compensation for the current block by using the motion vector to which the final motion refinement is applied or using the motion vector selected according to the merge index, thereby generating a final prediction block for the current block. The video decoding device then decodes the residual block and sums the final prediction block and the residual block to generate a reconstructed block for the current block.
[0204] When performing template matching-based motion refinement or template matching-based motion vector candidate reordering on the motion vector candidate of the current block, the video decoding apparatus performs the following processing. In some embodiments, when the prediction method of the current block is not a sub-block-based prediction method, the following process may be applied.
[0205] Figure 13 FIG. 4 is a diagram illustrating filtering performed on a template area according to one embodiment of the present disclosure.
[0206] In one embodiment, before performing template matching, the video decoding device may perform filtering on the reconstructed template area adjacent to the current block (ie, the template area of the current block). Figure 13 As shown in the top diagram of FIG, the template area of the current block split by the block is described below. The video decoding device can perform filtering at the partition boundary within the template area, such as Figure 13 as shown in the bottom diagram of .
[0207] In one embodiment, if the boundary receiving filtering is a horizontal boundary, the video decoding apparatus applies filtering to the template region (TM(i, j)) of the current block in the vertical direction as shown in Equation 4.
[0208] [Formula 4]
[0209]
[0210] In addition, if the boundary receiving the filtering is a vertical boundary, the video decoding apparatus applies the filtering to the template region (TM(i, j)) of the current block in the horizontal direction, as shown in Equation 5.
[0211] [Formula 5]
[0212]
[0213] In Equation 4 or Equation 5, the filter f(k) may be a smoothing filter having coefficients {1 / 4, 1 / 2, 1 / 4}.
[0214] As another embodiment, the video decoding apparatus applies a filter to the template region (TM(i, j)) of the current block, as shown in Equation 6.
[0215] [Formula 6]
[0216] TM(i,j)=M(i,j)×TM(i,j)
[0217] In equation 6, M(i, j) represents the matrix that performs pixel-by-pixel operations. Figure 14 As shown in the bottom illustration of , M(i, j) may have a value of 0 at locations around the partition boundary within the template region of the current block and a value of 1 at the remaining locations within the template region of the current block.
[0218] In addition, when Figure 14 The bottom shows the configuration M(i, j), for the corresponding template TM of the reference block used for template matching ref (i, j), the video decoding device may perform filtering as shown in Equation 6.
[0219] In one embodiment, for the template area of the current block, the video decoding device may divide the template area into sub-areas of size c×d, such as Figure 15 As shown, filtering can then be performed on the boundaries between the sub-regions. In this case, c and d can be adaptively determined based on the size, aspect ratio, etc. of the current block.
[0220] For example, if the difference between adjacent samples along the boundary of each sub-region is greater than a preset threshold, the video decoding device may apply filtering according to Equation 4 or Equation 5 to the boundary. The video decoding device may calculate the difference between adjacent samples by using one of the measurements such as SAD, MSE, etc. The threshold may be a predetermined value based on an agreement between the video encoding device and the video decoding device. Alternatively, the threshold may be an average value of the samples in the template region of the current block.
[0221] In one embodiment, by using the filtered template region TM(i, j) of the current block and the filtered corresponding template region TM of the reference block ref (i, j), the video decoding device calculates the template matching cost TM according to formula 7 cost The video decoding apparatus may calculate the template matching cost by using one of measurements such as SAD, MSE, etc.
[0222] [Formula 7]
[0223]
[0224] As another embodiment, for the template region of the current block, the video decoding apparatus may calculate the template matching cost according to Equation 8 or Equation 9. Since filtering is applied in Equation 8 or Equation 9, filtering according to Equations 4, 5, and 6 may not be applied separately to the template region TM(i, j) of the current block.
[0225] [Formula 8]
[0226] TM cost =∑ i,j |∑ k (f(k)×TM(i,j))-TM ref (i,j)|
[0227] [Formula 9]
[0228] TM cost =∑ i,j |M(i,j)×TM(i,j)-M(i,j)×TM ref (i,j)|
[0229] In Formula 8, f(k) can be derived using the same method as that used to derive f(·) in Formula 4 or Formula 5. Furthermore, in Formula 9, M(i, j) can be derived using the same method as that used to derive M(i, j) in Formula 6.
[0230] In the following, see Figure 16 and Figure 17 Describes a method for encoding / reconstructing the current block by using template matching in merge mode.
[0231] Figure 16 The present invention is a flowchart of a method for encoding a current block by a video encoding device according to at least one embodiment of the present disclosure.
[0232] The video encoding apparatus generates a motion vector candidate list including a preset number of motion vector candidates (S1600). Here, each motion vector candidate may be a unidirectional motion vector or a bidirectional motion vector.
[0233] The video encoding apparatus performs initial motion refinement on motion vector candidates ( S1602 ).
[0234] The video encoding apparatus may perform initial motion refinement by using template matching and / or BM.
[0235] The video encoding device performs initial motion refinement on the current block. The video encoding device refines each motion vector candidate using template matching between a template of the current block and a corresponding template in a reference block. Optionally, the video encoding device performs initial motion refinement on subblocks. In this case, the subblocks may be generated by segmenting the current block. The video encoding device refines each motion vector candidate using template matching between a template adjacent to each subblock and a corresponding template in a reference subblock.
[0236] The video encoding device may omit the initial motion refinement.
[0237] When performing initial motion refinement using template matching, the video encoding apparatus may apply filtering to block partition boundaries within the template region of the current block.
[0238] The video encoding apparatus reorders motion vector candidates ( S1604 ).
[0239] The video encoding device may generate a reference block based on the motion information of each motion vector candidate and then calculate a template matching cost between a template of the current block and a corresponding template of the reference block. The video encoding device compares the template matching costs of the motion vector candidates and reorders the motion vector candidates in the merge candidate list in order of increasing template matching costs.
[0240] The video encoding apparatus may omit reordering motion vector candidates.
[0241] When reordering motion vector candidates using template matching, the video encoding apparatus may apply filtering to block partition boundaries within a template region of the current block.
[0242] The video encoding apparatus determines a merge index indicating one of the motion vector candidates (S1606). In terms of rate-distortion optimization, the video encoding apparatus may determine the merge index.
[0243] The video encoding apparatus selects motion information of the current block from the motion vector candidates using the merge index (S1608).
[0244] When the motion information of the current block is a set of bidirectional motion vectors, the video encoding apparatus may perform final refinement of the bidirectional motion vectors by using at least one of sub-block based bilateral matching and BDOF.
[0245] The video encoding apparatus generates a prediction block of the current block by using the selected motion information (S1610).
[0246] The video encoding apparatus encodes the merge index (S1612).
[0247] The video encoding device then subtracts the predicted block from the original block of the current block to generate a residual block. The video encoding device transforms / quantizes / encodes the residual block to generate a bitstream and sends the generated bitstream to the video decoding device.
[0248] Figure 17 The present invention is a flowchart of a method for reconstructing a current block by a video decoding device according to at least one embodiment of the present disclosure.
[0249] A video decoding apparatus decodes a merge index of a current block from a bitstream (1700).
[0250] The video decoding apparatus generates a motion vector candidate list including a preset number of motion vector candidates (1702). Here, each motion vector candidate may be a unidirectional motion vector or a bidirectional motion vector.
[0251] The video decoding apparatus performs initial motion refinement on the motion vector candidates ( 1704 ).
[0252] The video decoding apparatus may perform initial motion refinement by using template matching and / or BM.
[0253] The video decoding device performs initial motion refinement on the current block. The video decoding device refines each motion vector candidate by using template matching between a template of the current block and a corresponding template of a reference block. Alternatively, the video decoding device performs initial motion refinement on subblocks. In this case, the subblocks may be generated by segmenting the current block. The video decoding device refines each motion vector candidate by using template matching between a template adjacent to each subblock and a corresponding template of a reference subblock.
[0254] The video decoding device may omit the initial motion refinement.
[0255] When performing initial motion refinement using template matching, the video decoding apparatus may apply filtering to block partition boundaries within the template region of the current block.
[0256] The video decoding apparatus reorders the motion vector candidates ( 1706 ).
[0257] The video decoding device generates a reference block based on the motion information of each motion vector candidate, and then calculates the template matching cost between the template of the current block and the corresponding template of the reference block. The video decoding device compares the template matching costs of the motion vector candidates. The video decoding device reorders the motion vector candidates in the merge candidate list in order of increasing template matching cost.
[0258] The video decoding apparatus may omit reordering motion vector candidates.
[0259] When reordering motion vector candidates using template matching, the video decoding apparatus may apply filtering to block partition boundaries within a template region of the current block.
[0260] The video decoding apparatus selects motion information of the current block from the motion vector candidates using the merge index ( 1708 ).
[0261] When the motion information of the current block is a bidirectional motion vector set, the video decoding apparatus performs final refinement of the bidirectional motion vector by using at least one of subblock-based bilateral matching and BDOF.
[0262] The video decoding apparatus generates a prediction block of the current block by using the selected motion information (1710).
[0263] The video decoding device decodes / dequantizes / detransforms the bitstream to generate a residual block. The video decoding device then sums the predicted block and the residual block to generate a reconstructed block of the current block.
[0264] Although the steps in the various flowcharts are described as being performed sequentially, these steps merely illustrate the technical concepts of some embodiments of the present disclosure. Therefore, a person skilled in the art can perform the steps by changing the order described in the various figures or by performing two or more steps in parallel. Therefore, the steps in the various flowcharts are not limited to the chronological order shown.
[0265] It should be understood that the above description presents illustrative embodiments that can be implemented in various other ways. The functions described in some embodiments can be implemented by hardware, software, firmware, and / or a combination thereof. It should also be understood that the functional components described in this disclosure are marked with "... unit" to strongly emphasize their possibility of independent implementation.
[0266] Meanwhile, the various methods or functions described in some embodiments may be implemented as instructions stored in a non-transitory recording medium that can be read and executed by one or more processors. For example, the non-transitory recording medium may include various types of recording devices in which data is stored in a form readable by a computer system. For example, the non-transitory recording medium may include a storage medium such as an erasable programmable read-only memory (EPROM), a flash drive, an optical drive, a magnetic hard drive, and a solid-state drive (SSD).
[0267] Although the embodiments of the present disclosure have been described for illustrative purposes, it will be appreciated by those skilled in the art that various modifications, additions, and substitutions are possible without departing from the concept and scope of the present disclosure. Therefore, for the sake of brevity and clarity, the embodiments of the present disclosure have been described. The scope of the technical concept of the embodiments of the present disclosure is not limited by the shown limitations. Therefore, it will be appreciated by those skilled in the art that the scope of the present disclosure should not be limited by the embodiments explicitly described above, but by the claims and their equivalents.
[0268] (reference number)
[0269] 124: Inter-frame predictor
[0270] 544: Interframe predictor
[0271] 902: Prediction unit determiner
[0272] 904: Forecasting Technology Determiner
[0273] 906: Prediction Mode Determiner
[0274] 908: Predictive Executor.
[0275] Cross-citation to related applications
[0276] This application claims priority to and the benefit of Korean Patent Application No. 10-2022-0165720, filed on December 1, 2022, and Korean Patent Application No. 10-2023-0162973, filed on November 22, 2023, the entire contents of each of which are incorporated herein by reference.
Claims
1. A method for reconstructing a current block by a video decoding device, the method comprising: Decode the merge index of the current block from the bitstream; generating a motion vector candidate list for the current block, wherein the motion vector candidate list includes a preset number of motion vector candidates (motion vector candidates), and any of the motion vector candidates (one motion vector candidate) is a unidirectional motion vector or a bidirectional motion vector; performing initial motion refinement on the motion vector candidate; reordering the motion vector candidates; Using the merge index, selecting the motion information of the current block from the motion vector candidates; and generating a prediction block for the current block using the selected motion information, Wherein, performing the initial motion refinement or reordering the motion vector candidates comprises: when using template matching, Filtering is applied to block partition boundaries within the template region of the current block.
2. The method according to claim 1, further comprising: When the motion information of the current block indicates the bidirectional motion vector, The bidirectional motion vector is finally refined using at least one technique of sub-block based bilateral matching (bilateral matching) and bidirectional optical flow.
3. The method according to claim 1, wherein Performing the initial motion refinement includes: when the motion vector candidate is the unidirectional motion vector, A reference block indicated by the unidirectional motion vector is derived, and the unidirectional motion vector is refined using template matching between a template of the current block and a corresponding template of the reference block.
4. The method according to claim 3, wherein: Performing the initial motion refinement includes: determining a search area within a reference picture including the reference block based on a size or an aspect ratio of the current block; and The template matching is performed within the search area.
5. The method according to claim 3, wherein: The template of the current block represents a template within a reconstruction area adjacent to the current block, and the corresponding template of the reference block represents a template at a corresponding position adjacent to the reference block.
6. The method according to claim 1, wherein Performing the initial motion refinement includes: when the motion vector candidate is the bidirectional motion vector, A reference block indicated by the bidirectional motion vector is derived, and the bidirectional motion vector is refined using template matching between a template of the current block and a corresponding template of the reference block.
7. The method according to claim 1, wherein Performing the initial motion refinement includes: when the motion vector candidate is the bidirectional motion vector, A reference block indicated by the bidirectional motion vector is derived, the bidirectional motion vector is refined using template matching between a template of the current block and a corresponding template of the reference block, and the refined bidirectional motion vector is further refined using bilateral matching between the reference blocks.
8. The method according to claim 1, wherein Reordering the motion vector candidates includes: generating a reference block according to the motion information of the motion vector candidate; generating a template matching cost between the template of the current block and the corresponding template of the reference block; and The motion vector candidates are reordered in order of increasing template matching costs of the motion vector candidates.
9. The method according to claim 8, wherein Generating the template matching cost includes: when the motion vector candidate is a bidirectional motion vector, For multiple reference blocks indicated by the bidirectional motion vector, the template matching cost between the template of the current block and the corresponding template of each reference block is calculated, and the template matching cost of the motion vector candidate is calculated by averaging the template matching costs of multiple reference blocks.
10. The method according to claim 1, wherein Applying the filtering to the block partition boundaries comprises: When the block partition boundary is a horizontal boundary, the template region of the current block is filtered in a vertical direction, and when the block partition boundary is a vertical boundary, the template region of the current block is filtered in a horizontal direction.
11. The method according to claim 10, wherein: Performing the initial motion refinement or the reordering of the motion vector candidates comprises: generating a reference block based on the motion information of the motion vector candidate; and A template matching cost is calculated between the filtered template region of the current block and a corresponding template region of the reference block.
12. The method according to claim 1, wherein Applying the filtering to the block partition boundaries comprises: A template region of the current block is multiplied by a matrix that performs a pixel-by-pixel operation, wherein the matrix has values of 0 at positions adjacent to the block partition boundary and values of 1 at remaining positions within the template region of the current block.
13. The method according to claim 12, wherein: Performing the initial motion refinement or the reordering of the motion vector candidates comprises: generating a reference block based on the motion information of the motion vector candidate; multiplying the corresponding template area of the reference block by the matrix; and A template matching cost is calculated between the template region of the current block multiplied by the matrix and the corresponding template region of the reference block multiplied by the matrix.
14. A method for encoding a current block by a video encoding device, the method comprising: generating a motion vector candidate list for the current block, wherein the motion vector candidate list includes a preset number of motion vector candidates (motion vector candidates), and any motion vector candidate (one motion vector candidate) is a unidirectional motion vector or a bidirectional motion vector; performing initial motion refinement on the motion vector candidate; reordering the motion vector candidates; determining a merge index indicating one of the motion vector candidates; Using the merge index, selecting the motion information of the current block from the motion vector candidates; and generating a prediction block for the current block using the selected motion information, Wherein, performing the initial motion refinement or reordering the motion vector candidates comprises: when using template matching, Filtering is applied to block partition boundaries within the template region of the current block.
15. The method according to claim 14, further comprising: The merge index is encoded.
16. The method according to claim 14, further comprising, when the motion information of the current block indicates a bidirectional motion vector: The bidirectional motion vector is finally refined using at least one technique of sub-block based bilateral matching (bilateral matching) and bidirectional optical flow.
17. A computer-readable recording medium storing a bitstream generated by a video encoding method, the video encoding method comprising: generating a motion vector candidate list of the current block, wherein the motion vector candidate list includes a preset number of motion vector candidates (motion vector candidates), and any motion vector candidate (one motion vector candidate) is a unidirectional motion vector or a bidirectional motion vector; performing initial motion refinement on the motion vector candidate; reordering the motion vector candidates; determining a merge index indicating one of the motion vector candidates; Using the merge index, selecting the motion information of the current block from the motion vector candidates; and generating a prediction block for the current block using the selected motion information, Wherein, performing the initial motion refinement or reordering the motion vector candidates comprises: when using template matching, Filtering is applied to block partition boundaries within the template region of the current block.
Citation Information
Patent Citations
Batteries, battery modules, electrical equipment and battery manufacturing methods
KR1020220165720A
Lipid compounds and their applications in nucleic acid delivery
KR1020230162973A