Video encoding method and apparatus using template adjustment for displacement symbol prediction

By adaptively adjusting the principle of minimizing template size and template matching costs, the problem of high bit consumption of block vector difference and motion vector difference in existing video encoding technologies is solved, and encoding efficiency and video quality are improved.

CN120303926APending Publication Date: 2025-07-11HYUNDAI MOTOR CO LTD +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202380082622.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-10-18
Filing Date
2023-10-20
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing video encoding technology requires the consumption of additional bit resources when transmitting symbols of block vector difference (BVD) and motion vector difference (MVD), resulting in inefficient encoding.

Method used

The symbols of block vector difference and motion vector difference are derived by adaptively adjusting the template size, and the optimal template is selected using the principle of minimizing template matching cost, reducing the cost of symbol transmission bits.

Benefits of technology

Improves video encoding efficiency and quality, reduces bit consumption for symbol transmission, and improves encoding efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120303926A_ABST
    Figure CN120303926A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a video coding method and device using template adjustment for displacement symbol prediction. In this embodiment, an image decoding device decodes a vector predictor and a magnitude of a vector difference for a vector used for prediction of a current block. An image decoding apparatus generates a vector difference candidate by combining an available symbol combination of a vector difference and a magnitude of the vector difference. An image decoding apparatus generates a vector candidate by summing a vector predictor and a vector difference candidate, and then generates a reference block candidate corresponding to the vector candidate within a picture. The image decoding apparatus adjusts the template of the current block and the template of the reference block candidate based on the magnitude of the vector difference. The image decoding apparatus calculates a template matching cost between the adjusted template of the current block and the adjusted template of the reference block candidate, and then determines a prediction block of the current block by selecting a template that minimizes the template matching cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a video encoding method and apparatus using template adjustment for displacement symbol prediction. Background Art

[0002] The statements in this section merely provide background information related to the present disclosure and do not necessarily constitute prior art.

[0003] Since video data has a large amount of data compared to audio or still image data, video data requires a large amount of hardware resources (including memory) to store or transmit the video data without processing for compression.

[0004] Thus, an encoder is typically used to compress and store or transmit video data. A decoder receives the compressed video data, decompresses the received compressed video data, and plays the decompressed video data. Video compression techniques include H.264 / Advanced Video Coding (AVC), High Efficiency Video Coding (HEVC), and Versatile Video Coding (VVC), which have about 30% or more improved coding efficiency compared to HEVC.

[0005] However, as the image size, resolution, and frame rate gradually increase, the amount of data to be encoded also increases. Thus, there is a need to provide a new compression technique with higher coding efficiency and improved image enhancement effect than existing compression techniques.

[0006] The Intra Block Copy (IBC) technique uses a reference block searched within the current picture as a prediction block for the current block, and the inter prediction technique uses a reference block searched within a reference picture as a prediction block for the current block. The IBC technique and the inter prediction technique represent the displacement between the current block and the reference block using a block vector (BV) and a motion vector (MV), respectively. To effectively transmit the block vector or the motion vector, instead of transmitting the vector itself, a predictor is calculated, and the difference between the vector and the predictor is transmitted. For example, the encoder may generate a block vector predictor (BVP) for the block vector to be transmitted, calculate a block vector difference (BVD) by subtracting the block vector predictor from the block vector, and transmit the BVD to the decoder. Similarly, the encoder may generate a motion vector predictor (MVP) for the motion vector for transmission, calculate a motion vector difference (MVD) by subtracting the motion vector predictor from the motion vector, and may transmit the MVD to the decoder.

[0007] When sending non-zero values of BVD or MVD, the prior art sends flags indicating the signs of their respective components (i.e., the horizontal and vertical components). Since the sign indication flags are transmitted separately for the horizontal and vertical directions, a total of two bits of resources are consumed to transmit the signs of each BVD or MVD. Therefore, in order to improve video coding efficiency and enhance video quality, a method for efficiently deriving signs during the transmission of BVD and MVD is needed. Summary of the Invention

[0008] [Technical Problem]

[0009] The present disclosure attempts to provide a video coding method and apparatus for effectively deriving the signs of displacement information (such as block vector difference, motion vector difference, etc.) by adaptively adjusting the template size according to the magnitude of the displacement.

[0010] [Technical Solution]

[0011] At least one aspect of the present invention provides a method for reconstructing a current block by a video decoding device. The method includes decoding, from a bitstream, a vector predictor for a vector used to predict the current block. The method also includes decoding, from the bitstream, the magnitude of a vector difference. The method also includes generating a vector difference candidate by combining the magnitude of the vector difference with possible sign combinations of the vector difference. The method also includes generating a vector candidate by summing the vector predictor and the vector difference candidate, and then generating a reference block candidate corresponding to the vector candidate within a picture. The method also includes adjusting the template of the reference block candidate and the template of the current block based on the magnitude of the vector difference. The method also includes calculating a template matching cost between the adjusted template of the current block and the adjusted template of the reference block candidate. The method further includes determining a predicted block for the current block by selecting a template that minimizes the template matching cost.

[0012] Another aspect of the present invention provides a method for encoding a current block by a video encoding device. The method includes determining a vector for predicting the current block and determining a vector predictor of the current block. The method further includes generating a magnitude of a vector difference by subtracting the vector predictor from the vector. The method further includes generating a vector difference candidate by combining the magnitude of the vector difference with possible sign combinations of the vector difference. The method further includes generating a vector candidate by summing the vector predictor and the vector difference candidate, and then generating a reference block candidate corresponding to the vector candidate within a picture. The method further includes adjusting a template of the reference block candidate and a template of the current block based on the magnitude of the vector difference. The method further includes calculating a template matching cost between the adjusted template of the current block and the adjusted template of the reference block candidate. The method further includes determining a first prediction block of the current block by selecting a template that minimizes the template matching cost.

[0013] Another aspect of the present invention provides a computer-readable recording medium storing a bitstream generated by a video encoding method. The video encoding method includes determining a vector for predicting a current block and determining a vector predictor of the current block. The video encoding method further includes generating a magnitude of a vector difference by subtracting the vector predictor from the vector. The video encoding method further includes generating a vector difference candidate by combining the magnitude of the vector difference with possible sign combinations of the vector difference. The video encoding method further includes generating a vector candidate by summing the vector predictor and the vector difference candidate, and then generating a reference block candidate corresponding to the vector candidate within a picture. The video encoding method further includes adjusting a template of the reference block candidate and a template of the current block based on the magnitude of the vector difference. The video encoding method further includes calculating a template matching cost between the adjusted template of the current block and the adjusted template of the reference block candidate. The video encoding method further includes determining a prediction block of the current block by selecting a template that minimizes the template matching cost.

[0014] [Advantageous Effects]

[0015] As described above, the present disclosure provides a video encoding method and device that effectively derive the sign of displacement information (such as block vector difference, motion vector difference, etc.) by adaptively adjusting the template size according to the magnitude of displacement. Thereby, the video encoding method and device improve video encoding efficiency and video quality. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 A block diagram of a video encoding device for implementing the technology of the present invention.

[0017] Figure 2 A method of dividing a block using a quadtree plus binary tree plus ternary tree (QTBTTT) structure is shown.

[0018] Figure 3a and Figure 3b shows multiple intra prediction modes including a wide-angle intra prediction mode.

[0019] Figure 4 shows neighboring blocks of a current block.

[0020] Figure 5 is a block diagram of a video decoding device that can implement the technology of the present invention.

[0021] Figure 6 is a diagram showing a prediction technique for a block vector difference (BVD) sign based on template matching.

[0022] Figure 7 is a diagram showing a template.

[0023] Figure 8 is a diagram for explaining a prediction technique for the positive / negative sign of a motion vector difference (MVD) based on template matching.

[0024] Figure 9a and Figure 9b is a diagram showing a template matching cost as a function of the size of a block vector difference (BVD).

[0025] Figure 10 is a diagram showing a template matching cost as a function of template adjustment according to at least one embodiment of the present disclosure.

[0026] Figure 11a and Figure 11b is a diagram showing the setting of a template of a reference point according to at least one embodiment of the present disclosure.

[0027] Figure 12a and Figure 12b is a diagram showing an overlapping template.

[0028] Figure 13 is a diagram showing the adjustment of a region for calculating a template matching cost according to at least one embodiment of the present disclosure.

[0029] Figure 14 is a flowchart of a method for a video encoding device to encode a current block according to at least one embodiment of the present disclosure.

[0030] Figure 15 is a flowchart of a method for a video decoding device to reconstruct a current block according to at least one embodiment of the present invention. Detailed Description

[0031] In the following, some embodiments of the present disclosure will be described in detail with reference to the accompanying drawings. In the following description, the same reference numerals denote the same elements, although the elements are shown in different drawings. Further, in the following description of some embodiments, for the purposes of clarity and conciseness, detailed descriptions of related known components and functions may be omitted when it is considered that they would obscure the subject matter of the present disclosure.

[0032] Figure 1 is a block diagram of a video encoding device that can implement the technology of the present disclosure. In the following, with reference to Figure 1 the illustration of, the video encoding device and the components of the device will be described.

[0033] The encoding device may include an image splitter 110, a predictor 120, a subtractor 130, a transformer 140, a quantizer 145, a rearrangement unit 150, an entropy encoder 155, an inverse quantizer 160, an inverse transformer 165, an adder 170, a loop filter unit 180, and a memory 190.

[0034] Each component of the encoding device may be implemented as hardware or software or implemented as a combination of hardware and software. Further, the functions of each component may be implemented as software, and a microprocessor may also be implemented to execute the functions of the software corresponding to each component.

[0035] A video is composed of one or more sequences including a plurality of images. Each picture is divided into a plurality of regions, and encoding is performed on each region. For example, a picture is divided into one or more tiles and / or slices. Here, one or more tiles may be defined as a tile group. Each tile and / or slice is divided into one or more coding tree units (CTUs). Additionally, each CTU is divided into one or more coding units (CUs) through a tree structure. Information applied to each coding unit (CU) is encoded as the syntax of the CU, and information commonly applied to the CUs included in one CTU is encoded as the syntax of the CTU. Further, information commonly applied to all blocks in one strip is encoded as the syntax of the strip header, and information applied to all blocks constituting one or more pictures is encoded as a picture parameter set (PPS) or a picture header. Additionally, information commonly referred to by multiple pictures is encoded into a sequence parameter set (SPS). Further, information commonly referred to by one or more SPSs is encoded into a video parameter set (VPS). Further, information commonly applied to one tile or a tile group may also be encoded as the syntax of a tile or a tile group header. The syntax included in the SPS, PPS, slice header, tile or tile group header may be referred to as high-level syntax.

[0036] The image splitter 110 determines the size of the coding tree unit (CTU). Information regarding the size of the CTU (CTU size) is encoded as the syntax of the SPS or PPS and is delivered to the video decoding device.

[0037] The image splitter 110 splits each picture constituting the video into a plurality of coding tree units (CTUs) having a predetermined size, and then recursively splits the CTUs by using a tree structure. The leaf nodes in the tree structure become coding units (CUs), which are the basic units for coding.

[0038] The tree structure can be a quadtree (QT), in which a higher node (or parent node) is divided into four lower nodes (or child nodes) having the same size. The tree structure can also be a binary tree (BT), in which a higher node is divided into two lower nodes. The tree structure can also be a ternary tree (TT), in which a higher node is divided into three lower nodes at a ratio of 1:2:1. The tree structure can also be a structure in which two or more of the QT structure, BT structure, and TT structure are mixed. For example, a quadtree plus binary tree (QTBT) structure can be used or a quadtree plus binary tree ternary tree (QTBTTT) structure can be used. Here, a binary tree ternary tree (BTTT) is added to the tree structure to be called a multi-type tree (MTT).

[0039] Figure 2 is a diagram for describing a method of splitting a block by using a QTBTTT structure.

[0040] As Figure 2 shown, the CTU can first be divided into a QT structure. The quadtree splitting can be recursive until the size of the split block reaches the minimum block size (MinQTSize) of the leaf nodes allowed in the QT. A first flag (QT_split_flag) indicating whether each node of the QT structure is divided into four lower nodes is encoded by the entropy encoder 155 and sent to the video decoding device. When the leaf node of the QT is not larger than the maximum block size (MaxBTSize) of the root node allowed in the BT, the leaf node can be further divided into at least one of the BT structure or the TT structure. There can be multiple splitting directions in the BT structure and / or TT structure. For example, there can be two directions, that is, the direction in which the block of the corresponding node is horizontally split and the direction in which the block of the corresponding node is vertically split. As Figure 2 shown, when the MTT splitting starts, a second flag (mtt_split_flag) indicating whether the node is split, and a flag additionally indicating the splitting direction (vertical or horizontal) and / or a flag indicating the splitting type (binary or ternary) if the node is split are encoded by the entropy encoder 155 and signaled to the video decoding device.

[0041] Optionally, before encoding a first flag (QT_split_flag) indicating whether each node is split into four lower-layer nodes, a CU split flag (split_cu_flag) indicating whether a node is split may also be encoded. When the value of the CU split flag (split_cu_flag) indicates that each node is not split, the block of the corresponding node becomes a leaf node in the split tree structure and becomes a CU that is a basic unit of encoding. When the value of the CU split flag (split_cu_flag) indicates that each node is split, the video encoding device first starts encoding the first flag by the above scheme.

[0042] When QTBT is used as another embodiment of the tree structure, there may be two types, that is, a type in which the block of the corresponding node is horizontally split into two blocks of the same size (i.e., symmetric horizontal split) and a type in which the block of the corresponding node is vertically split into two blocks of the same size (i.e., symmetric vertical split). A split flag (split_flag) indicating whether each node of the BT structure is split into lower-layer blocks and split type information indicating the split type are encoded by the entropy encoder 155 and delivered to the video decoding device. At the same time, a type in which the block of the corresponding node is split into two asymmetric blocks may be additionally presented. The asymmetric form may include a form in which the block of the corresponding node is split into two rectangular blocks having a size ratio of 1:3, or may also include a form in which the block of the corresponding node is split in a diagonal direction.

[0043] The CU may have different sizes according to QTBT or QTBTTT splitting from the CTU. Hereinafter, the block corresponding to the CU to be encoded or decoded (i.e., the leaf node of QTBTTT) is referred to as the "current block". Due to the use of QTBTTT splitting, the shape of the current block may be a rectangular shape in addition to a square shape.

[0044] The predictor 120 predicts the current block to generate a prediction block. The predictor 120 includes an intra predictor 122 and an inter predictor 124.

[0045] Generally, each of the current blocks in a picture may be predictively coded. Generally, prediction of the current block may be performed by using an intra prediction technique (using data from the picture including the current block) or an inter prediction technique (using data from the picture compiled before the picture including the current block). Inter prediction includes both uni-directional prediction and bi-directional prediction.

[0046] The intra predictor 122 predicts the pixels in the current block by using pixels (reference pixels) located on neighbors of the current block in the current picture including the current block. According to the prediction direction, there are multiple intra prediction modes. For example, as Figure 3aAs shown, multiple intra prediction modes may include 2 non-directional modes including a planar mode and a DC mode, and may include 65 directional modes. Adjacent pixels and arithmetic expressions to be used are defined differently according to each prediction mode.

[0047] To perform effective directional prediction on a current block having a rectangular shape, directional modes shown by dotted arrows in Figure 3b (modes #67 to #80, intra prediction modes #-1 to #-14) may be additionally used. The directional modes may be referred to as "wide-angle intra prediction modes". In Figure 3b , the arrows indicate corresponding reference samples for prediction and do not represent the prediction direction. The prediction direction is opposite to the direction shown by the arrows. When the current block has a rectangular shape, the wide-angle intra prediction mode is a mode that performs prediction in a direction opposite to a specific directional mode without additional bit transmission. In this case, in the wide-angle intra prediction mode, some wide-angle intra prediction modes available for the current block may be determined by the ratio of the width and height of the current block having a rectangular shape. For example, when the current block has a rectangular shape with a height smaller than the width, wide-angle intra prediction modes (intra prediction modes #67 to #80) having an angle smaller than 45 degrees are available. When the current block has a rectangular shape with a width greater than the height, wide-angle intra prediction modes having an angle greater than -135 degrees may be used.

[0048] The intra predictor 122 may determine the intra prediction to be used for encoding the current block. In some embodiments, the intra predictor 122 may encode the current block by using multiple intra prediction modes and may also select an appropriate intra prediction mode to be used from test modes. For example, the intra predictor 122 may calculate rate-distortion values by using rate-distortion analysis for multiple tested intra prediction modes and may also select an intra prediction mode having the best rate-distortion characteristics in the test modes.

[0049] The intra predictor 122 selects one intra prediction mode from multiple intra prediction modes and predicts the current block body by using adjacent pixels (reference pixels) and an arithmetic expression determined according to the selected intra prediction mode. Information about the selected intra prediction mode is encoded by the entropy encoder 155 and delivered to the video decoding device.

[0050] The inter - frame predictor 124 generates a prediction block for a current block by using motion - compensation processing. The inter - frame predictor 124 searches for a block most similar to the current block in a reference picture that was encoded and decoded earlier than the current picture, and generates a prediction block for the current block by using the searched - for block. Additionally, a motion vector (MV) is generated, which corresponds to the displacement between the current block in the current picture and the prediction block in the reference picture. Generally, motion estimation is performed on the luminance component, and the motion vector calculated based on the luminance component is used for both the luminance component and the chrominance components. Motion information including information about the reference picture and information about the motion vector used to predict the current block is encoded by the entropy encoder 155 and delivered to the video decoding device.

[0051] The inter - frame predictor 124 may also perform interpolation on the reference picture or reference block in order to increase the prediction accuracy. In other words, sub - sampling between two consecutive integer samples is interpolated by applying filter coefficients to a plurality of consecutive integer samples including the two integer samples. When performing the process of searching for a block most similar to the current block on the interpolated reference picture, the motion vector representation may be in decimal - unit precision rather than integer - sampling - unit precision. The precision or resolution of the motion vector may be set differently for each target region to be encoded (e.g., units such as slices, tiles, CTUs, CUs, etc.). When this adaptive motion - vector resolution (AMVR) is applied, information about the motion - vector resolution to be applied to each target region should be signaled. For example, when the target region is a CU, information about the motion - vector resolution applied to each CU is signaled. The information about the motion - vector resolution may be information representing the precision of the motion - vector difference that will be described below.

[0052] Meanwhile, the inter-frame predictor 124 can perform inter-frame prediction by using bidirectional prediction. In the case of bidirectional prediction, two reference pictures and two motion vectors representing the block positions in each reference picture that are most similar to the current block are used. The inter-frame predictor 124 selects a first reference picture and a second reference picture from reference picture list 0 (RefPicList0) and reference picture list 1 (RefPicList1), respectively. The inter-frame predictor 124 also searches for the block in the corresponding reference picture that is most similar to the current block to generate a first reference block and a second reference block. Additionally, a predicted block for the current block is generated by averaging or weighted averaging the first reference block and the second reference block. Additionally, motion information including information about the two reference pictures used for predicting the current block and information about the two motion vectors is delivered to the entropy encoder 155. Here, reference picture list 0 can be composed of pictures among the pre-reconstructed pictures that are before the current picture in display order, and reference picture list 1 can be composed of pictures among the pre-reconstructed pictures that are after the current picture in display order. However, although not particularly limited thereto, pre-reconstructed pictures that are after the current picture in display order can be additionally included in reference picture list 0. Conversely, pre-reconstructed pictures that are before the current picture can also be additionally included in reference picture list 1.

[0053] To minimize the number of bits consumed for encoding motion information, various methods can be used.

[0054] For example, when the reference picture and motion vector of the current block are the same as those of an adjacent block, information identifying the adjacent block is encoded to deliver the motion information of the current block to the video decoding device. This method is called the merge mode.

[0055] In the merge mode, the inter-frame predictor 124 selects a predetermined number of merge candidates (hereinafter referred to as "merge candidates") from the adjacent blocks of the current block.

[0056] As shown in Figure 4 , all or some of the left block A0, lower left block A1, upper block B0, upper right block B1, and upper left block B2 adjacent to the current block in the current picture can be used as adjacent blocks for deriving merge candidates. Further, blocks other than the current picture in which the current block is located within the reference picture (which can be the same as or different from the reference picture used for predicting the current block) can also be used as merge candidates. For example, blocks located at the same position as the current block within the reference picture or blocks adjacent to the blocks located at the same position can additionally be used as merge candidates. If the number of merge candidates selected by the method described above is less than the preset number, zero vectors are added to the merge candidates.

[0057] The inter-frame predictor 124 configures a merge list including a predetermined number of merge candidates by using neighboring blocks. A merge candidate to be used as motion information of a current block is selected from the merge candidates included in the merge list, and merge index information for identifying the selected candidate is generated. The generated merge index information is encoded by the entropy encoder 155 and delivered to the video decoding device.

[0058] The merge skip mode is a special case of the merge mode. After quantization, when all transform coefficients for entropy coding are close to zero, only the neighboring block selection information is transmitted without transmitting the residual signal. By using the merge skip mode, relatively high coding efficiency can be achieved for images with slight motion, still images, screen content images, and so on.

[0059] Hereinafter, the merge mode and the merge skip mode are collectively referred to as the merge / skip mode.

[0060] Another method for encoding motion information is the advanced motion vector prediction (AMVP) mode.

[0061] In the AMVP mode, the inter-frame predictor 124 derives motion vector predictor candidates for the motion vector of the current block by using neighboring blocks of the current block. As neighboring blocks for deriving the motion vector predictor candidates, all or some of the left block A0, the lower left block A1, the upper block B0, the upper right block B1, and the upper left block B2 adjacent to the current block in the current picture shown in Figure 4 can be used. In addition, blocks other than the current picture in which the current block is located within the reference picture (which may be the same as or different from the reference picture for predicting the current block) can also be used as neighboring blocks for deriving the motion vector predictor candidates. For example, a block co-located with the current block within the reference picture or a block adjacent to the co-located block can be used. If the number of motion vector candidates selected by the above method is less than a preset number, a zero vector is added to the motion vector candidates.

[0062] The inter-frame predictor 124 derives motion vector predictor candidates by using the motion vectors of neighboring blocks, and determines a motion vector predictor for the motion vector of the current block by using the motion vector predictor candidates. In addition, a motion vector difference is calculated by subtracting the motion vector predictor from the motion vector of the current block.

[0063] A motion vector predictor can be obtained by applying a predefined function (e.g., central value and average value calculation, etc.) to motion vector predictor candidates. In this case, the video decoding device also knows the predefined function. In addition, since the neighboring blocks used to derive the motion vector predictor candidates are blocks for which encoding and decoding have been completed, the video decoding device may also already know the motion vectors of the neighboring blocks. Therefore, the video encoding device does not need to encode the information for identifying the motion vector predictor candidates. Thus, in this case, the information about the motion vector difference and the information about the reference picture used to predict the current block are encoded.

[0064] Meanwhile, the motion vector predictor can also be determined by a scheme of selecting any one of the motion vector predictor candidates. In this case, the information for identifying the selected motion vector predictor candidate is additionally encoded jointly with the information about the motion vector difference and the information about the reference picture used to predict the current block.

[0065] The subtractor 130 generates a residual block by subtracting the predicted block generated by the intra predictor 122 or the inter predictor 124 from the current block.

[0066] The transformer 140 converts the residual signal in the residual block having pixel values in the spatial domain into transform coefficients in the frequency domain. The transformer 140 can transform the residual signal in the residual block by using the total size of the residual block as the transform unit, or can also divide the residual block into a plurality of sub - blocks and can perform the transformation by using the sub - blocks as the transform unit. Optionally, the residual block is divided into two sub - blocks, namely a transform region and a non - transform region, to transform the residual signal by using only the transform region sub - block as the transform unit. Here, the transform region sub - block can be one of two rectangular blocks having a size ratio of 1:1 based on the horizontal axis (or vertical axis). In this case, a flag (cu_sbt_flag) indicates that only the sub - block is transformed, and the directionality (vertical / horizontal) information (cu_sbt_horizontal_flag) and / or the position information (cu_sbt_pos_flag) are encoded by the entropy encoder 155 and signaled to the video decoding device. In addition, the size of the transform region sub - block can have a size ratio of 1:3 based on the horizontal axis (or vertical axis). In this case, a flag (cu_sbt_quad_flag) for the corresponding segmentation is additionally encoded by the entropy encoder 155 and signaled to the video decoding device.

[0067] Meanwhile, the transformer 140 can perform transformations on the residual blocks separately in the horizontal and vertical directions. For the transformations, different types of transformation functions or transformation matrices can be used. For example, a pair of transformation functions for horizontal and vertical transformations can be defined as a multiple transformation set (MTS). The transformer 140 can select a pair of transformation functions with the highest transformation efficiency in the MTS and can transform the residual blocks in each of the horizontal and vertical directions. The information (mts_idx) of the pair of transformation functions in the MTS is encoded by the entropy encoder 155 and signaled to the video decoding device.

[0068] The quantizer 145 quantizes the transform coefficients output from the transformer 140 using quantization parameters and outputs the quantized transform coefficients to the entropy encoder 155. The quantizer 145 can also immediately quantize the relevant residual blocks without performing a transformation for any block or frame. The quantizer 145 can also apply different quantization coefficients (scaling values) according to the positions of the transform coefficients in the transform block. The quantization matrix applied to the quantized transform coefficients arranged in a 2D layout can be encoded and signaled to the video decoding device.

[0069] The rearrangement unit 150 can perform rearrangement of the coefficient values for the quantized residual values.

[0070] The rearrangement unit 150 can change a 2D coefficient array into a 1D coefficient sequence by using coefficient scanning. For example, the rearrangement unit 150 can scan the DC coefficient into high-frequency domain coefficients by using zigzag scanning or diagonal scanning to output an ID coefficient sequence. According to the size of the transform unit and the intra prediction mode, vertical scanning that scans the 2D coefficient array along the column direction and horizontal scanning that scans the 2D block type coefficients along the row direction can also be used instead of zigzag scanning. In other words, according to the size of the transform unit and the intra prediction mode, the scanning method to be used can be determined among zigzag scanning, diagonal scanning, vertical scanning, and horizontal scanning.

[0071] The entropy encoder 155 generates a bitstream by encoding the sequence of ID quantized transform coefficients output from the rearrangement unit 150 using various coding schemes including context-based adaptive binary arithmetic coding (CABAC), exponential Golomb, etc.

[0072] In addition, the entropy encoder 155 encodes information related to block partitioning (such as CTU size, CTU partition flag, QT partition flag, MTT partition type, MTT partition direction, etc.) to allow the video decoding device to equivalently partition blocks to the video encoding device. In addition, the entropy encoder 155 encodes information regarding the prediction type indicating whether the current block is encoded by intra prediction or inter prediction. The entropy encoder 155 encodes intra prediction information (i.e., information regarding intra prediction mode) or inter prediction information (in the case of merge mode, merge index, and in the case of AMVP mode, information regarding reference picture index and motion vector difference) according to the prediction type. In addition, the entropy encoder 155 encodes information related to quantization (i.e., information regarding quantization parameter and information regarding quantization matrix).

[0073] The inverse quantizer 160 dequantizes the quantized transform coefficients output from the quantizer 145 to generate transform coefficients. The inverse transformer 165 transforms the transform coefficients output from the inverse quantizer 160 from the frequency domain to the spatial domain to reconstruct the residual block.

[0074] The adder 170 adds the reconstructed residual block and the prediction block generated by the predictor 120 to reconstruct the current block. When performing intra prediction on the next-order block, the pixels in the reconstructed current block can be used as reference pixels.

[0075] The loop filter unit 180 performs filtering on the reconstructed pixels in order to reduce block artifacts, ringing artifacts, blur artifacts, etc. that occur due to block-based prediction and transform / quantization. The loop filter unit 180 as a loop filter may include all or some of a deblocking filter 182, a sample adaptive offset (SAO) filter 184, and an adaptive loop filter (ALF) 186.

[0076] The deblocking filter 182 filters the boundaries between the reconstructed blocks to remove block effects that occur due to block-based encoding / decoding, and the SAO filter 184 and the ALF 186 perform additional filtering on the filtered video for deblocking. The SAO filter 184 and the ALF 186 are filters for compensating the difference between the reconstructed pixels and the original pixels, which occurs due to lossy decoding. The SAO filter 184 applies an offset as a CTU unit to enhance subjective image quality and encoding efficiency. On the other hand, the ALF 186 performs block-based filtering and applies different filters by dividing the degree of the boundary and variation amount of the corresponding block to compensate for distortion. Information regarding the filter coefficients to be used for the ALF can be encoded and signaled to the video decoding device.

[0077] The reconstructed blocks filtered by the deblocking filter 182, the SAO filter 184, and the ALF 186 are stored in the memory 190. When all the blocks in a picture are reconstructed, the reconstructed picture can be used as a reference picture for inter prediction of blocks within a picture to be encoded later.

[0078] The video encoding device may store the bitstream of the encoded video data in a non - transitory storage medium or transmit the bitstream to a video decoding device via a communication network.

[0079] Figure 5 is a functional block diagram of a video decoding device that can implement the technology of the present disclosure. Hereinafter, with reference to Figure 5 , the video decoding device and components of the device will be described.

[0080] The video decoding device may include an entropy decoder 510, a rearrangement unit 515, an inverse quantizer 520, an inverse transform unit 530, a predictor 540, an adder 550, a loop filter unit 560, and a memory 570.

[0081] Similar to Figure 1 the video encoding device, each component of the video decoding device may be implemented as hardware or software or implemented as a combination of hardware and software. In addition, the function of each component may be implemented as software, and a microprocessor may also be implemented to execute the functions of the software corresponding to each component.

[0082] The entropy decoder 510 extracts information related to block partitioning and extracts prediction information and information about the residual signal required to reconstruct the current block by decoding the bitstream generated by the video encoding device to determine the current block to be decoded.

[0083] The entropy decoder 510 determines the size of the CTU by extracting information about the CTU size from the sequence parameter set (SPS) or the picture parameter set (PPS), and divides the picture into CTUs with the determined size. Additionally, the CTU is determined as the highest layer of the tree structure, i.e., the root node, and the partitioning information of the CTU can be extracted to partition the CTU using the tree structure.

[0084] For example, when partitioning the CTU using the QTBTTT structure, first, the first flag (QT_split_flag) related to the partitioning of the QT is extracted to divide each node into four lower - layer nodes. In addition, for the nodes corresponding to the leaf nodes of the QT, the second flag (mtt_split_flag) related to the partitioning of the MTT, the partitioning direction (vertical / horizontal), and / or the partitioning type (binary / ternary) are extracted to divide the corresponding leaf nodes into the MTT structure. As a result, each node below the leaf nodes of the QT is recursively divided into the BT or TT structure.

[0085] As another embodiment, when splitting a CTU by using a QTBTTT structure, a CU split flag (split_cu_flag) indicating whether a CU is split is extracted. When the corresponding block is split, a first flag (QT_split_flag) may also be extracted. During the splitting process, for each node, zero or more recursive MTT splits may occur after zero or more recursive QT splits. For example, for a CTU, an MTT split may occur immediately, or conversely, only multiple QT splits may occur.

[0086] For another example, when splitting a CTU by using a QTBT structure, a first flag (QT_split_flag) related to the splitting of QT is extracted to split each node into four lower-layer nodes. In addition, a split flag (split_flag) indicating whether the node corresponding to the leaf node of QT is further split into BT and split direction information are extracted.

[0087] Meanwhile, when the entropy decoder 510 determines the current block to be decoded by using tree-structured splitting, the entropy decoder 510 extracts information about a prediction type indicating whether the current block is intra-predicted or inter-predicted. When the prediction type information indicates intra-prediction, the entropy decoder 510 extracts a syntax element for intra-prediction information (intra-prediction mode) for the current block. When the prediction type information indicates inter-prediction, the entropy decoder 510 extracts information representing a syntax element for inter-prediction information (i.e., a motion vector and a reference picture to which the motion vector refers).

[0088] In addition, the entropy decoder 510 extracts quantization-related information and extracts information about the quantized transform coefficients of the current block as information about the residual signal.

[0089] The rearrangement unit 515 may again change a sequence of 1D quantized transform coefficients entropy-decoded by the entropy decoder 510 into a 2D coefficient array (i.e., a block) in an order opposite to the coefficient scan order performed by the video coding device.

[0090] The inverse quantizer 520 dequantizes the quantized transform coefficients and dequantizes the quantized transform coefficients by using a quantization parameter. The inverse quantizer 520 may also apply different quantization coefficients (scaling values) to the quantized transform coefficients arranged in a 2D layout. The inverse quantizer 520 may perform inverse quantization by applying a matrix (scaling value) of quantization coefficients from the video coding device to the 2D array of the quantized transform coefficients.

[0091] The inverse transformer 530 generates a residual block for the current block by reconstructing a residual signal by inverse-transforming the dequantized transform coefficients from the frequency domain to the spatial domain.

[0092] In addition, when the inverse transform unit 530 inverse-transforms a partial region (sub-block) of a transform block, the inverse transform unit 530 extracts a flag (cu_sbt_flag) indicating that only the sub-block of the transform block is transformed, direction (vertical / horizontal) information (cu_sbt_horizontal_flag) of the sub-block, and / or position information (cu_sbt_pos_flag) of the sub-block. The inverse transform unit 530 also inverse-transforms the transform coefficients of the corresponding sub-block from the frequency domain to the spatial domain to reconstruct the residual signal and fills the un-inverse-transformed region with a value of "0" as the residual signal to generate a final residual block for the current block.

[0093] In addition, when applying MTS, the inverse transform unit 530 determines a transform function or transform matrix applied in each of the horizontal and vertical directions by using MTS information (mts_idx) signaled from the video coding device. The inverse transform unit 530 also inverse-transforms the transform coefficients in the transform block in the horizontal and vertical directions by using the determined transform function.

[0094] The predictor 540 may include an intra predictor 542 and an inter predictor 544. The intra predictor 542 is activated when the prediction type of the current block is intra prediction, and the inter predictor 544 is activated when the prediction type of the current block is inter prediction.

[0095] The intra predictor 542 determines the intra prediction mode of the current block among a plurality of intra prediction modes according to the syntax element of the intra prediction mode extracted from the entropy decoder 510. The intra predictor 542 also predicts the current block by using neighboring reference pixels of the current block according to the intra prediction mode.

[0096] The inter predictor 544 determines a motion vector of the current block and a reference picture for motion vector reference by using the syntax element for the inter prediction mode extracted from the entropy decoder 510.

[0097] The adder 550 reconstructs the current block by adding the residual block output from the inverse transform unit 530 and the prediction block output from the inter predictor 544 or the intra predictor 542. When performing intra prediction on a block to be decoded later, the pixels in the reconstructed current block are used as reference pixels.

[0098] The loop filter unit 560 serving as a loop filter may include a deblocking filter 562, a SAO filter 564, and an ALF 566. The deblocking filter 562 performs deblocking filtering on the boundaries between the reconstructed blocks to remove the blocking artifacts that occur due to block-based decoding. The SAO filter 564 and the ALF 566 perform additional filtering on the reconstructed blocks after deblocking filtering to compensate for the difference between the reconstructed pixels and the original pixels, which occurs due to lossy coding. The filter coefficients of the ALF are determined by using the information about the filter coefficients decoded from the bitstream.

[0099] The reconstructed blocks filtered by the deblocking filter 562, the SAO filter 564, and the ALF 566 are stored in the memory 570. When all the blocks in a picture are reconstructed, the reconstructed picture can be used as a reference picture for inter prediction of the blocks within the pictures to be encoded later.

[0100] In some embodiments, the present invention relates to encoding and decoding video images as described above. More specifically, the present disclosure provides a video encoding method and apparatus that effectively derives the sign of displacement information (such as block vector difference, motion vector difference, etc.) by adaptively adjusting the template size according to the magnitude of the displacement.

[0101] The following embodiments may be executed by the predictor 120 in a video encoding device. The following embodiments may also be executed by the predictor 540 in a video decoding device.

[0102] When encoding a current block, the video encoding device may generate signaling information associated with the present embodiment in terms of optimizing rate distortion. The video encoding device may encode the signaling information using the entropy encoder 155 and send the encoded signaling information to the video decoding device. The video decoding device may decode the signaling information associated with the decoding of the current block from the bitstream using the entropy decoder 510.

[0103] In the following description, the term "target block" may be used interchangeably with the current block or the coding unit (CU), or may refer to a certain region of the coding unit.

[0104] In addition, a value of a flag being true indicates when the flag is set to 1. Additionally, a value of a flag being false indicates when the flag is set to 0.

[0105] I. Predicting the Sign of BVD or MVD

[0106] Hereinafter, BVD (block vector difference) is used in the IBC AMVP mode (advanced motion vector prediction mode), and MVD (motion vector difference) is used in the AMVP mode of inter prediction.

[0107] As described above, in order to effectively transmit a block vector or a motion vector, instead of transmitting the vector as a whole, a prediction factor is calculated, and then the difference between the vector and the prediction factor is transmitted. For example, a video encoding device may generate a block vector prediction factor (BVP) for a block vector to be transmitted, calculate a block vector difference (BVD) by subtracting the block vector prediction factor from the block vector, and then transmit the BVD to a video decoding device. In the IBC AMVP mode, the video decoding device may generate a candidate list and then generate the BVP from the candidate list.

[0108] Similarly, a video encoding device may generate a motion vector prediction factor (MVP) for a motion vector to be transmitted, calculate a motion vector difference (MVD) by subtracting the motion vector prediction factor from the motion vector, and transmit the MVD to a video decoding device. In the AMVP mode of inter prediction, the video decoding device may generate a candidate list and then generate the MVP from the candidate list.

[0109] In the prior art, the BVD or MVD may be signaled / parsed according to the syntax shown in Table 1. If the value of the BVD or MVD is non-zero, then flags indicating the signs of each of the horizontal and vertical components are transmitted.

[0110] [Table 1]

[0111]

[0112] Table 1 is the syntax for motion vectors but is equally applicable to block vectors. In Table 1, mvd_sign_flag[0] indicates the horizontal (or vertical) component sign, and mvd_sign_flag[1] indicates the vertical (or horizontal) component sign. A total of two resource bits are consumed to transmit the signs of the BVD or MVD.

[0113] Hereinafter, for ease of description, the present disclosure regarding the BVD is described. However, the present disclosure is equally applicable to the MVD. Therefore, although for ease of description, the following description only relates to the BVD, however, this description should be interpreted as generally applicable to the BVD or MVD.

[0114] Figure 6 A diagram for showing a prediction technique for the sign of the BVD based on template matching.

[0115] In order to eliminate such a bit cost for transmitting the sign of the BVD, there is a technique for predicting the sign of the BVD by using a template. As Figure 6As shown, the present technology compares the template matching costs calculated from multiple BVD candidates based on selectable symbols and selects the BVD with the minimum cost. In this case, the template matching cost represents the difference between the template of the current block and the templates of the reference block candidates. The template matching cost can be calculated based on a loss function (such as the sum of absolute differences (SAD)).

[0116] Figure 7 is a diagram showing the template.

[0117] As Figure 7 shown, the template represents a part of the region adjacent to the upper side and the left side of the current block. The template size and shape represent the reconstructed adjacent region of pixels with a height and width of 'd' adjacent to the block. The template size and shape can be preset in advance. In such a case, the template of the current block can represent a part of the region adjacent to the upper side and the left side of the current block, and the template of the reference block candidate can represent a part of the region adjacent to the upper side and the left side of the reference block candidate.

[0118] Similar to the BVD symbol prediction, a prediction technique for the symbols of the MVD can be implemented, as Figure 8 shown.

[0119] Figure 9a and Figure 9b is a diagram showing the template matching cost as a function of the magnitude of the BVD.

[0120] As described above, the technique for deriving the symbol of the BVD compares the template matching costs obtained from multiple BVD candidates by using template comparison (as in the example of Figure 6 ) and selects the BVD with the minimum cost. In this case, for the symbol information of the BVD, there are at most four candidates (i.e., (+, +), (+, -), (-, -), and (-, -)), so there can be four BVD candidate items of the same number. The template matching cost represents the difference between the template of the current block and the template indicated by the BVD candidate. By using the prior art, as the magnitude of the BVD decreases, the spatial distance between the templates with different BVD symbols decreases, which cannot significantly distinguish the templates, as shown in the examples of Figure 9a and Figure 9b . In addition, when the magnitude of the BVD is small, the pixels within the template may substantially overlap. Therefore, with the prior art, as the magnitude of the BVD becomes smaller, it is difficult to accurately obtain the final block vector by using the template matching cost comparison.

[0121] The following embodiments are described with respect to a video decoding device, but can be implemented in the same or similar manner by a video encoding device.

[0122] II. Embodiments of the Present Disclosure

[0123] Figure 10 It is a diagram showing the template matching cost of a function adjusted as a template according to at least one embodiment of the present disclosure.

[0124] As Figure 9b shown, the problems of the prior art can be solved by adjusting the size of the template according to the amplitude of the BVD.

[0125] The implementation related to adjusting the template size is described below.

[0126] <Implementation 1> A method for adjusting the template size based on the amplitude of the BVD

[0127] Figure 11a And Figure 11b is a diagram showing the setting of the template of the reference point according to at least one embodiment of the present disclosure.

[0128] In this embodiment, the video decoding device may divide the template into an upper template and a left template for adjusting the size of the template, and then set different reference points for the upper template and the left template. As Figure 11a shown, when setting the template by using the reference point formed by the pixels at the upper left of the current block, the upper template may be set to a predetermined length extending from the left end of the current block, and the left template may be set to a predetermined length extending from the upper end of the current block. As Figure 11b shown, when setting the template by using the reference point formed by the center of the current block, the horizontal center of the upper template may be aligned with the current block and set to a predetermined extension length, and the vertical center of the left template may be aligned with the current block and set to a predetermined extension length. Here, the horizontal center refers to the vertical line that balances the block from one side to the other, and the vertical center refers to the horizontal line that balances the block from top to bottom.

[0129] Although the following method describes the template with the upper left side of the current block as the reference point, the same method can be applied to the template with the center of the current block as the reference point.

[0130] In Figure 11a and Figure 11b example, L A and D A respectively represent the length and thickness of the upper template above the current block, and L L and D L respectively represent the length and thickness of the left template. The length of the upper template and the length of the left template can be set according to the width and height of the current block respectively, as shown in Equation 1.

[0131] [Equation 1]

[0132] L A = α A·W

[0133] L L = α L ·H

[0134] Thus, in this implementation, the video decoding device can adjust the template size by changing the template parameter α A , D A , α L , D L . In the following description, "dX" defines the size of the horizontal component of the BVD, and 'dY' defines the size of its vertical component. The video decoding device can adjust the template size by (Implementation 1-1) using a representative value or by (Implementation 1-2) using each component. Alternatively, (Implementation 1-3) the video decoding device can adjust the calculation area of the template matching cost by using each component.

[0135] <Implementation 1-1> Adjusting the Template Size Using a Representative Value

[0136] In this implementation, the video decoding device sets a representative value R, and then uses the representative value to set the size of the upper template to be equal to the size of the left template. For example, α = α A = α L and D = D A = D L . According to this embodiment, the video decoding device can set the representative value R according to the sizes of dX and dY. If the dX and dY values are non-zero, R can be set by using one of the following methods.

[0137] Set R to the smaller of the dX and dY values, R = min(dX, dY)

[0138] Set R to the larger of the dX and dY values, R = max(dX, dY)

[0139] Set R to the average of the dX and dY values, R = (dX + dY) / 2

[0140] Set R to the sum of the dX and dY values, R = dX + dY

[0141] If one of the dX and dY values is zero, R can be set to the value of the non-zero component. If both the dX and dY values are zero, no symbol information is required. Therefore, this case is not considered.

[0142] When setting the value of R, the video decoding device can adjust the template size by selecting at least one of the following.

[0143] First, the video decoding device can adjust the size of the template by adjusting the α value. The video decoding device can apply the value of R to predefined interval criteria r1, r2, ..., r n to set the value of α. In this case, the number n of interval criteria can vary according to the implementation, such as 1, 2, 3, ..., etc. Therefore, the interval criteria r1, r2, ..., r n can also vary according to the implementation. In addition, the value of α adjusted according to the interval can also change according to the implementation. For example, if the number n of interval criteria = 1 and the corresponding set interval criterion r1 = 4, then α can be adjusted as shown in Equation 2.

[0144] [Equation 2]

[0145]

[0146] Next, in the same way as adjusting the α value, the video decoding device can adjust the size of the template by adjusting the D value. For example, if the number n of interval criteria = 2 and the corresponding set interval criteria r1 = 4 and r2 = 16, then D can be adjusted as shown in Equation 3.

[0147] [Equation 3]

[0148]

[0149] For example, if the representative value R is set to the smaller of the dX and dY values, and both the α and D values are adjusted, the template can be resized according to the sizes of the horizontal and vertical components of the BVD, as shown in Table 2.

[0150] [Table 2]

[0151] dX dY R α d 3 2 2 2 4 1 10 1 2 4 8 12 8 1 3 16 0 16 1 2

[0152] Here, it is assumed that the intervals for adjusting the α and D values depend on the embodiments described in Equation 2 and Equation 3.

[0153] <Implementation 1-2> Template adjustment using each component

[0154] In this implementation, the video decoding device resizes the upper template based on the horizontal component of the BVD and resizes the left template based on the vertical component of the BVD. Optionally, the video decoding device can resize the upper template based on the vertical component of the BVD and resize the left template based on the horizontal component of the BVD.

[0155] Similar to the method of adjusting the α and D values by using intervals in Embodiment 1-1, the video decoding device can adjust the upper template parameters α A and D A, and adjust the left template parameter α according to dY L and D L . For example, the template parameter α A , D A , α L and D L The interval-related settings of can be defined as shown in Equation 4.

[0156] [Equation 4]

[0157]

[0158] In this case, the size of the template can be adjusted according to the horizontal and vertical component sizes of the BVD, as shown in Table 3.

[0159] [Table 3]

[0160] dX dY <![CDATA[α A > <![CDATA[D A > <![CDATA[α L > <![CDATA[D L > 3 2 2 4 2 4 1 10 2 4 1 2 8 12 1 2 1 2 16 0 1 2 2 4

[0161] This implementation can be achieved as follows. The video decoding device can adjust the upper template parameter α according to dY A and D A , and adjust the left template parameter α according to dX L and D L .

[0162] <Implementation 1-3> Adjustment of the calculation area of the template matching cost using each component value Figure 12a and Figure 12b is a diagram showing the overlapping templates.

[0163] When the amplitude of the BVD is small, the templates can overlap each other in the horizontal direction and / or the vertical direction, as shown in Figure 12a and Figure 12b . When the templates overlap, many unnecessary calculations may be performed during the process of calculating the template matching cost. In this implementation, the video decoding device adjusts the area in the template used to calculate the template matching cost to eliminate these unnecessary calculations.

[0164] Figure 13 is a diagram showing the adjustment of the area for calculating the template matching cost according to at least one embodiment of the present disclosure.

[0165] In one embodiment, the video decoding device can adjust the area for calculating the template matching cost based on the components of the BVD as follows. For example, the area for calculating the template matching cost can be adjusted to two shortened areas with a size of 2·dX measured from the left and right sides of the upper template, respectively, and two shortened areas with a size of 2·dY measured from the upper and lower sides of the left template, respectively.

[0166] As Figure 13As shown, the video decoding device may adjust the horizontal calculation area of the upper template according to the value of dX and adjust the vertical calculation area of the left template according to the value of dY. If the value of dX or dY is zero, there is no need to adjust the calculation area in the relevant template. If the value of dX is zero, the entire upper template can be used, and if the value of dY is zero, the entire left template can be used. In addition, if the value of 2·dX or 2·dY is not an integer, the video decoding device may convert 2·dX or 2·dY to an integer through operations such as rounding, raising, lowering, discarding, etc., and then adjust the area by using the converted value of 2·dX or 2·dY.

[0167] <Implementation 2> Determine whether to apply the method of the present disclosure

[0168] The present disclosure may (Implementation 2-1) always be applied to replace the process of encoding / decoding the symbols of BVD, or the present disclosure may (Implementation 2-2) be selectively applied according to the signaling of the flag.

[0169] <Implementation 2-1> Always apply the present disclosure

[0170] In this implementation, the video decoding device may always apply the adjustment of the template to derive the BVD symbol as an alternative to decoding the BVD symbol. Here, the adjustment of the template may be (Implementation 1-1 or Implementation 1-2) the adjustment of the template size or (Implementation 1-3) the adjustment of the calculation area of the template matching cost.

[0171] When the template adjustment is always applied, the syntax may change as shown in Table 4.

[0172] [Table 4]

[0173]

[0174]

[0175] Compared with Table 1 according to the prior art, Table 2 according to this implementation may omit the information of transmitting mvd_sign_flag[0] and mvd_sign_flag[1].

[0176] <Implementation 2-2> Apply the present disclosure according to the transmitted flag.

[0177] In this implementation, the video decoding device selectively applies the adjustment of the template to derive the BVD symbol. If at least one of the horizontal and vertical components of the BVD is non-zero, the video encoding device may send a flag indicating whether to adjust the template, the template_adjustment_flag (hereinafter, "template adjustment flag") to the video decoding device. Here, the adjustment of the template may be (Implementation 1-1 or Implementation 1-2) the adjustment of the template size or (Implementation 1-3) the adjustment of the template matching cost calculation area.

[0178] If the template adjustment flag is true, the video decoding device adjusts the template to derive the symbol of the BVD. As Figure 6 shown, if the template adjustment flag is false, the video decoding device may use a template of a preset size to derive the symbol of the BVD. The syntax according to this embodiment may be changed as shown in Table 5.

[0179] [Table 5]

[0180]

[0181]

[0182] In Table 5, if both the horizontal and vertical components of the BVD are zero (i.e., if the amplitude of the BVD is zero), there is no need to derive the symbol of the BVD. Therefore, the template adjustment flag is not signaled.

[0183] Now referring to Figure 14 and Figure 15 , a template adjustment method for obtaining the symbol of the BVD or MVD is described.

[0184] Figure 14 is a flowchart of a method for encoding a current block by a video encoding device according to at least one embodiment of the present invention.

[0185] The video encoding device determines a vector for predicting the current block (S1400). Here, the vector may be a block vector or a motion vector. The video encoding device may determine the vector in terms of rate-distortion optimization.

[0186] The video encoding device determines a vector predictor for the current block (S1402).

[0187] For example, after generating the candidate list, in terms of rate-distortion optimization, the video encoding device may determine the vector predictor from the candidate list.

[0188] The video encoding device may subtract the vector predictor from the vector to generate the magnitude of the vector difference (S1404).

[0189] The video encoding device may combine the magnitude of the vector difference and the possible symbol combinations of the vector difference to generate a vector difference candidate (S1406).

[0190] When generating a vector candidate by summing a vector predictor and a vector difference candidate, the video encoding device generates a reference block candidate corresponding to the vector candidate in the picture (S1408).

[0191] At this time, if the vector is a block vector, the picture may be the current picture including the current block, and if the vector is a motion vector, the picture may be a reference picture of the current block.

[0192] The video encoding device adjusts the template of the reference block candidate and the template of the current block based on the magnitude of the vector difference (S1410).

[0193] Here, the adjustment of the template may be an adjustment of the template size (Implementation Method 1-1 or Implementation Method 1-2) or an adjustment of the calculation area of the template matching cost (Implementation Method 1-3).

[0194] The video encoding device calculates the template matching cost between the adjusted template of the current block and the adjusted template of the reference block candidate (S1412).

[0195] The video encoding device determines the predicted block of the current block based on the template matching cost. By selecting the template that minimizes the template matching cost and selecting the vector difference candidate and the reference block candidate corresponding to the selected template, the video encoding device can determine the predicted block of the current block.

[0196] The video encoding device encodes the vector predictor and the magnitude of the vector difference. For example, the video encoding device may encode the index indicating the vector predictor in the candidate list.

[0197] In addition, the video encoding device may subtract the predicted block from the original block of the current block to generate a residual block, and then encode the residual block.

[0198] On the one hand, if the horizontal component of the vector difference is non-zero or the vertical component of the vector difference is non-zero, the video encoding device may perform the steps of generating a vector difference candidate through the steps of determining the predicted block (S1406 to S1414). On the other hand, if both the horizontal component and the vertical component of the vector difference are zero, the video encoding device may not deduce the sign of the vector difference.

[0199] Figure 14 The flowchart shown in corresponds to Implementation Method 2-1, and Implementation Method 2-1 continuously performs the adjustment of the template for deriving the vector symbol. The video encoding device may selectively perform the adjustment of the template according to Implementation Method 2-2. For this purpose, the video encoding device may further Figure 14The steps shown below perform the following operations. For ease of description, according to Figure 14 the prediction block is defined as the first prediction block of the current block.

[0200] The video encoding device sets the template of the reference block candidate and the template of the current block to a preset size. The video encoding device calculates the template matching cost between the template of the current block and the template of the reference block candidate. The video encoding device determines the second prediction block of the current block by selecting the template that minimizes the template matching cost.

[0201] The video encoding device determines a template adjustment flag based on the first prediction block and the second prediction block. Here, the template adjustment flag indicates whether to adjust the template of the reference block candidate and the template of the current block. In terms of rate-distortion optimization, the video encoding device may determine the template adjustment flag. For example, if the first prediction block is the best, the video encoding device sets the template adjustment flag to true. On the other hand, if the second prediction block is the best, the video encoding device may set the template adjustment flag to false.

[0202] The video encoding device encodes the vector prediction factor and the magnitude of the vector difference. In addition, the video encoding device encodes the template adjustment flag.

[0203] On the one hand, if the horizontal component of the vector difference is non-zero or the vertical component of the vector difference is non-zero, the video encoding device may perform the process associated with determining the template adjustment flag. On the other hand, if both the horizontal component and the vertical component of the vector difference are zero, the video encoding device may not derive the sign of the vector difference.

[0204] Figure 15 is a flowchart of a method for a video decoding device to reconstruct a current block according to at least one embodiment of the present invention.

[0205] The video decoding device decodes a vector prediction factor (S1500) from a bitstream for predicting a vector of the current block. Here, the vector may be a block vector or a motion vector. For example, after generating a candidate list, the video decoding device may use the decoded index to determine a vector predictor from the candidate list.

[0206] The video decoding device decodes the magnitude of the vector difference (S1502) from the bitstream.

[0207] The video decoding device may combine the magnitude of the vector difference and the possible signs of the vector difference to generate vector difference candidates (S1504).

[0208] The video decoding device generates vector candidates by summing the vector prediction factor and the vector difference candidates, and then generates reference block candidates corresponding to the vector candidates in the picture (S1506).

[0209] In this case, if the vector is a block vector, the picture may be the current picture including the current block, and if the vector is a motion vector, the picture may be a reference picture of the current block.

[0210] The video decoding device adjusts the template of the reference block candidate and the template of the current block based on the magnitude of the vector difference (S1508).

[0211] Here, the adjustment of the template may be an adjustment of the template size (Implementation Method 1-1 or Implementation Method 1-2) or an adjustment of the calculation area of the template matching cost (Implementation Method 1-3).

[0212] The video decoding device calculates the template matching cost between the adjusted template of the current block and the adjusted template of the reference block candidate (S1510).

[0213] The video decoding device determines the predicted block of the current block by selecting the template that minimizes the template matching cost (S1512). By selecting the template that minimizes the template matching cost and selecting the corresponding vector difference candidate and reference block candidate of the selected template, the video decoding device can determine the predicted block of the current block.

[0214] In addition, the video decoding device may decode the residual block from the bitstream and then add the residual block and the predicted block to generate the reconstructed block of the current block.

[0215] On the one hand, if the horizontal component of the vector difference is non-zero or the vertical component of the vector difference is non-zero, the video decoding device may perform the step of generating the vector difference candidate by the step of determining the predicted block (S1504 to S1512). On the other hand, if both the horizontal component and the vertical component of the vector difference are zero, the video decoding device may not derive the sign of the vector difference.

[0216] Figure 15 The flowchart shown in corresponds to Embodiment 2-1, and Embodiment 2-1 continuously performs the adjustment of the template for deriving the vector sign. The video decoding device may selectively apply the adjustment of the template according to Implementation Method 2-2. For this purpose, in addition to Figure 15 the steps shown in, the video decoding device may also perform the following operations.

[0217] For example, after decoding the magnitude of the vector difference (step S1502), the video decoding device decodes the template adjustment flag from the bitstream. Here, the template adjustment flag indicates whether to adjust the template of the reference block candidate and the template of the current block.

[0218] The video decoding device checks the template adjustment flag.

[0219] If the template adjustment flag is true, the video decoding device performs the steps of generating vector difference candidates to determining a prediction block (S1504 to S1512).

[0220] Meanwhile, if the template adjustment flag is false, the video decoding device generates vector difference candidates by combining the magnitude of the vector difference and possible sign combinations of the vector difference. The video decoding device generates vector candidates by summing a vector prediction factor and the vector difference candidates, and subsequently generates reference block candidates corresponding to the vector candidates in the picture. The video decoding device sets the template of the reference block candidates and the template of the current block to a preset size. The video decoding device calculates a template matching cost between the template of the current block and the template of the reference block candidates. The video decoding device determines the prediction block of the current block by selecting the template that minimizes the template matching cost.

[0221] In addition, the video decoding device may decode a residual block from a bitstream and then add the residual block and the prediction block to generate a reconstructed block of the current block.

[0222] On the one hand, if the horizontal component of the vector difference is non-zero or the vertical component of the vector difference is non-zero, the video decoding device may perform processing associated with decoding the template adjustment flag. On the other hand, if both the horizontal component and the vertical component of the vector difference are zero, the video encoding device may not derive the sign of the vector difference.

[0223] Although the steps in the respective flowcharts are described as being executed sequentially, these steps merely illustrate the technical concepts of some embodiments of the present disclosure. Therefore, those of ordinary skill in the art to which the present disclosure pertains may execute the steps by changing the order described in the respective drawings or by executing two or more steps in parallel. Thus, the steps in the respective flowcharts are not limited to the shown chronological order.

[0224] It should be understood that the above description presents illustrative embodiments that can be implemented in different other ways. The functions described in some embodiments can be implemented by hardware, software, firmware, and / or combinations thereof. It should also be understood that the functional components described in the present disclosure are marked with “… unit” to strongly emphasize the possibility of their independent implementation.

[0225] Meanwhile, the various methods or functions described in some embodiments can be implemented as instructions stored in a non-transitory recording medium that can be read and executed by one or more processors. For example, the non-transitory recording medium may include various types of recording devices, in which data is stored in a form readable by a computer system. For example, the non-transitory recording medium may include storage media such as erasable programmable read-only memory (EPROM), flash drives, optical disc drives, magnetic hard disk drives, and solid state drives (SSD), etc.

[0226] Although embodiments of the present disclosure have been described for illustrative purposes, those of ordinary skill in the art to which the present disclosure pertains should recognize that various modifications, additions, and substitutions are possible without departing from the concept and scope of the present disclosure. Therefore, for the sake of brevity and clarity, embodiments of the present disclosure have been described. The scope of the technical concept of the embodiments of the present disclosure is not limited by the illustrations. Accordingly, those of ordinary skill in the art to which the present disclosure pertains should understand that the scope of the present disclosure should not be limited by the embodiments described explicitly above, but rather by the claims and their equivalents.

[0227] (Reference numeral)

[0228] 120: Predictor

[0229] 155: Entropy encoder

[0230] 510: Entropy decoder

[0231] 540: Predictor.

[0232] Cross - reference to related applications

[0233] This application claims the priority and benefits of Korean Patent Application No. 10 - 2022 - 0168876, filed on December 6, 2022, and Korean Patent Application No. 10 - 2023 - 0139622, filed on October 18, 2023, the entire contents of each of which are incorporated herein by reference.

Claims

1. A method for a video decoding device to reconstruct a current block, the method comprising: Decoding, from a bitstream, a vector prediction factor (vector prediction factor) for a vector used to predict the current block; Decoding, from the bitstream, an amplitude of a vector difference (difference); Generating a vector difference candidate by combining the amplitude of the vector difference with possible signs of the vector difference; Generating a vector candidate by summing the vector prediction factor and the vector difference candidate, and then generating a reference block candidate corresponding to the vector candidate within a picture; Adjusting a template of the reference block candidate and a template of the current block based on the amplitude of the vector difference; Calculating a template matching cost between the adjusted template of the current block and the adjusted template of the reference block candidate; And Determining a predicted block of the current block by selecting a template that minimizes the template matching cost.

2. The method according to claim 1, wherein, The vector includes: A block vector or a motion vector, Wherein, when the vector is the block vector, the picture is a current picture containing the current block, and Wherein, when the vector is the motion vector, the picture is a reference picture of the current block.

3. The method according to claim 1, wherein, The template of the reference block candidate and the template of the current block include: An upper template and a left template, the upper template having a size defined by a length and a thickness of the upper template, and the left template having a size defined by a length and a thickness of the left template.

4. The method according to claim 3, wherein Adjusting the template of the reference block candidate and the template of the current block includes: Selecting a representative value (representative value) based on the amplitude of the vector difference, and then adjusting the sizes of the upper template and the left template to be equal by using the representative value.

5. The method according to claim 4, wherein, Adjusting the template of the reference block candidate and the template of the current block includes: when the vector difference has a non-zero horizontal component and a non-zero vertical component: Setting the representative value to: the smaller of the horizontal component and the vertical component of the vector difference, or the larger of the horizontal component and the vertical component of the vector difference, or an average value of the horizontal component and the vertical component of the vector difference, or a sum of the horizontal component and the vertical component of the vector difference.

6. The method according to claim 3, wherein, Adjusting the template of the reference block candidate and the template of the current block includes: Adjusting the size of the upper template based on the horizontal component of the vector difference and adjusting the size of the left template based on the vertical component of the vector difference, or Adjusting the size of the left template based on the horizontal component of the vector difference and adjusting the size of the upper template based on the vertical component of the vector difference.

7. The method according to claim 3, wherein Adjusting the template of the reference block candidate and the template of the current block includes: Adjusting a calculation region of the template matching cost in the upper template and the left template based on the amplitude of the vector difference.

8. The method according to claim 7, wherein Adjusting the template of the reference block candidate and the template of the current block includes: For a preset horizontal length based on a horizontal component of the vector difference and a preset vertical length based on a vertical component of the vector difference, regions respectively extending the preset horizontal length from a left side and a right side of the upper template and regions respectively extending the preset vertical length from an upper side and a lower side of the left template are used as the calculation regions for the template matching cost.

9. The method according to claim 1, further comprising: decoding a template adjustment flag from the bitstream, the template adjustment flag indicating whether to adjust a template of the reference block candidate and a template of the current block; and checking the template adjustment flag, wherein when the template adjustment flag is true, the method continues to execute from generating the vector difference candidate to determining the prediction block.

10. The method according to claim 9, further comprising, when the template adjustment flag is false: generating the vector difference candidate by combining the magnitude of the vector difference with the possible sign combinations of the vector difference; generating the vector candidate by summing the vector prediction factor and the vector difference candidate, and then generating the reference block candidate corresponding to the vector candidate within the picture; setting the template of the reference block candidate and the template of the current block to a preset size; calculating a template matching cost between the template of the current block and the template of the reference block candidate; and determining the prediction block of the current block by selecting a template that minimizes the template matching cost.

11. A method for encoding a current block by a video encoding device, the method comprising: determining a vector for predicting the current block; determining a vector prediction factor (vector prediction factor) of the current block; generating a magnitude of a vector difference (difference) by subtracting the vector prediction factor from the vector; generating a vector difference candidate by combining the magnitude of the vector difference with possible sign combinations of the vector difference; generating a vector candidate by summing the vector prediction factor and the vector difference candidate, and then generating a reference block candidate corresponding to the vector candidate within the picture; adjusting the template of the reference block candidate and the template of the current block based on the magnitude of the vector difference; calculating a template matching cost between the adjusted template of the current block and the adjusted template of the reference block candidate; and determining a first prediction block of the current block by selecting a template that minimizes the template matching cost.

12. The method according to claim 11, further comprising: encoding the vector prediction factor and the magnitude of the vector difference.

13. The method according to claim 11, wherein The vector includes: a block vector or a motion vector, wherein when the vector is the block vector, the picture is the current picture containing the current block, and wherein when the vector is the motion vector, the picture is a reference picture of the current block.

14. The method according to claim 11, further comprising: setting the template of the reference block candidate and the template of the current block to a preset size; calculating a template matching cost between the template of the current block and the template of the reference block candidate; and The second prediction block of the current block is determined by selecting a template that minimizes the template matching cost.

15. The method according to claim 14, comprising: Determining a template adjustment flag based on the first prediction block and the second prediction block, the template adjustment flag indicating whether to adjust the template of the reference block candidate and the template of the current block; Encoding the vector prediction factor and the magnitude of the vector difference; And Encoding the template adjustment flag.

16. A computer-readable recording medium storing a bitstream generated by a video coding method, the video coding method comprising: Determining a vector for predicting a current block; Determining a vector prediction factor (vector prediction factor) of the current block; Generating a magnitude of a vector difference (difference) by subtracting the vector prediction factor from the vector; Generating a vector difference candidate by combining the magnitude of the vector difference with possible signs of the vector difference; Generating a vector candidate by summing the vector prediction factor and the vector difference candidate, and then generating a reference block candidate corresponding to the vector candidate within a picture; Adjusting the template of the reference block candidate and the template of the current block based on the magnitude of the vector difference; Calculating a template matching cost between the adjusted template of the current block and the adjusted template of the reference block candidate; And Determining a prediction block of the current block by selecting a template that minimizes the template matching cost.

Citation Information

Patent Citations

  • Gas emulsion combustion device

    KR1020220168876A

  • Device and method to authorize user based on video data

    KR1020230139622A