Method and apparatus for video coding and decoding using motion compensation filter adaptively in affine model-based prediction

By adaptively adjusting the control point motion vector of the affine model, accurately determining the pixel position of the reference block of the sub-block and applying a motion compensation filter, the problems of high computational complexity and insufficient prediction accuracy in the sub-block prediction of the affine model are solved, and more efficient video encoding and decoding and better video quality are achieved.

CN120266473APending Publication Date: 2025-07-04HYUNDAI MOTOR CO LTD +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202380081744.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-11-23
Filing Date
2023-11-24
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

When existing video encoding and decoding technologies use affine models to predict sub-blockwise, there are problems such as high computational complexity and insufficient prediction accuracy, resulting in poor encoding and decoding efficiency and video quality.

Method used

By adaptively adjusting the control point motion vector of the affine model, the pixel position of the reference block of each sub-block is accurately determined, and a motion compensation filter is used to predict, reducing the computational complexity and improving prediction accuracy.

Benefits of technology

Improves video encoding and decoding efficiency, enhances video quality, and reduces computational complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120266473A_ABST
    Figure CN120266473A_ABST
Patent Text Reader

Abstract

The present embodiment provides a video encoding and decoding method and apparatus that adaptively utilizes a motion compensation filter in affine model-based prediction. In this embodiment, an image decoding device decodes control point motion vector information and affine model information of a current block. The image decoding device determines a form of an affine model based on the affine model information, and derives a control point motion vector of the current block based on the affine model and the control point motion vector information. An image decoding device generates a motion vector of a sub-block by using a control point motion vector of a current block, and acquires a geometric model parameter of the sub-block. An image decoding device generates a reference block from a reference picture by using motion vectors of sub-blocks, and calculates positions of reference pixels in the reference block on the basis of geometric model parameters. An image decoding device applies an interpolation filter to a reference pixel to generate a prediction value of a sub-block, thereby generating a prediction block of a current block.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a video encoding and decoding method and apparatus for adaptively utilizing a motion compensation filter in affine model-based prediction. Background Art

[0002] The statements in this section merely provide background technical information related to the present invention and do not necessarily constitute prior art.

[0003] Since video data has a large amount of data compared to audio data or still image data, video data requires a large amount of hardware resources (including memory) to store or transmit uncompressed video data.

[0004] Accordingly, an encoder is typically used to compress and store or transmit video data. A decoder receives the compressed video data, decompresses the received compressed video data, and plays the decompressed video data. Video compression techniques include H.264 / Advanced Video Coding (AVC), High Efficiency Video Coding (HEVC), and Versatile Video Coding (VVC), and the Versatile Video Coding (VVC) has a coding / decoding efficiency that is approximately 30% or more higher than that of HEVC.

[0005] However, due to the gradual increase in image size, resolution, and frame rate, the amount of data to be encoded also increases. Accordingly, there is a need to provide a new compression technique with higher coding / decoding efficiency and improved image enhancement effects compared to existing compression techniques.

[0006] To improve the prediction performance in video compression, affine model-based prediction is used to perform encoding and decoding. To improve the encoding / decoding efficiency using the affine model, geometric relationships are derived and modeled for object signals or background signals in the video that change in space and time with the movement of the camera or object, and then the modeled relationships are applied to reference signals and original signals. To reduce the computational complexity of the affine model and reduce the encoding / decoding complexity during motion compensation, a geometric model is calculated according to a method of deriving a control point motion vector based on adjacent motion vectors. In addition, motion compensation of the current block is performed by dividing the current block into sub-blocks and calculating the motion vector of each sub-block. Compared with the method of performing motion compensation on a pixel-by-pixel basis, the method of performing sub-block-by-sub-block motion compensation based on the affine model reduces the computational complexity, but this sub-block-by-sub-block motion compensation suffers from deteriorated prediction accuracy. Therefore, in order to improve video encoding and decoding efficiency and enhance video quality, a method for improving prediction accuracy when performing sub-block-by-sub-block prediction using the affine model is needed. Summary of the Invention

[0007] Technical Problem

[0008] The present invention is dedicated to providing a video encoding and decoding method and apparatus, which adaptively adjusts the pixel positions of the reference blocks of each sub-block based on the control point motion vectors of the affine model when predicting the current block in units of sub-blocks based on the affine model. The video encoding and decoding method and apparatus apply a motion compensation filter to the adjusted pixel positions.

[0009] Technical solution

[0010] At least one aspect of the present invention provides a method for a video decoding apparatus to decode a current block. The method includes decoding affine model information and control point motion vector information of the current block from a bitstream. The affine model information indicates the form of the affine model, and the control point motion vector information includes a prediction method for the control point motion vector of the current block and a control point motion vector difference. The method further includes determining the form of the affine model based on the affine model information. The method further includes deriving the control point motion vector of the current block based on the affine model and the control point motion vector information. The method further includes generating a motion vector of a sub-block by using the control point motion vector of the current block, where the sub-block is generated based on the partitioning of the current block. The method further includes obtaining geometric model parameters of the sub-block, where the geometric model parameters are defined based on the control point motion vector of the sub-block and the motion vector of the sub-block. The method further includes generating a reference block according to a reference picture by using the motion vector of the sub-block. The method further includes calculating the position of a reference pixel in the reference block based on the geometric model parameters. The method further includes generating a prediction value of the sub-block by applying an interpolation filter to the reference pixel to generate a prediction block of the current block.

[0011] Another aspect of the present invention provides a method for a video encoding apparatus to encode a current block. The method includes determining affine model information of the current block and a prediction method for the control point motion vector. The affine model information indicates the form of the affine model. The method further includes determining the form of the affine model based on the affine model information. The method further includes deriving the control point motion vector of the current block based on the affine model. The method further includes generating a motion vector of a sub-block by using the control point motion vector of the current block, where the sub-block is generated based on the partitioning of the current block. The method further includes determining geometric model parameters of the sub-block, where the geometric model parameters are defined based on the control point motion vector of the sub-block and the motion vector of the sub-block. The method further includes generating a reference block according to a reference picture by using the motion vector of the sub-block. The method further includes calculating the position of a reference pixel in the reference block based on the geometric model parameters. The method further includes generating a prediction value of the sub-block by applying an interpolation filter to the reference pixel to generate a prediction block of the current block. The method further includes encoding the affine model information and the prediction method for the control point motion vector.

[0012] Another aspect of the present invention provides a computer-readable recording medium that stores a bitstream generated by a video encoding method. The video encoding method includes determining affine model information of a current block and a prediction method for controlling a motion vector of a control point. The affine model information indicates a form of the affine model. The video encoding method further includes determining the form of the affine model based on the affine model information. The video encoding method further includes deriving a motion vector of a control point of the current block based on the affine model. The video encoding method further includes generating a motion vector of a sub-block by using the motion vector of the control point of the current block, where the sub-block is generated based on a partition of the current block. The video encoding method further includes determining geometric model parameters of the sub-block, where the geometric model parameters are defined based on the motion vector of the control point of the sub-block and the motion vector of the sub-block. The video encoding method further includes generating a reference block from a reference picture by using the motion vector of the sub-block. The video encoding method further includes calculating a position of a reference pixel in the reference block based on the geometric model parameters. The video encoding method further includes generating a predicted value of the sub-block by applying an interpolation filter to the reference pixel to generate a predicted block of the current block. The video encoding method further includes encoding the affine model information and the prediction method for the motion vector of the control point.

[0013] Advantageous Effects

[0014] As described above, the present invention provides a video encoding and decoding method and apparatus that adaptively adjust the pixel positions of reference blocks of each sub-block based on the motion vector of the control point of an affine model when performing sub-block-by-sub-block prediction based on the affine model for a current block. The video encoding and decoding method and apparatus apply a motion compensation filter to the adjusted pixel positions. Therefore, the video encoding and decoding method and apparatus improve the video encoding and decoding efficiency and enhance the video quality. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 is a block diagram of a video encoding apparatus that can implement the technology of the present invention.

[0016] Figure 2 illustrates a method of partitioning blocks by using a quadtree plus binary tree plus ternary tree (QTBTTT) structure.

[0017] Figure 3a and Figure 3b illustrates a plurality of intra prediction modes including a wide-angle intra prediction mode.

[0018] Figure 4 illustrates adjacent blocks of a current block.

[0019] Figure 5 is a block diagram of a video decoding apparatus that can implement the technology of the present invention.

[0020] Figure 6Is a schematic diagram showing the per-pixel calculation of an affine model according to at least one embodiment of the present invention.

[0021] Figure 7 Is a schematic diagram showing the calculation of an affine model by using a central vector according to at least one embodiment of the present invention.

[0022] Figure 8 Is a schematic diagram showing the per-sub-block calculation of an affine model according to at least one embodiment of the present invention.

[0023] Figure 9 Is a schematic diagram showing a part of a video decoding device for prediction based on an affine model according to at least one embodiment of the present invention.

[0024] Figure 10a And Figure 10b Is a schematic diagram showing motion compensation at fractional position pixels according to at least one embodiment of the present invention.

[0025] Figure 11a And Figure 11b Is a schematic diagram showing motion compensation at fractional position pixels according to another embodiment of the present invention.

[0026] Figure 12 Is a flowchart of a method for encoding a current block by a video encoding device according to at least one embodiment of the present invention.

[0027] Figure 13 Is a flowchart of a method for reconstructing a current block by a video decoding device according to at least one embodiment of the present invention. Detailed Description

[0028] Hereinafter, some embodiments of the present invention will be described in detail with reference to the accompanying illustrative drawings. In the following description, the same reference numerals denote the same elements, although the elements are shown in different drawings. In addition, in the following description of some embodiments, when the detailed description of related known components and functions is considered to obscure the subject matter of the present invention, the detailed description of the related known components and functions may be omitted for clarity and conciseness.

[0029] Figure 1 Is a block diagram of a video encoding device that can implement the technology of the present invention. Hereinafter, with reference to Figure 1 The illustration of, the video encoding device and the components of the device are described.

[0030] The encoding device may include: an image splitter 110, a predictor 120, a subtractor 130, a transformer 140, a quantizer 145, a rearrangement unit 150, an entropy encoder 155, an inverse quantizer 160, an inverse transformer 165, an adder 170, a loop filter unit 180, and a memory 190.

[0031] Each component of the encoding device may be implemented as hardware or software, or as a combination of hardware and software. Additionally, the functions of each component may be implemented as software, and the microprocessor may also be implemented to execute the functions of the software corresponding to each component.

[0032] A video consists of one or more sequences including multiple images. Each image is segmented into multiple regions, and encoding is performed on each region. For example, an image is segmented into one or more tiles or / and slices. Here, one or more tiles may be defined as a tile group. Each tile or / and slice is segmented into one or more coding tree units (CTUs). Additionally, each CTU is segmented into one or more coding units (CUs) through a tree structure. The information applied to each coding unit (CU) is encoded into the syntax of the CU, and the information applied to the CUs included in one CTU is encoded into the syntax of the CTU. Additionally, the information applied to all blocks in one slice is encoded into the syntax of the slice header, while the information applied to all blocks constituting one or more images is encoded into the Picture Parameter Set (PPS) or the picture header. Furthermore, the information commonly referred to by multiple images is encoded into the Sequence Parameter Set (SPS). Additionally, the information commonly referred to by one or more SPSs is encoded into the Video Parameter Set (VPS). Moreover, the information applied to one tile or tile group may also be encoded into the syntax of the tile or tile group header. The syntax included in the SPS, PPS, slice header, tile or tile group header may be referred to as high-level syntax.

[0033] The image splitter 110 determines the size of the coding tree unit (CTU). The information regarding the size (CTU size) of the CTU is encoded into the syntax of the SPS or PPS and is transmitted to the video decoding device.

[0034] The image splitter 110 segments each image constituting the video into multiple coding tree units (CTUs) of a predetermined size, and then recursively segments the CTUs by using a tree structure. The leaf nodes in the tree structure become coding units (CUs), and the CUs are the basic units of encoding.

[0035] The tree structure can be a quadtree (QT), where a higher node (or parent node) is divided into four lower nodes (or child nodes) of the same size. The tree structure can also be a binary tree (BT), where a higher node is divided into two lower nodes. The tree structure can also be a ternary tree (TT), where a higher node is divided into three lower nodes in a 1:2:1 ratio. The tree structure can also be a structure that mixes two or more of the QT structure, BT structure, and TT structure. For example, a quadtree plus binarytree (QTBT) structure can be used, or a quadtree plus binarytreeternarytree (QTBTTT) structure can be used. Here, the binarytreeternarytree (BTTT) is added to the tree structure to form a multiple-type tree (MTT).

[0036] Figure 2 is a schematic diagram for describing a method of dividing a block by using the QTBTTT structure.

[0037] As Figure 2 shown, the CTU can first be divided into a QT structure. The quadtree division can be recursive until the size of the divided block reaches the minimum block size (MinQTSize) of the leaf nodes allowed in the QT. The entropy encoder 155 encodes a first flag (QT_split_flag) indicating whether each node of the QT structure is divided into four lower nodes and signals it to the video decoding device. When the leaf node of the QT is not larger than the maximum block size (MaxBTSize) of the root node allowed in the BT, the leaf node can be further divided into at least one of the BT structure or the TT structure. There can be multiple division directions in the BT structure and / or TT structure. For example, there can be two directions, namely, the direction of horizontally dividing the block of the corresponding node and the direction of vertically dividing the block of the corresponding node. As Figure 2 shown, when the MTT division starts, the entropy encoder 155 encodes a second flag (mtt_split_flag) indicating whether the node is divided, and a flag indicating the division direction (vertical or horizontal) and / or a flag indicating the division type (binary or ternary) in the case where the node is divided, and signals it to the video decoding device.

[0038] Alternatively, before encoding a first flag (QT_split_flag) indicating whether each node is split into four lower-layer nodes, a CU split flag (split_cu_flag) indicating whether a node is split may also be encoded. When the value of the CU split flag (split_cu_flag) indicates that each node is not split, the block of the corresponding node becomes a leaf node in the split tree structure and becomes a CU, which is the basic unit of encoding. When the value of the CU split flag (split_cu_flag) indicates that each node is split, the video encoding device starts encoding the first flag first with the above scheme.

[0039] When QTBT is used as another example of a tree structure, there can be two types, that is, a type in which the block of the corresponding node is horizontally split into two blocks of the same size (i.e., symmetric horizontal split) and a type in which the block of the corresponding node is vertically split into two blocks of the same size (i.e., symmetric vertical split). The entropy encoder 155 encodes a split flag (split_flag) indicating whether each node of the BT structure is split into lower-layer blocks and split type information indicating the split type, and transmits it to the video decoding device. On the other hand, there can be another type in which the block of the corresponding node is split into two asymmetric blocks. The asymmetric form can include a form in which the block of the corresponding node is split into two rectangular blocks with a size ratio of 1:3, or can also include a form in which the block of the corresponding node is split in the diagonal direction.

[0040] A CU can have various sizes according to the QTBT or QTBTTT split from the CTU. Hereinafter, the block corresponding to the CU to be encoded or decoded (i.e., the leaf node of the QTBTTT) is referred to as the "current block". When QTBTTT split is adopted, in addition to the square shape, the shape of the current block can also be a rectangular shape.

[0041] The predictor 120 predicts the current block to generate a predicted block. The predictor 120 includes an intra predictor 122 and an inter predictor 124.

[0042] Generally, each of the current blocks in the image can be predictively encoded. Generally, the prediction of the current block can be performed by using an intra prediction technique (which uses data from the image including the current block) or an inter prediction technique (which uses data from an image encoded before the image including the current block). Inter prediction includes both uni-directional prediction and bi-directional prediction.

[0043] The intra predictor 122 predicts the pixels in the current block by using the pixels (reference pixels) adjacent to the current block in the current image including the current block. According to the prediction direction, there are multiple intra prediction modes. For example, as Figure 3aAs shown, multiple intra-prediction modes may include two non-directional modes including the Planar mode and the DC mode, and may include 65 directional modes. The neighboring pixels and algorithm equations to be used are defined differently according to each prediction mode.

[0044] For efficient directional prediction of a current block having a rectangular shape, the directional modes shown by the dashed arrows in Figure 3b may be additionally used (modes #67 to #80, intra-prediction modes #-1 to #-14). The directional modes may be referred to as "wide angle intra-prediction modes". In Figure 3b the arrows indicate the corresponding reference samples for prediction, rather than representing the prediction direction. The prediction direction is opposite to the direction indicated by the arrows. When the current block has a rectangular shape, the wide angle intra-prediction mode is a mode that performs prediction in the direction opposite to a specific directional mode without additional bit transmission. In this case, in the wide angle intra-prediction mode, some wide angle intra-prediction modes available for the current block may be determined by the ratio of the width to the height of the current block having a rectangular shape. For example, when the current block has a rectangular shape with a height less than the width, the wide angle intra-prediction modes having an angle less than 45 degrees (intra-prediction modes #67 to #80) are available. When the current block has a rectangular shape with a width greater than the height, the wide angle intra-prediction modes having an angle greater than -135 degrees are available.

[0045] The intra-predictor 122 may determine the intra-prediction to be used for encoding the current block. In some examples, the intra-predictor 122 may encode the current block by utilizing multiple intra-prediction modes, and may also select an appropriate intra-prediction mode to be used from test modes. For example, the intra-predictor 122 may calculate rate-distortion values by utilizing rate-distortion analysis of multiple tested intra-prediction modes, and may also select the intra-prediction mode having the best rate-distortion characteristics in the test modes.

[0046] The intra-predictor 122 selects one intra-prediction mode from multiple intra-prediction modes, and predicts the current block by utilizing the neighboring pixels (reference pixels) and algorithm equations determined according to the selected intra-prediction mode. The entropy encoder 155 encodes the information about the selected intra-prediction mode and transmits it to the video decoding device.

[0047] The inter - frame predictor 124 generates a predicted block of the current block by using motion - compensation processing. The inter - frame predictor 124 searches for the block most similar to the current block in a reference image that has been encoded and decoded earlier than the current image, and generates a predicted block of the current block by using the searched - for block. In addition, a motion vector (MV) is generated, which corresponds to the displacement between the current block in the current image and the predicted block in the reference image. Generally, motion estimation is performed on the luma component, and the motion vector calculated based on the luma component is used for both the luma component and the chroma component. The entropy encoder 155 encodes the motion information including the information of the reference image and the information about the motion vector used for predicting the current block, and transmits it to the video decoding device.

[0048] The inter - frame predictor 124 may also perform interpolation of the reference image or reference block to increase the prediction accuracy. In other words, sub - samples are interpolated between two consecutive integer samples by applying filter coefficients to a plurality of consecutive integer samples including two integer samples. When performing the process of searching for the block most similar to the current block on the interpolated reference image, the motion vector can represent fractional - unit accuracy rather than integer - sample - unit accuracy. For each target region to be encoded, such as units like slices, tiles, CTUs, CUs, etc., the accuracy or resolution of the motion vector can be set differently. When applying such adaptive motion vector resolution (AMVR), information about the motion vector resolution to be applied to each target region should be signaled. For example, when the target region is a CU, information about the motion vector resolution applied to each CU is signaled. The information about the motion vector resolution can be information representing the accuracy of the motion - vector difference described below.

[0049] On the other hand, the inter-frame predictor 124 can perform inter-frame prediction by using bidirectional prediction. In the case of bidirectional prediction, two reference images and two motion vectors representing the positions of the blocks most similar to the current block in each reference image are used. The inter-frame predictor 124 selects the first reference image and the second reference image from reference picture list 0 (RefPicList0) and reference picture list 1 (RefPicList1), respectively. The inter-frame predictor 124 also searches for the blocks most similar to the current block in the corresponding reference images to generate the first reference block and the second reference block. In addition, a predicted block of the current block is generated by averaging or weighted averaging the first reference block and the second reference block. In addition, motion information including information about the two reference images used for predicting the current block and including information about the two motion vectors is transmitted to the entropy encoder 155. Here, reference picture list 0 may be composed of the images in the pre-reconstructed images that are before the current image in the display order, and reference picture list 1 may be composed of the images in the pre-reconstructed images that are after the current image in the display order. However, although not particularly limited thereto, the pre-reconstructed images after the current image in the display order may be additionally included in reference picture list 0. Conversely, the pre-reconstructed images before the current image may also be additionally included in reference picture list 1.

[0050] To minimize the amount of bits consumed for encoding the motion information, various methods can be used.

[0051] For example, when the reference image and the motion vector of the current block are the same as those of an adjacent block, the information of the adjacent block that can be identified is encoded to transmit the motion information of the current block to the video decoding device. This method is called the merge mode.

[0052] In the merge mode, the inter-frame predictor 124 selects a predetermined number of merge candidates (hereinafter referred to as "merge candidates") from the adjacent blocks of the current block.

[0053] As the adjacent blocks for deriving the merge candidates, all or some of the left block A0, the lower left block A1, the upper block B0, the upper right block B1, and the upper left block B2 adjacent to the current block in the current image can be used, as Figure 4 shown. In addition, in addition to the current image where the current block is located, the blocks in the reference image (which may be the same as or different from the reference image used for predicting the current block) can also be used as merge candidates. For example, the co-located block of the current block in the reference image or the block adjacent to the co-located block can be additionally used as a merge candidate. If the number of merge candidates selected by the above method is less than the preset number, zero vectors are added to the merge candidates.

[0054] The inter-frame predictor 124 configures a merge list including a predetermined number of merge candidates by using neighboring blocks. A merge candidate to be used as the motion information of the current block is selected from among the merge candidates included in the merge list, and merge index information for identifying the selected candidate is generated. The generated merge index information is encoded by the entropy encoder 155 and transmitted to the video decoding device.

[0055] The merge skip mode is a special case of the merge mode. After quantization, when all the transform coefficients for entropy coding are close to zero, only the neighboring block selection information is transmitted without transmitting the residual signal. By using the merge skip mode, relatively high coding efficiency can be achieved for images with slight motion, still images, screen content images, etc.

[0056] Thereafter, the merge mode and the merge skip mode are collectively referred to as the merge / skip mode.

[0057] Another method for encoding motion information is the advanced motion vector prediction (AMVP) mode.

[0058] In the AMVP mode, the inter-frame predictor 124 derives motion vector prediction candidates for the motion vector of the current block by using neighboring blocks of the current block. As neighboring blocks for deriving motion vector prediction candidates, all or some of the left block A0, lower left block A1, upper block B0, upper right block B1, and upper left block B2 adjacent to the current block in the current image shown in Figure 4 can be used. In addition, in addition to the current image where the current block is located, blocks in a reference image (which may be the same as or different from the reference image used to predict the current block) can also be used as neighboring blocks for deriving motion vector prediction candidates. For example, the co-located block of the current block in the reference image or a block adjacent to the co-located block can be used. If the number of motion vector candidates selected by the above method is less than a preset number, a zero vector is added to the motion vector candidates.

[0059] The inter-frame predictor 124 derives motion vector prediction candidates by using the motion vectors of neighboring blocks, and determines the motion vector prediction of the motion vector of the current block by using the motion vector prediction candidates. In addition, the motion vector difference is calculated by subtracting the motion vector prediction from the motion vector of the current block.

[0060] Motion vector prediction can be obtained by applying predefined functions (e.g., median and average calculations, etc.) to motion vector prediction candidates. In this case, the video decoding device also knows the predefined functions. Additionally, since the neighboring blocks used to derive the motion vector prediction candidates are blocks that have already been encoded and decoded, the video decoding device may also already know the motion vectors of the neighboring blocks. Therefore, the video encoding device does not need to encode the information for identifying the motion vector prediction candidates. Accordingly, in this case, the information about the motion vector difference and the information about the reference image used to predict the current block are encoded.

[0061] On the other hand, motion vector prediction can also be determined by a scheme of selecting any one of the motion vector prediction candidates. In this case, the information for identifying the selected motion vector prediction candidate is additionally encoded together with the information about the motion vector difference and the information about the reference image used to predict the current block.

[0062] The subtractor 130 generates a residual block by subtracting the prediction block generated by the intra predictor 122 or the inter predictor 124 from the current block.

[0063] The transformer 140 transforms the residual signal in the residual block having pixel values in the spatial domain into transform coefficients in the frequency domain. The transformer 140 can transform the residual signal in the residual block by using the entire size of the residual block as the transform unit, or the residual block can also be divided into multiple sub - blocks, and the transform can be performed by using the sub - blocks as the transform unit. Alternatively, the residual block is divided into two sub - blocks, namely a transform region and a non - transform region, to transform the residual signal by using only the transform region sub - block as the transform unit. Here, the transform region sub - block can be one of two rectangular blocks having a size ratio of 1:1 based on the horizontal axis (or vertical axis). In this case, the entropy encoder 155 encodes a flag (cu_sbt_flag) indicating only the transform sub - block, and the direction (vertical / horizontal) information (cu_sbt_horizontal_flag) and / or the position information (cu_sbt_pos_flag), and signals it to the video decoding device. Additionally, the size of the transform region sub - block can have a size ratio of 1:3 based on the horizontal axis (or vertical axis). In this case, the entropy encoder 155 additionally encodes a flag (cu_sbt_quad_flag) for dividing the corresponding segmentation and signals it to the video decoding device.

[0064] On the other hand, the transformer 140 can perform the transformation of the residual block separately in the horizontal direction and the vertical direction. For this transformation, various types of transformation functions or transformation matrices can be used. For example, a pair of transformation functions for horizontal transformation and vertical transformation can be defined as a multiple transform set (MTS). The transformer 140 can select a pair of transformation functions with the highest transformation efficiency in the MTS, and can transform the residual block on each of the horizontal direction and the vertical direction. The entropy encoder 155 encodes the information (mts_idx) regarding the pair of transformation functions in the MTS and signals it to the video decoding device.

[0065] The quantizer 145 quantizes the transform coefficients output from the transformer 140 using quantization parameters and outputs the quantized transform coefficients to the entropy encoder 155. The quantizer 145 can also quantize the relevant residual block immediately without transforming any block or frame. The quantizer 145 can also apply different quantization coefficients (scaling values) according to the positions of the transform coefficients in the transform block. The quantization matrix applied to the quantized transform coefficients arranged in two dimensions can be encoded and signaled to the video decoding device.

[0066] The rearrangement unit 150 can perform rearrangement of the coefficient values on the quantized residual values.

[0067] The rearrangement unit 150 can change the 2D coefficient array into a 1D coefficient sequence by using coefficient scanning. For example, the rearrangement unit 150 can use a zig-zag scan or a diagonal scan to scan the coefficients from the DC coefficient to the high-frequency region to output a 1D coefficient sequence. According to the size of the transform unit and the intra prediction mode, a vertical scan that scans the 2D coefficient array in the column direction and a horizontal scan that scans the 2D block type coefficients in the row direction can also be used instead of the zig-zag scan. In other words, according to the size of the transform unit and the intra prediction mode, the scan method to be used can be determined among the zig-zag scan, the diagonal scan, the vertical scan, and the horizontal scan.

[0068] The entropy encoder 155 encodes the sequence of 1D quantized transform coefficients output from the rearrangement unit 150 by using various coding schemes including context-based adaptive binary arithmetic coding (CABAC), exponential Golomb, etc. to generate a bitstream.

[0069] In addition, the entropy encoder 155 encodes information related to block partitioning (e.g., CTU size, CTU partitioning flag, QT partitioning flag, MTT partitioning type, and MTT partitioning direction, etc.) so that the video decoding device can partition blocks in the same way as the video encoding device. In addition, the entropy encoder 155 encodes information about the prediction type indicating whether the current block is encoded by intra prediction or inter prediction. The entropy encoder 155 encodes intra prediction information (i.e., information about the intra prediction mode) or inter prediction information (merge index in the case of the merge mode, and information about the reference image index and motion vector difference in the case of the AMVP mode) according to the prediction type. In addition, the entropy encoder 155 encodes information related to quantization (i.e., information about the quantization parameter and information about the quantization matrix).

[0070] The inverse quantizer 160 inverse quantizes the quantized transform coefficients output from the quantizer 145 to generate transform coefficients. The inverse transformer 165 transforms the transform coefficients output from the inverse quantizer 160 from the frequency domain to the spatial domain to reconstruct the residual block.

[0071] The adder 170 adds the reconstructed residual block and the prediction block generated by the predictor 120 to reconstruct the current block. When performing intra prediction on the next block, the pixels in the reconstructed current block are used as reference pixels.

[0072] The loop filter unit 180 performs filtering on the reconstructed pixels to reduce block artifacts, ringing artifacts, blurring artifacts, etc. that occur due to block-based prediction and transform / quantization. The loop filter unit 180, as an in-loop filter, may include all or some of a deblocking filter 182, a sample adaptive offset (SAO) filter 184, and an adaptive loop filter (ALF) 186.

[0073] The deblocking filter 182 filters the boundaries between the reconstructed blocks to remove blocking artifacts that occur due to block-based coding / decoding, and the SAO filter 184 and the ALF 186 perform additional filtering on the deblocked video. The SAO filter 184 and the ALF 186 are filters for compensating the difference between the reconstructed pixels and the original pixels that occur due to lossy coding. The SAO filter 184 applies an offset in units of CTUs to enhance the subjective image quality and coding efficiency. On the other hand, the ALF 186 performs block-based filtering and applies different filters by dividing the boundaries and the degree of variation of the corresponding blocks to compensate for distortion. Information about the filter coefficients to be used for the ALF can be encoded and signaled to the video decoding device.

[0074] The reconstructed blocks filtered by the deblocking filter 182, the SAO filter 184, and the ALF 186 are stored in the memory 190. When all the blocks in an image are reconstructed, the reconstructed image can be used as a reference image for inter prediction of the blocks within the image to be encoded subsequently.

[0075] The video encoding device can store the bitstream of the encoded video data in a non-volatile storage medium or transmit the bitstream to the video decoding device through a communication network.

[0076] Figure 5 is a functional block diagram of a video decoding device that can implement the technology of the present invention. Hereinafter, with reference to Figure 5 , the video decoding device and the components of the device are described.

[0077] The video decoding device may include an entropy decoder 510, a rearrangement unit 515, an inverse quantizer 520, an inverse transform unit 530, a predictor 540, an adder 550, a loop filter unit 560, and a memory 570.

[0078] Similar to Figure 1 the video encoding device, each component of the video decoding device can be implemented as hardware or software, or implemented as a combination of hardware and software. In addition, the functions of each component can be implemented as software, and the microprocessor can also be implemented to execute the functions of the software corresponding to each component.

[0079] The entropy decoder 510 extracts information related to block partitioning by decoding the bitstream generated by the video encoding device to determine the current block to be decoded, and extracts the prediction information and information about the residual signal required for reconstructing the current block.

[0080] The entropy decoder 510 determines the size of a coding tree unit (CTU) by extracting information about the CTU size from a sequence parameter set (SPS) or a picture parameter set (PPS), and divides the picture into CTUs with the determined size. In addition, the CTU is determined as the top layer (i.e., the root node) of a tree structure, and the segmentation information of the CTU is extracted to divide the CTU by using the tree structure.

[0081] For example, when dividing a CTU by using a QTBTTT structure, first, a first flag (QT_split_flag) related to the segmentation of a quantization tree (QT) is extracted to divide each node into four lower-layer nodes. In addition, a second flag (mtt_split_flag) related to the segmentation of a multi-tree transform (MTT), a segmentation direction (vertical / horizontal), and / or a segmentation type (binary / trinary) are extracted for a node corresponding to a leaf node of the QT to divide the corresponding leaf node into an MTT structure. As a result, each node below the leaf node of the QT is recursively divided into a binary tree (BT) or a ternary tree (TT) structure.

[0082] As another example, when dividing a CTU by using a QTBTTT structure, a CU segmentation flag (split_cu_flag) indicating whether to divide a coding unit (CU) is extracted. When dividing the corresponding block, the first flag (QT_split_flag) may also be extracted. During the division process, for each node, zero or more recursive MTT divisions may occur after zero or more recursive QT divisions. For example, for a CTU, the MTT division may occur immediately, or conversely, only multiple QT divisions may occur.

[0083] As another example, when dividing a CTU by using a QTBT structure, a first flag (QT_split_flag) related to the segmentation of a QT is extracted to divide each node into four lower-layer nodes. In addition, a segmentation flag (split_flag) indicating whether to further divide the node corresponding to the leaf node of the QT into a BT and segmentation direction information are extracted.

[0084] On the other hand, when the entropy decoder 510 determines a current block to be decoded by using the segmentation of a tree structure, the entropy decoder 510 extracts information about a prediction type indicating whether the current block is intra-frame predicted or inter-frame predicted. When the prediction type information indicates intra-frame prediction, the entropy decoder 510 extracts a syntax element for intra-frame prediction information (intra-frame prediction mode) of the current block. When the prediction type information indicates inter-frame prediction, the entropy decoder 510 extracts information about syntax elements representing inter-frame prediction information, that is, a motion vector and a reference picture to which the motion vector refers.

[0085] In addition, the entropy decoder 510 extracts quantization-related information and extracts information on the quantized transform coefficients of the current block as information on the residual signal.

[0086] The rearrangement unit 515 can change the sequence of 1D quantized transform coefficients entropy-decoded by the entropy decoder 510 back into a 2D coefficient array (i.e., a block) in the reverse order of the coefficient scan order performed by the video coding device.

[0087] The inverse quantizer 520 inverse-quantizes the quantized transform coefficients and inverse-quantizes the quantized transform coefficients by using the quantization parameter. The inverse quantizer 520 can also apply different quantization coefficients (scaling values) to the quantized transform coefficients arranged in 2D. The inverse quantizer 520 can perform inverse quantization by applying a matrix of quantization coefficients (scaling values) from the video coding device to the 2D array of quantized transform coefficients.

[0088] The inverse transformer 530 reconstructs the residual signal by inverse-transforming the inverse-quantized transform coefficients from the frequency domain to the spatial domain to generate a residual block of the current block.

[0089] In addition, when the inverse transformer 530 inverse-transforms a partial region (sub-block) of the transform block, the inverse transformer 530 extracts a flag (cu_sbt_flag) for inverse-transforming only the sub-block of the transform block, direction (vertical / horizontal) information (cu_sbt_horizontal_flag) of the sub-block, and / or position information (cu_sbt_pos_flag) of the sub-block. The inverse transformer 530 also inverse-transforms the transform coefficients of the corresponding sub-block from the frequency domain to the spatial domain to reconstruct the residual signal, and fills the un-inverse-transformed region with the value "0" as the residual signal to generate the final residual block of the current block.

[0090] In addition, when applying MTS, the inverse transformer 530 determines a transform function or transform matrix to be applied in each of the horizontal and vertical directions by using the MTS information (mts_idx) signaled from the video coding device. The inverse transformer 530 also performs inverse transformation on the transform coefficients in the transform block in the horizontal and vertical directions by using the determined transform function.

[0091] The predictor 540 can include an intra predictor 542 and an inter predictor 544. When the prediction type of the current block is intra prediction, the intra predictor 542 is activated, and when the prediction type of the current block is inter prediction, the inter predictor 544 is activated.

[0092] The intra predictor 542 determines the intra prediction mode of the current block among multiple intra prediction modes according to the syntax element of the intra prediction mode extracted from the entropy decoder 510. The intra predictor 542 also predicts the current block by using the adjacent reference pixels of the current block according to the intra prediction mode.

[0093] The inter-frame predictor 544 determines the motion vector of the current block and the reference image for motion vector reference by using the syntax element of the inter-frame prediction mode extracted from the entropy decoder 510.

[0094] The adder 550 reconstructs the current block by adding the residual block output from the inverse transformer 530 to the prediction block output from the inter-frame predictor 544 or the intra-frame predictor 542. When performing intra-frame prediction on the block to be decoded subsequently, the pixels in the reconstructed current block are used as reference pixels.

[0095] The loop filter unit 560 as an in-loop filter may include a deblocking filter 562, a SAO filter 564, and an ALF 566. The deblocking filter 562 performs deblocking filtering on the boundaries between the reconstructed blocks to remove block artifacts that occur due to block-based unit decoding. The SAO filter 564 and the ALF 566 perform additional filtering on the reconstructed blocks after deblocking filtering to compensate for the difference between the reconstructed pixels and the original pixels that occurs due to lossy coding. The filter coefficients of the ALF are determined by using the information about the filter coefficients decoded from the bitstream.

[0096] The reconstructed blocks filtered by the deblocking filter 562, the SAO filter 564, and the ALF 566 are stored in the memory 570. When all the blocks in an image are reconstructed, the reconstructed image can be used as a reference image for inter-frame prediction of the blocks in the image to be encoded subsequently.

[0097] In some embodiments, the present invention relates to encoding and decoding video images as described above. More specifically, the present invention provides a video encoding and decoding method and apparatus that adaptively adjusts the pixel positions of the reference blocks of each sub-block based on the control point motion vectors of the affine model when performing per-sub-block prediction based on the affine model for the current block. The video encoding and decoding method and apparatus apply a motion compensation filter to the adjusted pixel positions.

[0098] The following embodiments may be performed by the inter-frame predictor 124 in a video encoding apparatus. The following embodiments may also be performed by the inter-frame predictor 544 in a video decoding apparatus.

[0099] The video encoding apparatus may generate signaling information associated with the present embodiment from the perspective of optimizing rate distortion when encoding the current block. The video encoding apparatus may encode the signaling information using the entropy encoder 155 and send the encoded signaling information to the video decoding apparatus. The video decoding apparatus may decode the signaling information associated with the decoding of the current block from the bitstream using the entropy decoder 510.

[0100] In the following description, the term "target block" may be used interchangeably with the current block or coding unit (CU), or may refer to some regions of the coding unit.

[0101] In addition, a value of true for a flag indicates the case where the flag is set to 1. Further, a value of false for a flag indicates the case where the flag is set to 0.

[0102] I. Affine Model-Based Prediction

[0103] As a measure to improve coding / decoding efficiency, after following the movement of a camera or an object in space and time, the affine model processes a changing object signal or a changing background signal in a video to derive the geometric relationship between the object signal or background signal and the camera or object, thereby modeling the relationship and applying the modeled relationship to a reference signal and an original signal. Theoretically, if the relationship is perfectly derived as a three-dimensional affine model representing the same object, perfect prediction can be made using this relationship. However, such perfect prediction is only achievable in theory. In video coding and decoding, prediction errors can be compensated by using a prediction signal based on modeled prediction and using a differential signal. The affine model-based prediction method has the effect that the prediction accuracy can bring about improved coding / decoding efficiency, but for the calculation of the affine model, it suffers from an increase in computational complexity.

[0104] Figure 6 is a schematic diagram showing the per-pixel calculation of an affine model according to at least one embodiment of the present invention.

[0105] To reduce computational complexity, Figure 6 an embodiment of uses the possible sameness or similarity of pixel values and motion information between the current block and adjacent blocks. The video decoding device uses the motion vectors of adjacent blocks at positions corresponding to the vertices (A, B, C) of the current block to predict the corresponding control point motion vectors (CPMVs). The video decoding device uses the control point motion vectors to model the geometric transformation relationship, i.e., the affine model, between the current block and the prediction block, and then performs the prediction of the current block based on the modeled transformation relationship. Figure 6 shows a 6-parameter model using control point motion vectors at three control points A, B, and C, but some embodiments may employ a 4-parameter model using control point motion vectors at two control points A and B or A and C. According to the 4-parameter model and the 6-parameter model, by using the control point motion vectors and the position of each pixel, the motion vectors (mv x ,mv y) can be respectively represented as shown in Equation 1 and Equation 2.

[0106] [Equation 1]

[0107]

[0108] [Equation 2]

[0109]

[0110] Here, W and H represent the width and height of the current block. (cpmv ix , cpmv iy ) represents the motion vector of the i-th control point. The predicted value for each pixel of the current block can be predicted by using the motion vector calculated according to Equation 1 or Equation 2. Although CPMV4 is not included in the affine model, CPMV4 can be defined for position D, as shown in the example of Figure 6 . Since the pixel at position D is not yet reconstructed, the implementation can use the motion vector of the co-located pixel in the reference picture for CPMV4.

[0111] In Figure 6 's example, CPMV represents the control point motion vector predictor in the affine AMVP mode (affine advanced motion vector prediction mode). In the affine merge mode, the motion vector difference (MVD) is not transmitted, so CPMV is the same as the control point motion vector (CPMV). In the affine AMVP mode, the motion vector difference is transmitted, so CPMV can be calculated by summing CPMV and MVD. Figure 6 's example shows a method of generating a control point motion vector predictor using the structure in the affine AMVP mode. The affine merge mode and the affine AMVP mode are described below.

[0112] Figure 7 is a schematic diagram showing the calculation of the affine model using the center vector according to at least one implementation of the present invention.

[0113] In addition, in order to reduce the computational complexity in some implementations, the video decoding device can perform block-by-block prediction by using each control point motion vector as the center vector, as shown in Figure 7 . At this time, by assuming that the current block has four control point motion vectors, the current block can be divided into four blocks. In the method shown in Figure 7 , the block with each control point motion vector as the center vector can be predicted according to the common motion vector. This can reduce the computational complexity, although the prediction accuracy may be lower compared to the implementation that performs per-pixel calculation.

[0114] Figure 8It is a schematic diagram showing the per-sub-block calculation of the affine model according to at least one embodiment of the present invention.

[0115] In another embodiment, in order to reduce the computational complexity, the video decoding device may perform per-sub-block prediction for each sub-block, as Figure 8 shown. If the horizontal or vertical size of the current block is greater than the horizontal or vertical size of the sub-block, the video decoding device may divide the current block into sub-blocks. The video decoding device may derive the control point motion vectors at each vertex position of the divided sub-blocks by using the control point motion vectors of the current block, and then derive the representative motion vectors of the control point motion vectors of each sub-block. The video decoding device generates a prediction block in units of sub-blocks by using the derived representative motion vectors, and then combines the per-sub-block prediction blocks to generate a first prediction block of the current block.

[0116] As another example, the video decoding device calculates the motion vectors for each sub-block by substituting the center position of each sub-block into (x, y) of Equation 1 or Equation 2. Here, the center position may be the actual center point of the sub-block, or may be the sample position at the lower right side of the center point. For example, for a sub-block with a size of 4×4 and the coordinates of the upper left sample being (0, 0), the center position of the sub-block may be (1.5, 1.5) or (2, 2). The video decoding device generates a prediction block in units of sub-blocks by using the derived motion vectors, and then combines the prediction blocks in units of sub-blocks to generate a first prediction block of the current block.

[0117] In some embodiments, the video decoding device may subject the first prediction block to filtering to generate a second prediction block. Then, the video decoding device may generate a final prediction block of the current block by using one of the first prediction block and the second prediction block.

[0118] To reduce the number of bits required for encoding the control point motion vectors, some embodiments may adopt the conventional inter-frame prediction (translation motion prediction) methods described above, namely the affine merge mode and the affine AMVP mode. Hereinafter, the affine merge mode and the affine AMVP mode are referred to as the control point motion vector prediction methods.

[0119] As an example, in the affine merge mode, the inter-frame predictor 124 of the video encoding device forms a list of a predefined number (e.g., 5) of affine merge candidates. First, the video encoding device derives the inherited affine merge candidates from the adjacent blocks of the target block. For example, the video encoding device generates a merge candidate list by deriving a predefined number of inherited affine merge candidates from Figure 4 the adjacent samples (A0, A1, B0, B1, B2) of the target block shown. Each of the inherited affine merge candidates included in the candidate list corresponds to a combination of two or three CPMVs.

[0120] The video encoding device derives inherited affine merge candidates from the control point motion vectors of adjacent blocks of a target block predicted in the affine mode. Some embodiments may limit the number of merge candidates derived from adjacent blocks predicted in the affine mode. For example, the video encoding device may derive two inherited affine merge candidates, one of A0 and A1 plus one of B0, B1, and B2, from adjacent blocks predicted in the affine mode. The priority may be in the order of A0, A1, then B0, B1, and B2.

[0121] On the other hand, if the total number of merge candidates is three or more, the video encoding device may additionally derive the insufficiently constructed affine merge candidates from the translational motion vectors of adjacent blocks, as Figure 6 shown in the example of.

[0122] The video encoding device derives control point motion vectors CPMV1, CPMV2, or CPMV3 from each of the adjacent block groups {B2, B3, A2}, the adjacent block group {B1, B0}, and the adjacent block group {A1, A0}. As an example, the priority within each adjacent block group may be in the order of B2, B3, A2, the order of B1, B0, and the order of A1, A0. In addition, the video encoding device derives another control point motion vector CPMV4 from the co-located block C0 in the reference picture. The video encoding device combines two or three of the four control point motion vectors to additionally generate the insufficiently constructed affine merge candidates. The priority of combination is as follows. List the elements within each group in the following order: the upper left, upper right, and lower left control point motion vectors.

[0123] {CPMV1, CPMV2, CPMV3}, {CPMV1, CPMV2, CPMV4}, {CPMV1, CPMV3, CPMV4},

[0124] {CPMV2, CPMV3, CPMV4}, {CPMV1, CPMV2}, {CPMV1, CPMV3}

[0125] If the merge candidate list cannot be filled by using the inherited affine merge candidates and the constructed affine merge candidates, the video encoding device may add a zero motion vector as a candidate.

[0126] The video encoding device selects a merge candidate from a merge candidate list from the perspective of optimizing the encoding and decoding efficiency, and determines a merge index indicating the selected merge candidate. The video encoding device performs affine motion prediction on a target block by using the selected merge candidate. If the merge candidate consists of two control point motion vectors, affine motion prediction is performed by using a 4-parameter model. On the other hand, if the merge candidate consists of three control point motion vectors, affine motion prediction is performed by using a 6-parameter model. The video encoding device encodes the merge index and signals the encoded merge index to the video decoding device.

[0127] The video decoding device decodes the merge index. The inter-frame predictor 544 of the video decoding device forms a merge candidate list in the same manner as the video encoding device, and performs affine motion prediction by using the control point motion vectors corresponding to the merge candidate indicated by the merge index.

[0128] As another example, in the affine AMVP mode, from the perspective of optimizing the encoding and decoding efficiency, the inter-frame predictor 124 of the video encoding device determines the form of the affine model for the target block and its accompanying actual control point motion vectors. For each control point, the video encoding device calculates the motion vector difference (MVD) and then encodes the MVD of each control point, where the MVD is the difference between the actual control point motion vector and each control point motion vector predictor (the MVP of each control point). To derive each control point motion vector predictor, the inter-frame predictor 124 forms a list of a predefined number (e.g., 2) of affine AMVP candidates. If the target block is of the 4-parameter type, the candidates included in the list each consist of a pair of two control point motion vectors. On the other hand, if the target block is of the 6-parameter type, the candidates in the list each consist of a set of three control point motion vectors.

[0129] A method for forming a candidate list in the affine AMVP mode is described below. The affine AMVP candidate list can be derived in a manner similar to the method for forming the affine merge candidate list described above.

[0130] The video encoding device checks whether the reference picture of the inherited affine AMVP candidate is the same as the reference picture of the current block. Here, the inherited affine AMVP candidate can be a block predicted in the affine mode among the adjacent blocks (A0, A1, B0, B1, B2) of the target block shown in Figure 4 as in the foregoing affine merge mode.

[0131] If the reference picture of the inherited affine AMVP candidate is the same as the reference picture of the current block, the video encoding device adds the corresponding inherited affine AMVP candidate.

[0132] On the other hand, if the reference picture of the inherited affine AMVP candidate is different from the reference picture of the current block, the video coding device checks whether the reference pictures of all CPMVs of the constructed affine AMVP candidate are the same as the reference picture of the current block. Here, all CPMVs of the constructed affine AMVP candidate can be derived from the motion vectors of adjacent samples shown in Figure 6 in the same way as in the above-mentioned affine merge mode. If the reference pictures of all CPMVs of the constructed affine AMVP candidate are the same as the reference picture of the current block, the video coding device adds the corresponding constructed affine AMVP candidate.

[0133] At this time, the affine model form of the target block needs to be considered. The video coding device checks whether the affine model of the target block is of the 4-parameter type. If so, two control point motion vectors, i.e., the motion vectors of the upper left and upper right control points of the target block, are derived by using the affine model of the adjacent blocks. If the affine model of the target block is of the 6-parameter type, the video coding device uses the affine model of the adjacent blocks to derive three control point motion vectors, i.e., the motion vectors of the upper left, upper right, and lower left control points of the target block.

[0134] If the reference pictures of all CPMVs are different from the reference picture of the current block, the video coding device adds a translational motion vector as an affine AMVP candidate.

[0135] If the candidate list cannot be filled, that is, even after using all the above steps, the preset number of candidates is not satisfied, the video coding device adds zero motion vectors for the affine AMVP candidates.

[0136] The video coding device selects a candidate from the affine AMVP list and determines the candidate index indicating the selected candidate. At this time, each of the control point motion vectors of the selected candidate corresponds to each of the control point motion vector predictors. From the perspective of optimizing the coding and decoding efficiency, the video coding device determines the actual control point motion vectors for each control point in the target block, and then calculates the MVD between the actual control point motion vectors and the control point motion vector predictors. The video coding device encodes the affine model form, candidate index, and MVD of each control point of the target block, and signals the encoded affine model form, encoded candidate index, and encoded MVD of each control point of the target block to the video decoding device.

[0137] The video decoding device decodes the affine model form, candidate index, and MVD for each control point. The inter-frame predictor 544 of the video decoding device generates an affine AMVP list in the same manner as the video encoding device, and selects a candidate indicated by the candidate index from the affine AMVP list. The video decoding device sums the motion vector predictor for each control point of the selected candidate and the corresponding MVD to reconstruct the motion vector for each control point. The video decoding device uses the reconstructed control point motion vectors to perform affine motion prediction.

[0138] The following embodiments are described with respect to a video decoding device, but they can also be implemented by a video encoding device in the same or similar manner.

[0139] II. Embodiments according to the present invention

[0140] Figure 9 is a schematic diagram showing a part of a video decoding device for prediction based on an affine model according to at least one embodiment of the present invention.

[0141] A video decoding device according to some embodiments of the present invention may determine an affine model, perform prediction of a current block based on the determined affine model, and finally generate a reconstructed block of the current block. It can be executed by the entropy decoder 510, inter-frame predictor 544, and adder 550 of the video decoding device Figure 9 shown. On the other hand, the same operations as shown in Figure 9 can be executed by the image segmenter 110, predictor 120, and adder 170 of the video encoding device. In this case, the video decoding device may utilize the encoded information parsed from the bitstream, while the video encoding device may utilize the encoded information set elsewhere at a higher level in terms of minimizing rate distortion. Hereinafter, for ease of description, the embodiments are described centering on the video decoding device.

[0142] As Figure 9 shown, the inter-frame predictor 544 may include all or part of an affine model determiner 910, a control point motion vector generator 920, a motion vector generator 930, a reference block pixel position calculator 940, and a prediction executor 950.

[0143] In some embodiments, the video encoding device transmits all or part of an affine model application flag, information about the affine model, information about the control point motion vectors, and a residual block. Here, the affine model application flag indicates whether the current block is a prediction block based on prediction using an affine model.

[0144] The entropy decoder 510 may decode all or part of an affine model application flag, affine model information, information on a control point motion vector prediction method, a residual block, and a reference picture index from a bitstream transmitted by a video encoding device.

[0145] The affine model application flag indicates whether the current block is a block predicted by affine model-based prediction. The affine prediction model information indicates the form of the affine model, i.e., whether it is a 4-parameter model or a 6-parameter model. The control point motion vector information includes information on control point motion vectors depending on the affine model, such as the number of control point motion vectors, the control point motion vector prediction method, the control point motion vector difference, etc. Here, the control point motion vector prediction method may be an affine merge mode or an affine AMVP mode.

[0146] If the affine model application flag regarding affine model-based prediction is true, the affine model determiner 910 may use the affine model information as a basis for determining the form of the affine model to derive the number of control point motion vectors.

[0147] The control point motion vector generator 920 predicts the control point motion vectors for the current block. Based on the relevant syntax and the prediction mode information of neighboring blocks decoded for the current block, the control point motion vector generator 920 determines the number of control point motion vectors in the current block and generates corresponding control point motion vector predictors. Here, the relevant syntax includes the control point motion vector prediction method, the reference picture index, etc.

[0148] After generating the control point motion vector predictors, the control point motion vector generator 920 adds the control point motion vector difference to the control point motion vector predictors, thereby calculating the control point motion vectors. If the control point motion vectors of the affine model have been transmitted in the affine merge mode, the process of adding the control point motion vector difference to the control point motion vector predictors may be omitted. Here, the affine merge mode refers to a method of determining the control point motion vectors in the same manner as neighboring vectors without a motion vector difference.

[0149] The motion vector generator 930 calculates the motion vectors by using the control point motion vectors in the units for the current block when calculating the motion vectors of the current block. For example, if a sub-block is a unit for calculating the motion vectors, the motion vectors may be calculated on a sub-block basis.

[0150] The reference block pixel position calculator 940 calculates the positions of the reference pixels of the reference block based on the precision of the motion vector. At this time, the integer position and the fractional position of the reference pixel can be calculated. The prediction executor 950 performs interpolation filtering on the reference pixels at the integer position. The prediction executor 950 performs motion compensation by applying an interpolation filter depending on the fractional position of the reference pixel to the reference pixel, thereby generating a predicted block of the current block. Then, the prediction executor 950 can sum the predicted block and the decoded residual block to generate a reconstructed block of the current block.

[0151] Figure 10a and Figure 10b is a schematic diagram showing motion compensation at fractional position pixels according to at least one embodiment of the present invention.

[0152] In an example case, the sub-block has a size of 4×4, and the motion vector of the sub-block is calculated based on the control point motion vector, as Figure 10a shown. By using the calculated motion vector, the video decoding device determines the integer position pixels of the reference block, as Figure 10b shown, and performs interpolation filtering on the integer position pixels to perform motion compensation based on the fractional position pixels. In other words, by using interpolation filtering, the video decoding device changes the integer position pixels of the reference block to fractional position pixels according to the precision of the motion vector. In Figure 10b the example of, the video decoding device calculates the integer position pixels and the fractional position pixels according to Equation 3.

[0153] [Equation 3]

[0154] xIntL = xSb + (mvLX[0] >> IFR) + xL - srRange

[0155] yIntL = ySb + (mvLY[0] >> IFR) + yL - srRange

[0156] xFracL = mvLX[0] & (2IFR – 1)

[0157] yFracL = mvLY[0] & (2IFR – 1)

[0158] In Equation 3, (xSb, ySb) are the integer coordinates of the upper left pixel of the sub-block. (mvLX[0], mvLY[0]) is the motion vector of the sub-block, and the Interpolation Filter Resolution (IFR) represents the precision of the motion vector. For example, if the prediction precision is 1 / 16 of a pixel, then the IFR is 4. (mvLX[0] >> IFR, mvLY[0] >> IFR) represents the integer part of the motion vector, and (xSb + (mvLX[0] >> IFR), ySb + (mvLY[0] >> IFR)) represents the integer coordinates of the upper left pixel of the reference block. (xL, yL) is the integer offset representing the position of the pixel within the reference block. For example, for a sub-block of size 4×4, 0 ≤ xL, yL ≤ 3. srRange represents the motion offset between the reference picture and the current picture, which can typically be 0. (xIntL, yIntL) represents the integer position of each pixel in the reference block. As shown in the examples of Figure 10a and Figure 10b , the integer positions of the pixels are consecutive.

[0159] On the other hand, when the present invention operates in the Intra Block Copy (IBC) mode, for the search range used for applying IBC, srRange is used to correct the search range. If the present invention operates in inter-frame prediction, then srRange becomes 0.

[0160] (xFracL, yFracL) represents the fractional part of the motion vector of the sub-block.

[0161] The interpolation filter can be determined according to the protocol between the video encoding device and the video decoding device. In this case, an interpolation filter with preset filter coefficients according to the prediction direction, inter-frame prediction mode, fractional position, etc. is used. Alternatively, the video encoding device can calculate the filter coefficients and can send the calculated filter coefficients to the video decoding device. The video decoding device can perform interpolation filtering by using the decoded filter coefficients.

[0162] On the other hand, for sub-blocks with a common motion vector, xFracL and yFracL are determined to be the same value. Therefore, as in the example of Figure 10b , the integer position pixels are compensated with the value of the pixels located at the same distance from the integer position pixels. The integer position pixels become the pixels represented by the hatched lines.

[0163] In an embodiment according to the present invention, after deriving geometric model parameters for each sub-block based on control point motion vectors, or after decoding geometric model parameters transmitted by a video encoding device, a video decoding device adaptively changes the integer positions of reference pixels using the geometric model parameters.

[0164] Figure 11a and Figure 11b is a schematic diagram showing motion compensation at fractional position pixels according to another embodiment of the present invention.

[0165] In one example, a video decoding device derives geometric model parameters (SubAPx, SubAPy) based on control point motion vectors, and uses the derived geometric model parameters to perform motion compensation for a sub-block by changing the integer positions of reference pixels.

[0166] By using Equation 1 or Equation 2, a video decoding device calculates motion vectors at the vertices of a sub-block according to the control point motion vector of the current block. The video decoding device uses the calculated motion vectors as the control point motion vector Sub_CPMV of the sub-block.

[0167] Sub_CPMV1, Sub_CPMV2, and Sub_CPMV3 represent the control point motion vectors used when deriving geometric model parameters for an arbitrary sub-block. Sub_CPMV1, Sub_CPMV2, and Sub_CPMV3 represent the upper left side CPMV, upper right side CPMV, and lower left side CPMV of the sub-block, respectively. In Figure 11a the example of, Sub_CPMV1, Sub_CPMV2, and Sub_CPMV3 represent the control point motion vectors of the upper left side sub-block. Therefore, Sub_CPMV1 of the upper left side sub-block is equal to CPMV1 of the current block. Additionally, Sub_CPMV2 of the upper right side sub-block is equal to CPMV2 of the current block, and Sub_CPMV3 of the lower left side sub-block is equal to CPMV3 of the current block.

[0168] In one example, based on the control point motion vector of a sub-block, a video decoding device derives geometric model parameters (SubAPx, SubAPy) of the sub-block according to Equation 4.

[0169] [Equation 4]

[0170]

[0171] In Equation 4, the horizontal components of Sub_CPMV1 and Sub_CPMV2 are used to calculate the horizontal component SubAPx of the geometric model parameter. For example, if the values of the control point motion vectors SubCPMV2x, Sub_CPMV1x, and mvLX[0] of the sub-block are different and mvLX[0] is non-zero, the video decoding device derives the horizontal component SubAPx of the geometric model parameter of the sub-block according to Equation 4. Otherwise, SubAPx is set to 1. Additionally, the vertical components of Sub_CPMV1 and Sub_CPMV3 are used to calculate the vertical component SubAPy of the geometric model parameter. If the values of the control point motion vectors SubCPMV3y, Sub_CPMV1y, and mvLY[0] of the sub-block are different and mvLY[0] is non-zero, the video decoding device derives the vertical component SubAPy of the geometric model parameter of the sub-block according to Equation 4. Otherwise, SubAPy is set to 1. According to the cropping function, the horizontal and vertical components of the geometric model parameter each have a value greater than or equal to 1.

[0172] As another example, based on the control point motion vectors of the sub-block, the video decoding device derives the geometric model parameters (SubAPx, SubAPx) of the sub-block according to Equation 5.

[0173] [Equation 5]

[0174]

[0175] In Equation 5, the horizontal components of Sub_CPMV1 and Sub_CPMV2 are used to calculate the horizontal component SubAPx of the geometric model parameter. For example, if the values of the control point motion vectors SubCPMV2x, Sub_CPMV1x, and mvLX[0] of the sub-block are different and mvLX[0] is non-zero, the video decoding device derives the horizontal component SubAPx of the geometric model parameter of the sub-block according to Equation 5. Additionally, the vertical components of Sub_CPMV1 and Sub_CPMV3 are used to calculate the vertical component SubAPy of the geometric model parameter. If the values of the control point motion vectors SubCPMV3y, Sub_CPMV1y, and mvLY[0] of the sub-block are different and mvLY[0] is non-zero, the video decoding device derives the vertical component SubAPy of the geometric model parameter of the sub-block according to Equation 5. Depending on the cropping function, the horizontal and vertical components of the geometric model parameter each have a value greater than or equal to 1.

[0176] The video decoding device uses the derived geometric model parameters to determine the integer position pixels of the reference block, as shown in the examples of Figure 11a and Figure 11b and performs interpolation filtering on those pixels to perform motion compensation based on fractional position pixels. Figure 11bThe example of [[ID=]] shows interpolation in the horizontal direction with the vertical direction component fixed. On the other hand, the video decoding device calculates the integer position pixels and the fractional position pixels according to Equation 6.

[0177] [Equation 6]

[0178] xIntL = xSb + (mvLX[0] >> IFR) + xL × SubAPx – srRange

[0179] yIntL = ySb + (mvLY[0] >> IFR) + yL × SubAPy - srRange

[0180] xFracL = mvLX[0] & (2IFR – 1)

[0181] yFracL = mvLY[0] & (2IFR – 1)

[0182] In Equation 6, as described above, (xSb + (mvLX[0] >> IFR), ySb + (mvLY[0] >> IFR)) represents the integer coordinates of the upper left pixel of the reference block. (xL, yL) is the integer offset representing the position of the pixel within the reference block, which is scaled to (xL × SubAPx, yL × SubAPy) by the geometric model parameters. Therefore, (xIntL, yIntL) represents the integer position of each pixel in the reference block. Compared with the consecutive integer position pixels in the example of [[ID=]] Figure 10a and Figure 10b , the example of [[ID=]] Figure 11a and Figure 11b shows integer position pixels that may not be consecutive due to the application of (SubAPx, SubAPy) according to Equation 6. Therefore, the reference block containing the integer position pixels according to [[ID=]] Figure 11a can be larger than the reference block containing the integer position pixels according to the example of [[ID=]] Figure 10a .

[0183] On the other hand, the interpolation filter can be determined according to the protocol between the video encoding device and the video decoding device. The interpolation filter used has preset filter coefficients according to the prediction direction, inter-frame prediction mode, fractional position, etc. Alternatively, the video encoding device can calculate the filter coefficients and can send the calculated filter coefficients to the video decoding device. The video decoding device can perform interpolation filtering by using the decoded filter coefficients. Figure 11b The example of interpolation in the horizontal direction is described, but interpolation in the vertical direction can also be performed in the same way.

[0184] As another example, the video encoding device may calculate the optimal geometric model parameters (SubAPx, SubAPy) and may send the calculated geometric model parameters to the video decoding device. For transmission efficiency, the video encoding device may organize a set of geometric model parameters (SubAPx, SubAPy) of sub-blocks included in the current block into a geometric model parameter list, and may send an index indicating the geometric model parameters in the list for each sub-block to the video decoding device. The video decoding device decodes the index and derives the geometric model parameters of each sub-block from the list according to the index. The video decoding device may perform interpolation filtering on the sub-block by using the derived geometric model parameters and Equation 6. Alternatively, a mapping table may be preset in the video encoding device and the video decoding device in the form of geometric model parameters according to the size of the current block, the size of the reference block, and the difference of the control point motion vectors of the current block. The video encoding device and the video decoding device may derive a mapping index according to the sharing method, and then may perform interpolation filtering by using the geometric model parameters indicated by the mapping index.

[0185] As another example, when the fallback mode is applied to the affine model, the video decoding device may limit the integer positions calculated according to the geometric model parameters within the range in the fallback mode. The fallback mode restricts the region to which the control point motion vector of the current block is applied to a preset range. To limit the application region within the preset range in the fallback mode, the video decoding device may forcefully change the integer positions outside the preset range to the last integer position within the restricted range. Alternatively, when the fallback mode is applied and the interpolation according to the method of the present invention results in a violation of the preset range in the fallback mode, the video decoding device sets SubAPx and SubAPy to 1 and then performs interpolation filtering. Alternatively, when the fallback mode is applied, the video decoding device unconditionally sets SubAPx and SubAPy to 1 regardless of the preset range in the fallback mode and then performs interpolation filtering.

[0186] Hereinafter, with reference to Figure 12 and Figure 13 , a method for encoding and decoding a current block according to prediction based on an affine model will be described.

[0187] Figure 12 is a flowchart of a method for encoding a current block by a video encoding device according to at least one embodiment of the present invention.

[0188] The video encoding device determines the affine model information of the current block and the prediction method for controlling the motion vector of the points (S1200). Here, the affine model information refers to the form of the affine model, that is, whether it is a 4-parameter model or a 6-parameter model. The prediction method for controlling the motion vector of the points refers to the affine merge mode or the affine AMVP mode. From the aspect of rate-distortion optimization, the affine model information and the prediction method for controlling the motion vector of the points can be determined.

[0189] The video encoding device determines the form of the affine model based on the affine model information (S1202).

[0190] The video encoding device derives the motion vector of the control points of the current block based on the affine model (S1204).

[0191] The video encoding device uses the motion vector of the control points of the current block to generate the motion vector of the sub-blocks (S1206). Here, the sub-blocks are generated by partitioning the current block.

[0192] The video encoding device determines the geometric model parameters of the sub-blocks (S1208). Here, the geometric model parameters are defined as shown in Equation 4 based on the motion vector of the control points of the sub-blocks, the motion vector of the sub-blocks, and the accuracy of the motion vector of the sub-blocks.

[0193] For example, the video encoding device calculates the motion vector of the control points of the sub-blocks based on the motion vector of the control points of the current block according to the affine model. The video encoding device derives the geometric model parameters by using the motion vector of the sub-blocks and the motion vector of the control points of the sub-blocks.

[0194] The video encoding device generates a reference block from the reference picture by using the motion vector of the sub-blocks (S1210).

[0195] The video encoding device calculates the position of the reference pixels in the reference block based on the geometric model parameters (S1212).

[0196] The video encoding device applies an interpolation filter to the reference pixels to generate a predicted value of the sub-block, thereby generating a predicted block of the current block (S1214).

[0197] The video encoding device encodes the affine model information and the prediction method for controlling the motion vector of the points (S1216).

[0198] The video encoding device can derive an index indicating the geometric model parameters from a preset list of geometric model parameters and encode the derived index. In addition, the video encoding device can generate a residual block by subtracting the predicted block from the original block of the current block and encode the generated residual block.

[0199] Figure 13 is a flowchart of a method for reconstructing a current block by a video decoding device according to at least one embodiment of the present invention.

[0200] The video decoding device decodes the affine model information and the control point motion vector information of the current block from the bitstream (S1300). Here, the affine model information refers to the form of the affine model, that is, whether it is a 4-parameter model or a 6-parameter model. The control point motion vector information includes the prediction method for the control point motion vector and the control point motion vector difference. The prediction method for the control point motion vector refers to the affine merge mode or the affine AMVP mode.

[0201] The video decoding device determines the form of the affine model based on the affine model information (S1302).

[0202] The video decoding device derives the control point motion vector of the current block based on the affine model and the control point motion vector information (S1304).

[0203] The video decoding device uses the control point motion vector of the current block to generate the motion vector of the sub-block (S1306). Here, the sub-blocks are generated by partitioning the current block.

[0204] The video decoding device obtains the geometric model parameters of the sub-block (S1308). Here, the geometric model parameters are defined as shown in Equation 4 based on the control point motion vector of the sub-block, the motion vector of the sub-block, and the accuracy of the motion vector of the sub-block.

[0205] For example, the video decoding device calculates the control point motion vector of the sub-block based on the affine model according to the control point motion vector of the current block. The video decoding device derives the geometric model parameters by using the motion vector of the sub-block and the control point motion vector of the sub-block.

[0206] As another example, the video decoding device decodes an index indicating the geometric model parameters from the bitstream. The video decoding device derives the geometric model parameters based on the decoded index from a preset list of geometric model parameters.

[0207] The video decoding device generates a reference block from the reference picture by using the motion vector of the sub-block (S1310).

[0208] The video decoding device calculates the position of the reference pixel in the reference block based on the geometric model parameters (S1312).

[0209] The video decoding device applies an interpolation filter to the reference pixel to generate a predicted value of the sub-block, thereby generating a predicted block of the current block (S1314).

[0210] The video decoding device decodes the residual block from the bitstream. Then, the video decoding device can sum the predicted block and the decoded residual block to generate a reconstructed block of the current block.

[0211] Although the steps in the respective flowcharts described are executed sequentially, these steps merely illustrate the technical ideas of some embodiments of the present invention. Therefore, those of ordinary skill in the art to which the present invention pertains can execute the steps by changing the order described in the respective drawings or by executing two or more steps in parallel. Accordingly, the steps in the respective flowcharts are not limited to the order shown in the order of occurrence.

[0212] It should be understood that the above description presents illustrative embodiments that can be implemented in various other ways. The functions described in some embodiments can be implemented by hardware, software, firmware, and / or combinations thereof. It should also be understood that the functional components described in the present invention are labeled as "…… unit" to highlight the possibility of their independent implementation.

[0213] On the other hand, the various methods or functions described in some embodiments can be implemented as instructions stored in a non-volatile recording medium, which can be read and executed by one or more processors. The non-volatile recording medium can include, for example, various types of recording devices that store data in a form readable by a computer system. For example, the non-volatile recording medium can include storage media such as erasable programmable read-only memory (EPROM), flash drives, optical disk drives, magnetic hard disk drives, and solid-state drives (SSD), etc.

[0214] Although the exemplary embodiments of the present invention have been described for illustrative purposes, those of ordinary skill in the art to which the present invention pertains should understand that various modifications, additions, and substitutions can be made without departing from the spirit and scope of the present invention. Therefore, the embodiments of the present invention have been described for the sake of brevity and clarity. The scope of the technical ideas of the embodiments of the present invention is not limited by the illustration. Accordingly, those of ordinary skill in the art to which the present invention pertains should understand that the scope of the present invention should not be limited by the embodiments clearly described above, but by the claims and their equivalents.

[0215] Reference Numerals

[0216] 124: Inter-Frame Predictor

[0217] 544: Inter-Frame Predictor

[0218] 920: Control Point Motion Vector Generator

[0219] 930: Motion Vector Generator

[0220] 940: Reference Block Pixel Position Calculator

[0221] 950: Prediction Executor.

[0222] Cross-Reference to Related Applications

[0223] This application claims the priority and benefit of Korean Patent Application No. 10-2022-0162309, filed on November 29, 2022, and Korean Patent Application No. 10-2023-0163967, filed on November 23, 2023, the entire contents of each of which are incorporated herein by reference.

Claims

1. A method for decoding a current block by a video decoding device, the method comprising: Decoding affine model information and control point motion vector information of the current block from a bitstream, the affine model information indicating a form of an affine model, and the control point motion vector information including a prediction method for a control point motion vector of the current block and a control point motion vector difference; Determining the form of the affine model based on the affine model information; Deriving a control point motion vector of the current block based on the affine model and the control point motion vector information; Generating a motion vector of a sub-block by using the control point motion vector of the current block, the sub-block being generated based on a partition of the current block; Obtaining geometric model parameters of the sub-block, the geometric model parameters being defined based on the control point motion vector of the sub-block and the motion vector of the sub-block; Generating a reference block according to a reference picture by using the motion vector of the sub-block; Calculating positions of reference pixels in the reference block based on the geometric model parameters; And Generating a prediction value of the sub-block by applying an interpolation filter to the reference pixels to generate a prediction block of the current block.

2. The method according to claim 1, wherein, Obtaining the geometric model parameters includes: Calculating a control point motion vector of the sub-block based on the control point motion vector of the current block according to the affine model; and Deriving the geometric model parameters by using the motion vector of the sub-block and the control point motion vector of the sub-block.

3. The method according to claim 2, wherein, Obtaining the geometric model parameters includes: Decoding an index indicating the geometric model parameters from the bitstream; and Deriving the geometric model parameters from a preset list of geometric model parameters according to the index.

4. The method according to claim 1, wherein The geometric model parameters include a horizontal component and a vertical component, and the horizontal component and the vertical component have a minimum value of 1.

5. The method according to claim 1, wherein Calculating the positions of the reference pixels includes: Scaling an integer offset representing a position of a representative pixel within the reference block by using the geometric model parameters; and Calculating an integer position of the reference pixel by using the scaled integer offset and an integer part of the motion vector.

6. The method according to claim 5, wherein, Generating the prediction block includes: Generating a fractional part of the motion vector by using a precision of the motion vector of the sub-block; Obtaining an interpolation filter depending on the fractional part; and Applying the interpolation filter to the reference pixels existing at the integer positions.

7. The method according to claim 6, wherein, Obtaining the interpolation filter includes: Setting the interpolation filter to a preset interpolation filter.

8. The method according to claim 6, wherein Obtaining the interpolation filter includes: Decoding coefficients of the interpolation filter from the bitstream.

9. The method according to claim 5, wherein, Calculating the positions of the reference pixels includes: Restricting the integer position within a preset range in response to an application of a fallback mode that restricts a region to be subjected to the control point motion vector of the current block to a preset range.

10. A method for encoding a current block by a video encoding device, the method comprising: Determining affine model information of the current block and a prediction method for a control point motion vector, the affine model information indicating a form of an affine model; Determining the form of the affine model based on the affine model information; Deriving a control point motion vector of the current block based on the affine model; Generating a motion vector of a sub-block by using the control point motion vector of the current block, the sub-block being generated based on a partition of the current block; Determining geometric model parameters of the sub-block, the geometric model parameters being defined based on the control point motion vector of the sub-block and the motion vector of the sub-block; Generating a reference block based on a reference picture by utilizing motion vectors of sub - blocks; Calculating the positions of reference pixels in the reference block based on geometric model parameters; Generating a prediction value of the sub - block by applying an interpolation filter to the reference pixels to generate a prediction block of the current block; And Encoding affine model information and a prediction method for control - point motion vectors.

11. The method according to claim 10, wherein, Determining geometric model parameters includes: Calculating the control - point motion vectors of the sub - blocks based on the control - point motion vectors of the current block according to the affine model; and Deriving geometric model parameters by utilizing the motion vectors of the sub - blocks and the control - point motion vectors of the sub - blocks.

12. The method according to claim 11, further comprising: Deriving an index indicating geometric model parameters from a preset list of geometric model parameters; And Encoding the index.

13. The method according to claim 10, wherein, Calculating the positions of reference pixels includes: Scaling an integer offset representing the position of a pixel within the reference block by utilizing geometric model parameters; and Calculating the integer position of the reference pixel by utilizing the scaled integer offset and the integer part of the motion vector.

14. The method according to claim 13, wherein, Generating a prediction block includes: Generating a fractional part of the motion vector by utilizing the precision of the motion vector of the sub - block; Obtaining an interpolation filter depending on the fractional part; and Applying the interpolation filter to the reference pixels existing at the integer positions.

15. The method according to claim 14, further comprising: Encoding the coefficients of the interpolation filter.

16. A computer - readable recording medium storing a bitstream generated by a video encoding method, the video encoding method comprising: Determining affine model information of a current block and a prediction method for control - point motion vectors, the affine model information indicating the form of the affine model; Determining the form of the affine model based on the affine model information; Deriving control - point motion vectors of the current block based on the affine model; Generating motion vectors of sub - blocks by utilizing the control - point motion vectors of the current block, the sub - blocks being generated based on the partitioning of the current block; Determining geometric model parameters of the sub - blocks, the geometric model parameters being defined based on the control - point motion vectors of the sub - blocks and the motion vectors of the sub - blocks; Generating a reference block based on a reference picture by utilizing the motion vectors of the sub - blocks; Calculating the positions of reference pixels in the reference block based on the geometric model parameters; Generating a prediction block of the current block by applying an interpolation filter to the reference pixels to generate prediction values of the sub - blocks; And Encoding affine model information and a prediction method for control - point motion vectors.

Citation Information

Patent Citations

  • Safety Handle For Blind Operation

    KR1020220162309A

  • Apparatus for reducing moisture of front opening unified pod in load port module and semiconductor process device comprising the same

    KR1020230163967A