Method and apparatus for inter prediction in video processing system

By adopting affine motion prediction method in video encoding, the motion vector difference between control points and adjacent blocks is used to improve the encoding efficiency of high-resolution images, the problem of high-resolution image transmission and storage costs is solved, and more efficient video encoding is achieved.

CN120343281APending Publication Date: 2025-07-18HONOR DEVICE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411238681.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2018-04-13
Filing Date
2019-04-11
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The prior art uses high-resolution and high-quality images to process high-quality images, and the transmission and storage cost of image data is high, and video encoding efficiency needs to be improved.

Method used

Affine motion prediction method is adopted to derive the control point (CP) of the current block, and to derive the motion vector predictor (MVP) and motion vector difference (MVD) using adjacent blocks, and to derive more accurate sample unit motion vectors based on the difference value, reducing the amount of motion vector data to improve coding efficiency.

Benefits of technology

It significantly improves inter prediction efficiency, reduces the amount of motion vector data, improves the overall encoding efficiency, and reduces transmission and storage costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120343281A_ABST
    Figure CN120343281A_ABST
Patent Text Reader

Abstract

The invention relates to a method and an apparatus for inter prediction in a video processing system. Disclosed is a method for inter prediction, comprising the steps of deriving a control point (CP) for a current block, where the CPs include a first CP and a second CP, deriving a first motion vector predictor (MVP) for the first CP and a second MVP for the second CP based on a neighboring block of the current block, decoding a first motion vector difference (MVD) for the first CP, decoding a difference (DMVD) of two MVDs for the second CP, and decoding a second MVP for the second CP based on a neighboring block of the current block. And deriving a first motion vector (MV) for the first CP based on the first MVP and the first MVD, deriving a second MV for the second CP based on the second MVP and the DMVD for the second CP, and generating a prediction block for the current block based on the first MV and the second MV.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the patent application with the application number 201980032706.7 (PCT / KR2019 / 004334), international filing date of April 11, 2019, and invention title of "Method and Apparatus for Inter-Frame Prediction in a Video Processing System", which was filed on November 16, 2020. Technical Field

[0002] The present disclosure generally relates to video coding technology, and more particularly, to an inter-frame prediction method and apparatus in a video processing system. Background Art

[0003] The demand for high-resolution and high-quality images such as high-definition (HD) images and ultra-high-definition (UHD) images is increasing in various fields. Since the image data has high resolution and high quality, the amount of information or bits to be transmitted increases relative to traditional image data. Therefore, when transmitting image data using a medium such as a traditional wired / wireless broadband line or storing image data using an existing storage medium, the transmission cost and storage cost increase.

[0004] Therefore, an efficient image compression technology for effectively transmitting, storing, and reproducing information of high-resolution and high-quality images is needed. Summary of the Invention

[0005] The technical object of the present disclosure is to provide a method and apparatus for increasing video coding efficiency.

[0006] Another technical object of the present disclosure is to provide a method and apparatus for processing an image using affine motion prediction.

[0007] Another technical object of the present disclosure is to provide a method and apparatus for performing inter-frame prediction based on a sample unit motion vector.

[0008] Another technical object of the present disclosure is to provide a method and apparatus for deriving a sample unit motion vector based on a motion vector of a control point for a current block.

[0009] Another technical object of the present disclosure is to provide a method and apparatus for improving coding efficiency by using a difference between differences of motion vector differences of control points of a current block.

[0010] Another technical object of the present disclosure is to provide a method and apparatus for deriving a motion vector predictor of another control point based on a motion vector predictor of a control point.

[0011] Another technical problem of the present disclosure is to provide a method and apparatus for deriving a motion vector predictor of a control point based on a motion vector of a reference region adjacent to the control point.

[0012] An embodiment of the present disclosure provides an inter prediction method executed by a decoding device. The inter prediction method includes: deriving control points (CPs) for a current block, where the CPs include a first CP and a second CP; deriving a first motion vector predictor (MVP) for the first CP and a second MVP for the second CP based on neighboring blocks of the current block, decoding a first motion vector difference (MVD) for the first CP, decoding a difference (DMVD) between two MVDs for the second CP, deriving a first motion vector (MV) for the first CP based on the first MVP and the first MVD, deriving a second MV for the second CP based on the second MVP and the DMVD for the second CP, and generating a prediction block for the current block based on the first MV and the second MV, where the DMVD for the second CP represents the difference between the first MVD and the second MVD for the second CP.

[0013] According to another example of the present disclosure, a video encoding method executed by an encoding device is provided. The encoding method includes: deriving control points (CPs) for a current block, where the CPs include a first CP and a second CP; deriving a first motion vector predictor (MVP) for the first CP and a second MVP for the second CP based on neighboring blocks of the current block, deriving a first motion vector difference (MVD) for the first CP, deriving a difference (DMVD) between two MVDs for the second CP, and encoding image information including information about the first MVD and information about the DMVD for the second CP to output a bitstream, where the DMVD for the second CP represents the difference between the first MVD and the second MVD for the second CP.

[0014] According to still another embodiment of the present disclosure, a decoding device for executing an inter prediction method is provided. The decoding device includes an entropy decoder that decodes a first motion vector difference (MVD) for the first CP and a difference (DMVD) between two MVDs for the second CP; and a predictor that derives control points (CPs) for a current block, including a first CP and a second CP, derives a first motion vector predictor (MVP) for the first CP and a second MVP for the second CP based on neighboring blocks of the current block, derives a first motion vector (MV) for the first CP based on the first MVP and the first MVD, derives a second MV for the second CP based on the second MVP and the DMVD for the second CP, and generates a prediction block for the current block based on the first MV and the second MV, where the DMVD for the second CP represents the difference between the first MVD and the second MVD for the second CP.

[0015] According to yet another embodiment of the present disclosure, there is provided an encoding apparatus that performs video encoding. The encoding apparatus includes a predictor that derives control points (CPs) for a current block, including a first CP and a second CP, derives a first motion vector predictor (MVP) for the first CP and a second MVP for the second CP based on neighboring blocks of the current block, derives a first motion vector difference (MVD) for the first CP, and derives a difference (DMVD) between two MVDs for the second CP, and an entropy encoder that encodes picture information including information about the first MVD and information about the DMVD for the second CP to output a bitstream, where the DMVD for the second CP represents the difference between the first MVD and a second MVD for the second CP.

[0016] According to the present disclosure, a more accurate sample unit motion vector for a current block can be derived, and the inter prediction efficiency can be significantly increased.

[0017] According to the present disclosure, a motion vector of a sample of a current block can be effectively derived based on motion vectors of control points for the current block.

[0018] According to the present disclosure, by sending the difference between motion vectors of control points for a current block and / or the difference between the differences of motion vectors, the amount of data for motion vectors of control points can be removed or reduced, and the overall encoding efficiency can be improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 is a block diagram schematically illustrating a video encoding apparatus according to an embodiment of the present disclosure.

[0020] Figure 2 is a block diagram schematically illustrating a video decoding apparatus according to an embodiment of the present disclosure.

[0021] Figure 3 Illustratively represents a content streaming system according to an embodiment of the present disclosure.

[0022] Figure 4 is a flowchart for illustrating a method of deriving a motion vector prediction value from neighboring blocks according to an embodiment of the present disclosure.

[0023] Figure 5 Illustratively represents an affine motion model according to an embodiment of the present disclosure.

[0024] Figure 6 Illustratively represents a simplified affine motion model according to an embodiment of the present disclosure.

[0025] Figure 7 is a diagram for describing a method of deriving a motion vector predictor at a control point according to an embodiment of the present disclosure.

[0026] Figure 8 Illustratively shows two CPs for a four-parameter affine motion model according to an embodiment of the present disclosure.

[0027] Figure 9 Illustratively shows a case of additionally using a median in a four-parameter affine motion model according to an embodiment of the present disclosure.

[0028] Figure 10 Illustratively shows three CPs for a six-parameter affine motion model according to an embodiment of the present disclosure.

[0029] Figure 11 Schematically shows a video encoding method of an encoding device according to the present disclosure.

[0030] Figure 12 Schematically illustrates an inter-frame prediction method of a decoding device according to the present disclosure. Detailed implementation

[0031] Although the present disclosure may be modified in various forms, specific embodiments of the present disclosure will be described in detail and illustrated in the drawings. However, this is not intended to limit the present disclosure to specific embodiments. The terms used in this specification are only used to describe specific embodiments and are not intended to limit the technical idea of the present disclosure. Unless the context clearly indicates otherwise, the singular form may include the plural form. Terms such as "including (or comprising)" and "having (or being provided with)" are intended to indicate the presence of features, numbers, steps, operations, components, parts, or combinations thereof written in the following description, and thus should not be construed as precluding the possibility of the existence or addition of one or more different features, numbers, steps, operations, components, parts, or combinations thereof.

[0032] Meanwhile, for the convenience of explaining different specific functions, the configurations in the drawings described in the present disclosure are independently drawn in a video encoding device / decoding device, but this does not mean that the configuration is implemented by independent hardware or independent software. For example, two or more configurations may be combined to form a single configuration, and one configuration may be divided into multiple configurations. Embodiments in which configurations are combined and / or configurations are divided without departing from the concept of the present disclosure also belong to the present disclosure.

[0033] In the present disclosure, the terms " / " and "," should be interpreted as indicating "and / or". For example, the expression "A / B" may mean "A and / or B", and "A, B" may mean "A and / or B". In addition, "A / B / C" may mean "at least one of A, B, and / or C". Additionally, "A, B, C" may mean "at least one of A, B, and / or C".

[0034] In addition, in the present disclosure, the term "or" should be construed to indicate "and / or". For example, the expression "A or B" may include 1) only A, 2) only B, and / or 3) both A and B. In other words, the term "or" in this document may be construed to indicate "additionally or alternatively".

[0035] The present disclosure may be modified in various forms, and its specific embodiments will be described and illustrated in the accompanying drawings. However, these embodiments are not intended to limit the present disclosure. The terms used in the following description are only for describing specific embodiments and are not intended to limit the present disclosure. Singular expressions include plural expressions as long as they are clearly and differently read. Terms such as "including" and "having" are intended to indicate the presence of features, numbers, steps, operations, elements, components, or combinations thereof used in the following description, and thus it should be understood that there is no exclusion of the possibility of the presence or addition of one or more different features, numbers, steps, operations, elements, components, or combinations thereof.

[0036] In addition, for the convenience of explaining different specific functions, the elements in the drawings described in this embodiment are independently drawn, which does not mean that these elements are implemented by independent hardware or independent software. For example, two or more of the elements may be combined to form a single element, or one element may be divided into multiple elements. Embodiments in which the elements are combined and / or divided without departing from the concept of this embodiment belong to the present disclosure.

[0037] The following description can be applied to the technical field of processing videos, images, or pictures. For example, the methods or exemplary embodiments disclosed in the following description may be associated with the disclosure of the common video coding (VVC) standard (Recommendation ITU-T H.266), the next-generation video / image coding standard after VVC, or the standards before VVC (e.g., the high efficiency video coding (HEVC) standard (Recommendation ITU-T H.265), etc.).

[0038] Hereinafter, examples of this embodiment will be described in detail with reference to the accompanying drawings. In addition, throughout the drawings, like reference numerals are used to indicate like elements, and the same description of like elements will be omitted.

[0039] In the present disclosure, a video may mean a collection of a series of images over time. Generally, a picture means a unit representing an image at a specific time, and a slice is a unit that forms a part of a picture. A picture may be composed of multiple slices, and the terms picture and slice may be mixed with each other as needed according to the occasion.

[0040] A pixel or pel can mean the smallest unit that makes up a picture (or image). Additionally, the term "sample" can be used as a term corresponding to a pixel. A sample can generally represent a pixel or the value of a pixel, can represent only the pixel (pixel value) of the luminance component, and can represent only the pixel (pixel value) of the chrominance component.

[0041] A unit indicates a basic unit of image processing. A unit can include at least one of a specific region and information related to that region. Optionally, a unit can be mixed with terms such as a block, a region, etc. Typically, an M×N block can represent a set of samples or transform coefficients arranged in M columns and N rows.

[0042] Figure 1 is a block diagram briefly illustrating the structure of an encoding device according to an embodiment of the present disclosure. Hereinafter, the encoding / decoding device can include a video encoding / decoding device and / or an image encoding / decoding device, and the video encoding / decoding device can be used as a concept including the image encoding / decoding device, or the image encoding / decoding device can be used as a concept including the video encoding / decoding device.

[0043] Reference Figure 1 , the video encoding device 100 can include a picture splitter 105, a predictor 110, a residual processor 120, an entropy encoder 130, an adder 140, a filter 150, and a memory 160. The residual processor 120 can include a subtractor 121, a transformer 122, a quantizer 123, a reorderer 124, an inverse quantizer 125, and an inverse transformer 126.

[0044] The picture splitter 105 can separate an input picture into at least one processing unit.

[0045] In an example, the processing unit can be referred to as a coding unit (CU). In this case, the coding unit can be recursively separated from the largest coding unit (LCU) according to a quadtree binary tree (QTBT) structure. For example, one coding unit can be separated into multiple coding units of a deeper depth based on a quadtree structure, a binary tree structure, and / or a ternary tree structure. In this case, for example, the quadtree structure can be applied first, and the binary tree structure and the ternary tree structure can be applied later. Alternatively, the binary tree structure / ternary tree structure can be applied first. The encoding process according to the present embodiment can be performed based on the final coding unit that is no longer further separated. In this case, the largest coding unit can be used as the final coding unit based on image characteristics such as coding efficiency, or the coding unit can be recursively separated into coding units of a lower depth as needed and the coding unit with the optimal size can be used as the final coding unit. Here, the encoding process can include processes such as prediction, transformation, and reconstruction, which will be described later.

[0046] In another example, the processing unit may include an encoding unit (CU), a prediction unit (PU), or a transformer (TU). The encoding unit may be separated from the largest coding unit (LCU) into coding units of a deeper depth according to a quadtree structure. In this case, the largest coding unit may be directly used as the final coding unit based on image characteristics such as coding efficiency, or the coding unit may be recursively separated into coding units of a deeper depth as needed, and the coding unit with the optimal size may be used as the final coding unit. When the smallest coding unit (SCU) is set, the coding unit may not be separated into coding units smaller than the smallest coding unit. Here, the final coding unit refers to the coding unit that is split or separated into a prediction unit or a transformer. The prediction unit is a unit split from the coding unit and may be a unit for sample prediction. Here, the prediction unit may be divided into sub-blocks. The transformer may be divided from the coding unit according to a quadtree structure, and the transformer may be a unit for deriving transform coefficients and / or a unit for deriving a residual signal from the transform coefficients. Hereinafter, the coding unit may be referred to as a coding block (CB), the prediction unit may be referred to as a prediction block (PB), and the transformer may be referred to as a transform block (TB). The prediction block or the prediction unit may refer to a specific area in the form of a block in a picture and includes an array of prediction samples. In addition, the transform block or the transformer may refer to a specific area in the form of a block in a picture and includes an array of transform coefficients or residual samples.

[0047] The predictor 110 may perform prediction on a block to be processed (hereinafter, it may represent a current block or a residual block) and may generate a prediction block including prediction samples for the current block. The unit for performing prediction in the predictor 110 may be a coding block, or may be a transform block, or may be a prediction block.

[0048] The predictor 110 may determine whether to apply intra prediction or inter prediction to the current block. For example, the predictor 110 may determine whether to apply intra prediction or inter prediction in units of CUs.

[0049] In the case of intra prediction, the predictor 110 may derive a prediction sample for a current block based on reference samples outside the current block in a picture (hereinafter, the current picture) to which the current block belongs. In this case, the predictor 110 may derive a prediction sample based on an average value or interpolation of neighboring reference samples of the current block (case (i)), or may derive a prediction sample based on reference samples existing in a specific (prediction) direction with respect to the prediction sample among the neighboring reference samples of the current block (case (ii)). Case (i) may be referred to as a non-directional mode or a non-angular mode, and case (ii) may be referred to as a directional mode or an angular mode. In intra prediction, the prediction mode may include, as examples, 33 directional modes and at least two non-directional modes. The non-directional modes may include a DC mode and a planar mode. The predictor 110 may determine a prediction mode to be applied to the current block by using a prediction mode applied to an adjacent block.

[0050] In the case of inter prediction, the predictor 110 may derive a prediction sample for a current block based on samples specified by a motion vector on a reference picture. The predictor 110 may derive a prediction sample for the current block by applying any one of a skip mode, a merge mode, and a motion vector prediction (MVP) mode. In the case of the skip mode and the merge mode, the predictor 110 may use motion information of an adjacent block as motion information of the current block. In the case of the skip mode, different from the merge mode, a difference (residual) between a prediction sample and an original sample is not transmitted. In the case of the MVP mode, a motion vector of an adjacent block is used as a motion vector predictor to derive a motion vector of the current block.

[0051] In the case of inter prediction, adjacent blocks may include spatially adjacent blocks existing in the current picture and temporally adjacent blocks existing in a reference picture. A reference picture including temporally adjacent blocks may also be referred to as a collocated picture (colPic). Motion information may include a motion vector and a reference picture index. Information such as prediction mode information and motion information may be (entropy) encoded and then output in the form of a bitstream.

[0052] When motion information of a temporally adjacent block is used in the skip mode and the merge mode, the highest picture in a reference picture list may be used as a reference picture. Reference pictures included in the reference picture list may be aligned based on a picture order count (POC) difference between the current picture and the corresponding reference picture. The POC corresponds to a display order and may be distinguished from an encoding order.

[0053] The subtractor 121 generates a residual sample, which is a difference between an original sample and a prediction sample. If the skip mode is applied, a residual sample may not be generated as described above.

[0054] The transformer 122 transforms the residual samples in units of transform blocks to generate transform coefficients. The transformer 122 may perform the transformation based on the size of the corresponding transform block and the prediction mode applied to a prediction block or a coded block that spatially overlaps with the transform block. For example, if intra prediction is applied to a prediction block or a coded block that overlaps with the transform block and the transform block is a 4×4 residual array, a discrete sine transform (DST) transform kernel may be used to transform the residual samples, and in other cases, a discrete cosine transform (DCT) transform kernel may be used to transform the residual samples.

[0055] The quantizer 123 may quantize the transform coefficients to generate quantized transform coefficients.

[0056] The re-arranger 124 re-arranges the quantized transform coefficients. The re-arranger 124 may re-arrange the quantized transform coefficients in block form into a one-dimensional vector by a coefficient scanning method. Although the re-arranger 124 is described as a separate component, the re-arranger 124 may be part of the quantizer 123.

[0057] The entropy encoder 130 may perform entropy encoding on the quantized transform coefficients. The entropy encoding may include encoding methods such as exponential Golomb, context-adaptive variable length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), etc. In addition to the quantized transform coefficients, the entropy encoder 130 may also perform encoding on the information required for video reconstruction (such as syntax element values, etc.) together or separately according to the entropy encoding or according to a pre-configured method. The entropy encoding information may be sent or stored in units of network abstraction layer (NAL) in the form of a bitstream. The bitstream may be sent via a network or stored in a digital storage medium. Here, the network may include a broadcast network or a communication network, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SDD, etc.

[0058] The de-quantizer 125 de-quantizes the values (transform coefficients) quantized by the quantizer 123, and the inverse transformer 126 inverse-transforms the values de-quantized by the de-quantizer 125 to generate residual samples.

[0059] The adder 140 adds the residual samples to the prediction samples to reconstruct a picture. The residual samples may be added to the prediction samples in units of blocks to generate reconstructed blocks. Although the adder 140 is described as a separate component, the adder 140 may be part of the predictor 110. Additionally, the adder 140 may be referred to as a reconstructor or a reconstructed block generator.

[0060] Filter 150 may apply deblocking filtering and / or sample adaptive offset to the reconstructed picture. Artifacts at block boundaries in the reconstructed picture or distortions during quantization may be corrected by deblocking filtering and / or sample adaptive offset. After deblocking filtering is completed, sample adaptive offset may be applied on a sample-by-sample basis. Filter 150 may apply an adaptive loop filter (ALF) to the reconstructed picture. The ALF may be applied to the reconstructed picture to which deblocking filtering and / or sample adaptive offset have been applied.

[0061] Memory 160 may store the reconstructed picture (decoded picture) or information required for encoding / decoding. Here, the reconstructed picture may be the reconstructed picture filtered by Filter 150. The stored reconstructed picture may be used as a reference picture for (inter-frame) prediction of other pictures. For example, Memory 160 may store the (reference) picture for inter-frame prediction. Here, the picture for inter-frame prediction may be specified according to a reference picture set or a reference picture list.

[0062] Figure 2 is a block diagram briefly illustrating a video / image decoding apparatus according to an embodiment of the present disclosure.

[0063] Hereinafter, the video decoding apparatus may include an image decoding apparatus.

[0064] Reference Figure 2 , the video decoding apparatus 200 may include an entropy decoder 210, a residual processor 220, a predictor 230, an adder 240, a filter 250, and a memory 260. The residual processor 220 may include a rearranger 221, an inverse quantizer 222, and an inverse transformer 223.

[0065] In addition, although not depicted, the video decoding apparatus 200 may include a receiver for receiving a bitstream including video information. The receiver may be configured as a separate module or may be included in the entropy decoder 210.

[0066] When a bitstream including video / image information is input, the video decoding apparatus 200 may reconstruct video / image / picture in association with the process of processing video information in the video encoding apparatus.

[0067] For example, the video decoding apparatus 200 may perform video decoding using the processing units applied in the video encoding apparatus. Therefore, the processing unit blocks for video decoding may be, for example, encoding units, and in another example, encoding units, prediction units, or transformers. Encoding units may be separated from the largest coding unit according to a quadtree structure and / or a binary tree structure and / or a ternary tree structure.

[0068] In some cases, a prediction unit and a transformer may be further used. In such cases, a prediction block is a block derived from or split from an encoding unit and may be a unit of sample prediction. Here, the prediction unit may be divided into sub-blocks. The transformer may be separated from the encoding unit according to a quadtree structure and may be a unit for deriving transform coefficients or a unit for deriving a residual signal from transform coefficients.

[0069] The entropy decoder 210 may parse the bitstream to output information required for video reconstruction or picture reconstruction. For example, the entropy decoder 210 may decode the information in the bitstream based on an encoding method such as exponential Golomb coding, CAVLC, CABAC, etc., and may output the values of the syntax elements required for video reconstruction and the quantization values of the transform coefficients regarding the residuals.

[0070] More specifically, the CABAC entropy decoding method may receive bins corresponding to each syntax element in the bitstream, determine a context model using the decoding target syntax element information and the decoding information of adjacent blocks and the decoding target block or the information of the symbols / bins decoded in the previous step, predict the bin generation probability according to the determined context model, and perform arithmetic decoding of the bins to generate symbols corresponding to each syntax element value. Here, the CABAC entropy decoding method may update the context model using the information of the symbols / bins decoded by the context model decoding for the next symbol / bin after determining the context model.

[0071] The information regarding prediction in the information decoded by the entropy decoder 210 may be provided to the predictor 250, and the residual values (i.e., the quantized transform coefficients) that have been entropy decoded by the entropy decoder 210 may be input to the rearranger 221.

[0072] The rearranger 221 may rearrange the quantized transform coefficients into a two-dimensional block form. The rearranger 221 may perform a rearrangement corresponding to the coefficient scan performed by the encoding device. Although the rearranger 221 is described as a separate component, the rearranger 221 may be a part of the inverse quantizer 222.

[0073] The inverse quantizer 222 may inverse quantize the quantized transform coefficients based on the (inverse) quantization parameters to output transform coefficients. In this case, the information for deriving the quantization parameters may be signaled from the encoding device.

[0074] The inverse transformer 223 may perform an inverse transform on the transform coefficients to derive residual samples.

[0075] The predictor 230 may perform a prediction on the current block and may generate a prediction block including prediction samples for the current block. The unit of prediction performed in the predictor 230 may be an encoding block, or may be a transform block or may be a prediction block.

[0076] Predictor 230 may determine whether to apply intra prediction or inter prediction based on information about the prediction. In this case, the unit for determining which one will be used between intra prediction and inter prediction may be different from the unit for generating prediction samples. Additionally, the unit for generating prediction samples may also be different in inter prediction and intra prediction. For example, it may be determined on a CU-by-CU basis which one of inter prediction and intra prediction will be applied. Furthermore, for example, in inter prediction, prediction samples may be generated by determining a prediction mode on a PU-by-PU basis, and in intra prediction, prediction samples may be generated on a TU-by-TU basis by determining a prediction mode on a PU-by-PU basis.

[0077] In the case of intra prediction, predictor 230 may derive prediction samples for the current block based on neighboring reference samples in the current picture. Predictor 230 may derive prediction samples for the current block by applying a directional mode or a non-directional mode based on the neighboring reference samples of the current block. In this case, the prediction mode to be applied to the current block may be determined by using the intra prediction mode of adjacent blocks.

[0078] In the case of inter prediction, predictor 230 may derive prediction samples for the current block based on the samples specified in the reference picture according to the motion vector. Predictor 230 may use one of the skip mode, merge mode, and MVP mode to derive prediction samples for the current block. Here, the motion information (e.g., motion vector and information about the reference picture index) required for the inter prediction of the current block provided by the video coding device may be obtained or derived based on the information about the prediction.

[0079] In the skip mode and merge mode, the motion information of adjacent blocks may be used as the motion information of the current block. Here, the adjacent blocks may include spatially adjacent blocks and temporally adjacent blocks.

[0080] Predictor 230 may use the motion information of available adjacent blocks to construct a merge candidate list, and use the information indicated by the merge index on the merge candidate list as the motion vector of the current block. The merge index may be signaled by the coding device. The motion information may include a motion vector and a reference picture. In the skip mode and merge mode, when using the motion information of temporally adjacent blocks, the first sorted picture in the reference picture list may be used as the reference picture.

[0081] In the case of the skip mode, different from the merge mode, the difference (residual) between the prediction sample and the original sample is not sent.

[0082] In the case of the MVP mode, the motion vectors of adjacent blocks can be used as motion vector predictors to derive the motion vector of the current block. Here, the adjacent blocks can include spatially adjacent blocks and temporally adjacent blocks.

[0083] When applying the merge mode, for example, the motion vectors of the reconstructed spatially adjacent blocks and / or the motion vectors corresponding to the Col blocks as temporally adjacent blocks can be used to generate a merge candidate list. The motion vector of the candidate block selected from the merge candidate list is used as the motion vector of the current block in the merge mode. The above information about prediction can include a merge index that indicates the candidate block with the best motion vector selected from the candidate blocks included in the merge candidate list. Here, the predictor 230 can derive the motion vector of the current block using the merge index.

[0084] When applying the MVP (Motion Vector Prediction) mode as another example, the motion vectors of the reconstructed spatially adjacent blocks and / or the motion vectors corresponding to the Col blocks as temporally adjacent blocks can be used to generate a motion vector predictor candidate list. That is, the motion vectors of the reconstructed spatially adjacent blocks and / or the motion vectors corresponding to the Col blocks as temporally adjacent blocks can be used as motion vector candidates. The above information about prediction can include a predicted motion vector index that indicates the best motion vector selected from the motion vector candidates included in the list. Here, the predictor 230 can select the predicted motion vector of the current block from the motion vector candidates included in the motion vector candidate list using the motion vector index. The predictor of the encoding device can obtain the motion vector difference (MVD) between the motion vector of the current block and the motion vector predictor, encode the MVD, and output the encoded MVD in the form of a bitstream. That is, the MVD can be obtained by subtracting the motion vector predictor from the motion vector of the current block. Here, the predictor 230 can obtain the motion vector included in the information about prediction, and derive the motion vector of the current block by adding the motion vector difference to the motion vector predictor. In addition, the predictor can obtain or derive a reference picture index indicating the reference picture from the above information about prediction.

[0085] The adder 240 can add the residual samples to the predicted samples to reconstruct the current block or the current picture. The adder 240 can reconstruct the current picture by adding the residual samples to the predicted samples in units of blocks. When applying the skip mode, no residual is sent, and thus the predicted samples can become the reconstructed samples. Although the adder 240 is described as a separate component, the adder 240 can be part of the predictor 230. In addition, the adder 240 can be referred to as a reconstructor or a reconstructed block generator.

[0086] Filter 250 may apply deblocking filter, sample adaptive offset, and / or ALF to the reconstructed picture. Here, the sample adaptive offset may be applied on a sample-by-sample basis after the deblocking filter. ALF may be applied after the deblocking filter and / or after applying the sample adaptive offset.

[0087] Memory 260 may store the reconstructed picture (decoded picture) or information required for decoding. Here, the reconstructed picture may be the reconstructed picture filtered by filter 250. For example, memory 260 may store pictures for inter prediction. Here, the pictures for inter prediction may be specified according to a reference picture set or a reference picture list. The reconstructed picture may be used as a reference picture for other pictures. Memory 260 may output the reconstructed pictures in the output order.

[0088] In addition, as described above, when performing video coding, prediction is performed to improve the compression efficiency. Thus, a prediction block including prediction samples for a current block that is a block to be encoded (i.e., an encoding target block) may be generated. Here, the prediction block includes prediction samples in the spatial domain (or pixel domain). The prediction block is derived in the same manner in the encoding device and the decoding device, and the encoding device may signal information about the residual between the original block and the prediction block (residual information) instead of the original sample values of the original block to the decoding device, thereby improving the image coding efficiency. The decoding device may derive a residual block including residual samples based on the residual information, add the residual block to the prediction block to generate a reconstructed block including reconstructed samples, and generate a reconstructed picture including the reconstructed block.

[0089] The residual information may be generated through a transformation and quantization process. For example, the encoding device may derive a residual block between the original block and the prediction block, perform a transformation process on the residual samples (residual sample array) included in the residual block to derive transform coefficients, perform a quantization process on the transform coefficients to derive quantized transform coefficients, and may signal (through a bitstream) the relevant residual information to the decoding device. Here, the residual information may include value information, position information, transformation technique, transformation kernel, and quantization parameters of the quantized transform coefficients, etc. The decoding device may perform an inverse quantization / inverse transformation process based on the residual information and may derive residual samples (or a residual block). The decoding device may generate a reconstructed picture based on the prediction block and the residual block. Additionally, for future inter prediction of reference pictures, the encoding device may also perform inverse quantization / inverse transformation on the quantized transform coefficients to derive a residual block and generate a reconstructed picture based on the residual block.

[0090] Figure 3 Illustratively shows a content streaming system according to an embodiment of the present disclosure.

[0091] Reference Figure 3, the embodiments described in this disclosure can be embodied and executed on a processor, microprocessor, controller, or chip. For example, the functional units shown in each drawing can be embodied and executed on a computer, processor, microprocessor, controller, or chip. In this case, the information for the embodiments (e.g., information about instructions) or algorithms can be stored in a digital storage medium.

[0092] In addition, the decoding device and encoding device applied in this disclosure can be included in a multimedia broadcast transceiver, a mobile communication terminal, a home theater video device, a digital cinema video device, a surveillance camera, a video chat device, a real-time communication device such as video communication, a mobile streaming device, a storage medium, a portable camera, a video-on-demand (VoD) service providing device, an over-the-top (OTT) video device, an Internet streaming service providing device, a three-dimensional (3D) video device, a video phone video device, and a medical video device, and can be used to process video signals or data signals. For example, an over-the-top (OTT) video device can include a game console, a Blu-ray player, an Internet access TV, a home theater system, a smart phone, a tablet, a digital video recorder (DVR), and so on.

[0093] In addition, the processing method applied in this disclosure can be generated in the form of a program executable by a computer and stored in a computer-readable recording medium. Multimedia data having a data structure according to this disclosure can also be stored in a computer-readable recording medium. The computer-readable recording medium includes various storage devices and distributed storage devices in which computer-readable data is stored. The computer-readable recording medium can include, for example, a Blu-ray Disc (BD), a Universal Serial Bus (USB), a ROM, a PROM, an EPROM, an EEPROM, a RAM, a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device. In addition, the computer-readable recording medium includes a medium embodied in a carrier wave (e.g., transmission via the Internet). Additionally, the bitstream generated by the encoding method can be stored in a computer-readable recording medium or transmitted via a wired or wireless communication network.

[0094] In addition, the embodiments of this disclosure can be embodied as a computer program product by program code, and the program code can be executed on a computer by the embodiments of this disclosure. The program code can be stored on a computer-readable carrier.

[0095] The content streaming system applied in this disclosure can mainly include an encoding server, a streaming server, a web server, a media storage device, a user device, and a multimedia input device.

[0096] The encoding server is used to compress the content input from a multimedia input device (such as a smart phone, camera, portable video camera, etc.) into digital data, generate a bitstream, and send it to the streaming server. As another example, in the case where a multimedia input device such as a smart phone, camera, portable video camera, etc. directly generates a bitstream, the encoding server can be omitted.

[0097] The bitstream can be generated by the encoding method or bitstream generation method applied in the present disclosure. And the streaming server can temporarily store the bitstream during the process of sending or receiving the bitstream.

[0098] The streaming server sends the multimedia data to the user device via a web server based on the user's request. The web server serves as a tool to notify the user of what services are available. When the user requests the service they want, the web server transmits the request to the streaming server, and the streaming server sends the multimedia data to the user. In this regard, the content streaming system can include a separate control server, and in this case, the control server serves as controlling the commands / responses between the respective devices in the content streaming system.

[0099] The streaming server can receive content from a media storage device and / or an encoding server. For example, in the case of receiving content from the encoding server, the content can be received in real time. In this case, the streaming server can store the bitstream for a predetermined period of time to smoothly provide the streaming service.

[0100] For example, the user device can include a mobile phone, smart phone, laptop computer, digital broadcast terminal, personal digital assistant (PDA), portable multimedia player (PMP), navigation device, tablet PC, tablet computer, ultrabook, wearable device (such as a watch-type terminal (smart watch), glasses-type terminal (smart glasses), head-mounted display (HMD)), digital TV, desktop computer, digital signage, etc.

[0101] Each server in the content streaming system can operate as a distributed server, and in this case, the data received by each server can be processed in a distributed manner.

[0102] Hereinafter, it will be described in detail with reference to Figure 1 and Figure 2 the inter-frame prediction method described.

[0103] Figure 4 is a flowchart for illustrating a method of deriving a motion vector prediction value from neighboring blocks according to an embodiment of the present disclosure.

[0104] In the case of the motion vector prediction (MVP) mode, the encoder predicts the motion vector according to the type of the prediction block and sends the difference between the best motion vector and the predicted value to the decoder. In this case, the encoder sends the motion vector difference, neighboring block information, reference index, etc. to the decoder. Here, the MVP mode can also be referred to as the advanced motion vector prediction (AMVP) mode.

[0105] The encoder can construct a prediction candidate list for motion vector prediction, and the prediction candidate list can include at least one of a spatial candidate block and a temporal candidate block.

[0106] First, the encoder can search for spatial candidate blocks for motion vector prediction and insert them into the prediction candidate list (S410). For the process of constructing the spatial candidate blocks, a method of constructing conventional spatial merge candidates in inter prediction according to the merge mode can be applied.

[0107] The encoder can check whether the number of spatial candidate blocks is less than two (S420).

[0108] In the case where the number of spatial candidate blocks is less than two as a result of the check, the encoder can search for temporal candidate blocks and insert them into the prediction candidate list (S430). At this time, in the case where no temporal candidate blocks are available, the encoder can use a zero motion vector as the motion vector prediction value (S440). For the process of constructing the temporal candidate blocks, a method of constructing conventional temporal merge candidates in inter prediction according to the merge mode can be applied.

[0109] On the other hand, in the case where the number of spatial candidate blocks is equal to or greater than two as a result of the check, the encoder can end the construction of the prediction candidate list and select the block with the minimum cost from among the candidate blocks. The encoder can determine the motion vector of the selected candidate block as the motion vector prediction value of the current block and obtain the motion vector difference by using the motion vector prediction value. The motion vector difference obtained in this way can be sent to the decoder.

[0110] Figure 5 Illustratively represents an affine motion model according to an embodiment of the present disclosure.

[0111] The affine mode can be one of various prediction modes in inter prediction, and the affine mode can also be referred to as the affine motion mode or the sub-block motion prediction mode. The affine mode can refer to a mode of performing an affine motion prediction method using an affine motion model.

[0112] The affine motion prediction method can derive the motion vector of a sample unit by using two or more motion vectors in the current block. In other words, the affine motion prediction method can improve the coding efficiency by determining the motion vector not in units of blocks but in units of samples.

[0113] The general motion model may include a translation model, and motion estimation (ME) and motion compensation (MC) are performed based on the translation model that effectively represents simple motion. However, the translation model may not be effectively applied to complex motions in natural videos, such as zooming in, zooming out, rotating, and other irregular motions. Therefore, embodiments of the present disclosure may use an affine motion model that can be effectively applied to complex motions.

[0114] Reference Figure 5 , the affine motion model may include four motion models, but they are exemplary motion models, and the scope of the present disclosure is not limited thereto. The above four motions may include translation, scaling, rotation, and shear. Here, the motion models for translation, scaling, and rotation may be referred to as a simplified affine motion model.

[0115] Figure 6 Illustratively represents a simplified affine motion model according to an embodiment of the present disclosure.

[0116] In affine motion prediction, control points (CPs) may be defined to use the affine motion model, and the motion vectors of sub-blocks or sample units included in a block may be determined by using two or more control point motion vectors (CPMVs). Here, a set of motion vectors of sample units or a set of motion vectors of sub-blocks may be referred to as an affine motion vector field (affine MVF).

[0117] Reference Figure 6 , the simplified affine motion model may mean a model for determining the motion vectors of sample units or sub-blocks by using CPMV according to two CPs, and may also be referred to as a 4-parameter affine model. In Figure 6 , v0 and v1 may represent two CPMVs, and each arrow in the sub-block may represent the motion vector of the sub-block unit.

[0118] In other words, in the encoding / decoding process, the affine motion vector field may be determined in sample units or sub-block units. Here, the sample unit may refer to a pixel unit, and the sub-block unit may refer to a defined block unit. When the affine motion vector field is determined in sample units, the motion vectors may be obtained based on each pixel value, and in the case of block units, the motion vectors of the corresponding blocks may be obtained based on the center pixel value of the blocks.

[0119] Figure 7 Is a diagram for describing a method of deriving a motion vector predictor at control points according to an embodiment of the present disclosure.

[0120] The affine mode may include an affine merge mode and an affine motion vector prediction (MVP) mode. The affine merge mode may be referred to as a sub-block merge mode, and the affine MVP mode may be referred to as an affine inter-frame mode.

[0121] In the affine MVP mode, the CPMV of the current block may be derived based on a control point motion vector predictor (CPMVP) and a control point motion vector difference. In other words, the encoding device may determine the CPMVP of the CPMV of the current block, derive the CPMVD that is the difference between the CPMV of the current block and the CPMVP, and signal information about the CPMVP and information about the CPMVD to the decoding device. Here, the affine MVP mode may construct an affine MVP candidate list based on neighboring blocks, and the affine MVP candidate list may be referred to as a CPMVP candidate list. In addition, the information about the CPMVP may include an index indicating a block or a motion vector to be referenced from among the affine MVP candidate list.

[0122] Reference Figure 7 , the motion vector of the control point at the upper left sample position of the current block may be represented as v0, the motion vector of the control point at the upper right sample position may be represented as v1, the motion vector of the control point at the lower left sample position may be represented as v2, and the motion vector of the control point at the lower right sample position may be represented as v3.

[0123] For example, if two control points are used in the affine mode and the two control points are located at the upper left sample position and the upper right sample position, the motion vector of the sample unit or the sub-block unit may be derived based on the motion vectors v0 and v1.

[0124] The motion vector v0 may be derived based on at least one motion vector of neighboring blocks A, B, and C at the upper left sample position. Here, the neighboring block A may represent a block located above the upper left of the upper left sample position of the current block, the neighboring block B may represent a block located at the top of the upper left sample position of the current block, and the neighboring block C may represent a block located to the left of the upper left sample position of the current block.

[0125] The motion vector v1 may be derived based on at least one motion vector of neighboring blocks D and E at the upper right sample position. Here, the neighboring block D may represent a block located at the top of the upper right sample position of the current block, and the neighboring block E may represent a block located above the upper right of the upper right sample position of the current block.

[0126] For example, if three control points are used in the affine mode and the three control points are located at the upper left sample position, the upper right sample position, and the lower left sample position, the motion vector of the sample unit or the sub-block unit may be derived based on the motion vectors v0, v1, and v2. In other words, the motion vector v2 may be further used.

[0127] The motion vector v2 can be derived based on at least one motion vector of neighboring blocks F and G at the lower-left sample position. Here, the neighboring block F can represent a block located to the left of the lower-left sample position of the current block, and the neighboring block G can represent a block located diagonally below and to the left of the lower-left sample position of the current block.

[0128] The affine MVP mode can derive a CPMVP candidate list based on neighboring blocks and select the CPMVP pair with the highest correlation among the CPMVP candidate list as the CPMV of the current block. Information about the above CPMV can include an index indicating the CPMVP pair selected from the CPMVP candidate list.

[0129] Figure 8 Illustratively represent two CPs for a 4-parameter affine motion model according to an embodiment of the present disclosure.

[0130] Embodiments of the present disclosure can use two CPs. The two CPs can be located at the upper-left sample position and the upper-right sample position of the current block, respectively. Here, the CP located at the upper-left sample position can be represented as CP0, and the CP located at the upper-right sample position can be represented as CP1, and the motion vector at CP0 can be represented as mv0 and the motion vector at CP1 can be represented as mv1. The coordinates of each control point (CP i ) can be defined as (x i , y i ), i=0, 1 , and the motion vector at each control point can be represented as mvi=(v xi , v yi ), i=0, 1 .

[0131] In an embodiment of the present disclosure, the CP located at the upper-left sample position can be represented as CP1, and the CP located at the upper-right sample position can be represented as CP0. In this case, the following process can be similarly performed considering the switching positions of CP0 and CP1.

[0132] For example, if the width of the current block is W and its height is H, assuming the coordinates of the lower-left sample position of the current block are (0, 0), the coordinates of CP0 can be represented as (0, H), and the coordinates of CP1 can be represented as (W, H). Here, W and H can have different values, but can also have the same value, and the reference (0, 0) can be set differently.

[0133] As Figure 8As shown, since an affine motion model using two motion vectors based on two CPs uses four parameters according to the two motion vectors in an affine motion prediction method, it may be referred to as a 4-parameter affine motion model or a simplified affine motion model.

[0134] In an embodiment of the present disclosure, the motion vector of a sample unit may be determined by an affine motion vector field (affine MVF) and the position of the sample. The affine motion vector field may represent the motion vector of the sample unit based on two motion vectors according to two CPs. In other words, when the sample position is (x, y) as shown in Equation 1, the affine motion vector field may derive the motion vector (v x , vy) of the corresponding sample.

[0135] [Equation 1]

[0136]

[0137] In Equation 1, v0x and v0y may mean the (x, y) coordinate components of the motion vector mv0 at CP0, while v1x and v1y may mean the (x, y) coordinate components of the motion vector mv1 at CP1. Similarly, w may mean the width of the current block.

[0138] Meanwhile, Equation 1 representing the affine motion model is only an example, and the equation for representing the affine motion model is not limited to Equation 1. For example, in some cases, the signs of each coefficient disclosed in Equation 1 may be changed according to the signs of Equation 1.

[0139] In other words, according to an embodiment of the present disclosure, a reference block among the temporal and / or spatial neighboring blocks of the current block may be determined, and the motion vector of the reference block may be used as a motion vector predictor for the current block, and the motion vector of the current block may be expressed by the motion vector predictor and the motion vector difference. In addition, an embodiment of the present disclosure may signal the index of the motion vector predictor and the motion vector difference.

[0140] According to an embodiment of the present disclosure, during encoding, two motion vector differences according to two CPs may be derived based on two motion vectors according to two CPs and two motion vector predictors according to two CPs, and during decoding, two motion vectors according to two CPs may be derived based on two motion vector predictors according to two CPs and two motion vector differences according to two CPs. In other words, the motion vector at each CP may be composed of the sum of the motion vector predictor and the motion vector difference, as shown in Equation 2, and is similar to when using the motion vector prediction (MVP) mode or the advanced motion vector prediction (AMVP) mode.

[0141] [Equation 2]

[0142]

[0143] In Equation 2, mvp0 and mvp1 may represent motion vector predictors (MVPs) at each of CP0 and CP1, and mvd0 and mvd1 may represent motion vector differences (MVDs) at each of CP0 and CP1. Herein, mvp may be referred to as CPMVP, and mvd may be referred to as CPMVD.

[0144] Thus, the inter - frame prediction method according to the affine mode according to an embodiment of the present disclosure may encode and decode an index and motion vector differences (mvd0 and mvd1) at each CP. In other words, according to an embodiment of the present disclosure, motion vectors at CP0 and CP1 may be derived based on mvd0 and mvd1 at each of CP0 and CP1 of the current block and mvp0 and mvp1 according to the index, and inter - frame prediction may be performed by deriving the motion vector of the sample unit based on the motion vectors at CP0 and CP1.

[0145] The inter - frame prediction method according to another embodiment of the present disclosure may use one of the differences of motion vector differences (DMVD) according to two CPs. In other words, in another embodiment, when there are mvd0 and mvd1 according to CP0 and CP1, inter - frame prediction may be performed by encoding and decoding mvd0 and mvd1, the difference between mvd0 and mvd1, and one of the indices at each CP. More specifically, in another embodiment of the present disclosure, mvd0 and DMVD (mvd0 - mvd1) may be signaled, and mvd1 and DMVD (mvd0 - mvd1) may be signaled.

[0146] That is, in another embodiment of the present disclosure, the motion vector differences (mvd0 and mvd1) at CP0 and CP1 may be derived respectively based on the motion vector difference (mvd0 or mvd1) at CP0 or CP1 and the difference of the motion vector differences (mvd0 and mvd1) of CP0 and CP1, and the motion vectors at CP0 and CP1 may be derived respectively based on the motion vectors (mvp0 and mvp1) of CP0 and CP1 indicated by the index together with such motion vector differences, and inter - frame prediction may be performed by deriving the motion vector of the sample unit based on the motion vectors at CP0 and CP1.

[0147] Here, data on the difference of the two motion vector differences (DMVD) is closer to the zero motion vector (zero MV) than data on the normal motion vector difference, and encoding can be performed more efficiently compared to the case of other embodiments according to the present disclosure.

[0148] Figure 9Illustratively shows a case where a median value is additionally used in a four-parameter affine motion model according to an embodiment of the present disclosure.

[0149] In inter-frame prediction according to an affine mode according to an embodiment of the present disclosure, a median predictor of a motion vector predictor can be used to perform adaptive motion vector coding.

[0150] Reference Figure 9 , according to an embodiment of the present disclosure, two CPs can be used, and information about CPs at different positions can be derived based on the two CPs. Here, the information related to the two CPs can be the same as the two CPs of Figure 8 . Additionally, a CP at another position derived based on the two CPs can be referred to as CP2 and can be located at the lower left sample position of the current block.

[0151] The information about the CP at another position can include the motion vector predictor (mvp2) of CP2, and mvp2 can be derived in two ways.

[0152] One method of deriving mvp2 is as follows. mvp2 can be derived based on the motion vector predictor mvp0 of CP0 and the motion vector predictor mvp1 of CP1, and can be derived as shown in Equation 3.

[0153] [Equation 3]

[0154]

[0155] In Equation 3, mvp 0x and mvp 0y can mean the (x, y) coordinate components of the motion vector predictor (mvp0) at CP0, mvp 1x and mvp 1y can mean the (x, y) coordinate components of the motion vector predictor (mvp1) at CP1, while mvp 2x and mvp 2y can mean the (x, y) coordinate components of the motion vector predictor (mvp2) at CP2. Additionally, h can represent the height of the current block, and w can represent the width of the current block.

[0156] Another method of deriving mvp2 is as follows. mvp2 can be derived based on the neighboring blocks of CP2. Referring to Figure 9 , CP2 can be located at the lower left sample position of the current block, and mvp2 can be derived based on neighboring block A or neighboring block B of CP2. More specifically, mvp2 can be selected as one of the motion vectors of neighboring block A and neighboring block B.

[0157] In an embodiment of the present disclosure, mvp2 can be derived, and a median value can be derived based on mvp0, mvp1, and mvp2. Here, the median value can mean the value that is located in the middle in ascending or descending order among a plurality of values. Therefore, the median value can be selected from mvp0, mvp1, and mvp2.

[0158] According to an embodiment of the present disclosure, when the median value is equal to mvp0, the motion vector difference (mvd0) of CP0 and DMVD (mvd0 - mvd1) can be signaled for inter - frame prediction, and when the median value is equal to mvp1, the motion vector difference (mvd1) of CP1 and DMVD (mvd0 - mvd1) can be signaled for inter - frame prediction. When the median value is equal to mvp2, either the case where the median value is equal to mvp0 or the case where the median value is equal to mvp1 can be followed, and this can be predefined.

[0159] Here, the above - mentioned process can be performed for each of the x and y components of the motion vector predictor. In other words, the median value can be derived for the x component and the y component separately. In this case, when the x component of the median value is the same as the x component of mvp0, only the x component of the motion vector can be encoded and decoded according to the above - mentioned case where the median value is equal to mvp0, and when the y component of the median value is the same as the y component of mvp1, only the y component of the motion vector can be encoded and decoded according to the above - mentioned case where the median value is the same as mvp1. If the x component and / or y component of the median value is equal to the x component and / or y component of mvp2, the x component and / or y component of the motion vector can be encoded and decoded according to one of the above - mentioned cases where the median value is equal to mvp0 and the median value is equal to mvp1, which can be predefined.

[0160] Figure 10 Illustratively shows three CPs for a 6 - parameter affine motion model according to an embodiment of the present disclosure.

[0161] Embodiments of the present disclosure can use three CPs. The three CPs can be respectively located at the upper - left sample position, the upper - right sample position, and the lower - left sample position of the current block. Here, the CP located at the upper - left sample position can be denoted as CP0, the CP located at the upper - right sample position can be denoted as CP1, and the CP that can be located at the lower - left sample position can be denoted as CP2, and the motion vector at CP0 can be denoted as mv0, the motion vector at CP1 can be denoted as mv1, and the motion vector at CP2 can be denoted as mv2. The coordinates of each control point (CPi) can be defined as (x i , y i ), i=0, 1, 2 , and the motion vector at each control point can be denoted as mv i =( v x i,v yi ), i=0, 1, 2 。

[0162] In an embodiment of the present disclosure, three CPs may be respectively distributed at the upper left sample position, the upper right sample position, and the lower left sample position, but these three CPs may be positioned differently from this. For example, CP0 may be located at the upper right sample position, CP1 may be located at the upper left sample position, and CP2 may be located at the lower left sample position, but their positions are not limited thereto. In this case, the following processing may be similarly performed in consideration of the positions of each CP.

[0163] For example, if the width of the current block is W and its height is H, assuming the coordinates of the lower left sample position of the current block are (0, 0), the coordinates of CP0 may be represented as (0, H), the coordinates of CP1 may be represented as (W, H), and the coordinates of CP2 may be represented as (0, 0). Here, W and H may have different values, but may also have the same value, and the reference (0, 0) may be set differently.

[0164] As shown in Figure 10 , since an affine motion model using three motion vectors according to three CPs uses six parameters according to the three motion vectors in the affine motion prediction method, it may be referred to as a 6-parameter affine motion model.

[0165] In an embodiment of the present disclosure, the motion vector of a sample unit may be determined by an affine motion vector field (affine MVF) and the position of the sample. The affine motion vector field may represent the motion vector of the sample unit based on three motion vectors according to three CPs.

[0166] According to an embodiment of the present disclosure, a reference block among the temporal and / or spatial neighboring blocks of the current block may be determined, and the motion vector of the reference block may be used as a motion vector predictor of the current block, and the motion vector of the current block may be represented by the motion vector predictor and the motion vector difference. In addition, an embodiment of the present disclosure may signal the indexes of the motion vector predictor and the motion vector difference.

[0167] According to an embodiment of the present disclosure, three motion vector differences according to three CPs may be derived based on three motion vectors according to three CPs and three motion vector predictors according to three CPs during encoding, and during decoding, three motion vectors according to three CPs may be derived based on three motion vector predictors according to three CPs and three motion vector differences according to three CPs. In other words, the motion vector at each CP may be composed of the sum of the motion vector predictor and the motion vector difference, as shown in Equation 4.

[0168] [Equation 4]

[0169]

[0170] In Equation 4, mvp0, mvp1, and mvp2 may represent motion vector predictors (MVPs) at each of CP0, CP1, and CP2, and mvd0, mvd1, and mvd2 may represent motion vector differences (MVDs) at each of CP0, CP1, and CP2. Herein, mvp may be referred to as CPMVP, and mvd may be referred to as CPMVD.

[0171] Accordingly, an inter prediction method according to an affine mode according to an embodiment of the present disclosure may encode and decode an index and motion vector differences (mvd0, mvd1, and mvd2) at each CP. In other words, according to an embodiment of the present disclosure, based on mvd0, mvd1, and mvd2 at each of CP0, CP1, and CP2 of a current block and mvp0, mvp1, and mvp2 according to an index, motion vectors at CP0, CP1, and CP2 may be derived, and inter prediction may be performed based on the motion vectors at CP0, CP1, and CP2 to derive a motion vector of a sample unit.

[0172] An affine motion prediction method according to another embodiment of the present disclosure may use one of three motion vector differences, a difference (DMVD, a difference between two MVDs) between the motion vector difference and another motion vector difference (DMVD), and a difference (DMVD) between the motion vector difference and yet another motion vector difference.

[0173] More specifically, in another embodiment of the present disclosure, when mvd0, mvd1, and mvd2 are derived according to three CPs, inter prediction may be performed by signaling mvd0 and two DMVDs (mvd0 - mvd1 and mvd0 - mvd2), by signaling mvd1 and two DMVDs (mvd0 - mvd1 and mvd1 - mvd2), or by signaling mvd2 and two DMVDs (mvd0 - mvd2 and mvd1 - mvd2). Herein, for convenience, one of the two DMVDs may be referred to as a first DMVD (DMVD1), and the other DMVD may be referred to as a second DMVD (DMVD2). Additionally, when encoding and decoding mvd0 and two DMVDs (mvd0 - mvd1 and mvd0 - mvd2), mvd0 may be referred to as the MVD of CP0, DMVD (mvd0 - mvd1) may be referred to as the DMVD of CP1, and DMVD (mvd0 - mvd2) may be referred to as the DMVD of CP2.

[0174] That is, according to another embodiment of the present disclosure, three MVDs (e.g., MVD0, MVD1, and MVD2) can be derived based on one MVD and two DMVDs, and the motion vectors at each of the three CPs can be derived based on a motion vector predictor according to the index for the three CPs (e.g., CP0, CP1, and CP2) and together with such an MVD, and the inter-frame prediction can be performed by deriving the motion vector of the sample unit based on the motion vectors at the three CPs.

[0175] In reference Figure 10 In another embodiment of the present disclosure described, the method using the median predictor described in reference Figure 9 can be adaptively applied to effectively perform motion vector coding, and in this case, the process of deriving MVP2 can be omitted from the method described in reference Figure 9 described.

[0176] Hereinafter, in the description of the present disclosure, CP0, CP1, and CP2 may be respectively represented as the first CP, the second CP, and the third CP, and the motion vector (MV), motion vector predictor (MVP), and motion vector difference (MVD) according to each CP may also be represented in a similar manner as described above.

[0177] Figure 11 Schematically shows a video coding method of an encoding device according to the present disclosure.

[0178] Figure 11 The method disclosed in Figure 1 can be executed by the encoding device disclosed in Figure 11 . For example, S1100 to S1120 in

[0179] can be executed by the predictor of the encoding device; and S1130 can be executed by the entropy encoder of the encoding device.

[0180] The encoding device derives the control points (CPs) of the current block (S1100). When affine motion prediction is applied to the current block, the encoding device can derive the CPs, and depending on the embodiment, the number of CPs can be two or three.

[0181] For example, when there are three CPs, the CPs may be located at the upper-left sample position, the upper-right sample position, and the lower-left sample position of the current block, respectively. And if the height and width of the current block are H and W, respectively, and the coordinate components of the lower-left sample position are (0, 0), the coordinate components of the CPs may be (0, H), (W, H), and (0, 0), respectively.

[0182] The encoding device derives the MVPs for the CPs (S1110). For example, when the number of derived CPs is two, the encoding device may obtain two motion vectors. For example, when the number of derived CPs is three, the encoding device may obtain three motion vectors. The MVPs for the CPs may be derived based on neighboring blocks, and the detailed description thereof has been referred to above Figure 7 and Figure 9 for a detailed description.

[0183] For example, when the first CP and the second CP are derived, the encoding device may derive the first MVP for the first CP and the second MVP for the second CP based on the neighboring blocks of the current block. And when the third CP is further derived, the encoding device may further derive the third MVP based on the neighboring blocks of the current block.

[0184] For example, when the first CP, the second CP, and the third CP are derived, the encoding device may derive the third MVP for the third CP based on the first MVP for the first CP and the second MVP for the second CP, and may also derive the third MVP based on the motion vectors of the neighboring blocks of the third CP.

[0185] The encoding device derives at least one of a motion vector difference (MVD) and a difference of two MVDs (DMVD) (S1120). The motion vector difference (MVD) may be derived based on the motion vector (MV) and the motion vector predictor (MVP). For this purpose, the encoding device may also derive the motion vectors of each CP. The difference of motion vector differences (DMVD) may be derived based on multiple motion vector differences.

[0186] For example, when there are two CPs, the encoding device may derive one MVD and one DMVD. The encoding device may derive two MVDs from the motion vectors of the two CPs and the two MVPs. Additionally, one of the two MVDs to be encoded may be selected, and the difference of the two MVDs (DMVD) may be derived based on the selected one.

[0187] For example, when the first CP and the second CP are derived, the encoding device may derive the first MVD for the first CP and the DMVD for the second CP. Here, the DMVD for the second CP may represent the difference between the first MVD and the second MVD for the second CP, and the first MVD may be the reference.

[0188] For example, when there are three CPs, the encoding device may derive one MVD and two DMVDs. The encoding device may derive three MVDs from the motion vectors of the three CPs and the reference block and may select any one of the three MVDs to be encoded. Additionally, the encoding device may derive the difference between the selected one MVD and another MVD (DMVD1) and the difference between the selected one MVD and yet another MVD (DMVD2).

[0189] For example, when further deriving a third CP, the encoding device may derive a first MVD for the first CP, a DMVD for the second CP, and a DMVD for the third CP. Here, the DMVD for the second CP may represent the difference between the first MVD and the second MVD of the second CP, and the DMVD for the third CP may represent the difference between the first MVD and the third MVD of the third CP, and the first MVD may be used as a reference.

[0190] For example, when further deriving a third CP, the encoding device may derive a median based on the first MVP, the second MVP, and the third MVP, and in this case, the third MVD and the DMVD for the third CP may not be derived. The detailed description thereof has been referred to above Figure 9 for its detailed description.

[0191] The encoding device encodes based on one MVD and at least one DMVD and outputs a bitstream (S1130). For inter prediction, the encoding device may generate and output a bitstream for the current block, which includes one MVD, at least one DMVD, and an index for the motion vector predictor.

[0192] For example, when there are two CPs, the encoding device may generate a bitstream for the current block, which includes the indexes of the motion vector predictors of the two CPs and the motion vector difference of any one of the two CPs, and the difference between the motion vector differences of the two CPs.

[0193] For example, when deriving a first CP and a second CP, the encoding device may output a bitstream by encoding the image information including information about the first CP and information about the DMVD for the second CP.

[0194] For example, if there are three CPs, the encoding device may generate a bitstream for the current block, which includes the indexes of the motion vector predictors for the three CPs and the motion vector difference for one of the three CPs, the difference between the motion vector difference for this CP and the motion vector difference for another CP, and the difference between the motion vector difference for this CP and the motion vector difference for yet another CP.

[0195] For example, when further deriving a third CP, the encoding device may further include information on DMVD for the third CP in the image information, and may encode the image information to output a bitstream.

[0196] For example, when further deriving a third CP and deriving a median value, the encoding device may output a bitstream by encoding image information including information on a first MVD and information on DMVD for a second CP, and may not further include information on DMVD for the third CP in the image information.

[0197] The bitstream generated and output by the encoding device may be sent to the decoding device via a network or a storage medium.

[0198] Figure 12 Schematically illustrate the inter-frame prediction method of the decoding device according to the present disclosure.

[0199] Figure 12 The method disclosed in may be performed by Figure 2 the decoding device disclosed in. For example, Figure 12 S1200, S1210, S1230, and S1240 in may be performed by the predictor of the decoding device, and S1220 may be performed by the entropy decoder of the decoding device. Here, S1220 may be performed before S1200 and S1210.

[0200] The decoding device derives the control point (CP) of the current block (S1200). When affine motion prediction is applied to the current block, the decoding device may derive the CP, and depending on the embodiment, the number of CPs may be two or three.

[0201] For example, when there are two CPs, the CPs may be located at the upper left sample position and the upper right sample position of the current block respectively, and if the height and width of the current block are H and W respectively, and the coordinate components of the lower left sample position are (0, 0), the coordinate components of the CPs may be (0, H) and (W, H) respectively.

[0202] For example, when there are three CPs, the CPs may be located at the upper left sample position, the upper right sample position, and the lower left sample position of the current block respectively, and if the height and width of the current block are H and W respectively, and the coordinate components of the lower left sample position are (0, 0), the coordinate components of the CPs may be (0, H), (W, H), and (0, 0).

[0203] The decoding device derives MVPs for the CPs (S1210). For example, when the number of derived CPs is two, the decoding device may obtain two motion vectors. For example, when the number of derived CPs is three, the decoding device may obtain three motion vectors. The MVPs for the CPs may be derived based on neighboring blocks, and the detailed description thereof has been referred to above Figure 7 and Figure 9 for a detailed description.

[0204] For example, when the first CP and the second CP are derived, the decoding device may derive a first MVP for the first CP and a second MVP for the second CP based on the neighboring blocks of the current block, and when a third CP is further derived, the decoding device may further derive a third MVP based on the neighboring blocks of the current block.

[0205] For example, when the first CP, the second CP, and the third CP are derived, the decoding device may derive a third MVP for the third CP based on the first MVP for the first CP and the second MVP for the second CP, and derive the third MVP based on the motion vectors of the neighboring blocks of the third CP.

[0206] The decoding device decodes one MVD and at least one DMVD (S1220). The decoding device may obtain one MVD and at least one DMVD by decoding one MVD and at least one DMVD based on the received bitstream. Here, the bitstream may include an index of the motion vector predictor for the CP. The bitstream may be received from the encoding device via a network or a storage medium.

[0207] For example, when there are two CPs, the decoding device may decode one MVD and one DMVD, and when there are three CPs, the decoding device may decode one MVD and two DMVDs. Here, the DMVD may mean the difference between two MVPs.

[0208] For example, when the first CP and the second CP are derived, the decoding device may decode a first MVD for the first CP and decode the DMVD for the second CP. Here, the DMVD for the second CP may represent the difference between the first MVD and the second MVD for the second CP.

[0209] For example, when a third CP is further derived, the decoding device may decode a first MVD for the first CP and decode the DMVD for the second CP and the DMVD for the third CP. Here, the DMVD for the second CP may represent the difference between the first MVD and the second MVD for the second CP, and the DMVD for the third CP may represent the difference between the first MVD and the third MVD for the third CP.

[0210] For example, if a third CP is further derived and the median value is used, the decoding device can decode the first MVD for the first CP and decode the DMVD for the second CP. Here, the DMVD for the third CP may not be decoded.

[0211] The decoding device derives a motion vector for a CP based on the MVP for the CP, one MVD, and at least one DMVD (S1230). The motion vector for a CP can be derived based on the motion vector difference (MVD) and the motion vector predictor (MVP), and the motion vector difference can be derived based on the difference between the motion vector differences (DMVD).

[0212] For example, when there are two CPs, the decoding device can receive one MVD and one DMVD, and based on them, two MVDs can be derived according to the two CPs. The decoding device can receive the indexes for the two CPs together, and based on them, two MVPs can be derived. The decoding device can derive the motion vectors for the two CPs based on the two MVDs and the two MVPs respectively.

[0213] For example, when the first CP and the second CP are derived, the decoding device can derive the first MV based on the first MVD and the first MVP, derive the second MVD for the second CP based on the first MVD and the DMVD for the second CP, and derive the second MV based on the second MVD and the second MVP.

[0214] For example, when there are three CPs, the decoding device can receive one MVD and two DMVDs, and based on them, three MVDs can be derived according to the three CPs. The decoding device can receive the indexes for the three CPs together, and based on them, three MVPs can be derived. The decoding device can derive the motion vectors for the three CPs based on the three MVDs and the three MVPs respectively.

[0215] For example, if a third CP is further derived and the third MVP is derived, the decoding device can derive the third MVD for the third CP based on the first MVD and the DMVD for the third CP, and derive the third MV based on the third MVD and the third MVP.

[0216] For example, if a third CP is further derived and the median value is the same as the first MVP, the decoding device can derive the first MV based on the first MVD and the first MVP, derive the second MVD for the second CP based on the first MVD and the DMVD for the second CP, and derive the second MV based on the second MVD and the second MVP.

[0217] For example, if a third CP is further derived and the median is the same as the second MVP, the decoding device may derive a second MV based on the second MVD and the second MVP, derive a first MVD for the first CP based on the second MVD and the DMVD for the first CP, and derive a first MV based on the first MVD and the first MVP.

[0218] For example, if a third CP is further derived and the median is the same as the third MVP, the decoding device may derive the first MV and the second MV according to either the case where the median is the same as the first MVP or the case where the median is the same as the second MVP. Determining either the case where the median is the same as the first MVP or the case where the median is the same as the second MVP may be predefined. The detailed description thereof has been referred to above Figure 9 in its detailed description.

[0219] The decoding device generates a prediction block for the current block based on the motion vector (S1240). The decoding device may derive an affine motion vector field (affine MVF) based on the motion vectors of the corresponding CPs, and based on them, may derive the motion vectors of sample units to perform inter prediction.

[0220] In the above embodiments, the method is explained based on a flowchart by means of a series of steps or blocks, but the present disclosure is not limited to the order of the steps, and may be in a different order or steps from the above steps, or a certain step may be executed simultaneously with another step. In addition, those of ordinary skill in the art can understand that the steps shown in the flowchart are not exclusive, and another step may be incorporated or one or more steps in the flowchart may be deleted without affecting the scope of the present disclosure.

[0221] The above method according to the present disclosure may be implemented in software form, and the encoding device and / or decoding device according to the present disclosure may be included in a device for image processing such as a television, a computer, a smart phone, a set-top box, a display device, etc.

[0222] When the embodiments in the present disclosure are embodied by software, the above method may be embodied as a module (procedure, function, etc.) for performing the above functions. These modules may be stored in a memory and may be executed by a processor. The memory may be inside or outside the processor and may be connected to the processor in various well-known ways. The processor may include an application specific integrated circuit (ASIC), other chip sets, logic circuits, and / or data processing devices. The memory may include a read only memory (ROM), a random access memory (RAM), a flash memory, a memory card, a storage medium, and / or other storage devices.

Claims

1. An image decoding method performed by a decoding device, the method comprising: Obtaining residual information from a bitstream; Deriving a first motion vector predictor for a first control point of the current block, a second motion vector predictor for a second control point of the current block, and a third motion vector predictor for a third control point of the current block based on neighboring blocks of the current block, wherein the first control point is located at the upper left position of the current block, the second control point is located at the upper right position of the current block, and the third control point is located at the lower left position of the current block; Decoding information regarding a first motion vector difference for the first control point to derive the first motion vector difference; Decoding information regarding a difference between two motion vector differences for the second control point, the difference between the two motion vector differences being the difference between a second motion vector difference for the second control point and the first motion vector difference, to derive the difference between the two motion vector differences for the second control point; Decoding information regarding a difference between two motion vector differences for the third control point, the difference between the two motion vector differences being the difference between a third motion vector difference for the third control point and the first motion vector difference, to derive the difference between the two motion vector differences for the third control point; Deriving a first motion vector for the first control point based on the first motion vector predictor and the first motion vector difference; Deriving a second motion vector for the second control point based on the second motion vector predictor and the second motion vector difference; Deriving a third motion vector for the third control point based on the third motion vector predictor and the third motion vector difference; Generating a prediction sample for the current block based on the first motion vector, the second motion vector, and the third motion vector; Generating a residual sample for the current block based on the residual information; Generating a reconstructed sample based on the prediction sample and the residual sample; and Applying deblocking filtering to the reconstructed sample, wherein the second motion vector difference is derived based on the first motion vector difference and the difference between the two motion vector differences for the second control point, and wherein the third motion vector difference is derived based on the first motion vector difference and the difference between the two motion vector differences for the third control point.

2. An image encoding method performed by an encoding device, the method comprising: Deriving a first motion vector predictor for a first control point of the current block, a second motion vector predictor for a second control point of the current block, and a third motion vector predictor for a third control point of the current block based on neighboring blocks of the current block, wherein the first control point is located at the upper left position of the current block, the second control point is located at the upper right position of the current block, and the third control point is located at the lower left position of the current block; Deriving a first motion vector difference for the first control point; Deriving a difference between two motion vector differences for the second control point based on a second motion vector difference for the second control point and the first motion vector difference; Derive a difference between two motion vector differences for the third control point based on the third motion vector difference for the third control point and the first motion vector difference; Generate a prediction sample for the current block based on a first motion vector, a second motion vector, and a third motion vector; Generate a residual sample for the current block based on the prediction sample, and Encode image information including information about the first motion vector difference, information about the difference between two motion vector differences for the second control point, information about the difference between two motion vector differences for the third control point, and residual information related to the residual sample to output a bitstream, wherein the second motion vector difference is derived based on the second motion vector for the second control point and the second motion vector predictor for the second control point, and wherein the third motion vector difference is derived based on the third motion vector for the third control point and the third motion vector predictor for the third control point.

3. A non - transitory computer - readable storage medium storing a bitstream generated by an image encoding method, the method comprising: Derive a first motion vector predictor for a first control point of the current block, a second motion vector predictor for a second control point of the current block, and a third motion vector predictor for a third control point of the current block based on neighboring blocks of the current block, wherein the first control point is located at the upper - left position of the current block, the second control point is located at the upper - right position of the current block, and the third control point is located at the lower - left position of the current block; Derive a first motion vector difference for the first control point; Derive a difference between two motion vector differences for the second control point based on the second motion vector difference for the second control point and the first motion vector difference; Derive a difference between two motion vector differences for the third control point based on the third motion vector difference for the third control point and the first motion vector difference; Generate a prediction sample for the current block based on a first motion vector, a second motion vector, and a third motion vector; Generate a residual sample for the current block based on the prediction sample, and Encode image information including information about the first motion vector difference, information about the difference between two motion vector differences for the second control point, information about the difference between two motion vector differences for the third control point, and residual information related to the residual sample to output a bitstream, wherein the second motion vector difference is derived based on the second motion vector for the second control point and the second motion vector predictor for the second control point, and wherein the third motion vector difference is derived based on the third motion vector for the third control point and the third motion vector predictor for the third control point.

4. A method for transmitting data of an image, the method comprising: Obtain a bitstream for the image, wherein the bitstream is generated based on: deriving a first motion vector predictor for a first control point of the current block, a second motion vector predictor for a second control point of the current block, and a third motion vector predictor for a third control point of the current block based on neighboring blocks of the current block; deriving a first motion vector difference for the first control point; deriving a difference between two motion vector differences for the second control point based on the second motion vector difference for the second control point and the first motion vector difference; deriving a difference between two motion vector differences for the third control point based on the third motion vector difference for the third control point and the first motion vector difference; generating a prediction sample for the current block based on a first motion vector, a second motion vector, and a third motion vector; generating a residual sample for the current block based on the prediction sample; and encoding image information including information about the first motion vector difference, information about the difference between the two motion vector differences for the second control point, information about the difference between the two motion vector differences for the third control point, and residual information related to the residual sample; and Transmit the data including the bitstream, wherein the first control point is located at the upper left position of the current block, the second control point is located at the upper right position of the current block, and the third control point is located at the lower left position of the current block, wherein the second motion vector difference is derived based on the second motion vector for the second control point and the second motion vector predictor for the second control point, and wherein the third motion vector difference is derived based on the third motion vector for the third control point and the third motion vector predictor for the third control point.