Method and apparatus for inter prediction in video processing system
By adopting the affine motion prediction method in video encoding, using the difference between the control point and the motion vector difference for inter-frame prediction, the problem of high-resolution image transmission and storage costs is solved, and more efficient encoding and transmission is achieved.
Patent Information
- Application Number
- CN202411239398.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2018-04-13
- Filing Date
- 2019-04-11
- Publication Date
- 2025-07-18
AI Technical Summary
The prior art has high transmission and storage costs when processing high-resolution and high-quality images, and requires improving video encoding efficiency.
Affine motion prediction method is adopted, by derive the control point (CP) of the current block, the motion vector predictor (MVP) and the motion vector difference (MVD) are derived based on the adjacent block, and inter prediction is used to reduce the amount of motion vector data to improve coding efficiency.
Improve the accuracy and encoding efficiency of inter-frame prediction, reduce the amount of data, and reduce the transmission and storage costs.
Smart Images

Figure CN120343282A_ABST
Abstract
Description
[0001] This application is a divisional application of the patent application with the application number 201980032706.7 (PCT / KR2019 / 004334), the international filing date of which is April 11, 2019, and the invention title of which is "Method and Apparatus for Inter-Frame Prediction in a Video Processing System", and which was filed on November 16, 2020. Technical Field
[0002] The present disclosure generally relates to video coding techniques, and more particularly, to an inter-frame prediction method and apparatus in a video processing system. Background Art
[0003] The demand for high-resolution and high-quality images such as high-definition (HD) images and ultra-high-definition (UHD) images is increasing in various fields. Since the image data has high resolution and high quality, the amount of information or bits to be transmitted increases relative to conventional image data. Therefore, when transmitting image data using a medium such as a conventional wired / wireless broadband line or storing image data using an existing storage medium, the transmission cost and storage cost increase.
[0004] Therefore, an efficient image compression technique for effectively transmitting, storing, and reproducing information of high-resolution and high-quality images is needed. Summary of the Invention
[0005] The technical object of the present disclosure is to provide a method and apparatus for increasing video coding efficiency.
[0006] Another technical object of the present disclosure is to provide a method and apparatus for processing an image using affine motion prediction.
[0007] Another technical object of the present disclosure is to provide a method and apparatus for performing inter-frame prediction based on a sample unit motion vector.
[0008] Another technical object of the present disclosure is to provide a method and apparatus for deriving a sample unit motion vector based on a motion vector of a control point for a current block.
[0009] Another technical object of the present disclosure is to provide a method and apparatus for improving coding efficiency by using a difference between differences of motion vector differences of control points of a current block.
[0010] Another technical object of the present disclosure is to provide a method and apparatus for deriving a motion vector predictor of another control point based on a control point motion vector predictor.
[0011] Another technical problem of the present disclosure is to provide a method and apparatus for deriving a control point motion vector predictor based on a motion vector of a reference region adjacent to a control point.
[0012] Embodiments of the present disclosure provide an inter-frame prediction method performed by a decoding device. The inter-frame prediction method includes: deriving control points (CPs) for a current block, where the CPs include a first CP and a second CP; deriving a first motion vector predictor (MVP) for the first CP and a second MVP for the second CP based on neighboring blocks of the current block, decoding a first motion vector difference (MVD) for the first CP, decoding a difference (DMVD) between two MVDs for the second CP, deriving a first motion vector (MV) for the first CP based on the first MVP and the first MVD, deriving a second MV for the second CP based on the second MVP and the DMVD for the second CP, and generating a prediction block for the current block based on the first MV and the second MV, where the DMVD for the second CP represents the difference between the first MVD and the second MVD for the second CP.
[0013] According to another example of the present disclosure, a video encoding method performed by an encoding device is provided. The encoding method includes: deriving control points (CPs) for a current block, where the CPs include a first CP and a second CP; deriving a first motion vector predictor (MVP) for the first CP and a second MVP for the second CP based on neighboring blocks of the current block, deriving a first motion vector difference (MVD) for the first CP, deriving a difference (DMVD) between two MVDs for the second CP, and encoding image information including information about the first MVD and information about the DMVD for the second CP to output a bitstream, where the DMVD for the second CP represents the difference between the first MVD and the second MVD for the second CP.
[0014] According to still another embodiment of the present disclosure, a decoding device that performs an inter-frame prediction method is provided. The decoding device includes an entropy decoder that decodes a first motion vector difference (MVD) for the first CP and a difference (DMVD) between two MVDs for the second CP; and a predictor that derives control points (CPs) for a current block, including a first CP and a second CP, derives a first motion vector predictor (MVP) for the first CP and a second MVP for the second CP based on neighboring blocks of the current block, derives a first motion vector (MV) for the first CP based on the first MVP and the first MVD, derives a second MV for the second CP based on the second MVP and the DMVD for the second CP, and generates a prediction block for the current block based on the first MV and the second MV, where the DMVD for the second CP represents the difference between the first MVD and the second MVD for the second CP.
[0015] According to yet another embodiment of the present disclosure, there is provided an encoding apparatus for performing video encoding. The encoding apparatus includes a predictor that derives control points (CPs) for a current block, including a first CP and a second CP, derives a first motion vector predictor (MVP) for the first CP and a second MVP for the second CP based on neighboring blocks of the current block, derives a first motion vector difference (MVD) for the first CP, and derives a difference (DMVD) between two MVDs for the second CP, and an entropy encoder that encodes image information including information about the first MVD and information about the DMVD for the second CP to output a bitstream, where the DMVD for the second CP represents the difference between the first MVD and a second MVD for the second CP.
[0016] According to the present disclosure, a more accurate sample unit motion vector for a current block can be derived, and the inter-frame prediction efficiency can be significantly increased.
[0017] According to the present disclosure, the motion vector of a sample of a current block can be effectively derived based on the motion vectors of the control points for the current block.
[0018] According to the present disclosure, by sending the difference between the motion vectors of the control points for a current block and / or the difference between the motion vector differences, the data amount of the motion vectors for the control points can be removed or reduced, and the overall encoding efficiency can be improved. Description of the Drawings
[0019] Figure 1 is a block diagram schematically illustrating a video encoding apparatus according to an embodiment of the present disclosure.
[0020] Figure 2 is a block diagram schematically illustrating a video decoding apparatus according to an embodiment of the present disclosure.
[0021] Figure 3 Illustratively represents a content streaming system according to an embodiment of the present disclosure.
[0022] Figure 4 is a flowchart for illustrating a method of deriving a motion vector prediction value from neighboring blocks according to an embodiment of the present disclosure.
[0023] Figure 5 Illustratively represents an affine motion model according to an embodiment of the present disclosure.
[0024] Figure 6 Illustratively represents a simplified affine motion model according to an embodiment of the present disclosure.
[0025] Figure 7 is a diagram for describing a method of deriving a motion vector predictor at a control point according to an embodiment of the present disclosure.
[0026] Figure 8 Illustratively shows two CPs for a 4-parameter affine motion model according to an embodiment of the present disclosure.
[0027] Figure 9 Illustratively shows a case where a median is additionally used in a 4-parameter affine motion model according to an embodiment of the present disclosure.
[0028] Figure 10 Illustratively shows three CPs for a 6-parameter affine motion model according to an embodiment of the present disclosure.
[0029] Figure 11 Schematically shows a video encoding method of an encoding device according to the present disclosure.
[0030] Figure 12 Schematically illustrates an inter-frame prediction method of a decoding device according to the present disclosure. Detailed implementation
[0031] Although the present disclosure may be modified in various forms, specific embodiments of the present disclosure will be described in detail and illustrated in the drawings. However, this is not intended to limit the present disclosure to specific embodiments. The terms used in this specification are only used to describe specific embodiments and are not intended to limit the technical idea of the present disclosure. Unless the context clearly indicates otherwise, the singular form may include the plural form. Terms such as "comprising (or including)" and "having (or being provided with)" are intended to indicate the presence of features, numbers, steps, operations, components, parts, or combinations thereof written in the following description, and thus should not be construed as precluding the possibility of the presence or addition of one or more different features, numbers, steps, operations, components, parts, or combinations thereof.
[0032] Meanwhile, for ease of explaining different specific functions, the configurations in the drawings described in the present disclosure are drawn independently in a video encoding device / decoding device, but this does not mean that the configuration is implemented by independent hardware or independent software. For example, two or more configurations may be combined to form a single configuration, and one configuration may be divided into multiple configurations. Embodiments in which configurations are combined and / or configurations are divided without departing from the concept of the present disclosure also belong to the present disclosure.
[0033] In the present disclosure, the terms " / " and "," should be interpreted as indicating "and / or". For example, the expression "A / B" may mean "A and / or B", and "A, B" may mean "A and / or B". In addition, "A / B / C" may mean "at least one of A, B, and / or C". Further, "A, B, C" may mean "at least one of A, B, and / or C".
[0034] In addition, in the present disclosure, the term "or" should be interpreted to indicate "and / or". For example, the expression "A or B" may include 1) only A, 2) only B, and / or 3) both A and B. In other words, the term "or" in this document may be interpreted to indicate "additionally or alternatively."
[0035] The present disclosure can be modified in various forms, and its specific embodiments will be described and illustrated in the drawings. However, these embodiments are not intended to limit the present disclosure. The terms used in the following description are only for describing specific embodiments and are not intended to limit the present disclosure. A singular expression includes a plural expression as long as it is clearly read differently. Terms such as "including" and "having" are intended to indicate the presence of the features, numbers, steps, operations, elements, components, or combinations thereof used in the following description, and thus it should be understood that there is no exclusion of the possibility of the presence or addition of one or more different features, numbers, steps, operations, elements, components, or combinations thereof.
[0036] In addition, for the convenience of explaining different specific functions, the elements in the drawings described in this embodiment are drawn independently, which does not mean that these elements are implemented by independent hardware or independent software. For example, two or more of the elements may be combined to form a single element, or one element may be divided into multiple elements. Embodiments in which the elements are combined and / or divided without departing from the concept of this embodiment belong to the present disclosure.
[0037] The following description can be applied to the technical field of processing video, images, or pictures. For example, the methods or exemplary embodiments disclosed in the following description may be associated with the disclosure of the common video coding (VVC) standard (Recommendation ITU-T H.266), the next-generation video / image coding standard after VVC, or the standards before VVC (e.g., the high efficiency video coding (HEVC) standard (Recommendation ITU-T H.265), etc.).
[0038] Hereinafter, examples of this embodiment will be described in detail with reference to the drawings. In addition, throughout the drawings, like reference numerals are used to indicate like elements, and the same description of like elements will be omitted.
[0039] In the present disclosure, video may mean a collection of a series of images over time. Generally, a picture means a unit representing an image at a specific time, and a slice is a unit that forms a part of a picture. A picture may be composed of multiple slices, and the terms picture and slice may be mixed with each other as needed according to the occasion.
[0040] A pixel or pel can mean the smallest unit that makes up a picture (or image). Additionally, the term "sample" can be used as a term corresponding to a pixel. A sample can generally represent a pixel or the value of a pixel, can represent only the pixel (pixel value) of the luminance component, and can represent only the pixel (pixel value) of the chrominance component.
[0041] A unit indicates a basic unit of image processing. A unit can include at least one of a specific region and information related to that region. Optionally, a unit can be mixed with terms such as block, region, etc. Typically, an M×N block can represent a set of samples or transform coefficients arranged in M columns and N rows.
[0042] Figure 1 is a block diagram briefly illustrating the structure of an encoding device according to an embodiment of the present disclosure. Hereinafter, the encoding / decoding device can include a video encoding / decoding device and / or an image encoding / decoding device, and the video encoding / decoding device can be used as a concept including the image encoding / decoding device, or the image encoding / decoding device can be used as a concept including the video encoding / decoding device.
[0043] Reference Figure 1 , the video encoding device 100 can include a picture splitter 105, a predictor 110, a residual processor 120, an entropy encoder 130, an adder 140, a filter 150, and a memory 160. The residual processor 120 can include a subtractor 121, a transformer 122, a quantizer 123, a reorderer 124, an inverse quantizer 125, and an inverse transformer 126.
[0044] The picture splitter 105 can separate an input picture into at least one processing unit.
[0045] In an example, the processing unit can be referred to as a coding unit (CU). In this case, the coding unit can be recursively separated from the largest coding unit (LCU) according to a quadtree binary tree (QTBT) structure. For example, one coding unit can be separated into multiple coding units of a deeper depth based on a quadtree structure, a binary tree structure, and / or a ternary tree structure. In this case, for example, the quadtree structure can be applied first, and the binary tree structure and the ternary tree structure can be applied later. Alternatively, the binary tree structure / ternary tree structure can be applied first. The encoding process according to the present embodiment can be performed based on the final coding unit that is no longer further separated. In this case, the largest coding unit can be used as the final coding unit based on image characteristics such as coding efficiency, or the coding unit can be recursively separated into coding units of a lower depth as needed and the coding unit with the optimal size can be used as the final coding unit. Here, the encoding process can include processes such as prediction, transformation, and reconstruction, which will be described later.
[0046] In another example, the processing unit may include an encoding unit (CU), a prediction unit (PU), or a transform unit (TU). The encoding unit may be separated from a largest coding unit (LCU) into coding units of a deeper depth according to a quadtree structure. In this case, the largest coding unit may be directly used as a final coding unit based on encoding efficiency or the like according to image characteristics, or the coding unit may be recursively separated into coding units of a deeper depth as needed, and the coding unit having an optimal size may be used as the final coding unit. When a smallest coding unit (SCU) is set, the coding unit may not be separated into a coding unit smaller than the smallest coding unit. Here, the final coding unit refers to a coding unit that is divided or separated into a prediction unit or a transform unit. The prediction unit is a unit divided from the coding unit and may be a unit for sample prediction. Here, the prediction unit may be divided into sub-blocks. The transform unit may be divided from the coding unit according to a quadtree structure, and the transform unit may be a unit for deriving transform coefficients and / or a unit for deriving a residual signal from the transform coefficients. Hereinafter, the coding unit may be referred to as a coding block (CB), the prediction unit may be referred to as a prediction block (PB), and the transform unit may be referred to as a transform block (TB). The prediction block or the prediction unit may refer to a specific area in the form of a block in a picture and includes an array of prediction samples. Additionally, the transform block or the transform unit may refer to a specific area in the form of a block in a picture and includes an array of transform coefficients or residual samples.
[0047] The predictor 110 may perform prediction on a block to be processed (hereinafter, it may represent a current block or a residual block) and may generate a prediction block including prediction samples for the current block. The unit for performing prediction in the predictor 110 may be a coding block, or may be a transform block, or may be a prediction block.
[0048] The predictor 110 may determine whether to apply intra prediction or inter prediction to the current block. For example, the predictor 110 may determine whether to apply intra prediction or inter prediction in units of CUs.
[0049] In the case of intra prediction, the predictor 110 may derive a prediction sample for a current block based on reference samples outside the current block in a picture (hereinafter, the current picture) to which the current block belongs. In this case, the predictor 110 may derive the prediction sample based on an average value or interpolation of neighboring reference samples of the current block (case (i)), or may derive the prediction sample based on reference samples existing in a specific (prediction) direction with respect to the prediction sample among the neighboring reference samples of the current block (case (ii)). Case (i) may be referred to as a non-directional mode or a non-angular mode, and case (ii) may be referred to as a directional mode or an angular mode. In intra prediction, the prediction mode may include, as an example, 33 directional modes and at least two non-directional modes. The non-directional modes may include a DC mode and a planar mode. The predictor 110 may determine a prediction mode to be applied to the current block by using a prediction mode applied to an adjacent block.
[0050] In the case of inter prediction, the predictor 110 may derive a prediction sample for a current block based on samples specified by a motion vector on a reference picture. The predictor 110 may derive a prediction sample for the current block by applying any one of a skip mode, a merge mode, and a motion vector prediction (MVP) mode. In the case of the skip mode and the merge mode, the predictor 110 may use motion information of an adjacent block as motion information of the current block. In the case of the skip mode, different from the merge mode, a difference (residual) between the prediction sample and the original sample is not transmitted. In the case of the MVP mode, a motion vector of an adjacent block is used as a motion vector predictor to derive a motion vector of the current block.
[0051] In the case of inter prediction, adjacent blocks may include spatially adjacent blocks existing in the current picture and temporally adjacent blocks existing in a reference picture. A reference picture including temporally adjacent blocks may also be referred to as a collocated picture (colPic). Motion information may include a motion vector and a reference picture index. Information such as prediction mode information and motion information may be (entropy) encoded and then output in the form of a bitstream.
[0052] When motion information of a temporally adjacent block is used in the skip mode and the merge mode, the highest picture in the reference picture list may be used as a reference picture. Reference pictures included in the reference picture list may be aligned based on a picture order count (POC) difference between the current picture and the corresponding reference picture. The POC corresponds to a display order and may be distinguished from an encoding order.
[0053] The subtractor 121 generates a residual sample, which is a difference between the original sample and the prediction sample. If the skip mode is applied, the residual sample may not be generated as described above.
[0054] The transformer 122 transforms residual samples in units of transform blocks to generate transform coefficients. The transformer 122 may perform the transformation based on the size of the corresponding transform block and the prediction mode applied to a prediction block or a coded block that spatially overlaps with the transform block. For example, if intra prediction is applied to a prediction block or a coded block that overlaps with the transform block and the transform block is a 4×4 residual array, a discrete sine transform (DST) transform kernel may be used to transform the residual samples, and in other cases, a discrete cosine transform (DCT) transform kernel may be used to transform the residual samples.
[0055] The quantizer 123 may quantize the transform coefficients to generate quantized transform coefficients.
[0056] The reorderer 124 reorders the quantized transform coefficients. The reorderer 124 may reorder the quantized transform coefficients in block form into a one-dimensional vector by a coefficient scanning method. Although the reorderer 124 is described as a separate component, the reorderer 124 may be part of the quantizer 123.
[0057] The entropy encoder 130 may perform entropy coding on the quantized transform coefficients. The entropy coding may include coding methods such as exponential Golomb, context-adaptive variable length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), etc. In addition to the quantized transform coefficients, the entropy encoder 130 may also perform coding on the information required for video reconstruction (such as syntax element values, etc.) together or separately according to the entropy coding or according to a pre-configured method. The entropy-coded information may be sent or stored in units of network abstraction layer (NAL) in the form of a bitstream. The bitstream may be sent via a network or stored in a digital storage medium. Here, the network may include a broadcast network or a communication network, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SDD, etc.
[0058] The inverse quantizer 125 inverse-quantizes the values (transform coefficients) quantized by the quantizer 123, and the inverse transformer 126 inverse-transforms the values inverse-quantized by the inverse quantizer 125 to generate residual samples.
[0059] The adder 140 adds the residual samples to the prediction samples to reconstruct a picture. The residual samples may be added to the prediction samples in units of blocks to generate reconstructed blocks. Although the adder 140 is described as a separate component, the adder 140 may be part of the predictor 110. Additionally, the adder 140 may be referred to as a reconstructor or a reconstructed block generator.
[0060] Filter 150 may apply deblocking filtering and / or sample adaptive offset to the reconstructed picture. Artifacts at the block boundaries in the reconstructed picture or distortions during quantization may be corrected by deblocking filtering and / or sample adaptive offset. After the deblocking filtering is completed, the sample adaptive offset may be applied on a sample-by-sample basis. Filter 150 may apply an adaptive loop filter (ALF) to the reconstructed picture. The ALF may be applied to the reconstructed picture to which deblocking filtering and / or sample adaptive offset has been applied.
[0061] Memory 160 may store the reconstructed picture (decoded picture) or information required for encoding / decoding. Here, the reconstructed picture may be the reconstructed picture filtered by Filter 150. The stored reconstructed picture may be used as a reference picture for (inter-frame) prediction of other pictures. For example, Memory 160 may store the (reference) picture for inter-frame prediction. Here, the picture for inter-frame prediction may be specified according to a reference picture set or a reference picture list.
[0062] Figure 2 is a block diagram briefly illustrating a video / image decoding apparatus according to an embodiment of the present disclosure.
[0063] Hereinafter, the video decoding apparatus may include an image decoding apparatus.
[0064] Reference Figure 2 , video decoding apparatus 200 may include entropy decoder 210, residual processor 220, predictor 230, adder 240, filter 250, and memory 260. Residual processor 220 may include rearranger 221, dequantizer 222, and inverse transformator 223.
[0065] In addition, although not depicted, video decoding apparatus 200 may include a receiver for receiving a bitstream including video information. The receiver may be configured as a separate module or may be included in entropy decoder 210.
[0066] When inputting a bitstream including video / image information, video decoding apparatus 200 may reconstruct video / image / picture in association with the process of processing video information in a video encoding apparatus.
[0067] For example, video decoding apparatus 200 may perform video decoding using the processing units applied in a video encoding apparatus. Thus, the processing unit blocks for video decoding may be, for example, encoding units, and in another example, encoding units, prediction units, or transformers. Encoding units may be separated from the largest coding unit according to a quadtree structure and / or a binary tree structure and / or a ternary tree structure.
[0068] In some cases, a prediction unit and a transformer can be further used. In such cases, the prediction block is a block derived from or segmented from the coding unit and can be a unit of sample prediction. Here, the prediction unit can be divided into sub-blocks. The transformer can be separated from the coding unit according to a quadtree structure and can be a unit for deriving transform coefficients or a unit for deriving a residual signal from transform coefficients.
[0069] The entropy decoder 210 can parse the bitstream to output information required for video reconstruction or picture reconstruction. For example, the entropy decoder 210 can decode the information in the bitstream based on coding methods such as exponential Golomb coding, CAVLC, CABAC, etc., and can output the values of the syntax elements required for video reconstruction and the quantization values of the transform coefficients regarding the residuals.
[0070] More specifically, the CABAC entropy decoding method can receive the bins corresponding to each syntax element in the bitstream, use the decoding target syntax element information and the decoding information of adjacent blocks and the decoding target block or the information of the symbols / bins decoded in the previous step to determine the context model, predict the bin generation probability according to the determined context model, and perform arithmetic decoding of the bins to generate symbols corresponding to each syntax element value. Here, the CABAC entropy decoding method can use the information of the symbols / bins decoded by the context model decoding for the next symbol / bin to update the context model after determining the context model.
[0071] The information regarding prediction in the information decoded by the entropy decoder 210 can be provided to the predictor 250, and the residual values (i.e., the quantized transform coefficients) that have been entropy decoded by the entropy decoder 210 can be input to the rearranger 221.
[0072] The rearranger 221 can rearrange the quantized transform coefficients into a two-dimensional block form. The rearranger 221 can perform a rearrangement corresponding to the coefficient scanning performed by the encoding device. Although the rearranger 221 is described as a separate component, the rearranger 221 can be part of the inverse quantizer 222.
[0073] The inverse quantizer 222 can inverse quantize the quantized transform coefficients based on the (inverse) quantization parameters to output the transform coefficients. In this case, the information for deriving the quantization parameters can be signaled from the encoding device.
[0074] The inverse transformer 223 can perform an inverse transformation on the transform coefficients to derive the residual samples.
[0075] The predictor 230 can perform prediction on the current block and can generate a prediction block including prediction samples for the current block. The unit of prediction performed in the predictor 230 can be a coding block, or can be a transform block or can be a prediction block.
[0076] The predictor 230 may determine whether to apply intra prediction or inter prediction based on information about the prediction. In this case, the unit for determining which one to use between intra prediction and inter prediction may be different from the unit for generating prediction samples. In addition, the unit for generating prediction samples may also be different in inter prediction and intra prediction. For example, it may be determined on a CU-by-CU basis which one to apply between inter prediction and intra prediction. Further, for example, in inter prediction, prediction samples may be generated by determining a prediction mode on a PU-by-PU basis, and in intra prediction, prediction samples may be generated on a TU-by-TU basis by determining a prediction mode on a PU-by-PU basis.
[0077] In the case of intra prediction, the predictor 230 may derive prediction samples for the current block based on neighboring reference samples in the current picture. The predictor 230 may derive prediction samples for the current block by applying a directional mode or a non-directional mode based on neighboring reference samples of the current block. In this case, the prediction mode to be applied to the current block may be determined by using the intra prediction mode of adjacent blocks.
[0078] In the case of inter prediction, the predictor 230 may derive prediction samples for the current block based on samples specified in a reference picture according to a motion vector. The predictor 230 may use one of a skip mode, a merge mode, and an MVP mode to derive prediction samples for the current block. Here, the motion information (e.g., motion vector and information about a reference picture index) required for inter prediction of the current block provided by the video coding device may be obtained or derived based on information about the prediction.
[0079] In the skip mode and the merge mode, the motion information of adjacent blocks may be used as the motion information of the current block. Here, the adjacent blocks may include spatially adjacent blocks and temporally adjacent blocks.
[0080] The predictor 230 may use the motion information of available adjacent blocks to construct a merge candidate list, and use the information indicated by a merge index on the merge candidate list as the motion vector of the current block. The merge index may be signaled by the coding device. The motion information may include a motion vector and a reference picture. In the skip mode and the merge mode, when using the motion information of temporally adjacent blocks, the first sorted picture in the reference picture list may be used as the reference picture.
[0081] In the case of the skip mode, different from the merge mode, the difference (residual) between the prediction samples and the original samples is not sent.
[0082] In the case of the MVP mode, the motion vectors of adjacent blocks can be used as motion vector predictors to derive the motion vector of the current block. Here, the adjacent blocks can include spatially adjacent blocks and temporally adjacent blocks.
[0083] When applying the merge mode, for example, the motion vectors of the reconstructed spatially adjacent blocks and / or the motion vectors corresponding to the Col blocks as temporally adjacent blocks can be used to generate a merge candidate list. The motion vector of the candidate block selected from the merge candidate list is used as the motion vector of the current block in the merge mode. The above information about prediction can include a merge index that indicates the candidate block with the best motion vector selected from the candidate blocks included in the merge candidate list. Here, the predictor 230 can derive the motion vector of the current block using the merge index.
[0084] When applying the MVP (Motion Vector Prediction) mode as another example, the motion vectors of the reconstructed spatially adjacent blocks and / or the motion vectors corresponding to the Col blocks as temporally adjacent blocks can be used to generate a motion vector predictor candidate list. That is, the motion vectors of the reconstructed spatially adjacent blocks and / or the motion vectors corresponding to the Col blocks as temporally adjacent blocks can be used as motion vector candidates. The above information about prediction can include a predicted motion vector index that indicates the best motion vector selected from the motion vector candidates included in the list. Here, the predictor 230 can select the predicted motion vector of the current block from the motion vector candidates included in the motion vector candidate list using the motion vector index. The predictor of the encoding device can obtain the motion vector difference (MVD) between the motion vector of the current block and the motion vector predictor, encode the MVD, and output the encoded MVD in the form of a bitstream. That is, the MVD can be obtained by subtracting the motion vector predictor from the motion vector of the current block. Here, the predictor 230 can acquire the motion vector included in the information about prediction, and derive the motion vector of the current block by adding the motion vector difference to the motion vector predictor. In addition, the predictor can obtain or derive a reference picture index indicating the reference picture from the above information about prediction.
[0085] The adder 240 can add the residual samples to the predicted samples to reconstruct the current block or the current picture. The adder 240 can reconstruct the current picture by adding the residual samples to the predicted samples in units of blocks. When applying the skip mode, no residual is sent, and thus the predicted samples can become the reconstructed samples. Although the adder 240 is described as a separate component, the adder 240 can be a part of the predictor 230. In addition, the adder 240 can be referred to as a reconstructor or a reconstructed block generator.
[0086] Filter 250 may apply deblocking filter, sample adaptive offset, and / or ALF to the reconstructed picture. Here, the sample adaptive offset may be applied on a sample-by-sample basis after deblocking filter. ALF may be applied after deblocking filter and / or applying the sample adaptive offset.
[0087] Memory 260 may store the reconstructed picture (decoded picture) or information required for decoding. Here, the reconstructed picture may be the reconstructed picture filtered by filter 250. For example, memory 260 may store pictures for inter prediction. Here, the pictures for inter prediction may be specified according to a reference picture set or a reference picture list. The reconstructed picture may be used as a reference picture for other pictures. Memory 260 may output the reconstructed pictures in the output order.
[0088] In addition, as described above, when performing video encoding, prediction is performed to improve compression efficiency. Thus, a prediction block including prediction samples for a current block that is a block to be encoded (i.e., an encoding target block) may be generated. Here, the prediction block includes prediction samples in the spatial domain (or pixel domain). The prediction block is derived in the same manner in the encoding device and the decoding device, and the encoding device may signal information about the residual between the original block and the prediction block (residual information) instead of the original sample values of the original block to the decoding device, thereby improving image encoding efficiency. The decoding device may derive a residual block including residual samples based on the residual information, add the residual block to the prediction block to generate a reconstructed block including reconstructed samples, and generate a reconstructed picture including the reconstructed block.
[0089] The residual information may be generated through a transform and quantization process. For example, the encoding device may derive a residual block between the original block and the prediction block, perform a transform process on the residual samples (residual sample array) included in the residual block to derive transform coefficients, perform a quantization process on the transform coefficients to derive quantized transform coefficients, and may signal the relevant residual information to the decoding device (through a bitstream). Here, the residual information may include value information, position information, transform technique, transform kernel, and quantization parameter of the quantized transform coefficients, etc. The decoding device may perform an inverse quantization / inverse transform process based on the residual information and may derive residual samples (or a residual block). The decoding device may generate a reconstructed picture based on the prediction block and the residual block. In addition, for inter prediction of future reference pictures, the encoding device may also inverse quantize / inverse transform the quantized transform coefficients to derive a residual block and generate a reconstructed picture based on the residual block.
[0090] Figure 3 Illustratively shows a content streaming system according to an embodiment of the present disclosure.
[0091] Reference Figure 3, the embodiments described in this disclosure can be embodied and executed on a processor, microprocessor, controller, or chip. For example, the functional units shown in each drawing can be embodied and executed on a computer, processor, microprocessor, controller, or chip. In this case, the information for the embodiments (e.g., information about instructions) or algorithms can be stored in a digital storage medium.
[0092] In addition, the decoding device and encoding device applied in this disclosure can be included in a multimedia broadcast transceiver, a mobile communication terminal, a home theater video device, a digital cinema video device, a surveillance camera, a video chat device, a real-time communication device such as video communication, a mobile streaming device, a storage medium, a portable camera, a video-on-demand (VoD) service providing device, an over-the-top (OTT) video device, an Internet streaming service providing device, a three-dimensional (3D) video device, a video phone video device, and a medical video device, and can be used to process video signals or data signals. For example, an over-the-top (OTT) video device can include a game console, a Blu-ray player, an Internet access TV, a home theater system, a smartphone, a tablet, a digital video recorder (DVR), and so on.
[0093] In addition, the processing method applied in this disclosure can be generated in the form of a program executable by a computer and stored in a computer-readable recording medium. Multimedia data having a data structure according to this disclosure can also be stored in a computer-readable recording medium. The computer-readable recording medium includes various storage devices and distributed storage devices in which computer-readable data is stored. The computer-readable recording medium can include, for example, a Blu-ray Disc (BD), a Universal Serial Bus (USB), a ROM, a PROM, an EPROM, an EEPROM, a RAM, a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device. In addition, the computer-readable recording medium includes a medium embodied in a carrier wave (e.g., transmission via the Internet). Additionally, the bitstream generated by an encoding method can be stored in a computer-readable recording medium or transmitted via a wired or wireless communication network.
[0094] Furthermore, the embodiments of this disclosure can be embodied as a computer program product by program code, and the program code can be executed on a computer by the embodiments of this disclosure. The program code can be stored on a computer-readable carrier.
[0095] The content streaming system applied in this disclosure can mainly include an encoding server, a streaming server, a web server, a media storage device, a user device, and a multimedia input device.
[0096] The encoding server is used to compress the content input from a multimedia input device (such as a smart phone, a camera, a portable video camera, etc.) into digital data, generate a bitstream, and send it to the streaming server. As another example, in the case where a multimedia input device such as a smart phone, a camera, a portable video camera, etc. directly generates a bitstream, the encoding server can be omitted.
[0097] The bitstream can be generated by the encoding method or the bitstream generation method applied in the present disclosure. And the streaming server can temporarily store the bitstream during the process of sending or receiving the bitstream.
[0098] The streaming server sends the multimedia data to the user device through the web server based on the user's request. The web server is used as a tool to notify the user of what services exist. When the user requests the service the user wants, the web server transmits the request to the streaming server, and the streaming server sends the multimedia data to the user. In this regard, the content streaming system can include a separate control server, and in this case, the control server is used to control the commands / responses between the respective devices in the content streaming system.
[0099] The streaming server can receive the content from the media storage device and / or the encoding server. For example, in the case of receiving the content from the encoding server, the content can be received in real time. In this case, the streaming server can store the bitstream for a predetermined period of time to smoothly provide the streaming service.
[0100] For example, the user device can include a mobile phone, a smart phone, a laptop computer, a digital broadcast terminal, a personal digital assistant (PDA), a portable multimedia player (PMP), a navigation device, a tablet PC, a tablet computer, a superbook, a wearable device (such as a watch-type terminal (smart watch), a glasses-type terminal (smart glasses), a head-mounted display (HMD)), a digital TV, a desktop computer, a digital sign, etc.
[0101] Each server in the content streaming system can operate as a distributed server, and in this case, the data received by each server can be processed in a distributed manner.
[0102] Hereinafter, the inter-frame prediction method described with reference to Figure 1 and Figure 2 will be described in detail.
[0103] Figure 4 is a flowchart for illustrating a method of deriving a motion vector prediction value from neighboring blocks according to an embodiment of the present disclosure.
[0104] In the case of the motion vector prediction (MVP) mode, the encoder predicts the motion vector according to the type of the prediction block, and sends the difference between the best motion vector and the predicted value to the decoder. In this case, the encoder sends the motion vector difference, neighboring block information, reference index, etc. to the decoder. Here, the MVP mode can also be referred to as the advanced motion vector prediction (AMVP) mode.
[0105] The encoder can construct a prediction candidate list for motion vector prediction, and the prediction candidate list can include at least one of a spatial candidate block and a temporal candidate block.
[0106] First, the encoder can search for a spatial candidate block for motion vector prediction and insert it into the prediction candidate list (S410). For the process of constructing the spatial candidate block, a method of constructing a conventional spatial merge candidate in inter-frame prediction according to the merge mode can be applied.
[0107] The encoder can check whether the number of spatial candidate blocks is less than two (S420).
[0108] In the case where the number of spatial candidate blocks is less than two as a result of the check, the encoder can search for a temporal candidate block and insert it into the prediction candidate list (S430). At this time, in the case where no temporal candidate block is available, the encoder can use a zero motion vector as the motion vector prediction value (S440). For the process of constructing the temporal candidate block, a method of constructing a conventional temporal merge candidate in inter-frame prediction according to the merge mode can be applied.
[0109] On the other hand, in the case where the number of spatial candidate blocks is equal to or greater than two as a result of the check, the encoder can end the construction of the prediction candidate list and select a block with the minimum cost from among the candidate blocks. The encoder can determine the motion vector of the selected candidate block as the motion vector prediction value of the current block, and obtain the motion vector difference by using the motion vector prediction value. The motion vector difference obtained in this way can be sent to the decoder.
[0110] Figure 5 Illustratively represents an affine motion model according to an embodiment of the present disclosure.
[0111] The affine mode can be one of various prediction modes in inter-frame prediction, and the affine mode can also be referred to as the affine motion mode or the sub-block motion prediction mode. The affine mode can refer to a mode that performs an affine motion prediction method using an affine motion model.
[0112] The affine motion prediction method can derive the motion vector of a sample unit by using two or more motion vectors in the current block. In other words, the affine motion prediction method can improve the coding efficiency by determining the motion vector in units of samples rather than in units of blocks.
[0113] The general motion model may include a translation model, and motion estimation (ME) and motion compensation (MC) are performed based on the translation model that effectively represents simple motion. However, the translation model may not be effectively applicable to complex motions in natural videos, such as zooming in, zooming out, rotating, and other irregular motions. Therefore, embodiments of the present disclosure may use an affine motion model that can be effectively applied to complex motions.
[0114] Reference Figure 5 , the affine motion model may include four motion models, but they are exemplary motion models, and the scope of the present disclosure is not limited thereto. The above four motions may include translation, scaling, rotation, and shear. Here, the motion models for translation, scaling, and rotation may be referred to as a simplified affine motion model.
[0115] Figure 6 Illustratively represents a simplified affine motion model according to an embodiment of the present disclosure.
[0116] In affine motion prediction, control points (CPs) may be defined to use the affine motion model, and the motion vectors of sub-blocks or sample units included in a block may be determined by using two or more control point motion vectors (CPMVs). Here, the set of motion vectors of sample units or the set of motion vectors of sub-blocks may be referred to as an affine motion vector field (affine MVF).
[0117] Reference Figure 6 , the simplified affine motion model may mean a model for determining the motion vectors of sample units or sub-blocks using CPMV according to two CPs, and may also be referred to as a 4-parameter affine model. In Figure 6 , v0 and v1 may represent two CPMVs, and each arrow in the sub-block may represent the motion vector of the sub-block unit.
[0118] In other words, during the encoding / decoding process, the affine motion vector field may be determined in sample units or sub-block units. Here, the sample unit may refer to a pixel unit, and the sub-block unit may refer to a predefined block unit. When the affine motion vector field is determined in sample units, the motion vectors may be obtained based on each pixel value, and in the case of block units, the motion vectors of the corresponding blocks may be obtained based on the center pixel value of the blocks.
[0119] Figure 7 Is a diagram for describing a method of deriving a motion vector predictor at control points according to an embodiment of the present disclosure.
[0120] The affine mode may include an affine merge mode and an affine motion vector prediction (MVP) mode. The affine merge mode may be referred to as a sub-block merge mode, and the affine MVP mode may be referred to as an affine inter mode.
[0121] In the affine MVP mode, the CPMV of the current block may be derived based on a control point motion vector predictor (CPMVP) and a control point motion vector difference. In other words, the encoding device may determine the CPMVP of the CPMV of the current block, derive the CPMVD which is the difference between the CPMV of the current block and the CPMVP, and signal the information about the CPMVP and the information about the CPMVD to the decoding device. Here, the affine MVP mode may construct an affine MVP candidate list based on neighboring blocks, and the affine MVP candidate list may be referred to as a CPMVP candidate list. In addition, the information about the CPMVP may include an index indicating a block or a motion vector to be referenced from among the affine MVP candidate list.
[0122] Reference Figure 7 , the motion vector of the control point at the upper left sample position of the current block may be denoted as v0, the motion vector of the control point at the upper right sample position may be denoted as v1, the motion vector of the control point at the lower left sample position may be denoted as v2, and the motion vector of the control point at the lower right sample position may be denoted as v3.
[0123] For example, if two control points are used in the affine mode and the two control points are located at the upper left sample position and the upper right sample position, the motion vector of the sample unit or sub-block unit may be derived based on the motion vectors v0 and v1.
[0124] The motion vector v0 may be derived based on at least one motion vector of neighboring blocks A, B, and C at the upper left sample position. Here, the neighboring block A may represent the block located diagonally above the upper left sample position of the current block, the neighboring block B may represent the block located at the top of the upper left sample position of the current block, and the neighboring block C may represent the block located to the left of the upper left sample position of the current block.
[0125] The motion vector v1 may be derived based on at least one motion vector of neighboring blocks D and E at the upper right sample position. Here, the neighboring block D may represent the block located at the top of the upper right sample position of the current block, and the neighboring block E may represent the block located diagonally above the upper right sample position of the current block.
[0126] For example, if three control points are used in the affine mode and the three control points are located at the upper left sample position, the upper right sample position, and the lower left sample position, the motion vector of the sample unit or sub-block unit may be derived based on the motion vectors v0, v1, and v2. In other words, the motion vector v2 may be further used.
[0127] The motion vector v2 can be derived based on at least one motion vector of neighboring blocks F and G at the lower-left sample position. Here, the neighboring block F can represent a block located to the left of the lower-left sample position of the current block, and the neighboring block G can represent a block located diagonally below and to the left of the lower-left sample position of the current block.
[0128] The affine MVP mode can derive a CPMVP candidate list based on neighboring blocks and select the CPMVP pair with the highest correlation among the CPMVP candidate list as the CPMV of the current block. Information about the above CPMV can include an index indicating the CPMVP pair selected from the CPMVP candidate list.
[0129] Figure 8 Illustratively represent two CPs for a four-parameter affine motion model according to an embodiment of the present disclosure.
[0130] Embodiments of the present disclosure can use two CPs. The two CPs can be located at the upper-left sample position and the upper-right sample position of the current block, respectively. Here, the CP located at the upper-left sample position can be represented as CP0, and the CP located at the upper-right sample position can be represented as CP1. The motion vector at CP0 can be represented as mv0 and the motion vector at CP1 can be represented as mv1. The coordinates of each control point (CP i ) can be defined as (x i , y i ), i=0, 1 , and the motion vector at each control point can be represented as mvi = (v xi , v yi ), i=0, 1 .
[0131] In embodiments of the present disclosure, the CP located at the upper-left sample position can be represented as CP1, and the CP located at the upper-right sample position can be represented as CP0. In this case, the following process can be similarly performed considering the switching positions of CP0 and CP1.
[0132] For example, if the width of the current block is W and its height is H, assuming the coordinates of the lower-left sample position of the current block are (0, 0), the coordinates of CP0 can be represented as (0, H), and the coordinates of CP1 can be represented as (W, H). Here, W and H can have different values, but can also have the same value, and the reference (0, 0) can be set differently.
[0133] As Figure 8As shown, since an affine motion model using two motion vectors based on two CPs uses four parameters according to the two motion vectors in an affine motion prediction method, it may be referred to as a 4-parameter affine motion model or a simplified affine motion model.
[0134] In an embodiment of the present disclosure, the motion vector of a sample unit may be determined by an affine motion vector field (affine MVF) and the position of the sample. The affine motion vector field may represent the motion vector of the sample unit based on two motion vectors according to two CPs. In other words, when the sample position is (x, y) as shown in Equation 1, the affine motion vector field may derive the motion vector (v x , vy) of the corresponding sample.
[0135] [Equation 1]
[0136]
[0137] In Equation 1, v0x and v0y may mean the (x, y) coordinate components of the motion vector mv0 at CP0, while v1x and v1y may mean the (x, y) coordinate components of the motion vector mv1 at CP1. Similarly, w may mean the width of the current block.
[0138] Meanwhile, Equation 1 representing the affine motion model is only an example, and the equation for representing the affine motion model is not limited to Equation 1. For example, in some cases, the signs of each coefficient disclosed in Equation 1 may be changed according to the signs of Equation 1.
[0139] In other words, according to an embodiment of the present disclosure, a reference block among the temporal and / or spatial neighboring blocks of the current block may be determined, and the motion vector of the reference block may be used as a motion vector predictor of the current block, and the motion vector of the current block may be expressed by the motion vector predictor and the motion vector difference. In addition, an embodiment of the present disclosure may signal the index of the motion vector predictor and the motion vector difference.
[0140] According to an embodiment of the present disclosure, at the time of encoding, a difference between two motion vectors according to two CPs may be derived based on two motion vectors according to two CPs and two motion vector predictors according to two CPs, and at the time of decoding, two motion vectors according to two CPs may be derived based on two motion vector predictors according to two CPs and a difference between two motion vectors according to two CPs. In other words, the motion vector at each CP may be composed of the sum of the motion vector predictor and the motion vector difference, as shown in Equation 2, and is similar to when using the motion vector prediction (MVP) mode or the advanced motion vector prediction (AMVP) mode.
[0141] [Equation 2]
[0142]
[0143] In Equation 2, mvp0 and mvp1 may represent motion vector predictors (MVPs) at each of CP0 and CP1, and mvd0 and mvd1 may represent motion vector differences (MVDs) at each of CP0 and CP1. Herein, mvp may be referred to as CPMVP, and mvd may be referred to as CPMVD.
[0144] Accordingly, the inter-frame prediction method according to the affine mode according to an embodiment of the present disclosure may encode and decode an index and motion vector differences (mvd0 and mvd1) at each CP. In other words, according to an embodiment of the present disclosure, motion vectors at CP0 and CP1 may be derived based on mvd0 and mvd1 at each of CP0 and CP1 of a current block and mvp0 and mvp1 according to an index, and inter-frame prediction may be performed by deriving a motion vector of a sample unit based on the motion vectors at CP0 and CP1.
[0145] The inter-frame prediction method according to another embodiment of the present disclosure may use one of a motion vector difference according to two CPs and a difference of two MVDs (DMVD). In other words, in another embodiment, when there are mvd0 and mvd1 according to CP0 and CP1, inter-frame prediction may be performed by encoding and decoding mvd0 and mvd1, a difference between mvd0 and mvd1, and one of the indexes at each CP. More specifically, in another embodiment of the present disclosure, mvd0 and DMVD (mvd0 - mvd1) may be signaled, and mvd1 and DMVD (mvd0 - mvd1) may be signaled.
[0146] That is, in another embodiment of the present disclosure, motion vector differences (mvd0 and mvd1) at CP0 and CP1 may be derived respectively based on a motion vector difference (mvd0 or mvd1) at CP0 or CP1 and a difference between motion vector differences (mvd0 and mvd1) of CP0 and CP1, and motion vectors at CP0 and CP1 may be derived respectively based on motion vectors (mvp0 and mvp1) of CP0 and CP1 indicated by an index together with such motion vector differences, and inter-frame prediction may be performed by deriving a motion vector of a sample unit based on the motion vectors at CP0 and CP1.
[0147] Here, data on a difference of two motion vector differences (DMVD) is closer to a zero motion vector (zero MV) than data on a normal motion vector difference, and encoding can be performed more effectively as compared with the case of other embodiments according to the present disclosure.
[0148] Figure 9Illustratively shows a case where a median value is additionally used in a 4-parameter affine motion model according to an embodiment of the present disclosure.
[0149] In inter prediction according to an affine mode according to an embodiment of the present disclosure, a median predictor of a motion vector predictor may be used to perform adaptive motion vector coding.
[0150] Reference Figure 9 , according to an embodiment of the present disclosure, two CPs may be used, and information about CPs at different positions may be derived based on the two CPs. Here, the information related to the two CPs may be the same as the two CPs of Figure 8 . Additionally, a CP at another position derived based on the two CPs may be referred to as CP2 and may be located at the lower left sample position of the current block.
[0151] The information about the CP at another position may include a motion vector predictor (mvp2) of CP2, and mvp2 may be derived in two ways.
[0152] One method of deriving mvp2 is as follows. mvp2 may be derived based on a motion vector predictor mvp0 of CP0 and a motion vector predictor mvp1 of CP1, and may be derived as shown in Equation 3.
[0153] [Equation 3]
[0154]
[0155] In Equation 3, mvp 0x and mvp 0y may mean the (x, y) coordinate components of the motion vector predictor (mvp0) at CP0, mvp 1x and mvp 1y may mean the (x, y) coordinate components of the motion vector predictor (mvp1) at CP1, while mvp 2x and mvp 2y may mean the (x, y) coordinate components of the motion vector predictor (mvp2) at CP2. Additionally, h may represent the height of the current block, and w may represent the width of the current block.
[0156] Another method of deriving mvp2 is as follows. mvp2 may be derived based on neighboring blocks of CP2. Referring to Figure 9 , CP2 may be located at the lower left sample position of the current block, and mvp2 may be derived based on neighboring block A or neighboring block B of CP2. More specifically, mvp2 may be selected as one of the motion vectors of neighboring block A and neighboring block B.
[0157] In an embodiment of the present disclosure, mvp2 can be derived, and a median value can be derived based on mvp0, mvp1, and mvp2. Here, the median value can mean the value located in the middle in terms of size among a plurality of values. Therefore, the median value can be selected from mvp0, mvp1, and mvp2.
[0158] According to an embodiment of the present disclosure, when the median value is equal to mvp0, the motion vector difference (mvd0) of CP0 and DMVD (mvd0 - mvd1) can be signaled for inter-frame prediction, and when the median value is equal to mvp1, the motion vector difference (mvd1) of CP1 and DMVD (mvd0 - mvd1) can be signaled for inter-frame prediction. When the median value is equal to mvp2, either the case where the median value is equal to mvp0 or the case where the median value is equal to mvp1 can be followed, and this can be predetermined.
[0159] Here, the above process can be performed for each of the x and y components of the motion vector predictor. In other words, the median value can be derived for the x component and the y component separately. In this case, when the x component of the median value is the same as the x component of mvp0, according to the above case where the median value is equal to mvp0, only the x component of the motion vector can be encoded and decoded, and when the y component of the median value is the same as the y component of mvp1, according to the above case where the median value is the same as mvp1, only the y component of the motion vector can be encoded and decoded. If the x component and / or y component of the median value is equal to the x component and / or y component of mvp2, the x component and / or y component of the motion vector can be encoded and decoded according to one of the above cases where the median value is equal to mvp0 and the above case where the median value is equal to mvp1, which can be predefined.
[0160] Figure 10 Illustratively show three CPs for a 6-parameter affine motion model according to an embodiment of the present disclosure.
[0161] Embodiments of the present disclosure can use three CPs. The three CPs can be respectively positioned at the upper left sample position, the upper right sample position, and the lower left sample position of the current block. Here, the CP located at the upper left sample position can be denoted as CP0, the CP located at the upper right sample position can be denoted as CP1, and the CP that can be located at the lower left sample position can be denoted as CP2, and the motion vector at CP0 can be denoted as mv0, the motion vector at CP1 can be denoted as mv1, and the motion vector at CP2 can be denoted as mv2. The coordinates (CPi) of each control point can be defined as (x i , y i ), i=0, 1, 2 , and the motion vector at each control point can be denoted as mv i =( v x i,v yi ), i=0, 1, 2 。
[0162] In embodiments of the present disclosure, three CPs may be respectively distributed at the upper left sample position, the upper right sample position, and the lower left sample position, but these three CPs may be positioned differently from this. For example, CP0 may be located at the upper right sample position, CP1 may be located at the upper left sample position, and CP2 may be located at the lower left sample position, but their positions are not limited thereto. In this case, the following processing may be similarly performed in consideration of the positions of each CP.
[0163] For example, if the width of the current block is W and its height is H, assuming the coordinates of the lower left sample position of the current block are (0, 0), the coordinates of CP0 may be represented as (0, H), the coordinates of CP1 may be represented as (W, H), and the coordinates of CP2 may be represented as (0, 0). Here, W and H may have different values, but may also have the same value, and the reference (0, 0) may be set differently.
[0164] As shown in Figure 10 , since an affine motion model using three motion vectors according to three CPs uses six parameters according to the three motion vectors in the affine motion prediction method, it may be referred to as a 6-parameter affine motion model.
[0165] In embodiments of the present disclosure, the motion vector of a sample unit may be determined by an affine motion vector field (affine MVF) and the position of the sample. The affine motion vector field may represent the motion vector of the sample unit based on three motion vectors according to three CPs.
[0166] According to embodiments of the present disclosure, a reference block among temporal and / or spatial neighboring blocks of the current block may be determined, and the motion vector of the reference block may be used as a motion vector predictor of the current block, and the motion vector of the current block may be represented by the motion vector predictor and the motion vector difference. In addition, embodiments of the present disclosure may signal the indices of the motion vector predictor and the motion vector difference.
[0167] According to embodiments of the present disclosure, three motion vector differences according to three CPs may be derived based on three motion vectors according to three CPs and three motion vector predictors according to three CPs during encoding, and during decoding, three motion vectors according to three CPs may be derived based on three motion vector predictors according to three CPs and three motion vector differences according to three CPs. In other words, the motion vector at each CP may be composed of the sum of the motion vector predictor and the motion vector difference, as shown in Equation 4.
[0168] [Equation 4]
[0169]
[0170] In Equation 4, mvp0, mvp1, and mvp2 may represent motion vector predictors (MVPs) at each of CP0, CP1, and CP2, and mvd0, mvd1, and mvd2 may represent motion vector differences (MVDs) at each of CP0, CP1, and CP2. Herein, mvp may be referred to as CPMVP, and mvd may be referred to as CPMVD.
[0171] Therefore, the inter - frame prediction method according to the affine mode according to an embodiment of the present disclosure may encode and decode an index and motion vector differences (mvd0, mvd1, and mvd2) at each CP. In other words, according to an embodiment of the present disclosure, based on mvd0, mvd1, and mvd2 at each of CP0, CP1, and CP2 of a current block and mvp0, mvp1, and mvp2 according to an index, motion vectors at CP0, CP1, and CP2 may be derived, and inter - frame prediction may be performed based on the motion vectors at CP0, CP1, and CP2 to derive the motion vector of a sample unit.
[0172] An affine motion prediction method according to another embodiment of the present disclosure may use one of the three motion vector differences, a difference between the motion vector difference and another motion vector difference (DMVD, difference of two MVDs), and a difference between the motion vector difference and yet another motion vector difference (DMVD).
[0173] More specifically, in another embodiment of the present disclosure, when mvd0, mvd1, and mvd2 are derived according to three CPs, inter - frame prediction may be performed by signaling mvd0 and two DMVDs (mvd0 - mvd1 and mvd0 - mvd2), by signaling mvd1 and two DMVDs (mvd0 - mvd1 and mvd1 - mvd2), or by signaling mvd2 and two DMVDs (mvd0 - mvd2 and mvd1 - mvd2). Herein, for convenience, one of the two DMVDs may be referred to as the first DMVD (DMVD1), and the other DMVD may be referred to as the second DMVD (DMVD2). Additionally, when encoding and decoding mvd0 and two DMVDs (mvd0 - mvd1 and mvd0 - mvd2), mvd0 may be referred to as the MVD of CP0, DMVD (mvd0 - mvd1) may be referred to as the DMVD of CP1, and DMVD (mvd0 - mvd2) may be referred to as the DMVD of CP2.
[0174] That is, according to another embodiment of the present disclosure, three MVDs (e.g., MVD0, MVD1, and MVD2) can be derived based on one MVD and two DMVDs, and the motion vectors at each of the three CPs can be derived based on a motion vector predictor according to the indices for the three CPs (e.g., CP0, CP1, and CP2) and together with such an MVD, and the inter-frame prediction can be performed by deriving the motion vector of a sample unit based on the motion vectors at the three CPs.
[0175] In another embodiment of the present disclosure described with reference Figure 10 to, the method using the median predictor described with reference Figure 9 to can be adaptively applied to effectively perform motion vector coding, and in this case, the process of deriving MVP2 can be omitted from the method described with reference Figure 9 to.
[0176] Hereinafter, in the description of the present disclosure, CP0, CP1, and CP2 may be represented as the first CP, the second CP, and the third CP respectively, and the motion vector (MV), motion vector predictor (MVP), and motion vector difference (MVD) according to each CP may also be represented in a similar manner as above.
[0177] Figure 11 Schematically shows a video coding method of an encoding device according to the present disclosure.
[0178] Figure 11 The method disclosed in Figure 1 can be performed by the encoding device disclosed in Figure 11 . For example, S1100 to S1120 in
[0179] can be performed by the predictor of the encoding device; and S1130 can be performed by the entropy encoder of the encoding device.
[0180] The encoding device derives control points (CPs) of a current block (S1100). When affine motion prediction is applied to the current block, the encoding device can derive the CPs, and depending on the embodiment, the number of CPs can be two or three.
[0181] For example, when there are three CPs, the CPs may be located at the upper left sample position, the upper right sample position, and the lower left sample position of the current block respectively. And if the height and width of the current block are H and W respectively, and the coordinate components of the lower left sample position are (0, 0), then the coordinate components of the CPs may be (0, H), (W, H), and (0, 0) respectively.
[0182] The encoding device derives the MVP for the CP (S1110). For example, when the number of derived CPs is two, the encoding device can obtain two motion vectors. For example, when the number of derived CPs is three, the encoding device can obtain three motion vectors. The MVP for the CP can be derived based on neighboring blocks, and the detailed description has been referred to above Figure 7 and Figure 9 described in detail.
[0183] For example, when the first CP and the second CP are derived, the encoding device can derive the first MVP for the first CP and the second MVP for the second CP based on the neighboring blocks of the current block. And when the third CP is further derived, the encoding device can further derive the third MVP based on the neighboring blocks of the current block.
[0184] For example, when the first CP, the second CP, and the third CP are derived, the encoding device can derive the third MVP for the third CP based on the first MVP for the first CP and the second MVP for the second CP, and can also derive the third MVP based on the motion vectors of the neighboring blocks of the third CP.
[0185] The encoding device derives at least one of a motion vector difference (MVD) and a difference of two MVDs (DMVD) (S1120). The motion vector difference (MVD) can be derived based on the motion vector (MV) and the motion vector predictor (MVP). For this purpose, the encoding device can also derive the motion vectors of each CP. The difference of the motion vector differences (DMVD) can be derived based on multiple motion vector differences.
[0186] For example, when there are two CPs, the encoding device can derive one MVD and one DMVD. The encoding device can derive two MVDs from the motion vectors of the two CPs and the two MVPs. Additionally, one of the two MVDs to be encoded can be selected, and the difference of the two MVDs (DMVD) can be derived based on the selected one.
[0187] For example, when the first CP and the second CP are derived, the encoding device can derive the first MVD for the first CP and the DMVD for the second CP. Here, the DMVD for the second CP can represent the difference between the first MVD and the second MVD for the second CP, and the first MVD can be the reference.
[0188] For example, when there are three CPs, the encoding device may derive one MVD and two DMVDs. The encoding device may derive three MVDs from the motion vectors of the three CPs and the reference block from among three MVPs, and may select any one of the three MVDs to be encoded. Additionally, the encoding device may derive the difference between one selected MVD and another MVD (DMVD1) and the difference between one selected MVD and yet another MVD (DMVD2).
[0189] For example, when a third CP is further derived, the encoding device may derive a first MVD for the first CP, a DMVD for the second CP, and a DMVD for the third CP. Here, the DMVD for the second CP may represent the difference between the first MVD and the second MVD of the second CP, and the DMVD for the third CP may represent the difference between the first MVD and the third MVD of the third CP, and the first MVD may be used as a reference.
[0190] For example, when a third CP is further derived, the encoding device may derive a median based on the first MVP, the second MVP, and the third MVP, and in this case, the third MVD and the DMVD for the third CP may not be derived. The detailed description thereof has been referred to above Figure 9 and described in detail.
[0191] The encoding device encodes based on one MVD and at least one DMVD and outputs a bitstream (S1130). For inter prediction, the encoding device may generate and output a bitstream for the current block, which includes one MVD, at least one DMVD, and an index for the motion vector predictor.
[0192] For example, when there are two CPs, the encoding device may generate a bitstream for the current block, which includes the indexes of the motion vector predictors of the two CPs and the motion vector difference of any one of the two CPs, and the difference between the motion vector differences of the two CPs.
[0193] For example, when the first CP and the second CP are derived, the encoding device may output a bitstream by encoding the image information including the information about the first CP and the information about the DMVD for the second CP.
[0194] For example, if there are three CPs, the encoding device may generate a bitstream for the current block, which includes the indexes of the motion vector predictors for the three CPs and the motion vector difference for one of the three CPs, the difference between the motion vector difference for this CP and the motion vector difference for another CP, and the difference between the motion vector difference for this CP and the motion vector difference for yet another CP.
[0195] For example, when further deriving a third CP, the encoding device may further include information on DMVD for the third CP in the image information, and may encode the image information to output a bitstream.
[0196] For example, when further deriving a third CP and deriving a median value, the encoding device may output a bitstream by encoding image information including information on a first MVD and information on DMVD for a second CP, and may not further include information on DMVD for the third CP in the image information.
[0197] The bitstream generated and output by the encoding device may be sent to the decoding device via a network or a storage medium.
[0198] Figure 12 Schematically illustrate the inter-frame prediction method of the decoding device according to the present disclosure.
[0199] Figure 12 The method disclosed in may be performed by Figure 2 the decoding device disclosed in. For example, Figure 12 S1200, S1210, S1230, and S1240 in may be performed by the predictor of the decoding device, and S1220 may be performed by the entropy decoder of the decoding device. Here, S1220 may be performed before S1200 and S1210.
[0200] The decoding device derives control points (CPs) of the current block (S1200). When affine motion prediction is applied to the current block, the decoding device may derive CPs, and depending on the embodiment, the number of CPs may be two or three.
[0201] For example, when there are two CPs, the CPs may be located at the upper left sample position and the upper right sample position of the current block respectively, and if the height and width of the current block are H and W respectively, and the coordinate components of the lower left sample position are (0, 0), the coordinate components of the CPs may be (0, H) and (W, H) respectively.
[0202] For example, when there are three CPs, the CPs may be located at the upper left sample position, the upper right sample position, and the lower left sample position of the current block respectively, and if the height and width of the current block are H and W respectively, and the coordinate components of the lower left sample position are (0, 0), the coordinate components of the CPs may be (0, H), (W, H), and (0, 0).
[0203] The decoding device derives MVPs for the CPs (S1210). For example, when the number of derived CPs is two, the decoding device may obtain two motion vectors. For example, when the number of derived CPs is three, the decoding device may obtain three motion vectors. The MVPs for the CPs may be derived based on neighboring blocks, and the detailed description thereof has been referred to above Figure 7 and Figure 9 for a detailed description.
[0204] For example, when deriving a first CP and a second CP, the decoding device may derive a first MVP for the first CP and a second MVP for the second CP based on neighboring blocks of the current block, and when a third CP is further derived, the decoding device may further derive a third MVP based on neighboring blocks of the current block.
[0205] For example, when deriving a first CP, a second CP, and a third CP, the decoding device may derive a third MVP for the third CP based on the first MVP for the first CP and the second MVP for the second CP, and derive the third MVP based on the motion vectors of neighboring blocks of the third CP.
[0206] The decoding device decodes one MVD and at least one DMVD (S1220). The decoding device may obtain one MVD and at least one DMVD by decoding one MVD and at least one DMVD based on the received bitstream. Here, the bitstream may include indices of motion vector predictors for the CPs. The bitstream may be received from the encoding device via a network or a storage medium.
[0207] For example, when there are two CPs, the decoding device may decode one MVD and one DMVD, and when there are three CPs, the decoding device may decode one MVD and two DMVDs. Here, the DMVD may mean the difference between two MVDs.
[0208] For example, when deriving a first CP and a second CP, the decoding device may decode a first MVD for the first CP and decode the DMVD for the second CP. Here, the DMVD for the second CP may represent the difference between the first MVD and the second MVD for the second CP.
[0209] For example, when a third CP is further derived, the decoding device may decode a first MVD for the first CP and decode the DMVD for the second CP and the DMVD for the third CP. Here, the DMVD for the second CP may represent the difference between the first MVD and the second MVD for the second CP, and the DMVD for the third CP may represent the difference between the first MVD and the third MVD for the third CP.
[0210] For example, if a third CP is further derived and the median value is used, the decoding device may decode the first MVD for the first CP and the DMVD for the second CP. Here, the DMVD for the third CP may not be decoded.
[0211] The decoding device derives a motion vector for a CP based on the MVP for the CP, one MVD, and at least one DMVD (S1230). The motion vector for a CP may be derived based on a motion vector difference (MVD) and a motion vector predictor (MVP), and the motion vector difference may be derived based on the difference between motion vector differences (DMVD).
[0212] For example, when there are two CPs, the decoding device may receive one MVD and one DMVD, and based on them, two MVDs may be derived according to the two CPs. The decoding device may receive the indexes for the two CPs together, and based on them, two MVPs may be derived. The decoding device may derive the motion vectors for the two CPs based on the two MVDs and the two MVPs respectively.
[0213] For example, when the first CP and the second CP are derived, the decoding device may derive the first MV based on the first MVD and the first MVP, derive the second MVD for the second CP based on the first MVD and the DMVD for the second CP, and derive the second MV based on the second MVD and the second MVP.
[0214] For example, when there are three CPs, the decoding device may receive one MVD and two DMVDs, and based on them, three MVDs may be derived according to the three CPs. The decoding device may receive the indexes for the three CPs together, and based on them, three MVPs may be derived. The decoding device may derive the motion vectors for the three CPs based on the three MVDs and the three MVPs respectively.
[0215] For example, if a third CP is further derived and the third MVP is derived, the decoding device may derive the third MVD for the third CP based on the first MVD and the DMVD for the third CP, and derive the third MV based on the third MVD and the third MVP.
[0216] For example, if a third CP is further derived and the median value is the same as the first MVP, the decoding device may derive the first MV based on the first MVD and the first MVP, derive the second MVD for the second CP based on the first MVD and the DMVD for the second CP, and derive the second MV based on the second MVD and the second MVP.
[0217] For example, if a third CP is further derived and the median value is the same as the second MVP, the decoding device may derive a second MV based on the second MVD and the second MVP, derive a first MVD for the first CP based on the second MVD and the DMVD for the first CP, and derive a first MV based on the first MVD and the first MVP.
[0218] For example, if a third CP is further derived and the median value is the same as the third MVP, the decoding device may derive the first MV and the second MV according to either the case where the median value is the same as the first MVP or the case where the median value is the same as the second MVP. Determining either the case where the median value is the same as the first MVP or the case where the median value is the same as the second MVP may be predefined. The detailed description thereof has been referred to above Figure 9 and is described in detail therein.
[0219] The decoding device generates a prediction block for the current block based on the motion vector (S1240). The decoding device may derive an affine motion vector field (affine MVF) based on the motion vectors of the corresponding CPs, and based on them, may derive the motion vectors of the sample units to perform inter prediction.
[0220] In the above embodiments, the method is explained based on a flowchart by means of a series of steps or blocks, but the present disclosure is not limited to the order of the steps, and a certain step may be performed in a different order or steps from the above steps, or simultaneously with another step. In addition, those of ordinary skill in the art can understand that the steps shown in the flowchart are not exclusive, and another step may be incorporated or one or more steps in the flowchart may be deleted without affecting the scope of the present disclosure.
[0221] The above method according to the present disclosure may be implemented in software form, and the encoding device and / or decoding device according to the present disclosure may be included in a device for image processing such as a television, a computer, a smart phone, a set-top box, a display device, etc.
[0222] When the embodiments in the present disclosure are embodied by software, the above method may be embodied as modules (processes, functions, etc.) for performing the above functions. These modules may be stored in a memory and may be executed by a processor. The memory may be inside or outside the processor and may be connected to the processor in various well-known ways. The processor may include an application specific integrated circuit (ASIC), other chip sets, logic circuits, and / or data processing devices. The memory may include a read only memory (ROM), a random access memory (RAM), a flash memory, a memory card, a storage medium, and / or other storage devices.
Claims
1. A decoding device for image decoding, the decoding device comprising: a memory; and at least one processor, the at least one processor being connected to the memory and configured to: derive a first motion vector predictor for a first control point of the current block and a second motion vector predictor for a second control point of the current block based on neighboring blocks of the current block, wherein the first control point is located at the upper left sample position of the current block, and the second control point is located at the upper right sample position of the current block; decode information about a first motion vector difference for the first control point to derive the first motion vector difference; decode information about a difference between two motion vector differences for the second control point to derive the difference between the two motion vector differences for the second control point; derive a first motion vector for the first control point based on the first motion vector predictor and the first motion vector difference; derive a second motion vector for the second control point based on the second motion vector predictor and the difference between the two motion vector differences for the second control point; and generate prediction samples for the current block based on the first motion vector and the second motion vector, wherein the difference between the two motion vector differences for the second control point is the difference between the second motion vector difference for the second control point and the first motion vector difference, and wherein the at least one processor is further configured to: derive the second motion vector difference for the second control point based on the first motion vector difference and the difference between the two motion vector differences for the second control point; and derive the second motion vector difference for the second control point based on the second motion vector predictor and the second motion vector difference derived based on the difference between the two motion vector differences for the second control point.
2. An encoding device for image encoding, the encoding device comprising: a memory; and at least one processor, the at least one processor being connected to the memory and configured to: derive a first motion vector predictor for a first control point of the current block and a second motion vector predictor for a second control point of the current block based on neighboring blocks of the current block, wherein the first control point is located at the upper left sample position of the current block, and the second control point is located at the upper right sample position of the current block; derive a first motion vector difference for the first control point; derive a difference between two motion vector differences for the second control point; and encode image information including information about the first motion vector difference and information about the difference between the two motion vector differences for the second control point to output a bitstream, wherein the difference between the two motion vector differences for the second control point is the difference between the second motion vector difference for the second control point and the first motion vector difference, and wherein the at least one processor is further configured to: Derive a second motion vector difference for the second control point based on the second motion vector for the second control point and the second motion vector predictor for the second control point; and Derive the difference between the two motion vector differences for the second control point based on the second motion vector difference and the first motion vector difference.
3. An apparatus for transmitting data for an image, the apparatus comprising: At least one processor configured to obtain a bitstream for the image, wherein the bitstream is generated based on: deriving a first motion vector predictor for a first control point of the current block and a second motion vector predictor for a second control point of the current block based on neighboring blocks of the current block, wherein the first control point is located at the upper left position of the current block and the second control point is located at the upper right position of the current block; deriving a first motion vector difference for the first control point; deriving the difference between two motion vector differences for the second control point; and encoding image information including information about the first motion vector difference and information about the difference between the two motion vector differences for the second control point; and A transmitter configured to transmit the data including the bitstream, wherein the difference between the two motion vector differences for the second control point is the difference between the second motion vector difference for the second control point and the first motion vector difference, and wherein deriving the difference between the two motion vector differences for the second control point includes: Deriving the second motion vector difference for the second control point based on the second motion vector for the second control point and the second motion vector predictor for the second control point; and Deriving the difference between the two motion vector differences for the second control point based on the second motion vector difference and the first motion vector difference.