Method and apparatus for processing a video signal using affine prediction
By adjusting the resolution of the affine motion vector and optimizing the entropy encoding method in affine prediction, the problems of low affine prediction accuracy and insufficient encoding efficiency in the prior art are solved, and more efficient video signal encoding is achieved.
Patent Information
- Application Number
- CN201980045121.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2018-07-02
- Filing Date
- 2019-07-02
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2039-07-02
AI Technical Summary
The prior art is difficult to effectively improve the accuracy of affine prediction, especially when encoding a multidimensional transformed video signal, there is a problem of low compression efficiency.
By controlling the resolution of the affine motion vector used in affine prediction, combining the derivation of syntax elements and binary methods, the entropy encoding method of the motion model is optimized, thereby improving the accuracy and encoding efficiency of affine prediction.
It realizes the accuracy and compression efficiency of affine motion prediction, and enhances the encoding capability of high spatial resolution and high frame rate video content.
Smart Images

Figure CN112385230B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to a method and apparatus for processing a video signal using affine prediction, and more particularly, to a method and apparatus for processing a video signal by controlling a resolution of an affine motion vector used in affine prediction. Background Art
[0002] Compression coding refers to a series of signal processing techniques for transmitting digital information over a communication line, or techniques for storing information in a form suitable for a storage medium. Media including pictures, images, audio, etc. can be targets for compression coding, and in particular, techniques for performing compression coding on pictures are called video image compression.
[0003] It is generally considered that next-generation video content has characteristics of high spatial resolution, high frame rate, and high dimensionality of scene representation. To process such content, it will lead to a sharp increase in memory storage capacity, memory access rate, and processing power.
[0004] Therefore, there is a need to design an encoding tool for efficiently processing next-generation video content. Summary of the Invention
[0005] Technical Problem
[0006] An object of the present disclosure is to propose a method for controlling a resolution of an affine motion vector used in affine prediction to improve the accuracy of affine prediction.
[0007] In addition, an object of the present disclosure is to propose an entropy coding method that depends on the unique statistics of a motion model rather than a constant entropy coding method when performing entropy coding on an MVD.
[0008] The technical objects to be achieved in the present disclosure are not limited to the above technical objects, and other technical objects not described above can be clearly understood by those of ordinary skill in the art to which the present disclosure pertains according to the following description.
[0009] Technical Solution
[0010] In one aspect of the present disclosure, a method of processing a video signal using affine prediction may include the following steps: checking whether affine prediction is applied to a current block; if affine prediction is applied as a result of the check, obtaining at least one syntax element indicating a resolution of a motion vector difference used in the affine prediction; deriving a control point motion vector of the current block based on the at least one syntax element; deriving a motion vector of each of a plurality of sub-blocks included in the current block based on the control point motion vector; and generating a prediction sample of the current block using the motion vector of each sub-block.
[0011] Preferably, the step of obtaining at least one syntax element may include the following steps: obtaining a first syntax element indicating whether the resolution of the motion vector difference is a preset default resolution; and when the resolution of the motion vector difference is not the default resolution, obtaining a second syntax element indicating the resolution of the motion vector difference among the remaining resolutions other than the default resolution.
[0012] Preferably, the default resolution may be preset to 1 / 4 pixel precision.
[0013] Preferably, each of the remaining resolutions may include at least one of integer pixel precision, 4 pixel precision, 1 / 8 pixel precision, or 1 / 16 pixel precision.
[0014] Preferably, the step of deriving the control point motion vector may include the following steps: using at least one syntax element to determine the resolution of the motion vector difference; and obtaining the motion vector difference based on the resolution of the motion vector difference.
[0015] Preferably, the step of obtaining the motion vector difference may include the following steps: obtaining a flag indicating whether the motion vector difference is greater than 0; and when the motion vector difference is greater than 0, obtaining a flag indicating whether the motion vector difference is greater than a predefined specific value.
[0016] Preferably, when the motion vector difference is greater than 0 and less than or equal to the predefined specific value, the motion vector difference may be binary-coded using an exponential Golomb code of order 1. When the motion vector difference is greater than the predefined specific value, the motion vector difference may be binary-coded using a truncated binary coding method.
[0017] In another aspect of the present disclosure, a device for processing a video signal using affine prediction may include: an affine prediction mode recognition unit configured to check whether affine prediction is applied to a current block; a syntax element acquisition unit configured to obtain at least one syntax element indicating a resolution of a motion vector difference used in the affine prediction if, as a result of the check, affine prediction is applied; a control point motion vector derivation unit configured to derive a control point motion vector of the current block based on the at least one syntax element; a sub-block motion vector derivation unit configured to derive a motion vector of each of a plurality of sub-blocks included in the current block based on the control point motion vector; and a prediction sample generation unit configured to generate a prediction sample of the current block using the motion vector of each sub-block.
[0018] Preferably, the syntax element acquisition unit may be configured to obtain a first syntax element indicating whether the resolution of the motion vector difference is a preset default resolution, and when the resolution of the motion vector difference is not the default resolution, obtain a second syntax element indicating the resolution of the motion vector difference among the remaining resolutions other than the default resolution.
[0019] Preferably, the default resolution may be preset to 1 / 4 pixel precision.
[0020] Preferably, each of the remaining resolutions may include at least one of integer pixel precision, 4 pixel precision, 1 / 8 pixel precision, or 1 / 16 pixel precision.
[0021] Preferably, the control point motion vector derivation unit may be configured to use the at least one syntax element to determine the resolution of the motion vector difference and obtain the motion vector difference based on the resolution of the motion vector difference.
[0022] Preferably, the control point motion vector derivation unit may be configured to obtain a flag indicating whether the motion vector difference is greater than 0, and when the motion vector difference is greater than 0, obtain a flag indicating whether the motion vector difference is greater than a predefined specific value.
[0023] Preferably, when the motion vector difference is greater than 0 and less than or equal to the predefined specific value, the motion vector difference may be binary-coded using an exponential Golomb code of order 1. When the motion vector difference is greater than the predefined specific value, the motion vector difference may be binary-coded using a truncated binary coding method.
[0024] Beneficial effects
[0025] According to an embodiment of the present disclosure, the accuracy of affine motion prediction can be improved and the compression efficiency can be improved by controlling the motion vector accuracy of control points used in affine prediction.
[0026] In addition, according to an embodiment of the present disclosure, the coding efficiency and compression performance can be improved by adaptively setting the binarization method for each partitioned MVD region.
[0027] The effects obtainable in the present disclosure are not limited to the above effects, and those of ordinary skill in the art to which the present disclosure pertains can clearly understand other technical effects not described above based on the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] To assist in understanding the present disclosure, the drawings included as part of the detailed description provide embodiments of the present disclosure and describe the technical features of the present disclosure together with the detailed description.
[0029] Figure 1 is a schematic block diagram of an encoding device that performs encoding of video / image signals as an embodiment to which the present disclosure is applied.
[0030] Figure 2 is a schematic block diagram of a decoding device that performs decoding of video / image signals as an embodiment to which the present disclosure is applied.
[0031] Figure 3 is a diagram showing an example of a multi-type tree structure to which the present disclosure can be applied.
[0032] Figure 4 is a diagram showing a signaling mechanism of partition information of a quadtree having a nested multi-type tree structure as an embodiment to which the present disclosure can be applied.
[0033] Figure 5 is a diagram showing a method of dividing a CTU into a plurality of CUs based on a quadtree and a nested multi-type tree structure as an embodiment to which the present disclosure can be applied.
[0034] Figure 6 is a diagram showing a method for restricting quadtree splitting as an embodiment to which the present disclosure can be applied.
[0035] Figure 7 is a diagram showing redundant partition patterns that may occur in binary tree partitioning and quadtree partitioning as an embodiment to which the present disclosure can be applied.
[0036] Figure 8 and Figure 9 are diagrams showing a video / image coding method based on inter-frame prediction according to an embodiment of the present disclosure and an inter-frame prediction unit in an encoding device according to an embodiment of the present disclosure.
[0037] Figure 10 and Figure 11FIG. is a diagram showing an inter - frame prediction - based video / image decoding method according to an embodiment of the present disclosure and an inter - frame prediction unit in a decoding device according to an embodiment of the present disclosure.
[0038] Figure 12 FIG. is a diagram for describing adjacent blocks used in a merge mode or a skip mode to which an embodiment of the present disclosure is applied.
[0039] Figure 13 FIG. is a flowchart showing a method for configuring a merge candidate list according to an embodiment of the present disclosure.
[0040] Figure 14 FIG. is a flowchart showing a method for configuring a merge candidate list according to an embodiment of the present disclosure.
[0041] Figure 15 FIG. shows an example of a motion model according to an embodiment of the present disclosure.
[0042] Figure 16 FIG. shows an example of a control - point motion vector for affine motion prediction according to an embodiment of the present disclosure.
[0043] Figure 17 FIG. shows an example of motion vectors for each sub - block of a block for which affine motion prediction according to an embodiment of the present disclosure has been applied.
[0044] Figure 18 FIG. shows an example of adjacent blocks for affine motion prediction in an affine merge mode according to an embodiment of the present disclosure.
[0045] Figure 19 FIG. shows an example of a block for which affine motion prediction is performed using adjacent blocks for which affine motion prediction according to an embodiment of the present disclosure has been applied.
[0046] Figure 20 FIG. is a diagram for describing a method for generating a merge candidate list using peripheral affine - coded blocks according to an embodiment of the present disclosure.
[0047] Figure 21 and Figure 22 FIG. is a diagram for describing a method for configuring an affine merge candidate list using adjacent blocks predicted by affine coding according to an embodiment of the present disclosure.
[0048] Figure 23 FIG. shows an example of adjacent blocks for affine motion prediction in an affine inter - frame mode according to an embodiment of the present disclosure.
[0049] Figure 24Shows an example of adjacent blocks for affine motion prediction in the affine inter - frame mode according to an embodiment of the present disclosure.
[0050] Figure 25 And Figure 26 Is a diagram showing a method of deriving motion vector candidates using the motion information of adjacent blocks in the affine inter - frame mode according to an embodiment of the present disclosure.
[0051] Figure 27 Shows an example of a method of deriving an affine motion vector field in units of sub - blocks according to an embodiment of the present disclosure.
[0052] Figure 28 Shows a method of generating a prediction block and a motion vector in inter - frame prediction in which an affine motion model according to an embodiment of the present disclosure has been applied.
[0053] Figure 29 Is a diagram showing a method of performing motion compensation based on motion vectors of control points according to an embodiment of the present disclosure.
[0054] Figure 30 Is a diagram showing a method of performing motion compensation based on motion vectors of control points in an irregular block according to an embodiment of the present disclosure.
[0055] Figure 31 Is a diagram showing a method of performing motion compensation based on motion vectors of control points in an irregular block according to an embodiment of the present disclosure.
[0056] Figures 32 to 38 Is a diagram showing a method of performing motion compensation based on motion vectors of control points in an irregular block according to an embodiment of the present disclosure.
[0057] Figure 39 Shows an overall coding structure for deriving motion vectors according to an embodiment of the present disclosure.
[0058] Figure 40 Shows an example of an MVD coding structure according to an embodiment of the present disclosure.
[0059] Figure 41 Shows an example of an MVD coding structure according to an embodiment of the present disclosure.
[0060] Figure 42 Shows an example of an MVD coding structure according to an embodiment of the present disclosure.
[0061] Figure 43 Shows an example of an MVD coding structure according to an embodiment of the present disclosure.
[0062] Figure 44 It is a diagram showing a method for deriving affine motion vector difference information according to an embodiment to which the present disclosure is applied.
[0063] Figure 45 It is a diagram showing an encoding structure of a motion vector difference according to an embodiment to which the present disclosure is applied.
[0064] Figure 46 It is a diagram showing a method for deriving an affine motion vector based on precision information according to an embodiment of the present disclosure.
[0065] Figure 47 It is a diagram showing an encoding structure of a motion vector difference according to an embodiment to which the present disclosure is applied.
[0066] Figure 48 It is a flowchart showing a method for generating an inter prediction block based on affine prediction according to an embodiment to which the present disclosure is applied.
[0067] Figure 49 It is a diagram showing an inter prediction device based on affine prediction according to an embodiment to which the present disclosure is applied.
[0068] Figure 50 It shows a video coding system to which the present disclosure is applied.
[0069] Figure 51 It is an embodiment to which the present disclosure is applied and shows the structure of a content streaming system. Detailed Embodiments
[0070] Hereinafter, preferred embodiments of the present disclosure will be described in detail with reference to the accompanying drawings. It will be Figure 1 The detailed description to be disclosed together is intended to describe some embodiments of the present disclosure and is not intended to describe the only embodiments of the present disclosure. The following detailed description includes more details in order to provide a comprehensive understanding of the present disclosure. However, those skilled in the art will understand that the present disclosure can be implemented without such more details.
[0071] In some cases, in order to avoid obscuring the concepts of the present disclosure, known structures and devices are omitted or may be shown in block diagram form based on the core functions of each structure and device.
[0072] Although most of the terms used in the present disclosure are selected from general terms widely used in the art, the applicant has arbitrarily selected some terms and explained their meanings in detail in the following description as needed. Therefore, the present disclosure should be understood in terms of the intended meanings of the terms rather than their simple names or meanings.
[0073] Specific terms used in the following description have been provided to assist in understanding the present disclosure, and the use of these specific terms can be changed in various forms without departing from the technical spirit of the present disclosure. For example, signals, data, samples, pictures, frames, blocks, etc. can be appropriately replaced and interpreted in each encoding process.
[0074] In this specification, a "processing unit" refers to a unit in which encoding / decoding processing such as prediction, transformation, and / or quantization is performed. Hereinafter, for ease of description, the processing unit may be referred to as a "processing block" or a "block".
[0075] In addition, the processing unit can be interpreted to include a unit for the luminance component and a unit for the chrominance component. For example, the processing unit may correspond to a coding tree unit (CTU), a coding unit (CU), a prediction unit (PU), or a transformation unit (TU).
[0076] Furthermore, the processing unit can be interpreted as a unit for the luminance component or a unit for the chrominance component. For example, the processing unit may correspond to a coding tree block (CTB), a coding block (CB), a prediction unit PU, or a transformation block (TB) for the luminance component. In addition, the processing unit may correspond to a CTB, a CB, a PU, or a TB for the chrominance component. Furthermore, the processing unit is not limited thereto and can be interpreted to include a unit for the luminance component and a unit for the chrominance component.
[0077] In addition, the processing unit does not have to be limited to a square block and can be configured in a polygonal shape having three or more vertices.
[0078] In addition, in this specification, a pixel is referred to as a sample. Additionally, using a sample can mean using a pixel value or the like.
[0079] Figure 1 is a schematic block diagram of an encoding device that encodes a video / image signal according to an embodiment to which the present disclosure is applied.
[0080] Refer to Figure 1, the encoding device 100 may be configured to include an image splitter 110, a subtractor 115, a transformer 120, a quantizer 130, a dequantizer 140, an inverse transformer 150, an adder 155, a filter 160, a memory 170, an inter-frame predictor 180, an intra-frame predictor 185, and an entropy encoder 190. The inter-frame predictor 180 and the intra-frame predictor 185 may be collectively referred to as a predictor. In other words, the predictor may include the inter-frame predictor 180 and the intra-frame predictor 185. The transformer 120, the quantizer 130, the dequantizer 140, and the inverse transformer 150 may be included in a residual processor. The residual processor may further include the subtractor 115. In one embodiment, the image splitter 110, the subtractor 115, the transformer 120, the quantizer 130, the dequantizer 140, the inverse transformer 150, the adder 155, the filter 160, the inter-frame predictor 180, the intra-frame predictor 185, and the entropy encoder 190 may be configured as one hardware component (e.g., an encoder or a processor). Additionally, the memory 170 may include a decoded picture buffer (DPB) and may be implemented by a digital storage medium.
[0081] The image splitter 110 may divide an input image (or picture or frame) input to the encoding device 100 into one or more processing units. For example, the processing unit may be referred to as a coding unit (CU). In this case, the coding unit may be recursively divided from a coding tree unit (CTU) or a largest coding unit (LCU) based on a quadtree binary tree (QTBT) structure. For example, based on a quadtree structure and / or a binary tree structure, one coding unit may be divided into multiple coding units with a deeper depth. In this case, for example, the quadtree structure may be applied first, and then the binary tree structure may be applied. Alternatively, the binary tree structure may be applied first. The encoding process according to the present disclosure may be performed based on the final coding unit that is no longer divided. In this case, the largest coding unit may be directly used as the final coding unit based on the coding efficiency according to the image characteristics, or if necessary, the coding unit may be recursively divided into deeper coding units. Thus, the coding unit with an optimal size may be used as the final coding unit. In this case, the encoding process may include processes such as prediction, transformation, or reconstruction as will be described later. For another example, the processing unit may further include a prediction unit (PU) or a transformation unit (TU). In this case, each of the prediction unit and the transformation unit may be divided or partitioned from each final coding unit. The prediction unit may be a unit for sample prediction, and the transformation unit may be a unit from which transformation coefficients are derived and / or a unit in which a residual signal is derived from the transformation coefficients.
[0082] Depending on the context, a unit may be used interchangeably with a block or a region. In general, an M×N block may indicate a set of samples or a set of transform coefficients configured with M columns and N rows. Usually, a sample may indicate a pixel or the value of a pixel, and may indicate only the pixel / pixel value of the luminance component or only the pixel / pixel value of the chrominance component. In samples, a picture (or image) may be used as a term corresponding to a pixel or a pel.
[0083] The encoding device 100 may generate a residual signal (residual block or residual sample array) by subtracting the prediction signal (prediction block or prediction sample array) output from the inter-frame predictor 180 or the intra-frame predictor 185 from the input image signal (original block or original sample array). The generated residual signal is sent to the transformer 120. In this case, as shown in the figure, the unit in the encoding device 100 where the prediction signal (prediction block or prediction sample array) is subtracted from the input image signal (original block or original sample array) may be referred to as the subtractor 115. The predictor may perform prediction on the block to be processed (hereinafter referred to as the current block) and may generate a prediction block including prediction samples for the current block. The predictor may determine whether to apply intra-frame prediction or inter-frame prediction in the current block or CU unit. The predictor may generate various pieces of information about the prediction, such as the prediction mode information to be described later in the description of each prediction mode, and may send this information to the entropy encoder 190. The information about the prediction may be encoded in the entropy encoder 190 and may be output in the form of a bitstream.
[0084] The intra-frame predictor 185 may predict the current block by referring to the samples within the current picture. The position of the samples referred to may be adjacent to the current block or may be spaced apart from the current block according to the prediction mode. In intra-frame prediction, the prediction mode may include a plurality of non-angle modes and a plurality of angle modes. The non-angle modes may include, for example, the DC mode and the planar mode. For example, according to the fineness of the prediction direction, the angle mode may include 33 angle prediction modes or 65 angle prediction modes. In this case, for example, depending on the configuration, an angle prediction mode with more or fewer than 33 angle prediction modes or 65 angle prediction modes may be used. The intra-frame predictor 185 may use the prediction mode applied to the adjacent block to determine the prediction mode applied to the current block.
[0085] The inter - frame predictor 180 can derive a prediction block of a current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. In this case, in order to reduce the amount of motion information transmitted in the inter - frame prediction mode, based on the correlation of motion information between adjacent blocks and the current block, the motion information can be predicted as a block, a sub - block, or a sample unit. The motion information can include a motion vector and a reference picture index. The motion information can also include inter - frame prediction direction (L0 prediction, L1 prediction, Bi prediction) information. In the case of inter - frame prediction, adjacent blocks can include spatially adjacent blocks within the current picture and temporally adjacent blocks within the reference picture. The reference picture including the reference block and the reference picture including the temporally adjacent block can be the same or different. The temporally adjacent block can be referred to by the following names: collocated reference block or collocated CU (colCU). The reference picture including the temporally adjacent block can be referred to as a collocated picture (colPic). For example, the inter - frame predictor 180 can construct a motion information candidate list based on adjacent blocks, and can generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. The inter - frame prediction can be performed based on various prediction modes. For example, in the case of the skip mode and the merge mode, the inter - frame predictor 180 can use the motion information of adjacent blocks as the motion information of the current block. In the case of the skip mode, different from the merge mode, the residual signal may not be transmitted. In the case of the motion vector prediction (MVP) mode, the motion vector of an adjacent block can be used as a motion vector predictor. The motion vector of the current block can be indicated by signaling a motion vector difference.
[0086] The prediction signal generated by the inter - frame predictor 180 or the intra - frame predictor 185 can be used to generate a reconstructed signal or a residual signal.
[0087] The transformer 120 can generate transform coefficients by applying a transform scheme to the residual signal. For example, the transform scheme can include at least one of a discrete cosine transform (DCT), a discrete sine transform (DST), a Karhunen - Loève transform (KLT), a graph - based transform (GBT), or a conditional non - linear transform (CNT). In this case, GBT means a transform obtained from a graph if the relationship information between pixels is represented as a graph. CNT refers to a transform obtained based on a prediction signal generated using all previously reconstructed pixels. In addition, the transform process can be applied to a square pixel block of the same size, or can be applied to a block with a variable size rather than a square.
[0088] Quantizer 130 may quantize the transform coefficients and send them to entropy encoder 190. Entropy encoder 190 may encode the quantized signal (information about the quantized transform coefficients) and output it in the form of a bitstream. The information about the quantized transform coefficients may be referred to as residual information. Quantizer 130 may rearrange the quantized transform coefficients in block form into a one-dimensional vector based on the coefficient scan sequence, and may generate information about the quantized transform coefficients based on the quantized transform coefficients in one-dimensional vector form. Entropy encoder 190 may perform various encoding methods such as exponential Golomb, context-adaptive variable length coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC). In addition to the quantized transform coefficients, entropy encoder 190 may also encode, together or separately, information necessary for video / image reconstruction (e.g., values of syntax elements). The encoded information (e.g., encoded video / image information) may be sent or stored in units of network abstraction layer (NAL) units in the form of a bitstream. The bitstream may be sent through a network or may be stored in a digital storage medium. In this case, the network may include a broadcast network and / or a communication network. The digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD. A transmitter (not shown) for sending the signal output by entropy encoder 190 and / or a memory (not shown) for storing the signal may be configured as internal / external elements of encoding device 100, or the transmitter may be an element of entropy encoder 190.
[0089] The quantized transform coefficients output by quantizer 130 may be used to generate a prediction signal. For example, a residual signal may be reconstructed by applying dequantization and inverse transformation to the quantized transform coefficients via dequantizer 140 and inverse transformer 150 within a loop. Adder 155 may add the reconstructed residual signal to the prediction signal output by inter-frame predictor 180 or intra-frame predictor 185, thereby generating a reconstructed signal (reconstructed image, reconstructed block, or reconstructed sample array). If there is no residual for the target block to be processed as in the case where the skip mode has been applied, the predicted block may be used as the reconstructed block. Adder 155 may be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal may be used for intra-frame prediction of the next target block within the current picture, and may be used for inter-frame prediction of the next picture through filtering, which will be described later.
[0090] Filter 160 can improve subjective / objective picture quality by applying filtering to the reconstructed signal. For example, filter 160 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture. The modified reconstructed picture can be stored in memory 170, more specifically, in the DPB of memory 170. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filtering, and bilateral filtering. Filter 160 can generate various information for filtering (as will be described later in the description of each filtering method), and can send them to entropy encoder 190. The filtering information can be encoded by entropy encoder 190 and output in the form of a bitstream.
[0091] The modified reconstructed picture sent to memory 170 can be used as a reference picture in inter-frame predictor 180. If inter-frame prediction is applied, the encoding device can avoid prediction mismatches in encoding device 100 and the decoding device, and can improve encoding efficiency.
[0092] The DPB of memory 170 can store the modified reconstructed picture to use it as a reference picture in inter-frame predictor 180. Memory 170 can store the motion information of the blocks in which the motion information in the current picture is derived (or encoded) and / or the motion information of the blocks in the already reconstructed picture. The stored motion information can be forwarded to inter-frame predictor 180 to be used as the motion information of spatially adjacent blocks or temporally adjacent blocks. Memory 170 can store the reconstructed samples of the reconstructed blocks in the current picture and forward them to intra-frame predictor 185.
[0093] Figure 2 is an embodiment of applying the present disclosure and is a schematic block diagram of a decoding device for decoding video / image signals.
[0094] Referring to Figure 2 , decoding device 200 can be configured to include entropy decoder 210, dequantizer 220, inverse transformer 230, adder 235, filter 240, memory 250, inter-frame predictor 260, and intra-frame predictor 265. Inter-frame predictor 260 and intra-frame predictor 265 can be collectively referred to as predictors. That is, the predictor can include inter-frame predictor 180 and intra-frame predictor 185. Dequantizer 220 and inverse transformer 230 can be collectively referred to as residual processors. That is, the residual processor can include dequantizer 220 and inverse transformer 230. Entropy decoder 210, dequantizer 220, inverse transformer 230, adder 235, filter 240, inter-frame predictor 260, and intra-frame predictor 265 can be configured as one hardware component (e.g., a decoder or a processor) according to one embodiment. In addition, memory 250 can include a decoded picture buffer (DPB) and can be implemented by a digital storage medium.
[0095] When the input includes a bitstream of video / image information, the decoding device 200 may reconstruct an image according to the processing of the video / image information in the Figure 1 encoding device. For example, the decoding device 200 may use the processing unit applied in the encoding device to perform decoding. Thus, for example, the processing unit for decoding may be an encoding unit. The encoding unit may be split from the coding tree unit or the largest coding unit according to a quadtree structure and / or a binary tree structure. In addition, the reconstructed image signal decoded and output by the decoding device 200 may be played back by a playback device.
[0096] The decoding device 200 may receive the signal output by the Figure 1 encoding device in the form of a bitstream. The received signal may be decoded by the entropy decoder 210. For example, the entropy decoder 210 may derive information (e.g., video / image information) for image reconstruction (or picture reconstruction) by parsing the bitstream. For example, the entropy decoder 210 may decode the information in the bitstream based on an encoding method such as exponential Golomb, CAVLC, or CABAC, and may output the value of the syntax element for image reconstruction or the quantization value of the transform coefficient of the residual. More specifically, in the CABAC entropy decoding method, a binary number (bin) corresponding to each syntax element may be received from the bitstream, a context model may be determined using the decoding target syntax element information and the decoding information of adjacent and decoding target blocks or the information of the symbol / binary number decoded in the previous step, the probability of the binary number occurrence may be predicted based on the determined context model, and a symbol corresponding to the value of each syntax element may be generated by performing arithmetic decoding on the binary number. In this case, in the CABAC entropy decoding method, after determining the context model, the context model may be updated using the information of the symbol / binary number decoded by the context model for the next symbol / binary number. The information about prediction among the information decoded in the entropy decoder 2110 may be provided to the predictors (inter-frame predictor 260 and intra-frame predictor 265). The parameter information (i.e., quantization transform coefficient) related to the residual value for which entropy decoding has been performed in the entropy decoder 210 may be input to the dequantizer 220. In addition, the information about filtering among the information decoded in the entropy decoder 210 may be provided to the filter 240. In addition, a receiver (not shown) that receives the signal output by the encoding device may be further configured as an internal / external element of the decoding device 200, or the receiver may be an element of the entropy decoder 210.
[0097] The dequantizer 220 may dequantize the quantized transform coefficients and output the transform coefficients. The dequantizer 220 may rearrange the quantized transform coefficients in a two-dimensional block form. In this case, the rearrangement may be performed based on the coefficient scan order executed in the encoding device. The dequantizer 220 may perform dequantization on the quantized transform coefficients using quantization parameters (e.g., quantization step information) and may obtain the transform coefficients.
[0098] The inverse transformer 230 may output a residual signal (residual block or residual sample array) by applying an inverse transform to the transform coefficients.
[0099] The predictor may perform prediction on the current block and may generate a prediction block including prediction samples for the current block. The predictor may determine whether to apply intra prediction or inter prediction to the current block based on the information about prediction output by the entropy decoder 210 and may determine a specific intra / inter prediction mode.
[0100] The intra predictor 265 may predict the current block by referring to samples within the current picture. Depending on the prediction mode, the position of the reference samples may be adjacent to or spaced apart from the current block. In intra prediction, the prediction mode may include a plurality of non-angle modes and a plurality of angle modes. The intra predictor 265 may use the prediction mode applied to an adjacent block to determine the prediction mode applied to the current block.
[0101] The inter predictor 260 may derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. In this case, in order to reduce the amount of motion information transmitted in the inter prediction mode, based on the correlation of the motion information between an adjacent block and the current block, the motion information may be predicted as a block, a sub-block, or a sample unit. The motion information may include a motion vector and a reference picture index. The motion information may further include inter prediction direction (L0 prediction, L1 prediction, and Bi prediction) information. In the case of inter prediction, the adjacent blocks may include spatially adjacent blocks within the current picture and temporally adjacent blocks within the reference picture. For example, the inter predictor 260 may configure a motion information candidate list based on the adjacent blocks and may derive the motion vector and / or the reference picture index for the current block based on the received candidate selection information. The inter prediction may be performed based on various prediction modes. The information about prediction may include information indicating the mode of inter prediction for the current block.
[0102] The adder 235 may generate a reconstructed signal (reconstructed picture, reconstructed block, or reconstructed sample array) by adding the obtained residual signal to the prediction signal (prediction block or prediction sample array) output by the inter predictor 260 or the intra predictor 265. If there is no residual for the processing target block as in the case where the skip mode has been applied, the prediction block may be used as the reconstructed block.
[0103] The adder 235 may be referred to as a reconstructor or a reconstruction block generator. The generated reconstruction signal may be used for intra prediction of the next processing target block within the current picture, and may be used for inter prediction of the next picture through filtering, which will be described later.
[0104] The filter 240 may improve the subjective / objective picture quality by applying filtering to the reconstruction signal. For example, the filter 240 may generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture, and may send the modified reconstructed picture to the memory 250, more specifically, to the DPB of the memory 250. The various filtering methods may include, for example, deblocking filtering, sample adaptive offset (SAO), adaptive loop filtering (ALF), and bilateral filtering.
[0105] The (modified) reconstructed picture stored in the DPB of the memory 250 may be used as a reference picture in the inter predictor 260. The memory 250 may store the motion information of the blocks in which the motion information in the current picture is derived (or decoded) and / or the motion information of the blocks in the reconstructed picture. The stored motion information may be forwarded to the inter predictor 260 to be used as the motion information of spatially adjacent blocks or temporally adjacent blocks. The memory 170 may store the reconstructed samples of the reconstructed blocks in the current picture and forward them to the intra predictor 265.
[0106] In the present disclosure, the embodiments described in the filter 160, the inter predictor 180, and the intra predictor 185 of the encoding device 100 may be applied to the filter 240, the inter predictor 260, and the intra predictor 265 of the decoding device 200, respectively, identically, or in a corresponding manner.
[0107] Block Partitioning
[0108] The video / image encoding method according to the present disclosure may be performed based on various detailed techniques, and each of the various detailed techniques is described below. It will be apparent to those skilled in the art that the techniques described herein may be associated with related processes such as prediction, residual processing ((inverse) transformation and (de) quantization, etc.), syntax element encoding, filtering, partitioning / splitting, etc. in the video / image encoding / decoding processes described above and / or below.
[0109] The block partitioning process according to the present disclosure may be performed in the image divider 110 of the above-described encoding device, and the partitioning-related information may be (encoded) processed in the entropy encoder 190 and forwarded to the decoding device in the form of a bitstream. The entropy decoder 210 of the decoding device may obtain the block partitioning structure of the current picture based on the partitioning-related information obtained from the bitstream, and based on this, a series of processes for image decoding (such as prediction, residual processing, block reconstruction, and loop filtering, etc.) may be performed.
[0110] Partition the picture into CTUs
[0111] A picture may be divided into a sequence of coding tree units (CTUs). A CTU may correspond to a coding tree block (CTB). Alternatively, a CTU may include a coding tree block of luminance samples and two coding tree blocks of corresponding chrominance samples. In other words, for a picture including three types of sample arrays, a CTU may include an N×N block of luminance samples and two corresponding samples of chrominance samples.
[0112] The maximum supported size of a CTU for encoding and prediction may be different from the maximum supported size of a CTU for transformation. For example, the maximum supported size of the luminance block in a CTU may be 128×128.
[0113] Partition the CTUs using a tree structure
[0114] A CTU may be divided into CUs based on a quadtree (QT) structure. The quadtree structure may be referred to as a quaternary structure. This is to reflect various local features. In addition, in the present disclosure, a CTU may be divided based on a multi-type tree structure partitioning including a binary tree (BT), a ternary tree (TT), and a quadtree. Hereinafter, the QTBT structure may include a quadtree and a binary tree structure, and the QTBTTT may include a partitioning structure based on a binary tree and a ternary tree. Alternatively, the QTBT structure may also include a partitioning structure based on a quadtree, a binary tree, and a ternary tree. In the coding tree structure, a CU may have a square shape or a rectangular shape. First, a CTU may be divided into a quadtree structure. And then, the leaf nodes of the quadtree structure may be further divided by a multi-type tree structure.
[0115] Figure 3 It is a diagram showing an example of a multi-type tree structure to which an embodiment of the present disclosure may be applied.
[0116] In one embodiment of the present disclosure, the multi-type tree structure may include as Figure 3The four split types shown. These four split types can include vertical binary split (SPLIT_BT_VER), horizontal binary split (SPLIT_BT_HOR), vertical ternary split (SPLIT_TT_VER), and horizontal ternary split (SPLIT_TT_HOR). The leaf nodes of the multi-type tree structure can be referred to as CUs. Such CUs can be used in prediction and transformation processes. In the present disclosure, generally, CUs, PUs, and TUs can have the same block size. However, in the case where the maximum supported transform length is less than the width or height of the color component, CUs and TUs can have different block sizes.
[0117] Figure 4 FIG. is a diagram showing a signaling mechanism of partition split information of a quadtree having a nested multi-type tree structure to which embodiments of the present disclosure can be applied.
[0118] Here, a CTU can be regarded as the root of the quadtree and is initially partitioned into a quadtree structure. Each quadtree leaf node can then be further partitioned into a multi-type tree structure. In this multi-type tree structure, a first flag (e.g., mtt_split_cu_flag) is signaled to indicate whether the corresponding node is further partitioned. In the case where the corresponding node is further partitioned, a second flag (e.g., mtt_split_cu_vertical_flag) is signaled to indicate the split direction. Then, a third flag (e.g., mtt_split_cu_binary_flag) is signaled to indicate whether the split type is a binary split or a ternary split. For example, based on mtt_split_cu_vertical_flag and mtt_split_cu_binary_flag, a multi-type tree split mode (MttSplitMode) can be derived as shown in Table 1 below.
[0119] [Table 1]
[0120] MttSplitMode mtt_split_cu_vertical_flag mtt_split_cu_binary_flag SPLIT_TT_HOR 0 0 SPLIT_BT_HOR 0 1 SPLIT_TT_VER 1 0 SPLIT_BT_VER 1 1
[0121] Figure 5 FIG. is a diagram showing a method of partitioning a CTU into a plurality of CUs based on a quadtree and a nested multi-type tree structure to which embodiments of the present disclosure can be applied.
[0122] Here, the bold block edges represent quadtree partitioning, and the remaining edges represent multi-type tree partitioning. Quadtree partitioning with nested multi-type trees can provide a content-adaptive coding tree structure. A CU can correspond to a coding block (CB). Alternatively, a CU can include a coding block for luminance samples and two coding blocks for corresponding chrominance samples. The size of a CU can be as large as a CTU or as small as 4×4 (in terms of luminance samples). For example, in the case of the 4:2:0 color format (or chrominance format), the maximum chrominance CB size can be 64×64, and the minimum chrominance CB size can be 2×2.
[0123] In the present disclosure, for example, the maximum supported luminance TB size can be 64×64, and the maximum supported chrominance TB size can be 32×32. In the case where the width or height of a CB partitioned according to the tree structure is greater than the maximum transform width or height, the CB can be further partitioned until the TB size limits in the horizontal and vertical directions are automatically (or implicitly) satisfied.
[0124] Furthermore, for a quadtree coding tree scheme with nested multi-types, the following parameters can be defined or identified as SPS syntax elements.
[0125] - CTU size: the size of the root node of the quadtree
[0126] - MinQTSize: the minimum allowed quadtree leaf node size
[0127] - MaxBtSize: the maximum allowed binary tree root node size
[0128] - MaxTtSize: the maximum allowed ternary tree root node size
[0129] - MaxMttDepth: the maximum allowed hierarchical depth of the multi-type tree split from the quadtree leaf
[0130] - MinBtSize: the minimum allowed binary tree leaf node size
[0131] - MinTtSize: the minimum allowed ternary tree leaf node size
[0132] As an example of a quadtree coding tree scheme with a nested multi-type tree, the CTU size can be set to a block of 128×128 luma samples and 64×64 for two corresponding chroma samples (in 4:2:0 chroma samples). In this case, MinOTSize can be set to 16×16, MaxBtSize can be set to 128×128, MaxTtSzie can be set to 64×64, MinBtSize and MinTtSize (for both width and height) can be set to 4×4, and MaxMttDepth can be set to 4. A quadtree partition can be applied to the CTU and quadtree leaf nodes can be generated. The quadtree leaf nodes can be referred to as leaf QT nodes. The size of the quadtree leaf nodes can range from a size of 16×16 (i.e., MinOTSize) to a size of 128×128 (i.e., CTU size). In the case where the leaf QT node is 128×128, the leaf QT node may not be partitioned into a binary tree / trinary tree. This is because, even if the leaf QT node is partitioned, the leaf QT node exceeds MaxBtsize and MaxTtszie (i.e., 64×64). In other cases, the leaf QT node can be additionally partitioned into a multi-type tree. Thus, the leaf QT node can be the root node of the multi-type tree and the leaf QT node can have a multi-type tree depth (mttDepth) value of 0. In the case where the multi-type tree depth reaches MaxMttdepth (e.g., 4), additional partitioning may no longer be considered. In the case where the width of the multi-type tree node is equal to MinBtSize and less than or equal to 2×MinTtSize, additional horizontal partitioning may no longer be considered. In the case where the height of the multi-type tree node is equal to MinBtSize and less than or equal to 2×MinTtSize, additional vertical partitioning may no longer be considered.
[0133] Figure 6 is a diagram showing a method for restricting trinary tree splitting as an embodiment to which the present disclosure can be applied.
[0134] Referring to Figure 6 , in order to support a 64×64 luma block and 32×32 chroma pipeline design in a hardware decoder, TT splitting can be restricted in certain cases. For example, in the case where the width or height of a luma coding block is greater than a predetermined specific value (e.g., 32, 64), as Figure 6 shown, TT splitting can be restricted.
[0135] In the present disclosure, the coding tree scheme can support that the luminance and chrominance blocks have their respective block tree structures. For P slices and B slices, the luminance and chrominance CTBs in a single CTU can be restricted to have the same coding tree structure. However, for I slices, the luminance and chrominance blocks can have their respective separate block tree structures. In the case of applying the separate block tree mode, the luminance CTB can be partitioned into CUs based on a specific coding tree structure, and the chrominance CTB can be partitioned into chrominance CUs based on a different coding tree structure. This may mean that the CUs in I slices can include coded blocks of a chrominance component or coded blocks of two chrominance components, and the CUs in P slices or B slices can include blocks of three color components.
[0136] In the above "Partitioning CTUs Using a Tree Structure", a quadtree coding tree scheme with nested multi-type trees is described, but the structure for partitioning CUs is not limited thereto. For example, the BT structure and the TT structure can be interpreted as concepts included in the multi-partition tree (MPT) structure, and it can be interpreted that the CUs are partitioned by the QT structure and the MPT structure. In an example of partitioning CUs by the QT structure and the MPT structure, a syntax element (e.g., MPT_split_type) including information about the number of blocks into which the leaf nodes of the QT structure are partitioned and a syntax element (e.g., MPT_split_mode) including information about the direction in which the leaf nodes of the QT structure are partitioned in the vertical and horizontal directions can be signaled, and the splitting structure can be determined.
[0137] In another example, the CUs can be partitioned in a method different from the QT structure, the BT structure, or the TT structure. That is, different from partitioning a CU of a lower layer depth into a 1 / 4 size of a CU of a higher layer depth according to the QT structure, partitioning a CU of a lower layer depth into a 1 / 2 size of a CU of a higher layer depth according to the BT structure, or partitioning a CU of a lower layer depth into a 1 / 4 size or 1 / 2 size of a CU of a higher layer depth according to the TT structure, in some cases, a CU of a lower layer depth is partitioned into a 1 / 5, 1 / 3, 3 / 8, 3 / 5, 2 / 3, or 5 / 8 size of a CU of a higher layer depth, but the method of partitioning CUs is not limited thereto.
[0138] In the case where a part of a tree node block exceeds the bottom or right picture boundary, the corresponding tree node block can be restricted so that all samples of all coded CUs are within the picture boundary. In this case, for example, the following splitting rules can be applied.
[0139] - If a part of the tree node block exceeds both the bottom and right picture boundaries,
[0140] - If the block is a QT node and the size of the block is greater than the minimum QT size, the block is forced to be split using the QT split mode.
[0141] - Otherwise, the block is forced to be split using the SPLIT_BT_HOR mode
[0142] - Otherwise, if a part of the tree node block exceeds the bottom picture boundary,
[0143] - If the block is a QT node, and the size of the block is greater than the minimum QT size, and the size of the block is greater than the maximum BT size, the block is forced to be split using the QT split mode.
[0144] - Otherwise, if the block is a QT node, and the size of the block is greater than the minimum QT size, and the size of the block is less than or equal to the maximum BT size, the block is forced to be split using the QT split mode or the SPLIT_BT_HOR mode.
[0145] - Otherwise (the block is a BTT node, or the size of the block is less than or equal to the minimum QT size), the block is forced to be split using the SPLIT_BT_HOR mode.
[0146] - Otherwise, if a part of the tree node block exceeds the right picture boundary,
[0147] - If the block is a QT node, and the size of the block is greater than the minimum QT size, and the size of the block is greater than the maximum BT size, the block is forced to be split using the QT split mode.
[0148] - Otherwise, if the block is a QT node, and the size of the block is greater than the minimum QT size, and the size of the block is less than or equal to the maximum BT size, the block is forced to be split using the QT split mode or the SPLIT_BT_VER mode.
[0149] - Otherwise (the block is a BTT node, or the size of the block is less than or equal to the minimum QT size), the block is forced to be split using the SPLIT_BT_VER mode.
[0150] The quadtree coding block structure accompanying the multi-type tree can provide a very flexible block partitioning structure. Due to the split types supported by the multi-type tree, in some cases, different split modes may result in the same coding block structure. The generation of redundant split modes is restricted to reduce the data volume of the partitioning information. It will be described with reference to the following drawings.
[0151] Figure 7 is a diagram showing redundant partitioning patterns that may occur in binary tree partitioning and ternary tree partitioning as embodiments to which the present disclosure can be applied.
[0152] As Figure 7As shown, the two-level consecutive binary partitioning in one direction has the same coded block structure as the binary partitioning for the middle partition after ternary partitioning. In this case, the binary tree partitioning (along a given direction) for the middle partition of the ternary tree partitioning can be restricted. This restriction can be applied to the CUs of all pictures. When a specific partitioning is restricted, the signaling of syntax elements can be modified by reflecting this restricted situation, and the number of bits signaled for the partition can be reduced by the modified signaling. For example, similar to the example shown in Figure 7 when the binary tree partitioning for the middle partition of a CU is restricted, the syntax element mtt_split_cu_binary_flag indicating whether the partitioning is binary or ternary partitioning may not be signaled, and the decoder can infer the value as 0.
[0153] Prediction
[0154] To reconstruct the current processing unit in which decoding is performed, the decoded portions of the current picture or other pictures including the current processing unit can be used.
[0155] A picture that uses only the current picture for reconstruction (i.e., performs intra prediction) can be called an intra picture or I picture (slice), a picture (slice) that uses at most one motion vector and reference index to predict each unit can be called a predictive picture or P picture (slice), and a picture (slice) that uses at most two motion vectors and reference indexes can be called a bi-predictive picture or B picture (slice).
[0156] Intra prediction refers to a prediction method for deriving the current processing block from the data elements (e.g., sample values, etc.) of the same decoded picture (or slice). In other words, intra prediction refers to a method of predicting the pixel values of the current processing block by referring to the reconstructed region in the current picture.
[0157] Inter prediction will be described in more detail below.
[0158] Inter-frame prediction
[0159] Inter prediction refers to a prediction method for deriving the current processing block based on the data elements (e.g., sample values or motion vectors) of pictures other than the current picture. In other words, inter prediction refers to a method of predicting the pixel values of the current processing block by referring to the reconstructed regions in other reconstructed pictures other than the current picture.
[0160] Inter prediction (inter-picture prediction), as a technique for eliminating redundancy existing between pictures, is mainly performed through motion estimation and motion compensation.
[0161] In the present disclosure, a detailed description of the inter-frame prediction method described above in Figure 1 and Figure 2 is carried out, and the decoder can be represented as the inter-frame prediction unit in the Figure 10 video / image decoding method based on inter-frame prediction and Figure 11 the decoding device to be described below. In addition, the encoder can be represented as the inter-frame prediction unit in the Figure 8 video / image encoding method based on inter-frame prediction and Figure 9 the encoding device to be described below. Additionally, the encoded data of Figure 8 and Figure 9 can be stored in the form of a bitstream.
[0162] The prediction unit of the encoding device / decoding device can derive prediction samples by performing inter-frame prediction in units of blocks. Inter-frame prediction can be represented as prediction derived by a method that depends on data elements (e.g., sample values or motion information) of pictures other than the current picture. When inter-frame prediction is applied to the current block, a prediction block (prediction sample array) for the current block can be derived based on a reference block (reference sample array) specified by a motion vector on the reference picture indicated by the reference picture index.
[0163] In this case, in order to reduce the amount of motion information transmitted in the inter-frame prediction mode, based on the correlation of motion information between adjacent blocks and the current block, the motion information of the current block can be predicted in units of blocks, sub-blocks, or samples. The motion information can include a motion vector and a reference picture index. The motion information can also include information on the inter-frame prediction type (such as L0 prediction, L1 prediction, and Bi prediction, etc.).
[0164] In the case of applying inter-frame prediction, adjacent blocks can include spatially adjacent blocks present in the current picture and temporally adjacent blocks present in the reference picture. The reference picture including the reference block and the reference picture including the temporally adjacent blocks can be the same as or different from each other. Temporally adjacent blocks can be referred to by names such as collocated reference blocks, collocated CUs (colCUs), etc., and the reference picture including the temporally adjacent blocks can be referred to as a collocated picture (colPic). For example, a motion information candidate list can be configured based on the adjacent blocks of the current block, and a flag or index information indicating which candidate is selected (used) can be signaled to facilitate the derivation of the motion vector and / or reference picture index of the current block.
[0165] Inter-frame prediction can be performed based on various prediction modes. For example, in the case of the skip mode and the merge mode, the motion information of the current block can be the same as that of the selected neighboring block. In the case of the skip mode, a residual signal may not be sent as in the merge mode. In the case of the motion vector prediction (MVP) mode, the motion vector of the selected neighboring block can be used as a motion vector predictor, and the motion vector difference can be signaled. In this case, the motion vector of the current block can be derived by using the sum of the motion vector predictor and the motion vector difference.
[0166] Figure 8 And Figure 9 FIG. is a diagram illustrating an inter-frame prediction unit in a video / image encoding method based on inter-frame prediction according to an embodiment of the present disclosure and an encoding device according to an embodiment of the present disclosure.
[0167] Refer to Figure 8 And Figure 9 , S801 can be executed by the inter-frame prediction unit (inter-frame predictor) 180 of the encoding device and S802 can be executed by the residual processing unit of the encoding device. Specifically, S802 can be executed by the subtractor 115 of the encoding device. In S803, the prediction information can be derived by the inter-frame prediction unit (inter-frame predictor) 180 and encoded by the entropy encoder 190. In S803, the residual information can be derived by the residual processing unit and encoded by the entropy encoder 190. The residual information is information about the residual samples. The residual information can include information about the quantized transform coefficients of the residual samples.
[0168] As described above, the residual samples can be derived as transform coefficients by the transformer 120 of the encoding device, and the transform coefficients can be derived as quantized transform coefficients by the quantizer 130. Information about the quantized transform coefficients can be encoded by the entropy encoder 190 through a residual encoding process.
[0169] The encoding device performs inter-frame prediction (S801) for the current block. The encoding device can derive the inter-frame prediction mode and motion information of the current block and generate prediction samples of the current block. Here, the inter-frame prediction mode determination process, the motion information derivation process, and the prediction sample generation process can be performed simultaneously, and any one of the processes can be executed earlier than the other processes. For example, the inter-frame predictor 180 of the encoding device can include a prediction mode determination unit 181, a motion information derivation unit 182, and a prediction sample derivation unit 183, and the prediction mode determination unit 181 can determine the prediction mode for the current block, the motion information derivation unit 182 can derive the motion information of the current block, and the prediction sample derivation unit 183 can derive the prediction samples of the current block.
[0170] For example, an inter prediction unit (inter predictor) 180 of an encoding device may search for a block similar to a current block in a predetermined area (search area) of a reference picture through motion estimation, and derive a reference block with the smallest difference from the current block or with a difference equal to or less than a predetermined criterion. A reference picture index indicating the reference picture where the reference block is located may be derived based on this, and a motion vector may be derived based on the positional difference between the reference block and the current block. The encoding device may determine a mode to be applied to the current block among various prediction modes. The encoding device may compare the rate-distortion (RD) cost for various prediction modes and determine an optimal prediction mode for the current block.
[0171] For example, when a skip mode or a merge mode is applied to the current block, the encoding device may configure a merging candidate list to be described below, and derive a reference block with the smallest difference from the current block or with a difference equal to or less than a predetermined criterion among the reference blocks indicated by the merging candidates included in the merging candidate list. In this case, a merging candidate associated with the derived reference block may be selected, and merging index information indicating the selected merging candidate may be generated and signaled to the decoding device. The motion information of the current block may be derived by using the motion information of the selected merging candidate.
[0172] As another example, when an (A)MVP mode is applied to the current block, the encoding device may configure an (A)MVP candidate list to be described below, and use the motion vector of a selected mvp candidate among the motion vector prediction value (mvp) candidates included in the (A)MVP candidate list as the mvp of the current block. In this case, for example, the motion vector indicating the reference block derived through motion estimation may be used as the motion vector of the current block. Also, the mvp candidate having the smallest difference from the motion vector of the current block among the mvp candidates may become the selected mvp candidate. A motion vector difference (MVD) that is the difference obtained by subtracting the mvp from the motion vector of the current block may be derived. In this case, information about the MVD may be signaled to the decoding device. Further, when the (A)MVP mode is applied, the value of the reference picture index may be configured as reference picture index information and signaled separately to the decoding device.
[0173] The encoding device may derive a residual sample based on a prediction sample (S802). The encoding device may derive the residual sample by comparing the original sample of the current block with the prediction sample.
[0174] The encoding device encodes image information including prediction information and residual information (S803). The encoding device may output the encoded image information in the form of a bitstream. The prediction information may include information about prediction mode information (e.g., skip flag, merge flag, or mode index, etc.) and information about motion information as information related to the prediction process. The information about motion information may include candidate selection information as information for deriving a motion vector (e.g., merge index, mvp flag, or mvp index). In addition, the information about motion information may include information about MVD and / or reference picture index information.
[0175] In addition, the information about motion information may include information indicating whether L0 prediction, L1 prediction, or bilateral prediction is applied. The residual information is information about residual samples. The residual information may include information about quantization transform coefficients of the residual samples.
[0176] The output bitstream may be stored in a (digital) storage medium and transmitted to the decoding device or transmitted to the decoding device via a network.
[0177] In addition, as described above, the encoding device may generate a reconstructed picture (including reconstructed samples and reconstructed blocks) based on reference samples and residual samples. This is to derive the same prediction result as the prediction result performed by the decoding device, and thus, the encoding efficiency can be improved. Therefore, the encoding device may store the reconstructed picture (or reconstructed samples or reconstructed blocks) in the memory and use the reconstructed picture as a reference picture. As described above, the loop filter process may be further applied to the reconstructed picture.
[0178] Figure 10 and Figure 11 are diagrams showing an inter-frame prediction-based video / image decoding method according to an embodiment of the present disclosure and an inter-frame prediction unit in a decoding device according to an embodiment of the present disclosure.
[0179] Referring to Figure 10 and Figure 11 The decoding device may perform operations corresponding to the operations performed by the encoding device. The decoding device may perform prediction for the current block based on the received prediction information and derive prediction samples.
[0180] S1001 to S1003 can be performed by the inter - frame prediction unit (inter - frame predictor) 260 of the decoding device, and the entropy decoder 210 of the decoding device can obtain the residual information of S1004 from the bitstream. The residual processing unit of the decoding device can derive the residual samples for the current block based on the residual information. Specifically, the de - quantizer 220 of the residual processing unit can derive the transform coefficients by performing de - quantization according to the quantization transform coefficients derived based on the residual information, and the inverse transformer 230 of the residual processing unit can derive the residual samples for the current block by performing an inverse transform on the transform coefficients. S1005 can be performed by the adder 235 or the reconstruction unit of the decoding device.
[0181] Specifically, the decoding device can determine the prediction mode for the current block based on the received prediction information (S1001). The decoding device can determine which inter - frame prediction mode to apply to the current block based on the prediction mode information in the prediction information.
[0182] For example, it can be determined whether to apply the merge mode or the (A)MVP mode to the current block based on the merge flag. Alternatively, one of various inter - frame prediction mode candidates can be selected based on the mode index. The inter - frame prediction mode candidates can include the skip mode, the merge mode, and / or the (A)MVP mode, or can include various inter - frame prediction modes to be described below.
[0183] The decoding device derives the motion information of the current block based on the determined inter - frame prediction mode (S1002). For example, when the skip mode or the merge mode is applied to the current block, the decoding device can configure a merge candidate list to be described below and select one merge candidate among the merge candidates included in the merge candidate list. The selection can be performed based on the selection information (merge index). The motion information of the current block can be derived by using the motion information of the selected merge candidate. The motion information of the selected merge candidate can be used as the motion information of the current block.
[0184] As another example, when the (A)MVP mode is applied to the current block, the decoding device can configure an (A)MVP candidate list to be described below and use the motion vector of the selected mvp candidate among the motion vector prediction value (mvp) candidates included in the (A)MVP candidate list as the mvp of the current block. The selection can be performed based on the selection information (mvp flag or mvp index). In this case, the MVD of the current block can be derived based on the information about the MVD, and the motion vector of the current block can be derived based on the mvp and the MVD of the current block. In addition, the reference picture index of the current block can be derived based on the reference picture index information. The picture indicated by the reference picture index in the reference picture list for the current block can be derived as the reference picture referred to for the inter - frame prediction of the current block.
[0185] In addition, the motion information of the current block can be derived without candidate list configuration as described below, and in this case, the motion information of the current block can be derived according to the process disclosed in the prediction mode to be described below. In this case, the candidate list configuration can be omitted.
[0186] The decoding device can generate a prediction sample for the current block based on the motion information of the current block (S1003). In this case, the reference picture can be derived based on the reference picture index of the current block, and the prediction sample of the current block can be derived by using the samples of the reference block indicated by the motion vector of the current block on the reference picture. In this case, as described below, in some cases, a prediction sample filtering process for all or some of the prediction samples of the current block can be further performed.
[0187] For example, the inter prediction unit (inter predictor) 260 of the decoding device can include a prediction mode determination unit 261, a motion information derivation unit 262, and a prediction sample derivation unit 263. The prediction mode determination unit 261 can determine the prediction mode for the current block based on the received prediction mode information. The motion information derivation unit 262 can derive the motion information (motion vector and / or reference picture index) of the current block based on the information about the received motion information. And the prediction sample derivation unit 263 can derive the prediction sample of the current block.
[0188] The decoding device generates a residual sample for the current block based on the received residual information (S1004). The decoding device can generate a reconstructed sample for the current block based on the prediction sample and the residual sample, and generate a reconstructed picture based on the generated reconstructed sample (S1005). Thereafter, as described above, the in-loop filtering process can be further applied to the reconstructed picture.
[0189] As described above, the inter prediction process can include an inter prediction mode determination step, a motion information derivation step according to the determined prediction mode, and a prediction execution (prediction sample generation) step based on the derived motion information.
[0190] Determination of the inter-frame prediction mode
[0191] Various inter-frame prediction modes can be used to predict a current block in a picture. For example, various modes including a merge mode, a skip mode, an MVP mode, and an affine mode can be used. A decoder side motion vector refinement (DMVR) mode, an adaptive motion vector resolution (AMVR) mode, etc. can be further used as an ancillary mode. The affine mode can be referred to as an affine motion prediction mode. The MVP mode can be referred to as an advanced motion vector prediction (AMVP) mode.
[0192] Prediction mode information indicating an inter-frame prediction mode of a current block can be signaled from an encoding device to a decoding device. The prediction mode information can be included in a bitstream and received by the decoding device. The prediction mode information can include index information indicating one of a plurality of candidate modes. Alternatively, the inter-frame prediction mode can be indicated by hierarchical signaling of flag information. In this case, the prediction mode information can include one or more flags.
[0193] For example, it can be indicated whether to apply the skip mode by signaling a skip flag, and when the skip mode is not applied, it can be indicated whether to apply the merge mode by signaling a merge flag, and when the merge mode is not applied, it is indicated to apply the MVP mode or flags for additional discrimination can be further signaled. The affine mode can be signaled as an independent mode, or as a dependent mode of the merge mode or the MVP mode. For example, the affine mode can be configured as a candidate in a merge candidate list or an MVP candidate list described below.
[0194] Derive motion information according to the inter-frame prediction mode
[0195] Inter - frame prediction can be performed by using the motion information of the current block. The encoding device can derive the optimal motion information for the current block through a motion estimation process. For example, the encoding device can search for similar reference blocks with high correlation in a predetermined search range in the reference picture in units of fractional pixels by using the original block in the original picture for the current block, and derive the motion information based on the searched reference block. The similarity of the block can be derived based on the difference of the sample values based on the phase. For example, the similarity of the block can be calculated based on the SAD between the current block (or the template of the current block) and the reference block (or the template of the reference block). In this case, the motion information can be derived based on the reference block with the minimum SAD in the search area. The derived motion information can be signaled to the decoding device according to various methods based on the inter - frame prediction mode.
[0196] Merge mode and skip mode
[0197] Figure 12 is a diagram for describing adjacent blocks used in the merge mode or skip mode as an embodiment of applying the present disclosure.
[0198] When the merge mode is applied, the motion information of the current prediction block is not directly sent, and the motion information of the current prediction block is derived by using the motion information of the adjacent prediction blocks. Therefore, flag information indicating the use of the merge mode and a merge index indicating which adjacent prediction block to use are sent to indicate the motion information of the current prediction block.
[0199] The encoder can search for merge candidate blocks for deriving the motion information of the current prediction block to facilitate the execution of the merge mode. For example, up to five merge candidate blocks can be used, but the present disclosure is not limited thereto. Additionally, the maximum number of merge candidate blocks can be sent in the slice header (or tile group header), and the present disclosure is not limited thereto. After finding the merge candidate blocks, the encoder can generate a merge candidate list and select the merge candidate block with the minimum cost among the merge candidate blocks as the final merge candidate block.
[0200] The present disclosure provides various embodiments for the merge candidate blocks constituting the merge candidate list.
[0201] As a merge candidate list, for example, five merge candidate blocks can be used. For example, four spatial merge candidates and one temporal merge candidate can be used. As a specific example, in the case of spatial merge candidates, Figure 12 the blocks shown can be used as spatial merge candidates.
[0202] Figure 13FIG. is a flowchart showing a method for configuring a merge candidate list according to an embodiment of the present disclosure.
[0203] Referring Figure 13 , an encoding device (encoder / decoder) inserts spatial merge candidates obtained by searching for spatially adjacent blocks of a current block into a merge candidate list (S1301). For example, spatially adjacent blocks may include the lower left adjacent block, left adjacent block, upper right adjacent block, upper adjacent block, and upper left adjacent block of the current block. However, this is an example, and in addition to the above spatially adjacent blocks, additional adjacent blocks including the right adjacent block, lower adjacent block, and lower right adjacent block, etc. may be used as spatially adjacent blocks. The encoding device may derive available blocks by searching for spatially adjacent blocks based on priority, and may derive the motion information of the detected blocks as spatial merge candidates. For example, the encoder and decoder may search for the five blocks shown Figure 12 in the order of A1, B1, B0, A0, and B2, index the available candidates in sequence, and configure the indexed candidates as a merge candidate list.
[0204] The encoding device inserts temporal merge candidates obtained by searching for temporally adjacent blocks of the current block into the merge candidate list (S1302). The temporally adjacent blocks may be located on a reference picture that is a picture different from the current picture in which the current block is located. The reference picture on which the temporally adjacent blocks are located may be referred to as a collocated picture or col picture. The temporally adjacent blocks may be searched for on the col picture in the order of the lower right adjacent block and the lower right middle block of the co-located block for the current block.
[0205] In addition, when applying motion data compression, specific motion information may be stored as representative motion information in the col picture for each predetermined storage unit. In this case, it is not necessary to store the motion information of all blocks in the predetermined storage unit, and as a result, a motion data compression effect can be obtained. In this case, the above-mentioned predetermined storage unit may be determined in advance for each 16×16 sample unit or 8×8 sample unit, or the size information of the predetermined storage unit may be signaled from the encoder to the decoder. When applying motion data compression, the motion information of the temporally adjacent blocks may be replaced with the representative motion information of the predetermined storage unit in which the temporally adjacent blocks are located.
[0206] In other words, in this case, in terms of implementation, a temporal merge candidate may be derived based on motion information of a prediction block (instead of a prediction block located at coordinates of a temporally adjacent block) that overlays a position where coordinates (top-left sample position) of temporally adjacent blocks are arithmetically right-shifted by a predetermined value and then arithmetically left-shifted. For example, when a predetermined storage unit is a 2n×2n sample unit, if coordinates of a temporally adjacent block are (xTnb, yTnb), motion information of a prediction block located at ((xTnb>>n)<<n), (yTnb>>n)<<n)), which is a modified position, may be used for a temporal merge candidate.
[0207] Specifically, for example, when a predetermined storage unit is a 16×16 sample unit, if coordinates of a temporally adjacent block are (xTnb, yTnb), motion information of a prediction block located at ((xTnb>>4)<<4), (yTnb>>4)<<4)), which is a modified position, may be used for a temporal merge candidate. Alternatively, for example, when a predetermined storage unit is an 8×8 sample unit, if coordinates of a temporally adjacent block are (xTnb, yTnb), motion information of a prediction block located at ((xTnb>>3)<<3), (yTnb>>3)<<3)), which is a modified position, may be used for a temporal merge candidate.
[0208] The encoding device may check whether the number of current merge candidates is less than the maximum number of merge candidates (S1303). The maximum number of merge candidates may be predefined or signaled from the encoder to the decoder. For example, the encoder may generate information on the maximum number of merge candidates, encode the generated information, and transmit the encoded information to the decoder in the form of a bitstream. When the maximum number of merge candidates is completely exhausted, subsequent candidate addition processes may not be performed.
[0209] As a result of the check, when the number of current merge candidates is less than the maximum number of merge candidates, the encoding device inserts an additional merge candidate into the merge candidate list (S1304). The additional merge candidate may include, for example, an ATMVP, a combined bi-predictive merge candidate (when the slice type of the current slice is type B), and / or a zero vector merge candidate.
[0210] As a result of the check, when the number of current merge candidates is not less than the maximum number of merge candidates, the encoding device may terminate the configuration of the merge candidate list. In this case, the encoder may select an optimal merge candidate among the merge candidates constituting the merge candidate list based on rate-distortion (RD) cost, and signal selection information (e.g., a merge index) indicating the selected merge candidate to the decoder. The decoder may select an optimal merge candidate based on the merge candidate list and the selection information.
[0211] The motion information of the selected merge candidate can be used as the motion information of the current block as described above, and the predicted samples of the current block can be derived based on the motion information of the current block. The encoder can derive the residual samples of the current block based on the predicted samples and signal the residual information for the residual samples to the decoder. The decoder can generate the reconstructed samples according to the residual samples and the predicted samples derived based on the residual information as described above, and generate the reconstructed picture based on the generated reconstructed samples.
[0212] When the skip mode is applied, the motion information of the current block can be derived by the same method as in the case of applying the merge mode as described above. However, when the skip mode is applied, the residual signal for the corresponding block is omitted, and as a result, the predicted samples can be directly used as the reconstructed samples.
[0213] MVP mode
[0214] Figure 14 is a flowchart showing a method for configuring a merge candidate list according to an embodiment to which the present disclosure is applied.
[0215] When the motion vector prediction (MVP) mode is applied, the motion vector prediction value (mvp) candidate list can be generated by using the motion vectors of the reconstructed spatially adjacent blocks (e.g., the adjacent blocks described above) and / or the motion vectors corresponding to the temporally adjacent blocks (or Col blocks). In other words, the motion vectors of the reconstructed spatially adjacent blocks and / or the motion vectors corresponding to the temporally adjacent blocks can be used as the motion vector prediction value candidates. Figure 12 The information about the prediction can include selection information (e.g., an MVP flag or an MVP index) indicating the selected optimal motion vector prediction value candidate among the motion vector prediction value candidates included in the list. In this case, the predictor can select the motion vector prediction value of the current block among the motion vector prediction value candidates included in the motion vector candidate list by using the selection information. The predictor of the encoding device can obtain the motion vector difference (MVD) between the motion vector of the current block and the motion vector prediction value, and encode the obtained MVD and output the encoded MVD in the form of a bitstream. In other words, the MVD can be obtained by the value obtained by subtracting the motion vector prediction value from the motion vector of the current block. In this case, the predictor of the decoding device can obtain the motion vector difference included in the information about the prediction and derive the motion vector of the current block by adding the motion vector difference and the motion vector prediction value. The predictor of the decoding device can obtain or derive the reference picture index indicating the reference picture from the information about the prediction. For example, the motion vector prediction value candidate list can be configured as
[0216] shown. Figure 14 shown.
[0217] Affine motion prediction
[0218] Figure 15 An example of a motion model according to an embodiment of the present disclosure is shown.
[0219] Conventional image compression techniques (e.g., High Efficiency Video Coding (HEVC)) use one motion vector to represent the motion of an encoded block. Although the optimal motion in units of blocks can be represented for each block in a method using one motion vector, the optimal motion may not actually be the optimal motion for each pixel. Therefore, if the optimal motion vector is determined in units of pixels, the coding efficiency will be improved. Accordingly, an embodiment of the present disclosure describes a motion prediction method for encoding or decoding a video signal using a multi-motion model. In particular, motion vectors at two to four control points can be used to represent the motion vector in each pixel unit or sub-block unit of a block. A prediction scheme using motion vectors of such multiple control points can be referred to as affine motion prediction, affine prediction, etc.
[0220] An affine motion model according to an embodiment of the present disclosure can represent four motion models, such as Figure 15 the four motion models shown. An affine motion model that represents three motions (translation, scaling, and rotation) among the motions that can represent an affine motion model is called a similarity (or simplified) affine motion model. When describing the embodiments of the present disclosure, for ease of description, the similarity (or simplified) affine motion model is mainly described, but the present disclosure is not limited thereto.
[0221] Figure 16 An example of control point motion vectors for affine motion prediction according to an embodiment of the present disclosure is shown.
[0222] As Figure 16 shown, affine motion prediction can use a pair of two control point motion vectors (CPMV) v_0 and v_1 to determine the motion vector of a pixel position (or sub-block) included in a block. In this case, a set of motion vectors can be referred to as an affine motion vector field (MVF). In this case, the following Equation 1 can be used to determine the affine motion vector field.
[0223] [Equation 1]
[0224]
[0225] In Equation 1, v_0 (v_0 = {v_0x, v_0y}) indicates the motion vector CPMV0 at the first control point at the upper left position of the current block 1300. v_1 (v_1 = {v_1x, v_1y}) indicates the motion vector CPMV1 at the second control point at the upper right position of the current block 1300. In addition, w indicates the width of the current block 1300. v (v = {v_x, v_y}) indicates the motion vector at the {x, y} position. The motion vector in the sub-block (or pixel) unit can be derived using Equation 1. In one embodiment, the motion vector precision can be rounded to 1 / 16 precision.
[0226] Figure 17 An example of the motion vector for each sub-block of a block to which affine motion prediction according to an embodiment of the present disclosure has been applied is shown.
[0227] Referring Figure 17 , in the encoding or decoding process, the affine motion vector field (MVF) can be determined in units of pixels or in units of blocks. That is, in affine motion prediction, the motion vector of the current block can be derived in units of pixels or in units of sub-blocks.
[0228] If the affine motion vector field is determined in units of pixels, the motion vector can be obtained based on each pixel value. If the affine motion vector field is determined in units of blocks, the motion vector of the block can be obtained based on the center pixel value of the corresponding block. In the present disclosure, as Figure 17 shown, the case where the affine motion vector field (MVF) is determined in units of 4×4 blocks is assumed. However, this is for ease of description and does not limit the embodiments of the present disclosure. Figure 17 An example of the case where the coding block is composed of 16×16 samples and the affine motion vector field (MVF) is determined in units of 4×4 sized blocks is shown.
[0229] Affine motion prediction may include an affine merge mode (or AF_MERGE) and an affine inter-frame mode (or AF_INTER). The AF_INTER mode may include an AF_4_INTER mode using a four-parameter based motion model and an AF_6_INTER mode using a six-parameter based motion model.
[0230] Affine merge mode
[0231] AF_MERGE determines the control point motion vector (CPMV) according to the affine motion model of adjacent blocks encoded as affine motion prediction. The adjacent blocks encoded in affine order can be used for AF_MERGE. When one or more adjacent blocks are encoded as affine motion prediction, the current block can be encoded as AF_MERGE.
[0232] That is to say, if the affine merge mode is applied, the CPMV of adjacent blocks can be used to derive the CPMV of the current block. In this case, the CPMV of adjacent blocks can be used as the CPMV of the current block without any change. The CPMV of adjacent blocks is modified based on the size of adjacent blocks and the size of the current block, and it can be used as the CPMV of the current block.
[0233] Figure 18 An example of adjacent blocks used in affine motion prediction in the affine merge mode according to an embodiment of the present disclosure is shown.
[0234] In the affine merge (AF_MERGE) mode, the encoder can perform encoding as in the following process.
[0235] Step - 1: Scan adjacent blocks A1810, B 1820, C 1830, D1840, and E 1850 of the current encoded block 1800 in alphabetical order. The block that is first encoded in the affine prediction mode according to the scan order is determined as the candidate block for affine merge (AF_MERGE)
[0236] Step 2: Determine the affine motion model using the control point motion vector (CPMV) of the determined candidate block
[0237] Step 3: Determine the control point motion vector (CPMV) of the current block 1800 according to the affine motion model of the candidate block, and determine the MVF of the current block 1800.
[0238] Figure 19 An example of a block that performs affine motion prediction using adjacent blocks that have applied affine motion prediction according to an embodiment of the present disclosure is shown.
[0239] For example, as Figure 19 shown, if block A1920 is encoded in the affine mode, after determining block A1920 as the candidate block, the control point motion vector (CPMV) of block A1920 (e.g., v2 and v3) can be used to derive the affine motion model, and the control point motion vectors (CPMV) v0 and v1 of the current block 1900 can be determined. The affine motion vector field (MVF) of the current block 1900 can be determined based on the control point motion vector (CPMV) of the current block 1900, and encoding can be performed.
[0240] Figure 20 is a diagram for describing a method of generating a merge candidate list using a peripheral affine coding block according to an embodiment of the present disclosure.
[0241] Refer to Figure 20, if an affine merge candidate is used to determine a CPMV pair, candidates such as Figure 20 shown can be used. In Figure 20 , it is assumed that the scanning order of the candidate list is set to A, B, C, D, and E. However, the present disclosure is not limited thereto, and the scanning order can be preset in various orders.
[0242] In one embodiment, if the number of candidates encoded in an affine mode (or affine prediction) available in adjacent blocks (i.e., A, B, C, D, and E) is 0, the affine merge mode for the current block can be skipped. If the number of available affine candidates is one (e.g., A), the motion model of the corresponding candidate can be used to derive the control point motion vectors CPMV_0 and CPMV_1 of the current block. In this case, the index indicating the corresponding candidate may not be necessary (or encoded). If the number of available affine candidates is two or more, two candidates can be configured as a candidate list for AF_MERGE in the scanning order. In this case, candidate selection information such as an index indicating the selected candidate within the candidate list can be signaled. The selection information can be a flag or index information, and can be labeled as AF_MERGE_flag and AF_merge_idx, etc.
[0243] In one embodiment of the present disclosure, motion compensation for the current block can be performed based on the size of the sub-blocks. In this case, the size of the sub-blocks of the affine block (current block) is derived. If both the width and height of the sub-blocks are greater than 4 luminance samples, motion vectors for each sub-block are derived, and motion compensation based on DCT-IF (1 / 16 pixel for luminance and 1 / 32 for chrominance) can be invoked for the sub-block. Otherwise, motion compensation based on an enhanced bilinear interpolation filter is invoked for the entire affine block.
[0244] In one embodiment of the present disclosure, when the merge / skip flag is true and both the width and height of the CU are greater than or equal to 8, an affine flag at the CU level is signaled in the bitstream to indicate whether the affine merge mode is used. And when the CU is encoded as AF_MERGE, a merge candidate index with a maximum value of 5 is signaled to specify which motion information candidate in the affine merge candidate list is used for CUA.
[0245] Figure 21 and Figure 22 are diagrams for describing a method of configuring an affine merge candidate list using adjacent blocks encoded by affine prediction according to one embodiment of the present disclosure.
[0246] Referring to Figure 21, the affine merge candidate list is constructed according to the following steps.
[0247] 1) Insert model-based affine candidates
[0248] Model-based affine candidates mean that the candidates are derived from valid adjacent reconstructed blocks encoded in the affine mode. As Figure 21 shown, the scanning order of the candidate blocks is from left (A), up (b), upper right (C), lower left (D) to upper left (E).
[0249] If the adjacent lower left block A is encoded in the 6-parameter affine mode, the motion vectors v_4, v_5, and v_6 of the upper left corner, upper right corner, and lower left corner of the CU containing block A are obtained. And the motion vectors v_0, v_1, and v_2 of the upper left corner of the current CU are calculated according to v_4, v_5, and v_6 through the six-parameter affine model.
[0250] If the adjacent lower left block A is encoded in the 4-parameter affine mode, the motion vectors v_4 and v_5 of the upper left corner and upper right corner of the CU containing block A are obtained. And the motion vectors v_0 and v_1 of the upper left corner of the current CU are calculated according to v_4 and v_5 through the 4-parameter affine model.
[0251] 2) Insert control point-based affine candidates
[0252] Refer to Figure 21 , control point-based candidates mean that candidates are constructed by combining the adjacent motion information of each control point.
[0253] First, the motion information for the control points is derived from the specified spatial adjacent terms and temporal adjacent terms shown in Figure 21 . CP_k (k = 1, 2, 3, 4) represents the kth control point. A, B, C, D, E, F, and G are used to predict the spatial positions of CP_k (k = 1, 2, 3); H is used to predict the temporal position of CP4.
[0254] The coordinates of CP_1, CP_2, CP_3, and CP_4 are (0, 0), (W, 0), (H, 0), and (W, H) respectively, where W and H are the width and height of the current block.
[0255] The motion information of each control point is obtained according to the following priority order.
[0256] For CP_1, check the priority A → B → C. If available, use A. Otherwise, if B is available, use B. If both A and B are unavailable, use C. If all three candidates are unavailable, the motion information of CP1 cannot be obtained.
[0257] For CP_2, check the priority E → D;
[0258] For CP_3, the check priority is G → F;
[0259] For CP_4, use H.
[0260] Secondly, use a combination of control points to construct a motion model.
[0261] The motion vectors of two control points are required to calculate the transformation parameters in the 4-parameter affine model. Two control points can be selected from one of the following six combinations ({CP_1, CP_4}, {CP_2, CP_3}, {CP_1, CP_2}, {CP_2, CP_4}, {CP_1, CP_3}, {CP_3, CP_4}). For example, using the CP_1 and CP_2 control points to construct a 4-parameter affine motion model is denoted as affine(CP_1, CP_2).
[0262] The motion vectors of three control points are required to calculate the transformation parameters in the 6-parameter affine model. Three control points can be selected from one of the following four combinations ({CP_1, CP_2, CP_4}, {CP_1, CP_2, CP_3}, {CP_2, CP_3, CP_4}, {CP_1, CP_3, CP_4}). For example, using the CP_1, CP_2, and CP_3 control points to construct a 6-parameter affine motion model is denoted as affine(CP_1, CP_2, CP_3).
[0263] In addition, in one embodiment of the present disclosure, in the affine merge mode, if there is an affine merge candidate, it can always be regarded as the six-parameter affine mode.
[0264] Affine inter-frame mode
[0265] Figure 23 An example of adjacent blocks used in affine motion prediction in the affine inter-frame mode according to one embodiment of the present disclosure is shown.
[0266] Refer to Figure 23 , affine motion prediction may include an affine merge mode (or AF_MERGE) and an affine inter-frame mode (or AF_INTER). In the affine inter-frame mode (AF_INTER), after determining the two control point motion vector predictions (CPMVP) and CPMV, the control point motion vector difference CPMVD corresponding to the difference can be sent from the encoder to the decoder. The specific process of encoding the affine inter-frame mode AF_INTER can be as follows.
[0267] Step 1: Determine two pairs of CPMVP candidates
[0268] Step 1.1: Determine the maximum value of the twelfth CPMVP candidate combination (refer to Equation 2 below)
[0269] [Equation 2]
[0270] {(v0, v1, v2) | v0 = {v A , v B , v c}, v1 = {v D , v E}, v2 = {v F , v G}}
[0271] In Equation 2, v_0 indicates the motion vector CPMV0 at the upper left control point 2310 of the current block 2300. v_1 indicates the motion vector CPMV1 at the upper right control point 2311 of the current block 2300. v_2 indicates the motion vector CPMV2 at the lower left control point 2312 of the current block 2300. v_A represents the motion vector of the adjacent block A 2320 that is adjacent to the upper left of the upper left control point 2310 of the current block 2300. v_B indicates the motion vector of the adjacent block B 2322 that is adjacent to the upper side of the upper left control point 2310 of the current block 2300. v_C indicates the motion vector of the adjacent block C 2324 that is adjacent to the left side of the upper left control point 2310 of the current block 2300. v_D indicates the motion vector of the adjacent block D 2326 that is adjacent to the upper side of the upper right control point 2311 of the current block 2300. v_E indicates the motion vector of the adjacent block E 2328 that is adjacent to the upper right of the upper right control point 2311 of the current block 2300. v_F indicates the motion vector of the adjacent block F 2330 that is adjacent to the left side of the lower left control point 2312 of the current block 2300. v_G indicates the motion vector of the adjacent block G 2332 that is adjacent to the lower left of the lower left control point 2312 of the current block 2300.
[0272] Step 1.2: Use the top two candidates sorted based on the smaller difference (DV) in the CPMVP candidate combination (refer to Equation 3 below)
[0273] [Equation 3]
[0274] DV = |(v 1x - v 0x ) * h - (v 2y - v 0y ) * w| + |(v 1y - v 0y ) * h + (v 2x - v 0x ) * w|
[0275] v_0x indicates the x-axis element of the motion vector V0 or CPMV0 at the upper left control point 2310 of the current block 2300. v_1x indicates the x-axis element of the motion vector V1 or CPMV1 at the upper right control point 2311 of the current block 2300. v_2x indicates the x-axis element of the motion vector V_2 or CPMV_2 at the lower left control point 2312 of the current block 2300. v_0y indicates the y-axis element of the motion vector V_0 or CPMV_0 at the upper left control point 2310 of the current block 2300. v_1y indicates the y-axis element of the motion vector V1 or CPMV1 at the upper right control point 2311 of the current block 2300. v_2y indicates the y-axis element of the motion vector v_2 or CPMV_2 at the lower left control point 2312 of the current block 2300. W indicates the width of the current block 2300. H indicates the height of the current block 2300.
[0276] Step 2: When the control point motion vector prediction value (CPMVP) for a candidate is less than 2, use the AMVP candidate list
[0277] Step 3: Determine the control point motion vector prediction value (CPMVP) for each of the two candidates, and optimally select the candidate and CPMV with the smaller value by comparing the RD costs
[0278] Step 4: Send the index corresponding to the optimal candidate and the control point motion vector difference (CPMVD)
[0279] In one embodiment of the present disclosure, in AF_INTER, a construction process for CPMVP candidates is provided. Similar to AMVP, the number of candidates is 2, and a concurrent signal notifies the index indicating the position of the candidate list.
[0280] The construction process of the CPMVP candidate list is as follows:
[0281] 1) Scan adjacent blocks to check whether they are encoded as affine motion prediction. If the scanned block is encoded as an affine prediction, derive the motion vector pair of the current block from the affine motion model of the scanned adjacent block until the number of candidates is 2.
[0282] 2) If the number of candidates is less than two, perform the candidate construction process. Additionally, in one embodiment of the present disclosure, a four-parameter (two control points) affine inter-frame mode is used to predict content using a motion model of zooming in / out and rotation. As Figure 16 shown, the affine motion field of the block is described by two control point motion vectors.
[0283] The motion vector field (MVF) of the block is described by Equation 1 described above.
[0284] In the prior art, the Advanced Motion Vector Prediction (AMVP) mode requires signaling of the Motion Vector Prediction (MVP) index and the Motion Vector Difference (MVD). When the AMVP mode is applied in the present disclosure, the affine_flag is signaled to indicate whether affine prediction is used. If affine prediction is applied, the syntax of inter_dir, ref_idx, mvp_index, and two MVDs (mvd_x and mvd_y) is signaled. An affine MVP pair candidate list containing two affine MVP pairs is generated. The signaled mvp_index is used to select one of them. The affine MVP pair is generated from two affine MVP candidates. One is a spatial inherited affine candidate, and the other is a corner derived affine candidate. If the neighboring CU is coded in the affine mode, the spatial inherited affine candidate can be generated. The affine motion model of the neighboring affine coded block is used to generate the motion vectors of two control point MVP pairs. The MVs of the two control point MVP pairs of the spatial inherited affine candidate are derived by using the following formula.
[0285] [Equation 4]
[0286] V 0x = V B0x +(V B2_x - V B0x )*(posCurCU_Y - posRefCU_Y) / RefCU_height+(V B1x - V B0x )*(posCurCU_X - posRefCU_X) / RefCU_width
[0287] [Equation 5]
[0288] V 0y = V B0y +(V B2_y - V B0y )*(posCurCU_Y - posRefCU_Y) / RefCU_height+(V B1y - V B0y )*(posCurCU_X - posRefCU_X) / RefCU_width
[0289] Wherein, V_B0, V_B1, and V_B2 can be replaced by the upper left MV, upper right MV, and lower left MV of any reference / neighboring CU, (posCurCU_X, posCurCU_Y) is the position of the upper left sample of the current CU relative to the upper left sample of the frame, and (posRefCU_X, posRefCU_Y) is the position of the upper left sample of the reference / neighboring CU relative to the upper left sample of the frame.
[0290] [Equation 6]
[0291] V 1x = V B0x + (V B1x - V B0x ) * CU_width / RefCU_width
[0292] [Equation 7]
[0293] V 1y = V B0y + (V B1y - V B0y ) * CU_width / RefCU_width
[0294] Figure 24 shows an example of adjacent blocks for affine motion prediction in the affine inter - frame mode according to an embodiment of the present disclosure.
[0295] Referring to Figure 24 , if the number of MVP pairs is less than 2, angular - derived affine candidates are used. As shown in Figure 24 , adjacent motion vectors are used to derive affine MVP pairs. For the first angular - derived affine candidate, the first available MV in set A (A0, A1, and A2) and the first available MV in set B (B0 and B1) are used to construct the first MVP pair. For the second angular - derived affine candidate, the first available MV in set A and the first available MV in set C (C0 and C1) are used to calculate the MV of the upper - right control point. The first available MV in set A and the calculated MV of the upper - right control point are the second MVP pair.
[0296] In an embodiment of the present disclosure, two (three) candidate sets with two (three) candidates {mv_0, mv_1} ({mv_0, mv_1, mv_2}) are used to predict two (three) control points of the affine motion model. Given motion vector difference vectors mvd_0, mvd_1, mvd_2, the control points can be calculated using the following formula.
[0297] [Equation 8]
[0298]
[0299]
[0300]
[0301] Figure 25 and Figure 26It is a diagram showing a method of deriving motion vector candidates using motion information of adjacent blocks in an affine inter - frame mode according to an embodiment of the present disclosure.
[0302] The affine candidate list is sequentially added by expanding the affine motion from spatially adjacent blocks (extrapolated affine candidates), combinations of motion vectors from spatially adjacent blocks (virtual affine candidates), and HEVC motion vector prediction (MVP) candidates until there are two affine MVPs in the candidate list. The construction of the candidate set is as follows:
[0303] 1. Derive up to two different sets of affine MV prediction values from the affine motion of adjacent blocks. Check adjacent blocks A0, A1, B0, B1, and B2 as Figure 25 shown. If an adjacent block is encoded using an affine motion model and its reference frame is the same as that of the current block, then derive the MVs at two (for a 4 - parameter affine model) or three (for a 6 - parameter affine model) control points of the current block from the affine model of that adjacent item.
[0304] 2. Figure 29 Adjacent blocks for generating the virtual affine candidate set are shown. The adjacent MVs are divided into three groups: S_0 = {mv_A, mv_B, mv_C}, S_1 = {mv_D, mv_E}, and S_2 = {mv_F, mv_G}. mv_0 is the first MV in S0 for the reference picture that is the same as the current block's reference; mv_1 is the first MV in S1 for the same reference picture as the current block; and mv_2 is the first MV in S2 for the same reference picture as the current block.
[0305] If only mv_0 and mv_1 can be found, then mv_2 can be derived by using the following formula.
[0306] [Equation 9]
[0307]
[0308] Referring to Equation 9, the current block size is W×H.
[0309] If only mv_0 and mv_2 can be found, then mv_1 is derived by using the following formula.
[0310] [Equation 10]
[0311]
[0312] In an embodiment of the present disclosure, affine inter - frame prediction can be performed in the following order.
[0313] · Input: Affine motion parameters, reference picture samples
[0314] · Output: Predicted block of the CU
[0315] · Processing
[0316] – Derive the sub - block size of the affine block
[0317] – If both the width and height of the sub - block are greater than 4 luma samples,
[0318] – For each sub - block
[0319] – Derive the motion vector for the sub - block.
[0320] – Invoke DCT - IF - based motion compensation for the sub - block (for luma at 1 / 16 pixel and for chroma at 1 / 32)
[0321] – Otherwise, invoke motion compensation based on an enhanced bilinear interpolation filter for the entire affine block
[0322] In addition, in one embodiment of the present disclosure, when the merge / skip flag is false and both the width and height of the CU are greater than or equal to 8, signal an affine flag at the CU level in the bitstream to indicate whether to use the affine inter - frame mode. And when the CU is encoded in the affine inter - frame mode, signal a model flag to specify whether to use a 4 - parameter affine model or a 6 - parameter affine model for the CU. If the model flag is true, apply the AF_6_INTER mode (6 - parameter affine model) and parse 3 MVDs; otherwise, apply the AF_4_INTER mode (4 - parameter affine model) and parse 2 MVDs.
[0323] In the AF_4_INTER mode, similar to the affine merge mode, construct an affine motion vector pair extrapolated from adjacent blocks encoded in the affine mode and insert it into the candidate list first.
[0324] After that, if the size of the candidate list is less than 4, use adjacent blocks to construct candidates with motion vector pairs \(\{(v_0, v_1)|v0 = \{v_A, v_B, v_c\}, v_1 = \{v_D, v_E\}\}\). As Figure 22 shown, select \(v_0\) from the motion vectors of blocks A, B, or C. Scale the motion vectors from adjacent blocks according to the relationship between the reference list and the POCs of the references for adjacent blocks, the POC of the reference for the current CU, and the POC of the current CU. And the method of selecting \(v_1\) from adjacent blocks D and E is similar. When the candidate list is greater than 4, first sort the candidates according to the consistency of adjacent motion vectors (the similarity of the two motion vectors in a paired candidate) and only retain the top four candidates.
[0325] If the number of candidates in the candidate list is less than 4, the list is padded with motion vector pairs formed by copying each of the AMVP candidates.
[0326] In the AF_6_INTER mode, similar to the affine merge mode, an affine motion vector triple extrapolated from adjacent blocks encoded in the affine mode is constructed and inserted into the candidate list first.
[0327] After that, if the size of the candidate list is less than 4, candidates with a motion vector triple (((v_0, v_1, v_2)|v0 = {v_A, v_B, v_C}, v1 = {v_D, v_E}, v2 = {v_G, v_H}}) are constructed using adjacent blocks. As Figure 22 shown, v_0 is selected from the motion vectors of blocks A, B, or C. The motion vectors from adjacent blocks are scaled according to the relationship between the reference lists and the POCs of the references for adjacent blocks, the POC of the reference for the current CU, and the POC of the current CU. And the method of selecting v_1 from adjacent blocks D and E and v_2 from F and G is similar. When the candidate list is greater than 4, the candidates are sorted first according to the consistency of adjacent motion vectors (the similarity of two motion vectors in the triple candidate), and only the top four candidates are retained.
[0328] If the number of candidates in the candidate list is less than 4, the list is padded with motion vector triples formed by copying each of the AMVP candidates.
[0329] After deriving the CPMV of the current CU, according to the number of affine parameters, the MVF of the current CU is generated according to Equation 11 for the 4-parameter affine model and according to Equation 12 for the 6-parameter affine model.
[0330] [Equation 11]
[0331]
[0332] [Equation 12]
[0333]
[0334] The sub-block size M×N is derived as shown in Equation 13 below, where MvPre is the motion vector fraction accuracy (1 / 16).
[0335] [Equation 13]
[0336]
[0337] After the derivation from Equation 12, if necessary, M and N should be adjusted downward so that they are divisors of w and h, respectively. If M or N is less than 8, WIF is applied; otherwise, sub-block based affine motion compensation is applied.
[0338] Figure 27 An example of a method for deriving an affine motion vector field in units of sub-blocks according to an embodiment of the present disclosure is shown.
[0339] Referring to Figure 27 , in order to derive the motion vector of each M×N sub-block, as Figure 27 shown, the motion vector of the central sample of each sub-block is calculated according to Equation 11 or Equation 12 and rounded to a fractional accuracy of 1 / 16. Then, the SHVC upsampling interpolation filter is applied to generate the prediction of each sub-block using the derived motion vector.
[0340] The SHVC upsampling interpolation filter having the same filter length and normalization factor as the HEVC motion compensation interpolation filter is used as the motion compensation interpolation filter for additional fractional pel positions. The chrominance component motion vector accuracy is 1 / 32 samples, and an additional interpolation filter for the 1 / 32 pel fractional position is derived by using the average value of the filters at two adjacent 1 / 16 pel fractional positions.
[0341] The AF_MERGE mode is selected on the encoder side in a manner similar to performing conventional merge mode selection. First, a candidate list is constructed, and the minimum RD cost within the candidates is selected and compared with the RD costs of other inter-frame modes. The comparison result determines whether to apply AF_MERGE.
[0342] For the AF_4_INTER mode, RD cost check is used to determine which motion vector pair candidate to select as the control point motion vector prediction (CPMVP) of the current CU. After determining the CPMVP of the current affine CU, affine motion estimation is applied and the control point motion vector (CPMV) is found. Then the difference between the CPMV and the CPMVP is determined.
[0343] On the encoder side, the AF_6_INTER mode is verified only when the AF_MERGE or AF_4_INTER mode is selected as the best mode in the previous mode selection stage.
[0344] In one embodiment of the present disclosure, the affine inter - frame (affine AMVP) mode can be executed as follows:
[0345] 1) AFFINE_MERGE_IMPROVE: This improvement attempts to find an adjacent block with the largest coding unit size as an affine merge candidate, rather than finding the first adjacent block in the affine mode.
[0346] 2) AFFINE_AMVP_IMPROVE: Similar to the conventional AMVP process, adjacent blocks in the affine mode are added to the affine AMVP candidate list.
[0347] The detailed construction process of the affine AMVP candidate list is as follows.
[0348] First, check whether the lower - left adjacent block is using an affine motion model and has the same reference index as the current reference index. If not, then check the left adjacent block in the same way. If not, check whether the lower - left adjacent block is using an affine motion model and has a different reference index. If so, add the scaled affine motion vector to the reference picture list. If not, check the left adjacent block in the same way.
[0349] Second, then check the upper - right adjacent block, the upper adjacent block, and the upper - left adjacent block in the same way.
[0350] If we find two candidates after the above process, we will complete the construction of the affine AMVP candidate list. If we do not find two candidates, the original process in the JEM software will be executed to construct the affine AMVP candidate list.
[0351] 3) AFFINE_SIX_PARAM: In addition to the four - parameter affine motion model, a six - parameter affine motion model is added as an additional model.
[0352] The six - parameter affine motion model is derived by using the following formula.
[0353] [Equation 14]
[0354]
[0355] Since there are six parameters in the above motion model, three motion vectors at the upper - left position MV_0, the upper - right position MV_1, and the lower - left position MV_2 are required to determine the model. The three motion vectors are determined in a similar way to the two motion vectors in the four - parameter affine motion model. It should be noted that the affine model merge is always set to the six - parameter affine motion model.
[0356] 4) AFFINE_CLIP_REMOVE: Remove the motion vector constraints for all affine motion vectors. Let the motion compensation process handle the motion vector constraints on its own.
[0357] Affine motion model
[0358] As described above, various affine motion models can be used or considered in affine inter prediction. For example, the affine motion model can represent four motions as Figure 15 shown. An affine motion model that represents three of the motions (translation, scaling, and rotation) among the motions that can be represented by the affine motion model can be called a similarity (or simplified) affine motion model. The number of CPMVs derived according to which affine motion model is used and / or the method of deriving the sample / sub-block unit MV of the current block can be different.
[0359] A motion model can be used. In AF_INTER, in addition to the existing four-parameter motion model in JEM, a six-parameter motion model is proposed. The following Equation 15 describes the six-parameter affine motion model.
[0360] [Equation 15]
[0361] x′ = a * x + b * y + c
[0362] y′ = d * x + e * y + f
[0363] Here, the coefficients a, b, c, d, e, and f are affine motion parameters, and (x, y) and (x′, y′) are the pixel position coordinates before and after the transformation of the affine motion model. To use the affine motion model in video coding, if CPMV0, CPMV1, and CPMV2 are the MVs of CP0 (upper left), CP1 (upper right), and CP2 (lower left), Equation 16 can be described as:
[0364] [Equation 16]
[0365]
[0366] where CPMV_0 = {v_0x, v_0y}, CPMV_1 = {v_1x, v_1y}, CPMV_2 = {v_2x, v_2y}, and w and h are the width and height of the coding block, respectively. Equation 16 describes the motion vector field (MVF) of the block.
[0367] Parse the flag at the CU level to indicate whether to use a four-parameter or six-parameter affine motion model when adjacent blocks are coded as affine prediction. If the adjacent blocks are not coded as affine prediction, skip the flag and use the four-parameter model for affine prediction. In other words, the six-parameter model is considered under the condition that one or more adjacent blocks are coded as an affine motion model. When it comes to the number of CPMVDs, signal two CPMVDs and three CPMVDs for the four-parameter and six-parameter affine motion models respectively.
[0368] In addition, in one embodiment of the present disclosure, pattern-matched motion vector refinement can be used. In the pattern-matched motion vector derivation of JEM (named PMMVD in the JEM encoder description and simply referred to as PMVD in this document), the decoder needs to evaluate multiple motion vector (MV) candidates to determine the starting MV candidate for the CU-level search. In the sub-CU-level search, several MV candidates are added in addition to the best CU-level MV. The decoder needs to evaluate these MV candidates to find the best MV, which requires a large amount of memory bandwidth. In the proposed pattern-matched motion vector refinement (PMVR), the concepts of template matching and bilateral matching in PMVD of JEM are adopted. A PMVR_flag is signaled when the skip mode or merge mode is selected to indicate whether PMVR is enabled. To significantly reduce the memory bandwidth requirement compared to PMVD, an MV candidate list is generated, and if PMVR is applied, the starting MV candidate index is explicitly signaled.
[0369] The candidate list is generated by using the merge candidate list generation process, but sub-CU merge candidates such as affine candidates and ATMVP candidates are excluded. For bilateral matching, only unilateral prediction MV candidates are included. The bilateral prediction MV candidate is divided into two unilateral prediction MV candidates. In addition, similar MV candidates (MV difference less than a predefined threshold) are also removed. For the CU-level search, diamond search MV refinement is performed starting from the signaled MV candidates.
[0370] The sub-CU-level search is only enabled for the bilateral match merge mode. To reduce the memory bandwidth, only the MVs determined from the CU-level search are evaluated. The search window for the sub-CU-level search for all sub-CUs is the same as the search window for the CU-level search. Therefore, no additional bandwidth is required for the sub-CU-level search.
[0371] Template matching is also used to refine the MVP in the AMVP mode. In the AMVP mode, two MVPs are generated by using the HEVC MVP generation process, and a signal is sent to notify an MVP index to select one of them. The selected MVP is refined by using template matching in the PMVR. If adaptive motion vector resolution (AMVR) is applied, the MVP is rounded to the corresponding precision before the template matching refinement. This refinement process is called Pattern-Matched Motion Vector Prediction Refinement (PMVPR). In the remainder of this document, if not otherwise specified, the PMVR includes template matching PMVR, bilateral matching PMVR, and PMVPR.
[0372] To reduce the memory bandwidth requirement, the PMVR is disabled for 4×4, 4×8, and 8×4 CUs. To further reduce the memory bandwidth requirement, the search range of {template matching, bilateral matching} for CU regions equal to 64 is reduced to {±2, ±4}, and the search range of {template matching, bilateral matching} for CU regions greater than 64 is reduced to {±6, ±8}. Compared with the worst case in HEVC, by using all the above methods described in this PMVR section, the required memory bandwidth is reduced from 45.9x in the PMVD of JEM-7.0 to 3.1x in the PMVR.
[0373] Application technology when using affine in non-QT blocks
[0374] Figure 28 A method for generating a prediction block and a motion vector in an inter prediction in which an affine motion model according to an embodiment of the present disclosure has been applied is shown.
[0375] Figure 28 A relational expression for deriving a motion vector in the case of applying an affine motion model is shown. The motion vector can be derived based on Equation 17 below.
[0376] [Equation 17]
[0377]
[0378] In this case, v_x indicates the x-component of the sample unit motion vector of the (x, y) coordinate samples within the current block. v_y represents the y-component of the sample unit motion vector of the (x, y) coordinate samples within the current block. That is, (v_x, v_y) becomes the sample unit motion vector of the (x, y) coordinate samples. In this case, a, b, c, d, e, and f indicate parameters of a relational expression for deriving the sample unit motion vector of the (x, y) coordinate from the control points (CPs) of the current block. The CP can be represented as a control pixel. The parameters can be derived from the motion information of the CPs of each PU sent in units of PUs. The relational expression for deriving the sample unit motion vector from the motion information of the CPs can be applied to each sample of the block, and can be derived based on the relative positions of the x-axis and y-axis of each sample as the positions of the samples within the reference image. The sample unit motion vector can be derived differently according to the size, asymmetry or symmetry, and block position, etc. of the block in the QTBT (TT) block partitioning structure, and its detailed implementation is shown by Figures 29 to 38 shown.
[0379] Figure 29 FIG. is a diagram showing a method of performing motion compensation based on a control point motion vector according to an embodiment of the present disclosure.
[0380] Referring to Figure 29 , the case where the current block is a 2N×2N block is assumed and described. For example, the motion vector of the upper left sample within the current block can be considered as v_0. In addition, using the samples of adjacent blocks adjacent to the current block as CPs, the motion vectors of the CPs can be considered as v_1 and v_2. That is, assuming that the width and height of the current block are S and the coordinates at the upper left sample position of the current block are (xp, yp), the coordinates of CP0 among the CPs can be considered as (xp, yp), the coordinates of CP1 can be considered as (xp + S, yp), and the coordinates of CP2 can be considered as (xp, yp + S). The motion vector of CP0 can be considered as v_0, the motion vector of CP1 can be considered as v_1, and the motion vector of CP2 can be considered as v_2. The sample unit motion vector can be derived using the motion vectors of the CPs. The sample unit motion vector can be derived based on Equation 18 below.
[0381] [Equation 18]
[0382]
[0383] In this case, v_x and v_y respectively indicate the x-component and y-component of the motion vector for the sample at the (x, y) coordinates within the current block. v_x0 and v_y0 respectively indicate the x-component and y-component of the motion vector v_0 for CP0. v_x1 and v_y1 respectively indicate the x-component and y-component of the motion vector v_1 for CP1. v_x2 and v_y2 indicate the x-component and y-component of the motion vector v_2 for CP2. The motion vector of the sample within the current block can be derived based on the relative position within the current block according to a relational expression for deriving the motion vector of the sample unit, such as Equation 18.
[0384] Figure 30 FIG. is a diagram showing a method of performing motion compensation based on motion vectors of control points in an irregular block according to an embodiment of the present disclosure.
[0385] Figure 30 The CPs of a block partitioned into N×2N are shown. The same method as the partition type 2N×2N can be used to drive the relational expression for deriving the motion vector of the sample unit within the current block. During the process of deriving the relational expression, a width value suitable for the shape of the current block can be used. To derive the motion vector of the sample unit, three CPs can be derived. The positions of the CPs can be adjusted as Figure 30 shown. That is, assuming that the width and height of the current block are S / 2 and S, and the coordinates of the current block at the upper left sample position are (xp, yp), the coordinates of CP0 in the CPs can be (xp, yp), the coordinates of CP1 can be (xp + S / 2, yp), and the coordinates of CP2 can be (xp, yp + S). The motion vector of the sample unit can be derived based on Equation 19 below.
[0386] [Equation 19]
[0387]
[0388] In this case, vx and vy respectively indicate the x-component and y-component of the motion vector for the sample at the (x, y) coordinates within the current block. v_x0 and v_y0 respectively indicate the x-component and y-component of the motion vector v_0 for CP0. v_x1 and v_y1 respectively indicate the x-component and y-component of the motion vector v_1 for CP1. v_x2 and v_y2 respectively indicate the x-component and y-component of the motion vector v_2 for CP2. Equation 3 indicates the relational expression for deriving the motion vector of the sample unit where the width of the current block is considered to be S / 2. The motion vector of the sample within the current block partitioned from the CU based on the partition type N×2N can be derived based on the relative position within the current block according to a relational expression for deriving the motion vector of the sample unit, such as Equation 19.
[0389] Figure 31A diagram showing a method of performing motion compensation based on motion vectors of control points in an irregular block according to an embodiment of the present disclosure.
[0390] Figure 31 A block partitioned based on the partition type 2N×N is shown. To derive the sample unit motion vector, three CPs can be derived. As Figure 31 shown by adjusting the positions of the CPs, the height of the current block can be adjusted to S / 2 based on Figure 31 the shape of the current block shown. That is, assuming that the width and height of the current block are S and S / 2 and the coordinates of the current block at the upper left sample position are (xp, yp), the coordinates of CP0 among the CPs can be (xp, yp), the coordinates of CP1 can be (xp + S, yp), and the coordinates of CP2 can be (xp, yp + S / 2). The sample unit motion vector can be derived based on Equation 20 below.
[0391] [Equation 20]
[0392]
[0393] In this case, v_x and v_y respectively indicate the x component and y component of the motion vector for the sample at the (x, y) coordinates within the current block. v_x0 and v_y0 respectively indicate the x component and y component of the motion vector v_0 for CP0. v_x1 and v_y1 respectively indicate the x component and y component of the motion vector v_1 for CP1. v_x2 and v_y2 respectively indicate the x component and y component of the motion vector v_2 for CP2. Equation 4 indicates the relational expression for deriving the sample unit motion vector in which the height of the current block is considered to be S / 2. Based on the relational expression for deriving the sample unit motion vector such as Equation 4.18, the motion vector of each sample within the current block partitioned from the CU based on the partition type 2N×N can be derived based on the relative position within the current block.
[0394] Figures 32 to 38 A diagram showing a method of performing motion compensation based on motion vectors of control points in an irregular block according to an embodiment of the present disclosure.
[0395] Figure 32 The CPs of an asymmetric current block are shown. As Figure 32 shown, the width and height of the asymmetric current block can be considered to be W and H. To derive the sample unit motion vector, three CPs for each current block can be derived. As Figure 32As shown, the coordinates of the CP can be adjusted according to the width and height based on the shape of the current block. That is, assuming that the width and height of the current block are W and H and the coordinates of each current block at the upper left sample position are (xp, yp), the coordinates of CP0 in the CP can be set to (xp, yp), the coordinates of CP1 can be set to (xp + W, yp), and the coordinates of CP2 can be set to (xp, yp + H). In this case, the sample unit motion vector within the current block can be derived based on Equation 21 below.
[0396] [Equation 21]
[0397]
[0398] In this case, v_x and v_y respectively indicate the x-component and y-component of the motion vector for the sample at the (x, y) coordinates within the current block. v_x0 and v_y0 respectively indicate the x-component and y-component of the motion vector v_0 for CP0. v_x1 and v_y1 respectively indicate the x-component and y-component of the motion vector v_1 for CP1. v_x2 and v_y2 indicate the x-component and y-component of the motion vector v_2 for CP2. Equation 21 indicates the relational expression for deriving the sample unit motion vector in which the width and height of the asymmetric current block have been considered.
[0399] Furthermore, according to the present disclosure, in order to reduce the amount of data of the motion information for the CP indicated in units of blocks, motion information prediction candidates for at least one CP can be selected based on the motion information of adjacent blocks or adjacent samples of the current block. The motion information prediction candidates may be referred to as affine motion information candidates or affine motion vector candidates. The affine motion information candidates may include, for example, references to Figures 33 to 38 the disclosed content.
[0400] MVD coding
[0401] The current state-of-the-art video coding standards use motion vectors and their motion vector prediction values to generate a motion vector difference (MVD). The MVD can be more formally defined as the difference between the motion vector and the motion vector prediction value. Similar to the motion vector, the MVD has an x-component and a y-component corresponding to the motion in the x (horizontal) direction and y (vertical) direction. The MVD is an attribute that is only available when encoding an encoding unit using the (advanced) motion vector prediction ((A)MVP) mode.
[0402] Once the MVD is determined, it is then encoded using entropy techniques. Video standards rely on using the MVD as one of their possible methods to exploit redundancy in motion vectors and achieve compression. At the decoder, the motion vector difference (MVD) is decoded before the motion vectors of the coding units are decoded. By encoding the MVD instead of the actual motion vectors, the redundancy between the motion vectors and their predicted values can be exploited, thus improving the compression efficiency.
[0403] The input to the MVD coding stage at the decoder is simply the encoded MVD binary number that has been parsed for decoding. The input to the MVD coding stage at the encoder is the actual MVD value, and additionally a flag ("imv" flag) indicating the resolution for MVD coding. This flag is used to decide whether the MVD should be represented as 1 pixel (or picture element), 4 pixels, or quarter pixels.
[0404] Figure 39 Shows an overall coding structure for deriving motion vectors according to one embodiment of the present disclosure.
[0405] Refer to Figure 39 , first check whether the coding unit is in the merge mode (S3901).
[0406] If the coding unit is in the merge mode, the affine flag and the merge index are parsed for decoding (S3902).
[0407] If the coding unit is not in the merge mode, it will be in the AMVP mode. In the AMVP mode, first the list information is parsed (i.e., whether to use list 0, list 1, or both lists) (S3903).
[0408] Then, the affine flag is parsed (S3904). After that, it is checked whether the parsed affine flag is true or false (S3905).
[0409] If it is true, the parse_MVD_LT and parse_MVD_RT corresponding to the left (LT) and right (RT) MVDs are processed (S3906). If the affine flag is false, the MVD is processed (S3907). The affine motion model in the special case of AMVP will be described in detail below.
[0410] Figure 40 Shows an example of the MVD coding structure according to one embodiment of the present disclosure.
[0411] Refer to Figure 40 , first, the MVD greater than zero flags for the horizontal component (MVDxGT0) and the vertical component (MVDyGT0) are parsed (S4001).
[0412] After that, it is checked whether the parsed data for the horizontal component is greater than zero (i.e., MVDxGT0) (S4002). If the MVDxGT0 flag is true (i.e., MVDxGT0 equals "1"), then the parsed horizontal component is greater than one (i.e., MVDxGT1) (S4002). If MVDxGT0 is not true (i.e., MVDxGT0 equals "0"), then the MVDxGT1 data is not parsed.
[0413] Then, similar steps are performed for the vertical component (S4003 and S4004).
[0414] Thereafter, the parsed MVD data can be further processed in the blocks labeled MVDx_Rem_Level and MVDy_Rem_Level in order to obtain the reconstructed MVD (S4005 and S4006).
[0415] Figure 41 An example of an MVD coding structure according to an embodiment of the present disclosure is shown.
[0416] Figure 41 Shows how the decoder further processes Figure 40 the data in the block MVDx_Rem_Level in to decode the MVDx component. If the decoding flag indicating that the parsed data will be greater than zero (i.e., MVDxGT0) is true (S4101), and the decoding flag indicating that the parsed data will be greater than 1 (i.e., MVDxGT1) is true (S4102), then the exponential Golomb (EG) code of the first order is then used to decode the binary number corresponding to the parsed MVDx component (S4103). The input to the EG will be the binary number containing the absolute value minus two (absolutemin two) (i.e., Abs-2) MVD value and the Golomb of order 1.
[0417] Then the sign information is parsed by decoding the bypass binary number (bypass bin) containing this information (S4104). If the value of the decoded bypass binary number is 1, then a negative sign is appended to the decoded MVDx. However, if the value of the decoded bypass binary number is 0, then the decoded MVD is indicated as a positive value. If MVDxGR0 is true but MVDxGR1 is not true, it indicates that the absolute value of the MVDx being decoded is 1. Then the sign information is parsed and updated. However, if MVDxGR0 is false, then the reconstructed MVDx is 0.
[0418] In the following Figure 42A similar process for decoding MVDy (i.e., MVDy_Rem_Level) at the decoder is shown.
[0419] Figure 42 An example of an MVD coding structure according to an embodiment of the present disclosure is shown.
[0420] Referring to Figure 42 , if the decoding flag indicating that the parsed MVDy is greater than zero (i.e., MVDyGT0) is true (S4201), then the flag MVDyGR1 is checked (S4202).
[0421] If both MVDyGR0 and MVDyGR1 are true, the parsed MVD data is decoded using an EG code with an input of a binary number including absolute value minus 2 (Abs - 2) MVD and order 1 (S4203). Subsequently, the symbol information is parsed and decoded to obtain the decoded MVDy (S4204). If MVDyGR0 is true but MVDyGR1 is false, the absolute vertical value is regarded as +1 or -1. Then, the symbol information is parsed and decoded in a similar manner as above to obtain the decoded MVDy. If the MVDyGR0 flag is false, then MVDy is zero.
[0422] Figure 43 An example of an MVD coding structure according to an embodiment of the present disclosure is shown.
[0423] Referring to Figure 43 , at the encoder, the signed MVD value will be encoded. Similar to Figure 41 , by checking the absolute values of the horizontal and vertical components, binary numbers greater than zero (i.e., MVDxGR0 and MVDyGR0) are encoded for the x - component and y - component (S4301, S4311). Then, the greater - than - one flags (i.e., MVDxGR1 and MVDyGR1) are encoded for the horizontal component and vertical component (S4302, S4312). Thereafter, similar to the decoder, the absolute MVD value is encoded, and then the horizontal component and vertical component are encoded in sequence.
[0424] For horizontal MVD coding, if the absolute horizontal MVD component is greater than zero (i.e., MVDxGR0) and greater than 1 (i.e., MVDxGR1), then the (absolute value - 2) is encoded using a first - order EG code (S4303). Thereafter, the symbol information is encoded using a bypass binary number (S4304). If MVDxGR0 is true and MVDxGR1 is not true, only the symbol information is encoded. If MVDxGR0 is not true, then MVDx is zero. The same process is repeated to encode MVDy (S4313, S4314).
[0425] Affine coding
[0426] Previous video coding standards only considered translational motion models. However, potential motion can include effects such as scaling, rotation, panning, and other irregular motions. To capture the characteristics of such motion, the latest video coding standards have introduced affine motion coding, whereby the irregular characteristics of motion information can be captured using a 4-parameter or 6-parameter affine motion model.
[0427] If a 4-parameter model is used, 2 control points are generated, and if a 6-parameter model is used, 3 control points are used. As previously described Figure 16 illustrates the concept of affine motion more clearly. By using a 4-parameter model, the current block is encoded using two control point motion vectors given by v_0 (cpmv_0) and v1 (cpmv_1).
[0428] Once these control points are derived, the MVF for each of the 4×4 sub-blocks is described by Equation 22 below.
[0429] [Equation 22]
[0430]
[0431] where (v_0x, v_0y) is the motion vector of the upper left control point, and (v_1x, v_1y) is the motion vector of the upper right control point. As previously mentioned, as Figure 27 shown, the motion vector of each 4×4 sub-block is calculated by deriving the motion vector of the central sample of each sub-block.
[0432] Embodiment 1: MVD Precision (MVD PRECISION)
[0433] Affine coding can be used in both merge mode and (A)MVP mode. As described above, in AMVP mode, affine coding can use two or three control points depending on the motion model used. Therefore, there can be two or three motion vector differences (MVDs). In other words, if two control points are used according to the motion model, the MVD for at least one of the upper left (LT) and / or upper right (RT) control points can be encoded. If three control points are used, the MVD for at least one of the upper left (LT), upper right (RT), and / or lower left (LB) control points can be encoded.
[0434] In the decoder, the MVD is decoded before the motion vector in the coding unit is finally determined. In this case, the accuracy of affine prediction (or affine motion prediction) can depend on the accuracy of the control point motion vectors. As a result, the accuracy of affine prediction can depend on the accuracy of MVD coding.
[0435] However, in conventional image compression techniques, if affine prediction is applied, the MVD is encoded only with 1 / 4 pixel (or pel, fraction) accuracy (or precision, resolution).
[0436] In other words, the efficiency of affine coding may greatly depend on the high accuracy of the control point motion vectors and the high accuracy of the motion vectors of the center samples of each subsequent sub-block. In addition, the foregoing relations (e.g., 1, 11, 12, 16, and 22) for deriving motion vectors can provide accuracy far higher than 1 / 16 pixel. For example, if 1 / 16 pixel accuracy is used, the values calculated in the foregoing relations can be rounded to 1 / 16 pixel accuracy. This is useful because a motion compensation interpolation filter operating at 1 / 16 pixel accuracy can be applied to easily generate predicted samples for each sub-block using the derived motion vectors.
[0437] After motion compensation, the motion vectors of each sub-block with high accuracy can be rounded and stored with the same accuracy as ordinary motion vectors. The MVD is calculated based on the difference between the predicted value and the actual motion vector, and the initial calculation can be maintained with 1 / 16 pixel precision. However, in conventional image compression techniques, if affine prediction is applied, the MVD accuracy is reduced to 1 / 4 pixel and encoded. To more accurately decode the motion vectors, if a high accuracy can still be maintained despite the application of affine prediction, the accuracy of affine prediction can be increased, and the compression efficiency can be improved.
[0438] In conventional compression techniques, ordinary MVDs (i.e., MVDs that are not affine predictions) are processed (encoded or transmitted) with 1 / 4 pixel, 1 pixel (i.e., integer pixel), or 4 pixel accuracy. In addition, the encoder / decoder controls this accuracy using a precision flag (or syntax element). However, as described above, in affine prediction, the MVD is stored only with 1 / 4 pixel accuracy. Therefore, the present disclosure proposes a method for increasing the accuracy of the MVD in order to increase the accuracy of affine prediction.
[0439] In the present disclosure, for ease of description, if affine prediction is applied, the MVD can be represented as an affine MVD.
[0440] Figure 44 is a diagram showing a method for deriving affine motion vector difference information according to an embodiment to which the present disclosure is applied.
[0441] Referring to Figure 44 , for ease of description, the decoder is mainly described, but the present disclosure is not limited thereto. The method for signaling affine motion vector difference information can be applied to the encoder in substantially the same manner. In addition, in Figure 44In this case, it is assumed and described that two control points at the upper left and upper right positions are used for affine prediction. However, the present disclosure is not limited thereto, and can be applied substantially in the same manner to the case where three control points at the lower left, upper left, and upper right positions are used for affine prediction.
[0442] The decoder checks whether the merge mode is applied to the current block (S4401). If the merge mode is applied to the current block, the decoder parses the affine flag indicating whether affine prediction is applied to the current block and / or the merge index indicating the candidate applied to the current block within the merge candidate list (S4402).
[0443] The decoder parses the reference list index (or prediction list index) indicating the reference direction (or prediction direction, reference list) of the current block (S4403).
[0444] The decoder parses the affine flag indicating whether affine prediction is applied to the current block (S4404).
[0445] The decoder checks whether affine prediction is applied to the current block based on the affine flag value (S4405).
[0446] If affine prediction is not applied to the current block, the decoder parses the MVD of the current block (S4406).
[0447] In one embodiment of the present disclosure, if affine prediction is applied to the current block, the decoder may parse the precision flag (or precision index) and perform a check process on the precision.
[0448] Specifically, if affine prediction is applied to the current block, the decoder parses the MVD precision flag (S4407). In this case, the MVD precision flag (or affine MVD precision flag) indicates whether the adaptive affine MVD precision mode is applied. In one embodiment, if the adaptive affine MVD precision mode is applied, the affine MVD may be derived with another precision other than the predefined default (or basic) precision. If the adaptive affine MVD precision mode is applied, the affine MVD may be derived with the predefined default precision. In one embodiment, the predefined default precision may be 1 / 4 pixel precision. Another precision other than the predefined default precision may include at least one of integer pixel, 4 pixels, 1 / 8 pixel, and / or 1 / 16 pixel precision.
[0449] The decoder checks whether to apply the adaptive affine MVD precision mode based on the MVD precision flag value (S4408). If the adaptive affine MVD precision mode is applied, the decoder derives the MVDs for two control points with a precision other than the default precision (S4409). In one embodiment, if the adaptive affine MVD precision mode is applied, i.e., if a precision other than the default precision is applied, the encoder may send a syntax element indicating a specific precision among preset precisions to the decoder.
[0450] If the adaptive affine MVD precision mode is not applied, the decoder derives the MVDs for two control points with the default precision (S4410).
[0451] In one embodiment, the precision for the affine MVD may be signaled via a bitstream. To this end, the encoder may signal a higher-level syntax element to the decoder. For example, the higher-level syntax element may be signaled via a sequence parameter set, a picture parameter set, a slice header (or a tile group header), etc. Further, for example, the encoder may generate set_affine_MVD_precision_flag and may signal set_affine_MVD_precision_flag to the decoder. In this case, set_affine_MVD_precision_flag represents the higher-level syntax element indicating the precision of the affine MVD.
[0452] For example, set_affine_MVD_precision_flag may indicate whether the precision of the affine MVD is a predefined default (or base) precision (e.g., 1 / 4 pixel precision). If the predefined default precision is not applied, set_affine_MVD_precision_flag may include additional precision information. The additional precision information may be signaled from the encoder to the decoder. That is, the encoder may send a syntax element indicating whether the precision of the affine MVD is the predefined default precision (e.g., 1 / 4 pixel precision) to the decoder. If the precision of the affine MVD is not the predefined default precision, the encoder may send a syntax element indicating the detailed precision of the affine MVD to the decoder. For example, the detailed precision of the affine MVD may include at least one of integer pixel, 1 / 4 pixel, 1 / 8 pixel, or 1 / 16 pixel precision.
[0453] Alternatively, for example, the syntax element may indicate whether to send the affine MVD with a higher precision.
[0454] In one embodiment, the position of the syntax header can be summarized as high_level_parameter_set() according to Table 2. Additionally, in one embodiment, auxiliary syntax elements can be used as syntax elements (index or flag) for indicating detailed precision.
[0455] [Table 2]
[0456]
[0457] In Table 2, when set_affine_MVD_precision_flag is 1, this can indicate that set_affine_MVD_precision_flag exists in the slice header of non-IDR pictures of the coded video sequence (CVS). Additionally, when set_affine_MVD_precision_flag is 0, this can indicate that set_affine_precision_flag does not exist in the slice header and the adaptive affine MVD according to this embodiment is not used in the CVS.
[0458] Furthermore, in one embodiment, syntax elements for indicating detailed precision information can be signaled additionally. For example, a syntax structure according to Table 3 can be defined.
[0459] [Table 3]
[0460]
[0461] In Table 3, slice_affine_mvd_precision_idx represents a syntax element for indicating the specific (detailed) precision of the affine MVD. In the present disclosure, the name of slice_affine_mvd_precision_idx is not restricted, and a syntax element for indicating the specific precision of the affine MVD can be represented as a flag. Additionally, in Table 3, it is assumed that a syntax element for indicating the specific precision of the affine MVD is included in the slice segment header, but the present disclosure is not limited thereto. The slice segment header can include syntax at various levels. For example, a syntax element for indicating the specific (detailed) precision of the affine MVD can be included in the coded tree unit syntax or the coded unit syntax.
[0462] In one embodiment, when slice_affine_mvd_precision_idx is 0, this can indicate the default MVD precision of 1 / 4 pixel. Similarly, the index value 1 can indicate the MVD precision of 1 / 8 pixel, and the index value 2 can indicate the MVD precision of 1 / 16 pixel.
[0463] Embodiment 2: Entropy and GOLOMB parameters
[0464] One embodiment of the present disclosure proposes a method of using features in which the MVD statistic is changed. Specifically, the MVD statistic of an inter-coded block depends on the motion model for translational motion only. However, the statistics of the affine MVD are different from those of a normal AMVP block because the prediction unit (or coded block or coding unit) coded in the affine mode uses an affine motion model in which various other motions are considered. This means that the same entropy coding method and / or parameters should not be generally used to code the MVDs of all blocks.
[0465] As referred to above Figures 41 to 43 In the conventional image compression technique, when the absolute values in the horizontal and vertical directions of the MVD are greater than 1, decoding is performed using an exponential Golomb code of order 1. The exponential Golomb code can be very effective when representing the number of similar patterns or a set of numbers without restricting the maximum number that can be represented.
[0466] The degree of the exponential Golomb code (which may hereinafter be referred to as the Golomb degree) includes the possibility of the symbol occurring. In the conventional image compression technique, the degree 1 is used regardless of the distribution of the MVD values. However, in the case of affine motion, the same method need not be maintained. Therefore, the present disclosure proposes an exponential Golomb code having a degree depending on the range of the affine MVD value. In one embodiment, the encoder / decoder may use the same method as the method of Figure 45 to select the division of the range of the MVD, but the present disclosure is not limited thereto. Histogram analysis can be very useful in determining the range of the absolute value of the MVD. The most frequent values can be grouped, and each sub-region (or range) of the MVD can be coded using another Golomb degree.
[0467] This can be very effective because similar MVD values can be coded using the same degree. Specifically, in affine motion, the control points on the left and right can have a very close relationship. The encoder / decoder can use the statistics of any one control point to determine the most likely region (or range) of the other control point, and can select various Golomb degrees based on the most likely region (or range).
[0468] One embodiment of the present disclosure proposes an entropy coding method that depends on the unique statistic of a motion model when performing entropy coding for MVD, rather than a constant entropy coding method. This will be described with reference to the accompanying drawings.
[0469] Figure 45 FIG. is a diagram showing the coding structure of a motion vector difference according to an embodiment of the present disclosure.
[0470] Referring to Figure 45 , for ease of description, the decoder is mainly described, but the present disclosure is not limited thereto. The method of signaling affine motion vector difference information can be applied to the encoder in substantially the same manner.
[0471] In one embodiment of the present disclosure, when the MVD value is greater than 0, the decoder can divide the MVD greater than 0 based on a given integer N value without limitation, such as conventional MVDxGR1 and MVDyGR1. In addition, N can be determined based on the distribution of the MVD value.
[0472] Specifically, the decoder checks the syntax elements (flags) MVDxGR_0 and MVDyGR_0 that indicate whether the MVD value is greater than 0 (S4501, S4511). When the MVDxGR_0 and / or MVDyGR_0 value is 0, the MVD value in each direction (horizontal or vertical direction) is considered to be 0.
[0473] When the values of MVDxGR_0 and MVDyGR_0 are 1, the decoder checks the MVDxGR_N and MVDyGR_N syntax elements (flags) (S4502, S4512). When the MVDxGR_N and / or MVDyGR_N value is 1, the decoder uses Golomb series k1 (i.e., series 1) to decode (or parse) the MVD value in each direction based on the exponential Golomb code of the input with an absolute value of -N - 1 (Abs - N - 1) (S4503, S4513).
[0474] When the MVDxGR_N and / or MVDyGR_N value is 0, the decoder decodes (or parses) the MVD value in each direction using an exponential Golomb code based on another series other than Golomb series k1 (S4504, S4514). In one embodiment, the exponential Golomb binarization of Golomb series k2 (i.e., series 2) can be used to encode / decode the corresponding absolute value greater than 0 and less than or equal to N.
[0475] The decoder decodes (or parses) the sign of the MVD in each direction (S4505, S4515).
[0476] In addition, in one embodiment, the encoder / decoder may apply different binarizations to each of the parts divided into 0 and N. For example, the encoder / decoder may use the exponential Golomb code to encode the absolute value greater than 0 and less than N, and may use truncated binary (TB) (or truncated unary binarization) to encode the absolute value greater than N.
[0477] Embodiment 3: Precision control of MVD and entropy and GOLOMB parameters
[0478] One embodiment of the present disclosure proposes a method of combining two embodiments (Embodiment 1 and Embodiment 2). In other words, one embodiment of the present disclosure may include features of combining the above two embodiments. Specifically, one embodiment of the present disclosure proposes a method of integrating precision information for MVD and entropy coding.
[0479] Figure 46 is a diagram showing a method of deriving an affine motion vector based on precision information according to one embodiment of the present disclosure.
[0480] Referring to Figure 46 , for ease of description, the decoder is mainly described, but the present disclosure is not limited thereto. The method of signaling the affine motion vector difference information may be applied to the encoder in substantially the same manner.
[0481] If the precision control function has been activated, the decoder parses the syntax element (S4601) indicating a specific precision. In Figure 46 , the syntax element is represented as a precision index, but is not limited to this name.
[0482] The decoder parses the MVD value in the horizontal / vertical direction based on the precision checked in step S4601 (S4602).
[0483] In one embodiment, the precision index may indicate a high precision such as 1 / 16 pixel or 1 / 8 pixel, and may indicate a low precision such as an integer pixel or 4 pixels. For example, if the syntax element (e.g., set_affine_MVD_precision_flag) indicating whether to apply the adaptive affine precision mode is true, the decoder may additionally check the syntax element (e.g., slice_affine_mvd_precision_idx) indicating a specific precision. The decoder may determine the precision of the MVD encoded based on the syntax element indicating a specific precision. In addition, the decoder may parse the MVD information in the horizontal / vertical direction based on the determined precision.
[0484] In one embodiment, when decoding the MVD based on the determined precision, the method described in Embodiment 2 may be applied. If high precision is applied, when the MVD value in the horizontal and / or vertical direction is greater than 0, the decoder may parse MVDx_GR_N and / or MVDy_GR_N. As described above, the decoder may apply the first binarization when the absolute value is greater than N, and may apply the second binarization (or binarization method) when the absolute value is less than or equal to N. For example, the decoder may use the exponential Golomb code of order 1 as the first binarization, and may use the truncated binary (TB) (or truncated unary binarization) as the second binarization. If low precision (e.g., 1 / 4, 1, or 4 pixel precision) is applied, the decoder may use a third binarization to perform MVD decoding. For example, the decoder may use the truncated unary binarization as the third binarization.
[0485] Figure 47 is a diagram showing the coding structure of the motion vector difference according to an embodiment to which the present disclosure is applied.
[0486] Refer to Figure 47 , for ease of description, the decoder is mainly described, but the present disclosure is not limited thereto. The method of signaling the affine motion vector difference information may be applied to the encoder in substantially the same manner.
[0487] In one embodiment of the present disclosure, when the MVD value is greater than 0, the decoder may divide the MVD value greater than 0 based on a given integer N value without being restricted as in the conventional MVDxGR1 and MVDyGR1. In addition, N may be determined based on the distribution of the MVD values.
[0488] Specifically, the decoder checks the syntax elements (flags) MVDxGR_0 and MVDyGR_0 indicating whether the MVD value is a syntax element greater than 0 (S4701, S4711). When the MVDxGR_0 and / or MVDyGR_0 value is 0, the MVD value in each direction (horizontal and / or vertical direction) is considered 0.
[0489] When the values of MVDxGR_0 and MVDyGR_0 are 1, the decoder checks whether the MVD precision of the current block is higher than a predefined precision (S4702, S4711). For example, the predefined precision may be 1 pixel, 1 / 4 pixel, or 1 / 8 pixel precision.
[0490] When the current MVD precision is higher than the predefined precision, the decoder checks the MVDxGR_N and MVDyGR_N syntax elements (flags) (S4703, S4713).
[0491] When the MVDxGR_N and / or MVDyGR_N value is 1, the decoder decodes (or parses) the MVD value in each direction using the first binarization (or binarization method) (S4704, S4714). For example, the first binarization may be the exponential Golomb code method with a Golomb order of k1 (i.e., order 1). That is, the decoder may use the Golomb order k1 to decode (or parse) the MVD value in each direction based on the exponential Golomb code with the absolute value -N (Abs-N) as the input.
[0492] When the MVDxGR_N and / or MVDyGR_N value is 0, the decoder decodes (or parses) the MVD value in each direction using the second binarization (S4705, S4715). For example, the second binarization may be the exponential Golomb code using another order other than the Golomb order k1, and may be truncated binary (TB) (or truncated unary binarization).
[0493] When the current MVD precision is less than or equal to a predefined precision, the decoder decodes (or parses) the MVD value in each direction using the third binarization (S4706, S4716). For example, the third binarization may be the exponential Golomb code using another order other than the Golomb order k1, and may be truncated binary (TB) (or truncated unary binarization).
[0494] The decoder decodes (or parses) the sign of the MVD in each direction (S4707, S4717).
[0495] For ease of description, the foregoing embodiments of the present disclosure have been divided and described, but the present disclosure is not limited thereto. That is, the described embodiments 1 to 3 may be executed independently, and one or more embodiments may be combined and executed.
[0496] Figure 48 is a flowchart showing a method for generating an inter prediction block based on affine prediction according to an embodiment of the present disclosure.
[0497] Referring to Figure 48 , for ease of description, the decoder is mainly described, but the present disclosure is not limited thereto. The method for generating an inter prediction block according to an embodiment of the present disclosure may be executed in the same manner in both the encoder and the decoder.
[0498] The decoder checks whether affine prediction (or affine motion prediction) is applied to the current block (S4801).
[0499] If, as a result of the check, affine prediction is applied, the decoder obtains at least one syntax element indicating the resolution (or precision or accuracy) of the motion vector difference for affine prediction (S4802).
[0500] The decoder derives the control point motion vector of the current block based on the at least one syntax element (S4803).
[0501] The decoder derives the motion vector of each of the plurality of sub - blocks included in the current block based on the control point motion vector (S4804).
[0502] The decoder uses the motion vector of each sub - block to generate the prediction sample of the current block (S4805).
[0503] As described above, step S4802 may include the step of obtaining a first syntax element indicating whether the resolution of the motion vector difference is a preset default resolution, and, if the resolution of the motion vector difference is not the default resolution, obtaining a second syntax element indicating the resolution of the motion vector difference among the remaining resolutions other than the default resolution.
[0504] In addition, as described above, the default resolution may be preset to 1 / 4 pixel precision.
[0505] In addition, as described above, each of the remaining resolutions may include at least one of integer pixel precision, 1 / 4 pixel precision, 1 / 8 pixel precision, or 1 / 16 pixel precision.
[0506] In addition, as described above, step S4804 may further include the steps of: determining the resolution of the motion vector difference using at least one syntax element; and obtaining the motion vector difference based on the resolution of the motion vector difference.
[0507] In addition, as described above, the step of obtaining the motion vector difference may further include the step of obtaining a flag indicating whether the motion vector difference is greater than 0, and, when the motion vector difference is greater than 0, obtaining a flag indicating whether the motion vector difference is greater than a predefined specific value.
[0508] In addition, as described above, when the motion vector difference is greater than 0 and less than or equal to the predefined specific value, the motion vector difference may be binarized using an exponential Golomb code of order 1. When the motion vector difference is greater than the predefined specific value, the motion vector difference may be binarized using a truncated binary method.
[0509] Figure 49 FIG. is a diagram illustrating an inter - frame prediction apparatus based on affine prediction according to an embodiment to which the present disclosure is applied.
[0510] InFigure 49 In order to facilitate the description, the inter-frame prediction unit is shown as a block, but the inter-frame prediction unit may be implemented as a structure included in an encoder and / or a decoder.
[0511] Referring to Figure 49 , the inter-frame prediction unit implements Figures 8 to 48 the functions, processes, and / or methods proposed in
[0512] The affine prediction mode recognition unit 4901 checks whether affine prediction is applied to the current block.
[0513] If, as a result of the check, affine prediction is applied, the syntax element acquisition unit 4902 acquires at least one syntax element indicating the resolution of the motion vector difference for affine prediction.
[0514] The control point motion vector derivation unit 4903 derives the control point motion vector of the current block based on the at least one syntax element.
[0515] The sub-block motion vector derivation unit 4904 derives the motion vector of each of the plurality of sub-blocks included in the current block based on the control point motion vector.
[0516] The prediction sample generation unit 4905 generates a prediction sample of the current block using the motion vector of each sub-block.
[0517] As described above, the syntax element acquisition unit 4902 may acquire a first syntax element indicating whether the resolution of the motion vector difference is a preset default resolution, and if the resolution of the motion vector difference is not the default resolution, may acquire a second syntax element indicating the resolution of the motion vector difference among the remaining resolutions other than the default resolution.
[0518] In addition, as described above, the default resolution is preset to 1 / 4 pixel accuracy.
[0519] In addition, as described above, each of the remaining resolutions may include at least one of integer pixel accuracy, 4 pixel accuracy, 1 / 8 pixel accuracy, or 1 / 16 pixel accuracy.
[0520] In addition, as described above, the control point motion vector derivation unit 4903 may use the at least one syntax element to determine the resolution of the motion vector difference, and may obtain the motion vector difference based on the resolution of the motion vector difference.
[0521] In addition, as described above, the control point motion vector derivation unit 4903 can obtain a flag indicating whether the motion vector difference is greater than 0, and when the motion vector difference is greater than 0, can obtain a flag indicating whether the motion vector difference is greater than a predefined specific value.
[0522] In addition, as described above, when the motion vector difference is greater than 0 and less than or equal to a predefined specific value, the exponential Golomb code of order 1 can be used to binaryize the motion vector difference. When the motion vector difference is greater than the predefined specific value, the truncated binaryization method can be used to binaryize the motion vector difference.
[0523] Figure 50 A video coding system to which the present disclosure is applied is shown.
[0524] The video coding system may include a source device and a receiving device. The source device may forward the encoded video / image information or data to the receiving device in a file or stream format via a digital storage medium or a network.
[0525] The source device may include a video source, an encoding device, and a transmitter. The receiving device may include a receiver, a decoding device, and a renderer. The encoding device may be referred to as a video / image encoding device, and the decoding device may be referred to as a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display unit, and the display unit may be configured as an independent device or an external component.
[0526] The video source may obtain video / images through processes such as capture, synthesis, or generation. The video source may include a video / image capture device and / or a video / image generation device. For example, the video / image capture device may include one or more cameras and a video / image archive including previously captured video / images, etc. The video / image generation device may include, for example, a computer, a tablet, and a smart phone, and may (electrically) generate video / images. For example, virtual video / images may be generated by a computer, and in this case, the video / image capture process may be replaced by a process of generating relevant data.
[0527] The encoding device may encode the input video / image. The encoding device may perform a series of processes including prediction, transformation, quantization, etc. for compression and encoding efficiency.
[0528] The transmitter can forward the encoded video / image information or data output in bitstream format to the receiver of the receiving device via a digital storage medium or network in file or stream format. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD. The transmitter can include elements for generating a media file in a predetermined file format and can include elements for transmitting via a broadcast / communication network. The receiver can extract the bitstream and forward it to a decoding device.
[0529] The decoding device can perform a series of processes including dequantization, inverse transformation, and prediction corresponding to the operations of the encoding device, and decode the video / image.
[0530] The renderer can render the decoded video / image. The rendered video / image can be displayed via a display unit.
[0531] Figure 51 is a configuration diagram of a content streaming system to which one embodiment of the present disclosure is applied.
[0532] Referring to Figure 51 , the content streaming system to which the present disclosure is applied can include an encoding server, a streaming server, a Web server, a media storage device, a user device, and a multimedia input device.
[0533] The encoding server is used to compress the content input from a multimedia input device such as a smartphone, a camera, and a camcorder into digital data, generate a bitstream, and send the bitstream to the streaming server. As another example, when a multimedia input device such as a smartphone, a camera, and a camcorder directly generates a bitstream, the encoding server can be omitted.
[0534] The bitstream can be generated by applying the encoding method or bitstream generation method of the present disclosure, and the streaming server can temporarily store the bitstream during the process of sending or receiving the bitstream.
[0535] The streaming server sends multimedia data to the user device based on a user request via the Web server, and the Web server serves as a medium for notifying the user of the service. When the user sends a request for a desired service to the Web server, the Web server passes the request to the streaming server, and the streaming server sends the multimedia data to the user. Here, the content streaming system can include an additional control server, and in this case, the control server is used to control commands / responses between devices in the content streaming system.
[0536] The streaming server can receive content from a media storage device and / or an encoding server. For example, when receiving content from the encoding server, the streaming server can receive the content in real time. In this case, the streaming server can store the bitstream for a predetermined time to facilitate the provision of a smooth streaming service.
[0537] Examples of user equipment may include cellular phones, smartphones, laptop computers, digital broadcast terminals, PDAs (Personal Digital Assistants), PMPs (Portable Multimedia Players), navigation devices, slate computers, tablet computers, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, and HMDs (Head-Mounted Displays)), digital TVs, desktop computers, and digital signage, etc.
[0538] Each server in the content streaming system can operate as a distributed server, and in this case, the data received by each server can be processed in a distributed manner.
[0539] The embodiments described in this disclosure can be implemented and executed on a processor, microprocessor, controller, or chip. For example, the functional units shown in the drawings can be implemented and executed on a computer, processor, microprocessor, controller, or chip.
[0540] In addition, the decoder and encoder to which this disclosure is applied can be included in a multimedia broadcast transmission and reception device, a mobile communication terminal, a home theater video device, a digital movie video device, a camera for surveillance, a video conferencing device, a real-time communication device such as video communication, a mobile streaming device, a storage medium, a portable camera, a video-on-demand (VoD) service providing device, an over-the-top (OTT) video device, an Internet streaming service providing device, a three-dimensional (3D) video device, a video phone device, and a medical video device, and can be used to process video signals or data signals. For example, an OTT video device can include a game console, a Blu-ray player, an Internet-connected TV, a home theater system, a smartphone, a tablet computer, and a digital video recorder (DVR).
[0541] In addition, the processing method applying the present disclosure can be generated in the form of a program executed by a computer, and can be stored in a computer-readable recording medium. Multimedia data having a data structure according to the present disclosure can also be stored in a computer-readable recording medium. The computer-readable recording medium includes all types of storage devices in which computer-readable data is stored. For example, the computer-readable recording medium can include Blu-ray Disc (BD), Universal Serial Bus (USB), ROM, PROM, EPROM, EEPROM, RAM, CD-ROM, magnetic tape, floppy disk, and optical data storage device. In addition, the computer-readable recording medium includes a medium implemented in the form of a carrier (e.g., transmission via the Internet). In addition, the bit stream generated using an encoding method can be stored in a computer-readable recording medium, or can be transmitted over wired and wireless communication networks.
[0542] In addition, an embodiment of the present disclosure can be implemented as a computer program product using program code. The program code according to an embodiment of the present disclosure can be executed by a computer. The program code can be stored on a computer-readable carrier.
[0543] In the foregoing embodiments, the elements and features of the present disclosure have been combined in a specific form. Unless otherwise explicitly described, each element or feature can be regarded as optional. Each element or feature can be implemented in a form not combined with other elements or features. In addition, some elements and / or features can be combined to form an embodiment of the present disclosure. The operation order described in the embodiments of the present disclosure can be changed. Some elements or features of one embodiment can be included in another embodiment, or can be replaced by corresponding elements or features of another embodiment. Obviously, embodiments can be constructed by combining claims that do not have an explicit reference relationship in the claims, or can be included as new claims by amendment after the application is filed.
[0544] Embodiments according to the present disclosure can be implemented in various ways such as, for example, hardware, firmware, software, or a combination thereof. In the case of implementation by hardware, one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), processors, controllers, microcontrollers, and microprocessors, etc. can be used to implement embodiments of the present disclosure.
[0545] In the case of implementation by firmware or software, embodiments of the present disclosure can be implemented in the form of modules, procedures, or functions for performing the above functions or operations. The software code can be stored in a memory and driven by a processor. The memory can be located inside or outside the processor, and can exchange data with the processor in various known ways.
[0546] It will be apparent to those skilled in the art that the present disclosure may be implemented in other specific forms without departing from the basic characteristics thereof. Therefore, the detailed description should not be construed as restrictive, but should be construed as illustrative in all aspects. The scope of the present disclosure should be determined by a reasonable analysis of the appended claims, and all changes within the equivalent scope of the present disclosure are included within the scope of the present disclosure.
[0547] [Industrial Applicability]
[0548] The above preferred embodiments of the present disclosure have been disclosed for illustrative purposes, and those skilled in the art may make improvements, changes, substitutions or additions to various other embodiments without departing from the technical spirit and scope disclosed in the appended claims.
Claims
1. A method for processing a video signal using affine prediction, the method comprising the following steps: Obtaining a first syntax element indicating whether the affine prediction is applied to a current block; Obtaining a second syntax element indicating whether an adaptive motion vector difference precision mode is applied; Based on the first syntax element indicating that the affine prediction is applied to the current block and the second syntax element indicating that the adaptive motion vector difference precision mode is applied, obtaining a third syntax element indicating whether a precision of a motion vector difference of a control point motion vector CPMV used in the affine prediction is a predefined precision; Based on the third syntax element, obtaining a fourth syntax element including information about the precision of the motion vector difference, wherein the fourth syntax element indicates at least one precision among a plurality of different precisions of the motion vector difference; Deriving a control point motion vector of the current block based on at least one of the third syntax element and the fourth syntax element; Deriving a motion vector of each of a plurality of sub-blocks included in the current block based on the control point motion vector; and Generating a prediction sample of the current block based on the motion vector of each of the sub-blocks, wherein the at least one precision of the motion vector difference of the CPMV used in the affine prediction includes 1 / 16 pixel precision, wherein when the motion vector difference is greater than a predefined specific value, a truncated binary quantization method is used to binaryize the motion vector difference, and wherein when a predefined precision including 1 / 4 pixel precision, 1 pixel precision, or 4 pixel precision is applied, a truncated binary quantization method is used to binaryize the motion vector difference.
2. The method according to claim 1, Among them, The predefined precision is preset to 1 / 4 pixel precision.
3. The method according to claim 2, Among them, Each of the remaining precisions other than the predefined precision includes integer pixel precision, 4 pixel precision, or 1 / 8 pixel precision.
4. The method according to claim 1, Among them, The step of deriving the control point motion vector includes the following steps: Obtaining the motion vector difference based on the precision of the motion vector difference.
5. The method according to claim 4, Among them, The step of obtaining the motion vector difference includes the following steps: Obtaining a flag including information about whether the motion vector difference is greater than 0; and When the motion vector difference is greater than 0, obtaining a flag including information about whether the motion vector difference is greater than the predefined specific value.
6. The method according to claim 5, Among them, When the motion vector difference is greater than 0 and less than or equal to the predefined specific value, a first-order exponential Golomb code is used to binaryize the motion vector difference.
7. A method for encoding a video signal using affine prediction, the method comprising the following steps: Determining whether the affine prediction is applied to a current block; Determining whether an adaptive motion vector difference precision mode is applied; Determine whether the precision of the motion vector difference of the control point motion vector (CPMV) is a predefined precision based on the affine prediction being applied to the current block and the adaptive motion vector difference precision mode being applied; Determine at least one precision among multiple different pixel precisions of the motion vector difference based on information indicating whether the precision of the motion vector difference is the predefined precision; Derive the control point motion vector of the current block based on at least one of the information indicating whether the precision of the motion vector difference is the predefined precision and the information about the at least one precision; Derive the motion vector of each of the multiple sub-blocks included in the current block based on the control point motion vector; Generate a prediction sample of the current block based on the motion vector of each sub-block; Generate a first syntax element indicating whether the affine prediction is applied to the current block; Generate a second syntax element indicating whether the adaptive motion vector difference precision mode is applied; Generate a third syntax element indicating whether the precision of the motion vector difference is the predefined precision based on the first syntax element indicating that the affine prediction is applied to the current block and the second syntax element indicating that the adaptive motion vector difference precision mode is applied; And Generate a fourth syntax element indicating the at least one precision of the motion vector difference based on the third syntax element, wherein the at least one precision of the motion vector difference of the CPMV used in the affine prediction includes 1 / 16 pixel precision, wherein when the motion vector difference is greater than a predefined specific value, the motion vector difference is binary-coded using a truncated binary coding method, and wherein when applying a predefined precision including 1 / 4 pixel precision, 1 pixel precision, or 4 pixel precision, the motion vector difference is binary-coded using a truncated binary coding method.
8. A computer-readable storage medium storing encoded picture data, the encoded picture data being generated by performing the following steps: Determine whether affine prediction is applied to a current block; Determine whether the adaptive motion vector difference precision mode is applied; Determine whether the precision of the motion vector difference of the control point motion vector (CPMV) is a predefined precision based on the affine prediction being applied to the current block and the adaptive motion vector difference precision mode being applied; Determine at least one precision among multiple different pixel precisions of the motion vector difference based on information indicating whether the precision of the motion vector difference is the predefined precision; Derive the control point motion vector of the current block based on at least one of the information indicating whether the precision of the motion vector difference is the predefined precision and the information about the at least one precision; Derive the motion vector of each of the multiple sub-blocks included in the current block based on the control point motion vector; Generate a prediction sample of the current block based on the motion vector of each sub-block; Generate a first syntax element indicating whether the affine prediction is applied to the current block; Generate a second syntax element indicating whether to apply the adaptive motion vector difference precision mode; Based on the first syntax element indicating that the affine prediction is applied to the current block and the second syntax element indicating that the adaptive motion vector difference precision mode is applied, generate a third syntax element indicating whether the precision of the motion vector difference is the predefined precision; And Based on the third syntax element, generate a fourth syntax element indicating at least one precision of the motion vector difference, wherein at least one precision of the motion vector difference of the CPMV used in the affine prediction includes 1 / 16 pixel precision, wherein when the motion vector difference is greater than a predefined specific value, the motion vector difference is binary-coded using a truncated binary coding method, and wherein when applying a predefined precision including 1 / 4 pixel precision, 1 pixel precision, or 4 pixel precision, the motion vector difference is binary-coded using a truncated binary coding method.
9. A method for transmitting data, the data including a bitstream related to a video signal, the transmitting method comprising the steps of: Obtain the bitstream related to the video signal; and Transmit the data including the bitstream, wherein the bitstream is generated by performing the following steps: Determine whether affine prediction is applied to a current block; Determine whether to apply the adaptive motion vector difference precision mode; Based on the affine prediction being applied to the current block and the adaptive motion vector difference precision mode being applied, determine whether the precision of the motion vector difference of the control point motion vector CPMV is the predefined precision; Based on the information indicating whether the precision of the motion vector difference is the predefined precision, determine at least one precision among multiple different pixel precisions of the motion vector difference; Based on at least one of the information indicating whether the precision of the motion vector difference is the predefined precision and the information about the at least one precision, derive the control point motion vector of the current block; Based on the control point motion vector, derive the motion vector of each of the multiple sub-blocks included in the current block; Generate a prediction sample of the current block based on the motion vector of each sub-block; Generate a first syntax element indicating whether the affine prediction is applied to the current block; Generate a second syntax element indicating whether to apply the adaptive motion vector difference precision mode; Based on the first syntax element indicating that the affine prediction is applied to the current block and the second syntax element indicating that the adaptive motion vector difference precision mode is applied, generate a third syntax element indicating whether the precision of the motion vector difference is the predefined precision; and Based on the third syntax element, generate a fourth syntax element indicating at least one precision of the motion vector difference, wherein at least one precision of the motion vector difference of the CPMV used in the affine prediction includes 1 / 16 pixel precision, Among them, when the motion vector difference is greater than a predefined specific value, the motion vector difference is binary-coded using a truncated binary coding method, and Among them, when applying a predefined precision including 1 / 4 pixel precision, 1 pixel precision, or 4 pixel precision, the motion vector difference is binary-coded using a truncated binary coding method.
Citation Information
Patent Citations
Adaptive motion vector precision for video coding
US20180098089A1
AMVR-based image coding method and apparatus in image coding system
US20180270485A1
AMVR-based image coding method and apparatus in image coding system
WO2017052009A1
Cited By
Video signal decoding and encoding method and storage medium
CN120614463A