Decoding equipment, encoding equipment, and data transmission equipment
By updating the motion information of the target block during image encoding, the problem of transmission and storage costs for high-resolution images is solved, and encoding efficiency and image quality are improved.
Patent Information
- Application Number
- CN202310703397.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2016-12-05
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2036-12-05
AI Technical Summary
The increased cost of transmitting and storing high-resolution and high-quality images, the low efficiency of existing image coding techniques, and the insufficient accuracy of motion information in inter-frame prediction all contribute to low coding efficiency.
The motion information update process is optimized by calculating the corrected motion information after the target block is decoded, and updating the motion information of the target block based on the corrected motion information. Inter-frame prediction and encoding are performed using an entropy decoder, predictor, and memory.
It improves the overall efficiency of image coding, reduces the distortion propagation of motion information between neighboring blocks, and enhances coding quality.
Smart Images

Figure CN116527884B_ABST
Abstract
Description
[0001] This application is a divisional application of the original invention patent application No. 201680091905.1 (International Application No.: PCT / KR2016 / 014167, Application Date: December 5, 2016, Invention Title: Method and Apparatus for Decoding Images in an Image Encoding System). Technical Field
[0002] This invention relates to techniques for image encoding, and more specifically, to a method and apparatus for decoding images in an image encoding system. Background Technology
[0003] The demand for high-resolution, high-quality images (e.g., HD (High Definition) and UHD (Ultra-High Definition) images) is constantly increasing across various fields. Because image data is high-resolution and high-quality, the amount of information or bits that needs to be transmitted increases compared to traditional image data. Therefore, transmission and storage costs increase when using media such as traditional wired / wireless broadband lines to transmit image data or when using existing storage media to store image data.
[0004] Therefore, an efficient image compression technology is needed to effectively send, store, and reproduce information from high-resolution and high-quality images. Summary of the Invention
[0005] Technical issues
[0006] This invention provides a method and apparatus for improving image coding efficiency.
[0007] The present invention also provides a method and apparatus for inter-frame prediction, which updates the motion information of a target block.
[0008] The present invention also provides a method and apparatus for calculating corrected motion information of a target block after the decoding process of the target block and updating it based on the corrected motion information.
[0009] The present invention also provides a method and apparatus for using updated motion information of the target block for motion information of the next block adjacent to the target block.
[0010] Technical solution
[0011] In one aspect, an image decoding method executed by a decoding device is provided. This image decoding method includes the following steps: obtaining inter-frame prediction information about a target block via a bitstream; deriving motion information of the target block based on the inter-frame prediction information; deriving prediction samples by performing inter-frame prediction on the target block based on the motion information; generating a reconstructed block based on the prediction samples; deriving corrected motion information of the target block based on the reconstructed block; and updating the motion information of the target block based on the corrected motion information.
[0012] On the other hand, a decoding apparatus for performing image decoding is provided. The decoding apparatus includes: an entropy decoder configured to obtain inter-frame prediction information about a target block via a bitstream; a predictor configured to derive motion information of the target block based on the inter-frame prediction information, derive prediction samples by performing inter-frame prediction for the target block based on the motion information, generate a reconstructed block of the target block based on the prediction samples, and derive corrected motion information of the target block based on the reconstructed block; and a memory configured to update the motion information of the target block based on the corrected motion information.
[0013] On the other hand, a video coding method performed by an encoding device is provided. The method includes the following steps: generating motion information of a target block; deriving prediction samples by performing inter-frame prediction for the target block based on the motion information; generating a reconstructed block based on the prediction samples; generating corrected motion information of the target block based on the reconstructed block; and updating the motion information of the target block based on the corrected motion information.
[0014] In another aspect, a video encoding apparatus is provided. The encoding apparatus includes: a prediction unit configured to generate motion information of a target block, derive prediction samples by performing inter-frame prediction of the target block based on the motion information, generate a reconstructed block based on the prediction samples, and generate corrected motion information of the target block based on the reconstructed block; and a memory configured to update the motion information of the target block based on the corrected motion information.
[0015] Beneficial effects
[0016] According to the present invention, after the decoding process of the target block, the corrected motion information of the target block is calculated and updated to more accurate motion information, thereby improving the overall coding efficiency.
[0017] According to the present invention, the motion information of the next block adjacent to the target block can be derived based on the updated motion information of the target block, and the propagation of distortion can be reduced, thereby improving the overall coding efficiency. Attached Figure Description
[0018] Figure 1 This is a schematic diagram illustrating the configuration of a video encoding apparatus to which this disclosure applies.
[0019] Figure 2 This is a schematic diagram illustrating the configuration of a video decoding apparatus to which this disclosure applies.
[0020] Figure 3 Examples are shown for performing inter-frame prediction based on unidirectional motion information and for performing inter-frame prediction based on bidirectional motion information applied to the target block.
[0021] Figure 4 An example of the coding process, including a method for updating the motion information of the target block, is shown.
[0022] Figure 5 An example of the decoding process, including methods for updating motion information of the target block, is shown.
[0023] Figure 6 This illustrates an example of a method by which a decoding device updates the motion information of a target block using a block matching method.
[0024] Figure 7 The reference image is shown and can be used in the corrected motion information.
[0025] Figure 8 This example shows a list of potential merge blocks that are adjacent to the target block when the target block's motion information is updated to the target block's corrected motion information.
[0026] Figure 9 This example shows a list of potential merge blocks that are adjacent to the target block when the target block's motion information is updated to the target block's corrected motion information.
[0027] Figure 10 This example shows a list of motion vector predictors for the next block adjacent to the target block when the motion information of the target block is updated based on the corrected motion information of the target block.
[0028] Figure 11 This example shows a list of motion vector predictors for the next block adjacent to the target block when the motion information of the target block is updated based on the corrected motion information of the target block.
[0029] Figure 12 A video encoding method according to the encoding device of the present invention is illustrated schematically.
[0030] Figure 13 A video decoding method according to the decoding device of the present invention is illustrated schematically. Detailed Implementation
[0031] This disclosure may be modified in various forms, and specific embodiments thereof will be described and illustrated in the accompanying drawings. However, these embodiments are not intended to limit this disclosure. The terminology used in the following description is for the purpose of describing particular embodiments only and is not intended to limit this disclosure. Singular expressions include plural expressions, provided that they are interpreted differently. Terms such as “comprising” and “having” are intended to indicate the presence of the features, quantities, steps, operations, elements, components, or combinations thereof used in the following description, and therefore it should be understood that the possibility of having or adding one or more different features, quantities, steps, operations, elements, components, or combinations thereof is not excluded.
[0032] On the other hand, the elements in the accompanying drawings described in this disclosure are drawn independently for the purpose of illustrating different specific functions, and are not intended to imply that the elements are specifically implemented by independent hardware or independent software. For example, two or more elements may be combined to form a single element, or a single element may be divided into multiple elements. Embodiments in which elements are combined and / or divided are part of this disclosure without departing from its concept.
[0033] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the accompanying drawings. Furthermore, throughout the drawings, similar reference numerals are used to indicate similar elements, and identical descriptions of similar elements will be omitted.
[0034] In this specification, generally, a frame refers to a unit of an image representing a specific time, and a slice is a unit that constitutes a part of a frame. A frame may consist of multiple slices, and the terms frame and slice may be mixed together when necessary.
[0035] A pixel can refer to the smallest unit that makes up a picture (or image). Additionally, "sample" can be used as the corresponding term. A sample can typically represent a pixel or a pixel value; it can represent a pixel containing only the luminance component (pixel value) or a pixel containing only the chrominance component (pixel value).
[0036] A unit refers to the basic unit of image processing. A unit may include a specific region and at least one of the information associated with that region. Optionally, a unit may be combined with terms such as block, region, etc. Typically, an M×N block may represent a set of samples or transform coefficients arranged in M columns and N rows.
[0037] Figure 1 The structure of the video encoding apparatus to which this disclosure applies is briefly illustrated.
[0038] Reference Figure 1 The video encoding device 100 includes a screen splitter 105, a predictor 110, a subtractor 115, a transformer 120, a quantizer 125, a rearranger 130, an entropy encoder 135, an inverse quantizer 140, an inverse transformer 145, an adder 150, a filter 255, and a memory 160.
[0039] The screen splitter 105 can split the input screen into at least one processing unit. Here, the processing unit can be a coding unit (CU), a prediction unit (PU), or a transform unit (TU). A coding unit is a block of encoded units, and the maximum coding unit (LCU) can be split into deeper coding units according to a quadtree structure. In this case, the maximum coding unit can be used as the final coding unit, or the coding unit can be recursively split into deeper coding units as needed, and the coding unit with the optimal size can be used as the final coding unit based on coding efficiency according to video characteristics. When a minimum coding unit (SCU) is set, the coding unit cannot be split into coding units smaller than the minimum coding unit. Here, the final coding unit refers to the coding unit that has been segmented or split into a predictor or a transform. A prediction unit is a block segmented from a coding unit block and can be a block of sample prediction units. Here, the prediction unit can be divided into sub-blocks. A transform block can be split from a coding unit block according to a quadtree structure and can be a block of units for deriving transform coefficients and / or a block of units for deriving residual signals from transform coefficients.
[0040] Hereinafter, the coding unit may be referred to as the coding block (CB), the prediction unit may be referred to as the prediction block (PB), and the transform unit may be referred to as the transform block (TB).
[0041] A prediction block or prediction unit may refer to a specific region in the image that has a block shape, and may include an array of prediction samples. Similarly, a transform block or transform unit may refer to a specific region in the image that has a block shape, and may include an array of transform coefficients or residual samples.
[0042] Predictor 110 can perform predictions on the target block (hereinafter, the current block) and can generate a prediction block that includes prediction samples of the current block. The unit of prediction performed in predictor 110 can be a coded block, a transform block, or a prediction block.
[0043] Predictor 110 can determine whether to apply intra-frame prediction or inter-frame prediction to the current block. For example, predictor 110 can determine whether to apply intra-frame prediction or inter-frame prediction on a CU-by-CU basis.
[0044] In the case of intra-frame prediction, predictor 110 may derive the prediction sample for the current block based on reference samples outside the current block in the frame to which the current block belongs (hereinafter, the current frame). In this case, predictor 110 may derive the prediction sample based on the average or interpolation of the neighboring reference samples of the current block (case (i)), or may derive the prediction sample based on reference samples among the neighboring reference samples of the current block that indicate the existence of the prediction sample in a specific (prediction) direction (case (ii)). Case (i) may be referred to as a non-directional mode or a non-angular mode, and case (ii) may be referred to as a directional mode or an angular mode. In intra-frame prediction, as an example, the prediction modes may include 33 directional modes and at least two non-directional modes. Non-directional modes may include DC mode and planar mode. Predictor 110 may use the prediction modes applied to neighboring blocks to determine the prediction mode to be applied to the current block.
[0045] In the case of inter-frame prediction, predictor 110 can derive the predicted sample for the current block based on samples specified by motion vectors on a reference frame. Predictor 110 can derive the predicted sample for the current block by applying any of the following modes: skip mode, merge mode, and motion vector prediction (MVP). In skip mode and merge mode, predictor 110 can use motion information from neighboring blocks as motion information for the current block. In skip mode, unlike in merge mode, the difference (residual) between the predicted sample and the original sample is not sent. In MVP mode, the motion vectors of neighboring blocks are used as motion vector predictors, and therefore used as motion vector predictors for the current block to derive the motion vector for the current block.
[0046] In the case of inter-frame prediction, neighboring blocks can include spatially neighboring blocks existing in the current frame and temporally neighboring blocks existing in the reference frame. The reference frame, which includes temporally neighboring blocks, can also be referred to as a col-pic. Motion information can include motion vectors and reference frame indices. Information such as prediction mode information and motion information can be (entropy-encoded) and then output as a bitstream.
[0047] When using motion information from time-proximity blocks in skip and merge modes, the highest frame in the reference frame list can be used as the reference frame. Reference frames included in the reference frame list are aligned based on the frame order count (POC) difference between the current frame and the corresponding reference frame. The POC corresponds to the display order and is distinguishable from the encoding order.
[0048] Subtractor 115 generates a residual sample as the difference between the original sample and the predicted sample. When the skip mode is applied, the residual sample may not be generated as described above.
[0049] Transformer 120 transforms residual samples on a block-by-block basis to generate transform coefficients. Transformer 120 can perform the transform based on the size of the corresponding transform block and the prediction mode applied to the coded or prediction blocks that spatially overlap with the transform block. For example, when intra-frame prediction is applied to the coded or prediction blocks that overlap with the transform block and the transform block is a 4×4 residual array, the residual samples can be transformed using Discrete Sine Transform (DST); otherwise, Discrete Cosine Transform (DCT) is used.
[0050] Quantizer 125 can quantize the transform coefficients to generate quantized transform coefficients.
[0051] Rearranger 130 rearranges the quantized transform coefficients. Rearranger 130 can rearrange the block-form quantized transform coefficients into a one-dimensional vector using a coefficient sweep method. Although rearranger 130 is described as a separate component, it can be part of quantizer 125.
[0052] The entropy encoder 135 can perform entropy coding on the quantized transform coefficients. Entropy coding can include coding methods such as Golomb, Context Adaptive Variable Length Coding (CAVLC), and Context Adaptive Binary Arithmetic Coding (CABAC). In addition to the quantized transform coefficients, the entropy encoder 135 can encode information required for video reconstruction (e.g., syntactic element values) together or separately. The entropy-coded information can be transmitted or stored in bitstream form at the Network Abstraction Layer (NAL).
[0053] Inverse quantizer 140 performs inverse quantization on the value (transformation coefficient) quantized by quantizer 125, and inverse transformer 145 performs inverse transformation on the value inverse quantized by inverse quantizer 140 to generate residual samples.
[0054] Adder 150 adds the residual samples to the predicted samples to reconstruct the image. The residual samples can be added to the predicted samples in blocks to generate reconstructed blocks. Although adder 150 is described as a separate component, adder 150 can be part of predictor 110.
[0055] Filter 155 can apply deblocking filtering and / or adaptive sample offset to the reconstructed image. Artifacts at block boundaries or distortions in quantization in the reconstructed image can be corrected by deblocking filtering and / or adaptive sample offset. Adaptive sample offset can be applied on a sample-by-sample basis after deblocking filtering is completed. Filter 155 can apply an adaptive loop filter (ALF) to the reconstructed image. ALF can be applied to the reconstructed image after deblocking filtering and / or adaptive sample offset has been applied.
[0056] Memory 160 can store reconstructed frames or information required for encoding / decoding. Here, the reconstructed frame can be a frame reconstructed by filtering 155. The stored reconstructed frame can be used as a reference frame for (inter-frame) prediction of other frames. For example, memory 160 can store (reference) frames for inter-frame prediction. Here, the frame used for inter-frame prediction can be specified according to a set or list of reference frames.
[0057] Figure 2 The structure of the video decoding apparatus to which this disclosure applies is briefly illustrated.
[0058] Reference Figure 2 The video decoding device 200 includes an entropy decoder 210, a rearranger 220, an inverse quantizer 230, an inverse transformer 240, a predictor 250, an adder 260, a filter 270, and a memory 280.
[0059] When a bitstream containing video information is input, the video decoding device 200 can reconstruct the video by associating it with the process of processing the video information in the video encoding device.
[0060] For example, the video decoding apparatus 200 can use the processing unit applied in the video encoding apparatus to perform video decoding. Therefore, the processing unit block for video decoding can be an encoding unit block, a prediction unit block, or a transform unit block. As a decoding unit block, the encoding unit block can be split from the largest encoding unit block according to a quadtree structure. As a block split from the encoding unit block, the prediction unit block can be a sample prediction unit block. In this case, the prediction unit block can be divided into sub-blocks. As an encoding unit block, the transform unit block can be split according to a quadtree structure and can be a unit block used to derive transform coefficients or a unit block used to derive the residual signal from the transform coefficients.
[0061] The entropy decoder 210 can parse a bitstream to output information required for video reconstruction or picture reconstruction. For example, the entropy decoder 210 can decode information in the bitstream based on encoding methods such as exponential Golomb coding, CAVLC, CABAC, etc., and can output the values of the syntax elements required for video reconstruction as well as the quantized values of the transform coefficients with respect to the residuals.
[0062] More specifically, the CABAC entropy decoding method receives bins corresponding to each syntactic element in the bitstream, uses information about the target syntactic element and the decoding information of neighboring and target blocks, or information about symbols / bins decoded in previous steps, to determine a context model. Based on the determined context model, it predicts the bin generation probability and performs arithmetic decoding of the bins to generate symbols corresponding to each syntactic element value. Here, the CABAC entropy decoding method can update the context model after determining it using information about symbols / bins decoded for the next symbol / bin.
[0063] Information about prediction from the information decoded in the entropy decoder 210 can be provided to the predictor 250, and the residual values (i.e., quantized transform coefficients) of the entropy decoder 210 after entropy decoding can be input to the rearranger 220.
[0064] Rearranger 220 can rearrange the quantized transform coefficients into a two-dimensional block form. Rearranger 220 can perform rearrangements corresponding to the coefficient scans performed by the encoding device. Although rearranger 220 is described as a separate component, rearranger 220 can be part of quantizer 230.
[0065] The inverse quantizer 230 can inverse quantize the quantized transform coefficients based on the (inverse)quantization parameters to output transform coefficients. In this case, the information used to derive the quantization parameters can be signaled from the encoding device.
[0066] The inverse transformer 240 can perform an inverse transformation on the transformation coefficients to derive the residual samples.
[0067] Predictor 250 can perform predictions on the current block and generate a prediction block that includes prediction samples of the current block. The unit of prediction performed in predictor 250 can be a coded block, a transform block, or a prediction block.
[0068] Predictor 250 can determine whether to apply intra-frame prediction or inter-frame prediction based on information about the prediction. In this case, the unit used to determine which one to use between intra-frame and inter-frame prediction may differ from the unit used to generate the prediction samples. Furthermore, the units used to generate the prediction samples may also differ between inter-frame and intra-frame prediction. For example, which one to use between inter-frame and intra-frame prediction can be determined using CU units. Additionally, for example, in inter-frame prediction, prediction samples can be generated by determining the prediction mode in PU units, while in intra-frame prediction, prediction samples can be generated in TU units by determining the prediction mode in PU units.
[0069] In the case of intra-frame prediction, predictor 250 can derive the prediction sample for the current block based on neighboring reference samples in the current frame. Predictor 250 can derive the prediction sample for the current block by applying either a directional or non-directional mode based on neighboring reference samples. In this case, the intra-frame prediction mode of neighboring blocks can be used to determine the prediction mode to be applied to the current block.
[0070] In the case of inter-frame prediction, predictor 250 can derive prediction samples for the current block based on samples specified in the reference frame according to motion vectors. Predictor 250 can use one of skip mode, merge mode, and MVP mode to derive prediction samples for the current block. Here, motion information (e.g., motion vectors and information about the reference frame index) required for inter-frame prediction of the current block provided by the video coding device can be obtained or derived based on information about the prediction.
[0071] In skip and merge modes, motion information from neighboring blocks can be used as motion information for the current block. Here, neighboring blocks can include spatially neighboring blocks and temporally neighboring blocks.
[0072] Predictor 250 can construct a merge candidate list using motion information of available neighboring blocks and use the information indicated by the merge index on the merge candidate list as the motion vector of the current block. The merge index can be signaled by the encoding device. The motion information may include motion vectors and reference frames. When using motion information of temporally neighboring blocks in skip mode and merge mode, the highest frame in the reference frame list can be used as the reference frame.
[0073] In the skip mode, unlike the merge mode, the difference (residual) between the predicted sample and the original sample is not sent.
[0074] In the MVP mode, the motion vectors of neighboring blocks can be used as motion vector predictors to derive the motion vector of the current block. Here, neighboring blocks can include spatially neighboring blocks and temporally neighboring blocks.
[0075] When a merge mode is applied, for example, a merge candidate list can be generated using the motion vectors of reconstructed spatially neighboring blocks and / or the motion vectors corresponding to Col blocks that are temporally neighboring blocks. In merge mode, the motion vectors of candidate blocks selected from the merge candidate list are used as the motion vector of the current block. The aforementioned information about the prediction may include a merge index indicating the candidate block with the best motion vector selected from the candidate blocks included in the merge candidate list. Here, the predictor 250 can use the merge index to derive the motion vector of the current block.
[0076] When applying the MVP (Motion Vector Prediction) mode as another example, a motion vector predictor candidate list can be generated using the motion vectors of the reconstructed spatially neighboring blocks and / or the motion vectors corresponding to the Col blocks, which are temporally neighboring blocks. That is, the motion vectors of the reconstructed spatially neighboring blocks and / or the motion vectors corresponding to the Col blocks, which are temporally neighboring blocks, can be used as motion vector candidates. The aforementioned prediction information may include a predicted motion vector index indicating the best motion vector selected from the motion vector candidates included in the list. Here, predictor 250 can use the motion vector index to select the predicted motion vector for the current block from the motion vector candidates included in the motion vector candidate list. The predictor of the encoding device can obtain the motion vector difference (MVD) between the motion vector of the current block and the motion vector predictor, encode the MVD, and output the encoded MVD in the form of a bitstream. That is, the MVD can be obtained by subtracting the motion vector predictor from the motion vector of the current block. Here, predictor 250 can obtain the motion vector included in the prediction information and derive the motion vector of the current block by adding the motion vector difference to the motion vector predictor. Additionally, the predictor can obtain or derive a reference frame index indicating a reference frame from the aforementioned prediction information.
[0077] Adder 260 can reconstruct the current block or current frame by adding the residual sample to the predicted sample. Adder 260 can reconstruct the current frame by adding the residual sample to the predicted sample on a block-by-block basis. When skip mode is applied, the residual is not sent, so the predicted sample becomes the reconstructed sample. Although adder 260 is described as a separate component, adder 260 can be part of predictor 250.
[0078] Filter 270 can apply deblocking filtering, adaptive sample offset, and / or ALF to the reconstructed image. Here, adaptive sample offset can be applied on a sample-by-sample basis after deblocking filtering. ALF can be applied after deblocking filtering and / or applying adaptive sample offset.
[0079] Memory 280 can store reconstructed frames or information required for decoding. Here, the reconstructed frames can be frames reconstructed by filter 270. For example, memory 280 can store frames used for inter-frame prediction. Here, frames used for inter-frame prediction can be specified according to a set or list of reference frames. The reconstructed frames can be used as reference frames for other frames. Memory 280 can output the reconstructed frames in the output order.
[0080] As described above, when performing inter-frame prediction on a target block, the motion information of the target block can be generated, encoded, and output by applying skip mode, merge mode, or Adaptive Motion Vector Prediction (AMVP) mode. In this case, due to the block-by-block encoding process, distortion can be computed and included in the motion information of the target block, thus potentially not perfectly reflecting the motion information of the reconstructed block representing the target block. Specifically, when the merge mode is applied to the target block, the accuracy of the target block's motion information may degrade. That is, there may be a significant difference between the predicted block derived from the target block's motion information and the reconstructed block of the target block. In this case, the motion information of the target block is used in the decoding process of the next block adjacent to the target block, and distortion may propagate, thereby degrading the overall coding efficiency.
[0081] Therefore, this invention proposes a method in which, after the decoding process of the target block, the corrected motion information of the target block is calculated based on the derived reconstructed block, and the motion information of the target block is updated based on the corrected motion information, so that the next block adjacent to the target block can derive more accurate motion information. This improves the overall coding efficiency.
[0082] Figure 3 Examples are shown for performing inter-frame prediction based on unidirectional motion information and for performing inter-frame prediction based on bidirectional motion information applied to the target block. Bidirectional motion information may include an L0 reference frame index and an L0 motion vector, and an L1 reference frame index and an L1 motion vector. Unidirectional motion information may include an L0 reference frame index and an L0 motion vector or an L1 reference frame index and an L1 motion vector. L0 indicates the reference frame list L0 (list 0), and L1 indicates the reference frame list L1 (list 1). During image encoding, methods for inter-frame prediction may include deriving motion information through motion estimation and motion compensation. Figure 3 As shown, motion estimation can be indicated during the process of deriving a block that matches the target block for a reference frame encoded prior to the encoding of the target frame, which includes the target block. The block that matches the target block can be defined as a reference block, and the positional difference between the target block and the reference block, derived by assuming that the target frame includes the same reference block as the target block, can be defined as the motion vector of the target block. Motion compensation can include a unidirectional method that derives and uses one reference block, and a bidirectional method that derives and uses two reference blocks.
[0083] Information including the motion vector of the target block and information about the reference frame can be defined as motion information of the target block. The reference frame information may include a list of reference frames and a reference frame index indicating the reference frames included in the list. The encoding device may store the motion information of the target block for the next block or the frame to be encoded immediately following the target block after it has been encoded. The stored motion information may be used in methods representing the motion information of the next block or the frame to be encoded immediately following the target block. These methods may include: a merge mode, which indexes the motion information of the target block and sends the motion information of the next block adjacent to the target block as an index; and an AMVP mode, which uses only the difference between the motion vector of the next block and the motion vector of the target block to represent the motion vector of the next block adjacent to the target block.
[0084] This invention proposes a method for updating motion information of a target frame or target block after the encoding process. Motion information used in inter-frame prediction can be stored when inter-frame prediction and encoding are performed on a target block. However, the motion information may include distortions that occur during the process calculated using block matching methods, and because the motion information may be a value selected for a target block rate-distortion (RD) optimization process, it may not perfectly reflect the actual motion of the target block. Since the motion information of the target block can be used not only in the inter-frame prediction process of the target block, but also affects the frame encoded after the encoding of the next target frame adjacent to the target block (including the frame encoded after the target block) and the target frame including both the next block and the target block, the overall coding efficiency may be degraded.
[0085] Figure 4 An example of an encoding process including a method for updating motion information of a target block is shown. The encoding device encodes the target block (step S400). The encoding device can derive prediction samples by performing inter-frame prediction on the target block and generate a reconstructed block of the target block based on the prediction samples. The encoding device can derive motion information of the target block used for performing inter-frame prediction and generate inter-frame prediction information of the target block including this motion information. The motion information may be referred to as first motion information.
[0086] After the encoding process of the target block is performed, the encoding device calculates the motion information corrected for the reconstructed block based on the target block (step S410). The encoding device can calculate the corrected motion information of the target block using various methods. The corrected motion information can be referred to as the second motion information. At least one method, including direct methods such as optical flow (OF) methods, block matching methods, frequency domain methods, etc., and indirect methods such as singularity matching methods, methods using statistical properties, etc., can be applied to this method. In addition, direct methods and indirect methods can be applied simultaneously. The details of block matching methods and OF methods will be described below.
[0087] Due to the nature of motion information calculation for target blocks, encoding devices may find it difficult to calculate more accurate motion information using only the encoded information of the target blocks. Therefore, for example, the encoding device may calculate the corrected motion information after executing the encoding process for the target frame that includes the target blocks, rather than after the encoding process for the target blocks themselves. That is, the encoding device can execute the encoding process for the target frame and calculate the corrected motion information based on more information compared to immediately following the encoding process for the target blocks.
[0088] Furthermore, the encoding device can also generate motion vector update difference information indicating the difference between the existing motion vector included in the (first) motion information and the corrected motion information, and encode and output it. The motion vector update difference information can be sent in units of PU.
[0089] The encoding device determines whether to update the motion information of the target block (step S420). The encoding device can determine whether to update the (first) motion information by comparing the accuracy between the (first) motion information and the corrected motion information. For example, the encoding device can determine whether to update the (first) motion information by using the amount of difference between the image derived using each motion information through motion compensation and the original image. In other words, the encoding device can determine whether to update by comparing the amount of data of the residual signal between the reference block derived based on each motion information and the original block of the target block.
[0090] In step S420, if it is determined that the motion information of the target block needs to be updated, the encoding device updates the motion information of the target block based on the corrected motion information and stores the corrected motion information (step S430). For example, if the amount of data in the residual signal between a specific reference block and the original block derived based on the corrected motion information is less than the amount of data in the residual signal between the reference block and the original block derived based on the (first) motion information, the encoding device may update the (first) motion information based on the corrected motion information. In this case, the encoding device can update the motion information of the target block by replacing the (first) motion information with the corrected motion information and store the updated motion information that includes only the corrected motion information. Alternatively, the encoding device can update the motion information of the target block by adding the corrected motion information to the (first) motion information and store the updated motion information that includes both the (first) motion information and the corrected motion information.
[0091] Furthermore, in step S420, if it is determined that the motion information of the target block will not be updated, the encoding device stores the (first) motion information (step S440). For example, if the amount of data of the residual signal between a specific reference block and the original block derived based on the corrected motion information is not less than the amount of data of the residual signal between the reference block and the original block derived based on the (first) motion information, the encoding device may store the (first) motion information.
[0092] On the other hand, although not shown, the coding device may update the motion information of the target block based on the modified motion information, in the case where the modified motion information is derived without determining the update by comparing the (first) motion information used in inter-frame prediction with the modified motion information.
[0093] Furthermore, since the decoding device cannot be used with the original image, it can receive additional information indicating whether the target block has been updated. That is, the encoding device can generate and encode this additional information indicating whether an update has been made, and output it via a bitstream. For example, this additional information indicating whether an update has been made can be called an update flag. A 1 for the update block indicates that the motion information has been updated, while a 0 for the update block indicates that the motion information has not been updated. For example, the update flag can be sent in units of PU. Alternatively, the update flag can be sent in units of CU, CTU, or slice, and can be sent at a higher level, such as units of Picture Parameter Set (PPS) or Sequence Parameter Set (SPS).
[0094] Furthermore, the decoding device can determine whether to update based on the reconstructed block of the target block by comparing the (first) motion information of the target block with the corrected motion information, without receiving an update flag. If it is determined that the motion information of the target block has been updated, the decoding device can update the motion information of the target block based on the corrected motion information and store the updated motion information. Alternatively, if the update is determined by deriving the corrected motion information without comparing the motion information used in inter-frame prediction with the corrected motion information, the decoding device can update the motion information of the target block based on the corrected motion information. In this case, the decoding device can update the motion information of the target block by replacing the (first) motion information with the corrected motion information and store the updated motion information including only the corrected motion information. Alternatively, the decoding device can update the motion information of the target block by adding the corrected motion information to the (first) motion information and store the updated motion information including both the (first) motion information and the corrected motion information.
[0095] When the motion information update process is performed after the encoding process of the target block, the encoding process of the next block adjacent to the target block in the next encoding order can be performed.
[0096] Figure 5An example of a decoding process, including a method for updating motion information of a target block, is shown. The decoding process can be performed in a manner similar to the encoding process described above. The decoding device decodes the target block (step S500). When inter-frame prediction is applied to the target block, the decoding device can obtain inter-frame prediction information for the target block from the bitstream. The decoding device can deduce motion information of the target block based on the information used for inter-frame prediction and deduce prediction samples by performing inter-frame prediction on the target block. The motion information can be referred to as first motion information. The decoding device can generate a reconstructed block of the target block based on the prediction samples.
[0097] The decoding device calculates the corrected motion information of the target block (step S510). The decoding device can calculate the corrected motion information using various methods. The corrected motion information can be referred to as the second motion information. At least one method, including direct methods such as optical flow (OF) methods, block matching methods, and frequency domain methods, and indirect methods such as singularity matching methods and methods using statistical properties, can be applied to this method. In addition, direct and indirect methods can be applied simultaneously. The details of block matching methods and OF methods will be described below.
[0098] Due to the nature of motion information calculation for the target block, the decoding device may find it difficult to calculate more accurate motion information using only the information decoded from the target block. Therefore, for example, the decoding device may calculate the corrected motion information after performing the decoding process for the target image that includes the target block, rather than after the decoding process for the target block itself. That is, the decoding device can perform the decoding process for the target image and calculate the corrected motion information based on more information compared to immediately following the decoding process for the target block.
[0099] Furthermore, the decoding device can also obtain motion vector update difference information via the bit stream, which generates the difference between the existing motion vector included in the (first) motion information and the corrected motion information. The motion vector update difference information can be transmitted in units of PU. In this case, the decoding device can derive the corrected motion information by summing the (first) motion information and the obtained motion vector update difference information, instead of independently calculating the corrected motion information using the aforementioned direct and indirect methods. That is, the decoding device can derive the existing motion vector using the motion vector predictor (MVP) of the corrected motion information and derive the corrected motion vector by adding the motion vector update difference information to the existing motion vector.
[0100] The decoding device determines whether to update the motion information of the target block (step S520). The decoding device can determine whether to update the (first) motion information by comparing the accuracy between the (first) motion information and the corrected motion information. For example, the decoding device can determine whether to update the (first) motion information by using the amount of difference between the reference block derived using each motion information through motion compensation and the reconstructed block of the target block. In other words, the encoding device can determine whether to update by comparing the amount of data of the residual signal between the reference block derived based on each motion information and the original block of the target block.
[0101] In addition, the decoding device receives supplementary information from the encoding device indicating whether an update has been made, and determines whether to update the motion information of the target block based on this supplementary information. For example, the supplementary information indicating whether an update has been made can be called an update flag. An update block value of 1 indicates that the motion information has been updated, while an update block value of 0 indicates that the motion information has not been updated. For example, the update flag can be sent in units of PU.
[0102] In step S520, if it is determined that the motion information of the target block needs to be updated, the decoding device updates the motion information of the target block based on the corrected motion information and stores the updated motion information (step S530). For example, if the amount of data in the residual signal between a specific reference block and a reconstructed block derived based on the corrected motion information is less than the amount of data in the residual signal between the reference block and the reconstructed block derived based on the (first) motion information, the decoding device may update the (first) motion information based on the corrected motion information. In this case, the decoding device can update the motion information of the target block by replacing the (first) motion information with the corrected motion information and store the updated motion information that includes only the corrected motion information. Alternatively, the decoding device can update the motion information of the target block by adding the corrected motion information to the (first) motion information and store the updated motion information that includes both the (first) motion information and the corrected motion information.
[0103] Furthermore, in step S520, if it is determined that the motion information of the target block will not be updated, the decoding device stores the (first) motion information (step S540). For example, if the amount of data of the residual signal between a specific reference block and the reconstructed block derived based on the corrected motion information is not less than the amount of data of the residual signal between the reference block and the reconstructed block derived based on the (first) motion information, the decoding device may store the (first) motion information.
[0104] On the other hand, although not shown, the decoding device can update the motion information of the target block based on the modified motion information, without determining the update by comparing the (first) motion information with the modified motion information, in the case of deriving the modified motion information.
[0105] Furthermore, the stored motion information and the motion information used for transmission can have different resolutions. In other words, the units of the motion vectors included in the corrected motion information and the units of the motion vectors included in the motion information derived from the information used for inter-frame prediction can be different. For example, the units of the motion vectors in the motion information derived from the information used for inter-frame prediction can have 1 / 4 sample units, while the units of the motion vectors in the corrected motion information calculated in the decoding device can have 1 / 8 or 1 / 16 sample units. The decoding device can adjust the resolution to the required storage resolution during the processing of calculating the corrected motion information, or adjust the resolution through calculations (e.g., rounding, multiplication, etc.) during the processing of storing the corrected motion information. When the resolution of the motion vectors in the corrected motion information is higher than that of the motion vectors in the motion information derived from the information used for inter-frame prediction, the accuracy of scaling operations for calculating temporally adjacent motion information can be increased. In addition, in the case of decoding devices with different resolutions for transmitting motion information and for internal operations, there is an effect that decoding can be performed according to internal operation standards.
[0106] Furthermore, when the block matching method is applied in the method for calculating the corrected motion information, the corrected motion information can be derived as follows. The block matching method can be represented as the motion estimation method used in the encoding device.
[0107] The encoding device measures the degree of distortion based on the difference between the phase accumulated samples of the target block and the reference block and uses these as a cost function. Then, it derives the reference block that is most similar to the target block. That is, the encoding device can derive the corrected motion information of the target block based on the reference block with the smallest residual to the reconstructed block (or original block) of the target block. The reference block with the smallest residual can be called the specific reference block. The sum of absolute differences (SAD) and mean squared error (MSE) can be used as functions to represent the differences between samples. The SAD, which measures the degree of distortion based on the absolute value of the difference between the phase accumulated samples of the target block and the reference block, can be derived based on the following formula.
[0108] [Formula 1]
[0109]
[0110] In this article, Block cur (i,j) represents the reconstructed sample (or original sample) of the coordinates (i,j) in the reconstructed block (or original block) of the target block. ref (i,j) represents the reconstructed sample of the (i,j) coordinates in the reference block, width is the width of the reconstructed block (or the original block), and height represents the height of the reconstructed block (or the original block).
[0111] In addition, the MSE, which measures the degree of distortion by the square of the difference between the phase accumulation samples of the target block and the reference block, can be derived based on the following formula.
[0112] [Equation 2]
[0113]
[0114] In this article, Block cur (i,j) represents the reconstructed sample (or original sample) of the coordinates (i,j) in the reconstructed block (or original block) of the target block. ref (i,j) represents the reconstructed sample of the (i,j) coordinates in the reference block, width is the width of the reconstructed block (or the original block), and height represents the height of the reconstructed block (or the original block).
[0115] The computational complexity of the method used to calculate the corrected motion information can be flexibly varied depending on the search range used to search for a specific reference block. Therefore, when calculating the corrected motion information of the target block using a block matching method, the decoding device can search only for reference blocks that are within a predetermined region from the reference block derived from the (first) motion information used in the decoding process of the target block, thereby maintaining low computational complexity.
[0116] Figure 6 An example of a method by which a decoding device updates the motion information of a target block using a block matching method is shown. The decoding device decodes the target block (step S600). When inter-frame prediction is applied to the target block, the decoding device can obtain the inter-frame prediction information for the target block from the bitstream. Since the method for updating the motion information of the target block applied to the encoding device should be applied to the decoding device in the same way, the encoding device and the decoding device can calculate the corrected motion information of the target block based on the reconstructed block of the target block.
[0117] The decoding device performs motion estimation in a reference frame using the reconstructed blocks for the corrected motion information (step S610). The decoding device can deduce a reference frame for motion information derived based on information used for inter-frame prediction as a specific reference frame for the corrected motion information, and detect the specific reference frame among the reference blocks in the reference frame. The specific reference block can be the reference block with the smallest sum of absolute values of the differences between the samples of the reference blocks and the reconstructed blocks (i.e., the sum of absolute differences (SAD)). In addition, the decoding device can limit the search area for detecting the specific reference block to a predetermined area from the reference block represented by the motion information derived based on information used for inter-frame prediction. In other words, the decoding device can deduce the reference block with the smallest SAD with the reconstructed block among the reference blocks located within a predetermined area from the reference block as the specific reference block. The decoding device can perform motion estimation only within a predetermined area from the reference block, thereby increasing the reliability of the corrected motion information while reducing computational complexity.
[0118] If the size of the reconstructed block of the target block is larger than a specific size, in order to accurately calculate the corrected motion information, the decoding device divides the reconstructed block into smaller blocks smaller than the specific size and calculates more detailed motion information based on each smaller block (step S620). The specific size can be pre-configured, and the divided blocks can be called sub-reconstructed blocks. The decoding device can divide the reconstructed block into multiple sub-reconstructed blocks and derive a specific sub-reference block on a per-sub-reconstructed block basis within the reference frame of the corrected motion information, and use the derived specific sub-reference block as the specific reference block. The decoding device can calculate the corrected motion information based on the specific reference block.
[0119] Furthermore, when using the optical flow (OF) method to calculate the corrected motion information, the corrected motion information can be derived as follows. The OF method is based on the assumption that the velocity of objects in the target block is uniform and that the sample values representing the object do not change in the image. Given that the object moves δx on the x-axis and δy on the y-axis within time δt, the following formula can be established.
[0120] [Formula 3]
[0121] I(x,y,t)=I(x+δx,y+δy,t+δt)
[0122] In this paper, I(x,y,t) indicates the sample value representing the reconstructed sample at the (x,y) position of the object included in the target block at time t.
[0123] In Equation 3, when expanding the right-hand side of the Taylor series, the following equation can be derived.
[0124] [Formula 4]
[0125]
[0126] When equation 4 is satisfied, the following equation can be established.
[0127] [Formula 5]
[0128]
[0129] Rewriting equation 5, we can derive the following equation.
[0130] [Formula 6]
[0131]
[0132]
[0133] In this article, v x It is the vector component of the calculated motion vector on the x-axis, v y This is the vector component of the calculated motion vector on the y-axis. The decoding device can derive the partial derivatives of the object on the x, y, and t axes, and derive the motion vector (v) at the current position (i.e., the position of the reconstructed sample) by applying the derived partial derivatives to this formula. x ,v y In this case, for example, the decoding device may configure the reconstructed samples in the reconstructed block representing the target block of the object to be included in a 3×3 region unit, and calculate the motion vector (v) with the left-hand side close to 0 by applying the formula to the reconstructed samples. x ,v y ).
[0134] The storage format of the corrected motion information for the target block calculated by the encoding device can be various. Since it can be used for frames encoded after the encoding process of the next block adjacent to the target block or the target frame included in the target block, it may be advantageous to store the corrected motion information in the same format as the motion information used to predict the target block. The motion information used for prediction may include information about whether it is a single or double prediction, a reference frame index, and motion vectors. The corrected motion information can be calculated to have the same format as the motion information used for prediction.
[0135] Dual prediction is possible when only one reference frame of the target frame exists, i.e., except when the frame order count (POC) of the target frame is 1. Generally, prediction performance is higher when the encoding device performs dual prediction compared to single prediction. Therefore, when the target frame is available for dual prediction, the encoding device can use dual prediction to calculate and update the corrected motion information of the target block to improve the overall coding efficiency of the image by propagating accurate motion information, as described above. However, even when the target frame is available for dual prediction, i.e., even when the POC of the target frame is not 1, occlusion can occur, and dual prediction may not be possible. Occlusion is determined to occur when the sum of the absolute values of the differences between the samples of the predicted samples derived from the corrected motion information based on motion compensation and the samples of the reconstructed block (or original block) of the target block exceeds a predetermined threshold. For example, when the corrected motion information for the target block is double-predicted motion information, if the sum of the absolute values of the differences between the predicted samples derived from the L0 motion vector included in the corrected motion information and the samples of the reconstructed block (or original block) of the target block exceeds a predetermined threshold, occlusion can be determined based on the prediction of L0. Conversely, if the sum of the absolute values of the differences between the predicted samples derived from the L1 motion vector included in the corrected motion information and the samples of the reconstructed block (or original block) of the target block exceeds a predetermined threshold, occlusion can be determined based on the prediction of L1. In the case of occlusion occurring through the prediction of either L0 or L1, the encoding device can derive the corrected motion information as single-predicted motion information, in addition to the prediction information in the list of occlusions.
[0136] In cases where occlusion occurs due to predictions of either L0 or L1, or where corrected motion information can be used to calculate the motion information, a method for updating the target block's motion information can be applied using motion information from neighboring blocks. Alternatively, a method for not updating the target block's motion information can be applied based on the corrected motion information. For the purpose of storing and propagating accurate motion information, in the above scenario, the method of not updating the target block's motion information based on the corrected motion information may be more appropriate.
[0137] The L0 and L1 reference frame indices of the modified motion information can indicate one of the reference frames (L0 or L1) in the reference frame list. The reference frame indicated by the reference frame index of the modified motion information can be referred to as a specific reference frame. The method for selecting one of several reference frames in the reference frame list can be described as follows.
[0138] For example, the encoding device may select the most recently encoded reference frame from the reference frames included in the reference frame list. That is, the encoding device may generate a reference frame index included in the corrected motion information, which indicates the most recently encoded reference frame from the reference frames included in the reference frame list.
[0139] Additionally, the encoding device can select a reference frame from the reference frame list that has the closest Frame Order Count (POC) to the current frame. That is, the encoding device can generate a reference frame index included in the corrected motion information, which indicates the reference frame among the reference frames included in the reference frame list whose absolute difference between the POC and the POC of the target frame including the target block is the smallest.
[0140] In addition, the encoding device can select a reference frame that belongs to a lower layer in the hierarchical structure from the reference frame list. The reference frame belonging to the lower layer can be an I-slice or a reference frame encoded by applying a low quantization parameter (QP).
[0141] Additionally, the encoding device can select the reference frame with the highest reliability for motion compensation from the reference frame list. That is, the encoding device can derive a specific reference block of the reconstructed block (or original block) of the target block from the reference blocks in the reference frames included in the reference frame list and generate a reference frame index included in the corrected motion information, which indicates the reference frame including the derived specific reference block.
[0142] The methods described above for generating the reference frame index for the corrected motion information can be applied independently or in combination.
[0143] The motion vectors included in the modified motion information can be derived by at least one method, including OF method, block matching method, frequency domain method, etc., and may need to have motion vectors in units of at least a minimum block.
[0144] Figure 7 The reference image is shown and can be used in the corrected motion information.
[0145] Reference Figure 7When the target frame has a POC value of 3, dual prediction can be performed on the target frame. Frames with a POC value of 2 and a POC value of 4 can be derived as reference frames for the target frame. The L0 of the target frame can include frames with a POC value of 2 and frames with a POC value of 0, and the corresponding values of the L0 reference frame indices can be 0 and 1. Similarly, the L1 of the target frame can include frames with a POC value of 4 and frames with a POC value of 8, and the corresponding values of the L1 reference frame indices can be 0 and 1. In this case, the encoding device can select the reference frame for the corrected motion information of the target frame as the reference frame with the smallest absolute difference between the POC and the target frame among the reference frames in the list of reference frames. Furthermore, the encoding device can derive the motion vectors of the corrected motion information in the target frame in 4×4 blocks. The OF method can be applied to the method of deriving motion vectors. Motion vectors derived based on reference frames with a POC value of 2 can be stored in the L0 motion information included in the corrected motion information, and motion vectors derived based on reference frames with a POC value of 4 can be stored in the L1 motion information included in the corrected motion information.
[0146] Additionally, refer to Figure 7 When the target frame has a POC value of 8, single prediction can be performed on the target frame, and a frame with a POC value of 0 can be derived as the reference frame for the target frame. The L0 of the target frame can include the reference frame with a POC value of 0, and the value of the L0 reference frame index corresponding to the reference frame can be 0. In this case, the encoding device can select the specific reference frame for the corrected motion information of the target frame as the reference frame with the smallest absolute value of the difference between the POC and the POC of the target frame among the reference frames included in L0. In addition, the encoding device can derive the motion vector of the corrected motion information in the current frame in blocks of 4×4 size. The OF method can be applied to the method of deriving motion vectors. The block matching method can be applied to the method of deriving motion vectors, and the motion vector derived based on the reference frame with a POC value of 0 can be stored in the L0 motion information included in the corrected motion information.
[0147] Referring to the above implementation method, the reference frame for correcting the motion information of the target frame can be derived as the reference frame that indicates the frame with the smallest absolute value of the difference between the POC and the POC of the target frame among the reference frames included in the reference frame list, and the motion vector of the corrected motion information can be calculated in 4×4 blocks. When the target frame is a frame with a POC of 8, since the target frame is... Figure 7In the shown frame, inter-frame prediction is performed first, so single prediction is used. However, when the target frame has a POC of 3, dual prediction is available, and the corrected motion information of the target frame can be determined as dual-predicted motion information. Furthermore, even when the target frame has a POC of 3, if the corrected motion information is calculated and updated on a block-by-block basis in the target frame, the encoding device can determine to select a motion information format that includes both dual-predicted and single-predicted motion information on a block-by-block basis.
[0148] Furthermore, unlike the embodiments described above, a reference frame index for modified motion information can be generated to indicate the reference frame selected by the hierarchical structure, rather than the reference frame included in the reference frame list whose absolute difference between the POC and the POC of the target frame is the smallest. For example, if the target frame is a frame with a POC of 5, the L0 reference frame index for the modified motion information of the target frame can indicate a frame with a POC value of 0, rather than a frame with a POC value of 4. During timing after the encoding process of the target frame, most of the most recently encoded reference frames can indicate the reference frame whose absolute difference between the POC and the POC of the target frame is the smallest. However, similar to the case of a frame with a POC value of 6, the most recently encoded reference frame is a reference frame with a POC value of 2, and the frame with the smallest absolute difference between the POC and the POC of the target frame is a frame with a POC value of 4. Therefore, it is possible for different reference frames to be indicated.
[0149] When the motion information of the target block is updated to the corrected motion information using the method described above, the corrected motion information can be used to perform inter-frame prediction of the next block adjacent to the target block (the target block may be referred to as the first block, and the next block may be referred to as the second block). Methods for using the corrected motion information in the inter-frame prediction mode of the next block may include indexing and transmitting the corrected motion information, and representing the motion vector of the next block as the difference between the motion vector of the target block and the motion vector of the next block. The method of indexing and transmitting the corrected motion information can be used when applying a merging mode to the next block in the inter-frame prediction mode, and the method of representing the motion vector of the next block as the difference between the motion vector of the target block and the motion vector of the next block can be used when applying an AMVP mode to the next block in the inter-frame prediction mode.
[0150] Figure 8 This example shows a list of potential merge blocks next to the target block when the target block's motion information is updated to the target block's corrected motion information. (See also...) Figure 8The encoding device can be configured to include a list of candidate blocks to merge next to the target block. The encoding device can send a merge index indicating the block among the blocks included in the candidate list whose motion information is most similar to that of the next block. In this case, updated motion information of the target block can be used for the motion information of spatially neighboring candidate blocks of the next block, or alternatively, for the motion information of temporally neighboring candidate blocks. The motion information of the target block may include revised motion information of the target block. For example... Figure 8 As shown, the motion information of spatially neighboring candidate blocks can represent one of the stored motion information of the blocks adjacent to the next block at positions A1, B1, B0, A0, and B2. The motion information of the target block encoded before the encoding of the next block can be updated, and the updated motion information can affect the encoding process of the next block. In cases where erroneous motion information propagates to the next block through a merging pattern, the error may accumulate and propagate, but this can be mitigated by updating the motion information in each block using the method described above.
[0151] Figure 9 This illustrates an example of a merge candidate list for the next block adjacent to the target block when the target block's motion information is updated to a revised motion information. The (first) motion information and the revised motion information used to predict the target block adjacent to the next block can be stored. In this case, the encoding device can configure the target block indicating the newly calculated revised motion information of the target block in the merge candidate list as a separate merge candidate block. Figure 9 As shown, a target block indicating the corrected motion information of the target block can be inserted into the merge candidate list, next to the target block indicating the target block next to the target block whose (first) motion information was applied during the encoding process. Additionally, the priority of the merge candidate list can be changed, and the number of insertion positions can be configured differently.
[0152] Although A1 and B1 are shown as examples of target blocks for deriving the corrected motion information, this is only an example. Corrected motion information can also be derived for the remaining A0, B0, B2, T0, and T1, and the merge candidate list can be configured based on it.
[0153] For example, the next block may be located on a different screen than the target block, and the time-nearest candidate blocks of the next block may be included in the merge candidate list. Motion information of the time-nearest candidate blocks can be selected using the merge index of the next block. In this case, the motion information of the time-nearest candidate blocks can be updated to corrected motion information using the method described above, and the time-nearest candidate blocks indicating the updated motion information can be inserted into the merge candidate list of the next block. Alternatively, the (first) motion information and corrected motion information of the time-nearest candidate blocks can be stored separately using the method described above. In this case, during the generation of the merge candidate list of the next block, the time-nearest candidate blocks indicating the corrected motion information can be inserted as additional candidates into the merge candidate list.
[0154] Unlike Figure 9 The content shown can be used to apply the AMVP mode to the next block. When applying the AMVP mode to the next block, the encoding device can generate a list of motion vector predictor candidates based on the motion information of the neighboring blocks of the next block.
[0155] Figure 10 This illustrates an example of a motion vector predictor candidate list for the next block adjacent to the target block when the motion information of the target block is updated based on the corrected motion information of the target block. The encoding device can select motion information from the motion information included in the motion vector predictor candidate list according to specific conditions and send a motion vector difference (MVD) value indicating the index of the selected motion information, where the motion vector of the selected motion information is the motion vector of the next block's motion information. Similar to the method described above for inserting blocks indicating updated motion information of the next adjacent block into the spatial or temporal merge candidate block in the merge candidate list in merge mode, the process of generating a list of motion information of neighboring blocks including the AMVP mode can be applied to the method of inserting updated motion information of neighboring blocks as spatial or temporal motion vector predictor (MVP) candidates. Figure 10 As shown, the encoding device can detect whether the updated motion information of the target block A0 and target block A1 next to the next block conforms to a predefined specific condition in the order of the arrow direction, and deduce the first detected updated motion information conforming to the specific condition as MVP A to be included in the motion predictor candidate list.
[0156] In addition, the encoding device can detect whether the updated motion information of the next target block B0, target block B1 and target block B2 in the order of the arrow direction conforms to a predefined specific condition, and deduce the first detected updated motion information conforming to the specific condition as the MVP B to be included in the motion predictor candidate list.
[0157] Furthermore, the next block can be located in a different frame than the target block, and the time MVP candidate for the next block can be included in the motion vector predictor candidate list. For example... Figure 10 As shown, the encoding device can detect whether a predefined specific condition is met by following the order of updated motion information from target block T0 to target block T1 in the reference frame of the next block, and deduce the first detected updated motion information according to the specific condition as MVP Col to be included in the motion predictor candidate list. If MVP A, MVP B, and / or MVP Col are replaced by updated motion information instead of corrected motion information, the updated motion information can be used as a motion vector predictor candidate for the next block.
[0158] Furthermore, if the number of MVP candidates in the derived motion predictor candidate list is less than a certain number, the encoding device derives a zero vector as the MVP zero and includes it in the motion predictor candidate list.
[0159] Figure 11 This illustrates an example of a motion vector predictor candidate list for the next block adjacent to the target block when the motion information of the target block is updated based on the corrected motion information of the target block. When the motion information of the target block is updated motion information including both corrected and existing motion information, the encoding device can detect whether the existing motion information of the next adjacent target blocks A0 and A1 conforms to a specific condition in the order of the arrow directions, and deduce the first detected existing motion information conforming to the specific condition as MVP A to be included in the motion predictor candidate list.
[0160] In addition, the encoding device can detect whether the corrected motion information of target block A0 and target block A1 conforms to specific conditions in the order of the arrow direction, and deduce the first detected corrected motion information conforming to specific conditions as the updated MVP A included in the motion predictor candidate list.
[0161] In addition, the encoding device can detect whether the existing motion information of the target block B0, target block B1 and target block B2 next to the next block conforms to a predefined specific condition in the order of the arrow direction, and deduce the existing motion information that conforms to the specific condition first as the MVP B to be included in the motion predictor candidate list.
[0162] In addition, the encoding device can detect whether the corrected motion information of target block B0, target block B1 and target block B2 are in accordance with specific conditions in the order of the arrow direction, and deduce the first detected updated motion information in accordance with specific conditions as the updated MVP B included in the motion predictor candidate list.
[0163] Furthermore, the encoding device can detect whether the motion information of the target block T0 in the reference frame of the next block is in accordance with a specific condition, and deduce the first detected motion information in accordance with the specific condition as the MVP Col to be included in the motion predictor candidate list.
[0164] In addition, the encoding device can detect whether the motion information is corrected according to a specific condition in the order of the corrected motion information of the target block T0 in the reference frame of the next block to the corrected motion information of the target block T1, and deduce the first detected corrected motion information according to the specific condition as the MVP Col included in the motion predictor candidate list.
[0165] Furthermore, if the number of MVP candidates in the derived motion predictor candidate list is less than a certain number, the encoding device derives a zero vector as the MVP zero and includes it in the motion predictor candidate list.
[0166] Figure 12 A video encoding method according to the encoding device of the present invention is illustrated schematically. Figure 12 The method shown can be derived from Figure 1 The encoding device shown performs this. In a detailed example... Figure 12 Steps S1200 to S1230 can be executed by the prediction unit of the encoding device, and step S1240 can be executed by the memory of the encoding device.
[0167] The encoding device generates motion information for the target block (step S1200). The encoding device may apply inter-frame prediction to the target block. When applying inter-frame prediction to the target block, the encoding device may generate motion information for the target block by applying at least one of skip mode, merge mode, and adaptive motion vector prediction (AMVP) mode. The motion information may be referred to as first motion information. In the skip mode and merge mode, the encoding device may generate motion information for the target block based on the motion information of the target block's neighboring blocks. The motion information may include motion vectors and reference frame indices. The motion information may be dual-prediction motion information or single-prediction motion information. Dual-prediction motion information may include L0 reference frame index and L0 motion vector and L1 reference frame index and L1 motion vector, and unidirectional motion information may include L0 reference frame index and L0 motion vector or L1 reference frame index and L1 motion vector. L0 indicates reference frame list L0 (list 0), and L1 indicates reference frame list L1 (list 1).
[0168] In AMVP mode, the encoding device can use the motion vectors of the target block's neighboring blocks as motion vector predictors (MVPs) to derive the target block's motion vectors and generate motion information including the motion vectors and a reference frame index of the motion vectors.
[0169] The encoding device derives prediction samples by performing inter-frame prediction of the target block based on motion information (step S1210). The encoding device can generate prediction samples of the target block based on the reference frame index and motion vector included in the motion information.
[0170] The encoding device generates a reconstructed block based on the prediction samples (step S1220). The encoding device may generate a reconstructed block of the target block based on the prediction samples, or generate a residual signal of the target block and generate a reconstructed block of the target block based on the residual signal and the prediction samples.
[0171] The encoding device generates corrected motion information for the target block based on the reconstructed block (step S1230). The encoding device can calculate the corrected reference frame index and the corrected motion vector of the specific reference frame, which indicate the corrected motion information, using various methods. The corrected motion information can be referred to as second motion information. At least one method, including direct methods such as optical flow (OF) methods, block matching methods, and frequency domain methods, and indirect methods such as singularity matching methods and methods using statistical properties, can be applied to this method. In addition, direct and indirect methods can be applied simultaneously.
[0172] For example, the encoding device can generate corrected motion information using a block matching method. In this case, the encoding device can measure the degree of distortion based on the difference between the phase accumulation samples of the reconstructed block of the target block and the reference block, and use these as a cost function. Then, it detects a specific reference block of the reconstructed block. The encoding device can generate corrected motion information based on the detected specific reference block. In other words, the encoding device can generate corrected motion information, which includes a corrected reference frame index indicating a specific reference index and a corrected motion vector indicating a specific reference block in the specific reference frame. The encoding device can detect the reference block in the specific reference frame that has the smallest sum of the absolute values (or squared values) of the differences between the phases of the reconstructed block of the target block as the specific reference block, and derive the corrected motion information based on this specific reference block. As a method to represent the sum of the absolute values of the differences, the sum of absolute differences (SAD) can be used. In this case, Equation 1 above can be used to calculate the sum of the absolute values of the differences. Alternatively, as a method to represent the sum of the absolute values of the differences, the mean squared error (MSE) can be used. In this case, Equation 2 above can be used to calculate the sum of the absolute values of the differences.
[0173] Furthermore, the specific reference frame for the corrected motion information can be derived as the reference frame indicated by the reference frame index included in the (first) motion information, and the search area for detecting a specific reference block can be limited to a reference block located within a predetermined area that is a reference block derived from the reference frame based on the motion vectors related to the reference frame included in the (first) motion information. That is, the encoding device can deduce the reference block located within the predetermined area that has the smallest SAD with the reconstructed block among the reference blocks derived from the reference frame based on the motion vectors related to the reference frame included in the (first) motion information as the specific reference block.
[0174] Furthermore, if the size of the reconstructed block is larger than a predetermined size, the reconstructed block can be divided into multiple sub-reconstructed blocks, and a specific sub-reconstructed block can be derived in a specific reference frame on a unit basis. In this case, the encoding device can derive a specific reference block based on the derived specific sub-reconstructed blocks.
[0175] As another example, the encoding device can generate corrected motion information using the OF method. In this case, the encoding device can calculate the motion vector of the corrected motion information of the target block based on the assumptions that the velocity of the objects in the target block is uniform and that the sample values representing the objects do not change in the image. The motion vector can be calculated using Equation 6 above. The region indicating the samples of the objects included in the target block can be configured as a 3×3 area.
[0176] When calculating the corrected motion information using the method described above, the encoding device can calculate the corrected motion information in the same format as the (first) motion information. That is, the encoding device can calculate the corrected motion information in the same format as the motion information, between the dual-predictive motion information and the single-predictive motion information. The dual-predictive motion information may include an L0 reference frame index and an L0 motion vector, and an L1 reference frame index and an L1 motion vector. The single-predictive motion information may include an L0 reference frame index and an L0 motion vector, or an L1 reference frame index and an L1 motion vector. L0 indicates the reference frame list L0 (list 0), and L1 indicates the reference frame list L1 (list 1).
[0177] Furthermore, dual prediction can be performed on a target frame including a target block, and the encoding device can calculate the corrected motion information as dual-predicted motion information using the method described above. However, after the corrected motion information is calculated as dual-predicted motion information, if occlusion occurs on one of the specific reference blocks derived from the dual-predicted motion information, the encoding device can derive single-predicted motion information from the corrected motion information, in addition to the motion information of the reference frame list including the specific reference frame of the occluded reference block. Whether occlusion has occurred can be determined when the difference between samples of the phase of the reconstructed blocks of the reference block and the target block derived from the corrected motion information is greater than a certain threshold. The threshold can be pre-configured.
[0178] For example, the encoding device calculates the corrected motion information as double-predicted motion information, and if the difference between the samples of the phases of the reconstructed blocks of a specific reference block and the target block derived from the L0 motion vector and L0 reference frame index based on the corrected motion information is greater than a pre-configured threshold, the encoding device can derive the corrected motion information as single-predicted motion information including the L1 motion vector and L1 reference frame index.
[0179] For example, the encoding device calculates the corrected motion information as double-predicted motion information, and if the difference between the phase samples of the reconstructed blocks of a specific reference block and the target block derived from the L1 motion vector and L1 reference frame index based on the corrected motion information is greater than a pre-configured threshold, the encoding device can derive the corrected motion information as single-predicted motion information including the L0 motion vector and L0 reference frame index.
[0180] Furthermore, if the difference between the phase samples of a specific reference block and the reconstructed block of the target block derived based on the L0 motion vector and the L0 reference frame index is greater than a pre-configured threshold, and if the difference between the phase samples of a specific reference block and the reconstructed block of the target block derived based on the L1 motion vector and the L1 reference frame index is greater than a pre-configured threshold, the encoding device may derive motion information of the neighboring blocks of the target block as corrected motion information, or may not calculate corrected motion information.
[0181] Furthermore, the encoding device can select, through various methods, a specific reference frame indicated by the modified reference frame index included in the modified motion information.
[0182] For example, the encoding device can select the most recently encoded reference screen from the reference screens included in the reference screen list L0 and generate an L0 reference screen index indicating the correction of the reference screen. Additionally, the encoding device can select the most recently encoded reference screen from the reference screens included in the reference screen list L1 and generate an L1 reference screen index indicating the correction of the reference screen.
[0183] For example, the encoding device can select the reference frame with the smallest absolute value of the difference between the frame order count (POC) and the POC of the target frame from the reference frames included in the L0 reference frame index, and generate an L0 reference frame index indicating the correction of the reference frame. Additionally, the encoding device can select the reference frame with the smallest absolute value of the difference between the frame order count (POC) and the POC of the target frame from the reference frames included in the L1 reference frame index, and generate an L1 reference frame index indicating the correction of the reference frame.
[0184] For example, the encoding device can select the lowest-level reference frame in the hierarchical structure from the reference frames included in the various reference frame lists, and generate a reference frame index indicating the correction of the reference frame. The lowest-level reference frame can be an I-slice or a reference frame encoded by applying low quantization parameters (QP).
[0185] For example, the encoding device can select a reference frame from the reference frame list that includes the most reliable reference block with motion compensation, and generate a reference frame index indicating corrections to the reference frame. In other words, the encoding device can deduce a specific reference block for the reconstruction block of the target block based on the reference frames included in the reference frame list, and generate a reference frame index indicating corrections to a specific reference frame including the deduced specific reference block.
[0186] The methods described above for generating corrected reference frame indexes for motion information can be applied independently or in combination.
[0187] Additionally, although not shown, the encoding device can generate corrected motion information for the target block based on the original block of the target block. The encoding device can derive a specific reference block of the original block from reference blocks included in the reference frame, and generate corrected motion information indicating the derived specific reference block.
[0188] The encoding device updates the motion information of the target block based on the corrected motion information (step S1240). The encoding device can store the corrected motion information and update the motion information of the target block. The encoding device can update the motion information of the target block by replacing the motion information used to predict the target block with the corrected motion information. Alternatively, the encoding device can update the motion information of the target block by storing all the motion information used to predict the target block and the corrected motion information. The updated motion information can be used for the motion information of the next block adjacent to the target block. For example, if a merge mode is applied to the next block adjacent to the target block, the merge candidate list of the next block may include the target block. When the motion information of the target block is stored by replacing the motion information used to predict the target block with the corrected motion information, the merge candidate list of the next block may include the target block indicating the corrected motion information. In addition, since all the motion information used to predict the target block and the corrected motion information are stored in the motion information of the target block, the merge candidate list of the next block may include the target block indicating the motion information used to predict the target block and the target block indicating the corrected motion information. The target block indicating the corrected motion information may be inserted into the merge candidate list as a spatially adjacent candidate block or as a temporally adjacent candidate block.
[0189] For example, applying the AMVP mode to the next block adjacent to the target block is similar to the method described above in the merge mode, where updated motion information of the target block adjacent to the next block is inserted into the merge candidate list as spatial or temporal proximity motion information. Similarly, the method of inserting updated motion information of neighboring blocks into the motion vector predictor candidate list of the next block can be applied as spatial or temporal motion vector predictor candidates. That is, the encoding device can generate a motion vector predictor candidate list that includes updated motion information of the target block adjacent to the next block.
[0190] For example, if the updated motion information for the next neighboring target block only includes the corrected motion information, the motion vector predictor candidate list can include the corrected motion information as a spatial motion vector predictor candidate.
[0191] For example, if the updated motion information of the next neighboring target block includes the corrected motion information and the existing motion information of the target block, the motion vector predictor candidate list may include existing motion information selected from the existing motion information of the next neighboring target block according to specific conditions, and corrected motion information selected from the corrected motion information of the next neighboring target block according to specific conditions, as corresponding spatial motion vector predictor candidates.
[0192] Furthermore, the motion vector predictor candidate list for the next block may include updated motion information of co-position blocks in the reference frame of the next block that are at the same position as the next block, as well as updated motion information of neighboring blocks at the same position as the blocks in the temporal motion vector predictor candidate list. For example, if the updated motion information of co-position blocks at the same position as the next block and the updated motion information of neighboring blocks at the same position only include the corrected motion information of each block, the motion vector predictor candidate list may include the corrected motion information as part of the temporal motion vector predictor candidate list.
[0193] For example, when the updated motion information includes the corresponding corrected motion information of the block at the same location, the neighboring blocks of the block at the same location, and the existing motion information of the block at the same location, the motion vector predictor candidate list may include the corrected motion information selected according to specific conditions from the corrected motion information and the existing motion information selected according to specific conditions from the existing motion information, as separate spatial motion vector predictor candidates.
[0194] Furthermore, the stored motion information and the motion information used for transmission can have different resolutions. For example, the encoding device encodes and transmits the information used for motion information, and stores modified motion information for motion information adjacent to the target block. The unit of the motion vector in the motion information can be a 1 / 4 sample unit, and the unit of the motion vector in the modified motion information can represent a 1 / 8 sample or a 1 / 16 sample unit.
[0195] Furthermore, although not shown, the encoding device can determine whether to update the (first) motion information of the target block by comparing the (first) motion information performed on the reconstructed block of the target block with the corrected motion information. For example, the encoding device can determine whether to update the target block by comparing the amount of data of the residual signal of the reconstructed block of the specific reference block and the target block derived based on the corrected motion information with the amount of data of the residual signal of the reference block and the reconstructed block derived based on the motion information. If the amount of data of the residual signal of the specific reference block and the reconstructed block is small, the encoding device can determine to update the motion information and the target block. Alternatively, if the amount of data of the residual signal of the specific reference block and the reconstructed block is not small, the encoding device can determine not to update the motion information and the target block.
[0196] Furthermore, the encoding device can determine whether to update the (first) motion information of the target block by comparing the motion information performed on the original block of the target block with the modified motion information. For example, the encoding device can determine whether to update the target block by comparing the amount of data in the residual signals of the original blocks of a specific reference block and the target block derived based on the modified motion information with the amount of data in the residual signals of the reference block and the original blocks derived based on the motion information. If the amount of data in the residual signals of the specific reference block and the original blocks is small, the encoding device can determine to update the motion information and the target block. Conversely, if the amount of data in the residual signals of the specific reference block and the original blocks is not small, the encoding device can determine not to update the motion information and the target block.
[0197] In addition, the encoding device can generate and encode additional information indicating whether an update has occurred and output it via a bitstream. For example, this additional information indicating whether an update has occurred may be called an update flag. An update block of 1 indicates that motion information has been updated, while an update block of 0 indicates that motion information has not been updated. For example, update flags can be sent in units of PU. Alternatively, update flags can be sent in units of CU, CTU, or slices, and can be sent at a higher level, such as in units of Picture Parameter Set (PPS) or Sequence Parameter Set (SPS).
[0198] In addition, the encoding device can generate motion vector update difference information indicating the difference between the existing motion vector and the corrected motion information, and encode and output it. The motion vector update difference information can be sent in units of PU.
[0199] Although not shown, the encoding device can encode and output information about the residual samples of the target block. This information may include the transform coefficients of the residual samples.
[0200] Figure 13 A video decoding method according to the decoding device of the present invention is illustrated schematically. Figure 13 The method shown can be derived from Figure 2 The decoding device shown performs this operation. In a detailed example... Figure 13 Step S1300 can be executed by the entropy decoding unit of the decoding device, steps S1310 to S1340 can be executed by the prediction unit of the decoding device, and step S1350 can be executed by the memory of the decoding device.
[0201] The decoding device obtains inter-frame prediction information for the target block via the bitstream (step S1300). Inter-frame prediction or intra-frame prediction can be applied to the target block. When applying inter-frame prediction to the target block, the decoding device obtains inter-frame prediction information for the target block via the bitstream. Additionally, the decoding device obtains motion vector update difference information, indicating the difference between the existing motion vectors and the corrected motion information of the target block, via the bitstream. Furthermore, the decoding device obtains additional information via the bitstream indicating whether to update the target block. For example, the additional information indicating whether to update may be referred to as an update flag.
[0202] The decoding device derives motion information of the target block based on information about inter-frame prediction (step S1310). This motion information may be referred to as first motion information. The information about inter-frame prediction may represent the mode applied to the target block among skip mode, merge mode, and adaptive motion vector prediction (AMVP) mode. When skip mode or merge mode is applied to the target block, the decoding device may generate a merge candidate list including neighboring blocks of the target block and obtain a merge index indicating the neighboring blocks included in the merge candidate list. The merge index may be included in the information about inter-frame prediction. The decoding device may derive the motion information of the neighboring blocks indicated by the merge index as the motion information of the target block.
[0203] When AMVP mode is applied to the target block, similar to merge mode, the decoding device can generate a list based on the target block's neighboring blocks. The decoding device can generate indices indicating the neighboring blocks included in the generated list, as well as the motion vector difference (MVD) between the motion vector of the neighboring block indicated by the index and the motion vector of the target block. The index and MVD can be included in information about inter-frame prediction. The decoding device can generate motion information for the target block based on the motion vectors and MVD of the neighboring blocks indicated by the index.
[0204] Motion information may include motion vectors and reference frame indices. Motion information can be dual-predictive motion information or single-predictive motion information. Dual-predictive motion information may include the L0 reference frame index and L0 motion vector, and the L1 reference frame index and L1 motion vector. Unidirectional motion information may include either the L0 reference frame index and L0 motion vector or the L1 reference frame index and L1 motion vector. L0 indicates the reference frame list L0 (list 0), and L1 indicates the reference frame list L1 (list 1).
[0205] The decoding device derives prediction samples by performing inter-frame prediction of the target block based on motion information (step S1320). The decoding device can generate prediction samples of the target block based on the reference frame index and motion vector included in the motion information.
[0206] The decoding device generates a reconstructed block based on the predicted samples (step S1330). When a skip mode is applied to the target block, the decoding device can generate a reconstructed block of the target block based on the predicted samples. When a merge mode or AMVP mode is applied to the target block, the decoding device can generate a residual signal of the target block from the bitstream and generate a reconstructed block of the target block based on the residual signal and the predicted samples.
[0207] The encoding device derives the corrected motion information of the target block based on the reconstructed block (step S1340). The decoding device can calculate the corrected reference frame index and the corrected motion vector of the specific reference frame, which indicate the corrected motion information, using various methods. The corrected motion information can be referred to as second motion information. At least one method, including direct methods such as optical flow (OF) methods, block matching methods, and frequency domain methods, and indirect methods such as singularity matching methods and methods using statistical properties, can be applied to this method. In addition, direct and indirect methods can be applied simultaneously.
[0208] For example, a decoding device can generate corrected motion information using a block matching method. In this case, the decoding device measures the degree of distortion by accumulating the differences between samples of the phase of the reconstructed block and the reference block based on the target block and uses these as a cost function. Then, it detects a specific reference block of the reconstructed block. The decoding device can generate corrected motion information based on the detected specific reference block. In other words, the decoding device can generate corrected motion information that includes a corrected reference frame index indicating a specific reference index and a corrected motion vector indicating a specific reference block in the specific reference frame. The decoding device can detect the reference block in the reference block of the specific reference frame that has the smallest sum of the absolute values (or squared values) of the differences between samples of the phase of the reconstructed block based on the target block as the specific reference block, and derive the corrected motion information based on this specific reference block. As a method to represent the sum of the absolute values of the differences, the sum of absolute differences (SAD) can be used. In this case, the sum of the absolute values of the differences can be calculated using Equation 1 above. Alternatively, as a method to represent the sum of the absolute values of the differences, the mean squared error (MSE) can be used. In this case, the sum of the absolute values of the differences can be calculated using Equation 2 above.
[0209] Furthermore, the specific reference frame for the corrected motion information can be derived as a reference frame indicated by the reference frame index included in the (first) motion information. The search area for detecting a specific reference block can be limited to a reference block located within a predetermined area that is a reference block derived from the reference block in the specific reference frame based on the motion vector related to the reference frame included in the (first) motion information. That is, the decoding device can deduce the reference block located within the predetermined area that has the smallest SAD (Short-Adjusted Aspect Ratio) with the reconstructed block as the specific reference block.
[0210] Furthermore, if the size of the reconstructed block is larger than a predetermined size, the reconstructed block can be divided into multiple sub-reconstructed blocks, and specific sub-reconstructed blocks can be defined in a specific reference frame. In this case, the decoding device can derive a specific reference block based on the derived specific sub-reconstructed blocks.
[0211] For example, the decoding device can generate corrected motion information using the OF method. In this case, the decoding device can calculate the corrected motion vector of the target block based on the assumptions that the velocity of the objects in the target block is uniform and that the sample values representing the objects do not change in the image. The motion vector can be calculated using Equation 6 above. The region indicating the samples of the objects included in the target block can be configured as a 3×3 area.
[0212] When calculating the corrected motion information using the method described above, the decoding device can calculate the corrected motion information in the same format as the (first) motion information. That is, the decoding device can calculate the corrected motion information in the same format as the motion information, between the dual-predictive motion information and the single-predictive motion information. The dual-predictive motion information may include an L0 reference frame index and an L0 motion vector, and an L1 reference frame index and an L1 motion vector. The unidirectional motion information may include an L0 reference frame index and an L0 motion vector, or an L1 reference frame index and an L1 motion vector. L0 indicates the reference frame list L0 (list 0), and L1 indicates the reference frame list L1 (list 1).
[0213] Furthermore, dual prediction can be performed on a target frame including a target block, and the decoding device can calculate the corrected motion information as dual-predicted motion information using the method described above. However, after the corrected motion information is calculated as dual-predicted motion information, if occlusion occurs on one of the reference blocks derived from the dual-predicted motion information, the decoding device can derive the corrected motion information as single-predicted motion information, in addition to the motion information of the reference frame list including the specific reference frame of the occluded reference block. Whether occlusion has occurred can be determined as having occurred when the difference between the phase samples of the reconstructed blocks of the specific reference block and the target block derived from the corrected motion information is greater than a certain threshold. The threshold can be pre-configured.
[0214] For example, the decoding device calculates the corrected motion information as double-predicted motion information, and the encoding device can derive the corrected motion information as single-predicted motion information including L1 motion vector and L1 reference frame index when the difference between the samples of the phase of the reconstructed block of a specific reference block and the target block derived from the L0 motion vector and L0 reference frame index based on the corrected motion information is greater than a pre-configured threshold.
[0215] For example, the decoding device calculates the corrected motion information as dual-predicted motion information, and if the difference between the phase samples of the reconstructed blocks of a specific reference block and the target block derived from the L1 motion vector and L1 reference frame index based on the corrected motion information is greater than a pre-configured threshold, the decoding device can derive the corrected motion information as single-predicted motion information including the L0 motion vector and L0 reference frame index.
[0216] Furthermore, if the difference between the phase samples of a specific reference block and the reconstructed block of the target block derived from the L0 motion vector and the L0 reference frame index is greater than a pre-configured threshold, and if the difference between the phase samples of a specific reference block and the reconstructed block of the target block derived from the L1 motion vector and the L1 reference frame index is greater than a pre-configured threshold, the decoding device may deduce the motion information of the neighboring blocks of the target block as corrected motion information, or may not calculate the corrected motion information.
[0217] Furthermore, the decoding device can select a specific reference frame indicated by the corrected reference frame index, which is included in the corrected motion information, through various methods.
[0218] For example, the decoding device can select the most recently encoded reference screen from the reference screens included in the reference screen list L0, and generate an L0 reference screen index indicating the correction of the reference screen. Additionally, the decoding device can select the most recently encoded reference screen from the reference screens included in the reference screen list L1, and generate an L1 reference screen index indicating the correction of the reference screen.
[0219] For example, the decoding device can select the reference frame with the smallest absolute value of the difference between the Frame Order Count (POC) and the POC of the target frame from the reference frames included in the L0 reference frame index, and generate an L0 reference frame index indicating the correction of the reference frame. Additionally, the decoding device can select the reference frame with the smallest absolute value of the difference between the Frame Order Count (POC) and the POC of the target frame from the reference frames included in the L1 reference frame index, and generate an L1 reference frame index indicating the correction of the reference frame.
[0220] For example, the decoding device can select the lowest-level reference frame in the hierarchical structure from the reference frames included in the various reference frame lists, and generate a reference frame index indicating the correction of the reference frame. The lowest-level reference frame can be an I-slice or a reference frame encoded by applying low quantization parameters (QP).
[0221] For example, the decoding device can select the reference frame with the highest reliability, including motion compensation, from the reference frame list and generate a reference frame index indicating the correction of the reference frame. In other words, the decoding device can deduce a specific reference block of the reconstructed block of the target block based on the reference frames included in the reference frame list and generate a reference frame index indicating the correction of the specific reference frame including the deduced specific reference block.
[0222] The methods described above for generating corrected reference frame indexes for motion information can be applied independently or in combination.
[0223] Furthermore, the decoding device can obtain motion vector update difference information, which indicates the difference between the existing motion vector and the corrected motion vector of the target block, via a bit stream. In this case, the decoding device can derive the corrected motion information by summing the (first) motion information and the motion vector update difference information of the target block. The motion vector update difference information can be transmitted in units of PU as described above.
[0224] The decoding device updates the motion information of the target block based on the corrected motion information (step S1350). The decoding device can store the corrected motion information and update the motion information of the target block. The decoding device can update the motion information of the target block by replacing the motion information used to predict the target block with the corrected motion information. Alternatively, the decoding device can update the motion information of the target block by storing all the motion information used to predict the target block and the corrected motion information. The updated motion information can be used for the motion information of the next block adjacent to the target block.
[0225] For example, when applying a merge mode to the next block adjacent to the target block, the merge candidate list for the next block may include the target block. If the motion information of the target block is stored by replacing the motion information used to predict the target block with corrected motion information, the merge candidate list for the next block may include target blocks indicating the corrected motion information. Alternatively, if all motion information used to predict the target block and the corrected motion information are stored in the target block's motion information, the merge candidate list for the next block may include target blocks indicating the motion information used to predict the target block, as well as target blocks indicating the corrected motion information. Target blocks indicating corrected motion information may be inserted into the merge candidate list as spatially adjacent candidate blocks or as temporally adjacent candidate blocks.
[0226] For example, applying the AMVP mode to the next block adjacent to the target block is similar to the method described above in the merge mode, where updated motion information of the target block adjacent to the next block is inserted into the merge candidate list as spatial or temporal proximity motion information. This can be achieved by inserting updated motion information of neighboring blocks into the motion vector predictor candidate list of the next block as spatial or temporal motion vector predictor candidates. In other words, the decoding device can generate a motion vector predictor candidate list that includes updated motion information of the target block adjacent to the next block.
[0227] For example, if the updated motion information for the next neighboring target block only includes the corrected motion information, the motion vector predictor candidate list can include the corrected motion information as a spatial motion vector predictor candidate.
[0228] For example, if the updated motion information of the next neighboring target block includes the corrected motion information and the existing motion information of the target block, the motion vector predictor candidate list may include the corrected motion information selected according to specific conditions from the corrected motion information of the next neighboring target block and the existing motion information selected according to specific conditions from the existing motion information of the next neighboring target block, as corresponding spatial motion vector predictor candidates.
[0229] Furthermore, the motion vector predictor candidate list for the next block may include updated motion information of co-position blocks in the reference frame of the next block that are at the same position as the next block, as well as updated motion information of neighboring blocks at the same position as the blocks in the temporal motion vector predictor candidate list. For example, if the updated motion information of co-position blocks at the same position as the next block and the updated motion information of neighboring blocks at the same position only include the corrected motion information of each block, the motion vector predictor candidate list may include the corrected motion information as part of the temporal motion vector predictor candidate list.
[0230] For example, when the updated motion information includes the corresponding corrected motion information of the block at the same location, the neighboring blocks of the block at the same location, and the existing motion information of the block at the same location, the motion vector predictor candidate list may include the corrected motion information selected according to specific conditions from the corrected motion information and the existing motion information selected according to specific conditions from the existing motion information as separate spatial motion vector predictor candidates.
[0231] Furthermore, the stored motion information and the motion information used for transmission can have different resolutions. For example, the decoding device decodes and transmits the information used for motion information, and stores corrected motion information for motion information adjacent to the target block. The unit of the motion vector in the motion information can be a 1 / 4 sample unit, and the unit of the motion vector in the corrected motion information can represent a 1 / 8 sample or a 1 / 16 sample unit.
[0232] Furthermore, although not shown, the decoding device can determine whether to update the (first) motion information of the target block by performing a comparison process between the (first) motion information and the corrected motion information based on the reconstructed block of the target block. For example, the decoding device can determine whether to update the target block by comparing the amount of data of the residual signal of the reconstructed block of the specific reference block and the target block derived based on the corrected motion information with the amount of data of the residual signal of the reference block and the reconstructed block derived based on the motion information. Among these data amounts, if the amount of data of the residual signal of the specific reference block and the reconstructed block is small, the decoding device can determine to update the motion information and the target block. Alternatively, if the amount of data of the residual signal of the specific reference block and the reconstructed block is not small, the decoding device can determine not to update the motion information and the target block.
[0233] Furthermore, the decoding device can obtain additional information indicating whether the target block has been updated via the bitstream. For example, this additional information indicating whether an update has been made can be called an update flag. A value of 1 for the update block indicates that the motion information has been updated, while a value of 0 indicates that the motion information has not been updated. For example, the update flag can be sent in units of PU. Alternatively, the update flag can be sent in units of CU, CTU, or slice, and can be sent at a higher level, such as in units of Picture Parameter Set (PPS) or Sequence Parameter Set (SPS).
[0234] According to the present invention described above, after the decoding process of the target block, the corrected motion information of the target block is calculated and updated to more accurate motion information, thereby improving the overall coding efficiency.
[0235] Furthermore, according to the present invention, the motion information of the next block adjacent to the target block can be derived based on the updated motion information of the target block, and the propagation of distortion can be reduced, thereby improving the overall coding efficiency.
[0236] In the above embodiments, the method is described as a series of steps or blocks based on the flowchart. However, this disclosure is not limited to the order of these steps. Some steps may be performed simultaneously or in a different order than described above. Furthermore, those skilled in the art will understand that the steps shown in the flowchart are not exclusive. It will be understood that other steps may be included or one or more steps in the flowchart may be deleted without affecting the scope of this disclosure.
[0237] The method described above can be implemented in software. The encoding and / or decoding apparatus according to this disclosure can be included in apparatus for performing image processing, for example, a TV, computer, smartphone, set-top box, or display device.
[0238] When the embodiments of this disclosure are implemented in software, the above methods can be implemented by modules (processes, functions, etc.) that perform the above functions. These modules can be stored in memory and executed by a processor. The memory can be internal or external to the processor, and the memory can be connected to the processor using various well-known means. The processor may include application-specific integrated circuits (ASICs), other chipsets, logic circuits, and / or data processing devices. The memory may include ROM (read-only memory), RAM (random access memory), flash memory, memory cards, storage media, and / or other storage devices.
Claims
1. A decoding device for image decoding, the decoding device comprising: Memory; as well as At least one processor connected to the memory, the at least one processor being configured to: Derive the motion information of the first target block; The corrected motion information of the first target block is derived based on the motion information of the first target block and the dual prediction; When the merge mode is applied to the second target block, the merge candidate list of the second target block is configured based on the spatial neighbor block and the temporal neighbor block of the second target block; Receive the merge index of the second target block; One of the merge candidates that constitutes the merge candidate list is selected based on the merge index; The motion information of the second target block is derived based on the selected merging candidates; as well as Prediction samples for the second target block are generated by performing inter-frame prediction based on the motion information of the second target block. The merged candidates include spatial motion information candidates and temporal motion information candidates. Specifically, the candidate temporal motion information is derived based on the temporal proximity block, and Specifically, when the time neighbor block of the second target block corresponds to the first target block, the time motion information candidate of the second target block is derived based on the corrected motion information of the first target block.
2. An encoding device for image encoding, the encoding device comprising: Memory; as well as At least one processor connected to the memory, the at least one processor being configured to: Derive the motion information of the first target block; The corrected motion information of the first target block is derived based on the motion information of the first target block and the dual prediction; When the merge mode is applied to the second target block, the merge candidate list of the second target block is configured based on the spatial neighbor block and the temporal neighbor block of the second target block; Select one of the merge candidates that constitutes the merge candidate list; Generate the merge index of the selected merge candidate as an indication of the second target block; as well as The image information, including the merged index, is encoded. The merging candidates include spatial motion information candidates and temporal motion information candidates. The temporal motion information candidates are derived based on the temporally neighboring blocks. Specifically, when the time neighbor block of the second target block corresponds to the first target block, the time motion information candidate of the second target block is derived based on the corrected motion information of the first target block.
3. An apparatus for transmitting image data, the apparatus comprising: At least one processor is configured to obtain a bitstream of the image, wherein the bitstream is generated based on the following steps: deriving motion information of a first target block; deriving corrected motion information of the first target block based on the motion information and dual prediction of the first target block; configuring a merge candidate list of the second target block based on spatial and temporal neighbor blocks when a merge mode is applied to the second target block; selecting one of the merge candidates constituting the merge candidate list; generating a merge index of the second target block indicating the selected merge candidate; and encoding image information including the merge index; and A transmitter configured to transmit the data comprising the bit stream. The merging candidates include spatial motion information candidates and temporal motion information candidates. The temporal motion information candidates are derived based on the temporally neighboring blocks. Specifically, when the time neighbor block of the second target block corresponds to the first target block, the time motion information candidate of the second target block is derived based on the corrected motion information of the first target block.
Citation Information
Patent Citations
Motion vector prediction and refinement
CN102668562A
An apparatus, a method and a computer program for video coding and decoding
CN103891291A