Image coding method, digital storage medium, and data transmission method
By calculating and updating the corrected motion information of the target block during image decoding, the inter-frame prediction method is improved, which solves the problem of high-cost transmission and storage of high-resolution images, improves coding efficiency, and reduces distortion propagation.
Patent Information
- Application Number
- CN202310707482.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2016-12-05
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2036-12-05
AI Technical Summary
The increased cost of transmitting and storing high-resolution and high-quality images makes it difficult for existing technologies to effectively compress image data.
By calculating the corrected motion information of the target block and updating the motion information during the image decoding process, the inter-frame prediction method is improved, and the distortion propagation during the encoding process is reduced.
It improves image coding efficiency, reduces distortion propagation during the coding process, and enhances overall coding efficiency.
Smart Images

Figure CN116527885B_ABST
Abstract
Description
[0001] This application is a divisional application of the original application No. 201680091905.1 (International Application No. PCT / KR2016 / 014167, filed on December 5, 2016, entitled "Method and apparatus for decoding an image in an image coding system"). TECHNICAL FIELD TECHNICAL FIELD
[0002] The present application relates to a technology for image coding, and more particularly, to a method and apparatus for decoding an image in an image coding system. BACKGROUND
[0003] There is an increasing demand for high-resolution, high-quality images (e.g., HD (High Definition) images and UHD (Ultra High Definition) images) in various fields. Since image data has high resolution and high quality, the amount of information or bits to be transmitted increases relative to conventional image data. Therefore, when image data is transmitted using a medium such as a conventional wired / wireless broadband line or stored using an existing storage medium, the transmission cost and storage cost thereof increase.
[0004] Therefore, there is a need for an efficient image compression technology to effectively transmit, store, and reproduce information of high-resolution and high-quality images. SUMMARY
[0005] TECHNICAL PROBLEM
[0006] The present application provides a method and apparatus for improving image coding efficiency.
[0007] The present application also provides a method and apparatus for inter prediction, which updates motion information of a target block.
[0008] The present application also provides a method and apparatus for calculating corrected motion information of a target block after a decoding process of the target block and updating based on the corrected motion information.
[0009] The present application also provides a method and apparatus for using updated motion information of a target block for motion information of a next block neighboring the target block.
[0010] TECHNICAL SOLUTION
[0011] In one aspect, an image decoding method performed by a decoding device is provided. The image decoding method includes the steps of obtaining information about inter prediction of a target block through a bitstream, deriving motion information of the target block based on the information about the inter prediction, deriving a prediction sample by performing inter prediction for the target block based on the motion information, generating a reconstructed block based on the prediction sample, deriving corrected motion information of the target block based on the reconstructed block, and updating the motion information of the target block based on the corrected motion information.
[0012] In another aspect, a decoding device for performing image decoding is provided. The decoding device includes an entropy decoder configured to obtain information about inter prediction of a target block through a bitstream, a predictor configured to derive motion information of the target block based on the information about the inter prediction, derive a prediction sample by performing inter prediction for the target block based on the motion information, generate a reconstructed block of the target block based on the prediction sample, and derive corrected motion information of the target block based on the reconstructed block, and a memory configured to update the motion information of the target block based on the corrected motion information.
[0013] In another aspect, a video encoding method performed by an encoding device is provided. The method includes the steps of generating motion information of a target block, deriving a prediction sample by performing inter prediction for the target block based on the motion information, generating a reconstructed block based on the prediction sample, generating corrected motion information of the target block based on the reconstructed block, and updating the motion information of the target block based on the corrected motion information.
[0014] In another aspect, a video encoding device is provided. The encoding device includes a prediction unit configured to generate motion information of a target block, derive a prediction sample by performing inter prediction for the target block based on the motion information, generate a reconstructed block based on the prediction sample, and generate corrected motion information of the target block based on the reconstructed block, and a memory configured to update the motion information of the target block based on the corrected motion information.
[0015] Advantageous Effects
[0016] According to the present application, after a decoding process of a target block, corrected motion information of the target block is calculated and can be updated to more accurate motion information, whereby overall coding efficiency can be improved.
[0017] According to the present application, motion information of a next block adjacent to the target block can be derived based on the updated motion information of the target block, and propagation of distortion can be reduced, whereby overall coding efficiency can be improved. BRIEF DESCRIPTION OF DRAWINGS
[0018] Figure 1 is a schematic diagram showing a configuration of a video encoding apparatus to which the present disclosure is applicable.
[0019] Figure 2 is a schematic diagram showing a configuration of a video decoding apparatus to which the present disclosure is applicable.
[0020] Figure 3 Examples showing a case where inter prediction is performed based on uni-directional motion information and a case where inter prediction is performed based on bi-directional motion information applied to a target block.
[0021] Figure 4 Examples showing an encoding process including a method of updating motion information of a target block.
[0022] Figure 5 Examples showing a decoding process including a method of updating motion information of a target block.
[0023] Figure 6 Examples showing a method of updating motion information of a target block by a block matching method by a decoding device.
[0024] Figure 7 Reference pictures usable in the updated motion information are shown.
[0025] Figure 8 Examples showing a merge candidate list of a next block neighboring the target block in a case where the motion information of the target block is updated to the updated motion information of the target block.
[0026] Figure 9 Examples showing a merge candidate list of a next block neighboring the target block in a case where the motion information of the target block is updated to the updated motion information of the target block.
[0027] Figure 10 Examples showing a motion vector predictor candidate list of a next block neighboring the target block in a case where the motion information of the target block is updated based on the updated motion information of the target block.
[0028] Figure 11 Examples showing a motion vector predictor candidate list of a next block neighboring the target block in a case where the motion information of the target block is updated based on the updated motion information of the target block.
[0029] Figure 12 A video encoding method of an encoding device according to the present application is schematically shown.
[0030] Figure 13 A video decoding method of a decoding device according to the present application is schematically shown. DETAILED DESCRIPTION
[0031] The present disclosure can be modified in various forms, and specific embodiments thereof will be described and illustrated in the accompanying drawings. However, these embodiments are not intended to limit the present disclosure. The terms used in the following description are used to describe specific embodiments only and are not intended to limit the present disclosure. Singular expressions include plural expressions as long as they are clearly different from the context. Terms such as "include" and "have" are intended to indicate that there is a feature, number, step, operation, element, component, or a combination thereof described in the following description, and it should be understood that the possibility of adding one or more different features, numbers, steps, operations, elements, components, or a combination thereof is not excluded.
[0032] On the other hand, the elements in the drawings described in the present disclosure are independently drawn for the convenience of describing different specific functions, and are not intended to mean that the elements are embodied by independent hardware or independent software. For example, two or more of the elements can be combined to form a single element, or one element can be divided into a plurality of elements. Embodiments in which elements are combined and / or divided do not depart from the concept of the present disclosure.
[0033] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the accompanying drawings. In addition, similar reference numerals are used throughout the drawings to refer to similar elements, and the same description about similar elements will be omitted.
[0034] In the present specification, generally, a picture refers to a unit representing an image at a specific time, and a slice refers to a unit constituting a part of a picture. One picture can be composed of a plurality of slices, and the terms picture and slice can be mixed with each other as necessary.
[0035] A pixel can mean a minimum unit constituting one picture (or image). In addition, "sample" can be used as a term corresponding to a pixel. A sample can generally represent a pixel or a value of a pixel, can represent only a pixel (pixel value) of a luminance component, and can represent only a pixel (pixel value) of a chrominance component.
[0036] A unit indicates a basic unit of image processing. The unit can include at least one of a specific region and information related to the region. Alternatively, the unit can be mixed with terms such as a block, a region, etc. In a typical case, an MxN block can represent a set of samples or transform coefficients arranged in M columns and N rows.
[0037] Figure 1 The structure of a video encoding apparatus to which the present disclosure is applicable is briefly shown.
[0038] Referring to Figure 1 The video encoding apparatus 100 includes a picture partitioner 105, a predictor 110, a subtractor 115, a transformer 120, a quantizer 125, a rearranger 130, an entropy encoder 135, an inverse quantizer 140, an inverse transformer 145, an adder 150, a filter 255, and a memory 160.
[0039] The picture partitioner 105 can split an input picture into at least one processing unit. Here, the processing unit can be a coding unit (CU), a prediction unit (PU), or a transform unit (TU). The coding unit is a unit block for coding, and a largest coding unit (LCU) can be split into coding units of a deeper depth according to a quad tree structure. In this case, the largest coding unit can be used as a final coding unit, or the coding unit can be recursively split into coding units of a deeper depth as necessary and can be used as a final coding unit having an optimal size based on coding efficiency according to a video characteristic. When a smallest coding unit (SCU) is set, the coding unit cannot be split into a coding unit smaller than the smallest coding unit. Here, the final coding unit refers to a coding unit that is partitioned or split into a predictor or a transformer. The prediction unit is a block partitioned from the coding unit block, and can be a unit block for sample prediction. Here, the prediction unit can be divided into sub-blocks. The transform block can be split from the coding unit block according to a quad tree structure, and can be a unit block for deriving transform coefficients and / or a unit block for deriving a residual signal from the transform coefficients.
[0040] Hereinafter, the coding unit can be referred to as a coding block (CB), the prediction unit can be referred to as a prediction block (PB), and the transform unit can be referred to as a transform block (TB).
[0041] The prediction block or the prediction unit can mean a specific region having a block shape in a picture, and can include an array of predicted samples. Also, the transform block or the transform unit can mean a specific region having a block shape in a picture, and can include an array of transform coefficients or residual samples.
[0042] The predictor 110 can perform prediction on a processing target block (hereinafter, a current block), and can generate a prediction block including predicted samples of the current block. The unit of prediction performed in the predictor 110 can be a coding block, or can be a transform block, or can be a prediction block.
[0043] The predictor 110 can determine whether to apply intra prediction or to apply inter prediction on the current block. For example, the predictor 110 can determine whether to apply the intra prediction or the inter prediction in units of a CU.
[0044] In the case of intra prediction, the predictor 110 can derive prediction samples of the current block based on reference samples outside the current block in a picture to which the current block belongs (hereinafter, the current picture). In this case, the predictor 110 can derive the prediction samples based on an average or interpolation of neighboring reference samples of the current block (case (i)), or can derive the prediction samples based on a reference sample existing in a particular (prediction) direction with respect to the prediction sample among the neighboring reference samples of the current block (case (ii)). Case (i) can be referred to as a non-directional mode or a non-angular mode, and case (ii) can be referred to as a directional mode or an angular mode. In intra prediction, as an example, the prediction modes can include 33 directional modes and at least two non-directional modes. The non-directional modes can include a DC mode and a planar mode. The predictor 110 can determine a prediction mode to be applied to the current block using a prediction mode applied to a neighboring block.
[0045] In the case of inter prediction, the predictor 110 can derive prediction samples of the current block based on samples specified by a motion vector on a reference picture. The predictor 110 can derive the prediction samples of the current block by applying any one of a skip mode, a merge mode, and a motion vector prediction (MVP) mode. In the case of the skip mode and the merge mode, the predictor 110 can use motion information of a neighboring block as motion information of the current block. In the case of the skip mode, unlike in the merge mode, a difference (residual) between the prediction samples and original samples is not transmitted. In the case of the MVP mode, a motion vector of a neighboring block is used as a motion vector predictor, and thus as a motion vector predictor of the current block to derive a motion vector of the current block.
[0046] In the case of inter prediction, the neighboring blocks can include spatial neighboring blocks existing in the current picture and temporal neighboring blocks existing in the reference picture. The reference picture including the temporal neighboring blocks can also be referred to as a collocated picture (colPic). The motion information can include a motion vector and a reference picture index. Information such as prediction mode information and motion information can be (entropy) encoded and then output in the form of a bitstream.
[0047] When using motion information of the temporal neighboring blocks in the skip mode and the merge mode, a highest picture in a reference picture list can be used as a reference picture. The reference pictures included in the reference picture list can be aligned based on a picture order count (POC) difference between the current picture and the corresponding reference picture. The POC corresponds to a display order and can be distinguished from an encoding order.
[0048] The subtractor 115 generates residual samples that are differences between the original samples and the prediction samples. When the skip mode is applied, the residual samples can not be generated as described above.
[0049] The transformer 120 transforms the residual samples in transform block units to generate transform coefficients. The transformer 120 can perform a transform based on a size of a corresponding transform block and a prediction mode applied to a coding block or a prediction block spatially overlapping the transform block. For example, residual samples can be transformed using a discrete sine transform (DST) when intra prediction is applied to a coding block or a prediction block overlapping the transform block and the transform block is a 4x4 residual array, and transformed using a discrete cosine transform (DCT) in other cases.
[0050] The quantizer 125 can quantize the transform coefficients to generate quantized transform coefficients.
[0051] The rearranger 130 rearranges the quantized transform coefficients. The rearranger 130 can rearrange the quantized transform coefficients in a block form into a one-dimensional vector through a coefficient scanning method. Although the rearranger 130 is described as a separate component, the rearranger 130 can be a part of the quantizer 125.
[0052] The entropy encoder 135 can perform entropy encoding on the quantized transform coefficients. The entropy encoding can include an encoding method such as exponential Golomb, context adaptive variable length coding (CAVLC), context adaptive binary arithmetic coding (CABAC), or the like. In addition to the quantized transform coefficients, the entropy encoder 135 can perform encoding on information (e.g., syntax element values, etc.) required for video reconstruction, together or separately. The entropy-encoded information can be transmitted or stored in a network abstraction layer (NAL) unit in a bitstream form.
[0053] The inverse quantizer 140 inverse-quantizes values (transform coefficients) quantized by the quantizer 125, and the inverse transformer 145 inverse-transforms values inverse-quantized by the inverse quantizer 140 to generate residual samples.
[0054] The adder 150 adds the residual samples to the prediction samples to reconstruct a picture. The residual samples can be added to the prediction samples in block units to generate reconstructed blocks. Although the adder 150 is described as a separate component, the adder 150 can be a part of the predictor 110.
[0055] The filter 155 can apply deblocking filtering and / or sample adaptive offset to the reconstructed picture. Artifacts at block boundaries in the reconstructed picture or distortion in quantization can be corrected through deblocking filtering and / or sample adaptive offset. The sample adaptive offset can be applied in a sample unit after deblocking filtering is completed. The filter 155 can apply an adaptive loop filter (ALF) to the reconstructed picture. The ALF can be applied to the reconstructed picture to which deblocking filtering and / or sample adaptive offset have been applied.
[0056] The memory 160 can store a reconstructed picture or information required for encoding / decoding. Here, the reconstructed picture can be a reconstructed picture filtered by the filter 155. The stored reconstructed picture can be used as a reference picture for (inter) prediction of other pictures. For example, the memory 160 can store a (reference) picture for inter prediction. Here, the picture for inter prediction can be designated according to a reference picture set or a reference picture list.
[0057] Figure 2 The structure of a video decoding apparatus to which the present disclosure is applicable is briefly shown.
[0058] Referring to Figure 2 The video decoding apparatus 200 includes an entropy decoder 210, a rearranger 220, an inverse quantizer 230, an inverse transformer 240, a predictor 250, an adder 260, a filter 270, and a memory 280.
[0059] When a bitstream including video information is input, the video decoding apparatus 200 can reconstruct a video in association with a process of processing video information in a video encoding apparatus.
[0060] For example, the video decoding apparatus 200 can perform video decoding using a processing unit applied in a video encoding apparatus. Accordingly, a processing unit block of video decoding can be a coding unit block, a prediction unit block, or a transform unit block. As a unit block of decoding, the coding unit block can be split from a maximum coding unit block according to a quad tree structure. As a block split from the coding unit block, the prediction unit block can be a unit block of sample prediction. In this case, the prediction unit block can be divided into sub-blocks. As a coding unit block, the transform unit block can be split according to a quad tree structure, and can be a unit block for deriving a transform coefficient or a unit block for deriving a residual signal from a transform coefficient.
[0061] The entropy decoder 210 can parse a bitstream to output information required for video reconstruction or picture reconstruction. For example, the entropy decoder 210 can decode information in a bitstream based on an encoding method such as exponential Golomb coding, CAVLC, CABAC, or the like, and can output values of syntax elements required for video reconstruction and quantized values of transform coefficients with respect to a residual.
[0062] More specifically, the CABAC entropy decoding method can receive bins corresponding to respective syntax elements in a bitstream, determine a context model using the information of the decoding target syntax element and the decoding information of neighboring and decoding target blocks or the information of the symbols / bins decoded in the previous step, predict a bin generation probability according to the determined context model, and perform arithmetic decoding of the bins to generate symbols corresponding to respective syntax element values. Here, the CABAC entropy decoding method can update the context model using the information of the symbols / bins decoded for the context model of the next symbol / bin after determining the context model.
[0063] Among the information decoded in the entropy decoder 210, information about prediction can be provided to the predictor 250, and the residual values (i.e., quantized transform coefficients) on which the entropy decoder 210 has performed entropy decoding can be input to the rearranger 220.
[0064] The rearranger 220 can rearrange the quantized transform coefficients into a two-dimensional block form. The rearranger 220 can perform rearrangement corresponding to the coefficient scan performed by the encoding apparatus. Although the rearranger 220 is described as a separate component, the rearranger 220 can be a part of the quantizer 230.
[0065] The inverse quantizer 230 can inverse quantize the quantized transform coefficients based on a (de)quantization parameter to output transform coefficients. In this case, information for deriving the quantization parameter can be signaled from the encoding apparatus.
[0066] The inverse transformer 240 can inverse transform the transform coefficients to derive residual samples.
[0067] The predictor 250 can perform prediction on a current block, and can generate a prediction block including prediction samples of the current block. The unit of prediction performed in the predictor 250 can be a coding block or can be a transform block or can be a prediction block.
[0068] The predictor 250 can determine whether to apply intra prediction or inter prediction based on the information about prediction. In this case, the unit for determining which one between the intra prediction and the inter prediction will be used can be different from the unit for generating prediction samples. Also, in the inter prediction and the intra prediction, the unit for generating prediction samples can also be different. For example, which one between the inter prediction and the intra prediction will be applied can be determined in a CU unit. Further, for example, in the inter prediction, prediction samples can be generated by determining a prediction mode in a PU unit, and in the intra prediction, prediction samples can be generated in a TU unit by determining a prediction mode in a PU unit.
[0069] In the case of intra prediction, the predictor 250 can derive the prediction samples of the current block based on neighboring reference samples in the current picture. The predictor 250 can derive the prediction samples of the current block by applying a directional mode or a non-directional mode based on the neighboring reference samples of the current block. In this case, the intra prediction mode of the neighboring block can be used to determine the prediction mode to be applied to the current block.
[0070] In the case of inter prediction, the predictor 250 can derive the prediction samples of the current block based on samples specified in the reference picture according to a motion vector. The predictor 250 can derive the prediction samples of the current block using one of a skip mode, a merge mode, and an MVP mode. Here, the motion information (e.g., a motion vector and information about a reference picture index) required for inter prediction of the current block provided by the video encoding apparatus can be acquired or derived based on the information about prediction.
[0071] In the skip mode and the merge mode, the motion information of the neighboring block can be used as the motion information of the current block. Here, the neighboring block can include a spatial neighboring block and a temporal neighboring block.
[0072] The predictor 250 can construct a merge candidate list using the motion information of the available neighboring blocks and use the information indicated by a merge index on the merge candidate list as the motion vector of the current block. The merge index can be signaled by the encoding apparatus. The motion information can include a motion vector and a reference picture. When the motion information of the temporal neighboring block is used in the skip mode and the merge mode, the highest picture in the reference picture list can be used as the reference picture.
[0073] In the case of the skip mode, unlike the merge mode, the difference (residual) between the prediction samples and the original samples is not transmitted.
[0074] In the case of the MVP mode, the motion vector of the neighboring block can be used as a motion vector predictor to derive the motion vector of the current block. Here, the neighboring block can include a spatial neighboring block and a temporal neighboring block.
[0075] When the merge mode is applied, for example, a merge candidate list can be generated using the motion vector of the reconstructed spatial neighboring block and / or the motion vector corresponding to the Col block which is the temporal neighboring block. In the merge mode, the motion vector of the candidate block selected from the merge candidate list is used as the motion vector of the current block. The above-described information about prediction can include a merge index indicating the candidate block having the best motion vector selected from the candidate blocks included in the merge candidate list. Here, the predictor 250 can use the merge index to derive the motion vector of the current block.
[0076] When the MVP (Motion Vector Prediction) mode is applied as another example, a motion vector predictor candidate list can be generated using the motion vector of the reconstructed spatial neighboring block and / or the motion vector corresponding to the Col block which is a temporal neighboring block. That is, the motion vector of the reconstructed spatial neighboring block and / or the motion vector corresponding to the Col block which is a temporal neighboring block can be used as a motion vector candidate. The above-described information on prediction can include a prediction motion vector index indicating a best motion vector selected from the motion vector candidates included in the list. Here, the predictor 250 can select a prediction motion vector of the current block using the motion vector index from the motion vector candidates included in the motion vector candidate list. The predictor of the encoding apparatus can obtain a motion vector difference (MVD) between the motion vector of the current block and the motion vector predictor, encode the MVD, and output the encoded MVD in the form of a bitstream. That is, the MVD can be obtained by subtracting the motion vector predictor from the motion vector of the current block. Here, the predictor 250 can obtain the motion vector included in the information on prediction and derive the motion vector of the current block by adding the motion vector difference to the motion vector predictor. In addition, the predictor can obtain or derive a reference picture index indicating a reference picture from the above-described information on prediction.
[0077] The adder 260 can add the residual samples to the prediction samples to reconstruct the current block or the current picture. The adder 260 can reconstruct the current picture by adding the residual samples to the prediction samples in a unit of block. When the skip mode is applied, the residual is not transmitted, and thus the prediction samples can become the reconstructed samples. Although the adder 260 is described as a separate component, the adder 260 can be a part of the predictor 250.
[0078] The filter 270 can apply deblocking filtering, sample adaptive offset, and / or ALF to the reconstructed picture. Here, the sample adaptive offset can be applied in a unit of sample after deblocking filtering. The ALF can be applied after deblocking filtering and / or applying the sample adaptive offset.
[0079] The memory 280 can store the reconstructed picture or information required for decoding. Here, the reconstructed picture can be the reconstructed picture filtered by the filter 270. For example, the memory 280 can store a picture used for inter prediction. Here, the picture used for inter prediction can be specified according to a reference picture set or a reference picture list. The reconstructed picture can be used as a reference picture of other pictures. The memory 280 can output the reconstructed picture in an output order.
[0080] As described above, in a case where inter prediction is performed with respect to a target block, motion information of the target block can be generated by applying a skip mode, a merge mode, or an adaptive motion vector prediction (AMVP) mode, and encoded and output. In this case, in the motion information of the target block, due to a process of being encoded in a unit of a block, a distortion can be calculated and included, and thus, motion information indicating a reconstructed block of the target block can not be perfectly reflected. In particular, in a case where the merge mode is applied to the target block, accuracy of the motion information of the target block can be deteriorated. That is, there is a great difference between a predicted block derived through the motion information of the target block and the reconstructed block of the target block. In this case, the motion information of the target block is used in a decoding process of a next block neighboring the target block, and the distortion can be propagated, and thus, overall encoding efficiency can be deteriorated.
[0081] Accordingly, in the present application, a method is proposed in which corrected motion information of a target block is calculated based on a derived reconstructed block after a decoding process of the target block, the motion information of the target block is updated based on the corrected motion information, so that a next block neighboring the target block can derive more accurate motion information. Thus, overall encoding efficiency can be improved.
[0082] Figure 3 Examples of a case where inter prediction is performed based on uni-directional motion information and a case where inter prediction is performed based on bi-directional motion information applied to a target block are shown. The bi-directional motion information can include an L0 reference picture index and an L0 motion vector and an L1 reference picture index and an L1 motion vector, and the uni-directional motion information can include the L0 reference picture index and the L0 motion vector or the L1 reference picture index and the L1 motion vector. The L0 indicates a reference picture list L0 (list 0), and the L1 indicates a reference picture list L1 (list 1). In an image encoding process, a method for inter prediction can include deriving motion information through motion estimation and motion compensation. As shown, the motion estimation can be indicated in a process of deriving a block matching the target block with respect to a reference block encoded before encoding of a target picture including the target block. The block matching the target block can be defined as a reference block, and a position difference between the target block and the reference block derived by assuming that the same reference block as the target block is included in the target picture can be defined as a motion vector of the target block. The motion compensation can include a uni-directional method of deriving and using one reference block and a bi-directional method of deriving and using two reference blocks. Figure 3
[0083] Information including a motion vector of the target block and information of the reference picture can be defined as motion information of the target block. The information of the reference picture can include a reference picture list and a reference picture index indicating a reference picture included in the reference picture list. The encoding apparatus can store the motion information of the target block for a next block neighboring the target block to be encoded after the target block is encoded or a picture to be encoded after the target picture is encoded. The stored motion information can be used in a method of representing the motion information of the next block to be encoded after the target block or the picture to be encoded after the target block. The method can include a merge mode, a method of indexing the motion information of the target block and transmitting the motion information of the next block neighboring the target block as an index, and an AMVP mode, a method of representing the motion vector of the next block neighboring the target block only with a difference between the motion vector of the next block and the motion vector of the target block.
[0084] The present application proposes a method of updating motion information of a target picture or a target block after an encoding process of the target picture or the target block. When inter prediction is performed and encoded for the target block, motion information used in the inter prediction can be stored. However, the motion information can include distortion occurring during a process calculated by a block matching method, and since the motion information can be a value selected for the target block through a rate-distortion (RD) optimization process, it can be difficult to perfectly reflect actual motion of the target block. Since the motion information of the target block can be used not only in an inter prediction process of the target block but also affect a picture encoded after an encoding of a target picture including a next block neighboring the target block encoded after the target block, a picture encoded after an encoding of the target picture, and the target picture including the next block and the target block, overall encoding efficiency can be deteriorated.
[0085] Figure 4 An example of an encoding process including a method of updating motion information of a target block is shown. An encoding apparatus encodes a target block (step S400). The encoding apparatus can derive a prediction sample by performing inter prediction for the target block and generate a reconstructed block of the target block based on the prediction sample. The encoding apparatus can derive motion information of the target block for performing the inter prediction and generate information of inter prediction of the target block including the motion information. The motion information can be referred to as first motion information.
[0086] After the encoding process of the target block is performed, the encoding apparatus calculates motion information corrected based on the reconstructed block of the target block (step S410). The encoding apparatus can calculate the corrected motion information of the target block through various methods. The corrected motion information can be referred to as second motion information. At least one method including a direct method such as an optical flow (OF) method, a block matching method, a frequency domain method, etc. and an indirect method such as a singularity matching method, a method using statistical properties, etc. can be applied to the method. In addition, the direct method and the indirect method can be applied at the same time. Details of the block matching method and the OF method will be described below.
[0087] Due to the characteristics of the motion information calculation of the target block, it can be difficult for the encoding apparatus to calculate more accurate motion information using only the encoded information of the target block, and thus, for example, the encoding apparatus can calculate the corrected motion information after performing the encoding process of the target picture including the target block, not after the encoding process of the target block. That is, the encoding apparatus can perform the encoding process of the target picture and calculate the corrected motion information based on more information than immediately after performing the encoding process of the target block.
[0088] In addition, the encoding apparatus can further generate and encode and output motion vector update difference information indicating a difference between the existing motion vector included in the (first) motion information and the corrected motion information. The motion vector update difference information can be transmitted in units of PUs.
[0089] The encoding apparatus determines whether to update the motion information of the target block (step S420). The encoding apparatus can determine whether to update the (first) motion information through a comparison of the accuracy between the (first) motion information and the corrected motion information. For example, the encoding apparatus can determine whether to update the (first) motion information using the amount of difference between the images derived using each of the motion information through motion compensation and the original image. In other words, the encoding apparatus can determine whether to update by comparing the data amount of the residual signal between the reference block derived based on each of the motion information and the original block of the target block.
[0090] In step S420, in the case where it is determined to update the motion information of the target block, the encoding apparatus updates the motion information of the target block based on the corrected motion information and stores the corrected motion information (step S430). For example, in the case where the data amount of the residual signal between the specific reference block derived based on the corrected motion information and the original block is smaller than that of the reference block derived based on the (first) motion information and the original block, the encoding apparatus can update the (first) motion information based on the corrected motion information. In this case, the encoding apparatus can update the motion information of the target block by replacing the (first) motion information with the corrected motion information and store the updated motion information including only the corrected motion information. In addition, the encoding apparatus can update the motion information of the target block by adding the corrected motion information to the (first) motion information and store the updated motion information including the (first) motion information and the corrected motion information.
[0091] Further, in a case where it is determined not to update the motion information of the target block in step S420, the encoding apparatus stores the (first) motion information (step S440). For example, in a case where the amount of data of the residual signal between the particular reference block derived based on the corrected motion information and the original block is not smaller than the amount of data of the residual signal between the reference block derived based on the (first) motion information and the original block, the encoding apparatus can store the (first) motion information.
[0092] On the other hand, although not shown, in a case where the update is determined without the comparison between the (first) motion information used in the inter prediction and the corrected motion information in deriving the corrected motion information, the encoding apparatus can update the motion information of the target block based on the corrected motion information.
[0093] Further, since the decoding apparatus is not available to use the original picture, the decoding apparatus can receive additional information indicating whether or not to update the target block. That is, the encoding apparatus can generate and encode the additional information indicating whether or not to update, and output it through the bitstream. For example, the additional information indicating whether or not to update can be referred to as an update flag. A case where the update flag is 1 can indicate that the motion information is updated, and a case where the update flag is 0 can indicate that the motion information is not updated. For example, the update flag can be transmitted in a unit of a PU. Alternatively, the update flag can be transmitted in a unit of a CU, in a unit of a CTU, or in a unit of a slice, and can be transmitted through a higher level such as a unit of a picture parameter set (PPS) or a unit of a sequence parameter set (SPS).
[0094] In addition, the decoding apparatus can determine whether or not to update based on the comparison between the (first) motion information and the corrected motion information of the target block through the reconstructed block of the target block without receiving the update flag, and in a case where it is determined that the motion information of the target block is updated, the decoding apparatus can update the motion information of the target block based on the corrected motion information and store the updated motion information. In addition, in a case where the update is determined without the comparison between the motion information used in the inter prediction and the corrected motion information in deriving the corrected motion information, the decoding apparatus can update the motion information of the target block based on the corrected motion information. In this case, the decoding apparatus can update the motion information of the target block by replacing the (first) motion information with the corrected motion information and store the updated motion information including only the corrected motion information. In addition, the decoding apparatus can update the motion information of the target block by adding the corrected motion information to the (first) motion information and store the updated motion information including the (first) motion information and the corrected motion information.
[0095] In a case where the process of updating the motion information is performed after the encoding process of the target block, the encoding process of a next block neighboring the target block in a next encoding order of the target block can be performed.
[0096] Figure 5An example of a decoding process showing a method of updating the motion information of a target block is shown. The decoding process can be performed in a similar manner as the encoding process described above. The decoding device decodes the target block (step S500). In the case where inter prediction is applied to the target block, the decoding device can obtain the information for inter prediction of the target block through the bitstream. The decoding device can derive the motion information of the target block based on the information for inter prediction and derive the prediction samples by performing inter prediction for the target block. The motion information can be referred to as first motion information. The decoding device can generate the reconstructed block of the target block based on the prediction samples.
[0097] The decoding device calculates the corrected motion information of the target block (step S510). The decoding device can calculate the corrected motion information through various methods. The corrected motion information can be referred to as second motion information. At least one method including direct methods such as an optical flow (OF) method, a block matching method, a frequency domain method, etc. and indirect methods such as a singular point matching method, a method using statistical properties, etc. can be applied to the method. In addition, the direct methods and the indirect methods can be applied at the same time. Details of the block matching method and the OF method will be described below.
[0098] Due to the characteristics of the motion information calculation of the target block, it can be difficult for the decoding device to calculate more accurate motion information using only the decoded information of the target block, and thus, for example, the decoding device can calculate the corrected motion information after performing the decoding process of the target picture including the target block, not after the decoding process of the target block. That is, the decoding device can perform the decoding process of the target picture and calculate the corrected motion information based on more information than immediately after performing the decoding process of the target block.
[0099] In addition, the decoding device can also obtain motion vector update difference information indicating a difference between the existing motion vector included in the (first) motion information and the corrected motion information through the bitstream. The motion vector update difference information can be transmitted in units of PUs. In this case, the decoding device can not independently calculate the corrected motion information through the direct methods and the indirect methods described above, but derive the corrected motion information by summing the (first) motion information and the obtained motion vector update difference information. That is, the decoding device can derive the existing motion vector using a motion vector predictor (MVP) of the corrected motion information and derive the corrected motion vector by adding the motion vector update difference information to the existing motion vector.
[0100] The decoding apparatus determines whether to update the motion information of the target block (step S520). The decoding apparatus can determine whether to update the (first) motion information by comparing the accuracy between the (first) motion information and the corrected motion information. For example, the decoding apparatus can determine whether to update the (first) motion information using the amount of difference between the reference block derived using each motion information through motion compensation and the reconstructed block of the target block. In other words, the encoding apparatus can determine whether to update by comparing the amount of data of the residual signal between the reference block derived based on each motion information and the original block of the target block.
[0101] In addition, the decoding apparatus receives additional information indicating whether to update from the encoding apparatus and determines whether to update the motion information of the target block based on the additional information. For example, the additional information indicating whether to update can be referred to as an update flag. The case where the update flag is 1 can indicate that the motion information is updated, and the case where the update flag is 0 can indicate that the motion information is not updated. For example, the update flag can be transmitted in units of PUs.
[0102] In step S520, in the case where it is determined to update the motion information of the target block, the decoding apparatus updates the motion information of the target block based on the corrected motion information and stores the updated motion information (step S530). For example, in the case where the amount of data of the residual signal between the specific reference block derived based on the corrected motion information and the reconstructed block is smaller than the amount of data of the residual signal between the reference block derived based on the (first) motion information and the reconstructed block, the decoding apparatus can update the (first) motion information based on the corrected motion information. In this case, the decoding apparatus can update the motion information of the target block by replacing the (first) motion information with the corrected motion information and store the updated motion information including only the corrected motion information. In addition, the decoding apparatus can update the motion information of the target block by adding the corrected motion information to the (first) motion information and store the updated motion information including the (first) motion information and the corrected motion information.
[0103] In addition, in step S520, in the case where it is determined not to update the motion information of the target block, the decoding apparatus stores the (first) motion information (step S540). For example, in the case where the amount of data of the residual signal between the specific reference block derived based on the corrected motion information and the reconstructed block is not smaller than the amount of data of the residual signal between the reference block derived based on the (first) motion information and the reconstructed block, the decoding apparatus can store the (first) motion information.
[0104] On the other hand, although not shown, in the case where the corrected motion information is derived without determining the update through the comparison between the (first) motion information and the corrected motion information, the decoding apparatus can update the motion information of the target block based on the corrected motion information.
[0105] Further, the stored motion information and the motion information for transmission can have different resolutions. In other words, the unit of the motion vector included in the corrected motion information and the unit of the motion vector included in the motion information derived based on the information for inter prediction can be different. For example, the unit of the motion vector of the motion information derived based on the information for inter prediction can have a unit of 1 / 4 fractional sample, and the unit of the motion vector of the corrected motion information calculated in the decoding device can have a unit of 1 / 8 fractional sample or 1 / 16 fractional sample. The decoding device can adjust the resolution to the resolution required for storage in the process of calculating the corrected motion information, or adjust the resolution by calculation (e.g., rounding, multiplication, etc.) in the process of storing the corrected motion information. In the case where the resolution of the motion vector of the corrected motion information is higher than that of the motion vector of the motion information derived based on the information for inter prediction, accuracy can increase for the scaling operation for the temporal neighboring motion information calculation. In addition, in the case of a decoding device in which the resolutions for transmission motion information and for internal operation are different, there is an effect that decoding can be performed according to the internal operation standard.
[0106] Further, in the case where a block matching method is applied among the methods for calculating the corrected motion information, the corrected motion information can be derived as follows. The block matching method can be expressed as a motion estimation method used in the encoding device.
[0107] The encoding device can derive the reference block most similar to the target block by measuring the degree of distortion from the difference between the accumulated samples of the phases of the target block and the reference block and using these as a cost function. That is, the encoding device can derive the corrected motion information of the target block based on the reference block having the smallest residual from the reconstructed block (or the original block) of the target block. The reference block having the smallest residual can be referred to as a specific reference block. The sum of absolute differences (SAD) and the mean square error (MSE) can be used to express a function of the difference between samples. The SAD, which measures the degree of distortion from the absolute value of the difference between the accumulated samples of the phases of the target block and the reference block, can be derived based on the following equation.
[0108] [Equation 1]
[0109]
[0110] Herein, Block cur (i,j) denotes a reconstructed sample (or an original sample) of the (i,j) coordinate in the reconstructed block (or the original block) of the target block, Block ref (i,j) denotes a reconstructed sample of the (i,j) coordinate in the reference block, width is the width of the reconstructed block (or the original block), and height denotes the height of the reconstructed block (or the original block).
[0111] In addition, the MSE measuring the degree of distortion according to the square value of the difference between the phase accumulated samples of the target block and the reference block can be derived based on the following equation.
[0112] [Equation 2]
[0113]
[0114] Herein, Block cur (i,j) denotes a reconstructed sample (or an original sample) of the (i,j) coordinate in the reconstructed block (or the original block) of the target block, Block ref (i,j) denotes a reconstructed sample of the (i,j) coordinate in the reference block, width is the width of the reconstructed block (or the original block), and height denotes the height of the reconstructed block (or the original block).
[0115] The computational complexity of the method for calculating the corrected motion information can be flexibly changed according to the search range for searching for a specific reference block. Accordingly, in the case of calculating the corrected motion information of the target block using the block matching method, the decoding apparatus can search only for a reference block included in a predetermined region from the reference block derived by using (first) motion information used in the decoding process of the target block, and thus low computational complexity can be maintained.
[0116] Figure 6 An example of a method of updating the motion information of a target block by a decoding apparatus through a block matching method is shown. The decoding apparatus decodes a target block (step S600). In the case of applying inter prediction to the target block, the decoding apparatus can obtain information for inter prediction of the target block through a bitstream. Since the method of updating the motion information of a target block applied to an encoding apparatus should be applied to a decoding apparatus in the same manner, the encoding apparatus and the decoding apparatus can calculate the corrected motion information of the target block based on the reconstructed block of the target block.
[0117] The decoding apparatus performs motion estimation in the reference picture for the corrected motion information using the reconstructed block (step S610). The decoding apparatus can derive a reference picture based on the motion information derived for the inter prediction as a specific reference picture of the corrected motion information, and detect the specific reference picture among the reference blocks in the reference picture. The specific reference block can be a reference block having the smallest sum of absolute differences (SAD) between samples of the reconstructed block and samples of the specific reference block. In addition, the decoding apparatus can limit a search area for detecting the specific reference block to a predetermined area from the reference block indicated by the motion information derived for the inter prediction. In other words, the decoding apparatus can derive a reference block having the smallest SAD between the reconstructed block among the reference blocks located within the predetermined area from the reference block as the specific reference block. The decoding apparatus can perform the motion estimation only within the predetermined area from the reference block, thereby increasing the reliability of the corrected motion information while reducing the computational complexity.
[0118] In a case where the size of the reconstructed block of the target block is greater than a specific size, in order to carefully calculate the corrected motion information, the decoding apparatus divides the reconstructed block into small blocks having a size smaller than the specific size and calculates more detailed motion information based on each of the small blocks (step S620). The specific size can be pre-configured, and the divided blocks can be referred to as sub-reconstructed blocks. The decoding apparatus can divide the reconstructed block into a plurality of sub-reconstructed blocks and derive specific sub-reference blocks within the reference picture of the corrected motion information in units of the sub-reconstructed blocks based on the sub-reconstructed blocks, and derive the derived specific sub-reference blocks as the specific reference block. The decoding apparatus can calculate the corrected motion information based on the specific reference block.
[0119] Further, in a case where an optical flow (OF) method is applied among the methods for calculating the corrected motion information, the corrected motion information can be derived as follows. The OF method can calculate based on an assumption that the speed of an object in the target block is uniform and the sample value of a sample representing the object does not change in an image. In a case where the object moves by δx in an x-axis and moves by δy in a y-axis within a time δt, the following equation can be established.
[0120] [Equation 3]
[0121] I(x, t, t) = I(x + δx, y + δy, t + δt)
[0122] Herein, I(x, y, t) indicates a sample value of a reconstructed sample at an (x, y) position representing an object at a time t included in the target block.
[0123] In Equation 3, when the right term is expanded in a Taylor series, the following equation can be derived.
[0124] [Equation 4]
[0125]
[0126] When Equation 4 is satisfied, the following equation can be established.
[0127] [Equation 5]
[0128]
[0129] Rewriting Equation 5, the following equation can be derived.
[0130] [Equation 6]
[0131]
[0132]
[0133] Herein, v x is a vector component of the calculated motion vector in the x-axis, v y is a vector component of the calculated motion vector in the y-axis. The decoding device can derive partial derivative values of the object in the x-axis, the y-axis, and the t-axis, and derive a motion vector (v x ,v y ) of the current position (i.e., the position of the reconstructed sample) by applying the derived partial derivative values to the equation. In this case, for example, the decoding device can configure the reconstructed samples in the reconstructed block representing the target block of the object to include reconstructed samples in a 3x3 size region unit, and calculate a motion vector (v x ,v y ) in which the left term approaches 0 by applying the equation to the reconstructed samples.
[0134] The storage format of the corrected motion information of the target block calculated by the encoding device can have various formats, and since it can be used for a picture encoded after the encoding process of the target picture included in the next block adjacent to the target block or the target block, it can be beneficial that the corrected motion information can be stored in the same format as the motion information used for prediction of the target block. The motion information used for prediction can include information on whether it is single prediction or double prediction, a reference picture index, and a motion vector. The corrected motion information can be calculated to have the same format as the motion information used for prediction.
[0135] In a case where there is only one reference picture of the target picture, i.e., in a case where the value of the picture order count (POC) of the target picture is 1, bi-prediction can be performed. In general, the prediction performance can be higher in a case where the encoding apparatus performs bi-prediction than in a case where the encoding apparatus performs uni-prediction. Thus, in a case where the target picture can be used to perform bi-prediction, in order to improve the overall coding efficiency of the image by propagating accurate motion information, the encoding apparatus can calculate and update the corrected motion information of the target block using bi-prediction by the above-described method. However, even in a case where the target picture can be used to perform bi-prediction, i.e., even in a case where the value of the POC of the target picture is not 1, an occlusion can occur and a case where bi-prediction cannot be performed can occur. When the motion is compensated based on the corrected motion information and the sum of absolute values of differences between samples of a prediction sample derived based on the corrected motion information by motion compensation and samples of the reconstructed block (or the original block) of the target block is greater than a predetermined threshold, it can be determined that the occlusion occurs. For example, in a case where the corrected motion information of the target block is bi-prediction motion information, when the sum of absolute values of differences between samples of a prediction sample derived based on the L0 motion vector included in the corrected motion information and samples of the reconstructed block (or the original block) of the target block is greater than a predetermined threshold, according to the prediction of L0, it can be determined that the occlusion occurs. Also, when the sum of absolute values of differences between samples of a prediction sample derived based on the L1 motion vector included in the corrected motion information and samples of the reconstructed block (or the original block) of the target block is greater than a predetermined threshold, according to the prediction of L1, it can be determined that the occlusion occurs. In a case where the occlusion occurs through the prediction of any one of L0 and L1, the encoding apparatus can derive the corrected motion information as uni-prediction motion information, in addition to the information for prediction of the list in which the occlusion occurs.
[0136] In a case where the occlusion occurs through the prediction of any one of L0 and L1 or can be used to calculate the corrected motion information, a method of updating the motion information of the target block can be applied using the motion information of the neighboring block of the target block, or alternatively, a method of not updating the motion information of the target block based on the corrected motion information can be applied. In order to store and propagate accurate motion information, in the above-described case, the method of not updating the motion information of the target block based on the corrected motion information can be more appropriate.
[0137] The L0 reference picture index and the L1 reference picture index of the corrected motion information can indicate one of the reference pictures (L0 or L1) in the reference picture list. The reference picture indicated by the reference picture index of the corrected motion information can be referred to as a specific reference picture. The method of selecting one of a number of reference pictures in the reference picture list can be as described below.
[0138] For example, the encoding apparatus can select a reference picture having a closest picture order count (POC) to a POC of a target picture including a target block among reference pictures included in a reference picture list. That is, the encoding apparatus can generate a reference picture index included in the modified motion information, which indicates a reference picture having a closest POC to a POC of a target picture including a target block among reference pictures included in a reference picture list.
[0139] In addition, the encoding apparatus can select a reference picture having a closest POC to a POC of a target picture including a target block among reference pictures included in a reference picture list. That is, the encoding apparatus can generate a reference picture index included in the modified motion information, which indicates a reference picture having a closest POC to a POC of a target picture including a target block among reference pictures included in a reference picture list.
[0140] In addition, the encoding apparatus can select a reference picture having a closest POC to a POC of a target picture including a target block among reference pictures included in a reference picture list. That is, the encoding apparatus can generate a reference picture index included in the modified motion information, which indicates a reference picture having a closest POC to a POC of a target picture including a target block among reference pictures included in a reference picture list.
[0141] In addition, the encoding apparatus can select a reference picture having a closest POC to a POC of a target picture including a target block among reference pictures included in a reference picture list. That is, the encoding apparatus can generate a reference picture index included in the modified motion information, which indicates a reference picture having a closest POC to a POC of a target picture including a target block among reference pictures included in a reference picture list.
[0142] The above-described methods of generating a reference picture index of the modified motion information can be independently applied, or the methods can be combined and applied.
[0143] The motion vector included in the modified motion information can be derived by at least one method including an OF method, a block matching method, a frequency domain method, etc., and can need to have a motion vector in units of at least a minimum block.
[0144] Figure 7 A reference picture usable in the modified motion information is shown.
[0145] Reference Figure 7In a case where the target picture is a picture with a POC value of 3, bi-prediction can be performed with respect to the target picture, and a picture with a POC value of 2 and a picture with a POC value of 4 can be derived as reference pictures of the target picture. L0 of the target picture can include the picture with a POC value of 2 and a picture with a POC value of 0, and respective corresponding values of L0 reference picture indexes can be 0 and 1. In addition, L1 of the target picture can include the picture with a POC value of 4 and a picture with a POC value of 8, and respective corresponding values of L1 reference picture indexes can be 0 and 1. In this case, the encoding apparatus can select a specific reference picture of the modified motion information of the target picture as a reference picture having the smallest absolute value of a difference between a POC and a POC of the target picture among the reference pictures in each reference picture list. Further, the encoding apparatus can derive a motion vector of the modified motion information in the target picture in units of a 4x4 size block. An OF method can be applied to a method of deriving the motion vector. A motion vector derived based on the reference picture with a POC value of 2 can be stored in L0 motion information included in the modified motion information, and a motion vector derived based on the reference picture with a POC value of 4 can be stored in L1 motion information included in the modified motion information.
[0146] In addition, referring to Figure 7 In a case where the target picture is a picture with a POC value of 8, the target picture can perform uni-prediction, and a picture with a POC value of 0 can be derived as a reference picture of the target picture. L0 of the target picture can include the reference picture with a POC value of 0, and a value of an L0 reference picture index corresponding to the reference picture can be 0. In this case, the encoding apparatus can select a specific reference picture of the modified motion information of the target picture as a reference picture having the smallest absolute value of a difference between a POC and a POC of the target picture among the reference pictures included in L0. Further, the encoding apparatus can derive a motion vector of the modified motion information in the current picture in units of a 4x4 size block. An OF method can be applied to a method of deriving the motion vector. A block matching method can be applied to a method of deriving the motion vector, and a motion vector derived based on the reference picture with a POC value of 0 can be stored in L0 motion information included in the modified motion information.
[0147] Referring to the above-described embodiments, a reference picture of modified motion information of a target picture can be derived as a reference picture having the smallest absolute value of a difference between a POC and a POC of the target picture among reference pictures included in a reference picture list, and a motion vector of the modified motion information can be calculated in units of a 4x4 size block. In a case where the target picture is a picture with a POC of 8, since the target picture is Figure 7The picture in which the inter prediction is first performed among the pictures shown, so the single prediction is performed, but in the case where the target picture is the picture of which the POC is 3, the bi prediction is available and the modified motion information of the target picture can be determined as the bi prediction motion information. Also, even in the case where the target picture is the picture of which the POC is 3, in the case where the modified motion information is calculated and updated in the target picture in units of blocks, the encoding apparatus can determine to select one motion information format including the bi prediction motion information and the single prediction motion information in units of blocks.
[0148] Also, unlike the above-described embodiment, the reference picture index of the modified motion information can be generated to indicate the reference picture selected by the hierarchical structure, not the reference picture of which the absolute value of the difference of the POC from the POC of the target picture is the smallest among the reference pictures included in the reference picture list. For example, in the case where the target picture is the picture of which the POC is 5, the L0 reference picture index of the modified motion information of the target picture can indicate the picture of which the POC value is 0, not the picture of which the POC value is 4. At the timing after the encoding process of the target picture is performed, the most recently encoded reference picture can indicate the reference picture of which the absolute value of the difference of the POC from the POC of the target picture is the smallest. However, similarly to the case of the picture of which the POC value is 6, the most recently encoded reference picture is the reference picture of which the POC value is 2, and the picture of which the absolute value of the difference of the POC from the POC of the target picture is the smallest is the picture of which the POC value is 4, thus the case of indicating a different reference picture can occur.
[0149] In the case where the motion information of the target block is updated to the modified motion information by the above-described method, the inter prediction of the next block neighboring the target block can be performed using the modified motion information (the target block can be referred to as a first block, and the next block can be referred to as a second block). The method of using the modified motion information in the inter prediction mode of the next block can include a method of indexing and transmitting the modified motion information and a method of expressing the motion vector of the next block as a difference value between the motion vector of the target block and the motion vector of the next block. The method of indexing and transmitting the modified motion information can be a method in the case where the merge mode is applied to the next block among the inter prediction modes, and the method of expressing the motion vector of the next block as a difference value between the motion vector of the target block and the motion vector of the next block can be a method in the case where the AMVP mode is applied to the next block among the inter prediction modes.
[0150] Figure 8 An example of showing a merge candidate list of the next block neighboring the target block in the case where the motion information of the target block is updated to the modified motion information of the target block is shown. Referring to FIG. 10, the merge candidate list of the next block is shown in the case where the target block is the picture of which the POC is 3 and the next block is the picture of which the POC is 4. Figure 8The encoding device can be configured to include a list of candidate blocks to merge next to the target block. The encoding device can send a merge index indicating the block among the blocks included in the candidate list whose motion information is most similar to that of the next block. In this case, updated motion information of the target block can be used for the motion information of spatially neighboring candidate blocks of the next block, or alternatively, for the motion information of temporally neighboring candidate blocks. The motion information of the target block may include revised motion information of the target block. For example... Figure 8 As shown, the motion information of spatially neighboring candidate blocks can represent one of the stored motion information of the blocks adjacent to the next block at positions A1, B1, B0, A0, and B2. The motion information of the target block encoded before the encoding of the next block can be updated, and the updated motion information can affect the encoding process of the next block. In cases where erroneous motion information propagates to the next block through a merging pattern, the error may accumulate and propagate, but this can be mitigated by updating the motion information in each block using the method described above.
[0151] Figure 9 This illustrates an example of a merge candidate list for the next block adjacent to the target block when the target block's motion information is updated to a revised motion information. The (first) motion information and the revised motion information used to predict the target block adjacent to the next block can be stored. In this case, the encoding device can configure the target block indicating the newly calculated revised motion information of the target block in the merge candidate list as a separate merge candidate block. Figure 9 As shown, a target block indicating the corrected motion information of the target block can be inserted into the merge candidate list, next to the target block indicating the target block next to the target block whose (first) motion information was applied during the encoding process. Additionally, the priority of the merge candidate list can be changed, and the number of insertion positions can be configured differently.
[0152] Although A1 and B1 are shown as examples of target blocks for deriving the corrected motion information, this is only an example. Corrected motion information can also be derived for the remaining A0, B0, B2, T0, and T1, and the merge candidate list can be configured based on it.
[0153] For example, the next block can be located on a different picture than the target block, and temporal neighboring candidate blocks of the next block can be included in the merge candidate list. The motion information of the temporal neighboring candidate blocks can be selected by a merge index of the next block. In this case, the motion information of the temporal neighboring candidate blocks can be updated to the corrected motion information by the above-described method, and the temporal neighboring candidate blocks indicating the updated motion information can be inserted in the merge candidate list of the next block. In addition, the (first) motion information and the corrected motion information of the temporal neighboring candidate blocks can be separately stored by the above-described method. In this case, in the process of generating the merge candidate list of the next block, the temporal neighboring candidate blocks indicating the corrected motion information can be inserted as additional candidates of the merge candidate list.
[0154] Unlike Figure 9 As illustrated, the encoding device can detect whether the updated motion information of the target blocks A0 and A1 neighboring the next block follows the pre-defined specific condition in the order of the arrow direction, and derive the first detected updated motion information following the specific condition as the MVP A included in the motion predictor candidate list.
[0155] Figure 10 An example of a motion vector predictor candidate list of a next block neighboring a target block in a case where the motion information of the target block is updated based on the corrected motion information of the target block is illustrated. The encoding device can select motion information among the motion information included in the motion vector predictor candidate list according to a specific condition, and transmit a motion vector difference (MVD) value indicating an index of the selected motion information, the motion vector of which is the motion vector of the motion information of the next block. Similar to the above-described method of inserting a block indicating the updated motion information of the neighboring next block in the merge candidate list in the merge mode or the spatial or temporal merge candidate block, the process of generating a list including the motion information of the neighboring block in the AMVP mode can be applied to the method of inserting the updated motion information of the neighboring block as a spatial or temporal motion vector predictor (MVP) candidate. As illustrated, Figure 10 As illustrated, the encoding device can detect whether the updated motion information of the target blocks A0 and A1 neighboring the next block follows the pre-defined specific condition in the order of the arrow direction, and derive the first detected updated motion information following the specific condition as the MVP A included in the motion predictor candidate list.
[0156] In addition, the encoding device can detect whether the updated motion information of the target blocks B0, B1, and B2 neighboring the next block follows the pre-defined specific condition in the order of the arrow direction, and derive the first detected updated motion information following the specific condition as the MVP B included in the motion predictor candidate list.
[0157] In addition, the next block can be located on a different picture than the target block, and temporal MVP candidates of the next block can be included in the motion vector predictor candidate list. AsFigure 10 As shown, the encoding device can detect whether the updated motion information of the target block T0 in the reference picture of the next block to the updated motion information of the target block T1 in the order of the arrow direction, and derive the first detected updated motion information in accordance with the specific condition as the MVP Col included in the motion predictor candidate list. In the case that the MVP A, the MVP B and / or the MVP Col is updated motion information by being replaced by the updated motion information, the updated motion information can be used for the motion vector predictor candidate of the next block.
[0158] In addition, in the case that the number of the MVP candidates in the derived motion predictor candidate list is less than a specific number, the encoding device derives a zero vector as the MVP zero and includes it in the motion predictor candidate list.
[0159] Figure 11 An example of the motion vector predictor candidate list of the next block adjacent to the target block in the case that the motion information of the target block is updated based on the updated motion information of the target block. In the case that the motion information of the target block is the updated motion information including the updated motion information and the existing motion information, the encoding device can detect whether the existing motion information of the target block A0 and the target block A1 adjacent to the next block in accordance with the specific condition in the order of the arrow direction, and derive the first detected existing motion information in accordance with the specific condition as the MVP A included in the motion predictor candidate list.
[0160] In addition, the encoding device can detect whether the updated motion information of the target block A0 and the target block A1 in accordance with the specific condition in the order of the arrow direction, and derive the first detected updated motion information in accordance with the specific condition as the updated MVP A included in the motion predictor candidate list.
[0161] In addition, the encoding device can detect whether the updated motion information of the target block A0 and the target block A1 in accordance with the specific condition in the order of the arrow direction, and derive the first detected updated motion information in accordance with the specific condition as the updated MVP A included in the motion predictor candidate list.
[0162] In addition, the encoding device can detect whether the updated motion information of the target block A0 and the target block A1 in accordance with the specific condition in the order of the arrow direction, and derive the first detected updated motion information in accordance with the specific condition as the updated MVP A included in the motion predictor candidate list.
[0163] Further, the encoding device can detect whether the target block T0 in the reference picture of the next block is in accordance with the certain condition in the order of the existing motion information of the target block T0 to the existing motion information of the target block T1, and derive the first detected existing motion information in accordance with the certain condition as the MVP Col included in the motion predictor candidate list.
[0164] Further, the encoding device can detect whether the target block T0 in the reference picture of the next block is in accordance with the certain condition in the order of the existing motion information of the target block T0 to the existing motion information of the target block T1, and derive the first detected existing motion information in accordance with the certain condition as the MVP Col included in the motion predictor candidate list.
[0165] Further, in a case where the number of the MVP candidates in the derived motion predictor candidate list is less than a certain number, the encoding device derives a zero vector as the MVP zero and includes it in the motion predictor candidate list.
[0166] Figure 12 A video encoding method of an encoding device according to the present application is schematically shown. Figure 12 The shown method can be performed by Figure 1 The shown encoding device. In a detailed example, Figure 12 Steps S1200 to S1230 of the shown method can be performed by a prediction unit of the encoding device, and step S1240 can be performed by a memory of the encoding device.
[0167] The encoding device generates motion information of the target block (step S1200). The encoding device can apply inter prediction to the target block. In a case where the inter prediction is applied to the target block, the encoding device can generate the motion information of the target block by applying at least one of a skip mode, a merge mode, and an adaptive motion vector prediction (AMVP) mode. The motion information can be referred to as first motion information. In a case of the skip mode and the merge mode, the encoding device can generate the motion information of the target block based on motion information of neighboring blocks of the target block. The motion information can include a motion vector and a reference picture index. The motion information can be bi-prediction motion information or uni-prediction motion information. The bi-prediction motion information can include an L0 reference picture index and an L0 motion vector and an L1 reference picture index and an L1 motion vector, and the uni-prediction motion information can include the L0 reference picture index and the L0 motion vector or the L1 reference picture index and the L1 motion vector. L0 indicates a reference picture list L0 (list 0), and L1 indicates a reference picture list L1 (list 1).
[0168] In a case of the AMVP mode, the encoding device can derive a motion vector of the target block using motion vectors of neighboring blocks of the target block as motion vector predictors (MVPs), and generate the motion information including the motion vector and a reference picture index of the motion vector.
[0169] The encoding apparatus derives the prediction samples by performing inter prediction of the target block based on the motion information (step S1210). The encoding apparatus can generate the prediction samples of the target block based on the reference picture index and the motion vector included in the motion information.
[0170] The encoding apparatus generates the reconstructed block based on the prediction samples (step S1220). The encoding apparatus can generate the reconstructed block of the target block based on the prediction samples, or generate a residual signal of the target block and generate the reconstructed block of the target block based on the residual signal and the prediction samples.
[0171] The encoding apparatus generates the corrected motion information of the target block based on the reconstructed block (step S1230). The encoding apparatus can calculate the corrected reference picture index indicating a specific reference picture of the corrected motion information and the corrected motion vector of the specific reference picture by various methods. The corrected motion information can be referred to as second motion information. At least one method including a direct method such as an optical flow (OF) method, a block matching method, a frequency domain method, etc. and an indirect method such as a singularity matching method, a method using statistical properties, etc. can be applied to this method. In addition, the direct method and the indirect method can be applied at the same time.
[0172] For example, the encoding apparatus can generate the corrected motion information by the block matching method. In this case, the encoding apparatus can measure the degree of distortion from the difference between the phase accumulated samples of the reconstructed block of the target block and the reference block and use these as a cost function, and then detect a specific reference block of the reconstructed block. The encoding apparatus can generate the corrected motion information based on the detected specific reference block. In other words, the encoding apparatus can generate the corrected motion information including the corrected reference picture index indicating the specific reference index and the corrected motion vector indicating the specific reference block in the specific reference picture. The encoding apparatus can detect the reference block in which the sum of the absolute values (or square values) of the difference between the phases of the reconstructed block of the target block is the smallest as the specific reference block from among the reference blocks in the specific reference picture, and derive the corrected motion information based on the specific reference block. As a method of representing the sum of the absolute values of the difference, a sum of absolute differences (SAD) can be used. In this case, the sum of the absolute values of the difference can be calculated using Equation 1 described above. In addition, as a method of representing the sum of the absolute values of the difference, a mean square error (MSE) can be used. In this case, the sum of the absolute values of the difference can be calculated using Equation 2 described above.
[0173] In addition, the specific reference picture of the modified motion information can be derived as a reference picture indicated by a reference picture index included in the (first) motion information, and a search area for detecting the specific reference block can be limited to a reference block located within a predetermined area from a reference block derived in the reference picture based on a motion vector related to the reference picture included in the (first) motion information. That is, the encoding apparatus can derive, as the specific reference block, a reference block having the smallest SAD with the reconstructed block among reference blocks located within a predetermined area from a reference block derived in the reference picture based on a motion vector related to the reference picture included in the (first) motion information.
[0174] In addition, in a case where the size of the reconstructed block is greater than a predetermined size, the reconstructed block can be divided into a plurality of sub-reconstructed blocks, and a specific sub-reconstructed block can be derived in the specific reference picture in units of the sub-reconstructed block. In this case, the encoding apparatus can derive the specific reference block based on the derived specific sub-reconstructed block.
[0175] As another example, the encoding apparatus can generate the modified motion information through an OF method. In this case, the encoding apparatus can calculate a motion vector of the modified motion information of the target block based on an assumption that the speed of an object in the target block is uniform, and a sample value of a sample representing the object does not change in an image. The motion vector can be calculated through Equation 6 described above. An area indicating a sample of the object included in the target block can be configured as an area of 3x3 size.
[0176] In a case where the modified motion information is calculated through the above-described method, the encoding apparatus can calculate the modified motion information to have the same format as the (first) motion information. That is, the encoding apparatus can calculate the modified motion information to have the same format as the motion information between bi-prediction motion information and uni-prediction motion information. The bi-prediction motion information can include an L0 reference picture index and an L0 motion vector and an L1 reference picture index and an L1 motion vector, and the uni-prediction motion information can include an L0 reference picture index and an L0 motion vector or an L1 reference picture index and an L1 motion vector. L0 indicates a reference picture list L0 (List 0), and L1 indicates a reference picture list L1 (List 1).
[0177] In addition, bi-prediction can be used for performing in a target picture including a target block, and the encoding apparatus can calculate the modified motion information as bi-prediction motion information by the above-described method. However, in a case where an occlusion occurs on one of the particular reference blocks derived by the bi-prediction motion information after the modified motion information is calculated as the bi-prediction motion information, the encoding apparatus can derive the modified motion information as uni-prediction motion information in addition to motion information of a reference picture list of a particular reference picture including the particular reference block on which the occlusion occurs. It can be determined whether the occlusion occurs when a difference between samples of phases of a reconstructed block of the target block and a reference block derived based on the modified motion information is greater than a particular threshold value. The threshold value can be pre-configured.
[0178] For example, the encoding apparatus calculates the modified motion information as bi-prediction motion information, and in a case where a difference between samples of phases of a particular reference block and a reconstructed block of a target block derived based on an L0 motion vector and an L0 reference picture index based on the modified motion information is greater than a pre-configured threshold value, the encoding apparatus can derive the modified motion information as uni-prediction motion information including an L1 motion vector and an L1 reference picture index.
[0179] For another example, the encoding apparatus calculates the modified motion information as bi-prediction motion information, and in a case where a difference between samples of phases of a particular reference block and a reconstructed block of a target block derived based on an L1 motion vector and an L1 reference picture index based on the modified motion information is greater than a pre-configured threshold value, the encoding apparatus can derive the modified motion information as uni-prediction motion information including an L0 motion vector and an L0 reference picture index.
[0180] In addition, in a case where a difference between samples of phases of a particular reference block and a reconstructed block of a target block derived based on an L0 motion vector and an L0 reference picture index is greater than a pre-configured threshold value and in a case where a difference between samples of phases of a particular reference block and a reconstructed block of a target block derived based on an L1 motion vector and an L1 reference picture index is greater than a pre-configured threshold value, the encoding apparatus can derive motion information of a neighboring block of the target block as the modified motion information, or can not calculate the modified motion information.
[0181] Further, the encoding apparatus can select a particular reference picture indicated by a modified reference picture index included in the modified motion information by various methods.
[0182] For example, the encoding apparatus can select a most recently encoded reference picture among reference pictures included in a reference picture list L0 and generate a modified L0 reference picture index indicating the reference picture. In addition, the encoding apparatus can select a most recently encoded reference picture among reference pictures included in a reference picture list L1 and generate a modified L1 reference picture index indicating the reference picture.
[0183] For another example, the encoding apparatus can select, among the reference pictures included in the L0 reference picture list, a reference picture having a smallest absolute value of a difference in picture order count (POC) from a POC of a target picture, and generate the L0 reference picture list indicating a modification of the reference picture. Also, the encoding apparatus can select, among the reference pictures included in the L1 reference picture list, a reference picture having a smallest absolute value of a difference in picture order count (POC) from a POC of a target picture, and generate the L1 reference picture list indicating a modification of the reference picture.
[0184] For another example, the encoding apparatus can select, among the reference pictures included in each of the reference picture lists, a reference picture belonging to a lowest layer on a hierarchical structure, and generate the reference picture list indicating a modification of the reference picture. The reference picture belonging to the lowest layer can be an I slice or a reference picture encoded by applying a low quantization parameter (QP).
[0185] For another example, the encoding apparatus can select, among the reference pictures included in the reference picture list, a reference picture having a highest reliability of a motion-compensated reference block, and generate the reference picture list indicating a modification of the reference picture. In other words, the encoding apparatus can derive a specific reference block of a reconstructed block of a target block based on the reference pictures included in the reference picture list, and generate the reference picture list indicating a modification of a specific reference picture including the derived specific reference block.
[0186] The above-described methods of generating the modified reference picture list of the modified motion information can be independently applied, or the methods can be combinedly applied.
[0187] Also, although not shown, the encoding apparatus can generate the modified motion information of a target block based on an original block of the target block. The encoding apparatus can derive a specific reference block of the original block among reference blocks included in a reference picture, and generate the modified motion information indicating the derived specific reference block.
[0188] The encoding device updates the motion information of the target block based on the corrected motion information (step S1240). The encoding device can store the corrected motion information and update the motion information of the target block. The encoding device can update the motion information of the target block by replacing the motion information used to predict the target block with the corrected motion information. Alternatively, the encoding device can update the motion information of the target block by storing all the motion information used to predict the target block and the corrected motion information. The updated motion information can be used for the motion information of the next block adjacent to the target block. For example, if a merge mode is applied to the next block adjacent to the target block, the merge candidate list of the next block may include the target block. When the motion information of the target block is stored by replacing the motion information used to predict the target block with the corrected motion information, the merge candidate list of the next block may include the target block indicating the corrected motion information. In addition, since all the motion information used to predict the target block and the corrected motion information are stored in the motion information of the target block, the merge candidate list of the next block may include the target block indicating the motion information used to predict the target block and the target block indicating the corrected motion information. The target block indicating the corrected motion information may be inserted into the merge candidate list as a spatially adjacent candidate block or as a temporally adjacent candidate block.
[0189] For example, applying the AMVP mode to the next block adjacent to the target block is similar to the method described above in the merge mode, where updated motion information of the target block adjacent to the next block is inserted into the merge candidate list as spatial or temporal proximity motion information. Similarly, the method of inserting updated motion information of neighboring blocks into the motion vector predictor candidate list of the next block can be applied as spatial or temporal motion vector predictor candidates. That is, the encoding device can generate a motion vector predictor candidate list that includes updated motion information of the target block adjacent to the next block.
[0190] For example, if the updated motion information for the next neighboring target block only includes the corrected motion information, the motion vector predictor candidate list can include the corrected motion information as a spatial motion vector predictor candidate.
[0191] For example, if the updated motion information of the next neighboring target block includes the corrected motion information and the existing motion information of the target block, the motion vector predictor candidate list may include existing motion information selected from the existing motion information of the next neighboring target block according to specific conditions, and corrected motion information selected from the corrected motion information of the next neighboring target block according to specific conditions, as corresponding spatial motion vector predictor candidates.
[0192] Further, the motion vector predictor candidate list of the next block can include the updated motion information of the collocated block in the reference picture of the next block at the same position as the position of the next block and the updated motion information of the neighboring block of the same position block. For example, in a case where the updated motion information of the collocated block at the same position as the position of the next block and the updated motion information of the neighboring block of the same position block include only the corrected motion information of each block, the motion vector predictor candidate list can include the corrected motion information as the temporal motion vector predictor candidate list.
[0193] For another example, in a case where the updated motion information includes the respective corrected motion information of the same position block, the neighboring block of the same position block, and the existing motion information of the same position block, the motion vector predictor candidate list can include the corrected motion information selected from among the corrected motion information according to a specific condition and the existing motion information selected from among the existing motion information according to a specific condition, respectively, as separate spatial motion vector predictors.
[0194] Further, the stored motion information and the motion information for transmission can have different resolutions. For example, the encoding apparatus encodes and transmits information for the motion information and stores the corrected motion information with respect to the motion information adjacent to the target block, and the unit of the motion vector included in the motion information can have a unit of ¼ partial sample, and the unit of the motion vector of the corrected motion information can represent a unit of 1 / 8 partial sample or 1 / 16 partial sample.
[0195] Further, although not shown, the encoding apparatus can determine whether to update the (first) motion information of the target block by performing a comparison process between the (first) motion information and the corrected motion information based on the reconstructed block of the target block. For example, the encoding apparatus can determine whether to update the target block by comparing the data amount of the residual signal of the specific reference block and the reconstructed block of the target block derived based on the corrected motion information with the data amount of the residual signal of the reference block and the reconstructed block derived based on the motion information. Among these data amounts, in a case where the data amount of the residual signal of the specific reference block and the reconstructed block is smaller, the encoding apparatus can determine to update the motion information and the target block. In addition, among these data amounts, in a case where the data amount of the residual signal of the specific reference block and the reconstructed block is not smaller, the encoding apparatus can determine not to update the motion information and the target block.
[0196] In addition, the encoding apparatus can determine whether to update the (first) motion information of the target block by performing a comparison process between the original block of the target block and the modified motion information. For example, the encoding apparatus can determine whether to update the target block by comparing the amount of data of a certain reference block derived based on the modified motion information and the residual signal of the original block of the target block with the amount of data of a reference block derived based on the motion information and the residual signal of the original block. Among these amounts of data, in the case where the amount of data of the certain reference block and the residual signal of the original block is smaller, the encoding apparatus can determine to update the motion information and the target block. In addition, among these amounts of data, in the case where the amount of data of the certain reference block and the residual signal of the original block is not smaller, the encoding apparatus can determine not to update the motion information and the target block.
[0197] In addition, the encoding apparatus can generate and encode additional information indicating whether to update and output it through a bitstream. For example, the additional information indicating whether to update can be referred to as an update flag. The case where the update flag is 1 can indicate that the motion information is updated, and the case where the update flag is 0 can indicate that the motion information is not updated. For example, the update flag can be transmitted in units of a PU. Alternatively, the update flag can be transmitted in units of a CU, in units of a CTU, or in units of a slice, and can be transmitted through a higher level such as a unit of a picture parameter set (PPS) or a unit of a sequence parameter set (SPS).
[0198] In addition, the encoding apparatus can further generate and encode and output motion vector update difference information indicating a difference between an existing motion vector and the modified motion information. The motion vector update difference information can be transmitted in units of a PU.
[0199] Although not shown, the encoding apparatus can encode and output information about residual samples of the target block. The information about the residual samples can include transform coefficients of the residual samples.
[0200] Figure 13 A video decoding method of a decoding apparatus according to the present application is schematically shown. Figure 13 The method shown can be performed by Figure 2 The decoding apparatus shown. In a detailed example, Figure 13 Step S1300 of the decoding apparatus can be performed by an entropy decoding unit of the decoding apparatus, and steps S1310 to S1340 can be performed by a prediction unit of the decoding apparatus, and step S1350 can be performed by a memory of the decoding apparatus.
[0201] The decoding device obtains information for inter prediction of the target block through the bitstream (step S1300). Inter prediction or intra prediction can be applied to the target block. In a case where inter prediction is applied to the target block, the decoding device can obtain the information for inter prediction of the target block through the bitstream. In addition, the decoding device can obtain motion vector update difference information indicating a difference between an existing motion vector of the target block and the corrected motion information through the bitstream. Furthermore, the decoding device can obtain additional information whether to update the target block through the bitstream. For example, the additional information indicating whether to update can be referred to as an update flag.
[0202] The decoding device derives motion information of the target block based on the information for inter prediction (step S1310). The motion information can be referred to as first motion information. The information for inter prediction can indicate a mode applied to the target block among a skip mode, a merge mode, and an adaptive motion vector prediction (AMVP) mode. In a case where the skip mode or the merge mode is applied to the target block, the decoding device can generate a merge candidate list including neighboring blocks of the target block and obtain a merge index indicating a neighboring block among the neighboring blocks included in the merge candidate list. The merge index can be included in the information for inter prediction. The decoding device can derive motion information of the neighboring block indicated by the merge index as the motion information of the target block.
[0203] In a case where the AMVP mode is applied to the target block, the decoding device can generate a list based on neighboring blocks of the target block, similar to the merge mode. The decoding device can generate an index indicating a neighboring block among the neighboring blocks included in the generated list and a motion vector difference (MVD) between a motion vector of the neighboring block indicated by the index and a motion vector of the target block. The index and the MVD can be included in the information for inter prediction. The decoding device can generate the motion information of the target block based on the motion vector of the neighboring block indicated by the index and the MVD.
[0204] The motion information can include a motion vector and a reference picture index. The motion information can be bi-prediction motion information or uni-prediction motion information. The bi-prediction motion information can include an L0 reference picture index and an L0 motion vector and an L1 reference picture index and an L1 motion vector, and the uni-prediction motion information can include an L0 reference picture index and an L0 motion vector or an L1 reference picture index and an L1 motion vector. L0 indicates a reference picture list L0 (list 0), and L1 indicates a reference picture list L1 (list 1).
[0205] The decoding device derives a prediction sample by performing inter prediction of the target block based on the motion information (step S1320). The decoding device can generate a prediction sample of the target block based on a reference picture index and a motion vector included in the motion information.
[0206] The decoding device generates a reconstructed block based on the prediction sample (step S1330). In a case where the skip mode is applied to the target block, the decoding device can generate the reconstructed block of the target block based on the prediction sample. In a case where the merge mode or the AMVP mode is applied to the target block, the decoding device can generate a residual signal of the target block through the bitstream and generate the reconstructed block of the target block based on the residual signal and the prediction sample.
[0207] The encoding device derives the corrected motion information of the target block based on the reconstructed block (step S1340). The decoding device can calculate the corrected reference picture index indicating the corrected motion information of a certain reference picture and the corrected motion vector of the certain reference picture through various methods. The corrected motion information can be referred to as second motion information. At least one method including a direct method such as an optical flow (OF) method, a block matching method, a frequency domain method, etc. and an indirect method such as a singularity matching method, a method using statistical properties, etc. can be applied to this method. In addition, the direct method and the indirect method can be applied at the same time.
[0208] For example, the decoding device can generate the corrected motion information through the block matching method. In this case, the decoding device can measure a degree of distortion by accumulating a difference between samples according to phases of the reconstructed block of the target block and the reference block and use these as a cost function, and then detect a certain reference block of the reconstructed block. The decoding device can generate the corrected motion information based on the detected certain reference block. In other words, the decoding device can generate the corrected motion information including the corrected reference picture index indicating the certain reference index and the corrected motion vector indicating the certain reference block in the certain reference picture. The decoding device can detect a reference block in which a sum of absolute values (or square values) of differences between samples according to phases of the reconstructed block of the target block is the smallest as the certain reference block among the reference blocks in the certain reference picture, and derive the corrected motion information based on the certain reference block. As a method of representing the sum of absolute values of the differences, a sum of absolute differences (SAD) can be used. In this case, the sum of absolute values of the differences can be calculated using Equation 1 above. In addition, as a method of representing the sum of absolute values of the differences, a mean square error (MSE) can be used. In this case, the sum of absolute values of the differences can be calculated using Equation 2 above.
[0209] In addition, the particular reference picture of the modified motion information can be derived as a reference picture indicated by a reference picture index included in the (first) motion information, and a search area for detecting the particular reference block can be limited to reference blocks located within a predetermined area from the particular reference block derived in the particular reference picture based on a motion vector related to the reference picture included in the (first) motion information. That is, the decoding device can derive, as the particular reference block, a reference block having the smallest SAD with the reconstructed block among reference blocks located within a predetermined area from the particular reference block derived in the particular reference picture based on a motion vector related to the reference picture included in the motion information.
[0210] In addition, in a case where the size of the reconstructed block is greater than a predetermined size, the reconstructed block can be divided into a plurality of sub-reconstructed blocks, and a sub-reconstructed block can be specified in the particular reference picture in units of the sub-reconstructed blocks. In this case, the decoding device can derive the particular reference block based on the derived particular sub-reconstructed block.
[0211] For example, the decoding device can generate the modified motion information by the OF method. In this case, the decoding device can calculate the modified motion vector of the modified motion information of the target block based on an assumption that the speed of an object in the target block is uniform and a sample value of a sample representing the object does not change in an image. The motion vector can be calculated by Equation 6 described above. A region indicating samples of the object included in the target block can be configured as a region of 3x3 size.
[0212] In a case where the modified motion information is calculated by the above-described method, the decoding device can calculate the modified motion information to have the same format as the (first) motion information. That is, the decoding device can calculate the modified motion information to have the same format as the motion information between bi-prediction motion information and uni-prediction motion information. The bi-prediction motion information can include an L0 reference picture index and an L0 motion vector and an L1 reference picture index and an L1 motion vector, and the uni-prediction motion information can include an L0 reference picture index and an L0 motion vector or an L1 reference picture index and an L1 motion vector. L0 indicates a reference picture list L0 (List 0), and L1 indicates a reference picture list L1 (List 1).
[0213] In addition, bi-prediction can be used for performing in a target picture including the target block, and the decoding device can calculate the modified motion information as bi-prediction motion information by the above method. However, in a case where an occlusion occurs on one of the reference blocks derived by the bi-prediction motion information after the modified motion information is calculated as the bi-prediction motion information, the decoding device can derive the modified motion information as uni-prediction motion information in addition to the motion information of the reference picture list of the particular reference picture including the particular reference block where the occlusion occurs. Whether the occlusion occurs can be determined to occur when a difference between samples of phases of the particular reference block and the reconstructed block of the target block derived according to the modified motion information is greater than a particular threshold. The threshold can be pre-configured.
[0214] For example, the decoding device calculates the modified motion information as bi-prediction motion information, and in a case where a difference between samples of phases of the particular reference block and the reconstructed block of the target block derived according to the L0 motion vector and the L0 reference picture index based on the modified motion information is greater than a pre-configured threshold, the encoding device can derive the modified motion information as uni-prediction motion information including the L1 motion vector and the L1 reference picture index.
[0215] For another example, the decoding device calculates the modified motion information as bi-prediction motion information, and in a case where a difference between samples of phases of the particular reference block and the reconstructed block of the target block derived according to the L1 motion vector and the L1 reference picture index based on the modified motion information is greater than a pre-configured threshold, the decoding device can derive the modified motion information as uni-prediction motion information including the L0 motion vector and the L0 reference picture index.
[0216] In addition, in a case where a difference between samples of phases of the particular reference block and the reconstructed block of the target block derived according to the L0 motion vector and the L0 reference picture index is greater than a pre-configured threshold and in a case where a difference between samples of phases of the particular reference block and the reconstructed block of the target block derived according to the L1 motion vector and the L1 reference picture index is greater than a pre-configured threshold, the decoding device can derive motion information of a neighboring block of the target block as the modified motion information, or can not calculate the modified motion information.
[0217] Further, the decoding device can select the particular reference picture indicated by the modified reference picture index included in the modified motion information by various methods.
[0218] For example, the decoding device can select a most recently encoded reference picture among the reference pictures included in the reference picture list L0, and generate a modified L0 reference picture index indicating the reference picture. In addition, the decoding device can select a most recently encoded reference picture among the reference pictures included in the reference picture list L1, and generate a modified L1 reference picture index indicating the reference picture.
[0219] For another example, the decoding device can select, among the reference pictures included in the L0 reference picture index, a reference picture having a smallest absolute value of a difference in picture order count (POC) from a POC of the target picture, and generate the L0 reference picture index indicating a modification of the reference picture. Also, the decoding device can select, among the reference pictures included in the L1 reference picture index, a reference picture having a smallest absolute value of a difference in picture order count (POC) from a POC of the target picture, and generate the L1 reference picture index indicating a modification of the reference picture.
[0220] For another example, the decoding device can select, among the reference pictures included in each of the reference picture lists, a reference picture belonging to a lowest layer on a hierarchical structure, and generate the reference picture index indicating a modification of the reference picture. The reference picture belonging to the lowest layer can be an I slice or a reference picture encoded by applying a low quantization parameter (QP).
[0221] For another example, the decoding device can select, among the reference picture lists, a reference picture including a reference block having a highest reliability of motion compensation, and generate the reference picture index indicating a modification of the reference picture. In other words, the decoding device can derive a specific reference block of a reconstructed block of a target block based on a reference picture included in the reference picture lists, and generate the reference picture index indicating a modification of a specific reference picture including the derived specific reference block.
[0222] The above-described methods of generating the modified reference picture index of the modified motion information can be independently applied, or the methods can be combined and applied.
[0223] In addition, the decoding device can obtain, through the bitstream, motion vector update difference information indicating a difference between an existing motion vector of the target block and the modified motion vector. In this case, the decoding device can derive the modified motion information by summing the (first) motion information of the target block and the motion vector update difference information. The motion vector update difference information can be transmitted in units of the above-described PUs.
[0224] The decoding device updates the motion information of the target block based on the modified motion information (step S1350). The decoding device can store the modified motion information and update the motion information of the target block. The decoding device can update the motion information of the target block by replacing the motion information used to predict the target block with the modified motion information. Also, the decoding device can update the motion information of the target block by storing all of the motion information used to predict the target block and the modified motion information. The updated motion information can be used for motion information of a next block adjacent to the target block.
[0225] For example, in case that the updated motion information of the target block adjacent to the next block is composed of only the corrected motion information, the merge candidate list of the next block can include the corrected motion information as a spatial motion vector predictor candidate.
[0226] For example, in case that the updated motion information of the target block adjacent to the next block is composed of only the corrected motion information, the merge candidate list of the next block can include the corrected motion information as a spatial motion vector predictor candidate.
[0227] For example, in case that the updated motion information of the target block adjacent to the next block is composed of only the corrected motion information, the merge candidate list of the next block can include the corrected motion information as a spatial motion vector predictor candidate.
[0228] For example, in case that the updated motion information of the target block adjacent to the next block is composed of only the corrected motion information, the merge candidate list of the next block can include the corrected motion information as a spatial motion vector predictor candidate.
[0229] For example, in case that the updated motion information of the target block adjacent to the next block is composed of only the corrected motion information, the merge candidate list of the next block can include the corrected motion information as a spatial motion vector predictor candidate.
[0230] For example, in a case where the updated motion information includes a same position block, a corresponding corrected motion information of a neighboring block of the same position block, and an existing motion information of the same position block, the motion vector predictor candidate list can include the corrected motion information selected from among the corrected motion information according to a certain condition and the existing motion information selected from among the existing motion information according to the certain condition as separate spatial motion vector predictor candidates, respectively.
[0231] In addition, the stored motion information and the motion information for transmission can have different resolutions. For example, the decoding device decodes and transmits information for the motion information, and stores the corrected motion information with respect to the motion information neighboring the target block, and the unit of the motion vector included in the motion information can have a unit of ¼ partial sample, and the unit of the motion vector of the corrected motion information can represent a unit of 1 / 8 partial sample or 1 / 16 partial sample.
[0232] In addition, although not shown, the decoding device can determine whether to update the (first) motion information of the target block by performing a comparison process between the (first) motion information and the corrected motion information based on the reconstructed block of the target block. For example, the decoding device can determine whether to update the target block by comparing a data amount of a certain reference block derived based on the corrected motion information and a residual signal of the reconstructed block of the target block with a data amount of a reference block derived based on the motion information and a residual signal of the reconstructed block. Among the data amounts, in a case where the data amount of the certain reference block and the residual signal of the reconstructed block is smaller, the decoding device can determine to update the motion information and the target block. Also, among the data amounts, in a case where the data amount of the certain reference block and the residual signal of the reconstructed block is not smaller, the decoding device can determine not to update the motion information and the target block.
[0233] In addition, the decoding device can obtain additional information indicating whether to update the target block through a bitstream. For example, the additional information indicating whether to update can be referred to as an update flag. A case where the update flag is 1 can indicate that the motion information is updated, and a case where the update flag is 0 can indicate that the motion information is not updated. For example, the update flag can be transmitted in a unit of a PU. Alternatively, the update flag can be transmitted in a unit of a CU, in a unit of a CTU, or in a unit of a slice, and can be transmitted through a higher level such as a unit of a picture parameter set (PPS) or a unit of a sequence parameter set (SPS).
[0234] According to the present application described above, after a decoding process of a target block, corrected motion information of the target block is calculated, and can be updated to more accurate motion information, whereby overall coding efficiency can be improved.
[0235] In addition, according to the present application, motion information of a next block neighboring the target block can be derived based on the updated motion information of the target block, and propagation of distortion can be reduced, whereby overall coding efficiency can be improved.
[0236] In the above embodiments, the method is described as a series of steps or blocks based on the flowchart. However, this disclosure is not limited to the order of these steps. Some steps may be performed simultaneously or in a different order than described above. Furthermore, those skilled in the art will understand that the steps shown in the flowchart are not exclusive. It will be understood that other steps may be included or one or more steps in the flowchart may be deleted without affecting the scope of this disclosure.
[0237] The method described above can be implemented in software. The encoding and / or decoding apparatus according to this disclosure can be included in apparatus for performing image processing, for example, a TV, computer, smartphone, set-top box, or display device.
[0238] When the embodiments of this disclosure are implemented in software, the above methods can be implemented by modules (processes, functions, etc.) that perform the above functions. These modules can be stored in memory and executed by a processor. The memory can be internal or external to the processor, and the memory can be connected to the processor using various well-known means. The processor may include application-specific integrated circuits (ASICs), other chipsets, logic circuits, and / or data processing devices. The memory may include ROM (read-only memory), RAM (random access memory), flash memory, memory cards, storage media, and / or other storage devices.
Claims
1. An image decoding method performed by a decoding device, the image decoding method comprising the steps of: deriving motion information of a first target block; deriving corrected motion information of the first target block based on the motion information and bi-prediction; obtaining, from a bitstream, information on inter prediction of a second target block, wherein the information on inter prediction includes mode information indicating that a merge mode is applied to the second target block and merge index information of the second target block; determining that the merge mode is applied to the second target block based on the mode information; configuring a merge candidate list of the second target block based on spatial neighboring blocks and temporal neighboring blocks of the second target block; selecting one of merge candidates constituting the merge candidate list based on the merge index information; deriving motion information of the second target block based on the selected merge candidate; generating prediction samples of the second target block by performing the inter prediction based on the motion information of the second target block; and generating reconstructed samples based on the prediction samples, wherein the merge candidates include a spatial motion information candidate and a temporal motion information candidate, wherein the spatial motion information candidate is derived based on the spatial neighboring blocks and the temporal motion information candidate is derived based on the temporal neighboring blocks, wherein when the temporal neighboring blocks correspond to the first target block, the temporal motion information candidate is derived by using the corrected motion information of the first target block, and wherein when the spatial neighboring blocks correspond to the first target block, the spatial motion information candidate is derived by using the motion information of the first target block.
2. An image encoding method performed by an encoding device, the image encoding method comprising the steps of: deriving motion information of a first target block; deriving corrected motion information of the first target block based on the motion information and bi-prediction; determining that a merge mode is applied to a second target block; configuring a merge candidate list of the second target block based on spatial neighboring blocks and temporal neighboring blocks of the second target block; selecting one of merge candidates constituting the merge candidate list; generating mode information indicating that the merge mode is applied to the second target block; generating merge index information of the second target block indicating the selected merge candidate; and encoding information on inter prediction including the mode information and the merge index information, wherein the merge candidates include a spatial motion information candidate and a temporal motion information candidate, wherein the spatial motion information candidate is derived based on the spatial neighboring blocks and the temporal motion information candidate is derived based on the temporal neighboring blocks, wherein when the temporal neighboring blocks correspond to the first target block, the temporal motion information candidate is derived by using the corrected motion information of the first target block, and wherein when the spatial neighboring blocks correspond to the first target block, the spatial motion information candidate is derived by using the motion information of the first target block.
3. A transmission method of data of an image, the transmission method comprising the steps of: obtaining a bitstream of the image, wherein the bitstream is generated based on the following steps: deriving motion information of a first target block; deriving a corrected motion information of the first target block based on the motion information and bi-prediction; determining to apply a merge mode to a second target block; configuring a merge candidate list of the second target block based on spatial neighboring blocks and temporal neighboring blocks of the second target block; selecting one of the merge candidates constituting the merge candidate list; generating mode information indicating that the merge mode is applied to the second target block; generating merge index information of the second target block indicating the selected merge candidate; and encoding information on inter prediction including the mode information and the merge index information; and transmitting the data including the bitstream, wherein the merge candidate includes a spatial motion information candidate and a temporal motion information candidate, wherein the spatial motion information candidate is derived based on the spatial neighboring blocks and the temporal motion information candidate is derived based on the temporal neighboring blocks, wherein when the temporal neighboring blocks correspond to the first target block, the temporal motion information candidate is derived by using the corrected motion information of the first target block, and wherein when the spatial neighboring blocks correspond to the first target block, the spatial motion information candidate is derived by using the motion information of the first target block.
Citation Information
Patent Citations
Method of motion information coding
CN106031170A
Method for encoding and decoding image information and device using same
CN106101723A