Image Encoding / Decoding Method and Apparatus

The image encoding/decoding apparatus addresses performance gaps by correcting motion vector prediction values, enhancing coding efficiency through motion vector adjustment offsets and candidate list management.

JP7711293B2Active Publication Date: 2025-07-22INST OF IMAGE TECH INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2024185054
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2018-09-24
Filing Date
2024-10-21
Publication Date
2025-07-22
Estimated Expiration
2039-09-24

AI Technical Summary

Technical Problem

Conventional image encoding/decoding methods lack improvements in performance, particularly in image encoding and decoding, which are essential for efficient multimedia data processing in various systems.

Method used

An image encoding/decoding apparatus that corrects motion vector prediction values using an adjustment offset, constructing a motion information prediction candidate list, selecting a prediction candidate index, and deriving a prediction motion vector adjustment offset.

Benefits of technology

Enhances coding performance by efficiently obtaining predicted motion vectors, improving the efficiency of image processing systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007711293000001
    Figure 0007711293000001
  • Figure 0007711293000002
    Figure 0007711293000002
  • Figure 0007711293000003
    Figure 0007711293000003
Patent Text Reader

Abstract

To provide an image encoding and decoding method and device that correct a motion vector prediction value by using an adjustment offset.SOLUTION: An image encoding method includes the following steps of: constituting a motion information candidate list for an object block; selecting a prediction candidate index from the motion information candidate list; guiding an offset for motion vector adjustment; and restoring the motion vector of the object block through the prediction motion vector restored based on the offset.SELECTED DRAWING: Figure 19
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an image encoding / decoding method and apparatus.

Background Art

[0002] With the spread of the Internet and mobile terminals and the development of information and communication technologies, the use of multimedia data has increased rapidly. Therefore, in order to perform various services and operations through image prediction in various systems, the need to improve the performance and efficiency of image processing systems has increased significantly, but the results of research and development that can respond to such an atmosphere are insufficient.

[0003] Thus, in the conventional image encoding / decoding method and apparatus, improvement in performance for image processing, particularly image encoding or image decoding, is required.

Summary of the Invention

Problems to be Solved by the Invention

[0004] The present invention is for solving such problems, and an object thereof is to provide an image encoding / decoding apparatus that corrects a motion vector prediction value using an adjustment offset.

Means for Solving the Problems

[0005] A method for decoding an image according to an embodiment of the present invention for achieving the above object includes steps of constructing a motion information prediction candidate list for a target block, selecting a prediction candidate index, deriving a prediction motion vector adjustment offset, and restoring the motion information of the target block.

[0006] Here, the step of setting the motion information prediction candidate list may further include a step of including in the candidate group when candidates obtained based on already included candidates and offset information and candidates obtained based on new candidates and offset information do not overlap.

[0007] Here, the step of deriving the predicted motion vector adjustment offset may further include a step of being derived based on an offset application flag and / or offset selection information.

Advantages of the Invention

[0008] When using the inter-screen prediction according to the present invention described above, the coding performance can be improved by efficiently obtaining the predicted motion vector.

Brief Description of the Drawings

[0009]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Embodiments for Carrying Out the Invention

[0010] The image encoding / decoding method and apparatus according to the present invention can configure a prediction motion candidate list of a target block, derive a prediction motion vector based on a prediction candidate index from the motion candidate list, restore prediction motion vector adjustment offset information, and restore a motion vector of the target block based on the prediction motion vector and the prediction motion vector adjustment offset information.

[0011] In the image encoding / decoding method and apparatus according to the present invention, the motion candidate list can include at least one of a spatial candidate, a temporal candidate, a statistical candidate, or a combined candidate.

[0012] In the image encoding / decoding method and apparatus according to the present invention, the prediction motion vector adjustment offset can be determined based on at least one of an offset application flag or offset selection information.

[0013] In the image encoding / decoding method and apparatus according to the present invention, information regarding whether or not to assist the prediction motion vector adjustment offset information may be included in at least one of a sequence, a picture, a sub-picture, a slice, a tile, or a block.

[0014] In the image encoding / decoding method and apparatus according to the present invention, when the target block is encoded in the merge mode, the motion vector of the target block is restored using a zero vector, and when the target block is encoded in the competition mode, the motion vector of the target block can be restored using a motion vector difference value.

[0015] [Embodiment for Carrying Out the Invention] Since the present invention can be subjected to various changes and can have various embodiments, specific embodiments are illustrated in the drawings and described in detail in the detailed description. However, this is not intended to limit the present invention to the specific embodiments, and it should be understood to include all changes, equivalents, and alternatives included in the spirit and technical scope of the present invention.

[0016] Terms such as "first", "second", etc. can be used to describe various components, but these components should not be limited by the above terms. These terms are only used for the purpose of distinguishing one component from another. For example, without departing from the scope of the rights of the present invention, the first component can be named the second component, and similarly, the second component can also be named the first component. The term "and / or" includes a combination of a plurality of related described items or any of the plurality of related described items.

[0017] When a component is "connected" or "attached" to another component, it should be understood that it may be directly connected or attached to the other component, but there may also be another component intervening between them. In contrast, when a component is "directly connected" or "directly attached" to another component, it should be understood that there is no other component intervening between them.

[0018] The terms used in the present invention are merely used to describe specific embodiments and are not intended to limit the present invention. Singular expressions include plural expressions unless the context clearly dictates otherwise. In the present invention, terms such as "including" or "having" are intended to specify the presence of the features, numbers, steps, operations, components, parts, or combinations thereof described in the specification, and should not be construed as precluding the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.

[0019] Unless otherwise defined, all terms used herein, including technical or scientific terms, are to be construed as having the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Terms defined in commonly used dictionaries should be construed as having a meaning consistent with the context of the relevant art and should not be construed in an idealized or overly formal sense unless clearly defined in the present invention.

[0020] Generally, it can be composed of one or more color spaces according to the color format of the image. According to the color format, it can be composed of one or more pictures with a certain size, or one or more pictures with different sizes. As an example, color formats such as 4:4:4, 4:2:2, 4:2:0, Monochrome (composed of only Y) can be supported in the YCbCr color configuration. As an example, in the case of YCbCr4:2:0, it can be composed of one luminance component (in this example, Y) and two chrominance components (in this example, Cb / Cr). At this time, the composition ratio of the chrominance component and the luminance component can be 1:2 horizontally. As an example, in the case of 4:4:4, the horizontal and vertical can have the same composition ratio. When composed of one or more color spaces as in the above examples, the picture can be divided into each color space.

[0021] Images can be classified into I, P, B, etc. according to the image type (e.g., picture type, sub-picture type, slice type, tile type, block type, etc.). The I image type can mean an image that is encoded by itself without using a reference picture. The P image type can mean an image that is encoded using a reference picture but only allows forward prediction. The B image type can mean an image that is encoded using a reference picture and allows forward / backward prediction. However, depending on the encoding settings, some of the above types may be combined (combining P and B) or other configured image types may be supported.

[0022] The various encoding / decoding information generated in the present invention can be processed explicitly or implicitly. Here, explicit processing can generate encoding / decoding information in sequences, pictures, sub-pictures, slices, tiles, bricks, blocks, sub-blocks, etc., and record this in a bitstream. It can be understood that the decoder parses the relevant information in units at the same level as the encoder and restores it to the decoding information. Here, implicit processing can be understood as processing the encoding / decoding information by the same processes and rules in the encoder and the decoder.

[0023] FIG. 1 is a conceptual diagram showing an image encoding and decoding system according to an embodiment of the present invention.

[0024] Referring to FIG. 1, the image encoding device 105 and the decoding device 100 can be user terminals such as a personal computer (PC), a notebook computer, a portable terminal (PDA: Personal Digital Assistant), a portable multimedia player (PMP: Portable Multimedia Player), a PlayStation Portable (PSP), a wireless communication terminal, a smartphone, or a TV, or server terminals such as an application server or a service server, and include various devices equipped with a communication device such as a communication modem for communicating with various devices or a wired / wireless communication network, a memory (memory, 120, 125) for storing various programs and data for inter or intra prediction for encoding or decoding an image, or a processor (processor, 110, 115) for executing a program to perform operations and control, etc.

[0025] In addition, the image encoded into a bitstream by the image encoding device 105 can be transmitted to the image decoding device 100 via a wired or wireless communication network (Network) such as the Internet, a short-range wireless communication system, a wireless LAN network, a WiBro network, or a mobile communication network in real-time or non-real-time, or via various communication interfaces such as a cable or a universal serial bus (USB: Universal Serial Bus). Then, it is decoded by the image decoding device 100 and restored to an image so that it can be played back. Also, the image encoded into a bitstream by the image encoding device 105 can be transmitted from the image encoding device 105 to the image decoding device 100 via a computer-readable recording medium.

[0026] The above-described image encoding device and image decoding device can be separate devices respectively, but depending on the implementation, they may be made into one image encoding / decoding device. In that case, some configurations of the image encoding device are technical elements substantially the same as some configurations of the image decoding device, and can include at least the same structure or can be realized to perform at least the same function.

[0027] Therefore, in the following detailed explanations of the following technical elements and their operating principles, etc., duplicate explanations of corresponding technical elements will be omitted. Also, since the image decoding device corresponds to a computing device that applies the image encoding method performed by the image encoding device to decoding, in the following explanations, the image encoding device will be mainly described.

[0028] The computing device can include a memory that stores programs and software modules for implementing the image encoding method and / or the image decoding method, and a processor connected to the memory that executes the programs. Here, the image encoding device may sometimes be called an encoder, and the image decoding device may sometimes be called a decoder.

[0029] FIG. 2 is a block configuration diagram showing an image encoding device according to an embodiment of the present invention.

[0030] Referring to FIG. 2, the image encoding apparatus 20 can include a prediction unit 200, a subtraction unit 205, a conversion unit 210, a quantization unit 215, an inverse quantization unit 220, an inverse conversion unit 225, an addition unit 230, a filter unit 235, an encoded picture buffer 240, and an entropy encoding unit 245.

[0031] The prediction unit 200 can be implemented using a prediction module, which is a software module, and can generate a prediction block for a block to be encoded using an intra prediction method or an inter prediction method within the screen. The prediction unit 200 can predict a target block to be currently encoded in the image and generate a prediction block. In other words, the prediction unit 200 can predict the pixel value of each pixel of the target block to be encoded in the image by intra prediction or inter prediction, and generate a prediction block having the predicted pixel value of each generated pixel. Further, the prediction unit 200 can transmit information necessary for generating the prediction block, such as information regarding the prediction mode, such as the intra prediction mode or the inter prediction mode, to the encoding unit so that the encoding unit can encode the information regarding the prediction mode. At this time, the processing unit for which prediction is performed and the processing unit for which the prediction method and specific content are determined can be determined according to the encoding settings. For example, the prediction method, the prediction mode, etc. can be determined in the prediction unit, and the execution of the prediction can be performed in the conversion unit. Also, when using a specific encoding mode, it is also possible to directly encode the original block without generating a prediction block through the prediction unit and transmit it to the decoding unit.

[0032] The in-picture prediction unit can have a directional prediction mode such as a horizontal or vertical mode used according to the prediction direction, and a non-directional prediction mode such as DC or Planar that uses methods such as averaging or interpolation of reference pixels. An in-picture prediction mode candidate group can be configured through the directional and non-directional modes, and any of various candidates such as 35 prediction modes (33 directional + 2 non-directional), 67 prediction modes (65 directional + 2 non-directional), or 131 prediction modes (129 directional + 2 non-directional) can be used as the candidate group.

[0033] The in-picture prediction unit can include a reference pixel configuration unit, a reference pixel filter unit, a reference pixel interpolation unit, a prediction mode determination unit, a prediction block generation unit, and a prediction mode encoding unit. The reference pixel configuration unit can configure pixels adjacent to the target block that belong to adjacent blocks centered on the target block as reference pixels for in-picture prediction. Depending on the encoding settings, one of the most adjacent reference pixel lines can be configured as a reference pixel, or one of the other adjacent reference pixel lines can be configured as a reference pixel, and multiple reference pixel lines can be configured as reference pixels. If some of the reference pixels are not available, reference pixels can be generated using the available reference pixels, and if all are not available, reference pixels can be generated using a predetermined value (e.g., the median value of the pixel value range represented by the bit depth).

[0034] The reference pixel filter unit of the in-picture prediction unit can perform filtering on the reference pixels for the purpose of reducing the remaining degradation through the encoding process. At this time, the filter used can be a low-pass filter such as a 3-tap filter [1 / 4, 1 / 2, 1 / 4] or a 5-tap filter [2 / 16, 3 / 16, 6 / 16, 3 / 16, 2 / 16]. Based on the encoding information (e.g., block size, shape, prediction mode, etc.), the application or non-application of filtering and the type of filtering can be determined.

[0035] The reference pixel interpolation unit of the in-picture prediction unit can generate pixels in fractional units through the linear interpolation process of reference pixels according to the prediction mode, and based on the coding information, the interpolation filter to be applied can be determined. At this time, the interpolation filter to be used may include a 4-tap Cubic filter, a 4-tap Gaussian filter, a 6-tap Wiener filter, an 8-tap Kalman filter, and the like. Although interpolation is generally performed separately from the process of performing a low-pass filter, the filters applied to the two processes can also be integrated into one to perform a filtering process.

[0036] The prediction mode determination unit of the in-picture prediction unit can select at least one optimal prediction mode from the prediction mode candidate group in consideration of the coding cost, and the prediction block generation unit can generate a prediction block using the prediction mode. The optimal prediction mode can be encoded based on the predicted value by the prediction mode encoding unit. At this time, the prediction information can be adaptively encoded according to whether the predicted value holds or not.

[0037] In the in-picture prediction unit, the predicted value can be set as the MPM (Most Probable Mode), and a part of the modes among all the modes belonging to the prediction mode candidate group can be configured as the MPM candidate group. The MPM candidate group may include a predetermined prediction mode (for example, DC, Planar, vertical, horizontal, diagonal mode, etc.) or the prediction modes of spatially adjacent blocks (for example, left, upper, upper left, upper right, lower left blocks, etc.). Also, modes derived from the modes already included in the MPM candidate group (in the case of a directional mode, differences such as 1, -1, etc.) can be configured as the MPM candidate group.

[0038] There can be a priority order of prediction modes for the composition of the MPM candidate group. According to the priority order, the order included in the MPM candidate group can be determined. When only the number of the MPM candidate group (determined according to the number of prediction mode candidate groups) is satisfied according to the priority order, the composition of the MPM candidate group can be completed. At this time, the priority order can be determined in the order of the prediction modes of spatially adjacent blocks, a predetermined prediction mode, and the modes derived from the prediction modes first included in the MPM candidate group, but other variations are also possible.

[0039] For example, spatially adjacent blocks can be included in the candidate group in an order such as left - up - lower left - upper right - upper left blocks, and a predetermined prediction mode can be included in the candidate group in an order such as DC - Planar - vertical - horizontal modes. Six modes in total can be configured as the candidate group by including the modes obtained by adding +1, -1, etc. to the already included modes. Or, a total of seven modes can be configured as the candidate group by including them in the candidate group in a single priority order such as left - up - DC - Planar - lower left - upper right - upper left - (left + 1) - (left - 1) - (up + 1).

[0040] The subtraction unit 205 can generate a residual block by subtracting a predicted block from a target block. That is, the subtraction unit 205 calculates the difference between the pixel value of each pixel of the target block to be encoded and the predicted pixel value of each pixel of the predicted block generated through the prediction unit, and can generate a residual block which is a residual signal in the form of a block shape. Also, the subtraction unit 205 can generate a residual block based on units other than block units obtained through a block division unit described later.

[0041] The conversion unit 210 can convert a signal belonging to the spatial domain into a signal belonging to the frequency domain. The signal obtained through the conversion process is called a transformed coefficient. For example, a residual block having a residual signal transmitted from the subtraction unit can be converted to obtain a transformed block having transformed coefficients, but the input signal is determined according to the encoding settings and is not limited to the residual signal.

[0042] The conversion unit can convert a residual block using conversion techniques such as Hadamard Transform, DST Based-Transform (Discrete Sine Transform), DCT Based-Transform (Discrete Cosine Transform), etc., and is not limited thereto, and various conversion techniques obtained by improving and modifying this can be used.

[0043] At least one of the conversion techniques can be supported, and at least one detailed conversion technique can be supported for each conversion technique. At this time, the detailed conversion technique may be a conversion technique configured such that a part of the basis vectors is different for each conversion technique.

[0044] For example, in the case of DCT, one or more detailed conversion techniques among DCT-1 to DCT-8 can be supported, and in the case of DST, one or more detailed conversion techniques among DST-1 to DST-8 can be supported. A group of conversion technique candidates can be configured by constituting a part of the detailed conversion techniques. As an example, DCT-2, DCT-8, and DST-7 can be configured as a group of conversion technique candidates to perform conversion.

[0045] The conversion can be performed in the horizontal / vertical direction. For example, by performing one-dimensional conversion in the horizontal direction using the DCT-2 conversion technique and one-dimensional conversion in the vertical direction using the DST-7 conversion technique to perform a total two-dimensional conversion, the pixel values in the spatial domain can be converted into the frequency domain.

[0046] Can conversion be performed using a fixed single conversion technique, or can the conversion technique be adaptively selected according to the encoding settings and then the conversion be performed? At this time, in the adaptive case, the conversion technique can be selected using explicit or implicit methods. In the explicit case, the respective conversion technique selection information or conversion technique set selection information applied in the horizontal and vertical directions can be generated in units such as blocks. In the implicit case, the encoding settings can be defined according to the image type (I / P / B), color components, block size / shape / position, in-picture prediction mode, etc., and thereby a predetermined conversion technique can be selected.

[0047] Also, it is possible that some of the above conversions are omitted according to the encoding settings. That is, it means that one or more of the horizontal / vertical units can be omitted explicitly or implicitly.

[0048] The conversion unit can transmit the information necessary to generate the conversion block to the encoding unit so that the encoding unit encodes it, record the resulting information in the bitstream, and transmit it to the decoder. The decoding unit of the decoder can parse the information for this and use it in the inverse conversion process.

[0049] The quantization unit 215 can quantize the input signal. At this time, the signal obtained through the quantization process is called a quantized coefficient. For example, the residual block having the residual transform coefficients transmitted from the conversion unit can be quantized to obtain a quantized block having quantized coefficients, but the input signal is determined according to the encoding settings and is not limited to the residual transform coefficients.

[0050] The quantization unit can quantize the transformed residual block using quantization techniques such as dead zone uniform threshold quantization and quantization weighted matrix, and is not limited thereto. Various quantization techniques improved and modified therefrom can be used.

[0051] According to the symbolization setting, the quantization process can be omitted. For example, according to the symbolization setting (for example, the quantization parameter is 0, that is, a lossless compression environment), the quantization process (including the inverse process) can be omitted. As another example, when the compression performance by quantization is not exhibited according to the characteristics of the image, the quantization process can be omitted. At this time, the region in the quantization block (M×N) where the quantization process is omitted can be the entire region or a partial region (such as M / 2×N / 2, M×N / 2, M / 2×N), and the quantization omission selection information can be determined implicitly or explicitly.

[0052] The quantization unit can transmit the information necessary to generate the quantization block to the encoding unit to encode it, record the information thereby in the bit stream, and transmit it to the decoder. The decoding unit of the decoder can parse the information for this and use it for the inverse quantization process.

[0053] In the above example, it was described under the assumption that the residual block is transformed and quantized through the transformation unit and the quantization unit. However, the residual signal can be transformed to generate a residual block having transformation coefficients, and it is not necessary to perform the quantization process. It is possible not only to perform only the quantization process without transforming the residual signal of the residual block into transformation coefficients, but also not to perform both the transformation and the quantization process. This can be determined according to the setting of the encoder.

[0054] The inverse quantization unit 220 inverse quantizes the residual block quantized by the quantization unit 215. That is, the inverse quantization unit 220 inverse quantizes the quantized frequency coefficient sequence to generate a residual block having frequency coefficients.

[0055] The inverse transform unit 225 inverse-transforms the residual block inverse-quantized by the inverse quantization unit 220. That is, the inverse transform unit 225 inverse-transforms the frequency coefficients of the inverse-quantized residual block to generate a residual block having pixel values, that is, a restored residual block. Here, the inverse transform unit 225 can perform inverse transformation by using the inverse of the transformation method used in the transform unit 210.

[0056] The addition unit 230 adds the prediction block predicted by the prediction unit 200 and the residual block restored by the inverse transform unit 225 to restore the target block. The restored target block is stored in the coded picture buffer 240 as a reference picture (or reference block), and can be used as a reference picture when coding the next block or other backward blocks, or other pictures of the target block.

[0057] The filter unit 235 can include one or more post-processing filter processes such as a deblocking filter, SAO (Sample Adaptive Offset), and ALF (Adaptive Loop Filter). The deblocking filter can remove block distortion generated at the boundary between blocks from the restored picture. The ALF can perform filtering based on the value obtained by comparing the restored image after the block is filtered through the deblocking filter with the original image. The SAO can restore the offset difference from the original image for each pixel of the residual block to which the deblocking filter is applied. Such post-processing filters can be applied to the restored picture or block.

[0058] The coded picture buffer 240 can store the block or picture restored through the filter unit 235. The restored block or picture stored in the coded picture buffer 240 can be provided to the prediction unit 200 that performs intra prediction or inter prediction.

[0059] The entropy encoding unit 245 can scan at least one of the quantization coefficients, transform coefficients, or residual signals of the generated residual block according to at least one scan order (e.g., zigzag scan, vertical scan, horizontal scan, etc.) to generate a quantization coefficient sequence, a transform coefficient sequence, or a signal sequence, and can encode it using at least one entropy coding technique. At this time, the information about the scan order can be determined according to the encoding settings (e.g., image type, encoding mode, prediction mode, transform type, etc.), and can be determined implicitly or explicitly generate relevant information.

[0060] Also, encoded data including the encoded information transmitted from each component can be generated and output as a bitstream, which can be realized by a multiplexer (MUX). At this time, as the encoding technique, methods such as Exponential Golomb, CAVLC (Context-Adaptive Variable Length Coding), and CABAC (Context-Adaptive Binary Arithmetic Coding) can be used for encoding. It is not limited to this, and various encoding techniques obtained by improving and modifying this can be used.

[0061] When performing entropy encoding (assuming CABAC in this example) on syntax elements such as the residual block data and the information generated in the encoding / decoding process, the entropy encoding device can include a binarizer, a context modeler, and a binary arithmetic coder. At this time, the binary arithmetic coder can include a regular coding engine and a bypass coding engine.

[0062] Since the syntax elements input to the entropy encoding device may not be binary, when the syntax element is not binary, the binarization unit can binarize the syntax element and output a bin string consisting of 0 or 1. At this time, a bin indicates a bit consisting of 0 or 1 and can be encoded by the binary arithmetic coding unit. At this time, either the regular coding unit or the bypass coding unit can be selected based on the occurrence probabilities of 0 and 1, which can be determined based on the encoding / decoding settings. When the syntax element is data with the same frequency of 0 and 1, the bypass coding unit can be used; otherwise, the regular coding unit can be used.

[0063] Various methods can be used when performing binarization on the syntax element. For example, fixed length binarization, unary binarization, truncated rice binarization, K-th exp-golomb binarization, etc. can be used. Also, depending on the value range of the syntax element, signed binarization or unsigned binarization can be performed. The binarization process for the syntax elements generated in the present invention can be performed including not only the binarizations mentioned in the above examples but also other additional binarization methods.

[0064] FIG. 3 is a block configuration diagram showing an image decoding device according to an embodiment of the present invention.

[0065] Referring to FIG. 3, the image decoding device 30 can be configured to include an entropy decoding unit 305, a prediction unit 310, an inverse quantization unit 315, an inverse transform unit 320, an adder / subtractor 325, a filter 330, and a decoded picture buffer 335.

[0066] Also, the prediction unit 310 can be further configured to include an intra-prediction module and an inter-prediction module.

[0067] First, when the image bitstream transmitted from the image encoding device 20 is received, it can be transmitted to the entropy decoding unit 305.

[0068] The entropy decoding unit 305 can decode the bitstream to decode the decoded data including the quantized coefficients and the decoded information transmitted to each component.

[0069] The prediction unit 310 can generate a prediction block based on the data transmitted from the entropy decoding unit 305. At this time, based on the reference image stored in the decoded picture buffer 335, a reference picture list using the default configuration technique can also be configured.

[0070] The in-picture prediction unit can include a reference pixel configuration unit, a reference pixel filter unit, a reference pixel interpolation unit, a prediction block generation unit, and a prediction mode decoding unit. Some parts perform the same process as the encoder, and some parts can perform the reverse induction process.

[0071] The inverse quantization unit 315 can inverse-quantize the quantization conversion coefficients provided as a bitstream and decoded by the entropy decoding unit 305.

[0072] The inverse transformation unit 320 can apply an inverse DCT, an inverse integer transformation, or an inverse transformation technique of a similar concept to the transformation coefficients to generate a residual block.

[0073] At this time, the inverse quantization unit 315 and the inverse transformation unit 320 perform the reverse process of the processes performed by the transformation unit 210 and the quantization unit 215 of the image encoding device 20 described above, and can be realized in various ways. For example, the same processes and inverse transformations shared with the transformation unit 210 and the quantization unit 215 can also be used, and the transformation and quantization processes can be reversed using information on the transformation and quantization processes (e.g., transformation size, transformation shape, quantization type, etc.) from the image encoding device 20.

[0074] The residual block that has undergone the inverse quantization and inverse transformation processes can be added to the prediction block derived by the prediction unit 310 to generate a restored image block. Such addition can be performed by the adder / subtractor 325.

[0075] The filter 330 can also apply a deblocking filter to the restored image block to remove the blocking phenomenon as necessary, and another loop filter can be further used before and after the decoding process to improve the quality of the video.

[0076] The image block that has undergone restoration and filtering can be stored in the decoded picture buffer 335.

[0077] Although not shown, the image encoding / decoding apparatus can further include a picture partitioning unit and a block partitioning unit.

[0078] The picture partitioning unit can divide or partition a picture into at least one region based on a predetermined partitioning unit. Here, the partitioning unit can include sub-pictures, slices, tiles, bricks, blocks (e.g., maximum coding units), and the like.

[0079] A picture can be divided into one or more rows of tiles or one or more columns of tiles. At this time, a tile can be a block-based unit including a predetermined rectangular region of the picture. At this time, a tile can be divided into one or more bricks, and a brick can be composed of blocks in units of rows or columns of tiles.

[0080] A slice can have one or more configuration settings. One of them can be composed of bundles in scan order (e.g., blocks, bricks, tiles, etc.), one of them can be composed in a form including a rectangular region, and other additional definitions are possible.

[0081] The definition regarding the slice configuration can be explicitly related information that can occur, or can be implicitly defined. Not only for slices, but also the definition regarding the configuration of each division unit can be set in plural, and selection information regarding this can occur.

[0082] A slice can be composed of rectangular units such as blocks, bricks, tiles, etc., and slice position and size information can be expressed based on position information (for example, upper left position, lower right position, etc.) in the division unit.

[0083] In the present invention, it is assumed for explanation that a picture can be composed of one or more sub-pictures, a sub-picture can be composed of one or more slices or tiles or bricks, a slice can be composed of one or more tiles or bricks, and a tile can be composed of one or more bricks, but it is not limited thereto.

[0084] The division unit can be composed of an integer number of blocks, but is not limited thereto, and may be composed of a non-integer number. That is, when it is not composed of an integer number of blocks, at least one division unit can be composed of sub-blocks.

[0085] There can be division units such as sub-pictures and tiles other than rectangular-shaped slices, but the position and size information of the units can be expressed based on various methods.

[0086] For example, based on information such as the number of rectangular units, the number of columns or rows of rectangular units, information on whether the columns or rows of rectangular units are evenly divided, the width or height information of the columns or rows of rectangular units, the index information of rectangular units, etc., the position and size information of the rectangular-shaped units can be expressed.

[0087] In the case of sub-pictures and tiles, the position and size information of each unit can be expressed based on all or part of the above information, and based on this, it can be divided or partitioned into one or more units.

[0088] On one hand, it can be divided into blocks of various units and sizes via a block division unit. The basic encoding unit (or the maximum encoding unit, Coding Tree Unit, CTU) can be meant to be the basic (or starting) unit for prediction, transformation, quantization, etc. in the image encoding process. At this time, the basic encoding unit can be composed of one luminance basic encoding block (or the maximum encoding block, Coding Tree Block, CTB) and two basic chrominance encoding blocks according to the color format (in this example, YCbCr), and the size of each block can be determined according to the color format. An encoding block (Coding Block, CB) can be obtained according to the division process. The encoding block can be understood as a unit that has not been divided into more encoding blocks according to certain restrictions, and can be set as the starting unit for division into lower-level units. In the present invention, a block is not limited to a rectangular shape only, and can be understood as a broad concept including various shapes such as triangles and circles.

[0089] It should be understood that the content described later targets one color component, but can be changed and applied to other color components in proportion to the ratio according to the color format (for example, in the case of YCbCr4:2:0, the horizontal and vertical length ratios of the luminance component and the chrominance components are 2:1). Also, although block division dependent on other color components (for example, in Cb / Cr, dependent on the block division result of Y) is possible, it should be understood that block division independent of each color component is possible. Also, although a common block division setting (considering proportionality in length ratio) can be used, it is necessary to understand considering that individual block division settings are used according to the color component.

[0090] In the block division unit, a block can be represented by M×N, and the maximum and minimum values of each block Values can be obtained within a range. For example, when the maximum value of a block is set to 256×256 and the minimum value is set to 4×4, a block of size 2m×2n (in this example, m and n are integers from 2 to 8) or a block of size 2m×2m (in this example, m and n are integers from 2 to 128) or a block of size m×m (in this example, m and n are integers from 4 to 256) can be obtained. Here, m and n may or may not be the same, and one or more ranges of blocks such as the maximum and minimum values can occur.

[0091] For example, information such as the maximum and minimum sizes of a block can occur, and information such as the maximum and minimum sizes of a block in a partial division setting can occur. Here, in the former case, it is information on the range for the maximum and minimum sizes that can occur within an image, and in the latter case, it can be information on the maximum and minimum sizes that can occur based on a partial division setting. Here, the division setting can be defined by the image type (I / P / B), color component (such as YCbCr), block type (encoding / prediction / transformation / quantization, etc.), division type (Index or Type), division method (QT, BT, TT, etc. in the tree method, SI2, SI3, SI4, etc. in the Index method), etc.

[0092] In addition, there can be restrictions on the horizontal / vertical length ratio (block shape) that a block can have, and boundary value conditions for this can be set. At this time, only blocks with an arbitrary boundary value (k) or less / less than can be supported, and k can be defined based on the horizontal and vertical length ratio such as A / B (A is the longer or same value of the horizontal and vertical, B is the remaining value), and can be a real number of 1 or more such as 1.5, 2, 3, 4, etc. Restriction conditions regarding the shape of a single block in an image as in the above example can be supported, or one or more restriction conditions can be supported according to the division setting.

[0093] In summary, the presence or absence of support for block division can be determined by the range and conditions as described above and the division settings to be described later. For example, if the block conditions for supporting candidates (child blocks) by dividing a block (parent block) are satisfied, the division can be supported; if not, the division cannot be supported.

[0094] The block division unit can be set in relation to each component of the image encoding device and the decoding device, and the size and shape of the block can be determined through this process. At this time, the blocks to be set can be defined differently depending on the component. For example, in the case of the prediction unit, it can correspond to a prediction block (Prediction Block), in the case of the conversion unit, it can correspond to a transform block (Transform Block), and in the case of the quantization unit, it can correspond to a quantization block (Quantization Block). However, it is not limited to this, and block units by other components can be further defined. In the present invention, the case where the input and output of each component are rectangular in shape will be mainly described, but in some components, input / output of other shapes (for example, a right triangle, etc.) is possible.

[0095] The size and shape of the initial (or start) block of the block division unit can be determined from a higher unit. The initial block can be divided into blocks of a smaller size. When the optimal size and shape are determined by dividing the block, the block can be determined as the initial block of a lower unit. Here, the higher unit can be an encoding block, and the lower unit can be a prediction block or a transform block, but it is not limited to this, and various modified examples are possible. When the initial block of the lower unit is determined as in the above example, a division process for finding a block with the optimal size and shape like the higher unit can be performed.

[0096] In summary, the block splitting unit can split a basic encoding block (or maximum encoding block) into at least one encoding block, and can split the encoding block into at least one prediction block / transformation block / quantization block. Also, the prediction block can be split into at least one transformation block / quantization block, and the transformation block can be split into at least one quantization block. Here, some blocks may have a subordinate relationship with other blocks (i.e., defined by a higher unit and a lower unit), or may have an independent relationship. As an example, the prediction block may be a higher unit of the transformation block, or may be an independent unit from the transformation block, and various relationship settings are possible according to the type of block.

[0097] According to the encoding settings, the presence or absence of the combination of the higher unit and the lower unit can be determined. Here, the combination between units means performing the encoding process (e.g., prediction unit, transformation unit, inverse transformation unit, etc.) of the lower unit with the block (size and shape) of the higher unit without performing the split from the higher unit to the lower unit. That is, it can be meant that the splitting processes in multiple units are shared, and the splitting information is generated in one of the units (e.g., the higher unit).

[0098] As an example, (when the encoding block is combined with the prediction block and the transformation block), the prediction process, transformation, and inverse transformation processes can be performed with the encoding block.

[0099] As an example, (when the encoding block is combined with the prediction block), the prediction process can be performed with the encoding block, and the transformation and inverse transformation processes can be performed with a transformation block that is the same as or smaller than the encoding block.

[0100] As an example, (when the encoding block is combined with the transformation block), the prediction process can be performed with a prediction block that is the same as or smaller than the encoding block, and the transformation and inverse transformation processes can be performed with the encoding block.

[0101] As an example, (when the prediction block is combined with the transform block), the prediction process can be performed using a prediction block that is the same as or smaller than the encoding block, and the transform and inverse transform processes can be performed using the prediction block.

[0102] As an example, (when no block is combined), the prediction process can be performed using a prediction block that is the same as or smaller than the encoding block, and the transform and inverse transform processes can be performed using a transform block that is the same as or smaller than the encoding block.

[0103] Although various cases regarding the encoding, prediction, and transform blocks have been described in the above examples, the present invention is not limited thereto.

[0104] The combination between the units can support a fixed setting in the image or an adaptive setting considering various encoding elements. Here, the encoding elements can include an image type, a color component, an encoding mode (Intra / Inter), a segmentation setting, a block size / shape / position, a horizontal / vertical length ratio, prediction-related information (e.g., in-picture prediction mode, inter-picture prediction mode, etc.), transform-related information (e.g., transform technique selection information, etc.), quantization-related information (e.g., quantization region selection information, quantized transform coefficient encoding information, etc.), and the like.

[0105] As described above, when a block with an optimal size and shape is found, mode information (e.g., segmentation information, etc.) for this block can be generated. The mode information can be recorded in a bitstream together with information generated in the component to which the block belongs (e.g., prediction-related information and transform-related information, etc.) and transmitted to a decoder, where it can be parsed into units of the same level and used in the image decoding process.

[0106] Hereinafter, the segmentation method will be described. For the sake of convenience of explanation, it is assumed that the initial block has a square shape, but the present invention is not limited thereto because it can be applied identically or similarly when the initial block has a rectangular shape.

[0107] The block splitting unit can support various types of splitting. For example, it can support tree-based splitting or index-based splitting, and other methods can also be supported. Tree-based splitting can determine the splitting form with various types of information (such as whether to split, tree type, splitting direction, etc.), and index-based splitting can determine the splitting form with predetermined index information.

[0108] Figure 4 is an exemplary diagram showing various splitting forms that can be obtained by the block splitting unit of the present invention.

[0109] In this example, it is assumed that the splitting form as shown in Figure 4 is obtained by one execution (or process) of splitting, but it is not limited to this, and it is also possible to be obtained by multiple splitting operations. Also, additional splitting forms not shown in Figure 4 are possible.

[0110] (Tree-based splitting) In the tree-based splitting of the present invention, a quad tree (QT), a binary tree (BT), a ternary tree (TT), etc. can be supported. When one tree method is supported, it is called single-tree splitting, and when two or more tree methods are supported, it can be called multi-tree splitting.

[0111] In the case of QT, it means a method (n) in which the block is split into two in both the horizontal and vertical directions (i.e., four-way split), in the case of BT, it means a method (b to g) in which the block is split into two in one of the horizontal or vertical directions, and in the case of TT, it means a method (h to m) in which the block is split into three in one of the horizontal or vertical directions.

[0112] Here, in the case of QT, it is also possible to support a method (o, p) of dividing into four parts by limiting the division direction to either horizontal or vertical. Also, in the case of BT, it is possible to support only the method (b, c) with equal sizes, or only the method (d to g) with unequal sizes, or to support a mixture of both methods. Further, in the case of TT, it is possible to support only the method (h, j, k, m) having an arrangement where the division is biased in a specific direction (such as 1:1:2, 2:1:1 in the left → right or up → down direction), or only the method (i, l) arranged in the center (such as 1:2:1), or to support a mixture of both methods. Additionally, it is also possible to support a method (q) in which the division direction is divided into four parts (i.e., 16 parts) in both the horizontal and vertical directions.

[0113] And, among the tree methods, it is possible to support a method (b, d, e, h, i, j, o) of dividing into z parts by limiting to the horizontal division direction, or a method (c, f, g, k, l, m, p) of dividing into z parts by limiting to the vertical division direction, or to support a mixture of both methods. Here, z can be an integer of 2 or more such as 2, 3, 4.

[0114] In the present invention, an explanation will be given assuming that QT supports n, BT supports b and c, and TT supports i and l.

[0115] Depending on the encoding settings, one or more of the above tree division methods can be supported. For example, it is possible to support QT, or to support QT / BT, or to support QT / BT / TT.

[0116] The above example is an example where the basic tree split is QT, and BT and TT are included in the additional split method according to the presence or absence of other tree support, but various modifications are possible. At this time, information about the presence or absence of other tree support (such as bt_enabled_flag, tt_enabled_flag, bt_tt_enabled_flag, etc. It can have a value of 0 or 1, and if it is 0, there is no support, and if it is 1, there is support) is implicitly determined according to the encoding settings, or can be explicitly determined in units such as sequences, pictures, sub-pictures, slices, tiles, blocks, etc.

[0117] The split information may include information about whether or not to split (tree_part_flag, or qt_part_flag, bt_part_flag, tt_part_flag, bt_tt_part_flag. It can have a value of 0 or 1, and if it is 0, it is not split, and if it is 1, it is split). Also, according to the split method (BT and TT), information about the split direction (dir_part_flag, or bt_dir_part_flag, tt_dir_part_flag, bt_tt_dir_part_flag. It can have a value of 0 or 1, and if it is 0, it is <horizontal>, and if it is 1, it is <vertical>) can be added, and this can be information that can occur when splitting is performed.

[0118] When multiple tree splits are supported, various split information configurations are possible. Next, an example will be assumed and explained for how the split information is configured at one depth level (that is, although the supported split depth may be set to one or more and recursive splitting may be possible, for the purpose of convenience of explanation).

[0119] As an example (1), check the information about whether or not to split. At this time, if splitting is not performed, the splitting ends.

[0120] If a split is to be performed, check the selection information regarding the type of split (e.g., tree_idx. If it is 0, it is QT; if it is 1, it is BT; if it is 2, it is TT). At this time, further check the split direction information according to the selected split type, and proceed to the next step (if additional splitting is possible for reasons such as the split depth not reaching the maximum, start over from the beginning; if splitting is not possible, end the split).

[0121] As an example (2), check the information regarding whether to perform a split for some tree methods (QT) and proceed to the next step. At this time, if no split is performed, check the information regarding whether to perform a split for some tree methods (BT). At this time, if no split is performed, check the information regarding whether to perform a split for some tree methods (TT). At this time, if no split is performed, end the split.

[0122] If a split is performed for some tree methods (QT), proceed to the next step. Also, if a split is performed for some tree methods (BT), check the split direction information and proceed to the next step. Also, if a split is performed for some tree split methods (TT), check the split direction information and proceed to the next step.

[0123] As an example (3), check the information regarding whether to perform a split for some tree methods (QT). At this time, if no split is performed, check the information regarding whether to perform a split for some tree methods (BT and TT). At this time, if no split is performed, end the split.

[0124] If a split is performed for some tree methods (QT), proceed to the next step. Also, if a split is performed for some tree methods (BT and TT), check the split direction information and proceed to the next step.

[0125] The above examples can be cases where there is a priority order for tree splitting (Examples 2 and 3) or where there is no such priority order (Example 1), but various modified examples are possible. Also, in the above examples, the splitting at the current step is an example that explains the case where it is independent of the splitting result at the previous step, but it is also possible to have a setting where the splitting at the current step depends on the splitting result at the previous step.

[0126] For example, in Examples 1 to 3, if some tree-based splitting (QT) has been performed in the previous step and the process has moved to the current step, the same tree-based splitting (QT) can be supported in the current step.

[0127] On the other hand, if some tree-based splitting (QT) has not been performed in the previous step and some other tree-based splitting (BT or TT) has been performed and the process has moved to the current step, it is also possible to have a setting where some tree-based splitting (BT and TT) is supported in the subsequent steps including the current step, excluding some tree-based splitting (QT).

[0128] In the cases described above, it means that the tree configuration supported for block splitting can be adaptive, which also means that the above-described splitting information configuration can be configured differently. (The example described later assumes Example 3) That is, in the above example, if some tree-based splitting (QT) has not been performed in the previous step, the splitting process can be carried out in the current step without considering some tree-based splitting (QT). Also, the splitting information regarding the relevant tree-based method (for example, information on whether to split, splitting direction information, etc. In this example <qt>Then, information about whether or not to perform splitting can be removed and configured.

[0129] The above example pertains to an adaptive splitting information configuration for cases where block splitting is allowed (for example, the block size exists within a range between a maximum value and a minimum value, and the splitting depth of each tree-based method does not reach the maximum depth <allowable depth>, etc.), and an adaptive splitting information configuration is also possible for cases where block splitting is restricted (for example, the block size does not exist within the range between the maximum value and the minimum value, and the splitting depth of each tree-based method reaches the maximum depth, etc.).

[0130] As already described, in the present invention, tree-based splitting can be performed using a recursive method. For example, when the splitting flag of an encoded block with a splitting depth of k is 0, the encoding of the encoded block is performed on the encoded block with a splitting depth of k. When the splitting flag of the encoded block with a splitting depth of k is 1, the encoding of the encoded block is performed on N sub-encoded blocks with a splitting depth of k + 1 according to the splitting method (where N is an integer of 2 or more such as 2, 3, 4, etc.).

[0131] The sub-encoded blocks are again set as encoded blocks (k + 1), and through the above process, they can be split into sub-encoded blocks (k + 2). Such a hierarchical splitting method can be determined according to splitting settings such as the splitting range and the allowable splitting depth.

[0132] At this time, the bitstream structure for representing splitting information can be selected from one or more scanning methods. For example, the bitstream of splitting information can be configured based on the order of splitting depth, or the bitstream of splitting information can be configured based on whether or not to perform splitting.

[0133] For example, when based on the order of the split depth, it is a method of obtaining split information at the depth of the current level based on the first block, and then obtaining split information at the depth of the next level. When based on whether to split or not, it means a method of preferentially obtaining additional split information for the blocks split based on the first block, and other additional scanning methods can be considered.

[0134] Regardless of the type of tree, can the maximum block size and the minimum block size have common settings (or for all trees), or can they have individual settings according to each tree, or can they have common settings for two or more trees? At this time, the maximum block size can be set to be the same as or smaller than the maximum coded block. If the maximum block size according to a predetermined first tree is not the same as the maximum coded block, implicitly perform splitting using a predetermined second tree method until the maximum block size of the first tree is reached.

[0135] Regardless of the type of tree, can a common split depth be supported, or can individual split depths be supported according to each tree, or can a common split depth be supported for two or more trees? Or, for some trees, split depth may be supported, and for some trees, split depth may not be supported.

[0136] Explicit syntax elements for the said setting information can be supported, and some setting information may be implicitly defined.

[0137] (Index-based splitting)

[0138] In the index-based splitting of the present invention, methods such as the CSI (Constant Split Index) method and the VSI (Variable Split Index) method can be supported.

[0139] The CSI method can be a method in which k sub-blocks are obtained by dividing in a predetermined direction. At this time, k can be an integer of 2 or more such as 2, 3, 4. Specifically, regardless of the size and shape of the block, it can be a division method with a configuration in which the size and shape of the sub-blocks are determined based on the k value. Here, the predetermined direction can be one or a combination of two or more directions among the horizontal, vertical, diagonal (such as the upper left → lower right direction, or the lower left → upper right direction) directions.

[0140] The index-based CSI division method of the present invention can include candidates divided into z in either the horizontal or vertical direction. At this time, z is an integer of 2 or more such as 2, 3, 4, and among the horizontal or vertical lengths of each sub-block, one of them is the same, and the other one can be the same or different. The horizontal or vertical length ratio of the sub-blocks is A1:A2:...:AZ, and A1 to AZ can be integers of 1 or more such as 1, 2, 3.

[0141] Also, it can include candidates divided into x and y in the horizontal and vertical directions respectively. At this time, x and y can be integers of 1 or more such as 1, 2, 3, 4, but when x and y are both 1 (since a already exists), it can be restricted. FIG. 4 shows a case where the horizontal or vertical length ratio of each sub-block is the same, but it may also include candidates including different cases.

[0142] Also, it can include candidates divided into w in either one of some diagonal directions (upper left → lower right direction) or some diagonal directions (lower left → upper right direction), and w can be an integer of 2 or more such as 2, 3.

[0143] Referring to FIG. 4, the division forms can be classified into a symmetric division form (b) and an asymmetric division form (d, e) according to the length ratio of each sub-block, and can also be classified into a division form (k, m) biased in a specific direction and a division form (k) arranged in the center. The division form can be defined by various coding elements including not only the length ratio of the sub-blocks but also the form of the sub-blocks. Depending on the coding setting, the supported division forms may be implicitly or explicitly defined. Therefore, a candidate group in the index-based division method can be determined based on the supported division forms.

[0144] On the other hand, the VSI method can be a method in which one or more sub-blocks are obtained through division in a predetermined direction while the width w or height h of the sub-blocks is fixed. w and h can be integers of 1 or more such as 1, 2, 4, 8. Specifically, it can be a division method with a configuration in which the number of sub-blocks is determined based on the size and shape of the block and the value of w or n.

[0145] The index-based VSI division method of the present invention can include candidates divided by fixing either the horizontal or vertical length of the sub-blocks. Or, it can include candidates divided by fixing both the horizontal and vertical lengths of the sub-blocks. Since the horizontal or vertical length of the sub-blocks is fixed, it can have the feature of allowing equal division in the horizontal or vertical direction, but is not limited thereto.

[0146] If the block before division is M×N, and the horizontal length of the sub-block is fixed (w), or the vertical length is fixed (h), or both the horizontal and vertical lengths are fixed (w, h), the number of sub-blocks obtained can be (M*N) / w, (M*N) / h, (M*N) / w / h, respectively.

[0147] Depending on the coding setting, only the CSI method can be supported, or only the VSI method can be supported, or both methods can be supported, and information about the supported method can be implicitly or explicitly defined.

[0148] In the present invention, the description will be made assuming the case where the CSI method is supported.

[0149] According to the encoding setting, a candidate group can be configured including two or more candidates among the index divisions.

[0150] For example, candidate groups such as {a, b, c}, {a, b, c, n}, {a to g, n} can be configured, but it is an example where a block shape that is predicted to occur frequently based on general statistical characteristics, such as being divided into two in the horizontal or vertical direction, or being divided into two in both the horizontal and vertical directions, is configured as the candidate group.

[0151] Or, candidate groups such as {a, b}, {a, o}, {a, b, o} or {a, c}, {a, p}, {a, c, p} can be configured, and each includes candidates that are divided into two and four in the horizontal and vertical directions respectively. This can be an example where a block shape that is predicted to have many divisions in a specific direction is configured as the candidate group.

[0152] Or, candidate groups such as {a, o, p} or {a, n, q} can be configured, but it is an example where a block shape that is predicted to have many divisions with a size smaller than the block before division is configured as the candidate group.

[0153] Or, a candidate group such as {a, r, s} can be configured, but it is judged that the optimal division result obtained in a rectangular shape using another method (tree method) from the block before division, and it can be an example where a non-rectangular division form is configured as the candidate group.

[0154] As in the above examples, various candidate groups can be configured, and the configuration of one or more candidate groups can be supported in consideration of various encoding elements.

[0155] When the configuration of the candidate group is completed, various division information configurations are possible.

[0156] For example, index selection information can be generated from a candidate group including a non-divided candidate (a) and divided candidates (b to s).

[0157] Alternatively, information indicating whether or not to perform division (whether the division form is a) can be generated, and when division is performed (when it is not a), index selection information can be generated from a candidate group composed of divided candidates (b to s).

[0158] It is possible to configure division information in various ways other than the above description. Except for the information indicating whether or not to perform division, binary bits can be assigned to the indexes of each candidate in the candidate group using various methods such as fixed-length binarization and variable-length binarization. If the number of candidate groups is two, 1 bit is assigned to the index selection information, and if it is three or more, 1 bit or more can be assigned to the index selection information.

[0159] Different from the tree-based division method, the index-based division method can be a method of selectively configuring a candidate group with division forms predicted to occur frequently.

[0160] And, according to the number of supported candidate groups, the amount of bits for expressing index information can be increased. Therefore, it can be a method suitable for single-layer division (for example, the division depth is limited to 0) rather than the hierarchical division (recursive division) of the tree-based method. That is, it may be a method that supports one division operation, or the sub-blocks obtained through index-based division may be a method in which no further additional division is possible.

[0161] At this time, it can be meant that additional splitting into the same type of blocks having a smaller size is impossible (for example, the encoded block obtained by the index splitting method cannot be additionally split into encoded blocks), but a setting where additional splitting into other types of blocks is impossible (for example, not only from the encoded block to the encoded block, but also splitting into prediction blocks is impossible) is also possible. Of course, it is not limited to the above examples, and other modified examples are possible.

[0162] Next, consider the case where the block splitting setting is determined mainly by the type of block among the encoding elements.

[0163] First, an encoded block can be obtained through the splitting process. Here, for the splitting process, a tree-based splitting method can be used, and splitting form results such as a (no split), n (QT), b, c (BT), i, l (TT) in FIG. 4 can be obtained depending on the type of tree. Depending on the encoding setting, various combinations of each type of tree such as QT / QT + BT / QT + BT + TT are possible.

[0164] The example described later shows the process in which the prediction block and the transform block are finally partitioned based on the encoded block obtained through the above process, and assumes the case where the prediction, transform, and inverse transform processes are performed based on the size of each partitioned area.

[0165] As an example (1), the prediction block can be set as the size of the encoded block and the prediction process can be performed, and the transform block can be set as the size of the encoded block (or the prediction block) and the transform and inverse transform processes can be performed. In the case of the prediction block and the transform block, since they are set based on the encoded block, there is no separately generated splitting information.

[0166] As an example (2), the prediction block can be set to the size of the encoding block and the prediction process can be performed. In the case of the transform block, the transform block can be obtained through the splitting process based on the encoding block (or prediction block), and the transform and inverse transform processes can be performed based on the obtained size.

[0167] Here, for the splitting process, a tree-based splitting method can be used. Depending on the type of tree, splitting form results such as a (no split), b, c (BT), i, l (TT), n (QT), etc. in FIG. 4 can be obtained. Depending on the encoding settings, various combinations of types of each tree such as QT / BT / QT+BT / QT+BT+TT are possible.

[0168] Here, for the splitting process, an index-based splitting method can be used. Depending on the type of index, splitting form results such as a (no split), b, c, d, etc. in FIG. 4 can be obtained. Depending on the encoding settings, various candidate group configurations such as {a, b, c}, {a, b, c, d}, etc. are possible.

[0169] As an example (3), in the case of the prediction block, the prediction block can be obtained by performing a splitting process based on the encoding block, and the prediction process can be performed based on the obtained size. In the case of the transform block, the transform and inverse transform processes can be performed with the size of the encoding block set as it is. This example can apply when the prediction block and the transform block have an independent relationship with each other.

[0170] Here, for the splitting process, an index-based splitting method can be used. Depending on the type of index, splitting form results such as a (no split), b to g, n, r, s, etc. in FIG. 4 can be obtained. Depending on the encoding settings, various candidate group configurations such as {a, b, c, n}, {a to g, n}, {a, r, s}, etc. are possible.

[0171] As an example (4), in the case of a prediction block, a prediction block can be obtained by performing a splitting process based on an encoding block, and a prediction process can be performed based on the obtained size. In the case of a transform block, the transform and inverse transform processes can be performed with the size of the prediction block being set as it is. This example may be a case where the transform block is set with the size of the obtained prediction block as it is, or vice versa (the prediction block is set with the size of the transform block as it is).

[0172] Here, for the splitting process, a tree-based splitting method can be used, and depending on the type of tree, splitting forms such as a (no split), b, c (BT), n (QT), etc. in FIG. 4 can be obtained. Depending on the encoding settings, various combinations of types of each tree such as QT / BT / QT+BT are possible.

[0173] Here, for the splitting process, an index-based splitting method can be used, and depending on the type of index, splitting forms such as a (no split), b, c, n, o, p, etc. in FIG. 4 can be obtained. Depending on the encoding settings, various candidate group configurations such as {a, b}, {a, c}, {a, n}, {a, o}, {a, p}, {a, b, c}, {a, o, p}, {a, b, c, n}, {a, b, c, n, p}, etc. are possible. Also, among the index-based splitting methods, the VSI method may be used alone or in combination with the CSI method to form a candidate group.

[0174] As an example (5), in the case of a prediction block, a prediction block can be obtained by performing a splitting process based on an encoding block, and a prediction process can be performed based on the obtained size. Also, in the case of a transform block, a prediction block can be obtained by performing a splitting process based on the encoding block, and the transform and inverse transform processes can be performed based on the obtained size. This example can be a case where the encoding block is used as a basis to perform splitting for each of the prediction block and the transform block.

[0175] Here, for the splitting process, a tree-based splitting method and an index-based splitting method can be used, and candidate groups can be configured in the same or similar manner as in the fourth example.

[0176] The above examples illustrate some possible cases that can occur depending on whether the splitting processes of various types of blocks are shared or not, etc., but are not limited thereto, and various modified examples are possible. Also, not only the types of blocks, but also various coding elements can be considered to determine the block splitting settings.

[0177] At this time, the coding elements may include image type (I / P / B), color component (YCbCr), block size / shape / position, horizontal / vertical length ratio of the block, block type (encoded block, prediction block, transform block, quantization block, etc.), splitting state, coding mode (Intra / Inter), prediction-related information (intra-picture prediction mode, inter-picture prediction mode, etc.), transform-related information (transform technique selection information, etc.), quantization-related information (quantization region selection information, quantized transform coefficient coding information, etc.), and the like.

[0178] In the image coding method according to an embodiment of the present invention, the inter-picture prediction can be configured as follows. The inter-picture prediction of the prediction unit can include a reference picture construction step, a motion estimation step, a motion compensation step, a motion information determination step, and a motion information coding step. Also, the image coding apparatus can be configured to include a reference picture construction unit, a motion estimation unit, a motion compensation unit, a motion information determination unit, and a motion information coding unit that realize the reference picture construction step, the motion estimation step, the motion compensation step, the motion information determination step, and the motion information coding step. A part of the above-described processes can be omitted, or other processes can be added, and they can be changed in an order other than the described order.

[0179] In the image decoding method according to an embodiment of the present invention, the inter-picture prediction can be configured as follows. The inter-picture prediction of the prediction unit can include a motion information decoding step, a reference picture construction step, and a motion compensation step. Further, the image decoding apparatus can be configured to include a motion information decoding unit, a reference picture construction unit, and a motion compensation unit that implement the motion information decoding step, the reference picture construction step, and the motion compensation step. A part of the above-described process can be omitted, or other processes can be added, and the order can be changed to other orders instead of the described order.

[0180] Since the reference picture construction unit and the motion compensation unit of the image decoding apparatus perform the same role as the corresponding configurations of the image encoding apparatus, detailed descriptions are omitted, and the motion information decoding unit can be performed by using the method used in the motion information encoding unit in reverse. Here, the prediction block generated via the motion compensation unit can be transmitted to the addition unit.

[0181] FIG. 5 is an exemplary diagram showing various cases of obtaining a prediction block via the inter-picture prediction of the present invention.

[0182] Referring to FIG. 5, for one-way prediction, a prediction block (A. forward prediction) can be obtained from the previously encoded reference pictures T-1, T-2, or a prediction block (B. backward prediction) can be obtained from the subsequently encoded reference pictures T+1, T+2. For bidirectional prediction, prediction blocks (C, D) can be generated from a plurality of previously encoded reference pictures T-2 to T+2. Generally, the P picture type can support one-way prediction, and the B picture type can support bidirectional prediction.

[0183] As in the above example, the pictures referred to for encoding the current picture can be obtained from the memory, and a reference picture list can be constructed including reference pictures whose temporal order or display order is before the current picture T and reference pictures whose temporal order or display order is after the current picture T based on the current picture T.

[0184] Based on the current image, not only previous or subsequent images, but also inter-picture prediction E can be performed with the current image. Performing inter-picture prediction with the current image can be called non-directional prediction. This can be supported in the I picture type or can be supported in the P / B picture type. Depending on the encoding settings, the supported picture type can be determined. Performing inter-picture prediction with the current image generates a prediction block by utilizing spatial correlation, and it is only different from performing inter-picture prediction with other images for the purpose of utilizing temporal correlation. The prediction method (e.g., reference image, motion vector, etc.) can be the same.

[0185] Here, although the case where the picture types capable of performing inter-picture prediction are P and B pictures is assumed, it is also applicable to various other additional or alternative picture types. For example, a predetermined picture type can support only inter-picture prediction without supporting intra-picture prediction, can support only inter-picture prediction in a predetermined direction (backward direction), and can support only inter-picture prediction in a predetermined direction.

[0186] In the reference picture component, the reference pictures used for encoding the current picture can be configured and managed via the reference picture list. At least one reference picture list can be configured according to the encoding settings (e.g., picture type, prediction direction, etc.), and a prediction block can be generated from the reference pictures included in the reference picture list.

[0187] In the case of one-directional prediction, inter-picture prediction can be performed with at least one reference picture included in reference picture list 0 (L0) or reference picture list 1 (L1). Also, in the case of bi-directional prediction, inter-picture prediction can be performed with at least one reference picture included in the composite list LC generated by combining L0 and L1.

[0188] For example, one-direction prediction can be classified into forward prediction Pred_L0 using the forward reference picture list L0 and backward prediction Pred_L1 using the backward reference picture list L1. Bidirectional prediction Pred_BI can use all of the forward reference picture list L0 and the backward reference picture list L1.

[0189] Alternatively, copying the forward reference picture list L0 to the backward reference picture list L1 to perform two or more forward predictions can also be included in bidirectional prediction, and copying the backward reference picture list L1 to the forward reference picture list L0 to perform two or more backward predictions can also be included in bidirectional prediction.

[0190] The prediction direction can be represented by flag information indicating the direction (for example, inter_pred_idc. Assume that this value can be adjusted by predFlagL0, predFlagL1, and predFlagBI). predFlagL0 indicates whether it is forward prediction, and predFlagL1 indicates whether it is backward prediction. Bidirectional prediction can indicate the prediction situation via predFlagBI, or can indicate that predFlagL0 and predFlagL1 are simultaneously activated (for example, when each flag is 1).

[0191] In the present invention, the description will focus on the case of forward prediction using the forward reference picture list and one-direction prediction, but the same or modified application is also possible in the other cases.

[0192] Generally, a method can be used in which the encoder determines the optimal reference picture for the picture to be encoded and explicitly transmits information about the reference picture to the decoder. Therefore, the reference picture component can manage the picture list referred to for inter-picture prediction of the current picture and can set rules for reference picture management in consideration of the limited memory size.

[0193] The information to be transmitted can be defined as an RPS (Reference Picture Set). The pictures selected in the RPS are classified as reference pictures and stored in a memory (or DPB), and the pictures not selected in the RPS are classified as non-reference pictures and can be removed from the memory after a certain period of time. The memory can store a preset number of pictures (for example, 14, 15, 16 pictures or more), and the size of the memory can be set according to the level and image resolution.

[0194] FIG. 6 is an exemplary diagram showing the configuration of a reference picture list according to an embodiment of the present invention.

[0195] Referring to FIG. 6, generally, the reference pictures T-1 and T-2 existing before the current picture are assigned to L0, and the reference pictures T+1 and T+2 existing after the current picture are assigned to L1 and can be managed. When configuring L0, if the number of reference pictures allowed in L0 is not reached, the reference pictures in L1 can be assigned. Similarly, when configuring L1, if the number of reference pictures allowed in L1 is not reached, the reference pictures in L0 can be assigned.

[0196] Also, the current picture can be included in at least one reference picture list. For example, the current picture can be included in L0 or L1. A reference picture (or the current picture) with a time order of T can be added to the reference pictures before the current picture to configure L0, and a reference picture with a time order of T can be added to the reference pictures after the current picture to configure L1.

[0197] The configuration of the reference picture list can be determined according to the encoding settings.

[0198] The current picture can be managed through an individual memory that is not included in the reference picture list and is separated from the reference picture list, or the current picture can be included in at least one reference picture list for management.

[0199] For example, it can be determined by a signal (curr_pic_ref_enabled_flag) that indicates whether the current picture is included in the reference picture list. Here, the signal can be information determined implicitly or generated explicitly.

[0200] Specifically, when the signal is deactivated (for example, curr_pic_ref_enabled_flag = 0), the current picture is not included as a reference picture in any reference picture list. When the signal is activated (for example, curr_pic_ref_enabled_flag = 1), it is determined implicitly (for example, added only to L0, added only to L1, or added to both L0 and L1 simultaneously) whether the current picture is included in a predetermined reference picture list, or it can be determined by generating explicitly related signals (for example, curr_pic_ref_from_l0_flag, curr_pic_ref_from_l1_flag). The signal can be supported in units such as sequences, pictures, sub-pictures, slices, tiles, blocks, etc.

[0201] Here, the current picture can be located in the first or last order of the reference picture list as shown in FIG. 6, and the arrangement order in the list can be determined according to the encoding settings (for example, picture type information, etc.). For example, it can be located in the first for I type and in the last for P / B type, but not limited thereto, and other modified examples are possible.

[0202] Alternatively, an individual reference picture memory can be supported according to a signal (ibc_enabled_flag) that indicates whether to support block matching (or template matching) with the current picture. Here, the signal can be information determined implicitly or generated explicitly.

[0203] Specifically, when the signal is deactivated (e.g., ibc_enabled_flag = 0), it means not to assist block matching in the current picture. When the signal is activated (e.g., ibc_enabled_flag = 1), it means to assist block matching in the current picture and can assist the reference picture memory for this purpose. In this example, it is assumed to provide additional memory, but it is also possible to set to directly assist block matching with the existing memory assisted for the current picture without providing additional memory.

[0204] The reference picture component can include a reference picture interpolation part, and it can be determined whether to perform an interpolation process for pixels in fractional units according to the interpolation accuracy of inter-picture prediction. For example, when having interpolation accuracy in integer units, the reference picture interpolation process is omitted, and when having interpolation accuracy in fractional units, the reference picture interpolation process can be performed.

[0205] In the case of the interpolation filter used in the reference picture interpolation process, it can be implicitly determined according to the encoding settings, or can be explicitly determined from among a plurality of interpolation filters. The configuration of the plurality of interpolation filters can support fixed candidates or support adaptive candidates according to the encoding settings, and the number of candidates can be an integer such as 2, 3, 4 or more. The unit determined explicitly can be determined from among a sequence, a picture, a sub-picture, a slice, a tile, a level, a block, etc.

[0206] Here, the encoding settings can be determined by factors such as the picture type, color component, state information of the target block (e.g., block size, shape, horizontal / vertical length ratio, etc.), inter-picture prediction settings (e.g., motion information encoding mode, motion model selection information, motion vector accuracy selection information, reference picture, reference direction, etc.). The motion vector accuracy can mean the accuracy of the motion vector (i.e., pmv + mvd), but can also be replaced by the accuracy of the motion vector prediction value pmv or the motion vector difference value mvd.

[0207] Here, in the case of an interpolation filter, it can have a filter length of k - taps, where k can be an integer such as 2, 3, 4, 5, 6, 7, 8 or more. The filter coefficients can be derived from mathematical formulas having various coefficient characteristics such as Wiener filters and Kalman filters. The filter information used for interpolation (e.g., filter coefficients, tap information, etc.) can be implicitly defined or derived, or relevant information can be explicitly generated. At this time, the filter coefficients can be configured to include 0.

[0208] The interpolation filter can be applied to a predetermined pixel unit, but the predetermined pixel unit is not limited to integer or decimal units, or can be applied to both integer and decimal units.

[0209] For example, an interpolation filter can be applied to k integer - unit pixels adjacent in the horizontal or vertical direction centered on an interpolation - target pixel (i.e., a decimal - unit pixel).

[0210] Alternatively, an interpolation filter can be applied to p integer - unit pixels and q decimal - unit pixels (p + q = k) adjacent in the horizontal or vertical direction centered on the interpolation - target pixel. At this time, the accuracy of the decimal unit (e.g., 1 / 4 unit) referred to for interpolation can be expressed with the same or lower accuracy (e.g., 2 / 4 → 1 / 2) as the interpolation - target pixel.

[0211] When only one of the x and y components of the interpolation - target pixel is located at a decimal unit, interpolation can be performed based on k pixels adjacent in the component direction of the decimal unit (e.g., the x - axis is horizontal and the y - axis is vertical). When both the x and y components of the interpolation - target pixel are located at decimal units, linear interpolation can be performed based on x pixels adjacent in either the horizontal or vertical direction, and interpolation can be performed based on y pixels adjacent in the remaining direction. At this time, x and y are described as being the same as k, but are not limited thereto, and x and y may be different.

[0212] The interpolation target pixel can be obtained with an interpolation accuracy (1 / m), where m can be an integer of 1, 2, 4, 8, 16, 32 or more. The interpolation accuracy can be implicitly determined according to the encoding settings, or relevant information can be explicitly generated. The encoding settings can be defined based on image type, color component, reference picture, motion information encoding mode, motion model selection information, etc. The explicitly defined unit can be determined from among sequence, sub-picture, slice, tile, block, etc.

[0213] Through the above process, the interpolation filter setting of the target block can be obtained, and reference picture interpolation can be performed based on it. Also, based on the interpolation filter setting, a detailed interpolation filter setting can be obtained, and reference picture interpolation can be performed based on it. That is, assume that the obtained interpolation filter setting of the target block may be obtaining one fixed candidate, or may be obtaining the usability of multiple candidates. Assume that the interpolation accuracy in the example described later is in units of 1 / 16 (for example, there are 15 interpolation target pixels).

[0214] As an example of the detailed interpolation filter setting, for the first pixel unit at a predetermined position, one of the multiple candidates can be adaptively used, and for the second pixel unit at a predetermined position, a predetermined interpolation filter can be used.

[0215] The first pixel unit can be determined among the pixels of the whole fractional unit assisted by the interpolation accuracy (in this example, 1 / 16 to 15 / 16), the number of pixels included in the first pixel unit is a, and a can be determined between 0, 1, 2,... (m - 1).

[0216] The second pixel unit can include the pixels obtained by removing the first pixel unit from the pixels of the whole fractional unit, and the number of pixels included in the second pixel unit can be derived by removing the pixels of the first pixel unit from the total number of interpolation target pixels. In this example, an example of dividing the pixel unit into two is described, but it is not limited to this, and it is also possible to divide it into three or more.

[0217] For example, the first pixel unit can be a unit expressed in multiple units such as 1 / 2, 1 / 4, 1 / 8, etc. As an example, in the case of a 1 / 2 unit, it can be pixels at positions {8 / 16}, in the case of a 1 / 4 unit, it can be pixels at positions {4 / 16, 8 / 16, 12 / 16}, and in the case of a 1 / 8 unit, it can be pixels at positions {2 / 16, 4 / 16, 6 / 16, 8 / 16, 10 / 16, 12 / 16, 14 / 16}.

[0218] The setting of the detailed interpolation filter can be determined according to the encoding setting, and the encoding setting can be defined by the image type, color component, state information of the target block, inter-picture prediction setting, etc. Hereinafter, examples regarding the detailed interpolation filter setting by various encoding elements will be considered. For the convenience of explanation, it is assumed that a is 0 and a is c or more (c is an integer of 1 or more).

[0219] For example, in the case of color components (luminance, chrominance difference), a can be (1, 0), (0, 1), (1, 1). Or, in the case of the motion information encoding mode (merge mode, competition mode), a can be (1, 0), (0, 1), (1, 1), (1, 3). Or, in the case of the motion model selection information (moving motion, non-moving motion A, non-moving motion B), it can be (0, 1, 1), (1, 0, 0), (1, 1, 1), (1, 3, 7). Or, in the case of the motion vector accuracy (1 / 2, 1 / 4, 1 / 8), it can be (0, 0, 1), (0, 1, 0), (1, 0, 0), (0, 1, 1), (1, 0, 1), (1, 1, 0), (1, 1, 1). Or, in the case of the reference picture (current picture, other picture), it can be (0, 1), (1, 0), (1, 1).

[0220] As described above, any one of a plurality of interpolation accuracies can be selected to perform the interpolation process, and when the interpolation process with an adaptive interpolation accuracy is supported (for example, if adaptive_ref_resolution_enabled_flag.0, a predetermined interpolation accuracy is used, and if it is 1, any one of a plurality of interpolation accuracies is used), accuracy selection information (for example, ref_resolution_idx) can be generated.

[0221] The motion estimation and compensation process can be executed according to the interpolation accuracy, and the representation unit and storage unit for the motion vector can also be determined based on the interpolation accuracy.

[0222] For example, when the interpolation accuracy is 1 / 2 unit, the motion estimation and compensation process are performed in 1 / 2 unit, the motion vector is represented in 1 / 2 unit and can be used in the encoding process. Also, the motion vector is stored in 1 / 2 unit and can be referred to in the motion information encoding process of other blocks.

[0223] Or, when the interpolation accuracy is 1 / 8 unit, the motion estimation and compensation process are performed in 1 / 8 unit, the motion vector is represented in 1 / 8 unit and can be used in the encoding process, and can be stored in 1 / 8 unit.

[0224] Also, the motion estimation and compensation process and the motion vector can be executed, represented, and stored in units different from the interpolation accuracy, such as integer, 1 / 2, 1 / 4 units, etc., but this can be adaptively determined according to the inter-picture prediction method / settings (for example, motion estimation / compensation method, motion model selection information, motion information encoding mode, etc.).

[0225] As an example, when it is assumed that the interpolation accuracy is 1 / 8 unit, in the case of the moving motion model, the motion estimation and compensation process are performed in 1 / 4 unit, the motion vector is represented in 1 / 4 unit (in this example, the unit is assumed in the encoding process), and can be stored in 1 / 8 unit. In the case of a motion model other than moving, the motion estimation and compensation process are performed in 1 / 8 unit, the motion vector is represented in 1 / 4 unit, and can be stored in 1 / 8 unit.

[0226] As an example, when assuming that the interpolation accuracy is in units of 1 / 8, in the case of block matching, the motion estimation and compensation processes are performed in units of 1 / 4, the motion vectors are represented in units of 1 / 4, and can be stored in units of 1 / 8. In the case of template matching, the motion estimation and compensation processes are performed in units of 1 / 8, the motion vectors are represented in units of 1 / 8, and can be stored in units of 1 / 8.

[0227] As an example, when assuming that the interpolation accuracy is in units of 1 / 16, in the case of the competition mode, the motion estimation and compensation processes are performed in units of 1 / 4, the motion vectors are represented in units of 1 / 4, and can be stored in units of 1 / 16. In the case of the merge mode, the motion estimation and compensation processes are performed in units of 1 / 8, the motion vectors are represented in units of 1 / 4, and can be stored in units of 1 / 16. In the case of the skip mode, the motion estimation and compensation processes are performed in units of 1 / 16, the motion vectors are represented in units of 1 / 4, and can be stored in units of 1 / 16.

[0228] In summary, based on the inter-picture prediction method or setting, and the interpolation accuracy, the motion estimation and compensation, and the motion vector representation and storage units can be adaptively determined. Specifically, the motion estimation and compensation and the motion vector representation units can be adaptively determined according to the inter-picture prediction method or setting, and it is generally possible that the storage unit of the motion vector is determined according to the interpolation accuracy, but it is not limited thereto, and various deformation examples are possible. Also, in the above example, an example according to one category (for example, motion model selection information, motion estimation / compensation method, etc.) was given, but it is also possible that the setting is determined by mixing two or more categories.

[0229] Also, contrary to the interpolation accuracy information having a predetermined value or being selected from among a plurality of accuracies as described above, the reference picture interpolation accuracy can be determined based on the motion estimation and compensation settings supported according to the inter-picture prediction method or setting. For example, when supporting up to 1 / 8 unit in the case of the moving motion model and up to 1 / 16 unit in the case of the non-moving motion model, the interpolation process can be performed in accordance with the accuracy unit of the non-moving motion model having the highest accuracy.

[0230] That is, reference picture interpolation can be performed according to settings for supported accuracy information such as a motion movement model, a movement outside motion model, a competition mode, a merge mode, a skip mode, etc. In this case, the accuracy information can be determined implicitly or explicitly, and when related information is explicitly generated, it can be included in units such as a sequence, a picture, a sub-picture, a slice, a tile, a block, etc.

[0231] The motion estimation unit means a process of estimating (or searching) whether a target block has a high correlation with a predetermined block of a predetermined reference picture. The size and shape (M×N) of the target block for which prediction is performed can be obtained from the block division unit. As an example, the target block can be determined in the range of 4×4 to 128×128. Inter-picture prediction can generally be performed in units of prediction blocks, but can also be performed in units such as coded blocks and transform blocks according to the settings of the block division unit. Estimation is performed within the estimable range of the reference area, and at least one motion estimation method can be used. The order and conditions of estimation in pixel units can be defined by the motion estimation method.

[0232] Motion estimation can be performed based on a motion estimation method. For example, the area to be compared for the motion estimation process can be the target block in the case of block matching, and can be a predetermined area (template) set centered on the target block in the case of template matching. In the former case, a block with the highest possible correlation can be found within the estimable range of the target block and the reference area, and in the latter case, an area with the highest possible correlation can be found within the estimable range of the template defined according to the coding settings and the reference area.

[0233] Motion estimation can be performed based on a motion model. In addition to a translation motion model that only considers translation, additional motion models can be used for motion estimation and compensation. For example, motion estimation and compensation can be performed using a motion model that considers motions such as rotation, perspective, and zoom-in / zoom-out in addition to translation. This can assist in improving the encoding performance by generating prediction blocks that reflect the various types of motions that occur according to the regional characteristics of the image.

[0234] FIG. 7 is a conceptual diagram showing a motion model other than translation according to an embodiment of the present invention.

[0235] Referring to FIG. 7, as an example of a part of an affine model, an example of expressing motion information based on motion vectors V0 and V1 at a predetermined position is shown. Since motion can be expressed based on a plurality of motion vectors, accurate motion estimation and compensation are possible.

[0236] As in the above example, inter-screen prediction is performed based on a predefined motion model, but inter-screen prediction based on an additional motion model can also be supported. Here, it is assumed that the predefined motion model is a translation motion model and the additional motion model is an affine model, but it is not limited thereto, and various modifications are possible.

[0237] In the case of a translation motion model, it is assumed that motion information (assuming one-direction prediction) can be expressed based on one motion vector, and the control point (reference point) for indicating the motion information is the upper left coordinate, but it is not limited thereto.

[0238] In the case of a non-movement motion model, it can be expressed by movement information of various configurations. In this example, assume a configuration expressed by additional information added to one movement vector (based on the upper left coordinate). Some of the motion estimation and compensation mentioned in the examples described later may not be performed in block units, but may be performed in predetermined sub-block units. At this time, the size and position of the predetermined sub-block can be determined based on each motion model.

[0239] FIG. 8 is an exemplary diagram showing motion estimation in sub-block units according to an embodiment of the present invention. Specifically, it shows motion estimation in sub-block units by an affine model (two movement vectors).

[0240] In the case of a moving motion model, the movement vectors of pixel units included in the target block can be the same. That is, it can have a movement vector that is uniformly applied to pixel units, and motion estimation and compensation can be performed using one movement vector V0.

[0241] In the case of a non-movement motion model (affine model), the movement vectors of pixel units included in the target block may not be the same, and individual movement vectors of pixel units may be required. At this time, the movement vectors of pixel units or sub-block units can be derived based on the movement vectors V0 and V1 at predetermined control point positions of the target block, and motion estimation and compensation can be performed using the derived movement vectors.

[0242] For example, the movement vectors of sub-blocks or pixel units in the target block {for example, (Vx, Vy)} can be derived by the mathematical formulas Vx = (V1x - V0x) × x / M - (V1y - V0y) × y / N + V0x, Vy = (V1y - V0y) × x / M + (V1x - V0x) × y / N + V0y. In the above mathematical formulas, V0 {in this example, (V0x, V0y)} means the movement vector of the upper left side of the target block, and V1 {in this example, (V1x, V1y)} means the movement vector of the upper right side of the target block. Considering complexity, motion estimation and motion compensation of the non-movement motion model can be performed in sub-block units.

[0243] Here, the size of the sub-block (M×N) can be determined according to the encoding settings, and can have a fixed size or be set to an adaptive size. Here, M and N can be integers such as 2, 4, 8, 16 or more, and M and N may or may not be the same. The size of the sub-block can be explicitly generated in units such as sequences, pictures, sub-pictures, slices, tiles, bricks, etc. Or, it can be determined implicitly by a common agreement between the encoder and the decoder, or can be determined by the encoding settings.

[0244] Here, the encoding settings can be defined by one or more elements such as the state information of the target block, the image type, the color component, the inter-picture prediction setting information (motion information encoding mode, reference picture information, interpolation accuracy, motion model selection information, etc.).

[0245] The above example illustrates the process of inducing the size of the sub-block by a motion model other than a predetermined movement and performing motion estimation and compensation based on it. Motion estimation and compensation can be performed on a sub-block or pixel unit by a motion model as in the above example, and detailed description thereof will be omitted.

[0246] Hereinafter, various examples regarding motion information configured according to a motion model will be considered.

[0247] As an example, in the case of a motion model representing rotational motion, the translational motion of a block can be represented by one motion vector, and the rotational motion can be represented by rotation angle information. The rotation angle information can be measured based on a predetermined position (for example, the coordinates of the upper left side) as a reference (0 degrees), and can be represented by k candidates (k is an integer such as 1, 2, 3 or more) having a predetermined interval (for example, an angular difference value of 0 degrees, 11.25 degrees, 22.25 degrees, etc.) within a predetermined angle range (for example, between -90 degrees and 90 degrees).

[0248] Here, the rotation angle information can be encoded by itself during the motion information encoding process, or can be encoded (e.g., prediction + difference value information) based on the motion information of adjacent blocks (e.g., motion vectors, rotation angle information).

[0249] Alternatively, the translational motion of a block can be represented by one motion vector, and the rotational motion of the block can be represented by one or more additional motion vectors. At this time, the number of additional motion vectors can be 1, 2, or an integer greater than or equal to 1, and the control points of the additional motion vectors can be determined from the upper right, lower left, and lower right coordinates, or coordinates within other blocks can be set as the control points.

[0250] Here, the additional motion vectors can be encoded by themselves during the motion information encoding process, or can be encoded (e.g., prediction + difference value information) based on the motion information of adjacent blocks (e.g., motion vectors according to translational motion models or non-translational motion models), or can be encoded (e.g., prediction + difference value information) based on other motion vectors within the block representing rotational motion.

[0251] As an example, in the case of a motion model representing size adjustment or scaling motion such as zoom-in / zoom-out situations, the translational motion of a block can be represented by one motion vector, and the size adjustment motion can be represented by scaling information. The scaling information can be represented by scaling information indicating horizontal or vertical expansion or contraction with respect to a predetermined position (e.g., the coordinates of the upper left corner).

[0252] Here, scaling can be applied in at least one of the horizontal or vertical directions. Also, individual scaling information applied in the horizontal or vertical direction can be supported, or common scaling information applied in both directions can be supported. The width and height of the block to which the scaling is applied at a predetermined position (the coordinates of the upper left corner) can be added to determine the position for motion estimation and compensation.

[0253] Here, the scaling information can be encoded by itself during the motion information encoding process, or can be encoded (e.g., prediction + differential value information) based on the motion information of adjacent blocks (e.g., motion vectors, scaling information).

[0254] Alternatively, the moving motion of a block can be represented by one motion vector, and the size adjustment of the block can be represented by one or more additional motion vectors. At this time, the number of additional motion vectors can be 1, 2, or an integer greater than or equal to 2, and the control points of the additional motion vectors can be determined among the coordinates of the upper right, lower left, or lower right, or the coordinates within other blocks can be set as the control points.

[0255] Here, the additional motion vectors can be encoded by themselves during the motion information encoding process, or can be encoded (e.g., prediction + differential value information) based on the motion information of adjacent blocks (e.g., motion vectors according to a moving motion model or a non-moving motion model), or can be encoded (e.g., prediction + differential value) based on predetermined coordinates within the block (e.g., the coordinates of the lower right).

[0256] Although the case regarding the expression for indicating some motions in the above example has been described, it is also possible to be expressed by motion information for expressing a plurality of motions.

[0257] For example, in the case of a motion model expressing diverse and complex motions, the moving motion of a block can be represented by one motion vector, the rotational motion can be represented by rotation angle information, and the size adjustment can be represented by scaling information. The description regarding each motion can be induced by the examples described above, so the detailed description is omitted.

[0258] Alternatively, the movement of a block can be represented by one motion vector, and other movements of the block can be represented by one or more additional motion vectors. At this time, the number of additional motion vectors can be one, two, or an integer greater than or equal to two. The control points of the additional motion vectors can be determined among the upper right, lower left, and lower right coordinates, or coordinates within other blocks can be set as the control points.

[0259] Here, the additional motion vectors can be encoded by themselves during the motion information encoding process, or can be encoded (e.g., prediction + differential value information) based on the motion information of adjacent blocks (e.g., motion vectors according to a motion model or a non-motion model), or can be encoded (e.g., prediction + differential value information) based on other motion vectors within a block that represents various motions.

[0260] The above description can be related to the affine model, and the case where the number of additional motion vectors is one or two will be mainly described. To summarize, assuming that the number of motion vectors used by the motion model is one, two, or three, and can be regarded as an individual motion model according to the number of motion vectors used to represent the motion information. Also, assume that when there is one motion vector, it is a predefined motion model.

[0261] Multiple motion models for inter-picture prediction can be supported and determined by a signal (e.g., adaptive_motion_mode_enabled_flag) that indicates support for additional motion models. Here, if the signal is 0, the predefined motion model is supported; if the signal is 1, multiple motion models can be supported. The signal can be generated in units such as sequence, picture, sub-picture, slice, tile, block, brick, etc. However, if it cannot be confirmed separately, the value of the signal can be assigned according to the predefined settings. Alternatively, it can be determined implicitly whether to support it according to the encoding settings. Or, it may be determined whether it is implicit or explicit according to the encoding settings. Here, the encoding settings can be defined by one or more elements such as image type, image category (e.g., general image if 0, 360-degree image if 1), color component, etc.

[0262] Whether to support multiple motion models can be determined through the above process. Next, assume that there are two or more additionally supported motion models, and assume that it is determined that multiple motion models are supported in units such as sequence, picture, sub-picture, slice, tile, block, etc., but there can also be some exceptional configurations. In the example described later, assume that motion models A, B, and C are supportable, A is the basically supported motion model, and B and C are additionally supportable motion models.

[0263] Configuration information for the supported motion models can be generated in the above units. That is, configurations of supported motion models such as {A, B}, {A, C}, {A, B, C} are possible.

[0264] For example, indexes (0 to 2) can be assigned to the candidates of the above configurations and selected. If index 2 is selected, the motion model configuration {A, C} is determined; if index 3 is selected, the motion model configuration {A, B, C} is determined.

[0265] Alternatively, information indicating whether or not to support a predetermined motion model can be supported individually. That is, a support flag for B and a support flag for C can be generated. If both flags are 0, it may be the case where only A can be supported. This example may be an example in which processing is performed without generating information indicating whether or not to support a plurality of motion models.

[0266] When a candidate group of motion models to be supported is configured as in the above example, any one of the motion models in the candidate group can be explicitly determined and used in block units, or used implicitly.

[0267] Generally, the motion estimation unit can be a configuration existing in the encoding device, but can also be a configuration included in the decoding device according to a prediction method (for example, template matching, etc.). For example, in the case of template matching, motion estimation can be performed through the adjacent template of the target block in the decoder to obtain the motion information of the target block. At this time, motion estimation related information (for example, motion estimation range, motion estimation method <scan order>, etc.) can be implicitly determined or explicitly generated and included in units such as sequences, pictures, sub-pictures, slices, tiles, blocks.

[0268] The motion compensation unit means a process for obtaining, as a prediction block of the target block, data of a partial block of a predetermined reference picture determined through the motion estimation process. Specifically, based on the motion information (for example, reference picture information, motion vector information, etc.) obtained through the motion estimation process, a prediction block of the target block can be generated from at least one region (or block) of at least one reference picture.

[0269] Motion compensation can be performed based on the motion compensation method as follows.

[0270] In the case of block matching, based on the motion vectors (Vx, Vy) of the target block (M×N) explicitly obtained in the reference picture and the coordinates (Px + Vx, Py + Vy) obtained via the coordinates (Px, Py) of the upper left corner of the target block, the data in the area corresponding to M to the right and N downward can be compensated by the predicted block of the target block.

[0271] In the case of template matching, based on the motion vectors (Vx, Vy) of the target block (M×N) implicitly obtained in the reference picture and the coordinates (Px + Vx, Py + Vy) obtained via the coordinates (Px, Py) of the upper left corner of the target block, the data in the area corresponding to M to the right and N downward can be compensated by the predicted block of the target block.

[0272] Also, motion compensation can be performed based on the following motion model.

[0273] In the case of the translational motion model, based on any of the motion vectors (Vx, Vy) of the target block (M×N) explicitly obtained in the reference picture and the coordinates (Px + Vx, Py + Vy) obtained via the coordinates (Px, Py) of the upper left corner of the target block, the data in the area corresponding to M to the right and N downward can be compensated by the predicted block of the target block.

[0274] In the case of the non-translational motion model, based on the motion vectors (Vmx, Vny) of the m×n sub-blocks implicitly obtained via the multiple motion vectors (V0x, V0y), (V1x, V1y) of the target block (M×N) explicitly obtained in the reference picture and the coordinates (Pmx + Vnx, Pmy + Vny) obtained via the coordinates (Pmx, Pny) of the upper left corner of each sub-block, the data in the area corresponding to M / m to the right and N / n downward can be compensated by the predicted block of the sub-block. That is, the predicted blocks of the sub-blocks can be collected and compensated by the predicted block of the target block.

[0275] A process for selecting optimal motion information for a target block can be performed by a motion information determination unit. Generally, a rate-distortion technique that takes into account the distortion of a block (e.g., the Distortion between a target block and a restored block, such as SAD (Sum of Absolute Difference), SSD (Sum of Square Difference), etc.) and the amount of generated bits by the motion information can be used to determine optimal motion information in terms of encoding cost. A prediction block generated based on the motion information determined through the above process can be transmitted to a subtraction unit and an addition unit. Also, it can be a configuration included in a decoding device according to some prediction methods (e.g., template matching, etc.), and in this case, it can be determined based on the distortion of the block.

[0276] The motion information determination unit can consider inter-picture prediction related setting information such as a motion compensation method and a motion model. For example, when multiple motion compensation methods are supported, the motion compensation method selection information, the corresponding motion vector, reference picture information, etc. can be the optimal motion information. Or, when multiple motion models are supported, the motion model selection information, the corresponding motion vector, reference picture information, etc. can be the optimal motion information.

[0277] A motion information encoding unit can encode the motion information of the target block obtained through the motion information determination process. At this time, the motion information can be composed of information on the image and region referred to for the prediction of the target block. Specifically, it can be composed of information on the referred image (e.g., reference image information, etc.) and information on the referred region (e.g., motion vector information, etc.).

[0278] Also, inter-picture prediction related setting information (or selection information, etc., e.g., motion estimation / compensation method, motion model selection information, etc.) can also be included in the motion information of the target block. Based on the inter-picture prediction related setting, information on the reference image and region (e.g., the number of motion vectors, etc.) can be configured.

[0279] Information on a reference image and a reference region can be combined into one combination to encode motion information, and a combination of information on the reference image and the reference region can be configured as a motion information encoding mode.

[0280] Here, the information on the reference image and the reference region can be obtained based on adjacent blocks or predetermined information (e.g., an image encoded before or after with respect to the current picture, a zero motion vector, etc.).

[0281] The adjacent blocks belong to the same space as the target block, and can be categorized into the block <inter_blk_A> that is closest to the target block, the block <inter_blk_B> that is distantly adjacent and belongs to the same space as the target block, and the block <inter_blk_C> that belongs to a space different from the target block. Information on the reference image and the reference region can be obtained based on adjacent blocks (candidate blocks) belonging to at least one of these categories.

[0282] For example, the motion information of the target block can be encoded based on the motion information or reference picture information of the candidate block, or can be encoded based on information derived from the motion information or reference picture information of the candidate block (or information that has gone through an intermediate value, conversion process, etc.). That is, the motion information of the target block can be predicted from the candidate block and information related thereto can be encoded.

[0283] The motion information of the target block in the present invention can be encoded based on one or more motion information encoding modes. Here, various definitions are possible for the motion information encoding mode, and it can include one or more of a skip mode, a merge mode, a competition mode (Comp mode), etc.

[0284] Based on the above-mentioned template matching tmp, it can be combined with the motion information encoding mode, or supported by a separate motion information encoding mode, or included in all or part of the detailed configuration of the motion information encoding mode. This is premised on the case where it is defined to support template matching at a higher unit (e.g., picture, sub-picture, slice, etc.), but a flag regarding whether to support can be considered as a part of the elements in the setting of inter-picture prediction.

[0285] Based on the method ibc of performing block matching in the current picture described above, it can be combined with the motion information encoding mode, or supported by a separate motion information encoding mode, or included in all or part of the detailed configuration of the motion information encoding mode. This is premised on the case where it is defined to support block matching within the current picture at a higher unit, but a flag regarding whether to support can be considered as a part of the elements in the setting of inter-picture prediction.

[0286] Based on the above-mentioned motion model (affine), it can be combined with the motion information encoding mode, or supported by a separate motion information encoding mode, or included in all or part of the detailed configuration of the motion information encoding mode. This is premised on the case where it is defined to support a motion model other than movement at a higher unit, but a flag regarding whether to support can be considered as a part of the elements in the setting of inter-picture prediction.

[0287] For example, individual motion information encoding modes such as temp_inter, temp_tmp, temp_ibc, temp_affine, etc. can be supported. Or, combined motion information encoding modes such as temp_inter_tmp, temp_inter_ibc, temp_inter_affine, temp_inter_tmp_ibc, etc. can be supported. Or, it can be configured to include candidates based on templates, candidates based on the method of performing block matching within the current picture, and candidates based on affine among the groups of motion information prediction candidates that make up temp.

[0288] Here, temp can mean skip mode (skip), merge mode (merge), or competition mode (comp). As an example, in the case of the skip mode, movement information encoding modes such as skip_inter, skip_tmp, skip_ibc, skip_affine are supported, in the case of the merge mode, merge_inter, merge_tmp, merge_ibc, merge_affine are supported, and in the case of the competition mode, comp_inter, comp_tmp, comp_ibc, comp_affine are supported, which is the same as the above.

[0289] If the skip mode, merge mode, and competition mode are supported and the candidates considering the above elements are included in the movement information prediction candidate groups of each mode, one mode can be selected by a flag that distinguishes the skip mode, merge mode, and competition mode. As an example, if a flag indicating whether it is the skip mode is supported and has a value of 1, the skip mode is selected; if it has a value of 0, a flag indicating whether it is the merge mode is supported, and if it has a value of 1, the merge mode is selected; if it has a value of 0, the competition mode may be selected. And the movement information prediction candidate groups of each mode may include candidates based on inter, tmp, ibc, and affine.

[0290] Or, when multiple movement information encoding modes are supported under one common mode, in addition to the flag for selecting any one of the skip mode, merge mode, and competition mode, it is also possible to support an additional flag for distinguishing the detailed mode of the selected mode. As an example, when the merge mode is selected, it means that a flag for selecting from merge_inter, merge_tmp, merge_ibc, merge_affine, etc., which are the detailed modes related to the merge mode, is additionally supported. Or, a flag indicating whether it is merge_inter is supported, and when it is not merge_inter, it is also possible to support an additional flag for selecting from merge_tmp, merge_ibc, merge_affine, etc.

[0291] All or part of the motion information encoding mode candidates can be supported according to the symbolization setting. Here, the symbolization setting can be defined by one or more elements such as the state information of the target block, image type, image category, color component, inter-screen prediction support setting (for example, whether to support template matching, whether to support block matching in the current picture, motion model support elements other than movement), etc.

[0292] As an example, according to the size of the block, the supported motion information encoding mode can be determined. At this time, the range of the block size can be determined by the size of the first threshold (minimum value) or the size of the second threshold (maximum value), and the size of each threshold can be expressed by the width (W) and height (H) of the block as W, H, W×H, W*H. In the case of the size of the first threshold, W and H can be 4, 8, 16 or an integer greater than or equal to 16, and W*H can be 16, 32, 64 or an integer greater than or equal to 64. In the case of the size of the second threshold, W and H can be 16, 32, 64 or an integer greater than or equal to 64, and W*H can be 64, 128, 256 or an integer greater than or equal to 256. The above range can be determined by either one of the size of the first threshold or the size of the second threshold, or by both of them.

[0293] At this time, the size of the threshold can be fixed or adaptable according to the image (for example, image type, etc.). At this time, the size of the first threshold can be set based on the size of the minimum encoding block, minimum prediction block, minimum transformation block, etc., and the size of the second threshold can be set based on the size of the maximum encoding block, maximum prediction block, maximum transformation block, etc.

[0294] As an example, according to the image type, the supported motion information encoding mode can be determined. At this time, the I image type can include at least one of the skip mode, merge mode, and competition mode. At this time, a method of performing block matching (or template matching) on the current picture, an individual motion information encoding mode regarding the affine model (hereinafter referred to as the term "element") can be supported, or two or more elements can be combined to support the motion information encoding mode. Or, other elements can be configured as a motion information prediction candidate group for a predetermined motion information encoding mode.

[0295] The P / B image type can include at least one of the skip mode, merge mode, and competition mode. At this time, general inter-picture prediction, template matching, execution of block matching in the current picture, an individual motion information encoding mode regarding the affine model (hereinafter referred to as the term "element") can be supported, or two or more elements can be combined to support the motion information encoding mode. Or, other elements can be configured as a motion information prediction candidate group for a predetermined motion information encoding mode.

[0296] FIG. 9 is a flowchart showing the encoding of motion information according to an embodiment of the present invention.

[0297] Referring to FIG. 9, the motion information encoding mode of the target block can be confirmed (S900).

[0298] The motion information encoding mode can be defined by combining and setting predetermined information (for example, motion information, etc.) used for inter-picture prediction. The predetermined information can include one or more of a predicted motion vector, a motion vector difference value, a differential motion vector accuracy, reference image information, a reference direction, motion model information, information on the presence or absence of a residual component, and the like. The motion information encoding mode can include at least one of the skip mode, merge mode, and competition mode, and can also include other additional modes.

[0299] According to the motion information encoding mode, the composition and setting of explicitly generated information and implicitly defined information can be determined. At this time, in the case of explicit, each piece of information can be generated individually or in a combined form (for example, in an index form).

[0300] As an example, the skip mode or the merge mode can be defined based on a predicted motion vector, reference picture information, and a reference direction (one) predetermined index information. At this time, the index can be configured as one or more candidates, and each candidate can be set based on the motion information of a predetermined block (for example, configured as a combination of the predicted motion vector, reference picture information, and reference direction of the corresponding block). Also, the motion vector difference value can be implicitly processed (for example, as a zero vector).

[0301] As an example, the competition mode can be defined based on a predicted motion vector, reference picture information, and a reference direction (one or more) predetermined index information. At this time, the index can support each piece of information individually or support a combination of two or more pieces of information. At this time, the index can be configured as one or more candidates, and each candidate can be set based on the motion information of a predetermined block (for example, the predicted motion vector, etc.) or can be configured with a preset value (for example, processing the interval from the current picture as 1, 2, 3, etc., and processing the reference direction information as L0, L1, etc.) (for example, reference picture information, reference direction information, etc.). Also, the motion vector difference value can be explicitly generated, and differential motion vector accuracy information can also be additionally generated according to the motion vector difference value (for example, when it is not a zero vector).

[0302] As an example, in the skip mode, the presence or absence information of the residual component is implicitly processed (for example, processed as <cbf_flag = 0>), and the motion model can be implicitly processed with a predetermined value (for example, supporting a translational motion model).

[0303] As an example, in the merging mode or the competing mode, the presence or absence information of the residual component can be explicitly generated, and the motion model can be explicitly selected. At this time, can information for classifying the motion model be generated within one motion information encoding mode, or can a separate motion information encoding mode be supported?

[0304] When one index is supported in the above description, index selection information does not occur. When two or more indexes are supported, index selection information can occur.

[0305] In the present invention, a candidate group composed of two or more indexes is referred to as a motion information prediction candidate group. Also, depending on various predetermined information and settings, the existing motion information encoding mode can be changed, or a new motion information encoding mode can be supported.

[0306] Referring to FIG. 9, the predicted motion vector of the target block can be derived (S910).

[0307] The predicted motion vector can be determined to be a preset value, or can be selected from among those constituting a predicted motion vector candidate group. In the former case, it means the case of being implicitly determined. In the latter case, index information for selecting the predicted motion vector can be explicitly generated.

[0308] Here, in the case of the preset value, can it be set based on the motion vector of one block at a predetermined position (for example, blocks in the left, up, upper left, upper right, lower left directions, etc.), or can it be set based on the motion vectors of two or more blocks, or can it be set to a default value (for example, a zero vector, etc.)?

[0309] The predicted motion vectors may have the same or different candidate group configurations according to the motion information encoding mode. As an example, the skip mode / merge mode / competition mode can support a, b, and c predicted candidates respectively, and the number of each candidate can be configured to be the same or different. At this time, a, b, and c can be composed of integers greater than or equal to 1 such as 2, 3, 5, 6, 7. The candidate group configuration order, settings, etc. will be described through other embodiments. In the example described later, the candidate group configuration setting of the skip mode will be described assuming the same as the merge mode, but it is not limited to this, and some may have different configurations.

[0310] The predicted motion vector obtained based on the index information can be directly used for restoring the motion vector of the target block, or can be adjusted based on the reference direction and the distance between reference images.

[0311] For example, in the case of the competition mode, although the reference image information can be generated separately, when the distance between the current picture and the reference picture of the target mode is different from the distance between the picture including the predicted motion vector candidate block and the reference picture of the block, this can be adjusted to match the distance between the current picture and the reference picture of the target block.

[0312] Also, the number of motion vectors induced based on the motion model can be different from each other. That is, in addition to the motion vector of the control point position on the upper left side, the motion vectors of the control point positions on the upper right side and the lower left side can be further induced based on the motion model.

[0313] Referring to FIG. 9, the motion vector difference value of the target block can be restored (S920).

[0314] The motion vector difference value can be induced to a zero vector value in the skip mode and the merge mode, and the difference value information of each component can be restored in the competition mode.

[0315] The motion vector can be represented based on a preset predetermined accuracy (for example, set based on interpolation accuracy or set based on motion model selection information, etc.). Or, it can be represented based on any one of a plurality of accuracies, and predetermined information therefor can be generated. At this time, the predetermined information can be related to the selection of the motion vector accuracy.

[0316] For example, when one component of the motion vector is 32 / 16, motion component data for 32 can be obtained based on the default setting of the accuracy of 1 / 16 pixel unit. Or, based on the explicit setting of the accuracy of 2 pixel units, the motion component can be converted to 2 / 1, and motion component data for 2 can be obtained.

[0317] The above configuration can be a method for accuracy processing of a motion vector obtained by adding a predicted motion vector and a motion vector difference value having the same accuracy or obtained through the same accuracy conversion process.

[0318] Also, the motion vector difference value can be expressed according to a preset predetermined accuracy or can be expressed by any one of a plurality of accuracies. In the latter case, accuracy selection information can be generated.

[0319] The above accuracy can include at least one of 1 / 16, 1 / 8, 1 / 4, 1 / 2, 1, 2, 4 pixel units, and the number of candidate groups can be an integer of 1, 2, 3, 4, 5, or more. Here, the supported accuracy candidate group configuration (for example, classified by the number, candidates, etc.) can be explicitly defined in units such as sequence, picture, sub-picture, slice, tile, block, etc. Or, it can be implicitly defined according to the coding setting. Here, the coding setting can be defined by at least one element such as image type, reference image, color component, motion model selection information, etc.

[0320] Next, a case where accuracy candidates are configured according to various coding elements will be described. The accuracy will be described assuming an explanation regarding the motion vector difference value, but this is also applicable to motion vectors with similar or identical applications.

[0321] For example, in the case of a moving motion model, candidate group configurations such as {1 / 4, 1}, {1 / 4, 1 / 2}, {1 / 4, 1 / 2, 1}, {1 / 4, 1, 4}, {1 / 4, 1 / 2, 1, 4} are possible. In the case of a non-moving motion model, accuracy candidate group configurations such as {1 / 16, 1 / 4, 1}, {1 / 16, 1 / 8, 1}, {1 / 16, 1, 4}, {1 / 16, 1 / 4, 1, 2}, {1 / 16, 1 / 4, 1, 4} are possible. This assumes that the supported minimum accuracy is 1 / 4 pixel unit in the former case and 1 / 16 pixel unit in the latter case, and additional accuracies other than the minimum accuracy can be included in the candidate group.

[0322] Or, in the case where the reference image is a different picture, candidate group configurations such as {1 / 4, 1 / 2, 1}, {1 / 4, 1, 4}, {1 / 4, 1 / 2, 1, 4} are possible. In the case where the reference image is the current picture, candidate group configurations such as {1, 2}, {1, 4}, {1, 2, 4} are possible. At this time, in the latter case, it can be a configuration where interpolation in decimal units for block matching etc. is not performed. If interpolation in decimal units is performed, it can also be configured to include accuracy candidates in decimal units.

[0323] Or, in the case where the color component is the luminance component, candidate groups such as {1 / 4, 1 / 2, 1}, {1 / 4, 1, 4}, {1 / 4, 1 / 2, 1, 4} are possible. In the case where the color component is the chrominance component, candidate groups such as {1 / 8, 1 / 4}, {1 / 8, 1 / 2}, {1 / 8, 1}, {1 / 8, 1 / 4, 1 / 2}, {1 / 8, 1 / 2, 2}, {1 / 8, 1 / 4, 1 / 2, 2} are possible. At this time, in the latter case, a candidate group proportional to the luminance component can be configured according to the composition ratio of the color components (e.g., 4:2:0, 4:2:2, 4:4:4, etc.), or an individual candidate group can be configured.

[0324] The above example describes the case where there are multiple accuracy candidates for each encoding element, but it is also possible that there is only one candidate configuration (i.e., expressed with the already set minimum accuracy).

[0325] In summary, in the competitive mode, based on the differential motion vector accuracy, the motion vector difference value with the minimum accuracy can be restored, and in the skip mode and the merge mode, the motion vector difference value with a zero vector value can be induced.

[0326] Here, depending on the motion model, the number of restored motion vector difference values can be different from each other. That is, in addition to the motion vector difference value at the control point position on the upper left side, depending on the motion model, the motion vector difference values at the control point positions on the upper right side and the lower left side can be further induced.

[0327] Also, one differential motion vector accuracy can be applied to a plurality of motion vector difference values, or individual differential motion vector accuracies can be applied to each motion vector difference value. This can be determined according to the encoding settings.

[0328] Also, when at least one motion vector difference value is not 0, information regarding the differential motion vector accuracy can be generated, and when not, it can be omitted, but it is not limited to this.

[0329] Referring to FIG. 9, the motion vector of the target block can be restored (S930).

[0330] The predicted motion vector and the motion vector difference value obtained through the previous process of this step are obtained, and by adding these, the motion vector of the target block can be restored. At this time, if any one is selected from among a plurality of accuracies and restored, an accuracy unification process can be performed.

[0331] For example, the motion vector obtained by adding a predicted motion vector (related to motion vector accuracy) and a motion vector difference value can be restored based on accuracy selection information. Or, the motion vector difference value (related to differential motion vector accuracy) can be restored based on accuracy selection information, and a motion vector can be obtained by adding the restored motion vector difference value and the predicted motion vector.

[0332] Referring to FIG. 9, motion compensation can be performed based on the motion information of the target block (S940).

[0333] The motion information can include the predicted motion vector, reference picture information other than the motion vector difference value, reference direction, motion model information, etc. described above. In the case of the skip mode and the merge mode, some of the information (for example, reference picture information, reference direction, etc.) can be implicitly determined, and the competition mode can explicitly process the relevant information.

[0334] Based on the motion information obtained through the above process, motion compensation can be performed to obtain a predicted block.

[0335] Referring to FIG. 9, decoding of the residual component of the target block can be performed (S950).

[0336] By adding the residual component obtained through the above process and the predicted block, the target block can be restored. At this time, according to the motion information coding mode, explicit or implicit processing of the presence or absence information of the residual component is possible.

[0337] FIG. 10 is an arrangement diagram of a target block and blocks adjacent thereto according to an embodiment of the present invention.

[0338] Referring to FIG. 10, adjacent blocks <inter_blk_A> in the left, upper, upper-left, upper-right, lower-left directions, etc. centered on the target block, and blocks adjacent to the block corresponding to the target block in the central, left, right, upper, lower, upper-left, upper-right, lower-left, lower-right directions, etc. centered on the block corresponding to the target block in a space (Col_Pic) that is not temporally the same can be configured as candidate blocks for predicting the motion information (e.g., motion vector, etc.) of the target mode.

[0339] In the case of the said <inter_blk_A>, it can be an example where the direction of adjacent blocks is determined based on an encoding order such as raster scan or z-scan, and existing directions are removed according to various scan orders, or adjacent blocks in the right, lower, and lower-right directions can be further configured as candidate blocks.

[0340] Referring to FIG. 10, Col_Pic can be an image adjacent before or after the current image (e.g., when the interval between images is 1) based on the current image, and the corresponding block can be set to have the same position in the image as the target block.

[0341] Or, Col_Pic can be an image with an already defined interval between images (e.g., the interval between images is z. z is an integer of 1, 2, 3) based on the current image, and the corresponding block can be set at a position moved by a predetermined displacement vector (Disparity Vector, parallax vector) from a predetermined coordinate (e.g., the upper-left side) of the target block, but the displacement vector can be set to a predefined value.

[0342] Or, Col_Pic can be set based on the motion information (e.g., reference image) of adjacent blocks, the displacement vector is set based on the motion information (e.g., motion vector) of adjacent blocks, and the position of the corresponding block can be determined.

[0343] At this time, k adjacent blocks can be referenced, where k can be an integer of 1, 2, or more. If k is 2 or more, Col_Pic and the displacement vector can be obtained based on operations such as the maximum value, minimum value, median value, weighted average value, etc. of the motion information (e.g., reference image or motion vector) of the adjacent blocks. For example, the displacement vector can be set as the motion vector of the left or upper block, or can be set as the median value or average value of the motion vectors of the left and lower-left blocks.

[0344] The setting of the temporal candidates can be determined based on, for example, the setting of the motion information configuration. For example, depending on whether the motion information in block units or sub-block units to be included in the motion information prediction candidate group is configured in block units or sub-block units, the position of Col_Pic, the position of the corresponding block, etc. can be determined. As an example, when the motion information in sub-block units is obtained, the block at the position shifted by a predetermined displacement vector can be set as the position of the corresponding block.

[0345] The above example shows the case where the information regarding the position of Col_Pic and the corresponding block is implicitly determined, and the relevant information can also be generated explicitly in units such as sequences, pictures, slices, tile groups, tiles, bricks, etc.

[0346] On the other hand, when the divided unit partitions such as sub-pictures, slices, tiles, etc. are not limited to one picture (e.g., the current picture) but are shared among images, when the position of the block corresponding to Col_Pic is different from the divided unit to which the target block belongs (i.e., when it crosses or is outside the divided unit boundary to which the target block belongs), it can be determined based on the motion information of the subsequent candidates based on a predetermined priority order. Or, the position of the target block within the image can be set to be the same.

[0347] The motion information of the spatially adjacent blocks and temporally adjacent blocks described above can be included in the motion information prediction candidate group of the target block, and these are referred to as spatial candidates and temporal candidates. The spatial candidates and temporal candidates can support a and b, where a and b can be integers from 1 to 6 or more. At this time, a can be greater than or equal to b.

[0348] The spatial candidates and temporal candidates can support a fixed number, or a variable number (including, for example, 0) according to the composition of the preceding motion information prediction candidates.

[0349] Also, a block <inter_blk_C> that is not immediately adjacent to the target block can be considered in the composition of the candidate group.

[0350] As an example, the motion information of blocks located at a predetermined interval difference with respect to the target block (for example, the upper left coordinate) can be included in the candidate group. The interval can be p and q according to each component of the motion vector, where p and q can be integers greater than or equal to 2 such as 0, 2, 4, 8, 16, 32, and p and q can be the same or different. That is, assuming that the upper left coordinate of the target block is (m, n), the motion information of the blocks including the positions (m±p, n±q) can be included in the candidate group.

[0351] As an example, the motion information of the blocks whose encoding has been completed previously for the target block can be included in the candidate group. At this time, based on a predetermined encoding order (for example, raster scan, Z scan, etc.) with respect to the target block, a predetermined number of blocks whose encoding has been completed most recently can be considered as candidates, and the blocks with an older encoding order can be removed from the candidates according to the progress of encoding in a first-in first-out (FIFO) manner.

[0352] In the above example, it means the case of including the motion information of the non-closest block whose encoding has already been completed based on the target mode, and this is referred to as a statistical candidate.

[0353] At this time, as the reference block for obtaining statistical candidates, the target block can be set, but the upper block of the target block (for example, when the target block is the coded block <1> / prediction block <2>, the upper block can be the maximum coded block <1> / coded block <2>, etc.) can also be set. The statistical candidates can support k, and k can be an integer from 1 to 6 or more.

[0354] The motion information obtained by combining the motion information such as the spatial candidate, temporal candidate, and statistical candidate can be included in the candidate group. For example, a candidate obtained by applying a weighted average to each component of a plurality (two, three, or more) of motion information already included in the candidate group can be obtained, and a candidate obtained through processes such as the median, maximum value, and minimum value for each component of the plurality of motion information can be obtained. This is called a combined candidate, and r combined candidates can be supported, and r can be an integer of 1, 2, 3, or more.

[0355] When the candidates used for combination are not the same as the information of the reference image, the reference image information of the combined candidate can be set based on the reference image information of any of the candidates, or can be set as a predefined value.

[0356] The statistical candidates or combined candidates can support a fixed number, or can support a variable number according to the configuration of the previous motion information prediction candidates.

[0357] Also, default candidates with already set values such as (s, t) can be included in the candidate group, and a variable number (0, 1, or more integers) can be supported according to the configuration of the previous motion information prediction candidates. At this time, s and t can be set including 0, and can also be set based on the sizes of various blocks (for example, the horizontal or vertical length of the maximum coded / predicted / transformed block, the minimum coded / predicted / transformed block, etc.).

[0358] According to the motion information encoding mode, a motion information prediction candidate group can be formed from among the candidates, and may be formed including other additional candidates. Also, according to the motion information encoding mode, the configuration settings of the candidate group can be the same or different.

[0359] For example, the skip mode and the merge mode can commonly form a candidate group, and the competition mode can form an individual candidate group.

[0360] Here, the configuration settings of the candidate group can be defined by the category and position of the candidate block (for example, determined from among left / up / upper left / upper right / lower left directions, and the sub-block position where motion information is acquired within the determined direction, etc.), the number of candidate groups (for example, the total number, the maximum number by category, etc.), the candidate configuration method (for example, the priority by category, the priority within the category, etc.), and so on.

[0361] The number (total number) of the motion information prediction candidate group can be k, and k can be an integer from 1 to 6 or more. At this time, when the number of candidate groups is one, it means that candidate group selection information does not occur, and the motion information of the already defined candidate block is set as the predicted motion information. When the number of candidate groups is two or more, candidate group selection information can occur.

[0362] The category of the candidate block can be any one of inter_blk_A, inter_blk_B, and inter_blk_C. At this time, inter_blk_A can be the category that is basically included, and another category can be the category that is additionally supported, but is not limited thereto.

[0363] The above description can be related to the configuration of motion information prediction candidates for a motion model other than movement, and for a motion model other than movement (affine model), the same or similar candidate blocks can be used / referred to in the configuration of the candidate group. On the other hand, in the case of an affine model, since not only the number of motion vectors but also the motion characteristics are different from those of the moving motion model, it can have a configuration different from that of the candidate group configuration of the moving motion model.

[0364] For example, when the motion model of the candidate block is an affine model, the configuration of the motion vector set of the candidate block can be directly included in the candidate group. As an example, when the coordinates of the upper left and upper right are control points, the motion vector set of the coordinates of the upper left and upper right of the candidate block can be included in the candidate group as one candidate.

[0365] Or, when the motion model of the candidate block is a moving motion model, a combination of motion vectors of the candidate block set based on the position of the control point can be configured as one set and included in the candidate group. As an example, when the coordinates of the upper left, upper right, and lower left are control points, the motion vector of the upper left control point is predicted based on the motion vectors of the left, upper, and upper left blocks of the target block (for example, in the case of a moving motion model), the motion vector of the upper right control point is predicted based on the motion vectors of the upper and upper right blocks of the target block (for example, in the case of a moving motion model), and the motion vector of the lower left control point can be predicted based on the motion vectors of the left and lower left blocks of the target block. Although the case where the motion model of the candidate block set based on the position of the control point in the above example is a moving motion model has been described, even in the case of an affine model, it is also possible to obtain or derive only the motion vectors of the control point positions. That is, the motion vectors can be obtained or derived from among the motion vectors of the upper left, upper right, lower left, and lower right control points of the candidate block for the motion vectors of the upper left, upper right, and lower left control points of the target block.

[0366] In summary, when the motion model of the candidate block is an affine model, the set of motion vectors of the block can be included in the candidate group (A). When the motion model of the candidate block is a motion model, it is considered as a candidate for the motion vector of a predetermined control point of the target block, and the set of motion vectors obtained according to the combination of each control point can be included in the candidate group (B).

[0367] At this time, the candidate group can be configured using only one of the A or B methods, or the candidate group can be configured using both the A and B methods. And the A method can be configured first and the B method can be configured later, but it is not limited to this.

[0368] FIG. 11 shows an exemplary diagram related to statistical candidates according to an embodiment of the present invention.

[0369] Referring to FIG. 11, the blocks corresponding to A to L in FIG. 11 mean blocks for which encoding has been completed at a predetermined interval based on the target mode.

[0370] When the candidate block for the motion information of the target block is limited to the block adjacent to the target block, it may occur that various characteristics of the image are not reflected. For this reason, the blocks for which encoding has been completed before the current picture can also be considered as candidate blocks. The category inter_blk_B has already been mentioned in this regard. This is called statistical candidates.

[0371] Since the number of candidate blocks included in the motion information prediction candidate group is limited, efficient management of statistical candidates for this may be necessary.

[0372] (1) Blocks at preset positions having a predetermined distance interval based on the target block can be considered as candidate blocks.

[0373] As an example, based on a predetermined coordinate of the target block (for example, the coordinate of the upper left side, etc.), limiting to the x component, blocks having a certain interval can be considered as candidate blocks, and examples such as (-4, 0), (-8, 0), (-16, 0), (-32, 0) are possible.

[0374] As an example, based on a predetermined coordinate of the target block, limiting to the y component, blocks having a certain interval can be considered as candidate blocks, and examples such as (0, -4), (0, -8), (0, -16), (0, -32) are possible.

[0375] As an example, based on a predetermined coordinate of the target block, blocks having a certain non-zero interval as the x and y components can be considered as candidate blocks, and examples such as (-4, -4), (-4, -8), (-8, -4), (-16, 16), (16, -16) are possible. In this example, it is possible that the x component has a positive sign depending on the y component.

[0376] Candidate blocks considered as statistical candidates can also be limitedly determined according to the setting of encoding. For the example described later, refer to FIG. 11.

[0377] Candidate blocks can be classified into statistical candidates according to whether they belong to a predetermined unit. Here, the predetermined unit can be determined from among the maximum encoded block, block, tile, slice, subpicture, picture, etc.

[0378] For example, when selecting candidate blocks limited to blocks belonging to the same maximum encoded block as the target block, A to C can be the targets.

[0379] Or, when selecting candidate blocks limited to blocks belonging to the maximum encoded block to which the target block belongs and the block belonging to one maximum encoded block on the left, A to C, H, I can be the targets.

[0380] Or, when selecting candidate blocks by limiting them to the maximum coded block to which the target block belongs and the blocks belonging to the upper maximum coded block, A to C, E, and F can be the targets.

[0381] Or, when selecting candidate blocks by limiting them to the blocks belonging to the same slice as the target block, A to I can be the targets.

[0382] Candidate blocks can be classified into statistical candidates depending on whether they are located in a predetermined direction with respect to the target block. Here, the predetermined direction can be defined as, for example, the left, upper, upper left, upper right directions, etc.

[0383] For example, when selecting candidate blocks by limiting them to the blocks located in the left and upper directions with respect to the target block, B, C, F, I, and L may be the targets.

[0384] Or, when selecting candidate blocks by limiting them to the blocks located in the left, upper, upper left, and upper right directions with respect to the target block, A to L can be the targets.

[0385] As in the above example, even if candidate blocks considered as statistical candidates are selected, there can be a priority order for including them in the motion information prediction candidate group. That is, only the number supported by the statistical candidates can be selected from the priority order.

[0386] The priority order can support a predetermined order and can also be defined by various coding elements. The coding elements can be defined based on, for example, the distance between the candidate block and the target mode (e.g., whether it is a short distance. The distance between blocks can be confirmed based on the x and y components), the relative direction of the candidate block with respect to the target block (e.g., the left, upper, upper left, upper right directions. The order of left → up → upper right → left, etc.).

[0387] (2) Blocks that have been coded according to a predetermined coding order with respect to the target block can be considered as candidate blocks.

[0388] It should be understood that the examples described below assume the case of following raster scan (e.g., maximum coded block), Z scan (e.g., coded block, prediction block, etc.), or the content described below can be applied with modifications according to the scan order.

[0389] In the case of the example of (1) described above, similar to being able to support the priority for the configuration of the motion information prediction candidate group, in this embodiment as well, a predetermined priority can be supported. The priority can be determined by various coding elements, but for the sake of convenience of explanation, assume the case of being determined according to the coding order. In the examples described below, the candidate block and the priority will be considered together for explanation.

[0390] Candidate blocks considered as statistical candidates can be limitedly determined according to the coding settings.

[0391] Candidate blocks can be classified into statistical candidates depending on whether they belong to a predetermined unit. Here, the predetermined unit can be determined from among maximum coded block, brick, tile, slice, sub-picture, picture, etc.

[0392] For example, when selecting blocks belonging to the same picture as the target block as candidate blocks, candidate blocks and priorities such as J-D-E-F-G-K-L-H-I-A-B-C can be targeted.

[0393] Or, when selecting blocks belonging to the same slice as the target block as candidate blocks, candidate blocks and priorities such as D-E-F-G-H-I-A-B-C can be targeted.

[0394] Candidate blocks can be classified into statistical candidates depending on whether they are located in a predetermined direction with respect to the target block. Here, the predetermined direction can be determined as left, upward, etc.

[0395] For example, when selecting candidate blocks by limiting them to the blocks located to the left of the target block, candidate blocks such as K-L-H-I-A-B-C and their priorities can be targeted.

[0396] Or, when selecting candidate blocks by limiting them to the blocks located above the target block, candidate blocks such as E-F-A-B-C and their priorities can be targeted.

[0397] The above description is given assuming the case where candidate blocks and priorities are combined, but it is not limited to this, and various candidate blocks and priority settings are possible.

[0398] It has been mentioned that the priorities included in the motion information prediction candidate group can be supported through the descriptions of (1) and (2) above. One or more priorities can be supported, but whether one priority for the entire candidate blocks can be supported, or whether the candidate blocks are classified into two or more categories and individual priorities according to the categories can be supported. In the former case, it can be an example of selecting k candidate blocks with one priority, and in the latter case, it can be an example of selecting p and q candidate blocks (p + q = k) with each priority (for example, two categories). At this time, the category can be divided based on a predetermined coding element, and the coding element can include whether it belongs to a predetermined unit, whether the candidate block is located in a predetermined direction with respect to the target block, and the like.

[0399] An example of selecting candidate blocks according to whether they belong to a predetermined unit (for example, a division unit) described through (1) and (2) above has been explained. Apart from the division unit, candidate blocks can be selected by limiting them to the blocks belonging within a predetermined range based on the target mode. For example, it can be confirmed by boundary points that define ranges such as the minimum and maximum values of the x or y components like (x1, y1), (x2, y2), (x3, y3), etc. The values such as x1 to y3 can be integers such as 0, 4, 8, etc. (based on the absolute value standard).

[0400] Based on any one of the above-mentioned settings (1) or (2), statistical candidates can be supported, or statistical candidates in a form where (1) and (2) are mixed can be supported. It has been considered how to select candidate blocks for statistical candidates according to the above description. The management and update of statistical candidates and the case where they are included in the motion information prediction candidate group will be described later.

[0401] The composition of the motion information prediction candidate group is generally carried out in block units because it is highly likely that the motion information of adjacent blocks centered on the target block is the same or similar. That is, in the case of spatial candidates and temporal candidates, they can be composed based on the target mode.

[0402] On the other hand, since statistical candidates target positions where the candidate blocks are not adjacent to the target block, they can be composed based on a predetermined unit. The predetermined unit can be a higher-level block including the target block.

[0403] For example, when the target block is a prediction block, can statistical candidates be composed in prediction block units, or can statistical candidates be composed in coding block units?

[0404] Or, when the target block is a coding block, can statistical candidates be composed in coding block units, or can statistical candidates be composed in ancestor block units (for example, the depth information has a difference of 1 or more from the target block)? At this time, the ancestor block can also be a unit including the maximum coding block or a unit obtained based on the maximum coding block (for example, an integer multiple of the maximum coding block).

[0405] As in the above example, statistical candidates can be composed based on the target block or a higher-level unit. In the example described later, it is assumed that statistical candidates are composed based on the target block.

[0406] In addition, memory initialization for statistical candidates can be performed based on a predetermined unit. The predetermined unit can be determined from among a picture, a sub-picture, a slice, a tile, a brick, a block, etc. In the case of a block, it can be set based on the maximum coded block. For example, memory initialization for statistical candidates can be performed based on an integer number (1, 2, or an integer greater than or equal to 2) of maximum coded blocks, or based on the row or column unit of the maximum coded block.

[0407] Next, statistical candidate management and update will be described. Memory for managing statistical candidates can be prepared and can store motion information for a maximum of k blocks. At this time, k can be an integer from 1 to 6 or an integer greater than or equal to 6. The motion information stored in the memory can be determined from among a motion vector, a reference picture, a reference direction, etc. Here, the number of motion vectors (for example, 1 to 3, etc.) can be determined based on motion model selection information. For the sake of convenience of explanation, the following description will mainly focus on the case where a block that has been coded in a predetermined coding order with respect to a target block is considered as a candidate block.

[0408] (case1) The motion information of a block that was coded prior to the coding order can be included as a candidate in the order of precedence. Also, when updated by additional candidates when the maximum number is filled, the candidate in the previous order is removed and the order can be incremented by 1. In the following examples, x means a blank where the candidate configuration has not yet been formed.

[0409] (Example) 1order - [a, x, x, x, x, x, x, x, x, x] 2order - [a, b, x, x, x, x, x, x, x, x] ... 9order - [a, b, c, d, e, f, g, h, i, x] 10order - [a, b, c, d, e, f, g, h, i, j] → j is added. It is full. 11order - [b, c, d, e, f, g, h, i, j, k] → Append k. Remove the leading a and shift the order

[0410] (Case 2) If there is duplicate motion information, the existing candidates are removed, and the order of the existing candidates is advanced and adjusted.

[0411] At this time, duplication means that the motion information is the same. This can be defined by the motion information coding mode. As an example, in the case of the merge mode, it is possible to determine whether there is duplication based on the motion vector, reference picture, reference direction, etc., and in the case of the competition mode, it is possible to determine whether there is duplication based on the motion vector and reference picture. At this time, in the case of the competition mode, if the reference pictures of each candidate are the same, it is possible to determine whether there is duplication through comparison of the motion vectors, and if the reference pictures of each candidate are different, it is possible to determine whether there is duplication through comparison of the motion vectors scaled based on the reference pictures of each candidate, but it is not limited to this.

[0412] (Exemplification) 1order - [a, b, c, d, e, f, g, x, x, x] 2order - [a, b, c, d, e, f, g, d, x, x] → d is included, but there is duplication - [a, b, c, e, f, g, d, x, x, x] → Remove the existing d and shift the order

[0413] (Case 3) If there is duplicate motion information, the candidate can be marked and managed as a candidate to be stored long - term (long - term candidate) based on a predetermined condition.

[0414] A separate memory for the long - term candidate can be supported, and information such as the occurrence frequency can be additionally stored and updated in addition to the motion information. In the following exemplification, the memory for the candidate to be stored long - term can be expressed by ().

[0415] (Exemplification) 1order - [(), b, c, d, e, f, g, h, i, j, k]

[0416] 2 - order - [(), b, c, d, e, f, g, h, i, j, k] → e has the following order. Duplication - [(e), b, c, d, f, g, h, i, j, k] → Move to the beginning. Long - term marking

[0417] 3 - order - [(e), c, d, f, g, h, i, j, k, l] → Add l and remove b. ... 10 - order - [(e), l, m, n, o, p, q, r, s, t] → l has the following order. Duplication - [(e, l), m, n, o, p, q, r, s, t] → Long - term marking 11 - order - [(e, l), m, n, o, p, q, r, s, t] → l has the following order. Three duplications - [(l, e), m, n, o, p, q, r, s, t] → Change the order of the long - term candidates

[0418] The above examples show the cases when integrated with existing statistical candidates (short - term candidates), but they can be managed separately from short - term candidates, and the number of long - term candidates can be 0, 1, 2 or an integer greater than or equal to 2.

[0419] In the case of short - term candidates, they can be managed in a first - in - first - out manner based on the encoding order. In the case of long - term candidates, a configuration is possible where they can be managed according to the frequency based on the candidates with duplication occurring among short - term candidates. In the case of long - term candidates, the candidate order can be changed according to the frequency during the memory update process, but it is not limited to this.

[0420] The above examples can be some cases when the candidate blocks are classified into two or more categories and priorities according to the categories are supported. When constructing the motion information prediction candidate group including statistical candidates, a and b short - term and long - term candidates can be respectively constructed according to the priorities within each candidate. Here, a and b can be 0, 1, 2 or an integer greater than or equal to 2.

[0421] The above cases 1 to 3 are examples based on statistical candidates, and configurations of (case1 + case2) or (case1 + case3) are possible, and statistical candidates can be managed in various other modified and additional forms.

[0422] FIG. 12 is a conceptual diagram regarding statistical candidates according to a movement model outside movement in one embodiment of the present invention.

[0423] Statistical candidates can be candidates supported not only by a movement movement model but also by a movement model outside movement. In the case of a movement movement model, since the occurrence frequency is high, there are many blocks that can be referred to, such as spatial or temporal candidates. However, in the case of a movement model outside movement, since the occurrence frequency is low, statistical candidates other than spatial or temporal candidates may be required.

[0424] Referring to FIG. 12, a case (hmvp) is shown in which movement information {TV0, TV1} of a target block is predicted based on movement information {PV0, PV1} of a block <inter_blk_B> that belongs to the same space as the target block and is distantly adjacent.

[0425] Compared with the statistical candidates of the movement movement model described above, the number of motion vectors among the movement information stored in the memory may increase. Therefore, the basic concept regarding statistical candidates can be the same as or similar to that of the movement movement model, and related explanations can be induced with a configuration that varies the number of motion vectors. Next, the part regarding the differences due to the movement model outside movement will be described with emphasis later. In the example to be described later, a case of an affine model with three control points (for example, v0, v1, v2) is assumed.

[0426] In the case of an affine model, the movement information stored in the memory can be determined from a reference picture, a reference direction, and motion vectors v0, v1, v2. Here, the stored motion vectors can be stored in various forms. For example, the motion vectors can be stored as they are, or predetermined difference value information can be stored.

[0427] FIG. 13 is an exemplary diagram of the configuration of motion information of each control point position stored as a statistical candidate according to an embodiment of the present invention.

[0428] Referring to FIG. 13, in (a), the motion vectors of v0, v1, and v2 of the candidate block are represented by v0, v1, and v2, and in (b), the motion vector of v0 of the candidate block is represented by v0, and the motion vectors of v1 and v2 are represented by the difference values v*1 and v*2 from the motion vector of v0.

[0429] That is, the motion vectors of each control point position can be stored as they are, or the difference values from the motion vectors of other control point positions can be stored. This can be an example of a configuration considered in terms of memory management, and various modified examples are possible.

[0430] Whether or not to support statistical candidates can be explicitly determined in units such as sequences, pictures, slices, tiles, and blocks, or can be implicitly determined according to the encoding settings. Since the encoding settings can be defined by the various encoding elements described above, detailed descriptions are omitted.

[0431] Here, whether or not to support statistical candidates can be determined according to the motion information encoding mode. Alternatively, whether or not to support can be determined according to the motion model selection information. For example, statistical candidates can be supported among merge_inter, merge_ibc, merge_affine, comp_inter, comp_ibc, and comp_affine.

[0432] If statistical candidates are supported for both the moving motion model and the non-moving motion model, memory for a plurality of statistical candidates can be supported.

[0433] Next, a method for constructing a motion information prediction candidate group according to the motion information encoding mode will be described.

[0434] The set of motion information prediction candidates for the competition mode (hereinafter referred to as the competition mode candidate set) can include k candidates, and k can be an integer of 2, 3, 4, or more. The competition mode candidate set can include at least one of spatial candidates or temporal candidates.

[0435] Spatial candidates can be derived from at least one of the blocks adjacent to the reference block in the left, upper, upper-left, upper-right, lower-left directions, etc. centered on the reference block. Or, at least one candidate can be derived from the blocks adjacent to the left direction (left and lower-left blocks) and the blocks adjacent to the upper direction (upper-left, upper, and upper-right blocks). This setting will be assumed and described later.

[0436] There can be two or more priorities for constructing the candidate set. In the area adjacent to the left direction, the priority order of lower-left - left can be set, and in the area adjacent to the upper direction, the priority order of upper-right - upper - upper-left can be set.

[0437] The above example can be a configuration in which spatial candidates are derived only from the same blocks as the reference picture of the target block, but spatial candidates can also be derived through scaling processing (hereinafter indicated by *) based on the reference picture of the target block. In this case, in the area adjacent to the left direction, the priority order of left - lower-left - left* - lower-left* or left - lower-left - lower-left* - left* can be set, and in the area adjacent to the upper direction, the priority order of upper-right - upper - upper-left - upper-right* - upper* - upper-left* or upper-right - upper - upper-left - upper-left* - upper* - upper-right* can be set.

[0438] Temporal candidates can be derived from at least one of the blocks adjacent to the block corresponding to the reference block in the center and the left, right, upper, lower, upper-left, upper-right, lower-left, lower-right directions, etc. centered on the block corresponding to the reference block. There can be priorities for constructing the candidate set, and priorities such as center - lower-left - right - lower, lower-left - center - upper-left can be set.

[0439] When the total of the maximum allowable number of spatial candidates and the maximum allowable number of temporal candidates is less than the number of the candidate groups in the competition mode, regardless of the composition of the candidate groups of the spatial candidates, the temporal candidates can be included in the candidate groups.

[0440] Based on the priority order, the availability of each candidate block, and the maximum allowable number of temporal candidates (q, an integer between 1 and the number of the candidate groups in the competition mode), all or some of the candidates can be included in the candidate groups.

[0441] Here, when the maximum allowable number of spatial candidates is set to be the same as the number of the candidate groups in the merge mode, the temporal candidates cannot be included in the candidate groups. When the maximum allowable number is not satisfied only by the spatial candidates, it can be set that the temporal candidates can be included in the candidate groups. This example assumes the latter case.

[0442] Here, the motion vector of the temporal candidate can be obtained based on the motion vector of the candidate block and the distance interval between the current image and the reference image of the target block. The reference image of the temporal candidate can be obtained based on the distance interval between the current image and the reference image of the target block, or based on the reference image of the temporal candidate, or based on a predefined reference image (for example, the reference picture index is 0).

[0443] If the competition mode candidate groups are not satisfied by spatial candidates, temporal candidates, etc., the composition of the candidate groups can be completed through the default f candidates including the zero vector.

[0444] The competition mode focuses on the description of comp_inter. For comp_ibc or comp_affine, the candidate groups can be configured for the same or different candidates.

[0445] For example, in the case of comp_ibc, a candidate group can be constructed based on a predetermined candidate selected from spatial candidates, statistical candidates, combination candidates, default candidates, and the like. At this time, the candidate group can be constructed by prioritizing spatial candidates, and then in the order of statistical candidates - combination candidates - default candidates, etc., but it is not limited to this, and various orders are possible.

[0446] Alternatively, in the case of comp_affine, a candidate group can be constructed based on a predetermined candidate selected from spatial candidates, temporal candidates, statistical candidates, combination candidates, default candidates, and the like. Specifically, a motion vector set candidate of the candidate (for example, one block) can be configured as the candidate group, or a candidate in which motion vectors of candidates (for example, two or more blocks) based on the positions of control points are combined can be configured as the candidate group <2>. It is possible to include the candidate related to <1> in the candidate group first and then the candidate related to <2> in this order, but it is not limited to this.

[0447] Since the detailed description regarding the configuration of the various competition mode candidate groups described above can be induced in comp_inter, the detailed description is omitted.

[0448] In the process of constructing the competition mode candidate group, if there is duplicate motion information among the candidates included first, it is possible to include the candidate of the next priority in the candidate group while maintaining the candidate included first.

[0449] Here, before constructing the candidate group, the motion vectors existing in the candidate group can be scaled based on the distance between the reference picture and the current picture of the target block. For example, when the distance between the reference picture and the current picture of the target block is the same as the distance between the reference picture of the candidate block and the picture to which the candidate block belongs, the motion vector can be included in the candidate group, and if they are not the same, the motion vector scaled according to the distance between the reference picture and the current picture of the target block can be included in the candidate group.

[0450] At this time, the duplication means that the motion information is the same, which can be defined by the motion information coding mode. In the case of the competition mode, it is possible to determine whether there is duplication based on the motion vector, the reference picture, the reference direction, etc. For example, if at least one component of the motion vector is different, it can be determined that there is no duplication. The duplication confirmation process may generally be performed when a new candidate is included in the candidate group, but it is also possible to deform in the case of omission.

[0451] The motion information prediction candidate group for the merge mode (hereinafter, the merge mode candidate group) can include k candidates, and k can be an integer of 2, 3, 4, 5, 6 or more. The merge mode candidate group can include at least one of the spatial candidate or the temporal candidate.

[0452] The spatial candidate can be derived from at least one of the blocks adjacent to the reference block in the left, up, upper left, upper right, lower left directions, etc. There can be a priority order for candidate group configuration, and priority orders such as left - up - lower left - upper right - upper left, left - up - upper right - lower left - upper left, up - left - lower left - upper left - upper right can be set.

[0453] Based on the priority order, the usability of each candidate block (judged based on, for example, the coding mode, block position, etc.), and the maximum allowable number of spatial candidates (p, an integer between 1 and the number of the merge mode candidate group), all or part of the candidates can be included in the candidate group. Based on the maximum allowable number and the usability, it may not be included in the candidate group in the order of tl - tr - bl - t3 - l3. If the maximum allowable number is 4 and the usability of all candidate blocks is true, the motion information of tl is not included in the candidate group. If the usability of some candidate blocks is false, the motion information of tl can be included in the candidate group.

[0454] The temporal candidates can be derived from at least one of the adjacent blocks in the central, left, right, top, bottom, upper left, upper right, lower left, lower right directions, etc., centered on the block corresponding to the reference block. There can be a priority order for the composition of the candidate group, and priority orders such as central - lower left - right - bottom, lower left - central - upper left, etc. can be set.

[0455] Based on the above priority order, the usability of each candidate block, and the maximum allowable number of temporal candidates (q, an integer between 1 and the number of candidate groups in the merge mode), all or part of the candidates can be included in the candidate group.

[0456] Here, the motion vector of the temporal candidate can be obtained based on the motion vector of the candidate block, and the reference image of the temporal candidate can be obtained based on the reference image of the candidate block or can be obtained based on a predefined reference image (for example, the reference picture index is 0).

[0457] The priority order included in the merge mode candidate group can be set as spatial candidate - temporal candidate or vice versa, and it is also possible to support a priority order in which the spatial candidate and the temporal candidate are mixed. In this example, it is assumed that it is spatial candidate - temporal candidate.

[0458] In addition, statistical candidates or combined candidates can be further included in the merge mode candidate group. The statistical candidates and the combined candidates can be configured behind the spatial candidates and the temporal candidates, but are not limited to this, and various priority orders can be set.

[0459] The statistical candidate can manage up to n pieces of motion information, and z pieces of this motion information can be included in the merge mode candidate group as statistical candidates. z can be variable according to the candidate composition already included in the merge mode candidate group, can be 0, 1, 2, or an integer greater than or equal to that, and can be less than or equal to n.

[0460] Combination candidates can be derived by combining n candidates already included in the merge mode candidate group, where n can be an integer of 2, 3, 4, or more. The number (n) of the combination candidates can be information that explicitly occurs in units such as sequences, pictures, sub-pictures, slices, tiles, bricks, blocks, etc. Or, it can be implicitly determined according to the encoding settings. At this time, the encoding settings can be defined based on one or more factors such as the size, shape, position, image type, color components, etc. of the reference block.

[0461] Also, the number of the combination candidates can be determined based on the number of candidates that do not satisfy the merge mode candidate group. At this time, the number of candidates that do not satisfy the merge mode candidate group can be the difference value between the number of the merge mode candidate group and the number of candidates that have already been satisfied. That is, if the composition of the merge mode candidate group has already been completed, the combination candidates may not be added. If the composition of the merge mode candidate group is not completed, the combination candidates can be added, but if the number of candidates that satisfy the merge mode candidate group is one or less, the combination candidates are not added.

[0462] If the merge mode candidate group is not satisfied by spatial candidates, temporal candidates, statistical candidates, combination candidates, etc., the composition of the candidate group can be completed through default candidates including zero vectors.

[0463] The merge mode is centered on the description of merge_inter. In the case of merge_ibc or merge_affine, candidate groups can be formed for different or similar candidates.

[0464] For example, in the case of merge_ibc, candidate groups can be formed based on predetermined candidates selected from spatial candidates, statistical candidates, combination candidates, default candidates, etc. At this time, the candidate groups can be formed with spatial candidates prioritized, and then in the order of statistical candidates - combination candidates - default candidates, etc., but it is not limited to this, and various orders are possible.

[0465] Or, in the case of merge_affine, a candidate group can be constructed based on a predetermined candidate selected from spatial candidates, temporal candidates, statistical candidates, combination candidates, default candidates, etc. Specifically, a set of motion vector candidates of the candidate (for example, one block) can be configured as the candidate group <1>, or a candidate in which motion vectors of candidates based on the positions of control points (for example, two or more blocks) are combined can be configured as the candidate group <2>. It is possible, but not limited to, the order in which the candidate related to <1> is first included in the candidate group and then the candidate related to <2> is included.

[0466] Since the details regarding the construction of the various merge mode candidate groups described above can be derived in merge_inter, detailed description is omitted.

[0467] In the process of constructing the merge mode candidate group, if there is duplicate motion information among the previously included candidates, the candidates with the next priority can be included in the candidate group while maintaining the previously included candidates.

[0468] At this time, the duplication means that the motion information is the same, which can be defined by the motion information coding mode. In the case of the merge mode, it is possible to determine whether there is duplication based on the motion vector, reference picture, reference direction, etc. For example, if at least one component of the motion vector is different, it can be determined that there is no duplication. The duplication confirmation process is generally performed when a new candidate is included in the candidate group, but it is also possible to be deformed to the case of omission.

[0469] FIG. 14 is a flowchart regarding motion information coding according to an embodiment of the present invention. Specifically, it can be coding regarding the motion information of a target block in a competition mode.

[0470] A motion vector prediction candidate list for the target block can be generated (S1400). The above-described competition mode candidate group can mean a motion vector candidate list, and detailed description thereof will be omitted.

[0471] The motion vector difference value of the target block can be restored (S1410). The difference values regarding the x and y components of the motion vector can be restored individually, and the difference value of each component can have a value of 0 or more.

[0472] A prediction candidate index for the target mode can be selected from the motion vector prediction candidate list (S1420). Based on the motion vector obtained according to the candidate index from the candidate list, the motion vector prediction value of the target block can be derived. If one motion vector prediction value that has already been set can be derived, the selection process of the prediction candidate index and the index information can be omitted.

[0473] Differential motion vector accuracy information can be derived (S1430). The accuracy information commonly applied to the x and y components of the motion vector can be derived, or the accuracy information applied to each component can be derived. If the motion vector difference value is 0, the accuracy information can be omitted and this process can also be omitted.

[0474] An adjustment offset regarding the motion vector prediction value can be derived (S1440). The offset can be a value added or subtracted to / from the x or y component of the motion vector prediction value. The offset can assist only one of the x and y components, or can assist both the x and y components.

[0475] Assuming the motion vector prediction value is (pmv_x, pmv_y) and the adjustment offset is offset_x, offset_y, the motion vector prediction value can be adjusted (or obtained) to (pmv_x + offset_x, pmv_y + offset_y).

[0476] Here, the absolute values of offset_x and offset_y are each an integer of 0, 1, 2, or more, and can have values (such as 1, -1, +2, -2, etc.) where the sign information is considered together. Also, offset_x and offset_y can be determined based on a predetermined accuracy. The predetermined accuracy can be determined from among 1 / 16, 1 / 8, 1 / 4, 1 / 2, 1 pixel unit, etc., and can also be determined based on interpolation accuracy, motion vector accuracy, etc.

[0477] For example, when the interpolation accuracy is 1 / 4 pixel unit, it can be induced to 0, 1 / 4, -1 / 4, 2 / 4, -2 / 4, 3 / 4, -3 / 4, etc. by combining with the absolute value and sign information.

[0478] Here, a and b offset_x and offset_y can each support, and a and b can be integers of 0, 1, 2, or more. a and b can have fixed values or can have variable values. Also, a and b can have the same or different values.

[0479] Whether or not to support the adjustment offset for the motion vector prediction value can be explicitly supported in units such as sequence, picture, sub-picture, slice, tile, block, etc., or can be implicitly determined according to the coding settings. Also, the settings of the adjustment offset (such as value range, number, etc.) can be determined according to the coding settings.

[0480] The coding settings can be determined considering at least one of the coding elements such as image type, color component, state information of the target block, motion model selection information (such as whether it is a moving motion model), reference picture (such as whether it is the current picture), differential motion vector accuracy selection information (such as whether it is a predetermined unit among 1 / 4, 1 / 2, 1, 2, 4 units), etc.

[0481] For example, it can be determined whether to assist in adjusting the offset and the setting of the offset according to the size of the block. At this time, the size of the block can have the assistance and setting range determined by the size of the first threshold (minimum value) or the size of the second threshold (maximum value), and the size of each threshold can be expressed as W, H, W×H, or W*H in terms of the width (W) and height (H) of the block. In the case of the size of the first threshold, W and H can be 4, 8, 16, or an integer greater than or equal to 16, and W*H can be 16, 32, 64, or an integer greater than or equal to 64. In the case of the size of the second threshold, W and H can be 16, 32, 64, or an integer greater than or equal to 64, and W*H can be 64, 128, 256, or an integer greater than or equal to 256. The above range can be determined by either the size of the first threshold or the size of the second threshold, or by both.

[0482] At this time, the size of the threshold can be fixed or can be adaptive based on an image (for example, image type, etc.). At this time, the size of the first threshold can be set based on the size of the minimum coding block, the minimum prediction block, the minimum transform block, etc., and the size of the second threshold can be set based on the size of the maximum coding block, the maximum prediction block, the maximum transform block, etc.

[0483] Also, the adjustment offset can be applied to all candidates included in the motion information prediction candidate group or can be applied only to some candidates. In the example described later, it is assumed that the adjustment offset is applied to all candidates included in the candidate group, but the candidates to which the adjustment offset is applied can be selected between 0, 1, 2, and the maximum number of candidates in the candidate group.

[0484] If the assistance of the adjustment offset is not provided, this process and the adjustment offset information can be omitted.

[0485] The motion vector of the target block can be restored by adding the predicted motion vector value and the motion vector difference value (S1450). At this time, the process of unifying the predicted motion vector value or the motion vector difference value to the motion vector accuracy can precede, and the above-described motion vector scaling process can precede or can be performed in this process.

[0486] The above configurations and orders are merely examples and are not limited thereto, and various modifications are possible.

[0487] FIG. 15 is an exemplary diagram of the predicted motion vector candidates and the motion vector of the target block according to an embodiment of the present invention.

[0488] For convenience of explanation, two predicted motion vectors are supported, and it is assumed that a comparison of one component (either the x or y component) is made. Also, it is assumed that the interpolation accuracy or the motion vector accuracy is in units of 1 / 4 pixel. Also, it is assumed that the differential motion vector accuracy is supported in units of 1 / 4, 1, and 4 pixels (for example, when it is 1 / 4, it is binarized to <0>, when it is 1, it is binarized to <10>, and when it is 4, it is binarized to <11>). Also, it is assumed that the motion vector difference value is processed by single-term binarization with the sign information omitted (for example, 0: <0>, 1: <10>, 2: <110>, etc.).

[0489] Referring to FIG. 15, the actual motion vector (X) has a value of 2, candidate 1 (A) has a value of 0, and candidate 2 (B) has a value of 1 / 4.

[0490] (When differential motion vector accuracy is not supported) Since the distance (da) between A and X is 8 / 4 (9 bits) and the distance (db) between B and X is 7 / 4 (8 bits), from the viewpoint of the amount of bits generated, the prediction candidate can be selected as B.

[0491] (When differential motion vector accuracy is supported) Since da is 8 / 4, pixel unit accuracy (2 bits) and differential value information (3 bits) for 2 / 1 are generated, and a total of 5 bits can be generated. On the other hand, since db is 7 / 4, 1 / 4 pixel unit accuracy (1 bit) and differential value information (8 bits) for 7 / 4 are generated, and a total of 9 bits can be generated. From the perspective of bit amount generation, the prediction candidate can be selected as A.

[0492] When the differential motion vector accuracy is not supported as in the above example, it is advantageous that a candidate with a short distance interval from the motion vector of the target block is selected as the prediction candidate. When the differential motion vector accuracy is supported, it may be important that the prediction candidate is selected not only based on the distance interval from the motion vector of the target block but also based on the amount of information generated based on the accuracy information.

[0493] FIG. 16 is an exemplary diagram of a motion vector prediction candidate and the motion vector of a target block according to an embodiment of the present invention. The following assumes a case where the differential motion vector accuracy is supported.

[0494] Since da is 33 / 4, 1 / 4 pixel unit accuracy (1 bit) and differential value information (34 bits) for 33 / 4 are generated, and a total of 35 bits can be generated. On the other hand, since db is 21 / 4, 1 / 4 pixel unit accuracy (1 bit) and differential value information (22 bits) for 21 / 4 are generated, and a total of 23 bits can be generated. From the perspective of bit amount generation, the prediction candidate can be selected as B.

[0495] FIG. 17 is an exemplary diagram of a motion vector prediction candidate and the motion vector of a target block according to an embodiment of the present invention. The following assumes a case where the differential motion vector accuracy is supported and the adjustment offset information is supported.

[0496] In this example, assume that the adjustment offset has candidates of 0 and +1, and a flag (1 bit) and offset selection information (1 bit) for whether the adjustment offset is applied are generated.

[0497] Referring to FIG. 17, A1 and B1 can be motion vector prediction values obtained based on prediction candidate indexes, and A2 and B2 can be new motion vector prediction values obtained by correcting A1 and B1 with adjustment offsets. Assume that the distances between A1, A2, B1, B2 and X are da1, da2, db1, and db2, respectively.

[0498] (1) Since da1 is 33 / 4, 1 / 4 pixel unit accuracy (1 bit), offset application flag (1 bit), offset selection information (1 bit), and difference value information (34 bits) for 33 / 4 can be generated, and a total of 37 bits can be generated.

[0499] (2) Since da2 is 32 / 4, 4 pixel unit accuracy (2 bits), offset application flag (1 bit), offset selection information (1 bit), and difference value information (3 bits) for 2 / 1 can be generated, and a total of 7 bits can be generated.

[0500] (3) Since db1 is 21 / 4, 1 / 4 pixel unit accuracy (1 bit), offset application flag (1 bit), offset selection information (1 bit), and difference value information (22 bits) for 21 / 4 can be generated, and a total of 25 bits can be generated.

[0501] (4) Since db2 is 20 / 4, 1 pixel unit accuracy (2 bits), offset application flag (1 bit), offset selection information (1 bit), and difference value information (6 bits) for 5 / 1 can be generated, and a total of 10 bits can be generated.

[0502] From the perspective of bit amount generation, the prediction candidate can be selected as A, and the offset selection information can be selected as the 1st index (+1 in this example).

[0503] Considering the above example, when deriving a motion vector difference value based on a conventional motion vector prediction value, there are cases where many bits are generated due to the difference in a small amount of vector values. The above problem can be solved by adjusting the motion vector prediction value.

[0504] A predetermined flag can assist in applying an adjustment offset to the motion vector prediction value. The predetermined flag can be composed of an offset application flag, offset selection information, and the like.

[0505] (When only the offset application flag is supported)

[0506] When the offset application flag is 0, no offset is applied to the motion vector prediction value. When the offset application flag is 1, a preset offset can be added to or subtracted from the predicted motion vector value.

[0507] (When only the offset selection information is supported)

[0508] An offset set based on the offset selection information can be added to or subtracted from the predicted motion vector value.

[0509] (When the offset application flag and the offset selection information are supported)

[0510] When the offset application flag is 0, no offset is applied to the motion vector prediction value. When the offset application flag is 1, an offset set based on the offset selection information can be added to or subtracted from the predicted motion vector value. Assume this setting in the example described later.

[0511] On the other hand, the offset application flag and the offset selection information can be used as information that may be unnecessarily generated depending on the case. That is, even if no offset is applied, if it already has a zero value or the amount of information decreases based on the maximum accuracy information, the offset-related information can be rather inefficient. Therefore, it is necessary to support a setting where the offset-related information is not always explicitly generated but is implicitly generated according to a predetermined condition.

[0512] Next, assume that the motion-related encoding order proceeds to (restoring the motion vector difference value → obtaining the differential motion vector accuracy). In this example, assume that the differential motion vector accuracy is supported when the motion vector difference value is not zero in either the x or y component.

[0513] When the motion vector difference value is not zero, one of {1 / 4, 1 / 2, 1, 4} pixel units can be selected for the differential motion vector accuracy.

[0514] When the selection information belongs to a predetermined category, information regarding the adjustment offset can be implicitly omitted, and when the selection information does not belong to the predetermined category, information regarding the adjustment offset can be explicitly generated.

[0515] Here, the category can include one of the differential motion vector accuracy candidates, and various category configurations such as {1 / 4}, {1 / 4, 1 / 2}, {1 / 4, 1 / 2, 1} are possible. Here, the minimum accuracy can be included in the category.

[0516] For example, when the differential motion vector accuracy is 1 / 4 pixel unit (e.g., the minimum accuracy), the offset application flag is implicitly set to 0 (i.e., not applied), and when the differential motion vector accuracy is not 1 / 4 pixel unit, the offset application flag is explicitly generated. When the offset application flag is 1 (i.e., offset applied), offset selection information can be generated.

[0517] Or, when the differential motion vector accuracy is 4 pixel units (e.g., the maximum accuracy), the offset application flag is explicitly generated, and when the offset application flag is 1, offset selection information can be generated. Or, when the differential motion vector accuracy is not 4 pixel units, the offset information can be implicitly set to 0.

[0518] In the above example, it is assumed that when the differential motion vector indicates the minimum accuracy, the offset-related information is implicitly omitted, and when the differential motion vector indicates the maximum accuracy, the offset-related information is explicitly generated, but it is not limited to this.

[0519] FIG. 18 is an exemplary diagram regarding the arrangement of a plurality of motion vector prediction values according to an embodiment of the present invention.

[0520] When constructing the motion information prediction candidate group by the above-described example, the part regarding the duplication check was explained. Here, duplication means that the motion information is the same, and it was previously described that when at least one component of the motion vector is different, it can be determined that there is no duplication.

[0521] A plurality of candidates regarding the motion vector prediction value can be non-overlapping with each other through the duplication check process. However, when the component elements of the plurality of candidates are very similar (that is, the x or y component of each candidate exists within a predetermined range. The width or height of the predetermined range is 1, 2, or an integer greater than or equal to 1. Or the range can also be set based on the offset information), it can occur that the motion vector prediction value and the motion vector prediction value corrected by the offset overlap. Various settings are possible for this purpose.

[0522] As an example (C1), in the step of constructing the motion information prediction candidate group, a new candidate can be included in the candidate group when it does not overlap with the candidates already included. That is, it can be a configuration that can be included in the candidate group if it is the same as the above-described existing explanation and does not overlap only by comparing the candidates themselves.

[0523] As an example (C2), if the number of overlaps between a new candidate and the candidates obtained by adding an offset based on it (group_A) and the candidates already included and the candidates obtained by adding an offset based on them (group_B) is less than a predetermined number, the new candidate can be included in the candidate group. The predetermined number can be 0, 1, or an integer greater than or equal to 1. If the predetermined number is 0, even one overlap would mean that the new candidate may not be included in the candidate group. Or, as (C3), a predetermined offset (which is a different concept from the adjustment offset) can be added to the new candidate and then it can be included in the candidate group, and this offset can be a value that enables group_A to be formed without overlapping with group_B.

[0524] Referring to FIG. 18, a plurality of motion vectors belonging to categories A, B, C, and D (in this example, AX, BX, CX, DX, where X is 1 or 2. For example, A1 is a candidate that has been included in the candidate group earlier than A2) may satisfy the non - overlapping condition (when the candidate motion vectors are different even in one component).

[0525] In this example, assuming that - 1, 0, and 1 are supported for the x and y components of the offset respectively, the motion vector prediction value with the corrected offset can be represented by * in the drawing. Also, the dashed line (rectangle) means the range (such as group_A, group_B, etc.) that can be obtained by adding an offset around a predetermined motion vector prediction value.

[0526] In the case of category A, it can correspond to the situation where A1 and A2 do not overlap and group_A1 and group_A2 do not overlap.

[0527] In the case of categories B, C, and D, it can correspond to the situation where B1 / C1 / D1 and B2 / C2 / D2 do not overlap and group_B1 / C1 / D1 and group_B2 / C2 / D2 partially overlap.

[0528] Here, in the case of category B, it can be an example of forming a candidate group without taking special measures (C1).

[0529] Here, in the case of category C, C2 can be determined to have duplication in the duplication check step, and C3, which is next in priority, can be included in the candidate group (C2).

[0530] Here, in the case of category D, D2 is determined to have duplication in the duplication check step, and D2 can be corrected so that group_D2 does not overlap with group_D1 (that is, D3. D3 is not a motion vector existing in the priority order of candidate group formation) (C3).

[0531] Can it be applied to the setting in which a prediction mode candidate group is formed from the various categories, or various methods other than those mentioned can be applied.

[0532] FIG. 19 is a flowchart relating to encoding of motion information in a merge mode according to an embodiment of the present invention. Specifically, it can be encoding related to the motion information of a target block in the merge mode.

[0533] A motion information prediction candidate list for the target block can be generated (S1900). The merge mode candidate group described above can mean a motion information prediction candidate list, and detailed description thereof will be omitted.

[0534] A prediction candidate index for the target mode can be selected from the motion information prediction candidate list (S1910). Based on the motion information obtained according to the candidate index from the candidate list, a predicted value of the motion vector of the target block can be derived. If a predicted value of one piece of motion information that has already been set can be derived, the prediction candidate index selection process and index information can be omitted.

[0535] An adjustment offset for the motion vector prediction value can be derived (S1920). The offset can be a value added to or subtracted from the x or y component of the motion vector prediction value. The offset can assist only one of the x and y components, or can assist both the x and y components.

[0536] Since the adjustment offset in this process can be a concept identical or similar to the adjustment offset described above, a detailed description is omitted, and the parts regarding the differences will be described later.

[0537] Assuming that the motion vector prediction value is (pmv_x, pmv_y) and the adjustment offsets are offset_x and offset_y, the motion vector prediction value can be adjusted (or obtained) to (pmv_x + offset_x, pmv_y + offset_y).

[0538] Here, the absolute values of offset_x and offset_y can be integers such as 0, 1, 2, 4, 8, 16, 32, 64, 128, etc., and can have values considering the sign information together. Also, offset_x and offset_y can be determined based on a predetermined accuracy. The predetermined accuracy can be determined from among 1 / 16, 1 / 8, 1 / 4, 1 / 2, 1 pixel unit, etc.

[0539] For example, when the motion vector accuracy is 1 / 4 pixel unit, it can be derived to 0, 1 / 4, -1 / 4, 1 / 2, -1 / 2, 1, -1, 2, -2, etc. by combining the absolute value and the sign information.

[0540] Here, a and b offset_x and offset_y can assist respectively, and a and b can be integers such as 0, 1, 2, 4, 8, 16, 32, etc. a and b can have fixed values, or can have variable values. Also, a and b can have the same or different values.

[0541] Whether or not to assist in the adjustment offset for the motion vector prediction value can be explicitly assisted in units such as sequence, picture, sub-picture, slice, tile, block, etc., or can be implicitly determined according to the encoding settings. Also, the setting of the adjustment offset (for example, value range, number, etc.) can be determined according to the encoding settings.

[0542] The encoding settings can be determined in consideration of at least one of the encoding elements such as image type, color component, state information of the target block, motion model selection information (for example, whether it is a moving motion model), reference picture (for example, whether it is the current picture), etc.

[0543] For example, whether to assist in the adjustment offset and the setting of the offset can be determined according to the size of the block. At this time, the size of the block can have its support and setting range determined by the size of the first threshold (minimum value) or the size of the second threshold (maximum value), and the size of each threshold can be expressed as W, H, W×H, W*H in terms of the width (W) and height (H) of the block. In the case of the size of the first threshold, W and H can be 4, 8, 16 or an integer greater than or equal to that, and W*H can be 16, 32, 64 or an integer greater than or equal to that. In the case of the size of the second threshold, W and H can be 16, 32, 64 or an integer greater than or equal to that, and W*H can be 64, 128, 256 or an integer greater than or equal to that. The said range can be determined by either one of the size of the first threshold or the size of the second threshold, or by both of them.

[0544] At this time, the size of the threshold can be fixed or adaptable according to the image (for example, image type, etc.). At this time, the size of the first threshold can be set based on the size of the minimum coding block, minimum prediction block, minimum transform block, etc., and the size of the second threshold can also be set based on the size of the maximum coding block, maximum prediction block, maximum transform block, etc.

[0545] Also, the adjustment offset can be applied to all candidates included in the motion information prediction candidate group, or only to some candidates. In the example described later, it is assumed that the adjustment offset is applied to all candidates included in the candidate group, but the candidates to which the adjustment offset is applied can be selected between 0, 1, 2, and the maximum number of candidates in the candidate group.

[0546] A predetermined flag can assist in applying the adjustment offset to the predicted motion vector value. The predetermined flag can be configured via an offset application flag, offset absolute value information, offset sign information, and the like.

[0547] If the assistance of the adjustment offset is not provided, this process and the adjustment offset information can be omitted.

[0548] The motion vector of the target mode can be restored via the predicted motion vector value (S1930). Motion information other than the motion vector (for example, reference picture, reference direction, etc.) can be obtained based on the prediction candidate index.

[0549] The above configuration and order are only some examples and are not limited thereto, and various changes are possible. Since the background explanation of assisting the adjustment offset in the merge mode has been described above through various examples of the competition mode, a detailed explanation will be omitted.

[0550] The method according to the present invention can be realized in the form of program instructions executable via various computer means and can be recorded on a computer-readable medium. The computer-readable medium can include program instructions, data files, data structures, etc. alone or in combination. The program instructions recorded on the computer-readable medium are those specially designed and configured for the present invention or those known and usable by those skilled in the computer software art.

[0551] Examples of computer-readable media can include hardware devices specially configured to store and execute program instructions, such as ROM, RAM, flash memory, etc. Examples of program instructions can include not only machine language code created by a compiler, but also high-level language code that can be executed by a computer using an interpreter, etc. The hardware devices described above can be configured to operate as at least one software module for performing the operations of the present invention, and vice versa.

[0552] In addition, all or part of the configuration and functions of the above-described method or apparatus may be realized in combination or separately.

[0553] In the above, the preferred embodiments of the present invention have been described with reference thereto. However, those skilled in the art in the relevant technical field will understand that the present invention can be variously modified and changed without departing from the spirit and scope of the present invention described in the following claims.

Industrial Applicability

[0554] The present invention can be used for encoding / decoding images.< / qt>

Claims

1. An image decoding method for performing inter-picture prediction, comprising: constructing a merge candidate list for the target block, including a statistical merge candidate from a statistical merge candidate group and a temporal merge candidate from an array block of the target block; determining a motion vector from the merge candidate selected within the merge candidate list based on a candidate index; determining a predicted block of the target block based on the determined motion vector, wherein the statistical merge candidate group includes one or more statistical merge candidates, and the number of statistical merge candidates within the statistical merge candidate group cannot be greater than a predetermined maximum number of the statistical merge candidates; the statistical merge candidate indicates inter-picture prediction information of a block decoded before the target block, and the merge candidate list includes one or more statistical merge candidates within the statistical merge candidate group that are not immediately adjacent to the target block; the statistical merge candidate group is updated after completion of decoding of the target block by the following method: if the inter-picture prediction information of the target block does not overlap with any statistical merge candidate within the statistical merge candidate group, adding the inter-picture prediction information of the target block as the first statistical merge candidate of the statistical merge candidate group, deleting the last statistical merge candidate from the statistical merge candidate group, and updating the statistical merge candidate group by correcting the order of the remaining statistical merge candidates; if the inter-picture prediction information of the target block overlaps with any statistical merge candidate of the statistical merge candidate group, adding the inter-picture prediction information of the target block as the first statistical merge candidate of the statistical merge candidate group, deleting the overlapping statistical merge candidate from the statistical merge candidate group, and updating the statistical merge candidate group by correcting the order of the remaining statistical merge candidates; the statistical merge candidate group is initialized when the target block is the first block of a predetermined data unit; in the merge candidate list, the statistical merge candidate is included in an order after the temporal merge candidate; An image decoding method.

2. An image encoding method for performing inter-picture prediction, comprising: constructing a merge candidate list for the target block, including a statistical merge candidate from a statistical merge candidate group and a temporal merge candidate from an array block of the target block; Determining a motion vector from a merge candidate selected within the merge candidate list, and encoding the motion vector using a candidate index; Determining a predicted block of the target block based on the determined motion vector; and The statistical merge candidate group includes one or more statistical merge candidates, and the number of the statistical merge candidates within the statistical merge candidate group cannot be greater than a predetermined maximum number of the statistical merge candidates; The statistical merge candidate indicates inter-picture prediction information of a block encoded before the target block, and the merge candidate list includes one or more statistical merge candidates within the statistical merge candidate group that are not immediately adjacent to the target block; The statistical merge candidate group is updated in the following manner after completion of encoding of the target block; When the inter-picture prediction information of the target block does not overlap with any of the statistical merge candidates within the statistical merge candidate group, adding the inter-picture prediction information of the target block as a first statistical merge candidate of the statistical merge candidate group, deleting the last statistical merge candidate from the statistical merge candidate group, and updating the statistical merge candidate group by correcting the order of the remaining statistical merge candidates; When the inter-picture prediction information of the target block overlaps with any of the statistical merge candidates of the statistical merge candidate group, adding the inter-picture prediction information of the target block as the first statistical merge candidate of the statistical merge candidate group, deleting the overlapping statistical merge candidate from the statistical merge candidate group, and updating the statistical merge candidate group by correcting the order of the remaining statistical merge candidates; The statistical merge candidate group is initialized when the target block is a head block of a predetermined data unit; In the merge candidate list, the statistical merge candidates are included in an order after the temporal merge candidates; Method.

3. A method for transmitting a bitstream generated by an image encoding method, the method comprising: The image encoding method includes: Constructing a merge candidate list of the target block including a statistical merge candidate from a statistical merge candidate group and a temporal merge candidate from an array block of the target block; Determining a motion vector from a merge candidate selected within the merge candidate list, and encoding the motion vector using a candidate index; determining a predicted block of the target block based on the determined motion vector; the statistical merge candidate group includes one or more statistical merge candidates, and the number of the statistical merge candidates in the statistical merge candidate group cannot be greater than a predetermined maximum number of the statistical merge candidates; the statistical merge candidate indicates inter-picture prediction information of a block encoded before the target block, and the merge candidate list includes one or more statistical merge candidates in the statistical merge candidate group that are not immediately adjacent to the target block; the statistical merge candidate group is updated in the following manner after the encoding of the target block is completed; if the inter-picture prediction information of the target block does not overlap with any of the statistical merge candidates in the statistical merge candidate group, adding the inter-picture prediction information of the target block as the first statistical merge candidate of the statistical merge candidate group, deleting the last statistical merge candidate from the statistical merge candidate group, and updating the statistical merge candidate group by correcting the order of the remaining statistical merge candidates; if the inter-picture prediction information of the target block overlaps with any of the statistical merge candidates in the statistical merge candidate group, adding the inter-picture prediction information of the target block as the first statistical merge candidate of the statistical merge candidate group, deleting the overlapping statistical merge candidates from the statistical merge candidate group, and updating the statistical merge candidate group by correcting the order of the remaining statistical merge candidates; the statistical merge candidate group is initialized when the target block is the leading block of a predetermined data unit; in the merge candidate list, the statistical merge candidates are included in an order after the temporal merge candidates; Method.

Citation Information

Patent Citations

  • Offset temporal motion vector predictor (TMVP)

    US20180098085A1

  • Method and apparatus for encoding / decoding video signal

    WO2017164645A2

  • Offset vector identification of temporal motion vector predictor

    WO2018052986A1