Image encoding / decoding method and device
The image encoding/decoding device addresses performance issues in motion vector prediction by correcting motion vector offsets, leading to improved coding efficiency.
Patent Information
- Application Number
- JP2025114050
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2018-09-24
- Filing Date
- 2025-07-04
- Publication Date
- 2025-09-25
- Estimated Expiration
- 2039-09-24
AI Technical Summary
Existing image encoding/decoding methods lack performance improvements, particularly in motion vector prediction, which affects the efficiency of image processing systems.
An image encoding/decoding device that corrects motion vector prediction using an adjustment offset, constructing a motion information prediction candidate list, selecting a prediction candidate index, and deriving a predicted motion vector adjustment offset based on offset application flag and/or offset selection information.
Improves coding performance by efficiently obtaining a predicted motion vector, enhancing the efficiency of image processing systems.
Smart Images

Figure 2025138854000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an image encoding / decoding method and apparatus. [Background technology]
[0002] With the spread of the Internet and mobile terminals and the development of information and communication technology, the use of multimedia data is rapidly increasing. Therefore, in order to perform various services and operations through image prediction in various systems, the need for improving the performance and efficiency of image processing systems is greatly increasing, but the results of research and development that can respond to this trend are currently insufficient.
[0003] Thus, in the prior art image encoding / decoding methods and devices, there is a need for performance improvements in image processing, particularly image encoding or image decoding. Summary of the Invention [Problem to be solved by the invention]
[0004] The present invention has been made to solve the above problems, and its object is to provide an image encoding / decoding device that corrects a motion vector prediction value using an adjustment offset. [Means for solving the problem]
[0005] To achieve the above object, a method for decoding an image according to one embodiment of the present invention includes the steps of constructing a motion information prediction candidate list for a target block, selecting a prediction candidate index, deriving a predicted motion vector adjustment offset, and restoring motion information of the target block.
[0006] Here, the step of setting the motion information prediction candidate list may further include a step of including in the candidate group, if the candidates already included and the candidates obtained based on the offset information do not overlap with the new candidates and the candidates obtained based on the offset information.
[0007] Here, the step of deriving the predicted motion vector adjustment offset may further include deriving the offset based on an offset application flag and / or offset selection information. [Effects of the Invention]
[0008] When the inter-picture prediction according to the present invention is used, the coding performance can be improved by efficiently obtaining a predicted motion vector. [Brief explanation of the drawings]
[0009] [Figure 1] 1 is a conceptual diagram illustrating an image encoding and decoding system according to an embodiment of the present invention; [Figure 2] 1 is a block diagram showing the configuration of an image encoding device according to an embodiment of the present invention. [Figure 3] 1 is a block diagram showing the configuration of an image decoding device according to an embodiment of the present invention; [Figure 4] 10A to 10C are exemplary diagrams showing various division forms that can be obtained by the block division unit of the present invention; [Figure 5] 10A and 10B are diagrams illustrating various cases in which a prediction block is obtained through inter-frame prediction according to the present invention; [Figure 6] FIG. 10 is an exemplary diagram illustrating how to construct a reference picture list according to an embodiment of the present invention. [Figure 7] FIG. 2 is a conceptual diagram illustrating a non-moving motion model according to an embodiment of the present invention. [Figure 8] 1 is an example diagram illustrating sub-block-based motion estimation according to an embodiment of the present invention; [Figure 9] 10 is a flowchart illustrating encoding of motion information according to an embodiment of the present invention. [Figure 10] FIG. 2 is a layout diagram of a target block and its adjacent blocks according to an embodiment of the present invention. [Figure 11] FIG. 1 is an exemplary diagram of a statistical candidate according to an embodiment of the present invention. [Figure 12]FIG. 1 is a conceptual diagram of a statistical candidate based on a non-movement motion model according to an embodiment of the present invention. [Figure 13] 10 is a diagram illustrating an example of the configuration of motion information of each control point position stored as a statistical candidate according to an embodiment of the present invention; [Figure 14] 1 is a flowchart illustrating motion information encoding according to an embodiment of the present invention. [Figure 15] 1 is a diagram illustrating an example of motion vector predictor candidates and a motion vector of a current block according to an embodiment of the present invention; [Figure 16] 1 is a diagram illustrating an example of motion vector predictor candidates and a motion vector of a current block according to an embodiment of the present invention; [Figure 17] 1 is a diagram illustrating an example of motion vector predictor candidates and a motion vector of a current block according to an embodiment of the present invention; [Figure 18] 1 is an exemplary diagram illustrating an arrangement of a plurality of motion vector predictors according to an embodiment of the present invention; [Figure 19] 1 is a flowchart for encoding motion information in merge mode according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0010] The image encoding / decoding method and apparatus according to the present invention can construct a predicted motion candidate list for a target block, derive a predicted motion vector from the motion candidate list based on a predicted candidate index, restore predicted motion vector adjustment offset information, and restore a motion vector for the target block based on the predicted motion vector and the predicted motion vector adjustment offset information.
[0011] In the image encoding / decoding method and apparatus according to the present invention, the motion candidate list may include at least one of spatial candidates, temporal candidates, statistical candidates, or combined candidates.
[0012] In the image encoding / decoding method and apparatus according to the present invention, the predicted motion vector adjustment offset can be determined based on at least one of an offset application flag or offset selection information.
[0013] In the image encoding / decoding method and apparatus according to the present invention, the information on whether the predicted motion vector adjustment offset information is supported may be included in at least one of a sequence, a picture, a sub-picture, a slice, a tile, or a brick.
[0014] In the image encoding / decoding method and apparatus according to the present invention, if the target block is encoded in a merge mode, the motion vector of the target block is restored using a zero vector, and if the target block is encoded in a competitive mode, the motion vector of the target block can be restored using a motion vector differential value.
[0015] [Mode for carrying out the invention] Since the present invention can be modified in various ways and can have various embodiments, specific embodiments are illustrated in the drawings and described in detail in the detailed description, but it should be understood that the present invention is not limited to the specific embodiments, and includes all modifications, equivalents, and alternatives within the spirit and technical scope of the present invention.
[0016] The terms "first," "second," etc. may be used to describe various elements, but these elements should not be limited by these terms. These terms are used only to distinguish one element from another. For example, a first element can be termed a second element, and similarly, a second element can be termed a first element, without departing from the scope of the present invention. The term "and / or" includes a combination of two or more related listed items or any of two or more related listed items.
[0017] When a component is said to be "coupled" or "connected" to another component, it should be understood that the component may be directly coupled or connected to the other component, but there may be other components between them. Conversely, when a component is said to be "directly coupled" or "directly connected" to another component, it should be understood that there are no other components between them.
[0018] The terms used in the present invention are merely used to describe specific embodiments and do not limit the present invention. A singular expression includes a plural expression unless the context clearly indicates otherwise. In the present invention, the terms "comprise" or "have" and the like specify the presence of features, numbers, steps, operations, components, parts, or combinations thereof described in the specification, and should be understood not to preclude the presence or possibility of addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.
[0019] Unless otherwise defined, all terms used herein, including technical or scientific terms, are meant to be the same as commonly understood by a person of ordinary skill in the art to which the present invention pertains. Terms defined in commonly used dictionaries should be interpreted in accordance with the meaning they have in the context of the relevant art, and should not be interpreted in an ideal or overly formal sense unless explicitly defined in the present invention.
[0020] Typically, an image may be configured with one or more color spaces depending on its color format. Depending on the color format, it may be configured with one or more pictures of a fixed size or one or more pictures of different sizes. For example, in a YCbCr color configuration, color formats such as 4:4:4, 4:2:2, 4:2:0, and monochrome (composed of only Y) are supported. For example, in the case of YCbCr 4:2:0, it may be configured with one luminance component (Y in this example) and two chrominance components (Cb / Cr in this example). In this case, the horizontal / vertical ratio of the chrominance component to the luminance component may be 1:2. For example, in the case of 4:4:4, the horizontal and vertical ratios may be the same. When configured with one or more color spaces as in the above example, the picture may be divided into each color space.
[0021] Images can be classified into I, P, B, etc. depending on the image type (e.g., picture type, subpicture type, slice type, tile type, brick type, etc.), where I image type can mean an image that is coded by itself without using a reference picture, P image type can mean an image that is coded using a reference picture but allows only forward prediction, and B image type can mean an image that is coded using a reference picture and allows forward / backward prediction, but depending on the coding settings, some of the above types may be combined (P and B may be combined) or image types with other configurations may be supported.
[0022] Various encoding / decoding information generated in the present invention can be processed explicitly or implicitly. Here, explicit processing can be understood as generating encoding / decoding information in the form of a sequence, picture, subpicture, slice, tile, brick, block, subblock, etc. and recording it in a bitstream, and parsing related information in the decoder at the same level as the encoder to restore decoded information. Here, implicit processing can be understood as processing encoding / decoding information in the encoder and decoder using the same process or rule.
[0023] FIG. 1 is a conceptual diagram showing an image encoding and decoding system according to an embodiment of the present invention.
[0024] Referring to FIG. 1, the image encoding device 105 and the decoding device 100 may be user terminals such as a personal computer (PC), a notebook computer, a personal digital assistant (PDA), a portable multimedia player (PMP), a PlayStation Portable (PSP), a wireless communication terminal, a smartphone, or a TV, or may be server terminals such as an application server or a service server, and may include various devices including a communication device such as a communication modem for communicating with various devices or wired or wireless communication networks, memories (120, 125) for storing various programs and data for inter or intra prediction to encode or decode an image, or processors (110, 115) for executing programs, performing calculations, and control.
[0025] Furthermore, the image coded into a bitstream by the image coding device 105 can be transmitted to the image decoding device 100 in real time or non-real time via a wired or wireless communication network (Network) such as the Internet, a short-range wireless communication system, a wireless LAN network, a WiBro network, or a mobile communication network, or via various communication interfaces such as a cable or a Universal Serial Bus (USB), and can be restored and played back by being decoded by the image decoding device 100. Furthermore, the image coded into a bitstream by the image coding device 105 can be transmitted from the image coding device 105 to the image decoding device 100 via a computer-readable recording medium.
[0026] The image encoding device and the image decoding device may be separate devices, but may be implemented as a single image encoding / decoding device. In this case, some components of the image encoding device may be substantially the same technical elements as some components of the image decoding device, and may be implemented to include at least the same structure or perform at least the same functions.
[0027] Therefore, in the following detailed description of the technical elements and their operating principles, redundant descriptions of corresponding technical elements will be omitted. Also, since the image decoding device corresponds to a computing device that applies the image coding method performed in the image coding device to decoding, the following description will focus on the image coding device.
[0028] The computing device may include a memory for storing a program or software module for implementing the image encoding method and / or the image decoding method, and a processor coupled to the memory for executing the program, where the image encoding device may be referred to as an encoder and the image decoding device may be referred to as a decoder.
[0029] FIG. 2 is a block diagram showing the configuration of an image encoding device according to an embodiment of the present invention.
[0030] Referring to FIG. 2, the image encoding device 20 may include a prediction unit 200, a subtraction unit 205, a transformation unit 210, a quantization unit 215, an inverse quantization unit 220, an inverse transformation unit 225, an addition unit 230, a filter unit 235, an encoded picture buffer 240, and an entropy encoding unit 245.
[0031] The prediction unit 200 may be implemented using a prediction module, which is a software module, and may generate a predicted block for a block to be coded using intra prediction or inter prediction. The prediction unit 200 may generate a predicted block by predicting a target block to be currently coded in an image. In other words, the prediction unit 200 may predict pixel values of each pixel of a target block to be coded in an image using intra prediction or inter prediction, and generate a predicted block having predicated pixel values of each generated pixel. The prediction unit 200 may also transmit information required for generating a predicted block, such as information about a prediction mode (e.g., intra prediction mode or inter prediction mode), to the coding unit, allowing the coding unit to code the information about the prediction mode. In this case, the processing unit in which prediction is performed and the processing unit in which the prediction method and specific contents are determined may be determined according to coding settings. For example, the prediction method, prediction mode, etc. may be determined in prediction units, and prediction may be performed in transform units. Also, when a specific encoding mode is used, it is possible to directly encode the original block and transmit it to the decoder without generating a predicted block through the predictor.
[0032] The intra prediction unit may have directional prediction modes such as horizontal and vertical modes that are used according to the prediction direction, and non-directional prediction modes such as DC and planar that use methods such as averaging and interpolation of reference pixels. A group of intra prediction mode candidates may be configured based on the directional and non-directional modes, and any of various candidates such as 35 prediction modes (33 directional + 2 non-directional), 67 prediction modes (65 directional + 2 non-directional), or 131 prediction modes (129 directional + 2 non-directional) may be used as the candidate group.
[0033] The intra prediction unit may include a reference pixel construction unit, a reference pixel filter unit, a reference pixel interpolation unit, a prediction mode determination unit, a prediction block generation unit, and a prediction mode encoding unit. The reference pixel construction unit may configure pixels that belong to blocks adjacent to the current block as reference pixels for intra prediction. Depending on encoding settings, the reference pixel construction unit may configure one of the most adjacent reference pixel lines as reference pixels, or one of the other adjacent reference pixel lines as reference pixels, or may configure multiple reference pixel lines as reference pixels. If some of the reference pixels are unavailable, the reference pixels may be generated using available reference pixels. If all of the reference pixels are unavailable, the reference pixels may be generated using a predetermined value (e.g., the median of a pixel value range represented by a bit depth).
[0034] The reference pixel filter unit of the intra prediction unit may filter reference pixels to reduce artifacts remaining after the encoding process. The filter used may be a low-pass filter such as a 3-tap filter [1 / 4, 1 / 2, 1 / 4] or a 5-tap filter [2 / 16, 3 / 16, 6 / 16, 3 / 16, 2 / 16]. Whether or not filtering is applied and the type of filtering may be determined based on encoding information (e.g., block size, shape, prediction mode, etc.).
[0035] The reference pixel interpolator of the intra prediction unit may generate decimal-unit pixels through a linear interpolation process of reference pixels according to a prediction mode, and may determine an interpolation filter to be applied based on coding information. The interpolation filters used may include a 4-tap cubic filter, a 4-tap Gaussian filter, a 6-tap Wiener filter, an 8-tap Kalman filter, etc. While interpolation is generally performed separately from the low-pass filtering process, the filtering process may also be performed by integrating the filters applied in the two processes into one.
[0036] A prediction mode determination unit of the intra prediction unit may select at least one optimal prediction mode from a group of prediction mode candidates in consideration of coding costs, and a prediction block generation unit may generate a prediction block using the selected prediction mode. A prediction mode encoding unit may encode the optimal prediction mode based on a prediction value. In this case, prediction information may be adaptively encoded depending on whether the prediction value is applicable or not.
[0037] The intra prediction unit may set the predicted value as an MPM (Most Probable Mode), and configure some modes from among all modes belonging to a prediction mode candidate group as an MPM candidate group. The MPM candidate group may include a predetermined prediction mode (e.g., DC, planar, vertical, horizontal, diagonal mode, etc.) or a prediction mode of a spatially adjacent block (e.g., left, top, top-left, top-right, bottom-left block, etc.). In addition, a mode derived from a mode already included in the MPM candidate group (in the case of a directional mode, a difference such as 1 or -1) may be configured as an MPM candidate group.
[0038] There may be a priority order of prediction modes for constructing an MPM candidate group. The order of inclusion in the MPM candidate group may be determined according to the priority order, and the construction of the MPM candidate group may be completed when the number of MPM candidates (determined according to the number of prediction mode candidate groups) is filled according to the priority order. In this case, the priority order may be determined in the order of prediction modes of spatially adjacent blocks, a predetermined prediction mode, and a mode derived from a prediction mode initially included in the MPM candidate group, but other variations are also possible.
[0039] For example, spatially adjacent blocks may be included in the candidate group in an order such as left-top-bottom-left-top-top-top left block, and predetermined prediction modes may be included in the candidate group in an order such as DC-Planar-vertical-horizontal mode, and a total of six modes may be configured as the candidate group by adding +1, -1, etc. to an already included mode. Alternatively, a total of seven modes may be configured as the candidate group by including them in a single priority order such as left-top-DC-Planar-bottom-left-top-top-top-(left+1)-(left-1)-(top+1).
[0040] The subtraction unit 205 may subtract a prediction block from a current block to generate a residual block. That is, the subtraction unit 205 may calculate the difference between the pixel value of each pixel of the current block to be coded and the predicted pixel value of each pixel of the prediction block generated via the prediction unit to generate a residual block, which is a block-shaped residual signal. The subtraction unit 205 may also generate a residual block based on a unit other than a block unit obtained via a block division unit (to be described later).
[0041] The transform unit 210 can transform a signal belonging to the spatial domain into a signal belonging to the frequency domain. A signal obtained through the transform process is called a transformed coefficient. For example, a transform block having transform coefficients can be obtained by transforming a residual block having a residual signal transmitted from the subtraction unit. The input signal is determined according to the coding setting and is not limited to a residual signal.
[0042] The transform unit may transform the residual block using a transform technique such as a Hadamard transform, a discrete sine transform (DST based-transform), or a discrete cosine transform (DCT based-transform), but is not limited to these, and various improved and modified transform techniques may be used.
[0043] At least one of the transformation techniques may be supported, and each of the transformation techniques may support at least one detailed transformation technique, where the detailed transformation techniques may be configured such that some of the basis vectors are different for each transformation technique.
[0044] For example, in the case of DCT, one or more detailed conversion techniques from DCT-1 to DCT-8 can be supported, and in the case of DST, one or more detailed conversion techniques from DST-1 to DST-8 can be supported. A group of candidate conversion techniques can be configured by configuring some of the detailed conversion techniques. For example, DCT-2, DCT-8, and DST-7 can be configured as candidate conversion techniques for conversion.
[0045] The transformation can be horizontal or vertical. For example, a spatial domain pixel value can be transformed into the frequency domain using a one-dimensional transform in the horizontal direction (using a DCT-2 transform technique) and a one-dimensional transform in the vertical direction (using a DST-7 transform technique), resulting in a total of two-dimensional transformation.
[0046] Transformation can be performed using a single fixed transform technique, or by adaptively selecting a transform technique according to encoding settings. In this case, in the adaptive case, the transform technique can be selected using an explicit or implicit method. In the explicit case, selection information for each transform technique or transform technique set applied in the horizontal and vertical directions can be generated in units such as blocks. In the implicit case, encoding settings can be defined according to the image type (I / P / B), color components, block size / shape / position, intra-frame prediction mode, etc., and a predetermined transform technique can be selected accordingly.
[0047] Also, some of the conversions may be omitted depending on the encoding settings, which means that one or more horizontal / vertical units may be omitted explicitly or implicitly.
[0048] The transform unit can transmit the information necessary to generate a transform block to the encoder to encode it, and then record the information into a bitstream and transmit it to the decoder, and the decoding unit of the decoder can parse the information and use it in the inverse transform process.
[0049] The quantization unit 215 may quantize an input signal. At this time, a signal obtained through the quantization process is called a quantized coefficient. For example, a residual block having a residual transform coefficient transmitted from a transform unit may be quantized to obtain a quantized block having a quantized coefficient. However, the input signal is determined according to a coding setting, and is not limited to a residual transform coefficient.
[0050] The quantization unit can quantize the transformed residual block using a quantization technique such as Dead Zone Uniform Threshold Quantization or Quantization Weighted Matrix, but is not limited to these, and various improved and modified quantization techniques can be used.
[0051] The quantization process can be omitted depending on the encoding settings. For example, the quantization process (including the inverse process) can be omitted depending on the encoding settings (e.g., a quantization parameter of 0, i.e., a lossless compression environment). As another example, the quantization process can be omitted if the compression performance of quantization is not exhibited depending on the characteristics of the image. In this case, the region in the quantization block (M×N) where the quantization process is omitted is the entire region or a partial region (M / 2×N / 2, M×N / 2, M / 2×N, etc.), and the quantization omission selection information can be implicitly or explicitly determined.
[0052] The quantization unit can transmit the information necessary to generate a quantization block to the encoding unit to encode it, and then record the information into a bitstream and transmit it to the decoder, and the decoding unit of the decoder can parse the information and use it in the inverse quantization process.
[0053] In the above example, the explanation is based on the assumption that the residual block is transformed and quantized through a transform unit and a quantization unit. However, the residual signal may be transformed to generate a residual block having transform coefficients without performing the quantization process. Alternatively, the residual signal of the residual block may be subjected to only the quantization process without converting it into transform coefficients, or both the transform and quantization processes may be omitted. This can be determined according to the settings of the encoder.
[0054] The inverse quantization unit 220 inverse quantizes the residual block quantized by the quantization unit 215. That is, the inverse quantization unit 220 inverse quantizes the quantized frequency coefficient sequence to generate a residual block having frequency coefficients.
[0055] The inverse transform unit 225 inversely transforms the residual block dequantized by the inverse quantization unit 220. That is, the inverse transform unit 225 inversely transforms the frequency coefficients of the dequantized residual block to generate a residual block having pixel values, i.e., a reconstructed residual block. Here, the inverse transform unit 225 can perform inverse transform by using the transform method used by the transform unit 210 in reverse.
[0056] The adder 230 reconstructs the current block by adding the prediction block predicted by the predictor 200 and the residual block reconstructed by the inverse transformer 225. The reconstructed current block is stored in the coding picture buffer 240 as a reference picture (or reference block), and can be used as a reference picture when encoding the next block of the current block or another block or picture following the current block.
[0057] The filter unit 235 may include one or more post-processing filter processes, such as a deblocking filter, a sample adaptive offset (SAO), and an adaptive loop filter (ALF). The deblocking filter can remove block artifacts that occur at boundaries between blocks from a reconstructed picture. The ALF can perform filtering based on a value obtained by comparing an image reconstructed after a block is filtered through a deblocking filter with an original image. The SAO can restore an offset difference between a residual block to which the deblocking filter is applied and an original image on a pixel-by-pixel basis. Such post-processing filters can be applied to a reconstructed picture or block.
[0058] The coded picture buffer 240 can store blocks or pictures reconstructed through the filter unit 235. The reconstructed blocks or pictures stored in the coded picture buffer 240 can be provided to the prediction unit 200, which performs intra prediction or inter prediction.
[0059] The entropy coding unit 245 may scan the quantized coefficients, transform coefficients, or residual signals of the generated residual block according to at least one scan order (e.g., zigzag scan, vertical scan, horizontal scan, etc.) to generate a quantized coefficient sequence, a transform coefficient sequence, or a signal sequence, and encode the quantized coefficient sequence using at least one entropy coding technique. In this case, information about the scan order may be determined according to coding settings (e.g., image type, coding mode, prediction mode, transform type, etc.), and may be implicitly determined or related information may be explicitly generated.
[0060] Also, coded data including coding information transmitted from each component can be generated and output as a bitstream, which can be achieved by a multiplexer (MUX). In this case, coding can be performed using methods such as Exponential Golomb, Context-Adaptive Variable Length Coding (CAVLC), and Context-Adaptive Binary Arithmetic Coding (CABAC). However, various coding methods that are improvements or modifications of these methods can also be used.
[0061] When performing entropy coding (assumed to be CABAC in this example) on syntax elements such as the residual block data and information generated during the encoding / decoding process, the entropy coding device may include a binarizer, a context modeler, and a binary arithmetic coder. In this case, the binary arithmetic coder may include a regular coding engine and a bypass coding engine.
[0062] Since the syntax elements input to the entropy coding device may not be binary values, if the syntax elements are not binary values, a binarization unit can binarize the syntax elements and output a bin string consisting of 0 or 1. In this case, a bin indicates a bit consisting of 0 or 1 and can be coded by a binary arithmetic coding unit. In this case, either a regular coding unit or a bypass coding unit can be selected based on the occurrence probability of 0 and 1, which can be determined based on coding / decoding settings. If the syntax elements are data with the same frequency of 0 and 1, a bypass coding unit can be used, and if not, a regular coding unit can be used.
[0063] Various methods can be used to binarize the syntax elements. For example, fixed length binarization, unary binarization, truncated rice binarization, K-th Exp-Golomb binarization, etc. can be used. Also, signed or unsigned binarization can be performed depending on the value range of the syntax elements. The binarization process for the syntax elements generated in the present invention can be performed not only by the binarization methods mentioned in the above examples, but also by other additional binarization methods.
[0064] FIG. 3 is a block diagram showing the configuration of an image decoding device according to an embodiment of the present invention.
[0065] Referring to FIG. 3, the image decoding device 30 may include an entropy decoding unit 305, a prediction unit 310, an inverse quantization unit 315, an inverse transform unit 320, an adder / subtractor 325, a filter 330, and a decoded picture buffer 335.
[0066] The prediction unit 310 can further include an intra-frame prediction module and an inter-frame prediction module.
[0067] First, when an image bitstream transmitted from the image encoding device 20 is received, it can be transmitted to the entropy decoding unit 305 .
[0068] The entropy decoding unit 305 can decode the bitstream to generate decoded data including quantized coefficients and decoding information to be transmitted to each component.
[0069] The prediction unit 310 may generate a prediction block based on data received from the entropy decoding unit 305. In this case, the prediction unit 310 may construct a reference picture list using a default construction technique based on reference images stored in the decoded picture buffer 335.
[0070] The intra-screen prediction unit may include a reference pixel construction unit, a reference pixel filter unit, a reference pixel interpolation unit, a prediction block generation unit, and a prediction mode decoding unit, some of which may perform the same process as the encoder, and some of which may perform the reverse induction process.
[0071] The inverse quantization unit 315 can inverse quantize the quantized transform coefficients provided as a bitstream and decoded by the entropy decoding unit 305 .
[0072] The inverse transform unit 320 may apply an inverse transform technique such as an inverse DCT, an inverse integer transform, or a similar concept to the transform coefficients to generate residual blocks.
[0073] In this case, the inverse quantization unit 315 and the inverse transform unit 320 can be realized in various ways by reversing the processes performed by the transform unit 210 and the quantization unit 215 of the image encoding device 20 described above. For example, they can use the same processes and inverse transforms shared by the transform unit 210 and the quantization unit 215, or can reverse the transform and quantization processes using information about the transform and quantization processes from the image encoding device 20 (e.g., transform size, transform shape, quantization type, etc.).
[0074] The residual block that has undergone the inverse quantization and inverse transform processes may be added to the prediction block derived by the prediction unit 310 to generate a reconstructed image block. This addition may be performed by the adder / subtractor 325.
[0075] The filter 330 may also apply a deblocking filter to the reconstructed image blocks to remove blocking artifacts if necessary, and a separate loop filter may also be used before or after the decoding process to further improve the video quality.
[0076] The reconstructed and filtered image blocks can be stored in the decoded picture buffer 335 .
[0077] Although not shown, the image encoding / decoding device may further include a picture dividing unit and a block dividing unit.
[0078] The picture division unit may divide or partition a picture into at least one region based on a predetermined division unit, where the division unit may include a sub-picture, a slice, a tile, a brick, a block (e.g., a maximum coding unit), etc.
[0079] A picture can be divided into one or more tile rows or one or more tile columns. A tile can be a block-based unit including a predetermined rectangular area of the picture. A tile can be divided into one or more bricks, and a brick can be composed of blocks in units of tile rows or columns.
[0080] A slice can have one or more configurations, one of which can be a bundle of scan orders (e.g., blocks, bricks, tiles, etc.), one of which can be a shape containing rectangular regions, and other additional definitions are possible.
[0081] The definition of the slice configuration may be explicitly related information or implicitly defined. The definition of the configuration of each division unit as well as the slice may be set in multiple ways, and selection information related thereto may be generated.
[0082] A slice can be configured in rectangular units such as blocks, bricks, or tiles, and slice position and size information can be expressed based on position information (e.g., upper left position, lower right position, etc.) for each division unit.
[0083] In this invention, we will assume that a picture can be made up of one or more sub-pictures, a sub-picture can be made up of one or more slices or tiles or bricks, a slice can be made up of one or more tiles or bricks, and a tile can be made up of one or more bricks, but this is not limited to this.
[0084] The division unit may be composed of an integer number of blocks, but is not limited thereto, and may be composed of a decimal number instead of an integer number. In other words, if the division unit is not composed of an integer number of blocks, at least one division unit may be composed of a sub-block.
[0085] There may be division units other than rectangular slices, such as sub-pictures and tiles, and information on the position and size of the units may be expressed in various ways.
[0086] For example, the position and size information of rectangular units can be expressed based on information on the number of rectangular units, information on the number of columns or rows of rectangular units, information on whether the columns or rows of rectangular units are evenly divided, information on the width or height of the columns or rows of rectangular units, index information of rectangular units, etc.
[0087] In the case of subpictures and tiles, the position and size information of each unit can be expressed based on all or part of the above information, and the image can be divided or partitioned into one or more units based on this.
[0088] Meanwhile, the image can be divided into blocks of various units and sizes through the block division unit. A basic coding unit (or largest coding unit, or coding tree unit, CTU) may refer to a basic (or starting) unit for prediction, transformation, quantization, etc. in the image coding process. In this case, the basic coding unit may be composed of one luminance basic coding block (or largest coding block, or coding tree block, CTB) and two basic chrominance coding blocks according to the color format (YCbCr in this example), and the size of each block may be determined according to the color format. A coding block (CB) can be obtained according to the division process. A coding block can be understood as a unit that is not divided into further coding blocks according to certain restrictions, and can be set as a starting unit for division into lower units. In the present invention, a block is not limited to a rectangular shape but can be understood as a broad concept including various shapes such as a triangle and a circle.
[0089] Although the following description focuses on one color component, it should be understood that it can be changed and applied to other color components in proportion to the ratio according to the color format (for example, in the case of YCbCr4:2:0, the horizontal to vertical length ratio of the luminance component and the chrominance component is 2:1). It should also be understood that although block division that depends on other color components (for example, in the case of Cb / Cr, depending on the block division result of Y) is possible, independent block division is also possible for each color component. It should also be understood that while one common block division setting (taking into consideration the proportionality to the length ratio) can be used, individual block division settings can be used depending on the color component.
[0090] In the block division part, the blocks can be expressed as M × N, and the maximum and minimum values of each block are The value can be obtained within the range. For example, if the maximum value of the block is 256x256 and the minimum value is 4x4, the size is 2 m ×2 n(in this example, m and n are integers from 2 to 8), or a block of size 2m x 2m (in this example, m and n are integers from 2 to 128), or a block of size m x m (in this example, m and n are integers from 4 to 256). Here, m and n may or may not be the same, and the ranges supported by the blocks, such as the maximum and minimum values, may occur more than once.
[0091] For example, information on the maximum and minimum block sizes may be generated, and information on the maximum and minimum block sizes in a partial partition setting may be generated. Here, the former may be information on the range of the maximum and minimum sizes that can be generated within an image, and the latter may be information on the maximum and minimum sizes that can be generated based on a partial partition setting. Here, the partition setting may be defined by an image type (I / P / B), color components (YCbCr, etc.), block type (encoding / prediction / transform / quantization, etc.), partition type (Index or Type), partition method (QT, BT, TT, etc. in the tree method, SI2, SI3, SI4, etc. in the index method), etc.
[0092] In addition, there may be restrictions on the width / height ratio (block shape) that a block can have, and a boundary value condition for this can be set. In this case, only blocks that are equal to or less than a given boundary value (k) can be supported, and k can be defined based on the width / height ratio such as A / B (A is the longer or same value of the width or height, and B is the remaining value), and can be a real number greater than or equal to 1, such as 1.5, 2, 3, or 4. As in the above example, a restriction on the shape of one block in an image is supported, or more than one restriction can be supported depending on the division setting.
[0093] In summary, whether or not block division is supported can be determined based on the range and conditions described above and the division settings described below, etc. For example, if a candidate (child block) resulting from division of a block (parent block) satisfies the supported block conditions, the division can be supported, and if not, the division cannot be supported.
[0094] The block division unit can be set in relation to each component of the image encoding device and decoding device, and the size and shape of the block can be determined through this process. At this time, the set block can be defined differently depending on the component, and can correspond to a prediction block in the case of a predictor, a transform block in the case of a transformer, a quantization block in the case of a quantizer, etc. However, without being limited thereto, block units can be further defined according to other components. In the present invention, the case where the input and output of each component are rectangular will be mainly described, but some components can have inputs / outputs of other shapes (e.g., right-angled triangles, etc.).
[0095] The size and shape of the initial (or starting) block of the block division unit can be determined from the upper unit. The initial block can be divided into blocks of smaller sizes, and once the optimal size and shape for the block division is determined, the block can be determined as the initial block of the lower unit. Here, the upper unit can be a coding block, and the lower unit can be a prediction block or a transformation block, but is not limited thereto, and various variations are possible. Once the initial block of the lower unit is determined as in the above example, a division process can be performed to find a block of the optimal size and shape, like the upper unit.
[0096] In summary, the block division unit can divide a basic coding block (or a maximum coding block) into at least one coding block, and can divide the coding block into at least one prediction block / transform block / quantization block. Furthermore, the prediction block can be divided into at least one transform block / quantization block, and the transform block can be divided into at least one quantization block. Here, some blocks may have a subordinate relationship (i.e., defined by a higher-order unit and a lower-order unit) with other blocks, or an independent relationship. For example, the prediction block may be a higher-order unit of the transform block, or may be a unit independent of the transform block. Various relationship settings are possible depending on the type of block.
[0097] Depending on the encoding setting, whether or not to combine the upper unit and the lower unit may be determined. Here, combining between units means that the upper unit is not divided into the lower unit, but the encoding process (e.g., prediction unit, transform unit, inverse transform unit, etc.) of the lower unit is performed using the block (size and shape) of the upper unit. In other words, it may mean that the division process of multiple units is shared, and the division information is generated in one of the units (e.g., the upper unit).
[0098] For example, prediction, transformation, and inverse transformation processes can be performed on a coding block (when the coding block is combined with a prediction block and a transformation block).
[0099] For example, a prediction process can be performed on a coding block (when the coding block is combined with a prediction block), and a transform and inverse transform process can be performed on a transform block that is the same size as or smaller than the coding block.
[0100] For example, a prediction process can be performed using a prediction block that is the same size as or smaller than the coding block (when the coding block is combined with a transform block), and a transform and inverse transform process can be performed on the coding block.
[0101] For example, a prediction process can be performed on a prediction block that is the same size as or smaller than the coding block (when the prediction block is combined with a transformation block), and a transformation and inverse transformation process can be performed on the prediction block.
[0102] As an example, the prediction process can be performed using a prediction block that is the same as or smaller than the coding block (if neither block is combined), and the transform and inverse transform processes can be performed using a transform block that is the same as or smaller than the coding block.
[0103] Although various cases regarding coding, prediction, and transformation blocks have been described in the above examples, the present invention is not limited to these.
[0104] The combination between the units may be supported by a fixed setting for an image or by an adaptive setting in consideration of various coding factors, which may include an image type, color components, a coding mode (Intra / Inter), a partition setting, a block size / shape / position, a width / height ratio, prediction-related information (e.g., an intra prediction mode, an inter prediction mode, etc.), transform-related information (e.g., transform technique selection information, etc.), and quantization-related information (e.g., quantization region selection information, quantized transform coefficient coding information, etc.).
[0105] As described above, when a block having an optimal size and shape is found, mode information (e.g., partition information, etc.) for the block can be generated. The mode information can be recorded in a bitstream together with information generated in the component to which the block belongs (e.g., prediction-related information, transformation-related information, etc.) and transmitted to a decoder, where it can be parsed into units of the same level and used in the image decoding process.
[0106] In the following, the division method will be described, and for convenience of explanation, it is assumed that the initial block is square in shape, but this is not limited to this, as it can be applied in the same or similar manner when the initial block is rectangular in shape.
[0107] The block division unit may support various types of division. For example, it may support tree-based division or index-based division, or other methods. Tree-based division may determine the division type based on various types of information (e.g., whether or not to divide, tree type, division direction, etc.), while index-based division may determine the division type based on predetermined index information.
[0108] FIG. 4 is an exemplary diagram showing various division forms that can be obtained by the block division unit of the present invention.
[0109] In this example, it is assumed that the division pattern shown in Figure 4 is obtained by one division execution (or process), but this is not limited to this and it can be obtained by multiple division operations. Also, additional division patterns not shown in Figure 4 are possible.
[0110] (tree-based partitioning) The tree-based partitioning of the present invention can support quad trees (QT), binary trees (BT), ternary trees (TT), etc. When one tree type is supported, it is called single-tree partitioning, and when two or more tree types are supported, it is called multi-tree partitioning.
[0111] QT refers to a method (n) in which a block is divided into two parts horizontally and vertically (i.e., into four parts), BT refers to a method (b to g) in which a block is divided into two parts in one of the horizontal or vertical directions, and TT refers to a method (h to m) in which a block is divided into three parts in one of the horizontal or vertical directions.
[0112] Here, in the case of QT, a four-quarter division scheme (o, p) can be supported, with the division direction limited to either horizontal or vertical. In addition, in the case of BT, only equal-sized schemes (b, c) or non-uniform-sized schemes (d to g) can be supported, or a combination of both schemes can be supported. In addition, in the case of TT, only schemes (h, j, k, m) with a biased division arrangement (e.g., 1:1:2, 2:1:1 in the left-to-right or top-to-bottom direction) can be supported, or only schemes (i, l) with a centered arrangement (e.g., 1:2:1) can be supported, or a combination of both schemes can be supported. In addition, a four-quarter division scheme (i.e., 16 divisions) in both horizontal and vertical directions (q) can also be supported.
[0113] The tree method may support z-division in the horizontal direction (b, d, e, h, i, j, o), z-division in the vertical direction (c, f, g, k, l, m, p), or a combination of both methods, where z is an integer equal to or greater than 2, such as 2, 3, or 4.
[0114] In the present invention, it is assumed that QT supports n, BT supports b and c, and TT supports i and l.
[0115] Depending on the encoding settings, one or more of the tree splitting methods can be supported, for example, QT can be supported, or QT / BT can be supported, or QT / BT / TT can be supported.
[0116] The above example shows a case where the basic tree division is QT, and BT and TT are included in the additional division method depending on whether other trees are supported, but various modifications are possible. In this case, information on whether other trees are supported (bt_enabled_flag, tt_enabled_flag, bt_tt_enabled_flag, etc., which can have a value of 0 or 1, where 0 indicates no support and 1 indicates support) can be implicitly determined according to the coding settings, or can be explicitly determined in units such as sequence, picture, subpicture, slice, tile, brick, etc.
[0117] The division information may include information on whether or not to divide (tree_part_flag, or qt_part_flag, bt_part_flag, tt_part_flag, or bt_tt_part_flag, which may have a value of 0 or 1, where 0 indicates no division and 1 indicates division). In addition, depending on the division method (BT or TT), information on the division direction (dir_part_flag, or bt_dir_part_flag, tt_dir_part_flag, or bt_tt_dir_part_flag, which may have a value of 0 or 1, where 0 indicates <horizontal> and 1 indicates <vertical>) may be added, and this may be information that can be generated when division is performed.
[0118] When multiple tree partitions are supported, various partition information configurations are possible. The following will be described as an example of how partition information is configured at one depth level (i.e., for the sake of convenience, although recursive partitioning may be possible if the supported partition depth is set to one or more).
[0119] As an example (1), information on whether or not division is required is checked. If division is not required, the division is terminated.
[0120] If splitting is to be performed, the selection information regarding the type of splitting (for example, tree_idx. If it is 0, it is QT, if it is 1, it is BT, if it is 2, it is TT) is checked. At this time, the splitting direction information is further checked according to the type of splitting selected, and the process moves to the next step (if additional splitting is possible because the splitting depth has not reached the maximum, it starts again from the beginning, and if splitting is not possible, it ends the splitting).
[0121] As an example (2), check the information as to whether or not the split is for a partial tree method (QT) and move to the next step. At this time, if the split is not to be performed, check the information as to whether or not the split is for a partial tree method (BT). At this time, if the split is not to be performed, check the information as to whether or not the split is for a partial tree method (TT). At this time, if the split is not to be performed, the split ends.
[0122] If partial tree splitting (QT) is to be performed, proceed to the next step. Also, if partial tree splitting (BT) is to be performed, check the split direction information and proceed to the next step. Also, if partial tree splitting (TT) is to be performed, check the split direction information and proceed to the next step.
[0123] As an example (3), check whether or not a split is required for some tree methods (QT). If splitting is not required, check whether or not a split is required for some tree methods (BT and TT). If splitting is not required, the splitting is terminated.
[0124] If a partial tree-based split (QT) is performed, proceed to the next step. If a partial tree-based split (BT and TT) is performed, check the split direction information and proceed to the next step.
[0125] The above examples may have tree splitting priority (examples 2 and 3) or not (example 1), but various variations are possible. Also, the above examples illustrate cases where the splitting of the current step is independent of the splitting results of the previous step, but it is also possible to set the splitting of the current step to depend on the splitting results of the previous step.
[0126] For example, in examples 1 to 3, if some tree-based splitting (QT) was performed in the previous step and was moved to the current step, the same tree-based splitting (QT) can be supported in the current step.
[0127] On the other hand, if some tree-based splits (QT) were not performed in the previous step, but other tree-based splits (BT or TT) were performed and then moved to the current step, it is possible to set the partial tree-based splits (BT and TT) to be supported in subsequent steps including the current step, except for the partial tree-based split (QT).
[0128] In the above case, it means that the tree structure supported by block division can be adaptive, and therefore the above division information structure can also be configured differently. (The following example assumes the third example.) In other words, in the above example, if division of some tree methods (QT) has not been performed in the previous step, the current step can perform the division process without considering some tree methods (QT). In addition, division information related to the related tree methods (e.g., information on whether or not there is a division, division direction information, etc.) is also used. In this example, <qt>In this case, the information about whether it is a split or not can be removed and configured.
[0129] The above example concerns adaptive partitioning information configuration when block partitioning is allowed (e.g., the block size is within the range between the maximum and minimum values, or the partitioning depth of each tree method does not reach the maximum depth <allowable depth>), but adaptive partitioning information configuration is also possible when block partitioning is restricted (e.g., the block size is not within the range between the maximum and minimum values, or the partitioning depth of each tree method reaches the maximum depth).
[0130] As already mentioned, in the present invention, tree-based partitioning can be performed using a recursive method. For example, if the partition flag of a coding block with a partition depth of k is 0, the coding block is coded using a coding block with a partition depth of k, and if the partition flag of a coding block with a partition depth of k is 1, the coding block is coded using N sub-coding blocks (where N is an integer equal to or greater than 2, such as 2, 3, or 4) with a partition depth of k+1 according to the partitioning method.
[0131] The sub-coding block is again set as a coding block (k+1) and can be divided into a sub-coding block (k+2) through the above process. Such a hierarchical division method can be determined according to division settings such as the division range and the allowable division depth.
[0132] In this case, the bitstream structure for expressing the partition information can be selected from one or more scanning methods. For example, the bitstream of the partition information can be configured based on the order of the partition depth, or based on whether or not there is a partition.
[0133] For example, when the order of division depth is used as the criterion, this is a method of obtaining division information at the current level depth based on the first block, and then obtaining division information at the next level depth.When the criterion is whether or not a block is divided, this means a method of preferentially obtaining additional division information for blocks divided based on the first block, and other additional scanning methods can be considered.
[0134] The maximum block size and the minimum block size can be set to a common setting regardless of the type of tree (or all trees), or they can be set individually for each tree, or they can be set to a common setting for two or more trees. In this case, the maximum block size can be set to be equal to or smaller than the maximum coding block. If the maximum block size of a given first tree is not the same as the maximum coding block, implicit division is performed using a given second tree method until the maximum block size of the first tree is reached.
[0135] A common partition depth may be supported regardless of the type of tree, or an individual partition depth may be supported for each tree, or a common partition depth may be supported for two or more trees, or a partition depth may be supported for some trees and not for other trees.
[0136] Explicit syntax elements for the configuration information can be supported, and some configuration information may be implicit.
[0137] (index-based partitioning)
[0138] In the index-based division of the present invention, a CSI (Constant Split Index) scheme and a VSI (Variable Split Index) scheme can be supported.
[0139] The CSI scheme may be a scheme in which k sub-blocks are obtained by division in a predetermined direction, where k may be an integer equal to or greater than 2, such as 2, 3, or 4. Specifically, the CSI scheme may be a division scheme in which the size and shape of the sub-blocks are determined based on the value of k, regardless of the size and shape of the block. Here, the predetermined direction may be one or a combination of two or more of horizontal, vertical, and diagonal directions (e.g., from upper left to lower right or from lower left to upper right).
[0140] The index-based CSI partitioning scheme of the present invention can include z partition candidates in either the horizontal or vertical direction, where z is an integer equal to or greater than 2, such as 2, 3, or 4, and one of the horizontal or vertical lengths of each sub-block may be the same, and the other may be the same or different. The horizontal or vertical length ratio of the sub-blocks is A1:A2:...:A Z A1 to A Z can be an integer equal to or greater than 1, such as 1, 2, 3, etc.
[0141] Also, candidates for division into x and y in the horizontal and vertical directions, respectively, may be included. In this case, x and y may be integers equal to or greater than 1, such as 1, 2, 3, or 4, but restrictions may be imposed if x and y are both 1 (because a already exists). Although Fig. 4 shows the case where the horizontal or vertical length ratio of each sub-block is the same, candidates including cases where they are different may also be included.
[0142] In addition, it may include candidates that are divided into w in either a partial diagonal direction (top left → bottom right) or a partial diagonal direction (bottom left → top right), where w may be an integer greater than or equal to 2, such as 2 or 3.
[0143] 4, the partitioning pattern can be classified into a symmetrical partitioning pattern (b) and an asymmetrical partitioning pattern (d, e) according to the length ratio of each sub-block, and into a partitioning pattern (k, m) biased in a specific direction and a partitioning pattern (k) arranged in the center. The partitioning pattern can be defined according to various coding factors including the sub-block shape as well as the length ratio of the sub-blocks, and the supported partitioning pattern can be implicitly or explicitly determined depending on the coding setting. Therefore, a group of candidates for the index-based partitioning scheme can be determined based on the supported partitioning pattern.
[0144] Meanwhile, the VSI method may be a method in which one or more sub-blocks are obtained by dividing a sub-block in a predetermined direction while the width w or height h of the sub-block is fixed, where w and h may be integers equal to or greater than 1, such as 1, 2, 4, or 8. In particular, the VSI method may be a division method in which the number of sub-blocks is determined based on the size and shape of the block and the value of w or n.
[0145] The index-based VSI partitioning method of the present invention may include candidates that are partitioned by fixing either the horizontal or vertical length of the sub-block, or may include candidates that are partitioned by fixing both the horizontal and vertical lengths of the sub-block. Since the horizontal or vertical length of the sub-block is fixed, it may have a feature that allows equal division in the horizontal or vertical direction, but is not limited thereto.
[0146] If the block before division is M×N and the horizontal length of the sub-block is fixed (w), or the vertical length is fixed (h), or the horizontal and vertical lengths are fixed (w, h), the number of sub-blocks obtained can be (M*N) / w, (M*N) / h, or (M*N) / w / h, respectively.
[0147] Depending on the coding configuration, only the CSI method may be supported, or only the VSI method may be supported, or both methods may be supported, and information about the supported methods may be implicitly or explicitly specified.
[0148] In the present invention, it is assumed that the CSI scheme is supported.
[0149] Depending on the encoding settings, the set of candidates may include two or more of the index splits.
[0150] For example, a candidate group such as {a, b, c}, {a, b, c, n}, or {a to g, n} can be constructed. However, this may be an example of constructing a candidate group based on block shapes that are predicted to occur frequently based on general statistical characteristics, such as block shapes that are divided into two in the horizontal or vertical direction, or divided into two in both the horizontal and vertical directions.
[0151] Alternatively, a candidate group such as {a, b}, {a, o}, {a, b, o} or {a, c}, {a, p}, {a, c, p} can be constructed, which includes candidates divided into two and four in the horizontal and vertical directions, respectively. This can be an example of constructing a candidate group based on block shapes that are predicted to be frequently divided in a specific direction.
[0152] Alternatively, a candidate group such as {a, o, p} or {a, n, q} can be constructed, but an example of constructing a candidate group may be a block shape that is predicted to result in many divisions having a size smaller than the block before division.
[0153] Alternatively, a candidate group such as {a, r, s} can be constructed, but it may be determined that the optimal division result, which can be obtained as a rectangular shape using another method (tree method) from the block before division, has been obtained, and a non-rectangular division form may be constructed as a candidate group.
[0154] As in the above example, various candidate group configurations are possible, and more than one candidate group configuration can be supported taking into account various coding factors.
[0155] Once the candidate set has been constructed, various partition information configurations are possible.
[0156] For example, index selection information can be generated from a candidate group including a candidate (a) that is not divided and candidates (b to s) that are divided.
[0157] Alternatively, information indicating whether or not a division is to be performed (whether the division type is a or not) can be generated, and if division is to be performed (if not a), index selection information can be generated from a candidate group consisting of candidates to be divided (b to s).
[0158] Various methods other than those described above can be used to configure the partition information, and binary bits can be assigned to the index of each candidate in the candidate group, except for the information indicating whether or not the candidate is partitioned, using various methods such as fixed-length binarization, variable-length binarization, etc. If the number of candidate groups is two, one bit can be assigned to the index selection information, and if the number is three or more, one or more bits can be assigned to the index selection information.
[0159] Unlike the tree-based partitioning method, the index-based partitioning method can be a method of selectively configuring candidate partitioning patterns that are predicted to occur frequently.
[0160] Also, since the number of bits for expressing index information can increase depending on the number of supported candidate groups, this method may be suitable for single-level division (e.g., division depth is limited to 0) rather than tree-based hierarchical division (recursive division). That is, it may be a method that supports one division operation, or a method in which sub-blocks obtained through index-based division cannot be further divided.
[0161] In this case, it may mean that further division into blocks of the same type having a smaller size is not possible (for example, a coding block obtained by an index division method cannot be further divided into coding blocks), but it may also mean that further division into blocks of other types is not possible (for example, division of a coding block into not only coding blocks but also prediction blocks is not possible). Of course, this is not limited to the above example, and other variations are possible.
[0162] Next, a case where block division is determined mainly based on the type of block among the coding elements will be considered.
[0163] First, coding blocks are obtained through a splitting process. Here, a tree-based splitting method can be used for the splitting process, and splitting results such as a (no split), n (QT), b, c (BT), i, l (TT) as shown in Figure 4 can be obtained depending on the tree type. Depending on the coding settings, various combinations of tree types such as QT, QT+BT, and QT+BT+TT are possible.
[0164] The example described below shows the process of finally dividing a prediction block and a transformation block based on the coding block obtained by the above process, and assumes that prediction, transformation, and inverse transformation processes are performed based on each division size.
[0165] For example (1), a prediction block may be set to the same size as the coding block to perform a prediction process, and a transformation block may be set to the same size as the coding block (or prediction block) to perform a transformation and inverse transformation process. Since the prediction block and the transformation block are set based on the coding block, no separate partition information is generated.
[0166] For example, (2) a prediction block may be set to the same size as the coding block and a prediction process may be performed. In the case of a transform block, a transform block may be obtained through a division process based on the coding block (or the prediction block), and a transform and inverse transform process may be performed based on the obtained size.
[0167] Here, a tree-based splitting method can be used for the splitting process, and splitting results such as a (no split), b, c (BT), i, l (TT), and n (QT) can be obtained depending on the tree type. Depending on the encoding settings, various combinations of tree types are possible, such as QT / BT / QT+BT / QT+BT+TT.
[0168] Here, the splitting process can use an index-based splitting method, and depending on the type of index, splitting results such as a (no split), b, c, and d in Figure 4 can be obtained. Depending on the encoding settings, various candidate sets such as {a, b, c} and {a, b, c, d} can be constructed.
[0169] As an example (3), in the case of a prediction block, a prediction block can be obtained by performing a division process based on a coding block, and a prediction process can be performed based on the obtained size. In the case of a transformation block, the size of the coding block can be set as is and transformation and inverse transformation processes can be performed. This example may correspond to a case where the prediction block and the transformation block have an independent relationship with each other.
[0170] Here, the splitting process can use an index-based splitting method, and splitting results such as a (no split), b to g, n, r, and s in Figure 4 can be obtained depending on the type of index. Depending on the encoding settings, various candidate groups can be configured, such as {a, b, c, n}, {a to g, n}, and {a, r, s}.
[0171] As an example (4), in the case of a prediction block, a prediction block can be obtained by performing a division process based on the coding block, and a prediction process can be performed based on the obtained size. In the case of a transformation block, the size of the prediction block can be set as is, and then the transformation and inverse transformation processes can be performed. In this example, the transformation block may be set as is the size of the obtained prediction block, or vice versa (the prediction block is set as is the size of the transformation block).
[0172] Here, a tree-based splitting method can be used for the splitting process, and splitting patterns such as a (no split), b, c (BT), and n (QT) can be obtained depending on the tree type. Depending on the encoding settings, various combinations of tree types such as QT / BT / QT+BT are possible.
[0173] Here, an index-based splitting method can be used for the splitting process, and splitting patterns such as a (no split), b, c, n, o, and p in Fig. 4 can be obtained depending on the type of index. Depending on the coding setting, various candidate sets can be configured, such as {a, b}, {a, c}, {a, n}, {a, o}, {a, p}, {a, b, c}, {a, o, p}, {a, b, c, n}, and {a, b, c, n, p}. Furthermore, among the index-based splitting methods, the VSI method may be used alone or in combination with the CSI method to configure the candidate sets.
[0174] As an example (5), in the case of a prediction block, a prediction block can be obtained by performing a division process based on a coding block, and a prediction process can be performed based on the obtained size. Also, in the case of a transformation block, a prediction block can be obtained by performing a division process based on a coding block, and a transformation process and an inverse transformation process can be performed based on the obtained size. This example may be a case where a prediction block and a transformation block are each divided based on a coding block.
[0175] Here, the division process can use a tree-based division method or an index-based division method, and the candidate group can be constructed in the same or similar manner as in Example 4.
[0176] The above examples illustrate some possible cases depending on whether the division process of each type of block is shared, but the present invention is not limited to these examples and various variations are possible. Furthermore, block division settings may be determined taking into consideration not only the type of block but also various coding factors.
[0177] In this case, the coding elements may include image type (I / P / B), color components (YCbCr), block size / shape / position, block horizontal / vertical length ratio, block type (coding block, prediction block, transform block, quantization block, etc.), division state, coding mode (Intra / Inter), prediction-related information (intra-frame prediction mode, inter-frame prediction mode, etc.), transformation-related information (transformation technique selection information, etc.), quantization-related information (quantization region selection information, quantized transformation coefficient coding information, etc.), etc.
[0178] In an image encoding method according to an embodiment of the present invention, inter prediction may be configured as follows. The inter prediction of the predictor may include a reference picture construction step, a motion estimation step, a motion compensation step, a motion information determination step, and a motion information encoding step. Also, the image encoding apparatus may be configured to include a reference picture construction unit, a motion estimation unit, a motion compensation unit, a motion information determination unit, and a motion information encoding unit that implement the reference picture construction step, the motion estimation step, the motion compensation step, the motion information determination step, and the motion information encoding step. Some of the above-described processes may be omitted, or other processes may be added, and the processes may be arranged in a different order than the order described above.
[0179] In an image decoding method according to an embodiment of the present invention, inter prediction may be configured as follows. The inter prediction of the predictor may include a motion information decoding step, a reference picture construction step, and a motion compensation step. Also, the image decoding apparatus may be configured to include a motion information decoding unit, a reference picture construction unit, and a motion compensation unit that implement the motion information decoding step, the reference picture construction step, and the motion compensation step. Some of the above-described processes may be omitted, or other processes may be added, and the processes may be arranged in a different order than the order described above.
[0180] The reference picture construction unit and motion compensation unit of the image decoding device perform the same functions as the corresponding units of the image encoding device, and therefore detailed description thereof will be omitted. The motion information decoding unit can be performed using the reverse method used in the motion information encoding unit. Here, the prediction block generated by the motion compensation unit can be transmitted to the adder.
[0181] FIG. 5 is an exemplary diagram showing various cases in which a prediction block is obtained through inter-frame prediction according to the present invention.
[0182] Referring to Figure 5, unidirectional prediction can obtain a predictive block (A. forward prediction) from previously coded reference pictures T-1 and T-2, or can obtain a predictive block (B. backward prediction) from subsequently coded reference pictures T+1 and T+2. Bidirectional prediction can generate predictive blocks (C, D) from multiple previously coded reference pictures T-2 to T+2. Generally, P picture types support unidirectional prediction, while B picture types support bidirectional prediction.
[0183] As in the above example, the pictures referenced for encoding the current picture can be obtained from memory, and a reference picture list can be constructed based on the current picture T, including reference pictures that are before the current picture in time order or display order and reference pictures that are after the current picture in time order or display order.
[0184] Inter prediction E can be performed on the current image as well as on previous or subsequent images based on the current image. Performing inter prediction on the current image can be called non-directional prediction. This can be supported in I image types or P / B image types. The supported image types can be determined according to encoding settings. Performing inter prediction on the current image generates a prediction block using spatial correlation, which is different from performing inter prediction on another image for the purpose of utilizing temporal correlation, and the prediction method (e.g., reference image, motion vector, etc.) can be the same.
[0185] Here, it is assumed that P and B pictures are image types capable of performing inter prediction, but the present invention is applicable to various additional or alternative image types. For example, a certain image type may not support intra prediction but may support only inter prediction, may support only inter prediction in a certain direction (backward), or may support only inter prediction in a certain direction.
[0186] The reference picture construction unit can construct and manage reference pictures used for encoding the current picture through a reference picture list. At least one reference picture list can be constructed according to encoding settings (e.g., image type, prediction direction, etc.), and a prediction block can be generated from the reference pictures included in the reference picture list.
[0187] In the case of unidirectional prediction, inter prediction can be performed using at least one reference picture included in reference picture list 0 (L0) or reference picture list 1 (L1). In addition, in the case of bidirectional prediction, inter prediction can be performed using at least one reference picture included in a composite list LC generated by combining L0 and L1.
[0188] For example, unidirectional prediction can be divided into forward prediction Pred_L0 using forward reference picture list L0 and backward prediction Pred_L1 using backward reference picture list L1. Bidirectional prediction Pred_BI can use both forward reference picture list L0 and backward reference picture list L1.
[0189] Alternatively, bidirectional prediction may also include copying a forward reference picture list L0 to a backward reference picture list L1 to perform two or more forward predictions, and bidirectional prediction may also include copying a backward reference picture list L1 to a forward reference picture list L0 to perform two or more backward predictions.
[0190] The prediction direction can be represented by flag information indicating the direction (e.g., inter_pred_idc, which is assumed to be adjustable by predFlagL0, predFlagL1, and predFlagBI). predFlagL0 indicates whether or not forward prediction is performed, and predFlagL1 indicates whether or not backward prediction is performed. Bidirectional prediction can be indicated by indicating the prediction status via predFlagBI, or by predFlagL0 and predFlagL1 being simultaneously activated (e.g., when each flag is 1).
[0191] Although the present invention will be described mainly in terms of a case where forward prediction using a forward reference picture list and unidirectional prediction are used, the present invention can be applied in the same manner or with modifications to other cases.
[0192] Generally, a method can be used in which an encoder determines an optimal reference picture for a picture to be coded and explicitly transmits information about the reference picture to a decoder. For this reason, a reference picture configuration unit can manage a picture list referenced for inter prediction of a current picture and set rules for reference picture management in consideration of limited memory size.
[0193] The transmitted information can be defined as an RPS (Reference Picture Set), and pictures selected in the RPS are classified as reference pictures and stored in memory (or DPB), while pictures not selected in the RPS are classified as non-reference pictures and removed from memory after a certain period of time. The memory can store a preset number of pictures (e.g., 14, 15, 16 pictures or more), and the size of the memory can be set according to the level and image resolution.
[0194] FIG. 6 is an example diagram illustrating how to configure a reference picture list according to an embodiment of the present invention.
[0195] 6, generally, reference pictures T-1 and T-2 existing before the current picture are allocated to L0, and reference pictures T+1 and T+2 existing after the current picture are allocated to L1 and managed. When constructing L0, if the number of reference pictures allowed for L0 is not met, reference pictures from L1 can be allocated. Similarly, when constructing L1, if the number of reference pictures allowed for L1 is not met, reference pictures from L0 can be allocated.
[0196] The current picture may also be included in at least one reference picture list. For example, the current picture may be included in L0 or L1, and L0 may be constructed by adding a reference picture (or the current picture) whose temporal order is T to the reference pictures before the current picture, and L1 may be constructed by adding a reference picture whose temporal order is T to the reference pictures after the current picture.
[0197] The configuration of the reference picture list can be determined depending on the coding settings.
[0198] The current picture may not be included in the reference picture list and may be managed via a separate memory separate from the reference picture list, or the current picture may be included in at least one reference picture list and managed.
[0199] For example, it can be determined by a signal (curr_pic_ref_enabled_flag) indicating whether the current picture is included in the reference picture list, where the signal can be information that is implicitly determined or explicitly generated.
[0200] In detail, when the signal is deactivated (e.g., curr_pic_ref_enabled_flag=0), the current picture is not included as a reference picture in any reference picture list, and when the signal is activated (e.g., curr_pic_ref_enabled_flag=1), whether the current picture is included in a given reference picture list is determined implicitly (e.g., added only to L0, added only to L1, or can be added to both L0 and L1 at the same time) or explicitly by generating related signals (e.g., curr_pic_ref_from_l0_flag, curr_pic_ref_from_l1_flag). The signals can be supported in units of sequence, picture, subpicture, slice, tile, brick, etc.
[0201] Here, the current picture may be located at the first or last position in the reference picture list as shown in Fig. 6, and the arrangement order in the list may be determined according to encoding settings (e.g., image type information, etc.). For example, the current picture may be located at the first position for an I type and at the last position for a P / B type, but is not limited thereto, and other variations are possible.
[0202] Alternatively, a separate reference picture memory can be supported depending on a signal (ibc_enabled_flag) indicating whether block matching (or template matching) is supported in the current picture, where the signal can be information that is implicitly determined or explicitly generated.
[0203] In particular, when the signal is inactivated (e.g., ibc_enabled_flag=0), it means that block matching is not supported in the current picture, and when the signal is activated (e.g., ibc_enabled_flag=1), block matching is supported in the current picture and a reference picture memory for this purpose can be supported. In this example, it is assumed that additional memory is provided, but it is also possible to set block matching to be directly supported in an existing memory supported for the current picture without providing additional memory.
[0204] The reference picture construction unit may include a reference picture interpolation unit, and may determine whether to perform an interpolation process for fractional pixels according to the interpolation accuracy of the inter prediction. For example, if the interpolation accuracy is integer-based, the reference picture interpolation process may be omitted, and if the interpolation accuracy is fractional, the reference picture interpolation process may be performed.
[0205] The interpolation filter used in the reference picture interpolation process may be implicitly determined according to encoding settings, or may be explicitly determined from among multiple interpolation filters. The configuration of the multiple interpolation filters may support fixed candidates or adaptive candidates according to encoding settings, and the number of candidates may be 2, 3, 4, or an integer greater than or equal to 2. The explicitly determined unit may be determined from among a sequence, a picture, a subpicture, a slice, a tile, a rebrick, a block, etc.
[0206] Here, the encoding settings may be determined by an image type, color components, state information of a current block (e.g., block size, shape, horizontal / vertical length ratio, etc.), inter-prediction settings (e.g., motion information encoding mode, motion model selection information, motion vector precision selection information, reference picture, reference direction, etc.), etc. The motion vector precision may refer to the precision of a motion vector (i.e., pmv+mvd), but may also be replaced with the precision of a motion vector predicted value pmv or a motion vector differential value mvd.
[0207] Here, the interpolation filter may have a filter length of k-tap, where k may be an integer of 2, 3, 4, 5, 6, 7, 8, or greater. The filter coefficients may be derived from mathematical expressions having various coefficient characteristics, such as a Wiener filter or a Kalman filter. Filter information (e.g., filter coefficients, tap information, etc.) used for interpolation may be implicitly determined or derived, or related information may be explicitly generated. In this case, the filter coefficients may be configured to include 0.
[0208] The interpolation filter can be applied to a predetermined pixel unit, which can be limited to integer or fractional units, or can be applied to integer and fractional units.
[0209] For example, an interpolation filter can be applied to k integer unit pixels horizontally or vertically adjacent to a pixel to be interpolated (that is, a fractional unit pixel).
[0210] Alternatively, an interpolation filter can be applied to p integer unit pixels and q fractional unit pixels (p+q=k) adjacent to the target pixel in the horizontal or vertical direction. In this case, the precision of the fractional unit referenced for interpolation (for example, 1 / 4 unit) can be expressed with the same precision as the target pixel or lower precision (for example, 2 / 4 → 1 / 2).
[0211] If only one of the x and y components of the pixel to be interpolated is located in decimal units, interpolation can be performed based on k pixels adjacent in the component direction of the decimal unit (for example, the x axis is horizontal and the y axis is vertical). If both the x and y components of the pixel to be interpolated are located in decimal units, linear interpolation can be performed based on x pixels adjacent in either the horizontal or vertical direction, and then interpolation can be performed based on y pixels adjacent in the remaining direction. In this case, x and y are the same as k, but this is not limited to the case, and x and y may be different.
[0212] The interpolation target pixel can be obtained with an interpolation accuracy (1 / m), where m can be an integer of 1, 2, 4, 8, 16, 32, or more. The interpolation accuracy can be implicitly determined according to encoding settings, or related information can be explicitly generated. The encoding settings can be defined based on an image type, color components, reference pictures, motion information encoding mode, motion model selection information, etc. The explicitly determined unit can be determined from among a sequence, a subpicture, a slice, a tile, a brick, etc.
[0213] Through the above process, an interpolation filter setting for the current block is obtained, and reference picture interpolation can be performed based on the obtained setting. Furthermore, a detailed interpolation filter setting based on the interpolation filter setting is obtained, and reference picture interpolation can be performed based on the detailed interpolation filter setting. That is, it is assumed that the obtained interpolation filter setting for the current block may be a single fixed candidate or may be a plurality of candidates available for use. In the example described below, it is assumed that the interpolation accuracy is 1 / 16 (e.g., 15 pixels to be interpolated).
[0214] As an example of the detailed interpolation filter setting, one of a plurality of candidates can be adaptively used for a first pixel unit at a predetermined position, and one predetermined interpolation filter can be used for a second pixel unit at a predetermined position.
[0215] The first pixel unit can be defined among the total number of decimal pixels (in this example, 1 / 16 to 15 / 16) supported by the interpolation accuracy, and the number of pixels included in the first pixel unit is a, where a can be defined between 0, 1, 2, ... (m-1).
[0216] The second pixel unit may include pixels obtained by subtracting the first pixel unit from the pixels of the entire decimal unit, and the number of pixels included in the second pixel unit may be derived by subtracting the pixels of the first pixel unit from the total number of pixels to be interpolated. In this example, the pixel unit is divided into two, but the present invention is not limited thereto and may be divided into three or more.
[0217] For example, the first pixel unit may be a unit expressed in multiples of 1 / 2, 1 / 4, 1 / 8, etc. As an example, the first pixel unit may be a pixel at 8 / 16 in the case of 1 / 2 units, {4 / 16, 8 / 16, 12 / 16} in the case of 1 / 4 units, and {2 / 16, 4 / 16, 6 / 16, 8 / 16, 10 / 16, 12 / 16, 14 / 16} in the case of 1 / 8 units.
[0218] The setting of the fine interpolation filter can be determined according to the coding setting, which can be defined by the image type, color components, state information of the current block, inter prediction setting, etc. Below, examples of the fine interpolation filter setting according to various coding elements will be considered. For convenience of explanation, it is assumed that a is 0 and a is equal to or greater than c (c is an integer equal to or greater than 1).
[0219] For example, in the case of color components (luminance, chrominance), a can be (1,0), (0,1), or (1,1). Or, in the case of motion information coding modes (merged mode, competitive mode), a can be (1,0), (0,1), (1,1), or (1,3). Or, in the case of motion model selection information (movement motion, non-movement motion A, non-movement motion B), a can be (0,1,1), (1,0,0), (1,1,1), or (1,3,7). Or, in the case of motion vector precision (1 / 2, 1 / 4, 1 / 8), a can be (0,0,1), (0,1,0), (1,0,0), (0,1,1), (1,0,1), (1,1,0), or (1,1,1). Or, in the case of reference pictures (current picture, other pictures), it can be (0,1), (1,0), (1,1).
[0220] As described above, the interpolation process can be performed by selecting one of multiple interpolation accuracies, and if an interpolation process based on adaptive interpolation accuracy is supported (e.g., if adaptive_ref_resolution_enabled_flag.0, a predetermined interpolation accuracy is used, and if 1, one of multiple interpolation accuracies is used), accuracy selection information (e.g., ref_resolution_idx) can be generated.
[0221] The motion estimation and compensation processes can be performed according to the interpolation accuracy, and the representation unit and storage unit for the motion vectors can also be determined based on the interpolation accuracy.
[0222] For example, if the interpolation accuracy is 1 / 2 unit, the motion estimation and compensation process is performed in 1 / 2 unit, and the motion vector is expressed in 1 / 2 unit and can be used in the encoding process. Also, the motion vector is stored in 1 / 2 unit and can be referenced in the motion information encoding process of other blocks.
[0223] Alternatively, if the interpolation accuracy is 1 / 8 unit, the motion estimation and compensation process is performed in 1 / 8 unit, and the motion vector is expressed in 1 / 8 unit, can be used in the encoding process, and can be stored in 1 / 8 unit.
[0224] In addition, the motion estimation and compensation process and motion vectors can be performed, represented, and stored in units different from the interpolation accuracy, such as integer, 1 / 2, or 1 / 4 units, which can be adaptively determined depending on the inter-frame prediction method / settings (e.g., motion estimation / compensation method, motion model selection information, motion information coding mode, etc.).
[0225] For example, assuming that the interpolation accuracy is 1 / 8, in the case of a motion model, the motion estimation and compensation process is performed in 1 / 4 units, and the motion vector is expressed in 1 / 4 units (this example assumes units in the encoding process) and can be stored in 1 / 8 units. In the case of a non-motion model, the motion estimation and compensation process is performed in 1 / 8 units, and the motion vector is expressed in 1 / 4 units and can be stored in 1 / 8 units.
[0226] For example, assuming that the interpolation accuracy is 1 / 8, in the case of block matching, the motion estimation and compensation process is performed in 1 / 4 units, and the motion vector is expressed in 1 / 4 units and can be stored in 1 / 8 units. In the case of template matching, the motion estimation and compensation process is performed in 1 / 8 units, and the motion vector is expressed in 1 / 8 units and can be stored in 1 / 8 units.
[0227] For example, assuming that the interpolation accuracy is 1 / 16, in competitive mode, the motion estimation and compensation process is performed in 1 / 4 units, and motion vectors are expressed in 1 / 4 units and can be stored in 1 / 16 units. In merge mode, the motion estimation and compensation process is performed in 1 / 8 units, and motion vectors are expressed in 1 / 4 units and can be stored in 1 / 16 units. In skip mode, the motion estimation and compensation process is performed in 1 / 16 units, and motion vectors are expressed in 1 / 4 units and can be stored in 1 / 16 units.
[0228] In summary, motion estimation and compensation, and motion vector representation and storage units can be adaptively determined based on an inter-prediction method or setting and interpolation accuracy. In particular, motion estimation and compensation and motion vector representation units can be adaptively determined according to an inter-prediction method or setting, and the motion vector storage unit can generally be determined according to interpolation accuracy, but this is not limited to this and various modified examples are possible. In addition, although the above example is based on one category (e.g., motion model selection information, motion estimation / compensation method, etc.), the settings can also be determined by mixing two or more categories.
[0229] Also, as described above, the interpolation accuracy information has a predetermined value or is selected from among a plurality of accuracies, but the reference picture interpolation accuracy can be determined based on a motion estimation and compensation setting supported according to an inter-frame prediction method or setting. For example, when a motion model supports up to 1 / 8 units and a non-motion model supports up to 1 / 16 units, an interpolation process can be performed according to the accuracy unit of the non-motion motion model having the highest accuracy.
[0230] That is, reference picture interpolation may be performed according to settings for supported precision information such as a motion model, a non-motion model, a competitive mode, a merge mode, a skip mode, etc. In this case, the precision information may be implicitly or explicitly determined, and when related information is explicitly generated, it may be included in units such as a sequence, a picture, a sub-picture, a slice, a tile, or a brick.
[0231] The motion estimation unit estimates (or searches) whether a target block has a high correlation with a predetermined block of a predetermined reference picture. The size and shape (M×N) of the target block to be predicted can be obtained from the block division unit. For example, the target block can be determined to be in the range of 4×4 to 128×128. Inter prediction is generally performed in units of prediction blocks, but can also be performed in units of coding blocks, transform blocks, etc. depending on the setting of the block division unit. Estimation is performed within an estimable range of the reference region, and at least one motion estimation method can be used. The motion estimation method can define the pixel-by-pixel estimation order and conditions.
[0232] Motion estimation may be performed based on a motion estimation method. For example, the area to be compared for the motion estimation process may be a target block in the case of block matching, or a predetermined area (template) set around the target block in the case of template matching. In the former case, a block having the highest correlation within an estimable range between the target block and the reference area may be found, and in the latter case, an area having the highest correlation within an estimable range between the template defined according to encoding settings and the reference area may be found.
[0233] Motion estimation may be performed based on a motion model. In addition to a translation motion model that considers only translation, motion estimation and compensation may be performed using additional motion models. For example, motion estimation and compensation may be performed using a motion model that considers not only translation but also rotation, perspective, zoom-in / zoom-out, and other motions. This can help improve coding performance by generating a prediction block that reflects the various types of motion that occur according to the regional characteristics of an image.
[0234] FIG. 7 is a conceptual diagram illustrating a non-movement motion model according to one embodiment of the present invention.
[0235] 7, an example of expressing motion information based on motion vectors V0 and V1 at a predetermined position is shown as an example of an affine model. Since motion can be expressed based on multiple motion vectors, accurate motion estimation and compensation are possible.
[0236] As in the above example, inter prediction is performed based on a predefined motion model, but inter prediction based on an additional motion model can also be supported. Here, it is assumed that the predefined motion model is a translation motion model and the additional motion model is an affine model, but the present invention is not limited to this and various modifications are possible.
[0237] In the case of a moving motion model, motion information (assuming one-way prediction) can be expressed based on one motion vector, and the control point (reference point) for indicating the motion information is assumed to be the upper left coordinate, but is not limited to this.
[0238] In the case of a motion model other than motion, it can be expressed by various configurations of motion information. In this example, a configuration is assumed in which one motion vector (based on the upper left coordinate) is expressed by additional information. Some motion estimation and compensation mentioned in the examples below may not be performed in units of blocks, but may be performed in units of predetermined sub-blocks. In this case, the size and position of the predetermined sub-block may be determined based on each motion model.
[0239] 8 is an exemplary diagram illustrating sub-block-based motion estimation according to an embodiment of the present invention, specifically, sub-block-based motion estimation using an affine model (two motion vectors).
[0240] In the case of a translation motion model, the motion vectors of the pixels included in the target block can be the same, i.e., a motion vector can be applied to each pixel at once, and motion estimation and compensation can be performed using a single motion vector V0.
[0241] In the case of a non-movement motion model (affine model), the pixel-by-pixel motion vectors included in the target block do not need to be the same, and individual pixel-by-pixel motion vectors may be required. In this case, pixel-by-pixel or sub-block-by-sub-block motion vectors can be derived based on the motion vectors V0 and V1 at predetermined control point positions of the target block, and motion estimation and compensation can be performed using the derived motion vectors.
[0242] For example, a motion vector in a sub-block or pixel unit within a target block {e.g., (V x , V y )} is V x =(V 1x -V 0x )×x / M-(V 1y -V 0y )×y / N+V 0x , V y =(V 1y -V 0y )×x / M+(V 1x -V 0x )×y / N+V 0y In the above formula, V0 {in this example, (V 0x , V 0y )} means the motion vector on the upper left side of the target block, and V1 {in this example, (V 1x , V 1y )} means the motion vector at the upper right side of the target block. In consideration of complexity, motion estimation and motion compensation of the non-motion motion model can be performed on a sub-block basis.
[0243] Here, the size of the sub-block (M×N) can be determined according to coding settings and can be set to a fixed size or an adaptive size. Here, M and N can be integers of 2, 4, 8, 16, or more, and M and N may be the same or different. The size of the sub-block can be explicitly generated in units of a sequence, a picture, a sub-picture, a slice, a tile, a brick, etc. Alternatively, it can be implicitly determined by a common agreement between an encoder and a decoder, or can be determined by coding settings.
[0244] Here, the coding settings can be defined by one or more elements such as the state information of the target block, the image type, the color components, and the inter-frame prediction setting information (motion information coding mode, reference picture information, interpolation accuracy, motion model selection information, etc.).
[0245] The above example describes a process of deriving the size of a sub-block according to a predetermined non-motion motion model and performing motion estimation and compensation based on the size of the sub-block. As in the above example, motion estimation and compensation can be performed in sub-block or pixel units according to a motion model, and detailed description thereof will be omitted.
[0246] Below, various examples of motion information configured according to a motion model are considered.
[0247] For example, in the case of a motion model representing rotational motion, a translational motion of a block can be represented by one motion vector, and the rotational motion can be represented by rotation angle information. The rotation angle information can be measured with a predetermined position (e.g., the upper left coordinate) as the reference (0 degrees), and can be represented by k candidates (k is an integer of 1, 2, 3, or more) having predetermined intervals (e.g., angle difference values of 0 degrees, 11.25 degrees, 22.25 degrees, etc.) within a predetermined angle range (e.g., between -90 degrees and 90 degrees).
[0248] Here, the rotation angle information can be coded by itself during the motion information coding process, or can be coded (for example, prediction+difference value information) based on motion information (for example, motion vectors, rotation angle information) of neighboring blocks.
[0249] Alternatively, one motion vector can represent the translational motion of a block, and one or more additional motion vectors can represent the rotational motion of the block, where the number of additional motion vectors can be 1, 2, or an integer greater than 1, and the control point of the additional motion vector can be determined from the upper right, lower left, and lower right coordinates, or other coordinates within the block can be set as the control point.
[0250] Here, the additional motion vector can be coded by itself during the motion information coding process, or can be coded (e.g., prediction + difference value information) based on the motion information of an adjacent block (e.g., motion vectors based on a motion model or a non-motion model), or can be coded (e.g., prediction + difference value information) based on other motion vectors within the block that represent rotational motion.
[0251] For example, in the case of a motion model that represents a scaling motion such as a zoom-in / zoom-out situation, a block movement motion can be represented by one motion vector, and the scaling motion can be represented by scaling information that indicates expansion or contraction in the horizontal or vertical direction based on a predetermined position (e.g., the coordinates of the upper left corner).
[0252] Here, scaling can be applied in at least one direction of horizontal or vertical. Individual scaling information applied to the horizontal or vertical directions can be supported, or scaling information applied commonly can be supported. The width and height of the scaled block can be added to a predetermined position (the coordinates of the upper left corner) to determine the position for motion estimation and compensation.
[0253] Here, the scaling information can be coded by itself during the motion information coding process, or can be coded (for example, prediction+difference value information) based on motion information (for example, motion vectors, scaling information) of neighboring blocks.
[0254] Alternatively, one motion vector can represent the translation of a block, and one or more additional motion vectors can represent the size adjustment of the block, where the number of additional motion vectors can be 1, 2, or an integer greater than 1, and the control points of the additional motion vectors can be determined among the coordinates of the upper right, lower left, and lower right, or other coordinates within the block can be set as the control points.
[0255] Here, the additional motion vector can be coded by itself during the motion information coding process, or can be coded (e.g., prediction + difference value information) based on the motion information of an adjacent block (e.g., a motion vector based on a motion model or a non-motion model), or can be coded (e.g., prediction + difference value) based on a predetermined coordinate within the block (e.g., the coordinate on the lower right side).
[0256] Although the above example has been described with respect to an expression for showing a part of a movement, it is also possible to express it using movement information for expressing a plurality of movements.
[0257] For example, in the case of a motion model that expresses various and complex motions, a translation motion of a block can be expressed by one motion vector, a rotation motion can be expressed by rotation angle information, and a size adjustment can be expressed by scaling information. The description of each motion can be guided by the above examples, so a detailed description will be omitted.
[0258] Alternatively, one motion vector can represent the translational motion of a block, and one or more additional motion vectors can represent other motions of the block, where the number of additional motion vectors can be one, two, or an integer greater than one, and the control points of the additional motion vectors can be determined among the coordinates of the upper right, lower left, and lower right, or coordinates within other blocks can be set as the control points.
[0259] Here, the additional motion vector can be coded by itself during the motion information coding process, or can be coded (e.g., prediction + difference value information) based on the motion information of adjacent blocks (e.g., motion vectors based on a motion model or a non-motion model), or can be coded (e.g., prediction + difference value information) based on other motion vectors within the block that represent various motions.
[0260] The above description may be about an affine model, and will be focused on the case where there are one or two additional motion vectors. In summary, it is assumed that the number of motion vectors used by a motion model is one, two, or three, and that each motion model can be considered as an individual motion model depending on the number of motion vectors used to represent motion information. It is also assumed that when there is one motion vector, it is a predefined motion model.
[0261] Multiple motion models for inter prediction may be supported, and support for an additional motion model may be determined by a signal (e.g., adaptive_motion_mode_enabled_flag) indicating support for an additional motion model. Here, if the signal is 0, a predefined motion model is supported, and if the signal is 1, multiple motion models may be supported. The signal may be generated in units of a sequence, picture, subpicture, slice, tile, brick, block, etc., and if it is not possible to separately determine the motion model, a signal value may be assigned according to a predefined setting. Alternatively, whether or not to support the motion model may be implicitly determined according to the encoding setting. Alternatively, whether the motion model is implicit or explicit may be determined according to the encoding setting. Here, the encoding setting may be defined by one or more elements such as an image type, an image type (e.g., 0 for a general image, 1 for a 360-degree image), color components, etc.
[0262] Through the above process, it can be determined whether multiple motion models are supported. Hereinafter, it is assumed that two or more additional motion models are supported, and that it has been determined that multiple motion models are supported in units of a sequence, a picture, a subpicture, a slice, a tile, a brick, etc. However, some exceptions may exist. In the example described below, it is assumed that motion models A, B, and C are supportable, with A being a motion model that is basically supported and B and C being motion models that can be additionally supported.
[0263] Configuration information for supported motion models can be generated in the above units, i.e., supported motion models such as {A, B}, {A, C}, and {A, B, C} can be configured.
[0264] For example, the configuration candidates can be assigned indices (0 to 2) to select them. If index 2 is selected, {A, C} can be determined as the supported motion model configuration, and if index 3 is selected, {A, B, C} can be determined as the supported motion model configuration.
[0265] Alternatively, information indicating whether a specific motion model is supported can be supported individually. That is, a support flag for B and a support flag for C can be generated. If both flags are 0, it may be that only A is supported. This example may be an example in which processing is performed without generating information indicating whether multiple motion models are supported.
[0266] Once a set of supported motion model candidates is constructed as in the above example, one of the motion models from the set of candidates can be explicitly determined and used on a block-by-block basis, or can be implicitly used.
[0267] Generally, the motion estimation unit may be a component present in the encoding device, but may also be a component that can be included in the decoding device according to a prediction method (e.g., template matching, etc.). For example, in the case of template matching, a decoder can perform motion estimation using a neighboring template of a current block to obtain motion information of the current block. In this case, motion estimation-related information (e.g., motion estimation range, motion estimation method (scan order), etc.) can be implicitly determined or explicitly generated and can be included in units such as a sequence, picture, subpicture, slice, tile, brick, etc.
[0268] The motion compensation unit refers to a process for obtaining data of a part of a block of a predetermined reference picture determined through a motion estimation process as a prediction block of a current block. Specifically, the motion compensation unit can generate a prediction block of a current block from at least one region (or block) of at least one reference picture based on motion information (e.g., reference picture information, motion vector information, etc.) obtained through the motion estimation process.
[0269] Motion compensation can be performed based on the motion compensation method as follows.
[0270] In the case of block matching, the motion vector (V x , V y ) and the coordinates of the upper left corner of the target block (P x , P y ) x +V x , P y +V y ) as a reference, data in the corresponding area M to the right and N below can be compensated for with the predicted block of the target block.
[0271] In the case of template matching, the motion vector (V x , V y ) and the coordinates of the upper left corner of the target block (P x , P y ) x +V x , P y +V y ) as a reference, data in the corresponding area M to the right and N below can be compensated for with the predicted block of the target block.
[0272] Also, motion compensation can be performed based on a motion model as follows.
[0273] In the case of the motion model, one of the motion vectors (V x , V y ) and the coordinates of the upper left corner of the target block (P x , P y ) x +V x , P y +V y ) as a reference, data in the corresponding area M to the right and N below can be compensated for with the predicted block of the target block.
[0274] For non-motion motion models, multiple motion vectors (V 0x , V 0y ), (V 1x , V 1y ) are implicitly obtained through the motion vectors (V mx , V ny ) and the coordinates of the upper left corner of each sub-block (P mx , P ny ) mx +, V nx , P my +, V ny ), the data in the corresponding area M / m to the right and N / n below can be compensated for with the predicted block of the sub-block. That is, the predicted blocks of the sub-blocks can be collected and compensated for with the predicted block of the target block.
[0275] The motion information determination unit may perform a process for selecting optimal motion information for a target block. Generally, optimal motion information in terms of coding cost may be determined using a rate-distortion technique that takes into account block distortion (e.g., distortion between a target block and a reconstructed block, such as SAD (Sum of Absolute Difference) or SSD (Sum of Square Difference)) and the amount of bits generated by the motion information. A predicted block generated based on the motion information determined through the above process may be transmitted to a subtraction unit and an addition unit. The motion information determination unit may also be configured to be included in a decoding device according to some prediction methods (e.g., template matching), in which case the motion information may be determined based on block distortion.
[0276] The motion information determiner may consider inter-frame prediction-related setting information such as a motion compensation method, a motion model, etc. For example, when multiple motion compensation methods are supported, motion compensation method selection information and the corresponding motion vectors, reference picture information, etc. may be the optimal motion information. Alternatively, when multiple motion models are supported, motion model selection information and the corresponding motion vectors, reference picture information, etc. may be the optimal motion information.
[0277] The motion information encoding unit may encode the motion information of the target block obtained through the motion information determination process. In this case, the motion information may be composed of information on an image and an area referenced for predicting the target block. In particular, the motion information may be composed of information on the referenced image (e.g., reference image information) and information on the referenced area (e.g., motion vector information).
[0278] In addition, inter-prediction related setting information (or selection information, such as a motion estimation / compensation method, motion model selection information, etc.) may be included in the motion information of the current block. Based on the inter-prediction related setting, information on the reference image and region (e.g., the number of motion vectors, etc.) may be configured.
[0279] The information on the reference image and the reference region can be configured as one combination to encode the motion information, and the combination of the information on the reference image and the reference region can be configured as a motion information encoding mode.
[0280] Here, information about the reference image and the reference region can be obtained based on neighboring blocks or predetermined information (for example, images coded before or after the current picture, a zero motion vector, etc.).
[0281] The adjacent block belongs to the same space as the target block and is the block closest to the target block.<inter_blk_A> , belong to the same space as the target block and are far adjacent blocks.<inter_blk_B> , and blocks that belong to a space that is not the same as the target block<inter_blk_C> The information about the reference image and the reference region can be obtained based on neighboring blocks (candidate blocks) that belong to at least one of these categories.
[0282] For example, the motion information of the target block can be coded based on the motion information of the candidate block or the reference picture information, or the motion information of the target block can be coded based on information derived from the motion information of the candidate block or the reference picture information (or information obtained through intermediate values, transformation processes, etc.). In other words, the motion information of the target block can be predicted from the candidate block and information about it can be coded.
[0283] In the present invention, motion information of a current block may be coded based on one or more motion information coding modes, which may be defined in various ways and may include one or more of a skip mode, a merge mode, a competition mode, etc.
[0284] Based on the above-mentioned template matching tmp, it may be combined with the motion information coding mode, or may be supported in a separate motion information coding mode, or may be included in all or some of the detailed configurations of the motion information coding mode. This is based on the assumption that template matching is defined to be supported in a higher unit (e.g., picture, sub-picture, slice, etc.), and a flag indicating whether or not template matching is supported may be considered as a factor in setting inter prediction.
[0285] Based on the above-described method ibc for performing block matching in the current picture, it may be combined with the motion information coding mode, supported as a separate motion information coding mode, or included in all or some of the detailed configurations of the motion information coding mode. This assumes that block matching within the current picture is determined to be supported in a higher unit, and a flag indicating whether it is supported may be considered as a factor in setting inter prediction.
[0286] Based on the above-mentioned motion model (affine), it may be combined with the motion information coding mode, or may be supported as a separate motion information coding mode, or may be included in all or some of the detailed configurations of the motion information coding mode. This assumes that a motion model other than motion is defined to be supported in a higher unit, and a flag indicating whether it is supported or not may be considered as a factor in setting inter prediction.
[0287] For example, individual motion information coding modes such as temp_inter, temp_tmp, temp_ibc, and temp_affine may be supported. Alternatively, combined motion information coding modes such as temp_inter_tmp, temp_inter_ibc, temp_inter_affine, and temp_inter_tmp_ibc may be supported. Alternatively, the motion information prediction candidates constituting temp may include template-based candidates, candidates based on a method of performing block matching within the current picture, and affine-based candidates.
[0288] Here, "temp" can refer to skip mode (skip), merge mode (merge), or competition mode (comp). For example, in the case of skip mode, motion information coding modes such as "skip_inter", "skip_tmp", "skip_ibc", and "skip_affine" are supported; in the case of merge mode, "merge_inter", "merge_tmp", "merge_ibc", and "merge_affine" are supported; and in the case of competition mode, "comp_inter", "comp_tmp", "comp_ibc", and "comp_affine" are supported.
[0289] If a skip mode, a merge mode, and a competitive mode are supported and a group of motion information prediction candidates for each mode includes candidates that take the above factors into consideration, one mode can be selected according to a flag distinguishing the skip mode, the merge mode, and the competitive mode. For example, if a flag indicating whether or not a skip mode is supported and has a value of 1, the skip mode is selected; if a flag indicating whether or not a merge mode is supported and has a value of 0, the merge mode is selected; and if the flag indicates whether or not a merge mode is supported and has a value of 1, the merge mode is selected; and if the flag has a value of 0, the competitive mode is selected. In addition, the group of motion information prediction candidates for each mode may include candidates based on inter, tmp, ibc, and affine.
[0290] Alternatively, when a plurality of motion information coding modes are supported under one common mode, in addition to a flag for selecting one of the skip mode, merge mode, and competition mode, an additional flag for distinguishing a detailed mode of the selected mode may be supported. For example, when the merge mode is selected, a flag for selecting from merge_inter, merge_tmp, merge_ibc, merge_affine, etc., which are detailed modes related to the merge mode, may be additionally supported. Alternatively, a flag indicating whether or not it is merge_inter may be supported, and if it is not merge_inter, a flag for selecting from merge_tmp, merge_ibc, merge_affine, etc. may be additionally supported.
[0291] All or some of the motion information coding mode candidates may be supported depending on the coding setting, which may be defined by one or more factors such as state information of the current block, image type, image category, color components, inter-prediction support setting (e.g., whether template matching is supported, whether block matching is supported in the current picture, non-motion motion model support factor, etc.).
[0292] For example, supported motion information coding modes may be determined according to block size. In this case, a supported range of block size may be determined according to a first threshold size (minimum value) or a second threshold size (maximum value), and each threshold size may be expressed as W, H, W×H, or W*H, where W is the width (W) and height (H) of the block. For the first threshold size, W and H may be integers of 4, 8, 16, or more, and W*H may be an integer of 16, 32, 64, or more. For the second threshold size, W and H may be integers of 16, 32, 64, or more, and W*H may be an integer of 64, 128, 256, or more. The above range may be determined by either the first threshold size or the second threshold size, or by both.
[0293] At this time, the magnitude of the threshold may be fixed or adaptive depending on the image (e.g., image type, etc.), where the magnitude of the first threshold may be set based on the size of the smallest coding block, the smallest prediction block, the smallest transform block, etc., and the magnitude of the second threshold may be set based on the size of the largest coding block, the largest prediction block, the largest transform block, etc.
[0294] For example, supported motion information coding modes may be determined according to the image type. Here, the I-image type may include at least one of skip mode, merge mode, and competitive mode. Here, a method of performing block matching (or template matching) on the current picture, an individual motion information coding mode related to an affine model (hereinafter referred to as "element"), or a motion information coding mode in which two or more elements are combined may be supported. Alternatively, other elements may be configured as motion information prediction candidates in a predetermined motion information coding mode.
[0295] The P / B picture type may include at least one of a skip mode, a merge mode, and a competitive mode. In this case, individual motion information coding modes related to general inter-picture prediction, template matching, block matching in the current picture, and affine models (hereinafter referred to as elements) may be supported, or a motion information coding mode in which two or more elements are combined may be supported. Alternatively, other elements may be configured as motion information prediction candidates in a predetermined motion information coding mode.
[0296] FIG. 9 is a flowchart showing the coding of motion information according to an embodiment of the present invention.
[0297] Referring to FIG. 9, the motion information coding mode of the current block can be checked (S900).
[0298] A motion information coding mode can be defined by combining and setting predetermined information (e.g., motion information, etc.) used in inter-frame prediction. The predetermined information can include one or more of a predicted motion vector, a motion vector differential value, differential motion vector accuracy, reference image information, reference direction, motion model information, information on the presence or absence of a residual component, etc. The motion information coding mode can include at least one of a skip mode, a merge mode, and a competitive mode, and can also include additional modes.
[0299] The configuration and setting of explicitly generated information and implicitly defined information can be determined according to the motion information coding mode. In the case of explicit information, each piece of information can be generated individually or in a combined form (e.g., index form).
[0300] For example, the skip mode or merge mode may be defined based on (one) predetermined index information for a predicted motion vector, reference image information, and reference direction. The index may be configured as one or more candidates, and each candidate may be set based on the motion information of a predetermined block (e.g., a combination of a predicted motion vector, reference image information, and reference direction of the corresponding block). Furthermore, the motion vector difference value may be implicitly processed (e.g., a zero vector).
[0301] For example, in the competitive mode, a predicted motion vector, reference image information, and reference direction may be defined based on predetermined index information (one or more). The index may be supported individually for each piece of information, or may support a combination of two or more pieces of information. The index may be configured as one or more candidates, and each candidate may be set based on the motion information of a predetermined block (e.g., a predicted motion vector, etc.), or may be configured as a pre-set value (e.g., the distance from the current picture may be set to 1, 2, 3, etc., and the reference direction information may be set to L0, L1, etc.) (e.g., reference image information, reference direction information, etc.). A motion vector differential value may be explicitly generated, and differential motion vector accuracy information may be additionally generated according to the motion vector differential value (e.g., if the motion vector differential value is not a zero vector).
[0302] For example, in skip mode, the presence or absence of residual components is implicitly processed (e.g.,<cbf_flag=0> The motion model may be implicitly processed to a predetermined value (e.g., to support a parallel motion model).
[0303] For example, in the merge mode or the competitive mode, information on the presence or absence of residual components may be explicitly generated, and a motion model may be explicitly selected. In this case, information for distinguishing motion models may be generated within one motion information coding mode, or a separate motion information coding mode may be supported.
[0304] In the above description, if one index is supported, no index selection information is generated, and if two or more indexes are supported, index selection information can be generated.
[0305] In the present invention, a candidate group consisting of two or more indexes is called a motion information prediction candidate group. In addition, an existing motion information coding mode can be changed or a new motion information coding mode can be supported according to various predetermined information and settings.
[0306] Referring to FIG. 9, a predicted motion vector of a current block can be derived (S910).
[0307] The motion vector predictor may be set to a predefined value, or may be selected from a set of candidate motion vector predictors. In the former case, the motion vector predictor is implicitly determined, while in the latter case, index information for selecting the motion vector predictor may be explicitly generated.
[0308] Here, in the case of the previously set value, it can be set based on the motion vector of one block at a predetermined position (e.g., a block in the left, top, top-left, top-right, or bottom-left direction), or it can be set based on the motion vectors of two or more blocks, or it can be set to a default value (e.g., a zero vector).
[0309] The candidate group configuration setting of the predicted motion vector may be the same or different depending on the motion information coding mode. As an example, the skip mode / merge mode / competitive mode may support a, b, and c prediction candidates, respectively, and the number of candidates may be configured to be the same or different. In this case, a, b, and c may be integers equal to or greater than 1, such as 2, 3, 5, 6, and 7. The candidate group configuration order and configuration will be described in other embodiments. In the following example, the candidate group configuration setting of the skip mode will be described assuming the same case as that of the merge mode, but is not limited thereto, and some configurations may be different.
[0310] The predicted motion vector obtained based on the index information can be used as is to restore the motion vector of the current block, or can be adjusted based on the reference direction and the distance between the reference images.
[0311] For example, in the case of competitive mode, reference image information can be generated separately, but if the distance between the current picture and the reference picture of the target mode is different from the distance between the picture containing the predicted motion vector candidate block and the reference picture of the block, this can be adjusted to match the distance between the current picture and the reference picture of the target block.
[0312] Furthermore, the number of motion vectors derived based on the motion model may differ. That is, in addition to the motion vector at the upper left control point position, motion vectors at the upper right and lower left control point positions may be derived based on the motion model.
[0313] Referring to FIG. 9, the motion vector differential value of the current block can be restored (S920).
[0314] The motion vector difference value can be induced to zero vector value in skip mode and merge mode, and the difference value information of each component can be restored in competitive mode.
[0315] The motion vectors can be expressed based on a predetermined precision (e.g., based on interpolation precision or motion model selection information) that has already been set. Alternatively, the motion vectors can be expressed based on one of a plurality of precisions, and predetermined information for this purpose can be generated. In this case, the predetermined information can be related to the selection of the motion vector precision.
[0316] For example, when one component of a motion vector is 32 / 16, based on an implicit precision setting of 1 / 16 pixel units, motion component data for 32 can be obtained. Or, based on an explicit precision setting of 2 pixel units, the motion component can be converted to 2 / 1, and motion component data for 2 can be obtained.
[0317] The above configuration may be a method for precision processing on a motion vector obtained by adding a motion vector predictor and a motion vector differential value that have the same precision or are obtained through the same precision conversion process.
[0318] In addition, the motion vector difference value may be expressed according to a predetermined precision or one of a plurality of precisions, and in the latter case, precision selection information may be generated.
[0319] The precision may include at least one of 1 / 16, 1 / 8, 1 / 4, 1 / 2, 1, 2, and 4 pixel units, and the number of candidates may be an integer of 1, 2, 3, 4, 5, or more. Here, the supported precision candidate group configuration (e.g., classification into number, candidates, etc.) may be explicitly defined in units of sequence, picture, subpicture, slice, tile, brick, etc., or may be implicitly defined according to encoding settings. Here, the encoding settings may be defined by at least one element such as image type, reference image, color component, and motion model selection information.
[0320] Next, a case where precision candidates are configured according to various coding elements will be described, and the precision will be described assuming a description regarding motion vector differential values, but this can be similarly or identically applied to motion vectors.
[0321] For example, in the case of a mobile motion model, a candidate set can be configured such as {1 / 4, 1}, {1 / 4, 1 / 2}, {1 / 4, 1 / 2, 1}, {1 / 4, 1,4}, {1 / 4, 1 / 2, 1,4}, and in the case of a non-mobile motion model, a candidate set can be configured such as {1 / 16, 1 / 4, 1}, {1 / 16, 1 / 8, 1}, {1 / 16, 1,4}, {1 / 16, 1 / 4, 1,2}, {1 / 16, 1 / 4, 1,4}, etc. This assumes that the minimum supported accuracy is in 1 / 4 pixel units in the former case and in 1 / 16 pixel units in the latter case, and additional accuracy other than the minimum accuracy can be included in the candidate set.
[0322] Alternatively, when the reference image is a different picture, a candidate set can be configured as {1 / 4, 1 / 2, 1}, {1 / 4, 1, 4}, {1 / 4, 1 / 2, 1, 4}, etc., and when the reference image is the current picture, a candidate set can be configured as {1, 2}, {1, 4}, {1, 2, 4}, etc. In this case, the latter can be a configuration that does not perform fractional unit interpolation for block matching, etc., or can be a configuration that includes fractional unit precision candidates if fractional unit interpolation is performed.
[0323] Alternatively, if the color component is a luminance component, possible candidate sets include {1 / 4, 1 / 2, 1}, {1 / 4, 1, 4}, and {1 / 4, 1 / 2, 1, 4}, and if the color component is a chrominance component, possible candidate sets include {1 / 8, 1 / 4}, {1 / 8, 1 / 2}, {1 / 8, 1}, {1 / 8, 1 / 4, 1 / 2}, {1 / 8, 1 / 2, 2}, and {1 / 8, 1 / 4, 1 / 2, 2}, etc. In the latter case, a candidate set proportional to the luminance component can be formed according to the color component ratio (e.g., 4:2:0, 4:2:2, 4:4:4, etc.), or individual candidate sets can be formed.
[0324] Although the above example describes a case where there are multiple precision candidates for each coding element, it is also possible for there to be only one candidate configuration (ie, representation with the minimum precision that has already been set).
[0325] In summary, based on the differential motion vector accuracy in competitive mode, a motion vector differential value with minimum accuracy can be restored, and in skip mode and merge mode, a motion vector differential value with zero vector value can be derived.
[0326] Here, the number of restored motion vector differential values may differ depending on the motion model. That is, in addition to the motion vector differential value at the upper left control point position, motion vector differential values at the upper right and lower left control point positions may be further derived depending on the motion model.
[0327] Also, one differential motion vector precision can be applied to multiple motion vector differential values, or an individual differential motion vector precision can be applied to each motion vector differential value, which can be determined according to encoding settings.
[0328] Also, if at least one motion vector differential value is not 0, information on the differential motion vector accuracy may be generated, and if not, it may be omitted, but is not limited thereto.
[0329] Referring to FIG. 9, the motion vector of the current block can be restored (S930).
[0330] The motion vector of the current block can be restored by adding the predicted motion vector and the motion vector differential value obtained through the previous process of this step. In this case, if the motion vector is restored by selecting one of multiple precisions, a precision unification process can be performed.
[0331] For example, a motion vector obtained by adding a predicted motion vector and a motion vector differential value (related to motion vector precision) can be restored based on precision selection information, or a motion vector differential value (related to differential motion vector precision) can be restored based on precision selection information, and a motion vector can be obtained by adding the restored motion vector differential value and a predicted motion vector.
[0332] Referring to FIG. 9, motion compensation can be performed based on the motion information of the current block (S940).
[0333] The motion information may include reference image information, reference direction, motion model information, etc., other than the predicted motion vector and motion vector differential value described above. In the skip mode and merge mode, some of the information (e.g., reference image information, reference direction, etc.) may be implicitly determined, and in the competition mode, related information may be explicitly processed.
[0334] A prediction block can be obtained by performing motion compensation based on the motion information obtained through the above process.
[0335] Referring to FIG. 9, the residual components of the current block can be decoded (S950).
[0336] The current block can be reconstructed by adding the residual component obtained through the above process to the predicted block. In this case, the presence or absence of the residual component can be processed explicitly or implicitly depending on the motion information coding mode.
[0337] FIG. 10 is a layout diagram of a target block and its adjacent blocks according to one embodiment of the present invention.
[0338] Referring to FIG. 10, blocks adjacent to the target block in the left, upper, upper left, upper right, lower left, etc.<inter_blk_A> , and blocks adjacent to the target block in the center, left, right, top, bottom, top left, top right, bottom left, bottom right, etc. in a space (Col_Pic) that is not the same in time.<inter_blk_B> can be configured as a candidate block for prediction of motion information (eg, motion vectors) of the target mode.
[0339] The case of inter_blk_A can be an example in which the direction of adjacent blocks is determined based on an encoding order such as raster scan or z-scan, and existing directions are removed according to various scan orders, or adjacent blocks in the right, bottom, and bottom-right directions can be further configured as candidate blocks.
[0340] Referring to FIG. 10, Col_Pic can be an adjacent image before or after the current image (for example, when the interval between images is 1), and the corresponding block can be set to have the same position in the image as the target block.
[0341] Alternatively, Col_Pic can be an image in which the spacing between images is already defined based on the current image (for example, the spacing between images is z, where z is an integer of 1, 2, or 3), and the corresponding block can be set to a position moved by a predetermined displacement vector (disparity vector) from a predetermined coordinate (for example, the upper left side) of the target block, but the displacement vector can be set to a previously defined value.
[0342] Alternatively, Col_Pic can be set based on the motion information (e.g., reference image) of the neighboring blocks, and the displacement vector can be set based on the motion information (e.g., motion vector) of the neighboring blocks, and the position of the corresponding block can be determined.
[0343] In this case, k neighboring blocks can be referenced, where k can be an integer of 1, 2, or greater. If k is greater than or equal to 2, Col_Pic and the displacement vector can be obtained by calculating the maximum, minimum, median, or weighted average of the motion information (e.g., reference images or motion vectors) of the neighboring blocks. For example, the displacement vector can be set as the motion vector of the left or upper block, or as the median or average of the motion vectors of the left and lower-left blocks.
[0344] The setting of temporal candidates may be determined based on the setting of motion information configuration, etc. For example, the position of Col_Pic, the position of the corresponding block, etc. may be determined depending on whether the motion information to be included in the motion information prediction candidate group is configured in block units or sub-block units, such as block-unit motion information or sub-block-unit motion information. As an example, when sub-block-unit motion information is acquired, a block moved by a predetermined displacement vector may be set as the position of the corresponding block.
[0345] The above example shows a case where information about the location of Col_Pic and the corresponding block is implicitly determined, but explicit related information can also occur in units of sequence, picture, slice, tile group, tile, brick, etc.
[0346] Meanwhile, when a division unit partition such as a sub-picture, slice, or tile is not limited to one picture (e.g., the current picture) but is shared between images, if the position of the block corresponding to Col_Pic is different from the division unit to which the target block belongs (i.e., if it is adjacent to or outside the boundary of the division unit to which the target block belongs), it can be determined based on motion information of a successor candidate based on a predetermined priority, or can be set to be at the same position in the image as the target block.
[0347] The motion information of the spatially adjacent blocks and the temporally adjacent blocks can be included in a group of motion information prediction candidates for the target block, which are called spatial candidates and temporal candidates. The number of spatial candidates and temporal candidates supported can be a and b, where a and b can be integers from 1 to 6 or greater. In this case, a can be greater than or equal to b.
[0348] The number of spatial candidates and temporal candidates may be fixed or may be variable (eg, including 0) depending on the configuration of the preceding motion information prediction candidate.
[0349] Also, blocks that are not immediately adjacent to the target block<inter_blk_C> can be taken into consideration in constructing the candidate group.
[0350] For example, motion information of blocks located at a predetermined distance from the target block (e.g., the coordinates of the upper left corner) may be included in the candidate group. The distance may be p or q (according to each component of the motion vector), where p and q may be integers equal to or greater than 2, such as 0, 2, 4, 8, 16, or 32, and p and q may be the same or different. In other words, assuming that the coordinates of the upper left corner of the target block are (m, n), motion information of blocks including positions (m±p, n±q) may be included in the candidate group.
[0351] For example, motion information of blocks that have been coded before the target block may be included in the candidate group. In this case, a predetermined number of blocks that have been coded recently based on a predetermined coding order (e.g., raster scan, Z scan, etc.) may be considered as candidates based on the target block, and blocks with older coding order may be removed from the candidates as the coding progresses using a first-in, first-out method such as FIFO.
[0352] In the above example, this means that the candidate group includes motion information of a block that is not nearest to the target block and has already been coded based on the target mode, and therefore this is statistically referred to as a candidate.
[0353] In this case, the target block can be set as the reference block for obtaining statistical candidates, but the upper block of the target block (for example, the target block is an encoding block) <1> / predicted blocks <2> If , the upper block is the largest coded block <1> / encoding block <2> It is also possible to set k statistical candidates, where k can be an integer between 1 and 6 or greater.
[0354] The candidate group may include motion information obtained by combining motion information such as the spatial candidate, the temporal candidate, and the statistical candidate. For example, a candidate may be obtained by applying a weighted average to each component of a plurality of pieces of motion information (two, three, or more) already included in the candidate group, or a candidate may be obtained by taking a median, maximum, minimum, or the like for each component of a plurality of pieces of motion information. These candidates are called combination candidates, and r combination candidates may be supported, where r may be an integer of 1, 2, 3, or more.
[0355] If the candidates used for combination are not identical to the reference image information, the reference image information of the combination candidate can be set based on the reference image information of one of the candidates, or can be set to a previously defined value.
[0356] The statistical candidates or combination candidates may be supported in a fixed number or in a variable number depending on the configuration of preceding motion information prediction candidates.
[0357] In addition, default candidates having preset values such as (s, t) can be included in the candidate group, and a variable number (0, 1, or an integer greater than 1) can be supported depending on the configuration of the preceding motion information prediction candidate. In this case, s and t can be set to values including 0, and can also be set based on various block sizes (e.g., the horizontal or vertical length of the largest coding / prediction / transform block, the smallest coding / prediction / transform block, etc.).
[0358] A group of motion information prediction candidates may be configured from the candidates according to a motion information coding mode, and other additional candidates may be included. Also, the configuration of the group of candidates may be the same or different according to the motion information coding mode.
[0359] For example, the skip mode and merge mode may form a common candidate group, while the competitive mode may form a separate candidate group.
[0360] Here, the configuration of the candidate group can be defined by the category and position of the candidate block (e.g., determined from left / top / top-left / top-right / bottom-left directions, and the sub-block position from which motion information is obtained within the determined direction), the number of candidates (e.g., total number, maximum number per category, etc.), and the candidate configuration method (e.g., priority per category, priority within a category, etc.).
[0361] The number (total number) of motion information prediction candidate groups can be k, where k can be an integer of 1 to 6 or more. In this case, if the number of candidate groups is one, candidate group selection information is not generated, meaning that motion information of a predefined candidate block is set as predicted motion information, and if the number of candidate groups is two or more, candidate group selection information can be generated.
[0362] The category of the candidate block may be any one of inter_blk_A, inter_blk_B, and inter_blk_C, where inter_blk_A may be a category that is basically included, and other categories may be categories that are additionally supported, but are not limited thereto.
[0363] The above description may relate to the motion information prediction candidate configuration for a non-motion motion model, and the same or similar candidate blocks can also be used / referenced for the non-motion motion model (affine model) to configure the candidate group. Meanwhile, in the case of an affine model, since not only the number of motion vectors but also the motion characteristics are different from those of the motion model, the candidate group configuration can be different from that of the motion model.
[0364] For example, if the motion model of the candidate block is an affine model, the configuration of the motion vector set of the candidate block can be included in the candidate group as is. As an example, if the coordinates of the upper left and upper right sides of the candidate block are used as control points, the motion vector set of the coordinates of the upper left and upper right sides of the candidate block can be included in the candidate group as one candidate.
[0365] Alternatively, if the motion model of the candidate block is a moving motion model, a combination of motion vectors of the candidate block set based on the positions of the control points can be configured as a set and included in the candidate group. As an example, if the coordinates of the upper left, upper right, and lower left are used as control points, the motion vector of the upper left control point can be predicted based on the motion vectors of the left, upper, and upper left blocks of the target block (e.g., in the case of a moving motion model), the motion vector of the upper right control point can be predicted based on the motion vectors of the upper and upper right blocks of the target block (e.g., in the case of a moving motion model), and the motion vector of the lower left control point can be predicted based on the motion vectors of the left and lower left blocks of the target block. In the above example, we described a case where the motion model of the candidate block set based on the positions of the control points is a moving motion model, but even if it is an affine model, it is also possible to obtain or induce only the motion vectors of the control point positions. In other words, for the motion vectors of the control points on the upper left, upper right, and lower left of the target block, motion vectors can be obtained or derived from the control points on the upper left, upper right, lower left, and lower right of the candidate block.
[0366] In summary, if the motion model of the candidate block is an affine model, the motion vector set of the block can be included in the candidate group (A), and if the motion model of the candidate block is a translation motion model, the motion vector set that is considered as a candidate for the motion vector of a specified control point of the target block and obtained according to the combination of each control point can be included in the candidate group (B).
[0367] In this case, the candidate group can be constructed using only one of Method A or Method B, or can be constructed using both Method A and Method B. Method A can be constructed first and Method B can be constructed second, but this is not limiting.
[0368] FIG. 11 shows an exemplary diagram of statistical candidates according to one embodiment of the present invention.
[0369] Referring to FIG. 11, blocks A to L in FIG. 11 refer to coded blocks spaced apart by a predetermined interval based on the target mode.
[0370] When candidate blocks for motion information of a target block are limited to blocks adjacent to the target block, they may not reflect various characteristics of the image. For this reason, blocks that have been previously coded in the current picture can also be considered as candidate blocks. Inter_blk_B has already been mentioned as a category related to this. This is called a statistical candidate.
[0371] Since the number of candidate blocks included in the motion information prediction candidate set is limited, an efficient statistical candidate management for this purpose may be necessary.
[0372] (1) A block at a predetermined position that has a predetermined distance from the target block can be considered as a candidate block.
[0373] As an example, blocks with a certain interval can be considered as candidate blocks by limiting the x component based on a predetermined coordinate of the target block (e.g., the upper left coordinate), and examples include (-4, 0), (-8, 0), (-16, 0), and (-32, 0).
[0374] As an example, blocks with a certain interval can be considered as candidate blocks by limiting the y component based on a predetermined coordinate of the target block, and examples include (0, -4), (0, -8), (0, -16), and (0, -32).
[0375] For example, blocks having a certain interval between x and y components other than 0 based on a predetermined coordinate of the target block can be considered as candidate blocks, and examples include (-4, -4), (-4, -8), (-8, -4), (-16, 16), (16, -16), etc. In this example, the x component may have a positive sign depending on the y component.
[0376] Depending on the coding settings, the candidate blocks that are considered as statistical candidates may be restricted. For an example, see FIG.
[0377] A candidate block can be classified into statistical candidates depending on whether it belongs to a predetermined unit, which can be defined from among a largest coding block, a brick, a tile, a slice, a sub-picture, a picture, etc.
[0378] For example, when selecting candidate blocks by limiting them to blocks that belong to the same largest coding block as the target block, A to C can be the candidates.
[0379] Alternatively, when candidate blocks are selected by limiting them to the largest coded block to which the target block belongs and one largest coded block to the left, A to C, H, and I can be the candidates.
[0380] Alternatively, when candidate blocks are selected by limiting them to the largest coded block to which the current block belongs and the blocks that belong to the largest coded block above it, A to C, E, and F may be the candidates.
[0381] Alternatively, when selecting candidate blocks by limiting the blocks belonging to the same slice as the target block, A to I may be the candidates.
[0382] A candidate block can be classified as a statistical candidate depending on whether it is located in a predetermined direction relative to the target block, where the predetermined direction can be defined as left, top, top-left, top-right, etc.
[0383] For example, if you select candidate blocks by limiting them to blocks located to the left and above the target block, B, C, F, I, and L may be candidates.
[0384] Alternatively, when candidate blocks are selected by limiting them to blocks located to the left, above, upper left, and upper right of the target block, A to L may be the candidates.
[0385] As in the above example, even if a candidate block to be considered as a statistical candidate is selected, there may be a priority order for including it in the motion information prediction candidate group. That is, as many candidates as are supported by the statistical candidate can be selected from the priority order.
[0386] The priority can be determined in a predetermined order and can be determined by various coding factors, which can be defined based on the distance between the candidate block and the target block (e.g., whether the distance is short or not; the distance between blocks can be determined based on the x and y components), the relative direction of the candidate block relative to the target block (e.g., left, top, top-left, top-right directions, such as left → top → top-right → left), etc.
[0387] (2) Blocks that have been coded in a predetermined coding order based on the target block can be considered as candidate blocks.
[0388] It should be understood that the examples described below assume a raster scan (e.g., maximum coding block), a Z scan (e.g., coding block, predicted block, etc.), or the contents described below can be modified depending on the scan order.
[0389] In the above-described example (1), a priority order for forming a motion information prediction candidate group can be supported, and similarly, a predetermined priority order can be supported in this embodiment. The priority order can be determined according to various coding factors, but for convenience of explanation, it is assumed that the priority order is determined according to the coding order. In the following example, the candidate block and the priority order will be described together.
[0390] Depending on the coding settings, the candidate blocks that are considered as statistical candidates can be restricted.
[0391] A candidate block can be classified as a statistical candidate depending on whether it belongs to a predetermined unit, which can be determined from among a largest coding block, a brick, a tile, a slice, a sub-picture, a picture, etc.
[0392] For example, when a block belonging to the same picture as the current block is selected as a candidate block, candidate blocks and priorities such as JDEFGKLHIABC may be selected.
[0393] Alternatively, when a block belonging to the same slice as the target block is selected as a candidate block, the candidate blocks and priorities such as DEF, GHI, A, B, C, and the like can be selected.
[0394] A candidate block can be classified as a statistical candidate depending on whether it is located in a predetermined direction relative to the target block, where the predetermined direction can be defined as a leftward or upward direction.
[0395] For example, when candidate blocks are selected by limiting them to blocks located to the left of the target block, candidate blocks and priorities such as KL, H, A, B, C, and C may be selected.
[0396] Alternatively, when candidate blocks are selected by limiting the selection to blocks located above the target block, candidate blocks and priorities such as EF, A, B, C, and the like may be selected.
[0397] The above description is based on the assumption that the candidate blocks and the priorities are combined, but the present invention is not limited to this, and various candidate blocks and priorities can be set.
[0398] It has been mentioned in the above descriptions (1) and (2) that priority levels can be supported for motion information prediction candidates. One or more priority levels can be supported, including one priority level for all candidate blocks, or candidate blocks can be classified into two or more categories and individual priorities can be supported according to the categories. The former may be an example of selecting k candidate blocks according to one priority level, and the latter may be an example of selecting p or q (p + q = k) candidate blocks for each priority level (e.g., two categories). The categories can be classified based on a predetermined coding element, which may include whether the candidate block belongs to a predetermined unit, whether the candidate block is located in a predetermined direction relative to the target block, etc.
[0399] The example of selecting a candidate block based on whether it belongs to a predetermined unit (e.g., division unit) described above through (1) and (2) has been described. Apart from the division unit, candidate blocks can be selected by limiting them to blocks that belong to a predetermined range based on the target mode. For example, they can be identified by boundary points that define the range of the minimum and maximum values of the x or y components, such as (x1, y1), (x2, y2), (x3, y3), etc. The values of x1 to y3 can be integers such as 0, 4, 8, etc. (absolute value basis).
[0400] Statistical candidates can be supported based on either the setting of (1) or (2) above, or a statistical candidate that is a combination of (1) and (2) can be supported. The above explanation has been given to discuss how to select candidate blocks for statistical candidates. The management and update of statistical candidates and cases in which they are included in a motion information prediction candidate group will be described later.
[0401] The motion information prediction candidate group may generally be configured on a block-by-block basis because there is a high possibility that the motion information of blocks adjacent to the target block is the same or similar. That is, in the case of spatial candidates and temporal candidates, the group may be configured based on the target mode.
[0402] On the other hand, the statistical candidate can be configured based on a predetermined unit, which may be a block above the target block, since the candidate block is not adjacent to the target block.
[0403] For example, if the current block is a predicted block, statistical candidates can be formed in units of predicted blocks, or in units of coding blocks.
[0404] Alternatively, if the current block is a coding block, statistical candidates may be configured in units of coding blocks, or in units of ancestor blocks (e.g., blocks whose depth information differs from that of the current block by 1 or more). In this case, the ancestor blocks may include the largest coding block or may be units obtained based on the largest coding block (e.g., an integer multiple of the largest coding block).
[0405] As in the above example, statistical candidates can be constructed based on the current block or higher units, and in the examples described below, it is assumed that statistical candidates are constructed based on the current block.
[0406] Furthermore, memory initialization for statistical candidates may be performed based on a predetermined unit. The predetermined unit may be determined from among a picture, a subpicture, a slice, a tile, a brick, a block, etc. In the case of a block, the predetermined unit may be set based on the largest coding block. For example, memory initialization for statistical candidates may be performed based on an integer number (1, 2, or an integer greater than 1) of the largest coding block, a row or column unit of the largest coding block, etc.
[0407] Next, statistical candidate management and updating will be described. A memory for managing statistical candidates can be provided and can store motion information for up to k blocks. Here, k can be an integer between 1 and 6 or greater. The motion information stored in the memory can be determined from among motion vectors, reference pictures, reference directions, etc. Here, the number of motion vectors (e.g., 1 to 3) can be determined based on motion model selection information. For ease of explanation, the following description will focus on the case where blocks that have been coded according to a predetermined coding order based on a target block are considered as candidate blocks.
[0408] (Case 1) Motion information of a block that has been coded earlier according to the coding order may be included as a candidate in the earlier order. Also, when the maximum number of candidates is reached and an additional candidate is added, the candidate in the earlier order may be removed and the order may be increased by 1. In the following examples, x means a blank that has not yet been configured as a candidate.
[0409] (Example) 1order-[a, x, x, x, x, x, x, x, x, x] 2order-[a, b, x, x, x, x, x, x, x, x] ... 9order-[a, b, c, d, e, f, g, h, i, x] 10order-[a, b, c, d, e, f, g, h, i, j] → j added. Fully filled 11order-[b, c, d, e, f, g, h, i, j, k] → Add k. Remove leading a, move order.
[0410] (Case 2) If there is overlapping motion information, the existing candidates are removed and the order of the existing candidates is adjusted by moving them up.
[0411] Here, redundancy means that the motion information is the same. This can be defined according to the motion information coding mode. For example, in the merge mode, redundancy can be determined based on a motion vector, a reference picture, a reference direction, etc., and in the competitive mode, redundancy can be determined based on a motion vector and a reference picture. In this case, in the competitive mode, if the reference pictures of the candidates are the same, redundancy can be determined through a comparison of motion vectors. If the reference pictures of the candidates are different, redundancy can be determined through a comparison of motion vectors scaled based on the reference pictures of the candidates, but this is not limited to this.
[0412] (Example) 1order-[a, b, c, d, e, f, g, x, x, x] 2order-[a, b, c, d, e, f, g, d, x, x] → d is included, but duplicated -[a, b, c, e, f, g, d, x, x, x] → Remove existing d, move order
[0413] (Case 3) If there is overlapping motion information, the candidate can be marked and managed as a candidate to be stored for a long period of time (long-term candidate) based on a predetermined condition.
[0414] A separate memory for the long-term candidates can be supported, and information such as occurrence frequency in addition to motion information can be additionally stored and updated. In the following examples, the memory for long-term candidates can be represented by ().
[0415] (Example) 1order-[(), b, c, d, e, f, g, h, i, j, k]
[0416] 2order - [(), b, c, d, e, f, g, h, i, j, k] → e is the next order. Duplication -[(e), b, c, d, f, g, h, i, j, k] → Move to the beginning. Long-term marking
[0417] 3order-[(e), c, d, f, g, h, i, j, k, l] → add l and remove b. ... 10order-[(e), l, m, n, o, p, q, r, s, t] → l is the next order. -[(e,l),m,n,o,p,q,r,s,t] → long-term marking 11order-[(e, l), m, n, o, p, q, r, s, t] → l is the next order. 3 duplicates -[(l,e),m,n,o,p,q,r,s,t] → Change the order of long-term candidates
[0418] The above example shows a case where long-term candidates are managed in an integrated manner with existing statistical candidates (short-term candidates), but they can be managed separately from short-term candidates, and the number of long-term candidates can be 0, 1, 2 or an integer greater than or equal to 1.
[0419] Short-term candidates may be managed in a first-in-first-out manner based on the encoding order, while long-term candidates may be managed according to frequency based on candidates that have redundancy among short-term candidates. For long-term candidates, the candidate order may be changed according to frequency during memory update, but is not limited to this.
[0420] The above example may be a partial example in which candidate blocks are classified into two or more categories and priority according to the categories is supported. When a statistical candidate is included in the motion information prediction candidate group, a short-term candidate and b long-term candidate may be configured according to the priority within each candidate. Here, a and b may be integers of 0, 1, 2, or more.
[0421] The cases 1 to 3 are examples of statistical candidates, and the configuration of (case 1+case 2) or (case 1+case 3) is possible, and the statistical candidates can be managed in various other modified and added forms.
[0422] FIG. 12 is a conceptual diagram of statistical candidates based on a non-movement motion model according to an embodiment of the present invention.
[0423] The statistical candidate may be a candidate supported not only by a motion model but also by a non-motion model. In the case of a motion model, there are many blocks that can be referenced as spatial or temporal candidates because of their high occurrence frequency, but in the case of a non-motion model, there may be a case where a statistical candidate other than a spatial or temporal candidate is needed because of their low occurrence frequency.
[0424] Referring to FIG. 12, a block that belongs to the same space as the target block and is far adjacent to the target block is<inter_blk_B> 10 shows a case where motion information {TV0, TV1} of a current block is predicted based on motion information {PV0, PV1} of the current block (hmvp).
[0425] Compared to the statistical candidates of the above-mentioned mobile motion model, the number of motion vectors in the motion information stored in memory may be increased. Therefore, the basic concept of the statistical candidates may be the same as or similar to that of the mobile motion model, and related explanations can be provided by configuring the number of motion vectors to be different. Next, differences due to non-mobile motion models will be described below. In the example below, an affine model with three control points (e.g., v0, v1, and v2) will be assumed.
[0426] In the case of an affine model, the motion information stored in memory can be determined from among a reference picture, a reference direction, and motion vectors v0, v1, and v2. Here, the motion vectors can be stored in various forms. For example, the motion vectors can be stored as they are, or predetermined difference value information can be stored.
[0427] FIG. 13 is a diagram illustrating an example of the configuration of motion information for each control point position stored as a statistical candidate according to an embodiment of the present invention.
[0428] Referring to FIG. 13, (a) shows the motion vectors of candidate blocks v0, v1, and v2 expressed as v0, v1, and v2, and (b) shows the motion vector of candidate block v0 expressed as v0, and the motion vectors of v1 and v2 are expressed as the difference value v from the motion vector of v0. * 1, v * This shows the case where it is expressed as 2.
[0429] In other words, the motion vector of each control point position can be stored as is, or the difference value with the motion vector of another control point position can be stored. This is an example of a configuration that takes into consideration memory management, and various modifications are possible.
[0430] Whether or not a statistical candidate is supported can be determined by explicitly generating related information for each unit such as a sequence, a picture, a slice, a tile, a brick, etc., or by implicitly determining the support depending on the encoding settings. The encoding settings can be defined by the various encoding elements described above, and therefore, detailed description thereof will be omitted.
[0431] Here, whether or not a statistical candidate is supported may be determined according to a motion information coding mode, or whether or not a statistical candidate is supported may be determined according to motion model selection information. For example, a statistical candidate may be supported from among merge_inter, merge_ibc, merge_affine, comp_inter, comp_ibc, and comp_affine.
[0432] If statistical candidates are supported for both the moving motion model and the non-moving motion model, memory for multiple statistical candidates can be supported.
[0433] Next, a method for constructing a group of motion information prediction candidates according to the motion information coding mode will be described.
[0434] A group of motion information prediction candidates for a competitive mode (hereinafter, a competitive mode candidate group) may include k candidates, where k may be an integer of 2, 3, 4, or more. The competitive mode candidate group may include at least one of spatial candidates or temporal candidates.
[0435] The spatial candidate can be derived from at least one of the blocks adjacent to the reference block in the left, upper, upper left, upper right, lower left, etc. Alternatively, at least one candidate can be derived from the block adjacent to the left (left or lower left block) and the block adjacent to the upper side (upper left, upper or upper right block), which will be described later assuming this setting.
[0436] There can be more than one priority order for forming the candidate set: adjacent regions to the left can have a bottom-left-left priority order, and adjacent regions to the top can have a top-right-top-top-top-left priority order.
[0437] The above example may be a configuration in which spatial candidates are guided only in blocks that have the same reference picture as the current block. However, scaling processing based on the reference picture of the current block (described below) may be performed. * In this case, for the left-adjacent region, the spatial candidate can be induced via the left-bottom-left-left * -lower left * or left-lower left-lower left * -left * The order of priority can be set, and for adjacent areas in the upward direction, it is top right-top-top-top left-top right. * -above * -Top left * or top right-top-top-top left-top left * -above * -Top right * The order of priority can be set.
[0438] The temporal candidates can be derived from at least one of blocks adjacent to the block corresponding to the reference block in the center, left, right, top, bottom, top left, top right, bottom left, bottom right, etc. There can be a priority for configuring the candidate group, and priorities such as center-bottom left-right-bottom, bottom left-center-top left, etc. can be set.
[0439] If the setting is such that the sum of the maximum allowable number of spatial candidates and the maximum allowable number of temporal candidates is less than the number of competitive mode candidates, temporal candidates can be included in the candidate group regardless of the composition of the spatial candidate group.
[0440] Based on the priority and availability of each candidate block, and the maximum allowed number of temporal candidates (q, an integer between 1 and the number of competitive mode candidate groups), all or some of the candidates can be included in the candidate group.
[0441] Here, if the maximum allowable number of spatial candidates is set to be the same as the number of merging mode candidates, temporal candidates cannot be included in the candidate group, and if the maximum allowable number of spatial candidates is not met, temporal candidates can be included in the candidate group. This example assumes the latter case.
[0442] Here, the motion vector of the temporal candidate can be obtained based on the motion vector of the candidate block and the distance between the current image and the reference image of the target block, and the reference image of the temporal candidate can be obtained based on the distance between the current image and the reference image of the target block, or based on the reference image of the temporal candidate, or based on a pre-defined reference image (e.g., reference picture index is 0).
[0443] If the competitive mode candidate group is not filled with spatial candidates, temporal candidates, etc., the candidate group can be completed through default candidates including a zero vector.
[0444] The competitive mode is mainly explained with respect to comp_inter, and in the case of comp_ibc or comp_affine, a candidate group can be formed with similar or different candidates.
[0445] For example, in the case of comp_ibc, a candidate group can be configured based on predetermined candidates selected from spatial candidates, statistical candidates, combination candidates, default candidates, etc. In this case, the candidate group can be configured with spatial candidates as the priority, followed by statistical candidates, combination candidates, default candidates, etc. However, the candidate group is not limited to this and various orders are possible.
[0446] Alternatively, in the case of comp_affine, the candidate group may be configured based on predetermined candidates selected from spatial candidates, temporal candidates, statistical candidates, combination candidates, default candidates, etc. In particular, the candidate group may be configured based on motion vector set candidates of the candidate (e.g., one block). <1> Alternatively, candidates based on the positions of the control points (for example, two or more blocks) may be combined to form a candidate group. <2> It is possible to <1> The candidate for <2> The order in which candidates for are included is possible but not limited to:
[0447] Detailed explanations regarding the configuration of the various competition mode candidates mentioned above can be found in comp_inter, so detailed explanations will be omitted.
[0448] In the process of forming a competitive mode candidate group, if there is overlapping motion information among the candidates included earlier, the candidate with the next highest priority can be included in the candidate group while maintaining the candidate included earlier.
[0449] Here, before constructing the candidate set, the motion vectors in the candidate set may be scaled based on the distance between the reference picture of the current block and the current picture. For example, if the distance between the reference picture of the current block and the current picture is the same as the distance between the reference picture of the candidate block and the picture to which the candidate block belongs, the motion vector may be included in the candidate set. If the distances are not the same, the motion vector scaled according to the distance between the reference picture of the current block and the current picture may be included in the candidate set.
[0450] In this case, redundancy means that motion information is the same, and this can be defined depending on the motion information coding mode. In the competitive mode, redundancy can be determined based on motion vectors, reference pictures, reference directions, etc. For example, if at least one component of the motion vector is different, it can be determined that there is no redundancy. The redundancy checking process is generally performed when a new candidate is included in the candidate group, but it can also be omitted.
[0451] A set of motion information prediction candidates for a merge mode (hereinafter, a merge mode candidate set) may include k candidates, where k may be an integer of 2, 3, 4, 5, 6, or more. The merge mode candidate set may include at least one of spatial candidates or temporal candidates.
[0452] The spatial candidates can be derived from at least one of blocks adjacent to the reference block in the left, top, top-left, top-right, bottom-left directions, etc. Priorities for configuring the candidate group can be set, such as left-top-bottom-left-top-top-top-top-left, left-top-top-top-bottom-left-top, top-left-bottom-left-top-top-top-top.
[0453] Based on the priority, the availability of each candidate block (e.g., determined based on the coding mode, block position, etc.), and the maximum allowable number of spatial candidates (p, an integer between 1 and the number of merging mode candidate groups), all or some of the candidates may be included in the candidate group. Based on the maximum allowable number and availability, the candidates may not be included in the candidate group in the order tl-tr-bl-t3-l3; if the maximum allowable number is 4 and the availability of all candidate blocks is true, the motion information of tl is not included in the candidate group, but if the availability of some candidate blocks is false, the motion information of tl may be included in the candidate group.
[0454] The temporal candidates can be derived from at least one of blocks adjacent to the block corresponding to the reference block in the center, left, right, top, bottom, top left, top right, bottom left, bottom right, etc. There can be a priority for configuring the candidate group, and priorities such as center-bottom left-right-bottom, bottom left-center-top left, etc. can be set.
[0455] Based on the priority and availability of each candidate block, and the maximum allowed number of temporal candidates (q, an integer between 1 and the number of merging mode candidate groups), all or some of the candidates can be included in the candidate group.
[0456] Here, the motion vector of the temporal candidate can be obtained based on the motion vector of the candidate block, and the reference image of the temporal candidate can be obtained based on the reference image of the candidate block or based on an already defined reference image (for example, the reference picture index is 0).
[0457] The priority of the merge mode candidates can be set to spatial candidate-temporal candidate or vice versa, and a mixed priority of spatial candidate and temporal candidate can be supported. In this example, we will assume that the priority is spatial candidate-temporal candidate.
[0458] In addition, statistical candidates or combination candidates may be further included in the merging mode candidate group. The statistical candidates and combination candidates may be configured after the spatial candidates and the temporal candidates, but are not limited thereto, and various priorities may be set.
[0459] The statistical candidates can manage up to n pieces of motion information, of which z pieces of motion information can be included in the merging mode candidate group as statistical candidates, where z can vary depending on the candidate configuration already included in the merging mode candidate group, and can be an integer of 0, 1, 2, or more, and can be less than or equal to n.
[0460] A combination candidate can be derived by combining n candidates already included in the merging mode candidate group, where n can be an integer of 2, 3, 4, or more. The number (n) of combination candidates can be information that is explicitly generated in units of a sequence, a picture, a subpicture, a slice, a tile, a brick, a block, etc. Or, it can be implicitly determined according to the encoding setting. In this case, the encoding setting can be defined based on one or more factors such as the size, shape, position, image type, color components, etc. of the reference block.
[0461] In addition, the number of combination candidates may be determined based on the number of candidates that are not filled in the merging mode candidate group. In this case, the number of candidates that are not filled in the merging mode candidate group may be a difference value between the number of merging mode candidate groups and the number of candidates that have already been filled. In other words, if the configuration of the merging mode candidate group has already been completed, no combination candidate may be added. If the configuration of the merging mode candidate group has not been completed, a combination candidate may be added, but if the merging mode candidate group has one or less filled candidates, no combination candidate is added.
[0462] If the merging mode candidate group is not filled with spatial candidates, temporal candidates, statistical candidates, combination candidates, etc., the candidate group can be completed through default candidates including a zero vector.
[0463] The merge mode is mainly explained with respect to merge_inter, and in the case of merge_ibc or merge_affine, a candidate group can be formed from different or similar candidates.
[0464] For example, in the case of merge_ibc, a candidate group can be configured based on predetermined candidates selected from spatial candidates, statistical candidates, combination candidates, default candidates, etc. In this case, the candidate group can be configured with spatial candidates as the priority, followed by the statistical candidates, combination candidates, default candidates, etc. However, the candidate group is not limited to this and various orders are possible.
[0465] Alternatively, in the case of merge_affine, a candidate group may be configured based on predetermined candidates selected from spatial candidates, temporal candidates, statistical candidates, combination candidates, default candidates, etc. In particular, the candidate group may be configured based on motion vector set candidates of the candidate (e.g., one block). <1> Alternatively, candidates based on the positions of the control points (for example, two or more blocks) may be combined to form a candidate group. <2> It is possible to <1> The candidates for are included in the candidate set first, and then <2> The order in which candidates for are included is possible but not limited to:
[0466] Details regarding the configuration of the various merge mode candidates mentioned above can be derived from merge_inter, so a detailed description will be omitted.
[0467] In the process of constructing a group of merge mode candidates, if there is overlapping motion information among the previously included candidates, the previously included candidate can be maintained and the next highest priority candidate can be included in the group of candidates.
[0468] In this case, redundancy means that the motion information is the same, and this can be defined depending on the motion information coding mode. In the merge mode, whether or not there is redundancy can be determined based on the motion vector, reference picture, reference direction, etc. For example, if at least one component of the motion vector is different, it can be determined that there is no redundancy. The redundancy checking process is generally performed when a new candidate is included in the candidate group, but it can also be omitted.
[0469] 14 is a flowchart of motion information encoding according to an embodiment of the present invention, specifically, encoding of motion information of a current block in a competitive mode.
[0470] A motion vector prediction candidate list for the current block can be generated (S1400). The above-mentioned competitive mode candidate group can refer to the motion vector candidate list, and a detailed description thereof will be omitted.
[0471] The motion vector difference value of the target block can be restored (S1410). The difference values for the x and y components of the motion vector can be restored separately, and the difference value of each component can have a value of 0 or more.
[0472] A prediction candidate index for a target mode can be selected from the motion vector prediction candidate list (S1420). A motion vector predictor value for a target block can be derived based on the motion vector obtained from the candidate list according to the candidate index. If one pre-defined motion vector predictor value can be derived, the process of selecting a prediction candidate index and index information can be omitted.
[0473] The precision information of the differential motion vector can be derived (S1430). Precision information commonly applied to the x and y components of the motion vector can be derived, or precision information applied to each component can be derived. If the motion vector differential value is 0, the precision information can be omitted, and this process can also be omitted.
[0474] An adjustment offset for the motion vector predictor may be derived (S1440). The offset may be a value added to or subtracted from the x or y component of the motion vector predictor. The offset may be supported for only the x or y component, or for both the x and y components.
[0475] Assuming that the motion vector predictor is (pmv_x, pmv_y) and the adjusted offset is offset_x, offset_y, the motion vector predictor can be adjusted (or obtained) to (pmv_x+offset_x, pmv_y+offset_y).
[0476] Here, the absolute values of offset_x and offset_y are integers of 0, 1, 2, or more, and may have values (1, -1, +2, -2, etc.) that take into account sign information. Also, offset_x and offset_y may be determined based on a predetermined accuracy. The predetermined accuracy may be determined from 1 / 16, 1 / 8, 1 / 4, 1 / 2, 1 pixel, etc., and may be determined based on interpolation accuracy, motion vector accuracy, etc.
[0477] For example, if the interpolation accuracy is 1 / 4 pixel units, the absolute value and sign information can be combined to 0, 1 / 4, -1 / 4, 2 / 4, -2 / 4, 3 / 4, -3 / 4, etc.
[0478] Here, offset_x and offset_y can have a and b values, respectively, and a and b can be integers of 0, 1, 2 or more. a and b can have fixed values or variable values. Also, a and b can have the same or different values.
[0479] Whether or not an adjustment offset for a motion vector predictor is supported may be explicitly supported in units of a sequence, a picture, a subpicture, a slice, a tile, a brick, etc., or may be implicitly determined according to encoding settings. In addition, settings of the adjustment offset (e.g., a value range, the number, etc.) may be determined according to encoding settings.
[0480] The encoding settings can be determined taking into consideration at least one of the encoding elements such as image type, color components, state information of the target block, motion model selection information (e.g., whether it is a motion model or not), reference picture (e.g., whether it is the current picture or not), and differential motion vector precision selection information (e.g., whether it is a predetermined unit among 1 / 4, 1 / 2, 1, 2, and 4 units).
[0481] For example, whether or not to support an adjustment offset and how to set the offset may be determined depending on the size of the block. In this case, the support and setting range of the block size may be determined according to a first threshold size (minimum value) or a second threshold size (maximum value), and each threshold size may be expressed as W, H, W×H, or W*H, where W is the width (W) and H is the height (H) of the block. For the first threshold size, W and H may be integers of 4, 8, 16, or more, and W*H may be an integer of 16, 32, 64, or more. For the second threshold size, W and H may be integers of 16, 32, 64, or more, and W*H may be an integer of 64, 128, 256, or more. The above range may be determined by either the first threshold size or the second threshold size, or by both.
[0482] In this case, the threshold value may be fixed or adaptive based on the image (e.g., image type, etc.), where the first threshold value may be set based on the size of the smallest coding block, the smallest prediction block, the smallest transform block, etc., and the second threshold value may be set based on the size of the largest coding block, the largest prediction block, the largest transform block, etc.
[0483] In addition, the adjustment offset can be applied to all candidates included in the motion information prediction candidate group, or only to some candidates. In the example described below, it is assumed that the adjustment offset is applied to all candidates included in the candidate group, but the candidates to which the adjustment offset is applied can be selected from 0, 1, 2, or the maximum number of the candidate group.
[0484] If adjustment offset is not supported, this process and the adjustment offset information can be omitted.
[0485] The motion vector of the current block can be restored by adding the motion vector prediction value and the motion vector differential value (S1450). At this time, a process of unifying the motion vector prediction value or the motion vector differential value to the motion vector precision can precede the process, and the motion vector scaling process described above can precede or be performed during this process.
[0486] The above configurations and sequences are merely some examples, and are not limiting, and various modifications are possible.
[0487] FIG. 15 is a diagram illustrating motion vector predictor candidates and the motion vector of a current block according to an embodiment of the present invention.
[0488] For ease of explanation, it is assumed that two motion vector prediction candidates are supported and one component (either x or y component) is compared. It is also assumed that the interpolation accuracy or motion vector accuracy is in 1 / 4 pixel units. In addition, it is assumed that the differential motion vector accuracy is supported in 1 / 4, 1, and 4 pixel units (for example, when it is 1 / 4, <0> , when it is 1 <10> , when it is 4 <11> Also, it is assumed that the motion vector differential value is processed by unary binarization (for example, 0: <0> , 1: <10> , 2: <110> etc.)
[0489] Referring to FIG. 15, the actual motion vector (X) has a value of 2, candidate 1 (A) has a value of 0, and candidate 2 (B) has a value of 1 / 4.
[0490] (When differential motion vector accuracy is not supported) Since the distance (da) between A and X is 8 / 4 (9 bits), and the distance (db) between B and X is 7 / 4 (8 bits), from the viewpoint of bit generation, B can be selected as the prediction candidate.
[0491] (When differential motion vector accuracy is supported) Since da is 8 / 4, a total of 5 bits can be generated, including 1 pixel unit accuracy (2 bits) and 2 / 1 difference value information (3 bits). On the other hand, since db is 7 / 4, a total of 9 bits can be generated, including 1 / 4 pixel unit accuracy (1 bit) and 7 / 4 difference value information (8 bits). From the perspective of bit generation, prediction candidate A can be selected.
[0492] When differential motion vector accuracy is not supported, as in the example above, it is advantageous to select a candidate with a short distance from the motion vector of the target block as the prediction candidate, and when differential motion vector accuracy is supported, it may be important to select a prediction candidate based not only on the distance from the motion vector of the target block but also on the amount of information generated based on the accuracy information.
[0493] 16 is a diagram illustrating motion vector predictor candidates and motion vectors of a current block according to an embodiment of the present invention, assuming that differential motion vector accuracy is supported.
[0494] Since da is 3¾, a total of 35 bits can be generated, including 1 / 4 pixel accuracy (1 bit) and differential value information (34 bits) at 3¾. On the other hand, since db is 2¾, a total of 23 bits can be generated, including 1 / 4 pixel accuracy (1 bit) and differential value information (22 bits) at 2¾. From the perspective of bit generation, prediction candidate B can be selected.
[0495] 17 is a diagram illustrating an example of motion vector predictor candidates and a motion vector of a current block according to an embodiment of the present invention. Hereinafter, it is assumed that differential motion vector accuracy and adjusted offset information are supported.
[0496] In this example, it is assumed that the adjustment offset has two options, 0 and +1, and a flag (1 bit) indicating whether the adjustment offset has been applied and offset selection information (1 bit) are generated.
[0497] 17, A1 and B1 may be motion vector predictors obtained based on prediction candidate indexes, and A2 and B2 may be new motion vector predictors obtained by correcting the adjustment offsets for A1 and B1. Assume that the distances between A1, A2, B1, and B2 and X are da1, da2, db1, and db2, respectively.
[0498] (1) Since da1 is 33 / 4, a total of 37 bits can be generated, including 1 / 4 pixel unit accuracy (1 bit), offset application flag (1 bit), offset selection information (1 bit), and difference value information (34 bits) for 33 / 4.
[0499] (2) Since da2 is 32 / 4, a total of 7 bits can be generated, including 4 pixel unit accuracy (2 bits), offset application flag (1 bit), offset selection information (1 bit), and 2 / 1 difference value information (3 bits).
[0500] (3) Since db1 is 2 1 / 4, a total of 25 bits can be generated, including 1 / 4 pixel unit accuracy (1 bit), offset application flag (1 bit), offset selection information (1 bit), and difference value information (22 bits) for 2 1 / 4.
[0501] (4) Since db2 is 20 / 4, a total of 10 bits can be generated, including 1 pixel unit accuracy (2 bits), offset application flag (1 bit), offset selection information (1 bit), and 5 / 1 difference value information (6 bits).
[0502] From the viewpoint of bit generation, the prediction candidate can be selected as A, and the offset selection information can be selected as the first index (+1 in this example).
[0503] Considering the above example, when deriving a motion vector difference value based on a conventional motion vector predictor, a small difference in vector value may result in a large number of bits. This problem can be solved by adjusting the motion vector predictor.
[0504] A predetermined flag may be supported to apply an adjustment offset to a motion vector predicted value, and the predetermined flag may include an offset application flag, offset selection information, and the like.
[0505] (If only the offset application flag is supported)
[0506] When the offset application flag is 0, no offset is applied to the motion vector predicted value, and when the offset application flag is 1, a preset offset can be added to or subtracted from the predicted motion vector value.
[0507] (If only offset selection information is supported)
[0508] The offset set based on the offset selection information can be added to or subtracted from the predicted motion vector value.
[0509] (If offset application flag and offset selection information are supported)
[0510] When the offset application flag is 0, no offset is applied to the motion vector predicted value, and when the offset application flag is 1, an offset set based on the offset selection information can be added to or subtracted from the predicted motion vector value. This setting is assumed in the examples described later.
[0511] Meanwhile, the offset application flag and offset selection information can be used as information that is generated unnecessarily depending on the situation. That is, even if an offset is not applied, if the offset-related information already has a zero value or if the amount of information is reduced based on the maximum precision information, the offset-related information may be inefficient. Therefore, it is necessary to support a setting where the offset-related information is generated implicitly according to predetermined conditions, rather than being generated explicitly all the time.
[0512] Next, it is assumed that the motion-related coding order proceeds as follows: (restoration of motion vector differential value → acquisition of differential motion vector precision). In this example, it is assumed that differential motion vector precision is supported when at least one of the x and y components of the motion vector differential value is not 0.
[0513] If the motion vector differential value is not 0, the differential motion vector precision can be selected from among {1 / 4, 1 / 2, 1, 4} pixel units.
[0514] If the selected information belongs to a predetermined category, information about the adjustment offset can be implicitly omitted, and if the selected information does not belong to a predetermined category, information about the adjustment offset can be explicitly generated.
[0515] Here, the category may include one of the differential motion vector accuracy candidates, and various categories such as {1 / 4}, {1 / 4, 1 / 2}, {1 / 4, 1 / 2, 1}, etc. Here, the minimum accuracy may be included in the category.
[0516] For example, if the differential motion vector precision is in 1 / 4 pixel units (e.g., minimum precision), the offset application flag is implicitly set to 0 (i.e., not applied), if the differential motion vector precision is not in 1 / 4 pixel units, the offset application flag is explicitly generated, and if the offset application flag is 1 (i.e., offset applied), offset selection information can be generated.
[0517] Alternatively, if the differential motion vector precision is 4 pixel units (e.g., maximum precision), the offset application flag may be explicitly generated, and if the offset application flag is 1, the offset selection information may be generated. Alternatively, if the differential motion vector precision is not 4 pixel units, the offset information may be implicitly set to 0.
[0518] The above example assumes, but is not limited to, a case in which offset-related information is implicitly omitted when the differential motion vector points to the minimum precision, and offset-related information is explicitly generated when the differential motion vector points to the maximum precision.
[0519] FIG. 18 is an exemplary diagram illustrating the arrangement of a plurality of motion vector predictors according to an embodiment of the present invention.
[0520] The above example has been used to explain the overlap check when constructing a group of motion information prediction candidates. Here, overlap means that the motion information is the same, and it has been mentioned above that if at least one component of the motion vector is different, it can be determined that there is no overlap.
[0521] Multiple candidates for motion vector predictors may not overlap with each other through a redundancy check process. However, if the components of the multiple candidates are very similar (i.e., the x or y component of each candidate is within a predetermined range, and the width or height of the predetermined range is an integer of 1, 2, or more. Alternatively, the range can be set based on offset information), the motion vector predictor and the offset-corrected motion vector predictor may overlap. Various settings for this purpose are possible.
[0522] As an example (C1), in the motion information prediction candidate group construction step, a new candidate can be included in the candidate group if it does not overlap with any of the candidates already included. In other words, if it is the same as the previously described existing candidate and does not overlap with the candidate itself, it can be included in the candidate group.
[0523] As an example (C2), if a new candidate and a candidate (group_A) obtained by adding an offset thereto and a candidate (group_B) obtained by adding an offset thereto do not overlap by a predetermined number of times, they can be included in the candidate group. The predetermined number can be an integer of 0, 1, or more. If the predetermined number is 0, the new candidate may not be included in the candidate group if there is even one overlap. Alternatively, as (C3), a new candidate may be included in the candidate group by adding a predetermined offset (which is a different concept from the adjustment offset), and the offset may be a value that allows group_A to be configured so that it does not overlap with group_B.
[0524] Referring to FIG. 18, multiple motion vectors belonging to categories A, B, C, and D (in this example, AX, BX, CX, DX, and X are 1 and 2; for example, A1 is a candidate included in the candidate group before A2) may satisfy the non-overlapping condition (when the candidate motion vectors differ in even one component).
[0525] In this example, it is assumed that the offset supports -1, 0, and 1 for the x and y components, respectively, and the offset-corrected motion vector predicted value is represented by * in the drawing. Also, the dashed lines (rectangles) represent ranges (e.g., group_A, group_B, etc.) that can be obtained by adding an offset around a predetermined motion vector predicted value.
[0526] In the case of category A, it may be the case that A1 and A2 do not overlap, and group_A1 and group_A2 do not overlap.
[0527] In the case of categories B, C, and D, this may apply when B1 / C1 / D1 and B2 / C2 / D2 do not overlap, and group_B1 / C1 / D1 and group_B2 / C2 / D2 partially overlap.
[0528] Here, the B category may be an example of forming a candidate group without taking any special measures (C1).
[0529] Here, in the case of the C category, C2 is determined to have duplication in the duplication check step, and C3, which has the next highest priority, can be included in the candidate group (C2).
[0530] Here, in the case of the D category, it can be the case that D2 is determined to have redundancy in the redundancy check step, and D2 is corrected so that group_D2 does not overlap with group_D1 (i.e., D3. D3 is not a motion vector present in the priority of the candidate group configuration) (C3).
[0531] The present invention may be applied to a setting in which a group of prediction mode candidates is formed from the various categories, or various methods other than those mentioned above may be applied.
[0532] 19 is a flowchart illustrating a process for encoding motion information in a merge mode according to an embodiment of the present invention. Specifically, the process may involve encoding motion information of a current block in a merge mode.
[0533] A motion information prediction candidate list for the target block can be generated (S1900). The above-mentioned merging mode candidate group can refer to the motion information prediction candidate list, and a detailed description thereof will be omitted.
[0534] A prediction candidate index for a target mode can be selected from a motion information prediction candidate list (S1910). A motion vector prediction value for a target block can be derived based on motion information acquired from the candidate list according to the candidate index. If one pre-set motion information prediction value can be derived, the prediction candidate index selection process and index information can be omitted.
[0535] An adjustment offset for the motion vector predictor may be derived (S1920). The offset may be a value added to or subtracted from the x or y component of the motion vector predictor. The offset may be supported for only one of the x and y components, or for both the x and y components.
[0536] The adjustment offset in this process may be the same as or similar to the adjustment offset described above, so a detailed description will be omitted and differences will be described later.
[0537] Assuming that the motion vector predicted value is (pmv_x, pmv_y) and the adjusted offset is offset_x, offset_y, the motion vector predicted value can be adjusted (or obtained) to (pmv_x+offset_x, pmv_y+offset_y).
[0538] Here, the absolute values of offset_x and offset_y may be integers such as 0, 1, 2, 4, 8, 16, 32, 64, and 128, and may have values that take into account sign information. Also, offset_x and offset_y may be determined based on a predetermined precision. The predetermined precision may be determined from 1 / 16, 1 / 8, 1 / 4, 1 / 2, and 1 pixel units.
[0539] For example, if the motion vector precision is 1 / 4 pixel, the absolute value and sign information can be combined to 0, 1 / 4, -1 / 4, 1 / 2, -1 / 2, 1, -1, 2, -2, etc.
[0540] Here, offset_x and offset_y can have a and b values, respectively, and a and b can be integers such as 0, 1, 2, 4, 8, 16, and 32. a and b can have fixed values or variable values. Also, a and b can have the same or different values.
[0541] Whether or not an adjustment offset for a motion vector predictor is supported may be explicitly supported in units of a sequence, a picture, a subpicture, a slice, a tile, a brick, etc., or may be implicitly determined according to encoding settings. In addition, settings of the adjustment offset (e.g., a value range, the number, etc.) may be determined according to encoding settings.
[0542] The encoding settings can be determined taking into account at least one of the encoding factors such as image type, color components, state information of the target block, motion model selection information (e.g., whether it is a motion model), and reference picture (e.g., whether it is the current picture).
[0543] For example, whether or not to support an adjustment offset and how to set the offset may be determined depending on the size of the block. In this case, the support and setting range of the block size may be determined according to a first threshold size (minimum value) or a second threshold size (maximum value), and each threshold size may be expressed as W, H, W×H, or W*H, where W is the width (W) and height (H) of the block. For the first threshold size, W and H may be integers of 4, 8, 16, or more, and W*H may be an integer of 16, 32, 64, or more. For the second threshold size, W and H may be integers of 16, 32, 64, or more, and W*H may be an integer of 64, 128, 256, or more. The range may be determined by either the first threshold size or the second threshold size, or by both.
[0544] In this case, the threshold value may be fixed or adaptive depending on the image (e.g., image type, etc.). In this case, the first threshold value may be set based on the size of a minimum coding block, a minimum prediction block, a minimum transform block, etc., and the second threshold value may be set based on the size of a maximum coding block, a maximum prediction block, a maximum transform block, etc.
[0545] In addition, the adjustment offset can be applied to all candidates included in the motion information prediction candidate group, or only to some candidates. In the example described below, it is assumed that the adjustment offset is applied to all candidates included in the candidate group, but the candidates to which the adjustment offset is applied can be selected from 0, 1, 2, or the maximum number of the candidate group.
[0546] A predetermined flag for applying an adjusted offset to a motion vector predicted value can be supported, and the predetermined flag can be configured as an offset application flag, offset absolute value information, offset sign information, etc.
[0547] If adjustment offset is not supported, this process and adjustment offset information can be omitted.
[0548] A motion vector of the target mode can be reconstructed through the motion vector predictor (S1930).Motion information other than the motion vector (e.g., reference picture, reference direction, etc.) can be obtained based on the prediction candidate index.
[0549] The above configurations and sequences are merely examples, and are not limited thereto, and various modifications are possible. The background explanation for supporting the adjusted offset in the merge mode has been given above through various examples of the competition mode, so a detailed explanation will be omitted.
[0550] The methods of the present invention may be embodied in the form of program instructions that can be executed by various computer means and stored on computer-readable media. The computer-readable media may include, alone or in combination with program instructions, data files, data structures, and the like. The program instructions stored on the computer-readable media may be those specially designed and constructed for the purposes of the present invention, or those well known and available to those skilled in the computer software arts.
[0551] Examples of computer-readable media may include hardware devices specially configured to store and execute program instructions, such as ROM, RAM, flash memory, etc. Examples of program instructions may include not only machine language code produced by a compiler, but also high-level language code that can be executed by a computer using an interpreter, etc. The above-mentioned hardware devices may be configured to operate as at least one software module to perform the operations of the present invention, and vice versa.
[0552] Furthermore, the above-described methods or devices may be realized in such a manner that all or part of their components or functions are combined or separated.
[0553] While the present invention has been described above with reference to preferred embodiments, those skilled in the art will appreciate that various modifications and variations can be made to the present invention without departing from the spirit and scope of the invention as set forth in the following claims. [Industrial Applicability]
[0554] The present invention can be used to encode / decode images.< / qt>
Claims
1. constructing a list of motion candidate predictors for a current block; deriving a predicted motion vector from the motion candidate list based on a predicted candidate index; restoring predicted motion vector adjustment offset information; restoring a motion vector of a current block based on the predicted motion vector and the predicted motion vector adjustment offset information; The inter prediction method, wherein the motion candidate list includes at least one of spatial candidates, temporal candidates, statistical candidates, or combination candidates.
2. The inter prediction method according to claim 1 , wherein the predicted motion vector adjustment offset is determined based on at least one of an offset application flag or offset selection information.
3. The inter prediction method according to claim 1 , wherein the information on whether the predicted motion vector adjustment offset information is supported is included in at least one of a sequence, a picture, a sub-picture, a slice, a tile, or a brick.
4. If the current block is coded in a merge mode, the motion vector of the current block is restored using a zero vector; The inter prediction method according to claim 1 , wherein if the current block is coded in a competitive mode, the motion vector of the current block is reconstructed using a motion vector differential value.
Citation Information
Patent Citations
Offset temporal motion vector predictor (TMVP)
US20180098085A1
Method and apparatus for encoding / decoding video signal
WO2017164645A2
Offset vector identification of temporal motion vector predictor
WO2018052986A1
Method and apparatus for history-based motion vector prediction
WO2020018297A1